1 / 5100%
DATA SCREENING BASICS 1
Data Screening Basics
Hannah K. Smith
School of Behavioral Sciences, Liberty University
EDCO 735: Statistics
Dr. Robin Henson
September 14, 2025
Data Screening Basics
Prompt 1
When researchers use a z score and the standard normal table when finding percentile
ranks, they are assuming the data is following a normal distribution. However, if they are given
an extremely non-normal empirical frequency distribution, those assumptions fail. An example
of this can be looking at incomes. Most people fall within certain standard deviations, but there
are those few that make far more than the average person. This causes a heavily skewed curve,
and the percentiles that the normal table gives us will be exaggerated and we will not have a
clear idea at how rare the highest values are. In the same sense, if the distribution is heavily
skewed in the opposite direction due to extreme lows, the percentile ranks will underestimate
DATA SCREENING BASICS 2
how often the extremes actually show up (Field, 2018) and thus misrepresent the true
frequency of those cases. Basically, if a normal distribution table is used with an extremely non-
normal distribution or the wrong type of data, it could be misleading and we could have the
wrong idea in mind with what we are working with.
When that happens, researchers have some options. One approach a researcher can
take is to use empirical percentiles. This is when a researcher will figure out where a score falls
in relation to a certain value in the actual data set being used. Another approach could be
applying data transformations (e.g., log or arcsine) when it is appropriate (Warren, 2021).
Therefore, it may be worth it to just deal with the outliers individually rather than transforming
the entire data set is there are only a few causing the distortion (Warren, 2021). For a lot of
situations, an outlier can be removed or modified to help normalize the distribution shape and
reduce skewedness
(Warren, 2021).
Researchers could also encounter a Poisson distribution, which offer a good explanation
for uncommon or rare events such as incidences of defectives (Kumarasamy et al., 2025). You
could use them with counts, like how many sessions a client misses in therapy or how many
traffic accidents occur at an intersection in a week’s time. It is the average number of events in
the interval being observed. A researcher may also encounter a zero-inflated negative binomial
(ZINB) distribution, which is basically a lot of zeros showing up. If you look at the number of
times students skip class in a semester, there likely will be a lot of students that never skip class
and that would give multiple zero values and zero would be in excess then. The ZINB model
helps account for all students, the ones that skip and the ones that do not during that time
frame being observed.
Overall, when looking at data and the reasoning for it, researchers must decide which
tools are needed in analyzing the data. Certain models are better for certain needs at the time
than others. However, researchers should still report and disseminate real, unbiased data. This is
DATA SCREENING BASICS 3
where we can move into prompt 2, because data screening and distribution checks are
important to ensure we are reporting honest, accurate and uncompromised information
(Warner, 2021).
Prompt 2
Warner (2021) explains that data screening is needed to ensure the best information is
available to us by preventing compromised data quality and misinterpretation of results. Data
screening helps hold accountability against getting compromised data quality and
misinterpretation of findings and obtaining trustworthy data. Without screening, the small
issues can turn into much larger problems. An example would be when entering data, someone
may enter 300 instead of 30. That extra zero can easily be added and missed, and it can throw
off the averages and make the results look stronger or weaker than they really are. If you look at
it in terms of counseling research, a counselor could be providing a certain type of treatment
intervention, and an error could cause the researcher to come to a misleading conclusion. The
treatment could look far more effective than what it really is or look like it is not effective when
it really is. Data screening helps us catch problems early so that the results reported are as
accurate as possible.
Outliers can cause problems in statistics and are a big reason for data screening. They
can influence parameters, effect size, standard errors, confidence intervals, and tests statistics
(Warner, 2021). A way to identify univariate outliers is with a standard z score greater than 3.29
in absolute value for the distribution in question (Warner, 2021). This rule is based on personal
preference rather than reasoning, but it is often chosen and applied because they make sense in
multiple situations (Warner, 2021). Researchers can easily apply this rule without fear or as
much concern for if it is the right fit for what they are working on, but they should be mindful
still as it may not be the best rule to apply to every situation. However, this rule is a consistent
method to use and will likely serve their purpose well.
DATA SCREENING BASICS 4
Just because an outlier is present does not mean it should be deleted. This should only
take place under certain circumstances and be thoroughly thought through. The specific criteria
for deleting data should be specified before data is ever collected (Warner, 2021). If participants
gave poor quality responses to the data collection compromising the validity, then I could see a
justified reason for deleting that participant. If this happens, it should be reported as well as full
transparency is important. I would recommend being very mindful about this though because if
deletion occurs, the results could change. It could yield acceptable results if the missing data is
less than 5% but an argument could arise with that determination on the accuracy and bias of it
(Warner, 2021). Data screening ensures that analyses are using trustworthy data, prevents
misinterpretations, and promotes transparency. It is more than statistical practice. It can also be
looked at as an ethical responsibility to preserve the integrity if the findings.
DATA SCREENING BASICS 5
References
Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE.
Kumarasamy, P. V., Ramesh, T., Thottathil, A. T., & Thottathil, A. T. (2025). Acceptance—
Rejection criteria based on Birnbaum–Saunders lifetime distribution using intervened
poisson distribution. Quality and Reliability Engineering International., 41(4), 1183–
1194. https://doi.org/10.1002/qre.3719
Warner, R. M. (2021). Applied statistics II: Multivariable and Multivariate Techniques. Sage
Publications. ISBN: 978-1-5443-9872-3
Powered by TCPDF (www.tcpdf.org)
Students also viewed