DATA SCREENING 2
If Given an Extremely Non-Normal Empirical Frequency Distribution, Does It Make Sense
for a Researcher to Find the Percentile Rank Using the Z Score and the Standard Normal
Distribution Table? Why or Why Not? Describe Two Distributions (Other Than Normal)
That A Researcher Might Encounter in Data and When?
Data screening is carried out to ensure that the best information is available when the
results are presented. Finding issues with the way data is examined and impacted, like a non-
normal empirical frequency distribution, can be aided by data screening. A1non-normal empirical
frequency distribution does not follow a bell curve, has extreme values, and lacks symmetry
(Sainhani, 2012). The standard normal distribution, on the other hand, is always centered at zero
and has intervals that increase at one since it has a mean of zero and a standard deviation of one
(Warner, 2020a). A z score is represented by each integer on the horizontal axis (Liberty
University, 2025). The number of standard deviations is indicated by a z score. A z score
indicates how many standard deviations the data points are from the mean (Liberty University,
2025). For example, a z score of negative one is one standard deviation from the right of the
mean. Most importantly, a z score calculates how much area a z score is associated with, and this
can be done using a frequency distribution table (Liberty University, 2025). Since z scores are
specifically made to operate with normally distributed data, applying them to non-normal data
can lead to false interpretations (Andrades & Scores, 2021). In other words, as the name
suggests, the standard normal distribution table assumes a normal distribution (Warner, 2020a).
Some studies suggest converting z scores to standard scores to fix this problem. Changing a z
score to a standard score may allow for an easier reading of the data; however, it does not fix the
problem of non-normal distribution; it simply transforms the data to have a mean of 0 and a
standard deviation of 1, but the shape of the distribution remains the same, meaning the data
remains skewed (Andres & Scores, 2021; Emerson, 2017).
DATA SCREENING 3
Because of the aforementioned, it is not recommended to use the z score and the
conventional normal distribution table to determine the percentile rank of a piece of data.
Additionally, a researcher may come across distributions such as Poisson and binomial in data
because of the deviation from the normal distribution. Zero-inflated distributions, like Poisson
and binomial, exhibit a higher number of zero values compared to what a typical pattern would
predict (Warner, 2020b). For instance, more than 30% of the sample said they had never used
marijuana; among those who had, the frequency of use varied, with many claiming to use it nine
or more times per month. This distribution's non-normal form is evident (Warner, 2020b). The
Poisson distribution is suitable in situations where the data is distinct (counts) and not symmetric
like a normal distribution since it may be used to predict the number of events that occur during a
given time period (Warner, 2020a). In addition to being a crucial distribution in probability
theory, the Poisson distribution is also a helpful tool for examining random events, such as how
frequently an event occurs over time (Zhang et al., 2021). The binomial distribution is used when
there are only two possible outcomes for an event (like flipping a coin - heads or tails) (Warner,
2020a).1A binomial distribution can help identify the presence of outliers because it shows two
different groups (Ascari & Migliorati, 2021). "An outlier is a score that deviates significantly
from the mean" (Warner, 2020a, p. 144). Because extreme values affect the overall pattern of the
data, outliers cause a lot of issues in statistical analysis by causing erroneous inferences to be
formed from the analysis.
How Might the Absence of Data Screening Affect a Researcher’s Data Quality,
Their Interpretations of the Data, and Thereby Their Interpretations of Their Study’s
Findings? What Quantitative Rule May Be Used to Determine Univariate Outliers, and
Are there Situations in Which Deleting a Case/Participant May Be Justified? Explain.
7Based on the above information it can be seen that outliers can present a problem in
research. Indeed, how research is interpreted and explained also involves dealing with errors that
DATA SCREENING 4
can compromise findings. Data screening is crucial in interpreting findings. The absence of data
screening has significant impacts on the quality of data as it can lead to incorrect findings being
presented and overall diminishes improving the quality of data to prevent issues such as outliers
and missing values that can lead to problems in the research (Warner, 2020b). Outliers and
missing values can lead to bias in the research (Warner, 2020b). For example, while research
articles can be published with outliers and missing values, both can create concerns in creating
public-use data files and procedures to address these problems have been implemented (She &
Wu, 2019).
In research, the existence of statistical outliers is a common worry. Outliers have the
ability to skew the estimate of the relevant parameter and jeopardize the generalizability of study
results if they are disregarded or handled incorrectly (Wasylyshyn 7 El-Masri, 2019).
Researchers can use a range of statistical methods to help with the identification (Wasylyshyn &
El-Masri, 2019). The "1.5 times the Interquartile Range (IQR)" rule is the most widely used
quantitative criteria for identifying univariate outliers; the median + 3 × IQR and median – 3 ×
IQR are typically used to define the outside gates. The outer fences often contain 99 percent of
the scores. Outliers and extreme outliers can be found using boxplots' inner and outer fences
(Warner, 2020a). Outliers are identified if any of the scores are higher than the upper fence;
extreme outliers are identified if the scores are more than ±3 times the IQR from the median
(Warner, 2020a).
Outliers can misrepresent the results of an analysis if not properly addressed. Another
issue in research that can impact the results of a study is missing data. Missing data can
complicate research in several ways such as making a smaller sample size even smaller with
missing values (Warner, 2020b). Listwise deletion, which eliminates every observation with a
DATA SCREENING 5
missing value for any variable in the analysis and may result in a biased sample, can also be
caused by missing values. If, for instance, students with low grades are removed from a sample
utilized in the analysis because they declined to respond to certain questions regarding grades,
the remaining sample will primarily consist of students with higher grades. Due to bias, the
sample will not include responses from students with lower grades (Warner, 2020b).
There are situations where deleting a case or participant may be justified such as if there
is significant amount of missing data, it may be justified to remove this case to ensure the
integrity of the data analysis (Warner, 2020b). Nonetheless, removing a case or deleting data is
dependent on the type of missingness pattern. Missingness patterns usually fall between Type A
and Type B missingness (Warner, 2020b). Type A missingness refers to missing completely at
random and Type B missingness describes missing at random (Warner, 2020b). In Type A
missingness one variable can influence the other which can explain the missing value (Warner,
2020b). For example, suppose the Y axis representing the variable depression has missing
answers; the data for the other X variables such as sex, can affect the missing value as men may
refuse to answer the depression question (Warner, 2020b). The common approach is to remove
the rows that are missing (a listwise deletion) as the data is considered random, yet if a large
portion of data is missing it can negatively influence the data results. Data imputation techniques
like replacing large missing values with the mean, median is a better strategy for managing
missing values (Warner, 2020b).
·
DATA SCREENING 6
References
Ascari, R., & Migliorati, S. (2021). A new regression model for over dispersed binomial data
accounting for outliers and an excess of zeros. Statistics in Medicine, 40(17).
https://doi.org/10.1002/sim.9005
Emerson, R. W. (2017). Distribution of Scores around the Mean and the Purpose of z scores.
Journal of Visual Impairment & Blindness, 111(1), 90-92.
https://doi.org/10.1177/0145482X1711100111
Liberty University. (2025, Winter). EDCO:735: Statistics. Week four, lecture four: Z scores,
Standardizations, and the normal deviation.1https://learn.liberty.edu
Sainhani, K.L. (2012). Dealing with non-normal data. American Academy of Physical Medicine
and Rehabilitation, 4, 1001-1005. https://dx.doi.org/10.1016/j.pmrj.2012.10.013
She, X. & Wu, C. (2019). Validity and efficiency in analyzing ordinal responses with missing
observations.1Canadian Journal of Statistics, 48(2), 138-151.
https:doi1.org/10.1002/cjs.11523
Zhao, J., Zhang, F., Zhao, C., Wu, G.,Wang, H., & Cao, Xinyu. (2020). The properties and
application of Poisson distribution. Journal of Physics: Conference Series, 1550.
https://doi: 032109. 10.1088/1742-6596/1550/3/032109
Warner, R. M. (2020a).1Applied Statistics I1(3rd ed.). SAGE Publications, Inc.
Warner, R. M. (2020b). Applied Statistics II (3rd ed.). SAGE Publications, Inc.
Wasylyshyn, S.M., El-Masri, M.M. (2019). Univariate outliers: A conceptual overview for the
nurse researcher.1Canadian Journal of Nursing Research, 51(1):31-37.
https://doi:10.1177/0844562118786647