1 / 10100%
1
DATA SCREENING BASICS
Essay 4: Data Screening Basics
School of Behavioral Sciences, Liberty University
Author Note
I have no known conflict of interest to disclose.
2
DATA SCREENING BASICS
Essay 4: Data Screening Basics
In the realm of counseling and psychotherapy research, the pursuit of evidence-based
practice is paramount (Wampold & Imel, 2015). Effective interventions and treatment strategies
require not only clinical expertise but also a solid foundation in empirical research (Wampold &
Imel, 2015). Delving into the critical domain of data screening, a fundamental facet of empirical
research methodology, we focus specifically on its application within the context of non-normal
distributions which is a frequent occurrence within psychological research.
The investigation revolves around two pivotal themes, each carrying unique significance
within this framework. At the outset, a scrutiny of the practical effectiveness of percentile ranks
and z-scores when confronted with significantly non-normal empirical frequency distributions
illuminates their crucial role in enabling rigorous data analysis (Liu & Zhang, 2019; Mowbray et
al., 2018). Secondly, the complex terrain of data screening yields insights into its far-reaching
implications for data quality, interpretation, and, consequently, the findings of a study (Long-wen
& Zhao, 2022; Owuor et al., 2022).
These two themes intersect at the junction of statistical precision and psychological
inquiry, highlighting their profound relevance to the field of counseling and psychotherapy
research. The exploration is motivated by the aspiration to equip professionals and researchers
alike with essential tools to adeptly address the challenges posed by non-normal data, thereby
reinforcing the empirical integrity that underpins therapeutic practices. Exploring these themes
reveals the essential interplay between rigorous methodological approaches and the advancement
of mental health services. In an era where evidence-based practice is the cornerstone of effective
therapeutic interventions, understanding the nuances of statistical precision and data quality
becomes an indispensable asset in the realm of counseling and psychotherapy research.
3
DATA SCREENING BASICS
Addressing Non-Normal Distributions
If given an extremely non-normal empirical frequency distribution, does it make sense for a
researcher to find the percentile rank using the z score and the standard normal
distribution table? Why or why not?
In the realm of psychotherapy research, the issue of non-normal data distributions is a
recurring challenge (Warner, 2021a). Non-normal distributions significantly deviate from the
conventional bell-shaped curve that characterizes a normal distribution (Warner, 2021a). In
standard statistics, tools like the z score and the standard normal distribution table are tailored for
data that closely approximates a normal distribution, where the mean, median, and mode are
equal, and the data forms a symmetric pattern (Warner, 2021a). However, in psychotherapy
research, data often takes on non-normal characteristics due to the complexity of human behavior
and the intricate nature of psychological constructs (Wampold & Imel, 2015). Patient outcomes,
for instance, can exhibit skewed distributions and irregular patterns (Long-wen & Zhao, 2022).
Patient populations within psychotherapy research are diverse, each with its unique
characteristics and backgrounds. This diversity can lead to the emergence of non-normal
distributions of outcomes (Wampold & Imel, 2015). When dealing with such non-normal data,
researchers must explore alternative approaches to data analysis (Warner, 2021a). Conventional
methods like using the z score and standard normal distribution table may not be appropriate in
these circumstances due to their underlying assumptions of normality (Warner, 2021a).
Describe two distributions (other than normal) that a researcher might encounter in data
and when.
In psychotherapy research, two common non-normal distributions that researchers might
encounter are skewed distributions and heavy-tailed distributions (Wampold & Imel, 2015;
4
DATA SCREENING BASICS
Long-wen & Zhao, 2022). Skewed distributions are characterized by an asymmetry in their
shape, with the data points clustering more towards one tail than the other (Wampold & Imel,
2015). This phenomenon can be observed when studying various psychological phenomena, such
as the distribution of patients' response to a particular therapy (Long-wen & Zhao, 2022). Heavy-
tailed distributions, on the other hand, exhibit a higher frequency of extreme values or outliers
than a normal distribution would predict (Wampold & Imel, 2015). These distributions often
emerge when dealing with complex psychological constructs or diverse patient groups in
psychotherapy research (Long-wen & Zhao, 2022).
For instance, when examining the effectiveness of trauma-focused therapy on post-
traumatic stress disorder (PTSD) symptoms in a group of veterans, heavy-tailed distributions
may appear due to variations in the treatment responses among participants (Wampold & Imel,
2015; Long-wen & Zhao, 2022). The intricate nature of psychological constructs, coupled with
the diversity of patients and their unique backgrounds, contributes to the emergence of these non-
normal distributions in psychotherapy research (Wampold & Imel, 2015). Understanding these
distinct non-normal distributions is crucial for researchers to select appropriate statistical
methods and ensure the robustness of their analyses in the field of mental health services and
therapy outcomes (Long-wen & Zhao, 2022; Warner, 2021a; Wampold & Imel, 2015). This
comprehensive exploration of non-normal data distributions in psychotherapy research highlights
the importance of tailored statistical approaches to accommodate the unique characteristics of the
data (Warner, 2021a; Wampold & Imel, 2015; Long-wen & Zhao, 2022).
The Impact of Data Screening on Research Quality
How might the absence of data screening affect a researcher’s data quality, their
interpretations of the data, and thereby their interpretations of their study’s findings?
5
DATA SCREENING BASICS
Data screening is a critical step in ensuring the integrity and reliability of research
findings in any type of research (Warner, 2021a). It involves the initial examination of collected
data to identify errors, inconsistencies, outliers, or any other issues that might affect the quality
and reliability of the data (Warner, 2021a). The primary goal of data screening is to ensure that
the data used for analysis is accurate, complete, and suitable for the intended research or
statistical procedures (Warner, 2021a). Without proper checks and assessments carried out in data
screening, researchers risk compromising the quality of their data, which can have far-reaching
implications for the study's outcomes and interpretations (Warner, 2021a). For instance, in the
absence of data screening, incomplete or inaccurate data entries may go unnoticed, leading to
missing or erroneous information that skew the results of the study (Warner, 2021a). In
psychotherapy research, where the reliability and validity of data are paramount, overlooking
data screening can undermine the entire research process.
Moreover, failing to screen data can also affect the distribution of variables in the dataset.
Non-normal data distributions, as discussed earlier, are common in psychotherapy research
(Wampold & Imel, 2015). Without data screening, researchers may not identify the presence of
skewed or heavy-tailed distributions, which can mislead subsequent statistical analyses (Warner,
2021a). This, in turn, can impact the validity of the interpretations drawn from the data.
The consequences of inadequate data screening also extend to the interpretations of the
research findings. Misleading or flawed data can lead to incorrect conclusions and interpretations
of the study's results (Warner, 2021a). In psychotherapy research, where the findings often
inform clinical practice and interventions, inaccurate interpretations can have significant
implications for the delivery of mental health services (Wampold & Imel, 2015). Researchers and
clinicians may make decisions based on flawed findings, potentially jeopardizing the well-being
6
DATA SCREENING BASICS
of clients. To maintain the highest standards of empirical integrity, data screening is an essential
step that researchers should not overlook (Warner, 2021a).
What quantitative rule may be used to determine univariate outliers, and are there
situations in which deleting a case/participant may be justified? Explain.
Univariate outliers, or extreme data points that deviate substantially from the rest of the
data in a single variable, are a common concern in psychotherapy research (Warner, 2021b;
Mowbray et al., 2018). Detecting and handling these outliers is essential to maintain the accuracy
and reliability of research findings (Warner, 2021b). One quantitative rule often employed to
identify univariate outliers is the use of z-scores (Warner, 2021b; Mowbray et al., 2018). Z-
scores express how many standard deviations a data point is away from the mean of the
distribution. Typically, data points with z-scores exceeding a certain threshold, such as ±2 or ±3
standard deviations, are considered potential outliers (Warner, 2021b; Mowbray et al., 2018).
Identifying these outliers helps researchers pinpoint extreme observations that may skew the
overall data analysis.
The decision to delete a case or participant, however, based on the presence of univariate
outliers should be made judiciously and with a solid rationale (Warner, 2021b). Deleting data
points can impact the representativeness of the sample and the generalizability of the findings
(Warner, 2021b). It is generally advisable to thoroughly examine the outliers' nature and context
before making decisions such as this. Justifying the deletion of a case or participant may be
warranted in specific situations (Warner, 2021b; Mowbray et al., 2018). For instance, if a data
point is identified as an outlier due to a data entry error or an instrument malfunction, removing
it may enhance data accuracy (Warner, 2021b). However, the decision should be transparently
7
DATA SCREENING BASICS
documented in the research report, and sensitivity analyses should be conducted to assess the
robustness of the findings with and without the outliers (Warner, 2021b).
In psychotherapy research, where the quality of data directly impacts the validity of
therapeutic interventions, handling univariate outliers with care and complete transparency is
crucial (Wampold & Imel, 2015; Warner, 2021b). Researchers must prioritize data accuracy
while simultaneously safeguarding the integrity of their study samples and the validity of their
findings. This is imperative because the preceding discussions, such as discussion of p-hacking,
have underscored the significance of obtaining accurate and replicable results, which, in turn, are
crucial for advancing knowledge in critical research domains (Warner, 2021b; Zheng et al., 2020)
Conclusion
Through the examination of the intricate domains of non-normal data distributions and
data screening, it becomes evident that meticulous methodological precision is paramount in the
research process (Wampold & Imel, 2015; Warner, 2021a). In the context of research, delving
into the challenges posed by extremely non-normal empirical frequency distributions reveals the
limitations of traditional statistical tools like percentile ranks and z-scores when assumptions of
normality are violated (Warner, 2021a). As is often the case, research data may deviate
significantly from the bell-shaped curve of a normal distribution, primarily due to the diverse
characteristics of the studied populations and the complexity of the variables under
investigation (Wampold & Imel, 2015).
This exploration underscores the necessity of alternative data analysis approaches, such
as bootstrapping and non-parametric tests, which are better suited to handle non-normal data
structures and ensure the robustness of research findings (Wampold & Imel, 2015; Liu &
Zhang, 2019). By embracing these methods, researchers across various fields can enhance the
8
DATA SCREENING BASICS
validity and generalizability of their results, ultimately contributing to the advancement of
evidence-based practices in their respective domains (Wampold & Imel, 2015). Simultaneously,
the examination of the domain of data screening highlights the paramount importance of data
accuracy and integrity. The necessity of meticulous data screening procedures to identify and
rectify errors, outliers, and inconsistencies that could compromise the quality of research
outcomes has been emphasized (Warner, 2021b). Data screening serves not only as a technical
task but also as a foundational step that safeguards the reliability and reproducibility of research
results (Warner, 2021b).
In the context of counseling and psychotherapy research, where the well-being of
individuals is at stake, the implications of rigorous data screening and robust statistical analysis
cannot be overstated (Zheng et al., 2020). It is through these meticulous practices that
researchers uphold the empirical integrity that underpins therapeutic practices, ensuring that
evidence-based interventions and treatment strategies are built upon a solid foundation of
reliable data (Wampold & Imel, 2015; Zheng et al., 2020).
In sum, this exploration has shed light on the critical interplay between statistical
precision and research methodology, offering insights into how researchers can navigate the
complexities of non-normal data and uphold the highest standards of data quality. As research in
diverse fields continues to evolve, it is essential that researchers equip themselves with the
essential tools and methodologies to meet the unique challenges posed by non-normal data
distributions, ultimately advancing the knowledge base, and enhancing the quality of research
conducted across various domains.
9
DATA SCREENING BASICS
References
Liu, H. and Zhang, Z. (2019). A method of generating multivariate non-normal random numbers
with desired multivariate skewness and kurtosis. Behavior Research Methods, 52(3), 939-
946. https://doi.org/10.3758/s13428-019-01291-5
Long-wen, Z. and Zhao, Y. (2022). Hut‐based method for structural reliability considering the
non‐normal and unknown distributions. Quality and Reliability Engineering
International, 38(5), 2303-2323. https://doi.org/10.1002/qre.3076
Mowbray, F., Fox-Wasylyshyn, S., & El-Masri, M. (2018). Univariate outliers: a conceptual
overview for the nurse researcher. Canadian Journal of Nursing Research, 51(1), 31-37.
https://doi.org/10.1177/0844562118786647
Owuor, O., Benedict, T., & Kevin, O. (2022). Outlier detection technique for univariate normal
datasets. American Journal of Theoretical and Applied Statistics, 11(1), 1.
https://doi.org/10.11648/j.ajtas.20221101.11
Schrempp, M. (2018). Limit laws for the diameter of a set of random points from a distribution
supported by a smoothly bounded set. Extremes, 22(1), 167-191.
https://doi.org/10.1007/s10687-018-0309-9
Wampold, B. E., & Imel, Z. E. (2015). The great psychotherapy debate: The evidence for what
makes psychotherapy work (2nd ed.). Routledge/Taylor & Francis Group.
Warner, R. M. (2021a). Applied statistics I: Basic bivariate techniques (3rd ed.). Thousand Oaks,
CA: Sage Publications.
Warner, R. M. (2021b). Applied statistics II: Multivariable and multivariate techniques. Los
Angeles, CA: Sage Publications.
10
DATA SCREENING BASICS
Zheng, M., Marsh, J. K., Nickerson, J. V., & Kleinberg, S. (2020). How causal information
affects decisions. Cognitive Research: Principles and Implications, 5(6), 1-24.
https://doi.org/10.1186/s41235-020-0206-z
Students also viewed