1
BASIC STATISTICS
Essay 2: Basic Statistics
School of Behavioral Sciences, Liberty University
Author Note
I have no known conflict of interest to disclose.
2
BASIC STATISTICS
Essay 2: Basic Statistics
With its invaluable capacity to convert experiences into quantifiable data that permits
study and discovery, statistical analysis is a fundamental component of quantitative research in
many academic fields (Spiegelhalter, 2019). Though measures of central tendency, such as
means, medians, and modes, summarize data sets, data production is a complex process, and
with outliers in particular, these measures must be interpreted carefully (Condon et al., 2023).
Data points classified as outliers are those that stand out noticeably from the rest of the sample
(Condon et al., 2023; Warner, 2021b). Accurate and relevant conclusions can be drawn using
methods that recognize and control variability and outliers (Field, 2018; Warner, 2021a). Context
regarding inherent uncertainty in statistics is provided by sampling distribution and confidence
intervals (Calin-Jageman & Cumming, 2019; Howitt & Cramer, 2020). Ultimately, data literacy
shapes how researchers use metrics to tell stories and assess problems (Spiegelhalter, 2019).
Professionals from a number of disciplines can perform rigorous, moral work by learning the
fundamentals of statistics concerning central tendency, variability, outliers, and uncertainty
estimation (Spiegelhalter, 2019). When used carefully and competently, statistical data can
do more than just convey numbers; instead, it can facilitate valuable knowledge acquisition and
discoveries (Spiegelhalter, 2019).
Measure of Central Tendency and Sum of Squares
Provide an example of a scenario where the measures of central tendency are skewed as a
result of outliers. How can such a situation be identified and addressed?
When analyzing data sets, researchers use measures of central tendency including the
mean, median, and mode to describe the central position within the data. The mean involves
calculating the sum of all values divided by the number of data points (Warner, 2021a). The
3
BASIC STATISTICS
median represents the midpoint once data is ordered, while the mode indicates the most
frequently occurring value (Howitt & Cramer, 2020). While useful for summarizing data, these
measures can be significantly impacted by the presence of outliers (Field, 2018). In one relevant
example, Pett et al. (2019) discussed a study on client motivation for therapy where researchers
collected ratings on a motivation scale from 1 to 5, indicating from low to high motivation (p.
155). Out of 25 clients, 24 had motivation ratings between 2-4, but one client scored a 1 for the
lowest rating of motivation, an outlier that brought the mean motivation rating down to 2.8 (Pett
et al., 2019). However, the median rating of 3.5 more accurately reflected the central tendency
since most scores fell between 2 and 4 (Pett et al., 2019). Without considering the outlier's
impact, the researcher may draw inaccurate conclusions about typical client motivation levels
(Field, 2018; Warner, 2021a).
Skewed central tendency measures due to outliers can be identified by comparing the
mean, median, and mode, noting that substantial differences between these indicates potential
outliers (Field, 2018). Comparing dispersion measures like standard deviation also identifies
outliers (Warner, 2021a). To address skewed central tendency, researchers can use the median or
mode instead of the mean, transform data, selectively remove outliers, or use statistical methods
like bootstrapping (Howitt & Cramer, 2020; Pett et al., 2019). However, outliers should be
addressed carefully to avoid unnecessary data removal, with the goal to balance retaining
complete data with minimizing outlier distortion for valid interpretation (Field, 2018; Warner,
2021a). Graphical methods like box plots help visualize outliers and can supplement numerical
approaches (Warner, 2021a). By combining statistical tools and thoughtful analysis, researchers
can produce sound conclusions despite outliers (Field, 2018). This is a powerful illustration of a
quote by Nate Silver, “The numbers have no way of speaking for themselves. We speak for them.
4
BASIC STATISTICS
We imbue them with meaning” (Spiegelhalter, 2019, p. 1). This quotation highlights the
importance of not only identifying outliers statistically, but also interpreting their meaning within
the context of the data and research questions (Pett et al., 2019; Spiegelhalter, 2019). Robust
outlier detection integrated with critical analysis empowers researchers to derive meaningful
insights from data (Spiegelhalter, 2019)
Provide an example of what information SS provides us about a set of data. Under what
circumstances will the value of SS equal 0 and is it possible for SS to be negative?
Sum of squares (SS) indicates the variability within a dataset by quantifying the deviation
of each data point from the mean (Field, 2018). For example, in their research on client
motivation, Pett et al. (2019) presented data with an SS value of 3.5, demonstrating some
variability in motivation scores rather than identical levels for all clients. A higher SS would
indicate greater dispersion around the mean (Warner, 2021a). SS equals 0 when there is no
variability, which occurs when all data points are identical to the mean (Howitt & Cramer, 2020).
However, SS being 0 does not necessarily imply an error, just a lack of dispersion (Field, 2018).
In practical terms, this would mean that all participants in a study had identical pre- and post-
intervention anxiety scores, indicating a lack of treatment effect (Warner, 2021a). Finally, SS
cannot be negative, as squaring differences between data points and the mean always produces
positive values (Warner, 2021a). There is no circumstance where SS falls below 0, the lowest
possible measure of variance (Howitt & Cramer, 2020).
Sampling Distribution and Confidence Intervals (CI)
What is a sampling distribution? What does the knowledge of σ contribute to a
researcher’s understanding of the theoretical sampling distribution (regarding its
characteristics and shape)?
5
BASIC STATISTICS
A sampling distribution models the expected distribution of a statistic, such as the mean,
calculated repeatedly across random samples drawn from a population (Warner, 2021a). This
distribution is not just a theoretical construct but a pivotal tool in inferential statistics, allowing
researchers to gauge the variability inherent in statistical estimates derived from samples
(Condon, 2023). The central limit theorem further bolsters this concept, asserting that, regardless
of the population distribution, the sampling distribution of the sample mean will approximate a
normal distribution as the sample size increases (Spiegelhalter, 2019).
Knowledge of the population standard deviation (σ) plays a crucial role in understanding
the sampling distribution. Specifically, σ informs the shape and spread of the sampling
distribution (Howitt & Cramer, 2020). A lower σ suggests that the statistic will exhibit less
variability between samples, leading to a more concentrated distribution. Conversely, a higher σ
indicates greater variability, causing the distribution to spread out more (Pett et al., 2019). This
understanding is vital in counseling research, where variability in treatment outcomes can
significantly impact study interpretations and practical applications (Snider, 2022).
Furthermore, the standard error, derived from σ, quantifies the expected deviation of the
sample mean from the population mean (Warner, 2021a). This metric is essential for constructing
confidence intervals and hypothesis testing, providing a quantitative measure of precision and
reliability in statistical inference (Field, 2018; Warner, 2021a).
Contrast a t distribution from that of the standard normal distribution. In what ways
might N affect the CI?
In the context of confidence intervals, the choice of distribution is pivotal. The t
distribution, with its heavier tails, offers a more accurate estimation of uncertainty in population
parameters, particularly when dealing with small sample sizes typical in specialized fields like
6
BASIC STATISTICS
counseling and psychotherapy (Warner, 2021b). This distribution is preferable over the standard
normal distribution, especially in clinical trials and studies where the sample size is limited, and
the precision of parameter estimates is crucial (Snider, 2022).
The sample size (N) exerts a significant influence on the confidence interval's width
(Field, 2018; Warner, 2021b). Larger sample sizes tend to produce narrower confidence intervals,
reflecting higher precision and confidence in the estimate (Field, 2018). Conversely, smaller
sample sizes result in broader intervals, indicating greater uncertainty and lower confidence in
the estimate (Field, 2018; Warner, 2021b). This inverse relationship between N and confidence
interval width necessitates careful consideration in research design, particularly in counseling
research, where obtaining large samples can be challenging (Pett et al., 2019).
A comprehensive understanding of sampling distributions, including the influence of σ
and the choice between t and standard normal distributions, is imperative for researchers
(Warner, 2021a; Warner, 2021b). This knowledge is instrumental in determining appropriate
sample sizes, interpreting statistical results, and applying these results in practical settings (Field,
2018; Warner 2021b). Furthermore, the consideration of sample size in relation to confidence
interval precision underlines the importance of methodological rigor in research design,
especially in fields dealing with human subjects where ethical considerations and feasibility
often constrain sample sizes (Pett et al., 2019; Warner, 2021a).
Conclusion
Statistical analysis is an invaluable tool across research disciplines, but also requires
diligent application for valid insights. Measures of central tendency can meaningfully summarize
data, however outliers must be properly identified and managed to prevent distortion (Field,
2018; Warner, 2021a). Evaluating variability through sum of squares provides vital context on
7
BASIC STATISTICS
data dispersion (Pett et al., 2019). Likewise, sampling distribution and confidence intervals
characterize the inherent uncertainty in statistics, with parameters like standard deviation shaping
theoretical distribution and interpretation (Calin-Jageman & Cumming, 2019; Howitt & Cramer,
2020; Warner, 2021a). In fields like counseling research where sample sizes may be limited,
selecting appropriate distributions like the t distribution bolsters analysis (Snider, 2022; Warner,
2021b). Ultimately, quantitative literacy enables professionals to interrogate data rigorously yet
responsibly (Gärdenfors & Lombard, 2022; Spiegelhalter, 2019). Researchers must leverage
statistical tools to derive ethical, meaningful insights while acknowledging limitations. Blending
computational competence with thoughtful analysis empowers the discovery of robust evidence
to guide practice.
8
BASIC STATISTICS
References
Calin-Jageman, R. J., & Cumming, G. (2019). The new statistics for better science: Ask how
much, how uncertain, and what else is known. The American Statistician, 73(S1), 271-
280. https://doi.org/10.1080/00031305.2018.1518266
Condon, D. (2023). The mean may not mean what you think it means: The use and misuse of
measures of central tendency. The Journal of Applied Business and Economics, 25(4), 74-
88. https://doi.org/10.33423/jabe.v25i4.6341
Gärdenfors, P., & Lombard, M. (2022). The evolution of human causal cognition. In I. Hodder &
C. Renfrew (Eds.), The Oxford handbook of cognitive archaeology (pp. 406549384).
Oxford University Press. https://doi.org/10.1093/oxfordhb/9780192895950.013.6
Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE Publications.
Howitt, D., & Cramer, D. (2020). Understanding statistics in psychology with SPSS. Pearson.
Pett, M. A., Lackey, N. R., & Sullivan, J. J. (2019). Making sense of factor analysis: The use of
factor analysis for instrument development in health care research. SAGE Publications.
Snider, E (2022). Clinical wisdom in evidence based spiritual care and psychotherapy: What is
it? Pastoral Psychology 71(5), 639–651. https://doi.org/10.1007/s11089-022-01018-y
Spiegelhalter, D. (2019). The art of statistics: Learning from data. Penguin UK.
Warner, R. M. (2021a). Applied statistics I: Basic bivariate techniques (3rd ed.). Thousand Oaks,
CA: Sage Publications.
Warner, R. M. (2021b). Applied statistics II: Multivariable and multivariate techniques. Los
Angeles, CA: Sage Publications.
Powered by TCPDF (www.tcpdf.org)