Week 4 - Assignment: Apply the Normal Distribution
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 1/7
Week 4
BUS-7105 v3: Statistics I (7103872203)
Normal Distribution and the Central Limit Theorem
Determining whether or not a phenomenon exists, in a statistical sense, is based on the
principles of probability.
Given a set of data gathered during a study we might ask, “What is the probability of X?”
We use this decision rule to analyze the data.
More specifically, say a manager is interested in her workers’ intentions to quit their jobs.
Knowing this is important to planning staffing needs. She might suggest, based on
observation, or perhaps a literature review, that given the climate of the work
environment (e.g., hours, how demanding the jobs are, etc.) 20% of her employees would
quit their jobs if a reasonable alternative surfaced. She could ask her workers,
anonymously, if they are currently searching for another position and compare the results
to see if her 20% prediction is accurate. Armed with those data she might initiate training
to better support her workers.
Traditional parametric statistics tools are based on the assumption that all data are
“normally distributed.” In other words, they fall under a bell curve where there is the
probability that a few instances occur at the high end of the scale and a few at the low
end with the remaining instances occurring around the middle.
For example, the manager above may be interested also in levels of performance at work.
She may define this as the number of performance errors that her workers make on a
monthly basis. Her hypothesis might be that her workers’ performance is normally
distributed. She could observe her workers over a month and plot their error rates. If her
workers' error rates are “normally distributed,” there would be a few who commit more
than the average number of errors and a few who commit almost no errors. The remaining
people would fall around the average or the high point of the curve.
The normal distribution is perhaps the most common distribution used in business
applications. It is easy to recognize a normal distribution because when a histogram is
constructed from normally distributed data, the shape is "bell-shaped." The distribution
occurs when studying characteristics of people and animals, such as height, weight, and
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 2/7
IQ. It also arises when studying measurement errors as well as in many theoretic
situations related to hypothesis testing.
In the real world of research and data analysis, our samples seldom fall within a normal
distribution. However, the rules of statistics will work even in these skewed samples.
Most behavioral phenomena are normally distributed at the population level even if
samples may be somewhat skewed (e.g., too many outliers on one end, or tail, of the
distribution). Given the reality of a normal distribution at the population level, we can
generally trust our findings even in a skewed sample, provided it was randomly
assembled.
The normal distribution also arises in many business applications. For example, if a
histogram is made of the daily percentage changes in stock prices, it is usually normal. In a
production process, suppose that measurements are taken of a critical aspect, and then a
histogram is created. If the process is working properly, then the measurements should
be normally distributed. But if too much variation, such as poor machine adjustments, or
untrained operators, causes the process to produce defects, this can be detected in a
histogram that distorts a bell-shaped curve.
The following images are depictions of four common distributions found in sample
research. We noted in the introduction to Section 2 that distributions are fairly normal at
the population level, but not so much at the sample level. The images on the left are
histograms from datasets drawn from samples that follow somewhat normal distributions.
The image on the upper right represents something close to bimodal or having two
modes. It is possible that a dataset can have many modes. The histogram on the bottom
right represents a positive or right skew. The distribution is skewed to the right because
of a range of outlying values at the high, or positive, end of the scale. It is also possible to
have a left skew for which outliers at the low, or negative, end of the scale are pulling the
tail of the distribution to the left. An example of a right skew is measuring average salary
across an entire firm and the CEO is paid ten times the rest of the workers. An example of
a left skew is a typical grading distribution for a graduate classroom. Most people who
enroll in graduate programs are motivated and want to do well so they, collectively, earn a
lot of A’s and B’s. Typically grades of C and lower are not considered passing so there is
motivation to do well. Only in extreme instances do we see C’s, D’s, and F’s in the
graduate classroom.
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 3/7
Figure 3. Four common distributions found in sample research
One of the most important and best-known facts about a normal distribution is something
called the "Empirical Rule." The rule gives the approximate areas beneath the normal
curve. Given such a curve we assume that the total area under the curve is equal to 1 or
100%. The rule tells us that, given a normal (bell-shaped) distribution with mean m and
standard deviation s:
68.3% of the total area is between m - s and m + s
95.4% of the total area is between m - 2s and m + 2s
99.7% of the total area is between m - 3s and m + 3s
So, in a normal distribution, the standard deviation has a very specific meaning, as noted
above. But in reality, when working with non-normally distributed sample data, the
standard deviation is less readily interpretable and more useful in informing other
hypothesis testing formulas.
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 4/7
Figure 4. The normal distribution and the empirical rule.
The Central Limit Theorem tells us that a sampling distribution of the mean for an
independent and random variable will be normal if the sample size is large enough. In
other words, if we take a population of 100,000 people, for example, and pull every
possible sample of 30 from it, calculate the means of each sample and plot them under a
distribution curve, the result will be a near-normal distribution. There is plenty of
evidence that the math will work correctly. This principle is what gives us confidence that
traditional statistical tools will function meaningfully in sample research.
The question then becomes, how large a sample is large enough to make statistics work?
The answer to this depends on two circumstances. First, we must ask whether or not the
population is normally distributed? Research findings indicate that personality, and most
attitudinal, measures are near-normal at the population level. This may or may not hold
true for other measurable phenomena. Secondly, there are requirements for how
accurately the sample resembles the population. Random selection or assignment can
help with this, but the reality is that in sample research we often use who and what we
have available. Therefore, the question becomes more salient.
Most statisticians and textbook authors will indicate that a sample size of 30 (n=30) is an
adequate sample size when pulling from a population that is near-normally distributed. If
we know the population to be non-normal, we should then gather as many subjects as
possible to mitigate this.
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 5/7
Books and Resources for this Week
Frank, J., & Klar, B. (2016). Methods to
test for equality of two normal
distributions. Statistical Methods and
Applications, 25(4), 581-599. Link
Central Limit Theorem (CLT)Central Limit Theorem (CLT)
Mean Median Mode
Launch in a separate window
Be sure to review this week's resources carefully. You are expected to apply the
information from these resources when you prepare your assignments.
80 % 4 of 5 topics complete
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 6/7
Li, J. C.-H. (2016). Effect size measures
in a two-independent-samples case
with nonnormal and nonhomogeneous
data. Behavior Research Methods... Link
Introduction to Business Statistics (7th
ed.) External Learning Tool
NCU School of Business Best Practice
Guide for Quantitative Research Design
and Methods in Dissertationse Link
Week 4 - Assignment: Apply the Normal Distribution Assignment
Due June 26 at 11:59 PM
For this week’s assignment, you will present your answers to the following questions in a
formal paper format. Please separate each question with a short heading, e.g., normal
curve, bell curve, etc.
Begin with a brief introduction in which you explain the importance of normal
distribution.
Next, address the following questions in order:
Describe the characteristics of the normal curve and explain why the curve, in
sample distributions, never perfectly matches the normal curve.
Why is the bell curve used to represent the normal distribution? Why not a
different shape?
Why is the central limit theorem important in statistics?
What does the central limit theorem inform us about the sampling distribution of
the sample means?
Imagine that you recently took an exam for certification in your field. The certifying
agency has published the results of the exam and 75% of the test takers in your
group scored below the average. In a normal distribution, half of the scores would
6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)
https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 7/7
fall above the mean and the other half below. How can what the certifying agency
published be true?
Why do researchers use z-scores to determine probabilities? What are the
advantages to using z-scores?
Conclude with a brief discussion of how the concept of probability might affect research
that you might undertake in your dissertation project. In other words, how would a basic
understanding of probability concepts aid you in analyzing and interpreting data?
Length: 4 to 6 pages not including title page and reference page.
References: Include a minimum of 3 scholarly resources.
Your paper should demonstrate thoughtful consideration of the ideas and concepts
presented in the course and provide new thoughts and insights relating directly to this
topic. Your response should reflect scholarly writing and current APA standards. Be sure
to adhere to Northcentral University's Academic Integrity Policy.
Upload your document and click the Submit to Dropbox button.