Week 4 - Assignment: Apply the Normal Distribution

profileFila64
BUS-7105v3_StatisticsI7103872203-BUS-7105v3_StatisticsI7103872203.pdf

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 1/7

Week 4

BUS-7105 v3: Statistics I (7103872203)

Normal Distribution and the Central Limit Theorem

Determining whether or not a phenomenon exists, in a statistical sense, is based on the

principles of probability.

Given a set of data gathered during a study we might ask, “What is the probability of X?”

We use this decision rule to analyze the data.

More specifically, say a manager is interested in her workers’ intentions to quit their jobs.

Knowing this is important to planning staffing needs. She might suggest, based on

observation, or perhaps a literature review, that given the climate of the work

environment (e.g., hours, how demanding the jobs are, etc.) 20% of her employees would

quit their jobs if a reasonable alternative surfaced. She could ask her workers,

anonymously, if they are currently searching for another position and compare the results

to see if her 20% prediction is accurate. Armed with those data she might initiate training

to better support her workers.

Traditional parametric statistics tools are based on the assumption that all data are

“normally distributed.” In other words, they fall under a bell curve where there is the

probability that a few instances occur at the high end of the scale and a few at the low

end with the remaining instances occurring around the middle.

For example, the manager above may be interested also in levels of performance at work.

She may define this as the number of performance errors that her workers make on a

monthly basis. Her hypothesis might be that her workers’ performance is normally

distributed. She could observe her workers over a month and plot their error rates. If her

workers' error rates are “normally distributed,” there would be a few who commit more

than the average number of errors and a few who commit almost no errors. The remaining

people would fall around the average or the high point of the curve.

The normal distribution is perhaps the most common distribution used in business

applications. It is easy to recognize a normal distribution because when a histogram is

constructed from normally distributed data, the shape is "bell-shaped." The distribution

occurs when studying characteristics of people and animals, such as height, weight, and

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 2/7

IQ. It also arises when studying measurement errors as well as in many theoretic

situations related to hypothesis testing.

In the real world of research and data analysis, our samples seldom fall within a normal

distribution. However, the rules of statistics will work even in these skewed samples.

Most behavioral phenomena are normally distributed at the population level even if

samples may be somewhat skewed (e.g., too many outliers on one end, or tail, of the

distribution). Given the reality of a normal distribution at the population level, we can

generally trust our findings even in a skewed sample, provided it was randomly

assembled.

The normal distribution also arises in many business applications. For example, if a

histogram is made of the daily percentage changes in stock prices, it is usually normal. In a

production process, suppose that measurements are taken of a critical aspect, and then a

histogram is created. If the process is working properly, then the measurements should

be normally distributed. But if too much variation, such as poor machine adjustments, or

untrained operators, causes the process to produce defects, this can be detected in a

histogram that distorts a bell-shaped curve.

The following images are depictions of four common distributions found in sample

research. We noted in the introduction to Section 2 that distributions are fairly normal at

the population level, but not so much at the sample level. The images on the left are

histograms from datasets drawn from samples that follow somewhat normal distributions.

The image on the upper right represents something close to bimodal or having two

modes. It is possible that a dataset can have many modes. The histogram on the bottom

right represents a positive or right skew. The distribution is skewed to the right because

of a range of outlying values at the high, or positive, end of the scale. It is also possible to

have a left skew for which outliers at the low, or negative, end of the scale are pulling the

tail of the distribution to the left. An example of a right skew is measuring average salary

across an entire firm and the CEO is paid ten times the rest of the workers. An example of

a left skew is a typical grading distribution for a graduate classroom. Most people who

enroll in graduate programs are motivated and want to do well so they, collectively, earn a

lot of A’s and B’s. Typically grades of C and lower are not considered passing so there is

motivation to do well. Only in extreme instances do we see C’s, D’s, and F’s in the

graduate classroom.

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 3/7

Figure 3. Four common distributions found in sample research

One of the most important and best-known facts about a normal distribution is something

called the "Empirical Rule." The rule gives the approximate areas beneath the normal

curve. Given such a curve we assume that the total area under the curve is equal to 1 or

100%. The rule tells us that, given a normal (bell-shaped) distribution with mean m and

standard deviation s:

68.3% of the total area is between m - s and m + s

95.4% of the total area is between m - 2s and m + 2s

99.7% of the total area is between m - 3s and m + 3s

So, in a normal distribution, the standard deviation has a very specific meaning, as noted

above. But in reality, when working with non-normally distributed sample data, the

standard deviation is less readily interpretable and more useful in informing other

hypothesis testing formulas.

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 4/7

Figure 4. The normal distribution and the empirical rule.

The Central Limit Theorem tells us that a sampling distribution of the mean for an

independent and random variable will be normal if the sample size is large enough. In

other words, if we take a population of 100,000 people, for example, and pull every

possible sample of 30 from it, calculate the means of each sample and plot them under a

distribution curve, the result will be a near-normal distribution. There is plenty of

evidence that the math will work correctly. This principle is what gives us confidence that

traditional statistical tools will function meaningfully in sample research.

The question then becomes, how large a sample is large enough to make statistics work?

The answer to this depends on two circumstances. First, we must ask whether or not the

population is normally distributed? Research findings indicate that personality, and most

attitudinal, measures are near-normal at the population level. This may or may not hold

true for other measurable phenomena. Secondly, there are requirements for how

accurately the sample resembles the population. Random selection or assignment can

help with this, but the reality is that in sample research we often use who and what we

have available. Therefore, the question becomes more salient.

Most statisticians and textbook authors will indicate that a sample size of 30 (n=30) is an

adequate sample size when pulling from a population that is near-normally distributed. If

we know the population to be non-normal, we should then gather as many subjects as

possible to mitigate this.

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 5/7

Books and Resources for this Week

Frank, J., & Klar, B. (2016). Methods to

test for equality of two normal

distributions. Statistical Methods and

Applications, 25(4), 581-599. Link

Central Limit Theorem (CLT)Central Limit Theorem (CLT)

Mean Median Mode

Launch in a separate window

Be sure to review this week's resources carefully. You are expected to apply the

information from these resources when you prepare your assignments.

80 % 4 of 5 topics complete

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 6/7

Li, J. C.-H. (2016). Effect size measures

in a two-independent-samples case

with nonnormal and nonhomogeneous

data. Behavior Research Methods... Link

Introduction to Business Statistics (7th

ed.) External Learning Tool

NCU School of Business Best Practice

Guide for Quantitative Research Design

and Methods in Dissertationse Link

Week 4 - Assignment: Apply the Normal Distribution Assignment

Due June 26 at 11:59 PM

For this week’s assignment, you will present your answers to the following questions in a

formal paper format. Please separate each question with a short heading, e.g., normal

curve, bell curve, etc.

Begin with a brief introduction in which you explain the importance of normal

distribution.

Next, address the following questions in order:

Describe the characteristics of the normal curve and explain why the curve, in

sample distributions, never perfectly matches the normal curve.

Why is the bell curve used to represent the normal distribution? Why not a

different shape?

Why is the central limit theorem important in statistics?

What does the central limit theorem inform us about the sampling distribution of

the sample means?

Imagine that you recently took an exam for certification in your field. The certifying

agency has published the results of the exam and 75% of the test takers in your

group scored below the average. In a normal distribution, half of the scores would

6/23/22, 3:09 PM BUS-7105 v3: Statistics I (7103872203) - BUS-7105 v3: Statistics I (7103872203)

https://ncuone.ncu.edu/d2l/le/content/258948/printsyllabus/PrintSyllabus 7/7

fall above the mean and the other half below. How can what the certifying agency

published be true?

Why do researchers use z-scores to determine probabilities? What are the

advantages to using z-scores?

Conclude with a brief discussion of how the concept of probability might affect research

that you might undertake in your dissertation project. In other words, how would a basic

understanding of probability concepts aid you in analyzing and interpreting data?

Length: 4 to 6 pages not including title page and reference page.

References: Include a minimum of 3 scholarly resources.

Your paper should demonstrate thoughtful consideration of the ideas and concepts

presented in the course and provide new thoughts and insights relating directly to this

topic. Your response should reflect scholarly writing and current APA standards. Be sure

to adhere to Northcentral University's Academic Integrity Policy.

Upload your document and click the Submit to Dropbox button.