Statistics for Managers 2

profile5ToGo
bus308_chapter_03.pdf

3

Applying z to Groups

Learning Objectives

After reading this chapter, you should be able to:

• Describe the distribution of sample means.

• Explain the central limit theorem.

• Calculate and explain z-test results.

• Explain statistical significance.

• Determine the sample size required for a particular analysis.

• Calculate the z-test using Excel.

• Explain how decision errors can affect statistical analysis.

iStockphoto/Thinkstock

tan81004_03_c03_051-074.indd 51 2/22/13 3:31 PM

CHAPTER 3Section 3.1 Expanding the z Score Discussion

Chapter Outline

3.1 Expanding the z Score Discussion

3.2 The Distribution of Sample Means The Central Limit Theorem Sampling Error

3.3 The z-test Calculating the z-Test Interpreting the Value From the z-Test Another z-Test

3.4 Statistical Significance Determining Statistical Significance Interpreting z Another View of Significance Sampling Error as an Explanation of Difference More Confidence in the Sample Decision Errors

3.5 The z-Test Using Excel

Chapter Summary

3.1 Expanding the z Score Discussion

The z scores that were calculated in Chapter 2 allow someone to ask how an individual compares to a group. Knowing sales figures for the entire sales force, a manager can rely on z scores to ask what the likelihood is that one individual will have sales that exceed a particular value, or fall below a specified value, or occur between two values. However, in business settings the more important questions tend to be about groups rather than indi- viduals. For example, a manager is more likely to be curious about how the performance of a particular sales team compares to all salespersons’ than how one individual performs compared to a group. This chapter extends the z score discussion from Chapter 2, where individuals were compared to groups, to an analysis of how groups compare to popula- tions. Groups’ characteristics tend to reflect less variability than individuals manifest, a notion that will be particularly relevant to the discussion in this chapter.

The discussion in Chapter 2 included the point that many kinds of populations are nor- mally distributed. The companion to that statement, of course, is that some are not nor- mally distributed. Income data, for example, tend to reflect a distribution with right skew; most of us have salaries in the five-digit range, but the Bill Gateses and Warren Buffetts— with their other-worldly incomes—create salary distributions with positive (right) skew. The same is true of home prices where a few very expensive homes make the general distribution of home values not normal. Because prices of this sort lack of normality, it is typical to describe home prices and salaries in terms of median, rather than mean prices. Median prices are usually a more accurate descriptor of the characteristics of the popula- tion when whatever is being described reflects a skewed distribution.

tan81004_03_c03_051-074.indd 52 2/22/13 3:31 PM

CHAPTER 3Section 3.2 The Distribution of Sample Means

A related problem is that z score analysis took us to the values from Table 2.1 to determine the probability of the specified outcomes, and those Table 2.1 values assume normality. If it isn’t safe to assume normality, what adjustments are required? Help comes in the form of the distribution of sample means.

3.2 The Distribution of Sample Means

To this point, a normal distribution has referred to a distribution created by measuring every individual in a specified population so that plotting all the individual scores in a graph produces a bell-shaped (Gaussian) distribution. While a population based on measuring one person at a time may be the prevailing image of how a normal distribution is formed, technically a population just means that every member of the group is repre- sented, not necessarily that every individual is represented as an individual. For example, if every individual in a population becomes part of a sample drawn from the population, and each sample is represented by its mean score (M), collectively those sample means still constitute a population. The fact that individual scores have their influence as compo- nents of a sample rather than as isolated individuals is unimportant to whether the result is a population.

To illustrate this, suppose an executive search firm is selecting management candidates for a nationwide chain of retail grocery stores for which operations are being expanded. Because successful managers need to be able to analyze situations involving complex data sets, the human resources department for the grocery chain administers a test of analytical abil- ity to each management applicant. Collectively, the

scores from the candidates hired over the last several years constitute the population of managers’ analytical ability scores. Because the company is pleased with the performance of the current managers, perhaps those at the executive search firm decide that the most efficient way to select new store managers is to compare applicants’ scores to the current managers’ mean level of analytical ability.

If the mean level of analytical ability is determined by retrieving the managers’ scores from the chain’s database, one at a time, the symbol for the result would be the population parameter, m. As an alternative, current managers might be sampled in groups of 10, and the mean level of analytical ability for each group of 10 calculated plotted. If this is done until every manager is represented in one of the samples, the result will be a population distribution of sample means.

The Central Limit Theorem

You may be wondering aloud, “Why bother?” Since each individual must be measured to determine the sample mean anyway, why not just plot the individual measures and base the population on individual scores rather than going through the additional step of calculating sample means?

Key Terms: The distribution of sample means is a popula- tion based on sample means rather than on individual scores.

tan81004_03_c03_051-074.indd 53 2/22/13 3:31 PM

CHAPTER 3Section 3.2 The Distribution of Sample Means

Relying on sample means as the basic element of the population provides an important advantage related to the normality of the data. It is explained by the central limit theorem. That theorem, or logical statement, holds that if a population is sampled an infinite num- ber of times using a consistent sample size, plotting all the sample means will produce a distribution of sample means that will be normal. What is particularly important is that the distribution will be normal even if the original population of individual scores was not. When the unit of analysis is a sample rather than an individual, the workings of the cen- tral limit theorem solve the “what if the data are not normal?” problem. As a result, the values in Table 2.1 are relevant for analyses where samples, rather than individuals, are the issue, analyses that can occur without the usual leap of faith regarding data normality.

Can this important tendency be demonstrated? After all, there isn’t any way to draw an infinite number of samples from a population. But even plotting a less-than-infinite number of samples from the same population provides clues to how the central limit theorem works.

Referring back to the analytical ability of managers, perhaps the analytical ability scores of 10 management candidates are gathered. Scores on the instrument range from 1 to 10, and each candidate just happens to receive a different score. If the individual scores of 1 to 10 are plotted, the result is the distribution in Figure 3.1.

Figure 3.1: A frequency distribution for the scores 1 through 10, each score occurring once

No one would confuse the distribution in Figure 3.1 with a normal distribution. For one thing, it’s extremely platykurtic—something that is also apparent from the descriptive statistics:

• R 5 9 (10 – 1) • s 5 3.028

Recall that in a normal distribution, the standard deviation is about one-sixth of the range. Here, it is slightly more than one-third of the range.

Score Frequency

10

9

8

7

6

5

4

3

2

1

Score Values

1 2 3 4 5 6 7 8 9 10

tan81004_03_c03_051-074.indd 54 2/22/13 3:31 PM

CHAPTER 3Section 3.2 The Distribution of Sample Means

Using a procedure by Diekoff (1992), the workings of the central limit theorem can emerge even with a less-than-infinite number of samples. Suppose the HR people plotted a distri- bution based on all possible samples of n 5 2 rather than on just the 10 individual scores. Table 3.1 indicates there are 90 possible combinations of the 10 analytical ability scores ranging from 1 to 10 when n 5 2.

Table 3.1: All possible combinations of analytical ability scores ranging from 1–10

1, 2 2, 1 3, 1 4, 1 5, 1 6, 1 7, 1 8, 1 9, 1 10, 1

1, 3 2, 3 3, 2 4, 2 5, 2 6, 2 7, 2 8, 2 9, 2 10, 2

1, 4 2, 4 3, 4 4, 3 5, 3 6, 3 7, 3 8, 3 9, 3 10, 3

1, 5 2, 5 3, 5 4, 5 5, 4 6, 4 7, 4 8, 4 9, 4 10, 4

1, 6 2, 6 3, 6 4, 6 5, 6 6, 5 7, 5 8, 5 9, 5 10, 5

1, 7 2, 7 3, 7 4, 7 5, 7 6, 7 7, 6 8, 6 9, 6 10, 6

1, 8 2, 8 3, 8 4, 8 5, 8 6, 8 7, 8 8, 7 9, 7 10, 7

1, 9 2, 9 3, 9 4, 9 5, 9 6, 9 7, 9 8, 9 9, 8 10, 8

1, 10 2, 10 3, 10 4, 10 5, 10 6, 10 7, 10 8, 10 9, 10 10, 9

If a mean is calculated for each pair of the scores in Table 3.1 and then plotted in a fre- quency distribution, the result is Figure 3.2. This figure is a distribution of sample means.

Figure 3.2: A frequency distribution of possible combinations of n 5 2 analytical ability scores, 1–10

Ninety samples is well short of an infinite number, but the effect of the central limit theo- rem nevertheless begins to emerge. Although not a normal distribution, it is much closer to a normal distribution than Figure 3.1. Note that for the 90 sample means:

Score Frequency

10

9

8

7

6

5

4

3

2

1

Score Values

1.5 2 2.5 3 3.5 4 4.5 5 5.5 6 6.5 7 7.5 8 8.5 9 9.5

tan81004_03_c03_051-074.indd 55 2/22/13 3:31 PM

CHAPTER 3Section 3.2 The Distribution of Sample Means

• R 5 8 which is the difference between the largest sample mean (8 1 9)/2 or (9 1 8)/2 and the smallest sample mean (1 1 2)/2 or (2 1 1)/2

• s 5 1.926

The standard deviation is now less than one-quarter of the range rather than just over one- third. The distribution is less platykurtic than the distribution of individual scores.

What is it that makes Figure 3.2 more like a normal distribution than Figure 3.1? The scores that occur in the middle of the distribution occur with greater frequency than the scores in the tails. A glance at the figure indicates that the mean is 5.5. There are several combinations of scores that can produce M 5 5.5:

1 & 10, 2 & 9, 3 & 8, 4 & 7, 5 & 6, 6 & 5, 7 & 4, 8 & 3, 9 & 2, 10 & 1

The scores in the tails of the distribution occur with much less frequency. There are only 2 possible ways to have M 5 1.5, for example:

1 & 2, 2 & 1

Likewise, there are only two possible ways to have M 5 9.5:

9 & 10, 10 & 9

With repetitive sampling, the different frequencies of the several sample means begin to shape the distribution.

The Mean of the Distribution of Sample Means

The symbol used for a population mean to this point, m actually indicates the mean of a population formed from one score at a time. The mean of a population of sample means is indicated by mM, literally, a mean of means.

Note that summing the scores 1 through 10 and calculating a mean produces 5.5. If the scores 1 through 10 are the entire population of scores, m 5 5.5.

Figure 3.2 suggests that the mean of the distribution of the 90 sample means is also 5.5 (mM 5 5.5). This can be figured directly by calculating the mean of each pair of scores, and then determining the mean of all the means.

The point is this: A population based on all the individual scores and a distribution of sample means based on all possi- ble samples of n from the same data will have the same value, m 5 mM.

Review Question A: What are the requirements for a population?

tan81004_03_c03_051-074.indd 56 2/22/13 3:31 PM

CHAPTER 3Section 3.2 The Distribution of Sample Means

Variability in the Distribution of Sample Means

The standard deviation for the distribution of sample means was smaller than for the distribution of individual scores. This is because samples moderate the effect of extreme scores. In the distribution of individual scores there is 1 chance in 10 of selecting a candidate who has an analytical ability score of 10, for exam- ple, but there is no chance of selecting a sample for which analytical ability can be M 5 10. Any time 10 is part of the sample its effect will be diminished by whatever other score is in the sample. At the other end of the distribution, the same occurs in a sample including 1. There will always be less variability in a distribution of sample means than in a distribution based on one-at-a-time sampling.

The Standard Error of the Mean

The sigma (s) that indicates the standard deviation of a population is specific to a distri- bution based on individual scores. To indicate the standard deviation of the distribution of sample means an M is appended to sigma, sM (just as M was subscripted to m to indi- cate the mean). The formal name for sM is the standard error of the mean.

The word “error” in statistical language refers to unexplained variability rather than to having made a mistake. Several kinds of “standard errors” will be calculated in these chapters. All standard error val- ues measure data variability.

Sampling Error

Although the standard error of the mean doesn’t refer to a mistake per se, another kind of error, sampling error, does. In inferential statistics, a sample of management candidates is often important for what it reveals about all management candidates. This presumes that the sample accurately represents the population. When it doesn’t, the explanation is sampling error.

Samples reflect the population well when two conditions are satisfied:

• the samples must be relatively large, and • the sample is based on random selection.

The safety of large samples is explained by the law of large numbers. This mathematical principle indicates that as a proportion of the whole, errors diminish as the number of data points increases. The potential for serious sampling error diminishes as the size of the sample grows.

Key Terms: The standard error of the mean is the mea- sure of variability among the sample means that make up the distribution of sample means.

Key Terms: The law of large numbers indicates that errors decrease as a portion of the number of decisions as the num- ber increases.

tan81004_03_c03_051-074.indd 57 2/22/13 3:31 PM

CHAPTER 3Section 3.3 The z-Test

Random selection refers to a situation where every member of the population has an equal probability of being selected. A random sample of n 5 5 can be created from the 10 management applicants by:

• assigning each person a number, • placing the 10 numbers into a container,

and • shaking the container and drawing out

five numbers.

When samples are randomly selected, they differ from populations only by chance. When samples fail to capture some important characteristic of the pop- ulation (other than its size) there is sampling error. A sample can never exactly duplicate all the descriptive characteristics of the population, but sampling error will usually be minor if samples are relatively large and randomly selected.

Statistical analysis procedures tolerate some minor, random sampling error, but systematic sampling errors are another matter. Systematic sampling error occurs when the same error occurs time after time.

In 1936, the publishers of the Literary Digest, a prominent publication of the time, decided to predict the outcome of that year’s presidential election in the United States. To ensure an adequate sample size, millions of postcards were sent to people listed in telephone books and with car registrations. As an aside, the Harris and Gallup organizations typically get very accurate results with a few thousand, and sometimes a few hundred responses. Unfortunately, the decision to use telephone books and automobile registrations distorted the sample, identifying those of relative prosperity at the height of the Great Depression. The results indicated that Alf Landon would win and of course, Franklin Roosevelt was re-elected to a second term in a landslide. Roosevelt carried every state in the union except Maine and Vermont.

The problem was a systematic sampling error. The pollsters consistently and non-ran- domly selected from a group not representative of the entire population. With random selection and their large sample, chances are that they would have had a very accurate prediction of election results, but sample size alone wasn’t enough to salvage the effort.

3.3 The z-Test

To summarize, a distribution of sample means is a distribution based not on individual scores, but on the means of samples repeatedly drawn from the same population. The central limit theorem indicates that when a population is based on samples rather than individual scores, the resulting population will be normal, whatever the nature of the original population of individual scores.

Because the central limit theorem ensures a normal distribution, the values in Table 2.1 can answer questions about groups that were asked about individuals in Chapter 2. This represents a distinction between the z-test and the z score that came up in the last chapter.

Key Terms: Sampling error occurs when sample statistics differ from population param- eters. Random selection (all have an equal chance of selec- tion) and large samples mini- mize sampling error.

tan81004_03_c03_051-074.indd 58 2/22/13 3:31 PM

CHAPTER 3Section 3.3 The z-Test

Where

z 5 the calculated value of z M 5 the sample mean

mM 5 the mean of the distribution of sample means sM 5 the standard error of the mean

Note that in addition to substituting M for x, and mM for m the denominator now indicates data variability in the distribution of sample means. Like the z score, the z-test can answer a variety of questions. For example, someone might ask in reference to the group of man- agement candidates:

• If a group of applicants is randomly selected from a population of applicants with analytical ability score mM 5 ___, and sM 5 ___, what is the probability that the selected group will have a mean level of analytical ability that is the same as or higher than the population?

• What is the probability that a group with mean level of analytical ability lower than a particular value would be selected from a population with mM 5 ___ and sM 5 ___?

For either of those questions, the procedure would be to calculate a value of z, locate the associated table value in Table 2.1, and interpret the result just as we did for z scores in Chapter 2. Perhaps the most common application of the z-test, however, is to deter- mine whether a sample is characteristic of the population to which it is compared. When samples are selected from any population it is impossible for the sample to have charac- teristics identical to those of the population. However, the sample will usually be similar enough that we can safely conclude that it probably was drawn at random from the partic- ular population. In such instances, any difference between the sample and the population is attributed to sampling variability and declared to be statistically insignificant.

With the z score, the issue was how an individual compares to a group—either a sample or a population. With the z-test, the question is how sample groups compare to populations. The procedure is the same: Calculate a value of z, and then use the table to interpret that value. The z transformation formula for population data from Chapter 2 was the following:

z 5 x 2 m

s

If instead of comparing an individual score to the mean of the population, the compari- son is of a sample mean to the mean of the distribution of sample means, the following is the result:

Formula 3.1 z 5 M 2 mM

sM

tan81004_03_c03_051-074.indd 59 2/22/13 3:31 PM

CHAPTER 3Section 3.3 The z-Test

Where

sM 5 the standard error of the mean s 5 the population standard deviation n 5 the number in the group

For example, if the standard deviation for sales among a group of 35 sales associates is $4,750 for a particular month, then the standard error of the mean for a group of 35 sales associates is

sM 5 s

"n

sM 5 4750

"35 5

4750 5.916

5 802.907

The standard error of the mean for sales data is 802.907. That means that in a distribution based on repetitive samples of sales associates, the standard deviation of all those mean scores is 802.907.

Note that to calculate a z-test the value of s must be available. That problem will be rec- tified in Chapter 4, but in the meantime, let’s say that the group of sales associates for whom the standard error of the mean was just calculated have mean sales of $33,452 for the month (M 5 $33,452). Is this group representative of sales associates nationally who have sales of $31,119?

Formula 3.2 sM 5 s

"n

On the other hand, when the sample is so different from the population to which it is compared that differences can’t be explained by sampling variability, the sample is said to represent some population other than the one to which it was compared. The difference between the sample and the population in that case is statistically significant.

Calculating the z-Test

Note that some components of the z-test must be provided. Either mM or m must be indi- cated, for example, and since no one is likely to have the mean scores for an infinite num- ber of samples, we cannot directly calculate the standard error of the mean, sM. Either the standard error of the mean will have to be given, or it will need to be determined as follows:

tan81004_03_c03_051-074.indd 60 2/22/13 3:31 PM

CHAPTER 3Section 3.3 The z-Test

• M 5 $33,452 • Since m 5 $31,119, mM 5 $31,119 (remember that m 5 mM)

• sM 5 $802.907

z 5 M 2 mM

sM 5

33452 2 31119 802.907

5 2.906

Interpreting the Value From the z-Test

The z 5 2.906 is a value of z similar to those calculated in Chapter 2, except that here the value is a gauge of how much a sample mean (M) differs from the mean of a population of samples (mM) rather than how an individual (x) differs from either M or m. If the z value is rounded to 2.91, Table 2.1 indicates that of the upper half of the standard normal distribution, a proportion of .4982 out of .5 occurs between z 5 2.91 and the mean of the distribution.

• If the question is what is the probability of select- ing a sample of sales associates with mean sales for the month of $33,452 or lower from a popula- tion with a mean of $31,119, remembering that Table 2.1 has proportion values for just half of the area under the normal distribution curve, the answer is:

.4982 1 .5 5 .9982

• If the question is what proportion of the distribu- tion is beyond z 5 2.91, the answer is:

.5 2 .4982 5 .0018, or 18 ten thousands of the distribution.

The location of z 5 2.91 in the standard normal distribution is indicated in Figure 3.3.

Review Question B: How do the measures of central tendency and variability compare when one population is based on individual scores and another is based on sample means?

tan81004_03_c03_051-074.indd 61 2/22/13 3:31 PM

CHAPTER 3Section 3.3 The z-Test

Figure 3.3: The probability of selecting a sample of sales associates with mean sales of $33,452 or lower from a population with mean sales of $31,119 and a standard error of the mean of 802.907

Variability between group means tends to be small relative to the variability between indi- viduals. Recall that the measure of data variability in the distribution of sample means is the standard error of the mean. Because the standard error of the mean is always substan- tially smaller than the standard deviation (it is determined by dividing the standard devi- ation by the square root of the number of values), relatively minor differences between the sample mean and the population mean result in large values of z.

Another z-Test

Suppose that 10 management applicants have a mean analytical ability score of 5.50, while for the organization as a whole, analytical ability parameters are m 5 6.650 and s 5 1.817.

What is the probability that this sample is representative of that population? Note that question is whether it is probable that a sample with these characteristics could be one of those randomly selected from a population with m 5 6.650 and s 5 1.817. Before calculating the z value, the standard error of the mean must be determined:

sM 5 s

"n 5

1.817

"10 5 .575

With the standard error of the mean, and noting that m 5 mM, z can then be calculated:

z 5 M 2 mM

sM 5

5.50 2 6.650 .575

5 2.0

The table value for z 5 2.0 is .4772

p = .5 + .4982 = .9982

–3 –2 –1 +3+2+10z value

z = 2.19 (.4982)

p = .4982p = .5

z = M - µ M = 33452 – 31119 = 2.906, or 2.91

802.907 M

Review Question C: What is meant by “sta- tistically significant”?

tan81004_03_c03_051-074.indd 62 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

The question was about the probability of selecting a sample with M 5 5.50 from a popu- lation with mM 5 6.650. The z 5 2.0 and the associated table value of .4772 indicate that the probability is nearly p 5 .98 (.4772 1 .5) that any sample drawn from this population will have analytical ability scores higher than M 5 5.50. Remember that multiplying the prob- ability value by 100 indicates the percentage of the distribution involved; 97.72% of the distribution involves samples with mean scores higher than 5.50. There is only about 2% chance that a sample with a mean analytical ability score of 5.5 or lower can be randomly drawn from a population with mM 5 6.650.

3.4 Statistical Significance

Earlier the point was made that besides providing answers to questions about the pro-portion of a distribution above a particular point, between two points, and so on, calculating the z-test also allows judgments about statistical significance, a concept that comes up for the first time with the z-test, but transcends that test. In fact, all of the tests that come up in the balance of the book involve questions about statistical significance. The common element among the several tests is that statistical significance means results are unlikely to have occurred by chance.

In the case of the z-test the issue is at what point is the sample mean (M) so distant from the mean of the distribution of sample means (mM) that the sample is probably not one of the samples that make up the specified population? To put it another way, at what point does the value of z suggest that the sample probably belongs to a population other than the one to which it was compared? For the sample in the problem completed just above, M 5 5.50. Given the standard error of the mean (sM 5 .575) by which the difference is divided, is M sufficiently distant from mM 5 6.650 that the sample represents a population other than the population of all managers in this particular field? Does M 5 5.50 represent some other population, such as the population of the general public?

Determining Statistical Significance

Ronald Fisher, a statistician in the world’s first department of statistical analysis at Uni- versity College London coined the term “statistically significant.” In the effort to estab- lish a standard, Fisher adopted what is basically an arbitrary criterion. Fisher suggested that if an outcome will occur by chance only 5% of the time or less, it could be viewed as “statistically significant.” In terms of probability, any outcome with p 5 .05, or less, is statistically significant.

To place Fisher’s standard in a z-test context, samples with means similar to the popula- tion mean are the most likely. This was illustrated in Figure 3.2. Because the tails in a nor- mal distribution extend infinitely outward in either direction, in theory any sample mean can be one of the samples in the distribution of sample means, but as the sample becomes more distant from the population mean, this becomes less probable.

tan81004_03_c03_051-074.indd 63 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

Part of the complexity in this discussion is that populations are not discrete; they over- lap other populations. That most extreme 5% of the distribution is statistically significant because it is probably more representative of some other, overlapping distribution. In our problem, the population that represents analytical ability scores in management applicants likely overlaps with analytical ability scores among the population of the general public, with the population of clerical employees, and perhaps a number of other populations.

Because normal distributions are symmetrical, that most extreme 5% of outcomes is divided into halves indicating that the most extreme 21/2% in either tail of the distribution are the outcomes that are statistically significant. The middle 95% of outcomes are not statistically significant. Given the nature of the population, those outcomes are likely to occur as just random outcomes.

Interpreting z

The symmetrical nature of the normal distribution meant that Table 2.1 needed to provide the values for only half of the distribution, since the values for the other half are the same. Examining Table 2.1, what value of z includes 50 2 2.5 5 47.5% of the distribution and so excludes the most extreme 21/2%? Note that z 5 1.96 occurs at a point where there is .475 of .5 of the distribution between that point and the mean of the distribution. Put the other way, z 5 1.96 and beyond indicates the domain of statistically significant outcomes in a z-test. At this point, the sign of the value doesn’t matter. Whether z is positive or negative only indicates the side of the distribution in which the outcome occurs.

The problem worked earlier comparing 10 management applicants with M 5 5.50 to the population of management personnel with mM 5 6.650 yielded z 5 22.0. Since that 22.0 is more extreme than z 5 21.96, the result is statistically significant. The sample of 10 probably was not drawn from a population with mM 5 6.650. These 10 people are more characteristic of some population other than a population with mM 5 6.650 analytical abil- ity scores.

Note that the table value associated with z 5 22.0 indicates that .4772 out of half (.5) of the distribution occurs between z 5 22.0 and the mean of the population. That means that only .0228 of the samples in this distribution have scores lower than M 5 5.50. Although some individuals within the sample scored well, the HR people might want to reexamine their recruitment efforts. As a group, these are not the people they want to hire. Their ana- lytical ability scores are too low.

Another View of Significance

Whether p 5 .05, .01, .001, or whatever it is, Fisher picked a point, and said essentially, “anything beyond this level of probability isn’t likely to have occurred by chance.” Not everyone agrees that there has to be such a standard. The author’s statistics professor at Texas A&M University took the position that what is “significant” depends upon circum- stances. His approach was to calculate the probability that an event could occur by chance, and then let consumers make their own decision about whether it’s significant.

tan81004_03_c03_051-074.indd 64 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

Where

n 5 the required sample size z 5 the value of z that corresponds to how certain we wish to be of the result. Since 6z 5 1.96 includes the middle 95% of the distribution, using that value in the for- mula provides p 5 .95 that the sample emulates the population. If .99 certainty is required, z 5 2.58. s 5 the standard deviation of the population. If the population standard deviation isn’t available, a sample standard deviation (s) can be substituted, although the estimate will lose some precision. variation from s 5 the variability allowed from s or s.

Formula 3.3 n 5 a 1z 2 1s 2 variation from s

b 2

Dr. Smith is in good company. Rather than indicating that a result was, or was not, statisti- cally significant, many of the statistical packages that professionals use simply determine the probability than an event could have occurred by chance and leave the interpretation to those who intend to use the results.

Sampling Error as an Explanation of Difference

Fisher chose the most extreme 5% of outcomes as those that were statistically significant because he recognized that in virtually every z-test there will be some difference between M and mM so that z has some value other than 0. When the differences fall short of sta- tistical significance (z , 1.96), how are they explained? The answer is sampling error. Because no sample can exactly emulate the population, most samples in the distribution of sample means will have descriptive characteristics that are different from those of the population—M ≠ mM, for example. When a difference is not large enough to be statisti- cally significant, the difference between M and mM is explained by sampling error. As it was with the standard error of the mean, “error” does not refer to having made a mistake in the usual sense. It means that there is some variability between sample and population that is not explained by anything except that the sample does not exactly duplicate the essential characteristics of the population.

More Confidence in the Sample

Although there is always going to be some degree of sampling error, it can be minimized. Larger samples are likely to have less sampling error than small samples, all other things being equal, but it can be difficult to define “large.” Sprinthall (2000) explains an approach to determining the needed sample size based on the answers to two questions:

1. How much certainty must there be that the sample is like the population? 2. How much error can be tolerated?

The formula is as follows:

tan81004_03_c03_051-074.indd 65 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

For example, suppose that the manager at a large manufacturing plant wants to select a sample that will allow her to analyze unexcused employee absences. She calculates a stan- dard deviation of the number of unexcused absences by all employees for the prior year and finds it to be 3.462 days. The manager is willing for the sample data to digress from plant-wide data by 1 point and wishes to be .95 confident of the result.

With

s 5 3.462 z 5 1.96

n 5 a 1z 2 1s 2 variation from s

b 2

n 5 a 11.96 2 13.462 2 1.0

b 2

5 approximately 46 people

With p 5 .95, a random sample of about 46 people will provide a sample within a point of the standard deviation for the entire plant.

Changing the conditions can dramatically affect the required sample size. If the man- ager needs to be within a half-point of the population standard deviation and wishes for p 5 .99, note the impact on the result:

n 5 a 12.58 2 13.462 2 .5

b 2

5 approximately 319 people

Smaller variance allowances and greater certainty of emulating the population translate into larger required samples, but this is what we would expect. On the other hand, the for- mula provides some protection against choosing a larger sample than necessary to com- plete the needed analysis. Formula 3.3 can help strike a balance. Samples that are very large can be time-consuming and expensive to work with. Samples that are very small may not reflect the essential characteristics of the population, making it difficult to have confidence in the results.

Decision Errors

The fact that determining statistical significance is based on the probability that an event did not occur by chance suggests that there is some risk involved in the decision. Because statistical analysis is based on probabilities rather than certainties, any statistical decision involves the potential for error. The errors are of two kinds. There are alpha, or type I, errors, which occur when a result is determined to be statistically significant and further analysis reveals that the decision was incorrect. And there are beta, or type II, errors, which occur when a test indicates that the outcome is just a random, nonsignificant result,

tan81004_03_c03_051-074.indd 66 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

but further testing would indicate that in fact the result is statistically significant. Decision errors can be best understood in the context of a problem. Consider the following:

Alpha (α), or Type I, Errors

An employment specialist is screening applicants for positions as assemblers on an assem- bly line where electronics components are put together. The primary criterion for assem- blers is manual dexterity, so each applicant must complete a timed manual dexterity test. The mean manual dexterity score for all hired assemblers is m 5 45.377 with a standard deviation of s 5 5.519. In response to a news spot that indicates that the electronics com- pany is hiring, 17 people charter a bus and come as a group from a neighboring town to take the test. Their mean level of manual dexterity is M 5 42.244. Do they meet the manual dexterity requirement?

• The standard error of the mean is sM 5 s"n 5 5.519/"17 5 1.339. • The value of z 5 (M 2 mM)/sM 5 (42.244 2 45.377)/1.339 5 22.340.

The z 5 22.340 corresponds to p 5 .4904 in Table 2.1. Only .5 2 .4904 5 .0096 of the popu- lation of hired assemblers would have an average score of 42.244 or lower. This z value is more extreme than z 5 1.96, which corresponds to the standard p 5 .95. z 5 22.340 actu- ally corresponds to p 5 .4904 3 2 5 .9808. This indicates that this group has significantly lower manual dexterity scores than the hired assemblers have. Although some individu- als within the applicant group may be qualified, as a group they are significantly lower than the company standard. It appears that collectively, at least, they do not meet the hiring requirement.

The test indicates that, as a group, these applicants have significantly lower manual dex- terity than assemblers have. It appears that in terms of manual dexterity they belong not to the population that includes successful assemblers, but to some population that has lower manual dexterity. However, perhaps the bus these people chartered broke down and they spent the night stranded on some remote road. Perhaps retesting them after they have had rest and something to eat would indicate that they aren’t any different in manual dexterity than the employees. If this is the case, the initial decision constitutes an alpha error—the significantly lower manual dexterity scores were false.

The probability of an alpha error in any statistical test is identical to the criterion for statis- tical significance for the test. If the test is completed with p 5 .05 level for significance, the potential for alpha error is .05 (a 5 .05). This should make sense since that most extreme 5% of the distribution is actually part of the distribution, albeit the most distant part. Sometimes that most distant part of the distribution is erroneously assumed to be a part of some other distribution, which is when type I, or alpha, errors occur.

If the test is described in terms of the potential for alpha error rather than the significance level, the indication is that a 5 .05. The “p” refers to the level selected for determining statistical significance. The “a” refers to the potential for a type I error. But they should be thought of as two different ways to describe the same thing.

tan81004_03_c03_051-074.indd 67 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

Beta (b) or Type II Errors

In a second scenario, consider another group of aspiring components assemblers. They too take the manual dexterity test and as a group of 12 score M 5 43.872. How do they compare to the employed assemblers?

• Their standard error of the mean is sM 5 s/"n 5 5.519/"12 5 1.593.

• The value of z 5 (M 2 mM)/sM 5 (43.872 2 45.377)/1.593 5 20.945.

The z value indicates that as a group, these applicants are not significantly different from the hired assemblers. They appear to meet the hiring requirement. However, perhaps it comes to light that one of the applicants had a friend responsible for conducting the company’s testing, and that friend supplied a copy of the manual dexterity test so that the applicants could surreptitiously practice before being given the test. When they are retested with an unfamiliar form of the test, they are found to be significantly lower in manual dexterity than the

employees are. A decision based on the first result constitutes a type II, or beta (b), error—a finding that indicated a nonsignificant difference was in error.

The Relationship Between α and b

The only time a type I error can occur is when test results indicate statistical significance. If the result isn’t judged statistically significant, there is no potential for a type I error. On the other hand, the only time a type II error can occur is when the decision is that the result is not significant. There is no situation in which both errors can occur.

Although the errors are mutually exclusive, the decisions an analyst makes can prompt one type of decision error to be more probable than the other. This is usually the course taken when one error has the potential to be more damaging than the other. In those instances, it makes sense to protect against the more damaging error. The complication is that diminishing the potential for one type of error increases the probability that the other will occur.

The reason for this dependence goes back to the overlapping populations that were referred to ear- lier. Often the data people rely on to make decisions are not the perfect indicators that the best decision- making requires. For example, if sales data are used to determine promotions of sales representatives to team leaders, perhaps a sale that actually occurred on the 30th of the previous month isn’t entered until the 1st of the succeeding month, distorting what the salesperson actually accomplished for either month. In a perfectly objective world, the person evaluated may be doing better or worse than the evaluation indicates. If a number of people are evaluated, those who are the best among those judged not ready for promotion might actually be doing better than those who barely meet the criteria for promotion. This is illustrated in Figure 3.4.

Review Question D: You are in the busi- ness of certifying forklift operators for manufacturing com- panies. If you certify someone as compe- tent who is actually no more competent than the population generally, what type of error has occurred?

Key Terms: Alpha, or type I, errors occur when results are erroneously found statistically significant. Beta, or type II, errors occur when results are erroneously found not significant.

tan81004_03_c03_051-074.indd 68 2/22/13 3:31 PM

CHAPTER 3Section 3.4 Statistical Significance

Figure 3.4: The populations of the promoted and the not promoted

In the figure,

• The upper horizontal line represents the population of executives recom- mended for promotion.

• The lower horizontal line is the population of those not recommended for promotion.

• Because of imperfect measurement of the characteristics necessary for promo- tion, the two horizontal lines overlap.

• The vertical red line represents the minimal standard for promotion.

If the red line that represents the standard for promotion is moved to the right so that the standards for promotion become more rigorous, there will be fewer type II errors (some- times called “false negatives”), but there will be more type I errors (“false positives”). If the standard is made less rigorous by moving it to the left to protect against type I errors, those errors will be reduced, but there will be more type II errors.

Is one error more damaging than the other? Do analysts have a preference for one type of error? The answer, of course, depends upon circumstances and especially on the impact that a decision error has on the people involved. If promoting someone who isn’t qualified can have a serious detrimental impact on the company, the decision should be to protect against type I errors. If failing to promote someone who is qualified means that the person goes elsewhere instead of remaining to make an important contribution to the company, the decision might be made to protect against type II errors.

Perhaps a committee is evaluating educational programs, and the committee deems the program at university A to be significantly better than its competitors’ programs. If the result is that the graduates from university A receive preferential hiring, and the differ- ence really is not statistically significant after all, there has been a type I error.

On the other hand, perhaps management trainees who take a particular seminar in group dynamics make substantially better personnel decisions than those who didn’t take the course. If administrators in the organization fail to recognize that those who took the seminar consistently make better decisions than others, then there has been a type II error.

Type I Errors

Type II Errors Promoted Population

Not Promoted Population

Promoted Standard

tan81004_03_c03_051-074.indd 69 2/22/13 3:31 PM

CHAPTER 3Section 3.5 The z-Test Using Excel

So which error is the more serious depends upon circumstances, but statisticians may have their own bias. Power in statistical testing is described in terms of the likelihood of a type II error. The most powerful tests are those for which type II errors are the least common. The power of a statistical test is symbolically indicated by the expression power 5 1 2 b. The complication is that although we always know the probability that a type I error has occurred—it is the same as the criterion for statistical significance—we do not know the probability of a type II error. In the case of either error, the only way to check is to gather new data and run the analysis again to check for a consistent result.

3.5 The z-Test Using Excel

Although there’s an option for a z-test in the Data Analysis package in Excel, it’s a different z-test than the one explained here. To complete this test in Excel requires programming in some formulas, but they are not difficult to complete, and even without a dedicated test the program will help with the repetitive calculations that are needed to produce the mean and standard deviation of the sample.

Suppose that a project manager is keeping track of the hours the members of a project team commit to a particular project during the week so that she can bill their hours toward the right project. For the eight members of the team, the following hours are logged:

13.5, 18, 22.375, 25.240, 26, 29.331, 30, 30

For projects similar to this one that the company has taken on in the past, team members have expended a mean number of hours of 19.500 per week with a standard deviation of 4.525. Is the number of hours that this project has consumed significantly different from the amount of time that other, similar projects have required? The solution will be com- pleted in Excel with the formulas that need to be entered designated in red. The cells that are involved in the calculations are indicated in blue:

1. Enter the number of hours worked data into a spreadsheet in cells A12A8. 2. Have Excel calculate the mean number of hours worked by the team by enter-

ing the formula 5average(A1:A8) in cell A9. 3. In cell A11 determine the standard error of the mean (sM) by dividing the

population standard deviation, which is given above as s 5 4.525, by the square root of the number of team members (8). The commands in Excel are 5 4.525/sqrt(8). Recall that the forward slash (/) is the sign that indicates division in Excel. The “sqrt” is the abbreviation for square root, which will be performed on whatever number follows in parentheses—8 in this instance.

4. Determine the value of z for the problem in cell A13 by entering the command 5(A9-19.5)/A11 in that cell. This function will take the difference between the sample mean, which is the value in cell A9, and the population mean, which is 19.5, and divide the difference by the standard error of the mean that was calculated in cell A11. Figure 4.5 is a screenshot of how the display will appear before executing this step by pressing “Enter.”

Key Terms: Power in statisti- cal testing refers to the likeli- hood of type II errors; powerful tests minimize type II errors.

tan81004_03_c03_051-074.indd 70 2/22/13 3:31 PM

CHAPTER 3Chapter Summary

Figure 3.5: Calculating a z-test in Excel

The result is z 5 3.004. If the test is completed at p 5 .05, and for which the critical value from the table is z 5 1.96, these 8 people contributed significantly more time to this project during the week than other, similar projects in the past have required.

Chapter Summary

The z-test provides a good introduction to formal statistical testing. It’s an uncompli-cated test that still manages to involve many of the same issues that come up in the more advanced tests including, most prominently, statistical significance (Objective 4). The test provides an important advantage over the z scores in Chapter 2 because, gener- ally speaking, analysts are more interested in the performance of groups than of single individuals. Individual performance can be highly variable. Groups’ performances are a good deal more stable, which means that a manager is more likely to ask whether an incentive program is effective for an entire sales staff than whether the program is effec- tive for one individual. The z-test provides a mechanism for evaluating how samples compare to populations.

The z-test is based on the distribution of sample means (Objective 1), which is a popu- lation based on the means of samples of individuals, rather than on the individuals in isolation. The point of creating a population based on sample means is that the central limit theorem indicates that such a population will be normally distributed, whether or not the original distribution of individual scores was normal (Objective 2). Because the distribution of sample means is a normal distribution with all the characteristics that define normality, the values in Table 2.1 that were used with z scores in Chapter 2 have application here as well.

tan81004_03_c03_051-074.indd 71 2/22/13 3:31 PM

CHAPTER 3Answers to Review Questions

Although Table 2.1 values allow the same questions to be asked of groups that were asked of individuals in the prior chapter (what is the probability that a group will score above a point, below a point, between two points, and so on), questions about whether a particu- lar sample is likely to belong to a defined population can also be raised. When a sample is probably not representative of a particular population, the sample is said to be signifi- cantly different from that population (Objective 4).

Statistical significance will come up repeatedly in this book. It is relevant to many dis- cussions besides those related to z-test. Because any decision about whether a result is statistically significant is based on a probability, further analysis can sometimes indicate that an initial decision was incorrect. The decision errors are of two kinds, type I errors that find significance erroneously, and type II errors that erroneously find no significant difference (Objective 7). Minimizing the potential for one type of error inevitably increases the potential for the other.

Samples have the best chance of emulating the characteristics of the populations from which they are drawn when they are large and randomly selected. “Large,” however, is a relative term, and how large a sample must be to have a good chance of capturing the relevant characteristics can be difficult to determine. There are many formulas and proce- dures for determining needed sample size. The formula introduced in this chapter makes determining sample size a function of the level of certainty required that the sample does represent the population, and the level of divergence from the population the analyst is willing to tolerate (Objective 5). The formula provides a good way to answer the question “how big is big enough?” where sample size is concerned.

The data analysis pack in Excel provides commands for many of the basic statistical pro- cedures that analysts use. Unfortunately, the z-test introduced in this chapter is not one of them, but the procedure is not difficult to program into a spreadsheet, particularly with the help of the square root, mean, and sample standard deviation commands that Excel does provide (Objective 6).

Answers to Review Questions

A. A population requires only that every individual is in the defined group. It doesn’t require that each individual be represented as an individual.

B. For the two populations the measures of central tendency will have the same value, m 5 mM. The measures of variability will differ, however, s ≠ sM. Indi- vidual scores vary more than sample means.

C. “Statistically significant” means that an outcome is not random; it is not likely to have occurred by chance.

D. This is an example of a type I error. The individual is judged significantly more competent than the population as a whole, but turns out to be not competent. This is an example of a “false positive.”

tan81004_03_c03_051-074.indd 72 2/22/13 3:31 PM

CHAPTER 3Management Application Exercises

Chapter Formulas

Formula 3.1 z 5 M 2 mM

sM This is the formula for the z-test. It allows one to determine

whether a sample is characteristic of, or significantly differ- ent from, a population.

Formula 3.2 sM 5 s

"n If a value for the population standard deviation (s) is avail-

able, one can calculate the standard error of the mean (sM) with this formula.

Formula 3.3 n 5 a 1z 2 1s 2 variation from s

b 2

Management Application Exercises

Unless otherwise stated, use p 5 .05 in all your answers.

1. How will the variability in sales compare between the individual members of a sales team and the mean sales figures for several sales teams?

2. If monthly utility costs are available for all those in a particular county, and utility users are sampled repeatedly 30 at a time, what can be said about the distribution of the resulting means?

3. If mM 5 $155 for the utility costs mentioned in item 2 with sM 5 $15.73, what is the probability that a random sample of 30 customers will have M 5 $140 or lower?

4. The vice president for a company that installs monitored security systems has job satisfaction scores for those who monitor residential alarms. The mean job satisfac- tion score is 32.956 with a standard error of the mean of 5.924.

a. What is the probability of randomly selecting a sample with a job satisfac- tion mean of 35.0 or higher?

b. If a group with M 5 35.0 is actually selected, is it significantly different from the population?

5. The clerical staff in a law office has the following job performance scores: 25, 37, 38, 43, 44, 48, 51. If the mean level of performance for all clerical staff is 33.255, are those in this law office characteristic of that population? Test at p 5 .05.

6. The standard deviation for a major intelligence test is s 5 15.0. If in a given year the test is administered to 347 people, what is the value of the standard error of the mean?

7. If a researcher wishes to gather a sample of people who have intelligence scores that differ from the national standard deviation of 15 by no more than 3 points, with .95 confidence, how large must the sample be? How large must the sample be if it is to vary from the national standard deviation by no more than 2 points?

tan81004_03_c03_051-074.indd 73 2/22/13 3:31 PM

CHAPTER 3Key Terms

8. Regarding item 7, what steps could the research take to reduce the size of the needed sample?

9. In a particular university, a graduate program in management requires GRE quan- titative scores of 500 or better. This year’s entering class has n 5 16 and M 5 625. Is it characteristic of a national population of graduate students for whom m 5 500 with s 5 100? What’s the probability that a group of applicants selected at random would have M 5 500 or better?

10. Members of a sales staff complete a survey that measures their levels of optimism and score as follows: 11, 14, 14, 16, 19, 20, 22, 23, 27, 30. If the population standard deviation is 4.554:

a. What’s the value of the standard error of the mean? b. What’s the z value for a z-test with this group if mM 5 26.0? c. If 26.0 is the mean for all employed adults, is this group of sales personnel

significantly different? d. Complete this problem using Excel. Refer to Figure 3.5 for help.

Key Terms

• The central limit theorem holds that instead of forming a population by measuring every individual in the population and entering their scores one at a time, repeatedly sampling and entering the sample means will result in a normal distribution.

• The distribution of sample means is a population based on sample means rather than on individual scores.

• The standard error of the mean is the measure of variability among the sample means that make up the distribution of sample means. It is the population standard deviation of all the sample means.

• Sampling error occurs when the characteristics of the sample differ from those of the population. Random selection, selecting individuals to the sample so that each has an equal chance of selection, along with relatively large samples will minimize sampling error.

• Systematic sampling error occurs when the same error occurs each time the sample is constructed. It is systematic sampling error that constitutes sampling bias.

• The z-test is a statistical test used to determine the probability that a sample is charac- teristic of a particular population. When the probability that the sample belongs to the specified population is p 5 .05 or less, the sample is said to be statistically significant.

• The law of large numbers indicates that as a mathematical principle as a proportion of the whole, errors diminish as the number of data points increases. As a practical matter, this means that larger samples are likely to include less sampling error than small samples.

• Statistical decisions are based on what is most likely to occur. Since what is most likely to occur is not always what does occur, there are errors made in statistical analysis. Alpha, or type I, errors occur when results are erroneously found to be statistically significant. Beta, or type II, errors occur when results are erroneously found to be not significant.

• Power in statistical testing refers to the likelihood of type II errors; powerful tests minimize type II errors.

tan81004_03_c03_051-074.indd 74 2/22/13 3:31 PM