BUS 308 Statistics for Managers (10 question Quiz).

profile5ToGo
bus308_chapter_04.pdf

4

t-Tests: Comparing Groups to Populations and Groups to Groups

Learning Objectives

After reading this chapter, you should be able to:

• Explain the advantage of the one-sample t over the z-test.

• Compare the one-sample t-test to the independent samples t-test.

• Distinguish between one-tailed and two-tailed t-tests.

• Explain hypothesis testing in statistical analysis.

• Calculate effect size statistics for t-tests.

• Discuss research applications for t-tests.

iStockphoto/Thinkstock

tan81004_04_c04_075-102.indd 75 2/22/13 3:39 PM

CHAPTER 4Section 4.1 The z-Test's Limitations

Chapter Overview

4.1 The z-Test’s Limitations

4.2 Estimating the Standard Error of the Mean

4.3 The One-Sample t-Test Degrees of Freedom and the One-Sample t-Test Calculating the One-Sample t-Test Interpreting t-Test Results Another Example The One-Sample t-Test in Excel The One-Tailed Test

4.4 Hypothesis Testing The Null and Alternate Hypotheses in a One-Tailed Test

4.5 The Independent Samples t-Test The Population Based on Difference Scores The Test Statistic for the Independent Samples t-Test Assumptions Associated With the Independent Samples t-Test The Independent Samples t-Test on Excel

4.6 Determining Practical Significance

Chapter Summary

Introduction

The z-test in Chapter 3 introduced us to the idea of statistical significance. To varying degrees, all who work in business deal with quantitative data and need to be able to distinguish between outcomes that probably occurred by chance and those that are likely to emerge each time the data are gathered and the analysis completed—outcomes that are statistically significant. For example, when sales data indicate that for a particular period one retail outlet had better sales than other similar stores, a sales manager may wish to determine whether this is a random difference, something that occurred by chance, or whether the more successful outlet is likely to continue to lead in sales. The z-test can answer such questions. It is the first of a number of statistical tests that can help those in business make important decisions.

4.1 The z-Test’s Limitations

The z-test has important limitations, however. Recall that the z-test has this form: z 5 (M – mM)/sM. Both mM and sM are the characteristics of populations—parameter values. Because the mean of the distribution of sample means, mM, has the same value as m, and population values are often published, sM can usually be accessed. On the other hand, the standard error of the mean for the population, sM, frequently isn’t available and can be fairly difficult to determine.

tan81004_04_c04_075-102.indd 76 2/22/13 3:39 PM

CHAPTER 4Section 4.2 Estimating the Standard Error of the Mean

Where

SEM 5 the estimated standard error of the mean s 5 the standard deviation of the sample n 5 the number in the sample

Gosset knew, however, that sometimes even the most carefully selected samples can dif- fer from the population. The difference between s and s, or between M and m are the evidence for what we called sampling error in Chapter 3.

When the sample isn’t like the population, the sampling error results in statistics that are poor estimates of parameter values. Gosset’s solution was to build in a correction for sampling error, an adjustment that is greatest when the risk of sampling error is likewise

Formula 4.1 SEM 5 s/"n

A second limitation of the z-test is that this procedure allows just one type of comparison, that of a sample to a population. It does not allow for samples to be compared to samples, for example. For example, what if the sales manager of a regional office wishes to compare vehicle sales of two auto dealerships? The z-test cannot respond to such questions.

The man who solved these problems is someone who should resonate particularly with business students. Earning degrees in mathematics and chemistry from Oxford Univer- sity in 1897 and 1899, William Sealy Gosset went to work for Guinness Brewing. His job was quality control, and he studied ways to make sure that day-to-day brewing remained consistent with the Guinness standard. A man of remarkable ability, his work led Gossett to procedures for quantifying product quality and then for testing the consistency of the quality over time. Gosset wanted to publish his results for the benefit of others, but Guin- ness had a “no publishing” policy to protect its trade secrets. Caring little for who received credit for the work, Gosset published his procedures anonymously under the pseudonym, “student.” He named the tests he developed to check product quality the “t-tests.” In tra- ditional statistics textbooks, there are still references to “Student’s t.”

4.2 Estimating the Standard Error of the Mean

The first problem Gosset solved was how to work around the need for the population standard error of the mean, sM, that the z-test requires. Although sM can be calculated by dividing the population standard deviation (s) by the square root of the number of data points (sM 5 s/ "n ), that’s small consolation if the value of s isn’t available.

Gosset reasoned that if samples were large and the individuals in the sample were ran- domly selected, the standard deviation of the sample, s, provides a reasonably accurate estimate of the value of the population standard deviation, s. Using s as one of the pos- sible values of s, Gosset decided to estimate the standard error of the mean. Instead of sM 5 s/ "n , he proceeded this way:

tan81004_04_c04_075-102.indd 77 2/22/13 3:39 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

Where

t 5 the calculated value of t M 5 the sample mean mM 5 the mean of the distribution of sample means, a value that will always

equal to the population mean SEM 5 the estimated standard error of the mean, which is based on the sample

standard deviation, s

In addition to the denominator in the test statistic, there is another fundamental difference between z and t. In Chapter 2 we noted that although there are many normal distributions, there is just one standard normal distribution, the z distribution. It is the one for which char- acteristics are described in Table 2.1. Because there is just one z distribution, a single value of z indicates outcomes that occur beyond the middle 95% of the distribution. By the most

common standard, z 5 61.96 indicates a statistically significant outcome.

There isn’t just one distribution of t, however. Because each distribution has different characteris- tics, each has a different critical value that indicates whether an outcome is statistically significant.

Degrees of Freedom and the One-Sample t-Test

Recall that degrees of freedom (df ) are the number of scores free to vary when the final value of some characteristic is known. For the one sample t-test, degrees of freedom are the number of scores in the sample minus 1, n 2 1.

• If the sample size (n) 5 10, df 5 9. • If n 5 30, df 5 29, and so on.

Key Terms: A critical value is a value from a table of such values indicating a statistically significant result.

Formula 4.2 t 5 M 2 mM

SEM

the greatest. Other things equal, a sample is least likely to reflect the characteristics of the population when it is smallest. Sample size was indirectly the index Gosset used to make the adjustment, as we shall see below.

4.3 The One-Sample t-Test

The one-sample t answers the same question the z-test answered: Does the sample belong to the population it is compared to? For that reason, the test statistics look very similar, except that the estimated standard error of the mean (SEM) is in the place of the population standard error of the mean, sM. That substitution makes the formula for the one-sample t-test look this way:

tan81004_04_c04_075-102.indd 78 2/22/13 3:39 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

Each different number of degrees of freedom, from 1 to ∞ (infinity), defines a different t distribution. The various degrees of freedom values each have their unique number, a particular critical value, to indicate whether the calculated t is statistically significant. Although the values are different for each different t distribution, the different critical values are all interpreted the same way. Plus or minus the critical value for p 5 .05 and the specific degrees of freedom for the problem indicates the range of t values containing the middle 95% of that particular distribution.

As we noted earlier, the risk that a sample won’t accurately reflect the population is great- est when the samples are smallest. Correspondingly, the critical values that a calculated t must meet to be statistically significant are largest for the smallest samples. As df increase, the critical values decline until ultimately, when df 5 ∞, the critical value for t matches that for z, as can be seen in the last line of the first column of Table 4.1. The critical value for a two-tailed t-test with p 5 .05 and df 5 ∞ is 1.96. Although a sample with an infinite num- ber of degrees of freedom isn’t possible in our world, the point still has value; it reminds us that as degrees of freedom increase, the distributions of t become increasingly similar to that of z. Note that the critical value for a two-tailed t-test with p 5 .05 and df 5 30 is around 2, which is close to 1.96. That is why the t-test is often recommended for samples sizes below 30. Past that point, if the data lend themselves to a z-test, it can be used.

Table 4.1: t Distribution critical values

df Two-tailed tests One-tailed tests

p 5 .05 p 5 .01 p 5 .05 p 5 .01

1 12.706 63.657 6.314 31.821

2 4.303 9.925 2.920 6.965

3 3.182 5.841 2.353 4.541

4 2.776 4.604 2.132 3.747

5 2.571 4.032 2.015 3.365

6 2.447 3.707 1.943 3.143

7 2.365 3.499 1.895 2.998

8 2.306 3.355 1.860 2.896

9 2.262 3.250 1.833 2.821

10 2.228 3.169 1.812 2.764

11 2.201 3.106 1.796 2.718

12 2.179 3.055 1.782 2.681

13 2.160 3.012 1.771 2.650

(continued)

tan81004_04_c04_075-102.indd 79 2/22/13 3:39 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

Table 4.1: t Distribution critical values (continued)

df Two-tailed tests One-tailed tests

14 2.145 2.977 1.761 2.624

15 2.131 2.947 1.753 2.602

16 2.120 2.921 1.746 2.583

17 2.110 2.898 1.740 2.567

18 2.101 2.878 1.734 2.552

19 2.093 2.861 1.729 2.539

20 2.086 2.845 1.725 2.528

21 2.080 2.831 1.721 2.518

22 2.074 2.819 1.717 2.508

23 2.069 2.807 1.714 2.500

24 2.064 2.797 1.711 2.492

25 2.060 2.787 1.708 2.485

26 2.056 2.779 1.706 2.479

27 2.052 2.771 1.703 2.473

28 2.048 2.763 1.701 2.467

29 2.045 2.756 1.699 2.462

30 2.042 2.750 1.697 2.457

 ∞ 1.96 2.576 1.645 2.326

Source: Critical Values of the t distribution (2011). Retrieved from http://shazam.econ.ubc.ca/intro/critval.htm.

The first column in the table is for degrees of freedom. The second and third columns indicate different criteria for statistical significance. Although we have relied on p 5 .05 (the third column), which is probably the most common standard, probably the next most common standard is p 5 .01. When testing at that level, just 1% of the most extreme out- comes are considered significant. The more rigorous standard minimizes the potential for erroneously finding a result to be significant (type I error).

Calculating the One-Sample t-Test

The steps for completing the one-sample t-test are the following:

1. Calculate the sample mean and the sample standard deviation. 2. Calculate the estimated standard error of the mean. 3. Calculate the value of t. 4. Compare the calculated t to the critical value for t for df 5 n 2 1.

tan81004_04_c04_075-102.indd 80 2/22/13 3:39 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

To illustrate, suppose that an area manager for a chain of elec- tronics retail stores is interested in whether sales for the 10 locations she oversees are at par with national sales figures for a particular month. She is informed that for that particular month, the national sales average is $268,000. She gathers the sales figures from the store managers, in thousands of dollars of sales, and they are as follows for stores 1–10, respectively:

223, 261, 295, 290, 337, 354, 236, 338, 240, 420

1. Verify that for this sample, the mean (M) 5 299.4, and the sample standard deviation (s) 5 62.713.

Recall that since m 5 268,000, mM 5 268,000, or in thousands, 268.

If the mean sales for the 10 stores and the mean for all stores in the country are used to create a bar graph in Excel, the result is Figure 4.1.

Figure 4.1: The mean for 10-store sample versus national sales for the month

There is certainly a visual difference between the region and the country. Is the difference great enough to be statistically significant? The one- sample t-test answers that question.

2. With the sample standard deviation and the sample size, SEM can be estimated with Formula 4.1:

Sample

Sales for the Month

Nation

350,000

300,000

250,000

200,000

150,000

100,000

50,000

0

Review Question B: Noting the conclusion we drew with the 10 outlets compared to national sales data for the month, what decision error might have occurred, type I or type II?

Review Question A: How are degrees of freedom for a t-test related to the critical values that determine when a result is statis- tically significant?

tan81004_04_c04_075-102.indd 81 2/22/13 3:39 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

If these were z-test results, the next step would be to compare the calculated value to 1.96, since the z distribution involves just 1 population with just 1 value indicating the standard for statistical significance. With the t-test, every time the degrees of freedom change, a different distribution is involved. Each t-distribution has its own critical value indicating the point at which a calcu- lated value of t becomes statistically significant.

4. In Table 4.1, Part A, “Critical Values for the Two-Tailed Test” move down the first (df) column to 9, and then over to the third column for testing at .05.

The critical value for t with df 5 9 is 2.262 written this way, t.05(9) 5 2.262.

Interpreting t-Test Results

Comparing the sales of the 10 stores to the national average resulted in t 5 1.583. Just as with z, when the calculated value is equal to, or larger than, the Table 4.1 value, t.05(9) 5 2.262, the outcome is statistically significant. The sales for these 10 stores were not significantly different from national data.

Clearly, the mean for the 10 stores is higher than the national mean. It was $299,400 com- pared with $268,000. But the difference isn’t statistically significant. The result indicates that a sample with M 5 $299,400 was probably one of those making up a distribution of sample means with mM 5 $268,000.

Another Example

In October 2011, in the midst of a recession, unemployment was high in the United States and underemployment was also a problem. Nationally, the length of the average work- week was 34.3 hours (Shierholz, 2011). Suppose that for a particular county, the lengths of the average workweek for specific industries were as follows:

t 5 M 2 mM

SEM 5

299.4 2 268 19.832

5 1.583

3. With SEM, M, and mM, the t statistic is,

SEM 5 s/"n

SEM 5 62.713/"10 5 19.832

tan81004_04_c04_075-102.indd 82 2/22/13 3:39 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

• farm workers: 28.2 hours • retail workers: 34.7 hours • clerical and office workers: 29.5 hours • manufacturing workers: 32.5 hours • food industry workers: 26.5 hours • housekeeping and yard maintenance: 25.8 hours

Is the length of the workweek in the county significantly different from the national work- week? Figure 4.2 is a visual presentation of the difference.

Figure 4.2: The length of the workweek: country versus county

1. Assuming that there were equal numbers of employees in each group, verify that for these employees M 5 29.533 and s 5 3.476.

Note that mM is 34.3.

2. SEM 5 s/"n 5 3.476/"6 5 3.476/2.449 5 1.419

3. t 5 M 2 mM

SEM 5

29.533 2 34.3 1.419

5 23.359

4. t.05(5) 5 2.571

The calculated t has a greater absolute value than the critical value from the table. At p 5 .05 level and df 5 5, the difference between the mean number of hours employees in the county work and the mean for all workers in the country is statistically signifi- cant. The negative sign indicates that length of the workweek in this county is signifi- cantly shorter than the national average, which means that underemployment is more pronounced in this county.

National

Hours in the Employment Week

County

40

35

30

25

20

15

10

5

0

tan81004_04_c04_075-102.indd 83 2/22/13 3:40 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

The One-Sample t-Test in Excel

While there is no procedure in Excel’s data analysis package for the one-sample t, it is not a difficult procedure to set up. One reason is it is based on statistics that Excel can readily produce. By way of an example, a roadside fruit stand has dollar sales in the following amounts for 10 successive days:

235, 340, 228, 430, 378, 394, 285, 312, 374, 425

A review of data for the same period over the last several years indicates that the annual mean for the period 5 $305. Are sales for these latest 10 days significantly different from historical figures?

In a blank Excel spreadsheet enter the following. Note that as always, the quotation marks indicate what is to be entered but are not part of the entry or command.

• Enter the label dollar sales in cell A1. Widen column A to accommodate the label and the output that will be produced below by placing the cursor on the vertical line to the right of “A” in the column, and then right-click and hold while you move the mouse to the right a few spaces.

• From cell A2 to A11 enter the daily sales data. • In cell C1 enter the label population mean 5. • In cell C2 enter the label sample mean 5. • In cell C3 enter the label standard deviation 5. • In cell C4 enter the label n 5. • In cell C5 enter the label square root of n 5. • In cell C6 enter the label standard error of the mean 5. • In cell C7 enter the label t 5. • In cell C8 enter the label critical value of t 5. • Widen column C enough to accommodate these labels. • At the top of page, click Data and then Data Analysis at the far right. • From the window highlight Descriptive Statistics and click OK. • The Input Range is A2:A11. • Click the Output Range button and enter A13. • Select Summary Statistics and click OK. • In cell D1 enter 305, which is the population mean. • From the results of the summary statistics note that the same mean is 340.1,

the sample mean. Enter that value in cell D2. • Enter the sample standard deviation, 73.196, in cell D3. • In cell D4 enter 10—the number of days for which we have data. • In cell D5 the entry is 5SQRT(D4). This will take the square root of 10. • In cell D6 the entry is 5D3/D5. • In cell D7 the entry is 5(D2-D1)/D6. • In cell D8 enter 2.262, the critical value of t at .05 with 9 degrees of freedom.

The screen-shot that is Figure 4.3 is the result of the above.

tan81004_04_c04_075-102.indd 84 2/22/13 3:40 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

Figure 4.3: A one-sample t-Test in Excel

Note that although the difference between the sample mean and the population mean seems substantial, the calculated t is not statistically significant. The difference between means is overwhelmed by the variability within the sample, variability gauged by the sample standard deviation, and the resulting estimated standard error of the mean. With a relatively large value in the denominator of the ratio, the difference between the means is not great enough to result in a statistically significant t value.

The One-Tailed Test

With two-tailed tests, whether the z or t values are positive or negative isn’t relevant to whether the result is statistically significant. Outcomes can be in either tail of the distribu- tion. As long as the value without regard to the sign is as large as, or larger than, the criti- cal value, the result is significant. Whether a sample is different from a population doesn’t specify the direction of the difference, but sometimes the issue is whether a sample mean is significantly greater than a population mean, and sometimes the question is whether it is significantly lower.

With that wording, the analysis becomes a one-tailed test. Language such as “more stressed” or “less productive” predicts how the sample is expected to differ from the popu- lation rather than just suggesting that difference. It specifies the tail of the distribution in which the difference will occur.

tan81004_04_c04_075-102.indd 85 2/22/13 3:40 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

The “Statistically Significant” Region in the One-Tailed Test

At issue in one- versus two-tailed tests is whether we can look to either tail of the distribution, or just one, for a significant difference.

• In two-tailed tests (is there a “difference?”), the most extreme 5% of the distribution is divided equally between the tails of the distribution, 2½% in each.

Distribution “A” in Figure 4.1 indicates a two- tailed test.

• In one-tailed tests (“does the sample have a significantly higher or signifi- cantly lower mean than the population?”) the entire 5% rejection region occurs in just one tail of the distribution.

For example, in an agricultural supply company, fertilizer sales over a 16 week period averaged 18.559 % of total sales, and the standard deviation was 7.928. An industry news- letter indicates that on average, statewide fertilizer sales were 15.002% of total sales in similar companies. As a percentage of total sales, are fertilizer sales significantly higher in this company?

With

M 5 18.559 mM 5 15.002

s 5 7.928 SEM 5 s/"n 5 7.928/"16 5 1.982

t 5 M 2 mM

SEM 5

18.550 2 15.002 1.982

51.795

Asking whether fertilizer sales, as a percentage of total sales at this company, were significantly higher than statewide sales makes this a one-tailed test. Rather than a rejection region equally divided between the two tails of the distribution, as it would be if the question were simply whether sales at this company were different from those statewide (Figure 4.4A), the entire rejection region is in the upper tail

Key Terms: A one-tailed test is one for which the direction of the difference between means (positive or negative) is pre- dicted, which allows for a greater rejection region in that one tail of the distribution.

Review Question C: The CEO of an auto- motive parts manu- facturer asks whether sales for the month are better than for past years. Is the request for a one- or two-tailed test?

tan81004_04_c04_075-102.indd 86 2/22/13 3:40 PM

CHAPTER 4Section 4.3 The One-Sample t-Test

of the distribution (Figure 4.4B). Note that doing this means that less extreme values of t reach statistical significance.

Figure 4.4: Two- and one-tailed t-tests

With df 5 15 (16 2 1), the critical value for two-tailed tests is t.05(15) 5 2.131, and a t value of 1.795 would not be statistically significant. But for a one-tailed test the critical value is t.05(15) 5 1.753, so a t value of 1.795 is significant. The percentage of fertilizer to total sales at this company is significantly greater the state average.

With an easier significance standard to meet for one-tailed tests, why don’t people use them exclusively? First of all, one-tailed tests presume that we can predict the direction of the difference. Often we don’t know. Even if there is a reason to suspect the direction, if something unusual occurs and there is an extreme value of t in the opposite direction of the one predicted, there is no way to detect the statistically significant result since the entire significance region is in the opposite tail. For these reasons, some analysts do not bother with one-tailed tests.

Lowest 2½% Highest 2½%

Highest 5%

A. Statistically Significant Regions in a Two-tailed Test.

B. Statistically Significant Regions in a One-tailed Test where Sample is Predicted to be Higher than the Population.

tan81004_04_c04_075-102.indd 87 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

4.4 Hypothesis Testing In statistical testing, two predictions cover all possible outcomes. Either the result will be significant or not. For the one-sample t-test,

• the null hypothesis predicts that the result is not significant. It is the hypothe- sis of no difference, and it indicates that the mean of the population from which the sample was drawn (m1) has the same value as the mean of the population to which it is compared (m2). The null hypothesis is written this way: Ho: m1 5 m2.

A statistically significant t value indicates that the sample probably doesn’t represent the population it was compared to; it is characteristic of some other population.

• This alternate hypothesis has this form: HA: m1 ≠ m2.

A statistically significant t prompts “rejecting the null hypothesis.” When a result is not statistically significant, rather than accepting the null hypothesis, the decision is to “fail to reject the null hypothesis.” This is because unless all population parameters are known with certainty, it would be extremely difficult to prove that m1 5 m2. Since statistics are about probabilities, a t value not great enough to reject the null hypothesis does not neces- sarily mean that m1 5 m2.

The Null and Alternate Hypotheses in a One-Tailed Test

In one-tailed tests the null hypothesis is still Ho: m1 5 m2; the prediction is still of no difference. The alternate hypothesis changes, however, the specific prediction. In the fertilizer example the alternate hypothesis was HA: m1 . m2, meaning that the mean of the population from which the sample was drawn was predicted to have a greater value than the pop- ulation to which the sample is compared. Had the prediction been that the outlet’s fertilizer average sales were lower than those statewide, the hypothesis would have been HA: m1 , m2.

4.5 The Independent Samples t-Test

Gosset’s one-sample t-test represents an important development to anyone who needs to make a judgment about whether a particular sample is characteristic of a specified population. Unlike the z-test, the t-test can be completed without a value for the popula- tion standard error of the mean (sM). The next test, Gosset’s independent samples t-test, requires no population values at all. Everything the test requires can be determined from the sample data.

Rather than “is the sample significantly different from the population?” perhaps the ques- tion is “are two samples significantly different from each other?” Suppose an engineer for a

Key Terms: All results of sta- tistical tests are covered by two possibilities: Either the results are significant, or they aren’t. The null hypothesis predicts a nonsignificant outcome. The alternate hypothesis predicts significance.

tan81004_04_c04_075-102.indd 88 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

tire manufacturer examines tread wear on two com- peting premium brands of tires and wishes to know whether the differences are statistically significant. In that case, the independent variable (IV) is brand, which is a nominal variable, and the dependent variable (DV) is tread wear, which can be measured on an interval or ratio scale. Questions that com- pare two samples to each other are the domain of the independent samples t-test.

When there is more than one group involved, whether the groups are related is relevant to the kind of analysis that can be performed. In this test, “inde-

pendent” means that the two samples involve separate groups of people; those who are members of one group cannot also be members of the other.

The Population Based on Difference Scores

The z- and one-sample t-tests were based on the distribution of sample means. The vari- ability we could expect among the individual samples that made up the distribution of sample means was the basis for posing questions about whether a particular sample was likely drawn from a specified population.

The independent samples t-test is based on another population, one created from “differ- ence scores.” Rather than sampling one group at a time and then plotting the individual means of each sample to create the distribution of sample means, this population is cre- ated by:

• repeatedly selecting two samples of the same size, • computing the mean for each, • subtracting the second mean from the first (M1 2 M2) and • plotting the differences between all pairs of means to create a frequency

distribution.

If this process were repeated an infinite number of times, the result would be a popula- tion that could be described as a distribution of difference scores. When the mean of the first sample is greater than the mean of the second sample (M1 . M2), the difference between the pair of sample means will be positive, a result plotted in the right half of the distribution. When M1 , M2, the difference is negative, a result plotted in the left half of the distribution. Most of the differences will be minor—differences plotted near the middle of the distribution.

The mean for the distribution of difference scores is symbolized mM12M2. The subscript,

M12M2, reminds us that the mean is based on the differences between pairs of sample means. Like the z distribution, the value of the distribution of difference scores is always 0, but not because it’s mathematically manipulated to be so. In spite of the fact that there is typically some difference between the first sample mean and the second, over many pairs the positive differences counterbalance the negative differences. If all those samples are truly drawn from the same population, the mean of all those M1 – M2 differences will be zero, mM12M2 5 0.

Key Terms: The independent samples t-test gauges whether two samples belong to popula- tions with the same mean. It’s based on the distribution of difference scores, a population of the differences between the means of all possible pairs of samples. Variability is measured by the standard error of the difference.

tan81004_04_c04_075-102.indd 89 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

Where

SEM1 2 5 the square of the estimated standard error of the mean for sample 1

SEM2 2 5 the square of the estimate standard error of the mean for sample 2

SEM 5 s/"n

Since the independent sample t-test involves two samples, the standard error of the mean must be calculated for both samples so that the standard error of the difference, SEd, can combine variance from both groups. For this reason, the SEd statistic is sometimes called an “estimate of pooled variance.”

The Hypotheses in the Independent Samples t-Test

The null hypothesis and alternate hypotheses in an independent t-test look the same as they do for the one-sample t but they’re explained differently:

Ho: m1 5 m2

HA: m1 ≠ m2

Formula 4.3 SEd 5 "SEM1 2 1 SEM2 2

Using the Distribution of Difference Scores to Explain Results

So whatever the population involved, the mean of the distribution of difference scores is always zero. The importance of the distribution of difference scores to the independent samples t-test is this: It indicates how much M1 2 M2 difference can be expected when random pairs of samples are drawn from the same population and their means compared. When the difference between the two samples in question exceeds the differences that are likely to occur by chance, the explanation is that the samples probably do not represent the same population.

The Standard Error of the Difference

However, the t statistic is not just the difference between the sample means. Like the z-test and the one-sample t-test, the t value is a ratio score. In this case, the ratio is based on the difference between the sample means (M1 2 M2) compared with how much variability there is within the two samples. The difference within the samples is a value measured by the standard error of the difference, sM12M2. Technically, sM12M2 is the standard devia- tion of all possible differences between the means of all pairs of groups that make up the distribution of difference scores, but for the independent t-test, the value is estimated with SEd. The SEd statistic is the estimated standard error of the difference, and it is calculated as follows:

tan81004_04_c04_075-102.indd 90 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

Where

M1 5 the mean of the first sample M2 5 the mean of the second sample SEd 5 the estimated standard error of the difference

For the independent samples t-test, the degrees of freedom for determining the critical t value are calculated as the sum of the number of scores in both samples minus two, n1 1 n2 2 2. So far, we have used samples of equal size.

• If the sample sizes (n) 5 10, df 5 10 1 10 2 2 5 18. • If n 5 15, df 5 15 1 15 2 2 5 28, and so on.

The steps for calculating the independent t-test are as follows:

1. Calculate the mean and standard deviation for both groups. 2. Calculate the standard error of the mean for both groups. 3. Calculate the standard error of the difference. 4. Solve for t. 5. Compare the calculated value of t to the critical value of t for n1 1 n2 2 2

degrees of freedom.

The null hypothesis predicts that the mean of the population represented by the first sam- ple is the same as the mean of the population represented by the second. In other words, the samples could have been drawn from or represent the same population.

The alternate hypothesis predicts that this is not the case, that the samples have been drawn from different populations. If the independent samples t-test is one-tailed, the alternate hypothesis indicates the direction of the predicted difference, just as with the one-sample t-test:

HA: m1,m2., or

HA: m1.m2.

The Test Statistic for the Independent Samples t-Test

With M1 2 M2 in the numerator and the estimate of the standard error of the difference in the denominator, the test statistic is:

Formula 4.4 t 5 M1 2 M2

SEd

tan81004_04_c04_075-102.indd 91 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

Formula 4.5 SEd 5 Å a 1n1 2 12s21 1 1n2 2 12s22

1n1 1 n2 2 22 a 1

n1 1

1 n2 b

An Independent Samples t-Test Example

An employment specialist is interested in whether optimism about the future differs between those who are employed and those who are unemployed. Two groups of individ- uals, one employed and one unemployed, are asked to complete a survey that measures optimism. In this case, the independent variable is whether the participant is employed or unemployed (a nominal variable), and the dependent variable is level of optimism expressed through answers to the survey questions (an interval or ratio scale). The data are as follows:

Unemployed: 5, 7, 7, 8, 11, 14, 15

Employed: 7, 10, 12, 15, 15, 16, 17

1. M1 5 9.571, s1 5 3.823, n1 5 7

M2 5 13.143, s2 5 3.625, n2 5 7

2. SEM1 5 s1/ "n1 5 3.823/"7 5 1.445

SEM2 5 s2/ "n1 5 3.625/"7 5 1.370

3. SEd 5 "1SEM12 1 SEM22 2 5 "11.4452 1 1.3702 2 5 1.991

4. t 5 M1 2 M2

SEd 5

9.571 2 13.143 1.991

5 21.794

5. df 5 n1 1 n2 – 2 5 7 1 7 – 2 5 12

6. The critical value for t.05(12) 5 2.179 (two-tailed test).

In this instance, the difference in optimism between the employed and the unemployed is not statistically significant, and the proper decision is to fail to reject the null hypothesis.

SEd for Unequal Samples

Formula 4.3 for the standard error of the difference assumes equal sample sizes. The qual- ity is indicated by the way both standard errors of the mean are treated equally in the formula—they are both simply squared and then added. When the sample sizes differ, the formula for the standard error of the difference must be adjusted to accommodate the inequality so that a sample with size of, say, n 5 8, doesn’t have the same impact on the resulting SEd value (and then the value of the t statistic) as one that has n 5 12. Formula 4.5 adjusts for sample sizes (n values) that are different for each group.

tan81004_04_c04_075-102.indd 92 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

Although the formula is lengthy it isn’t difficult. It involves just two statistics, which are the variance (the square of the standard deviation) and the size of each sample. Take care to follow the order of mathematical operations (work in parentheses first, for example) and the formula is quite straightforward. The steps for completing the SEd for unequal sample sizes are as follows:

1. Determine each group’s variance (s2). As the notation indicates, it is just the standard deviation, squared. If it is done in Excel, the command is VAR.

2. Take the number in the first group minus 1 (n1 2 1) and multiply the result by the variance for group 1 (s21).

3. Repeat step 2 for group 2. 4. Add the results of steps 2 and 3 together and divide their sum by the number

in the 2 groups, minus 2 (n1 1 n2 22). 5. Add the result of 1/n1 and 1/n2. 6. Multiple the result of step 4 and the result of step 5. 7. The last step is a square root of the result of step 6.

If Formula 4.5 is used for t-tests with equal sample sizes it will produce the same answer as Formula 4.3. It is generally avoided except for unequal n values because it involves more calculations, but the results are not different. When an independent samples t-test is done on Excel, no adjustment needs to be made. The software was written to use the formula for unequal sample sizes for everything and so it makes the adjustment automatically.

The Independent Samples t-Test with Unequal Sample Sizes

Referring back to the problem on the level of optimism among the employed versus the unemployed, suppose that the unemployed person with the score of 15 is no longer will- ing to participate in the study and insists that the score be deleted. The data in that case will be the following:

Unemployed: 5, 7, 7, 8, 11, 14

Employed: 7, 10, 12, 15, 15, 16, 17

M1 5 8.667, s1 5 3.266, n1 5 6

Group 2 is unchanged, M2 5 13.143, s2 5 3.625, n2 5 7

tan81004_04_c04_075-102.indd 93 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

SEd 5 Å a 1n1 2 12s21 1 1n2 2 12s22

1n1 1 n2 2 22 b a 1

n1 1

1 n2 b

SEd 5 Å a 16 2 123.2662 1 17 2 123.6252

16 1 7 2 22 b a 1 6

1 1 7 b

SEd 5 Å a 15210.667 1 16213.141

1112 b 1.167 1 .1432

SEd 5 "112.016 3 .312 5 1.930

t 5 M1 2 M2

SEd 5

8.667 2 13.143 1.930

5 22.319

The critical value for t.05(11) 5 2.201 (two-tailed test).

How is it that the problem that was earlier not statistically significant becomes significant with one individual removed from group one? First of all, note that it happens to have been the most extreme score that was eliminated. Eliminating that one most extreme score impacts the outcome in two ways. In this particular instance it increases the difference between means—the mean of the first group was already lower than that of the second group, and eliminating the highest score from the lower group increased that difference.

Second, recall that the t-score is a ratio. Whether it’s based on the one-sample t or the inde- pendent samples t-test, the t-value is a ratio of the difference between (either the sample and the population, or between two samples) to the difference within (measured by either the standard error of the mean or the standard error of the difference). Either increasing the difference between or decreasing the difference within makes the t-value larger. In this case, eliminating that one score did both. The result is that even with a larger critical value to meet because df 5 11 rather than 12, the calculated t value exceeds the table value and is statistically significant at p 5 .05.

If the adjustment for unequal sample sizes had not been made by using Formula 5.5, and SEd were calculated for this problem as though the sample sizes were equal, what would its value have been?

SEM1 5 s1/ "n1 5 3.266/ "6 51.333 SEM2 5 s2/ "n2 5 3.625/ "7 51.370—the same as before, of course

SEd 5 "1SEM12 1 SEM22 2 5 "11.3332 1 1.3702 2 5 1.911

It is not a large adjustment compared to the SEd 5 1.930 from the unequal sample sizes formula, but neither was the change in sample size very large (from n 5 7 to n 5 6), and that is what Formula 4.5 is designed to do: adjust for sample size differences.

tan81004_04_c04_075-102.indd 94 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

Assumptions Associated With the Independent Samples t-Test

Each statistical test is based on certain assumptions or conditions. Those associated with the independent samples t-test are as follows:

1. The samples are independent. 2. The participants in each group are randomly selected. 3. The two samples have similar variability, or homoscedasticity. Although

there are tests to indicate the point at which homoscedasticity has been vio- lated, Excel does not provide one. We will just say that when measures of variability (R, s, or s2) are quite different from one group to the other, homoscedasticity should not be assumed, and there is an adjustment that should be made when completing a problem with Excel.

4. The measures of the dependent variable (the variables on which the groups are being compared) need to be continuous, which requires an interval or ratio scale.

Although conditions 1 and 4 are fairly strict require- ments, the t-test is fairly robust in the face of viola- tions to conditions 2 and 3. It takes something pretty obviously wrong with the way the subjects are selected, or with the comparative variability of the two samples, to make the independent samples t-test inappropriate.

The Independent Samples t-Test in Excel

Suppose that an automotive parts and accessories chain is experimenting with a new sales promotion. Two similar stores are selected for the experiment. For store 1, nothing changes. This store constitutes the control group. For store 2, the treatment group, the pro- motion is implemented. Sales in hundreds of dollars over a five-day period are as follows:

Control: 2, 5, 2, 4, 7, 1, 2, 3, 4, 5

Treatment: 6, 6, 7, 10, 12, 9, 6, 5, 5, 7

The expectation that sales will be higher in the treatment group makes this a one-tailed test; the alternate hypothesis is m1 , m2 Using Excel to determine whether differences are statistically significant, the process is as follows:

1. Create a data set in Excel. With a column for each group, enter the names con- trol and treatment in cells A1 and B1 respectively.

2. Enter the scores in their respective columns, beginning in A2 and working down for the control group and B2 and then down for the treatment group.

3. Click, the Data tab, and then Data Analysis. Since the range values for the 2 groups are fairly similar, 5 for the control

group and 7 for the treatment group, we will assume equal variances. 4. Select t-Test: Two-Sample Assuming Equal Variances. 5. Click OK.

Key Terms: Data are homosce- dastic when they are distributed similarly across multiple groups.

tan81004_04_c04_075-102.indd 95 2/22/13 3:40 PM

CHAPTER 4Section 4.5 The Independent Samples t-Test

6. The language that Excel uses refers to the groups as “variables.” The control group will be variable 1. Indicate that the range for the data in this group is A2:A11. Indicate that the range for variable 2, the treatment group, is B2:B11.

7. Indicate that the hypothesized mean difference is 0 by entering 0 in the box. The 0 indicates that if the promotion has no effect, the two groups will have the same mean, which is the null hypothesis, m1 5 m2.

8. Select a range for the output, say C15, and click OK. 9. Widen column C so that all the output shows and the result is Figure 4.5.

Figure 4.5: The independent samples t-test in Excel

The output provides descriptive statistics including the means and variances for each group, as well as a statistic that Excel calls the “pooled variance” that we didn’t calculate longhand, but the t value is the same as it would have been had the calculations been done longhand.

• Excel provides the critical value for both one- and two-tailed tests, but rather than indicating whether the result is statistically significant, the program calculates the probability that this value of t could have occurred by chance if the samples both came from the same population.

• Any time p 5 .05 or less, the result is statistically significant. The lower the p value in the result, the less likely the difference between the groups is to have occurred by chance.

• The data indicate that both the one-tailed and two-tailed tests have p values less than .05. The results for both the one- and two-tailed are significant for these data.

tan81004_04_c04_075-102.indd 96 2/22/13 3:40 PM

CHAPTER 4Section 4.6 Determining Practical Significance

Where

M1 2 M2 is the absolute value of the difference between the means of the two samples, and sdv is the standard deviation for all the data/scores in both groups.

For the sales promotion problem, the means of the two samples were:

M1 5 3.50

M2 5 7.30

The denominator in Cohen’s d is the standard deviation of all 20 scores together. Don’t make the mistake of trying to just calculate a mean for the standard deviations from the two groups. The standard deviation of all 20 scores together can be quite different from the standard deviations of the two groups, averaged.

sdv 5 2.817

d 5 1M1 2 M2 2

Sdv

d 5 13.50 2 7.302

2.817

d 5 1.349

It is the absolute value of this statistic that matters. It is an indicator of how much impact the independent variable (whether the sales promotion was implemented) had on the dependent variable (sales). Cohen’s interpretations of the d value are as follows:

Review Question D: If a production manager wishes to compare the performance of the day shift with that of the swing shift in terms of number of units manufactured, what is the indepen- dent variable and what type of data scale is it? What is the dependent variable and what type of data scale is it?

Formula 4.6 d 5 M1 2 M2

Sdv

4.6 Determining Practical Significance

Results that are statistically significant are not always important. Strictly speaking, “statistically significant” means that a result isn’t likely to have occurred by chance; it is not a random occurrence. But significance does not indicate necessarily that the results are important. To deal with the question of whether an outcome has practical relevance, one option is to calculate an effect size. One of the more common effect sizes for the independent samples t-test is Cohen’s d. For the sales promotion problem in Figure 4.5 Cohen’s d is determined with Formula 4.6 as follows,

tan81004_04_c04_075-102.indd 97 2/22/13 3:40 PM

CHAPTER 4Chapter Summary

d 5 .2, a “small” effect d 5 .5 a “medium” effect d 5 .8 a “large” effect

These guidelines are “point indicators” rather than ranges of values. Cohen leaves unanswered the question of how to interpret, for example, a value between .2 and .5, or between .5 and .8. However, with d 5 1.349, the effect of the promotion is “large.”

In situations where t is not significant, there is no need to calculate Cohen’s d because any difference between means is chalked up to sampling error.

Chapter Summary t-tests have noted benefits over z-tests. In the one-sample test, samples can be compared to populations without the need for determining sM, which may not be available. The independent samples t-test doesn’t require any parameters at all to examine two inde- pendent samples for significant differences. The independent samples t is widely used in research. It has been a test of tremendous importance (Objective 1).

In hypothesis testing there are two predictions relevant to t-tests. In the case of the one- sample test, either the sample represents the population to which it is compared, or it does not. In the case of the independent t, either both samples belong to populations with the same mean, or they do not (Objective 2). The hypotheses simplify the way we report statistical results, although the use of one- versus two-tailed tests adds another element. One-tailed tests provide an alternate hypothesis that is directional. It predicts how the mean of the population represented by the first group will differ from the mean of the population represented by the second, not just that the means differ (Objectives 3 and 4).

There is an important harmony in each of the statistical procedures discussed so far.

1. The numerator for each formula involves a difference score. • For the z score the numerator was x 2 M. • For the z-test the numerator was M 2 mM. • For the one-sample t-test, the numerator was also M 2 mM. • For the independent samples t-test, the numerator was M1 2 M2

2. The denominator for each formula involves a measure of data variability. • For the z score the denominator was s. • For the z-test the denominator was sM. • For the one-sample t-test, the numerator was SEM. • For the independent samples t-test, the denominator is SEd (Objective 2).

Be careful not to confuse the type of t-test with the type of hypothesis. The one-sample t-test compares the sample to the population. The one-tailed t-test makes a prediction about how the first group differs from the second (Objective 3).

Key Terms: Statistical sig- nificance means only that a result is probably not random. To determine whether the result is important, effect size is calculated. Cohen’s d is one of several effect size measures. It indicates the practical impor- tance of a significant outcome.

tan81004_04_c04_075-102.indd 98 2/22/13 3:40 PM

CHAPTER 4Chapter Formulas

Two questions can be involved in an independent t-test. The first is whether differences between samples are statistically significant, a question addressed by the value of t. A second question is whether the difference has practical importance, something to which effect sizes respond. Cohen’s d answers this second question. Cohen’s d indicates how large the effect of the independent variable is on the dependent variable (Objective 5).

Answers to Review Questions

A. As degrees of freedom increase, the critical values decline. B. Since the t-value indicated that the result wasn’t statistically significant, the

only possible decision error is a Type II, or beta error. C. Asking whether the sales were “better than” makes this a one-tailed test. D. The independent variable is the shift—day or swing, which is a nominal scale.

The dependent variable is the number of units completed, which is a ratio scale.

Chapter Formulas

Formula 4.1 SEm 5 s/"n The estimated standard error of the mean

Formula 4.2 t 5 1M 2 mm 2

SEm The one-sample t-test

Formula 4.3 SEd 5 "1SEm12 1 SEm22 2 The estimated standard error of the difference for equal sample sizes

Formula 4.4 t 5 1M1 2 M2 2

SEd The independent samples t-test

Formula 4.5 SEd 5 Å a 1n1 2 12s21 1 1n2 2 12s22

1n1 1 n2 2 22 b a 1

n1 1

1 n2 b

Formula 4.6 d 5 M1 2 M2

Sdv Cohen’s d, a measure of effect size

Management Application Exercises

Unless otherwise stated, use p 5 .05 in all your answers.

1. A human resource specialist interviewed eight applicants for a position. Their interview scores are: 67, 55, 88, 74, 69, 81, 72, 70.

a. What is the value of SEM? b. Are the scores of this group representative of the population of applicants

for whom m 5 66.0? c. Does the human resource specialist need to interview more applicants or

should she choose from among the eight applicants already interviewed?

The standard error of the difference when sample sizes are unequal

tan81004_04_c04_075-102.indd 99 2/22/13 3:40 PM

CHAPTER 4Management Application Exercises

2. Another group of interviewees scored 52, 58, 64, 65, 67, 68, 69, 70. a. Are their scores significantly different from those in item 1? b. Is this group significantly better or significantly worse than the first group?

3. A shift manager at a manufacturing plant gathers data on the number of units workers assemble during 2 different shifts over 10 different days. If the number of units assembled by each shift varies greatly from day to day, what impact will that have on the likelihood of a significant difference between the two shifts?

4. A filling station operator tracks the number of gallons sold per hour during her eight-hour shift. They are as follows: 375, 400, 425, 425, 490, 500, 510, 530. If the mean number of gallons sold per hour for a week is m 5 500, are the data from this shift representative of the whole week?

5. If an analyst compares the ages of managers and assistant managers in retail outlets to the ages of floor personnel and finds that age differences not signifi- cant, how are the existing differences between the mean ages of the two groups explained?

6. An advertising agency is commissioned to develop a store display advertising a new detergent. To gauge the effectiveness of the display, the number of detergent boxes sold in six stores where the display was placed (Group 1) is compared to the number of detergent boxes sold in six similar stores where the display was not present (Group 2). The data are as follows:

Group 1: 13, 15, 12, 17, 14, 14. Group 2: 10, 12, 12, 11, 13, 9.

a. Is the difference significant? b. Write out the null and the alternate hypotheses for this problem. c. What is Cohen’s value, and what does it indicate in this case? d. Was the store display effective in boosting the sales of the new detergent?

7. A medical group wants to analyze differences in patients’ attitudes about the care they receive when a percentage of their bills is rebated. Group A receives a rebate; group B does not.

Group A: 23, 24, 27, 27, 29, 32, 35 Group B: 16, 17, 17, 18, 19, 19

a. Are the differences statistically significant? b. What formula for the standard error of the difference must be used? c. What will Cohen’s d explain?

8. With reference to problem 7: a. What are the IV and DV? b. What is the type of data scale for the IV and the DV? c. What is the most appropriate statistical test to investigate this question?

tan81004_04_c04_075-102.indd 100 2/22/13 3:40 PM

CHAPTER 4Key Terms

9. The owner of a small bakery has installed solar panels on the roof in the hope of reducing electricity costs. For 12 months before the installation and 12 months immediately after, the monthly costs were as follows in hundreds of dollars:

Before: 13, 12, 14, 14, 13, 17, 15, 16, 15, 17, 16, 19 After: 9, 9, 11, 12, 12, 11, 13, 14, 13, 11, 14, 13

a. What are the null and alternate hypotheses? b. Is the appropriate test for this question one- or two-tailed? c. Is t significant? d. Has there been a significant reduction in electricity costs? e. Is the result important (d)?

10. Ten concession sales employees who sell sandwiches at a sports event have sales of M 5 144.55, s 5 12.57. Ten others who sell Mexican food have sales of M 5 159.97 and s 5 11.10.

a. Do those who sell Mexican food have significantly greater sales? b. How much of the difference in sales can be explained by the type of food?

Key Terms

• A critical value is a value from a table of critical values that indicates the point at which a calculated value is no longer a random outcome; it is statistically significant.

• An independent samples t-test is a test of whether two samples belong to popula- tions with the same mean.

• The distribution of difference scores is a population based on the differences between the means of all possible pairs of samples belonging to a common popu- lation and what the independent samples t-test is based on.

• The one-tailed test is one for which the direction of the result is predicted. Rather than a hypothesis that the sample is significantly different from the population, for example, circumstances might favor predicting that the sample has a mean signifi- cantly greater than that of the population. In such instances, one looks for a differ- ence only in one tail of the population—the right tail for significantly greater, or the left tail for significantly less, thus the name.

• All results of statistical tests are covered by two possibilities: Either the results are significant, or they aren’t. The null hypothesis predicts a nonsignificant outcome. The alternate hypothesis predicts significance.

• Data variability within the two groups involved in an independent t-test is measured by the standard error of the difference.

• A statistically significant result is unlikely to have occurred by chance, but it does not necessarily have practical importance. Effect sizes indicate the real-world impor- tance of the result.

• The independent t-test is based on the assumption that data in both groups are distributed similarly. The technical word for that similarity is that the data are homoscedastic.

• Cohen’s d is one of several effect size measures. It indicates the practical impor- tance of a significant independent t-test outcome.

tan81004_04_c04_075-102.indd 101 2/22/13 3:40 PM

tan81004_04_c04_075-102.indd 102 2/22/13 3:40 PM