06244 - 4 Pages within 8hrs
The Standard NormalDistribution and z Scores
Keren Su/Corbis
Chapter Learning Objectives
After reading this chapter, you should be able to do the following:
1. Identify the characteristics of the standard normal distribution.
2. Demonstrate the use of the z transformation.
3. Determine the percent of a population above a point, below a point, andbetween two points on the horizontal axis of a normal distribution.
4. Calculate z scores using Excel.
5. Describe alternative standard scores.
6. Demonstrate the use of the modified standard score.
Introduction
The data that describe characteristics of groups come from either samples or populations,explained in the first two chapters. By way of reminder, recall that populations include allpossible members of any specified group. All university students, all psychology majors, allresidents of Orange County, and all left-handed male tennis players in their 20s are eachdescriptions of a population. We rely on Greek letters, such as µ for the mean and σ for thestandard deviation, to distinguish population parameters from the statistics that describesamples. (The word parameter indicates a characteristic of a population.) Remove one ormore individuals from any population, and the resulting group is a sample.
As we were describing populations, we noted that some are “normally distributed.” Thesecharacteristics indicate normality: (a) data distributions are symmetrical, (b) all the measuresof central tendency have very similar values, and (c) the value of the standard deviation isabout one-sixth of the range.
Data normality does not simply mean that the frequency distribution will appear as a bell-shaped curve; it means that predictable proportions of the entire population will occur inspecified regions of the distribution, and this holds for all normal data distributions. Forexample, the region under a normal curve from the mean of the population to one standarddeviation below the mean always includes 34.13% of the area under the curve. Because normal distributions are symmetrical, from the mean to one standard deviation above themean also includes 34.13%, so from +1σ or −1σ includes about 68.26% of the area under thecurve in any normally distributed population. As long as the data are normally distributed,those percentages hold true. Since many mental characteristics are normally distributed,researchers can know a good deal about such a characteristic without actually gathering thedata and doing the analysis. Whether the characteristic is intelligence, achievementmotivation, anxiety, or any other normally distributed characteristics, the proportion of thedistribution within +1 or −1 standard deviation from the mean will be the same:
· If a particular intelligence scale has µ = 100 and σ = 15, about 68% of any generalpopulation will have intelligence scores between 85 and 115.
· Likewise, if an achievement motivation scale has µ = 40 and σ = 8, about 2/3 of anypopulation will have achievement motivation scores from 32 to 48.
· And for an anxiety measure with µ = 25 and σ= 5, about 68% of any generalpopulation will have scores between 20 and 30.
The consistency in the way so many characteristics are distributed affords a good deal ofinterpretive power. Anyone who needs information about the likelihood of individualsscoring in certain areas of a distribution has an advantage when data are normally distributed.In addition to the 68% of any general population likely to score between +1σ and −1σ,
· from µ to +2σ is about 47.72% of the population, so about 95% (2 × 47.72) of thepeople in any general population will have intelligence scores between 70 (100 − 30)and 130 (100 + 30).
· from +3σ (49.87%) to −3σ includes nearly everyone in any normally distributedpopulation (2 × 49.87 = 99.74).
These observations emphasize that, sometimes, isolated bits of data can be quite informative.When a 12-year-old with an intelligence score of 170 pops up on YouTube, it is immediatelyapparent that this is a very unusual child. An intelligence score of that magnitude is about4.667σ (170 − 100 = 70; 70 ÷ 15 = 4.667) beyond the mean of the general population. If from+3σ to −3σ includes more than 99% of the population, from +4.667σ to −4.667σ mustinclude all but the utmost extreme scores. We obtain an even better context for how common(or uncommon) particular measures may be when we can determine the precise probability oftheir occurrence.
This is a self-assessment and will not affect your grade. You may only take this pre-test once.
Test Ch 3: The Standard Normal Distribution and z Scores
Top of Form
1. The z transformation changes any raw score into a z score so that it fits the standard normal distribution.
· a. TRUE
· b. FALSE
2. Calculating z scores doesn’t alter the distribution; it just makes them fit a distribution where the mean is 0 and the standard deviation is 1.0.
· a. FALSE
· b. TRUE
3. We can apply the z transformation to sample data even when there is reason to believe that the population from which the sample was drawn is not normally distributed.
· a. TRUE
· b. FALSE
4. Raw scores can be determined if the mean and standard deviation are available.
· a. FALSE
· b. TRUE
5. Individual scores in the standard normal distribution are called z scores.
· a. TRUE
· b. FALSE
Finish
Bottom of Form
3.1 A Primer in Probability
Scholars, data analysts, and in fact people on the whole are rarely interested in outcomes thatoccur every time. If everyone had an intelligence score of 170, no one would pay anyattention to someone with such a score. The fact that we know it to be uncommon is whatpiques our curiosity.
If we are not interested in events that always occur, neither do we closely follow events thatnever occur. If no one had ever had an intelligence score of 170, probably no one wouldwonder about what such a score means for the person who has it. The things that occur someof the time, however, intrigue us. The “some of the time” indicates that the event has some probability, or likelihood, of occurrence.
· What is the probability that those newlyweds will divorce?
· How likely is Germany to win the World Cup?
· What is the probability that an earthquake will occur on a particular day for someonewho lives near the San Andreas Fault?
· What is the probability of an IRS audit for one taxpayer?
Joseph Sohm/Visions of America/Corbis
The probability that something willoccur, such as how likely it is thatour favorite baseball team will winthe World Series, intrigues us and isan important component in thedecision-making process.
Because all of the items listed have happened in thepast and because their occurrence is important to atleast someone, people are interested in theprobability of those occurrences whether or notthey use the language of probability. When statednumerically, probability values range from 0 to 1.0.Something with a probability of zero (p = 0) neveroccurs. On the other hand, p = 1.0 indicates that theevent occurs every time, and p = 0.5 indicates thatthe event occurs 50% of the time.
As that last point indicates, percentages can beconverted to probability values. Dividing the percentage of times an event occurs by 100indicates the associated probability of the event.
Returning to the intelligence scores, we see thatbecause about 68% of the population has intelligence scores between 85 and 115, theprobability (p) that someone selected at random from the general population will have a scoresomewhere between 85 and 115 is 0.68 (68.26/100, if the result is rounded to two decimalplaces).
What is the probability that someone selected at random from the general population willhave an intelligence score of 100 or lower? Because 100 is the mean for intelligence scores,and because 50% of the population occur at the mean or below, p = 0.5.
What is the probability that someone selected at random will have an intelligence scorehigher than 115? First, we noted earlier that 34.13% of the population falls between themean, µ, and one standard deviation above the mean at σ = +1.0 in any normally distributedpopulation. In terms of intelligence score values, that is the region between scores of 100 and115. Since 50% of any normally distributed population will occur at the mean and above, ifwe subtract from 50% that portion between the mean and one standard deviation above themean, the remainder will be the portion of the distribution above 115: 50% − 34.13% =15.87%; that is, 15.87% of all intelligence scores in a normally distributed population willoccur above 115. Dividing by 100 (15.87/100 = 0.1587) and rounding the result to twodecimal places produces the probability p = 0.16.
By the same logic, because a score of 85 is one standard deviation below the mean, theprobability p = 0.16 means that someone selected at random from the population will scorebelow 85. If we combine the two outcomes, the probability is p = 0.32 that someone from thepopulation will score either below 85 or above 115.
Consider the number line shown in Figure 3.1.
Figure 3.1: Standard deviations for intelligence scores
The number line shows the portion of scores that fall within two standarddeviations above and below the mean. If M = 100, we can know the probabilityof someone scoring below 85 or above 115.
If this number line represents all intelligence scores ranging from two standard deviationsbelow to two standard deviations above the mean, we can see the percentages of thepopulation that will have scores in the designated areas. Using the percentages and dividingby 100 indicates the probability of a score in any of the designated areas.
Recall that the lowest probability for any value is zero (p = 0). If p = 0, then the event oroutcome never occurs. There is no such thing as a negative probability.
3.2 The Standard Normal Distribution
Not all populations are normally distributed. Home sales are usually reported in terms of the medianprice of a home, and salary data are likewise reported as median values. Those cases use the mediansbecause the related populations are very unlikely to be normally distributed and, as a measure ofcentral tendency, medians are less affected by extreme values than are means. A few very high salariesor home values create positive skew in the resulting distribution. In contrast, when it comes to, say,mental characteristics such as intelligence, achievement motivation, problem-solving ability, verbalaptitude, reading comprehension, and so on, population data are often normally distributed.
Hello Lovely/Corbis
When evaluating information aboutpeople’s characteristics, keep in mindthat data are often normally distributed.
Although there are many normal distributions all havingthe same proportions, each has different descriptivevalues. An intelligence test might have µ = 100 and σ =15 points. A nationally administered reading test mighthave a mean of 60 and a standard deviation of 8. Thesedifferent parameters can make it difficult to compare oneindividual’s performance across multiple measures. Asone author noted regarding scores from the WechslerIntelligence Test for Children (WISC), “A raw score of 5on one [sub]test will not have the same meaning as a rawscore 5 on another [sub]test” (Brock, 2010).
One way to resolve this interpretation problem is toconvert the scores from different distributions into acommon metric, or measurement system. If researchersalter scores from different distributions so that they both fit the same distribution, they can comparescores directly. A researcher can compare them directly to determine, for example, on which test anindividual scored highest. Such comparisons are one of the purposes of the standard normaldistribution.
The standard normal distribution looks like all other normal distributions—from the mean to +1standard deviation includes 34.13% of the distribution, for example. What separates it from the othersis that in the standard normal distribution, the mean is always 0, and the standard deviation is always1.0 (Figure 3.2). Other distributions may have fixed values for their means and standard deviations, buthere µ is always 0 and σ is always 1.0.
Figure 3.2: The standard normal distribution
In the standard normal distribution, the mean is always 0, and the standard deviation isalways 1.0.
The Standard Normal, or z, Distribution
Although various normal distributions have different means and standard deviations, they all mirroreach other in terms of how much of their populations occur in particular regions. The standard normaldistribution’s advantage is that the proportions of the whole that occur in the various regions of thedistribution have been calculated. That means that if data from any normal distribution are made toconform to the standard normal distribution, we can answer questions about what is likely to occur invirtually any area of the distribution, such as how likely it is to score 2.5 standard deviations below themean on a particular test, or what percentage of the entire population will likely occur between twospecified points. All such questions can be answered when adapting normal data to the characteristicsof the standard normal distribution.
Individual scores in the standard normal distribution are called z scores, which is why the standardnormal distribution is often called “the z distribution.” The formula used to turn scores from anynormal distribution into scores that conform to the standard normal distribution is the ztransformation:
Formula 3.1
z=x−Ms
where z is a score in the standard normal distribution, x is the score from the original distribution(often called a “raw” score), M is the mean of the scores before the original distribution, and s is thestandard deviation of the scores from the original distribution.
Because normality is characteristic of only very large groups, samples will rarely be normal. However,we can apply the z transformation to sample data when there is reason to believe that the populationfrom which the sample was drawn is normally distributed. This is what Formula 3 reflects. The M and s indicate that the data involved are sample data. In those situations where an analyst has access topopulation data—a social worker has all the data for those served by Head Start in a particular county,for example—µ replaces M and σ replaces s in the formula. With either sample or population data, thetransformation is from data that can have any mean and standard deviation to a distribution where themean will always equal 0 and the standard deviation will always equal 1.0.
To turn raw scores into z scores, perform the following steps:
1. Determine the mean and standard deviation for the data set.
2. Subtract the mean of the data set from each score to be transformed.
3. Divide the difference by the standard deviation of the data set.
For example, consider a psychologist interested in the level of apathy among potential voters regardingmental health issues that affect the community. Scores on the S ummary o f WH o’s A pathetic T est (theSoWHAT for short), an apathy measure, are gathered for 10 registered voters:
5, 6, 9, 11, 15, 15, 17, 20, 22, 25
What’s the z score for someone who has an apathy score of 11?
· Verify that for these 10 scores, M = 14.5 and s = 6.737.
· The z score equivalent for an apathy score of 11 is
z=x−Ms=11−14.56.737=−0.5195
An apathy score of 11 translates into a z score of −0.5195. Because the mean of the z distribution is 0and the standard deviation in the z distribution is 1.0, where would a score of −0.5195 occur on thehorizontal axis of the data distribution? It would be a little over half a standard deviation below themean, right? Figure 3.3 shows the z distribution and the point about where a raw score of 11 occurs inthis distribution once it is transformed into a z score.
It is important to know that the z transformation does not make data normal. Calculating z scores doesnot alter the distribution; it just makes them fit a distribution where the mean is 0 and the standarddeviation is 1.0. Evaluating skew and kurtosis must allow the analyst to assume that the data arenormal before using the z transformation.
With a mean of 0 in the standard normal distribution, half of all z scores—all the scores below themean—are going to be negative. A raw score of 11 from the SoWHAT data is lower than the mean,which was M = 14.5, so it has a negative z value (−0.5195).
Try It!: #1
How many standard deviations from themean of the distribution is a z score of 1.5?
Besides indicating by its sign whether the zscore is above or below the mean, the value ofthe z score indicates how far from the mean the zscore is in standard deviations. If a score had a zvalue of 1.0, it would indicate that the score isone standard deviation above the mean. The zscore for the raw score of 11 was −0.5195,indicating that it is just over half a standarddeviation below the mean. This ease of interpretation is one of the great values of z scores: the sign ofthe score indicates whether the associated raw score was above or below the mean, and the value ofthe score indicates how far from the mean the raw score falls, in standard deviation units (Fischer andMilfont, 2010).
Figure 3.3: Location of a score on the z distribution
Half of all z scores will fall below the mean, resulting in a negative value. A score of z= −0.5195 is slightly less than one-half a standard deviation below the mean.
Comparing Scores from Different Instruments
Consider another application of the standard normal distribution. A counselor has intelligence andreading scores for the same person and wishes to know on which measure the individual scored higher.Table 3.1 shows the data for the two tests. On the intelligence test, the individual scored 105, and onthe reading test, the individual scored 62.
Table 3.1: Reading and intelligence test results
|
Test |
Mean |
Standard deviation |
|
Intelligence |
100 |
15 |
|
Reading |
60 |
8 |
If the counselor transforms both scores to make them fit the standard normal distribution, they can becompared directly.
The z for the intelligence score is
z=x−Ms=105−10015=0.333
The z for the reading test score is
z=x−Ms=62−608=0.250
The intelligence score of 105 and the reading score of 62 are difficult to compare because they belongto different distributions with different means and standard deviations. When both are transformed tofit the standard normal distribution, an analyst can directly compare scores. The larger z value forintelligence makes it clear that individual scored higher in intelligence than in reading.
Expanding the Use of the z Distribution
Because the standard normal distribution is a normal distribution, we know that predictableproportions of its population will occur in specific areas. As we noted earlier, however, thoseproportions are known in great detail for the z distribution because this population is so often used toanswer detailed questions about the likelihood of particular outcomes. Table 3.2 indicates how muchof the entire population is above or below all of the most commonly occurring values of z. So, bytransforming scores from other distributions to fit the z distribution, we can use what we know aboutthis population to answer questions about scores from any normal distribution.
Not all tables for z values are alike. Probably as a matter of the developer’s preference, some tablesindicate the percentage of the population below a point. Some indicate the percentage between a pointand the mean of the distribution. Some indicate the probability of scoring in a particular area, and soon. This particular table indicates the proportion of the population between the specified value of z andthe mean of the distribution. (Table 3.2 is listed as Table B.1 in Appendix B.)
Table 3.2: The z table
|
|
0.00 |
0.01 |
0.02 |
0.03 |
0.04 |
0.05 |
0.06 |
0.07 |
0.08 |
0.09 |
|
0.0 |
0.0000 |
0.0040 |
0.0080 |
0.0120 |
0.0160 |
0.0199 |
0.0239 |
0.0279 |
0.0319 |
0.0359 |
|
0.1 |
0.0398 |
0.0438 |
0.0478 |
0.0517 |
0.0557 |
0.0596 |
0.0636 |
0.0675 |
0.0714 |
0.0753 |
|
0.2 |
0.0793 |
0.0832 |
0.0871 |
0.0910 |
0.0948 |
0.0987 |
0.1026 |
0.1064 |
0.1103 |
0.1141 |
|
0.3 |
0.1179 |
0.1217 |
0.1255 |
0.1293 |
0.1331 |
0.1368 |
0.1406 |
0.1443 |
0.1480 |
0.1517 |
|
0.4 |
0.1554 |
0.1591 |
0.1628 |
0.1664 |
0.1700 |
0.1736 |
0.1772 |
0.1808 |
0.1844 |
0.1879 |
|
0.5 |
0.1915 |
0.1950 |
0.1985 |
0.2019 |
0.2054 |
0.2088 |
0.2123 |
0.2157 |
0.2190 |
0.2224 |
|
0.6 |
0.2257 |
0.2291 |
0.2324 |
0.2357 |
0.2389 |
0.2422 |
0.2454 |
0.2486 |
0.2517 |
0.2549 |
|
0.7 |
0.2580 |
0.2611 |
0.2642 |
0.2673 |
0.2704 |
0.2734 |
0.2764 |
0.2794 |
0.2823 |
0.2852 |
|
0.8 |
0.2881 |
0.2910 |
0.2939 |
0.2967 |
0.2995 |
0.3023 |
0.3051 |
0.3078 |
0.3106 |
0.3133 |
|
0.9 |
0.3159 |
0.3186 |
0.3212 |
0.3238 |
0.3264 |
0.3289 |
0.3315 |
0.3340 |
0.3365 |
0.3389 |
|
1.0 |
0.3413 |
0.3438 |
0.3461 |
0.3485 |
0.3508 |
0.3531 |
0.3554 |
0.3577 |
0.3599 |
0.3621 |
|
1.1 |
0.3643 |
0.3665 |
0.3686 |
0.3708 |
0.3729 |
0.3749 |
0.3770 |
0.3790 |
0.3810 |
0.3830 |
|
1.2 |
0.3849 |
0.3869 |
0.3888 |
0.3907 |
0.3925 |
0.3944 |
0.3962 |
0.3980 |
0.3997 |
0.4015 |
|
1.3 |
0.4032 |
0.4049 |
0.4066 |
0.4082 |
0.4099 |
0.4115 |
0.4131 |
0.4147 |
0.4162 |
0.4177 |
|
1.4 |
0.4192 |
0.4207 |
0.4222 |
0.4236 |
0.4251 |
0.4265 |
0.4279 |
0.4292 |
0.4306 |
0.4319 |
|
1.5 |
0.4332 |
0.4345 |
0.4357 |
0.4370 |
0.4382 |
0.4394 |
0.4406 |
0.4418 |
0.4429 |
0.4441 |
|
1.6 |
0.4452 |
0.4463 |
0.4474 |
0.4484 |
0.4495 |
0.4505 |
0.4515 |
0.4525 |
0.4535 |
0.4545 |
|
1.7 |
0.4554 |
0.4564 |
0.4573 |
0.4582 |
0.4591 |
0.4599 |
0.4608 |
0.4616 |
0.4625 |
0.4633 |
|
1.8 |
0.4641 |
0.4649 |
0.4656 |
0.4664 |
0.4671 |
0.4678 |
0.4686 |
0.4693 |
0.4699 |
0.4706 |
|
1.9 |
0.4713 |
0.4719 |
0.4726 |
0.4732 |
0.4738 |
0.4744 |
0.4750 |
0.4756 |
0.4761 |
0.4767 |
|
2.0 |
0.4772 |
0.4778 |
0.4783 |
0.4788 |
0.4793 |
0.4798 |
0.4803 |
0.4808 |
0.4812 |
0.4817 |
|
2.1 |
0.4821 |
0.4826 |
0.4830 |
0.4834 |
0.4838 |
0.4842 |
0.4846 |
0.4850 |
0.4854 |
0.4857 |
|
2.2 |
0.4861 |
0.4864 |
0.4868 |
0.4871 |
0.4875 |
0.4878 |
0.4881 |
0.4884 |
0.4887 |
0.4890 |
|
2.3 |
0.4893 |
0.4896 |
0.4898 |
0.4901 |
0.4904 |
0.4906 |
0.4909 |
0.4911 |
0.4913 |
0.4916 |
|
2.4 |
0.4918 |
0.4920 |
0.4922 |
0.4925 |
0.4927 |
0.4929 |
0.4931 |
0.4932 |
0.4934 |
0.4936 |
|
2.5 |
0.4938 |
0.4940 |
0.4941 |
0.4943 |
0.4945 |
0.4946 |
0.4948 |
0.4949 |
0.4951 |
0.4952 |
|
2.6 |
0.4953 |
0.4955 |
0.4956 |
0.4957 |
0.4959 |
0.4960 |
0.4961 |
0.4962 |
0.4963 |
0.4964 |
|
2.7 |
0.4965 |
0.4966 |
0.4967 |
0.4968 |
0.4969 |
0.4970 |
0.4971 |
0.4972 |
0.4973 |
0.4974 |
|
2.8 |
0.4974 |
0.4975 |
0.4976 |
0.4977 |
0.4977 |
0.4978 |
0.4979 |
0.4979 |
0.4980 |
0.4981 |
|
2.9 |
0.4981 |
0.4982 |
0.4982 |
0.4983 |
0.4984 |
0.4984 |
0.4985 |
0.4985 |
0.4986 |
0.4986 |
|
3.0 |
0.4987 |
0.4987 |
0.4987 |
0.4988 |
0.4988 |
0.4989 |
0.4989 |
0.4989 |
0.4990 |
0.4990 |
Source: StatSoft. (2011). Electronic Statistics Textbook. Tulsa, OK: StatSoft. Retrieved from http://www.statsoft.com/textbook/distribution-tables/#z
The z value calculation for a SoWHAT score of 11 rounded to 4 decimal values for the sake of theillustration. The table rounds z values to just two decimals, so from this point forward, round z valuesto two decimals when using the table. Rounding makes the z value for a raw score of 11 = −0.52.
To interpret the z score, read the whole numbers and the tenths (the tenths are the first value to theright of the decimal) vertically down the left margin of the table. For the hundredths (the second valueto the right of the decimal), move from left to right across the columns at the top of the table.
1. Read down the left margin to the line indicating 0.5.
2. Read across the top to the column indicating 0.02.
3. The table value where row and column intersect is 0.1985. This value is the proportion (out ofa total of 1.0) of any normally distributed population that will occur between z = 0.52 and thepopulation’s mean.
4. To determine the percentage of the distribution between z = −0.52 and the mean, multiply thetable value by 100: 100 × 0.1985 = 19.85% of the distribution is between −0.52 and thepopulation mean.
Note that all the z values in Table 3.2 are positive. Our z score from the SoWHAT score was actuallynegative (z = −0.52). Since the mean of the standard normal distribution is z = 0, the z value for anyscore below the mean will be negative. However, the negative values pose no problem because allnormal distributions are symmetrical, so the proportion of a normal population between z = −0.52 andthe mean will be the same as that between z = 0.52 and the mean. We simply look up the proportionfor the appropriate value of z, remembering that when z is negative, it is a proportion to the left of themean rather than to the right.
Try It!: #2
Table 3.2 has table values only for positive z scores. How do we interpret the valuewhen z turns out to be negative?
To state all this as a principle, because normaldistributions are symmetrical, z scores with thesame absolute value (the same numbers withoutregard to the sign) include the same proportionsbetween their values and the mean of thedistribution. For this reason, the z table indicatesonly the proportions for half the distribution. Inthe case of Table 3.2, that half is the positive(right) half of the distribution.
Because 50% of the distribution occurs eitherside of the mean, if 19.85% of the distribution is from a z = −0.52 back to the mean, the balance of theleft (negative) half of the distribution must occur below a z score of −0.52. That proportion is 50 −19.85 = 30.15%, as the number line illustrates:
Working in the other direction: if the question is what percentage of the population will score 11 orlower on the SoWHAT, the answer is 50 − 19.85 = 30.15%.
If instead someone asks what the probability of scoring at or below 11 (30.15%) is, we must turn thepercentage back into a probability: 30.15 / 100 = 0.3015, or p = 0.3015 of scoring at or below 11.
Try It!: #3
What is the largest possible value for z?
Note that the language above is “11 or lower,”and “at or below.” The characteristics of thenormal curve allow us to determine thepercentage between points, but not at a discretepoint. Technically, a particular point has nowidth and so no associated percentage.
Converting z Scores to Percentage
Now that we have learned how to transform scores from other distributions to fit the z distribution, wewill take a further look at how we can convert scores on opposite sides of the mean and scores with thesame sign to percentages.
Two Scores on Opposite Sides of the Mean
If 5 and 25 are the most extreme apathy scores gathered in the sample of SoWHAT scores, we mightask what percentage of the entire distribution will score between 5 and 25. Because those were thelowest and highest scores, the answer should be 100%, correct? Remember that the collected data werea sample:
5, 6, 9, 11, 15, 15, 17, 20, 22, 25
Although everyone in the sample scored between 5 and 25, it is entirely possible, even probable, thatsomeone in the larger population will have a more extreme score. Using the z distribution, we candetermine how probable by following these steps:
1. Convert both 5 and 25 into z scores.
2. Determine the table values for both z scores.
3. Turn the table values into percentages.
4. Add the percentages together.
The z score formula is
z=x−Ms
Allowing that the subscript to each z indicates the raw score and that M = 14.5 and s = 6.737 from thesample data produces the following calculations:
z5=5−14.56.737=−1.410,
for which the table value is 0.4207,
which corresponds to a percentage of 42.07% (0.4207 × 100).
z25=25−14.56.737=1.559=1.56,
which has a table value of 0.4406.
Expressed as a percentage, the value is 44.06% (0.4406 × 100).
Adding the two percentages together to determine the total percentage between them produces thefollowing:
42.07 + 44.06 = 86.13% from 5 to 25.
Clearly, these scores do not equal 100%. The results indicate that in the population for which thesedata are a sample, about 13.87% (100 − 86.13) will score either lower than 5 or higher than 25. Figure3.4 indicates this result.
Figure 3.4: Areas under the normal curve below z = −1.41 andbeyond z = 1.56
In this distribution, z values that fall below −1.41 or above +1.56 (raw scores below 5 orabove 25) are considered extreme scores, comprising only about 13.87% of thepopulation.
The answer to this problem underscores two important concepts. First, remember that we are dealingwith sample data, and the sample will never exactly duplicate a population. The second, more subtlepoint reveals that there is no point at which we can be confident that no one will produce a moreextreme score. The curve represents this fact by extending the tails (the endpoints of the curve)outward in either direction along the horizontal axis. Although the gap between tail and axis narrowsconstantly, the tails never touch the axis (the 50-cent word is that the tails are “asymptotic” to thehorizontal axis). The application means a value of z will never account for 100% of the distribution.
z Scores with the Same Sign
The previous example raised the question about the percentage of the distribution between z scores onopposite sides of the mean—two z scores where one was positive (z = 1.56) and the other negative (z =−1.41). Perhaps the researcher has a question about the percentage of the distribution betweenSoWHAT scores of 15 and 20. When M = 14.5, both of these raw scores are higher than the mean andboth will result in positive z values. When two z scores have the same sign, determining the percentageof the distribution between them requires that we complete the following steps:
1. Calculate z scores for the raw scores.
2. Determine the table values for each z.
3. Subtract the smaller proportion from the larger.
4. Convert the result into a percentage by multiplying by 100.
z=x−Ms
z15=15−14.56.737=0.0742,
or 0.07, for which the table value is 0.0279.
The 0.0279 is the proportion of the distribution from z = 0.07 and the mean of the distribution. For araw score of 20,
z20=20−14.56.737=0.8164,
or 0.82, which corresponds to p = 0.2939.
This is the proportion of the distribution between z = 0.82 and the mean of the distribution.
When the z scores are on opposite sides of the mean, as they were in our first example, determining theproportion of the distribution between them was a simple matter of adding the two table values. Whenboth z scores are on the same side of the distribution, however, their table values overlap. To determinethe proportion between two values of z with the same sign, take the proportion between the larger(absolute) value and the mean minus the proportion from the smaller (absolute) value to the mean:0.2939 − 0.0279 = 0.2660. Multiplying that by 100 produces the percentage: 100 × 0.2660 = 26.6% ofthe distribution will score between 15 and 20.
Figure 3.5 illustrates this result.
Figure 3.5: Areas under the curve between z = 0.07 and z = 0.82
The percentage of scores between two z values with the same sign is determined bycalculating the difference between the smaller z score table value and the larger one,then multiplying the result by 100.
When trying to answer a question about the percentage of the distribution in a particular area, drawinga simple diagram like Figure 3.5 helps make the question less abstract.
Apply It! Attention to Detail
gerenme/iStock/Thinkstock
A psychological services company administers a test thatmeasures the respondent’s attention to detail. The company’sclients are employers in a variety of organizations that requirepeople with good analytical skills. Respondents who score inthe lowest ranges of the scale are indifferent to potentiallyimportant details. Those who score in the highest ranges tendto fixate on details that may be unimportant to an outcome.Individuals who meet the qualification on this particular testscore in the range from 3.80 to 4.30. Data for those who havetaken the test in the past indicate that M = 4.00 and s = 0.120.For researchers, the initial question is, “Of those who take thetest, what proportion are rejected because either they areinattentive to important details or they become focused on thewrong details?” In terms of the z distribution, the equivalentquestions are the following:
a. What proportion of those who took the test in the past failed to meet the minimumqualification for attention to relevant detail? In other words, what proportion scoredlower than 3.80?
b. What proportion of test-takers scored higher than 4.30?
Regarding question (a), to determine the value of z, the following apply:
x = 3.80
M = 4.00
s = 0.12
Since
z=x−Ms=(3.80−4.00)/0.120
z3.80 = −1.67
The z score table (Table 3.2) indicates that a proportion of 0.4525 of the entire population willfall between this z score and the mean of the distribution. However, the researchers’ interest isin the proportion below this point. Therefore,
0.5 − 0.4525 = 0.0475
In other words, a proportion of 0.0475 occurs below x = 3.80. Stated as a percentage, 4.75% ofthe candidates will score below 3.80 on the test.
For the proportion above 4.30,
z4.30 = (4.30 − 4.00) / 0.12 = 2.5
Table 3.2 indicates that
· this z score corresponds to a proportion of 0.4938, indicating that, as a percentage,49.38% of the population occurs between a score of 4.30 and the mean of thedistribution, and
· the percentage above this point will be 50 − 49.38 = 0.62, or 0.62%, of those who takethe test score at 4.30 or beyond.
Apply It! boxes written by Shawn Murphy
Comparing Data from Different Tests
Earlier chapters discussed how test scores from two different instruments with different means andstandard deviations can be compared. Perhaps a juvenile gang member under court-ordered counselingis required to complete two different assessments: one measuring aggression and one social alienation.The gang member scores 39 on the aggression test and 15 on the alienation test. Table 3.3 shows themeans and standard deviations of the two tests.
Table 3.3: Test results for aggression and social alienation
|
Test |
Mean |
Standard deviation |
|
Aggression measure |
32.554 |
5.824 |
|
Social alienation |
12.917 |
2.674 |
In both cases, the gang member scored higher than average on both aggression and social alienation.For which measure is the score the most extreme?
Because the two tests have different means and standard deviations, comparing the raw scores directlyis not helpful. However, employing the z transformation allows both scores to fit a distribution wherethe mean is 0 and the standard deviation is 1.0. The raw scores may not reveal much, the z scores canbe directly compared. Recall that
z=x−Ms
Calculating z for the aggression score produces:
z39=39−32.5545.824=1.107
Then calculate the z for social alienation:
z15=15−12.9172.674=0.779
Doug Menuez/Photodisc/Thinkstock
Using z scores enables researchers tobetter understand test results measuringaggression and social alienation injuvenile gang members.
Interpreting Multiple z Values
Since the question is which of the juvenile’s two testscores is the more extreme, we have no need for tablevalues—only the value of z. As both z values arepositive, the z for aggression is more extreme than thatfor social alienation. Performing the z transformationallows us to note that the aggression value is 1.107standard deviations from the mean of the distribution.Alienation, meanwhile, is just 0.779 standard deviationsfrom its mean. Practically, the results show thisindividual is more aggressive than alienated. As long asraw scores, means, and standard deviations are available,researchers can use z to make direct comparison of verydifferent qualities, in this case, aggression and socialalienation in the same individual.
Another Comparison
Psychologist Lewis Terman developed the Stanford-Binet test, which measures children’s intelligence.Suppose a psychologist is similarly interested in giftedness among children. Because unusual verbalability often seems to accompany superior intelligence in gifted children, the psychologist measuresboth characteristics for a group of subjects. One particular subject scores 140 on intelligence and 55.0on verbal ability. Table 3.4 lists the descriptive data for each test.
Table 3.4: Test results for intelligence and verbal ability
|
Test |
Mean |
Standard deviation |
|
Intelligence |
100 |
15 |
|
Verbal ability measure |
40 |
5.451 |
As in the previous example, the researcher must convert scores into z scores before they can bedirectly compared.
For the intelligence score, the z score is calculated as:
z=x−Ms
z140=140−10015=2.667
For the verbal ability measure, the z score is calculated as:
z=x−Ms
z55=55−405.451=2.752
The z scores indicate that both test scores are about the same distance from their respective means.This makes it more difficult to glance at the raw scores and know which is higher. But because bothhave been transformed into z scores, the two measures now belong to a common distribution, and theresearcher can see that the verbal ability measure is slightly higher than the intelligence score.
Determining How Much of the Distribution Occurs Under Particular Areas of theCurve
If we draw a distribution and clarify what is at issue, questions about how much of the distribution isabove a point, below a point, or between two points do not require researchers to observe formal rules.For the sake of order and clarity, however, the flowchart in Figure 3.6 provides some direction foranswering different questions a researcher might ask about proportions within a distribution.
Figure 3.6: Flowchart to address questions pertaining to adistribution
Use the steps illustrated in the flowchart to resolve questions about the proportionswithin a population.
Try It!: #4
Figuratively speaking, how does the ztransformation allow you to compareapples to oranges?
The list of steps must seem like a great deal toremember. In fact, the better course whenconfronted with a z score problem is to sketchout a distribution to produce something likeFigures 3.3 and 3.4. The visual displays helpclarify the question and suggest the steps neededto answer it.
The Normal Curve
00:00
00:00
3.3 z Scores, Percentile Ranks, and Other Standard Scores
Our task to this point has been to transform raw scores into z scores and then to percentagesor proportions of the distribution in specified areas. If the percentages are already available,but neither the raw data nor the related descriptive statistics are, Table 3.2 (the z table) allowsus to work backward to determine the z value—even without the mean and standarddeviation for the data.
Let us assume that published data indicate that only 1% of the population has intelligencescores above 140. What z score does this represent?
1. Because Table 3.2 lists proportions, the first step is to turn the percentage into aproportion: 1% is 1/100, which is the same as a proportion of 0.01.
2. Recall that Table 3.2 indicates the proportion of a normal population between a particular value of z and the mean for half (0.5) of the distribution. Therefore, we needa z value which includes all but that most extreme 0.01, which will be the z value for aproportion of 0.50 − 0.01 = 0.49. A z value for a proportion of 0.49, will be the valuethat includes the 49% of the distribution, which means it excludes the highest 1% ofthe distribution.
3. Table 3.2 does not list a proportion of exactly 0.49, but it does list 0.4901, which isvery close. Reading leftward from the proportion to the margin and also vertically tothe column heading, the associated z value for 0.4901 is 2.33. If data were gathered forintelligence scores, a z = 2.33 excludes close to the top 0.01 or 1%.
To state this more directly, when viewed as z scores, any intelligence score where z > 2.33 issomewhere among the top 1% of all intelligence scores. Figure 3.7 illustrates this proportion.
Figure 3.7: The value of z associated with a particularproportion
A normal distribution curve that shows where the highest 1% of scores fallwithin a given population.
Converting z Scores to Percentile Ranks
Chapter 2 introduced percentile scores. Recall that percentiles indicate the point belowwhich a specified percentage of the group occurs. For example, 73% of the distributionoccurs at or below the point defined by the 73rd percentile, and so on. Because researcherscan use the table values associated with z scores to determine the percentage of thedistribution occurring below a point, it is not difficult to take one more step and turn thatpercentage into a percentile score. For example, because
· z = 1.0 includes 34.13% between that point and the mean and
· that part of the distribution from the mean downward is 50%, then
· 34.13% + 50% = 84.13% of scores are at or below z = 1.0; therefore, z = 1.0 occurs atthe 84th percentile.
Although percentile scores can be easily determined from the table values that are associatedwith z scores, note an important difference between percentile scores and z scores. The zscore is one of several standard scores. Standard scores are all equal-interval scores—theinterval between consecutive integers is constant, which means that in terms of data scale,standard scores are interval scale. The increase in whatever is measured from z = −1.5 to z =−1.0 is the same as it is from z = 0.3 to z = 0.8. The increase is 0.5 in either case.
This interval scale does not apply to percentile scores. Because these scores indicate thepercentage of scores below a point rather than reflecting a direct measure of somecharacteristic, the distances between consecutive scores differ widely in various parts of thedistribution. Most of the data in any normal distribution are in the middle portion, wherescores have the greatest frequency. The frequency with which scores occur diminishes asscores become more distant from the mean, something reflected in the curves in frequencydistributions that are vertically highest in the middle and then decline as they extend outwardto the two tails. Note the comparison between percentiles and z scores in Figure 3.8.
Figure 3.8: z scores and percentile scores
A comparison of z scores and percentiles for a normal distribution shows thatthe majority of scores are found within the 50th percentile. Meanwhile, thefrequency of scores above the 99th and below the 1st percentiles is low in anormal distribution.
As a result of high frequency in the middle of the distribution, in any normal distribution thedifference between consecutive percentile scores is always much smaller near the middle ofthe distribution (between the 50th and 51st percentiles, for example) than betweenconsecutive percentile scores in the tails (between the 10th and 11th, or the 90th and 91stpercentiles, for example). This characteristic has important implications. The differencebetween scoring at the 50th and 51st percentile score on something like the Beck DepressionInventory is almost inconsequential compared to the difference between the 90th and 91stpercentile, a much greater difference. Percentile scores are ordinal scale, whereas z scores areinterval scale.
Converting z to Other Standard Scores
Part of the appeal of the z score is that it enables researchers to readily determine relativeperformance. A positive z value indicates that the individual has scored in the upper half ofthe distribution. Someone who scores one standard deviation beyond the mean, as we notedearlier, has scored at the 84th percentile, and so on. The z scores belong to a family ofmeasures called “standard scores.” They have in common these characteristics: a) a fixedmean and standard deviation and b) equal intervals between consecutive data points.
Another standard score is the t score. It is used in the place of z scores when those reportingthem prefer not to report negative scores, which of course are half of all possible z values.After calculating the z value, a researcher can easily change it to a t score. In fact this is truefor any score that has a fixed mean and standard deviation, whether it is a standard score like t or, for example, a Graduate Record Exam (GRE) score (see Table 3.5), which also has afixed mean and standard deviation.
Table 3.5: Comparison of t scores and GRE scores
|
|
Mean |
Standard deviation |
|
t scores |
50 |
10 |
|
Graduate Record Exam |
500 |
100 |
Try It!: #5
What makes a score a standard score?
Either score can be derived from z. Toconvert from z to t, for example, simplymultiply z by 10 and add 50. So, if z = 1.75,then
t = 10 × 1.75 = 17.5 + 50 = 67.5
z = 1.75 is the same as t = 67.5.
For GRE, we would multiply z by 100 and add 500:
GRE = 100 × 1.75 = 175.0 + 500 = 675
z = 1.75 is the same as GRE= 675 (and as t = 67.5).
Although more common in educational than in psychological testing and research, normalcurve equivalent scores (NCE) and “stanine” scores (standard nine-point scale) are alsoexamples of standard scores. Like z and t, each is equal-interval, and both have fixed meansand standard deviations.
3.4 Using Excel to Perform the z Score Transformation
monkeybusinessimages/iStock/Thinkstock
Evaluating achievement motivationscores can give researchers valuableinformation about the relationshipbetween poverty and achievement inschools.
The z score transformation is a fairly simple formula. Asa result, to program it into Excel and transform an entiredata set into z scores is not difficult. In fact, theapplication offers several ways to do this, but thischapter will explore just one. It involves programmingthe z score transformation formula directly into the datasheet.
A researcher interested in the relationship betweenpoverty and achievement motivation among secondary-school-aged young people gathers data from a group ofstudents whose families qualify for free and reduced-price lunches at school. The achievement motivationscores are as follows:
4, 5, 7, 7, 8, 9, 9, 9, 10, 13
To use Excel to transform those data into their z score equivalents, follow these steps:
1. List the data in Excel in Column B, with the label “Ach Mot” in B1.
2. Enter the 10 scores into cells B2 to B11.
3. In cell B12, enter the formula =average(B2:B11). (Note: Virtually all spreadsheets, includingExcel, have shortcuts for the more common calculations, such as means and standard deviations.A user can enter the formula, as we have done here, or use a shortcut. Shortcut procedures vary,however, depending on the operating system and the version of Excel. Excel for Mac, forexample, allows users to enter the data, position the cursor where they desire the statistic’s valueto appear, and then double-click the name of the desired statistic under the Formula tab.)
c. The equal sign indicates to Excel that a formula follows.
c. The command average will provide the arithmetic mean.
c. When several cells are to be included in the function, they are placed in parentheses ( ).When the cells are consecutive, the colon (:) indicates that all cells from B2 to B11 are tobe included in the function.
1. Press Enter.
1. In cell A12, enter the label “mean =.” The value in cell B12 will be 8.1, the mean of theachievement motivation scores.
1. In cell B13, enter the formula =stdev(B2:B11). Note that stdev is the Excel abbreviation for“sample standard deviation.” For Mac users, the abbreviation is stdev.s.
1. Press Enter.
1. In cell A13, enter the label std dev =. The value in cell B13 will be 2.558211, the standarddeviation of the scores.
1. In cell C1, enter the label equiv z.
1. In cell C2, enter the formula =(B2_8.1)/2.558 and press Enter. Consistent with the z scoretransformation, this formula subtracts the mean from the raw score in cell B2 and then dividesthe result by the standard deviation, 2.558, which we rounded to three decimals.
1. Repeat that operation for all the other scores as shown next:
k. With the cursor in cell C2, click and drag the cursor down from C2 to C11 so that cells C2to C11 are highlighted.
k. In the Editing section at the top of the page near the right side is a Fill command with adown-arrow at the left (for Macs, the command is on the left side, below the Home tab).Click the down-arrow to the right of the Fill command, then click Down. This action willrepeat the result in C2 for the other nine cells, adjusting for the different test scores in eachcell.
Figure 3.9 shows how the spreadsheet will look after Step 11, with the z score equivalents of all theoriginal achievement motivation scores displayed to the right of the original scores.
Figure 3.9: Raw scores transformed to z scores in Excel
Excel converts raw scores to z scores using a simple formula.
Source: Microsoft Excel. Used with permission from Microsoft.
Using Excel to Perform the z Score Transformation
00:00
00:00
3.5 Using z Scores to Determine Other Measures
Occasionally, a researcher has access to z scores and the mean and standard deviation but notthe original raw scores. Formula 3.1 used the raw score (x), the mean (M), and the standarddeviation (s) to determine the value of z, but actually, any three of the values in the formulacan be used to determine the value of the fourth. Just as Formula 3.1 uses x, M, and s todetermine z, we could use z, M, and s to derive x. Altering Formula 3.1 to determine thevalue of something other than z involves a little algebra but is not difficult.
Determining the Raw Score
To determine the raw score, follow these steps:
1. Because z = (x − M)/s, swap the terms before and after the equal sign so that (x − M)/s= z.
2. To eliminate the s in the denominator of the first term, multiply both sides by s so thatit disappears from the first term and emerges in the second: x − M = sz.
3. To isolate x, add M to both sides of the equation so that x = sz + M.
Returning to the Excel problem, if the z scores and descriptive statistics are available, we candetermine the raw score for which z = −1.603 as follows:
If M = 8.10, s = 2.558, and x = z × s + M, substituting the values we have produces x = (−1.603)(2.558) + 8.10 = 3.9995, which rounds to 4.0.
Checking the earlier data reveals that 4 was indeed the raw score for which z = −1.603.
Determining the Standard Deviation
If the raw scores, the mean, and z are available, but s is lacking, z = (x − M)/s, so (x − M)/s = z. Taking the reciprocal of each half of the equation—which means inverting the term so that(x − M)/s becomes s/(x − M) and z/1 becomes 1/z, giving us s/(x − M) = 1/z. Multiplying bothsides by (x − M) yields s = (x − M).
Using the data from the Excel problem again, for the first participant, x = 4, M = 8.10, and z =−1.603. According to the adjusted formula,
s = (x − M)/z
therefore, substituting the values we have produces the following:
s=4−8.1−1.603=2.5577
which rounds to 2.558, the standard deviation value for the original data set.
Determining the Mean
It would be very unusual for a researcher to have the z scores, the standard deviation, and theoriginal achievement motivation score but not have the mean. However, just to complete theset, the mean can be determined from the other three values as follows:
Because
z = (x − M)/s
if both halves of the equation are multiplied by s, then s appears in the first termand disappears from the second. The result is
sz = x − M
If M is then added to both sides, M appears in the first term and is eliminated in thesecond. The result is
sz + M = x.
If sz is then subtracted from both sides, it is eliminated from the first term andadded to the second. The result is
M = x − sz
For the first participant, the achievement motivation score x = 4.0, the z score =−1.603, and s = 2.558. The mean for the test can be determined as follows:
M = x − sz, or
M = 4 − (2.558 × −1.603) = 8.10, which was the original mean.
Maintaining Fixed Means and Standard Deviation
One of the characteristics of widely used standardized tests is that their mean and standarddeviation values remain the same over time. The major intelligence tests, for example, have afixed mean of 100 and a standard deviation of 15, even though the Stanford-Binet andWechsler tests have been revised several times. When a test is revised and updated, do themeans and standard deviations likewise change? In fact, they do. Flynn and Weiss (2007)documented significant increases in intelligence scores over a 70-year period, but to make thescores comparable over time, psychologists use what are called modified standard scores.The modified standard score allows those working with the test to gather data that have anymean and any standard deviation and then adjust them so that they conform to predeterminedvalues. This process follows these steps:
1. Gather data with the new instrument.
2. Determine the equivalent z scores for test-takers’ raw scores.
3. Apply the formula.
Formula 3.2 is used to modify a score so that the mean and standard deviation for thepopulation of scores take on specified values.
Formula 3.2
MSS = (sspec)(z)+ Mspec
where
MSS = the modified standard score,
sspec = the specified standard deviation, and
Mspec = the specified mean.
Note that this formula is the same used to transform z scores into t scores. By way of anexample, perhaps a psychologist has developed what she has labeled the Brief IntelligenceTest (BIT). To compare results to tests her colleagues have used traditionally, she wants theBIT’s descriptive characteristics to conform to those of the more established tests. For eightparticipants, the BIT scores are as follows:
22, 25, 26, 29, 29, 32, 32, 35
Of course, no one norms an intelligence test on only eight people. The potential for what wewill later call sampling error is too great. Still, to illustrate the process, we will assume thesample scores are valid.
Verify that M = 28.75 and s = 4.268.
For the participant with an intelligence score of 22, the corresponding z value is
z=x−Ms=22−28.754.268=−1.582
To determine that participant’s score on an instrument with a mean of 100 and a standarddeviation of 15, the psychologist will apply the formula:
MSS = (sspec)(z) + Mspec
= (15)(=1.582) + 100
= 76.276
Although the original BIT score was 22, using the z transformation and modified standardscore procedures makes the BIT score conform to the mean and standard deviation of a moreestablished test. Among scores for which the mean is 100 and the standard deviation is 15,the modified standard score for the BIT score of 22 is 76.276.
Writing Up Statistics
Although z scores are an important part of data analysis, like the raw scores that researchersgather in their work, z scores often do not appear in research reports. Reports list the meansand standard deviations, but often omit the raw scores and their z score transformation. In astudy of the weight-gain side effect that anti-psychotic drugs might have on adolescents,Overbeek (2012), however, used a combination of height and weight to determine body-mass-index (BMI) scores for each subject, and then transformed the BMI scores into easier-to-interpret z scores. A z score near 0 indicated that given the subject’s height, weight wasprobably appropriate. A positive z score indicated that the individual might be overweight,negative z scores indicated underweight, and so on. Overbeek also used z scores to index theweight-gain data over the course of the study.
Brown (2012), too, used z scores. In his study, they offered a way to counter the effect thatgrade inflation has on college students’ class rankings. He posited that class rankings are lessinformative than they once were because weaker students in departments where courseworkis easier are ranked ahead of students who have a higher level of academic aptitude butcompete in more demanding programs. Brown’s solution was to use the z transformationwithin departments to indicate how much above or below the mean students were in theirindividual programs.
Summary and Resources
Chapter Summary
Normal distributions are unimodal and symmetrical, and their standard deviations tend to beabout one-sixth of the range. Although not all data are normally distributed, many of themental characteristics that psychologists and social scientists measure are normal. Becausethe proportions of the population that occur in specified ranges remains constant in normallydistributed populations, we can have some confidence about how scores will be arrayed evenbefore we view a display of the data.
What the standard normal distribution, or z distribution, does is capitalize on the consistencyin normally distributed populations by offering one distribution by which all other normalpopulations can be referenced. In this distribution, where the mean is always 0 and thestandard deviation is 1.0 (Objective 1), table values indicate the proportions of the populationlikely to occur anywhere along its range. By transforming raw scores (Objective 2) from anynormal population so that they fit this z distribution, we can take advantage of how well thecharacteristics of this distribution are known and answer important questions about data fromany population (Objective 3) in terms of z:
· For example, when someone scores at a particular level, we can ask what proportion ofthe entire population is likely to score below (or above) that point.
· When most of the people in a particular group score between two points, we can askwhat proportion of the entire population will score between (or outside) those points.
Because the z score transformation is a relatively simple formula, programming Excel toproduce the z equivalents for any set of scores (Objective 4) is simple and can be helpful withlarge data sets.
The z is one of several standard scores in fairly common use. Whether z or some other, allstandard scores indicate how distant one individual’s score is from the mean of thedistribution. Rather than providing an absolute measure of some characteristic, standardscores are normative, meaning that they indicate the level of what is measured relative toothers in the same population. Those who prefer not to deal in negative values (whichcharacterize half of the z distribution) can employ t scores. In all material respects, t is thesame as z, except that the mean is 50 and the standard deviation is 10.
The modified standard score (Objective 6) enhances standard scores’ ability to communicatean individual’s standing relative to a population. Researchers often use standard scores toreport the data from standardized tests, but these tests are revised from time to time, whichcan affect the test means and standard deviations. To ensure stability, the modified standardscore uses the z transformation as a way to maintain constant descriptive characteristics, evenas the instrument used to measure it, or even the characteristic measured, changes with time.
In the incremental nature of statistics books, each chapter prefaces the next. Chapters 1–3 area prelude to Chapter 4. With all our effort to label, display, and describe data sets, the focusin the discussion of z scores and the other topics has been primarily about analyzing theperformance of individuals. Behavioral scientists, however, are generally much moreinterested in asking questions about groups. Analyzing how those in a sample compare tothose in the entire population is the focus of Chapter 4. It will do so by expanding discussionof the z distribution.
The math and the logic involved in Chapter 4 will be much the same. If the discussion in thischapter makes sense, the material in Chapter 4 will not be difficult. Still, it is a good idea toreview the Chapter 3 material and recalculate the sample problems, as repetition has value.
Chapter 3 Flashcards
Key Terms
modified standard scores
percentile
probability
standard normal distribution
standard scores
t score
z score
z transformation
Review Questions
Answers to the odd-numbered questions are provided in Appendix A.
1. A researcher is interested in people’s resistance to change. For the dogmatism scale(DS), data for 10 participants are as follows:
28, 28, 29, 29, 32, 33, 35, 36, 39, 42
a. What is the z score for someone who has a DS score of 28?
a. Will a z score for a raw score of 35 be positive or negative? How do you know?
a. How many standard deviations is a score of 28 from the mean?
a. What will be the z value of a raw score of 33.1?
a. Since there are no scores below 28, does it make any sense to calculate z for a rawscore of 25, for example? Shouldn’t such a score have a zero probability ofoccurring?
1. Examining the relationship between recreational activity and level of optimism amongsenior citizens, a psychologist develops the Recreation Activity Test (RAT). Scores for8 participants are as follows:
11, 11, 14, 14, 14, 17, 18, 22
b. What is the z score for someone with RAT = 23?
b. Why isn’t 0 the answer to 2a, since none of the participants scored 23?
b. What is the z score for someone with RAT = 15.125?
b. Explain the answer to 2c.
b. How does z allow one to compare tests with entirely different means and standarddeviations?
1. One participant has RAT = 11. The same individual is administered a ConsistentApproval Test (CAT) and scores 45. The CAT data, including that participant’s score,are as follows:
42, 45, 48, 49, 55, 58, 62, 64
c. Which score is higher, the RAT or the CAT?
c. Why isn’t the answer to 3a automatically CAT, since it has the higher mean value?
1. Researchers developed the ANxious, Gnawing Stress Test (ANGST) to measure emotional stability among law-enforcement professionals. A random sample of policepatrol officers yielded the following scores:
54, 58, 61, 64, 75, 81, 82, 85
d. What proportion of the population will score 81 or higher?
d. What proportion will score 60 or higher?
d. If x > 75 is the cutoff for “highly stressed,” what is the probability that someone,selected at random, will be highly stressed?
d. What is the probability of scoring between 60 and 81?
1. Using the data in Question 4, what percentage of the population will score lower than55?
1. What is the t equivalent to a z score for someone with an ANGST score of 64?
f. What is the mean of the t distribution?
f. Why is t sometimes preferred over z?
f. If z = 2.5, what is t?
1. Refer to the data in Questions 3 and 4: For an individual who scores 60 on RAT and 78on ANGST, which is the higher score?
1. In any standard normal distribution, determine the following:
h. What percentage of scores will occur below z = 0?
h. What is the probability of a positive value of z?
h. What percentage of scores will occur between ±1.96 z?
1. If someone scores z = 1.0, what is the corresponding percentile rank?
1. What percentile rank is z = 0? What measure of central tendency represents the 50thpercentile?
1. A psychologist wishes to maintain a mean of 25 and a standard deviation of 5 for a testdeveloped to measure compulsive behavior. On a revised test, the scores are asfollows:
14, 17, 19, 19, 22, 27, 28, 29
k. What is the modified standard score for the person who scored 17 on the revisedinstrument?
k. What is the probability of scoring 17 or lower according to the eight scores?
1. Given the data in Question 11,
l. What is the z equivalent of a raw score of 28?
l. What is the probability of scoring somewhere from 14 to 29?
1. Draw a normal distribution and identify where z = −1.17 and z = +2.53 are located.What percentage of the population occurs between these two z values?
Answers to Try It! Questions
1. A z of 1.5 indicates that the associated raw score is 1.5 standard deviations (thedenominator in z) from the mean.
2. Because the standard normal distribution (the z distribution) is normal, the distributionis symmetrical. The proportion of the distribution between a value of z and the meanwill be the same for a negative z as it is for a positive z with the same numerical value.
3. Do not refer to Table 3.2 for help with this one. The table’s highest score is z = 3.09,but in fact z has no upper limit. In theory, the tails in the z distribution never actuallytouch the horizontal axis of the graph, which means that there exists always at least thepossibility of scores higher (or lower) than any already measured.
4. One of the values of the z transformation is that scores that have any descriptivecharacteristics can be recalibrated so that they fit a distribution where the mean is 0 andthe standard deviation is 1.0. By doing so, scores from any variety of sources can becompared directly after converting them to z scores. The only requirement is that theybe normally distributed.
5. Standard scores are equal-interval scores with a fixed mean and standard deviation,thereby allowing the magnitude of the score to indicate how an individual compares toall others for whom scores are available.
4
Applying z to Groups
Victor Faile/Corbis
Chapter Learning Objectives:
After reading this chapter, you should be able to do the following:
1. Describe the distribution of sample means.
2. Explain the central limit theorem.
3. Analyze the relationship between sample size and confidence in normality.
4. Calculate and explain z test results.
5. Explain statistical significance.
6. Calculate and explain confidence intervals.
7. Explain how decision errors can affect statistical analysis.
8. Calculate the z test using Excel.
Introduction
As we noted at the end of Chapter 3, researchers are generally more interested in groups thanin individuals. Individuals can be highly variable, and what occurs with one is not necessarilya good indicator of what to expect from someone else. What occurs in groups, on the otherhand, can be very helpful in understanding the nature of the entire population. A Googlesearch indicates that the suicide rate is higher among dentists than it is among those of manyother professions. If we wanted to experiment with some therapy designed to relievedepressive symptoms among dentists, we would be more confident observing how a group of50 dentists responds than in examining results from just one. This chapter will use thematerial from the first three chapters to begin analyzing people in groups.
Noting that many of the characteristics that interest behavioral scientists are normallydistributed in a population implies that some characteristics are not. Since samples can neverexactly emulate their populations, it may not be clear in the midst of a particular study whendata are normally distributed. This uncertainty potentially poses a problem: we may wish touse the z transformation and Table B.1 of Appendix B in our analysis, but Table B.1 is basedon the normality assumption. If the data are not normal, where does that leave the relatedanalysis?
This is a self-assessment and will not affect your grade. You may only take this pre-test once.
Test Ch 4: Applying z to Groups
Top of Form
1. The potential for sampling error diminishes as the size of the sample grows.
· a. TRUE
· b. FALSE
2. Populations based on samples are more inclined to normality than populations based on individual scores.
· a. FALSE
· b. TRUE
3. Determining normality requires all of the scores in a population.
· a. TRUE
· b. FALSE
4. The z test produces a z value based on individual scores rather than on sample means.
· a. TRUE
· b. FALSE
5. A statistically significant result is one that is unlikely to have occurred by chance.
· a. FALSE
· b. TRUE
Finish
Bottom of Form
4.1 Distribution of Sample Means
iStockphoto/Thinkstock
A population is allmembers of a definedgroup, such as all voters ina county.
What options do researchers have if they are suspiciousabout data normality? One important answer is the distribution of sample means, so named because the scoresthat constitute the distribution are the means of samplesrather than individual scores.
Note that the descriptor population means all possiblemembers of a defined group. Recall that the frequencydistribution—the bell-shaped curve representing thepopulation—was a figure based on the individual measuressampled one subject at a time. In discussing the frequencydistribution, we assumed that we would measure eachindividual on some trait, and then plot each individual score.Instead of selecting each individual in a population one at atime, suppose a researcher
1. selects a group with a specified size;
2. calculates the sample mean (M) for each group;
3. plots the value of M (rather than the value of eachscore) in a frequency distribution;
4. and continues doing this until the population is exhausted.
How would plotting group scores rather than individual ones affect the distribution? Wouldthe end result still be a population? The answer to the second question is yes: because everymember is included, it is still a population. Whether a population is measured individually oras members of a group is incidental, as long as all are included.
Perhaps researchers are interested in language development among young children and wishto measure mean length of utterance (MLU) in a county population. Whether the researchersmeasure and plot MLU for each child in a county’s Head Start program or plot the meanMLU for every group of 25 in the program, the result is population data for Head Startlearners for that county.
The Central Limit Theorem
The answer to the question “how would the distribution be affected?” is a little moreinvolved, but it is important to nearly everything we do in statistical analysis. It involveswhat is called the central limit theorem:
If a population is sampled an infinite number of times using sample size n and themean (M) of each sample is determined, then the multiple M measures will take onthe characteristics of a normal distribution, whether or not the original populationof individuals is normal.
Take a minute to absorb this. A population of an infinite number of sample means drawnfrom one population will reflect a normal distribution whatever the nature of the originaldistribution. A healthy skepticism prompts at least two questions: 1) How would we provewhether this is true since no one can gather an infinite number of samples? and 2) Why doessampling in groups rather than as individuals affect normality?
Although prove is too strong a word, we can at least provide evidence for the effect of thecentral limit theorem using an example. Perhaps a psychologist is working with 10 people ontheir resistance to change, their level of dogmatism. Technically, because 10 constitutes theentire group, the population is N = 10. Recall that N refers to the number in a population.Even with a small population we cannot have an infinite number of samples, of course, butfor the sake of the illustration we will assume that
· dogmatism scores are available for each of the 10 people;
· the data are interval scale;
· the scores range from 1 to 10; and
· each person receives a different score.
So with N = 10, the scores are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10. Figure 4.1 depicts a frequencydistribution of those 10 scores.
The distribution in Figure 4.1 is not normal. With R = 10 − 1 = 9 and s = 3.028 (a calculationworth checking), the distribution is extremely platykurtic (i.e., flatter than normal); the rangeis less than 3 times the value of the standard deviation rather than the approximately 6 timesassociated with normal distributions. There is either no mode or there are 10 modes, neitherof which suggests normality. We can illustrate the workings of the central limit theorem witha procedure Diekhoff (1992) used. We will use samples of n = 2, and make the examplemanageable by using one sample for each possible combination of scores in samples of n = 2from the population, rather than an infinite number of samples.
Figure 4.1: A frequency distribution for the scores 1through 10: Each score occurring once
A frequency distribution of ten scores, each with a different value. This type ofdistribution, which is not normal, is highly platykurtic.
Table 4.1 lists all the possible combinations of two scores from values 1–10. Ninetycombinations of the 10 dogmatism scores are possible. The larger the sample size, the morereadily it demonstrates the tendency toward normality, but all combinations of (for example)three scores would result in a very large table.
Table 4.1: All possible combinations of the integers 1–10
|
|
|||||||||
|
1, 2 |
2, 1 |
3, 1 |
4, 1 |
5, 1 |
6, 1 |
7, 1 |
8, 1 |
9, 1 |
10, 1 |
|
1, 3 |
2, 3 |
3, 2 |
4, 2 |
5, 2 |
6, 2 |
7, 2 |
8, 2 |
9, 2 |
10, 2 |
|
1, 4 |
2, 4 |
3, 4 |
4, 3 |
5, 3 |
6, 3 |
7, 3 |
8, 3 |
9, 3 |
10, 3 |
|
1, 5 |
2, 5 |
3, 5 |
4, 5 |
5, 4 |
6, 4 |
7, 4 |
8, 4 |
9, 4 |
10, 4 |
|
1, 6 |
2, 6 |
3, 6 |
4, 6 |
5, 6 |
6, 5 |
7, 5 |
8, 5 |
9, 5 |
10, 5 |
|
1, 7 |
2, 7 |
3, 7 |
4, 7 |
5, 7 |
6, 7 |
7, 6 |
8, 6 |
9, 6 |
10, 6 |
|
1, 8 |
2, 8 |
3, 8 |
4, 8 |
5, 8 |
6, 8 |
7, 8 |
8, 7 |
9, 7 |
10, 7 |
|
1, 9 |
2, 9 |
3, 9 |
4, 9 |
5, 9 |
6, 9 |
7, 9 |
8, 9 |
9, 8 |
10, 8 |
|
1, 10 |
2, 10 |
3, 10 |
4, 10 |
5, 10 |
6, 10 |
7, 10 |
8, 10 |
9, 10 |
10, 9 |
For each possible pair of scores, if we calculate a mean and plot the value in a frequencydistribution as a test of the central limit theorem, the result is Figure 4.2. Because the entiredistribution is based on sample means, Figure 4.2 is a distribution of sample means. Strictlyspeaking, this distribution is not normal, but although based on precisely the same data,Figure 4.2’s distribution is a good deal more normal than the distribution in Figure 4.1.
Figure 4.2: A frequency distribution of the means of allpossible pairs of scores 1 through 10
A distribution of the means of each possible pair of scores with values between1 and 10. This distribution is not normal, but has more normality than thedistribution shown in Figure 4.1.
Mean of the Distribution of Sample Means
The symbol used for a population mean to this point, µ, is actually the symbol for apopulation mean formed from one score at a time. To distinguish between the mean of thepopulation of individual scores and the mean of the population of sample means, we’llsubscript µ with an M: µM. This symbol indicates a population mean (µ) based on samplemeans (M).
Try It!: #1
Why is there less variability in thedistribution of sample means than in adistribution of individual scores?
With a distribution of just 90 sample means,Figure 4.2 shows nothing like an infinitenumber, of course, but it is instructivenevertheless. The mean of the scores 1through 10 is 5.5 (µ = 5.5). Study Figure4.2 for a moment. What is the mean of thatdistribution? The mean of our distributionof sample means is also 5.5: µM = 5.5. Thepoint is this: When the same data are usedto create two distributions—one apopulation based on individual scores andthe other a distribution of sample means—the two population means will have the samevalue, or, symbolically stated: µ = µM.
Describing the distribution as “normal” is a stretch, but Figure 4.2 is certainly more normalthan Figure 4.1. For one thing, rather than the perfectly flat distribution that occurs when allthe scores have the same frequency, some scores have greater frequency than others. Thesample means near the middle of the distribution in Figure 4.2 occur more frequently than thesample means at either the extreme right or left.
Why are extreme scores less likely than scores near the middle of the distribution? It isbecause many combinations of scores can produce the mean values in the middle of thedistribution, but comparatively few combinations can produce the values in the tails of thedistribution. With repetitive sampling, the mean scores that can be produced by multiplecombinations increase in frequency and the more extreme scores occur only occasionally,which the next section illustrates.
Variability in the Distribution of Sample Means
In the original distribution of 10 scores (Figure 4.1), what is the probability that someonecould randomly select one score (x) that happens to have a value of 1? Because there are 10scores, and just one score of 1, the probability is p = 1/10 = 0.1, right? By the same token,what is the probability of selecting x = 10? It is the same, p = 0.1.
Moving to the distribution based on 90 scores (Figure 4.2), what is the probability ofselecting a sample of n = 2 that will have M = 1.0? Is there any probability of selecting twoscores out of the 10 that will have M = 1.0? Because there is only one value of 1, there is noway to select two values with M = 1.0. As soon as a score of 1 is averaged with any otherscore in the group, all of which are greater than one, the result is M > 1. That is why thelowest possible mean score in Figure 4.2 is 1.5, which can only occur when 1 and 2 are in thesame sample.
The same thing occurs in the upper end of the distribution. The probability of selecting agroup of n = 2 with M = 10 is also zero (p = 0) because all other scores have lower valuesthan 10. For the 90 possible combinations, the highest possible mean score is 9.5, which canoccur only when the 10 and the 9 happen to be in the same sample.
The point is that variability in group scores is always less than the variability in individualscores. A related point is that the impact of the most extreme scores in a distributiondiminishes when they are included in samples with less extreme scores. Applied, theseprinciples mean that a researcher examining, for example, problem-solving ability among agroup of subjects can afford to be less concerned about the impact of one extremely low orone extremely high score as the size of the group increases. Larger group sizes minimize theeffect of extreme scores.
Standard Error of the Mean
Recall that the sigma, σ, indicates a population’s standard deviation. Specifically, σ indicatesa standard deviation from a population of individual scores. The symbol for the standarddeviation in a distribution of sample means is σM and as the subscript M suggests, it measuresvariability among the sample means. The formal name for σM is the standard error of themean.
In the language of statistics, error as in standard error of the mean refers to unexplainedvariability. As we move through the different procedures, we will calculate other standarderrors, which all have this in common: all are measures of unexplained data variability.
Earlier, we noted that whether charting the distribution of individual scores or the distributionof sample means, the means of the two distributions will always be equal: µ = µM. Is itlogical to expect the same from the measures of variability; in other words, will σ = σM? Thefact that the distribution of individual scores always has more variability than the distributionof sample means answers this question. Symbolically speaking, σ > σM, something that thedogmatism data show.
The standard deviation of the 10 original scores (1, 2, 3, 4, 5, 6, 7, 8, 9, 10) is σ = 2.872. Notethat this instance and the calculation of the standard error of the mean just below deal withpopulations and the formula must, therefore, employ N, rather than n − 1. Elsewhere in thebook, however, the formula will always be n − 1.
The standard error of the mean can be calculated by taking the standard deviation of themean scores of each of those 90 samples which constituted the distribution of sample means.The calculation is a little laborious, and happily, not a pattern that must be followed later, butthe value is σM = 1.915.
So, as predicted, σ has a larger value than σM. The smaller value for σM reflects the way lessextreme scores moderate the more extreme scores when they occur in the same sample.
Sampling Error
Although the standard error of the mean does not per se refer to a mistake, another kind oferror, sampling error, does. In inferential statistics, samples are important for what theyreveal about populations. However, information from the sample is helpful for drawinginferences only when the sample accurately represents the population. The degree to whichthe sample does not represent the population is the degree of sampling error.
Samples reflect the population with the greatest fidelity when two prerequisites are met: 1)the sample must be relatively large, and 2) the sample must be based on random selection.
The safety of large samples is explained by the law of large numbers. According to thismathematical principle, as a proportion of the whole, errors diminish as the number of datapoints increases. The potential for serious sampling error diminishes as the size of the samplegrows. The text earlier referred to this principle in noting that the distorting effect of extremescores diminishes as sample size grows.
Random selection refers to a situation where every member of the population has an equalprobability of being selected. Random selection contrasts with what are called conveniencesamples, samples that are used intact because they are handy. A sociology professor who usesthe students in a particular section of his class is relying on a type of convenience sampleknown as a nonrandom sample.
A random sample of n = 5 could be created from the 10 people being treated for dogmaticbehavior by assigning each person a number, placing the 10 numbers into a paper bag,shaking the bag well, and without looking, drawing out five numbers.
The result would be a randomly selected sample. When randomly selected, samples differfrom populations only by chance. They will differ, of course, but the differences are less andless important as sample size grows.
National Atlas
Systematic sampling error providesresults drastically different from theactual outcome. Such a samplingerror occurred before the 1936presidential election, when a studypredicted a win by Alf Landon. Theelection results, displayed in themap, were drastically different.
If the sample should fail to capture some importantcharacteristic of the population other than its size,the problem is sampling error. The importantcharacteristic might be the mean, for example, andwhen M ≠ µ, sampling error has occurred. In fact,some sampling error always occurs because asample can never exactly duplicate all thedescriptive characteristics of the population, butsampling error will usually be minor if samples arerelatively large and randomly selected.
Statistical analysis procedures tolerate minor,random sampling error, but systematic samplingerror is another matter. Systematic sampling erroroccurs when the same mistake is made time aftertime.
In 1936, the publishers of Literary Digest, aprominent publication of the time, decided topredict the outcome of that year’s presidential election in the United States. To ensure thatsample size would not pose a problem, they sent out millions of postcards to registeredvoters. Literary Digest seemed to at least have met the requirement for a relatively largesample, because the Harris and Gallup polling organizations typically obtain very accurateresults with a few thousand, and sometimes just a few hundred, responses. Unfortunately, thepublishers decided to rely on telephone books and automobile registrations to locate pollrecipients. Consider the historical setting. At the height of the Great Depression, twoindicators of relative prosperity identified voters: a telephone in the home and a currentlyregistered car. The study proved disastrous for the magazine’s reputation. Poll resultsindicated that Alf Landon would win, but of course Franklin Roosevelt was elected in alandslide to a second term, carrying every state in the union except Maine and Vermont.
The problem was systematic sampling error. The voters were consistently and nonrandomlyselected from groups not representative of the entire population. If they had been randomlyselected, chances are that with the large sample size, the study would have predicted the election results accurately, but the sample size alone was not enough to compensate for theerror.
4.2 The z Test
The distribution of sample means is a distribution based not on individual scores but on themeans of samples of the same size repeatedly drawn from a population. The central limittheorem assures that such a distribution will be normal. Consequently, if the z score formulafrom Chapter 3 is adjusted to accommodate groups rather than individual scores, Table B.1answers all the same questions about groups that it did about individuals in Chapter 3.
Recall that the z score formula (3.1) had the following form:
z=x−Ms
If the following substitutions are made:
· M for x, so that the focus is on a sample mean rather than on an individual score;
· µM for M to shift from the sample mean to the mean of the distribution of samplemeans;
· σM for σ so that the measure of variability is for the distribution rather than the sample;
then the result is the z test:
Formula 4.1
z=M−μMσM
The z test produces a z value for groups rather than individual scores. Just as it did forindividual scores, z indicates how distant a particular sample mean is from the mean of thedistribution of sample means.
Note the similarities in Formulas 3.1 and 4.1:
· Both formulas produce values of z.
· Both numerators call for subtractions that result in difference scores.
· Both denominators measure data variability.
Calculating the z Test
When calculating z scores, as shown in Chapter 3, everything that is needed (x, M, and s) canbe determined from the sample:
z=x−Ms
Values needed for the z test, however, are often not as easy to determine. Because µM = µ,one of those two parameters must be provided. The standard error of the mean (σM) can alsopresent a problem. No one wishing to complete a z test is going to have the mean scores forthat infinite number of samples that make up the distribution of sample means. So calculatingthe standard deviation of those means, which is what the standard error of the meanrepresents, is not an option. It is possible, however, to determine the value of the standarderror of the mean without having to calculate a standard deviation for an indeterminatenumber of scores. If a researcher does not have σM (the standard error of the mean is . . .) butdoes know the population standard deviation (σ), σM can be found as follows:
Formula 4.2
σM=σN√
where
σM = the standard error of the mean
σ = the population standard deviation
N = the number in the group
So for a group of 100 with the value of σ as 15, then σM is
σM=σN‾‾√
σM=15100‾‾‾‾√=1510=1.5
This approach affords only a partial solution to the standard error of the mean, however,because it still requires at least σ.
Chapter 5 will explain a way around the problem of determining σM, but in the meantime,consider the following example. A marriage and family counselor has access to somenational data on the frequency of negative verbal comments exchanged between couples introubled marriages. The counselor finds the following:
· Couples in troubled marriages tend to have 11 negative exchanges per week, with astandard deviation of 4.755.
· A study of 45 couples who have filed for divorce in the area where the counselor hasher practice reveals that the mean number of negative comments per week is 12.865.
· Given the national data, the counselor wants to know the probability that a randomlyselected group of couples from that population will have as many negative exchangesas the counselors’ clients, or more.
Although the counselor’s question is about groups rather than individuals, the problem ismuch like a z score problem. The counselor knows the following: µ and (because µM = µ) µM= 11.0, σ = 4.755, N = 45, and M = 12.865.
Jeffrey Hamilton/DigitalVision/Thinkstock
Marriage and family counselors canuse existing data to help them learnmore about their clients.
The standard error of the mean is
σM=σN‾‾√=4.75545‾‾‾√=0.709
And z is
z=M−μMσM=12.865−11.00.709=2.630
Comparing M to µM indicates that the counselor’sgroup has a higher number of negative verbalexchanges per week than the number nationallyamong couples with troubled marriages: 12.865 is ahigher value than 11.0. In the following section, wewill discuss what else can be determined from the analysis.
Interpreting the Value from the z Test
The result from the marriage and family counselor’s group in the previous section is a valueof z just like those that were calculated in Chapter 3, except that the value indicates howmuch a sample mean (M) differs from the mean of a population of samples (µM),instead ofhow an individual (x) differs from either a sample mean (M) or a population mean (µ). TheTable B.1 value indicates that 0.4957 out of 0.5 occurs between a value for z = 2.63 and themean of the distribution. So among the population of couples with troubled marriages,49.57% will have negative verbal exchanges somewhere between the level of this group(12.865 per week) and the mean of the national population, 11.0 per week. But the questionconcerns the probability that a group of clients selected at random would have 12.865negative comments per week, or more. Because 49.57% will have 12.865 or fewer negativeexchanges per week, just 0.43% (50% − 49.57%) will have 12.865 negative comments perweek or more. Stated as a probability, p = 0.0043 that a group of individuals in troubledmarriages will have 12.865 negative exchanges per week or more.
Figure 4.3 depicts this result.
Figure 4.3: The probability of selecting a sample with M =12.865 or higher from a population with µM = 11.0
The probability of selecting a sample with M = 12.865 or higher is indicated bydetermining the z equivalent of a sample with M = 12.865 and then determiningthe proportion of the distribution at that point and higher in the population. Theproportion is indicated in red.
Note some important differences between this z and those calculated in Chapter 3.
· The difference between the mean of the population (µM = 11.0) and the sample mean(M = 12.865) is really quite modest, but the z value (z = 2.630) is comparativelyextreme. Recall that ±2z includes 95% of the distribution, and z = 2.630 is substantiallybeyond that.
· The reason for the rather large value of z is the quite small standard error of the mean,0.709. That value reminds us that variability in populations based on groups is smallcompared to variability based on individual scores, and it does not take much of adifference between the sample mean (M) and the mean of the distribution of samplemeans (µM) to produce an extreme value of z.
Another z Test
To evaluate the impact of group therapy on a group of juvenile offenders, a psychologistconstructs a study in which the level of social alienation among juvenile offenders attendingcourt-mandated counseling sessions is compared to the level of social alienation measured ina national sample. For a group of 15 juveniles who are the psychologist’s clients, socialalienation (SA) after 6 months of counseling has M = 13.554. Nationally, SA is μ = 14.500with σ = 2.734. What percentage of all randomly selected groups of juvenile offenders willhave mean levels of SA 13.554 or lower? The psychologist knows that µ, and therefore, µM =14.500, σ = 2.734, N = 15, and M = 13.554. First, calculate the standard error of the mean:
σM=σ/N‾‾√=2.734/15‾‾‾√=0.706
Then determine the value of z:
z=M−μMσM=13.554−14.5000.706=−1.340
The table value for z = −1.3450 is 0.4099.
The researcher wants to know what percentage of all juvenile offender groups will havemean SA scores 13.554 or lower. Since the national mean SA for juvenile offenders is14.500, we know that 50% of all groups will be 14.500 or lower. The proportion with 13.554or lower is 0.5 minus the table value for z = −1.340: 0.5 − 0.4099 = 0.0901. Multiplying thatproportion by 100 will indicate the percentage at or below that point: 0.0901 × 100 = 9.01%.
It looks as though the counselor’s therapy sessions are reasonably effective. This groupmanifests a level of social alienation considerably lower than what is represented in thenational population of juvenile offenders. Is the result just a chance outcome? How could the researcher know?
4.3 Statistical Significance
Like the z score problems in Chapter 3, the z test is a ratio of the difference (M − μM in thenumerator) compared to data variability (σM in the denominator). A large ratio indicates thatthe score (in the z score problem) or the sample mean (in the case of the z test) is quite distantfrom the mean to which it is compared.
With increasing values of z, is there a point at which the sample mean (M) becomes sodifferent from the mean of the distribution of sample means (μM) that a researcher shouldconclude that the sample mean is more characteristic of some distribution other than the oneto which it is compared? In the z test, when the sample is more characteristic of some otherpopulation rather than the one to which it is compared, it is statistically significant. To put itanother way, a statistically significant result is one that is extreme enough—sufficientlydistant from that to which it is compared—that it is unlikely to have occurred by chance.
In the first z test problem, we proceeded as though the sample of those who had filed fordivorce was a subgroup of all couples with troubled marriages. What if the sample is actuallymore characteristic of some other distribution, say, a population of couples for whom divorceis imminent? Can large values of z reflect the fact that the sample actually represents apopulation different from the one to which it was compared?
Consider another example before we answer this question. Those in a college honorsprogram are probably adults. If researchers are interested in studying intelligence, would it bereasonable to expect that the members of this group represent what is characteristic of alladults? From the standpoint of age (and in the absence of child prodigies), those honorsstudents are probably all adults, but in terms of intelligence, they probably are not typical.Perhaps they are more representative of the population of intellectually gifted adults than ofadults in general.
The individuals in every sample belong to many different populations. The couples on theverge of divorce belong to
· the population of married people;
· the population of adults;
· the population of adults in the particular state;
· the population of adults in the particular county;
· the population of couples with troubled marriages, and so on.
One of the questions the z test helps answer is whether a particular sample is mostcharacteristic of the population to which it is compared, or whether the sample is more likesome other population. The magnitude of the z value is the key to the answer.
Statistical Significance and Probability
In the case of the z test, an outcome is statistically significant when it meets these conditions:
· It is so unlike the population to which it is compared that statistically, at least, itrepresents some other population.
· The random selection of a sample from the population with the particular value of μMwould almost always result in a less extreme difference between M and µM.
So, at what point is an outcome nonrandom? Fisher (1925), who created the term statisticallysignificant, made the answer a matter of probability. If the probability that an outcome (in ourcase, the value of z) occurred by chance is p = 0.05 or less, the outcome is probably notrandom; it is statistically significant.
Try It!: #2
What does the term statisticallysignificant mean?
Although p = 0.05 is probably the mostcommon, other probability levels have alsobeen used to indicate statistical significance.Reviewing journal articles indicatesstatistical testing done at p = 0.01, p =0.001, and occasionally, even p = 0.1. It isup to the person doing the analysis to statethe level chosen to indicate statisticalsignificance (before conducting the test, bythe way). Because we can use the z test andthe z score table to calculate the probability of an occurrence (in addition to the other thingswe can do to determine the percentage of the population above a point, below a point, andbetween points), we can also use the table to determine whether an outcome is statisticallysignificant. In the first z test we completed, we compared the mean number of negative verbalexchanges in a sample of couples on the verge of divorce to the mean level of negativeexchanges among those identified as the population of couples with “troubled” marriages andfound that z = 2.630. The table value indicates that the probability of randomly selecting asample of couples that would have M = 12.865 or more negative verbal exchanges per weekwas p = 0.0043. At less than p = 0.05, that outcome is unlikely to have occurred by chance. Itis statistically significant.
The second z test dealt with how the group of juveniles in therapy compared to a nationalpopulation of juvenile offenders. For that problem, z = −1.340, and the table value for that zwas 0.4099. We determined that the probability of a group scoring M = 13.554 or lower was p = 0.0901 (shown in Figure 4.4).
Figure 4.4: The probability of selecting a sample withsocial alienation scores of M = 13.554 or lower from apopulation with mean SA of µM = 14.5
Visual representation of the probability of selecting a random sample ofindividuals with social alienation scores of 13.554 or lower, when thepopulation mean is 14.5.
Determining Significance Without the Table
Remember that ±z = 1.0 includes about 68% of the z distribution, so the probability ofrandomly selecting an outcome that occurs in the ±z = 1.0 area is p = 0.68. Nothing in thatregion is going to be statistically significant because those z values indicate results that arecharacteristic of the distribution as a whole. The uncharacteristic events are the significantones, and Fisher’s standard of p = 0.05 indicates that the key is a z value that excludes onlythe most extreme 5% of the distribution.
Recall that normal distributions are symmetrical. That 5% exclusion means that the mostextreme 2.5% of outcomes in the lower tail and the most extreme 2.5% of outcomes in theupper tail are the regions that include statistically significant outcomes. Because Table B.1provides proportions for only the upper half of the distribution, the z value, which includesall but the extreme 2.5% of outcomes, will be the point at which results become statisticallysignificant. If 2.5% needs to be the percentage excluded, 47.5% is the percentage included.As a proportion, 47.5% is expressed as 0.475.
· From Table B.1, find the z value which includes 0.475 from that point to the mean ofthe distribution.
· Because z = 1.96 includes 0.475 of the distribution, ± that value will include 0.95 ofthe distribution (2 × 0.475 = 0.95).
· Any time a z test produces a z = ±1.96 or greater, the result is statistically significant at p = 0.05.
For example, if the registrar at a university had a group of students applying for admission toa graduate program, and the admissions test scores for that group resulted in a value of, say, z= 1.98 compared to all applicants, the registrar would know straight away that those studentshave scores significantly greater than those to whom they were compared. Such a differenceis not likely to be an artifact of sampling variability.
Apply It! Confidence in the Claim
A parent is looking at private high schools for his child. A particular high schoolclaims that last year, its students scored above average on the math and verbalsections of the SAT. The parent, who knows something about statistical analysis,decides to test this claim.
Jack Hollingsworth/Thinkstock
The parent finds the nationwide results for lastyear’s SAT scores. The mean math SAT scorewas 500, with a standard deviation of 100. Themean verbal score was also 500, with astandard deviation of 100. The parent asks tosee the high school’s study. According to thestudy, the high school looked at SAT scoresfrom a random sample of 40 students for thatsame year. The mean math score was 515 andthe mean verbal score was 510.
The parent intends to test whether the high school scores represent a significantlydifferent population from the national scores. The z test will provide an answer. If thevalue of z could occur by chance with a probability p = 0.05 or less, the parent willview the scores from the particular high school as something that probably was not arandom outcome, an artifact of sampling variability. Since the math scores were themost different from the population scores, the parent examines them first.
Math Scores
µ, and therefore, µM = 500
σ = 100
N = 40
M = 515
Calculate the standard error of the mean:
σM=σ/N‾‾√=100/40‾‾‾√=15.811
Then determine the value of z:
z=M−μMσM=515−50015.811=0.949
The table value for z = 0.949 is 0.3289.
The probability of scoring 515 or higher is 0.5 − 0.3289 = 0.1711, which issubstantially greater than the 0.05 that indicates a statistically significant result.Although these particular high-schoolers scored higher on the SAT-math than thenational population—the difference between the sample and population meansanswers that question—they did not score significantly higher. This much differencecould have occurred by chance. The parent’s best explanation for the difference issampling variability rather than any systematic performance at this high school that isbetter than students are performing nationally. Since the math scores manifested thegreater difference, the parent finds little use in calculating the outcome for the verbalscores. There too, the result will not be statistically significant.
Apply It! boxes written by Shawn Murphy
Another View of Significance
The z = ±1.96 indicator for statistical significance assumes that the most extreme 5% ofoutcomes are statistically significant. There are alternatives. A p = 0.01 standard excludes themost extreme 1% of outcomes and p = 0.001 excludes the most extreme 0.1%. All of thesestandards for statistical significance are somewhat arbitrary. Fisher selected a point on thedistribution and said essentially, “Anything beyond this level of probability is unlikely tohave occurred by chance.” Not everyone agrees such a standard must exist. Dr. WilliamSmith, who chaired the department of statistics at Texas A&M University during the author’sstudent days, took the position that what is “significant” depends upon circumstances. Hisapproach was to calculate the probability that an event could occur by chance, and then letconsumers make their own decision about whether such an outcome is significant.
Dr. Smith is in good company. Rather than indicating that a result is or is not statisticallysignificant, many of the statistical packages that professionals use today simply determine theprobability than an event could have occurred by chance, and leave it to the person doing theanalysis to interpret the outcome. Fisher’s position fills the need to have a standard forstatistical significance in the absence of some other criterion, but the researcher may at timeshave a better sense than p = 0.05 of what is worthy of note.
Sampling Error as an Explanation of Difference
Virtually every z test will have some difference between M and µM, which means that z willhave some value other than 0. When the differences fall short of statistical significance (z <1.96), how can the researcher explain them? The answer could be sampling error. Because nosample can exactly emulate the population, most samples in the distribution of sample meanswill have a sample mean different from the population mean. In the SAT example, those fromthe particular high school the parent was examining scored better than the nationalpopulation, but the difference was not large enough to be statistically significant. Such adifference might reflect the fact that those selected for the sample group just happened to begenerally above the mean of the distribution. In the example concerning the number ofnegative exchanges in marriage, some of the difference may also be due to sampling error,but that factor alone is insufficient to explain the difference between M and µM.
More Confidence in the Sample
The foregoing discussion underscores the importance of having confidence in the samplefrom the start. Even though samples can never mirror populations exactly, large, randomlyselected samples minimize sampling error. However, it can be difficult to define large.
Consider the following example. An instructor wants to gauge the impact that a servicelearning course has on students’ attitudes toward community service. The university researchoffice has surveyed students’ interest in service learning and from the scores on theinstrument has determined a standard deviation of 8.294. While recognizing that the samplewill not have the same characteristics as the population, the instructor decides that the resultswill still be very helpful if the variability in the sample data digress from university-widedata by no more than 2 points. The instructor wishes to be 0.95 confident of the result.
Those two criteria, the amount of variation from the population and the required level ofconfidence, provide the parameters for determining minimum sample size (Sprinthall, 2000).Sprinthall’s formula follows:
Formula 4.3
n=((z)(σ)vatarion from σ)2
where
n = the required sample size
z = the value of z that corresponds to how certain of the result a researcher wishesto be. Because ±z = 1.96 includes the middle 95% of the distribution, using thatvalue in the formula provides p = 0.95 that the sample emulates the population. If0.99 certainty is required, z = 2.58.
σ = the standard deviation of the population. If the population standard deviation isunavailable, a sample standard deviation (s) can be substituted, although theestimate will lose some precision.
variation from σ = the amount a researcher is willing to allow the sample standarddeviation, s, to vary from the population standard deviation, σ.
Using this formula in our example results in σ = 8.294 and z = 1.96, so
n=((1.96)(8.294)2.0)2= approximately 66
With p = 0.95, which is the level of confidence required, a random sample of at least 66people will provide a sample whose characteristics are within 2 points of the populationstandard deviation of students’ interest in service learning.
Try It!: #3
If a result is not statistically significant,how is the difference between M andµM explained?
Changing the conditions can dramaticallyaffect the required sample size. If theinstructor needs to be within one point ofthe population standard deviation anddesires p = 0.99, note the impact on theresult:
n=((2.58)(8.294)1.0)2= approximately 458 people
The example indicates the impact on sample size of requiring both less variation from thepopulation and more certainty. It indicates the constant tension between how certain and howprecise we need to be on the one hand and the size of the needed sample on the other, and theresults can be dramatic. Increasing the level of certainty or requiring less error both requirelarger sample sizes, but at least Formula 4.3 can help strike a balance between the two.Samples that are very large can be time-consuming and expensive for researchers. Samplesthat are very small may not reflect the essential characteristics of the population, makinggeneralizing the results a problem.
Decision Errors
Statistical significance is based on the probability that an event could occur by chance, andinterpreting outcomes based on probabilities carries a risk. Consider the followingpossibilities related to the chapter’s examples:
· Is it not possible, however unlikely, that a researcher could accidentally sample thecouples in the distribution who have the most negative exchanges? Maybe they do notbelong to a distinct population at all. Maybe these couples are just from the mostextreme portion of the population of all married couples.
· On the other hand, is it not also possible that those juvenile offenders in the group-therapy study belonged to a population of juvenile offenders with significantly lowersocial alienation scores, but because they actually had extremely high SA scores beforethe therapy, their improvement brought them to just slightly lower, rather thansignificantly lower than the national mean?
Any statistical decision involves the risk of one or the other—but never both—in the sameanalysis. Because statistical decisions are based on probabilities rather than certainties, anystatistical decision based on a probability can result in a decision error. There are two typesof decision errors, and they are mutually exclusive.
Type I Errors
Type I errors in statistical testing occur when an outcome is initially determined to bestatistically significant, but further research and testing indicate that it is not. In other words,the first, errant conclusion is an anomaly that fails to hold up under further scrutiny. Theprobability of this error is defined by the level at which the testing occurs. If the criterion forstatistical significance is p = 0.05, and the result is deemed statistically significant, theprobability of a type I error is 0.05. Because type I error is also called alpha (α) error, thesignificance level of a test is sometimes noted in terms of the risk of alpha error, α = 0.05,rather than p = 0.05. It means the same thing, except that the author has chosen to indicatethe probability of type I error rather than referring directly to the criterion for statisticalsignificance. At p = 0.05, or α = 0.05, for every 100 times someone concludes that a result isstatistically significant, a type I error will occur an average of 5 times.
Although those most extreme outcomes are the least likely to occur, that most extreme 5% ofthe distribution is still part of the distribution in question. Outcomes in that area of thedistribution hold the greatest potential for a type I error. Two other items may be notedregarding type I error:
· The only time a type I error is possible is when a result is deemed statisticallysignificant. If there is no statistically significant outcome, the probability of a type I(alpha) error is zero.
· In a particular significant finding, a researcher has no way to know whether a type Ierror has occurred. Gathering new data and repeating the analysis is the only way tocheck, which is why replication studies are so important.
Type II Errors
In a z test, type II errors occur when the sample is actually characteristic of some population other than the distribution of sample means to which it was compared, but the statisticaltesting (z < 1.96) suggests no significant difference. This type of decision error is also calleda beta (ß) error.
Type II errors would present little problem if the populations involved were completelyseparate, but they often have important similarities. The population of all high schoolstudents taking the SAT probably bears a number of similarities to the population of high-achieving students. The more the populations overlap, the more likely decision errorsbecome.
Although the level at which the statistical test is conducted (often p = 0.05) defines thelikelihood of a type I error, the probability of a type II error is more elusive, and in fact wenever know the exact probability of committing this error, although some statistical tests aremore prone to it than others. Note also:
· The only time a type II error can occur is when a result is determined not to be statistically significant.
· In a particular analysis where the result appears to be not significant, a researcher hasno way to know whether a decision has resulted in a type II error.
Decision Errors and Power
Try It!: #4
An analysis results in a finding that isstatistically significant at p = 0.05.What is the probability of a type IIerror?
Is one error more damaging than the other?Do analysts have a preference for one typeof error? The answer, of course, dependsupon circumstances and especially on theimpact that a decision error has on thepeople involved.
Perhaps a committee is evaluatingcertification programs for mental healthprofessionals, and it deems the program atUniversity A to be significantly better than the competing programs. If the result is that thegraduates from University A receive preferential hiring, but the difference among programsis not statistically significant after all, the study has made a type I error.
On the other hand, perhaps a client has a serious illness and comes to a health professionalfor a diagnosis. If the health professional fails to recognize that the client is not healthy andmisses the condition that is affecting the client’s well-being, a type II error has occurred.
In short, which kind of error is more serious depends upon circumstances, but statisticiansmay have their own preference. Power in statistical testing is described in terms of thelikelihood of a type II error. The most powerful tests are associated with the fewest betaerrors. The power of a statistical test is symbolically indicated as 1 − ß.
Which Type of Errors to Use
Although type I and type II errors cannot both occur in the same analysis, the probability ofone affects the likelihood of the other. Mental-health professionals ordinarily must pass somesort of licensing requirement, perhaps in the form of a test. Like most professional licensingtests, it is most likely a good, but certainly not perfect, indicator of who is competent. Figure4.5 shows the possible decision errors.
Figure 4.5: Types of decision errors by outcome
The flowchart depicts the path of each type of decision error.
If type I error is thought to be the greater problem, the licensing body might simply raise therequired test score. This tactic would probably reduce the number of incompetent people whoare licensed. The companion problem it creates, however, is excluding more professionalswho actually are competent but because they do not test well fail to demonstrate theircompetence on the required measure—a type II error. This inherent connection between thetwo kinds of decision errors is why someone has to make a decision about which is the moredamaging.
In testing professionals like airline pilots and surgeons, the decision is straightforward. Usually it favors excluding some who are competent (therefore committing a type II error)rather than risk licensing some who are not competent (committing a type I error). Thepotential cost to the well-being of others is too great to do otherwise. In other circumstances, the greatest good is less clear.
4.4 The Confidence Interval
When the results from a z test are statistically significant, the sample best represents apopulation other than the one to which it was compared. In these instances, the mean of thesample (M) is in fact an estimate of the value of that other population mean. Because it is adiscrete value, M is called a point estimate of the population mean µM. For a variety ofreasons related to the nature of the sample and the way it was selected, M might not be a veryaccurate point estimate of that other population mean, however. The confidence interval (CI)provides a way to determine how precisely M estimates µM.
When a z value is significant, the result indicates that the sample represents a populationother than the one to which it was compared. The confidence interval allows us to estimatethe mean of that other population, or at least a range of values within which µM is likely tooccur. If the z test is not significant, we have no need for the confidence interval because ouranalysis indicates that the sample belongs to the population to which it was compared.Calculating the confidence interval requires values from the z test. The formula is
Formula 4.4
CI = ±z(σM) + M
where
CI = the interval within which the population mean is expected to occur
z = the table value that reflects the level at which the z testing wasconducted. For p = 0.05, z = 1.96.
σM = the value of the standard error of the mean from the z test
M = the value of the sample mean
For the study of couples engaged in negative verbal exchanges, the result was significant,indicating that the sample probably belongs to some distribution of sample means other thanthe one to which it was compared. The confidence interval will establish a range of valueswithin which the mean for that other population probably occurs.
CI = ±z(σM) + M
CI = ±1.96(0.709) + 12.865
CI = ±1.390 + 12.865 = 11.475, 14.255
With a probability of 0.95, the population which the sample represents has a mean (µM) valuesomewhere between 11.475 and 14.255.
Note that the level of probability is one of the factors affecting the size of the confidenceinterval. If we wish to be more certain of capturing the population mean, we can use a 0.99confidence interval instead of 0.95, and substitute z = 2.58 for z = 1.96 in the formula.Recalculating the confidence interval for p = 0.99,
CI = ±z(σM) + M
CI = ±2.58(0.709) + 12.865
CI = ±1.829 + 12.865 = 11.036, 14.694
Try It!: #5
What does the confidence interval for zdetermine?
A greater level of certainty of capturing thetrue mean of the distribution represented bythe sample requires a wider confidenceinterval.
The other factor that affects the width of theconfidence interval is the standard error ofthe mean, σM, which measures the amountof variability in the distribution of samplemeans. More data variability translates intoa larger standard error of the mean, which makes the confidence interval larger.
Researchers have no need for a confidence interval unless the z test results are statisticallysignificant. The reason can be illustrated by completing a confidence interval for the socialalienation (SA) problem. Recall that σM = 0.706 and M = 13.554 for that example.
CI = ±z(σM) + M
CI = ±1.96(0.706) + 13.554
CI = ±1.384 + 13.554 = 12.170, 14.938
Note that this confidence interval includes within its range the value of the originalpopulation, 14.500. That is because with a nonsignificant z value, a researcher wouldconclude that the population that the sample represented is likely the same population towhich it was compared. A nonsignificant z test value will always produce a confidenceinterval that includes the original population mean.
Apply It! How Long is Too Long?
A psychologist working with the military has studied post-traumatic stress disorder(PTSD) and is convinced of the relationship between length of deployment and theincidence of PTSD. The currently used scale for PTSD has been normed on servicepersonnel and yields μ = 54.375 with σ = 6.037 for all who have been deployed up to60 days. Using data for a sample of 50 individuals who have all been deployed for atleast 4 months, the psychologist has determined that M = 56.989. The psychologistwants to know whether this particular sample still represents the population of servicepersonnel deployed up to 60 days. Anxious about a type I error, or what is sometimescalled a false-positive error, the psychologist decides to test at α = 0.01.
µ and therefore µM = 54.375
σ = 6.037
n = 50
The sample mean, M = 56.989
First, using σ and n, the psychologist must calculate the standard error of the mean:
σM=σ/n‾‾√=6.037/50‾‾‾√=0.854
Then to determine z:
z = (M − μM)/σM = (56.989 − 54.375)/0.854 = 3.061
Recall that the standard for statistical significance at α = 0.01 is an absolute value of zof 2.58.We say “absolute value” because at this point, the issue is not whether thesample manifests significantly more PTSD or significantly less; the question is: doesthe sample represent the population? Had the value of z been −3.061, the psychologistwould draw the same conclusion about whether the sample represents the populationto which it was compared: that is, that it does not. In the case of a negative z value forthe sample compared to the population, the psychologist would have concluded thatfor some reason, this sample belongs to a population with significantly lower PTSDthan the one to which it was actually compared. As it is, however, the sample betterrepresents some population with a mean PTSD score significantly higher than thepopulation to which the sample was compared.
As we noted earlier in the chapter, z values are ratios of the difference to variability.In the case of the z test, it is a ratio of the difference between sample and populationmeans to the standard error of the mean, which is the relevant measure of datavariability in the z test. When that ratio is 1.96 or larger, the result is significant at α =0.05. When the ratio is 2.58 or larger, the result is significant at α = 0.01. Large ratiosare the result either of larger differences between sample and population means, or ofcomparatively small measures of data variability. In turn, small measures of thestandard error of the mean occur when data have little variability (small standarddeviation values), or large samples, or both. Although the difference between themeans of the sample and the population does not appear to be very large, the result isstatistically significant because of a relatively small standard error of the mean.
Since the data indicate that this sample represents a population of service personnelwith a mean measure of PTSD larger than the 54.375, the psychologist can calculatethe confidence interval to estimate a range within which that new population mean islikely to fall. The appropriate value of z for the CI will be the same as the one used inthe z test, 2.58.
CI = ±z(σM) + M
CI = ±2.58(0.854) + 56.989
CI = ±2.203 + 56.989 = 54.786, 59.192
With p = 0.99 confidence, the mean level of PTSD for those deployed to a combatzone for 4+ months is somewhere from 54.786 to 59.192. In terms of PTSD, thesepersonnel are likely a different population than those deployed up to 2 months.
To solidify our understanding of the z test, statistical significance, and confidenceintervals, consider one more example. The population mean (μ) for intelligencescores is 100, with a standard deviation (σ) of 15. A group of recruits who hope towork as intelligence analysts in a security agency have the following intelligencescores: 105, 100, 110, 120, 105, 110, 115, and 125. Are they characteristic of thepopulation as a whole? If not, with 0.95 confidence, what is the mean of thepopulation that they do represent? The intelligence scores for the 8 subjects areshown graphically in Figure 4.6.
Figure 4.6: Intelligence scores for recruit population
The bar graph depicts the intelligence scores for each member of a groupof recruits seeking employment as intelligence analysts.
The sample mean is 111.250.
The standard error of the mean is calculated from
σM=σ/n‾‾√=15/8‾‾√=5.303.
For the z test, we have
z=M−μMσM=111.250−1005.303=2.121
Depicted in Figure 4.7, this result is in the extreme upper tail of the distribution,beyond the middle 95%. By the usual standard, the group of eight recruits hasintelligence scores that are statistically significant, compared to the population of alladults.
Figure 4.7 Intelligence z scores compared to thedistribution curve
A frequency distribution of the intelligence scores shows that therecruits’ scores are higher than the population of all adults, a statisticallysignificant result that may represent a Type I error.
Since the conclusion is that the result was statistically significant, the only decisionerror possible is a type I error. It is possible that the sample does belong to thepopulation of all adults, rather than some separate population with a higher meanintelligence, but the probability of such an occurrence is just p = 0.05.
Although in any significant finding there is a probability of type I error, note that it iscomparatively low. The p = 0.05 suggests that, in circumstances such as this one, thefinding of statistical significance will fail to hold up just one time in 20. The greaterprobability is that the sample belongs to a separate population. The confidenceinterval will help us bracket the mean of that other population:
CI.95 = ±z(σM) + M = (+/−1.96 × 5.303) + 111.250 = 100.856, 121.644
With 0.95 confidence, the mean of that alternate population is somewhere between100.856 and 121.644. The lower bound of that confidence interval suggests that thealternate population might have a mean quite similar to the population of all adults.The fact that the lower value is quite similar to the mean of the original populationreflects the fact that the z value was only modestly greater than the limit forsignificance, as well as the fact that the sample size was very small.
Apply It! boxes written by Shawn Murphy
4.5 The z Test Using Excel
Although Excel’s Data Analysis includes an option for a z test, it is a different z test than the one weperform here. To complete our test, we must program some formulas. We will use the followingsample data.
A social worker’s caseload includes eight people with the following annual incomes (in thousands):
13.5, 18, 22.375, 25.240, 26, 29.331, 30, 30.
If all social workers’ clients have an average annual income of 19.500 and a standard deviation of4.525, are this particular social worker’s clients significantly different? We will use Excel to performthe work as follows:
1. Enter the income data into a spreadsheet in cells A1–A8.
2. Enter the formula =average(A1:A8) in cell A9 to have Excel calculate the mean.
3. In cell A11, determine the standard error of the mean by dividing the population standarddeviation by the square root of the number. The commands in Excel are =4.525/sqrt(8).
4. Determine the z value in cell A13 by entering the commands =(A9-19.5)/A11. The part inparentheses is the numerator in the z ratio: M (cell A9) − µM.
Figure 4.8 shows a screenshot of the display just before you press Enter.
The result is z = 3.004. Testing at p = 0.05 (for which z = 1.96), these eight people have significantlydifferent incomes than the population of all social workers’ clients.
Figure 4.8: Calculating a z test in Excel
Use Excel to perform a z test by entering the formulas identified.
Source: Microsoft Excel. Used with permission from Microsoft.
The z Test Using Excel
The z Test Using Excel
00:00
00:00
Summary and Resources
Chapter Summary
While the z test is not common in statistical testing, it is not unknown. Its unique valuecomes from the way it introduces us to statistical testing. The z test provides a goodintroduction to formal statistical testing and although an uncomplicated test, it addressesmany of the same issues that more advanced tests do. Generally speaking, behavioralresearchers are much more interested in analyzing the performance of groups than of singleindividuals. We have many reasons to wonder whether this or that group truly represents thepopulation to which it is being compared. The z test provides a mechanism for evaluatingprogram quality (is this program as good as the rest?), therapeutic techniques (is thisintervention as effective as others?), improvement strategies, and indeed any situation wherewe wish to compare one group for whom we have data to an identified population.
The z test is based on the distribution of sample means (Objective 1), a population of themeans of samples rather than of individuals’ scores. The central limit theorem indicates thatsuch a distribution will be normal even if the distribution of individual scores is not(Objective 2). The normality allows the use of the z table to analyze how groups compare topopulations (Objectives 4 and 6). Because the sample data researchers analyze sometimes donot fit well with the population presumed to be the source, the z test provides a way todetermine whether the sample belongs to some other population, an outcome related to theconcept of statistical significance (Objective 5).
When the sample is determined to represent some other population, the sample mean is apoint estimate of the value of that other µM, but only an estimate. The confidence intervalprovides a range of values within which the mean of that other population will occur with aspecified probability (Objective 7). In doing so, the confidence interval indicates theprecision with which M estimates µM.
Inferential statistical analysis involves the risk of making an incorrect decision. Occasionally,results that appear significant in one test will not hold up when the study is repeated with newdata. On the other hand, further analysis sometimes overturns a nonsignificant finding. Thesetype I and type II errors, respectively, remind us that statistical decisions are based onprobabilities rather than certainties (Objective 8).
Small samples, no matter how carefully selected, cannot mirror all the relevant characteristicsof complex populations, and populations involving people are invariably complex. For thisreason, a procedure to determine the size of the sample needed to emulate the importantcharacteristics of the population has some utility (Objective 3). Formula 4.3 meets that need.
Statistical significance is a very important concept in educational analysis. When newprograms or strategies are instituted, we often look for ways to determine whether theprogram makes a difference. The z test helps answer some of these questions. Specifically, itindicates whether a result is likely to have occurred by chance, whether the outcome israndom. As important as the z test is as an introduction, it has limitations in that it requiresaccess to two parameters, µM and σM. Although a researcher can usually determinepopulation means, the standard error of the mean sometimes is not accessible. The t testsChapter 5 discusses provide a way around this difficulty.
This summary offers a good barometer of your grasp of Chapters 1–4. Although some of thematerial has probably been familiar, many of the ideas are likely new. If this review makessense, that is excellent. If some areas seem difficult, take some time to go back to the relevantsections and review. Statistical analysis is incremental, as we have stressed before, so it isimportant to understand what has been presented before continuing. Working the examplesjust below—repeatedly if needed—will help.
Chapter 4 Flashcards
Key Terms
central limit theorem
confidence interval
decision errors
distribution of sample means
law of large numbers
power
random selection
sampling error
standard error of the mean
statistically significant
systematic sampling error
type I errors, or alpha (α) errors
type II errors, or beta (ß) errors
z test
Review Questions
Answers to the odd-numbered questions are provided in Appendix A.
1. If all the psychologists working at a state mental hospital have an average age of 47.5years, what will be the value of µM created from such a population?
2. A researcher calculates the standard deviation of the psychologists’ ages in ReviewQuestion 1. If the researcher calculates the standard error of the mean for the distribution of sample means for the same data, which will have the greater value?Why is there a difference?
b. For which population is the standard error of the mean the variability measure?
b. What is the equivalent of the standard error of the mean for a population based onindividual scores?
1. The assistant vice president for personnel at a college has job-performance scores forall clerical staff, with a mean value of 32.956 and a standard error of the mean of5.924.
c. What is the probability of randomly selecting a sample with a job satisfactionmean of 35.0 or higher?
c. If a group with M = 35.0 is selected, is it significantly different from thepopulation?
c. If a group is significantly different from a population, which type of decision erroris possible?
1. The clerical staff in a large law office scores the following on job performance:
25, 37, 38, 43, 44, 48, 51
If the mean level of performance for all clerical staff is 33.255, with a standard error ofthe mean of 3.248, are those in the law office characteristic of that population? Test at p = 0.05.
1. What is the relationship between the level of probability for a statistical test and theprobability of type I error in the event of a significant finding? When a result isstatistically significant, what is the probability of type II error?
1. The standard deviation for a major intelligence test is σ = 15.0. If, in a given year, thetest is administered to 347 people, what is the value of the standard error of the mean?
1. An exclusive graduate program requires GRE Quantitative scores of 500 or better. Thisyear’s entering class have n = 16 and M = 625.
g. Are they characteristic of a national population of graduate students for whom µ =500 with σ = 100?
g. What is the probability that a group of 16 applicants selected at random wouldhave σ = 100 and M = 525 or better?
1. A researcher wishes to gather a sample of people who have intelligence scores thatdiffer from the national standard deviation of 15 by no more than 3 points, with 0.95confidence.
h. How large must the sample be?
h. What will happen to the required sample size if the researcher insists on being0.99 confident?
h. Explain the answer to 8b.
h. How large must the sample be if it is to vary from the national standard deviationby no more than 2 points?
1. A group of social workers take a measure of optimism and score as follows:
11, 14, 14, 16, 19, 20, 22, 23, 27, 30
If the population standard deviation is 4.554,
i. What is the value of the standard error of the mean?
i. What is the z value for a z test with this group if µM = 26.0?
i. If 26.0 is the mean for all employed adults, is this group of social workerssignificantly different?
i. Use Excel to complete this problem. Refer to Figure 4.5 for help.
1. If a z test result is not significant, why will a confidence interval for the populationmean contain the value of the population mean to which the sample was compared?
1. What factors will reduce the size of a confidence interval?
1. If someone is testing at p = 0.01 and the result is statistically significant, what is theprobability of a type I error? What is the probability of a type II error?
Answers to Try It! Questions
1. The distribution of sample means has less variability than a distribution of individualscores because sample means moderate the effect of extreme scores. The larger thesample, the more extreme scores are minimized as factors in data variability.
2. “Statistically significant” means that the calculated value, z in this case, is large enoughthat it is unlikely to have occurred by chance; it is probably not a random outcome.
3. When the difference between M and µM in a z test is not significant, the difference isattributed to random variability; the value of M is one of the possible values of samplesdrawn at random from the distribution of sample means.
4. A type II, or beta, error can occur only when a result is determined not statisticallysignificant. When the result is significant, the probability of ß = 0.
5. Calculated only for a statistically significant result, a confidence interval for the valueof z indicates a range of scores within which the population mean that the sample doesprobably represent occurs.
t Tests
Owen Franken/Corbis
Chapter Learning Objectives
After reading this chapter, you should be able to do the following:
1. Explain the advantage of the one-sample t test over the z test.
2. Compare the one-sample t test to the independent t test.
3. Distinguish between one-sample and one-tailed t tests.
4. Explain hypothesis testing in statistical analysis.
5. Determine practical significance.
6. Construct a confidence interval for the difference between the means.
7. Discuss research applications for the t tests.
Introduction
The z test (Chapter 4) involves more than just an expansion of the z score from individuals togroups. The z test also introduced us to statistical significance. Those who work withquantitative data need to be able to distinguish between outcomes that probably occurred bychance and those that are likely to emerge each time the data are gathered and analyzed. Forexample, data indicate that a group of clients, each one grieving the loss of a loved one,becomes more positive and peaceful with time. A therapist needs to know whether this wouldhave happened anyway with the passage of time, or whether it has something to do with thetreatment the therapist provided. The z test answers such questions.
The z test has important limitations, however. Its greatest difficulty is the requirement of avalue for the population standard error of the mean, σM. That value is not the type ofinformation that tends to come up just as a matter of course, and it may be inaccessible whenthe researcher lacks access to population data, including the population standard deviation.
The z test’s second limitation is allowing only one type of comparison: a sample to apopulation. What if two therapists want to compare their respective groups of grieving clientsto see if one type of grief counseling is better than the other? The z test does not allow thatcomparison, which is where William Sealy Gosset comes in.
Gosset worked for Guinness Brewing during the early part of the 20th century. Part of hisresponsibility was quality control, and he studied ways to make sure that day-to-day brewingremained consistent with Guinness’s standards. A man of remarkable ability, Gosset devisedprocedures for quantifying product quality and then testing the consistency of the qualityover time. To help him in his analyses, he developed the t tests.
Gosset recognized that his t tests had application well beyond studying the quality of beerand wanted to publish information about his tests so that others could benefit. A no-publishing policy at Guinness—instituted after a previous employee published what thecompany considered trade secrets—presented a roadblock. Believing that his research wouldnot compromise Guinness, Gosset published anyway under the pseudonym “Student.” Traditional statistics textbooks still contain references to “Student’s t.”
The pseudonym Gosset selected suggests his unpretentious nature. As his career progressed, he became associated with Karl Pearson and Ronald Fisher, who also made remarkablecontributions to statistical analysis. These were men of substantial ego in a runningprofessional and personal battle with each other, but Gosset maintained good relationshipswith both.
This is a self-assessment and will not affect your grade. You may only take this pre-test once.
Test Ch 5: t Tests
Top of Form
1. “Statistical” and “practical” significance are equivalent terms.
· a. FALSE
· b. TRUE
2. An independent t test is used to determine if a particular sample represents a specified population.
· a. FALSE
· b. TRUE
3. The null hypothesis predicts that there will not be a statistically significant result.
· a. FALSE
· b. TRUE
4. A question analyzed in a one-tailed t test is “is the sample significantly greater than the population?”
· a. TRUE
· b. FALSE
5. There are many t distributions.
· a. TRUE
· b. FALSE
Finish
5.1 Estimating the Standard Error of the Mean
The z test statistic requires two parameter (population) values: the mean (µM) and thestandard error of the mean (σM):
z=M−μMσM
Researchers usually have access to the population mean. Gosset’s problem was how to workaround the more elusive population standard error of the mean, σM, that the z test alsorequires.
Although σM can be calculated by dividing the population standard deviation by the squareroot of the number
σM=σ/n‾‾√
that offers small consolation if the population standard deviation (σ) is unavailable. How canresearchers proceed without a value for sigma?
The sample mean, M, is one of the possible values of the population mean, μ. Likewise, thesample standard deviation, s, is one of the possible values of the population standarddeviation, σ. Gosset reasoned that statistics from samples can provide quite accurateestimates of their equivalent population parameters, particularly when the samples are fairlylarge and randomly selected.
Samples can never exactly replicate the characteristics of populations, however, and thedifference is called “sampling error.” To adjust for sampling error, the difference between μand M or between σ and s, Gosset proposed a correction. Part of the genius of his solution isthat correction is graduated; with small samples, the adjustment is greatest when the risk ofsampling error is greatest.
To determine the standard error of the mean, Gosset used the sample standard deviation (s) asan estimate of the population standard deviation (σ). So instead of σM = σ/√n, he estimatedthe standard error of the mean, SEM, this way:
Formula 5.1
SEM=sn‾‾√
where
s = the standard deviation of the sample
n = the number in the sample
So, perhaps a researcher examining depression among the aged finds the followingdepression scores: 13, 17, 16, 17, 11, 13, 18, 15, 15, and 12. Rather than trying to determinethe standard deviation (σ) for depression in this population, the researcher can calculate thesample standard deviation (s) and divide it by the square root of 10 to determine theestimated standard deviation. That value is
2.359/10‾‾‾√=0.746.
5.2 One-Sample t Test
Using the estimated standard error of the mean (SEM) in the place of σM changes the test statistic for the one-sample t test to
Formula 5.2
t=M−μMSEM
where
t = the calculated value of t
M = the sample mean
µM = the mean of the distribution of sample means
SEM = the estimated standard error of the mean
The numerator of the one-sample t test is the same as the numerator for the z test becauseboth tests answer the same question: Is the sample characteristic of the population to which itis compared? The t test, however, frees researchers from needing to know the populationstandard error of the mean. Choosing to rely on statistics from samples instead of populationvalues, however, creates the risk of sampling error.
Besides substituting the estimated standard error of the mean (SEM) for the populationstandard error of the mean (σM), another difference between z and the t tests is that there isjust one z distribution but many t distributions. Because there is just one z distribution, onevalue, called a critical value, indicates statistical significance. In the z test, that value is z =1.96 (Table B.1). With many t distributions, each with its own characteristics, a differentcritical value indicates statistical significance for each distribution. Different t distributionsare defined by their degrees of freedom (df). For the one-sample t test, the degrees of freedomare the number of scores in the sample minus one: df = n − 1. So if the sample size (n) = 10,then df = 9; if n = 30, df = 29, and so on.
The risk of sampling error is greatest when samples are smallest because, of course, smallsamples have a hard time emulating population characteristics. So Gosset used df, whichdirectly correlates with sample size, as an index to the critical values which indicatestatistical significance. Small samples (and therefore small df) require a larger value of t toindicate statistical significance than when sample sizes (and df) are larger. As df increases,the critical values required for significance decline until ultimately, with a sample of infinitesize (which of course is possible only in theory), the critical values for t match that for z,1.96.
Calculating the One-Sample t
The steps for completing the one-sample t follow:
1. Calculate the sample mean, M (Formula 1.1) and the sample standard deviation, s(Formula 1.3).
2. Calculate the estimated standard error of the mean, SEM (Formula 5.1).
3. Calculate the value of t (Formula 5.2).
4. Compare the calculated t to the critical value for t for the appropriate df.
Consider a psychologist working for the police department in a major city who is interestedin job stress among law enforcement officers. The psychologist wants to know if job stress issignificantly different among police officers than it is in the general population, where stresshas a mean of 27.353. The psychologist administers the too l t est i nitia t iv e (“too-tite,” forshort) to 10 police officers at random. Their too-tite scores are as follows:
22, 26, 29, 29, 33, 35, 36, 38, 40, 42
First, verify that for this sample M = 33.0, and s = 6.412. Remember that the mean of apopulation based on individual values equals the mean of the distribution of sample means,or µ = µM (Chapter 4). So, because µ = 27.353, µM = 27.353. Next, estimate the standarderror of the mean:
SEM=s/n‾‾√
(Formula 5.1)
SEM=6.412/10‾‾‾√=2.028
Then calculate the value of t:
t = (M – µM) / SEM = (33.0 − 27.353) / 2.028 = 2.785;
10‾‾‾√=2.028
Finally, the value of the calculated t must be compared to the critical value for t, listed inTable 5.1 (Table B.2 in Appendix B). The researcher’s test happens to be a “two-tailed” test(explained later in the chapter), so consult the columns under “Two-tailed tests.”
Try It!: #1
In terms of degrees of freedom, which t distributions are accompanied by thehighest critical values for t?
Earlier, we noted that the greatest threats tosampling error occur when samples aresmallest. That reality is reflected in thecritical values in the table. Note howdramatic the change is from, say 2 to 3degrees of freedom compared to the changefrom 29 to 30.
The first column in the table is for degreesof freedom. In our example, n = 10, so df =9. The second column lists the criticalvalues of t when the test is conducted at p = 0.05. The p value indicates the probability that asignificant finding will result in a type I error. As we noted in Chapter 4, the probability of atype I or alpha error (α) is always the same as the probability level (p) at which the test isconducted. The 0.05 is the traditional level at which statistical tests are conducted. Forexample, a research article that does not state the alpha level usually assumes α = 0.05.
The third column has the critical values of t when testing at p (or α) = 0.01. For p = 0.05, thecritical value for df = 9 is 2.262. So as not to confuse the critical value of t with the calculatedvalue from the t test, it is helpful to write the critical value and note its df and the level ofprobability for the test as follows:
t0.05(9) = 2.262
The subscripts to t indicate that the test was conducted at p = 0.05 with 9 degrees of freedom.
Note that the critical value for testing at p = 0.01 with df = 9 is 3.250. The greater the criticalvalue, the less likely that such a calculated value of t would occur by chance, which meansthat the likelihood of a type I error is reduced with the more stringent p or alpha level. Thecompanion problem, as we noted in Chapter 4, is that the likelihood of a type II errorincreases correspondingly. This relationship between type I and type II errors places theresearcher on the horns of a dilemma: Which error poses the greater risk, type I or type II?
In a circumstance where the question is whether, for example, psychologists’ analyticalabilities are characteristic of the analytical abilities of all adults, failing to find that thepsychologists are unique when actually they are is a type II error. However, the type II errorprobably poses little risk, except perhaps for those attempting to identify the characteristicsof effective psychologists. On the other hand, failing to identify that violent offenders havesignificantly greater levels of social alienation than what prevails in the general populationmay be much more important. If social alienation leads to the commission of violent crimes,and a researcher sets an alpha level that is so stringent (0.01 or 0.001) that it does not identifythose with the greatest level of alienation as belonging to a distinct population, perhaps thereis no opportunity for intervention. In this latter situation, the researcher would probablyprefer a type I to a type II error.
Table 5.1: t distribution critical values
|
df |
Critical t value |
|||
|
|
Two-tailed tests |
One-tailed tests |
||
|
|
p = 0.05 |
p = 0.01 |
p = 0.05 |
p = 0.01 |
|
1 |
12.706 |
63.657 |
6.314 |
31.821 |
|
2 |
4.303 |
9.925 |
2.920 |
6.965 |
|
3 |
3.182 |
5.841 |
2.353 |
4.541 |
|
4 |
2.776 |
4.604 |
2.132 |
3.747 |
|
5 |
2.571 |
4.032 |
2.015 |
3.365 |
|
6 |
2.447 |
3.707 |
1.943 |
3.143 |
|
7 |
2.365 |
3.499 |
1.895 |
2.998 |
|
8 |
2.306 |
3.355 |
1.860 |
2.896 |
|
9 |
2.262 |
3.250 |
1.833 |
2.821 |
|
10 |
2.228 |
3.169 |
1.812 |
2.764 |
|
11 |
2.201 |
3.106 |
1.796 |
2.718 |
|
12 |
2.179 |
3.055 |
1.782 |
2.681 |
|
13 |
2.160 |
3.012 |
1.771 |
2.650 |
|
14 |
2.145 |
2.977 |
1.761 |
2.624 |
|
15 |
2.131 |
2.947 |
1.753 |
2.602 |
|
16 |
2.120 |
2.921 |
1.746 |
2.583 |
|
17 |
2.110 |
2.898 |
1.740 |
2.567 |
|
18 |
2.101 |
2.878 |
1.734 |
2.552 |
|
19 |
2.093 |
2.861 |
1.729 |
2.539 |
|
20 |
2.086 |
2.845 |
1.725 |
2.528 |
|
21 |
2.080 |
2.831 |
1.721 |
2.518 |
|
22 |
2.074 |
2.819 |
1.717 |
2.508 |
|
23 |
2.069 |
2.807 |
1.714 |
2.500 |
|
24 |
2.064 |
2.797 |
1.711 |
2.492 |
|
25 |
2.060 |
2.787 |
1.708 |
2.485 |
|
26 |
2.056 |
2.779 |
1.706 |
2.479 |
|
27 |
2.052 |
2.771 |
1.703 |
2.473 |
|
28 |
2.048 |
2.763 |
1.701 |
2.467 |
|
29 |
2.045 |
2.756 |
1.699 |
2.462 |
|
30 |
2.032 |
2.750 |
1.697 |
2.457 |
|
∞ |
1.960 |
2.576 |
1.645 |
2.326 |
Source: Critical values of the t distribution. Retrieved from shazam.econ.ubc.ca/intro/critval.htm
Interpreting t Test Results
If the outcomes from the police officers’ job-stress level test were z test results, the onlyquestion would be whether the calculated value is equal to or greater than 1.96. If so, theresult is significant. With t, the question is still whether the calculated value is equal to orgreater than a table value, but that value changes depending upon df. In the job-stress levelexample, the calculated t = 2.785. From the table of critical values, the value for p = 0.05 and df = 9 (t0.05(9)) is 2.262. Because the calculated value is larger than the table value, the resultis statistically significant, which in our example, indicates that police officers havesignificantly greater stress than people in the general population. Figure 5.1 illustrates thisresult, where the plus- and minus-1.96 values associated with statistical significance for a ztest are replaced with the plus- and minus-2.262 for a t test with 9 degrees of freedom. Thedifference in job stress between law enforcement officers and the general public probably didnot occur by chance.
According to these data, law enforcement officers represent a population with job stresshigher than the general public’s. Such a finding confirms a potential problem to decision-makers and allows those involved to design services for law enforcement personnel that canhelp remediate stress-related problems.
Another Example
Pursuing stress findings, the psychologist decides to compare social workers’ levels of stressto the job stress of the general public. For 12 randomly selected social workers, the stressscores are as follows:
19, 21, 22, 25, 27, 27, 32, 33, 35, 35, 36, 37
1. Verify that M = 29.083, and s = 6.374; µM is still 27.353.
2. Calculate the standard error:
SEM=s/n‾‾√=6.374/12‾‾‾√=1.840
3. Next, find the value of t:
t=M−μMSEM=29.083−27.3531.840=0.940
4. t0.05(11) = 2.201
Figure 5.1: Statistical significance in a t test with 9degrees of freedom
Because the sample size—and therefore the df—have changed, the critical value is differentfor this problem than it was for the example about stress among police officers. At p = 0.05and df = 11, the difference between social workers and the general population is notstatistically significant; the calculated value of t is less than the critical value from the table.Although as sample size (and so df) increase the critical value declines, a t = 0.940 is nevergoing to be statistically significant, as scanning down the table of critical values verifies. At p= 0.05, no critical value is as small as 0.940, and the critical values at p = 0.01 values arelarger yet. Even with a sample of infinite size, the lowest critical value when testing at p =0.05 is 1.96, the same as the critical value for the z test. In terms of stress, social workers areno different than people in the general population. In contrast to law enforcement officers,probably no special program or action needs to address the stress-related problems of thisparticular group.
The One-Tailed Test
Try It!: #2
If the degrees of freedom for a one-sample t test (df) = 14, what is theassociated sample size?
The z and t significance tests to this pointhave been what are called “two-tailedtests.” That description means that thesample mean is either significantly less thanthe population mean, or significantly greater. To put it differently, in a two-tailedtest, whether the calculated values of z or tare positive (placing the result in the uppertail of the distribution) or negative (in thelower tail of the distribution) does not matter. The only issue is whether the absolutecalculated value—the value without regard to the sign—is as large as or larger than thecritical value.
In the first t test example, if law enforcement officers’ too-tite scores had been the samedistance below the mean for the population as they were above (that is, if M = 21.706 insteadof 33.0 and the SEM had remained the same at 2.028), the t value would have been
t=M−μMSEM=21.706−27.3532.028=−2.785
While we would interpret this result differently—law enforcement officers are significantlyless, rather than more, stressed than the general population—the result would have beenstatistically significant nevertheless. In a two-tailed test, the sign of t does not matter. But inthe original problem, the police psychologist might have framed the question this way: Arepolice officers significantly more stressed than the general population?
With that wording, the analysis becomes a one-tailed test. Language such as more stressedor less stressed indicates a prediction about more precisely how the sample is expected todiffer from the population, rather than just asking whether the sample is different from thepopulation. Specifying the direction of the difference (more or less) indicates the tail of thedistribution in which the difference is expected to occur.
The Statistically Significant Region in the One-Tailed Test
To restate the concept, the way the question is framed determines whether the result canoccur in either, or just one, tail of the t distribution. In a two-tailed test conducted at p =0.05, for example, is the sample significantly different from the population?, the followingapply:
· A statistically significant result can occur in either the highest 2.5% of the distribution or the lowest 2.5% of the distribution.
· The sign of the calculated value of t does not matter.
Apply It! Testing a Pilot Program
Mandy Glinsbockel/Demotix/Corbis
An innovative teacher has implementedmeditation (she calls it “quiet time”) in herclassroom to relieve high stress, increase testscores, and improve behavior among studentswho participate. The school principal wants todetermine whether the quiet-time students aregetting better grades than the population of allstudents at the school. Note that when thequestion is whether the quiet time students aredoing better, not just performing differently, the comparison becomes a one-tailedtest. Because of widespread skepticism about the program from school boardmembers, the principal chooses a conservative p = 0.01 criterion for statisticalsignificance.
During the 2010 school year, 10% of the student population was chosen at random toparticipate in the quiet-time program. That year, the grade point average for thestudent population as a whole was 2.53. The GPAs for 12 randomly selected studentswho had been involved in the quiet-time program were as follows:
2.49, 2.85, 3.18, 2.82, 2.89, 2.50, 3.26, 3.20, 2.48, 2.93, 2.67, 2.47
For these 12 students, M = 2.812 and s = 0.295.
SEM=s/n‾‾√=0.295/12‾‾‾√=0.085
Note that because µ = 2.530, µM = 2.530.
t=M−μMSEM=2.812−2.5300.085=3.318
To compare the calculated t to the critical value for t, consult the columns in Table 5.1for one-tailed test (the final two columns in the table). For this analysis, with p = 0.01and df = 11, the critical value of t 0.01(11) = 2.718.
At p = .01 level and df = 11, the results are statistically significant because 3.318 >2.718. Therefore, at p = 0.01, students in the quiet-time program had significantlyhigher GPAs than those of the general middle-school student population.
Using statistics, the principal was able to present quantitative evidence that the quiet-time program is associated with higher student GPAs. To decrease the possibility oftype I error, the principal selected p = 0.01 rather than 0.05. Perhaps this analysis, aswell as other tests measuring stress and behavioral issues, can enable the principal toraise additional funds to expand the pilot program.
Apply It! boxes written by Shawn Murphy
Figure 5.2 indicates a two-tailed test.
Figure 5.2: Statistically significant regions in a two-tailedtest
In a one-tailed test conducted at p = 0.05, for example, is the sample significantly greater than the population?, the following apply:
· A statistically significant result can occur in only the highest 5% of the distribution.
· The sign of the calculated value of t is all important.
In the t test example comparing the police officers to the general public, the question “arepolice officers significantly more stressed than the general public?” would have been a one-tailed test, with a rejection region resembling the distribution in Figure 5.3.
Figure 5.3: Statistically significant regions in a one-tailedtest where the sample is predicted to be higher than thepopulation
Critical Values in the One-Tailed Test
Because the rejection region is divided between both tails of the distribution in two-tailedtests and placed entirely in one tail in one-tailed tests, the point at which a result becomesstatistically significant differs in one- and two-tailed tests, which means that they havedifferent critical values. When testing at p = 0.05, the critical value for a two-tailed test is thepoint which includes 0.475 out of each 0.5 of the distribution. The equivalent proportion for aone-tailed test includes 0.45 of the distribution. The z value that includes 0.45 of thedistribution is z = 1.65 (actually z = 1.645, which the table does not display). Recall that the zvalue that includes 0.475 of the distribution is z = 1.96. In our last example, where n = 10 (df= 9), the critical value for a two-tailed test (t0.05(9)) is 2.262. For the equivalent one-tailedtest, the value is 1.833. Table 5.1 shows the critical values for one-tailed tests.
Try It!: #3
A researcher asks whether intelligencescores are significantly different formath majors and journalism majors. Isthe researcher calling for a one-tailedor a two-tailed test?
The less-demanding critical value makes it“easier” to find statistical significance witha one-tailed test, as long as the difference isin the direction predicted; less extremevalues of t can be statistically significant inone-tailed tests. The downside to one-tailedtests is that the opposite tail contains norejection region. If something unexpectedoccurs and the value of t or z has the“wrong” (the unexpected) sign, there is noway to find statistical significance, nomatter how extreme the value. To protectagainst having to ignore the unexpectedoutcome, some researchers and analysts stay clear of one-tailed tests.
5.3 Hypothesis Testing
In statistical testing, the calculated value either is, or is not, significant. Those twopossibilities are represented in predictions called the null hypothesis and the alternatehypothesis. The null hypothesis predicts that the result will not be statistically significant:that there is no (null) difference between the population that the sample represents and thepopulation to which it is compared. The alternate hypothesis predicts that the result will bestatistically significant: that there is a difference.
· The null hypothesis for a z, or for a one-sample t test, is written as follows: H0: µ1 =µ2. Symbolically, the null hypothesis (H0) states that the mean of the population fromwhich the sample was drawn (µ1) has the same value as the population to which thesample is compared (µ2).
· The alternate hypothesis (HA) for two-tailed tests—we will address one-tailed testsbelow—is stated this way: HA: µ1 ≠ µ2. The mean of the population from which thesample was drawn (µ1) does not have the same value as the population to which thesample is compared (µ2).
In hypothesis testing, results are always stated in terms of the null hypothesis. When the z or t value is statistically significant, one “rejects the null hypothesis,” which means that thesample probably belongs to a population other than the one to which it was compared. Whena result is not statistically significant, one “fails to reject the null hypothesis.” Although wemight be tempted to say “accept the null hypothesis,” it would be extremely difficult to provethat µ1 = µ2. Because the calculated value is not great enough to reject the null hypothesisdoes not necessarily mean that µ1 = µ2. We know only that there is not enough of a differenceto let us reject H0.
The Null and Alternate Hypotheses in a One-Tailed Test
In one-tailed tests, the null hypothesis is the same as it is for two-tailed tests, H0: µ1 = µ2.The alternate hypothesis, however, reflects the “greater-than,” or “less-than” prediction. If thepsychologist researching job stress had predicted higher job stress for the law enforcementofficers than for the general population, the alternate hypothesis would have been HA: µ1 >µ2, i.e., the mean of the population represented by the sample is greater than the mean of thepopulation to which it is compared. In terms of the example, the population the lawenforcement officers represent is expected to have a higher mean stress level than the generalpopulation. Note that the first population mean, µ1, represents the mean of the populationfrom which the sample was drawn. Placing the mean of the population from which thesample was drawn first in the hypothesis matches the way the z and t tests are written, with Mpreceding µM in the numerator of the test statistic (M − µM). If the researcher predicts thesample to belong to a population where the mean is lower than the population to which it iscompared, the alternate hypothesis is HA: µ1 < µ2, the mean of the population represented bythe sample is less than the mean of the population to which it is compared.
Another Example
LWA-Dann Tardif/Corbis
How can an analyst compare thecompassion among a sample ofnurses’ assistants to the level ofcompassion among the generalpopulation?
A hospital administrator reads a study indicatingthat compassion for patients is higher amongnurse’s assistants than among health-careprofessionals generally and wishes to test thatfinding at San Juan General Hospital. The mean forall health professionals on the I ndex of CA reful R ecuperative E ffort (the I-CARE) is 19.600. For 10nursing assistants, I-CARE results are
15, 17, 20, 20, 23, 23, 23, 25, 27, 28
Do nursing assistants have more compassion thanother health-care professionals? Is the differencestatistically significant?
The null hypothesis for this analysis is H0: µ1 = µ2.
To test whether nurses’ assistants have morecompassion for patients than the population of health care professionals does, the alternatehypothesis is HA: µ1 > µ2.
To complete the test, follow these steps:
1. Verify that for nursing assistants, the I-CARE M = 22.1 and s = 4.149.
2. Note that
SEM=s/n‾‾√=4.149/10‾‾‾√=1.312
3. Calculate
t=M−μMSEM=22.1−19.6001.312=1.905
4. In the column for one-tailed tests, note that the critical value is t0.05(9) = 1.833.
· With the entire rejection region in one tail of the distribution, the calculated valueof t is more extreme than the critical value from the table.
· The results are statistically significant.
· Reject H0.
Try It!: #4
Had the I-CARE problem been a two-tailed test, (a) what would the alternatehypothesis have been, and (b) whatwould the correct decision have been?
Try It!: #5
If the t value in the I-CARE problemwere −1.905 instead of +1.905, wouldthe statistical decision change?
5.4 The Independent t Test
Gosset’s one-sample t test represents an important development to anyone who needs to make ajudgment about whether a particular sample represents a specified population. It relieves the researcherof the need to have that standard error of the mean parameter (σM) for the entire population, a valuewhich may not be accessible.
But if rather than does the sample belong to the population the question is whether two samples belongto different populations, the z and one-sample t tests offer no way to answer. Another approach isrequired. Suppose a psychologist working with substance abusers divides them into two samples, eachreceiving a different kind of therapy, and wishes to know which is better? Questions that compare twosamples to each other are the domain of the independent t test. In this test, Gosset designed aprocedure where everything needed to complete the test can be calculated from the sample data. Nopopulation parameters are necessary.
The word independent in the term independent t test refers to the nature of the samples. They mustinvolve separate groups; subjects who are in one group cannot also be in the other. (In other words,this procedure does not allow researchers to measure a group and then apply a different treatment andmeasure them again; such a procedure describes what is called the before/after t test.)
A Population Based on Difference Scores
The z test and the one-sample t test are based on the distribution of sample means. A frequencydistribution based on the means of samples makes it possible to ask questions about whether aparticular sample is likely to have been drawn from a particular population.
The independent t test is based on a population created from “difference scores.” Rather than samplingone group at a time and then plotting the sample mean to create the distribution of sample means, thispopulation is created by
· selecting two samples of the same size,
· computing the sample mean for each,
· subtracting the second mean from the first (M1 − M2), and
· repeating this with enough randomly selected pairs of samples until the result is a distributionof difference scores.
When the first sample mean has a higher value than the second sample mean (M1 > M2), the differencewill be positive and plotted in the right half of the distribution. When the second mean (M1 < M2) hasthe higher value, the difference is plotted in the left half of the distribution, and so on. The mean forthis distribution of difference scores is symbolized as µM1 − M2 with the subscript M1 − M2 reminding usthat this parameter is the mean of all differences between pairs of sample means. Although the meansof pairs of samples usually have some difference, over many pairs of samples, the positive differencescounter the negative differences and the mean of all those M1 − M2 differences will be 0, µM1 − M2 = 0.
In the z distribution, σ measured variability. In the distribution of sample means σM, the standard errorof the mean, measured data variability. In the distribution of differences, variability in the differencescores is measured with the standard error of the difference, σM1 − M2. In practice, however, thestandard error of the difference is estimated, just as the standard error of the mean was estimated forthe one-sample t test.
Using the Distribution of Difference Scores to Explain Results
As we noted earlier, the mean of the distribution of difference scores for any population is always 0.That population indicates how much M1 − M2 difference can be expected to occur by chance when twosamples are selected from the same population. When differences between sample means exceedexpectations—when the value of t in an independent t test is significant—the samples probablyrepresent different populations.
Perhaps a researcher selects two random samples of subjects from the population of all clericalworkers in a particular city. If the researcher measures subjects’ analytical ability, the result wouldprobably show some differences between the two groups, but they would likely be minor—twosamples drawn from the same population. If a facilitator presents analytical techniques in a series ofweekly seminars to those in one sample, and then the researcher measures analytical ability again after60 days, perhaps the difference between the two groups grows too great to be explained by samplingvariability. The evidence would be a statistically significant t value.
The Hypotheses in the Independent t Test
The null and alternate hypotheses for independent t tests look the same as for the one-sample t test, H0:µ1 = µ2 and H0: µ1 ≠ µ2. The distinction is they reflect two samples rather than just one. The nullhypothesis predicts that the mean of the population from which sample one was drawn has a value thesame as the mean of the population from which sample two was drawn. In other words, the samplescome from the same population.
The alternate hypothesis maintains that the samples were drawn from different populations. If theindependent t test is one-tailed, the alternate hypothesis indicates the direction of the predicteddifference, just as with the one-sample t test. Either H0: µ1 < µ2 or H0: µ1 > µ2.
The test statistic for the independent t test is
Formula 5.3
t=M1−M2SEM1−M2
The numerator in the test statistic for the independent t test is actually, M1 − M2 − µM1−M2, the mean ofthe first sample minus the mean of the second sample minus the mean of the distribution ofdifferences. But because µM1−M2 is always 0, that part of the term is usually omitted, leaving us with M1 − M2.
For the denominator of the test statistic, Gosset provided an estimated standard error of the difference(SEM1−M2). It takes the place of the population standard error of the difference σM1−M2, which wouldbe extremely difficult to calculate. That parameter value that would require the calculation of thestandard deviation (σ) of all differences between all possible pairs of sample means in the distributionof differences! Clearly, that value is better estimated than calculated directly.
Before actually using the independent t test, note the harmony in the tests covered to this point:
1. The numerator for each procedure, from the z score to the independent t, involves a differencescore.
· For the z score, the numerator is x − M.
· For the z test, the numerator is M − µM.
· For the one-sample t test, the numerator is also M − µM.
· Similarly, for the independent t test, the numerator is M1 − M2.
2. The denominator for each involves a measure of data variability.
· For the z score, the denominator is s.
· For the z test, the denominator is σM.
· For the one-sample t test, the denominator is SEM.
· The independent t also involves a measure of data variability, SEM1−M2.
The Variability Estimate
Because this test includes data from two samples, the variability statistic, SEM1 − M2, must include thevariance within both samples, making the estimated standard error of the difference a “measure ofpooled variance.” When the number in both samples is equal, the formula for the estimated standarderror of the difference is as follows:
Formula 5.4
SEM1−M2=(SEM1)2+(SEM2)2‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√
where,
SEM=s/n‾‾√
SEM12 = the square of the estimated standard error of the mean for sample 1
SEM22 = the square of the estimated standard error of the mean for sample 2
An Independent t Test Example
Suppose that a sports psychologist is interested in athletes’ willingness to take risks as they participatein their sport. With research to suggest that having a winning record prompts athletes to repeat whathas been successful in the past and to be inclined toward conventional play, the psychologist reasonsthat risk-taking will be lower in athletes with winning records than in those with losing records. Usingthe D emonstration A ssessment for R eady E valuation ( DARE ), the psychologist collects data from twoteams of eight volleyball players, the first with a winning record, the second with a losing record. Notethat the researcher’s hypothesis holds that risk-taking will be higher in the team with the losing record.As we noted earlier, making such a prediction has implications for the alternate hypothesis and theway the results are interpreted. Table 5.2 shows DARE scores for members of the two teams.
Table 5.2: DARE scores
|
Record |
Scores |
|
Winning |
37, 51, 57, 60, 66, 62, 68, 69 |
|
Losing |
34, 37, 44, 47, 51, 54, 57, 59 |
Because the researcher predicts that risk-taking will be higher in the team with the losing record, ifthat team represents the μ2 population, the hypotheses for this problem are:
· H0: µ1 = µ2
· HA: µ1 < µ2
The steps to completing the independent t test follow:
1. Calculate the means (M) and standard deviations (s) for each group.
2. Calculate the standard error of the mean (SEM) for each group.
3. Calculate the standard error of the difference, (SEM1−M2).
4. Calculate the value of t.
5. Compare the value of t to the critical value of t.
For the independent t test, degrees of freedom are n1 + n2 − 2, or the number in Group 1 plus thenumber in Group 2 minus 2. This calculation allows the degrees of freedom for the entire problem toreflect the degrees of freedom in each of the component groups.
The results follow. Note that the subscripts to M, s, and SEM indicate the group that the statisticrepresents: the team with the winning record (1) or the team with the losing record (2).
Hemera/ Thinkstock
If a sports psychologisthypothesizes that winning andlosing athletes havesignificantly different levels ofrisk-taking, what is thealternate hypothesis?
1. First, we calculate the mean and standard deviation of eachsample:
M1 = 58.75, s1 = 10.634,
M2 = 47.875, and s2 = 9.109.
2. Then using the standard deviation for each sample, wecalculate the associated standard error of the mean:
SEM1=s1/n1‾‾‾√=10.634/8‾‾√=3.760
SEM2=s2/n2‾‾‾√=9.109/8‾‾√=3.221
3. Knowing the standard error of the mean for each groupallows us to estimate the standard error of the difference:
SEM1−M2=(SEM1)2+(SEM2)2‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√
=3.7602+3.2212‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√
=4.951
4. With that value, we can calculate the independent t:
t=M1−M2SEM1−M2=58.75−47.8754.951=2.197
determine the number of degrees of freedom:
df = 8 + 8 − 2 = 14
and compare it to the critical value of t with 14 degrees of freedom:
t0.05(14) = 1.761
The critical value of t is from the portion of the table for one-tailed tests because the prediction wasthat the winning team would have a significantly lower (rather than just a different) level of risk-takingthan the losing team. The calculated t value exceeds the table value. Was the result statisticallysignificant?
Try It!: #6
Because the risk-taking problem produceda result that was not in the tail predictedand so was not statistically significant, is itappropriate to just run the test again as atwo-tailed test?
The researcher hypothesized that risk-takingwould be lower among the more successfulplayers (Group 1). A negative value of t wouldhave reflected that result, but as it turned out, thewinning group had the higher risk-taking scores.With a one-tailed test and the stated hypothesis,the entire 5% rejection region was in thenegative tail of the t distribution.
No positive value of t, no matter how extreme,can be significant when a one-tailed test is basedon the prediction of a negative t. This problemillustrates the gamble researchers take with one-tailed tests. If the result is in the direction opposite the predicted one, the magnitude of the value isunimportant. The proper statistical decision is to fail to reject the null hypothesis.
Apply It! The Science of Athletic Performance
A. Green/Corbis
Sports science is a discipline that studies theapplication of scientific principles and techniqueswith the aim of improving sporting performance. Asports scientist is interested in determining if acertain training method will change the 100-yard-dash times of sprinters. Does training with aparachute on a runner’s back (a small one, openedup behind the runner, creating a drag) affect sprinttimes? The obvious conclusion would be that itimproves times, but the scientist is concerned thatperhaps it alters the way sprinters run in a way thatleads to poorer running mechanics and slower times.
The sports scientist tests his theory on high schoolsprinters. No population data are available, so the scientist employs an independent, two-tailed t test, with p = 0.05. He uses a two-tailed test because his interest is only determining if thistraining has an effect. Whether the parachute can lower or raise the average sprint times isunknown.
Sixty high school sprinters are chosen at random from throughout the country. During the trackseason, 30 sprinters are randomly selected from the original 60 to incorporate parachutes intotheir normal training (Group 1). The remaining 30 (Group 2) do not. The null hypothesispredicts that the mean of the population represented by the first group is the same as the meanof the population of the second group:
H0: µ1 = µ2
At the end of the season, the researcher records the final 100-yard sprint times for each of the30 runners in the two groups. The means, standard deviations, and standard errors of the meanfor the two groups are listed in Table 5.3.
Table 5.3: Sprinter data for parachute experiment
|
|
Group 1 |
Group 2 |
|
Mean, M |
12.07 seconds |
12.21 seconds |
|
Standard deviations, s |
0.97 seconds |
0.94 seconds |
|
Standard errors of themean, SE |
s1/n1‾‾‾√=0.97/30‾‾‾√=0.177 |
s2/n2‾‾‾√=0.94/30‾‾‾√=0.172 |
The two samples are independent, and participants in each group were randomly selected. Asthe standard errors of the mean reflect, the two samples have similar variability. This similarityis important because one of the assumptions for an independent t test is that both groups havesimilar variability. The measure of pooled variance is calculated first.
SEM1−M2=(SEM1)2+(SEM2)2‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√
=0.1772+0.1722‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√
=0.247
t=M1−M2SEM1−M2
12.07−12.210.247
=−0.567
The critical value for t0.05(29) = 2.045 (two-tailed test).
In this instance, the difference is not statistically significant. The results indicate that trainingwith parachutes had no significant effect on the 100-yard-dash times of high school sprinters.Note that because it was a two-tailed test (i.e., was there an effect?), the sign of t is incidental.
Apply It! boxes written by Shawn Murphy
Another Independent t Test Example
A sociologist wants to know whether optimism about the future differs between those who areemployed and those who are unemployed. Data are collected using the Upward Scientific Differential(the Upside) and are listed in Table 5.4.
Table 5.4: Upside scores for unemployed and employed
|
Employment status |
Scores |
|
Unemployed |
5, 7, 7, 8, 11, 14, 15 |
|
Employed |
7, 10, 12, 15, 15, 16, 17 |
1. First, we calculate the mean and standard deviation of each sample:
M1 = 9.571, s1 = 3.823
M2 = 13.143, s2 = 3.625
2. Then, using the standard deviation for each sample, we calculate the associated standard error ofthe mean:
SEM1=s1/n1‾‾‾√=3.823/7‾‾√=1.445SEM2=s2/n2‾‾‾√=3.625/7‾‾√=1.370
3. Knowing the standard error of the mean for each group allows us to estimate the standard errorof the difference:
SEM1−M2=(SEM1)2+(SEM2)2‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√=1.4452+1.3702‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√=1.991
4. Using that value, we can calculate the independent t:
t=M1−M2SEM1−M2=9.571−13.1431.991=−1.794
determine the number of degrees of freedom:
df = 7 + 7 − 2 = 12
and compare it to the critical value of t with 12 degrees of freedom:
t0.05(12) = 2.179 (two-tailed test)
Try It!: #7
Because the question in the Upsideexample concerns whether optimism aboutthe future differs between the two groups,is it a one-tailed or a two-tailed test?
A calculated t value of −1.794 is not a largeenough difference to be statistically significant.Once again, the proper decision is to fail toreject. In the distribution of difference scores onwhich the independent t test is based, this muchdifference between the two groups could occurbecause of sampling error, as the distribution ofdifferences in Figure 5.4 suggests. Note that thearea for which we would reject the nullhypothesis is outside the area from −2.179 to+2.179. Both groups probably belong to thesame optimism population.
Figure 5.4: A nonsignificant difference
Standard Error of the Difference for Unequal Samples
Formula 5.4, to determine the standard error of the difference, is based on the assumption that the twosamples have the same numbers of subjects. The fact that the formula treats the two standard errors ofthe mean values the same bears this out. When the sample sizes differ, the t test statistic looks thesame, but the standard error of the difference has to accommodate the unequal sample sizes. Formula5.4 must be adjusted for n values that are different in each group:
Formula 5.5
SEM1−M2=[(n1−1)s21+(n2−1)s22(n1+n2−2)][1n1+1n2]‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾⎷
Although this formula looks complicated, it simply involves the number of scores in each group andthe square of the standard deviation (or variance) for each group. Note that the adjustment for unequalsample sizes is made in both the first and second bracketed terms. Be careful with the order ofmathematical operations; for example, remember to calculate within the parentheses first.
The steps for completing SEM1 −M2 for unequal sample sizes follow:
1. Determine each group’s variance (s2). As the notation suggests, the variance is the square of thestandard deviation. If your calculator lacks a standard deviation function and you determine thevariance longhand, delete the final step, calculating the square root; the result is the variance. Ifyour calculator does have a standard deviation function, square the result to find the variance. Tocalculate the variance in Excel, use the VAR command.
2. Take the Group 1 number minus 1 (n1 − 1) and multiply the result by the variance for Group 1(s21).
3. Repeat Step 2 for Group 2.
4. Add the results of Steps 2 and 3 together, and divide their sum by the number in the two groupsminus 2 (n1 + n2 − 2). Note that this is the degrees of freedom for the problem.
5. Add the result of 1 / n1 plus 1 / n2.
6. Multiply the result of Step 4 and the result of Step 5.
7. Find the square root of the result of Step 6.
The Independent t Test with Unequal Sample Sizes
Now, return to the problem of the level of optimism among the employed versus the unemployed.Suppose that the unemployed person who scored 15 on the Upside is no longer willing to be involvedin the study and insists that his data be deleted. The remaining data, in that case, are listed in Table 5.5.
Table 5.5: New set of Upside scores
|
Status |
Scores |
|
Unemployed |
5, 7, 7, 8, 11, 14 |
|
Employed |
7, 10, 12, 15, 15, 16, 17 |
From these data, we find M1 = 8.667 and s1 = 3.266, which makes s21 = 10.667.
Group 2 stays unchanged: M2 = 13.143, s2 = 3.625, and s22 = 13.143. The standard error of thedifference, then, is
SEM1−M2=[(n1−1)s21+(n2−1)s22(n1+n2−2)][1n1+1n2]‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾⎷=[(6−1)10.667+(7−1)13.143(6+7−2)][16+17]‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾⎷=1.929andt=M1−M2SEM1−M2=6.67−13.1431.929=−2.321
The critical value for the two-tailed test is t0.05(11) = 2.201.
Other things equal, eliminating a score reduces the potential for a significant result because the criticalvalue—the value that calculated t must meet or exceed—becomes slightly higher with fewer scores. Inthis case, removing the score makes the result significant. Why? There are two reasons. First, in thisparticular problem, taking out the highest score in the lower-scoring group increases the differencebetween means; the M1 − M2 difference becomes greater. Second, eliminating the highest score fromthe group also decreases the variability within the group. Less within-group variability results in asmaller value s21, which in turn, reduces the value of the standard error of the difference, SEM1−M2. Asmaller denominator in the t ratio and a greater M1 − M2 distance create a larger value of t. Theincrease is great enough that even with a larger critical value to meet—because df = 11 rather than 12—the calculated t value exceeds the table value and is statistically significant at p = 0.05.
How much difference did using Formula 5.5 instead of 5.4 make? If SEM1−M2 had been calculatedusing the formula for equal sample sizes, what would its value have been for the following?
SEM1=s1/n1‾‾‾√=3.266/6‾‾√=1.333SEM2=s2/n2‾‾‾√=3.625/7‾‾√=1.370−—this is the same as before, of courseSEM1−M2=(SEM1)2+(SEM2)2‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√=1.3332+1.3702‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√=1.912
The adjustment is not very significant compared to the SEM1 − M2 = 1.930 from the unequal samplesizes formula, but it is certainly more accurate. The difference is more apparent as sample sizesbecome more disparate.
Assumptions Associated with the Independent t Test
Each statistical test is based on certain assumptions or conditions. Those associated with theindependent t include the following:
1. The samples are independent.
2. The participants in each group are randomly selected.
3. The measures of the dependent variable must be at least interval scale.
4. The two samples have equivalent variability.
The chapter discussed the first three assumptions earlier. The technical word for the equivalentvariability in the final assumption is homogeneity of variance. The term does not refer to exactequality of variability which, of course, will rarely be the case. Homogeneity in this case means“similar” equality. Although there are tests to indicate the point at which homogeneity is violated,Excel does not provide one. We will say that when measures of variability (R, s, or s2) are quitedifferent from one group to the other, a researcher should not assume homogeneity of variance. In thatcase, when completing a problem with Excel, the researcher makes an adjustment.
Although Conditions 1 and 3 are fairly strict requirements, the t test is robust in the face of violationsto Conditions 2 and 4.
Running an Independent t Test Using Excel
An expert is interested in analyzing the effect of a mild stimulant on participants’ recall ability. Theexpert randomly selects 20 volunteers from a larger group and randomly assigned them to groups.Those in the first group receive a placebo; those in the second group receive a mild stimulant.Members of both groups are given a complex scenario from which they are required to retrieve details.The number of details that the volunteers in each group can correctly retrieve is listed in Table 5.6:
Table 5.6: Recall ability experimental data
|
Group |
Scores |
|
Placebo |
2, 5, 2, 4, 7, 1, 2, 3, 4, 5 |
|
Stimulant |
6, 6, 7, 10, 12, 9, 6, 5, 5, 7 |
Are the retrieval differences statistically significant? We will use Excel to answer the question:
1. Begin a data set in Excel by naming the variables “placebo” (cell A1) and “stimulant” (cell B1).
2. Enter the scores in their respective columns.
3. Click the Data tab, and then click Data Analysis. Because the ranges for the scores in the twosamples are fairly similar, 5 and 7 respectively, we can safely assume equal variances.
4. Select t Test: Two-Sample Assuming Equal Variances.
5. Click OK.
6. Excel names the Group 1 data that we put in column A “variable 1.” Indicate that the range forthe data in this group is A2:A11. Indicate that the range for “variable 2,” the data for Group 2, is B2:B11.
7. Indicate that the “hypothesized mean difference” is 0 by entering 0 in the box. This reflects thenull hypothesis, µ1 = µ2.
8. Select a range for the output—say, C15—and click OK. Expand column C so that all the outputshows results in Figure 5.5.
Figure 5.5: The independent t test in Excel
Source: Microsoft Excel. Used with permission from Microsoft.
The output provides descriptive statistics, including the means and variances for each group. Ratherthan indicate whether the result is or is not statistically significant, Excel calculates the probability thatthis value of the calculated value of t could have occurred by chance in the same population. Theprogram produces results for both one-tailed and two-tailed tests.
Whenever p = 0.05 or less, the result is statistically significant. The lower the p value in the result, theless likely it is that the difference between the groups occurred by chance. The data indicate that boththe one-tailed and two-tailed tests have p values lower than 0.05. Either test would produce astatistically significant outcome.
The Independent t test in Excel
00:00
00:00
Bottom of Form
5.5 The Confidence Interval of the Difference
When an independent t test result is statistically significant, it means that the two samplesprobably represent different populations, or, to put it another way, they belong to populationswith different values for their means (µ1 − µ2 ≠ 0). In such situations, the means of thesamples are estimates of their respective population means. The confidence interval, calledthe confidence interval of the difference for the independent t test, allows us to estimate thedifference between the means of those populations. The confidence interval of the differenceuses this formula:
Formula 5.6
CI0.95 = ±t (SEM1 − M2) + |M1 − M2|
where
CI0.95 = a 0.95 confidence interval is used when the t test is conducted at p = 0.05.It indicates that there is a 1 − 0.05 = 0.95 probability of capturing the true value ofthe difference between the means of the populations represented by the samples.
The ±t means that the table value is used twice, once as a positive value times thestandard error of the difference and a second time as a negative value.
The value of t used is from the t table (rather than the calculated t value), given the df for the problem and the p value at which the test was conducted. The t value for df = 18—the Excel problem—and p = 0.05 is 2.101.
SEM1 − M2 = the standard error of the difference, the denominator in the t ratio.
|M1 – M2|. The vertical brackets around the difference between the means indicatethe absolute (non-negative) value of the difference between the sample means.
A confidence interval for the Excel problem begins with calculating SEM1 − M2 since Exceldoes not show it with the other results:
1.
s21=3.389 and s=s2‾‾‾√, s=3.389‾‾‾‾‾‾√=1.841s22=5.344 and s=s2‾‾‾√, s=5.344‾‾‾‾‾‾√=2.312SEM1=1.841/10‾‾‾√=0.582SEM2=2.312/10‾‾‾√=0.731SEM1−M2=(SEM1)2+(SEM2)2‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√=5.822+0.7312‾‾‾‾‾‾‾‾‾‾‾‾‾‾‾√=0.934
2. Now, the confidence interval of the difference:
CI0.95 = ±t (SEM1 − M2) + |M1 − M2|
CI0.95 = +2.101(0.934) + (3.8) and −2.101(0.934) + (3.8)
CI0.95 = 5.762, 1.838
The absolute value of the M1 − M2 difference is just the difference between the means, avalue that is never negative.
The ±t (SEM1 − M2) results, added to the difference between the means, provide the range ofvalues within which the difference between the related population means will occur 95% ofthe time. Stated formally, the result is: With 0.95 confidence, the difference in the populationmeans represented by these samples is somewhere between 1.838 and 5.762.
The statistically significant t test value indicated that the two samples probably belong todistinct populations. The confidence interval of the difference estimates how much differenceis between those respective population means.
The size of the confidence interval depends upon three things:
1. the difference between the sample means
2. the critical value of t (which depends upon degrees of freedom and therefore samplesize)
3. the size of the standard error of the difference for the problem
5.6 Determining a Result’s Practical Significance
Results that are statistically significant are not always important. Remember that significanceindicates that a result is not likely to have occurred by chance, a non-random—but notnecessarily important—occurrence. How do we know if an outcome has practical relevance?Researchers calculate effect sizes to gauge how important the outcome is.
When a result is significant, calculate an effect size to indicate the outcome’s importance.One very useful effect-size calculation for the independent t test is omega-squared (ω2). Forthe recall ability (Excel) problem in Figure 5.5, omega squared is
Formula 5.7
ω2 = t2 − 1/(t2 + n1 + n2 − 1)
Where,
t = the calculated value of t from the problem −4.066
n1 = the number in the first group, 10
n2 = the number in the first group, 10
ω2 = t2 − 1 / (t2 + n1 + n2 − 1)
= (−4.066) 2 − 1 / [(−4.066)2] + 10 + 10 − 1
= 15.532 / 35.532 = 0.437
In Excel, the t test indicated that subjects’ use of a stimulant versus a placebo had astatistically significant effect on recall ability, so we know that the difference between thetwo groups is probably not random. The omega-squared value indicates how much of thedifference between the two groups’ recall ability can be attributed to the treatment.
The ω2 = 0.437 result suggests that about 44% of the variance between these two groups canbe explained by whether the subjects used a stimulant or a placebo. With nearly half of thevariance between groups explained by the treatment, researchers would judge this animportant result. Assume for the moment that this had been a one-tailed test because theresearcher asked, “Will those receiving a mild stimulant retrieve significantly more items thanthose receiving the placebo?” Further, assume that the calculated value of t = 1.734, whichwould be a statistically significant value for p = 0.05 and 18 degrees of freedom. Thecorresponding value of ω2 = 1.7342 − 1/(1,7342 − 10 + 10 − 1) = 0.091. While that result isstatistically significant (not random), we might question its practical importance. The use ofthe stimulant versus the placebo explains just 9% of the variance. The other 91% of the variance—the difference between the two groups—is unexplained.
Returning to the original Excel problem, we might wonder about the unexplained 56% of thevariance. Any study will have differences between the groups that are unrelated to thetreatment. Even randomly selected groups are unlikely to have identical levels of recallability to begin with, so some of the difference between the groups pre-existed the treatment.Omega squared tells us how much of the difference we can attribute to the treatment. In thiscase, about 44% of the difference results from the treatment, which—if the data were notcontrived to begin with—would be noteworthy. In situations where t is not significant, thereis no need to calculate an effect size, because any difference between means is attributed tosampling error rather than the treatment.
5.7 Writing Up Statistics
The independent t test is often used in social and behavioral science research. The particularlanguage researchers use varies, but the question is whether two groups that initiallyrepresented the same population differ sufficiently as a result of the independent variable thatafter the treatment, they represent different populations.
Interested in the organization of sensorimotor cortices, Longo, Long, and Haggard (2012)studied the tendency of people with a missing limb to have vivid experiences of the absentlimb and to consistently overestimate its size. With the presence or absence of the limb as theindependent variable, they used t test to compare the size of the remaining limb, a hand andfingers, for example, to the size of the missing limb. Whether the limb was absent at birth orlost sometime later, t test results showed subjects consistently overestimated the length offingers, and underestimated the distance between knuckles.
Bear and Babcock (2012), meanwhile, explored gender differences in negotiation. Usinggender as the independent variable, they studied the success of women versus men indifferent types of negotiations and found that when subjects viewed the topic as masculine,men outperformed women. When subjects viewed the negotiated topic as feminine, nosignificant differences resulted.
5.8 Statistical Tests, Variables, and Research Design
In any t test, the question that drives the analysis is whether the independent variable (IV) hasprompted sufficient change in the groups that they no longer represent a common population.In such questions, the t test is the analytical component of a research design, a formal planfor testing whether, in the case of the t test, two samples remain elements of the samepopulation. Recall from Chapter 1 that a research design is something of a “blueprint” for theresearchers to follow as they complete their study. Different research designs pursue differentkinds of research questions, but when researchers want to know whether an independentvariable has had so much impact that two samples no longer belong to the same parentpopulation, they often use a t test analysis.
The IV is the treatment used in a research problem. Because different groups receive differentlevels of the treatment, the IV defines the groups involved. In the recall-ability example, the IV was whether participants were administered the stimulant or the placebo. The dependentvariable (DV) is the measure that is the subject of analysis. In the recall example, the DV wasthe subject’s score, recall ability, which represented the number of items each participantremembered.
monkeybusinessimages/ iStock /Thinkstock
Experimental research isresearch in which membersof the groups in the studyare randomly selected andthe treatment is randomlyassigned. For example,experimental researchwould study the correlationbetween flu shots andworkdays missed byrandomly selecting a groupof participants andrandomly assigning half toreceive flu shots.
In an independent t test, the IV always involves a nominalscale variable. The DV is interval or ratio scale. Thefollowing questions reflect these requirements; each could beanswered with an independent t test:
· Do clinically depressed people (IV) intake morecaffeine (DV) than nonclinically depressed people?
· Does verbal praise (IV) prompt increased levels ofverbal interaction (DV) among seminar participants?
· Do prison psychologists and school psychologists(occupation, IV) differ in salary (DV)?
· Do those who do or do not receive flu shots (IV) misssignificantly different numbers of work days (DV)?
· Are verbal ability scores (DV) different for socialscience majors than for communications majors(major, IV)?
The way the groups are created also relates to researchdesign. Although we commonly hear references to “anexperiment,” truly experimental research involves the random selection of participants and also their randomassignment to groups. Although participants can berandomly selected, clinical depression cannot be randomlyassigned, so the first question in the list above could never bea true experiment. The second and third questions in the listcould be experimental designs if a group of participants israndomly selected and then half are randomly assigned toeither treatment. The last question, like the first, belongs todesigns that are nonexperimental because the independentvariable cannot be assigned, even if those in the two groupswere randomly selected.
Although the independent t test allows just one IV at a time, later we will encounterprocedures that allow for multiple IVs. A quasi-experimental design is one for which one ofthe IVs is assigned but another cannot be. If we are asking whether female and male studentsdiffered in recall ability after receiving either a stimulant or a placebo, the experiment wouldbe quasi-experimental. Although the stimulant can be randomly assigned, gender is beyondthe researcher’s control.
Summary and Resources
Chapter Summary
Gosset’s tests help to liberate the researcher. Using the one-sample test, researchers cancompare samples to populations without needing to determine σM, a value which is often notavailable (Objective 1). The independent t test requires none of the population values andallows the researcher to examine two independent samples for significant differences. Theindependent t is widely used in research (Objectives 2 and 7).
In hypothesis testing, two predictions are relevant to the t tests. In the case of the one-sampletest, either the sample represents the population to which it is compared, or it does not. In thecase of the independent t, either both samples belong to populations with the same mean, orthey do not. The hypotheses simplify the way we report statistical results, although the use ofone-tailed versus two-tailed tests adds another element. Recall that one-tailed tests provide analternate hypothesis that is directional. It predicts how the mean of the population representedby the first group will differ from the mean of the population represented by the second,instead of simply predicting a difference.
Be careful not to confuse the type of t test with the type of hypothesis. The one-sample t testcompares the sample to the population. The one-tailed t test makes a prediction about howthe first group differs from the second (Objectives 3 and 4).
When an independent t test result is significant, it indicates that the samples probablyrepresent populations with different means. The confidence interval of the differenceestimates the difference between the means of those two populations. As such it is an“interval estimate” to contrast with the “point estimate” that the sample means in the t testprovide (Objective 6).
Effect-size calculations (and other procedures you will encounter later in the book) keep theresearcher grounded with the reminder that a significant result only indicates that the result isnot likely to be a random occurrence. Omega-squared (ω2) estimates how much anindependent variable affects a dependent variable (Objective 5).
As Gosset’s tests liberated us from the need for parameters, Fisher’s analysis of variance(ANOVA) test is going to relieve us of the restriction to two groups. Chapter 6 discussesanalysis of variance. It is an important development in statistical analysis, but the logicinvolved is very consistent with t test.
Chapter 5 Flashcards
Key Terms
alternate hypothesis
confidence interval of the difference
critical value
distribution of difference scores
effect sizes
experimental research
homogeneity of variance
independent t test
nonexperimental research
null hypothesis
omega-squared
one-sample t test
one-tailed test
quasi-experimental research
standard error of the differences
two-tailed test
Review Questions
Answers to the odd-numbered questions are provided in Appendix A.
1. A group of clients have anxiety scores as follows:
67, 55, 88, 74, 69, 81, 72, 70
a. What is the value of SEM?
a. Is this group representative of the population of all who are treated for anxietydisorders for whom µ = 66.0?
1. Applicants to a graduate program in psychology have the following scores on theGraduate Record Exam, Quantitative portion:
375, 400, 425, 425, 490, 500, 510, 530
Do they represent the national population for whom µ = 500?
1. A second therapist working with those who have anxiety-related disorders has clientswith the following scores:
52, 58, 64, 67, 67, 69, 70, 71
Is this group significantly different from the group in Review Question 1?
1. The test statistics for the one-sample t test and the independent t test both includemeasures of within-group data variability. What are they?
1. A researcher conducts an independent t test using the 0.05 criterion for significance.
e. If H0 is rejected, what is the probability of alpha error? Of beta error?
e. What is the alternate hypothesis if the question is whether there is a significantdifference? What if the question is whether the first group is significantly lowerthan the second group?
1. A researcher is interested in whether significantly more questions are asked at a townhall meeting when the politician provides verbal reinforcement for each question. Thepolitician responds “thank you” each time someone in group 1 asks a question. Thegroup 2 subjects do not receive any response from the politician. The number ofquestions group participants asked in six 30-minute sessions are as follows:
Group 1: 13, 15, 12, 17, 14, 14
Group 2: 10, 12, 12, 11, 13, 9
f. Is the difference significant?
f. Write out the alternate hypothesis for this problem.
f. What will omega squared indicate? What is its value?
1. A researcher wants to analyze differences in patients’ attitudes about the care theyreceive when they receive a rebate on their bill. One group receives a rebate; thesecond does not.
g. What are the IV and DV?
g. The IV data in an independent t must have what scale?
g. The DV data in an independent t must have what scale?
1. What is the relationship between sample size and the likelihood of a significantfinding?
1. A social worker randomly selects 20 cases from among the unemployed and assignshalf to each of two groups. The first group is a control group and receives no specialtreatment. Those in the second group receive a weekly newsletter about jobopportunities. During the study, one of the people in group 2 disappears. The dataindicate the number of weeks unemployed. The question is whether group 2 people areunemployed for less time than the group 1 people:
Group 1: 10, 12, 14, 14, 15, 15, 15, 16, 16, 18
Group 2: 6, 9, 9, 10, 12, 12, 12, 13, 14
i. Is the test one-tailed or two-tailed?
i. What is the alternate hypothesis?
i. Is t significant? If so, with 0.95 confidence, what is the interval between themeans of the populations represented by these samples?
1. Twenty prison inmates who volunteer to participate in a study are given scholasticaptitude tests and found to have M = 44.552, s = 6.577. A group of 20 non-offendingadults randomly selected from a local mall is given the same test and has M = 49.979and s = 5.101. What is a 0.95 confidence interval of the difference between the meansof the populations represented by these two groups?
Answers to Try It! Questions
1. The critical values for t are highest in distributions with the fewest degrees of freedom.This adjusts for the fact that small distributions are less likely to emulate thecharacteristics of a normal population.
2. If degrees of freedom for a one-sample t test are 14, n = 15, because df = n −1.
3. As long as the question is whether the two groups are significantly different, the test istwo-tailed.
4. The alternate hypothesis would have been H0: µ1 ≠ µ2 and the decision would havebeen to fail to reject H0. With a critical value of 2.262 in the two-tailed test, thedifference would not have been statistically significant.
5. Yes, it would have changed the decision. Because the original problem was a one-tailed test (remember that the nurses’ assistants were predicted to have morecompassion than health professionals generally), there is no rejection region in theother tail of the distribution. The entire rejection region is in the positive tail.
6. The answer is no. Once an analysis has been completed it is unethical to, in effect,change the hypotheses after the fact. It is entirely appropriate, however, to gather newdata, frame new hypotheses, and repeat the study. In fact this type of replication is veryimportant to developing a body of research.
7. It is a two-tailed test. The word “differs” allows for a difference in either direction.