Psychology 302

profileHong666
Exampleofaone-sampleZ-test.pdf

Example of a one-sample Z-test

The previous lecture in Powerpoint explained the general procedure for conducting a hypothesis

test. Now let’s go through an example of a hypothesis test to see how it actually works.

Say you’ve developed a new drug that you think might influence people’s cognitive ability,

although you aren’t really sure what it will do. (Yes, this is a really bad study!) You give the

drug to a sample of N = 49 people and then test their IQ. Say their IQ turns out to be X = 106.

Based on decades of test development, we know that in the general U.S. population, the mean IQ

is  = 100 with a standard deviation of  = 15. So the research question that we want to address

is whether the people who take the “IQ Pill” will have IQ scores that are different from the

general population (suggesting that the pill had some effect), or whether they basically look the

same as everybody else in the country (suggesting that the pill didn’t do anything).

Let’s test this question by using the formal steps of hypothesis testing:

1.Generate H0 and HA

2.Select statistical procedure

3.Select 

4.Calculate observed statistic for your data

5.Determine critical statistic

6.Compare (4) and (5)

7.If (4) exceeds (5), reject H0

8.Otherwise, fail to reject H0

1. Generate H0 and HA

So what are H0 and HA? We don’t know whether the drug will make people’s IQs get higher or

lower, so we need to use a 2-tailed or non-directional hypothesis test. Since the mean in the

general population is 100, we will test whether the mean in our test group is equal to 100.

H0:  = 100 HA:   100

In this case,  refers to the mean of the population that our treatment group represents. This is a

theoretical population of people who have taken our IQ pill. Obviously, this population doesn’t

exist in reality, because the only people who have ever taken our pill are the 49 people who were

in our study. But IN THEORY, a whole lot of other people could also take this drug, and IN

THEORY, they should respond to it in the same way that our 49 participants respond. So we are

making an inference from our 49 participants to that very large group of people who in theory

could also have taken this drug.

2.Select statistical procedure

At this point, I’m just going to tell you that the statistical procedure that you will use is

something called the Z-test. This is a test that is appropriate for testing whether one sample is

different from some specified value when the value of the population standard deviation () is

known. Over the next few days, we will be introduced to some other statistical procedures that

are appropriate for different kinds of situations than this, and you will need to choose which one

is the best to use.

3.Select 

In many cases, I will just tell you what  level to use. But you should understand why a

researcher might choose one  level compared to another. Remember that  tells us the

probability of making a Type I error, or rejecting a null that should not have been rejected. Most

of the time, researchers in the social sciences use an  of .05, meaning that there is a 5% chance

of making a Type I error. If this type of mistake has particularly bad consequences in your study

(e.g., telling people that a very expensive medication is effective when in fact it is not), then you

may want to choose a more conservative  level, such as .01. For this example, let’s be boring

and go with  = .05.

4.Calculate observed statistic for your data

Now we need to compute what is called a test statistic for our data. For a z-test, the procedure is

to compare the mean of our sample to the mean that we would have expected if the null

hypotheses was true, divided by the standard error of the mean. The general formula looks like

this:

Zobt = 0

X

X

N

 where 0 is the mean in the null hypothesis

Using the numbers in our example, we get:

106 100 6 2.80

15 2.14 49

obt Z

   

5.Determine critical statistic

Of course, we don’t know whether our test is significant until we have something to compare our

obtained z to. So we need to compute the critical value of z. Then, if our observed z exceeds the

critical value, we will be able to reject the null. How do we get zcrit? Remember, we wanted to

have  = .05, and our test was non-directional (because we didn’t know whether the pill would

make people more intelligent or less intelligent). So we need to find the z-score that will put 5

percent of the z-distribution in the tails beyond that score. That z-score is  1.96.

.025 .025

-1.96 +1.96

6.Compare (4) and (5)

7.If (4) exceeds (5), reject H0

8.Otherwise, fail to reject H0

OK, so our zobt was 2.80 and our zcrit was  1.96. So what do we conclude? 2.80 is larger than

+1.96, so we can reject the null. Looking at the picture again, you can see that 2.80 is in the tail

beyond 1.96, so it is in the region of rejection.

Thus, our conclusion is that we can reject H0. But it is not enough to stop there. Read any

psychology journal and you will find that they NEVER actually say “we rejected the null

hypothesis.” Rejecting the null isn’t that interesting. What is interesting is what you can

conclude based on your rejected null. In this case, our conclusion is that there is evidence that

the “IQ Pill” that you developed has an effect on people’s IQ scores.

Would it be appropriate to say that there is evidence that people who took the IQ pill are

smarter? The group mean was 106, which is obviously higher than the national average of 100,

right?

This is a matter of some disagreement among statisticians. Technically, our null hypothesis was

set up so that we would have also been able to reject the null if people had scored six points

LOWER on the IQ test. Say the sample had scored 94 on the IQ test instead of 106. Then we

would get:

94 100 6 2.80

15 2.14 49

obt Z

     

-2.80 is beyond –1.96, so we would also be able to reject the null in this case. So with a non-

directional or two-tailed null hypotheses, you would be willing to reject the null hypothesis in

either direction. So technically you should not interpret the direction of the effect, just that there

was a difference.

However, in the real world, most researchers DO interpret the direction of the effect even if they

used a two-tailed hypothesis. But if you really wanted to predict a particular direction, you

should start things off a little bit differently…

.025 .025

-1.96 +1.96 2.80

Using a directional hypothesis test

What if we actually did think from the very beginning that taking this drug would make people

smarter? In that case, we would have set up a directional null hypothesis that would test the

theory that people who take the pill will have IQ scores that are higher than 100. Our hypotheses

would look like this:

H0:   100 HA:   100

Notice that the null hypothesis is that the mean is “less than or equal to” 100. This is because we

would only reject the null if people’s scores were GREATER than 100. If it turned out that

people who took the pill got scores of exactly 100 or even lower than 100, then clearly the drug

is not having the effect we expected it to, and we would not be able to reject the null. We will

only be able to reject the null if it turns out that people’s IQ scores after taking the drug are

GREATER than 100.

How else will this change our hypothesis test? Well, now we will only have a region of rejection

in the upper tail of the distribution, rather than having regions of rejection in both tails of the

distribution. That means that if we want to keep a = .05, we will end up putting all of that .05

probability in the upper tail. What is the z-score that puts .05 in tail beyond it? That value is

actually not on your table, but it’s right between two values that are on the table. The exact z-

score that will put .05 in the tail beyond is z = 1.645. So zcrit for  = .05, one-tailed, is 1.645.

Using a one-tailed test does not change the value of zobt, so that will still be 2.80. So can we

reject the null? Sure. 2.80 is greater than 1.645, so we can reject the null. This time, we are able

to conclude that taking the IQ pill is making people smarter.

Example of a non-significant result

What if the IQ pill had a very small effect in our sample? Say that we set up the 2-tailed, non-

directional null hypothesis that we used for the first example, except that this time the mean of

the sample was only X = 98. How will this change our test? Well, we’re back to the original hypotheses with:

H0:  = 100 HA:   100 and zcrit = 1.96.

But this time our zobt will also be affected:

98 100 2 .93

15 2.14 49

obt Z

     

What can we conclude in this case? Our zobt does not exceed zcrit , so we cannot reject the null.

The conclusion in this case is that there is no evidence that the IQ pill has an effect on people’s

IQ scores.

Notice that this is not the same as saying that the IQ pill does not affect people’s IQs. It is still

possible that it does affect IQ scores, but this particular study did not provide sufficient evidence

to say that it does. It’s like when OJ Simpson was acquitted of murdering his wife. We don’t

know for sure that he was innocent of the crime, but we do know that the evidence that was

presented at the trial was not sufficient to convince the jury beyond a reasonable doubt that he

actually committed the crime. There was not enough evidence to prove he was guilty, so it was

concluded that he was not guilty. See how this is not the same thing as proving that he was

innocent. Maybe OJ should have taken one of the IQ pills that you developed!

One-sample t-test

In the example with the IQ pill, we used a z-test to determine whether a group mean is

statistically different from a known population parameter.

With the IQ test, we knew that the standard deviation of the population was  = 15, because we

know that IQ tests have been developed over many years. We were able to use the z-test because

 was known.

But most of the time, we are testing things for which we do not know the standard deviation of

the population. If we don’t know the population standard deviation, we must estimate it using a

sample. Remember the formula for the standard deviation when we are estimating a population

value by using a sample:

2 ( )

1 X

X X s

N

 

 The symbol for this is a lower-case s, to indicate that it’s an estimate.

We had to divide by N-1 in the formula instead of N because if we didn’t then our estimate

would be biased. If you use the formula where you just divide by N instead of dividing my N-1,

on average the standard deviation that you compute will tend to underestimate the true

population value. To adjust for this bias, we divide by N-1. Doing this will make the standard

deviation estimate a little bigger, adjusting for the bias.

Notice that the bigger your N is, the less it matters that you have to divide by N-1. The

difference between 5 and 5 - 1 = 4 is pretty big. But the difference between 50000 and 50000 - 1

= 49999 is practically nothing at all. This is consistent with the idea that bigger samples will

give us more accurate estimates of the true population standard deviation.

Now if we want to do a null hypothesis test, the formal steps for doing it will be the same as

before, except that the formula for the obtained test statistic will use the sx instead of x. Also,

we will call this a t-statistic rather than a z-statistic. More on that in a minute:

tobt = 0 X

X

s

N



Notice that in the formula on the right side, the standard error of the estimate is still just X s

N .

We don’t need to divide by N-1 for this part, because we already did that when we computed sx.

The null hypothesis test will proceed in the same way that we did it before, except that now we

can no longer use the z-table to look up probabilities. When we use a t-statistic instead of a z-

statistic, we must look up probabilities in a t-table, rather than in a z-table.

Huh? Well, it turns out that the sampling distribution of the mean is not quite normal when you

use an estimated standard deviation, particularly when your estimate is based on a small sample.

The probabilities of the z-distribution (e.g., that .05 is beyond 1.645 in the upper tail) are not

quite accurate when you are using an estimated standard deviation.

Instead of looking up zcrit using the z-table, we need to look up our critical value using the t-

distribution. The t-distribution is pretty similar to the z-distribution except that it is a little bit

flatter in the middle, and has more area in the tails.

Also, the t-distribution differs depending on how big your sample is. More specifically, it varies

depending on how many degrees of freedom you have. For the t-test of a mean, degrees of

freedom is N-1. Why? Because the t-test uses an estimated standard deviation. And when you

estimate a standard deviation, you have to divide by N-1.

2 ( )

1 X

X X s

N

 

When we look up tcrit using the t-table, we will use a different line on the table depending on how

many degrees of freedom we have. Let’s try an example:

A researcher wanted to know whether smoking cigarettes reduces olfactory sensitivity (makes

your sense of smell worse). On a test of olfactory sensitivity, the mean is known to be 18 where

higher scores mean better sensitivity, so the researcher wants to see whether people who smoke

have olfactory sensitivity scores that are lower than 18. The researcher collects data from a

sample of 30 smokers and finds that they have a mean score of X =17.2 and a standard deviation of sx =1.52. Let’s go through the formal steps of hypothesis testing:

1.Generate H0 and HA

This is a one-tailed test where we think that scores will be lower, so the hypotheses are:

H0:   18 HA:   18

2.Select statistical procedure

this time, we need to use a one-sample t-test

3.Select 

We’ll be boring again and go with  = .05.

4.Calculate observed Z or t for your data

17.2 18.0 .8 2.88

1.52 .27 30

obt t

     

5.Determine critical Z or t

Now we need to break out the t-table. How many degrees of freedom do we have? 30 – 1 = 29.

Since this is a one-tailed test, we look in the t-table column for the one-tailed test. And since our

alternate hypothesis has a “less than” symbol, that means that we will need to use a negative

value for tcrit. With df = 29 and  = .05, tcrit will be -1.699.

6.Compare (4) and (5)

7.If Zobt or tobt exceeds Zcrit or tcrit, reject H0

8.Otherwise, fail to reject H0

Our obtained t of –2.88 is farther in the tail of the distribution than our critical t of –1.699, so we

can reject the null. Our conclusion is that smoking does reduce olfactory sensitivity.