psych case study & stat hw assignment
Presentation: Hypothesis Testing and Using Excel to Perform Z Tests
Before we get started, I’d like to review some concepts that you have covered so far in your course reading. First, let’s look at the difference between a population and a sample. A population is the complete set of people, animals or other objects that have something in common. A sample is a subgroup, or small part, of the population. The sample is representative of population, and researchers draw conclusions about a population based on information from a sample. When we talk about populations and samples, there are three concepts that are important to keep in mind. First, the mean for each is represented by a different symbol. Second, the standard deviation for each is also represented by a different symbol. Finally, we can talk about a distribution for a population or a distribution for a sample. Each of these distributions has certain characteristics. For this course, we usually assume that these distributions are normal distributions.
You have also read much about hypothesis testing, and the difference between the null hypothesis and the research, or alternative, hypothesis. The null hypothesis is the one that states that there is no difference between a given value and the corresponding population value itself. It is represented as H sub0. The research, or alternative, hypothesis states that there is some kind of difference between a given value and the corresponding population value itself. There are two important things to remember about the alternative hypothesis. First, it can be either directional or non-directional. A directional hypothesis results in a one-tailed hypothesis test; and a non-directional hypothesis results in a two-tailed hypothesis test. Second, starting now, it will be important to determine the alpha value associated with your hypothesis test. This is equal to the chance of making a Type I error, or rejecting the null hypothesis when it’s actually true. Alpha is always set ahead of time and is usually set to either .05, .01, or .001.
Hypotheses can be written out in words and in symbols. Here, I’m going to show you how to do both. Here is a sample research situation: A middle school counselor has noticed an increase in the number of referrals for ADHD in her school. She wonders whether this number is greater, or higher, than the number of referrals in the general population. She does a review of current research and finds that the general population of middle schools has a mean number of referrals equal to 3 per month and a standard deviation of 1. Based on this information, let’s write out our hypotheses using sentences and then using symbols.
My null hypothesis can be stated as seen here: The number of referrals at this school is equal to or less than the number of referrals in the general population. Using symbols, I can restate the null as seen here (mu is less than or equal to 3). It will help you to remember that the null hypothesis will always contain an equal sign, since it is the statement that there is no difference between your sample and population values—that, in essence, they are equal.
My research hypothesis can be stated as seen here: The number of referrals at this school is greater than the number of referrals in the general population. Using symbols, I can restate the research hypothesis as seen here (mu is greater than 3). This is an example of a directional hypothesis—that is, it predicts that the difference will be in a certain direction. This means that the we would use a one-tailed hypothesis test, seeking for a difference that is greater than the population average. Thus, we will be looking in the right-hand tail of the distribution for our critical values. It will help you to remember that the research hypothesis will never contain an equal sign, since it is the statement that there is some type of difference between your sample and population values—that, in other words, they are not equal. Now that we’ve reviewed some of this basic material, let’s move on to our special topic for this week.
It is usually the case that researchers run studies using a sample of participants instead of one individual. However, most of the work we’ve done up to this point has involved single scores only, such as one person’s number of minutes of exercise, or one person’s SAT score. In these situations, we have tested hypotheses based on a distribution of single scores. However, when we want to look at the performance of a sample compared to the population, we can not test a hypothesis based on a distribution of single scores—rather, we have to look at a distribution of sample means. What does this mean, and why is this true?
Imagine that you are trying to determine how an entire 4th grade class performed on an achievement test as compared to the national average of all 4th grade classes. As we have already learned, the best way to describe a sample of scores is to use summary statistics, such as the mean and standard deviation. So, you compute the mean and standard deviation of your sample 4th grade class. Then, you decide to see how your class did by comparing the mean class score to a distribution of individual children’s scores on the achievement test. What’s wrong with this decision? In everyday language, it’s a meaningless comparison. It would be a bit like using a yardstick to measure the width of a grain of rice. When trying to determine the chance that a certain individual score will occur, it makes sense to look at a distribution of individual scores. However, when trying to determine the chance of a certain mean score occurring that comes from a sample of more than one person, it makes sense to look at a distribution of other mean scores from samples with the same number of individuals in them.
This particular type of distribution is called the distribution of sample means. It is also commonly called the sampling distribution of sample means, or the sampling distribution of the mean. Regardless of what it’s called, however, it is important to remember that each number represented on the distribution is not the score of a single individual, as we have been working with up until now. Each number on the distribution of sample means is the mean of some sample of a given size, like samples consisting of 100 individuals each, for example.
In our example, let’s say our 4th grade class has 31 students. This means that the number or N of our sample is 31. We take the mean achievement test score for this 4th grade class, which turns out to be 89. In order to test how likely it is that this score would occur in the general population, and to test whether our class is at the national average, or is doing better or worse, we now have to compare the mean score of the 4th grade class to an appropriate distribution of sample means. This distribution is the distribution of mean scores of samples of size N = 31.
I want to address one point that often confuses people early on in this process. You might be thinking, “Surely each testing agency doesn’t publish distributions for every possible sample size, like 25 students or 138 students and every number in between. This is true. Any comparison distribution of sample means, at this point in our course, will be one big distribution of the means in the general population, like “4th grade classes”. When we begin our hypothesis test, however, the math that we use to compute certain values for this test takes into account the size of our sample as well as the population values, so you can use the same population information for samples of different sizes. We can use the same population information to ask questions about samples of size 31, like our example 4th grade class, or of size 100. We’ll get into these ideas in more detail later, but it helps to start thinking about it now.
Now that we’ve learned what the distribution of sample means is, let’s look at its characteristics. The following things are always true about any distribution of sample means: the mean of the distribution of sample means is the same as the population mean. The second characteristic explains in more detail what was mentioned in the previous slide about sample sizes. The variance of the distribution of sample means is equal to the population variance divided by the number in the sample, or N. Notice that the variance will change depending on the sample size. This is how we modify the population information to take into account a particular sample size, without constructing hundreds of separate distributions for samples of all possible sizes.
Third, the standard deviation of the distribution of sample means is called the standard error of the mean and is equal to the square root of the variance from number 2. Here you see the standard error of the mean written in two different but equivalent ways. Finally, the shape of the distribution of sample means is normal for samples that are greater than 30 or if the population is known to be normally distributed.
Hypothesis testing using sample means is carried out in much the same way as hypothesis testing for individual scores. The steps remain the same, but some changes need to be made in how we compute and compare scores. First, the comparison distribution becomes the distribution of sample means. It is this distribution that we will use to construct the null hypothesis for any research question. Second, instead of using a simple standard deviation, we have to compute the standard error of the mean, which is basically a standard deviation that takes into account the sample size. Finally, the values used for computing the test statistic will change to reflect our new type of distribution.
The test statistic we will be looking at in this presentation is useful for samples of size 30 or more, when the population information is known. It is called a Z test. Just as a Z score is useful for telling us where an individual score lies along a given distribution, the Z test is a way of doing the same thing, but for sample means instead of individual scores. The formula for calculating Z is very similar to the one you already know, and the steps of the hypothesis testing process remain virtually identical but with different values. Let’s look at how this process works.
Here we see the first two steps of the Z test. First, as before, write out your null and research hypotheses based on your research question. Remember to use evaluators such as “greater than” or “equal to” in your statements. Second, determine the characteristics of your comparison distribution. As we mentioned, the comparison distribution is now the distribution of sample means. Based on what we learned about this distribution, we know that the mean of the distribution of sample means is equal to the population mean, which is normally given. We also know that the standard deviation of the distribution of sample means is the standard error of the mean, so we compute this based on the formula we learned.
Third, we want to determine the cut-off score, or critical value, that marks the rejection region of our hypothesis test. Our final decision to reject the null hypothesis or not depends on where our sample’s mean lies in relation to this cut-off score. (Z stuff here) Fourth, we determine our sample’s score on the comparison distribution (sample’s Z stuff here). Finally, we can put all of our information together and decide whether or not to reject the null hypothesis. If our sample’s Z score lies outside of the cut-off in the expected direction, then we say that our results are unlikely if the null hypothesis were true, so we decide to reject the null hypothesis and support the research, or alternative, hypothesis. If our sample’s Z score lies within the acceptance region, then we decide not to reject the null hypothesis.
A common method used for talking about the significance of results is the p value. You will see this in almost every journal article that reports results of hypothesis tests. Your text briefly introduces you to P values, but let’s look at them a little more closely so that you’ll know what they are and what they mean. You will be required to compute them as part of your Excel homework. If you recall from an earlier presentation, the area under the normal curve is equal to one. We used z scores to compute the percentage or probability of certain values based on this property of the normal curve. If given any value along the curve, you can determine the probability of scores higher than this value, less than this value, or between this value and another value. For example, you may remember that the probability of getting a Z score higher than 0 is equal to 50%, or .5.
P values are exactly this—they represent the probability of a certain score occurring based on a given distribution, in this case the distribution of sample means. Here we see the steps for determining p values. First, determine the cut-off, or critical, p value, which the researcher sets at the beginning of the study. Most critical p values will equal .05, .01, or .001. Second, calculate your sample’s Z score in the usual way. Third, find the probability associated with your Z score using either a Normal Curve Table or a statistics computer program. If your sample result’s p value is less than the critical p value (of .05 or .01), then you can claim that your results are statistically significant. P-values will be covered further in the Excel portion of this presentation.
After you have completed the Z test and computed the p value of your result, you are ready to make a decision about whether to reject the null hypothesis or not. If your results are statistically significant, you can say the following: “Based on these results, we are able to reject the null hypothesis and support the research hypothesis that there is a true difference between the mean of our sample and the population mean.” If your results are not statistically significant, you can say: “Based on these results, there is not enough evidence to support the research hypothesis, so we retain the null hypothesis.”
We are now going to look at performing Z tests in Excel. In order to do this, we are going to use an example research scenario and data set. A family systems psychologist is interested in the amount of television that households in his town watch per day. He knows that the population mean for minutes of television watched per day per household in his state is 20. He wonders whether families in his town watch more television than the general population. He decides to interview a sample of 55 households in his town and record the number of minutes of television watched per day.
The first step is to create a results table in which we can enter the different values necessary to run the hypothesis test and its results. Here you see one such table with many different entries. We will go over these in steps throughout the rest of the presentation.
The first step of any hypothesis testing process is to state the null and alternative, or research, hypotheses. In our Excel table, we will use symbols. This will help us determine whether the test is one- or two-tailed, and if it’s one-tailed, the direction of the predicted difference. The psychologist wants to know whether the minutes of television families in his town watch is greater than the population mean, which is 75 minutes. Restating the question as a hypothesis, the research hypothesis is that families in this town watch more than 75 minutes of television per day. This is shown in the table as “mu is greater than 75.” The null hypothesis is the hypothesis that includes a statement that there is no difference between the two populations. This requires the use of the equal sign. Since our psychologist is only interested in whether the town’s mean is greater than the population mean, the null hypothesis also includes the less than symbol. We can now tell easily from looking at our symbols that our hypothesis is directional—it is only interested in one type of difference, a difference of “greater than 75.” So this tells us that our test is one-tailed. Specifically, we will be interested in the right-hand tail of the distribution, in values that are greater than the mean.
Next, we can fill in the information that we already know based on previous research and given in the research question. Using some of these values, we can compute the standard error of the mean, which is necessary because we are conducting a hypothesis test using the distribution of sample means. The easiest method to do this is to enter a formula manually, and the formula shown is the one we want to use in Excel. To input this formula, click in the desired cell and type the equal sign, then your population standard deviation (in this case, 20), then a forward slash, then the letters “SQRT” followed by the sample size in parentheses, in this case 55. “SQRT” is a function that tells Excel to compute the square root of the number that comes after it.
Hit enter, and the resulting value of about 2.697 shows up. Next, we want to calculate our sample mean and fill in our chosen alpha level. Compute the sample mean in the usual way, calculating the average of the raw sample data using the AVERAGE function. The psychologist wants to test this hypothesis at the .05 level, so we type this in the appropriate cell for the alpha level. Now it is time to compute the critical Z value, or cut-off score, that will mark our rejection region. Since our hypothesis test is one-tailed and is interested in a mean that is greater than the population mean, any Z score that is greater than our critical value will be considered statistically significant.
The Excel function that helps us compute z scores in this manner is the “normsinv” function. This function basically returns the Z score from the standard normal distribution that corresponds to a certain probability. The syntax for the function is “Normsinv” followed by the probability level in parentheses. The probability level has to do with the alpha level. With an alpha equal to .05 and a one-tailed test, there are two options depending on whether the test is left-tailed or right-tailed. If left-tailed, then the probability associated with the test is simply equal to alpha, or .05. However, if the test is right-tailed, the probability must include all of the area under the normal curve up to the last 5% in the tail. This is equal to 1 minus alpha, or .95 in this case. Since our test is right-tailed and our alpha is .05, we want to use a probability of .95 in our formula. It is important to remember this distinction between the two tails when running a hypothesis test. The procedure for a two-tailed test is similar except that you would divide alpha in half, then find the Z score associated with this probability in both tails of the distribution. However, in this section we will focus on one-tailed tests.
The resulting formula is “equals normisinv (.95)”. When we type this in to the appropriate cell, it returns a Z score of 1.64 as seen here. This is our critical value. If our sample’s Z score is greater than 1.64, we will be able to reject the null hypothesis that there is no difference between the means, and accept our research hypothesis that the families in Any Town, USA watch more television per day than families in the general population.
To compute this sample Z score, we type in the Z score formula with the values we already know. This formula is the sample mean minus the population mean divided by the standard error of the mean, just as we went over earlier in the presentation. So, filling in these values in Excel formula style, exactly as shown here, we have our sample Z score of 1.146. Of course, when filling in the formula in Excel, you can simply click on the cells containing the values instead of typing in the values yourself, as we’ve gone over before. We can now compare our two Z scores. Remember that in order to reject the null and support the research hypothesis, our sample’s Z score had to be greater than the critical Z cut-off score. Since 1.146 is not greater than 1.64, we can not reject the null hypothesis because our sample’s score does not lie in the rejection region. We can also say that our research hypothesis is not supported by the data.
Finally, we can calculate p value of our sample and compare it to the pre-determined value of p = .05. With p-values, unlike Z scores, you will always be searching for a sample p value that is less than the critical p-value, whether your test is in the right tail, the left tail, or both tails of the distribution. That is, the one-tailed hypothesis is tested at the .05 level of significance, so if our p value is less than .05, we can consider the results to be statistically significant. Well, we’ve already seen that the Z score of our sample does not support the research hypothesis, so it is pretty much guaranteed that our p value is going to be greater than .05 and not statistically significant. Let’s see if this is the case.
First, enter the critical p value in the appropriate cell as shown. We are running a one-tailed test at the .05 level, so this value is .05. Next, we want to calculate the sample’s p value to see if it is less than the critical p value. As with the Z score, the way this Excel formula is entered depends on which tail of the distribution you are interested in.
The function we want to use to compute our sample’s p value is the “normSdist” function. This function returns the area under the standard normal curve up to a certain Z score. You will recall that a couple of weeks ago you learned how to use the Normal Curve Tables to compute areas or percentages under the normal curve that correspond to certain Z scores. This function does exactly the same thing, but without a normal curve table. If your hypothesis test is left-tailed, the formula is entered as shown, ‘equals normsdist” followed by your sample’s Z score in parentheses. This will give the area under the normal curve that lies to the left of the sample’s Z score. Since our test is right-tailed, however, we are only interested in the area in the tail which is to the right of our sample’s Z score. Because the NormSDist function returns only the area to the left of the Z score, we must subtract the value from one in order to compute the area that is to the right of the Z score. This formula is entered as shown above, with an extra set of parentheses around the entire formula.
So, after entering our formula to compute the p value for a right-tailed test, we find that our sample’s p value is about .126 as shown. Our sample’s p value is greater than .05, not less than or equal to. This means that there is a low chance that our sample came from a distribution that is different from the distribution of the general population. In other words, based on this study, the psychologist can state that the results are not statistically significant and that the mean of minutes of TV watched in his town may not be truly different from the mean of the general population. The null hypothesis is retained, and the research hypothesis is not supported based on the data from this sample.
It is always important to remember that hypothesis tests are statements about the likelihood of certain statements being true or not. The results of hypothesis tests should never be presented as absolutes. Statements like “These results prove the research hypothesis” or “The null hypothesis is false” are not acceptable, because one study or even ten studies does not prove anything beyond the shadow of a doubt when dealing with statistics. Remember that a hypothesis test deals with probabilities, and there is always the possibility of a Type I or Type II error, of accepting or rejecting the wrong hypothesis when the other one is actually true. So be cautious and conscientious when presenting results, using statements like “The evidence supports the research hypothesis” or “Based on these findings, we can not reject the null hypothesis.”
Presentation: Confidence Intervals and Computing Confidence Intervals in Excel
In many instances, you may wish to use information from a sample to estimate a population mean that is unknown. This is how surveys usually work. For example, an organization that polls voters concerning their political opinions is trying to get an idea of the political opinions of the general population based on a smaller sample. This is often done with the presidential approval rating. About 1,000 people provide answers for this survey, and these results are then used to estimate how much the American public in general approves of the president. In this situation, the general approval rating in the population is unknown. The sample information helps provide a basic estimate. When the mean of a population is unknown, it can be proven that the best estimate for this population mean is the sample mean.
However, most social scientists are not comfortable with providing just one number as a point estimate of the population mean. Rather, it is often the case that researchers will report a range of values called a confidence interval. A confidence interval is a range of possible means that is likely to contain the population mean. This range is based on the scores from a sample, so, as you have probably figured out by now, confidence intervals are constructed using the distribution of sample means (which is covered in an earlier presentation). A confidence interval is also useful even if you know the population mean, to see whether or not the confidence interval based on your sample includes the actual population mean or not.
Confidence intervals are bound at the high and low ends by confidence limits. These are the numbers that mark the beginning and end of the interval. Confidence intervals can have different widths depending on how confident you want to be that your range contains the actual population mean: the wider the confidence interval, the more certain you can be that your range of means is likely to contain the population mean. A narrower interval is less likely to contain the true population mean. Why is this true?
Imagine that you are fishing over the side of a boat using a net. In particular, you’ve heard stories about a giant trout that lives in these waters, and you would like to increase your chances of catching this trout among all of the other fish in your net. You have a choice of two nets: one is somewhat large, and the other is really large. Which net is more likely to catch the giant trout among all of the other fish in your net? The larger the net, the more confident you can be that your net might contain that giant trout when you pull it out of the water. Confidence intervals work in basically the same way: the wider the interval, the more likely that it will contain the true population mean among all of the possible means in the interval.
When we construct a confidence interval using the information from a sample, we want to give an idea of the level or amount of confidence we have that our interval might contain the true population mean. In research, it is common to use numbers like 90%, 95% and 99% to talk about confidence intervals. That is, if we construct a 95% confidence interval, roughly speaking, we can say that we are 95% sure that this interval contains the true population mean. It doesn’t necessarily mean that it will. Notice that we are only 95% sure—there is still a 5% chance that our interval will not contain the population mean. Just because we have a bigger net doesn’t guarantee that we will catch the giant fish. Now, if we decided to construct a 99% confidence interval, the chances of our interval containing the population mean increase to the point that we can say we are now 99% sure that our interval contains the population mean. In other words, the 99% confidence interval is wider—it’s the really big net.
Now that we understand basically what a confidence interval is, let’s look at how to construct one. Remember that we are looking at an interval that contains a range of sample means, so we are working with the distribution of sample means. In our first example, the mean presidential approval rating from our sample of 1,000 people would be our starting point. Here you can see all of the information that you need in order to construct a confidence interval: the sample mean, the population standard deviation (which is given), and the desired confidence level. The example values we will use are mean presidential rating of 81, a sigma equal to 3, and a desired confidence level of 95%. In other words, we are constructing a 95% confidence interval.
Once we know these values, the next step is to compute the standard error of the mean, just like we did for the Z test. The formula is given again here. Our standard error of the mean based on our example is equal to 3 divided by 10, or .3. Next, we must compute the lower and upper limits of our confidence interval. Since we are constructing our confidence interval based on our sample mean, the sample mean value will be at the center of the interval. The confidence interval will range a certain distance below and above the sample mean. How far above and below depends on our confidence level. The Z scores that bound the 95% confidence interval are -1.96 and +1.96, and the Z scores that bound the 99% confidence interval are -2.57 and +2.57. To compute the 95% confidence limits, we simply multiply the positive Z score (1.96) by the standard error of the mean (.3). Then we subtract this value from the sample mean for the lower limit, and we add this value to the sample mean for the upper limit, as shown here. For the 99% confidence interval, follow the same steps, but use 2.57 instead of 1.96.
All that remains to do is to make some statement regarding your results. When talking about confidence intervals, it is appropriate to say that they represent the range of possible means likely to include the real population mean. So, after constructing a 95% confidence interval, it is appropriate to say that we can be 95% sure that this range contains the true population mean, but there is always the 5% chance that it doesn’t. The last part of this presentation will be learning how to compute confidence intervals in Excel.
Before we get there, though, I want to tie some things together for you. This week, we have looked at two approaches to answering research questions: hypothesis testing and confidence intervals. Both of these processes deal with samples and populations, using information from these to answer certain research questions. However, each process goes about this in a different way. It is helpful to understand this distinction as you move forward in statistics. When we compute a confidence interval, we first take data from a sample and then make a statement about the population. We are ESTIMATING the population mean based on the sample data. When we perform a hypothesis test (like a Z test), we first make a statement about the population (our hypothesis), and then take data from a sample and decide whether the statement seems reasonable or not. We are using our sample data to TEST our statement about the population. You can see that each of these methods approaches the process from a different end—one starts with the sample, the other starts with the population.
We will be using the same research question and data set that we looked at in the presentation for hypothesis testing and Z tests. To review the information, a family systems psychologist is interested in the amount of television that households in his town watch per day. He knows that the population standard deviation for minutes of television watched per day per household in his state is 20. He decides to interview a sample of 55 households in his town and record the number of minutes of television watched per day. He wants to use this information to construct a 95% confidence interval for mean number of minutes of television watched per day.
Here is a portion of the hypothetical sample data, as well as a table that the psychologist constructed in which to enter his results. The table contains cells for both the information that he already knows as well as the confidence limit information that he will compute in Excel.
We start off by filling in the information that is already known. The sample size is N=55 and the population standard deviation, or sigma, equals 20. Use the Excel AVERAGE function to compute the mean of the sample data, as we have done several times before. Finally, fill in your alpha level. When constructing a confidence interval, alpha is equal to 1 minus your confidence level expressed as a decimal. A 95 % confidence interval expressed as a decimal is .95, so 1 minus .95 equals .05. Put this value into the cell for alpha.
Now it is time to compute the confidence interval. We do this using the Excel function called Confidence. If you paid attention earlier when we went over computing the confidence interval by hand, you may notice that there is something missing from our table. There is no cell for the standard error of the mean or any type of Z score that is used to determine the confidence interval. This is because the Excel Confidence formula computes these values automatically from the given information. Let’s go over how to use this formula. As you can see, the formula requires that we enter alpha, the population standard deviation, and the sample size.
The easiest way to enter these values is to click on each cell in turn, typing a comma between each cell reference. So, first we click in the cell in which we want our value to appear. Type in “equals CONFIDENCE”. Excel will prompt you to enter the arguments for the formula. First, click on the cell containing alpha, then type a comma, then click on the cell containing the sigma value, another comma, and then the cell containing your sample size. Your final formula should appear as shown, with either numbers or the cell references containing the numbers in parentheses.
The resulting value appears in the cell when you press enter. This value represents your margin of error, that is, how far in each direction the confidence interval will extend from the sample mean. In order to actually determine what this interval will look like, we must compute the lower and upper limits. This is done simply by starting with the sample mean and either subtracting or adding the margin of error value to the mean. To compute the lower confidence limit, click in the cell under “lower limit” and type “equals E3-H3” and press enter. For the upper limit, simply type the same formula but change the minus to a plus.
Here we see the final table. Our 95% confidence interval for mean number of minutes of TV watched per day per household ranges from 72.8 minutes to 83.4 minutes. Our family systems psychologist can say that he can be about 95% certain that this range of means contains the true population mean. Remember that this is only an estimate, though, and that there is still a 5% risk that this range does not contain the true population mean.