U2D2-64 - Analyze, Discuss and specify the term Confidence Intervals....see details

profileCrdrAble34
Chapter2-BASICSTATISTICSSAMPLINGERRORANDCONFIDENCEINTERVALS.docx

Chapter 2 - BASIC STATISTICS, SAMPLING ERROR, AND CONFIDENCE INTERVALS

2.1 Introduction

The first few chapters of a typical introductory statistics book present simple methods for summarizing information about the distribution of scores on a single variable. It is assumed that readers understand that information about the distribution of scores for a quantitative variable, such as heart rate, can be summarized in the form of a frequency distribution table or a histogram and that readers are familiar with concepts such as central tendency and dispersion of scores. This chapter reviews the formulas for summary statistics that are most often used to describe central tendency and dispersion of scores in batches of data (including the mean, M, and standard deviation, s). These formulas provide instructions that can be used for by-hand computation of statistics such as the sample mean, M. A few numerical examples are provided to remind readers how these computations are done. The goal of this chapter is to lead students to think about the formula for each statistic (such as the sample mean, M). A thoughtful evaluation of each equation makes it clear what information each statistic is based on, the range of possible values for the statistic, and the patterns in the data that lead to large versus small values of the statistic.

Each statistic provides an answer to some question about the data. The sample mean, M, is one way to answer the question, What is a typical score value? It is instructive to try to imagine these questions from the point of view of the people who originally developed the statistical formulas and to recognize why they used the arithmetic operations that they did. For example, summing scores for all participants in a sample is a way of summarizing or combining information from all participants. Dividing a sum of scores by N corrects for the impact of sample size on the magnitude of this sum.

The notation used in this book is summarized in Table 2.1 . For example, the mean of scores in a sample batch of data is denoted by M. The (usually unknown) mean of the population that the researcher wants to estimate or make inferences about, using the sample value of M, is denoted by μ (Greek letter mu).

One of the greatest conceptual challenges for students who are taking a first course in statistics arises when the discussion moves beyond the behavior of single X scores and begins to consider how sample statistics (such as M) vary across different batches of data that are randomly sampled from the same population. On first passing through the material, students are often so preoccupied with the mechanics of computation that they lose sight of the questions about the data that the statistics are used to answer. This chapter discusses each formula as something more than just a recipe for computation; each formula can be understood as a meaningful sentence. The formula for a sample statistic (such as the sample mean, M) tells us what information in the data is taken into account when the sample statistic is calculated. Thinking about the formula and asking what will happen if the values of X increase in size or in number make it possible for students to answer questions such as the following: Under what circumstances (i.e., for what patterns in the data) will the value of this statistic be a large or a small number? What does it mean when the value of the statistic is large or when its value is small?

The basic research questions in this chapter will be illustrated by using a set of scores on heart rate (HR); these are contained in the file hr130.sav. For a variable such as HR, how can we describe a typical HR? We can answer this question by looking at measures of central tendency such as mean or median HR. How much does HR vary across persons? We can assess this by computing a variance and standard deviation for the HR scores in this small sample. How can we evaluate whether an individual person has an HR that is relatively high or low compared with other people’s HRs? When scores are normally distributed, we can answer questions about the location of an individual score relative to a distribution of scores by calculating a z score to provide a unit-free measure of distance of the individual HR score from the mean HR and using a table of the standard normal distribution to find areas under the normal distribution that correspond to distances from the mean. These areas can be interpreted as proportions and used to answer questions such as, Approximately what proportion of people in the sample had HR scores higher than a specific value such as 84?

Table 2.1 Notation for Sample Statistics and Population Parameters

a. The first notation listed for each sample statistic is the notation most commonly used in this book.

We will consider the issues that must be taken into account when we use the sample mean, M, for a small random sample to estimate the population mean, μ, for a larger population. In introductory statistics courses, students are introduced to the concept of sampling error , that is, variation in values of the sample mean, M, across different batches of data that are randomly sampled from the same population. Because of sampling error, the sample mean, M, for a single sample is not likely to be exactly correct as an estimate of μ, the unknown population mean. When researchers report a sample mean, M, it is important to include information about the magnitude of sampling error; this can be done by setting up a confidence interval (CI) . This chapter reviews the concepts that are involved in setting up and interpreting CIs.

2.2 Research Example: Description of a Sample of HR Scores

In the following discussion, the population of interest consists of 130 persons; each person has a score on HR, reported in beats per minute (bpm). Scores for this hypothetical population are contained in the data file hr130.sav. Shoemaker (1996) generated these hypothetical data so that sample statistics such as the sample mean, M, would correspond to the outcomes from an empirical study reported by Mackowiak, Wasserman, and Levine (1992). For the moment, it is useful to treat this set of 130 scores as the population of interest and to draw one small random sample (consisting of N = 9 cases) from this population. This will provide us with a way to evaluate how accurately a mean based on a random sample of N = 9 cases estimates the mean of the population from which the sample was selected. (In this case, we can easily find the actual population mean, μ, because we have HR data for the entire population of 130 persons.) IBM SPSS® Version 19 is used for examples in this book. SPSS has a procedure that allows the data analyst to select a random sample of cases from a data file; the data analyst can specify either the percentage of cases to be included in the sample (e.g., 10% of the cases in the file) or the number of cases (N) for the sample. In the following exercise, a random sample of N = 9 HR scores was selected from the population of 130 cases in the SPSS file hr130.sav.

Figure 2.1 shows the Data View for the SPSS worksheet for the hr130.sav file. Each row in this worksheet corresponds to scores for one participant. Each column in the SPSS worksheet corresponds to one variable. The first column gives each person’s HR in beats per minute (bpm).

Clicking on the tab near the bottom left corner of the worksheet shown in Figure 2.1 changes to the Variable View of the SPSS dataset, displayed in Figure 2.2 . In this view, the names of variables are listed in the first column. Other cells provide information about the nature of each variable—for example, variable type. In this dataset, HR is a numerical variable, and the variable type is “scale” (i.e., quantitative or approximately interval/ratio) level of measurement. HR is conventionally reported in whole numbers; the choice of “0” in the decimal points column for this variable instructs SPSS to include no digits after the decimal point when displaying scores for this variable.

Readers who have never used SPSS will find a brief introduction to SPSS in the appendix to this chapter; they may also want to consult an introductory user’s guide for SPSS, such as George and Mallery (2010).

Figure 2.1 The SPSS Data View for the First 23 Lines From the SPSS Data File hr130.sav

Figure 2.2 The Variable View for the SPSS Worksheet for hr130.sav

Prior to selection of a random sample, let’s look at the distribution of this population of 130 scores. A histogram can be generated for this set of scores by starting in the Data View worksheet, selecting the <Graphs> menu from the menu bar along the top of the SPSS Data View worksheet, and then selecting <Legacy Dialogs> and <Histogram> from the pull-down menus, as shown in Figure 2.3 .

Figure 2.4 shows the SPSS dialog window for the Histogram procedure. Initially, the names of all the variables in the file (in this example, there is only one variable, HR) appear in the left-hand panel, which shows the available variables. To designate HR as the variable for the histogram, highlight it with the cursor and click on the right-pointing arrow to move the variable name HR into the small window on the right-hand side under the heading Variable. (Notice that the variable named HR has a “ruler” icon associated with it. This ruler icon indicates that scores on this variable are scale [i.e., quantitative or interval/ratio] level of measurement.) To request a superimposed normal curve, click the check box for Display normal curve. Finally, to run the procedure, click the OK button in the upper right-hand corner of the Histogram dialog window. The output from this procedure appears in Figure 2.5 , along with the values for the population mean μ = 73.76 and population standard deviation σ = 7.06 for the entire population of 130 scores.

Figure 2.3 SPSS Menu Selections <Graphs> → <Legacy Dialogs> → <Histogram> to Open the Histogram Dialog Window

NOTE: IBM SPSS Version 19 was used for all examples in this book.

To select a random sample of size N = 9 from the entire population of 130 scores in the SPSS dataset hr130.sav, make the following menu selections, starting from the SPSS Data View worksheet, as shown in Figure 2.6 : <Data> → <Select Cases>. This opens the SPSS dialog window for Select Cases, which appears in Figure 2.7 . In the Select Cases dialog window, click the radio button for Random sample of cases. Then, click the Sample button; this opens the Select Cases: Random Sample dialog window in Figure 2.8 . Within this box under the heading Sample Size, click the radio button that corresponds to the word “Exactly” and enter in the desired sample size (9) and the number of cases in the entire population (130). The resulting SPSS command is, “Randomly select exactly 9 cases from the first 130 cases.” Click the Continue button to return to the main Select Cases dialog window. To save this random sample of N = 9 HR scores into a separate, smaller file, click on the radio button for “Copy selected cases to a new dataset” and provide a name for the dataset that will contain the new sample of nine cases—in this instance, hr9.sav. Then, click the OK button.

Figure 2.4 SPSS Histogram Dialog Window

Figure 2.5 Output: Histogram for the Entire Population of Heart Rate (HR) Scores in hr130.sav

Figure 2.6 SPSS Menu Selection for <Data> → <Select Cases>

Figure 2.7 SPSS Dialog Window for Select Cases

Figure 2.8 SPSS Dialog Window for Select Cases: Random Sample

When this was done, a random sample of nine cases was obtained; these nine HR scores appear in the first column of Table 2.2 . (The computation of the values in the second and third columns in Table 2.2 will be explained in later sections of this chapter.) Of course, if you give the same series of commands, you will obtain a different subset of nine scores as the random sample.

The next few sections show how to compute descriptive statistics for this sample of nine scores: the sample mean, M; the sample variance, s2; and the sample standard deviation, s. The last part of the chapter shows how this descriptive information about the sample can be used to help evaluate whether an individual HR score is relatively high or low, relative to other scores in the sample, and how to set up a CI estimate for μ using the information from the sample.

Table 2.2 Summary Statistics for Random Sample of N = 9 Heart Rate (HR) Scores

NOTES: Sample mean for HR: M = ∑ X/N = 658/9 = 73.11. Sample variance for HR: s2 = SS/(N − 1) = 244.89/8 = 30.61. Sample standard deviation for HR:

2.3 Sample Mean (M)

A sample mean provides information about the size of a “typical” score in a sample. The interpretation of a sample mean, M, can be worded in several different ways. A sample mean, M, corresponds to the center of a distribution of scores in a sample. It provides us with one kind of information about the size of a typical X score. Scores in a sample can be represented as X1, X2, …, Xn, where N is the number of observations or participants and Xi is the score for participant number i. For example, the HR score for a person with the SPSS case record number 2 in Figure 2.1 could be given as X2 = 69. Some textbooks, particularly those that offer more mathematical or advanced treatments of statistics, include subscripts on X scores; in this book, the i subscript is used only when omitting subscripts would create ambiguity about which scores are included in a computation. The sample mean, M, is obtained by summing all the X scores in a sample of N scores and dividing by N, the number of scores:

Adding the scores is a way of summarizing information across all participants. The size of ∑X depends on two things: the magnitudes of the individual X scores and N, the number of scores. If N is held constant and all X scores are positive, ∑X increases if the values of individual X scores are increased. Assuming all X scores are positive, ∑X also increases as N gets larger. To obtain a sample mean that represents the size of a typical score and that is independent of N, we have to correct for sample size by dividing ∑X by N, to yield M, our sample mean. Equation 2.1 is more than just instructions for computation. It is also a statement or “sentence” that tells us the following:

 

1. What information is the sample statistic M based on? It is based on the sum of the Xs and the N of cases in the sample.

2. Under what circumstances will the statistic (M) turn out to have a large or small value? M is large when the individual X scores are large and positive. Because we divide by N when computing M to correct for sample size, the magnitude of M is independent of N.

In this chapter, we explore what happens when we use a sample mean, M, based on a random sample of N = 9 cases to estimate the population mean μ (in this case, the entire set of 130 HR scores in the file hr130.sav is the population of interest). The sample of N = 9 randomly selected HR scores appears in the first column of Table 2.2 . For the set of the N = 9 HR scores shown in Table 2.2 , we can calculate the mean by hand:

(Note that the values of sample statistics are usually reported up to two decimal places unless the original X scores provide information that is accurate up to more than two decimal places.)

The SPSS Descriptive Statistics: Frequencies procedure was used to obtain the sample mean and other simple descriptive statistics for the set of scores in the file hr9.sav. On the Data View worksheet, find the Analyze option in the menu bar at the top of the worksheet and click on it. Select Descriptive Statistics from the pull-down menu that appears (as shown in Figure 2.9 ); this leads to another drop-down menu. Because we want to see a distribution of frequencies and also obtain simple descriptive statistics such as the sample mean, M, click on the Frequencies procedure from this second pull-down menu.

This series of menu selections displayed in Figure 2.9 opens the SPSS dialog window for the Descriptive Statistics: Frequencies procedure shown in Figure 2.10 . Move the variable name HR from the left-hand panel into the right-hand panel under the heading Variables to indicate that the Frequencies procedure will be performed on scores for the variable HR. Clicking the Statistics button at the bottom of the SPSS Frequencies dialog window opens up the Frequencies: Statistics dialog window; this contains a menu of basic descriptive statistics for quantitative variables (see Figure 2.11 ). Check box selections can be used to include or omit any of the statistics on this menu. In this example, the following sample statistics were selected: Under the heading Central Tendency, Mean and Sum were selected, and under the heading Dispersion, Standard deviation and Variance were selected. Click Continue to return to the main Frequencies dialog window. When all the desired menu selections have been made, click the OK button to run the analysis for the selected variable, HR. The results from this analysis appear in Figure 2.12 . The top panel of Figure 2.12 reports the requested summary statistics, and the bottom panel reports the table of frequencies for each score value included in the sample. The value for the sample mean that appears in the SPSS output in Figure 2.12 , M = 73.11, agrees with the numerical value obtained by the earlier calculation.

Figure 2.9 SPSS Menu Selections for the Descriptive Statistics and Frequencies Procedures Applied to the Random Sample of N = 9 Heart Rate Scores in the Dataset Named hrsample9.sav

How can this value of M = 73.11 be used? If we wanted to estimate or guess any one individual’s HR, in the absence of any other information, the best guess for any randomly selected individual member of this sample of N = 9 persons would be M = 73.11 bpm. Why do we say that the mean M is the “best” prediction for any randomly selected individual score in this sample? It is best because it is the estimate that makes the sum of the prediction errors (i.e., the XM differences) zero and minimizes the overall sum of squared prediction errors across all participants.

To see this, reexamine Table 2.2 . The second column of Table 2.2 shows the deviation of each score from the sample mean (XM), for each of the nine scores in the sample. This deviation from the mean is the prediction error that arises if M is used to estimate that person’s score; the magnitude of error is given by the difference XM, the person’s actual HR score minus the sample mean HR, M. For instance, if we use M to estimate Participant 1’s score, the prediction error for Case 1 is (70 − 73.11) = −3.11; that is, Participant 1’s actual HR score is 3.11 points below the estimated value of M = 73.11.

Figure 2.10 The SPSS Dialog Window for the Frequencies Procedure

Figure 2.11 The Frequencies: Statistics Window With Check Box Menu for Requested Descriptive Statistics

Figure 2.12 SPSS Output From Frequencies Procedure for the Sample of N = 9 Heart Rate Scores in the File hrsample9.sav Randomly Selected From the File hr130.sav

How can we summarize information about the magnitude of prediction error across persons in the sample? One approach that might initially seem reasonable is summing the XM deviations across all the persons in the sample. The sum of these deviations appears at the bottom of the second column of Table 2.2 . By definition, the sample mean, M, is the value for which the sum of the deviations across all the scores in a sample equals 0. In that sense, using M to estimate X for each person in the sample results in the smallest possible sum of prediction errors. It can be demonstrated that taking deviations of these X scores from any constant other than the sample mean, M, yields a sum of deviations that is not equal to 0. However, the fact that ∑(XM) always equals 0 for a sample of data makes this sum uninformative as summary information about dispersion of scores.

We can avoid the problem that the sum of the deviations always equals 0 in a simple manner: If we first square the prediction errors or deviations (i.e., if we square the XM value for each person, as shown in the third column of Table 2.2 ) and then sum these squared deviations, the resulting term ∑(XM)2 is a number that gets larger as the magnitudes of the deviations of individual X values from M increase.

There is a second sense in which M is the best predictor of HR for any randomly selected member of the sample. M is the value for which the sum of squared deviations (SS), ∑(XM)2, is minimized. The sample mean is the best predictor of any randomly selected person’s score because it is the estimate for which prediction errors sum to 0, and it is also the estimate that has the smallest sum of squared prediction errors. The term ordinary least squares (OLS) refers to this criterion; a statistic meets the criterion for best OLS estimator when it minimizes the sum of squared prediction errors.

This empirical demonstration 1 only shows that ∑(XM) = 0 for this particular batch of data. An empirical demonstration is not equivalent to a formal proof. Formal proofs for the claim that ∑(XM) = 0 and the claim that M is the value for which the SS, ∑(XM)2, is minimized are provided in mathematical statistics textbooks such as deGroot and Schervish (2001). The present textbook provides demonstrations rather than formal proofs.

Based on the preceding demonstration (and the proofs provided in mathematical statistics books), the mean is the best estimate for any individual score when we do not have any other information about the participant. Of course, if a researcher can obtain information about the participant’s drug use, smoking, age, gender, anxiety level, aerobic fitness, and other variables that may be predictive of HR (or that may influence HR), better estimates of an individual’s HR may be obtainable by using statistical analyses that take one or more of these predictor variables into account. Two other statistics are commonly used to describe the average or typical score in a sample: the mode and the median. The mode is simply the score value that occurs most often. This is not a very useful statistic for this small batch of sample data because each score value occurs only once; no single score value has a larger number of occurrences than other scores. The median is obtained by rank ordering the scores in the sample from lowest to highest and then counting the scores. Here is the set of nine scores from Figure 2.1 and Table 2.2 arranged in rank order:

[64, 69, 70, 71, 73, 74, 75, 80, 82]

The score that has half the scores above it and half the scores below it is the median; in this example, the median is 73. Because M is computed using ∑X, the inclusion of one or two extremely large individual X scores tends to increase the size of M. For instance, suppose that the minimum score of “64” was replaced by a much higher score of “190” in the set of nine scores above. The mean for this new set of nine scores would be given by

However, the median for this new set of nine scores with an added outlier of X = 190,

[69, 70, 71, 73, 74, 75, 80, 82, 190],

would change to 74, which is still quite close to the original median (without the outlier) of 73.

The preceding example demonstrates that the inclusion of one extremely high score typically has little effect on the size of the sample median. However, the presence of one extreme score can make a substantial difference in the size of the sample mean, M. In this sample of N = 9 scores, adding an extreme score of X = 190 raises the value of M from 73.11 to 87.11, but it changes the median by only one point. Thus, the mean is less “robust” to extreme scores or outliers than the median; that is, the value of a sample mean can be changed substantially by one or two extreme scores. It is not desirable for a sample statistic to change drastically because of the presence of one extreme score, of course. When researchers use statistics (such as the mean) that are not very robust to outliers, they need to pay attention to extreme scores when screening the data. Sometimes extreme scores are removed or recoded to avoid situations in which the data for one individual participant have a disproportionately large impact on the value of the mean (see Chapter 4 for a more detailed discussion of identification and treatment of outliers).

When scores are perfectly normally distributed, the mean, median, and mode are equal. However, when scores have nonnormal distributions (e.g., when the distribution of scores has a longer tail on the high end), these three indexes of central tendency are generally not equal. When the distribution of scores in a sample is nonnormal (or skewed), the researcher needs to consider which of these three indexes of central tendency is the most appropriate description of the center of a distribution of scores.

Despite the fact that the mean is not robust to the influence of outliers, the mean is more widely reported than the mode or median. The most extensively developed and widely used statistical methods, such as analysis of variance (ANOVA), use group means and deviations from group means as the basic building blocks for computations. ANOVA assumes that the scores on the quantitative outcome variable are normally distributed. When this assumption is satisfied, the use of the mean as a description of central tendency yields reasonable results.

2.4 Sum of Squared Deviations (SS) and Sample Variance (s2)

The question we want to answer when we compute a sample variance can be worded in several different ways. How much do scores differ among the members of a sample? How widely dispersed are the scores in a batch of data? How far do individual X scores tend to be from the sample mean M? The sample variance provides summary information about the distance of individual X scores from the mean of the sample. Let’s build the formula for the sample variance (denoted by s2) step by step.

First, we need to know the distance of each individual X score from the sample mean. To answer this question, a deviation from the mean is calculated for each score as follows (the i subscript indicates that this is done for each person in the sample—that is, for scores that correspond to person number i for i = 1, 2, 3, …, N). The deviation of person number i’s score from the sample mean is given by Equation 2.2 :

The value of this deviation for each person in the sample appears in the second column of Table 2.2 . The sign of this deviation tells us whether an individual person’s score is above M (if the deviation is positive) or below M (if the deviation is negative). The magnitude of the deviation tells us whether a score is relatively close to, or far from, the sample mean.

To obtain a numerical index of variance, we need to summarize information about distance from the mean across subjects. The most obvious approach to summarizing information across subjects would be to sum the deviations from the mean for all the scores in the sample:

As noted earlier, this sum turns out to be uninformative because, by definition, deviations from a sample mean in a batch of sample data sum to 0. We can avoid this problem by squaring the deviation for each subject and then summing the squared deviations. This SS is an important piece of information that appears in the formulas for many of the more advanced statistical analyses discussed later in this textbook:

What range of values can SS have? SS has a minimum possible value of 0; this occurs in situations where all the X scores in a sample are equal to each other and therefore also equal to M. (Because squaring a deviation must yield a positive number, and SS is a sum of squared deviations, SS cannot be a negative number.) The value of SS has no upper limit. Other factors being equal, SS tends to increase when

 

 

1. the number of squared deviations included in the sum increases, or

2. the individual XiM deviations get larger in absolute value.

A different version of the formula for SS is often given in introductory textbooks:

Equation 2.5 is a more convenient procedure for by-hand computation of the SS than is Equation 2.4 because it involves fewer arithmetic operations and results in less rounding error. This version of the formula also makes it clear that SS depends on both ∑X, the sum of the Xs, and ∑X2, the sum of the squared Xs. Formulas for more complex statistics often include these same terms: ∑X and ∑X2. When these terms (∑X and ∑X2) are included in a formula, their presence implies that the computation takes both the mean and the variance of X scores into account. These chunks of information are the essential building blocks for the computation of most of the statistics covered later in this book.

From Table 2.2 , the numerical result for SS = ∑(X – M)2 is 244.89.

How can the value of SS be used or interpreted? The minimum possible value of SS occurs when all the X scores are equal to each other and, therefore, equal to M. For example, in the set of scores [73, 73, 73, 73, 73], the SS term would equal 0. However, there is no upper limit, in practice, for the maximum value of SS. SS values tend to be larger when they are based on large numbers of deviations and when the individual X scores have large deviations from the mean, M. To interpret SS as information about variability, we need to correct for the fact that SS tends to be larger when the number of squared deviations included in the sum is large.

2.5 Degrees of Freedom (df) for a Sample Variance

It might seem logical to divide SS by N to correct for the fact that the size of SS gets larger as N increases. However, the computation (SS/N) produces a sample variance that is a biased estimate of the population variance; that is, the sample statistic SS/N tends to be smaller than σ2, the true population variance . This can be empirically demonstrated by taking hundreds of small samples from a population, computing a value of s2 for each sample by using the formula s2 = SS/N, and tabulating the obtained values of s2. When this experiment is performed, the average of the sample s2 values turns out to be smaller than the population variance, σ2. 2 This is called bias in the size of s2; s2 calculated as SS/N is smaller on average than σ2, and thus, it systematically underestimates σ2. SS/N is a biased estimate because the SS term is actually based on fewer than N independent pieces of information. How many independent pieces of information is the SS term actually based on?

Let’s reconsider the batch of HR scores for N = 9 people and the corresponding deviations from the mean; these deviations appear in column 2 of Table 2.2 . As mentioned earlier, for this batch of data, the sum of deviations from the sample mean equals 0; that is, ∑(XiM) = −3.11 − 2.11 + .89 + 6.89 − .11 + 1.89 + 8.89 − 9.11 − 4.11 = 0. In general, the sum of deviations of sample scores from the sample mean, ∑(Xi – M), always equals 0. Because of the constraint that ∑(XM) = 0, only the first N − 1 values (in this case, 8) of the XM deviation terms are “free to vary.” Once we know any eight deviations for this batch of data, we can deduce what the remaining ninth deviation must be; it has to be whatever value is needed to make ∑(XM) = 0. For example, once we know that the sum of the deviations from the mean for Persons 1 through 8 in this sample of nine HR scores is +4.11, we know that the deviation from the mean for the last remaining case must be −4.11. Therefore, we really have only N − 1 (in this case, 8) independent pieces of information about variability in our sample of 9 subjects. The last deviation does not provide new information. The number of independent pieces of information that a statistic is based on is called the degrees of freedom , or df . For a sample variance for a set of N scores, df = N − 1. The SS term is based on only N − 1 independent deviations from the sample mean.

It can be demonstrated empirically and proved formally that computing the sample variance by dividing the SS term by N results in a sample variance that systematically underestimates the true population variance. This underestimation or bias can be corrected by using the degrees of freedom as the divisor. The preferred (unbiased) formula for computation of a sample variance for a set of X scores is thus

Whenever a sample statistic is calculated using sums of squared deviations, it has an associated degrees of freedom that tells us how many independent deviations the statistic is based on. These df terms are used to compute statistics such as the sample variance and, later, to decide which distribution (in the family of t distributions , for example) should be used to look up critical values for statistical significance tests.

For this hypothetical batch of nine HR scores, the deviations from the mean appear in column 2 of Table 2.2 ; the squared deviations appear in column 3 of Table 2.2 ; the SS is 244.89; df = N − 1 = 8; and the sample variance, s2, is 244.89/8 = 30.61. This agrees with the value of the sample variance in the SPSS output from the Frequencies procedure in Figure 2.12 .

It is useful to think about situations that would make the sample variance s2 take on larger or smaller values. The smallest possible value of s2 occurs when all the scores in the sample have the same value; for example, the set of scores [73, 73, 73, 73, 73, 73, 73, 73, 73] would have a variance s2 = 0. The value of s2 would be larger for a sample in which individual deviations from the sample mean are relatively large, for example, [44, 52, 66, 97, 101, 119, 120, 135, 151], than for the set of scores [72, 73, 72, 71, 71, 74, 70, 73], where individual deviations from the mean are relatively small.

The value of the sample variance, s2, has a minimum of 0. There is, in practice, no fixed upper limit for values of s2; they increase as the distances between individual scores and the sample mean increase. The sample variance s2 = 30.61 is in “squared HR in beats per minute.” We will want to have information about dispersion that is in terms of HR (rather than HR squared); this next step in the development of sample statistics is discussed in Section 2.7. First, however, let’s consider an important question: Why is there variance? Why do researchers want to know about variance?

2.6 Why Is There Variance?

The best question ever asked by a student in my statistics class was, “Why is there variance?” This seemingly naive question is actually quite profound; it gets to the heart of research questions in behavioral, educational, medical, and social science research. The general question of why is there variance can be asked specifically about HR: Why do some people have higher and some people lower HR scores than average? Many factors may influence HR—for example, family history of cardiovascular disease, gender, smoking, anxiety, caffeine intake, and aerobic fitness. The initial question that we consider when we compute a variance for our sample scores is, How much variability of HR is there across the people in our study? In subsequent analyses, researchers try to account for at least some of this variability by noting that factors such as gender, smoking, anxiety, and caffeine use may be systematically related to and therefore predictive of HR. In other words, the question of why is there variance in HR can be partially answered by noting that people have varying exposure to all sorts of factors that may raise or lower HR, such as aerobic fitness, smoking, anxiety, and caffeine consumption. Because people experience different genetic and environmental influences, they have different HRs. A major goal of research is to try to identify the factors that predict (or possibly even causally influence) each individual person’s score on the variable of interest, such as HR.

Similar questions can be asked about all attributes that vary across people or other subjects of study; for example, Why do people have differing levels of anxiety, satisfaction with life, body weight, or salary?

The implicit model that underlies many of the analyses discussed later in this textbook is that an observed score can be broken down into components and that each component of the score is systematically associated with a different predictor variable. Consider Participant 7 (let’s call him Joe), with an HR of 82 bpm. If we have no information about Joe’s background, a reasonable initial guess would be that Joe’s HR is equal to the mean resting HR for the sample, M = 73.11. However, let’s assume that we know that Joe smokes cigarettes and that we know that cigarette smoking tends to increase HR by about 5 bpm. If Joe is a smoker, we might predict that his HR would be 5 points higher than the population mean of 73.11 (73.11, the overall mean, plus 5 points, the effect of smoking on HR, would yield a new estimate of 78.11 for Joe’s HR). Joe’s actual HR (82) is a little higher than this predicted value (78.11), which combines information about what is average for most people with information about the effect of smoking on HR. An estimate of HR that is based on information about only one predictor variable (in this example, smoking) probably will not be exactly correct because many other factors are likely to influence Joe’s HR (e.g., body weight, family history of cardiovascular disease, drug use). These other variables that are not included in the analysis are collectively called sources of “error.” The difference between Joe’s actual HR of 82 and his predicted HR of 78.11 (82 − 78.11 = +3.89) is a prediction error. Perhaps Joe’s HR is a little higher than we might predict based on overall average HR and Joe’s smoking status because Joe has poor aerobic fitness or was anxious when his HR was measured. It might be possible to reduce this prediction error to a smaller value if we had information about additional variables (such as aerobic fitness and anxiety) that are predictive of HR.

Because we do not know all the factors that influence or predict Joe’s HR, a predicted HR based on just a few variables is generally not exactly equal to Joe’s actual HR, although it may be a better estimate of his HR than we would have if we just used the sample mean to estimate his score.

Statistical analyses covered in later chapters will provide us with a way to “take scores apart” into components that represent how much of the HR score is associated with each predictor variable. In other words, we can “explain” why Joe’s HR of 82 is 8.89 points higher than the sample mean of 73.11 by identifying parts of Joe’s HR score that are associated with, and predictable from, specific variables such as smoking, aerobic fitness, and anxiety. More generally, a goal of statistical analysis is to show that we can predict whether individuals tend to have high or low scores on an outcome variable of interest (such as HR) from scores on a relatively small number of predictor variables. We want to explain or account for the variance in HR by showing that some components of each person’s HR score can be predicted from his or her scores on other variables.

2.7 Sample Standard Deviation (s)

An inconvenient property of the sample variance that was calculated in Section 2.5 (s2 = 30.61) is that it is given in squared HR rather than in the original units of measurement. The original scores were measures of HR in beats per minute, and it would be easier to talk about typical distances of individual scores from the mean if we had a measure of dispersion that was in the original units of measurement. To describe how far a typical subject’s HR is from the sample mean, it is helpful to convert the information about dispersion contained in the sample variance, s2, back into the original units of measurement (scores on HR rather than HR squared). To obtain an estimate of the sample standard deviation (s), we take the square root of the variance. The formula used to compute the sample standard deviation (which provides an unbiased estimate of the population standard deviation, s) is as follows:

For the set of N = 9 HR scores given above, the variance was 30.61; the sample standard deviation s is the square root of this value, 5.53. The sample standard deviation, s = 5.53, tells us something about typical distances of individual X scores from the mean, M. Note that the numerical estimate for the sample standard deviation, s, obtained from this computation agrees with the value of s reported in the SPSS output from the Frequencies procedure that appears in Figure 2.12 .

How can we use the information that we obtain from sample values of M and s? If we know that scores are normally distributed, and we have values for the sample mean and standard deviation, we can work out an approximate range that is likely to include most of the score values in the sample. Recall from Chapter 1 that in a normal distribution, about 95% of the scores lie within ±1.96 standard deviations from the mean. For a sample with M = 73.11 and s = 5.53, if we assume that HR scores are normally distributed, an estimated range that should include most of the values in the sample is obtained by finding M ± 1.96 × s. For this example, 73.11 ± (1.96 × 5.53) = 73.11 ± 10.84; this is a range from 62.27 to 83.95. These values are fairly close to the actual minimum (64) and maximum (82) for the sample. The approximation of range obtained by using M and s tends to work much better when the sample has a larger N of participants and when scores are normally distributed within the sample. What we know at this point is that the average for HR was about 73 bpm and that the range of HR in this sample was from 64 to 82 bpm. Later in the chapter, we will ask, How can we use this information from the sample (M and s) to estimate μ, the mean HR for the entire population?

However, several additional issues need to be considered before we take on the problem of making inferences about μ, the unknown population mean. These are discussed in the next few sections.

2.8 Assessment of Location of a Single X Score Relative to a Distribution of Scores

We can use the mean and standard deviation of a population, if these are known (μ and σ, respectively), or the mean and standard deviation for a sample (M and s, respectively) to evaluate the location of a single X score (relative to the other scores in a population or a sample).

First, let’s consider evaluating a single X score relative to a population for which the mean and standard deviation, μ and σ, respectively, are known. In real-life research situations, researchers rarely have this information. One clear example of a real-life situation where the values of μ and σ are known to researchers involves scores on standardized tests such as the Wechsler Adult Intelligence Scale (WAIS).

Suppose you are told that an individual person has received a score of 110 points on the WAIS. How can you interpret this score? To answer this question, you need to know several things. Does this score represent a high or a low score relative to other people who have taken the test? Is it far from the mean or close to the mean of the distribution of scores? Is it far enough above the mean to be considered “exceptional” or unusual? To evaluate the location of an individual score, you need information about the distribution of the other scores. If you have a detailed frequency table that shows exactly how many people obtained each possible score, you can work out an exact percentile rank (the percentage of test takers who got scores lower than 110) using procedures that are presented in detail in introductory statistics books. When the distribution of scores has a normal shape, a standard score or z score provides a good description of the location of that single score relative to other people’s scores without the requirement for complete information about the location of every other individual score.

In the general population, scores on the WAIS intelligence quotient (IQ) test have been scaled so that they are normally distributed with a mean μ = 100 and a standard deviation σ of 15. The first thing you might do to assess an individual score is to calculate the distance from the mean—that is, X − μ (in this example, 110 − 100 = +10 points). This result tells you that the score is above average (because the deviation has a positive sign). But it does not tell whether 10 points correspond to a large or a small distance from the mean when you consider the variability or dispersion of IQ scores in the population.

To obtain an index of distance from the mean that is “unit free” or standardized, we compute a z score; we divide the deviation from the mean (X − μ) by the standard deviation of population scores (σ) to find out the distance of the X score from the mean in number of standard deviations, as shown in Equation 2.8 :

If the z transformation is applied to every X score in a normally distributed population, the shape of the distribution of scores does not change, but the mean of the distribution is changed to 0 (because we have subtracted μ from each score), and the standard deviation is changed to 1 (because we have divided deviations from the mean by σ). Each z score now represents how far an X score is from the mean in “standard units”—that is, in terms of the number of standard deviations. The mapping of scores from a normally shaped distribution of raw scores, with a mean of 100 and a standard deviation of 15, to a standard normal distribution, with a mean of 0 and a standard deviation of 1, is illustrated in Figure 2.13 .

For a score of X = 110, z = (110 − 100)/15 = +.67. Thus, an X score of 110 IQ points corresponds to a z score of +.67, which corresponds to a distance of two thirds of a standard deviation above the population mean.

Recall from the description of the normal distribution in Chapter 1 that there is a fixed relationship between distance from the mean (given as a z score, i.e., numbers of standard deviations) and area under the normal distribution curve. We can deduce approximately what proportion or percentage of people in the population had IQ scores higher (or lower) than 110 points by (a) finding out how far a score of 110 is from the mean in standard score or z score units and (b) looking up the areas in the normal distribution that correspond to the z score distance from the mean.

Figure 2.13 Mapping of Scores From a Normal Distribution of Raw IQ Scores (With μ = 100 and σ = 15) to a Standard Normal Distribution (With μ = 0 and σ = 1)

The proportion of the area of the normal distribution that corresponds to outcomes greater than z = +.67 can be evaluated by looking up the area that corresponds to the obtained z value in the table of the standard normal distribution in Appendix A . The obtained value of z (+.67) and the corresponding areas appear in the three columns on the right-hand side of the first page of the standard normal distribution table, about eight lines from the top. Area C corresponds to the proportion of area under a normal curve that lies to the right of z = +.67; from the table, area C = .2514. Thus, about 25% of the area in the normal distribution lies above z = +.67. The areas for sections of the normal distribution are interpretable as proportions; if they are multiplied by 100, they can be interpreted as percentages. In this case, we can say that the proportion of the population that had z scores equal to or above +.67 and/or IQ scores equal to or above 110 points was .2514. Equivalently, we could say that 25.14% of the population had IQ scores equal to or above 110.

Note that the table in Appendix A can also be used to assess the proportion of cases that lie below z = +.67. The proportion of area in the lower half of the distribution (from z = –∞ to z = .00) is .50. The proportion of area that lies between z = .00 and z = +.67 is shown in column B (area = .2486) of the table. To find the total area below z = +.67, these two areas are summed: .5000 + .2486 = .7486. If this value is rounded to two decimal places and multiplied by 100 to convert the information into a percentage, it implies that about 75% of persons in the population had IQ scores below 110. This tells us that a score of 110 is above average, although it is not an extremely high score.

Consider another possible IQ score. If a person has an IQ score of 145, that person’s z score is (145 − 100)/15 = +3.00. This person scored 3 standard deviations above the mean. The proportion of the area of a normal distribution that lies above z = +3.00 is .0013. That is, only about 1 in 1,000 people have z scores greater than or equal to +3.00 (which would correspond to IQs greater than or equal to 145).

By convention, scores that fall in the most extreme 5% of a distribution are regarded as extreme, unusual, exceptional, or unlikely. (While 5% is the most common criterion for “extreme,” sometimes researchers choose to look at the most extreme 1% or .1%.) Because the most extreme 5% (combining the outcomes at both the upper and the lower extreme ends of the distribution) is so often used as a criterion for an “unusual” or “extreme” outcome, it is useful to remember that 2.5% of the area in a normal distribution lies below z = −1.96, and 2.5% of the area in a normal distribution lies above z = +1.96. When the areas in the upper and lower tails are combined, the most extreme 5% of the scores in a normal distribution correspond to z values ≤ −1.96 and ≥ +1.96. Thus, anyone whose score on a test yields a z score greater than 1.96 in absolute value might be judged “extreme” or unusual. For example, a person whose test score corresponds to a value of z that is greater than +1.96 is among the top 2.5% of all test scorers in the population.

2.9 A Shift in Level of Analysis: The Distribution of Values of M Across Many Samples From the Same Population

At this point in the discussion, we need to make a major shift in thinking. Up to this point, the discussion has examined the distributions of individual X scores in populations and in samples. We can describe the central tendency or average score by computing a mean; we describe the dispersion of individual X scores around the mean by computing a standard deviation. We now move to a different level of analysis: We will ask analogous questions about the behavior of the sample mean, M; that is, What is the average value of M across many samples, and how much does the value of M vary across samples? It may be helpful to imagine this as a sort of “thought experiment.”

In actual research situations, a researcher usually has only one sample. The researcher computes a mean and a variance for the data in that one sample, and often the researcher wants to use the mean and variance from one sample to make inferences about (or estimates of) the mean and variance of the population from which the sample was drawn.

Note, however, that the single sample mean, M, reported for a random sample of N = 9 cases from the hr130 file (M = 73.11) was not exactly equal to the population mean μ of 73.76 (in Figure 2.5 ). The difference M − μ (in this case, 73.11 − 73.76) represents an estimation error; if we used the sample mean value M = 73.11 to estimate the population mean of μ = 73.76, in this instance, our estimate will be off by 73.11 − 73.76 = −.65. It is instructive to stop and think, Why was the value of M in this one sample different from the value of μ?

It may be useful for the reader to repeat this sampling exercise. Using the <Data> → <Select Cases> → <Random> SPSS menu selections, as shown in Figures 2.6 and 2.7 earlier, each member of the class might draw a random sample of N = 9 cases from the file hr130.sav and compute the sample mean, M. If students report their values of M to the class, they will see that the value of M differs across their random samples. If the class sets up a histogram to summarize the values of M that are obtained by class members, this is a sampling distribution for M—that is, a set of different values for M that arise when many random samples of size N = 9 are selected from the same population. Why is it that no two students obtain the same answer for the value of M?

2.10 An Index of Amount of Sampling Error: The Standard Error of the Mean (σM)

Different samples drawn from the same population typically yield different values of M because of sampling error. Just by “luck of the draw,” some random samples contain one or more individuals with unusually low or high scores on HR; for those samples, the value of the sample mean, M, will be lower (or higher) than the population mean, μ. The question we want to answer is, How much do values of M, the sample mean, tend to vary across different random samples drawn from the same population, and how much do values of M tend to differ from the value of μ, the population mean that the researcher wants to estimate? It turns out that we can give a precise answer to this question. That is, we can quantify the magnitude of sampling error that arises when we take hundreds of different random samples (of the same size, N) from the same population. It is useful to have information about the magnitude of sampling error; we will need this information later in this chapter to set up CIs, and we will also use this information in later chapters to set up statistical significance tests.

The outcome for this distribution of values of M—that is, the sampling distribution of M—is predictable from the central limit theorem . A reasonable statement of this theorem is provided by Jaccard and Becker (2002):

 

Given a population [of individual X scores] with a mean of μ and a standard deviation of σ, the sampling distribution of the mean [M] has a mean of μ and a standard deviation [generally called the “[population] standard error,” σM] of and approaches a normal distribution as the sample size on which it is based, N, approaches infinity. (p. 189)

For example, an instructor using the entire dataset hr130.sav can compute the population mean μ = 73.76 and the population standard deviation σ = 7.062 for this population of 130 scores. If the instructor asks each student in the class to draw a random sample of N = 9 cases, the instructor can use the central limit theorem to predict the distribution of outcomes for M that will be obtained by class members. (This prediction will work well for large classes; e.g., in a class of 300 students, there are enough different values of the sample mean to obtain a good description of the sampling distribution; for classes smaller than 30 students, the outcomes may not match the predictions from the central limit theorem very closely.)

When hundreds of class members bring in their individual values of M, mean HR (each based on a different random sample of N = 9 cases), the instructor can confidently predict that when all these different values of M are evaluated as a set, they will be approximately normally distributed with a mean close to 73.76 bpm (the population mean) and with a standard deviation or standard error, σM, of bpm The middle 95% of the sampling distribution of M should lie within the range μ − 1.96σM and μ + 1.96σM; in this case, the instructor would predict that about 95% of the values of M obtained by class members should lie approximately within the range between 73.76 − 1.96 × 2.35 and 73.76 + 1.96 × 2.35, that is, mean HR between 69.15 and 78.37 bpm. On the other hand, about 2.5% of students are expected to obtain sample mean M values below 69.15, and about 2.5% of students are expected to obtain sample mean M values above 78.37. In other words, before the students go through all the work involved in actually drawing hundreds of samples and computing a mean M for each sample and then setting up a histogram and frequency table to summarize the values of M across the hundreds of class members, the instructor can anticipate the outcome; while the instructor cannot predict which individual students will obtain unusually high or low values of M, the instructor can make a fairly accurate prediction about the range of values of M that most students will obtain.

The fact that we can predict the outcome of this time-consuming experiment on the behavior of the sample statistic M based on the central limit theorem means that we do not, in practice, need to actually obtain hundreds of samples from the same population to estimate the magnitude of sampling error, σM. We only need to know the values of σ and N and to apply the central limit theorem to obtain fairly precise information about the typical magnitude of sampling error.

The difference between each individual student’s value of M and the population mean, μ, is attributable to sampling error. When we speak of sampling error, we do not mean that the individual student has necessarily done something wrong (although students could make mistakes while computing M from a set of scores). Rather, sampling error represents the differences between the values of M and μ that arise just by chance. When individual students carry out all the instructions for the assignment correctly, most students obtain values of M that differ from μ by relatively small amounts, and a few students obtain values of M that are quite far from μ.

Prior to this section, the statistics that have been discussed (such as the sample mean, M, and the sample standard deviation, s) have described the distribution of individual X scores. Beginning in this section, we use the population standard error of the mean, σM, to describe the variability of a sample statistic (M) across many samples. The standard error of the mean describes the variability of the distribution of values of M that would be obtained if a researcher took thousands of samples from one population, computed M for each sample, and then examined the distribution of values of M; this distribution of many different values of M is called the sampling distribution for M.

2.11 Effect of Sample Size (N) on the Magnitude of the Standard Error (σM)

When the instructor sets up a histogram of the M values for hundreds of students, the shape of this distribution is typically close to normal; the mean of the M values is close to μ, and the population mean, as well as the standard error (essentially, the standard deviation) of this distribution of M values, is close to the theoretical value given by

Refer back to Figure 2.5 to see the histogram for the entire population of 130 HR scores. Because this population of 130 observations is small, we can calculate the population mean μ = 73.76 and the population standard deviation σ = 7.062 (these statistics appeared along with the histogram in Figure 2.5 ). Suppose that each student in an extremely large class (500 class members) draws a sample of size N = 9 and computes a mean M for this sample; the values of M obtained by 500 members of the class would be normally distributed and centered at μ = 73.76, with

as shown in Figure 2.15 . When comparing the distribution of individual X scores in Figure 2.5 with the distribution of values of M based on 500 samples each with an N of 9 in Figure 2.15 , the key thing to note is that they are both centered at the same value of μ (in this case, 73.76), but the variance or dispersion of the distribution of M values is less than the variance of the individual X scores. In general, as N (the size of each sample) increases, the variance of the M values across samples decreases.

Recall that σM is computed as σ/√N. It is useful to examine this formula and to ask, Under what circumstances will σM be larger or smaller? For any fixed value of N, this equation says that as σ increases, σM also increases. In other words, when there is an increase in the variance of the original individual X scores, it is intuitively obvious that random samples are more likely to include extreme scores, and these extreme scores in the samples will produce sample values of M that are farther from μ.

For any fixed value of σ, as N increases, the value of σM will decrease. That is, as the number of cases (N) in each sample increases, the estimate of M for any individual sample tends to be closer to μ. This should seem intuitively reasonable; larger samples tend to yield sample means that are better estimates of μ—that is, values of M that tend to be closer to μ. When N − 1, σM = σ; that is, for samples of size 1, the standard error is the same as the standard deviation of the individual X scores.

Figures 2.14 through 2.17 illustrate that as the N per sample is increased, the dispersion of values of M in the sampling distributions continues to decrease in a predictable way. The numerical values of the standard errors for the histograms shown in Figures 2.14 through 2.17 are approximately equal to the theoretical values of σM computed from σ and N:

Figure 2.14 The Sampling Distribution of 500 Sample Means, Each Based on an N of 4, Drawn From the Population of 130 Heart Rate Scores in the hr130.sav Dataset

Figure 2.15 The Sampling Distribution of 500 Sample Means, Each Based on an N of 9, Drawn From the Population of 130 Heart Rate Scores in the hr130.sav Dataset

The standard error, σM, provides information about the predicted dispersion of sample means (values of M) around μ (just as σ provided information about the dispersion of individual X scores around M).

We want to know the typical magnitude of differences between M, an individual sample mean, and μ, the population mean, that we want to estimate using the value of M from a single sample. When we use M to estimate μ, the difference between these two values (M − μ) is an estimation error. Recall that σ, the standard deviation for a population of X scores, provides summary information about the distances between individual X scores and μ, the population mean. In a similar way, the standard error of the mean, σM, provides summary information about the distances between M and μ, and these distances correspond to the estimation error that arises when we use individual sample M values to try to estimate μ. We hope to make the magnitudes of estimation errors, and therefore the magnitude of σM, small. Information about the magnitudes of estimation errors helps us to evaluate how accurate or inaccurate our sample statistics are likely to be as estimates of population parameters. Information about the magnitude of sampling errors is used to set up CIs and to conduct statistical significance tests.

Because the sampling distribution of M has a normal shape (and σM is the “standard deviation” of this distribution) and we know from Chapter 1 ( Figure 1.4 ) that 95% of the area under a standard normal distribution lies between z = −1.96 and z = +1.96, we can reason that approximately 95% of the means of random samples of size N drawn from a normally distributed population of X scores, with a mean of μ and standard deviation of σ, should fall within a range given by μ = (1.96) × σM and μ + 1.96 × σM.

Figure 2.16 The Sampling Distribution of 500 Sample Means, Each Based on an N of 25, Drawn From the Population of 130 Heart Rate Scores in the hr130.sav Dataset

2.12 Sample Estimate of the Standard Error of the Mean (SEM)

The preceding section described the sampling distribution of M in situations where the value of the population standard deviation, σ, is known. In most research situations, the population mean and standard deviation are not known; instead, they are estimated by using information from the sample. We can estimate σ by using the sample value of the standard deviation; in this textbook, as in most other statistics textbooks, the sample standard deviation is denoted by s. Many journals, including those published by the American Psychological Association, use SD as the symbol for the sample standard deviations reported in journal articles.

Figure 2.17 The Sampling Distribution of 500 Sample Means, Each Based on an N of 64, Drawn From the Population of 130 Heart Rate Scores in the hr130.sav Dataset

Earlier in this chapter, we sidestepped the problem of working with populations whose characteristics are unknown by arbitrarily deciding that the set of 130 scores in the file named hr130.sav was the “population of interest.” For this dataset, the population mean, μ, and standard deviation, σ, can be obtained by having SPSS calculate these values for the entire set of 130 scores that are defined as the population of interest. However, in many real-life research problems, researchers do not have information about all the scores in the population of interest, and they do not know the population mean, μ, and standard deviation, σ. We now turn to the problem of evaluating the magnitude of prediction error in the more typical real-life situation, where a researcher has one sample of data of size N and can compute a sample mean, M, and a sample standard deviation, s, but does not know the values of the population parameters μ or σ. The researcher will want to estimate μ using the sample M from just one sample. The researcher wants to have a reasonably clear idea of the magnitude of estimation error that can be expected when the mean from one sample of size N is used to estimate μ, the mean of the corresponding population.

When σ, the population standard deviation, is not known, we cannot find the value of σM. Instead, we calculate an estimated standard error (SEM), using the sample standard deviation s to replace the unknown value of σ in the formula for the standard error of the mean, as follows (when σ is known):

When σ is unknown, we use s to estimate σ and relabel the resulting standard error to make it clear that it is now based on information about sample variability rather than population variability of scores:

The substitution of the sample statistic σ as an estimate of the population σ introduces additional sampling error. Because of this additional sampling error, we can no longer use the standard normal distribution to evaluate areas that correspond to distances from the mean. Instead, a family of distributions (called t distributions) is used to find areas that correspond to distances from the mean.

Thus, when σ is not known, we use the sample value of SEM to estimate σM, and because this substitution introduces additional sampling error, the shape of the sampling distribution changes from a normal distribution to a t distribution. When the standard deviation from a sample (s) is used to estimate σ, the sampling distribution of M has the following characteristics:

 

 

1. It is distributed as a t distribution with df = N − 1.

2. It is centered at μ.

3. The estimated standard error is

2.13 The Family of t Distributions

The family of “t” distributions is essentially a set of “modified” normal distributions, with a different t distribution for each value of df (or N). Like the standard normal distribution, a t distribution is scaled so that t values are unit free. As N and df decrease, assuming that other factors remain constant, the magnitude of sampling error increases, and the required amount of adjustment in distribution shape also increases. A t distribution (like a normal distribution) is bell shaped and symmetrical; however, as the N and df decrease, t distributions become flatter in the middle compared with a normal distribution, with thicker tails (they become platykurtic ). Thus, when we have a small df value, such as df = 3, the distance from the mean that corresponds to the middle 95% of the t distribution is larger than the corresponding distance in a normal distribution.

As the value of df increases, the shape of the t distribution becomes closer to that of a normal distribution; for df > 100, a t distribution is essentially identical to a normal distribution. Figure 2.18 shows t distributions for df values of 3, 6, and ∞. As df increases, the shape of the t distribution converges toward the normal distribution; a t distribution with df > 100 is essentially indistinguishable from a normal distribution.

Figure 2.18 Graph of the t Distribution for Three Different df Values (df = 3, 6, and Infinity, or ∞)

SOURCE: www.psychstat.missouristate.edu/introbook/sbk24.htm

For a research situation where the sample mean is based on N = 7 cases, df = N − 1 = 6. In this case, the sampling distribution of the mean would have the shape described by a t distribution with 6 df; a table for the distribution with df = 6 would be used to look up the values of t that cut off the top and bottom 2.5% of the area. The area that corresponds to the middle 95% of the t distribution with 6 df can be obtained either from the table of the t distribution in Appendix B or from the diagram in Figure 2.18 . When df = 6, 2.5% of the area in the t distribution lies below t = −2.45, 95% of the area lies between t = −2.45 and t = +2.45, and 2.5% of the area lies above t = +2.45.

2.14 Confidence Intervals

2.14.1 The General Form of a CI

When a single value of M in a sample is reported as an estimate of μ, it is called a point estimate. An interval estimate (CI) makes use of information about sampling error. A CI is reported by giving a lower limit and an upper limit for likely values of μ that correspond to some probability or level of confidence that, across many samples, the CI will include the actual population mean μ. The level of “confidence” is an arbitrarily selected probability, usually 90%, 95%, or 99%.

The computations for a CI make use of the reasoning, discussed in earlier sections, about the sampling error associated with values of M. On the basis of our knowledge about the sampling distribution of M, we can figure out a range of values around μ that will probably contain most of the sample means that would be obtained if we drew hundreds or thousands of samples from the population. SEM provides information about the typical magnitude of estimation error—that is, the typical distance between values of M and μ. Statistical theory tells us that (for values of df larger than 100) approximately 95% of obtained sample means will likely be within a range of about 1.96 SEM units on either side of μ.

μ.

When we set up a CI around an individual sample mean, M, we are essentially using some logical sleight of hand and saying that if values of M tend to be close to μ, then the unknown value of μ should be reasonably close to (most) sample values of M. However, the language used to interpret a CI is tricky. It is incorrect to say that a CI computed using data from a single sample has a 95% chance of including μ. (It either does or doesn’t.) We can say, however, that in the long run, approximately 95% of the CIs that are set up by applying these procedures to hundreds of samples from a normally distributed population with mean = μ will include the true population mean, μ, between the lower and the upper limits. (The other 5% of CIs will not contain μ.)

2.14.2 Setting Up a CI for M When σ Is Known

To set up a 95% CI to estimate the mean when σ, the population standard deviation, is known, the researcher needs to do the following:

 

 

1. Select a “level of confidence.” In the empirical example that follows, the level of confidence is set at 95%. In applications of CIs, 95% is the most commonly used level of confidence.

2. For a sample of N observations, calculate the sample statistic (such as M) that will be used to estimate the corresponding population parameter (μ).

3. Use the value of σ (the population standard deviation) and the sample size N to calculate σM.

4. When σ is known, use the standard normal distribution to look up the critical values ” of z that correspond to the middle 95% of the area in the standard normal distribution. These values can be obtained by looking at the table of the standard normal distribution in Appendix A . For a 95% level of confidence, from Appendix B , we find that the critical values of z that correspond to the middle 95% of the area are z = −1.96 and z = +1.96.

This provides the information necessary to calculate the lower and upper limits for a CI. In the equations below, LL stands for the lower limit (or boundary) of the CI, and UL stands for the upper limit (or boundary) of the CI. Because the level of confidence was set at 95%, the critical values of z, zcritical, were obtained by looking up the distance from the mean that corresponds to the middle 95% of the normal distribution. (If a 90% level of confidence is chosen, the z values that correspond to the middle 90% of the area under the normal distribution would be used.)

The lower and upper limits of a CI for a sample mean M correspond to the following:

As an example, suppose that a student researcher collects a sample of N = 25 scores on IQ for a random sample of people drawn from the population of students at Corinth College. The WAIS IQ test is known to have σ equal to 15. Suppose the student decides to set up a 95% CI. The student obtains a sample mean IQ, M, equal to 128.

The student needs to do the following:

 

1. Find the value of

2. Look up the critical values of z that correspond to the middle 95% of a standard normal distribution. From the table of the normal distribution in Appendix A , these critical values are z = –1.96 and z = +1.96.

3. Substitute the values for σM and zcritical into Equations 2.11 and 2.12 to obtain the following results:

What conclusions can the student draw about the mean IQ of the population (all students at Corinth College) from which the random sample was drawn? It would not be correct to say that “there is a 95% chance that the true population mean IQ, μ, for all Corinth College students lies between 122.12 and 133.88.” It would be correct to say that “the 95% CI around the sample mean lies between 122.12 and 133.88.” (Note that the value of 100, which corresponds to the mean, μ, for the general adult population, is not included in this 95% CI for a sample of students drawn from the population of all Corinth College students. It appears, therefore, that the population mean WAIS score for Corinth College students may be higher than the population mean IQ for the general adult population.)

To summarize, the 95% confidence level is not the probability that the true population mean, μ, lies within the CI that is based on data from one sample (μ either does lie in this interval or does not). The confidence level is better understood as a long-range prediction about the performance of CIs when these procedures for setting up CIs are followed. We expect that approximately 95% of the CIs that researchers obtain in the long run will include the true value of the population mean, μ. The other 5% of the CIs that researchers obtain using these procedures will not include μ.

2.14.3 Setting Up a CI for M When the Value of σ Is Not Known

In a typical research situation, the researcher does not know the values of μ and σ; instead, the researcher has values of M and s from just one sample of size N and wants to use this sample mean, M, to estimate μ. In Section 2.12, I explained that when σ is not known, we can use s to calculate an estimate of SEM. However, when we use SEM (rather than σM) to set up CIs, the use of SEM to estimate σM results in additional sampling error. To adjust for this additional sampling error, we use the t distribution with N − 1 degrees of freedom (rather than the normal distribution) to look up distances from the mean that correspond to the middle 95% of the area in the sampling distribution. When N is large (>100), the t distribution converges to the standard normal distribution; therefore, when samples are large (N > 100), the standard normal distribution can be used to obtain the critical values for a CI.

The formulas for the upper and lower limits of the CI when σ is not known, therefore, differ in two ways from the formulas for the CI when σ is known. First, when σ is unknown, we replace σM with SEM. Second, when σ is unknown and N < 100, we replace zcritical with tcritical, using a t distribution with N − 1 df to look up the critical values (for N ≥ 100, zcritical may be used).

For example, suppose that the researcher wants to set up a 95% CI using the sample mean data reported in an earlier section of this chapter with N = 9, M = 73.11, and s = 5.533 (sample statistics are from Figure 2.12 ). The procedure is as follows:

 

 

1. Find the value of

2. Find the tcritical values that correspond to the middle 95% of the area for a t distribution with df = N − 1 − 9 = 1 = 8. From the table of the distribution of t, using 8 df, in Appendix B , these are tcritical = −2.31 and tcritical = +2.31.

3. Substitute the values of M, tcritical, and SEM into the following equations:

Lower limit = M − [tcritical × SEM] = 73.11 − [2.31 × 1.844] = 73.11 − 4.26 = 68.85;

Upper limit = M + [tcritical × SEM] = 73.11 + [2.31 × 1.844] = 73.11 + 4.26 = 77.37.

What conclusions can the student draw about the mean HR of the population (all 130 cases in the file named hr130.sav) from which the random sample of N = 9 cases was drawn? The student can report that “the 95% CI for mean HR ranges from 68.85 to 77.37.” In this particular situation, we know what μ really is; the population mean HR for all 130 scores was 73.76 (from Figure 2.5 ). In this example, we know that the CI that was set up using information from the sample actually did include μ. (However, about 5% of the time, when a 95% level of confidence is used, the CI that is set up using sample data will not include μ.)

The sample mean, M, is not the only statistic that has a sampling distribution and a known standard error. The sampling distributions for many other statistics are known; thus, it is possible to identify an appropriate sampling distribution and to estimate the standard error and set up CIs for many other sample statistics, such as Pearson’s r.

2.14.4 Reporting CIs

On the basis of recommendations made by Wilkinson and Task Force on Statistical Inference (1999), the Publication Manual of the American Psychological Association (American Psychological Association [APA], 2009) states that CI information should be provided for major outcomes wherever possible. SPSS provides CI information for many, but not all, outcome statistics of interest. For some sample statistics and for effect sizes , researchers may need to calculate CIs by hand (Kline, 2004).

When we report CIs, such as a CI for a sample mean, we remind ourselves (and our readers) that the actual value of the population parameter that we are trying to estimate is generally unknown and that the values of sample statistics are influenced by sampling error. Note that it may be inappropriate to use CIs to make inferences about the means for any specific real-world population if the CIs are based on samples that are not representative of a specific, well-defined population of interest. As pointed out in Chapter 1 , the widespread use of convenience samples (rather than random samples from clearly defined populations) may lead to situations where the sample is not representative of any real-world population. It would be misleading to use sample statistics (such as the sample mean, M) to make inferences about the population mean, μ, for real-world populations if the members of the sample are not similar to, or representative of, that real-world population. At best, when researchers work with convenience samples, they can make inferences about hypothetical populations that have characteristics similar to those of the sample.

The results obtained from the analysis of a random sample of nine HR scores could be reported as follows:

 

Results

Using the SPSS random sampling procedure, a random sample of N = 9 cases was selected from the population of 130 scores in the hr130.sav data file. The scores in this sample appear in Table 2.2 . For this sample of nine cases, mean HR M = 73.11 beats per minute (bpm), with SD = 5.53 bpm. The 95% CI for the mean based on this sample had a lower limit of 68.85 and an upper limit of 77.37.

2.15 Summary

Many statistical analyses include relatively simple terms that summarize information across X scores, such as ∑X and ∑X2. It is helpful to recognize that whenever a formula includes ∑X, information about the mean of X is being taken into account; when terms involving ∑X2 are included, information about variance is included in the computations.

This chapter reviewed several basic concepts from introductory statistics:

 

 

1. The computation and interpretation of sample statistics, including the mean, variance, and standard deviation, were discussed.

2. A z score is used as a unit-free index of the distance of a single X score from the mean of a normal distribution of individual X scores. Because values of z have a fixed relationship to areas under the normal distribution curve, a z score can be used to answer questions such as, What proportion or percentage of cases have scores higher than X?

3. Sampling error arises because the value of a sample statistic such as M varies across samples when many random samples are drawn from the same population.

4. Given some assumptions (e.g., that the distribution of scores in the population of interest is normal in shape), it is possible to predict the shape, mean, and variance of the sampling distribution of M. When σ is known, the sampling distribution of M has the following known characteristics: It is normal in shape; the mean of the distribution of values of M corresponds to μ, the population mean; and the standard deviation or standard error that describes typical distances of sample mean values of M from μ is given by σ/N. When σ is not known and the researcher uses a sample standard deviation s to estimate σ, a second source of sampling error arises; we now have potential errors in estimation of σ using s as well as errors of estimation of μ using M. The magnitude of this additional sampling error depends on N, the size of the samples that are used to calculate M and s.

5. Additional sampling error arises when s is used to estimate σ. This additional sampling error requires us to refer to a different type of sampling distribution when we evaluate distances of individual M values from the center of the sampling distribution—that is, the family of t distributions (instead of the standard normal distribution).

6. The family of t distributions has a different distribution shape for each degree of freedom. As the df for the t distribution increases, the shape of the t distribution becomes closer to that of a standard normal distribution. When N (and therefore df) becomes greater than 100, the difference between the shape of the t and normal distributions becomes so small that distances from the mean can be evaluated using the normal distribution curve.

7. All these pieces of information come together in the formula for the CI. We can set up an “interval estimate” for μ based on the sample value of M and the amount of sampling error that is theoretically expected to occur.

8. Recent reporting guidelines for statistics (e.g., Wilkinson and the Task Force on Statistical Inference, 1999) recommend that CIs should be included for all important statistical outcomes in research reports wherever possible.

Appendix on SPSS

The examples in this textbook use IBM SPSS Version 19.0. Students who have never used SPSS (or programs that have similar capabilities) may need an introduction to SPSS, such as George and Mallery (2010). As with other statistical packages, students may either purchase a personal copy of the SPSS software and install it on a PC or use a version installed on their college or university computer network. When SPSS access has been established (either by installing a personal copy of SPSS on a PC or by doing whatever is necessary to access the college or university network version of SPSS), an SPSS® icon appears on the Windows desktop, or an SPSS for Windows folder can be opened by clicking on Start in the lower left corner of the computer screen and then on All Programs. When SPSS is started in this manner, the initial screen asks the user whether he or she wants to open an existing data file or type in new data.

When students want to work with existing SPSS data files, such as the SPSS data files on the website for this textbook, they can generally open these data files just by clicking on the SPSS data file; as long as the student has access to the SPSS program, SPSS data files will automatically be opened using this program. SPSS can save and read several different file formats. On the website that accompanies this textbook, each data file is available in two formats: as an SPSS system file (with a full file name of the form dataset.sav) and as an Excel file (with a file name of the form dataset.xls). Readers who use programs other than SPSS will need to use the drop-down menu that lists various “file types” to tell their program (such as SAS) to look for and open a file that is in Excel XLS format (rather than the default SAS format).

SPSS examples are presented in sufficient detail in this textbook so that students should be able to reproduce any of the analyses that are discussed. Some useful data-handling features of SPSS (such as procedures for handling missing data) are discussed in the context of statistical analyses, but this textbook does not provide a comprehensive treatment of the features in SPSS. Students who want a more comprehensive treatment of SPSS may consult books by Norusis and SPSS (2010a, 2010b). Note that the titles of recent books sometimes refer to SPSS as PASW, a name that applied only to Version 18 of SPSS.

Notes

1. Demonstrations do not constitute proofs; however, they require less lengthy explanations and less mathematical sophistication from the reader than proofs or formal mathematical derivations. Throughout this book, demonstrations are offered instead of proofs, but readers should be aware that a demonstration only shows that a result works using the specific numbers involved in the demonstration; it does not constitute a proof.

2. The population variance, σ2, is defined as σ2 = ∑(X − μ)2/N.

I have already commented that when we calculate a sample variance, s2, using the formula s2 = ∑(X – M)2/N − 1, we need to use N − 1 as the divisor to take into account the fact that we only have N − 1 independent deviations from the sample mean. However, a second problem arises when we calculate s2; that is, we calculate s2 using M, an estimate of μ that is also subject to sampling error.

Comprehension Questions

1.

Consider the following small set of scores. Each number represents the number of siblings reported by each of the N = 6 persons in the sample: X scores are [0, 1, 1, 1, 2, 7].

 

 

a.

Compute the mean (M) for this set of six scores.

 

 

b.

Compute the six deviations from the mean (XM), and list these six deviations.

 

 

c.

What is the sum of the six deviations from the mean you reported in (b)? Is this outcome a surprise?

 

 

d.

Now calculate the sum of squared deviations (SS) for this set of six scores.

 

 

e.

Compute the sample variance, s2, for this set of six scores.

 

 

f.

When you compute s2, why should you divide SS by (N − 1) rather than by N?

 

 

g.

Finally, compute the sample standard deviation (denoted by either s or SD).

2.

In your own words, what does an SS tell us about a set of data? Under what circumstances will the value of SS equal 0? Can SS ever be negative?

3.

For each of the following lists of scores, indicate whether the value of SS will be negative, 0, between 0 and +15, or greater than +15. (You do not need to actually calculate SS.) Sample A: X = [103, 156, 200, 300, 98] Sample B: X = [103, 103, 103, 103, 103, 103] Sample C: X = [101, 102, 103, 102, 101]

4.

For a variable that interests you, discuss why there is variance in scores on that variable. (In Chapter 2 , e.g., there is a discussion of factors that might create variance in heart rate, HR.)

5.

Assume that a population of thousands of people whose responses were used to develop the anxiety test had scores that were normally distributed with μ = 30 and σ = 10. What proportion of people in this population would have anxiety scores within each of the following ranges of scores?

 

a.

Below 20

 

b.

Above 30

 

c.

Between 10 and 50

 

d.

Below 10

 

e.

Below 50

 

f.

Above 50

 

g.

Either below 10 or above 50

 

     Assuming that a score in the top 5% of the distribution would be considered extremely anxious, would a person whose anxiety score was 50 be considered extremely anxious?

6.

What is a confidence interval (CI), and what information is required to set up a CI?

7.

What is a sampling distribution? What do we know about the shape and characteristics of the sampling distribution for M, the sample mean?

8.

What is SEM? What does the value of SEM tell you about the typical magnitude of sampling error?

 

a.

As s increases, how does the size of SEM change (assuming that N stays the same)?

 

b.

As N increases, how does the size of SEM change (assuming that s stays the same)?

9.

How is a t distribution similar to a standard normal distribution score? How is it different?

10.

Under what circumstances should a t distribution be used rather than the standard normal distribution to look up areas or probabilities associated with distances from the mean?

11.

Consider the following questions about CIs.     A researcher tests emotional intelligence (EI) for a random sample of children selected from a population of all students who are enrolled in a school for gifted children. The researcher wants to estimate the mean EI for the entire school. The population standard deviation, σ, for EI is not known.     Let’s suppose that a researcher wants to set up a 95% CI for IQ scores using the following information:

 

 

The sample mean M = 130.

 

 

The sample standard deviation s = 15.

 

 

The sample size N = 120.

 

 

The df = N − 1 = 119.

 

 

For the values given above, the limits of the 95% CI are as follows:

 

 

Lower limit = 130 − 1.96 × 1.37 = 127.31;

 

 

Upper limit = 130 + 1.96 × 1.37 = 132.69.

 

   The following exercises ask you to experiment to see how changing some of the values involved in computing the CI influences the width of the CI.

 

   Recalculate the CI above to see how the lower and upper limits (and the width of the CI) change as you vary the N in the sample (and leave all the other values the same).

 

a.

What are the upper and lower limits of the CI and the width of the 95% CI if all the other values remain the same (M = 130, s = 15) but you change the value of N to 16? For N = 16, lower limit = _________ and upper limit = ____________. Width (upper limit − lower limit) = ______________________. Note that when you change N, you need to change two things: the computed value of SEM and the degrees of freedom used to look up the critical values for t.

 

b.

What are the upper and lower limits of the CI and the width of the 95% CI if all the other values remain the same but you change the value of N to 25? For N = 25, lower limit = __________ and upper limit = ___________. Width (upper limit – lower limit) = _______________________.

 

c.

What are the upper and lower limits of the CI and the width of the 95% CI if all the other values remain the same (M = 130, s = 15) but you change the value of N to 49? For N = 49, lower limit = __________ and upper limit = ___________. Width (upper limit – lower limit) = ______________________.

 

d.

Based on the numbers you reported for sample size N of 16, 25, and 49, how does the width of the CI change as N (the number of cases in the sample) increases?

 

e.

What are the upper and lower limits and the width of this CI if you change the confidence level to 80% (and continue to use M = 130, s = 15, and N = 49)? For an 80% CI, lower limit = ________ and upper limit = __________. Width (upper limit – lower limit) = ______________________.

 

f.

What are the upper and lower limits and the width of the CI if you change the confidence level to 99% (continue to use M = 130, s = 15, and N = 49)? For a 99% CI, lower limit = ________ and upper limit = ___________. Width (upper limit – lower limit) = ______________________.

 

g.

How does increasing the level of confidence from 80% to 99% affect the width of the CI?

12.

Data Analysis Project:

 

   The N = 130 scores in the temphr.sav file are hypothetical data created by Shoemaker (1996) so that they yield results similar to those obtained in an actual study of temperature and HR (Mackowiak et al., 1992).

 

   Use the Temperature data in the temphr.sav file to do the following:

 

   Note that temperature in degrees Fahrenheit (tempf) can be converted into temperature in degrees centigrade (tempc) by the following: tempc = (tempf − 32)/1.8.

 

   The following analyses can be done on tempf, tempc, or both tempf and tempc.

 

a.

Find the sample mean, M; standard deviation, s; and standard error of the mean, SEM, for scores on temperature.

 

b.

Examine a histogram of scores on temperature. Is the shape of the distribution reasonably close to normal?

 

c.

Set up a 95% CI for the sample mean, using your values of M, s, and N (N = 130 in this dataset).

 

d.

The temperature that is popularly believed to be “average” or “healthy” is 98.6°F (or 37°C). Does the 95% CI based on this sample include the value 98.6, which is widely believed to represent an “average/healthy” temperature? What conclusion might you draw from this result?

(Warner 71-80)

Warner, Rebecca (Becky) (Margaret). Applied Statistics: From Bivariate Through Multivariate Techniques, 2nd Edition. SAGE Publications, Inc, 04/2012. VitalBook file.

The citation provided is a guideline. Please check each citation for accuracy before use.