Literature Review Final

profileTT24
Read-WarnerIChapters1315-16.pdf

Hayes, A. F. (2022). Introduc on to media on, modera on, and condi onal process analysis (3rd ed.). NewYork, NY: Guilford Press. CHAPTER 13 ONE-WAY BETWEEN-SUBJECTS ANALYSIS OF VARIANCE 13.1 RESEARCH SITUATIONS WHERE ONE-WAY ANOVA IS USED One-way between-subjects analysis of variance (usually called ANOVA) is used in situa ons where researchers compare means on a quan ta ve Y outcome variable across two or more groups. It is called analysis of variance because the goal is to par on or divide the variance of scores on the Y outcome into variance that can be predicted from group membership and variance that cannot be predicted from group membership. In prac ce this par on is made by compu ng sums of squares (SS). In an experiment, variance related to group membership, also called between-group variance, is related to the dosage or type of treatment received by each group. Variance within groups (variance that is not related to group membership) is called experimental error. Within-group variance is due to the influence of all other uncontrolled variables (other than the treatment in the study) on Y scores. Researchers generally hope that between-group variance or SS will be rela vely large and that within-group variance or SS will be small, because this outcome suggests that dosage or type of treatment affects the outcome variable Y. You saw an ANOVA table in the bivariate regression output in Chapter 11. However, when people refer to ANOVA, they usually refer to comparison of means across groups. The t test provides informa on about the distance between the means on a quan ta ve outcome variable for just two groups, whereas a one-way ANOVA compares means on a quan ta ve variable across any number of groups. The categorical predictor variable in an ANOVA may represent either naturally occurring groups (in a nonexperimental study) or groups formed by a researcher and then exposed to different interven ons (in an experiment). The term between-S (like the term independent samples) tells us that each par cipant is a member of one and only one group, that there are no repeated measures, and that par cipants are not matched or paired across samples. When data consist of repeated measures or paired or matched samples, repeated-measures or within-S ANOVA (discussed in a later chapter) is required. In ANOVA, the categorical predictor variable is called a factor. The groups are called levels of this factor. Levels of a factor in an experiment may represent different types of treatment and/or control groups or different dosage levels of the same treatment. In the hypothe cal research example introduced in Sec on 13.3, the factor is called “type of stress,” and the levels of this factor represent four different types of stress situa ons: type 1, no stress; type 2, cogni ve stress from a mental arithme c task; type 3, stressful social role play; and type 4, stress during a mock job interview. The outcome variable is a quan ta ve measure of anxiety. The ques on is whether mean anxiety differs across these four situa ons.

When there is just one categorical variable or factor, it is o en called Factor A; the number of levels or groups is denoted a. In Chapter 16, on factorial ANOVA, you will see that you can have more than one factor, and in those situa ons factors are o en named A, B, C, and so on. Comparisons among several group means could be made by calcula ng t tests for each pairwise comparison among the means of these four treatment groups. However, as described earlier, doing numerous significance tests leads to an inflated risk for Type I error. If a study includes k groups, there are k(k – 1)/2 pairs of means; thus, for a set of four groups, the researcher would need to do (4 × 3)/2 = 6 different t tests to make all possible pairwise comparisons. If α = .05 is used as the criterion for significance for each test, and the researcher conducts six significance tests, the probability that this set of six decisions contains at least one instance of Type I error is greater than .05. One way that ANOVA limits the risk for Type I error is by obtaining a single omnibus test that examines all possible comparisons among means in the study. Researchers o en want to examine selected pairwise comparisons of means as a follow-up analysis to obtain more informa on about the pa ern of differences among groups. 13.2 QUESTIONS IN ONE-WAY BETWEEN-S ANOVA The overall null hypothesis for one-way ANOVA is that the means of the k popula ons that correspond to the groups in the study are all equal: Other

(13.1) When each group has been exposed to different types or dosages of a treatment, as in a typical experiment, this null hypothesis corresponds to an assump on that the treatment has no effect on the outcome variable. The alterna ve hypothesis in this situa on is not that all popula on means are unequal; the alterna ve hypothesis is that there is at least one inequality between one pair of means in the set. The best ques on ever asked by a student in any of my sta s cs classes was decep vely simple: “Why is there variance?” In the hypothe cal experimental study described in the following sec on, the outcome variable is a self-report measure of anxiety, and the group membership variable is type of stress. We want to know, How much of the variance in anxiety can be predicted from type of stress? Is stress a major reason why anxiety scores differed among persons in this study? Why do some persons report more anxiety than other persons? To what extent are the differences in amount of anxiety systema cally associated with the independent variable (type of stress), and to what extent are differences in the amount of self-reported anxiety due to other factors (such as trait levels of anxiety, physiological arousal, drug use, sex, other anxiety-arousing events that each par cipant has experienced on the day of the study, etc.)?

In sta s cs, the term error usually does not mean the same thing as in everyday life. In everyday life, we use the word error to mean “mistake.” In ANOVA, the term error refers to the parts of scores that cannot be predicted from type of treatment or group membership. The part of anxiety scores that we cannot predict from type of stress is presumably due to the effects of other variables that we have not included in the study, such as personality, other upse ng events that may have happened to the person just before the study, recent use of drugs such as alcohol, caffeine, and tobacco, and possibly a mul tude of other unknown variables. Ques ons in ANOVA:

1. The first ques on in one-way ANOVA is this: When all group means are considered as a set, are there any significant differences between means? An overall F ra o will tell us whether there are any significant differences among the group means, but it does not tell us which specific means differ. It is possible that each group mean differs from every other group mean, but it is also possible that only one or a few pairs of means differ.

2. The second ques on in one-way ANOVA is this: Which specific pairs (or combina ons) of group means differ significantly? There are two ways to answer this ques on. A data analyst either decides which comparisons are of interest ahead of me and sets up planned contrasts or explores data using post hoc follow-up tests to which means differ significantly. Both approaches are discussed in this chapter.

 Planned contrasts (some mes just called contrasts and also called a priori comparisons) can be set up to examine a limited number of differences between means that the data analyst has decided ahead of me are of interest. These are called unprotected tests (that is, not protected against inflated risk for Type I error) because, except for limi ng the number of significance tests, there are no other correc ons for inflated risk for Type I error.

 Post hoc tests (such as the Tukey honestly significant difference [HSD] test) can be used to examine many or all the possible comparisons among means. These are called protected tests because most of them use more conserva ve per comparison criteria for sta s cal significance (like per comparison alpha [PCα] in the Bonferroni procedure).

Later in your study of sta s cs, you will discover that many analyses involve a similar approach: first, an omnibus test that includes all groups and/or all variables, then follow-up analyses to evaluate which groups or which variables show significant differences. 13.3 HYPOTHETICAL RESEARCH EXAMPLE Suppose that an experiment is done to compare the effects of four situa ons: Group 1 is tested in a “no-stress,” baseline situa on; Group 2 does a mental arithme c task; Group 3 does a stressful social role play; and Group 4 does a mock job interview. For this study, the X variable is a categorical variable with codes 1, 2, 3, and 4 that represent which of these four types of stress each par cipant received. This categorical X predictor variable is called a factor; in this case, the factor is called “type of stress”; the four levels of this factor correspond to no stress, mental arithme c, stressful role play, and a mock job interview. At the end of each session, the par cipants self-report their anxiety on a scale that ranges from 0 = no anxiety to 20 = extremely high anxiety. Scores on anxiety are, therefore, scores on a quan ta ve Y outcome

variable. Imagine that there is a convenience sample of N = 28 par cipants. (Capital N denotes the total number of par cipants in the study.) Imagine that par cipants were randomly assigned to one of the four levels of stress. This results in k = 4 groups with n = 7 par cipants in each group, for a total of N = 28 par cipants in the en re study. Lowercase n indicates the number of cases per group. The SPSS Data View worksheet that contains data for this imaginary study appears in Figure 13.1, and data are available in the SPSS file stress_anxiety.sav. The goal of data analysis is to find out:

1. Whether mean anxiety levels differed across these four situa ons. 2. Which situa ons elicited the highest and lowest anxiety. 3. Which treatment group means differed significantly from the baseline (no-stress)

condi on. 4. Whether mean anxiety differed among the mental arithme c, role play, and mock job

interview stress situa ons.

Figure 13.1 Data View Worksheet for Stress and Anxiety Study in stress_anxiety.sav

13.4 ASSUMPTIONS AND DATA SCREENING FOR ONE-WAY ANOVA The assump ons for one-way ANOVA are the same as those described for the independent-samples t test. The scores on the dependent variable must be quan ta ve. Observa ons must be independent of one another, both within and between groups. Ideally, scores should be approximately normally distributed within each group, and variances should be approximately equal across groups. ANOVA, like the t test, is robust against viola ons of the normality and equal variance assump ons if within-group n’s are reasonably large. Finally, there should not be extreme outliers. Preliminary screening involves the same procedures as for the t test: Histograms can be examined separately for each group to assess normality of distribu on shape; boxplots for groups can iden fy and poten al outliers within groups. The Levene test (or another test of homogeneity of variance) can be requested as part of the output and used to assess whether the homogeneity of variance assump on is violated. Because preliminary data screening for one-way between-S ANOVA uses the same procedures as those shown in Chapter 12, on the independent-samples t test, these procedures are not repeated here. 13.5 COMPUTATIONS FOR ONE-WAY BETWEEN-S ANOVA 13.5.1 Overview ANOVA begins with familiar sta s cs. For each group, we obtain M, s, SS, and n. Recall that a sum of squares or SS is obtained by finding M for the group of interest, compu ng a (Y – M) devia on for each individual Y score, squaring the devia on for each score, and summing the squared devia ons. For the independent-samples t test, we needed to find SS only for Groups 1 and 2. In ANOVA, several different forms of SS are obtained. SStotal is obtained by: Finding the grand mean for the en re data set, denoted MY. Obtaining the (Y – MY) devia on for every score in the data set. Squaring each devia on. Summing the squared devia ons. For a batch of data, recall that the sample variance s = SS/df. For SStotal, df = N – 1, where N is the total number of scores in the en re data set. We could use SStotal to find the total variance s for all Y scores in the study; however, to find out what propor on of variance in Y is related to group membership, it is more convenient to focus on SS than s. In ANOVA, an SS divided by its df is usually called a mean square (MS). A one-way ANOVA divides SStotal into two sources of variance, o en called SSbetween groups and SSwithin groups. The formulas to obtain the la er two SS terms can appear confusing, so let’s just focus on the informa on provided by each term.

SSbetween groups tells us how far the values of M1, M2,…, Mk are from the grand mean. If the group means are all exactly equal, SSbetween groups will be 0. SS terms can never be nega ve, and there is no fixed upper limit for values. A “large” value of SSbetween groups (also called SSbetween) tells us that:

 group means are far away from the grand mean, and/or  group means are far away from one another.

What informa on do we need to consider to decide whether SSbetween is “large”? First, we need to divide SS by its df. Devia ons of group means from the grand mean, like devia ons of individual scores from a sample mean, must sum to 0. If there are k group means, only the first k – 1 devia ons of group means from the grand mean are free to vary. Thus, for SSbetween, df = k – 1 (where k is the number of groups). Dividing an SS by its df corrects for the number of independent devia ons used to calculate the SS. An SS divided by its df is called a mean square. MSbetween is, in effect, the variance of the group means. For technical reasons, sta s cians do not refer to MS as a variance (but essen ally, that’s what it is). Second, we need to compare SSbetween with informa on about error variance or within-group variance. The error variance term is called SSwithin. There are several ways to compute SSwithin. The easiest way to think about it is this: First, find SS for the set of scores within each treatment group. For Treatment Group 1, find the group mean, M1; compute the devia on of each Y score in that group from M1; square the devia ons; and sum the squared devia ons. This yields SS1, and this tells us about varia on of scores within Group 1. For a study with k = 4 groups and n = 7 cases within each group, you obtain the following:

The df for MSwithin is the sum of df1, df2, df3, and df4; this can also be wri en as n1 + n2 + n3 + n4 – k, or N – k, where N is the total number of persons in the study and k is the number of groups. If SS1 (or the SS for any group) = 0, that tells us that all scores within Group 1 were equal to one another. As the value of SS1 gets larger, we have evidence that a sample of

people who received the same treatment have different score values, and these differences are due to other variables that influenced the outcome. In the hypothe cal study of stress and anxiety, anxiety scores may be influenced by recent drug use, depression, or events in the lab. A er you calculate SStotal, SSbetween, and SSwithin, you will find that this equality holds (as long as you have not made arithme c errors): Other

(13.2) This equa on describes the par on (division) of total varia on of Y into two sources of variance: differences among group means (SSbetween) and differences among scores within the same treatment groups (SSwithin). We hope that most of the varia on between groups is due to the different types or amounts of treatment received by groups, and we usually hope that SSbetween will be large. We know that SSwithin provides informa on about response differences among people who received the same type of treatment and that SSwithin tells us about magnitude of experimental error; we want SSwithin to be small. Recall that one of the effect sizes for t was η2 and that η2 was the propor on of variance of Y scores that is predictable from or related to group membership. In one-way ANOVA, η2 = SSbetween/SStotal. Thus, the SS terms provide effect size informa on. To obtain a sta s cal significance test, we set up an F ra o: Other

(13.3) Because it is a ra o of MS terms, F cannot be nega ve. F would be zero if all group means were equal. There is no fixed upper limit for values of F. To decide whether F is large enough to be sta s cally significant, we need to find a cri cal value of F from the table in Appendix C at the end of this book. The reject region for F is always one tailed (values in the top 5% of an F distribu on, for instance). To locate the cri cal value that corresponds to the top 5% of the distribu on, you need to know about df. The independent-samples t test required only one df term. An F ra o compares two different MS terms, and each of those MS terms has its own df, so we need to specify two different df terms: Other

(13.4)

Other

(13.5) where k is the number of groups and N is the total number of cases. To summarize: The by-hand computa on for one-way ANOVA (with k groups and a total of N observa ons) involves the following steps. Complete formulas are provided in the following sec ons.

1. Compute SSbetween, SSwithin, and SStotal. 2. Find effect size: η2 = SSbetween/SStotal. 3. Compute MSbetween by dividing SSbetween by its df, k – 1. 4. Compute MSwithin by dividing SSwithin by its df, N – k. 5. Compute an F ra o: MSbetween/MSwithin. 6. Compare this F value obtained with the cri cal value of F from a table of the F

distribu on with (k – 1) and (N – k) df (using the table in Appendix C at the end of the book that corresponds to the desired alpha level; for example, the first table provides cri cal values for α = .05). If the F value obtained exceeds the tabled cri cal value of F for the predetermined alpha level and the applicable degrees of freedom, reject the null hypothesis that all the popula on means are equal.

In prac ce, these computa ons are done by programs such as SPSS; you can decide whether the outcome is sta s cally significant by examining the p value for the F test and evaluate effect size by calcula ng an η2. 13.5.2 SSbetween: Informa on About Distances Among Group Means The following nota on will be used: Let k be the number of groups in the study. Let n1, n2,…, nk be the number of scores in Groups 1, 2,…, k. Let Yij be the score of subject j in Group i (i = 1, 2,…, k). Let M1, M2,…, Mk be the means of scores in Groups 1, 2,…, k. Let N be the total N in the en re study; N = n1 + n2 + … + nk. Let MY be the grand mean of all scores in the study (i.e., the total of all the individual scores, divided by N, the total number of scores). Once we have calculated the means of each individual group (M1, M2,…, Mk) and the grand mean MY, we can summarize informa on about the distances of the group means, Mj, from the grand mean, MY, by compu ng SSbetween as follows:

Other

(13.6) For the hypothe cal data in Figure 13.1, the mean anxiety scores for Groups 1 through 4 were as follows: M1 = 9.86, M2 = 14.29, M3 = 13.57, and M4 = 17.00. The grand mean on anxiety, MY, is 13.68. Each group had n = 7 scores. Therefore, for this study, Other

SSbetween ≈ 182 (this agrees with the value of SSbetween in the SPSS output presented in Figure 13.8 except for a small amount of rounding error). 13.5.3 SSwithin: Informa on About Variability of Scores Within Groups To summarize informa on about the variability of scores within each group, we compute MSwithin. For each group, for groups numbered i = 1, 2,…, k, we first find the sum of squared devia ons of scores rela ve to each group mean, SSi. The SS for scores within Group i is found by taking this sum: Other

(13.7) That is, for each of the k groups, find the devia on of each individual score from the group mean; square and sum these devia ons for all the scores in the group. These within-group SS terms for Groups 1, 2,…, k are summed across the k groups to obtain the total SSwithin: Other

(13.8) For this data set, we can find the SS term for Group 1 (for example) by taking the sum of the squared devia ons of each individual score in Group 1 from the mean of Group 1, M1. The values are shown for by-hand computa ons; it can be instruc ve to do this as a spreadsheet, entering the value of the group mean for each par cipant as a new

variable and compu ng the devia on of each score from its group mean and the squared devia on for each par cipant. Other

For the four groups of scores in the data set in Figure 13.1, these are the values of SS for each group: SS1 = 26.86, SS2 = 27.43, SS3 = 41.71, and SS4 = 26.00. Thus, the total value of SSwithin for this set of data is SSwithin = SS1 + SS2 + SS3 + SS4 = 26.86 + 27.43 + 41.71 + 26.00 = 122.00. 13.5.4 SStotal: Informa on About Total Variance in Y Scores We can also find SStotal; this involves taking the devia on of every individual score from the grand mean, squaring each devia on, and summing the squared devia ons across all scores and all groups: Other

(13.9) The grand mean MY = 13.68. The SStotal term includes 28 squared devia ons, one for each par cipant in the data set, as follows: Other

As noted earlier, the SSbetween and SSwithin terms will sum to SStotal: Other

(13.10) For these data, SStotal = 304, SSbetween = 182, and SSwithin = 122, so the sum of SSbetween and SSwithin equals SStotal (because of rounding error, these values differ slightly from the values that appear in the SPSS output in Sec on 13.13). 13.5.5 Conver ng Each SS to a Mean Square and Se ng Up an F Ra o An F ra o is a ra o of two mean squares. A mean square is the ra o of a sum of squares to its degrees of freedom, MS = SS/df. Note that the formula for a sample variance is

also SS/df. MS terms in ANOVA are similar to variances, but they are not called variances for technical reasons. The df terms for the two MS terms in a one-way between-S ANOVA are based on k, the number of groups, and N, the total number of scores in the en re study (where N = n1 + n2 + ··· + nk). The between-group SS was obtained by summing the devia ons of each of the k group means from the grand mean; only the first k – 1 of these devia ons are free to vary, so the between-groups df = k – 1, where k is the number of groups. Other

(13.11) In ANOVA, the mean square between groups is calculated by dividing SSbetween by its degrees of freedom: Other

(13.12) For the data in the hypothe cal study of stress and anxiety, SSbetween = 182, d etween = 4 – 1 = 3, and MSbetween = 182/3 = 60.7. The df for each SS within-group term is given by n – 1, where n is the number of par cipants in each group. Thus, in this example, SS1 had n – 1 or df = 6. When we form SSwithin, we add up SS1 + SS2 + ··· + SSk. There are (n – 1) df associated with each SS term, and there are k groups, so the total dfwithin = k × (n – 1). This can also be wri en as Other

(13.13) where N is the total number of scores (n1 + n2 + ··· + nk) and k is the number of groups. We obtain MSwithin by dividing SSwithin by its corresponding df: Other

(13.14) For the hypothe cal stress and anxiety data in Figure 13.1, MSwithin = 122/24 = 5.083. Finally, we can set up a test sta s c for the null hypothesis H0: μ1 = μ2 = ··· = μk by taking the ra o of MSbetween to MSwithin:

Other

(13.15)

Figure 13.2 Reject Region for F Distribu on With 3 and 24 df Using α = .05 For the stress and anxiety data, F = 60.702/5.083 = 11.94. This F ra o is evaluated using the F distribu on with (k – 1) and (N – k) df. For this data set, k = 4 and N = 28, so df values for the F ra o are 3 and 24. An F distribu on has a shape that differs from the normal or t distribu on. Because an F is a ra o of two mean squares and MS cannot be less than 0, the minimum possible value of F is 0. On the other hand, there is no fixed upper limit for the value of F. Therefore, the distribu on of F tends to be posi vely skewed, with a lower limit of 0, as in Figure 13.2. The reject region for significance tests with F ra os consists of only one tail (at the upper end of the distribu on). The first table in Appendix C at the end of the book shows the cri cal values of F for α = .05. The second and third tables in Appendix C provide cri cal values of F for α = .01 and α = .001. In the hypothe cal study of stress and anxiety, the F ra o has df equal to 3 and 24. Using α = .05, the cri cal value of F from the first table in Appendix C with df = 3 in the numerator (across the top of the table) and df = 24 in the denominator (along the le -hand side of the table) is 3.01. Thus, in this situa on, the α = .05 decision rule for evalua ng sta s cal significance is to reject H0 when values of F > +3.01 are obtained. A value of 3.01 cuts off the top 5% of the area in the right-hand tail of the F distribu on with df equal to 3 and 24, as shown in Figure 13.2. The obtained F = 11.94 would therefore be judged sta s cally significant.

13.6 PATTERNS OF SCORES AND MAGNITUDES OF SSBETWEEN AND SSWITHIN It is important to understand what informa on about pa ern in the data is contained in these SS and MS terms. SSbetween is a func on of the distances among the group means (M1, M2,…, Mk); the farther apart these group means are, the larger SSbetween tends to be. Most researchers hope to find significant differences among groups, and therefore, they want SSbetween (and F) to be rela vely large. SSwithin is the total of squared within-group devia ons of scores from group means. SSwithin would be 0 in the unlikely event that all scores within each group were equal to one another. The greater the variability of scores within each group, the larger the value of SSwithin. Consider the example shown in Table 13.1, which shows hypothe cal data for which SSbetween would be 0 (because all the group means are equal); however, SSwithin is not 0 (because the scores vary within groups). Table 13.2 shows data for which SSbetween is not 0 (group means differ) but SSwithin is 0 (scores do not vary within groups). Table 13.3 shows data for which both SSbetween and SSwithin are nonzero. Finally, Table 13.4 shows a pa ern of scores for which both SSbetween and SSwithin are 0. Table 13.1 Data for Which SSbetween Is 0 (Because All the Group Means Are Equal), but SSwithin Is Not 0 (Because Scores Vary Within Groups)

Table 13.2 Data for Which SSbetween Is Not 0 (Because Group Means Differ), but SSwithin Is 0 (Because Scores Do Not Vary Within Groups)

Table 13.3 Data for Which SSbetween and SSwithin Are Both Nonzero

Table 13.4 Data for Which Both SSwithin and SSbetween Equal 0

13.7 CONFIDENCE INTERVALS FOR GROUP MEANS Once we know the mean, variance, and n for each group, we can set up a confidence interval (CI) around the mean for each group or a CI for any difference between a pair of group means. Procedures for CIs were reviewed in Chapter 12, on the independent- samples t test, and are not repeated here. 13.8 EFFECT SIZES FOR ONE-WAY BETWEEN-S ANOVA By comparing the sizes of these SS terms that represent variability of scores between and within groups, we can make a summary statement about the compara ve size of the effects of the independent and extraneous variables. The propor on of the total variability (SStotal) that is due to between-group differences is given by Other

(13.16) In the context of a well-controlled experiment, these between-group differences in scores are, presumably, due primarily to the manipulated independent variable; in a nonexperimental study that compares naturally occurring groups, this propor on of variance is reported only to describe the magnitudes of differences between groups, and it is not interpreted as evidence of causality. An eta squared (η2) is an effect size index given as a propor on of variance; if η2 = .50, then 50% of the variance in the Yij scores is related to between-group differences. This is the same eta squared that was introduced in the previous chapter as an effect size index for the independent-samples t test; verbal labels that can be used to describe effect sizes are provided in Table 12.2. If the scores in a two-group t test are par oned into components using the logic just described here and then summarized by crea ng sums of squares, the η2 value obtained will be iden cal to the η2 that was calculated from the t and df terms. It is also possible to calculate eta squared from the F ra o and its df; this is useful when reading journal ar cles that report F tests without providing effect size informa on: Other

(13.17) An eta squared is interpreted as the propor on of variance in scores on the Y outcome variable that is predictable from group membership (i.e., from the score on X, the predictor variable). Suggested verbal labels for eta squared effect sizes were given in Table 12.2. One alterna ve effect size measure some mes used in ANOVA is called omega squared (ω2) (see Hays, 1994). The eta squared index describes the propor on of variance due to between-group differences in the sample, but it is a biased es mate of the propor on of variance that is theore cally due to differences among the popula ons. The ω2 index is

essen ally a (downwardly) adjusted version of eta squared that provides a more conserva ve es mate of variance among popula on means; however, eta squared is more widely used in sta s cal power analysis and as an effect size measure in the literature. Cohen’s f2 is yet another effect size, o en used in sta s cal power analysis. Cohen’s f2 = η2/(1 – η2). 13.9 STATISTICAL POWER ANALYSIS FOR ONE-WAY BETWEEN-S ANOVA Table 13.5 is an example of a sta s cal power table that can be used to make decisions about sample size when planning a one-way between-S ANOVA with k = 3 groups and α = .05. Using Table 13.5, given the number of groups, the number of par cipants, the predetermined alpha level, and the an cipated popula on effect size es mated by eta squared, the researcher can look up the minimum n of par cipants per group that is required to obtain various levels of sta s cal power. The researcher needs to make an educated guess: How large an effect is expected in the planned study? If similar studies have been conducted in the past, the eta squared values from past research can be used to es mate effect size; if not, the researcher may have to make a guess on the basis of less exact informa on. The researcher chooses the alpha level (usually .05), calculates d etween (which equals k – 1, where k is the number of groups in the study), and decides on the desired level of sta s cal power (usually .80, or 80%). Using this informa on, the researcher can use the tables in Cohen (1988) or in Jaccard and Becker (2009) to look up the minimum sample size per group that is needed to achieve the power of 80%. For example, using Table 13.5, for an alpha level of .05, a study with three groups and d etween = 2, a popula on eta squared value of .15, and a desired level of power of .80, the minimum number of par cipants required per group would be 19. Table 13.5 Sta s cal Power for One-Way Between-S ANOVA With k = 3 Groups Using α = .05

Java applets are available on the web for sta s cal power analysis; typically, if the user iden fies a Java applet that is appropriate for the specific analysis (such as between-S one-way ANOVA) and enters informa on about alpha, the number of groups, popula on effect size, and desired level of power, the applet provides the minimum per group sample size required to achieve the user-specified level of sta s cal power. 13.10 PLANNED CONTRASTS The idea behind planned contrasts is that the researcher iden fies a limited number of comparisons between group means before looking at the data. The test sta s c that is used for each comparison is essen ally iden cal to a t ra o, except that the denominator is usually based on the MSwithin for the en re ANOVA, rather than just the variances for the two groups involved in the comparison. Some mes an F is reported for the significance of each contrast, but F is equivalent to t2 in situa ons where only two group means are compared or where a contrast has only 1 df. For the means of Groups a and b, the null hypothesis for a simple contrast between Ma and Mb is as follows: Other

or Other

The test sta s c can be in the form of a t test: Other

(13.18) where n is the number of cases within each group in the ANOVA. (If the n’s are unequal across groups, then an average value of n is used; usually, this is the harmonic1 mean of n’s.) Note that this is essen ally equivalent to an ordinary t test. In a t test, the measure of within-group variability is s2p; in a one-way ANOVA, informa on about within-group variability is contained in the term MSwithin. In cases where an F is reported as a significance test for a contrast between a pair of group means, F is equivalent to t2. The df for this t test equal N – k, where N is the total number of cases in the en re study and k is the number of groups.

When a researcher uses planned contrasts, it is possible to make other kinds of comparisons that may be more complex in form than a simple pairwise comparison of means. For instance, suppose that the researcher has a study in which there are four groups; Group 1 receives a placebo, and Groups 2 to 4 all receive different an depressant drugs. One hypothesis that may be of interest is whether the average depression score combined across the three drug groups is significantly lower than the mean depression score in Group 1, the group that received only a placebo. The null hypothesis that corresponds to this comparison can be wri en in any of the following ways: Other

which can be stated: Other

Contrast coefficients can be used to test for specific pa erns, such as a linear trend (scores on the outcome variable might tend to increase linearly if Groups 1 through 5 correspond to equally spaced dosage levels of a drug): (–2, –1, 0, +1, +2). A curvilinear trend can also be tested; for instance, the researcher might expect to find that the highest scores on the outcome variable occur at moderate dosage levels of the independent variable. If the five groups received five equally spaced different levels of background noise and the researcher predicts the best task performance at a moderate level of noise, an appropriate set of contrast coefficients would be (–1, 0, +2, 0, –1). When a user specifies contrast coefficients, it is necessary to have one coefficient for each level or group in the ANOVA; if there are k groups, each contrast that is specified must include k coefficients. A user may specify more than one set of contrasts, although usually the number of contrasts does not exceed k – 1 (where k is the number of groups). The following simple guidelines are usually sufficient to understand what comparisons a given set of coefficients makes: Groups with posi ve coefficients are compared with groups with nega ve coefficients; groups that have coefficients of 0 are omi ed from such comparisons. It does not ma er which groups have posi ve versus nega ve coefficients; a difference can be detected by the contrast analysis whether or not the coefficients code for it in the direc on of the difference.

For contrast coefficients that represent trends, if you draw a graph that shows how the contrast coefficients change as a func on of group number (X), the line shows pictorially what type of trend the contrast coefficients will detect. Thus, if you plot the coefficients (–2, –1, 0, +1, +2) as a func on of the group numbers 1, 2, 3, 4, and 5, you can see that these coefficients test for a linear trend. The test will detect a linear trend whether it takes the form of an increase or a decrease in mean Y values across groups. When a researcher uses more than one set of contrasts, he or she may want to know whether those contrasts are logically independent, uncorrelated, or orthogonal. There is an easy way to check whether the contrasts implied by two sets of contrast coefficients are orthogonal or independent. Essen ally, to check for orthogonality, you just compute a (shortcut) version of a correla on between the two lists of coefficients. First, you list the coefficients for Contrasts 1 and 2 (make sure that each set of coefficients sums to 0, or this shortcut will not produce valid results).In words, this null hypothesis says that when we combine the means using certain weights (such as +1, –1/3, –1/3, and –1/3), the resul ng composite is predicted to have a value of 0. This is equivalent to saying that the mean outcome averaged or combined across Groups 2 to 4 (which received three different types of medica on) is equal to the mean outcome in Group 1 (which received no medica on). Weights that define a contrast among group means are called contrast coefficients. Usually, contrast coefficients are constrained to sum to 0, and the coefficients themselves are usually given as integers for reasons of simplicity. If we mul ply this set of contrast coefficients by 3 (to get rid of the frac ons), we obtain the following set of contrast coefficients that can be used to see if the combined mean of Groups 2 to 4 differs from the mean of Group 1 (+3, –1, –1, –1). If we reverse the signs, we obtain the set (–3, +1, +1, +1), which s ll corresponds to the same contrast. The F test for a contrast detects the magnitude, and not the direc on, of differences among group means; therefore, it does not ma er if the signs on a set of contrast coefficients are reversed. In SPSS, users can select an op on that allows them to enter a set of contrast coefficients to make many different types of comparisons among group means. To see some possible contrasts, imagine a situa on in which there are k = 5 groups. This set of contrast coefficients simply compares the means of Groups 1 and 5 (ignoring all the other groups): (+1, 0, 0, 0, –1). This set of contrast coefficients compares the combined mean of Groups 1 to 4 with the mean of Group 5: (+1, +1, +1, +1, –4). Contrast coefficients can be used to test for specific pa erns, such as a linear trend (scores on the outcome variable might tend to increase linearly if Groups 1 through 5 correspond to equally spaced dosage levels of a drug): (–2, –1, 0, +1, +2). A curvilinear trend can also be tested; for instance, the researcher might expect to find that the highest scores on the outcome variable occur at moderate dosage levels of the

independent variable. If the five groups received five equally spaced different levels of background noise and the researcher predicts the best task performance at a moderate level of noise, an appropriate set of contrast coefficients would be (–1, 0, +2, 0, –1). When a user specifies contrast coefficients, it is necessary to have one coefficient for each level or group in the ANOVA; if there are k groups, each contrast that is specified must include k coefficients. A user may specify more than one set of contrasts, although usually the number of contrasts does not exceed k – 1 (where k is the number of groups). The following simple guidelines are usually sufficient to understand what comparisons a given set of coefficients makes:

1. Groups with posi ve coefficients are compared with groups with nega ve coefficients; groups that have coefficients of 0 are omi ed from such comparisons.

2. It does not ma er which groups have posi ve versus nega ve coefficients; a difference can be detected by the contrast analysis whether or not the coefficients code for it in the direc on of the difference.

3. For contrast coefficients that represent trends, if you draw a graph that shows how the contrast coefficients change as a func on of group number (X), the line shows pictorially what type of trend the contrast coefficients will detect. Thus, if you plot the coefficients (–2, –1, 0, +1, +2) as a func on of the group numbers 1, 2, 3, 4, and 5, you can see that these coefficients test for a linear trend. The test will detect a linear trend whether it takes the form of an increase or a decrease in mean Y values across groups.

When a researcher uses more than one set of contrasts, he or she may want to know whether those contrasts are logically independent, uncorrelated, or orthogonal. There is an easy way to check whether the contrasts implied by two sets of contrast coefficients are orthogonal or independent. Essen ally, to check for orthogonality, you just compute a (shortcut) version of a correla on between the two lists of coefficients. First, you list the coefficients for Contrasts 1 and 2 (make sure that each set of coefficients sums to 0, or this shortcut will not produce valid results). Contrast 1: (–2, –1, 0, +1, +2) Contrast 2: (+1, –1, 0, 0, 0) You cross-mul ply each pair of corresponding coefficients (i.e., the coefficients that are applied to the same group) and then sum these cross products. In this example, you get

In this case, the sum of the cross products is –1. This means that the two contrasts above are not independent or orthogonal; some of the informa on that they contain about differences among means is redundant. Consider a second example that illustrates a situa on in which the two contrasts are orthogonal or independent:

In this second example, the curvilinear contrast is orthogonal to the linear trend contrast. In a one-way ANOVA with k groups, it is possible to have up to (k – 1) orthogonal contrasts. The preceding discussion of contrast coefficients assumed that the groups in the one-way ANOVA had equal n’s. When the n’s in the groups are unequal, it is necessary to adjust the values of the contrast coefficients so that they take unequal group size into account; this is done automa cally in programs such as SPSS. 13.11 POST HOC OR “PROTECTED” TESTS If the researcher wants to make all possible comparisons among groups or does not have a theore cal basis for choosing a limited number of comparisons before looking at the data, it is possible to use test procedures that limit the risk for Type I error by using “protected” tests. Protected tests use a more stringent criterion than would be used for planned contrasts in judging whether any given pair of means differs significantly. One method for se ng a more stringent test criterion is the Bonferroni procedure, described in Chapter 10. The Bonferroni procedure requires that the data analyst use a more conserva ve (smaller) alpha level to judge whether each individual comparison between group means is sta s cally significant. For instance, in a one-way ANOVA with k = 5 groups, there are k × (k – 1)/2 = 10 possible pairwise comparisons of group means. If the researcher wants to limit the overall experiment-wise risk for Type I error (EWα) for the en re set of 10 comparisons to .05, one possible way to achieve this is to set the PCα level for each individual significance test between means at αEW/(number of post hoc tests to be performed). For example, if the experimenter wants an experiment-wise α of .05 when doing k = 10 post hoc comparisons between groups, the alpha level for each individual test would be set at EWα/k, or .05/10, or .005 for each individual test. The t test could be calculated using the same formula as for an ordinary t test, but it would be judged significant only if its obtained p value were less than .005. The Bonferroni procedure is extremely conserva ve, and many researchers prefer less conserva ve methods of limi ng the risk for Type I error. (One way to make the Bonferroni procedure less conserva ve is to set the experiment-wise alpha to some higher value, such as .10.)

Dozens of post hoc or protected tests have been developed to make comparisons among means in ANOVA that were not predicted in advance. Some of these procedures are intended for use with a limited number of comparisons; other tests are used to make all possible pairwise comparisons among group means. Some of the be er known post hoc tests include the Scheffé test, the Newman-Keuls test, and the Tukey HSD test. The Tukey HSD test has become popular because it is moderately conserva ve and easy to apply; it can be used to perform all possible pairwise comparisons of means and is available as an op on in widely used computer programs such as SPSS. The menu for the SPSS one-way ANOVA procedure includes the Tukey HSD test as one of many op ons for post hoc tests; SPSS calls it the Tukey procedure. The Tukey HSD test (and several similar post hoc tests) uses a different method of limi ng the risk for Type I error. Essen ally, the Tukey HSD test uses the same formula as a t ra o, but the resul ng test ra o is labeled q rather than t, to remind the user that it should be evaluated using a different sampling distribu on. The Tukey HSD test and several related post hoc tests use cri cal values from a distribu on called the “Studen zed range sta s c,” and the test ra o is o en denoted by the le er q: Other

(13.19) where a and b denote any two groups a and b. Values of the q ra o are compared with cri cal values from tables of the Studen zed range sta s c (see the table in Appendix F at the end of the book). The Studen zed range sta s c is essen ally a modified version of the t distribu on. Like t, its distribu on depends on the numbers of subjects within groups, but the shape of this distribu on also depends on k, the number of groups. As the number of groups (k) increases, the number of pairwise comparisons also increases. To protect against inflated risk for Type I error, larger differences between group means are required for rejec on of the null hypothesis as k increases. The distribu on of the Studen zed range sta s c is broader and fla er than the t distribu on and has thicker tails; thus, when it is used to look up cri cal values of q that cut off the most extreme 5% of the area in the upper and lower tails, the cri cal values of q are larger than the corresponding cri cal values of t. This formula for the Tukey HSD test could be applied by compu ng a q ra o for each pair of sample means and then checking to see if the obtained q for each comparison exceeded the cri cal value of q from the table of the Studen zed range sta s c. However, in prac ce, a computa onal shortcut is o en preferred. The formula is rearranged so that the cutoff for judging a difference between groups to be sta s cally

significant is given in terms of differences between means rather than in terms of values of a q ra o. Other

(13.20) Then, if the obtained difference between any pair of means (such as Ma – Mb) is greater in absolute value than this HSD, this difference between means is judged sta s cally significant. An HSD criterion is computed by looking up the appropriate cri cal value of q, the Studen zed range sta s c, from a table of this distribu on (see the table in Appendix F). The cri cal q value is a func on of both n, the average number of subjects per group, and k, the number of groups in the overall one-way ANOVA. As in other test situa ons, most researchers use the cri cal value of q that corresponds to α = .05, two tailed. This cri cal q value obtained from the table is mul plied by the error term to yield HSD. This HSD is used as the criterion to judge each obtained difference between sample means. The researcher then computes the absolute value of the difference between each pair of group means (M1 – M2), (M1 – M3), and so forth. If the absolute value of a difference between group means exceeds the HSD value just calculated, then that pair of group means is judged to be significantly different. Then, if the obtained difference between any pair of means (such as Ma – Mb) is greater in absolute value than this HSD, this difference between means is judged sta s cally significant. An HSD criterion is computed by looking up the appropriate cri cal value of q, the Studen zed range sta s c, from a table of this distribu on (see the table in Appendix F). The cri cal q value is a func on of both n, the average number of subjects per group, and k, the number of groups in the overall one-way ANOVA. As in other test situa ons, most researchers use the cri cal value of q that corresponds to α = .05, two tailed. This cri cal q value obtained from the table is mul plied by the error term to yield HSD. This HSD is used as the criterion to judge each obtained difference between sample means. The researcher then computes the absolute value of the difference between each pair of group means (M1 – M2), (M1 – M3), and so forth. If the absolute value of a difference between group means exceeds the HSD value just calculated, then that pair of group means is judged to be significantly different. When a Tukey HSD test is requested from SPSS, SPSS provides a summary table that shows all possible pairwise comparisons of group means and reports whether each of these comparisons is significant. If the overall F for the one-way ANOVA is sta s cally significant, it implies that there should be at least one significant contrast among group

means. However, it is possible to have situa ons in which a significant overall F is followed by a set of post hoc tests that do not reveal any significant differences among means. This can happen because protected post hoc tests are somewhat more conserva ve and thus require slightly larger between-group differences as a basis for a decision that differences are sta s cally significant, than the overall one-way ANOVA. 13.12 ONE-WAY BETWEEN-S ANOVA IN SPSS To run the one-way between-S ANOVA procedure in SPSS, make the following menu selec ons from the menu bar at the top of the Data View worksheet, as shown in Figure 13.3: <Analyze> → <Compare Means> → <One-Way ANOVA>. This opens the dialog box in Figure 13.4. Enter the name of one (or several) dependent variables into the pane labeled “Dependent List”; enter the name of the categorical variable that provides group membership informa on into the box labeled “Factor.” For this example, addi onal windows were accessed by clicking on the bu ons marked Post Hoc, Contrasts, and Op ons. The screenshots that correspond to this series of dialog boxes appear in Figures 13.4 through 13.7.

Figure 13.5 One-Way ANOVA: Post Hoc Mul ple Comparisons Dialog Box

Figure 13.6 Specifica on of a Planned Contrast

The null hypothesis about a weighted linear composite of means that is represented by this set of contrast coefficients: Other

or Other

or Other

From the menu of post hoc tests, this example uses the one SPSS calls “Tukey” (this corresponds to the Tukey HSD test). To define a contrast that compares the mean of Group 1 (no stress) with the mean of the three stress treatment groups combined, these contrast coefficients are entered one at a me: +3, –1, –1, –1. From the list of op ons, “Descrip ve” sta s cs and “Homogeneity of variance test” were selected by placing checks in the boxes next to the names of these tests.

Figure 13.7 One-Way ANOVA: Op ons Dialog Box

13.13 OUTPUT FROM SPSS FOR ONE-WAY BETWEEN-S ANOVA The output for this one-way ANOVA is reported in Figure 13.8. The first panel provides descrip ve informa on about each of the groups: mean, standard devia on, n, a 95% CI for the mean, and so forth. The second panel shows the results for the Levene test of the homogeneity of variance assump on; this is an F ra o with (k – 1) and (N – k) df. The obtained F was not significant for this example; there was no evidence that the homogeneity of variance assump on had been violated. The third panel shows the ANOVA source table with the overall F; this was sta s cally significant, and this implies that there was at least one significant contrast between group means. In prac ce, a researcher would not report both planned contrasts and post hoc tests; however, both were presented for this example to show how they are obtained and reported. Figure 13.9 shows the output for the planned contrast that was specified by entering these contrast coefficients: (+3, –1, –1, –1). These contrast coefficients correspond to a test of the null hypothesis that the mean anxiety of the no-stress group (Group 1) was not significantly different from the mean anxiety of the three stress interven on groups (Groups 2–4) combined. SPSS reported a t test for this contrast (some textbooks and programs use an F test). This t test was sta s cally significant, and examina on of the group means indicated that the mean anxiety level was significantly higher for the three stress interven on groups combined, compared with the control group.

Figure 13.10 shows the results for the Tukey HSD tests that compared all possible pairs of group means. The table “Mul ple Comparisons” gives the difference between means for all possible pairs of means (note that each comparison appears twice; that is, Group a is compared with Group b, and in another row, Group b is compared with Group a). Examina on of the “Sig.” or p values indicates that several of the pairwise comparisons were significant at the .05 level. The results are displayed in a more easily readable form in the last panel under the heading “Homogeneous Subsets.” Each subset consists of group means that were not significantly different from one another using the Tukey test. The no-stress group was in a subset by itself; in other words, it had significantly lower mean anxiety than any of the three stress interven on groups. The second subset consisted of the stress role play and mental arithme c groups, which did not differ significantly in anxiety. The third subset consisted of the mental arithme c and mock job interview groups.

Figure 13.10 SPSS Output for Post Hoc Test (Tukey HSD)

Note that it is possible for a group to belong to more than one subset; the anxiety score for the mental arithme c group was not significantly different from the stress role play or the mock job interview groups. However, because the stress role play group differed significantly from the mock job interview group, these three groups did not form one subset. Note also that it is possible for all the Tukey HSD comparisons to be nonsignificant even when the overall F for the one-way ANOVA is sta s cally significant. This can happen because the Tukey HSD test requires a slightly larger difference between means to achieve significance. In this imaginary example, as in some research studies, the outcome measure (anxiety) is not a standardized test for which we have norms. The numbers by themselves do not tell us whether the mock job interview par cipants were moderately anxious or twitching, stu ering wrecks. Studies that use standardized measures can make comparisons with test norms to help readers understand whether the group differences were large enough to be of clinical or prac cal importance. Alterna vely, qualita ve data about the behavior of par cipants can also help readers understand how substan al the group differences were.

SPSS one-way ANOVA does not provide an effect size measure, but this can easily be calculated by hand. In this case, eta squared is found by taking the ra o SSbetween/SStotal from the ANOVA source table: η2 = .60.

Graphs used to represent group means with confidence intervals for one-way ANOVA are like those used in Chapter 12, on the independent-samples t test; menu selec ons are not repeated here. A bar chart with 95% CI error bars appears in Figure 13.11. 13.14 REPORTING RESULTS FROM ONE-WAY BETWEEN-S ANOVA Following is an example of a “Results” sec on for the one-way between-S ANOVA in the study of anxiety and stress. Results A one-way between-S ANOVA was done to compare the mean scores on an anxiety scale (0 = not at all anxious, 20 = extremely anxious) for par cipants who were randomly assigned to one of four groups: Group 1, control group/no stress; Group 2, mental arithme c; Group 3, stressful role play; and Group 4, mock job interview. Examina on of a histogram of anxiety scores indicated that the scores were approximately normally distributed with no extreme outliers. Prior to the analysis, the Levene test for homogeneity of variance was used to examine whether there were serious viola ons of the homogeneity of variance assump on across groups, but no significant viola on was found, F(3, 24) = .718, p = .72. The overall F for the one-way ANOVA was sta s cally significant, F(3, 24) = 11.94, p < .001. This corresponded to an effect size of η2 = .60; about 60% of the variance in anxiety scores was predictable from the type of stress interven on. This is a large effect. The means and standard devia ons for the four groups are shown in Table 13.6.

One planned contrast (comparing the mean of Group 1, no stress, with the combined means of Groups 2–4, the stress interven on groups) was performed. This contrast was tested using α = .05, two tailed; the t test that assumed equal variances was used because the homogeneity of variance assump on was not violated. For this contrast, t(24) = –5.18, p < .001. The mean anxiety score for the no-stress group (M = 9.86) was significantly lower than the mean anxiety score for the three combined stress interven on groups (M = 14.95).

In addi on, all possible pairwise comparisons were made using the Tukey HSD test. On the basis of this test (using α = .05), it was found that the no-stress group scored significantly lower on anxiety than all three stress interven on groups. The stressful role play (M = 13.57) was significantly less anxiety producing than the mock job interview (M = 17.00). The mental arithme c task produced a mean level of anxiety (M = 14.29) that was intermediate between the other stress condi ons, and it did not differ significantly from either the stress role play or the mock job interview. Overall, the mock job interview produced the highest levels of anxiety. Figure 13.11 shows a bar chart of group means with 95% CIs. Data analysts usually just report one type of follow-up analysis, either post hoc tests or planned contrasts, but not both. 13.15 ISSUES IN PLANNING A STUDY When an experiment is designed to compare treatment groups, the researcher needs to decide how many groups to include, how many par cipants to include in each group, how to assign par cipants to groups, and what types or dosages of treatments to administer to each group. In an experiment, the levels of the factor may correspond to different dosage levels of the same treatment variable (such as 0 mg caffeine, 100 mg caffeine, 200 mg caffeine) or to qualita vely different types of interven ons (as in the stress study, where the groups received, respec vely, no stress, mental arithme c, role play, or mock job interview stress interven ons). In addi on, it is necessary to think about issues of experimental control, as described in research methods textbooks; for example, in experimental designs, researchers need to avoid confounds of other variables with the treatment variable. Research methods textbooks (e.g., Cozby & Bates, 2017) provide more detailed discussion of design issues; a few guidelines are listed here.

a) If the treatment variable has a curvilinear rela on to the outcome variable, it is necessary to have a sufficient number of groups to describe this rela on accurately—at least 3 groups. On the other hand, it may not be prac cal or affordable to have a very large number of groups; if an absolute minimum n of 10 (or, be er, 30) par cipants per group are included, then a study that included 15 groups would require a total N of 150 (or 450) par cipants.

b) In studies of interven ons that may create expectancy effects, it is necessary to include one or several kinds of placebo and/or no-treatment/control groups for comparison. For instance, studies of the effect of biofeedback on heart rate (HR) some mes include a group that gets real biofeedback (a tone is turned on when HR increases), a group that gets noncon ngent feedback (a tone is turned on and off at random in a way that is unrelated to HR), and a group that receives instruc ons for relaxa on and sits quietly in the lab without any feedback (Burish, 1981).

c) Random assignment of par cipants to condi ons is desirable to try to ensure equivalence of the groups prior to treatment. For example, in a study where HR is the dependent variable, it would be desirable to have par cipants whose HRs

were equal across all groups prior to the administra on of any treatments. However, it cannot be assumed that random assignment will always succeed in crea ng equivalent groups. In studies where equivalence among groups prior to treatment is in doubt, the researcher should collect data on par cipant characteris cs and compare groups prior to treatment to verify whether the groups are equivalent.

d) It is crucial to make sure that no other variable is confounded with the treatment variable; the presence of a confound makes differences among group means uninterpretable.

The same factors that affect the size of the t ra o also affect the sizes of F ra os: distances among group means, the amount of variability of scores within each group, and the number, n, of par cipants per group. Other things being equal, an F ra o tends to be larger when there are large between-group differences among dosage levels (or par cipant characteris cs). As an example, a study that used three different noise levels in decibels as the treatment variable, for example, 35, 65, and 95 dB, would result in larger differences in group means on arousal and a larger F ra o than a study that looked, for example, at these three noise levels, which are closer together: 60, 65, and 70 dB. Like the t test, the F ra o involves a comparison of between-group and within- group variability; the selec on of homogeneous par cipants, standardiza on of tes ng condi ons, and control over extraneous variables will tend to reduce the magnitude of within-group variability of scores, which in turn tends to produce a larger F (or t) ra o. A researcher can o en increase the size of an F (or t) ra o by increasing the differences in the dosage levels of treatments given to groups and/or reducing the effects of extraneous variables through experimental control, and/or increasing the number of par cipants. In studies that involve comparisons of naturally occurring groups, such as age groups, a similar principle applies: A researcher is more likely to see age-related changes in mental processing speed in a study that compares ages 20, 50, and 80 than in a study that compares ages 20, 25, and 30. 13.16 SUMMARY One-way between-S ANOVA provides a method for comparison of more than two group means. However, the overall F test for the ANOVA does not provide enough informa on to completely describe the pa ern in the data. It is o en necessary to perform addi onal comparisons among specific group means to provide a complete descrip on of the pa ern of differences among group means. These can be a priori (also called planned contrast) comparisons if a limited number of differences are predicted in advance and a small number of significance tests are performed. If the researcher did not make predic ons in advance about differences between group means, then he or she may use protected or post hoc tests to do any follow-up comparisons; the Bonferroni procedure and the Tukey HSD test were described here, and many other post hoc procedures are available.

The most important concept from this chapter is the idea that a score can be divided into components (one part that is related to group membership or treatment effects and a second part that is due to the effects of all other “extraneous” variables that uniquely influence individual par cipants). Informa on about the rela ve sizes of these components can be summarized across all the par cipants in a study by compu ng the sum of squared devia ons (SS) for the between-group and within-group devia ons. On the basis of the SS values, it is possible to compute an effect size es mate (η2) that describes the propor on of variance predictable from group membership (or treatment variables) in the study. Researchers usually hope to design their studies in a manner that makes the propor on of explained variance reasonably high and that produces sta s cally significant differences among group means. However, researchers should remember that the propor on of variance due to group differences in the ar ficial world of research may not correspond to the “true” strength of the influence of the variable out in the “real world.” In experiments, we create an ar ficial world by holding some variables constant and by manipula ng the treatment variable; in nonexperimental research, we create an ar ficial world through our selec on of par cipants and measures. Research results should be interpreted and generalized cau ously. APPENDIX 13A: ANOVA MODEL AND DIVISION OF SCORES INTO COMPONENTS When we do a one-way ANOVA, the analysis involves par on of each score into two components: a component of the score that is associated with group membership and a component of the score that is not associated with group membership. The sums of squares (SS) summarize informa on about the magnitudes of these components across all scores, and a ra o of SS terms will es mate what propor on of variance in scores is associated with type or amount or treatment or other characteris cs that differ across groups. Examining propor on of variance associated with treatment group membership is one way of approaching the general research ques on: Why is there variance in the scores on the outcome variable?

Table 13.7 Par on of Heart Rate Scores in Hypothe cal Sex Differences in Heart Rate Data

CHAPTER 15 ONE-WAY REPEATED-MEASURES ANALYSIS OF VARIANCE 15.1 INTRODUCTION One-way repeated-measures analysis of variance (ANOVA) provides comparisons across more than two levels of a repeated-measures or within-S factor. The data set for the hypothe cal study of effect of stress on heart rate (HR) (withins.sav) used in the previous chapter on the paired-samples t test is used again here. In this hypothe cal study, HR was assessed for each par cipant under four different stress condi ons: baseline (no stress), pain, mental arithme c, and role play. This example illustrates a design problem that is discussed later in this chapter: problems with order effects. The paired-samples t test examined only the first two condi ons; repeated-measures ANOVA will be used to examine differences among means across all four condi ons. For repeated-measures ANOVA, the same variable must be measured (in the same units) at each point in me. It would not make sense to compare mean anxiety during baseline with mean blood pressure during mental arithme c. The levels of a within-S factor can be any of the following: Different types of treatment. The example in this chapter uses different types of stress. Different amounts or dosages of the same treatment. For example, a drug could be given in doses of 10 mg, 20 mg, 30 mg, and so forth. For both these situa ons, at least one baseline or no-treatment control group is usually included; return to baseline a er treatment may also be assessed. Some mes levels of the within-S factors correspond to different mes:

 Levels of a within-S factor can represent assessments made at different mes (in the absence of any interven on). A longitudinal study of anxiety in elementary school students might assess anxiety for the same par cipants in Grade 1, Grade 2, Grade 3, and so on.

 For a study that includes an interven on, similar to the pretest–pos est quasi- experimental design described earlier, baseline may be assessed at more than one me and/or treatment effects may be assessed at more than one me. If there is an interven on at Month 1 in a study, assessments of the outcome variable can be made at Months 3, 4, 5, and later to evaluate whether treatment effects wear off over me.

Repeated-measures ANOVA has many of the same advantages as the paired-samples t test. Smaller numbers of par cipants may be needed than for between-S studies, and the F test may have greater sta s cal power than the F test for between-S ANOVA. Effects of treatments and differences among mes are assessed with sta s cal control for individual differences among par cipants. Just as SEd was o en smaller than SEM1– M2, SS*error for repeated-measures ANOVA is usually smaller than SSerror for between- S ANOVA. (When numerical values differ, corresponding terms for repeated-measures ANOVA are shown with an asterisk, and terms for the between-S ANOVA are shown without an asterisk.) For both the paired-samples t test and repeated-measures ANOVA,

error terms are usually smaller because varia on in scores due to systema c individual differences among persons is removed from the error term. Repeated-measures ANOVA has many of the same poten al disadvantages as the correlated-samples t test (e.g., order effects, carryover effects, fa gue, prac ce effects). For some issues, such as order effects, counterbalancing requirements are more complicated. 15.2 NULL HYPOTHESIS FOR REPEATED-MEASURES ANOVA When there are more than two levels or groups of repeated measures, one-way repeated-measures ANOVA is used to test differences among means. The null hypothesis is that the mean for the dependent variable (such as heart rate) is the same across all

mes or treatments. For a within-S factor that has k levels, H0 is: Other

(15.1) For the example used in this chapter, with HR as the dependent variable, and four levels of the within-S factor, H0: μbaseline = μarithme c = μpain = μroleplay. As in other applica ons of ANOVA, the alterna ve hypothesis is that there is at least one sta s cally significant difference between a pair of means (or for at least one contrast of means). A significant overall F does not indicate which comparisons are significant. Contrasts or follow-up tests are required to evaluate the nature of group differences. 15.3 PRELIMINARY ASSESSMENT OF REPEATED-MEASURES DATA Data for repeated-measures ANOVA are organized as shown in Table 15.1, with one row for each case or par cipant. Each measure (made at different mes or under different treatment condi ons) is represented as a variable. Data for this hypothe cal study are in the SPSS file withins.sav. By now, calcula ng means for groups of scores should be almost a reflex. For the data in Table 15.1, you can compute mean HR for each type of stress (for each column of scores). This is familiar. The values of these means provide a preliminary look at possible treatment effects; the arithme c task elicited the lowest HR among the three types of stress, and role play elicited the highest. No ce that in this situa on, you can also calculate a set of means for each row of the data set. The row means provide informa on about individual differences in HR. No ce that Erin’s HR was the highest, and Chris’s the lowest, among persons. Mgrand (the grand mean for all of the HR scores) can be obtained by averaging either the four treatment group means or the six person means. SS for persons can be obtained by examining devia ons of person means from the grand mean.

Having informa on about mean HR for each par cipant makes it possible to sta s cally control for these individual differences in HR among par cipants. When we sta s cally control for person effects in repeated measures, we remove variance associated with differences among persons from the error term. When we have a set of means, we can evaluate how much the means differ from one another (or how much the means differ from the grand mean) by compu ng a sum of squares (a sum of squared differences of each mean from the grand mean). Recall that in a between-S ANOVA with k groups, the par on of SStotal was as follows:

Other

(15.2) For the repeated-measures data, an SS term is added to provide informa on about differences in the person means for the dependent variable scores (heart rate, in this example). The par on of SStotal in a one-way repeated-measures ANOVA with k treatments and n persons includes these three SS terms for a repeated-measures ANOVA. Other

(15.3) SSerror is used in the equa on for par on of SS in the between-S design; SS*error, which will have a different and usually smaller numerical value than SSerror, appears with an asterisk in the equa on for the repeated-measures design (this nota on is not generally used elsewhere). Similarly, dferror and df *error are used to dis nguish df for

error in between- versus within-S analyses. The asterisk draws your a en on to the fact that the numerical values of these terms will differ. SStotal and SStreatment have the same values for the between-S and within-S analyses. It follows that SSerror in the between-S ANOVA equals the sum of SS*error and SSpersons in the within-S ANOVA (Equa on 15.4). Also, dferror in between-S ANOVA is the sum of df *error and dfpersons in the within-S ANOVA (Equa on 15.5). Other

(15.4) Other

(15.5) Rearranging Equa on 15.4, we have: Other

(15.6) Recall that we usually want error SS to be small in ANOVA. Equa on 15.6 tells us that for the within-S or repeated-measures ANOVA, SS*error is obtained by subtrac ng (or par alling out) the varia on due to persons (SSpersons). Thus SS*error will be less than SSerror (as long as SSpersons > 0). Within-S ANOVA (almost always) has a smaller error sum of squares than between-S ANOVA. This happens because we can quan fy varia ons among persons in within-S ANOVA (SSpersons) and subtract or remove SSpersons from the error SS in between-S ANOVA. We can say that we are par alling out or sta s cally controlling for differences among persons (differences among persons in heart rate, in this example). Sta s cal control is a fundamental concept in further study of sta s cs. Up to this point in the book, sta s cs have been limited to bivariate situa ons (analyses that include only one X predictor and one Y outcome). When one or more addi onal variables (call this addi onal variable Z) are added to an analysis, there are many ways that X, Y, and Z can be related. Some mes sta s cal control for Z changes the apparent rela onship between X and Y. Varia on related to Z may be par alled out of one or both the other variables. Mul ple regression and other analyses provide addi onal ways to do sta s cal control. Most published research reports include more than two variables and usually include one or more forms of sta s cal control. In general, when you read sentences of the form “X was predic ve of Y, controlling for Z,” it tells you that associa ons or correla ons with Z have been removed from one or both of the variables and that the X, Y rela onship s ll holds with Z controlled. Controlling for Z can change the associa on between X and Y in different ways. In other situa ons, controlling for a Z

variable can make an X, Y correla on become weaker, disappear, or even change sign. In the repeated-measures ANOVA example in this chapter, controlling for Z (individual differences in HR) usually makes the associa on between type of stress (X) and HR (Y) stronger. 15.4 COMPUTATIONS FOR ONE-WAY REPEATED-MEASURES ANOVA To obtain a value for SStotal, as in other types of ANOVA, we subtract the grand mean Mgrand from each individual score, square the resul ng devia ons, and sum these squared devia ons (across all persons and treatments, in this case, across all 24 scores in the body of Table 15.1). Xij represents the score for Person i in Treatment Condi on j or Time j. The use of Xij in Equa on 15.7 indicates that devia ons from the grand mean are obtained for every individual score (the score for Person i, in Treatment Condi on j). N represents the number of par cipants, and k represents the number of levels (or treatments or mes) for the repeated-measures factor. N × k is the total number of scores. The grand mean, denoted Mgrand, is obtained by summing all the individual scores and dividing by the total number of scores. All other SS terms are based on devia ons of scores or means from Mgrand. Other

(15.7) Other

(15.8) Other

(15.9) For the data in Table 15.1 we would obtain: Other

To obtain a value for SStreatment, we need to find the devia on of each column mean from the grand mean, square these devia ons, and sum the squared devia ons across all k mes or treatments. Because there are N scores in each column, we mul ply that sum by N. MTj refers to the mean for Treatment Group j; this nota on tells us that a devia on from Mgrand is calculated for each of the k treatment group means. In the empirical example, the set of treatment group means refers to Mbaseline, Mpain, Marithme c, and Mroleplay.

Other

(15.10) where MTj is the mean score for treatment condi on j, Mgrand is the grand mean of all the scores, and N is the number of scores in each treatment condi on (in this example, N = 6). For d reatment, we have: Other

(15.11) where k is the number of levels of repeated-measures factor (number of treatments or

mes). When we apply Equa on 15.9 to the data in Table 15.1 we obtain: Other

To obtain a value for SSpersons, we need to find the devia on of each row mean (i.e., the mean across all scores for Person i) from the grand mean, square this devia on, and sum the squared devia ons across all N persons: Other

(15.12) where MPi is the mean of scores for Person i, Mgrand is the grand mean, and k is the number of treatment condi ons (in this example, k = 4).

For the person SS, the corresponding df term is: Other

(15.13) where N is the number of persons. In this case, when we apply Equa on 15.12 to the data in Table 15.1, we obtain: Other

Finally, the value of SS*error can be obtained by subtrac on: Other

(15.14) Other

(15.15) where N is the number of persons, and k is the number of levels of within-S factor. The df for error is denoted df * for the repeated-measures ANOVA to make it clear that this differs from df error in a between-S ANOVA. For the data in the preceding example, using the data from Table 15.1, we obtain the following: Other

In a one-way repeated-measures ANOVA with N subjects and k treatments, there is a total of N × k (in our empirical example, 6 × 4 = 24) observa ons. Here is a summary of the df values for the hypothe cal study of stress and heart rate:

Each SS is used to calculate a corresponding mean square (MS) term by dividing it by the corresponding df term: Other

An F ra o is set up to evaluate MS for treatments rela ve to MS*error: Other

(15.16) with (k – 1) d reatment and (N – 1) × (k – 1) df*error. For the one-way repeated-measures ANOVA of the scores in Table 15.1, we obtain F = 162.819/18.986 = 8.58, with 3 and 15 df. The cri cal value of F for α = .05 and (3, 15) df is 3.29. This result for this one-way repeated-measures ANOVA can therefore be judged sta s cally significant at the α = .05 level. This tells us that there must be at least one significant difference or contrast between treatment group means, but it does not tell us which pairs of means differ significantly. (It is not customary to set up an F ra o for person effects; person effects are viewed as a type of error variance.) Recall that a common effect size for between-S one-way ANOVA was η2. For between-S ANOVA: Other

(15.17)

The η2 for a one-way between-S ANOVA is interpreted as the propor on of variance in scores on the dependent variable (such as heart rate) that is predictable from (or associated with) treatment group membership. For the repeated-measures ANOVA, effect size can be assessed using par al η2: Other

(15.18) Par al eta squared answers the ques on, A er par alling out (or removing or controlling for) individual differences among persons, what percentage of the remaining variance is associated with differences among treatment groups? (No ce that SSpersons is not included in the divisor for par al η2 in Equa on 15.18; it has been par alled out or removed.) We could obtain an η2 (instead of a par al η2) for the repeated-measures ANOVA by including SSpersons in the divisor, as in Equa on 15.19. This would tell us what propor on of the total variance in HR scores was related to treatment condi on, not controlling for differences among persons: Other

(15.19) Unless SSpersons happens to be exactly 0, par al η2 will have a smaller divisor than η2, and par al η2 will be larger than η2. Par al η2 is interpreted as the propor on of variance in scores for the dependent variable that is related to or predictable from treatment group membership when variance due to individual differences has been par alled out (or sta s cally controlled). MS*error for a repeated-measures design is SS*error/df*error. When we compare MS*error with MSerror in a between-S analysis, both the numerator and divisor differ. SS*error is smaller than SSerror; however, df*error is also smaller than dferror. Whether MS*error for repeated-measures ANOVA is smaller or larger than MSerror for between-S ANOVA depends on whether the reduc on in SS*error is large enough to compensate for the reduc on in df. When MS*error is smaller than MSerror, repeated-measures ANOVA has greater sta s cal power than between-S ANOVA. 15.5 USE OF SPSS RELIABILITY PROCEDURE FOR ONE-WAY REPEATED-MEASURES ANOVA The SPSS general linear model (GLM) procedure is the usual way to obtain a repeated- measures ANOVA. For a simple one-way repeated-measures ANOVA, it is possible to use the SPSS reliability procedure. This is a useful hack: It is easier to set up the analysis and easier to read the output of the reliability procedure. Reliability also provides one kind

of informa on not given by GLM that may be useful: test of an assump on about Person × Treatment interac on. However, GLM provides more informa on about comparisons or contrasts involving group means, and the use of GLM is also demonstrated later in the chapter. Recall that measures for heart rate are correlated with each other across condi ons; these correla ons are evidence of consistent individual differences (some persons tended to have consistently higher heart rates than others). The SPSS reliability procedure was developed primarily to examine consistency or reliability. It is more o en used to assess consistency of responses to items on a self-report survey. (Because of this, the naming of sources of variances in the source table produced by the reliability procedure is not consistent with the names generally used in reports of repeated measures.) The SPSS menu selec ons to run repeated-measures ANOVA using the reliability procedure appear in Figure 15.1: <Analyze> → <Scale> → <Reliability Analysis>. In the main dialog box (Figure 15.2), enter the names of the variables that represent scores obtained in the treatment condi ons in the pane under the “Items” heading. For the hypothe cal study about stress and HR, these variables are HR_baseline, HR_pain, and so forth. (When the reliability procedure reports differences among means for “items,” these correspond to the differences among means for the levels of the within-S factor.) Click the Sta s cs bu on in Figure 15.3 to open the Reliability Analysis: Sta s cs dialog box and make the following selec ons: Under “Descrip ves for,” check the box for “Item” (this will display treatment group means). Under “Inter-Item,” check “Correla ons” (this will provide correla ons between HR_baseline and HR_pain, HR_pain and HR_arithme c, and all other pairs of HR measures). Under the “ANOVA Table” heading, click the radio bu on for “F test.” Click Con nue and OK.

The first part of the output in Figure 15.4 is Cronbach’s alpha reliability. (For now, you can ignore this; it is useful when you want to assess response consistency in mul ple- measure assessments such as self-report measures of a tude.) The second part of the output provides the mean for each treatment condi on. This tells you which groups had the highest and lowest mean heart rate. The third part provides correla ons among all four HR variables. If these correla ons are low (or, in the worst case, if any of them are nega ve), then a repeated-measures ANOVA provides li le or no advantage in sta s cal power. Moderate to high posi ve correla ons among measures are needed to obtain smaller error terms and improvement in sta s cal power. (For the paired-samples t test, Equa on 12.11 showed how the magnitude of correla ons between measures is related to the size of the error term.) If you plan to do future studies, finding small (or even nega ve) correla ons among measures would suggest that you are not likely to gain sta s cal power by using repeated measures.

The outcome of primary interest is the ANOVA source table; this appears with more informa ve labels in Figure 15.5. (Source refers to sources of variance.) The numerical values agree with those from earlier by-hand computa ons. You might think that an F ra o could be created for individual differences among persons by looking at MSpersons/MSerror. However, this is usually not reported; individual differences in an outcome variable such as HR are regarded as a type of error, and it would not make sense to try to generalize about the magnitude of individual differences in HR among all persons in some larger hypothe cal popula on on the basis of a small sample. Par al η2 for this example can be calculated by hand. Par al η2 = SStreatment /(SStreatment + SS*error) = 488.458/(488.458 + 284.792) = .63. A er controlling for (or par alling out) differences among persons in HR, 63% of the variance in HR was related to type of treatment.

A brief report of this analysis would be as follows:

A more complete report would include a 95% confidence interval (CI) for each value of M, either in the text or in an error bar graph. Follow-up tests should be performed to assess which pairs of means differ significantly. That informa on is not provided by the reliability procedure; it appears in the upcoming output from the GLM procedure. Analysis so far does not tell us whether pairs of treatments (e.g., pain vs. arithme c) have significantly different means. Addi onal informa on should be included when repor ng repeated-measures ANOVA.

 The “Method” sec on must specify treatment order and what was done to deal with possible order or carryover effects.

 Tests for poten al viola ons of assump ons for repeated-measures ANOVA should be conducted, and the results of those tests should be included. Many of these are the same as the evalua ons of assump ons for the paired-samples t test. Histograms and boxplots can be obtained to assess whether scores on HR_baseline, HR_pain, and so on, are normally distributed and whether extreme outliers are present. Sca erplots for each pair of variables (e.g., HR_baseline, HR_pain) are used to assess linearity and iden fy bivariate outliers. A checkbox in the Reliability Analysis dialog box in Figure 15.2 can be used to request Tukey’s test of addi vity (see Appendix 15A for details). The SPSS GLM procedure provides a test of another important assump on (Mauchly’s sphericity test).

 If major assump ons are violated, the F ra o may be inflated, and the p value may underes mate the true risk for Type I error. Later sec ons describe possible ways to deal with these problems.

 Follow-up tests are needed to iden fy which pairs of group means, or contrasts between group means, are sta s cally significant.

 Addi onal analyses may be needed to evaluate effects of par cipant dropout (also called a ri on, or some mes subject mortality) and other problems in longitudinal studies (not covered in this chapter).

 Poten al problems with repeated measures (e.g., order effects, fa gue) should be evaluated in the “Discussion” sec on.

Figure 15.5 Relabeled ANOVA Source Table From the SPSS Reliability Procedure

15.6 PARTITION OF SS IN BETWEEN-S VERSUS WITHIN-S ANOVA Let’s revisit the difference in the par on of SStotal for between-S and within-S one-way ANOVA by examining the numerical values of SS terms in Figure 15.5. The results for a one-way between-S ANOVA for this data set appear in Figure 15.6. For the between-S ANOVA that examines differences in mean HR across four levels of a factor called stress (Level 1, no stress/baseline; Level 2, pain induc on; Level 3, mental arithme c task; and Level 4, stressful social role play), F(3, 20) = 1.485, p = .249, η2 = .18. From the one-way between-S ANOVA, we would conclude that mean HR does not differ significantly across different levels or types of stress. The par on of SStotal in a one-way between-S ANOVA can be diagrammed as follows (see Figure 15.7).

Figure 15.6 Output From One-Way Between-S ANOVA Using Same Scores as in Within-S ANOVA The source table for the one-way repeated-measures analysis of variance is repeated here for ease of comparison (see Figure 15.8). The diagram in Figure 15.9 illustrates the way SStotal is por oned in repeated-measures ANOVA. First, SStotal is divided into SSbetween people (or persons) and SSwithin people. In the output, SSbetween people = 1,908.375 and SSwithin people = 773.250; these two SS terms sum to SStotal, 2,681.625.

Figure 15.9 Par on of SStotal in One-Way Repeated-Measures ANOVA

Source tables for repeated-measures ANOVA are divided into between- and within-S sec ons. SSbetween S tells us how much varia on in the dependent variable (such as HR) is associated with consistent differences among persons or subjects. This varia on is treated as a type of error that is removed before further breaking down sources of variance. SSwithin S is broken down into two components: SS for treatments and SS*error. The F ra o, MStreatment/MS*error, is used to evaluate whether treatment means differ significantly a er varia on because individual differences in the dependent variable (such as HR) have been removed from the scores. 15.7 ASSUMPTIONS FOR REPEATED-MEASURES ANOVA For the F ra o in repeated-measures ANOVA to provide accurate informa on about the true risk for Type I error (that is, for its p value to be believable), assump ons for repeated-measures ANOVA need to be reasonably well sa sfied. For a repeated-measures ANOVA, the first two assump ons are the same as those for the paired-samples t test. 15.7.1 Scores on Outcome Variables Are Quan ta ve and Approximately Normally Distributed Without Extreme Outliers Scores on the outcome variables should be quan ta ve and approximately normally distributed, and there should not be extreme outliers. The normality assump on can be evaluated by examining histograms; boxplots are one way to iden fy poten al outliers. 15.7.2 Rela onships Among the Repeated-Measures Variables Should Be Linear Without Bivariate Outliers For example, if you set up sca erplots between HR_baseline and HR_pain, HR_pain and HR_arithme c, and so forth, all sca erplots should show reasonably linear pa erns. Also, correla ons among all the repeated-measures variables should be posi ve and at least moderate. 15.7.3 Popula on Variances of Contrasts Should Be Equal (Sphericity Assump on) To understand this aspect of repeated measures, we need to think about the analysis in terms of contrasts. Each contrast can be the same as a difference score between two treatments. I will denote each contrast as C (instead of d, the nota on used in the paired-samples t test) to be consistent with nota on elsewhere in discussion of repeated measures. For example, Contrast 1 (C1) could correspond to differences of scores between baseline and pain. This is the same as the d difference score used in the paired-samples t test. C2 could be difference scores between baseline and arithme c, and C3 could be difference scores between baseline and role play. The number of contrasts in a set is (k – 1), where k is the number of levels for the within-S factor. Just as you used SPSS to compute a d or difference score for the paired-samples t test, you could create scores for each of these C variables. The null hypothesis for each C contrast is:

Other

(15.20) Another way to state the null hypothesis for the omnibus F test in repeated-measures ANOVA is that, as a set, the popula on means for the (k – 1) contrasts across levels are zero. (This is essen ally equivalent to the null hypothesis given in Equa on 15.1.) Repeated-measures ANOVA assumes homogeneity of variances for scores of these C variables. This resembles the homogeneity of variance test you have seen earlier for the between-S ANOVA; the homogeneity of variance null hypothesis is in which each σ2 term represented the popula on variance that corresponded to one group of scores. For the repeated-measures ANOVA, the homogeneity of variance assump on for the C contrasts can be stated as follows (Field, 2018): Other

(15.21) Equality among these variances is called sphericity. If the variances of the contrasts do not differ from one another significantly, then the F ra o and its corresponding p value should provide accurate informa on about the risk for Type I error. If the variances of contrasts are unequal, this violates an assump on required for the repeated-measures ANOVA; the p value may seriously underes mate the true risk for Type I error. SPSS provides a test of whether the sphericity assump on is violated (Mauchly’s sphericity test). Along with that test, values of epsilon (ε) are provided. Values of ε indicate the degree to which the sphericity assump on is violated. If there is no viola on, ε = 1; smaller values of ε indicate greater viola ons of the sphericity assump on. For the independent-samples t test and between-S ANOVA, viola ons of the homogeneity of variance test usually do not cause serious problems (provided that n is at least 30 per group). However, viola ons of the sphericity assump on for repeated measures are common, and viola on of this assump on does cause serious problems and should not be ignored. If Mauchly’s test indicates significant viola ons of the sphericity assump on, you should report versions of the F test that have downwardly adjusted df values; this is discussed further in the context of GLM output. As you have seen earlier, viola ons of assump ons o en create more problems when your sample sizes are small (for example, less than 30) than when your sample sizes are large. When sample sizes are very small, in fact, there just isn’t enough informa on to evaluate assump ons. However, tests to detect viola ons of assump ons usually have much more power to detect even minor departures from assump ons when sample sizes are large. I suggest that when you have a large sample (such as N = 100), you use a small p value (such as p < .001) as the criterion to decide whether a viola on of an assump on is sta s cally significant. If you have a rela vely small sample, such as N = 40, you might use a larger p value (such as p = .10 or .20) as a criterion

for deciding whether a viola on is sta s cally significant. Possible ways to remedy viola ons of sphericity are discussed in the context of the GLM output. 15.7.4 Assump on of No Person-by-Treatment Interac on An example of a viola on of the “no person-by-treatment interac on” assump on would be if Ann shows a decrease in heart rate from baseline to pain, Frank shows an increase, and Erin shows no change (in other words, if their reac ons to the treatments are different). The concept of interac on is discussed extensively in Chapter 16, on factorial ANOVA. You can obtain a test of viola on of this assump on (the Tukey test of nonaddi vity) from the SPSS reliability procedure; see Appendix 15A for further informa on. 15.8 CHOICES OF CONTRASTS IN GLM REPEATED MEASURES As you might have guessed from the preceding sec on, contrasts are an integral part of repeated-measures analyses. An example of contrast in the stress/HR study would be the difference between HR_pain and HR_baseline scores. In the paired-samples t example, this difference was called d. In repeated-measures ANOVA, a more common nota on for this kind of difference is C. If the mean of this contrast score is significantly different from 0, then mean HR during pain is significantly higher than mean HR during baseline. If the overall F for differences among treatments is sta s cally significant, further analysis is needed to evaluate where the significant differences are. It is possible that there is only one significance difference, or there might be several. The SPSS GLM procedure provides contrasts that can be used to evaluate differences. By default, it provides polynomial contrasts (discussed below). (Use of polynomial contrasts is rare in behavioral and social sciences.) The choice of contrast depends on the nature or pa ern of differences you expect to see across the set of means. A pull-down menu within the SPSS GLM procedure provides a place to specify choice of contrast. (One of the contrasts must be selected; however, you can ignore contrast results if none of them correspond to comparisons that are meaningful to you.) The number of contrasts that can be included in a set depends on the number of levels for the within-S factor. If the repeated-measures factor has k levels, a set of contrasts includes (k – 1) comparisons. 15.8.1 Simple Contrasts For a study that compares one control (or untreated) group with treatment groups, the usual ques on is whether each of the treatment groups has a mean that differs significantly from the control group mean. The stress/HR study in the file withins.sav is an example. The first level of the within-S factor corresponds to the baseline (no treatment or control) condi on, and the next three levels correspond to different types of stress. Either the first or last group can be iden fied as the control group when this contrast is used. The three null hypotheses for simple contrasts in the HR/stress study are:

Other

In words, the mean for each treatment condi on is compared with the mean for the baseline or control group. This set of contrasts does not provide informa on about all possible comparisons; for example, it does not include a test of whether μpain = μarithme c. 15.8.2 Repeated Contrasts These compare the mean of each level (except the last) with the mean of the subsequent level. This is useful in situa ons where a response is expected to change from Time 1 to Time 2, Time 2 to Time 3, and so forth. (Some changes could be increases and others decreases.) This type of contrast would be a reasonable choice for the systolic blood pressure study described in Comprehension Ques on 5 (using data from bpstudy.sav). Blood pressure was assessed during increasingly stressful condi ons: Time 1, before informed consent; Time 2, a er informed consent; Time 3, during task rehearsal; and Time 4, during a video-recorded social role play. The null hypotheses for this set of repeated contrasts are as follows: Other

In other words, mean HR during each condi on is compared with mean HR for the me before it. 15.8.3 Polynomial Contrasts Polynomial contrasts o en make more sense when the levels of the repeated-measures factor correspond to different amounts of treatment (rather than different types of treatment). If there are k levels for the repeated-measures factor, then (k – 1) polynomials will be tested. Consider the following hypothe cal study. Levels of the repeated measures correspond to different doses of the same drug (e.g., 10 mg, 20 mg, 30 mg). The research ques on is how pa ent response (such as anxiety) varies across levels of this drug. 15.8.3.1 Linear Trend A linear func on is of the form Y′ = b1X + b0. The values b0 and b are constants (i.e., the intercept and slope in a regression equa on). In this equa on the highest power of the X predictor variable is X1. This equa on defines a straight line (that has no curves) with either a posi ve or nega ve slope, as in Figure 15.10. At least two group means are required to assess linear trend.

15.8.3.2 Quadra c Trend A quadra c func on is of the form y = ax2 + bx + c, or in more familiar regression nota on, Y′ = b0 + b1 × X + b2X2; this includes X1 and X2 terms. This func on has one curve. To test for quadra c trend, there must be at least three group means. The curve may have a hump in the middle (as on the le side of Figure 15.11) or it can be U shaped (as on the right side of Figure 15.11).

These curves describe responses to some drugs. For example, consider a pain/sleep medica on given in doses of 10 mg (for Group 1), 20 mg (for Group 2) and 30 mg (for Group 3). Mean hours of sleep is the dependent variable. A possible outcome corresponds to the “hump”-shaped curve on the le side of Figure 15.11. Pa ents may sleep few hours when they receive only 10 mg (M1 is low). They may sleep more hours when they receive 20 mg (M2 is high). When they receive 30 mg, mean sleep me (M3) may be low. This would imply that the op mal dose of the drug is about 20 mg; lower and higher doses may yield fewer hours of sleep. It is possible for both linear and quadra c trends to be significant; Figure 15.12 shows an example. 15.8.3.3 Higher Order Polynomials In behavioral and social sciences, it is rare that cubic, quar c, or higher order trends are hypothesized or make sense theore cally. When any of these higher order trends turn out to be sta s cally significant, it is o en because a pa ern of means just happens to fall into that pa ern. For a cubic trend (Figure 15.13), the highest power of X is X3, and k, the number of

means, must be at least 4. The func on will have two curves. Quar c (X4) and quin c (X5) trends (Figure 15.14) are possible, but these are not likely to correspond to empirical data.

15.8.4 Other Contrasts Available in the SPSS GLM Procedure Other types of contrasts include devia on contrasts, which compare the mean of each level (except a reference category, such as a control group) with the mean of all levels (grand mean). Helmert contrasts compare the mean of each level of the factor (except the last) with the mean of subsequent levels. Difference contrasts compare the mean of each level (except the first) with the mean of previous levels. These are rarely useful. 15.9 SPSS GLM PROCEDURE FOR REPEATED-MEASURES ANOVA The SPSS GLM procedure provides more op ons than the SPSS reliability procedure for analysis of repeated measures. GLM provides contrasts among group means and a test for the sphericity assump on (it does not provide the test for Person × Treatment interac on available in the reliability procedure). The reliability procedure can do only one-way repeated measures, while the GLM procedure can handle addi onal complex analyses (with or without repeated measures). As noted earlier, GLM is an abbrevia on for general linear model. Almost all the sta s cs you will learn in early sta s cs courses are special cases of the GLM. Note that this differs from the generalized linear model (another op on on the pull-down menu). The SPSS menu selec ons to run repeated-measures ANOVA using GLM appear in Figure 15.15: <Analyze> → <General Linear Model> → <Repeated Measures>. Be careful to select <General Linear Model> (not <Generalized Linear Models>). This series of menu selec ons opens the Repeated Measures Define Factor(s) dialog box shown on the le in Figure 15.16. Ini ally, the default name for the repeated-measures factor is “factor1.” You can replace this with a more meaningful name for the within-S factor. In this example “factor1” was replaced by “stress” by typing that new label in place of “factor1.” The number of levels of the repeated-measures factor (in this example, k = 4) was placed in the box for “Number of Levels” immediately below the name of the repeated-measures factor. Clicking the Add bu on moves the new factor, named stress, into the box or window that contains the list of repeated-measures factors (this box is to the right of and below the label “Number of Levels”). (More than one repeated-measures factor could be specified; however, only one repeated-measures factor is included in this example.) Clicking the Define bu on opens a new dialog box, which appears on the right in Figure 15.16. In this second dialog box, the user tells SPSS the name(s) of the variable(s) that correspond to each level of the repeated- measures factor by clicking each blank name in the ini al list under the heading “Within-Subjects Variables (stress)” and using the right arrow to move each variable name into the list of levels. In this example, the first level of the repeated-measures factor corresponds to the variable named HR_baseline, the second level corresponds to HR_pain, the third level corresponds to HR_arithme c, and the fourth and last level corresponds to HR_roleplay. When all four of these variables have been specified as the levels of the within-S factor, the descrip on of the basic repeated-measures design is complete.

The EM Means (es mated marginal means) bu on was used to open the dialog box on the le in Figure 15.17. The names OVERALL (grand mean) and stress were moved into the “Display Means for” window. (There are other ways to request group means for stress; this method includes 95% confidence intervals.) The Op ons bu on in the Repeated Measures dialog box on the right side of Figure 15.16 was used to open the Repeated Measures: Op ons dialog window on the right in Figure 15.17, and es mates of effect size were requested by checkbox (this provides the values of par al η2). The Plots bu on in the Repeated Measures dialog box on the right side of Figure 15.16 opens the Repeated Measures: Profile Plots window (at le in Figure

15.18); move the name of the within-S factor into the box for “Horizontal Axis,” then click Add. This provides a line graph of cell means for each of the four levels of the stress factor. Make sure the “Include Error bars” op on is checked. The Contrasts bu on in the Repeated Measures dialog box on the right side of Figure 15.16 was used to request simple contrasts among levels of the repeated-measures factor (see Figure 15.19). In this example, simple contrasts were selected from the pull-down menu of different types of contrasts. In addi on, a radio bu on selec on specified the first category (which corresponds to HR during baseline, in this example) as the reference group for all comparisons. In other words, planned contrasts were done to see whether the mean HRs for Level 2 (pain), Level 3 (arithme c), and Level 4 (role play) each differed significantly from the mean HR on Level 1 of the repeated-measures factor. The Change bu on must be clicked for this selec on to replace the default choice of polynomials. Finally, click OK in the main dialog box to run all the requested analyses.

Figure 15.17 SPSS Repeated Measures Procedure: Op ons to Request Means and Effect Sizes

Figure 15.18 Request Profile Plot of Means

Figure 15.19 Request (Simple) for Repeated-Measures Contrasts 15.10 OUTPUT OF GLM REPEATED-MEASURES ANOVA The first part of GLM output for repeated measures is a mul variate analysis of variance (MANOVA) table that is not shown here; ignore this. MANOVA provides an alterna ve way to

test overall significance of differences among means on repeated measures; however, it is a more advanced topic. Examina on of M, SD, and confidence intervals (in Figure 15.20) and the profile plot (in Figure 15.21) provides informa on about the pa ern of outcomes. As in earlier analyses, the highest mean heart rate was observed during the social role play condi on and the lowest mean HR at baseline. SPSS reports a test of the sphericity assump on for repeated measures (see Figure 15.22). Analysts usually hope to see a nonsignificant result for Mauchly’s test of sphericity; in this example, “Sig.” or p = .985. A large p value tells us that the data do not provide evidence that this assump on is violated (a large p value for this test is good news). A large p value tells us that when we go on to look at the ANOVA source tables, we can look just at the first row, labeled “Sphericity Assumed.” For this example, p = .985 for the Mauchly test (Figure 15.22). There is not a significant viola on of the sphericity assump on. What if the p value for the Mauchly test is small, for example, p < .05? This would be evidence that the sphericity assump on is violated. If the sphericity assump on is violated, you could use MANOVA (if you are familiar with it), or you can use the downwardly adjusted df values in either the row labeled “Huynh-Feldt” or the row labeled “Greenhouse-Geisser” in the ANOVA source table. These df are usually not integer values. The value of the F ra o is the same in these lines as for the sphericity assumed condi on. However, the df that are used to evaluate the sta s cal significance of F are downwardly adjusted to compensate for viola on of sphericity. The Greenhouse-Geisser df are obtained by mul plying the sphericity assumed df by the Greenhouse-Geisser ε value that appeared in output for Mauchly’s sphericity test; Huynh-Feldt uses a different, slightly less conserva ve value of ε to adjust the df.

Figure 15.20 SPSS GLM Procedure Output: Mean, SD, and 95% CI for Heart Rate During Stress Condi ons

As in other situa ons where assump ons are tested, it is a good idea to use a larger p value as your criterion for significance for viola on of assump ons (Mauchly’s test) when N is small, for example, p < .20. A smaller p value can be used to judge the significance of Mauchly’s test when N is large, for example, p < .001. The dilemma we have when tes ng assump ons is that a viola on of assump ons is a more serious problem for small-N than for large-N data. For small- N data, we have poor sta s cal power to detect significant viola ons, whereas for large-N data, we have high sta s cal power for tests of viola ons of assump ons. Viola ons of assump ons are generally a greater problem when N is small. The primary output table of interest is “Tests of Within-Subjects Effects” (Figure 15.23). The within-S treatment variable is stress. You will see that there are four rows that provide values of sums of squares, df, and MS for stress. The first row within the stress por on of the table corresponds to the “sphericity assumed” situa on. If Mauchly’s test has a large p value (for example, p > .05), report F using the df numerator value in this line of the table. When the sphericity assump ons for repeated-measures ANOVA are sa sfied, the df for stress = k – 1,

where k is the number of levels for the stress factor. An F ra o has a df term for the numerator (stress) and the denominator (error); to find the df for error, when the sphericity assump on is met, look in the row “Error(stress) Sphericity Assumed”; the df error is (n – 1) × (k – 1), where n is the number of par cipants and k is the number of levels of stress. For the stress/HR example, you could say that the overall F for differences in mean heart rate across the four stress condi ons was sta s cally significant, F(3, 15) = 8.576, p = .001, with par al η2 = .632.

If Mauchly’s test of sphericity had been sta s cally significant, you could choose to report F along with the df for either the Greenhouse-Geisser or the Huynh-Feldt correc on. (The SS terms and F do not change, only the df terms, for these rows of the table.) For example, you could say, “The difference in mean heart rate across all four levels of stress was sta s cally significant, F(2.69, 13.451) = 8.576, p = .002; Greenhouse-Geisser-adjusted df values were used because of viola on of the sphericity assump on.” (No ce that Greenhouse-Geisser df for stress = [k – 1] × Greenhouse-Geisser ε, and df for error = (n – 1) × (k – 1) × Greenhouse-Geisser ε.) In this example, the Huynh-Feldt procedure would not reduce the df; Huynh-Feldt ε was 1. In this example, the overall F for stress is sta s cally significant whether the “sphericity assumed” or “Greenhouse-Geisser” df are used to evaluate the sta s cal significance of F. It is possible to find that F is significant using sphericity assumed df, but not significant when the Greenhouse-Geisser adjustment is made. Keep in mind that viola ons of the sphericity assump on are serious and that when they occur, you should use the adjusted df and p value (either Greenhouse-Geisser or Huynh-Feldt). Under some circumstances it is be er to use MANOVA tests, but those are not discussed here. Another part of the output is the “Tests of Between-Subjects Effects” table (Figure 15.24). These tests are generally not of interest. The null hypothesis here is that the popula on grand mean for heart rate = 0. There can be situa ons where a null hypothesis about the grand mean is of interest, but it is not useful for this example. (If the outcome variable had been number of pounds lost in diet programs, it would be of interest to ask whether the grand mean for this variable differed from 0.) Later in

this chapter, when more complicated design possibili es are briefly noted, you will see that you can add a between-S variable such as sex to a repeated-measures ANOVA. If you did that, the test for differences between male and female HR would appear in the “Tests of Between- Subjects Effects” table. Type III sum of squares is discussed in Appendix 16C in the following chapter, on factorial ANOVA. For this analysis, you do not need to state that Type III sum of squares was used. Recall that for between-S one-way ANOVA, differences between group means could be evaluated in either of these two ways:

 Planned contrasts (for between-S ANOVA, these are user specified by typing in contrast coefficients). In GLM, you make planned contrasts by selec ng one of the op ons from the pull-down menu for types of contrasts (e.g., simple, polynomial).

 Post hoc or protected tests (such as the Tukey honestly significant difference). These are not available for levels of within-S factors. However, you can do a small number of paired-samples t tests as a follow-up to evaluate differences among means on a repeated-measures factor, perhaps using the Bonferroni procedure to limit inflated risk for Type I error; see Sec on 15.11.

In GLM you can select among a set of defined contrasts (discussed in Sec on 15.8). O en you can find a contrast that provides the comparisons you want for your study. For the hypothe cal stress/HR study, I selected simple contrasts; this works well when you have a control group (Figure 15.25). The “reference group” or control group for this study was the first group (baseline condi on). Thus, the simple contrasts compared means for each of the three stress condi ons (Level 2, Level 3, and Level 4) with the mean HR at baseline (Level 1). The difference in mean heart rate between arithme c (Level 3) and baseline (Level 1) was not sta s cally significant, F(1, 5) = 2.231, p = .195, par al η2 = .309. The difference between pain (Level 2) and baseline (Level 1) was sta s cally significant, F(1, 5) = 14.817, p = .012, par al η2 = .748. The difference between role play (Level 4) and baseline (Level 1) was also sta s cally significant, F(1, 5) = 21.977, p = .005, η2 = .815. Addi onal comparisons (for example, comparison of the social role play and pain condi ons) can be obtained by examining paired- samples t tests for any pair of treatments.

Figure 15.25 Tests of Simple Within-Subjects Contrasts 15.11 PAIRED-SAMPLES T TESTS AS FOLLOW-UP Each set of predefined contrasts described in Sec on 15.8 includes only k – 1 contrasts; some mes researchers want to make addi onal comparisons, or different comparisons. In the stress and HR study, we might also want to know whether the different types of stress have significantly different means for HR. Paired-samples t tests can be used to make addi onal contrasts. When numerous significance tests are performed, the risk for Type I error is inflated. The Bonferroni procedure, introduced in Chapter 10, can be used here to correct for inflated risk for Type I error. Steps in the Bonferroni procedure are as follows:

 Count the number of tests that will be performed; let k = the number of tests in the set of paired-samples t tests. For this example, I will use k = 6 as the number of paired- samples t tests.

 Decide on the limit for Type I error that you want to have for the en re set of k significance tests. This is called the “experiment-wise” error rate, and it is denoted EWmα. For this example, I will set EWα = .10.

 Calculate the per comparison α, denoted PCα, as follows: PCα = EWα/k = .10/6 = .0167.  To implement the test, the α level used to evaluate the significance for each of the

individual paired-samples t tests is .0167. Output for the set of six paired-samples t tests (to compare each type of stress with baseline, and to compare each stress condi on with the other two stress condi ons) appears in Figure 15.26. Menu selec ons to obtained paired-samples t tests were shown earlier and are not repeated here. Results could be reported as follows:

Six comparisons were made between means. Mean HR in each stress condi on was compared with baseline, and mean HR in each stress condi on was compared with the other two stress condi ons. Bonferroni-corrected per comparison alphas that corresponded to an experiment- wise risk for Type I error of EWα = .10 were used. The per comparison α for each paired- samples t test was .0167.

15.12 RESULTS Prior to the “Results” sec on, earlier parts of the research report would explain the following: the number and types of stress, the outcome measure, the number of par cipants, and the order in which par cipants experienced the stress condi ons. (This example did not include counterbalancing to control for order effects.) The N of cases in this example (6) is inadequate to evaluate assump ons such as normality of distribu on of heart rate scores, linearity of associa ons between scores across condi ons, and so forth. An actual study would use a larger number of par cipants and would use counterbalancing (explained in a later sec on of this chapter) to control for order effects. With a reasonably large number of par cipants, it would make sense to examine histograms for HR measures at each me (to assess normality of distribu on shape) and to examine boxplots for outliers and to obtain sca erplots for HR_baseline with HR_pain, HR_pain with HR_arithme c, and so on, to assess linearity. Results Preliminary data screening did not indicate problems with assump ons of normality and linearity. Mauchly’s sphericity test did not indicate a significant viola on of the sphericity test. The assump ons for repeated-measures ANOVA appeared to be sa sfied. Therefore, corrected

df for the F test, based on the Huynh-Feldt or Greenhouse-Geisser procedure, were not used. There was a sta s cally significant difference in mean heart rate across the four condi ons in the study (baseline, pain, arithme c, and role play), F(3, 15) = 8.576, p = .001, par al η2 = .632. A er controlling for individual differences, about 63% of the variance in heart rate was associated with type of stress. Six follow-up paired-samples t tests were examined to evaluate which condi ons differed significantly in mean HR (see Figure 15.15). Each of the three stress condi ons was compared with baseline, and the three stress condi ons were compared with one another, as follows:

 Pain versus baseline  Arithme c versus baseline  Role play versus baseline  Arithme c versus pain  Role play versus arithme c  Pain versus role play

Six comparisons were made between means; mean HR in each stress condi on was compared with baseline, and mean HR in each stress condi on was compared with the other two stress condi ons. Bonferroni-corrected per comparison alphas that corresponded to an experiment- wise risk for Type I error of EWα = .10 were used; the per comparison α for each paired-samples t test was .0167. Mean HR was significantly higher in the pain condi on than in the baseline condi on, Md = 9, t(5) = 3.8549, p = .012, two tailed, η2 = .75. Mean HR was also significantly higher during the social role play than during baseline, Md = 11.33, t(5) = 4.688, p = .005, two tailed. None of the other four differences would be judged sta s cally significant using Bonferroni-corrected α’s. Addi onal informa on that should be provided, in text, table, or graph form, includes M and SD for all four stress condi ons and the 95% CI for Md (and possibly CIs for the means of each treatment group). In addi on, we have not yet considered the problem of order effects. 15.13 EFFECT SIZE In repeated-measures ANOVA, as in other ANOVA designs, effect size can be calculated; it will be a par al η2 because variance due to individual differences among persons is par alled out. SPSS reports par al η2 if effect size is requested in the Op ons dialog box for GLM. As noted earlier, for a one-way within-S ANOVA, par al η2 (i.e., an effect size that has the variance due to individual differences among persons par alled out) is found by taking: Other

(15.22) For the results reported in Figure 15.23, SStreatment or SSstress = 488.458 and SSerror = 284.792; therefore, par al η2 = 488.458/(488.458 + 284.792) = .632. Eta squared is interpreted as the propor on of variance in the scores that is predictable from the me or treatment factor (a er individual person differences have been removed from the scores). To say this another

way: A er individual differences in HR are removed from the data, about 63% of the remaining variance was related to type of treatment. 15.14 STATISTICAL POWER In many research situa ons, a repeated-measures ANOVA has be er sta s cal power than a between-S ANOVA. Repeated-measures ANOVA par als out or removes varia on due to stable individual differences among persons (individual differences in HR, in the empirical example presented here) so that these are not included in the error term used to assess differences among treatment group means. The within-S F ra o in Equa on 15.16 typically has a smaller SS*error term than the SSerror term in the divisor for the between-S ANOVA. However, the within-S F ra o also has a smaller df * term for the divisor. When the decrease in error variance (i.e., the difference in size between SS*error and SSerror) is large enough to more than offset the reduc on in degrees of freedom (df * is smaller than df), the repeated-measures ANOVA will have greater sta s cal power than the between-S ANOVA. On the basis of this reasoning, a repeated-measures ANOVA will o en, but not always, provide be er sta s cal power than an independent-samples ANOVA. When a researcher can make an educated guess about the popula on value of the par al η2 effect size, the following table can be used to evaluate the sample size (N, number of par cipants) needed to achieve reasonable sta s cal power. As in previous chapters, you can use the power table for α = .05 in Table 15.2 to look up required sample size as a func on of desired power (usually .80) and assumed popula on par al η2 effect size. The table can also be used to evaluate expected level of power as a func on of assumed effect size and planned sample size.

Later sec ons discuss the need to examine different treatment orders in some studies. When k different treatment orders are included, the total N is usually an integer mul ple of k. For example, if the study includes k = 3 treatment orders and the researcher wants at least 4 persons for each order, the total N of par cipants might be 12. 15.15 COUNTERBALANCING IN REPEATED-MEASURES STUDIES The use of a repeated-measures design can have some valuable advantages. When a researcher uses the same par cipants in all treatment condi ons, poten al confounds between type of treatment and par cipant characteris cs (such as age, anxiety level, and drug use) are avoided.

The same par cipants are tested in each treatment condi on. In addi on, the use of repeated- measures design makes it possible, at least in theory, to iden fy and remove variance in scores that is due to stable individual differences in the outcome variable (such as HR), and this may result in a smaller error term (and a larger t or F ra o) than a between-S design. In addi on, repeated measures can be more cost-effec ve; we obtain more data from each par cipant. This can be an important considera on when par cipants are difficult to recruit or require special training or background or prepara on. However, as noted in the introduc on, the use of repeated-measures design also gives rise to some problems that were not an issue in between-S designs, including order and carryover effects. In the hypothe cal example described so far, each par cipant’s HR was measured four

mes: Time 1: baseline Time 2: pain induc on Time 3: mental arithme c test Time 4: stressful social role play This simple design has a built-in confound between order of presenta on (first, second, third, fourth) and type of treatment (none, pain, arithme c, role play). If we observe the highest HR during the stressful role play, we cannot be certain whether this occurs because this is the most distressing of the four situa ons or because this treatment was the fourth in a series of (possibly increasingly unpleasant) experiences. Another situa on that illustrates poten al problems arising from order effects was provided by a magazine adver sement published several years ago. The adver sers (a company that sold rum) suggested the following taste test: First, taste a shot of whiskey (W); then, taste a shot of vodka (V); then, taste gin (G); and finally, taste rum (R). (Now, doesn’t that taste great?) It is important to unconfound (or balance) order of presenta on and type of treatment when designing repeated-measures studies. This requires that the researcher present the treatments in different orders. The presenta on of treatments in different orders is called counterbalancing. Type of treatment and order of presenta on are being “balanced,” that is, unconfounded. There are several ways this can be done. One method of counterbalancing, called complete counterbalancing, involves presen ng the treatments in all possible orders. This is generally not prac cal, par cularly for studies where the number of treatments (k) is large, because the number of possible orders in which k treatments can be presented is given by k! or “k factorial,” which is k × (k – 1) × (k – 2) × ··· × 1. Even if we have only k = 4 treatments, there are 4! = 4 × 3 × 2 × 1 = 24 different possible orders of presenta on. That would mean we need at least 24 par cipants (one for each possible

order), and perhaps we would want more than one person to experience each order of presenta on, and that would increase required sample size s ll further. It is some mes not necessary to run all possible orders to unconfound order with type of treatment. In fact, for k treatments, k treatment orders can be sufficient, provided that the treatment orders are worked out carefully. In se ng up treatment orders, we want to achieve two goals. First, we want each treatment to occur once in each ordinal posi on. For example, in the liquor-tas ng study, when we set up the order of beverages for different groups of par cipants, one group should taste rum first, one group should taste it second, one group should taste it third, and one group should taste it fourth (last). This ensures that the taste ra ngs for rum are not en rely confounded with the fact that it is the last of four alcoholic beverages. Ideally, we would also want to make sure that each treatment follows each other treatment just once; for example, one group should taste rum first, one group should taste rum immediately a er tas ng whiskey, one group should taste rum a er vodka, and the last group should taste rum a er gin. This controls for possible contrast effects; for example, rum might taste much more pleasant when it immediately follows a somewhat bi er-tas ng beverage (gin) than when it follows a more neutral-tas ng beverage (vodka). Treatment orders that control for both ordinal posi on and sequence of treatments are called La n squares. It can be somewhat challenging to work out a La n square for more than about three treatments; fortunately, textbooks on experimental methods o en provide tables of La n squares. For the four alcoholic beverages in this example, a La n square design would involve the following four orders of presenta on. Par cipants would be randomly assigned to four groups: Group 1 would receive the beverages in the first order, Group 2 in the second order, and so forth. Order 1: WVGR Order 2: VRWG Order 3: RGVW Order 4: GWRV In these four orders, each treatment (such as W, whiskey) appears once in each ordinal posi on (first, second, third, fourth). Also, each treatment, such as gin, follows each other treatment in just one of the four orders. In other words, order is not confounded with type of treatment. In addi on, the orders have been created to avoid having gin always follow the same other beverage; in Order 1, gin follows vodka; in Order 2, gin follows whiskey; in Order 3, gin follows rum; and in Order 4, gin does not follow any other drink. Using a La n square as in this example, it is possible to control for order effects (i.e., to make sure that they are not confounded with type of treatment) using as few as k different orders of presenta on for k treatments. However, there are addi onal poten al problems.

Some mes when treatments are presented in a series, if the me interval between treatments is too brief, the effects of one treatment do not have me to wear off before the next treatment is introduced; when this happens, we say there are “carryover effects.” In the beverage-tas ng example, there could be two kinds of carryover: One, the taste of a beverage such as gin may s ll be s mula ng the taste buds when the next “treatment” is introduced, and two, there could be cumula ve effects on judgment from the alcohol consump on. Carryover effects can be avoided by allowing a sufficiently long me interval between treatments so that the effect of each treatment has me to “wear off” before the next treatment is introduced or to do something to try to neutralize treatment effects (at a wine tas ng, for example, people may rinse their mouths with water or eat bread in between wines to clear the palate). In any study where par cipants perform the same task or behavior repeatedly, there can be other changes in behavior as a func on of me. If the task involves skill or knowledge, par cipants may improve their performance across trials with prac ce. If the task is arduous or dull, performance on the task may deteriorate across trials due to boredom or fa gue. In many studies where par cipants do a series of tasks and are exposed to a series of treatments, it becomes possible for them to see what features of the situa on the researcher has set out to vary systema cally; they may guess the purpose of the experiment, and they may try to do what they believe the researcher expects. Some self-report measures involve sensi za on; for example, when people fill out a physical symptom checklist over and over, they begin to no ce and report more physical symptoms over me. Measures can be reac ve; they can change the behavior they are only supposed to measure.1 Finally, if the repeated-measures study takes place over a rela vely long period of me (such as months or years), matura on of par cipants may alter their responses, outside events may occur in between the interven ons or treatments, and many par cipants may drop out of the study (they may move, lose interest, or even die). Campbell and Stanley’s (2001) classic monograph on research design points out these and many other poten al problems that may make the results of repeated-measures studies difficult to interpret, and it remains essen al reading for researchers who do repeated- measures research under less than ideal condi ons. 15.16 MORE COMPLEX DESIGNS Design elements from various chapters can o en be combined to create more complex designs. The ini al GLM Repeated Measures dialog box (Figure 15.27) has room to add other variables. A categorical variable can be entered into the “Between-Subjects Factor(s)” pane; sex is an example. The results of an analysis with sex added would tell us whether women and men differ in their mean responses to four different types of stress. The window for covariates can be used to enter a quan ta ve variable (such as anxiety scores). Assuming that the outcome variable (heart rate) is related to the covariate variable (anxiety), this would yield results in which we have par alled out not only person effects but also anxiety effects. You can think of most of the analyses you learn during early courses in sta s cs as building blocks that can be combined in various ways. Most published research reports include more than two variables. You need a solid founda on to understand how two variables are related (bivariate analyses). What you learn about bivariate analysis con nues to be relevant as you go on to more advanced techniques.

Figure 15.27 Addi onal Possible Variables for GLM Repeated Measures 15.17 SUMMARY Repeated-measures designs can offer substan al advantages to researchers. A repeated- measures design can make it possible to obtain more data from each par cipant and to sta s cally control for stable individual differences among par cipants (and remove that variance from the error term) when tes ng treatment effects. However, repeated-measures designs raise several problems. The study must be designed in a manner that avoids confounding order effects with treatment. Repeated-measures ANOVA involves assump ons about the pa ern of covariances among repeated-measures scores, and if those assump ons are violated, it may be necessary to use corrected degrees of freedom to assess the significance of F ra os (or to use MANOVA rather than repeated-measures ANOVA). Finally, the paired- samples t test and repeated-measures ANOVA may result in different conclusions about the nature of change than alterna ve methods covered in more advanced treatments (such as analysis of covariance [ANCOVA] using pretest scores as a covariate). When the choice of sta s cal analysis makes a substan al difference, the researcher needs to consider carefully whether repeated-measures ANOVA or ANCOVA provides a be er way of assessing change across me. Other methods, such as mul level modeling, provide flexible ways to assess changes in behaviors of individuals across me (Grimm & Ram, 2016).

CHAPTER 16 FACTORIAL ANALYSIS OF VARIANCE 16.1 RESEARCH SITUATIONS WHERE FACTORIAL DESIGN IS USED Factorial analysis of variance is used in research situa ons where two or more group membership variables (called “factors”) are combined and scores are obtained for a quan ta ve Y outcome variable such as hos lity. This chapter covers only between-subjects (between-S) designs; that is, each par cipant contributes a score in only one cell (each cell represents one combina on of the A and B treatment levels). For now, we will assume that the number of par cipants in each cell or group (n) is equal for all cells. The number of scores within each cell is denoted by n; thus, the total number of scores in the en re data set, N, is found by compu ng the product n × a × b. An example is a study in which Factor A is sex (with Groups A1 = female and A2 = male) and Factor B is level of crowding (with Groups B1 = low crowding, B2 = medium crowding, and B3 = high crowding). When all groups for Factor A are tested within all the condi ons for Factor B, there are six groups, as shown in Table 16.1. This is factorial design that is completely crossed. Factorial means that we have more than one factor. Completely crossed means that all levels of Factor A are combined with all levels of Factor B. The factorial designs in this chapter are completely between-S; that is, each person is tested in only one of these six groups. In this example, a researcher might recruit 30 male and 30 female par cipants. Men could be randomly assigned to a lab situa on with low, medium, or high crowding, and women would also be randomly assigned to low, medium, and high crowding. This would result in a total N of 60, with n = 10 persons in each of the six cells or groups iden fied in Table 16.1. For the Y variable hos lity we can ask, Are men more hos le than women? Are people more hos le in crowded environments? And do men and women react differently to crowding? For example, women may not show much increase in hos lity when exposed to high levels of crowding, whereas men may respond to high crowding with large increases in hos lity.

The terminology to describe factorial studies is as follows. Each of the categorical predictor variables is called a factor. The groups iden fied by each factor are called levels. Levels of a

factor can represent different naturally occurring group memberships (such as sex or age groups), different types of treatment, or different dosages of the same treatment. The number of levels for Factor A is denoted a; the number of levels for Factor B is called b. A factorial is o en described by naming the number of levels; for example, Table 16.1 shows a 2 × 3 factorial design with factors sex and crowding. More generally, we can describe a study as an a × b factorial. Addi onal Factors C (room temperature, cold vs. warm), D ( me in semester, early vs. late), and so forth, could be added; the study would be an a × b × c × d factorial. It is rare for studies to include large numbers of factors with large numbers of levels.1 Most examples in this chapter are 2 × 2 factorials. A factorial design is called two way if it includes two factors, three way if it includes three factors, and so forth. Factorial analysis of variance (ANOVA) is o en used in experiments (where at least one of the factors is a variable manipulated by the researcher), but factorial ANOVA can also be applied in nonexperimental research situa ons, that is, research situa ons where all factors correspond to naturally occurring groups rather than to different treatments administered by the researcher. Factorial ANOVA is a generaliza on of one-way ANOVA. In effect, we do three analyses simultaneously: a one-way ANOVA for Factor A, a one-way ANOVA for Factor B, and assessment of possible A × B interac on. 16.2 QUESTIONS IN FACTORIAL ANOVA The hypothe cal example of a 2 × 2 factorial design uses data in the file socialsupportstress.sav. In this imaginary study, Factor A is a naturally occurring group membership: amount of social support people receive in daily life (A1 = low social support, A2 = high social support). Factor B is also a naturally occurring group membership: amount of stress persons experience in daily life (B1 = low stress, B2 = high stress). (Either or both these variables could also be manipulated in a lab se ng; for example, da ng partner present or absent could be a social support manipula on, and a researcher could assign each par cipant to either a low- or high-stress task.) The quan ta ve outcome variable (Y) is the number of physical illness symptoms for each par cipant. There are n = 5 scores in each of the four groups, and thus the overall number of par cipants in the en re data set N = a × b × n = 2 × 2 × 5 = 20 scores. Table 16.2 reports a hypothe cal set of means for the groups included in this 2 × 2 factorial design. The research ques ons focus on the pa ern of means. Are there substan al differences in mean level of symptoms between people with low versus high social support? Are there substan al differences in mean level of symptoms between groups that report low versus high levels of stress? Is there a par cular combina on of circumstances (e.g., low social support and high stress) that predicts a much higher level of symptoms? As in one-way ANOVA, F ra os are used to assess whether differences among group means are sta s cally significant. Because the factors in this hypothe cal study correspond to naturally occurring group memberships rather than experimentally administered treatments, this is a nonexperimental or correla onal design, and results cannot be interpreted as evidence of causality.

In this chapter, we will see that a factorial ANOVA of the data in socialsupportstress.sav provides three significance tests: a significance test for the main effect of Factor A (social support), a significance test for the main effect of Factor B (stress), and a test for an interac on between the A and B factors (social support by stress). The ability to detect poten al interac ons between predictors is one of the major advantages of factorial ANOVA compared with one-way ANOVA designs. It is generally useful to ask how each new analysis compares with earlier, simpler analyses. What informa on do we obtain from a factorial ANOVA that we cannot obtain from a simpler one-way ANOVA? Consider for a moment what we could see in the data if we do not know anything about par cipant stress. If we have no informa on about stress (or choose to ignore the stress variable), we can only do a one-way ANOVA to compare mean symptoms across levels of social support. Results of that one-way ANOVA appear in Figure 16.1. This provides a baseline, something we can compare with the results of the two-way factorial. Figure 16.1 presents results for a one-way ANOVA for the data in socialsupportstress.sav that tests whether mean level of symptoms differs across levels of social support. The difference in symptoms in this one-way ANOVA, a mean of 8.30 for the low–social support group versus a mean of 5.00 for the high–social support group, was not sta s cally significant at the conven onal α = .05 level: F(1, 18) = 3.80, p = .067. Adding a second factor to ANOVA can be helpful in two ways. Adding a second factor makes it possible to detect a poten al interac on effect (in this example, an interac on between social support and stress as predictors of symptoms). In addi on, the error term (SSwithin) used to compute the divisor for F ra os can decrease substan ally when one or more addi onal factors or variables are included in an analysis. Later, when we obtain results for the two-way factorial, we can evaluate how adding the stress factor to the analysis changes our understanding of the effect of social support on symptoms. In the one-way ANOVA, we examine the effect of social support on symptoms, not controlling for

stress (not taking stress into account). In the factorial ANOVA, we examine the effect of social support on symptoms while sta s cally controlling for stress (taking stress into account).

16.3 NULL HYPOTHESES IN FACTORIAL ANOVA To assess the pa ern of outcomes in a two-way factorial ANOVA, we will test three separate null hypotheses. 16.3.1 First Null Hypothesis: Test of Main Effect for Factor A Other μ μ The first null hypothesis (main effect for Factor A) is that the popula on means on the quan ta ve Y outcome variable are equal across all levels of A. We will obtain informa on relevant to this null hypothesis by compu ng an SS term that is based on the observed sample means for Y for the A1 and A2 groups. (Of course, the A factor may include more than two groups.) If these sample means are “far apart” rela ve to within-group variability in scores, we conclude that there is a sta s cally significant difference for the A factor. As in one-way ANOVA, we evaluate whether sample means are far apart by examining an F ra o that compares MSbetween (in this case, MSA, a term that tells us how far apart the means of Y are for the A1 and A2 groups) with the mean square (MS) that summarizes the amount of variability of scores within cells, MSwithin (also called MSerror). To test the null hypothesis of no main effect for the A factor, we will compute an F ra o, FA: Other

16.3.2 Second Null Hypothesis: Test of Main Effect for Factor B Other

16.3.3 Third Null Hypothesis: Test of the A × B Interac on This null hypothesis can be wri en algebraically, but that requires understanding the algebraic model for ANOVA described in Appendix 16D. In prac ce it is easier to state this null hypothesis verbally: Other

To evaluate possible presence of an A × B interac on, examine the pa ern of means in the cells. This can be done using a line graph that represents cell means. Procedures to set up line graphs appear later in this chapter. When there is no A × B interac on, the lines in a graph of the cell means are parallel, as shown in Figure 16.2 (le ). When an interac on is present, the lines in the graph of cell means are not parallel, as in Figure 16.2 (right). Robert Rosenthal (1966) aptly described interac on as “different slopes for different folks.” If the lines in a graph of cell means are not parallel (as in Figure 16.2, right), and if the F ra o that corresponds to the interac on is sta s cally significant, we can say that Factor A (social support) interacts significantly with Factor B (stress) to predict different levels of symptoms. We can also say that social support moderates the associa on between stress and symptoms.

The line graphs in Figure 16.2 represent two different possible outcomes for the hypothe cal study of social support and stress as predictors of symptoms. If there is no interac on between social support and stress, then the difference in mean symptoms between the high-stress (B2)

and low-stress (B1) groups is the same for the A1 (low social support) group as it is for the A2 (high social support) group, as illustrated in the le panel of Figure 16.2. In this graph, both the A1 and A2 groups showed the same difference in mean symptoms between the low- and high- stress condi ons; that is, mean symptom score was 5 points higher in the high-stress condi on (B2) than in the low-stress condi on (B1). Thus, the difference in mean symptoms across levels of stress was the same for both low– and high–social support groups in the le panel of Figure 16.2. (If we saw this pa ern of means in data in an experimental factorial design, where the B factor represented a stress interven on administered by the experimenter, we might say that the effect of an experimentally manipulated increase in stress was the same for the A1 and A2 groups.) The 5-point difference between means corresponds to the slope of the line that we obtain by connec ng the points that represent cell means. On the other hand, the graph in the right panel of Figure 16.2 shows a possible interac on between social support and stress as predictors of symptoms. For the low–social support group (A1), there was a 10-point difference in mean symptoms between the low- and high-stress condi ons. For the high–social support group (A2), there was only a 1-point difference between mean symptoms in low- versus high-stress condi ons. Thus, the difference in mean symptoms across levels of stress was not the same for the two social support groups. An F test is needed to evaluate whether any departure from parallel lines no ced in the graph is sta s cally significant (i.e., whether the interac on is sta s cally significant). When a substan al interac on effect is present, the interpreta on of results o en focuses primarily on the interac on rather than on main effects. 16.4 SCREENING FOR VIOLATIONS OF ASSUMPTIONS The assump ons for a factorial ANOVA are essen ally the same as those for a one-way ANOVA:

1. Scores for the Y dependent variable must be quan ta ve. 2. Within each group (and cell), scores for the Y variable should be approximately normally

distributed with no extreme outliers. This can be assessed by examining separate histograms and boxplots for each cell.

3. In this chapter, it is assumed that the design is completely between-S (i.e., each par cipant provides data in one group or cell of the design). When data involve repeated measures or matched samples, we need to use repeated-measures analy c methods that take the correla ons among scores into account such as one-way repeated- measures ANOVA. As in other independent groups or between-S designs, we assume that the scores are obtained in a manner that also leads to independent observa ons within groups.

4. We assume that popula on variances are reasonably equal across groups. SPSS provides an op onal test (the Levene test) to assess whether this assump on of homogeneity of variance is violated. Factorial ANOVA is fairly robust to viola ons of the normality assump on and the homogeneity of variance assump on unless the numbers of cases in the cells are very small and/or unequal.

5. We also assume, for the moment, that the number of observa ons is equal in all cells. Appendix 16B discusses handling unequal n’s when calcula ng row means and column means from cell means; Appendix 16C discusses computa on of SS, MS, and F when cell n’s are unequal.

16.5 HYPOTHETICAL RESEARCH SITUATION The data in socialsupportstress.sav are hypothe cal results of a survey on social support, stress, and symptoms. There are scores on two predictor variables: Factor A, level of social support (1 = low, 2 = high) and Factor B, level of stress (1 = low, 2 = high). The dependent variable was number of symptoms of physical illness. Because the numbers of scores in the cells are equal across all four groups in this study (each group has n = 5 par cipants), this study does not raise problems of confounds between factors; it is an orthogonal factorial ANOVA (see Appendix 16C for discussion of calcula on of SS terms in nonorthogonal factorial ANOVA). We can ask three ques ons about these data. Do people with low social support have a different mean level of symptoms than people with high social support (Factor A)? Do people experiencing low levels of stress report a different mean level of symptoms than people experiencing high levels of stress (Factor B)? Does something different happen within groups than you might expect just from the main effects of Factor A and B (the A × B interac on)? If people who have high social support respond differently to high stress than people with low social support, this is an example of an interac on. On the basis of past research on social support, there are two different hypotheses about the combined effects of social support and stress on physical illness symptoms. The first theory corresponds to a no-interac on model (this can also be called a main effects–only model). The first theory predicts that stress and social support each independently influence symptoms but that there is no interac on between these predictors. Another way to say this is that stress increases symptoms for the same amount in the low– versus high–social support groups. The second theory predicts an interac on between social support and stress. Some researchers hypothesize that people who have high levels of social support are protected against, or buffered from, the effects of stress. On the basis of this theory we expect to see high levels of physical symptoms associated with higher levels of stress only for persons who have low levels of social support and not for persons who have high levels of social support. Examining the pa ern of cell means provides evidence that might be consistent with Theory 1 or Theory 2. The pa erns of cell means that would correspond to these two different hypotheses appear in Figure 16.2. The no-interac on model predicts that symptoms decrease as a func on of increasing social support and symptoms increase as a func on of increases in stress. The no-interac on model assumes that no ma er what level of social support a person has, the difference in physical symptoms between people with low and high stress should be about the same. The pa ern of cell means that would be consistent with the predic ons of this direct-effects (no-interac on) hypothesis is shown in Figure 16.2 on the le .

The buffering hypothesis about the effects of social support is an example of an interac on effect. The buffering hypothesis predicts that people with low levels of social support should show more symptoms under high stress than under low stress. However, some theories suggest that a high level of social support buffers (or protects against) the impact of stress, and therefore, people with high levels of social support are predicted to show li le or no difference in symptoms as a func on of level of stress. This is an example of Rosenthal’s (1966) “different slopes for different folks.” For the low–social support group, there is a large increase in symptoms when you contrast the high versus low stress levels. For the high–social support group, a much smaller change in symptoms is predicted between the low- and high-stress condi ons. A pa ern of cell means that would be consistent with the buffering (interac on) hypothesis is shown in Figure 16.2 on the right. A factorial ANOVA of the data provides the informa on needed to judge whether there is a sta s cally significant interac on between stress and social support as predictors of symptoms and, also, whether there are significant main effects of stress and social support on symptoms. 16.6 COMPUTATIONS FOR BETWEEN-S FACTORIAL ANOVA The goal is to summarize informa on about the pa ern of scores on Y in the data. To begin, we need to find the grand mean of Y, the Y mean for each level of A, the Y mean for each level of B, and the Y mean for each of the cells. All subsequent analyses are based on assessment of the differences among these means (compared with the variances of scores within cells). Recall that for a one-way ANOVA with Factor A, the par on of SStotal for the Y outcome variable is as follows. Three different names can be used to describe varia on of scores within cells in ANOVA (SSwithin, SSerror, and SSresidual) to describe varia on of scores within cells. Because these are all common designa ons, you should recognize that they refer to the same thing. Other

(16.1) Varia on among Y scores is divided into an SS that is associated with or predictable from the A factor variable and an SS that is not associated with or predictable from the A factor. When we add a B factor, we also need to add an A × B interac on term to the analysis. SStotal for an A × B factorial design is par oned or divided into the following components: Other

(16.2) To calculate SS terms, as in a one-way ANOVA, you begin by finding values of M for every group. There are numerous ways to define groups in this situa on. We need to examine the A1 versus A2 group means, the B1 versus B2 group means, and the differences among the four cell means. (You are not likely to do all this computa on by hand; however, it is important to understand what informa on each MS and F ra o provides.) For example:

 Find the grand mean for all Y scores in the en re data set (denoted MY), and SStotal, and

N (the total number of scores in the study).  The A factor divides the data into two groups: people in the A1 condi on versus people

in the A2 condi on. Find M for each of these groups of scores and denote these means MA1 and MA2 (these are also called “row means”). More generally, when a is the number of levels for Factor A, you will need to find means for each of the a groups. If a = 4, then you need es mates for MA1, MA2, MA3, and MA4.

 The B factor splits the data into two groups: people in the B1 condi on versus people in the B2 condi on. Find M for each of these groups; these column means are denoted MB1 and MB2. More generally, when b is the number of levels for Factor B, you need to find means for each of the b groups. If b = 3, then you need MB1, MB2, and MB3. MSB is informa on about the magnitude of differences among these B means. If MSB is large compared with MSerror, this is evidence inconsistent with the null hypothesis that corresponding popula on means are equal.

 Row means and column means are some mes also called marginal means because they are in the margins of the table.

If Factor A has a levels and Factor B has b levels, there are a × b cells. In this example, A has two levels and B has two levels, so there are 2 × 2 = 4 cells. Obtain M, SS, and n for each cell in the factorial design. Each cell corresponds to a combina on of levels of factors A and B, and subscripts are used to iden fy the mean for each cell or combina on (i.e., MA1B1, MA1B2, MA2B1, and MA2B2). Factorial ANOVA provides answers to three ques ons:

1. Is there a main effect for Factor A? Do the means for the A1 and A2 groups differ significantly? This is like a one-way ANOVA for only Factor A. As in previous ANOVA situa ons, we obtain informa on about differences among means by obtaining values for those means and an SS value within each A group. On the basis of SSA and dfA, we compute MSA. We set up an F ra o to test H0: μA1 = μA2 by dividing MSA by MSerror. Essen ally, we do a one-way ANOVA for Factor A, but at the same me, we examine differences of means across levels of B and the cells.

2. Do the means for the B1 and B2 groups differ significantly? This difference is called a main effect for Factor B. This is like a one-way ANOVA for only Factor B. We evaluate the difference among B group means by obtaining values for those means and an SS value within each B group. On the basis of SSB and dfB, we compute MSB. We set up an F ra o to test H0: μB1 = μB2 by dividing MSB by MSerror. Essen ally, this is a one-way ANOVA for Factor B, but at the same me, we also examine differences of means across levels of A and the cell means.

3. We then ask about the pa ern of differences among the four cell means. This pa ern of means provides informa on about a poten al interac on between Factors A and B.

A factorial ANOVA provides informa on about three ques ons:

1. Is the A × B interac on sta s cally significant? If so, how strong is the interac on, and what is the nature of the interac on? O en, when there is a sta s cally significant

interac on, the descrip on of outcomes of the study focuses primarily on the nature of the interac on.

2. Is the A main effect sta s cally significant? If so, how strong is the effect, and what is the nature of the differences in means across levels of A?

3. Is the B main effect sta s cally significant? If so, how strong is the effect, and what is the nature of the differences in means across levels of B?

In an orthogonal factorial design, the outcomes of these three ques ons are independent; that is, you can have any combina on of yes/no answers to the three ques ons about significance of the A × B interac on and the A and B main effects. Factorial design is orthogonal if the n’s are equal across all cells or unequal in a way that does not confound levels of A with levels of B. For beginning students, the SPSS default method of computa on of SS terms corrects for confounds automa cally. Advanced students should see Appendix 16C for further discussion. Equal n’s in cells are highly desirable, but factorial ANOVA can be performed even if n’s are unequal. Each of the three ques ons requires essen ally the same informa on as a one-way ANOVA. For each effect, it is customary to present an F test (to assess sta s cal significance). If F is nonsignificant, usually there is no further discussion of means of a factor. If F is significant, then some indica on of effect size (such as η2) should be presented, and the magnitudes and direc ons of the differences among group means should be discussed and interpreted. In par cular, note whether obtained differences between group means are in the predicted direc on. Follow-up tests to compare means across levels of factors may be useful for factors that have more than two levels. Analysis of simple main effects is introduced as an addi onal follow-up method. 16.7 COMPUTATION OF SS AND DF IN TWO-WAY FACTORIAL ANOVA Sums of squares and degrees of freedom for ANOVA are generally obtained using computerized calcula ons such as the SPSS GLM (general linear model) procedure. It can be useful to consider what informa on about scores is included in these computa ons, but instructors who only want students to be able to interpret computer output may wish to skip this sec on. The formulas for by-hand computa ons of a two-way factorial ANOVA are o en presented in a form that minimizes the number of arithme c opera ons needed, but in that form, it is o en not very clear how pa erns in the data are related to the numbers you obtain. A spreadsheet approach to the computa on of sums of squares makes it clearer how each SS term provides informa on about the different (theore cal) components that make up the scores. The opera ons described here yield the same numerical results as the formulas presented in introductory sta s cs textbooks. We can use three subscripts to denote each individual score. The first subscript (i) tells you which level of Factor A the score belongs to, the second subscript (j ) tells you which level of Factor B the score belongs to, and the third subscript (k) tells you which par cipant within that group the score belongs to. In general, an individual score is denoted by Yijk. For example, Y215 would indicate a score from Level 2 of Factor A, Level 1 of Factor B, for the fi h subject within that group or cell.

The subscripts in the SS formulas make explicit which score values and group means are included in each computa on. For SStotal, the ijk subscript makes it explicit that the sum includes the squared devia on of every score from the grand mean, summed across all rows (i), columns (j), and members within each cell (k). Appendix 16E shows values of all the devia ons and squared devia ons that are included in the sums in Equa ons 16.3, 16.4, and 16.5. Appendix 16E also explains the way these equa ons differ from computa onal formulas for SS in some other textbooks. Although the equa ons appear different, they yield the same numerical results. First consider how to find SStotal, which describes the total varia on of all individual Y scores in the study rela ve to the grand mean of scores (MY) in the study: Other

(16.3) This equa on should look familiar by now. You saw similar equa ons for SS terms used to compute sample variance, s2, and in one-way ANOVA. SStotal provides informa on about the total varia on in Y scores for all the par cipants in the study. The next two SS terms (SSA and SSB) tell us what part of that total varia on of scores is due to differences in mean Y values across levels of Factor A (social support), and levels of Factor B (stress). Other

(16.4) SSA provides informa on about the variance (or differences) among means for levels of A. If SSA is large, the Y means for levels of A are far apart. Other

(16.5) SSB provides informa on about the variance among means for levels of B. If SSB is large, the Y means for levels of B are far apart. The SS term for within, error, or residual is based on the devia ons of individual scores from their cell means. This can be wri en as in Equa on 16.6: Other

(16.6) It may be easier to think about finding SSwithin in two steps. In the first step, we can find an SS term for each cell, based on the set of n scores within each of the a × b cells. Then, we can add these SS terms across all the cells. For example, SSAB11 is found by obtaining the devia ons of the n individual scores in cell AB11 from the AB11 cell mean, squaring these devia ons, and

summing them. More generally, SSABij is found by obtaining the devia ons of the n individual scores in cell ABij from the ABij cell mean, squaring these devia ons, and summing them. In the social support and stress ANOVA example, there are two levels of social support and two levels of stress, so there is one SS term for each of the following four cells: SSAB11, SSAB12, SSAB21, and SSAB22. The total value of SSerror is the sum of these four separate within-cell SS values (i.e., SSerror = SSAB11 + SSAB12 + SSAB21 + SSAB22). The corresponding MS term (SSerror/dferror = MSerror) is used as the divisor for F ra os for A, B, and A × B. If SSwithin is large, then there is a lot of varia on in scores among persons who are in the same combina on of treatment condi ons. This varia on cannot be explained by Factor A or Factor B or their interac on. This varia on represents experimental error. In the social support and stress example, this within-group varia on tells us how much people who have the same levels of support and stress differ from one another in symptoms. They may differ on immune system func on, vaccina ons, personality characteris cs such as anxiety and neuro cism, me when they were asked to rate symptoms, and so forth. Experimental error can be reduced, at least in theory, by trying to hold these other variables constant or at least limit their influence. For instance, we can collect data from everyone during the same part of the year and perhaps include only par cipants who have had flu shots. The minimum possible value for SSerror would be 0, if all persons within each group had iden cal scores. That, of course, will not happen in real-world research. Researchers generally want one or more of the terms that correspond to predictor variables and interac ons (SSA, SSB, and SSA×B) to be large and SSerror to be small. Because SStotal = SSA + SSB + SSA×B + SSwithin, a er values of other SS terms are known, it is easy to find SSA×B by subtrac on. (SSA×B can be calculated directly; see Appendix 16E for details.) Other

(16.7) To find the degrees of freedom that correspond to each sum of squares, we use the informa on about the number of levels of the A factor, a; the number of levels of the B factor, b; the number of cases in each cell of the design, n; and the total number of scores, N:

The degrees of freedom for a factorial ANOVA are also addi ve, so you should check that N – 1 = (a – 1) + (b – 1) + (a – 1)(b – 1) + ab(n – 1). To find the mean square for each effect, divide the sum of squares by its corresponding degrees of freedom; this is done for all terms, except that (conven onally) a mean square is not reported for SStotal:

The F for each effect (A, B, and their interac on) is found by dividing the mean square for that effect by mean square error.

The following numerical results were obtained for the data in socialsupportstress.sav:

For each sum of squares, we calculate the corresponding mean square by dividing by the corresponding degrees of freedom:

Finally, we obtain an F ra o to test each of the null hypotheses by dividing the MS for each of the three effects (A, B, and A × B) by MSwithin:

Research reports used to include all the intermediate computa ons for factorial ANOVA summarized as a source table, as in Table 16.3. (Source refers to sources of variance.) Most journal ar cles no longer include intermediate results such as SS and MS. Reports typically include the F ra o and its associated df and p values, effect size es mates such as η2, tables or graphs of cell means or row and column means, with confidence intervals (CIs) for each es mated mean, and results of any follow-up analyses. The source tables produced by the GLM procedure in SPSS contain addi onal rows that are not shown in Table 16.3 and not usually included in research reports. For example, SPSS reports a sum of squares for the combined effects of A, B, and the interac on of A and B; this combined test (both main effects and the interac on effect) is rarely reported in journal ar cles. There is a difference in the way SPSS (and some other programs) labels SStotal compared with the nota on used here and in most other sta s cs textbooks. The term that is generally called SStotal is labeled “Corrected Total” in the SPSS output. The term that SPSS labels “SStotal” was calculated by taking Σ(Yijk – 0)2, that is, the sum of the squared devia ons of all scores from 0; this sum of squares can be used to test the null hypothesis that the grand mean for the en re study equals 0. This term is usually not of interest. When you read the source table from the SPSS GLM procedure, you can generally ignore the rows that are labeled “Corrected Model,” “Intercept,” and “Total.” Use the SS value in the “Corrected Total” row that SPSS labels as your value for SStotal for by-hand computa on of η2 effect size es mates.

16.8 EFFECT SIZE ESTIMATES FOR FACTORIAL ANOVA For each of the sources of variance in a two-way factorial ANOVA, simple η2 effect size es mates can be computed from the sums of squares:

For the socialsupportstress.sav data, on the basis of the SS and df values above, the effect size η2 for each of the three effects can also be computed by taking the ra o of SSeffect to SStotal:

Note that the η2 values above correspond to unusually large effect sizes. These hypothe cal data were inten onally constructed so that the effects would be large. When effect sizes are requested from the SPSS GLM procedure, a par al η2 is reported for each main effect and interac on instead of the simple η2 effect sizes that are defined by Equa ons

16.20 to 16.22. For a par al η2 for the effect of the A factor, the divisor is SSA + SSerror instead of SStotal; that is, the variance that can be accounted for by the main effect of B and the A × B interac on is removed when the par al η2 effect size is calculated. The par al η2 that describes the propor on of variance that can be predicted from A (social support) when the effects of B (stress) and the A × B interac on are sta s cally controlled (by being accounted for in the analysis) is calculated as follows:

Note that the simple η2 tells us what propor on of the total variance in scores on the Y outcome variable is predictable from each factor in the model, such as Factor A. The par al η2 tells us what propor on of the remaining variance in Y outcome variable scores is predictable from the A factor a er the variance associated with other predictors in the analysis (such as main effect for B and the A × B interac on) has been removed. These par al η2 values are typically larger than simple η2 values. Note that eta squared is only a descrip on of the propor on of variance in the Y outcome scores that is predictable from a factor, such as Factor A, in the sample. It can be used to describe the strength of the associa on between variables in a sample, but eta squared tends to overes mate the propor on of variance in Y that is predictable from Factor A in some broader popula on. Other effect size es mates, such as omega squared (ω2, described by Hays, 1994), may provide es mates of the popula on effect size. However, eta squared is more widely reported than omega squared as a sample effect size in journal ar cles, and eta squared is more widely used in sta s cal power tables. 16.9 STATISTICAL POWER As in other ANOVA designs, the basic issue is the minimum number of cases required in each cell (to test the significance of the interac on) or in each row or column (to test the A and B main effects) to have adequate sta s cal power. If interac ons are not predicted or of interest, then the researcher may be more concerned with the number of cases in each level of the A and B factors rather than with cell size. To make a reasonable judgment about the minimum number of par cipants required to have adequate sta s cal power, the researcher needs to make an educated guess about popula on effect size. If comparable past research reports F ra os, these can be used to compute simple es mated effect sizes (η2), using the formulas provided above. Table 16.4 (adapted from Jaccard & Becker, 2009) can be used to decide on a reasonable minimum number of scores per group using eta squared as the index of popula on effect size. These tables are for significance tests using an α level of .05. (Jaccard and Becker also provided sta s cal power tables for other levels of alpha and for other degrees of freedom.) The usual minimum level of power desired is .80. To use the sta s cal power table, the researcher needs to know the degrees of freedom for the effect to be tested and needs to assume a popula on eta squared value. Suppose that a researcher wants to test the significance of dosage level of

caffeine with three dosage levels using α = .05. The df for this main effect would be 2. If past research has reported effect sizes on the order of .15, then from Table 16.4, the researcher would locate the por on of the table for designs with df = 2, the column in the table for η2 = .15, and the row in the table for power = .80. The n given by the table for this combina on of df, η2, and power is n = 19. Thus, to have an 80% chance of detec ng an effect that accounts for 15% of the variance in scores for a factor that compares three groups (2 df ), a minimum of 19 par cipants are needed in each group (each dosage level of caffeine). Note that a larger number of cases may be needed to obtain reasonably narrow confidence intervals for each es mated group mean; it is o en desirable to have sample sizes larger than the minimum numbers suggested by power tables.

16.10 FOLLOW-UP TESTS 16.10.1 Nature of a Two-Way Interac on One of the most common ways to illustrate a significant interac on in a two-way factorial ANOVA is to graph the cell means (as in Figure 16.2) or to present a table of cell means (as in Table 16.2). Visual examina on of a graph makes it possible to “tell a story” about the nature of the outcomes. Given a 2 × 2 ANOVA in which Factor A represents two levels of social support and Factor B represents two levels of stress, a researcher might follow up a significant two-way interac on with an “analysis of simple main effects.” For example, the researcher could ask whether within each level of B, the means for cells that represent two different levels of A differ significantly (or, conversely, the researcher could ask whether separately, within each level of A, the two B group means differ significantly). If these comparisons are made on the basis of a priori theore cal predic ons, it may be appropriate to do planned comparisons; the error term used for these could be the MSwithin for the overall two-way factorial ANOVA. If comparisons among cell means are made post hoc (without any prior theore cally based predic ons), then it may be more appropriate to use protected tests (such as the Tukey honestly significant difference test) to make comparisons among means. The APA Task Force on Sta s cal Inference (Wilkinson & Task Force on Sta s cal Inference, APA Board of Scien fic Affairs, 1999) suggested that it is more appropriate to specify a limited number of planned contrasts in advance rather than to make all possible comparisons among cell means post hoc. Some mes, par cularly when a computer program does not include post hoc tests as an op on or when there are problems with homogeneity of variance assump ons, it is convenient to do independent-samples t tests to compare pairs of cell means and to use the Bonferroni

procedure to control overall risk for Type I error. The Bonferroni procedure, described in earlier chapters, is very simple: If a researcher wants to perform k number of t tests, with an overall experiment-wise error rate (EWα) of .05, then the per comparison alpha (PCα) that would be used to assess the significance of each individual t test would be set to PCα = EWα/k. Usually, EWα is set to .05 (although higher levels such as EWα = .10 or .20 may also be reasonable). Thus, if a researcher wanted to do three t tests among cell means as a follow-up to a significant interac on, he or she might require a p value less than .05/3 = .0167 for each t test as the criterion for sta s cal significance. 16.10.2 Nature of Main Effect Differences When a factor has only two levels or groups, and if the F for the main effect for that factor is sta s cally significant, no further tests are necessary to understand the nature of the differences between group means. However, when a factor has more than two levels or groups, the researcher may want to follow up a significant main effect for the factor with either planned contrasts or post hoc tests to compare means for different levels of that factor. It is important, of course, to note whether significant differences among means were in the predicted direc on. 16.11 FACTORIAL ANOVA USING THE SPSS GLM PROCEDURE The hypothe cal data in the file socialsupportstress.sav represent scores for a nonexperimental factorial study. The two factors were as follows: Factor A, amount of social support (A1 = low, A2 = high), and Factor B, stress (B1 = low, B2 = high). The outcome variable was number of physical symptoms reported on a checklist. To request the means, highlight the list of effects in the le -hand pane, and move the en re list into the right-hand pane that is headed “Display Means for,” as shown in Figure 16.5. This results in output of the grand mean, the mean for each row, each column, and each cell in this 2 × 2 factorial design. If a test of homogeneity of variance (the Levene test) is desired, check the box for “Homogeneity tests.” Close this window by clicking OK. It is useful to generate a graph of cell means; to do this, click the Plots bu on in the main GLM dialog box (Figure 16.4). This opens the Univariate: Plots dialog box, shown in Figure 16.6. To place the stress factor on the horizontal axis of the plot, move the factor name stress_b into the “Horizontal Axis” box; to define a separate line for each level of the social support factor, move the factor name socsup_b into the “Separate Lines” box. Click Add to request this plot.

Figure 16.3 Menu Selec ons for SPSS GLM Univariate Procedure

Figure 16.4 GLM Univariate Dialog Box

Figure 16.5 Op ons for GLM Request Means and Levene Test

Figure 16.6 GLM Univariate: Profile Plots Dialog Box

In addi on, a follow-up analysis was performed to compare mean symptoms between the low- and high-stress groups, separately for the low–social support group and the high–social support group. An analysis that compares means on one factor (such as stress) within the groups that are defined by a second factor (such as level of social support) is called an analysis of simple main effects. There are several ways to conduct an analysis of simple main effects. For the following example, the <Data> → <Split File> command was used, and social support was iden fied as the grouping variable for which separate analyses were requested. An independent-samples t test was performed to evaluate whether mean symptoms differed between the low- and high-stress condi ons, separately within the low– and high–social support groups. 16.12 SPSS OUTPUT Results of the Levene test for the homogeneity of variance assump on are displayed in Figure 16.7. This test did not indicate a significant viola on of the homogeneity of variance assump on: F(3, 16) = 1.23, p = .33. (Even if this test were sta s cally significant, ANOVA is robust to viola ons of this assump on.) The GLM source table produced by SPSS that summarizes the factorial ANOVA to predict symptoms from social support and stress appears in Figure 16.8. The row labeled “Corrected Model” represents an “omnibus” test: This test tells us whether the combined effects of A, B, and A × B are sta s cally significant. This omnibus F is rarely reported. The line labeled “Intercept” tests the null hypothesis that the grand mean for the outcome variable equals 0; this is rarely of interest. The lines that are labeled with the names of the factors (socsup_A, stress_B, and socsup_A × stress_B) represent the main effects and the interac on. The row labeled “Error” corresponds to the term that is usually called “within groups” in textbook descrip ons of ANOVA. The SStotal that is usually reported in textbook examples, and is used to calculate es mates of effect size, corresponds to the row of the table that SPSS labels “Corrected Total” (it is called “corrected” because the grand mean was “corrected for” or subtracted from each score before the terms were squared and summed). Finally, the row that SPSS designates “Total” corresponds to the sum of the squared scores on the outcome variable (ΣY2); this term is not usually of any interest (unless the researcher wants to test the null hypothesis that the grand mean of scores on the outcome variable Y equals 0).

The parts of the GLM source table you should focus on are retained in the edited version of the GLM source table in Figure 16.9. Tables of means (grand mean, row means, column means, and cell means) appear in Figures 16.10 and 16.11. For this hypothe cal data set, the obtained pa ern of cell means was similar to the pa ern predicted by the buffering hypothesis (illustrated in Figure 16.2, right panel). An increase in stress (from low to high) was associated with an increase in reported symptoms only for the low–social support group (A1). The high–social support group (A2) showed very li le increase in symptoms as stress increased; this outcome could be interpreted as consistent with the buffering hypothesis illustrated by the graph of cell means in Figure 16.2 on the right. It appears that high social support buffered or protected people against the effects of stress. To generate a bar graph for means in a factorial ANOVA, make the following SPSS menu selec ons: <Graphs> → <Legacy Dialogs> → <Bar>. Select the clustered type of bar chart and the radio bu on for “Summaries for groups of cases,” then click the Define bu on to open the dialog box in Figure 16.12.

Figure 16.12 Clustered Bar Chart Dialog Box

In the dialog box for Define Clustered Bar: Summaries for Groups of Cases (Figure 16.12), place the name of the dependent variable (symptom) in the box named “Variable” and choose the radio bu on for “Other sta s c (e.g., mean).” Note that in this bar chart, the height of each bar corresponds to the mean on the outcome variable for the corresponding group. Place the name of the stress factor (stress_b) in the box for “Define Clusters by” and the name of the social support factor (socsup_b) in the “Category Axis” window. To obtain error bars (i.e., 95% confidence intervals) for group means, click the Op ons bu on. In the Op ons dialog box (Figure 16.13), check “Display error bars.” In this example, the radio bu on selec on calls for error bars that correspond to the 95% CI for each group mean. To see the resul ng bar chart, click the Con nue bu on, then the OK bu on (see Figure 16.14). Analysts o en want to conduct an analysis of simple main effects, that is, to ask whether the mean level of symptoms differs between the high- and low-stress groups, separately for an analysis of data in the high–social support and low–social support groups. The analysis reported in Figure 16.15 provides this informa on. One-way ANOVA was performed to compare stress cell means, separately within the high–social support group and the low–social support group. (Independent-samples t tests would also be acceptable.)

Figure 16.13 Op ons Dialog Box: Request 95% CIs for Clustered Bar Chart

Figure 16.14 Clustered Bar Chart: Heights of Bars Correspond to Group Means for Symptoms

Figure 16.15 Analysis of Simple Main Effects: Differences in Mean Symptoms Between High- and Low-Stress Groups, Separately by Social Support Group 16.13 RESULTS Following is an example of a “Results” sec on for the 2 × 2 factorial ANOVA in the study of stress and social support. Results A 2 × 2 factorial ANOVA was performed using the SPSS GLM procedure to assess whether number of reported symptoms (Y) could be predicted from level of social support (A1 = low, A2 = high), level of stress (B1 = low, B2 = high), and the interac on between social support and stress. On the basis of the buffering hypothesis, it was expected that the high–social support group (A2) would show li le or no increase in symptoms at higher levels of stress, whereas the low–social support group (A1) was predicted to show substan ally higher levels of symptoms under high-stress condi ons than under low-stress condi ons. This was an orthogonal factorial design; each of the four cells had the same number of par cipants (n = 5). Preliminary data screening was done to assess whether the assump ons for ANOVA were seriously violated. Examina on of a histogram of scores on the outcome variable suggested that the symptom scores had a skewed distribu on; however, no data transforma on was applied. The Levene test indicated no significant viola on of the homogeneity of variance assump on. As predicted, there was a sta s cally significant social support by stress interac on: FA×B(1, 16) = 21.49, p < .001. The corresponding effect size es mate (η2 = .20) indicated a strong effect. The bar chart of cell means (in Figure 16.14) indicated that the low–social support/high-stress group had a much higher level of mean symptoms than the other three groups. The pa ern of cell means was consistent with the predic on made by the buffering hypothesis. For the low–social support group, the mean number of symptoms was greater in the high-stress condi on (M = 12.8) than in the low-stress condi on (M = 3.8). On the other hand, for persons high in social support, symptoms were not much greater in the high-stress condi on (M = 6.0) than in the low-stress condi on (M = 4.0). An analysis of simple main effects was done to assess whether the differences (between low and high levels of stress) were significant within the low– and high–social support groups. For the low–social support group, the difference between means for the low- versus high-stress groups was sta s cally significant, F(1, 8) = 54.73, p < .001. For the high–social support group, the difference between means for the low- versus high-stress groups was not significant, F(1, 8) = 5.00, p = .056. The nature of the obtained interac on was consistent with the predic on based on the buffering hypothesis. Par cipants with high levels of social support did not report significantly higher symptoms under high stress; par cipants with low levels of social support did report significantly higher symptoms under high stress.

There were also significant main effects: for social support, FA(1, 16) = 19.10, p < .001, with an associated η2 effect size es mate of .17; for stress, FB(1, 16) = 53.07, p < .001, with an es mated effect size of η2 = .48 (an extremely large effect size). A few further comments about results: In the one-way ANOVA that compared differences in mean symptoms across social support groups (in Figure 16.2), the social support factor was not sta s cally significant. In the factorial ANOVA, the social support factor is sta s cally significant. In this situa on, adding the second factor (stress) reduced the within group error variance. However, in a situa on where the interac on is sta s cally significant, the interpreta on is likely to focus more on the nature of the interac on and less on the nature of any main effects. 16.14 DESIGN DECISIONS AND MAGNITUDES OF SS TERMS

 For each factor, how many levels are needed?  What can you do to maximize between-group differences (selec on of naturally

occurring groups, dosage levels for amounts of treatment)?  What can you do to minimize experimental error?  What sample size is needed for adequate sta s cal power and reasonably narrow CIs?

Just as in previous chapters on the independent-samples t test and one-way ANOVA, in factorial ANOVA, we want to assess whether group means are far apart rela ve to the within-group variability of scores. The same factors that affected the size of t and F in these earlier analyses are relevant in factorial ANOVA. 16.14.1 Distances Between Group Means (Magnitudes of SSA and SSB) Other factors being equal, group means tend to be farther apart, and SSbetween and the F ra o for an effect tend to be larger, when the dosage levels of treatments administered to a group are different enough to produce detectable differences. For instance, if a researcher compares groups that receive 0 versus 300 mg of caffeine, the effects of caffeine on heart rate will probably be larger than if the researcher compares 0 versus 30 mg of caffeine. For comparison of preexis ng groups that differ on par cipant characteris cs, differences between group means tend to be larger (and F ra os tend to be larger) when the groups are chosen so that these differences are substan al; for example, a researcher has a be er chance of finding age- related differences in blood pressure when comparing groups that are age 20 versus age 70 than when comparing groups that are age 20 versus age 25. 16.14.2 Number of Scores Within Each Group or Cell Assuming a constant value of SS, MSwithin tends to become smaller when the n of par cipants within groups is increased (because MS = SS/df, and df for the within-group SS term increases as N increases). This, in turn, implies that the value of an F ra o o en tends to be higher as the total sample size N increases, assuming that all other aspects of the data (such as the distances between group means and the within-cell varia on of scores) remain the same. This corresponds to a commonsense intui on. Most of the sta s cal significance test sta s cs that you have encountered so far (such as the independent-samples t test and the F ra o in ANOVA)

tend to yield larger values as the sample size N is increased, assuming that other terms involved in the computa on of F (such as the distances between group means and the variability of scores within groups) remain the same and that the sample means are not exactly equal across groups. 16.14.3 Variability of Scores Within Groups or Cells (Magnitude of MSwithin) Other factors being equal, the variance of scores within groups (MSwithin) tends to be smaller, and therefore F ra os tend to be larger, when sources of error variance within groups can be controlled (e.g., through selec on of homogeneous par cipants, through standardiza on of tes ng or observa on methods, by holding extraneous variables constant so that they do not create varia ons in performance within groups). Including a blocking factor can be another way to reduce the variability of scores within groups. In addi on to the ability to detect poten al interac on effects, a factorial ANOVA design also offers researchers a possible way to reduce SSwithin, the variability of scores within groups, by blocking on subject characteris cs. For example, suppose that a researcher wants to study the effects of social support on symptoms, but another poten al predictor variable, stress, is also related to symptoms. The researcher can use stress as a “blocking factor” in the analysis. In this hypothe cal example involving a 2 × 2 factorial, the factor that represents social support (A1 = low, A2 = high) can be crossed with a second blocking factor, stress (B1 = low, B2 = high). We would say that the par cipants have been “blocked on level of stress.” Just as in studies that use t tests and one-way ANOVA, researchers who use factorial ANOVA designs can try to maximize the size of F ra os by increasing the differences between the treatment dosage levels or par cipant characteris cs that differen ate groups, by controlling extraneous sources of variance through experimental or sta s cal control, or by increasing the number of cases within groups. 16.15 SUMMARY This chapter demonstrated that adding a second factor to ANOVA (se ng up a factorial analysis of variance) can provide more informa on than a one-way ANOVA analysis. When a second factor is added, factorial ANOVA provides informa on about main effects for two factors and, in addi on, informa on about poten al interac ons between factors. Also, blocking on a second factor (such as level of stress) can reduce the variability of scores within cells (SSwithin); this, in turn, can lead to larger F ra os for the tests of main effects. Note that the F test for the main effect of social support was sta s cally significant in the factorial ANOVA reported in the “Results” sec on, even though it was not sta s cally significant in the one-way ANOVA reported at the beginning of the chapter. The addi on of the stress factor and the interac on term resulted in a smaller value of SSwithin in the factorial design (rela ve to the one-way ANOVA). The reduc on in degrees of freedom for the error term (df was reduced by 2) did not outweigh this reduc on in SSwithin, so in this situa on, MSwithin was smaller in the factorial ANOVA than in the one-way ANOVA, and the F ra o for the main effect of social support was larger in the factorial ANOVA than in the one-way ANOVA.

This chapter described the meaning of an interac on in a 2 × 2 factorial design. When an interac on is present, the observed cell means differ from the cell means that would be predicted by simply summing the es mates of the grand mean μ, row effect αi, and column effect βj for each cell. The significance test for the interac on allows the researcher to judge whether departures from the pa ern of cell means that would be predicted by a simple addi ve combina on of row and column effects are large enough to be judged sta s cally significant. A major goal of this chapter was to make it clear exactly how the pa ern of cell means is related to the presence or absence of an interac on. Descrip on of the nature of an interac on usually focuses on the pa ern of cell means (as summarized either in a graph or a table). We can expand on the basic factorial ANOVA design described in this chapter in many different ways. First of all, we can include more than two levels on any factor; for example, we could run a 2 × 3 or a 4 × 6 factorial ANOVA. If significant F’s are obtained for main effects on factors that have more than two levels, post hoc comparisons (or planned contrasts) may be used to assess which par cular levels or groups differed significantly. We can combine or cross more than two factors. When more than two factors are included in a factorial ANOVA, each one is typically designated by an uppercase le er (e.g., A, B, C, D). A fully crossed three-way factorial A × B × C ANOVA combines all levels of A with all levels of B and C. A possible disadvantage of three-way and higher-order factorials is that theories rarely predict three-way (or higher order) interac ons, and three- or four-way interac ons can be difficult to interpret. Furthermore, it may be difficult to fill all the cells (in a 3 × 2 × 5 factorial, we would need enough par cipants to fill 30 cells or groups). We can obtain repeated measures on one or more factors; for example, in a study that assesses the effect of different dosage levels of caffeine on male and female par cipants, we can expose each par cipant to every dosage level of the drug. We can combine group membership predictors (or factors) with con nuous predictors (usually called “covariates”) to see whether group means differ when we sta s cally control for scores on covariates; this type of analysis is called analysis of covariance (ANCOVA). We can measure mul ple outcome variables and ask whether the pa erns of means on a set of several outcome variables differ in ways that suggest main effects and/or interac on effects; when we include mul ple outcome measures, our analysis is called a mul variate analysis of variance (MANOVA). A complete treatment of all these forms of ANOVA (e.g., nested designs, mul ple factors, mixed models that include both between-S and repeated-measures factors) is beyond the scope of this textbook. More advanced textbooks provide detailed coverage of these various forms of ANOVA.