BIZZNA TUTOR
CONTROLS IN RESEARCH
Objective: Following completion of this class, the student will understand the critical need for controlling as many factors as possible and can discuss all of Campbell and Stanley’s (1963) threats to internal validity.
Did you ever hear the workplace phrase, control freak? While on the job, that is often a pejorative. However, the researcher must attempt to set up absolute control in every possible way. This ideal was expressed long ago: “If two situations are equal in every respect except for a factor present in one of the situations, any difference which appears between the two situations can be attributed to the factor. This statement is referred to as the” law of the single variable” (Galfo & Miller, 1970, p. 17). In the experimental and quasi-experimental designs presented earlier, we looked at the importance of using a control group and pretests, both designed to address the conditions of the law of the single variable. In the next section, we will examine selection of subjects and the way that best ensures equivalence of the groups at the outset.
There are many other controls that must be considered. Pretesting and post-testing must be at the same time of day because physiological and cognitive functions vary in a diurnal fashion. Some subjects will be more energetic and alert in the early morning while afternoons find them drowsy. Pretesting under one condition and post-testing under another introduces a variable that should not exist. By the same token, pretesting the control group in the afternoon and the experimental group in the morning is contraindicated.
In administering a telephone survey, the researcher must use a script so that all respondents are given precisely the same instruction. No variability is allowed, for then we have a potentially confounding factor, and the law of the single variable has been violated. A script should also be used in administering a paper-and-pencil survey for the same reason. Of course, one should also control for diurnal variations in this kind of data collection as well.
It would seem obvious that surveys given to two groups must be identical, without exception. While that seems obvious to you and me, I have had students produce surveys that were similar, but worded differently for the two groups that were to be compared. The researcher must be diligent in looking at ways to control these and other potential influences, other elements that would add a factor that might possibly be different and thus getting away form the law of the single variable ideal. This requires higher order thinking and it is just good common sense to seek the advice of others such as peers and/or mentors in brainstorming this problem. Keep in mind that we are seeking the ideal in terms of control and that the educational research literature is replete with less than perfect controls.
Campbell and Stanley (1963) addressed several factors that can jeopardize the validity of a research project through inadequate controls. They classify these as factors affecting internal validity and factors affecting external validity. Lack of internal validity would make the findings questionable or even not interpretable because of the potential effect of lack of certain kinds of controls. Failure to obtain external validity limits how broadly you can generalize your results. Because the focus of this section is on control, we will now look at those factors related to internal validity.
The first factor is history. What kinds of events occurred between the pretest and the post-test that could have added an effect to that of the intervention, thus complicating the experiment? A natural and obvious response is to add a control group so that the history of the two groups is similar. However, when as is usually the case the two groups are not in the same environment during the intervention, some aspect of history could confound the situation. If the groups are separate and one is exposed to an event that might be emotionally traumatic, the effect on that group could carry over to the post-test. If either group somehow experiences anything different, emotional or not, there becomes a question of violation of control. History could be a problem even when there is no pretest involved such as is the case in a couple of designs that we discuss. Certainly, even in survey research there might be events that occur between early returns and late returns that could affect perceptions. Surveyed subjects could interact with those not yet questioned. There could be a newspaper article directly addressing the topic in question that those surveyed earlier had no opportunity to read, but others did. To avoid that possibility, it would be necessary to administer the survey to all subjects within a very tight time frame.
The second factor is maturation. What might have changed over the passage of time? If it is a long term study, certainly, as the title suggests, maturing brings about some alterations that one would have to consider might affect scores on a test. If it is a study on the elderly, perhaps aging would be a more properly descriptive term than maturation. Campbell and Stanley (1963) add that over time such conditions as hunger, fatigue, and boredom might become factors that affect testing results and must be subsumed under the term maturation. If part of your sample is tested or surveyed at 7 a.m., another part at 11 a.m., and a third part at 5 p.m., certainly the passage of time will possibly have an effect. Part of your results would be due to natural biological diurnal variations in each subject and part could be a result of simply higher energy in the morning, hunger prior to lunch, and fatigue later in the day. We might never know what factors interacted with the intervention to bring about the results. Remember, we’re considering passing time and associated changes, not just the physiological process of maturing. Using a control group, if both groups experience precisely the same factors of maturation, can allow valid conclusions regarding the effects of the intervention.
The third factor to be considered as a potential problem with validity and something to control is the effect of testing itself. We know quite well that in multiple testing situations, the first score is usually lower than those following. Because of high stakes testing we regularly take students out of physical education, art, and music classes to teach to the test and to administer practice tests, knowing that this latter procedure will probably result in slightly higher scores. In GRE and SAT testing it is common for students upon retesting, with no intervention, to score higher the second time around. By the same token, in designs with a pretest, there is the possibility that the pretest might have an effect that interacts with the intervention to alter posttest results. In designs with a control group that also received a pretest, if everything else is equal we can make valid judgments on the intervention’s effect because we know that both groups were equally affected by the pretesting.
Campbell and Stanley’s (1963) fourth factor that requires controlling for internal validity is instrumentation or instrument decay, would seem to be more relevant to biological or other scientific studies that require instruments that might be a problem. An instrument that measures percentages of gases requires periodic calibration to ensure accuracy. Failure to do the calibration would result in erroneous readings and invalid measurements. But measurements are not always done with what we usually consider instruments. Humans can be the instruments as well. Interviewers in qualitative research can suffer from many conditions (e.g., illness, fatigue, etc.) that would negatively affect the quality of data obtained and this is a problem with instrumentation. They might become more skilled over weeks of collecting data and there would be an effect of instrumentation in that more accurate measurements would result toward the end. Any member of this class who has graded large numbers of essays knows that standards for evaluation sometimes undergo a slight shift from beginning to end, another instrumentation failure.
Fifth among the factors that can affect internal validity is regression. Considered in the most simplistic way, there is a tendency for subjects with extreme scores, high or low, to move towards the mean in subsequent testing. If in 100 3rd grade students, you take the top ten and bottom ten on a pretest on mathematics proficiency and give an intervention to the bottom ten, subsequent testing (post-testing) will almost always result in a lower score by the best students and a higher score by the worst. The problem is that if you had the foresight to use randomly assigned controls, the same pattern would occur without the intervention. Statistical testing could reveal the effect of the intervention if the control group was included in the study, but without it conclusions would be invalid.
Bias in subject assignment to groups is the sixth threat to internal validity. This can occur any time that random assignment isn’t done. A really bad example would be in a study dealing with the effect of some sort of in-service activity on teacher morale to seek volunteers for the experimental group and to assign non-volunteers to the control group. Volunteers, by their very nature are different and always suspect subjects. Don’t use volunteers in any study, experimental or otherwise. For experimental studies, random assignment is mandated. For other research, be careful to make sure that you can defend the selection procedures. In your descriptive research pilot study, you cannot assign across groups randomly because you’ll be comparing two different populations. However, for each population that you choose, you can do random assignment of subjects from each. If one population is special education teachers in Smith County and your comparison population is mothers of special education students in Smith County, you’d make two lists and randomly select from each, at the very least limiting subject bias to the best of your ability. Keep in mind that the larger the sample, the less potential for subject bias during random sampling.
Experimental mortality, defined for these purposes as differential loss of subjects from your groups, is definitely a threat to validity. Here’s an example. Let’s assume that you
Table 1. Sources of internal invalidity*.
|
|
History |
Maturation |
Testing |
Instrumentation |
Regression |
Selection |
Mortality |
|
One shot |
- |
- |
|
|
|
- |
- |
|
Single group pre and posttest |
- |
- |
- |
- |
? |
+ |
+ |
|
Static group comparison |
+ |
? |
+ |
+ |
+ |
- |
- |
|
Pre-posttest with control gp |
+ |
+ |
+ |
+ |
+ |
++ |
+ |
|
Posttest only with control gp |
+ |
+ |
+ |
+ |
+ |
+ |
+ |
|
Solomon four-group |
+ |
+ |
+ |
+ |
+ |
+ |
+ |
|
Nonrandom pre & post with control group |
+ |
+ |
+ |
+ |
? |
? |
+ |
|
Nonrandom time series & control group** |
+ |
+ |
+ |
? |
+ |
? |
+ |
*Adapted from Campbell & Stanley (1963). Plus is a strength, negative a weakness, and question marks indicate a possible concern.
**Not included in the review. Extrapolated from their simple time series with no control group.
randomly assigned boys and girls to a control group and an experimental group to examine the effects of studying sexist literature on attitudes towards women’s place in the corporate world. It happens that half of the girls in the experimental group were offended by what they saw as an attempt to brainwash them, so they refused to continue in the study. This is experimental mortality in action and obviously any findings from this study would be flawed. Large losses from a control group have the same effect of lessening control over the experiment and decreasing validity of the study.
A final threat to internal validity deals with the interaction of subject selection with other factors and its complexity is beyond the purview of this class. By the same token, Campbell and Stanley (1963) present a section on threats to external validity, basically how extensively you can generalize your findings. These are technical and really more appropriate for consideration in more sophisticated research than we’ll be addressing in this class. For our purposes, we will just assume that generalization can not be made beyond the population from which the sample was taken.
Note the strength of the three true experimental designs in the matrix on the preceding page. Also, it can be seen that the last two, quasi-experimental designs popular in educational research, have few weaknesses and this are a logical solution when random assignment of subjects is not possible.
Suggested reading
Campbell, D.T. & Stanley, J.C. (1963). Experimental and Quasi-experimental Designs for Research. Boston: Houghton Mifflin.