Discussion 1: Evidence-Based Practices: How Do You Know They Are Working?
Vol. 75. No. 3, />/>. 365-383.
©2009 Council for Exceptional Children.
Exceptional Childre
Determining Evidence-Based Practices in Special Education
BRYAN G. COOK University of Hawaii at Manoa
MELODY TANKERSLEY Kent State University
TIMOTHY J . LANDRUM University of Virginia
ABSTRACT: Determining evidence-based practices is a complicated enterprise that requires analyz-
ing the methodological quality and magnitude of the available research supporting specific prac-
tices. This article reviews criteria and procedures for identifying what works in the fields of clinical
psychology, school psychology, and general education; and it compares these systems with proposed
guidelines for determining evidence-based practices in special education. The authors then summa-
rize and analyze the approaches and findings of the 5 reviews presented in this issue. In these re-
views, prominent special education scholars applied the proposed quality indicators for
high-quality research and standards for evidence-based practice to bodies of empirical literature.
The article concludes by synthesizing these scholars' preliminary recommendations for refining the
proposed quality indicators and standards for evidence-based practices in special education, as well
as the process for applying them.
uch forces as the standards- based education movement, the mandated participation of stu- dents with disabilities in state proficiency testing, inclusion,
and the recognition that many students with dis- abilities are capable of higher levels of academic and social attainment than previously expected have driven an intensified focus on improving outcomes for students with disabilities. Perhaps because many factors that may inhibit the out- comes of students with disabilities are beyond the direct control of educators (e.g., poverty, limited resources, attitudes), special educators have
tended to focus their attention on one determi- nant of students' outcomes over which they have always exercised primary control—teaching prac- tices. Unfortunately, many teachers of students with disabilities have implemented teaching prac- tices shown to have little effect on student out- comes while eschewing many research-based practices (e.g., B. G. Cook & Schirmer, 2003; Kauffman, 1996). In an effort to bridge this re- search-to-practice gap, lawmakers have empha- sized practices that research has shown to be effective in such legislation as the No Child Left Behind Act of 2001 and the Individuals With Disabilities Education Act of 2004. Researchers
Exceptional Children 3 6 5
in special education have also conducted initial research on how to effectively support teachers in adopting and maintaining the use of research- based or evidence-based practices (see Wanzek & Vaughn, 2006, for a review of this research). De- spite the considerable interest in basing instruc- tional practices on research evidence, special educators have not yet established definitively which practices are or are not evidence-based or settled on a systematic process for determining evidence-based practices (EBPs).
E V I D E N C E - B A S E D P R A C T I C E S
All interventions are not equal; some are much more likely than others to positively affect student outcomes (Forness, Kavale, Blum, &C Lloyd, 1997). Simple logic appears to suggest that, in general, teachers should prioritize the use of in- structional practices that are most likely to bring about desired student outcomes. Although some contend that research cannot reliably determine which educational practices produce desired gains in student outcomes (e.g., Gallagher, 1998), we proceed under the positivist assumption that it can (Lloyd, PuUen, Tankersley, & Lloyd, 2006). The use of EBPs, or those practices shown by re- search to work, seems particularly imperative in special education. As Dammann and Vaughn (2001) suggested, whereas many nondisabled stu- dents make adequate progress under a variety of instructional conditions, students with disabilities require the most effective teaching techniques to succeed. However, advocating for implementing EBPs in special education begs two critical ques- tions: what are EBPs, and how can researchers identify them?
All interventions are not equal; some are much more likely than others to positively affect student outcomes.
Determining whether a practice is evidence- based involves a number of issues: What types of research designs should researchers consider? How many studies with converging findings are neces- sary to instill confidence that a practice is effec- tive? How methodologically rigorous must a
study be for the results to be meaningful? To what extent must an intervention affect student out- comes for researchers to consider it effective? Al- though other issues certainly affect the difficult business of determining EBPs, we limit our dis- cussion to these four issues—research design, quantity of research, methodological quality, and magnitude of effect.
RESEARCH DESIGN
A generally accepted tenet of educational research holds that research designs exhibiting experimen- tal control most appropriately address the ques- tion of whether a practice works (B. G. Gook, Tankersley, Gook, & Landrum, 2008). We recog- nize that no research design can completely rule out all alternative explanations for findings when conducted in the real-world settings of schools and classrooms; however, some designs do so more meaningfully than others. By using a con- trol group, randomly assigning participants to groups, and actively introducing the intervention to the experimental group, group experimental designs can produce reliable knowledge claims re- garding whether an intervention affects student outcomes (L. H. Gook, Gook, Landrum, & Tankersley, 2008). We are not implying that ex- perimental research is better than other research designs; rather, different types of research address different questions, and researchers should use them accordingly. Should true experiments be the only research design considered in determining EBPs? Gan quasi-experiments, single-subject re- search (SSR), correlational research, and qualita- tive research also meaningfully determine whether a practice works?
QUANTITY OE RESEARCH
The process of conducting educational research and accumulating knowledge from research is tentative and cumulative (Rumrill & Gook, 2001). Because of the recognized vagaries in con- ducting field-based educational research (Berliner, 2002), it seems unwise to place too much faith in the results of a single study regardless of its de- sign, effect size, or methodological rigor. Ger- tainly, as more studies with converging evidence accrue, research consumers can have greater confi- dence in those findings. But how many studies
3 6 6 Spring 2009
supporting a practice are sufficient to reasonably conclude that it works?
METHODOLOGICAL QUALITY
The methodological rigor with which a study is conducted affects the confidence that one can have in its findings. For example, evidence of ac- ceptable implementation fidelity seems to be a necessary feature of a trustworthy study. If the re- searchers did not implement the intervention as designed, they can draw no meaningful conclu- sion about the effectiveness of the practice. In- deed, Simmerman and Swanson (2001) reported that the presence of desirable methodological fea- tures in a study (e.g., controlling for teacher effects, iising appropriate units of analysis in ana- lyzing data, reporting psychometric properties of measurement tools) significantly corresponds with lower efFect sizes. Examining and accounting for the methodological quality of studies in deter- mining EBPs therefore appears important. Should researchers determine EBPs by using only studies of high methodological quality? What method- ological features are critically important for a high-quality study?
MAGNITUDE OE EEEECT
EBPs should have a considerable and meaning- ful—as opposed to trivial—positive effect on stu- dent outcomes. Researchers have traditionally gauged the impact of an intervention in group studies by using tests of statistical significance, which estimate the likelihood that differences be- tween the grotips occurred by chance. However, in part because of concerns that studies involving a large number of participants can yield statisti- cally significant findings even when outcomes may not be educationally meaningful, researchers have begun to report effect sizes (e.g., Cohen's d), which sample size does not affect, to help inter- pret an intervention's effect (American Psycholog- ical Association, 2001). Although Cohen (1988) suggested values for interpreting effect sizes as small, medium, and large, he was careful to point out that researchers should not consider these subjective guidelines to be absolute standards. How is the effect of an intervention best evalu- ated? If researchers use effect sizes to assess the impact of a practice, how large an effect is neces-
sary to indicate a meaningful change? If re- searchers use SSR studies to determine EBPs in special education, how should they evaluate the effect of the intervention?
The importance of using practices shown by research to be the most effective is by no means unique to special education (Odom et al., 2005). The medical field generally receives credit for pio- neering efforts in this area, with evidence-based medicine becoming prominent in the 1990s (see Sackett, Rosenberg, Cray, Haynes, & Richardson, 1996). Such other professions as clinical psychol- ogy (see Chambless et al., 1996, 1998); school psychology (see Kratochwill & Stoiber, 2002; Task Force on Evidence-Based Interventions in School Psychology, 2003); and general education (see What Works Clearinghouse, WWC, n.d.a) have followed suit, developing criteria and proce- dures for identifying EBPs in their fields. To con- textualize efforts to determine EBPs in special education, we briefly review the criteria and stan- dards for determining EBPs in these three fields.
D E T E R M I N I N G E V I D E N C E - B A S E D
P R A C T I C E S I N R E L A T E D F I E L D S
The Division 12 (Division of Clinical Psychol- ogy) Task Force on Promotion and Dissemination of Psychological Procedures (1995) delineated cri- teria, which Chambless et al. updated in 1996 and 1998, for well-established treatments and probably efficacious treatments in clinical psy- chology. Subsequently, in the field of school psy- chology. Division 16 and the Society for the Study of School Psychology Task Force developed a detailed system for coding and describing multi- ple aspects of research studies (Task Force on Evi- dence-Based Interventions in School Psychology, 2003). Instead of categorizing the degree to which iriterventions are evidence-based, the coding sys- tem generated by the school psychology team provides a detailed description of a research base, from which consumers "draw their own conclu- sions based on the evidence provided" regarding the sufficiency of research supporting an interven- tion (Kratochwill Sc Stoiber, 2002, p. 360).
In general education, the WWC, established in 2002 by the U.S. Department of Education's Institute of Education Sciences, rates reviewed
Exceptional Children
practices as having positive, potentially positive, mixed, no discernible, potentially negative, or negative effects (WWC, n.d.b). This section ex- amines how these three diverse approaches for identifying what works in fields closely related to special education treat the issues of research de- sign, quantity of research, methodological quality, and niagnitude of effect.
RESEARCH DESIGN
Clinical Psycholog. Chambless et al. (1998) considered only studies employing between-group experimental and SSR designs in determining both well-established treatments and probably ef- ficacious tteatments.
School Psychology. The Task Force on Evi- dence-Based Interventions in School Psychology (2003) aims to provide descriptions of group re- search, SSR, confirmatory program evaluation, and qualitative research (Kratochwill & Stoiber, 2002). Coding manuals are currently available for group research and SSR, but are being expanded to include criteria for qualitative research and confirmatory program evaluation (T. Kratochwill, personal communication, September 26, 2008). Kratochwill and Stoiber suggested that coding ndnexperimental research studies (i.e., qualitative and confirmatory program evaluation) provides information on a broad range of research relevant to consumers but do not indicate that these dif- fetent research designs contribute equally to de- termining whether a practice works.
General Education. The WWC (2008) con- siders only randomized controlled trials and quasi-experimental studies (i.e., quasi-expeti- ments with equating, regression discontinuity designs, and SSR) when determining the effec- tiveness of an intervention. The WWC classifies studies as meeting evidence standards, meeting evidence standards with reservations, or not meet- ing evidence standards. Only randomized con- trolled studies can meet evidence standards without reservation. Quasi-experimental studies that satisfy the WWC's methodological criteria, as well as randomized controlled studies with methodological limitations, can meet evidence standards with teservations. Methodological crite- ria for SSR and regression discontinuity designs
have been under development since September 2006 but are not yet available (WWC).
QUANTITY OF RESEARCH
Clinical Psychology. The Division 12 Task Force considers a psychological treatment well-es- tablished when at least two good between-group design experiments or nine SSR studies support it (Chambless et al., 1998). The clinical psychology task force considers a treatment to be possibly ef- ficacious when supported by at least (a) one group experiment that meets all methodological criteria for group experiments except the require- ment for multiple investigators, (b) two group ex- periments that produce superior outcomes in comparison with a wait-list control group, or (c) three SSR studies that meet all SSR criteria except the requirement fot multiple investigators.
School Psychology. Because the school psy- chology task force did not seek to categorize prac- tices regarding its effectiveness, it did not establish criteria related to the number of required studies for evidence-based classifications.
General Education. The WWC (n.d.b) re- quires at least one or two studies for a practice or curriculum to be considered as having positive, potentially positive, mixed, potentially negative, or negative effects. The specific number and type of studies required varies within and between these categories of effectiveness. For example, a positive effect requires two or more studies show- ing statistically significant positive effects, at least one of which meets WWC evidence standards without reservations, and no studies showing sta- tistically significant or substantively important negative effects. A potentially positive efFect, how- ever, requires at least one study showing a statisti- cally significant or substantively important positive effect, no studies showing statistically sig- nificant or substantively important negative ef- fects, and no more studies showing indeterminate effects than studies showing statistically signifi- cant or substantively important positive effects.
METHODOLOGICAL QUALITY
Clinical Psychology. In addition to stipulating that tesearchers must compare interventions with a placebo or other tteatment, the Division 12 cri- teria for well-established treatments require that
3 6 8 Spring 2009
researchers (a) conduct experiments with treat- ment manuals, (b) clearly describe participant characteristics, and (c) have two separate investi- gators or investigatory teams conduct supporting studies (Chambless et al., 1998). These standards are relaxed for possibly efficacious treatments. Group experiments that compare the treatment group with a wait-list control group and that the same investigators conduct may be considered for possibly efficacious practices, as can SSR studies that the same investigators conduct.
When evidence regarding the effects of an in- tervention is mixed, reviewers further assess the methodological quality of studies to determine which studies to weigh more heavily (Chambless et al., 1998). Chambless and HoUon (1998) rec- ommend assessing such methodological features as the following:
• The descriptions of samples use standard di- agnostic labels assigned from a structured di- agnostic interview.
• Outcome measures demonstrate acceptable reliability and validity in previous research.
• With the exception of simple procedures, the researchers follow a written treatment man- ual when delivering the intervention.
• Researchers avoid Type I error (e.g., adjust alpha level when conducting multiple statis- tical tests), control for pretest scores when comparing groups' posttest measures, and ad- just analysis and interpretation if differential attrition or participation rates exist between groups.
• A stable baseline, typically with at least three data points, is established in SSR.
School Psychology. Although the Division 16 procedures do not classify studies according to their methodological quality, reviewers do rate and describe a number of methodological fea- tures—which consumers use to make informed decisions about an intervention's evidence base and effectiveness (Kratochwill & Stoiber, 2002). Reviewers evaluate studies, regardless of design, by using multiple criteria along three dimensions: general characteristics, key evidence components, and other descriptive or supplemental features. For example, researchers rate the strength of eight key components for group research on a 4-point
scale. These key components are measurement, comparison group, statistical significance of out- comes, educational and clinical significance, im- plementation fidelity, replication, site of implementation, and follow-up assessment. In ad- dition to providing an overall rating for each component, reviewers record additional informa- tion for most components. Regarding the com- parison group, for example, reviewers select the type of comparison group from a list of options; rate their confidence in determining the type of comparison group (from very low to very high); indicate how the researchers counterbalanced change agents (by change agent, statistical, other); check how the researchers established group equivalence (e.g., random assignment, post hoc matched set, statistical matching, post hoc test for group equivalence); and check whether and how mortality was equivalent between groups.
General Education. The WWC (2008) speci- fies that for randomized controlled trials to meet evidence standards without reservations, (a) re- searchers must randomly assign participants to conditions; (b) overall and differential attrition must not be high; (c) no evidence of intervention contamination (e.g., changed expectancy, novelty, disruption, local history event) exists; and (d) re- searchers avoid a teacher-intervention confound by either assigning more than one teacher to each condition or by presenting evidence that teacher effects are negligible. The WWC uses similar, but less stringent, criteria for randomized controlled trials and quasi-experimental studies to meet evi- dence standards with reservations.
MAGNITUDE OF EFFECT
Glinical Psychology. For a group design study to support a well-established or possibly effica- cious treatment, Chambless et al. (1998) require that treatment groups achieve outcomes that are statistically significantly superior to a control group or equivalent to a comparison group that received a treatment that researchers had previ- ously determined to be well-established. With re- gard to SSR, Chambless and Hollon (1998) suggest that "evaluators . . . carefully examine data graphs and draw their own conclusions about the efficacy of the intervention" (p. 13).
Exceptional Children 3 6 9
School Psychology. Because the Division 16 Task Force (Task Force on Evidence-Based Inter- ventions in School Psychology, 2003) coding pro- cédures do not classify interventions in terms of their effectiveness, no criteria are specified regard- ing magnitude of effect. However, reviewers code study characteristics related to significance of out- comes: statistical significance, educational and clinical significance, and effect size for group studies; and visual analysis, effect size, and educa- tional and clinical significance for SSR.
General Education. The WWC (n.d.b) uses five categories to describe the magnitude of effect for reviewed studies: statistically significant posi- tiye effects, substantively important positive effects, indeterminate effects, substantively im- portant negative effects, and statistically signifi- cant negative effects. Substantively important effects are educationally meaningful although not statistically significant; the WWC suggests using an effect size of greater than ±0.25 as a cutoff for substantively important effects. Indeterminate ef- fects are neither statistically significant nor have effect sizes greater than ±0.25.
CRITIQUES OF PROCESSES FOR
DETERMINING WHAT WORKS IN
OTHER E I ELDS
Although it is difficiilt to disagree with the gen- eral notion that "evidence should play a role in educational practice" (Slavin, 2008, p. 47), con- troversy seems to follow closely on the heels of proposals for establishing EBPs. Indeed, Kendall (1998) likened EBPs to religion and politics as lightning rods for conflict. Elliott (1998) noted that criticisms of EBPs tend to fall into one of two categories: concerns about the general en- deavor of designating EBPs and disagreements with the particular standards and criteria used. Al- though the first category includes many impor- tant issues (e.g.. Can research conclusively identify any practice as truly effective? Will ap- proaches not labeled as evidence-based be disre- garded?), this article focuses here on critiques of specific features of the three processes reviewed.
Waehler, Kalodner, Wampold, and Lichten- berg (2000) noted that some have criticized the Division 12 criteria for determining empirically validated treatments in clinical psychology for re-
lying too heavily on randomized clinical trials, psychological diagnoses, and adherence to treat- ment manuals, as well as for being too lenient. Scholars in school psychology also took issue with the Division 16 coding procedures as overwhelm- ing and overly complex (Durlak, 2002; Levin, 2002; Nelson Si Epstein, 2002; Stoiber, 2002); as seeming to endorse research designs that do not permit making causal inferences (Nelson & Ep- stein); and for producing ambiguous, descriptive reports rather than designating EBPs (Wampold, 2002). Finally, some researchers have criticized the WWC's (2008) standards as relying too heav- ily on randomized controlled trials, which are ex- tremely difficult to conduct in school settings (Kingsbury, 2006); as overly rigorous, resulting in few practices with positive effects identified (caus- ing some to refer to the WWC as the "'nothing works' clearinghouse," Viadero & Huff, 2006, p. 8); and as politically influenced (Schoenfeld, 2006).
Criticism regarding criteria and standards for determining EBPs may be unavoidable. Establish- ing EBPs involves addressing a number of ques- tions that lack any unequivocally correct answers and about which different stakeholders are bound to disagree. For example, requiring a large num- ber of randomized controlled trials that meet stringent methodological criteria and report large effect sizes will produce a high degree of confi- dence in practices shown to be evidence-based. However, this approach may be unnecessarily stringent, potentially excluding meaningful stud- ies. Yet designating practices as evidence-based be- cause of one study or a few research studies of any design without stringent methodological stan- dards invites false positives.
The categorization of practices represents an- other contentious issue for which multiple valid approaches may exist. Using a dichotomous sys- tem for labeling practices (e.g., evidence-based or not evidence-based) provides straightforward input for prioritizing instructional practices. However, a binary categorization scheme may overlook the complexities involved in interpreting bodies of research literature as well as promote the unfounded view that practices are either com- pletely effective or completely ineffective. In con- trast, whereas in-depth descriptions of a research base might facilitate nuanced and comprehensive
3 7 O Spring 2009
understanding, they may be of limited practical use for practitioners seeking guidance on how to teach in their classrooms the following day.
Any approach to determining what works in special education will inevitably have limitations. This recognition does not suggest that endeavors to establish EBPs are destined to fail. Rather, the strength of a system for determining EBPs lies in matching criteria and standards with the collec- tive traditions, values, and goals of the field that will use it. Therefore, special educators should de- sign a system for determining what works in spe- cial education based on the unique characteristics and needs of their field. Odom et al. (2005) en- deavored to delineate the "devilish details" (p. 138) of guidelines for determining EBPs rooted in the history and research traditions of special education.
P R O P O S E D G U I D E L I N E S F O R
E V I D E N C E - B A S E D P R A C T I C E S
I N S P E C I A L E D U C A T I O N
As an initial step for basing practice on research, the Division for Research of the Gouncil for Ex- ceptional Ghildren, under the leadership of Sam Odom, commissioned a series of papers that pro- posed quality indicators (QIs; i.e., features present in high-quality research studies) for four different research designs: group experimental studies (Gersten et al., 2005); SSR (Horner et al., 2005); correlational research (Thompson, Diamond, McWilliam, Snyder, & Snyder, 2005); and quali- tative research (Brantlinger, Jimenez, Klingner, Pugach, & Richardson, 2005). Gersten et al. also proposed standards for determining EBPs on the basis of group experimental/quasi-experimental research, and Horner et al. proposed standards for determining EBPs on the basis of SSR. Gonsid- ered together, the proposed QIs and standards constitute initial guidelines for establishing EBPs in special education. The number of prominent special education researchers who developed the proposed criteria and standards and the incorpo- ration of feedback from special education re- searchers who discussed the proposed criteria and standards at a Research Project Director's Meeting (hosted by the Office of Special Education Pro-
grams; Odom et al., 2004) enhances their credi- bihty.
The following sections examine the proposed guidelines for determining EBPs in special educa- tion and compare the proposed guidelines in spe- cial education with the systems for determining what works in clinical psychology, school psychol- ogy, and general education in relation to research design, quantity of research, methodological qual- ity, and magnitude of effect.
RESEARCH DESIGN
We assume that because standards for EBPs were proposed only for group experimental and quasi- experimental research (Gersten et al., 2005) and SSR (Horner et al., 2005), these research designs are the only ones to consider in determining whether a practice in special education is evi- dence-based. The Division for Research Task Force probably based this decision on the unique ability of these designs to exhibit experimental control (Gook, Tankersley, Gook, & Landrum, 2008). Special education, clinical psychology, and general education share many similarities in their treatment of research design in determining EBPs. For example, all three fields consider group experimental studies in determining EBPs. Re- searchers can also consider practices as evidence- based in special education, as well-established in clinical psychology, and as having potentially pos- itive effects (but not as having positive effects) in general education on the basis of SSR. However, whereas Gersten et al. allowed for quasi-experi- mental studies to constitute the sole research sup- port for EBPs in special education, Ghambless et al. (1998) did not consider quasi-experimental re- search in establishing empirically validated thera- pies in clinical psychology, and the WWG (n.d.b) requires at least one true experiment to support practices with positive effects in general educa- tion.
QUANTITY OE RESEARCH
Gersten et al. (2005) required a minimum of two high-quality group studies or four acceptable- quality group studies to consider a practice evi- dence-based or promising in special education. These numbers are similar to the quantity of group-design studies required for determining
Exceptional Children 3 7 1
EBPs in clinical psychology and general educa- tion. For example, Chambless et al. (1998) required two or more group studies for a well- established treatment, and the WWC (n.d.b) calls for two or more group design studies, at least one of which must be a randomized controlled trial, to support practices with positive effects.
To consider a practice to be evidence-based in special education, Horner et al. (2005) speci- fied a minimum of five SSR studies that involve a total of at least 20 total participants and that at least three different researchers conduct across at least three different geographical locations. This number is somewhat less than the number of SSR studies (« = 9) that Chambless et al. (1998) re- quired to deem a treatment in clinical psychology well established. By contrast, the WWC (2008) considers SSR studies as quasi-experimental de- signs, which cannot alone constitute sufficient ev- idence to deem a practice as having positive effects.
METHODOLOGICAL QUALITY
Cersten et al. (2005) proposed four essential QIs for group experimental research in the areas of de- scribing participants, implementing interventions and describing comparison conditions, measuring outcomes, and analyzing data. Each QI subsumes a number of specific criteria that a study must meet for it to address the QI. For example, to meet the QI oí describing participants, a study must address these three criteria:
1. Was sufficient information provided to determine/confirm whether the partici- pants demonstrated the disability(ies) or difficulties presented?
2. Were appropriate procedures used to in- crease the likelihood that relevant charac- teristics of participants in the sample were comparable across conditions?
3. Was sufFicient information characterizing the interventionists or teachers provided? Did it indicate whether they were compa- rable across conditions? (Gersten et al., p. 152)
Cersten et al. (2005) also proposed eight de- sirable QIs related to attrition, reliability and data collectors, outcome measures beyond posttest, va-
lidity, detailed assessment of implementation fi- delity, nature of instruction in comparison condi- tion, audiotape or videotape excerpts regarding the intervention, and presentation of results. In addition to meeting all the essential QIs, high- quality group studies must address at least four of the desirable QIs. Acceptable studies must meet only one of the desirable QIs in addition to ad- dressing all but one of the essential QIs.
The QIs for group studies that Cersten et al. (2005) proposed are somewhat distinct from the criteria for high-quality group research used in other fields. For example, among the study fea- tures required for a high-quality group study in special education that the WWC (2008) does not require for a group study that meets evidence standards without reservations in general educa- tion are
• Detailed descriptions of participants, setting, and independent variable, and services pro- vided in the comparison group.
• The use of multiple outcome measures col- lected at appropriate times.
• Documentation of implementation fidelity
• Appropriate units of analysis (although WWC reviews must note misalignment be- tween units of assignment and units of analy- sis).
Among the features that the WWC requires for a study that meets evidence standards without reservations but that Gersten et al. does not re- quire for high-quality group studies are overall and differential attrition not severe or accounted for (although Cersten et al. included attrition as a desirable QI), and no intervention contamina- tion. Both sets of criteria for high-quality group studies require researchers to demonstrate the comparability of interventionists across condi- tions.
Horner et al. (2005) proposed QIs for SSR in special education in seven areas: describing par- ticipants and settings, dependent variable, inde- pendent variable, baseline, experimental control and internal validity, external validity, and social validity. Horner et al. proposed 21 criteria to as- sess the presence of these QIs. For example, to meet the dependent variable QI, a study must meet the following criteria:
3 7 2 Spring 2009
1. Dependent variahles are described with operational precision.
2. Each dependent variable is measured with a ptocedure that generates a quantifiable index.
3. Measurement of the dependent variable is valid and described with replicable preci- sion.
4. Dependent variables are measured repeat- edly over time.
5. Data are collected on the reliability or in- terobserver agreement (IOA) associated with each dependent variable, and IOA levels must meet minimal standards (e.g., IOA = 80%, Kappa = 60%). (Horner et al., p. 174)
Horner et al. (2005) indicated that reviewers use the QIs, "for detetmining if a study meets the 'acceptable' methodological rigor needed to be a credible example of SSR" (p. 173). Horner et al. do not explicitly state whether studies must meet all the QIs to be considered of acceptable methodological quality, although we infet that they must. The QIs for high-quality SSR studies in special education overlap somewhat with the criteria for studies that suppott empitically vali- dated treatments in clinical psychology. Both Chambless et al. (1998) and Horner et al. require that researchers clearly describe participant chat- actetistics and use an apptopriate SSR design. Chambless et al. tequite that researchers compare the intetvention with a placebo ot anothet treat- ment and conduct the intetvention by using tteatment manuals, whereas Horner et al. do not (although Horner et al. do require that tesearchers overtly measure fidelity of implementation of the independent variable). Horner et al. require a number of criteria that Chambless et al. do not call fot, such as desctiption of physical location, desctiption of the dependent variable with tepli- cable precision, acceptable levels of interobserver agreement regarding the dependent vatiable, and documentation of the external and social validity of the dependent vatiable.
MAGNITUDE OF EFFECT
For a practice to be considered evidence-based in special education, Cetsten et al. (2005) ptoposed
that the weighted effect size of group experimen- tal studies should be significantly greater than zero. We presume that this effect size derives from only those studies found to be acceptable or of high quality vis-a-vis the QIs. Fot promising prac- tices, Getsten et al. required that a 20% confi- dence intetval for the weighted effect size across studies be gteatet than zeto. In contrast, both clinical psychology (Chambless et al., 1998) and general education (WWC, n.d.b) use statistical significance as the standatd to judge whether group studies support well-established tteatments and ptactices with positive effects, tespectively. The WWC does consider effect sizes (e.g., d S 0.25) in the absence of statistically significant findings fot determining that a practice has po- tentially positive effects.
Horner et al. (2005) did not prescribe a par- ticular effect size needed fot SSR studies to sup- pott a practice. However, for the authors to consider a practice as evidence-based on the basis of SSR in special education, they required a docu- mented causal ot functional telationship between use of the ptactice and change in a socially impor- tant dependent vatiable. Horner et al. suggested that visual analysis "of the level, ttend, and vati- ability of petfonnance occurring during baseline and intervention conditions" (p. 171) establishes a functional relationship. Visual inspection of gtaphic displays of student behaviot involves the following:
1. Immediacy of effects following the onset and withdrawal of the practice.
2. Overlap of data points in adjacent phases.
3. Magnitude of change in the depen- dent variable.
4. Consistency of data patterns across conditions (Horner et al.).
Chambless and Hollon (1998) similarly sug- gested using visual inspection criteria to deter- mine the effect for SSR studies in clinical psychology. The WWC (2008) is developing guidelines, which are not yet available, for assess- ing the magnitude of effect in SSR studies.
Exceptional Children 3 7 3
A P P L I C A T I O N S O F P R O P O S E D
G U I D E L I N E S F O R D E T E R M I N I N G
E B P S I N S P E C I A L E D U C A T I O N
Gersten et al. (2005) suggested that their pro- posed criteria and standards for determining EBPs in special education were "merely a first step," which researchers should refine, "based on field- testing" (p. 163). In response, special education researchers have begun to use the proposed QIs and standards in reviews and analyses of research literature. For example, Browder, Wakeman, Spooner, Ahlgrim-Delzell, and Algozzine (2006) applied the QIs and standards for EBPs proposed by Cersten et al. (2005) and Horner et al. (2005) to 128 intervention studies (88 SSR studies and 40 group quasi-experimental studies) that investi- gated reading outcomes for individuals with sig- nificant cognitive disabilities. Browder et al. condensed the seven QIs and 21 criteria that Horner et al. proposed for SSR into four cate- gories:
• Dependent variable operationally defined and included data on reliability.
• Methods adequately described.
• Data collected on procedural fidelity.
• Baseline and experimental control (with par- ticular focus on, between, and within partici- pant replications).
Two coders independently coded the pres- ence of these four categories for all 88 SSR stud- ies. Interrater agreement was 100% in each category except procedural fidelity, for which in- terrater agreement was 93%. Fifty-six of the SSR studies met all four of Browder et al.'s (2006) cat- egories of QIs for SSR. From these studies, massed trial as well as systematic prompting met Horner et al.'s (2005) standards for an EBP (i.e., at least five supporting studies involving a mini- mum of 20 total participants, conducted by at least three different researchers in at least three different locations) for the outcomes oí sight- word vocabulary, picture vocabulary, and compre- hension. The researchers determined that time delay was also an EBP for sight-word vocabulary and fiuency and that pictures were an EBP for comprehension for the target population.
Browder et al. (2006) also clustered the four essential and eight desirable QIs that Gersten et
al. (2005) proposed for group research into four categories:
• Outcome measures—operationally defined and evidence of reliability and validity.
• Intervention clearly defined.
• Measure of procedural fidelity.
• Use of comparison group and intervention defined.
Because of the perceived level of judgment required to code these methodological categories, Browder et al. (2006) used a consensus model to establish reliability. In this consensus model, two coders discussed coding decisions until they reached agreement. Therefore, Browder et al. did not report interrater reliability (IRR). Only 2 of the 40 group studies met all four of Browder et al.'s categories for group studies, with no particu- lar practice having sufficient empirical support to be considered evidence-based.
In reviewing the empirical literature on inter- ventions aimed at improving self-advocacy for students with disabilities. Test, Fowler, Brewer, and Wood (2005) assessed the presence of the QIs that Gersten et al. (2005) proposed in 11 group experimental studies and that Horner et al. (2005) proposed in 11 SSR studies. Test et al. found high levels of IRR for coding the QIs in a subset of studies: means of 98.5% agreement for SSR and 98.7% agreement for group experimen- tal research. Test et al. reported that only one of the 11 SSR studies that they reviewed met all the QIs. Although most of the SSR studies met most QIs, only six sufficiently described how partici- pants were selected and only two described and measured procedural fidelity. Test et al. assessed 23 criteria for group studies, examining essential and desirable QIs together and including criteria regarding the conceptualization of a study (that Gersten et al. included in their QIs for research proposals). None of the group studies met all or all but one of Test et al.'s criteria. Among the cri- teria that few studies met: data collectors unfamil- iar with study conditions {n = 4), data collectors unfamiliar with participants (« = 4), documenta- tion of attrition {n = 3), clear descriptions of the difference between intervention and control (« = 3), and measures of procedural fidelity (n = 1). Test et al. did not apply Gersten et al.'s proposed
3 7 4 Spring 2009
Standards to determine whether any practices evaluated in the reviewed studies were evidence- based.
On a smaller scale, we applied the proposed Qls, as literally as possible, to two group experi- mental studies (B. C. Cook & Tankersley, 2007) and two SSR studies (Tankersley, Cook, & Cook, 2008). Although the small scope of these pilot projects limits their generalizability, our applica- tion of Cersten et al.'s (2005) criteria indicated that the group experimental studies that we re- viewed met 40% of the QI components; whereas the SSR studies that we reviewed met 48% of the QI components that Horner et al. (2005) pro- posed. We reported a moderately low IRR of .69 for SSR QI components (Tankersley et al.; we used a consensus model and did not assess IRR for the group Qls). We found that reliably deter- mining whether studies addressed many of the proposed Qls was difficult because of incomplete and ambiguous reporting in the articles reviewed and because of the lack of specificity and clarity (e.g., operationalized definitions) in the proposed Qls (B. C. Cook & Tankersley; Tankersley et al.).
It is encouraging that both Browder et al. (2006) and Test et al. (2005) applied the pro- posed Qls and reported high IRR in their coding. However, it is important to note that Browder et al. did not apply all the Qls. Furthermore, Test et al. made no distinction berween essential and de- sirable Qls for group studies and did not apply standards for determining EBPs. Thus, to our knowledge, no published studies have applied the specific Qls and standards for EBPs as Cersten et al. (2005) and Horner et al. (2005) proposed across an entire body of research literature. Clearly, to meaningfully determine the feasibility of applying the proposed Qls and standards and to identify aspects of the Qls and standards that researchers might fruitfully reflne, researchers should conduct additional field tests.
S U M M A R Y A N D A N A L Y S I S
O F F I V E F I E L D T E S T S O F
D E T E R M I N I N G E B P S I N
S P E C I A L E D U C A T I O N
We asked five teams of expert reviewers to faith- fully apply the Qls and standards for EBPs, pro-
posed by Cersten et al. (2005) and Horner et al. (2005), to bodies of research literature on inter- ventions relevant to their fields of expertise. Re- view teams evaluated the intervention literature on five interventions frequently used with stu- dents with disabilities: cognitive strategy instruc- tion (Montague &C Dietz, 2009); repeated reading (Chard, Ketterlin-Celler, Baker, Doabler, & Apichatabutra, 2009); self-regulated strategy de- velopment (Baker, Chard, Ketterlin-Celler, Apichatabutra, & Doabler, 2009); time delay (Browder, Ahlgrim-Delzell, Spooner, Mims, & Baker, 2009); and function-based interventions (Lane, Kalberg, & Shepcaro, 2009). This section summarizes and analyzes the approaches and findings of these five reviews with the goal of making preliminary recommendations for refin- ing the proposed Qls and standards for EBPs, as well as the process for applying them.
SCOPE OF REVIEW
Reviewers initially had to determine whether and how to delimit the scope of their review. As Brow- der et al. (2009) suggests, reviewers might delimit "the specific population of focus, the scope of the dependent variable to be considered . . ., and other aspects of the studies" (p. 360). Four of the five review teams identified a target population more specific than students with disabilities— Baker et al. (2009) and Chard et al. (2009) re- viewed studies involving students with and at risk for learning disabilities, Browder et al. focused on students with significant cognitive disabilities, and Lane et al. (2009) targeted students with or at risk for emotional and behavioral disorders. Only Lane et al. included an age parameter, re- viewing outcomes for secondary students only. The review teams varied in the degree to which they set parameters for dependent variables. Browder et al. reviewed only studies that specifi- cally assessed picture or word recognition. Mon- tague and Dietz (2009) and Baker et al. stated more general outcome parameters for their re- views—mathematical problem solving and writ- ing performance, respectively. Although Chard et al. and Lane et al. did not specify such outcome variables as inclusion criteria for their reviews, their interventions are associated with particular
Exceptional Children 3 7 5
outcome areas (i.e., reading for Ghard et al. and behavioral outcomes for Lane et al.).
The specific parameters that reviewers apply represent an important concern. Using overly broad parameters (e.g., students with or at risk for disabilities) may not address such critical questions as for whom the practice works with sufficient specificity for research consumers. Gonversely, overly narrow parameters may reduce the number of studies available and limit the implications of the review. Although we realize that a variety of sensible rationales exists, for focusing reviews on specific groups, in the absence of a compelling ra- tionale, we recommend that reviews focus on as broad a population as seems reasonable and meaningful and that authors carefully describe participants across studies reviewed to inform consumers about the population for whom the intervention has been shown to be efFective.
DETERMINING THE PRESENCE
OE QUALITY INDICATORS
A particular element of high-quality research is often neither completely present nor completely absent in a research report but instead is partially present. Recognizing this issue. Baker et al. (2009) and Ghard et al. (2009) collaboratively constructed 4-point rubrics for rating the pres- ence of Ql components for group experiments and SSR. The other review teams rated each com- ponent dichotomously, as met or not met. Be- cause relatively low IRR was associated with using the 4-point rubric, we recommend that future re- views use a dichotomous approach for classifying the presence of QIs, at least until reviewers refine a more detailed rubric that they can use with greater reliability.
Ultimately, the method of choice for identi- fying the presence of methodological QIs may be a philosophical issue. If the purpose of the reviews is to provide in-depth descriptions of a research base, the use of a rubric—perhaps supplemented with descriptions of the strengths and weaknesses of the literature base for each Ql—may be desir- able. Alternatively, if the main intent of the re- views is to yield a straightforward decision about whether a practice is evidence-based, the benefit of additional information gained by using a more detailed rating system may not be worth the cost
of extra time involved in assessing and reporting the information or the possibility of decreased IRR. Of course, the goals of providing in-depth information on a research base and categorizing practices as evidence-based are not mutually ex- clusive. Future reviewers in special education may want to provide descriptions, use a multiple-point rating system, and employ a "yes/no checklist" ap- proach, thereby generating reviews to serve differ- ent purposes for different audiences.
The Division for Research of the Gouncil for Exceptional Ghildren asked Gersten et al. (2005) and Horner et al. (2005) to identify and briefiy describe, not operationally define, sets of QIs (S. L. Odom, personal communication, April 7, 2006) in their development of the QIs. Accord- ingly, Gersten et al. and Horner et al. stated some of the QIs somewhat subjectively. For example, Horner et al. required that the dependent variable be practical and cost-effective but did not provide concrete guidelines for determining practicality or cost-effectiveness. Accordingly, many of the re- view teams interpreted, and in some cases modi- fied, the QIs for their reviews. For example. Lane et al. (2009) required that researchers explicitly describe the cost-effectiveness of their interven- tion. At times, review teams also expanded on the QIs. For instance, Browder et al. (2009) specified that not only must researchers overtly measure implementation fidelity but that they must also document a minimum level of 80%. Lane et al. also required that all components for the internal validity Ql for SSR studies be met as a precon- dition for the external validity Ql. In other situa- tions, review teams for this issue reduced the criteria for certain QIs (e.g.. Lane et al., 2009, and Montague & Dietz, 2009, set their criteria at 3 data points for baseline, as opposed to the 5 points that Horner et al. suggested). Browder et al. also adapted some of the SSR QIs for the spe- cific outcomes of their review (e.g., they defined the socially important change component as learn- ing at least five new words or pictures).
Browder et al. (2009) suggested that adapt- ing the QIs to optimize their applicability for the intervention being reviewed should be a critical component of each review. Determining whether the QIs can be sufficiently specific and opera- tionálized to yield reliable ratings yet flexible enough to apply meaningfully to a wide variety of
3 7 6 Spring 2009
TABLE 1
Summary of Single-Subject Research Quality Indicators Rated as Present
Quality Indicator
Participants/setting
Dependent variahle
Independent variahle
Baseline
Internal validity
External validity
Social validity
Total
Cbard Ketterlin-Geller,
Baker, Doabler, &
Apichatabutra (2009)
1/6
3/6
2/6
3/6
4/6
0/6 4/6
17/42, 40%
Lane, Kalberg, & Shepcaro (2009)
1/12
5/12
6/12
7/12
2/12
1/12
1/12
23/84, 27%
Browder, Ahlgrim-Dekell,
Spooner, Mims, &
Baker (2009)
28/30
30/30
26/30
29/30
30/30
29/30
28/30
200/210,95%
Montague & Dietz (2009)
5/5
1/5
0/5
5/5
5/5
5/5
5/5
26/35,1^%
Baker, Chard,
Ketterlin- Geller, Apichatabutra,
& Doabler (2009)
8/9
9/9
7/9
8/9
9/9
9/9
9/9
59/63,94%
Total
43/62, 69%
48/62, 77%
41/62, 66%
52/62, 84%
50/62,81%
44/62, 71%
47/62, 76%
325/434, 75%
studies will be a considerable challenge. Indeed, perhaps some freedom to adapt QIs may be ap- propriate for certain reviews. We are concerned, however, that giving review teams too much lati- tude to interpret and adapt QIs may, in some sit- uations, result in reviews that vary considerably in their rigor and findings.
EiNDINGS
Quality Indicators. Table 1 summarizes the number of studies meeting Horner et al.'s (2005) QIs for SSR, and Table 2 reports the same infor- mation for Cersten et al.'s (2005) QIs for group experimental research (Browder et al., 2009, and Lane et al., 2009, reviewed only SSR studies). The review teams reported widely discrepant findings as to how frequently the studies reviewed met the QIs. The proportion of QIs met in spe- cific reviews ranged from 27% to 95% for SSR studies and from 12.5% to 95% for group experi- mental studies. It is noteworthy that these consid- erable disparities were not associated with differences in the rating procedure used. That is, although they all used a dichotomous approach for identifying the presence of SSR QIs, Lane et al. found almost three fourths of QIs absent in the studies that they reviewed, whereas Browder et al. and Montague and Dietz (2009) indicated that almost all QIs were present in the studies that they reviewed. Moreover, both Baker et al.
(2009) and Chard et al. (2009) used a 4-point rubric to identify the presence of QIs. However, Chard et al. found that only 25% of the QIs for group experiments were present in the five studies that they reviewed, whereas Baker et al. reported that 95% were present in the five group studies that they reviewed. The disparities in identified QIs may simply reflect significant variation in the methodological quality of the bodies of literature reviewed. Since many of the QIs are not opera- tionally defined, another possibility is that review teams systematically varied in their interpretation of the QIs.
In comparison with the wide discrepancies of QIs met between reviews, the variance in specific QIs present across the studies reviewed was mini- mal. Of the SSR QIs, the most frequently met was baseline (achieved in 52 of a total of 62 SSR studies reviewed), whereas the least frequently met was independent variable (41 of the 62 SSR studies met this QI). For group experimental studies, the number of total studies that met a QI ranged from 5 (of 12 total studies reviewed) for independent variable/comparison condition to 7 for participants and outcome variable. Across the stud- ies reviewed, the SSR studies met a much higher proportion of QIs than group experiments did. This outcome may have occurred because of dif- ferences in the quality of the studies reviewed, dif- ferences in the rigor required by the two sets of
Exceptional Children
TABLE 2
Summary of Group Experimental Research Quality Indicators Rated as Present
Quality Indicator
Participants
Independent variable/comparison condition
Outcome measure
Data analysis
Total
Chard,
Ketterlin-Geller,
Baker,
Doabler, &
Apichatabutra
(2009)
1/5 1/5 2/5 1/5
5/20, 25%
Montague & Dietz (2009)
1/2
0/2
0/2
0/2
1/8, 12.5%
Baker,
Chard, Ketterlin-Geller,
Apichatabutra,
& Doabler
(2009)
5/5 4/5 5/5 5/5
19/20,95%
Total
im, 58% 5/12, 42% 7/12, 58% 6/12, 50% 25/48, 52%
QIs, or both. The disproportionately high num- ber of SSR studies teviewed, in comparison with group experiments, may indicate a telative dearth of group experiments in the special education lit- erature (Seethaler & Fuchs, 2005). The small number of group experiments conducted in spe- cial education appears to pose a particular con- cern for those wishing to establish EBPs in the field, given that the tesults of group experimental research figure prominently in this process.
The identification of components that re- searchers addressed least often can suggest areas of focus for futute teseatchets to imptove the methodological rigor of intetvention tesearch in the field of special education. The least fte- quently addressed component of the dependent variahle QI fot SSR studies appears to be appto- priate documentation of IRR. Issues related to implementation fidelity were clearly the primary reason that studies did not meet the independent variahle QI, with each team of reviewers report- ing that multiple SSR studies reviewed did not meet this component. For group studies, the least frequently addtessed component for the interven- tion/comparison condition QI was also implemen- tation fidelity. The primary shortcoming of group experiments for the outcome variahle QI appears to be not using multiple dependent measutes, at least one of which does not tightly align with the independent vatiable. And the sole reason that group experimental studies reviewed did not meet the data analysis QI was failure to teport effect sizes. It is impottant to note that this special issue reviewed a relatively small number of studies.
especially group experimental studies, and that the studies may not tepresent the latger pool of intervention research in special education, sug- gesting that these methodological concerns may not be generalizable.
Interrater Reliability. Unlike the proportion of QIs met, IRR did appear to vary according to the method used to täte the presence of the QIs. Cenetally, the three teviews that categotized QIs dichotomously (i.e., present or absent) tepotted relatively high levels of IRR. For example. Lane et al. (2009) reported 100% IRR for 15 of the 21 SSR components, with only one component falling below 8 3 % (IRR for the component change in dependent variahle is socially valid was 75%). Browder et al. (2009) reported a mean IRR across SSR QI components of 97%, with a tange ftom 83% to 100%. And Montague arid Dietz (2009) tepotted a mean IRR of 93% across QIs for SSR studies and 77% for group studies. In contrast. Baker et al. (2009) and Chard et al. (2009)—both of whom used a 4-point tubric to rate the presence of QIs—reported IRR of .36 and 62% for SSR studies and .53 and 77% for group studies. Although IRRs for these two re- views were much higher when allowing for 1- point disctepancies, the teliability for determining the presence of QIs appears to be meaningfully lower when using a 4-point tubtic, which is not unexpected, given that Chard et al. tepotted some difficulties in discriminating between the multiple tating levels. No systematic differences appear to exist fot IRR between SSR and group studies. In the thtee reviews that consideted both types of
3 7 8 Spring 2009
research. Baker et al. and Ghard et al. reported higher IRR for the group studies, whereas Mon- tague and Dietz indicated higher IRR for SSR studies.
RECOMMENDATIONS FOR REFINING
THE PROCESS
On the basis of their experiences applying Ger- sten et al.'s (2005) and Horner et al.'s (2005) QIs and standards for determining EBPs in special ed- ucation, the review teams for this issue made a number of recommendations for refining the pro- cess. The reviewers suggested adding some new QIs or making some of the existing QIs and their components more rigorous. For example, Ghard et al. (2009) proposed requiring researchers to de- scribe the theoretical or conceptual framework for the intervention reviewed (see also Browder et al., 2009). And Montague and Dietz (2009) advo- cated that researchers specify inclusion and exclu- sion criteria for selecting participants, assess treatment fidelity with at least two impartial ob- servers with interrater agreement of at least 80%, and report effect sizes for SSR. Moreover, both Ghard et al. and Montague and Dietz suggested making some of the desirable QIs for group ex- periments, such as documenting validity of mea- surement instruments and minimal attrition, essential QIs.
In contrast to these calls for additional or more rigorous QIs, Lane et al. (2009) suggested that some of the SSR QIs might be overly rigor- ous. They recommended, for example, that re- searchers reconsider the requirements for documenting the instruments and process used to determine the disability of participants and de- scribing the cost-effectiveness of the intervention in SSR studies. Lane et al. also advocated that the field consider requiring less than 100% of com- ponents for meeting a QI, perhaps using an 80% criterion.
Review teams also noted the need for greater operationalization of the QIs and their compo- nents. In particular, Montague and Dietz (2009) called for greater clarity with regard to what con- stitutes a typical intervention agent in SSR studies. Baker et al. (2009) provided another suggestion for improving the ability of reviewers to determine the presence of QIs in reports of research—furnish
opportunities for researchers, perhaps on Web sites linked to the Journal, to give additional, detailed information that might otherwise go unreported because of space limitations.
In regard to standards for EBPs, Lane et al. (2009) raised the issue of whether all QIs were equally important, and if not, whether they might be weighted differentially in determining EBPs (see also Montague & Dietz, 2009). Montague and Dietz also suggested that researchers might develop standards for determining when to con- sider an EBP evidence-based for subpopulations (e.g., how many studies involving students with a particular disability are necessary to demonstrate that the intervention is evidence-based for that population?).
These recommendations all appear to have merit and warrant further consideration while special educators work toward refining the process for determining EBPs in special education. How- ever, we also advise caution in revising the QIs and the process for establishing EBPs too readily or repeatedly. Special educators can and should refine the QIs and standards, perhaps periodically over time, to optimize their efficiency, reliability, and validity. For example, we endorse the idea that the QIs should be further operationalized—a process that the Gouncil for Exceptional Ghildren has undertaken (Bruno, 2007). However, no sin- gle set of QIs or standards will meet every pur- pose; and for the most part, the review teams found the application of the proposed QIs and standards feasible and meaningful. When the QIs and standards have been refined and vetted through what we envision as an iterative but lim- ited sequence of field trials, stability and consis- tency in the QIs and standards for EBPs in the field will be of significant importance.
C O N C L U S I O N
The authors of the five reviews in this topical issue took on a task that posed multiple chal- lenges. The review teams not only had to system- atically review a large number of studies, but they did so by using criteria that often required inter- pretation while they devised their own processes for field-testing the proposed QIs and standards for EBPs in special education. Not surprisingly.
Exceptional Children 3 7 9
this process was time-consuming—Browder et al. (2009) estimated that their review team devoted more than 400 hours to their review. The review- ers also no doubt found the review process diffi- cult because we asked them to apply the Qls literally. At times, literally applying the Qls may have seemed to highlight limitations in the re- search of respected colleagties. It is important to note that authors of previous research wrote the body of extant research without foreknowl- edge of the future standards of methodological rigor to which it might be held and that they con- formed to external requirements of the day (e.g., little emphasis on reporting effect sizes; the per- petual space limitations in journals). Nonetheless, the results of these reviews have provided the first large-scale application of the Qls and the stan- dards for EBPs in special education. We appreci- ate and applaud the work of the reviewers and the scholars who conducted the original research re- viewed, as well as the pioneering work of Cersten et al. (2005) and Horner et al. (2005).
Collectively, the application of Qls to deter- mine high-quality group research and SSR across five bodies of special education intervention re- search indicate the following:
• Approximately three quarters of the SSR Qls were present across studies reviewed, whereas approximately one half of group experimen- tal Qls were present.
• Considerable variability existed between re- views in the proportion of Qls met. The rat- ing procedure used did not appear to explain this variability.
• The IRR for rating Qls varied markedly be- tween reviews, although reviews using a di- chotomous yes/no scheme for identifying Qls tended to yield adequate IRR.
Reviewers also made a number of suggestions for refining the Qls, such as operationalizing them, adding and deleting particular components of some Qls, and weighting the Qls according to their importance. In addition to considering these and other technical matters (e.g.. Should reviews be restricted to articles published in peer-reviewed journals?), special education leaders will need to address some foundational issues regarding the need for and merits of determining EBPs in spe-
cial education so that they can garner the broad support of the special education community for this process.
The philosophical objections to EBPs that we have heard from special educators often parallel criticisms raised regarding the advent of evidence- based medicine. As described by Sackett et al. (1996), "criticism has ranged from evidence based medicine being old hat to it being a dangerous in- novation, perpetrated by the arrogant to . . . sup- press clinical freedom" (p. 71). Civen the documented research-to-practice gap (e.g., B. C. Cook & Schirmer, 2003), the claim that EBPs are old hat seems unwarranted in special education. As for concerns that EBPs in special education will force instruction to conform to an approved menu of interventions, we believe that EBPs will not and should not ever take the place of profes- sional judgment but can be used to inform and enhance the decision making of special education teachers. As Sackett et al. suggested for evidence- based medicine.
Good doctors use both individual clinical ex- pertise and the best available external evi- dence, and neither alone is enough. Without clinical expertise, practice risks becoming tyrannised by evidence, for even excellent external evidence may be inapplicable to or inappropriate for an individual patient. Without current best evidence, practice risks becoming rapidly out of date, to the detri- ment of patients, (p. 71)
Likewise, we in no way imagine evidence- based special educators being directed as to when and in what situations they can or cannot use par- ticular teaching practices. Instead, EBPs should interface with the professional wisdom of teachers to maximize the outcomes of students with dis- abilities (Cook, Tankersley, & Harjusola-Webb, 2008).
EBPs should interface with the
professional wisdom of teachers to maximize the outcomes of students with disabilities.
We concur, then, with Sackett et al.'s (1996) declaration that, "clinicians who fear top down cookbooks will find the advocates of evidence
3 8 O Spring 2009
based medicine [or special education] joining them at the barricades" (p. 72). However, al- though we recognize the dangers of overempha- sizing EBPs in a field premised on individualized instruction, we believe that special educators would be remiss if they did not make every effort to prioritize practices shown by our best research to result in meaningful improvements in student outcomes. Identifying practices that are evidence- based for students with disabilities is a necessary but insufficient step in a process that we hope will culminate in the consistent implementation of the most effective practices with fidelity, ulti- mately resulting in improved outcomes for stu- dents with disabilities.
R E F E R E N C E S
American Psychological Association. (2001). Publica- tion manual of the American Psychological Association (5th ed.). Washington, DC: Author.
Baker, S. K., Chard, D. J., Kerterlin-Geller, L. R., Apichatabutra, C , & Doabler, C. (2009). Teaching writing to at-risk students: The qualiry of evidence for self-regulated strategy development. Exceptional Chil- dren, 75,
Berliner, D. C. (2002). Educational research: The hard- est science of all. Educational Research, 31{8), 18-20.
Brandinger, E., Jimenez, R., Klingner, J., Pugach, M., ßc Richardson, V. (2005). Qualitative studies in special education. Exceptional Children, 71, 195-207.
Browder, D., Ahlgrim-Delzell, L., Spooner, E, Mims, P. J., & Baker, J. N. (2009). Using time delay to teach lit- eracy to students with severe developmental disabilities. Exceptional Children, 75, 343-364.
Browder, D. M., Wakeman, S. Y., Spooner, E., Ahlgrim-Delzell, L., & Algozzine, B. (2006). Research on reading instruction for individuals with significant cognitive disabilities. Exceptional Children, 72, 392- 408.
Bruno, R. (2007). CEC's evidence based practice effort Retrieved September 29, 2008. from htrp://education. u o r e g o n . e d u / g r a n t m a t t e r s / p d f / D R / S h o w c a s e / Bruno.ppr.
Chambless, D. L., Baker, M. J., Baucom, D. H., Beut- ler, L. E., Calhoun, K. S., Crirs-Christoph, P., et al. (1998). Update on empirically validated therapies, II. The Clinical Psychologist, 51, 3—16.
Chambless, D. L, & Hollon, S. D. (1998). Defming empirically supported therapies. Journal of Consulting and Clinical Psychology, 66, 7-18.
Chambless, D. L., Sanderson, W C , Shoham, V., Ben- nett Johnson, S., Pope, K. S., Crits-Christoph, P., et al. (1996). An update on empirically validated therapies. The Clinical Psychologist, 49, 5 - 1 8 .
Chard, D. J., Ketterlin-Celler, L. R., Baker, S. K., Doabler, C , & Apichatabutra, C. (2009). Repeated reading interventions for students with learning disabil- ities: Status of the evidence. Exceptional Children, 75, 263-281.
Cohen, J. (1988). Statistical power analysis for the be- havioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Ed- baum.
Cook, B. C , &c Schirmer, B. R. (Eds.). (2003). What is special about special education [Special issue]. The Journal of Special Education, 37(3).
Cook, B. G., & Tankersley, M. (2007). A preliminary examination to identify rhe presence of qualiry indica- tors in experimental research in special education. In J. Crockett, M. M. Cerber, & T. J. Landrum (Eds.), Achieving the radical reform of special education: Essays in honor of James M. Kauffman (pp. 189-212). Mahwah, NJ: Lawrence Erlbaum.
Cook, B. G., Tankersley, M., Cook, L., & Landrum, T. J. (2008). Evidence-based practices in special educa- tion: Some practical considerations. Intervention in School and Clinic, 44{T), 69-75.
Cook, B. G., Tankersley, M., & Harjusola-Webb, S. (2008). Evidence-based practice and professional wis- dom: Putting it all together. Intervention in School and Clinic, 44{2), 105-111.
Cook, L. H., Cook, B. G., Landrum, T J., «¿Tankers- ley, M. (2008). Examining the role of group experi- mental research in establishing evidenced-based practices. Intervention in School and Clinic, 44{2), 76-82.
Dammann, J. E., & Vaughn, S. (2001). Science and sanity in special education. Behavioral Disorders, 27, 21-29.
Dudak, J. A. (2002). Evaluating evidence-based inter- ventions in school psychology. School Psychology Quar- terly, 17, 475-482.
Elliott, R. (1998). Editor's introduction: A guide to empirically supported treatments controversy. Psy- chotherapy Research, 8, 115-125.
Forness, S. R., Kavale, K. A., Blum, I. M., & Lloyd, J. W. (1997). What works in special education and re- lated services: Using meta-analysis to guide practice. TEACHING Exceptional Children, 29, 4-9.
Exceptional Children 3 8 1
Gallagher, D. J. (1998). The scientific knowledge base of special education: Do we know what we think we know? Exceptional Children, 64, 493-502.
Gersten, R., Fuchs, L. S., Compton, D., Coyne, M., Greenwood, C , & Innocenti, M. S. (2005). Quality indicators for group experimental and quasi-experi- mental research in special education. Exceptional Chil- dren, 71, 149-164.
Horner, R. H., Carr, E. G., Halle, J., McGee, G., Odom, S., & Wolery, M. (2005). The use of single- subject research to identify evidence-hased practice in special education. Exceptional Children, 71, 165—179.
Individuals With Disabilities Education Act, 20 U.S.C. § 1400 et seq. (2004).
Kauffman, J. M. (1996). Research to practice issues. Behavioral Disorders, 22, 55—60.
Kendall, P. C. (1998). Empirically supported psycho- logical therapies. Journal of Consulting and Clinical Psy- chology, 66, 3-6.
Kingsbury, G. G. (2006). The medical research model: No magic formula. Educational Leadership, 63{6), 79-82.
Kratochwill, T R., & Stoiber, K. C. (2002). Evidence- based interventions in school psychology: Conceptual foundations of the Procedural and Coding Manual of Division 16 and the Society for the Study of School Psychology Task Force. School Psychology Quarterly, 17, 341-389.
Lane, K. L., Kalberg, J. R., & Shepcaro, J. C. (2009). An examination of the evidence base for function-based interventions for students with emotional or behavioral disorders attending middle and high schools. Excep- tional Children, 75, 321-340.
Levin, J. R. (2002). How to evaluate the evidence of evidence-based interventions. School Psychology Quar- terly, 17, 483-492.
Lloyd, J. W, Pullen, P C , Tankersley, M., & Lloyd, P A. (2006). Critical dimensions of experimental studies and research syntheses that help defme effective prac- tices. In B. G. Gook & B. R. Schirmer (Eds.), What is special about special education: The role of evidence-based practices (pp. 136-153). Austin, TX: PRO-ED.
Montague, M., & Dietz, S. (2009). Evaluating the evi- dence base for cognitive strategy instruction and math- ematical problem solving. Exceptional Children, 75, 285-302.
Nelson, J. R., & Epstein, M. H. (2002). Report on evi- dence-based interventions: Recommended next steps. School Psychology Quarterly, 17, 493-499.
No Child Left Behind, 20 U.S.C. § 16301 et seq. (2001).
Odom, S. L., Brantlinger, E., Gersten, R., Horner, R. H., Thompson, B., & Harris, K. (2004). Quality indi- cators for research in special education and guidelines for evidence-based practices: Executive summary. Retrieved September 29, 2008, from education, uoregon.edu/ grantmatters/pdf/DR/Exec_Summary.pdf
Odom, S. L., Brandinger, E., Gersten, R., Horner, R. H., Thompson, B., & Harris, K. R. (2005). Research in special education: Scientific methods and evidence- based practices. Exceptional Children, 71, 137—148.
Rumrill, P D., & Cook, B. G. (Eds.). (2001). Research in special education: Designs, methods and applications. Springfield, IL: Charles C Thomas.
Sackett, D. L., Rosenberg, W. M. C , Gray, J. A. M., Haynes, R. B., & Richardson, W. S. (1996). Evidence based medicine: What it is and what it isn't. British Medical Journal, 312, 71-72.
Schoenfeld, A. H. (2006). What doesn't work: The challenge and failure of the What Works Clearinghouse to conduct meaningful reviews of studies of mathemat- ics curricula. Educational Researcher, 35(2), 13-21.
Seethaler, P M., & Fuchs L. S. (2005). A drop in the bucket: Randomized controlled trials testing reading and math interventions. Learning Disabilities Research and Practice, 20(2), 98-102.
Simmerman, S., & Swanson, H. L. (2001). Treatment outcomes for students with learning disabilities: How important are internal and external validity? Journal of Learning Disabilities, 34, 221-236. Slavin, R. E. (2008). Evidence-based reform in educa- tion: Which evidence counts? Educational Researcher, 37, 47-50.
Stoiber, K. C. (2002). Revisiting efforts on construct- ing a knowledge base of evidence-based intervention within school psychology. School Psychology Quarterly, 17, 533-546.
Tankersley, M., Cook, B. G., & Cook, L. (2008). A preliminary examination to identify the presence of quality indicators in single-subject research. Education and Treatment of Children, 31(4), 523-548.
Task Force on Evidence-Based Interventions in School Psychology. (2003). Procedural and coding manual for review of evidence-based interventions. Division 16 of the American Psychological Association. Retrieved from www.indiana.edu/'-ebi/documents/_workingfiles/ EBImanuall.pdf
Task Force on Promotion and Dissemination of Psy- chological Procedures. (1995). Training in and dissemi- nation of empirically-validated psychological treatments. The Clinical Psychologist, 48, 3—23.
Spring 2009
Test, D. W., Fowlet, C. H., Btewet, D. M., & Wood, W. M. (2005). A content and methodological review of self-advocacy intervention studies. Exceptional Children, 72, 101-125.
Thompson, B., Diamond, K. E., McWilliam, R., Sny- der, P., & Snyder, S. W. (2005). Evaluating the quality of evidence from correlational research for evidence- based practice. Exceptional Children, 71, 181-194. Viadero, D., & Huff, D. J. (2006). "One stop" research shop seen as slow to yield views that educators can use. Education Week, 26(5), 8-9.
Waehler, C. A., Kalodner, C. R., Wampold, B. E., & Lichtenberg, J. W. (2000). Empirically supported treat- ments (ESTs) in perspective: Implications for counsel- ing psychology training. Counseling Psychologist, 28, 657-671.
Wampold, B. E. (2002). An examination of the bases of evidence-based interventions. School Psychology Quarterly, 17, 500-507.
Wanzek, J., & Vaughn, S. (2006). Bridging the re- search-to-practice gap: Maintaining the consistent im- plementation of research-based practices. In B. G. Cook & B. R. Schirmer (Eds.), What is special about special education: The role of evidence-based practices (pp. 165-174). Austin, TX: PRO-ED.
What Works Clearinghouse. (2008). What Works Clear- inghouse evidence standards for reviewing studies. Re- trieved September 23, 2008, from http://ies.ed. gov/ncee/wwc/pdf/study_standards_fmal.pdf What Works Clearinghouse, (n.d.a). Welcome to WWC. Retrieved September 23, 2008, from http://ies.ed. gov/ncee/wwc/
What Works Clearinghouse, (n.d.b). What Works Clear- inghouse intervention rating scheme. Retrieved Septem- ber 23, 2008, from http://ies.ed.gov/ncee/wwc/pdf/ rating_scheme.pdf.
A B O U T T H E A U T H O R S
BRYAN G. COOK (CEC HI Federation), Profes- sor, Department of Special Education, University of Hawaii, Honolulu, MELODY T A N K E R S L E Y (CEC OH Federation), Professor, Department of Special Education, Kent State University, Kent, Ohio. TIMOTHY J. LANDRUM (CEC VA Fed- eration), Senior Scientist, Department of Cur- riculum, Instruction, and Special Education, University of Virginia, Charlottesville.
Address correspondence to Bryan G. Cook, Uni- versity of Hawaii at Manoa, College of Educa- tion, Department of Special Education, 1776 University Ave., Wist Hall 117, Honolulu, HI 96822 (e-mail: [email protected]).
The authors thank the Division for Research of the Council for Exceptional Children for their support of this work and for its leadership in identifying and applying evidence-based practices in special education.
Manuscript received June 2008; accepted Septem- ber 2008.
Exceptional Children 3 8 3