Literature review on behavior analysis
Effect Size for Token Economy Use in Contemporary Classroom Settings: A Meta-Analysis of
Single-Case Research
Denise A. Soares University of Mississippi
Judith R. Harrison Rutgers University
Kimberly J. Vannest Texas A&M University
Susan S. McClelland University of Mississippi
Abstract. Recent meta-analyses of the effectiveness of token economies (TEs) report insufficient quality in the research or mixed effects in the results. This study examines the contemporary (post-Public Law 94-142) peer-reviewed published single-case research evaluating the effectiveness of TEs. The results are stratified across quality of demonstrated functional relationship using a nonparametric effect size (ES) that controls for undesirable baseline trends in the analysis. In addition, moderators (i.e., classroom setting, age of participant, outcomes, use of response cost, and use of verbal cueing) were analyzed. Eighty-eight AB phase contrasts were calculated from 28 studies (1980 –2014) representing 90 partici- pants and produced a weighted mean ES of 0.82 (SE � 0.03, 95% CI [0.77, 0.88]). Strong quality produced a combined weighted mean ES of 0.85 (SE � 0.642, 95% CI [0.74, 0.97]). Moderator analyses revealed that a TE was slightly more effective for youth between the ages of 6 and 15 years than for children between the ages of 3 and 5 years or when used with behavioral goals in comparison to academic goals. However, no difference was found when implemented in general or special education settings or with the inclusion of response cost or verbal cueing.
A token economy (TE) is one of a hand- ful of interventions found in classroom set- tings. Based on the well-established principles
of reinforcement described by Skinner (1931), a TE is a secondary reinforcement system (Alberto & Troutman, 2003), whereby inher-
Correspondence concerning this paper should be addressed to Denise A. Soares, University of Mississippi, P.O. Box 1848, 49 Guyton Drive University, MS 38677; e-mail: [email protected]
Copyright 2016 by the National Association of School Psychologists, ISSN 0279-6015, eISSN 2372-966x
School Psychology Review, 2016, Volume 45, No. 4, pp. 379 –399
379
ently neutral items (i.e., tokens) are awarded for the demonstration of targeted behaviors. Tokens are accumulated and exchanged for backup reinforcers valued by the student (Ka- zdin, 1971; Simonsen, Fairbanks, Briesch, Myers, & Sugai, 2008). A TE has historically been considered a best-practices behavior management strategy for use in schools (Fil- check, McNeil, Greco, & Bernard, 2004; Mat- son & Boisjoli, 2009) and is one intervention frequently implemented within the positive behavioral interventions and supports frame- work. However, the emphasis on meta-ana- lytic thinking (see Maggin, Chafouleas, God- dard, & Johnson, 2011) that evolved after the passage of the Individuals with Disabilities Education Improvement Act (2004) and No Child Left Behind Act (2001) has provoked questions regarding its effectiveness (Maggin et al., 2011). In the following sections, we describe the historical and current research on the use of TEs and gaps in the literature that are addressed by this meta-analysis.
TOKEN ECONOMY RESEARCH
Numerous individual studies have demonstrated successful application of TEs across populations and settings. Specifi- cally, TEs have produced positive effects for students with emotional and behavioral problems (Cavalier, Ferretti, & Hodges, 1997), intellectual disabilities (Millersmith, Weber, & McLaughlin, 2013), attention def- icit hyperactivity disorder (DuPaul & Wey- andt, 2006), learning disabilities (Higgins, Williams, & McLaughlin, 2001), and schizo- phrenia (Ulmer, 1976). The use of TEs has been effective not only in schools (Filcheck et al., 2004) but also in residential treatment cen- ters (Murray & Sefchik, 1992), mental health hospitals (Hopko, Lejuez, Lepage, Hopko, & McNeil, 2003), prisons or detention centers (Bippes, McLaughlin, & Williams, 1986), and colleges (Stilitz, 2009).
A TE has been deemed an effective in- tervention by two seminal reviews (Kazdin & Bootzin, 1972; Kazdin, 1982), two systematic reviews (Dickerson, Tenhual, & Green-Paden, 2005; Matson & Boisjoli, 2009), and one
meta-analysis (Maggin et al., 2011). Kazdin and Bootzin (1972) and Kazdin (1982) evalu- ated benefits of using a TE, such as immediate reinforcement of behavior to maintain perfor- mance across time. They also identified obsta- cles to its effective implementation, such as inadequate staff training, client resistance, cir- cumvention of contingencies, and lack of re- sponse. In 1982, Kazdin updated the original review, evaluating the progress in the field since 1972. The authors found research had uncovered solutions to previously identified obstacles such as individualizing tokens and backup reinforcers (i.e., frequency, value) to enhance responsiveness, revising methods of staff training, decreasing resistance, and emphasizing the need to maintain effects across time. However, the authors did not complete systematic literature reviews, as their goal was to identify obstacles and meth- ods of overcoming them and not to synthesize the literature. Thus, these two articles do not identify or appraise the state of the literature.
Two systematic literature reviews (Dickerson et al., 2005; Matson & Boisjoli, 2009) synthesized and reported information from source articles. Dickerson et al. (2005) evaluated the use of a TE to improve socially appropriate behaviors of individuals with mental health disorders in hospital settings. They reviewed 13 studies (group and single- case experimental design [SCED]) with 1,074 participants ranging from 18 to 55 years old; 29% were diagnosed with schizophrenia, 13% with psychotic disorder, and 57% with other mental illnesses. Results indicated that a TE was effective for increasing adaptive behav- iors such as work performance, social interac- tion, and the daily care skills of these patients. Similarly, Matson and Boisjoli (2009) evalu- ated the effects of a TE on the behaviors of individuals with autism and/or development disabilities. They reviewed 16 group and SCED studies conducted in multiple settings, such as schools, homes, summer camps, group homes, state hospitals, and a developmental center. The 164 participants ranged in age from 4 to 18 years; approximately 91% were children with intellectual disabilities and 8% were children with autism. Results indicated
School Psychology Review, 2016, Volume 45, No. 4
380
that a TE was associated with positive out- comes in social, behavioral, and academic ar- eas. Although these studies systematically re- viewed the literature, neither quantified the effect associated with the use of a TE in schools through meta-analytic procedures.
Maggin et al. (2011) raised questions regarding a TE as an evidence-based interven- tion in the only meta-analysis to date. They conducted a meta-analysis of SCED studies to extend findings from earlier reviews. Maggin et al. targeted behavioral outcomes and coded the studies utilizing the Protocol for Assessing Single-Subject Research Quality (PASS-RQ; Maggin & Chafouleas, 2010) developed by the authors, with indicators from the guide- lines of Horner et al. (2005) for quality re- search and the What Works Clearinghouse (WWC) standards (Kratochwill et al., 2010). In addition, Maggin et al. calculated four dif- ferent meta-analytic effect sizes (ESs): percent of nonoverlapping data (PND; Scruggs & Mastropieri, 2001) � 78.49%; improvement rate difference (IRD; Parker, Vannest, & Brown, 2009) � 51.47%; standardized mean difference (SMD; Busk & Serlin, 1992) � 8.02; and raw-data multilevel ES (RMD; Van den Noortgate & Onghena, 2003, 2008) � 8.74. Maggin et al. stated, “Three of the four effect size (ES) measures found a significant improve- ment” (p. 550). The authors reported that the three significant ESs were PND, SMD, and RMD (D. Maggin, personal communication, March 2, 2016). However, the authors con- cluded design quality did not meet WWC stan- dards for 70% of the included studies because of fewer than three opportunities to demonstrate an effect (n � 7) or fewer than three data points per phase (n � 10). Weaknesses were found in the description of measurement procedures used; the number of data points per phase; and the report- ing of treatment fidelity, interobserver agree- ment (IOA), and social validity. The message from these findings is that rigorous research with sufficient methodological quality is needed to support a TE as an evidence-based intervention. Additional information can be contributed to the findings of this meta-analysis through the inclu- sion of studies that evaluate both academic and
behavioral outcomes and stratification of results across design quality.
Building on earlier studies, these re- searchers provide valuable information, which represents the historical research in TE litera- ture. The studies (i.e., Kazdin, 1982; Kazdin & Bootzin, 1972) were characterized as literature reviews emphasizing many strengths of TEs and identifying obstacles; however, the au- thors provided summary findings and not spe- cific data from the source articles. The more recent reviews (Dickerson et al., 2005; Matson & Boisjoli, 2009) indicated a TE was associated with an increased effect on social, behavioral, and academic outcomes; however, neither re- view reported an ES, confidence intervals (CIs), or design quality. In addition, Dickerson et al. (2005) searched only one database, the National Library of Medicine’s PubMed, and Matson and Boisjoli (2009) reviewed “representative litera- ture.” Finally, Maggin et al. (2011) found large effects of a TE on behavioral outcomes; how- ever, the researchers contended that a TE could not be considered an evidence-based interven- tion because of the quality of the research. Ad- ditional research is needed so that a TE can potentially be considered an evidence-based strategy.
SCED AND QUALITY
SCED has a long history of use in the applied fields of education and human behav- ior and is particularly suited to school-based practices allowing the single subject to serve as his or her own control (Horner et al., 2005). As such, SCED is especially relevant to re- views of school-based use of TEs. In recent years, experts have developed quality indica- tors and standards for individual SCED studies that can be utilized to synthesize the method- ological rigor of a group of studies (Horner et al., 2005; Kratochwill et al., 2010). Horner et al. (2005) identified the criteria necessary to evaluate the quality of (a) research reporting (e.g., description of participants and settings, social validity, research questions) and (b) de- sign (e.g., dependent variable, independent variable, baseline, experimental control or in- ternal validity, external validity). In addition,
Main and Moderator Effects for Token Economies
381
Kratochwill et al. (2010) outlined specific cri- teria for a study to meet standards or meet standards with reservations based on (a) the number of phases per design, (b) the number of data points per phase, and (c) the percent of IOA that must be measured. On the basis of recommendations from experts in the field (i.e., Maggin, Briesch, & Chafouleas, 2013), combining conceptually relevant constructs, such as operational definitions of participants and settings (see Horner et al., 2005) and sufficient evidence to support a functional re- lationship between the dependent and inde- pendent variable (Kratochwill et al., 2010), results in a rigorous evaluation of SCED qual- ity. Thus, design quality can be evaluated based on these criteria, and outcomes can be stratified by quality.
EFFECT SIZES
With the emphasis on evidence-based interventions, the need for quantifying inter- vention effects achieved through SCED stud- ies has come to the forefront. In 2007, Parker and Hagan-Burke found over 40 ESs; how- ever, the field has not come to consensus for which is best (Manolov, Solanas, Sierra, & Evans, 2011). Tau-U (an ES from the fre- quently used Kendall’s � and Mann-Whitney U) is summarized by Parker, Vannest, Davis, and Sauber (2011) as “having statistical power that is flexible and can calculate trend only, nonoverlap between phases only, or a combi- nation of the two” (p. 291). Tau-U is a con- servative measure that offers important bene- fits of a “bottom-up” approach (Parker & Vannest, 2012), designed to explain the im- pact of changes at the individual phase con- trast on the overall effect. Benefits of Tau-U’s nonparametric bottom-up approach include (a) consistency with visual analysis; (b) applica- bility to short data series and simple designs; (c) appropriateness with any design; (d) char- acterization by strong statistical power (i.e., one of the strongest parametric tests; Parker et al., 2011, p. 288); (e) control in Phase A trend; and (f) usefulness at three levels—nonaggre- gated data from a single client, aggregated data from a complex design, and meta-analy-
ses (Parker et al., 2011). Tau-U allows for the calculation of CIs and p values. All data are used, reflecting the interventionist experimen- tal perspective that each data point reflects performance.
MODERATORS
Although a TE has been found to be effective, a comparison has never been made to determine in which environment, with which participants, and with which outcome measures (i.e., academic and/or behavioral) a TE is most successful. Moderator variables can account for variations across studies (e.g., characteristics of setting, participant, outcome, or implementation). Identifying the impact of these variables has the potential to increase the effectiveness and efficiency of a TE for edu- cators. Although numerous moderators could be hypothesized, to avoid Type I error (Fairch- ild & MacKinnon, 2009) in analyses (i.e., finding a moderator effect when none exists) and minimize the chances of inaccurate re- sults, we identified five potential moderators for which we have strong hypotheses: (a) classroom setting; (b) age of participant; (c) type of outcome, academic or behavioral; (d) use of response cost (RC); and (e) use of verbal cueing. Next, we describe our hypoth- eses and rationale for those hypotheses.
Some authors contend that a TE is most effective in settings with small teacher-to-stu- dent ratios, such as self-contained special ed- ucation classrooms (Center & Wascom, 1984; Kazdin & Geesey, 1980), with younger stu- dents (Filcheck et al., 2004), and with only behavioral outcomes (Himle, Woods, & Bu- naciu, 2008; Jones, Weber, & McLaughlin, 2013). These findings have the potential to limit its use in different types of classroom settings, with certain students, and to address specific behaviors. However, we hypothesize that setting does not change the effectiveness of a TE, as individual studies seem to indicate otherwise. A TE has been found to be effec- tive (a) in multiple instructional settings such as inclusive, general education, special educa- tion, and alternative settings (De Martini- Scully, Bray, & Kehle, 2000; Rhode, Jenson,
School Psychology Review, 2016, Volume 45, No. 4
382
& Reavis, 1993); (b) with differing age groups such as elementary-age students (Akin-Little & Little, 2004; Christensen, Young, & March- ant, 2004; Filcheck et al., 2004), junior high students (Carlson, Pelham, Milich, & Dixon, 1992, Cavalier et al., 1997; Feindler, Marriott, & Iwata, 1984; Heaton & Safer, 1982), and high school students (Schellenberg, Skok, & McLaughlin, 1991); and (c) to increase aca- demic (Klimas & McLaughlin, 2007; Salend, Tintle, & Balber, 1988; Sran & Borrero, 2010) as well as behavioral outcomes (Center & Wascom, 1984; De Martini-Scully et al., 2000). Thus, on the basis of these studies, we hypothesize that these potential variables do not moderate the effects of a TE.
Similarly, some authors contend that a TE can be a very complex intervention, which leads to lack of adoption in schools (Milten- berger, 2001; Rosen, Taylor, O’Leary, & Sanderson, 1990; Skinner, Cashwell, & Bunn, 1996). Adding procedures, such as RC or ver- bal cueing, increases the complexity of the intervention. RC is a procedure designed to decrease behavior by contingently withdraw- ing a specific amount of reinforcement follow- ing an inappropriate behavior or response (Ka- zdin, 1972). The impact of RC on a TE’s effectiveness is not clear. Some previous re- search has supported RC procedures in TE systems (Gresham, 1979; Rapport, Murphy, & Bailey, 1980; Witt & Elliot, 1982). Other stud- ies (e.g., Phillips, Phillips, Fixsen, & Wolf, 1971) found RC has harmful side effects, such as the opportunity for the implementer to over- penalize and the possibility of decreasing the incentive of demonstrating the target behavior. Verbal cueing (i.e., prompting) is another pro- cedure frequently added that increases the complexity of a TE. Similarly, some research findings regarding the effect of adding verbal cueing to a TE have suggested the procedure might increase effectiveness (Latham & Locke, 1991) and others suggested it might not (Balcazar, Hopkins, & Suarez, 1985; Kluger & DeNisi, 1996). We hypothesized that RC and verbal cueing are not necessary for a TE to be effective. This hypothesis is founded on some research suggesting that nei- ther RC nor verbal cueing is necessary in
hopes of reducing complexity of implementa- tion that might decrease adoption and use.
Although we have strong hypotheses re- garding the moderating nature of the previ- ously discussed variables, no meta-analysis of TEs has reported moderator analysis results. Maggin et al. (2011) conducted moderator analyses with participant characteristics and intervention features. However, they did not report results because of “the likely presence of family-wise error in these findings” (p. 547). Thus, we conducted a moderator anal- ysis for which we had strong a priori hy- potheses to decrease the risk of Type I error (McKillup, 2011).
PURPOSE
The current study addresses limitations to the prior literature and provides additional information to the field by (a) evaluating re- search design quality, (b) calculating ESs and CIs, (c) stratifying results by design quality, and (d) evaluating moderator analyses of peer- reviewed literature through 2014, with a focus on both academic and behavioral outcomes. Thus, this meta-analysis addressed the follow- ing research questions:
1. What is the research design quality of studies of the effectiveness of TEs across SCED studies?
2. What is the overall effect of TEs in public school classrooms?
3. Do the effects of TEs differ by design quality?
4. What are the effects of potential moderators?
METHOD
We conducted this study in four phases and organized the methods accordingly. Meth- ods are adapted from Bowman-Perrott, Burke, Zhang, and Zaini (2014). Details are included below for literature review and study selection, data extraction, ES and CI calculations, stratifi- cation across methodological quality, and mod- erator analyses. Procedures for each phase are described in detail.
Main and Moderator Effects for Token Economies
383
Phase 1: Literature Review and Study Selection
We applied standard methods identified by Cooper and Hedges (1994) to search the EBSCO Research Complete, Education Full Text, PsycINFO, and Education Resource In- formation Center (ERIC) electronic databases. Key words, Boolean strings, and truncated words used to conduct the search included “token economy,” “intervention,” “reinforce- ment,” “contingency management,” “system- atic positive reinforcement,” “tokens,” “oper- ant conditioning,” “applied behavior analysis,” “backup reinforcers,” “behavior therapy,” “points,” and/or “response cost.” In addition to the search of databases, we conducted a hand search for titles related to TEs or secondary reinforcement by reviewing the tables of con- tents for the years 1980 to 2014 in 21 journals on special education, school psychology, and be- havioral psychology (e.g., Journal of Positive Behavior Interventions, Behavior Therapy, Be- havioral Interventions, and The Journal of Spe- cial Education). We selected the year 1980 as a delimiter based on a desire to use classroom settings that were more likely to be inclusive of students with and without disabilities than class- rooms before or immediately after the passage of Public Law 94-142 in 1975. We conducted his- torical searches with all of the resulting screened articles, and the process was repeated to include studies with potential to meet the inclusion cri- teria. The initial search conducted by the first author yielded 1,833 results. Article titles and abstracts were screened based on inclusion cri- teria (described in the following subsection) re- sulting in the elimination of 1,436 studies. We eliminated 278 of the remaining 397 articles based on information in the article indicating the study was not conducted in a school setting, was exclusively a descriptive study, or was without peer review. Fifty-five articles remained. During the gathering of articles, reliability checks were conducted and assessed using simple percent of agreement (Sum of agreement/Total number of agreements � disagreements � 100; House, House, & Campbell, 1981). Initial agreement for article inclusion was 100%.
Inclusion Criteria We examined the full text of 55 studies
for potential inclusion in this meta-analysis. A TE was operationally defined as a program in which students earned tokens for identified academic skills (e.g., task engagement, accu- racy, and completion) or behaviors (e.g., dis- ruptive, out of seat) and then exchanged the earned tokens for backup reinforcers (Alberto & Troutman, 2003; Martin & Pear, 2003). The first author and a doctoral student indepen- dently coded the 55 articles for inclusion cri- teria in separate spreadsheets. We included studies if they (a) were published between the years 1980 and 2014, (b) occurred in U.S. public school classroom settings, (c) included school-age children (i.e., 3 to 21 years old), (d) were published in peer-reviewed journals, and (e) included SCED with published data in readable graphs. We elected to only include studies published in peer-reviewed journals as studies were reviewed and filtered for quality to maintain standards, scientific merit, and va- lidity (Voight & Hoogenboom, 2012). We as- sessed for publication bias statistically (see Publication Bias subsection).
Exclusion Criteria We excluded 27 studies: 6 addressed
multicomponent interventions for which the intervention data could not be disaggregated, 5 did not include a visual graph of data from which raw data could be digitized, 4 were not intervention studies, 4 were graduate theses, 3 included participants other than school-age children, 2 were set in international class- rooms, 1 was set in a classroom within a residential treatment center, 1 focused on the implementation of the TE by paraprofession- als and did not include intervention data for the children in the study, and 1 included the intervention in the baseline data. After exclu- sion of these studies, the literature search re- sulted in 28 SCED studies in which a TE was the intervention in a classroom setting with school-age children.
Publication Bias We examined the final set of selected
studies for publication bias (i.e., tendency of
School Psychology Review, 2016, Volume 45, No. 4
384
studies with null effects to not be published; Rosenthal & DiMatteo, 2001) using the Egg- er’s test (Egger, Smith, Schneider, & Minder, 1997). The Egger’s test evaluates Y inter- cept � 0 using linear regression of the effect against precision. The intercept for the Egg- er’s test (1.04; 90% CI [– 0.03, 2.11]; p � .52) indicated that statistically significant publica- tion bias was not found in this sample. Heter- ogeneity was measured using H and I2 (Hig- gins & Thompson, 2002), where H � 1.5 (95% CI [1.2, 1.8]) and I2 � 54.0% (95% CI [29.5, 70.0]), indicating less than notable het- erogeneity in the sample with H values above 1.5 considered not notable (Abramson, 2011).
Coding We coded for four purposes: (a) inclu-
sion and exclusion in the review, (b) dem- onstration of a functional relationship between the independent and dependent variables according to Horner et al. (2005) and the WWC standards for SCED (design quality; Kratochwill et al., 2010; see Table 1 for specific criteria), (c) moderator vari- ables, and (d) descriptive reporting of pro- cedural integrity or fidelity. Specific codes are listed in Tables 1 and 2. Each of the coded variables provided the basic data for the reliability analyses.
We assessed reliability of data coding by IOA checks between the two independent raters (the first author and a doctoral student). Coding-sheet training and trial coding were performed for agreement between the two rat- ers before reliability was calculated. Within a discussion format, the two raters identified one example and one nonexample of each coding variable. If difficulties in this task arose be- cause of lack of code clarity, the pair deliber- ated until clarity was achieved with 100% agreement. Official coding began when a min- imum acceptable value of IOA (�80%) was met (Hartmann, Barrios, & Wood, 2004). Each rater coded every variable in all articles independently as is commonly done in com- puting intercoder reliability in meta-analyses (Yeaton & Wortman, 1993).
We calculated Cohen’s � reliability agreement for coding using NCSS (Hintze, 2004) by entering the agreement– disagree- ment matrix for analysis. Kappa is a conser- vative measure of reliability and perhaps even underestimates agreement (Ary & Suen, 1989; Strijbos, Martens, Prins, & Jochems, 2006). Kappa was 0.96.
Phase 2: Design Quality Evaluation
We evaluated design quality based on published guidelines (Horner et al., 2005; Kratochwill et al., 2010) with a goal of strat- ifying results across study quality. Two raters (first and second authors) independently as- signed ratings of weak, medium, and strong to each included study. Weak ratings are equiv- alent to the “does not meet standards” cate- gory and include fewer than three opportuni- ties for demonstration of an effect and fewer than three data points per phase. Medium rat- ings are equivalent to the WWC “meets stan- dards with reservations” and include at least three opportunities for demonstration of an effect with at least three data points per phase for reversal or multiple-baseline designs and four for multielement designs and reporting of IOA. Strong ratings are equivalent to the WWC design standards that “meet criteria” and include at least three opportunities for demonstration of an effect with five or more data points per phase and reporting of IOA. Studies without three demonstrations of effect do not meet criteria (e.g., ABCD, ABBCC) as they do not provide three oppor- tunities for demonstration of an effect. Simul- taneous, multiple-probe, alternating-treatment (e.g., ABCBD), and multielement designs were coded as their underlying design (e.g., ABA, ABC). Horner et al. (2005) described reporting the level of treatment integrity as “highly desirable” (p. 174), but it was not a tenet described by Kratochwill et al. (2010) to meet the standards for design quality for SCED. Thus, we did not include it in our ratings of design quality, but we did report the number of studies that stated the percent of treatment integrity.
Main and Moderator Effects for Token Economies
385
T ab
le 1 .
T au
-U E
ff ec
t S iz
es fo
r E
ac h
S tu
d y
b y
Q u
al it
y an
d A
u th
o r
D es
ig n
Q ua
li ty
a A
ut ho
rs an
d Y
ea r
O ut
co m
e
D es
ig n
P ha
se A
B C
on tr
as ts
, n
P ar
ti ci
pa nt
s, n
T au
-U 95
% C
I A
ca d
or B
eh av
T yp
e
W ea
k F
il ch
ec k
et al
., 20
04 B
eh av
D is
ru pt
iv e
A B
A C
C M
1 17
0. 67
�0 .6
0, 1.
00
W ea
k H
im le
et al
., 20
08 B
eh av
D is
ru pt
iv e
M ul
ti el
em en
t N
4 4
0. 65
�0 .3
4, 1.
00
W ea
k K
az di
n &
G ee
se y,
19 80
A ca
d T
as k
en ga
ge m
en t
S im
ul ta
ne ou
s N
2 2
1. 00
�0 .7
8, 1.
00
W ea
k K
az di
n &
M as
ci te
ll i,
19 80
A ca
d T
as k
en ga
ge m
en t
S im
ul ta
ne ou
s N
2 2
0. 99
�0 .7
3, 1.
00
W ea
k K
li m
as &
M cL
au gh
li n,
20 07
A ca
d A
ss ig
nm en
t co
m pl
et io
n A
B C
M 3
1 1.
00 �0
.7 6,
1. 00
W
ea k
M il
le rs
m it
h et
al .,
20 13
A ca
d A
ss ig
nm en
t ac
cu ra
cy R
ev er
sa l
N 2
1 .9
7 �0
.3 6,
1. 00
W
ea k
R os
en be
rg ,
19 86
B eh
av D
is ru
pt iv
e A
B C
N 5
5 0.
99 �0
.6 9,
1. 00
W
ea k
S al
en d
& A
ll en
, 19
85 B
eh av
D is
ru pt
iv e
A B
C B
C N
2 2
1. 00
�0 .6
2, 1.
00
W ea
k S
ra n
& B
or re
ro ,
20 10
A ca
d T
as k
ac cu
ra cy
A B
C D
N 12
4 0.
40 �0
.2 2,
0. 55
W
ea k
S te
ve ns
, S
id en
er ,
R ee
ve ,
& S
id en
er ,
20 11
A ca
d T
as k
ac cu
ra cy
M ul
ti pr
ob e
de si
gn M
4 2
1. 00
�0 .6
0, 1.
00
W ea
k S
ul li
va n
& O
’L ea
ry ,
19 90
B eh
av D
is ru
pt iv
e A
B B
C C
N 2
1 1.
00 �0
.5 7,
1. 00
M
ed iu
m C
ar ne
tt et
al .,
20 14
B eh
av D
is ru
pt iv
e A
lt er
na ti
ng -t
re at
m en
t de
si gn
G 2
1 1.
0 �0
.4 5,
1. 00
M ed
iu m
D e
M ar
ti ni
-S cu
ll y
et al
., 20
00 B
eh av
D is
ru pt
iv e
M B
D N
2 2
0. 95
�0 .6
3, 1.
00
M ed
iu m
H ig
gi ns
et al
., 20
01 B
eh av
D is
ru pt
iv e
M B
D M
3 1
0. 98
�0 .6
3, 1.
00
M ed
iu m
Jo ne
s et
al .,
20 13
B eh
av D
is ru
pt iv
e A
B A
B N
2 2
.7 2
�0 .2
2, 1.
00
M ed
iu m
M cG
oe y
& D
uP au
l, 20
00 B
eh av
D is
ru pt
iv e
R ev
er sa
l M
4 4
0. 75
�0 .4
2, 1.
00
M ed
iu m
S im
on ,
A yl
lo n,
& M
il an
, 19
82 B
eh av
D is
ru pt
iv e
A B
C B
N 3
1 0.
78 �0
.3 8,
1. 00
M
ed iu
m S
m it
h &
F ow
le r,
19 84
B eh
av D
is ru
pt iv
e M
B D
N 6
6 0.
92 �0
.6 5,
1. 00
M
ed iu
m T
ho m
ps on
, M
cL au
gh li
n, &
D er
by ,
20 11
B eh
av D
is ru
pt iv
e M
B D
N 3
1 .9
5 �0
.5 8,
1. 00
M ed
iu m
T ru
ch li
ck a,
M cL
au gh
li n,
& S
w ai
n, 19
98 A
ca d
T as
k ac
cu ra
cy M
B D
M 3
3 0.
35 �0
.1 2,
0. 58
S tr
on g
C en
te r
& W
as co
m ,
19 84
A ca
d T
as k
ac cu
ra cy
R ev
er sa
l N
5 5
0. 71
�0 .4
8, 0.
94
S tr
on g
C on
ye rs
et al
., 20
04 B
eh av
D is
ru pt
iv e
A lt
er na
ti ng
-t re
at m
en t
de si
gn N
2 2
0. 83
�0 .5
2, 1.
00
S tr
on g
M ag
li o
& M
cL au
gh li
n, 19
81 B
eh av
D is
ru pt
iv e
R ev
er sa
l M
1 1
1. 00
�0 .7
9, 1.
00
(T ab
le 1
co nt
in ue
s)
School Psychology Review, 2016, Volume 45, No. 4
386
T ab
le 1 .
C o n
ti n
u ed
D es
ig n
Q ua
li ty
a A
ut ho
rs an
d Y
ea r
O ut
co m
e
D es
ig n
P ha
se A
B C
on tr
as ts
, n
P ar
ti ci
pa nt
s, n
T au
-U 95
% C
I A
ca d
or B
eh av
T yp
e
S tr
on g
M ot
tr am
, B
ra y,
K eh
le ,
B ro
ud y,
& Je
ns on
, 20
02 B
eh av
D is
ru pt
iv e
M B
D M
3 3
0. 99
�0 .8
3, 1.
00
S tr
on g
M us
se r,
B ra
y, K
eh le
, &
Je ns
on ,
20 01
B eh
av D
is ru
pt iv
e M
B D
M 3
3 0.
88 �0
.5 3,
1. 00
S tr
on g
R ei
tm an
et al
., 20
04 B
eh av
D is
ru pt
iv e
A lt
er na
ti ng
-t re
at m
en t
de si
gn N
3 3
0. 69
�0 .4
0, 0.
95
S tr
on g
S al
en d
& L
am b,
19 86
B eh
av D
is ru
pt iv
e R
ev er
sa l
M 2
9 1.
00 �0
.5 3,
1. 00
S
tr on
g S
al en
d, T
in tl
e, &
B al
be r,
19 88
A ca
d T
as k
en ga
ge m
en t
R ev
er sa
l N
2 2
1. 00
�0 .5
3, 1.
00
O ve
ra ll
88 90
0. 82
�0 .7
7, 0.
88
N o te
. A
B co
nt ra
st s
w er
e us
ed in
ef fe
ct si
ze ca
lc ul
at io
ns .
T he
po ss
ib le
ra ng
e fo
r C
I is
0 to
1. A
� ba
se li
ne ;
A ca
d �
ac ad
em ic
; B
� fi
rs t
in te
rv en
ti on
; C
� se
co nd
in te
rv en
ti on
; D
� th
ir d
in te
rv en
ti on
; B
eh av
� be
ha vi
or al
; C
I �
co nfi
de nc
e in
te rv
al ;
G �
au th
or s
in cl
ud ed
ge ne
ra li
za ti
on ph
as e,
M �
au th
or s
in cl
ud ed
m ai
nt en
an ce
ph as
e; M
B D
� m
ul ti
pl e-
ba se
li ne
de si
gn ;
N �
au th
or s
in cl
ud ed
ne it
he r
ge ne
ra li
za ti
on ph
as e
no r
m ai
nt en
an ce
ph as
e. a S tr
o n g
in di
ca te
s at
le as
t th
re e
de m
on st
ra ti
on s
of an
ef fe
ct w
it h
fi ve
or m
or e
da ta
po in
ts pe
r ph
as e
an d
re po
rt in
g of
in te
ro bs
er ve
r ag
re em
en t;
m ed
iu m
, at
le as
t th
re e
de m
on st
ra ti
on s
of ef
fe ct
w it
h th
re e
to fo
ur da
ta po
in ts
pe r
ph as
e an
d re
po rt
in g
of in
te ro
bs er
ve r
ag re
em en
t; an
d w
ea k,
fe w
er th
an th
re e
de m
on st
ra ti
on s
of an
ef fe
ct an
d th
re e
to fo
ur da
ta po
in ts
pe r
ph as
e.
Main and Moderator Effects for Token Economies
387
Phase 3: ES, CIs, and Stratification Calculation
To calculate the Tau-U, we extracted raw data from the graphs and figures of the included studies. We used a computer scanner and software program to view and assign data values electronically. Digitizing the data resulted in an exact reconstruction of the original graphic data providing numeric raw data to enable proper comparisons (Glass, 1976). All SCED graphs of included articles were digitized with GetData Graph Digitizer (Version 2.21).
ES and Visual Analysis We calculated Tau-U for each individ-
ual contrast between the baseline (e.g., A1) and the adjacent intervention contrast (e.g., B1) for each unit of analysis (i.e., student, class, behavior). Tau-U is derived from Ken- dall’s � and Mann-Whitney U (see Parker et al., 2011) and is calculated by merging trend and nonoverlap data. We completed calcula- tions using the online Tau-U calculator (Van- nest, Parker, & Gonen, 2011), selecting the analysis to “control for baseline trend,” for individual ES (i.e., an ES for each participant, behavior, or setting). The individual ESs were entered into the statistical program WinPEPI
for analysis to produce the “combined” omni- bus ES and 95% CI (Abramson, 2011). The algorithm for WinPEPI to calculate the overall ES is the weighted average of all individual ESs, with weights equaling the inverse of the variance (i.e., not standard error). We coded the Tau-U effects as small (0 – 0.65), medium (0.66 – 0.92), and large (0.93–1.00), which are equivalent to ranges recommended for non- overlap of all pairs (Parker et al., 2009) to compare Tau-U to effects garnered through visual analysis, our next step.
We visually analyzed data from each study to determine whether a functional rela- tionship existed between the independent and dependent variables based on recommenda- tions from Kratochwill et al. (2010). Compar- ing the ES to visual analysis in SCED studies increases the credibility of the ES (Parker & Vannest, 2012). This is in part because (a) visual analysis is the traditionally accepted approach to SCED analysis (Kratochwill et al., 2010) and (b) the documentation of concurrent validity between a visual analysis and an ES is a component of a bottom-up approach (Ninci et al., 2015; Parker & Vannest, 2012). The first and second authors visually analyzed the data following procedures recommended by Kratochwill et al. (2010) with additional
Table 2. Summary of Moderators
Study Characteristic Category
Studies, n
AB Contrasts, na
Participants, n Tau-U SE 95% CI z scoreb p valueb
Setting Special 16 47 47 0.89 0.05 �0.80, 0.98 General 12 41 43 0.86 0.07 �0.74, 0.98 0.04 .70
Age 3–5 years 6 29 22 0.64 0.06 �0.53, 0.75 6–15 years 22 64 68 0.91 0.04 �0.84, 0.98 2.33 .02*
Outcome Academic 8 29 21 0.89 0.08 �0.73, 1.00 Behavioral 20 59 73 0.93 0.07 �0.75, 1.00 0.04 .69
Response cost Yes 16 47 47 0.84 0.06 �0.75, 0.95
No 12 41 43 0.91 0.05 �0.81, 1.00 0.99 .32 Verbal cue Yes 9 25 20 0.83 0.06 �0.74, 0.92
No 19 63 70 0.88 0.28 �0.33, 1.00 0.20 .84
Note. A � baseline; B � intervention; CI � confidence interval. aNumber of AB contrasts used in effect size calculation. bReliable difference z-test scores and corresponding p values are reported for the moderators. * p .05.
School Psychology Review, 2016, Volume 45, No. 4
388
guidance from Lane and Gast (2014). First, we analyzed within-phase data by visually in- specting the (a) level (i.e., mean, median); (b) trend (i.e., baseline trend, slope of the best fitting line); and (c) variability (i.e., bounce) of the data around the line within each phase. Second, we analyzed between-phase data by visually inspecting the (a) immediacy of the effect of a TE by observing the level change between the data at the end of a phase (last three data points) and the beginning of the next phase (first three data points), (b) fre- quency of overlap between two phases by determining the number of data points in one phase that overlapped with the adjacent phase, and (c) consistency of the data across similar phases.
From these analyses, cumulatively, we conceptualized the effect of a TE derived from each study as follows: (a) no effect, with no evidence of a functional relation between the independent and dependent variables; (b) weak effect, with some evidence of a func- tional relation with latency between the intro- duction of the independent variable, variabil- ity of data in the baseline and/or intervention phases, overlap between adjacent phases, and variability of data in similar phases; (c) me- dium effect, with mixed evidence of a func- tional relation with either latency between the introduction of the independent variable, vari- ability of data in the baseline and/or interven- tion phases, overlap between adjacent phases, or variability of data in similar phases; or (d) strong effect, with clear evidence of a func- tional relationship demonstrated by a mean consistent level in the baseline indicating the need for intervention, evidence of an immedi- ate effect between phases with a positive trend, minimal variability in all phases, no overlap between phases, and clear consistency between similar data phases.
We individually completed these analy- ses, and initial agreement between observers was 95%. When there was disagreement be- tween the two, discussion occurred until a consensus was reached with 100% agreement. We then compared the magnitude of effect from visual analysis (i.e., no, weak, medium,
or strong) with the magnitude of effect from Tau-U (i.e., small, medium, or large).
Phase 4: Moderator Analyses
We coded moderator data for setting (special education, general education), age of participant (preschool children ages 3 to 5 years and school-age children ages 6 to 15 years), outcome (academic, such as task accu- racy or task engagement, or behavioral, such as out of seat or disruptive), and two proce- dural differences (i.e., RC and verbal cueing). Following the procedures of Bowman-Perrott et al. (2014), we analyzed moderator effects by dichotomously coding the moderator vari- ables within the studies and examining statis- tically significant differences between the ES (Tau-U) of studies within each category.
We calculated a reliable difference (i.e., difference that cannot be accounted for solely by chance) for each moderator pair to deter- mine if the differences were statistically significant using the following formula: (L1 – L2)/�[(SETau1
2) � (SETau2 2)], where
L1 is the first level of the moderator (e.g., academic outcome) and L2 is the second level of the moderator (e.g., behavioral outcome). Specifically, we compared effects for a TE in general and special education settings, for children ages 3 to 5 years (preschool age) and 6 to 15 years (school age), for academic and behavioral outcomes, for use or nonuse of RC, and for use or nonuse of verbal cueing. Reliable-difference z-test scores and p values are reported in the Results section.
RESULTS
We conducted this study in four phases and organized the results accordingly. Results are included below for the literature review and study selection, design quality, effect size and stratification, and the moderator analysis.
Phase 1: Literature Review and Study Selection
Twenty-eight studies met the inclusion criteria and included 90 students and 88 opportunities for demonstrations of effect
Main and Moderator Effects for Token Economies
389
(A1B1). Seventy-nine percent of the studies were implemented prior to 2005 when the quality indicators were published, and of those studies, 50% were published between 1980 and 1989. Forty-three percent of the studies were implemented in general education class- rooms, and 57% were implemented in special education classrooms. Seventy-nine percent of the studies included children ages 6 to 15 years, and 21% included children ranging from 3 to 5 years old. Seventy-one percent of the studies used behavioral outcome measures (e.g., disruptive, talking out, out of seat), and 29% used outcome measures of academic be- haviors (e.g., task accuracy, completion, en- gagement). RC was used in 16 studies (57%). Verbal cueing was used in nine studies (31%). Six designs (series phase; multiple baseline; alternating intervention design; reversal de- sign; simultaneous; multiple probe) were uti- lized in the studies (see Table 1). Ten studies (36%) reported using a follow-up or mainte- nance phase, and one (3%) reported on gener- alizing a TE to another setting. Seventeen studies (61%) did not include maintenance to monitor the use of a TE on the outcome mea- sure (see Table 1).
Phase 2: Design Quality
Each of the 28 studies was visually an- alyzed for quality using the rubric of internal validity described in the Method section. Eleven studies were rated as weak quality, nine as medium quality, and eight as strong quality. Of the 20 weak- and medium-quality studies, 16 had insufficient demonstrations of effect, 8 had at least three to four data points per phase, and 2 did not report IOA. Integrity data were collected in 10 studies. Integrity ranged from 30% to 100% (M � 86.14%, SD � 23.36%).
Phase 3: ES, CIs, and Stratification
Tau-U was calculated for 88 baseline versus intervention contrasts (A1 versus B1) controlling for baseline trend for the 28 stud- ies. The weighted mean Tau-U of the TE was 0.82 (SE � 0.03; 95% CI [0.77, 0.88]) and ranged from 0.35 to 1.00. Figure 1 illus-
trates the range of ESs and 95% CIs for each study. Categorically, 19 of the studies had a Tau-U of 0.80 or above, 7 studies were be- tween 0.50 and 0.79, and 2 studies fell be- low 0.50 (i.e., 0.20 and 0.49).
ES Stratified by Design Quality Stratifying the ESs by quality level pro-
duced the following scores: Eleven studies in the weak-quality range had a combined Tau-U of 0.77 (SE � 0.05; 95% CI [0.67, 0.87]). The nine studies with a medium-quality rating had a combined Tau-U of 0.84 (SE � 0.04; 95% CI [0.76, 0.93]). The eight studies categorized as strong quality had a combined Tau-U of 0.85 (SE � 0.06; 95% CI [0.74, 0.97]). Medium- and strong-quality studies produced a large Tau-U of 0.84 (SE � 0.03; 95% CI [0.78, 0.91]).
Visual Analyses We visually analyzed 88 baseline and
intervention phases and contrasts. Results in- dicated baseline trend in 82% of baseline phases (n � 72), variability in 73% of baseline and intervention phases (n � 65), a strong immediate effect in 75% of phase changes (n � 66), a weak to medium immediate effect in 15% of phase changes (n � 13), and no immediate effect in 10% of phase changes (n � 9). In addition, 40% (n � 35) included overlapping data in adjacent phases.
From these visual analyses, we deter- mined that 58% of phase contrasts (n � 51) represented strong effects, 13% (n � 11) rep- resented medium effects, 22% (n � 19) rep- resented weak effects, and 8% (n � 7) repre- sented no effect. These results showed a 75% agreement (n � 66) with Tau-U. In compari- son to visual analysis, Tau-U slightly over- rated the effect on 14 occasions, in which the author team judged the data to represent (a) no effect compared to small effect for Tau-U on six occasions, (b) a weak effect compared to a medium Tau-U on three occasions, and (c) a medium effect compared to a large Tau-U on five occasions. In comparison to visual analy- sis, Tau-U slightly underrated the effect on five occasions, in which the author team judged the data to represent (a) a medium
School Psychology Review, 2016, Volume 45, No. 4
390
effect to a small Tau-U on two occasions and (b) a strong effect to a medium Tau-U on three occasions.
Phase 4: Moderator Analyses
The results of five potential moderator analyses are presented as follows. (Also see Table 2.)
Setting Twelve studies with 43 participants
and 41 phase contrasts were coded for general education and 16 studies with 47 participants and phase contrasts were coded for special
education. Results indicated that studies in general education settings had a lower ES (0.86; SE � 0.07; 95% CI [0.74, 0.98]) than special education settings (ES � 0.89; SE � 0.05; 95% CI [0.80, 0.98]). When the two cate- gories were compared for parameter estimates, overlapping CIs indicated there may not be a statistically significant difference, which was confirmed by the values from the reliable-differ- ence formula: z � 0.04, p � .70.
Age The preschool category of children
ages 3 to 5 years contained six studies, 22
Figure 1. Forest Plot of ESs for 28 Included Studies and Overall ES
Note. ES � effect size; Min CI � minimum confidence interval; Max CI � maximum confidence interval.
Main and Moderator Effects for Token Economies
391
participants, and 29 phase contrasts. The school-age category for ages 6 to 15 years contained 22 studies, 68 participants, and 64 phase contrasts. Results indicated that studies for ages 3 to 5 years had a lower ES (0.64; SE � 0.06; 95% CI [0.53, 0.75]) than for ages 6 to 15 years (ES � 0.91; SE � 0.04; 95% CI [0.84, 0.98]). When the two categories were compared for parameter estimates, non- overlapping CIs were observed indicating that the moderator might be statistically signifi- cantly different. The values from the reliable- difference formula were z � 2.33 and p � .02, confirming statistically significant results.
Outcome Measures (Academic, Behavioral) Eight studies with 21 participants and 29
phase contrasts were coded as academic out- comes (e.g., task accuracy, task engagement; see Table 1). Twenty studies with 73 partici- pants and 59 phase contrasts were coded as behavioral outcomes (e.g., disruptive behav- ior, noncompliance). Results indicated that studies that included academic outcomes had a lower ES (0.89; SE � 0.08; 95% CI [0.73, 1.00]) than studies that included behav- ioral outcomes (ES � 0.93; SE � 0.07; 95% CI [0.75, 1.00]). When the two categories were compared for parameter estimates, over- lapping CIs indicated there may not be a sta- tistically significant difference, which was confirmed by the values from the reliable- difference formula: z � 0.04, p � .69.
Response Cost Sixteen studies that used RC with 47
participants and 47 phase contrasts were ag- gregated and compared with the 12 studies with 43 participants and 41 phase contrasts without RC. Results indicated that studies that included RC had a lower ES (0.84; SE � 0.06; 95% CI [0.75, 0.95]) than studies that did not include RC (ES � 0.91; SE � 0.05; 95% CI [0.81, 1.00]). When the two categories were compared for parameter estimates, overlap- ping CIs indicated there may not be a sta- tistically significant difference, which was confirmed by the values from the reliable- difference formula: z � 0.99, p � .32.
Verbal Cueing Nine studies that included verbal cueing
with 20 participants and 25 phase contrasts were aggregated and compared with 19 studies including 70 participants and 63 phase changes without verbal cueing. Results indi- cated that studies using verbal cueing pro- duced a lower ES (0.83; SE � 0.06; 95% CI [0.74, 0.92]) than studies without verbal cue- ing (ES � 0.88; SE � 0.28; 95% CI [0.33, 1.00]). When the two categories were compared for parameter estimates, overlap- ping CIs indicated there may not be a statisti- cally significant difference, which was con- firmed by the values from the reliable-differ- ence formula: z � 0.20, p � .84.
DISCUSSION
This meta-analysis synthesized and an- alyzed the findings from SCED studies to evaluate the effectiveness of TEs in public schools from 1980 to 2014. We specifically set out to (a) evaluate the quality of the design, (b) calculate ESs and CIs, (c) stratify the results across quality, and (d) evaluate moderator variables of peer-reviewed literature, with a focus on both academic outcomes (e.g., task accuracy, task engagement) and behavioral outcomes (e.g., disruptive behavior, noncom- pliance). Supporting and contributing infor- mation to previous findings (Dickerson et al., 2005; Matson & Boisjoli, 2009), these results suggest that a TE is an effective intervention, specifically for use in classroom settings.
We first sought to evaluate the quality of the design in SCED on TEs. We evaluated the design quality, and contrary to Maggin et al. (2011), who found 30% of studies with de- signs of medium to strong quality, we found 64% rated as medium to strong. The remain- der of studies were rated as low primarily because of insufficient demonstration of ef- fects, lack of reporting of sufficient informa- tion (e.g., IOA), and the number of data points per phase. We found a large majority of in- cluded studies reported IOA. However, only a third of the studies reported treatment fidelity. This difference can, at least partially, be ex- plained by the different methodologies uti-
School Psychology Review, 2016, Volume 45, No. 4
392
lized, which resulted in inclusion of unique studies. Maggin et al. included studies that had not been through a peer-review process in an attempt to eliminate publication bias. We elected to statistically address publication bias. Thus, Maggin et al. included 10 studies published prior to 1980, two dissertations, and one study presented at a conference, and our review did not include studies published prior to 1980 and only included studies published after the peer-review process. Thus, there were only four overlapping studies between this meta-analysis and the Maggin et al. meta- analysis. Therefore, it appears that the quality of the research is increasing with time and through the peer-review process.
The second question addressed by this study was the overall effectiveness of TEs. The overall Tau-U was large from 88 individ- ual ESs. Similar to Maggin et al. (2011), re- sults provide preliminary evidence that a TE is an effective intervention. However, it is im- portant to note that we chose to evaluate the effectiveness of a TE with Tau-U to increase the trustworthiness of our results as 82% of the included studies showed a positive baseline trend, which is a threat to internal validity addressed by Tau-U. These results suggest further analytic evidence that a TE is effective at reducing challenging behaviors. In addition, results imply that a TE is an intervention that can be utilized to increase academic readiness skills in school settings.
The third question addressed by this meta-analysis was whether effects differed across design quality. When stratified, the ES for studies considered medium or large quality was strong. Given the results, methodological quality did not appear to explain differences in the effectiveness of the intervention; rather, methodological quality seemed to only explain the extent to which the study could be repli- cated. From these results, it is apparent that although some design quality issues are evi- dent, a majority of the data and designs are sufficient for a TE to be preliminarily consid- ered an evidence-based intervention for imple- mentation in classrooms.
The fourth question addressed was po- tential moderators. Results supported our a
priori hypotheses that neither setting nor out- come nor addition of RC or verbal cueing moderated the effects of a TE. A TE appeared to be equally as effective in general and spe- cial education classes, to target academic and behavioral outcomes, and with and without RC or verbal cueing. However, it does seem that a TE is slightly more effective with older children than with younger children.
The finding that setting did not moderate the effects of a TE is important. If the TE is an intervention needed by a student who receives special education services, it appears that re- sults provide initial support for use of a TE in general or special education settings. How- ever, this finding is not consistent with the report of DuPaul, Eckert, and McGoey (1997), who found that interventions had a greater impact on behavior when they were imple- mented in special education classrooms as op- posed to implementation in general education classrooms. Thus, more research is needed to determine if a TE can be implemented effec- tively in both general and special education classrooms.
Data suggest the only statistically signif- icant moderator was the age of the participant. Results indicated a TE was more effective for children age 6 years and above. Nonoverlap- ping CIs were observed in this moderator, indicating the difference was statistically sig- nificant. A TE was effective with both groups of children; however, it appears to be most effective with older children. This finding might be contributed to the probability of older students understanding the procedures and being able to identify items or activities that would be motivating as backup reinforc- ers. This finding is consistent with previous research (Brumfield & Roberts, 1998; Shriver & Allen, 1997) and suggests that the age of the child may be predictive of compliance with this intervention. The implication for this seems to be that implementers need to spend additional time on training, modeling, and practice prior to use with younger children to increase the likelihood of the child under- standing the concept of backup reinforcers.
Furthermore, the finding that outcomes did not moderate the effect of a TE is mean-
Main and Moderator Effects for Token Economies
393
ingful. The equality of a TE for both academic (e.g., task accuracy, task engagement) and be- havioral outcomes implies a single interven- tion may be incorporated for dual targets, which is likely to simplify the procedures for increased efficiency. Thus, practitioners can use TEs to target various outcomes.
Past studies have evaluated the effects across procedural differences, and our findings contribute to mixed results in this area. Spe- cifically, our findings suggest that a TE is effective with or without the use of RC. This finding differs from those who have found RC increases effectiveness (Center & Wascom, 1984; De Martini-Scully et al., 2000). Further- more, while our findings indicate a TE with and without verbal cues is equally effective, Kluger and DeNisi (1996) found that to min- imize the negative effects of the use of verbal cues (e.g., decreases internal motivation, draws attention to negative), cueing should only be used during goal setting. In addition, the researchers found if more verbal cueing was needed, there was a lack of understanding of or confusion about the expected behaviors. These findings may be affected by the strength of the token for expected behaviors and the exchange for the secondary reinforcer. As such, practitioners are advised to consider the importance of the secondary reinforcer to the student and change the secondary reinforcer when it is no longer motivating. Further study is needed to clarify these findings; however, we are encouraged that practitioners might be able to simplify a TE by eliminating RC and verbal cueing.
Several findings of the visual analysis are noteworthy and slightly temper the confi- dence with which we can say definitively that the studies strongly support TEs. However, Tau-U addresses some of the concerns. First, a majority of the studies had a positive baseline trend. This finding could be problematic be- cause of potential issues with internal validity. However, while this trend influenced out- comes of visual analyses, it did not influence Tau-U, as we controlled for baseline trend within the analyses. Second, 67% of the phases had variable data, and 40% of adjacent phases had overlapping data points. These
findings indicated fluctuation in participant per- formance and affected our decisions regarding the functional relationship between a TE and the outcomes. However, 78.70% of calculated Tau-U and decisions made through visual analysis were in agreement, providing stron- ger support for our findings.
Although results should be interpreted with caution because of low design quality, the current meta-analysis extends the empiri- cal literature suggesting the potential for a TE to be effective in classroom settings. Findings indicate medium to large effects for a TE overall when used to achieve academic and behavioral outcomes and when used with or without RC or verbal cues. Aggregated results across design quality reveal preliminary sup- port for use of TEs in special and general education classrooms. Professionals serving children in classroom settings may be encour- aged by the finding that a TE does not have to include additional components (e.g., verbal cueing and RC), thus minimizing the com- plexity of the intervention.
Limitations
Limitations of the current meta-analysis should be mentioned. First, studies were only included that were published in peer-reviewed journals as we were interested in the design quality of published studies. We conducted the Egger’s test to evaluate the sample for publi- cation bias; however, caution should be given to interpretations because of the small number of included studies. Second, although Wolery (2013) and others suggested treatment integ- rity be included as a component of design quality, we elected not to include it in our measure to align coding with the current stan- dards for demonstrating a functional relation- ship between the dependent and independent variables (i.e., Horner et al., 2005; Kratochwill et al., 2010). Third, as with any nonparametric measure of effect, Tau-U has some limitations including ceiling effects, as demonstrated here with 9 studies hitting the upper ceiling and 23 studies with CIs that hit the upper ceiling. Using nonparametric effects, in this study, outweighed the limitations. Fourth, age group-
School Psychology Review, 2016, Volume 45, No. 4
394
ings may potentially have limited this study. Ages 6 to 15 years is a broad range, and results may vary within the category. Finally, al- though design quality was rated as medium to strong for a majority of our studies, results should be interpreted in light of the design quality weaknesses of the remaining studies.
Implications for Practice and Future Research
Educators and professionals who work with children in schools struggle for econom- ical (use of time, training, money, personnel, expertise) interventions. Within school-wide behavior programs, students frequently re- ceive tokens, often in the form of “bucks” that can be traded in at a school-level store. In addition, a similar system could be put in place in a classroom setting. A TE can be built within an individual student behavior contract (see Soares, Cegelka, & Payne, 2016). Stu- dents can earn tokens and work toward having a sufficient number of tokens to purchase a desired item. A TE is effective with and with- out RC. Therefore, the procedure can be 100% positive with no removal of reinforcers.
More and varied research on the effects of a TE is needed. Future studies should cal- culate more than one ES (Kratochwill et al., 2010) and examine the level of treatment in- tegrity correlated with intervention effective- ness. We found that only 10 studies out of 28 reported levels of treatment integrity, and Maggin et al. (2011) found that 2 of 24 studies reported treatment fidelity. When authors do not report treatment integrity, implementation is assumed. Low levels of treatment integrity in schools are concerning. However, reports of low treatment fidelity in educational settings are frequent (Becker & Domitrovich, 2011; Riley-Tillman & Eckert, 2001) and potentially decrease the likelihood of effectiveness. Our finding suggests that a TE is effective and based on implementation of the intervention as designed. Researchers are only beginning to evaluate the level of integrity that is needed to achieve effects with evidence-based interven- tions in educational settings outside of the tight controls of research (see Owens et al.,
2014). Thus, this is an area of caution and one in which much research is needed. In addition, research is needed to identify the levels of professional development and ongoing coach- ing required to maintain reliable implementa- tion and generalization across behaviors and settings. Furthermore, researchers need to re- fer to the current quality standards when de- signing and implementing SCED.
REFERENCES
References marked with an asterisk indicate studies included in the meta-analysis. Abramson, J. H. (2011). WINPEPI updated: Computer
programs for epidemiologists, and their teaching potential. Epidemiological Perspective & Innova- tions, 8, 1–9. doi:10.1186/1742-5573-8-1. Retrieved from http://archive.biomedcentral.com/1742-5573/ content/8/1/1
Akin-Little, K. A., & Little, S. G. (2004). Re-examining the over justification effect: A case study. Journal of Behavioral Education, 13, 179 –192. doi:10.1023/ B:JOBE.0000037628.81867.69
Alberto, A. A., & Troutman, A. C. (2003). Applied be- havior analysis for teacher (6th ed.). Upper Saddle River, NJ: Merrill Prentice Hall.
Ary, D., & Suen, H. K. (1989). Analyzing qualitative behavioral observational data. Mahwah, NJ: Law- rence Erlbaum Associates.
Balcazar, F., Hopkins, B. L., & Suarez, Y. (1985). A critical, objective review of performance feedback. Journal of Organizational Behavior Management, 7(3– 4), 65– 89. doi:10.1300/J075v07n03_05
Becker, K. D., & Domitrovich, C. E. (2011). The concep- tualization, integration, and support of evidence-based interventions in the schools. School Psychology Re- view, 40(4), 582–589. Retrieved from http://eric.ed. gov/?id�EJ962068
Bippes, R., McLaughlin, T. F., & Williams, R. L. (1986). A classroom token system in a detention center: Ef- fects for academic and social behavior. Techniques: A Journal for Remedial Education and Counseling, 2, 126 –132. Retrieved from http://psycnet.apa.org/ psycinfo/1987-14003-001
Bowman-Perrott, L., Burke, M., Zhang, N., & Zaini, S. (2014). Direct and collateral benefits of peer tutoring on social and behavioral outcomes: A meta-analysis of single-case studies. School Psychology Review, 43, 260 –285. Retrieved from http://bmo.sagepub.com/ content/early/2014/09/25/0145445514551383.abstract
Brumfield, B. D., & Roberts, M. W. (1998). A comparison of two measurements of child compliance with normal pre- school children. Journal of Clinical Child Psychology, 27(1), 109 –116. doi:10.1207/s15374424jccp2701_12
Busk, P. L., & Serlin, R. C. (1992). Meta-analysis for single-case research. In T. R. Kratochwill & J. R. Levin (Eds.), Single-case research design and analy- sis: New directions for psychology and education (pp. 187–212). Hillsdale, NJ: Erlbaum.
Carlson, C. L., Pelham, W. E., Milich, R., & Dixon, J. (1992). Single and combined effects of methylpheni- date and behavior therapy on the classroom perfor-
Main and Moderator Effects for Token Economies
395
mance of children with attention-deficit hyperactivity disorder. Journal of Abnormal Child Psychology, 20, 213–232. doi:10.1007/BF00916549
*Carnett, A., Raulston, T., Lang, R., Tostanoski, A., Lee, A., Sigafoos, J., & Machalicek, W. (2014). Effects of a perseverative interest-based token economy on chal- lenging and on-task behavior in a child with autism. Journal of Behavioral Education, 23(3), 368 –377. doi: 10.1007/s10864-014-9195-7
Cavalier, A. R., Ferretti, R. P., & Hodges, A. E. (1997). Self-management within a classroom token economy for students with learning disabilities. Research in De- velopmental Disabilities, 18, 167–178. doi:10.1016/ S0891-4222(96)00045-5
*Center, D. B., & Wascom, A. (1984). Transfer of reinforcers: A procedure for enhancing response cost. Educational and Psychological Research, 4, 19 –27. Retrieved from http://davidcenter.com/ documents/Publications/39.pdf
Christensen, L., Young, K. R., & Marchant, M. (2004). The effects of a peer-mediated positive behavior sup- port program on socially appropriate classroom behav- ior. Education & Treatment of Children, 27, 199 –234. Retrieved from http://www.jstor.org/stable/42900544
*Conyers, C., Miltenberger, R. G., Gubin, A., Barenz, R., Jurgens, M., Sailer, A., . . . Kopp, B. (2004). A comparison of response cost and differential rein- forcement of other behavior to reduce disruptive behavior in a preschool classroom. Journal of Ap- plied Behavior Analysis, 37, 411– 415. doi:10.1901/ jaba.2004.37-411
Cooper, H. M., & Hedges, L. V. (Eds.). (1994). The handbook of research synthesis. New York, NY: Rus- sell Sage Foundation.
*De Martini-Scully, D., Bray, M. A., & Kehle, T. J. (2000). A packaged intervention to reduce disruptive behaviors in general education students. Psychology in the Schools, 37, 149 –156. doi:10.1002/(SICI)1520- 6807(200003)37:2 149::AID-PITS6�3.0.CO;2-K
Dickerson, F. B., Tenhual, W. N., & Green-Paden, L. D. (2005). The token economy for schizophrenia: Review of the literature and recommendations for future re- search. Schizophrenia Research, 75, 405– 416. doi: 10.1016/j.schres.2004.08.026
DuPaul, G. J., Eckert, T. L., & McGoey, K. E. (1997). Interventions for students with attention-deficit/hyper- activity disorder: One size does not fit all. School Psychology Review, 26, 369 –381. Retrieved from http://www.nasponline.org/publications/periodicals/ spr/volume-26/volume-26-issue-3/interventions-for- students-with-attention-deficit/hyperactivity-disorder- one-size-does-not-fit-all
DuPaul, G. J., & Weyandt, L. L. (2006). School-based intervention for children with Attention Deficit Hyper- activity Disorder: Effects on academic, social, and behavioural functioning. International Journal of Dis- ability, Development and Education, 53(2), 161–176. doi:10.1080/10349120600716141
Egger, M., Smith, G. D., Schneider, M., & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ, 315(7109), 629 – 634. doi: 10.1136/bmj.315.7109.629
Fairchild, A. J., & MacKinnon, D. P. (2009). A general model for testing mediation and moderation effects. Prevention Science: The Official Journal of the Society
for Prevention Research, 10(2), 87–99. doi:10.1007/ s11121-008-0109-6
Feindler, E. L., Marriott, S. A., & Iwata, M. (1984). Group anger control training for junior high school delin- quents. Cognitive Therapy and Research, 8(3), 299 – 311. doi:10.1007/BF01173000
*Filcheck, H. A., McNeil, C. B., Greco, L. A., & Bernard, R. S. (2004). Using a whole-class token economy and coaching of teacher skills in a preschool classroom to manage disruptive behavior. Psychology in the Schools, 41, 351–361. doi:10.1002/pits.10168
GetData Graph Digitizer (Version 2.21) [Software]. Re- trieved from http://www.getdata-graph-digitizer.com/ download.php
Glass, G. V. (1976). Primary, secondary, and meta-anal- ysis of research. Educational Researcher, 5, 3– 8. doi: 10.3102/0013189x005010003
Gresham, F. M. (1979). Comparison of response cost and timeout in a special education setting. Journal of Special Education, 13, 199 –208. doi:10.1177/ 002246697901300211
Hartmann, D. P., Barrios, B. A., & Wood, D. D. (2004). Principles of behavioral observation. In S. N. Haynes & E. M. Heiby (Eds.), Comprehensive handbook of psychological assessment: Behavioral assessment (Vol. 3, pp. 108 –127). Hoboken, NJ: John Wiley and Sons.
Heaton, R. C., & Safer, D. J. (1982). Secondary school outcome following a junior high school behavioral program. Behavior Therapy, 13, 226 –231. doi: 10.1016/S0005-7894(82)80066-X
Higgins, J. P., & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medi- cine, 21(11), 1539 –1558. doi:10.1002/sim.1186
*Higgins, J. W., Williams, R. L., & McLaughlin, T. F. (2001). The effects of a token economy employing instructional consequences for a third-grade student with learning disabilities: A data-based case study. Education & Treatment of Children, 24, 99 –106. Re- trieved from http://www.jstor.org/stable/42899646
*Himle, M. B., Woods, D. W., & Bunaciu, L. (2008). Evaluating the role of contingency in differentially reinforced tic suppression. Journal of Applied Behav- ior Analysis, 41, 285–289. doi:10.1901/jaba.2008.41- 285
Hintze, J. (2004). NCSS 2004 [Computer software]. Re- trieved from http://www.ncss.com
Hopko, D. R., Lejuez, C. W., Lepage, J. P., Hopko, S. D., & McNeil, D. W. (2003). A brief behavioral activation treatment for depression: A randomized pilot trial within an inpatient psychiatric hospital. Behavior Modification, 27(4), 458 – 469. doi:10.1177/0145445503255489
Horner, R. H., Carr, E. G., Halle, J., McGee, J., Odom, S., & Wolery, M. (2005). The use of single-subject re- search to identify evidence-based practice in special education. Exceptional Children, 71, 165–179. doi: 10.1177/001440290507100203
House, A. E., House, B. J., & Campbell, M. B. (1981). Measures of interobserver agreement: Calculation for- mulas and distribution effects. Journal of Behavioral Assessment, 3, 37–57. doi:10.1007/BF01321350
Individuals with Disabilities Education Improvement Act, H.R. 1350, 108th Cong. (2004).
*Jones, M. N., Weber, K. P., & McLaughlin, T. F. (2013). No teacher left behind: Educating students with ASD and ADHD in the inclusion classroom.
School Psychology Review, 2016, Volume 45, No. 4
396
The Journal of Special Education Apprenticeship, 2(2), 1–22. Retrieved from http://josea.info/ archives/vol2no2/vol2no2-5-FT.pdf
Kazdin, A. E. (1971). The effect of response cost in sup- pressing behavior in a pre-psychotic retardate. Journal of Behavior Therapy and Experimental Psychiatry, 2, 137– 140. doi:10.1016/0005-7916(71)90029-2
Kazdin, A. E. (1972). Response cost: The removal of conditioned reinforcers for therapeutic change. Be- havior Therapy, 3, 533–546. doi:10.1016/S0005- 7894(72)80001-7
Kazdin, A. E. (1982). Single-case research designs: Meth- ods for clinical and applied settings. New York, NY: Oxford University Press.
Kazdin, A. E., & Bootzin, R. R. (1972). The token econ- omy: An evaluative review. Journal of Applied Behav- ior Analysis, 5, 343–372. doi:10.1901/jaba.1972.5-343
*Kazdin, A. E., & Geesey, S. (1980). Enhancing class- room attentiveness by preselection of back rein forcers in a token economy. Behavior Modification, 4, 98 – 114. doi:10.1177/014544558041006
*Kazdin, A. E., & Mascitelli, S. (1980). The opportunity to earn oneself off a token system as a reinforcer for attentive behavior. Behavior Therapy, 11, 68 –78. doi: 10.1016/s0005-7894(80)80037-2
*Klimas, A., & McLaughlin, T. F. (2007). The effects of a token economy system to improve social and aca- demic behavior with a rural primary aged child with disabilities. International Journal of Special Educa- tion, 22, 72–77. Retrieved from http://eric.ed.gov/ ?id�EJ814513
Kluger, A. N., & DeNisi, A. (1996). The effects of feed- back interventions on performance: A historical re- view, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119, 254 – 284. doi:10.1037/0033-2909.119.2.254
Kratochwill, T. R., Hitchcock, J., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., & Shadish, W. R. (2010). Single-case designs technical documentation. Retrieved from http://ies.ed.gov/ncee/wwc/pdf/wwc_ scd.pdf
Lane, J. D., & Gast, D. L. (2014). Visual analysis in single- case experimental design studies: Brief review and guide- lines. Neuropsychological Rehabilitation, 24(3– 4), 445– 463. doi:10.1080/09602011.2013.815636
Latham, G. P., & Locke, E. A. (1991). Self-regulation through goal setting. Organizational Behavior and Hu- man Decision Processes, 50, 212–247. doi:10.1016/ 0749-5978(91)90021-K
Maggin, D. M., Briesch, A. M., & Chafouleas, S. M. (2013). An application of the What Works Clearinghouse stan- dards for evaluating single-subject research: Self-man- agement interventions. Remedial and Special Education, 34, 44 –58. doi:10.1177/0741932511435176
Maggin, D. M., & Chafouleas, S. M. (2010). PASS-RQ: Protocol for assessing single-subject research quality. Unpublished research instrument.
Maggin, D. M., Chafouleas, S. M., Goddard, K. M., & Johnson, A. H. (2011). A systematic evaluation of token economies as a classroom management tool for students with challenging behavior. Journal of School Psychology, 49, 529 –554. doi:10.1016/j.jsp.2011.05.001
*Maglio, C., & McLaughlin, T. F. (1981). Effects of a token reinforcement system and teacher attention in reducing inappropriate verbalization with a junior high school student. Corrective and Social Psychiatry, 27,
140 –145. Retrieved from http://psycnet.apa.org/ psycinfo/1982-24369-001
Manolov, R., Solanas, A., Sierra, V., & Evans, J. J. (2011). Choosing among techniques for quantifying single-case intervention effectiveness. Behavior Ther- apy, 42(3), 533–545.
Martin, G., & Pear, J. (2003). Behavior modification: What it is and how to do it? (7th ed.). Upper Saddle River, NJ: Simon & Schuster.
Matson, J. L., & Boisjoli, J. A. (2009). The token econ- omy for children with intellectual disability and/or autism: A review. Research on Developmental Dis- abilities, 30, 240 –248. doi:10.1016/j.ridd.2008.04.001
*McGoey, K. E., & DuPaul, G. J. (2000). Token rein- forcement and response cost procedures: Reducing the disruptive behavior of preschool children. School Psychology Quarterly, 15, 330 –343. doi:10.1037/ h0088790
McKillup, S. (2011). Statistics explained: An introductory guide for life scientists. Cambridge, UK: Cambridge University Press.
*Millersmith, T., Weber, K. P., & McLaughlin, T. F. (2013). The use of token economy and a math manip- ulative for a child with moderate intellectual disabili- ties. International Journal of Basics and Applied Sci- ences, 1(3), 634 – 640. Retrieved from http://www. insikapub.com/
Miltenberger, R. G. (2001). Behavior modification: Prin- ciples and procedures (2nd ed.). Pacific Grove, CA: Brooks/Cole.
*Mottram, L. M., Bray, M. A., Kehle, T. J., Broudy, M., & Jenson, W. R. (2002). A classroom-based intervention to reduce disruptive behaviors. Journal of Applied School Psychology, 19, 65–74. doi: 10.1300/j370v19n01_05
Murray, L., & Sefchik, G. (1992). Regulating behavior management practices in residential treatment facili- ties. Children and Youth Services Review, 14(6), 519 – 539. doi:10.1016/0190-7409(92)90004-F
*Musser, E. H., Bray, M. A., Kehle, T. J., & Jenson, W. R. (2001). Reducing disruptive behaviors in students with serious emotional disturbance. School Psychology Re- view, 30, 294 –304. Retrieved from http://www. nasponline.org/publications/spr/abstract.aspx?ID�1590
Owens, J. S., Lyon, A. R., Brant, N. E., Masia-Warner, C., Nadeem, E., Spiel, C., & Wagner, M. (2014). Imple- mentation science in school mental health: Key con- structs in a developing research agenda. School Mental Health, 6(2), 99 –111. doi:10.1007/s12310-013-9115-3
Ninci, J., Neely, L. C., Hong, E. R., Boles, M. B., Gilli- land, W. D., Ganz, J. B., . . . Vannest, K. J. (2015). Meta-analysis of interventions to improve functional living skills for people with autism spectrum disorder. Review of Journal of Autism and Developmental Dis- orders. 2, 184 –198. doi:10.1007/s40489-014-0046-1
No Child Left Behind Act of 2001, 20 U.S.C. 70 § 6301 et seq (2001).
Parker, R., & Hagan-Burke, S. (2007). Useful effect size interpretations for single-case research. Behavior Ther- apy, 38, 95–105. doi:10.1016/j.beth.2006.05.002
Parker, R. I., & Vannest, K. J. (2012). Bottom up analysis of single-case research designs. Journal of Behavioral Education, 17(1), 254 –265. doi:10.1007/s10864-012- 9153-1
Parker, R. I., Vannest, K. J., & Brown, L. (2009). The “improvement rate difference” for single-case re-
Main and Moderator Effects for Token Economies
397
search. Exceptional Children, 75, 135–150. Retrieved from http://eric.ed.gov/?id�EJ842529
Parker, R. I., Vannest, K. J., Davis, J. L., & Sauber, S. B. (2011). Combining non-overlap and trend for single- case research: Tau-U. Behavior Therapy, 42, 284 –299. doi:10.1016/j.beth.2010.08.006
Phillips, E. L., Phillips, E. A., Fixsen, D. L., & Wolf, M. M. (1971). Achievement place: Modification of the behaviors of pre-delinquent boys within a token econ- omy. Journal of Applied Behavior Analysis, 4, 45–50. doi:10.1901/jaba.1971.4-45
Rapport, M. D., Murphy, A., & Bailey, J. S. (1980). The effects of a response cost treatment tactic on hyperac- tive children. Journal of School Psychology, 18, 98 – 111. doi:10.1016/0022-4405(80)90025-4
*Reitman, D., Murphy, M. A., Hupp, S. D. A., & O’Callaghan, P. M. (2004). Behavior change and per- ceptions of change: Evaluating the effectiveness of a token economy. Child & Family Behavior Therapy, 26(2), 17–36. doi:10.1300/J019v26n02_02
Rhode, G., Jenson, W. R., & Reavis, H. K. (1993). The tough kid book: Practical classroom management strategies. Longmont, CO: Sopris West.
Riley-Tillman, T. C., & Eckert, T. L. (2001). Generaliza- tion programming and school based consultation: An examination of consultees’ generalization of consulta- tion-related skills. Journal of Educational and Psycho- logical Consultation, 12, 217–241. doi:10.1207/ s1532768xjepc1203_03
Rosen, L. A., Taylor, S. A., O’Leary, S. G., & Sanderson, W. (1990). A survey of classroom management prac- tices. Journal of School Psychology, 28(3), 257–269. doi:10.1016/0022-4405(90)90016-Z
*Rosenberg, M. S. (1986). Maximizing the effectiveness of structured classroom management programs: Imple- menting rule-review procedures with disruptive and distractible students. Behavioral Disorders, 11, 239 – 248. Retrieved from http://www.jstor.org/stable/ 23882205
Rosenthal, R., & DiMatteo, M. R. (2001). Meta-analysis: Recent developments in quantitative methods for liter- ature reviews. Annual Review of Psychology, 52, 59 – 82. doi:10.1146/annurev.psych.52.1.59
*Salend, S. J., & Allen, E. M. (1985). Comparative effects of externally-managed response cost systems on inappropri- ate classroom behavior. Journal of School Psychology, 23, 59 – 67. doi:10.1016/0022-4405(85)90035-4
*Salend, S. J., & Lamb, E. A. (1986). Effectiveness of a group-managed interdependent contingency system. Learning Disability Quarterly, 9, 268 –273. doi: 10.2307/1510380
*Salend, S. J., Tintle, L., & Balber, H. (1988). Effects of a student-managed response cost system on the behav- ior of two mainstreamed students. The Elementary School Journal, 89, 89 –97. doi:10.1086/461564
Schellenberg, T., Skok, R., & McLaughlin, T. F. (1991). The effects of contingent free time on homework com- pletion in English with high school English students. Child & Family Behavior Therapy, 13(3), 1–12. doi: 10.1300/J019v13n03_01
Scruggs, T. E., & Mastropieri, M. A. (2001). How to summarize single-participant research: Ideas and ap- plication. Exceptionality, 9, 227–244. doi:10.1207/ S15327035EX0904_5
Shriver, M. D., & Allen, K. D. (1997). Defining child noncompliance: An examination of temporal parame-
ters. Journal of Applied Behavior Analysis, 30(1), 173– 176. doi:10.1901/jaba.1997.30-173
*Simon, S. J., Ayllon, T., & Milan, M. A. (1982). Behav- ioral compensation: Contrast like effects in the class- room. Behavior Modification, 6, 407– 420. doi: 10.1177/014544558263006
Simonsen, B., Fairbanks, S., Briesch, A., Myers, D., & Sugai, G. (2008). Evidence-based practices in class- room management: Considerations for research to practice. Education and Treatment of Children, 31, 351–380. doi:10.1353/etc.0.0007
Skinner, B. F. (1931). The concept of the reflex in the description of behavior. Journal of General Psychology, 5, 427– 458. doi:10.1080/00221309.1931.9918416
Skinner, C. H., Cashwell, C. S., & Bunn, M. S. (1996). Independent and interdependent group contingencies: Smoothing the rough waters. Special Services in the Schools, 12, 61–78. doi:10.1300/J008v12n01_04
*Smith, L. K., & Fowler, S. A. (1984). Positive peer pressure: The effects of peer monitoring on children’s disruptive behavior. Journal of Applied Behavior Anal- ysis, 17, 213–227. doi:10.1901/jaba.1984.17-213
Soares, D. A., Cegelka, W. J., & Payne, J. S. (2016). The token economy playbook: The ultimate guide to pro- moting superior performance and personal growth. San Diego, CA: University Readers.
*Sran, S. K., & Borrero, J. C. (2010). Assessing the value of choice in a token system. Journal of Ap- plied Behavior Analysis, 43, 553–557. doi:10.1901/ jaba.2010.43-553
*Stevens, C., Sidener, T. M., Reeve, S. A., & Sidener, D. W. (2011). Effects of behavior-specific and general praise on acquisition of tacts in children with pervasive develop- mental disorders. Research in Autism Spectrum Disor- ders, 5, 666 – 669. doi:10.1016/j.rasd.2010.08.003
Stilitz, I. (2009). A token economy of the early 19th century. Journal of Applied Behavior Analysis, 42(4), 925–926. doi:10.1901/jaba.2009.42-925
Strijbos, J., Martens, R., Prins, F., & Jochems, W. (2006). Content analysis: What are they talking about? Computers & Education, 46, 29 – 48. doi: 10.1016/j.compedu.2005.04.002
*Sullivan, M. A., & O’Leary, S. G. (1990). Maintenance following reward and cost token programs. Behavior Therapy, 21, 139 –149. doi:10.1016/s0005-7894(05) 80195-9
*Thompson, M. J., McLaughlin, T. F., & Derby, K. M. (2011). The use of differential reinforcement to de- crease the inappropriate verbalizations of a nine-year old girl with autism. Electronic Journal of Research in Educational Psychology, 9(1), 183–196. Retrieved from http://eric.ed.gov/?id�EJ926483
*Truchlicka, M., McLaughlin, T. F., & Swain, J. C. (1998). Effects of token reinforcement and response cost on the accuracy of spelling performance with middle-school special education students with behav- ior disorders. Behavioral Interventions, 13, 1–10. doi:10.1002/(SICI)1099-078X(199802)13:1 1::AID- BIN1�3.0.CO;2-Z
Ulmer, R. A. (1976). On the development of a token economy mental hospital treatment program. Wash- ington, DC: Hemisphere.
Witt, J. C., & Elliot, S. N. (1982). The response cost lottery: A time efficient and effective classroom inter- vention. Journal of School Psychology, 20(2), 155– 161. doi:10.1016/0022-4405(82)90009-7
School Psychology Review, 2016, Volume 45, No. 4
398
Wolery, M. (2013). A commentary single-case design technical document of the What Works Clearinghouse. Remedial and Special Education, 34(1), 39 – 43. doi: 10.1177/0741932512468038
Van den Noortgate, W., & Onghena, P. (2003). Hierar- chical linear models for the quantitative integration of effect sizes in single-case research. Behavior Research Methods, Instruments, & Computers, 35, 1–10. doi: 10.3758/bf03195492
Van den Noortgate, W., & Onghena, P. (2008). A multi- level meta-analysis of single-subject experimental design studies. Evidence-Based Communication As- sessment & Intervention, 2, 142–151. doi:10.1080/ 17489530802505362
Vannest, K.J., Parker, R.I., & Gonen, O. (2011). Single case research: Web based calculators for SCR analysis
(Version 1.0) [Web-based application]. College Sta- tion, TX: Texas A&M University. Retrieved from singlecaseresearch.org
Voight, M. L., & Hoogenboom, B. J. (2012). Publishing your work in a journal: Understanding the peer review process. International Journal of Sports Physical Ther- apy, 7(5), 452– 460. Retrieved from http://www.ncbi. nlm.nih.gov/pmc/articles/PMC3474310/
Yeaton, W. H., & Wortman, P. M. (1993). On the reli- ability of meta-analytic reviews: The role of intercoder agreement. Evaluation Review, 17(3), 292–309. doi: 10.1177/0193841X9301700303
Date Received: April 4, 2015 Date Accepted: August 26, 2015
Associate Editor: Lisa Bowman-Perrott
Denise A. Soares, PhD, is the Assistant Department Chair of Teacher Education, an assistant professor of special education, and Special Education Program Coordinator at the University of Mississippi. Her research interests include applied and practical expe- riences in academic and behavior interventions for at-risk students, as well as examining the efficacy of those interventions in classroom settings where teachers have competing time demands.
Judith R. Harrison, PhD, is an assistant professor in the Department of Educational Psychology–Special Education at Rutgers University in New Brunswick, New Jersey. Her research interests include the effectiveness, acceptability, and feasibility of assessment, interventions, and other services for youth with emotional and behavioral disorders and attention deficit hyperactivity disorder.
Kimberly J. Vannest, PhD, is a professor in the Department of Educational Psychology– Special Education at Texas A&M University. Her research interests are in determining effective interventions for children and youth with or at risk for emotional and behavioral disorders, including teacher behaviors and measurement.
Susan S. McClelland, PhD, is an associate professor of educational leadership and Chair of the Department of Teacher Education at the University of Mississippi. Her research interests include leadership for students with disabilities, literacy, school and organiza- tional culture, and issues relating to rural education.
Main and Moderator Effects for Token Economies
399
Copyright of School Psychology Review is the property of National Association of School Psychologists and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use.