Week 6 - Assignment: Examine the Impacts of Policies Implemented During the Great Recession and Week 7 - Assignment: Measure the Effects of Social Policies
Presidential Address: Making Federal Social Programs Work
Ron Haskins
One of the opportunities presented by being named President-elect of the Associa- tion for Public Policy Analysis and Management (APPAM) is the privilege of playing a major role in planning the annual conference. This opportunity includes the right to name the theme of the conference. Thus, after years of working to both under- stand and promote the use of social science evidence to improve the nation’s social policy, I was able to declare that we are in “the golden age of evidence-based policy” with some prospect that policy analysts and other scholars might notice. It would be difficult to exaggerate the pleasure I took in seeing in print, backed by the finest organization in the nation promoting the use of research to improve policy and management, the claim that evidence-based policy is an important and growing movement.
Another privilege of the APPAM Presidency is the opportunity to present the Presidential Address during the annual conference and to prepare a written version of the address for publication in APPAM’s distinguished Journal of Policy Analysis and Management. For me, these opportunities create a second and third chance to strengthen the theme that we are living in the golden age of evidence-based policy. Thus, I am seizing this opportunity and, as in my presidential address, intend to lay out the case that we are indeed in the golden age of evidence-based policy. After summarizing the case, I turn to an examination of where and how the focus on evidence can lead the nation’s social policy. For openers, let’s stipulate that though we may be in the golden age of evidence-based policy, it is not yet clear that the many branches of the evidence-based movement will actually lead to major improvements in the nation’s social policy. I am optimistic, but guarded. It would be hypocritical to make bold predictions about the impacts of the evidence-based policy movement on the nation’s social problems until we have better evidence about those impacts.
THE EVIDENCE-BASED MOVEMENT BRANCHES OUT
And make no mistake—the primary goal of the evidence-based movement is to reduce the nation’s social problems by improving the nation’s social programs. The straightforward meaning of this goal is that more evidence-based social pro- grams will reliably improve the national level of the social problem—delinquency, teen pregnancy, school dropout, college enrollment, economic equality, poverty, unemployment—they are designed to address. A major part of the progress of the evidence-based movement is that the basic tools to achieve and detect program impacts are well developed. These consist of highly specified intervention programs, a sophisticated and widely understood science of program evaluation, a growing but
Journal of Policy Analysis and Management, Vol. 36, No. 2, 276–302 (2017) C© 2017 by the Association for Public Policy Analysis and Management Published by Wiley Periodicals, Inc. View this article online at wileyonlinelibrary.com/journal/pam DOI:10.1002/pam.21983
Presidential Address: Making Federal Social Programs Work / 277
still skimpy understanding of how to scale up successful programs, and a number of other building blocks of expanding the number of programs that can successfully reduce the nation’s social problems.
In this article, I review the innovative building blocks that are leading to further progress of the evidence-based movement and its ongoing attack on the nation’s social problems. Taken as a whole, these building blocks not only reinforce the claim that we are in the golden age of evidence-based policy, but also show that innovative approaches to improving the nation’s social policy continue to appear. Moreover, most of these building blocks have already shown some success.
As we have seen, the truest measure of success will be that evidence-based social intervention programs play a role in reducing the national social problems they address. I do not think we yet have a convincing way to isolate the impact of, say, 1,000 evidence-based programs in operation throughout the nation in accounting for a change in a national social problem (Berlin, 2014). If, for example, the high school dropout rate had fallen by 50 percent over a certain period and there were more than 1,000 evidence-based programs in operation throughout the nation, the field would seem to be justified in claiming that evidence-based programs are contributing to progress. We do not have large-scale examples like this, but below I analyze an example of a case with more than 40 model programs addressed to reducing teen pregnancy. This example yields many implications for the future of scaling up evidence-based programs to attack the nation’s social problems. But first, let us review the building blocks of the evidence-based movement.
THE EXPANDING BUILDING BLOCKS OF EVIDENCE-BASED POLICY
Table 1 summarizes major characteristics of 14 of the most prominent and promis- ing building blocks of the evidence-based movement. In addition to this overview of selected building blocks, I turn now to a more detailed description of eight of these building blocks—tiered-evidence grantmaking, Results First, Pay for Success, be- havioral science (“nudge”) interventions, clearinghouses, nonprofit interest groups, administrative data, and the evidence-based policymaking commission—in order to provide a deeper understanding of how powerful these foundations of the movement are and why it is reasonable to claim that we are in the golden age of evidence-based policy.
Tiered Evidence
The tiered-evidence grant program is a method of distributing federal grant funds. The essence of the strategy is to give most of the funds in a given federal grant program, say 75 percent, only to organizations planning to implement model social programs that are evidence based. The remaining 25 percent of grant funds would be reserved for innovative programs that could, following rigorous evaluation, earn the status of evidence based if they are shown to produce impacts by an evaluation with a rigorous design. The two advantages of this strategy are, first, that most federal grant funds would be spent on model programs supported by rigorous evidence of producing impacts, and, second, that some but fewer grant funds would be spent on innovative programs that will over time produce a steady stream of new evidence- based social programs. As shown by the list of large federal grant programs in Table 2, which presents only a partial slice of federal grant programs, the potential field of application of the tiered-evidence strategy is enormous. According to OMB (2016), in 2015 the federal government gave $467.5 billion in grants to state and local governments, most of which could be spent on the tiered-evidence strategy. In my view, the tiered-evidence strategy is the most general and potentially powerful of all
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
278 / Presidential Address: Making Federal Social Programs Work
T a b
le 1
. T
h e
b u
il d
in g
b lo
ck s
o f
th e
ev id
en ce
-b a se
d p
o li
cy m
o ve
m en
t.
B u
il d
in g
b lo
ck A
p p
ro x im
a te
y ea
r st
a rt
ed O
ri g in
a to
r o
r ex
a m
p le
s D
es cr
ip ti
o n
T ie
re d
- ev
id en
ce p
ro g ra
m s
2 0
0 9
O b
a m
a a d
m in
is tr
a ti
o n
A st
ra te
g y
fo r
in cr
ea si
n g
th e
im p
a ct
o f
fe d
er a l
g ra
n t
d o
ll a rs
o n
so ci
a l
p ro
b le
m s
b y
sp en
d in
g a b
o u
t 7
5 %
o f
th e
fu n
d s
o n
p ro
g ra
m s
th a t
h a ve
ri g o
ro u
s ev
id en
ce o
f p
ro d
u ci
n g
im p
a ct
s a n
d a ro
u n
d 2
5 %
o f
th e
fu n
d s
o n
p ro
m is
in g
p ro
g ra
m s
th a t
h a ve
n o
t y et
b ee
n te
st ed
b y
ri g o
ro u
s d
es ig
n s
R es
u lt
s F
ir st
2 0
1 0
P ew
a n
d M
a cA
rt h
u r
F o
u n
d a ti
o n
s; G
a ry
V a n
la n
d in
g h
a m
fr o
m P
ew w
a s
th e
to p
a d
m in
is tr
a to
r a s
th e
p ro
g ra
m ex
p a n
d ed
d u
ri n
g it
s ea
rl y
y ea
rs
A p
ro g ra
m th
a t
h el
p s
st a te
s re
vi ew
th ei
r cu
rr en
t p
ro g ra
m s
in se
le ct
ed a re
a s
o f
so ci
a l
p o
li cy
, d
et er
m in
e w
h et
h er
th e
p ro
g ra
m s
m ee
t a
b en
ef it
-c o
st te
st ,
a n
d re
p la
ce o
n es
th a t
d o
n o
t w
it h
ev id
en ce
-b a se
d p
ro g ra
m s;
2 1
st a te
s p
a rt
ic ip
a te
P a y
fo r
S u
cc es
s F
ir st
P a y
fo r
S u
cc es
s ex
p er
im en
t in
th e
U n
it ed
S ta
te s,
ca ll
ed th
e A
d o
le sc
en t
B eh
a vi
o ra
l L
ea rn
in g
E x p
er ie
n ce
p ro
g ra
m ,
w a s
la u
n ch
ed a t
R ik
er s
Is la
n d
in N
ew Y
o rk
C it
y in
2 0
1 2
C o
n d
u ct
ed b
y a
te a m
th a t
in cl
u d
ed M
D R
C ,
N ew
Y o
rk C
it y
D ep
a rt
m en
t o
f C
o rr
ec ti
o n
s, O
sb o
rn e
A ss
o ci
a ti
o n
, F
ri en
d s
o f
Is la
n d
A ca
d em
y , a n
d V
er a
In st
it u
te o
f Ju
st ic
e; fi
n a n
ci n
g w
a s
p ro
vi d
ed b
y G
o ld
m a n
S a ch
s a n
d g u
a ra
n te
ed b
y B
lo o
m b
er g
P h
il a n
th ro
p ie
s
A n
in n
o va
ti ve
m et
h o
d o
f fi
n a n
ci n
g so
ci a l
in te
rv en
ti o
n p
ro g ra
m s
in w
h ic
h a n
in d
iv id
u a l,
co rp
o ra
te ,
o r
fo u
n d
a ti
o n
in ve
st o
r p
a y s
fo r
a so
ci a l
in te
rv en
ti o
n p
ro g ra
m ;
th e
p ro
g ra
m is
ri g o
ro u
sl y
ev a lu
a te
d a n
d th
e in
ve st
o r’
s fu
n d
in g
is re
tu rn
ed if
th e
p ro
g ra
m p
ro d
u ce
s b
en ef
it s
th a t
ju st
if y
th e
re tu
rn (p
er h
a p
s w
it h
in te
re st
)
B eh
a vi
o ra
l E
co n
o m
ic s
P ro
g ra
m s
A p
p ro
x im
a te
ly 1
9 8
0 s
B a se
d o
n re
se a rc
h in
p sy
ch o
lo g y , es
p ec
ia ll
y th
a t
co n
d u
ct ed
b y
D a n
ie l
K a h
n em
a n
a n
d A
m o
s T
ve rs
k y ;
a la
n d
m a rk
ev en
t in
th e
h is
to ry
o f
b eh
a vi
o ra
l ec
o n
o m
ic s
w a s
p u
b li
ca ti
o n
o f
N u
d ge
: Im
p ro
v in
g D
ec is
io n
s a b
o u
t H
ea lt
h ,
W ea
lt h
, a n
d H
a p
p in
es s
in 2
0 0
8 b
y R
ic h
a rd
T h
a le
r a n
d C
a ss
S u
n st
ei n
T h
e w
o rk
o f
th e
W h
it e
H o
u se
S o
ci a l
a n
d B
eh a vi
o ra
l S
ci en
ce T
ea m
(S B
S T
) is
a g o
o d
ex a m
p le
o f
h o
w b
eh a vi
o ra
l ec
o n
o m
ic s
p ro
g ra
m s
h a ve
in fl
u en
ce d
p o
li cy
. S
B S
T w
o rk
s w
it h
fe d
er a l
a g en
ci es
to d
es ig
n a n
d te
st in
te rv
en ti
o n
s b
a se
d o
n in
si g h
ts o
f b
eh a vi
o ra
l sc
ie n
ce to
im p
ro ve
g o
ve rn
m en
t ef
fi ci
en cy
a n
d in
cr ea
se th
e im
p a ct
o f
so ci
a l
p ro
g ra
m s.
S B
S T
h a s
p u
b li
sh ed
tw o
a n
n u
a l re
p o
rt s
w it
h 4
0 ex
p er
im en
ts b
a se
d o
n th
e p
ri n
ci p
le s
o f
b eh
a vi
o ra
l sc
ie n
ce a n
d p
er fo
rm ed
w it
h va
ri o
u s
fe d
er a l
a n
d st
a te
a g en
ci es
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 279
T a b
le 1
. C
o n
ti n
u ed
.
B u
il d
in g
b lo
ck A
p p
ro x im
a te
y ea
r st
a rt
ed O
ri g in
a to
r o
r ex
a m
p le
s D
es cr
ip ti
o n
C le
a ri
n g h
o u
se s
T h
e C
a li
fo rn
ia E
vi d
en ce
-B a se
d C
le a ri
n g h
o u
se fo
r C
h il
d W
el fa
re (C
E B
C ),
o n
e o
f th
e o
ld es
t cl
ea ri
n g h
o u
se s,
w a s
st a rt
ed in
2 0
0 4
T h
e C
a li
fo rn
ia E
vi d
en ce
-B a se
d C
le a ri
n g h
o u
se fo
r C
h il
d W
el fa
re (C
E B
C )
w a s
fo u
n d
ed b
y th
e C
a li
fo rn
ia D
ep a rt
m en
t o
f S
o ci
a l
S er
vi ce
s a n
d th
e C
h a d
w ic
k C
en te
r fo
r C
h il
d re
n a n
d F
a m
il ie
s
C le
a ri
n g h
o u
se s
co n
ta in
ex te
n si
ve in
fo rm
a ti
o n
a b
o u
t ev
a lu
a ti
o n
s o
f so
ci a l
in te
rv en
ti o
n p
ro g ra
m s.
T h
er e
a re
n in
e cl
ea ri
n g h
o u
se s
o n
so ci
a l
p ro
g ra
m s,
o n
e o
f w
h ic
h (t
h e
P ew
/M a cA
rt h
u r
R es
u lt
s F
ir st
C le
a ri
n g h
o u
se )
in cl
u d
es th
e co
n te
n t
o f
th e
o th
er ei
g h
t. T
h e
cl ea
ri n
g h
o u
se s
a re
d ev
o te
d to
a p
a rt
ic u
la r
ty p
e o
f so
ci a l
p ro
g ra
m su
ch a s
cr im
e a n
d d
el in
q u
en cy
, ch
il d
w el
fa re
, p
ro g ra
m s
fo r
y o
u th
, ed
u ca
ti o
n , a n
d so
fo rt
h . S
ev er
a l
o f
th e
cl ea
ri n
g h
o u
se s
p ro
vi d
e a n
o ve
ra ll
ra ti
n g
o f
th e
ef fe
ct iv
en es
s o
f th
e so
ci a l
in te
rv en
ti o
n p
ro g ra
m s
th ey
re vi
ew N
o n
p ro
fi t
in te
re st
g ro
u p
s
In te
re st
g ro
u p
s su
p p
o rt
in g
so ci
a l
p o
li cy
h a ve
b ee
n a
st a n
d a rd
fe a tu
re in
th e
U n
it ed
S ta
te s
fo r
m a n
y d
ec a d
es
R es
u lt
s fo
r A
m er
ic a ,
C o
a li
ti o
n fo
r E
vi d
en ce
-B a se
d P
o li
cy ,
se ve
ra l
fo u
n d
a ti
o n
s, a n
d m
a n
y o
th er
g ro
u p
s
T h
e su
p p
o rt
p ro
vi d
ed b
y n
o n
p ro
fi t
in te
re st
g ro
u p
s in
cl u
d es
fi n
a n
ci n
g , lo
b b
y in
g fe
d er
a l
a n
d st
a te
g o
ve rn
m en
ts ,
a n
a ly
si s
o f
so ci
a l
p ro
b le
m s,
co m
m u
n ic
a ti
o n
s, a n
d re
se a rc
h
A d
m in
is tr
a ti
ve d
a ta
A d
m in
is tr
a ti
ve d
a ta
se ts
h a ve
b ee
n a
st a n
d a rd
fe a tu
re o
f b
o th
g o
ve rn
m en
ta l
a n
d p
ri va
te en
ti ti
es fo
r m
a n
y d
ec a d
es
U se
fu l
d a ta
se ts
fo r
a d
va n
ci n
g ev
id en
ce -b
a se
d p
o li
cy in
cl u
d e
th e
U .S
. C
en su
s, th
e S
ta te
L o
n g it
u d
in a l
D a ta
S y st
em s,
in d
iv id
u a l
a n
d co
rp o
ra te
ta x
d a ta
, a n
d u
n em
p lo
y m
en t
in su
ra n
ce w
a g e
d a ta
A d
m in
is tr
a ti
ve d
a ta
se ts
co n
ta in
ex te
n si
ve in
fo rm
a ti
o n
o n
d em
o g ra
p h
ic ch
a ra
ct er
is ti
cs ,
ea rn
in g s,
co n
su m
p ti
o n
ex p
en d
it u
re s,
a rr
es ts
, in
ca rc
er a ti
o n
, h
o u
si n
g , ed
u ca
ti o
n ,
a n
d o
th er
in fo
rm a ti
o n
E vi
d en
ce -b
a se
d p
o li
cy m
a k
in g
co m
m is
si o
n
2 0
1 6
u n
d er
a b
ip a rt
is a n
vo te
o f
C o
n g re
ss S
p ea
k er
P a u
l R
y a n
o f
th e
U .S
. H
o u
se o
f R
ep re
se n
ta ti
ve a n
d U
.S .
S en
a to
r P
a tt
y M
u rr
a y
C o
m m
is si
o n
m em
b er
s h
a ve
b ee
n a p
p o
in te
d (1
5 m
em b
er s)
a n
d th
e C
o m
m is
si o
n is
n o
w m
ee ti
n g
a n
d h
o ld
in g
h ea
ri n
g s.
T h
e C
o m
m is
si o
n re
p o
rt ,
w h
ic h
w il
l a d
d re
ss p
ri va
cy is
su es
, a cc
es s
to fe
d er
a l,
st a te
, a n
d p
er h
a p
s p
ri va
te a d
m in
is tr
a ti
ve d
a ta
se ts
, li
n k
in g
d a ta
se ts
, a n
d th
e p
o ss
ib le
cr ea
ti o
n o
f a
fe d
er a l
cl ea
ri n
g h
o u
se co
n ta
in in
g a d
m in
is tr
a ti
ve d
a ta
se ts
, is
d u
e in
S ep
te m
b er
2 0
1 7
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
280 / Presidential Address: Making Federal Social Programs Work
T a b
le 1
. C
o n
ti n
u ed
.
B u
il d
in g
b lo
ck A
p p
ro x im
a te
y ea
r st
a rt
ed O
ri g in
a to
r o
r ex
a m
p le
s D
es cr
ip ti
o n
L ea
d er
sh ip
b y
O ff
ic e
o f
M a n
a g em
en t
a n
d B
u d
g et
(O M
B )
O M
B h
a s
b ee
n a
le a d
er in
m o
vi n
g a n
ev id
en ce
-b a se
d a g en
d a
si n
ce a t
le a st
th e
ea rl
y 2
0 0
0 s
R o
b er
t S
h ea
, P
et er
O rs
za g ,
R o
b er
t G
o rd
o n
, Je
ff L
ie b
m a n
, K
a th
y S
ta ck
, a n
d m
a n
y o
th er
s
O M
B ’s
se n
io r
o ff
ic ia
ls h
a ve
a lw
a y s
b ee
n le
a d
er s
in u
si n
g ri
g o
ro u
s ev
a lu
a ti
o n
s o
f so
ci a l
p ro
g ra
m s
to im
p ro
ve th
e im
p a ct
s o
f th
e n
a ti
o n
’s g ra
n t
p ro
g ra
m s.
S in
ce O
M B
im p
le m
en te
d th
e P
ro g ra
m A
ss es
sm en
t R
a ti
n g
T o
o l
(P A
R T
) in
th e
B u
sh a d
m in
is tr
a ti
o n
, O
M B
h a s
p la
y ed
th e
m o
st im
p o
rt a n
t ro
le in
le a d
in g
th e
ev id
en ce
-b a se
d m
o ve
m en
t in
th e
a d
m in
is tr
a ti
ve b
ra n
ch o
f th
e fe
d er
a l
g o
ve rn
m en
t. A
m o
n g
o th
er th
in g s,
O M
B p
u ts
p re
ss u
re o
n ex
ec u
ti ve
a g en
ci es
to ev
a lu
a te
th ei
r p
ro g ra
m s
a n
d u
se th
e re
su lt
s to
im p
ro ve
p ro
g ra
m s
o r
re a ll
o ca
te fu
n d
s In
st it
u te
o f
E d
u ca
ti o
n S
ci en
ce s
(I E
S )
2 0
0 2
U .S
. C
o n
g re
ss (G
ro ve
r “R
u ss
” W
h it
eh u
rs t,
fi rs
t d
ir ec
to r)
IE S
h a s
re vo
lu ti
o n
iz ed
th e
fi el
d o
f ed
u ca
ti o
n a l
re se
a rc
h b
y fu
n d
in g
ra n
d o
m a ss
ig n
m en
t st
u d
ie s
to d
et er
m in
e w
h et
h er
sp ec
if ic
ed u
ca ti
o n
p ro
g ra
m s
a n
d p
ra ct
ic es
w o
rk .
T h
ey a ls
o cr
ea te
d th
e W
h a t
W o
rk s
C le
a ri
n g h
o u
se ,
o n
e o
f th
e fi
rs t
cl ea
ri n
g h
o u
se s
th a t
p ro
vi d
ed in
fo rm
a ti
o n
a b
o u
t th
e su
cc es
s (o
r n
o t)
o f
ed u
ca ti
o n
a l
p ro
g ra
m s
a n
d p
ra ct
ic es
E va
lu a ti
o n
b y
fe d
er a l
a g en
ci es
In re
ce n
t y ea
rs ,
m o
st a g en
ci es
th a t
a d
m in
is te
r so
ci a l
p ro
g ra
m s
h a ve
h a d
so m
e m
o n
ey fo
r ev
a lu
a ti
o n
. B
u t
fu n
d s
fo r
ev a lu
a ti
o n
a re
g en
er a ll
y li
m it
ed
T h
er e
is n
o sp
ec if
ic p
er so
n o
r a d
m in
is tr
a ti
ve p
o si
ti o
n in
ch a rg
e o
f ev
a lu
a ti
o n
o f
so ci
a l
p ro
g ra
m s
b y
a ll
fe d
er a l
a g en
ci es
, a lt
h o
u g h
O M
B en
co u
ra g es
a g en
ci es
to ev
a lu
a te
th ei
r p
ro g ra
m s
A fe
w a u
th o
ri za
ti o
n o
r a p
p ro
p ri
a ti
o n
s b
il ls
o f
so ci
a l
p ro
g ra
m s
re q
u ir
e ev
a lu
a ti
o n
s a n
d p
ro vi
d e
fu n
d s.
B u
t m
o st
fe d
er a l
so ci
a l
p ro
g ra
m s
a re
n o
t ev
a lu
a te
d a n
d a g en
ci es
th a t
a d
m in
is te
r th
e p
ro g ra
m s
h a ve
li m
it ed
fu n
d s
fo r
ev a lu
a ti
o n
. T
h e
D ep
a rt
m en
t o
f L
a b
o r
(D O
L )
a p
p ea
rs to
h a ve
b ro
k en
fr o
m th
is p
a tt
er n
b y
a p
p o
in ti
n g
a C
h ie
f E
va lu
a ti
o n
O ff
ic er
a n
d p
ro vi
d in
g h
er o
ff ic
e w
it h
re so
u rc
es .
D O
L a ls
o h
a s
a C
o n
g re
ss io
n a ll
y a u
th o
ri ze
d a m
o u
n t
se t
a si
d e
o f
u p
to 0
.7 5
% o
f th
e D
O L
b u
d g et
fo r
ev a lu
a ti
o n
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 281
T a b
le 1
. C
o n
ti n
u ed
.
B u
il d
in g
b lo
ck A
p p
ro x im
a te
y ea
r st
a rt
ed O
ri g in
a to
r o
r ex
a m
p le
s D
es cr
ip ti
o n
R es
ea rc
h a n
d ev
a lu
a ti
o n
co m
p a n
ie s
F o
u n
d ed
in 1
9 4
8 a s
a d
ef en
se re
se a rc
h o
rg a n
iz a ti
o n
, th
e R
a n
d C
o rp
o ra
ti o
n so
o n
b ra
n ch
ed o
u t
to so
ci a l
p o
li cy
re se
a rc
h a n
d b
ec a m
e o
n e
o f
th e
fi rs
t p
ri va
te so
ci a l
re se
a rc
h o
rg a n
iz a ti
o n
s in
th e
n a ti
o n
. S
in ce
th en
m a n
y p
ri va
te co
m p
a n
ie s
h a ve
b ee
n cr
ea te
d to
co n
d u
ct so
ci a l
re se
a rc
h ,
in cl
u d
in g
p ro
g ra
m ev
a lu
a ti
o n
R a n
d T
h e
m a n
y co
m p
a n
ie s
th a t
n o
w co
n d
u ct
so ci
a l
re se
a rc
h a n
d p
ro g ra
m ev
a lu
a ti
o n
h a ve
a cq
u ir
ed g re
a t
ex p
er ti
se in
p la
n n
in g
a n
d co
n d
u ct
in g
so ci
a l
ex p
er im
en ts
. T
h ey
em p
lo y
re se
a rc
h d
es ig
n a n
d st
a ti
st ic
a l
ex p
er ts
a n
d h
a ve
d ev
el o
p ed
st a ff
s th
a t
a re
ca p
a b
le o
f im
p le
m en
ti n
g p
ro g ra
m ev
a lu
a ti
o n
s a n
d d
a ta
co ll
ec ti
o n
in m
a n
y co
m m
u n
it ie
s th
ro u
g h
o u
t th
e n
a ti
o n
. S
u ch
co m
p a n
ie s,
in a d
d it
io n
to R
a n
d ,
in cl
u d
e M
D R
C ,
M a th
em a ti
c P
o li
cy R
es ea
rc h
, A
b t,
W es
ta t,
a n
d o
th er
s
J- P
a l
2 0
0 3
A b
h ij
it B
a n
er je
e, E
st h
er D
u fl
o ,
a n
d S
en d
h il
M u
ll a in
a th
a n
J- P
a l,
a ls
o ca
ll ed
th e
A b
d u
l L
a ti
f Ja
m ee
l P
o ve
rt y
A ct
io n
L a b
, is
lo ca
te d
a t
th e
M a ss
a ch
u se
tt s
In st
it u
te o
f T
ec h
n o
lo g y . It
is co
m p
o se
d o
f 1
4 3
a ff
il ia
te d
p ro
fe ss
o rs
fr o
m 4
9 u
n iv
er si
ti es
. J-
P a l
a im
s to
re d
u ce
p o
ve rt
y b
y co
n d
u ct
in g
ri g o
ro u
s ev
a lu
a ti
o n
s o
f a n
ti p
o ve
rt y
p ro
g ra
m s
a ro
u n
d th
e w
o rl
d .
J- P
a l
a ls
o co
n d
u ct
s p
o ve
rt y
o u
tr ea
ch a n
d tr
a in
in g
w o
rl d
w id
e. T
h e
J- P
a l
w eb
si te
h a s
th e
o u
tc o
m es
o f
8 0
4 ra
n d
o m
iz ed
ev a lu
a ti
o n
s co
n d
u ct
ed b
y it
s a ff
il ia
te s
in 7
3 co
u n
tr ie
s F
o u
n d
a ti
o n
su p
p o
rt D
if fi
cu lt
to d
a te
, b
u t
fo r
a t
le a st
th e
la st
tw o
d ec
a d
es ,
fo u
n d
a ti
o n
s h
a ve
b ee
n p
ro vi
d in
g su
p p
o rt
fo r
q u
a li
ty p
ro g ra
m ev
a lu
a ti
o n
s, th
e h
ea rt
o f
ev id
en ce
-b a se
d p
o li
cy
W il
li a m
T .
G ra
n t,
A n
n ie
E .
C a se
y ,
L a u
ra a n
d Jo
h n
A rn
o ld
, P
ew T
ru st
s, Jo
h n
D .
a n
d K
a th
er in
e T
. M
a cA
rt h
u r,
E d
n a
M cC
o n
n el
l C
la rk
, a n
d m
a n
y o
th er
s
B ec
a u
se fo
u n
d a ti
o n
s h
a ve
fi n
a n
ce d
b o
th p
o li
cy -r
el ev
a n
t re
se a rc
h a n
d p
ro g ra
m ev
a lu
a ti
o n
s, th
ey a re
a vi
ta l
co g
in th
e ev
id en
ce -b
a se
d m
o ve
m en
t. S
o m
e fo
u n
d a ti
o n
s, su
ch a s
E d
n a
M cC
o n
n el
l C
la rk
, p
ro vi
d e
fu n
d s
fo r
o p
er a ti
n g
ex p
en se
s o
f lo
ca l
so ci
a l
in te
rv en
ti o
n p
ro g ra
m s
a n
d th
en re
q u
ir e
th em
to w
o rk
h a n
d -i
n -h
a n
d w
it h
a p
ro g ra
m ev
a lu
a to
r to
m a k
e su
re th
e lo
ca l
p ro
g ra
m is
d o
in g
ev er
y th
in g
p o
ss ib
le to
p ro
d u
ce p
ro g ra
m im
p a ct
s— a n
d to
k n
o w
w h
et h
er th
ey a re
p ro
d u
ci n
g im
p a ct
s
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
282 / Presidential Address: Making Federal Social Programs Work
Table 2. Federal grant programs that do or could fund evidence-based model programs.
� Elementary and Secondary Education Act (especially Title I)
� Higher Education Act
� Workforce Innovation and Opportunity Act (WIOA)
� Juvenile Justice and Delinquency Prevention programs
� Community Oriented Policing Services (COPS)
� Section 8 Housing Choice Vouchers and Project-Based Rental Assistance
� Substance Abuse and Mental Health Services
� Community Health Centers
� Maternal and Child Health
� Child Protection Programs (especially Titles IV-B & IV-C of the Social Security Act)
� Temporary Assistance for Needy Families (TANF)
� Community Services Block Grant
Note: These grant programs are only a fraction of all federal grant programs that could fund evidence- based model programs on a competitive basis.
the building blocks of the evidence-based movement. I reach this conclusion because federal agencies could use the tiered-evidence program to increase the effectiveness of all federal grant programs based on competitive grant awards.1 In the second part of this paper, I explore the Teen Pregnancy Prevention (TPP) program, the biggest tiered-evidence program to produce a round of evaluations to date.
Results First
Initiated in 2011 by the Pew and MacArthur Foundations, Results First is an am- bitious project for foundations, even foundations with the resources of Pew and MacArthur. The goal of Results First is to work with as many states as are interested and willing to identify state social programs addressed to a specific social problem that could be replaced by more effective programs. Thus, Results First aims to help states spend their grant funds on model programs that have strong evidence of suc- cess by identifying and terminating ineffective programs and replacing them with evidenced-based programs that have been shown to produce benefits that exceed program costs. In short, Results First is trying to do at the state and county level what the Obama administration tried to do at the federal level.
1 And current formula grant programs that distribute their funds according to a formula such as pro- portional to the poverty population residing in each state could be converted in whole or in part to competitive grant programs through the legislative process.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 283
The Pew-MacArthur team approaches states individually and offers to work with them to improve the effectiveness of their social programs. If states indicate an interest, the program is explained to them in a series of phone calls or visits that involve the Pew-MacArthur team and ideally both senior executive branch and leg- islative branch leaders in the state. If the state decides to join the initiative, they sign a letter of agreement with Pew-MacArthur that outlines the respective commit- ments of both the state and the Foundations. In order to participate, the state must have the financial commitment, the personnel and agencies that have the capacity to follow the Pew-MacArthur model, and a staff team that can lead the project’s implementation. Once the agreement has been signed, the state selects the social issue or issues they want to address. The issues selected by the state can include adult criminality or juvenile justice, child welfare, education at all levels, mental health, and substance abuse, although the Pew-MacArthur team is open to working on additional social issues.
To help the states conduct their literature review and locate programs that have evidence of success, the Pew-MacArthur team created a very useful tool by combin- ing information from eight clearinghouses that provide extensive details on model intervention programs in several areas of social policy. These eight clearinghouses will be discussed in more detail below, but here I want to emphasize that the Pew- MacArthur team had the foresight to realize how helpful these clearinghouses are for any group or individual wanting to review what is known about the impacts of model social programs in a given area of social policy. It is not difficult to imagine the hours saved by the 21 states participating in the Results First Initiative as they review the evidence on the social intervention programs being used by their own state and search for information on other model programs they could adopt. Even more important, by making it relatively easy to search for more effective programs by bringing the evidence from eight clearinghouses together in one web location, participating states are likely to do a much better job of finding the programs that best suit the goals and conditions in their own state.
Modeled in large part on the Washington State Institute for Public Policy (2016), the Results First approach with the states hinges on conducting benefit-cost analyses. More specifically, states begin by identifying all the existing state pro- grams tasked with solving the social problem they have selected. Typically, around 70 percent of the programs being used by states have no evidence of success. The states are more or less forced to confront this problem because they must conduct the literature reviews described above to determine whether research shows that their programs produce benefits that exceed costs. If not, states are encouraged to select model programs for which there is evidence that the program produces benefits that exceed costs. In some cases, states participating in the Results First initiative have faithfully implemented the Results First approach by actually ending programs that are not working and using the money saved to begin implement- ing programs that have strong evidence of success (Pew-MacArthur Results First Initiative, 2014).
Thus, the Pew-MacArthur Initiative has already demonstrated that with the right help and guidance, states can complete the politically difficult step of ending lousy programs and replacing them with programs that have achieved success as mea- sured by rigorous evaluations. Now the goal for assessment of the Pew-MacArthur results is to determine whether the new, evidence-based programs that states se- lect to replace failing programs actually improve outcomes. Improving outcomes at scale is, of course, the acid test of the entire evidence-based movement. So two cheers for now for the Pew-MacArthur Initiative, with the third cheer waiting on evi- dence that the Initiative can move the needle on the social problems states choose to attack.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
284 / Presidential Address: Making Federal Social Programs Work
Pay for Success
Pay for Success involves the traditional elements of trying to determine whether a model social program produces impacts plus a unique financing innovation. Specif- ically, an individual, business, or foundation investor puts up money to pay for the intervention. Then, if the program produces government savings, the govern- ment pays the investor back, including interest if the level of savings justifies the additional payment.
The details of the full Pay for Success approach, although they vary somewhat in practice, are straightforward. In most cases, due in large part to descriptive or demographic data, a social problem is known to be large and perhaps growing. In response to the developing social problem, there may be a body of social intervention programs that attempts to tackle the problem, usually with indifferent success. But as some promising interventions begin to appear, a potential investor may come to believe that implementing a particular intervention holds promise for successfully attacking the problem. An intermediary organization works with the investor and another organization, often a community-based entity, that has developed expertise in implementing a model program to reduce the social problem. The intermediary also selects a third organization to conduct a rigorous evaluation of the program’s success or lack thereof. The intermediary is the quarterback of the entire enterprise, working with the investor, the service provider, and the evaluator to ensure that the model program is well implemented and that the evaluation proceeds as planned. After the intervention has lasted for a period of time agreed to by all parties, the impact of the program and the determination of whether the impacts have been sufficient to save government money is made by the intermediary. If so, the investor is paid off in accord with the terms of the agreement and the agency benefitting from the intervention, as well as the individuals or families served by the intervention, can continue to reap the rewards in the future. If not, the investor loses their up-front funds, and the government is not out any funding.
From the perspective of the evidence-based movement, two features of Pay for Success are especially notable. First, the approach has the potential of expanding sources of funding for social intervention programs. Theoretically, government can- not lose. If the program does not work, it was the investor’s money that paid for the program and the investor loses their investment. If the program reduces the social problem, government saves money because of the benefits associated with the re- duction. For example, if a preschool program reduces the need for special education or grade retention, the local education agency saves money because special educa- tion programs usually cost about twice as much as regular education and it costs twice as much to repeat a grade as to take it one time. Thus, government either does not pay anything or the government pays off the investor with money that will be saved because of the program’s success. Second, a mandatory component of a well- conducted Pay for Success trial is that a rigorous evaluation determines whether the program worked and saved money. Not only does a rigorous evaluation assure that the determination of program success is accurate, but the use of rigorous evaluation strengthens respect for evaluation and demonstrates the function of evaluation as a vital component of the evidence-based movement.
According to the Pay for Success Learning Hub operated by the Nonprofit Finance Fund (2016), as of early 2016 there were 10 Pay for Success projects that had been launched in the United States, one that has been completed, and over 50 in the pipeline. Federal programs that have funds for Pay for Success projects include the Social Innovation Fund operated by the Corporation for National and Community Service (CNCS), the Veterans Administration in cooperation with CNCS, and several programs in the Department of Education (Lester, 2016). Some of the programs now in the pipeline have private investors, but it is not clear yet whether there will be an
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 285
abundance of private investors ready and willing to put up private funds to pay for social intervention programs on the chance that they turn a profit.
And here, a fascinating point made about Pay for Success by George Overholser (2015) of Third Sector Capital Partners comes into play. Overholser argues that the bulk of the money to pay for Pay for Success programs will not be put up by private investors looking for a profit. True, investment dollars will often come from private investors to handle program financing during the time between conceptualizing the intervention and the time findings are released after a year or more. At that point, if the program is successful, the entity with normal responsibility for the program will begin paying for it—and gladly because it will be saving money, probably over a long period of time. Meanwhile, the investment dollars can be recycled and used to support a new innovative program. As Overholser (2015) points out, “This recycling phenomenon is what makes it possible for a small amount of private loan capital to catalyze large amounts of government PFS [Pay for Success] payments” (p. 37).
Behavioral Economics Programs2
In their 2008 book Nudge: Improving Decisions about Health, Wealth, and Happiness, Richard Thaler and Cass Sunstein argue that humans are not always rational and explain many ways this fact can be exploited for the public good. The public good lies in the work of understanding the actual ways humans think and make decisions, especially under uncertainty, which in turn can be used to “nudge” us toward wiser decisions.
There is a line between encouraging wise decisions and thought control, the latter of which is implicated in the usual criticism of government functioning as big brother. This value-laden debate about whether it is appropriate for government to try to subtly influence people’s choices is too complex to discuss in any detail here. Even so, I simply observe that the research on the impacts of nudge-type programs in promoting better decisions by typical adolescents and adults and by government employees convinces me that the modest risk of inappropriate shaping of behavior by government is a small price to pay for the desirable outcomes for both individuals and government that nudge programs have been shown to produce.
A good way to gain an idea of what these programs are like and whether they can lead to improvements in decisionmaking is to examine the work of the White House Social and Behavioral Sciences Team (SBST). Started in 2014, the unit conducts experiments founded on behavioral principles, largely randomized controlled trials (RCTs), in cooperation with executive agencies including but not limited to the Departments of Agriculture, Education, Labor, Justice, Commerce, and Defense. In their two annual reports, the SBST has published the results of 40 experiments conducted with these agencies on issues ranging from how to promote government efficiency to how to convince members of the military to contribute to retirement savings accounts (SBST, 2015, 2016).
Consider the series of experiments the SBST has undertaken with the Department of Defense (DOD) to increase enrollment by service members in the federal govern- ment’s Thrift Savings Plan (TSP). In 2016, all military personnel who reported to a new military base (which happened about 640,000 times) were required to place
2 The background of how the branch of psychology that provided the founding ideas that led to behav- ioral economics was conceptualized by two Israelis, Daniel Kahneman (who eventually won the Nobel Prize) and Amos Tversky. Their story, and especially their work, is well told by Michael Lewis (2017). Kahneman’s work with Richard Thaler during the mid-1980s as Thaler was developing the founding ideas and experiments that were a major part of the origin of behavioral economics is covered in some detail by the biographical essay Kahneman (2002) wrote for the Nobel Prize Committee.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
286 / Presidential Address: Making Federal Social Programs Work
a “yes” or a “no” on a form to indicate whether they wanted to sign up for TSP. This modest intervention on forced choice resulted in an 8.3 percent increase in enrollment as compared with the procedure that did not require the “yes” or “no” box to be completed. As a result of this and previous experiments, the DOD will begin to make enrollment in TSP the default choice for all new service members. Under the default procedure, service members will be automatically enrolled in the TSP program unless they indicate on the form that they don’t want to participate by checking a box. Previous research involving other programs shows that the default procedure leads to substantial increases in enrollment.
The promise of the behavioral approach to basing policy on evidence is shown by the rapid spread of behavioral research, not only to the White House, but to prominent research and evaluation organizations as well. MDRC, for exam- ple, a widely respected evaluation company, has created a behavioral science re- search unit and has initiated a series of behavioral experiments with funding from the Department of Health and Human Services (HHS) and several founda- tions. In one intervention, initiated in 2010, MDRC launched 15 RCTs involving about 100,000 participants in seven states. According to MDRC (n.d.) the “interven- tions improved child care subsidy renewal rates and the use of quality-rated child care, boosted requests for child support modifications and payment consistency, and improved engagement in welfare and other social service appointments and activities.”
As these examples show, the behavioral/nudge approach to evidence-based policy has shown that it can have impacts on a broad range of behaviors. Moreover, the approach is especially welcome because most behavioral solutions are inexpensive. But as we have seen, a primary criticism of the approach is that government is deciding what people should do and using subtle messages and cues that interfere with free choice. As Michael Thomas, an economist at Utah State University, told Fox News in 2013, “nudging . . . assumes a small group of people in government know better about choices than the individuals making them” (Lott, 2013). So far, this argument against nudge programs has failed to disrupt their growth. Moreover, given the successes of nudge programs, many of which have been shown to save government money, it might be expected that nudge programs will continue to grow and find application in more and more settings.
Clearinghouses
One of the most useful tools that supports the evidence-based movement is clear- inghouses. An evidence-based clearinghouse is a website that contains extensive information about social intervention programs and especially their evaluations. Perhaps the most valuable function of these clearinghouses is to provide abundant information about programs and practices that have been evaluated, the results of those evaluations, and often a rating of how effective (or not) the program or practice has been shown to be by the evaluations. This information is well-suited to use by community groups and individuals (often social entrepreneurs or scholars) giving advice to these groups who are trying to identify programs they can adopt to fight a social problem at the local level.
There are nine clearinghouses, two of which are now inactive (but still available), that feature this type of information. The clearinghouses cover a range of social issues, including education, teen pregnancy, child abuse and neglect, drug abuse, youth development, and crime and delinquency. The Pew Foundation, as part of its Results First Initiative (see above), created a clearinghouse that contains the contents of the other eight clearinghouses—providing a new definition of one-stop shopping. The Pew-MacArthur team wanted to make use of clearinghouses as easy
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 287
as possible for the 21 states they are working with, but now their comprehensive clearinghouse is available to everyone.
A caution is in order about use of these clearinghouses. As I will argue below, an important goal of the tiered-evidence strategy is to tighten the criteria now used by HHS to identify evidence-based model programs. Social intervention programs that earn the title evidence-based should have a high probability of replicating, but as we will see they often do not replicate. In an evolving and young field such as evidence-based policy, the definition of evidence based is bound to fluctuate. Yet, the rigor with which the various clearinghouses consider programs to be evidence based (not all the clearinghouses use this term) appears to vary substantially. The criteria and procedures used by the Coalition for Evidence-Based Policy, one of the nine clearinghouses addressed to social policy, appear to be demanding and lead to designating relatively few programs as evidence based. By contrast, the criteria and procedures used by the National Registry of Evidence-based Programs and Practices run by the Substance Abuse and Mental Health Services Administration seem more generous and result in more programs and practices being determined to be “effective,” the National Registry’s highest rating. Hopefully, especially by using high probability of replicating as a criteria, the clearinghouses will evolve toward more stringent definitions of evidence based.
Nonprofit Interest Groups
Although we want to think about evidence-based policy as being a scientific issue above the political fray, the issue is in fact pursued in a highly politicized envi- ronment in both the nation’s capital and in many state capitals. In such an envi- ronment, a movement needs friends with money and influence. Fortunately, there are a host of organizations in the nonprofit world, and even a few in the for-profit world, that actively work to advance the cause of the evidence-based movement. The functions performed by these groups include lobbying for legislation that promotes evidence-based policy, providing funding for a host of activities related to promot- ing evidence-based policy, assisting with communications, and conducting analysis and research.
Even if the environment in the various capitals were less politicized, there are many cases in which advancing the evidence-based movement requires legislation. Legislation in turn requires good, practical ideas that can be translated to legislative language that will successfully make its way through congress and whatever presi- dential administration is in power at the time. These are all tasks that require great expertise, experience with the legislative process, and political connections. Wealthy private companies pay millions of dollars to buy help of this type—and often their purpose is only to stop new legislation or regulations—an easier task than passing new legislation or ensuring the administration writes a favorable regulation once legislation is passed.
An interesting example will provide a clear idea of how nonprofit interest groups are helping to promote the evidence-based movement. The example is provided by the Coalition for Evidence-Based Policy, founded in Washington, DC, by the lawyer Jon Baron in 2001.3 Having watched Baron and the Coalition operate on many occasions, and having participated myself in some of Baron’s plans for influenc- ing legislation, I judge the Coalition to have been one of the most important and influential of the nonprofits that put the evidence-based movement on the map in
3 The Coalition has now disbanded because Baron and his staff elected to join the Laura and John Arnold Foundation. However, the Coalition website is still active (http://coalition4evidence.org/).
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
288 / Presidential Address: Making Federal Social Programs Work
the early-to-mid 2000s. Two of the many important pieces of legislation the Coali- tion helped shape were the Investing in Innovation Fund and the Social Innovation Fund, both tiered-evidence initiatives pushed through Congress by the Obama ad- ministration in 2009 (Haskins & Margolis, 2015, chapters 4 and 5). Baron and the Coalition provided advice to the Office of Management and Budget officials working on the legislation both in-person and in writing and even provided draft legislative language as the legislation was being written. The Coalition for Evidence- Based Policy (2009) also released a detailed analysis about how tiered legislation should be written and why it was important. Some of the language in the Coali- tion’s document looked like it came off the same word processor as the one used to write the legislation. Given the importance of the tiered-evidence strategy to the evidence-based movement (see next), the Coalition can be said to have played a di- rect role in writing the legislation that authorized the first tiered-evidence programs in 2009 and thereby provided an immense boost to the evidence-based movement. This example demonstrates vividly how nonprofits serve a major role in analyzing policy issues, developing legislative ideas based on the analysis, and having the con- tacts and credibility with key officials to get their analysis translated into public policy.
Administrative Data
Using administrative data for research and evaluation has many advantages as com- pared with data that are collected for a given study. Administrative data are often available on entire populations; attrition and nonresponse are usually not problems; sample sizes are often enormous; and costs are usually minimal because the data already exist.
In addition to these characteristics, the nation is blessed with a host of adminis- trative data sets containing data on many issues. Executive agencies at both the state and national level have data on their benefit programs, sometimes over a period of decades on entire populations of recipients. These data sets can allow researchers, evaluators, and managers to have access to the types and amount of benefits indi- viduals receive over a period of years. The Census Bureau has data not only from annual censuses of population samples, but has results from the decennial cen- suses, which contain extensive data on nearly the entire population. The Internal Revenue Service has extensive data on income for both individuals and businesses. The National Directory of New Hires contains wage and unemployment insurance data from every state on most individuals who have a job. This list could be greatly expanded.
Thus, Table 3 displays a selected list of many of the federal data sets that could prove useful to researchers and program managers. These and many similar data sets controlled by government agencies at the state and federal levels have useful information that could be used to perform an array of research and evaluation func- tions. However, there are many barriers to access to these data for research and management purposes. Perhaps the most important is privacy concerns. The U.S. has a long tradition of privacy that applies to administrative records that contain in- formation about individual Americans and businesses. It seems wise for the research and evaluation communities to carefully observe the American tradition of respect for privacy and nondisclosure of private information. Fortunately, modern technol- ogy permits the deidentification of administrative records so that researchers and managers could have access to the data in specific administrative data systems and even to link systems without being able to identify individuals.
Of the many examples of research using administrative records that could be cited as examples to demonstrate how important administrative records can be, the work
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 289
Table 3. Examples of federal administrative data sets.
Data set name Controlling agency Description
National Vital Statistics System
Centers for Disease Control and Prevention
Births, deaths, marriages, divorces, and fetal deaths
Fiscal Year Disability Claim Data
Social Security Administration
Beneficiary claim and receipt information
National Incident Based Reporting System (NIBRS)
National Archive of Criminal Justice Data, a part of the Inter- University Consortium for Political and Social Research (ICPSR) at the University of Michigan
Types of offenses, victim and offender characteristics, and types and value of property stolen
Identity History Summary
Federal Bureau of Investigation
Date of arrest, arrest charge, and disposition of the arrest
American Community Survey
United States Census Bureau
National demographic, housing, economic, and social information
Decennial Census United States Census Bureau
Sex, age, race, relationship, and housing information of all persons living in the United States
Survey of Income and Program Participation (SIPP)
United States Census Bureau
Economic well-being, family dynamics, education, assets, health insurance, childcare, food security, and use of public benefit programs
State Longitudinal Data Systems (K-12 student records)
United States Department of Education
School enrollment history, demographic characteristics, program participation record, and test scores of every student
Individual and corporate tax data
Internal Revenue Service Income for all businesses and households in the United States
National Directory of New Hires (NDNH)
United States Department of Health & Human Services, Office of Child Support Enforcement
Employment information, unemployment insurance information, and quarterly wage data from every state
State-Level Unemployment Insurance Wage and Contribution Reports
Maintained by state-level unemployment insurance programs
Unemployment insurance payments and individual wage records
Longitudinal Employer–Household Dynamics Program
United States Census Bureau and States
Earnings history by job level, living and work arrangements of American workers, and firm characteristics
American Housing Survey
United States Census Bureau
Characteristics of American housing units and householders
GDP and the National Income and Product Account (NIPA) Historical Tables
United States Department of Commerce, Bureau of Economic Analysis
Macroeconomic statistics including GDP, national income, corporate profits, government receipts and expenditures, as well as personal income, savings, and consumption expenditures
Immigration Data & Statistics
Department of Homeland Security, Office of Immigration Statistics
Immigration enforcement actions (arrests, removals, and returns) and country of origin and demographic information for lawful permanent residents, refugees, naturalized persons, and apprehended aliens
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
290 / Presidential Address: Making Federal Social Programs Work
of Raj Chetty of Stanford and his team shows how enticing the possibilities are. Chetty worked with the Statistics of Income (SOI) Division at the Internal Revenue Service and gained access to income tax records. In one study, based on the earn- ings of American parents and their children from tax records, Chetty et al. (2016) traced the percentage of children whose income at age 30 was greater than their parents’ income at age 30 by year of birth. The analysis shows that the percent- age of children whose income exceeded that of their parents declined from over 90 percent for children born in 1940 to 50 percent for children born in 1985. These data justify the Chetty conclusion that the American Dream is “fading.” In another study, Chetty and his Harvard colleagues Nathaniel Hendren and Larry Katz (2015) examined the connection between the neighborhoods in which children grow up and their earnings as adults. Again using SOI tax data to obtain earnings, Chetty and Hendren examine more than five million families who moved to new neigh- borhoods when their children were at various ages. Using the adult earnings of children who lived permanently in particular neighborhoods when they were chil- dren as the yardstick of neighborhood quality, they find that children who moved to better neighborhoods as children had higher earnings as adults than children who moved to worse neighborhoods. The study also showed that the younger chil- dren were when they moved to better neighborhoods, the greater their gain in adult earnings. The findings on the positive effects of moving to better neighborhoods are not based on a random assignment study. But the suggestion that young children moving to better neighborhoods leads to better development and better outcomes in adulthood as measured by income provides a good suggestion for a study that could establish causation. Suggesting such causal studies can be one of the great advantages of using administrative data.
The Evidence-Based Policymaking Commission
If these Chetty studies are examples, increased access to and use of administrative data sets will have bountiful applications to evidence-based policy. Thus, it is not surprising that Republican Speaker Paul Ryan teamed up with Democratic Senator Patty Murray to establish a Commission on Evidence-Based Policymaking. Their charge to the Commission is to explore ways to use federal (and other) administra- tive data sets for use in research and program evaluation and perhaps to establish a clearinghouse that would facilitate use of the administrative data sets by researchers while ensuring the security and privacy of the data. The 15-member Commission has now been appointed and will issue its report in September 2017. Early indications are that the Commission is intent on reducing the barriers on access to admin- istrative data sets contingent on ensuring privacy, thereby promoting the kind of research epitomized by the work of Chetty and his team.
An equally important use of administrative data is to obtain outcome information on the long-term impacts of social intervention programs. We have already seen that it is expensive and difficult to follow over the long term people who partic- ipate in program evaluations. Attrition is an especially serious problem and can undermine the validity of any evaluation. But the data in administrative data sets— including educational attainment (years of schooling completed) and achievement (test scores), employment and unemployment, earnings, marital history, health in- cluding mental illness, and many others—can provide a robust look at the im- pacts of social intervention programs. If the Commission’s likely recommendations are implemented, and the implementation results in easier and more frequent ac- cess to federal and state data sets, we can expect that program evaluations will be able to provide more information to researchers and policymakers than in the past.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 291
TIERED EVIDENCE: THE PROSPECTS FOR SCALING UP SUCCESSFUL PROGRAMS
These eight examples, especially when combined with several additional exam- ples in Table 1, show that the evidence-based policy movement is healthy and growing. If all of these elements of the evidence-based movement continue to grow and prove effective, the evidence-based movement will have a good chance of leading to improved social policy. But to truly achieve success, the move- ment must begin to show that it can have impacts on the nation’s social prob- lems. The indicator of this success would be measurable impact on the nation’s most serious social problems, such as school dropout, teen pregnancy, low lev- els of post-secondary education, inequality, and so forth. What is wanted now is a close look at the way evidence-based policy can address national social problems.
In this section, I argue that the federal government is already building a path by which evidence-based policy can begin to exert the kind of impact that has the potential to make substantial progress against the nation’s leading social problems. I refer to the tiered-evidence strategy, which is reviewed in Haskins & Margolis (2015).
For anyone interested in the evidence-based movement, the July 2016 release of evaluations of the 41 TPP programs by the Office of Adolescent Health (OAH) was a major event.4 In accord with the tiered grant-making strategy, the release consisted of two types of evaluations. Tier 1 evaluations were replications of programs that had been shown by one or more rigorous evaluations to produce positive impacts on various measures of adolescent sexual behavior or pregnancy itself; Tier 2 was composed of programs that were judged to show promise but which had not yet been shown to produce impacts by a rigorous evaluation. To have 41 high-quality evaluations made public at one time is downright inspiring to those who are look- ing to the evidence-based movement as the way to revolutionize federal use of grant funds by spending them primarily on evidence-based programs. Moreover, 21 of the evaluations are featured in the American Journal of Public Health, a highly regarded scholarly journal, along with commentary and editorials about the specific evalu- ations and the general TPP program (Morabia, 2016). The entire process of TPP grant making, project implementation, project evaluation, and project communica- tion with researchers, policymakers, and the public provides a model for the future of federal grant making.
Two features of the way OAH is handling the release of evaluation information is especially impressive and hopefully will be followed by federal and state grant- making agencies in the future. First, the huge volume of information released about the evaluations is useful for understanding how a wide range of TPP programs is working. In addition to elaborate evaluation reports on several of the individual projects (Rotz et al., 2016), the July release consists of numerous documents that describe the evaluations and their results. Two especially useful documents prepared by OAH (2016d) include one that summarizes every study in Tier 1 and a second that summarizes every study in Tier 2. In addition, the extensive treatment of the original evaluation studies and the importance of the evaluations in the special issue of the American Journal of Public Health is a useful way of communicating with scholars, advocates, program operators, policymakers, and the public. The release, in other words, is a model of transparency and gives interested parties both a thorough report on many of the projects as well as helpful summary documents. The second admirable feature of the release is that, although OAH emphasizes the
4 Three of the 41 evaluations were late and were released to the public after the initial OAH July 2016 evaluation release and are included in this analysis of results.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
292 / Presidential Address: Making Federal Social Programs Work
Table 4. Reasons an evaluation was ruled inconclusive by the Office of Adolescent Health.
� High level of attrition
� Poor contrast between the experiences of experimental and control groups
� Children were too young to have had sexual experience before the end of the project
� Poor implementation quality
� Analysis or reporting did not meet research quality standards defined by the Department of Health and Human Services
◦ No correction for attrition
◦ Did not have baseline equivalence between experimental and control groups
◦ Matching procedure was not strong
positive outcomes by pulling them together in a separate document and in their press release, OAH (2016a) does not make extravagant claims for the impacts of either Tier of TPP programs, leaving it to readers to make their own determination of whether the outcomes are encouraging. OAH shows itself to be an effective and transparent government agency, not a cheerleader.
There were 19 evaluations of Tier 1 projects and 22 evaluations of Tier 2 projects. The evaluations measured one or more of seven outcomes, although all projects did not measure all outcomes. The seven outcomes were sexual initiation, recent sexual activity, number of sexual partners, frequency of sexual activity, contraception, sexually transmitted diseases (including HIV), and pregnancy (a few projects also measured births).
Unfortunately, a number of the evaluations had problems that caused OAH to judge them to be “an invalid test of the program” because, for example, they had high levels of attrition or because participating children were too young to have had sexual experience (Farb & Margolis, 2016), even by the end of the project (see Table 4 for all the reasons). Seven projects were assigned to the “invalid test” category from Tier 1 and six from Tier 2 (Table 4).
The review of results across the Tier 1 and Tier 2 projects begins with a general measure of the number and share of projects that produced at least one statistically significant impact on any of the seven measures of sexual behavior. This review is followed by a summary of impacts by Tier 1 and Tier 2 projects on each of the seven individual measures of sexual behavior.
In determining the share of Tier 1 and Tier 2 programs that had at least one impact, an issue arises of whether the invalid projects should be included in the analysis. Whether these projects are included in the denominator in calculating the share of programs that produced at least one significant impact makes a difference in our conclusions about how successful the Tier 1 and Tier 2 programs were. Using the total number of 19 projects in Tier 1 as the denominator in the calculation of the share of projects that had at least one impact, reveals that only 21 percent of the projects produced at least one impact (Table 5; top panel). The same calculation for the 22 projects in Tier 2 shows that 36 percent of the projects produced at least
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 293
Table 5. Summary of results for Tiers 1 and 2 of teen pregnancy prevention programs.
Tier
Category 1 2
Projects producing one or more impacts Number of projects evaluated 19 22 Number of evaluations that met standards 12 16 At least one impact 4 8
As a percent of all evaluations 21 36 As a percent of all evaluations that met OAH standards
33 50
Impacts on sex-related behaviors Impact on
Sexually transmitted infection 0/1 (0%) 0/1 (0%) Frequency of sexual activity 0/2 (0%) 0/0 (0%) Number of sexual partners 0/1 (0%) 0/5 (0%) Recent sexual activity 1/7 (14%) 1/11 (9%) Sexual initiation/abstinence 2/13 (15%) 3/15 (20%) Contraceptive use and/or consistency (including condoms)
1/14 (7%) 6/18 (33%)
Impacts on teen pregnancy Pregnancy 1/4 (25%) 4/8 (50%)
Source: OAH (2016d).
one impact. If we eliminate from the calculation the projects that OAH judged to be invalid, there are 12 rather than 19 Tier 1 and 16 rather than 22 Tier 2 projects. These smaller denominators in turn produce the results that four of 12 (33 percent) Tier 1 projects reported at least one positive impact on a behavior related to sex and eight of 16 (50 percent) Tier 2 projects reported at least one positive impact (Table 5, top panel).
Turning to the seven individual measures of sexual behavior, Table 5 (second panel) reports each measure for all the projects that reported that measure. One measure, acquiring a sexually transmitted disease, was collected by only one Tier 1 and one Tier 2 project; neither of the two projects had an impact on this measure. The frequency of sexual activity and the number of sexual partners were also collected by very few projects. Only two Tier 1 projects and zero Tier 2 projects collected data on frequency of sexual activity. Neither of the Tier 1 projects that reported on this measure found significant impacts. Similarly, only one Tier 1 project and five Tier 2 projects collected data on the number of sexual partners. No impacts on this measure were detected in any of these six projects.
Relatively more projects collected data on the last four measures reported in the second panel of Table 5.5 Recent sexual activity was collected by seven Tier 1 projects, only one (14 percent) of which had a significant positive impact;
5 Note that outcomes are tabulated on the basis of individual evaluations rather than on programs grouped by program model. Evaluating the body of evidence by program model yields a more complex picture. While Tier 2 programs are each a unique program model by definition, Tier 1 grantees used only 10 different model programs (from the HHS list of 44 evidence-based programs), four of which produced an impact in at least one instance. For example, the Teen Outreach Program (TOP) produced a positive impact on pregnancy in one evaluation, a positive impact on pregnancy for boys and a negative impact on pregnancy for girls in another, and no impact in five more evaluations.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
294 / Presidential Address: Making Federal Social Programs Work
eleven Tier 2 projects collected this measure and again only one (9 percent) produced a significant positive impact. Twenty-eight projects collected data on sexual initiation. Of the Tier 1 projects that collected this measure, two of 13 (15 percent) produced a positive impact while three of 15 (20 percent) Tier 2 projects produced a positive impact. On the very important measure of contra- ceptive use, one of 14 (7 percent) Tier 1 projects produced a significant pos- itive impact while six of 18 (33 percent) Tier 2 projects produced a positive impact.
Finally, perhaps the most important measure of success for TPP projects is a re- duction in teen pregnancy rates (Table 5, bottom panel). Only a minority of projects collected this measure, mostly because teens in both the experimental and control groups of most of the projects were so young that many kids in both groups had not yet started having sex or even engaging in sex-related behaviors. In these cases, tracking pregnancy would have required a longer follow-up than the scope of the grant period allowed. Even so, one of four (25 percent) projects in Tier 1 that re- ported this measure produced a positive impact while four of eight (50 percent) Tier 2 projects produced a significant impact. In all but one of these cases, the effect had faded by the 12 month follow-up.
WHAT NOW?
It is likely that some observers, including members of Congress and their staffers, will conclude that the results from the first wave of TPP programs are only modestly encouraging. Of the 41 projects that reported results, under the approach that ex- cludes “invalid” tests from the analysis, about a third of the evidence-based projects in Tier 1 were replicated by producing at least one positive impact on a behavior related to teen sex. Under the less generous approach that counts all the evalua- tions, only a fifth of the Tier 1 projects produced a positive impact. The results in Tier 2 were more encouraging—50 percent under the more generous approach and 36 percent under the more stringent approach produced at least one positive impact. Similarly, results for the individual measures were also mixed. In only three of the 14 cases did more than 20 percent of projects show positive impacts. In six of the 14 cases, all with very few projects reporting, no project produced a positive impact. On the other hand, the share of projects in Tier 2 reporting positive impacts on use of birth control and the share of projects in both Tiers 1 and 2 reporting positive impacts on pregnancy reduction—the two most important measures—ranged from 25 percent to 50 percent. As I argue below, it seems wise not to give an intervention program too much credit if it has only one significant impact.
In pondering these impacts, I raise six issues that are important, not just for reflecting on this remarkable set of evaluations, but also for a better understanding of the tiered-evidence strategy. I examine these issues and make recommendations because the tiered-evidence strategy has the potential to make the nation’s social policy more successful.
How Successful Are the Nation’s Social Programs?
The first issue is that rigorous evaluations of social programs show that most of them do not work. Head Start, one of the most celebrated social programs of the War on Poverty, was thought for years to produce sizeable impacts. But when Head Start was finally subjected to a random-assignment evaluation by order of Congress, it was shown to produce only modest impacts at the end of the preschool year
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 295
and no pattern of significant impacts thereafter (Puma et al., 2012).6 In fact, Jon Baron and Isabel Sawhill (2010) reviewed the evidence from RCTs on ten large-scale and popular federal social programs such as Head Start, Upward Bound, and 21st Century Learning Centers, which were widely believed to produce positive impacts. Their review found that in nine of the 10 cases, evaluations showed modest or no impacts. The same pattern occurs in fields other than social science, such as business and medicine. Jim Manzi (2012), in his book on use of RCTs in the medical, social science, and business sectors, reported that about 80 to 90 percent of the evaluations in all three sectors find null effects. The record of null effects from high-quality evaluations of intervention programs shows definitively that most interventions do not produce statistically significant impacts. Thus, the TPP results should not be surprising.
In fact, using the percentage of programs that produced at least one significant impact as the criterion, it could be argued that the TPP programs as a network were more successful than most social programs funded by government. If the typ- ical rate of successful programs for most interventions is 10 to 20 percent, under both approaches to counting the share of successful Tier 1 and Tier 2 programs that produced at least one significant impact, the level of projects with successful impacts was above 20 percent in all four cases, although the Tier 1 success rate under the less generous measure was only 21 percent. Based on calculations that drop the projects that OAH deemed to be invalid tests, the rates of success in Tier 1 and Tier 2 were 33 and 50 percent, respectively, both much higher than would be expected from previous rigorous evaluations of social programs. This way of think- ing about the results of the TPP evaluations is that the tiered-evidence approach is promising.
Role of Federal Agencies
Role as Teacher and Guide
Federal agencies, which must play an increasingly important role in creating suc- cessful grant programs, should learn from the OAH administration of TPP. One lesson is that even awarding federal funds and providing free advice and technical assistance with issues such as recruiting participants, implementing a curriculum, training, and conducting an evaluation do not guarantee uniformly high quality. It follows that it would be ideal if federal agencies could develop and hone their ability to select grantees that are relatively more likely to avoid the problems that lead to so many invalid tests in the TPP programs. Further, given the fine performance of OAH in administering the TPP projects, it is possible to imagine that federal agencies can develop expertise in helping grantees adopt effective model programs that produce above average rates of success. This is especially the case since federal agencies are at the beginning of administering evidence-based programs and can be expected to learn from experience the lessons of helping programs produce impacts.
Role in Improving Evaluations
Another important role for federal agencies is to become expert in program evalu- ation and to bring consistency and rigor to the evaluation of projects under their
6 However, other less rigorous studies have reported long-term impacts of Head Start on educational attainment at the secondary and postsecondary level, as well as on earnings and crime reduction (Deming, 2009; Garces, Thomas, & Currie, 2002).
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
296 / Presidential Address: Making Federal Social Programs Work
jurisdiction.7 Following the lead of the Department of Labor, all federal agencies that have jurisdiction over social programs should have a chief evaluation officer or the equivalent with responsibility for ensuring that most projects overseen by the agency receive high-quality evaluations on a routine basis. Agencies should identify and define a small set of outcomes, usually around four or five, and require every project to collect data on all those outcomes. Beyond identifying and defining the outcome variables, agencies should also impose commonality on how the variables are measured. Reporting on these measures should be a condition of giving a federal grant to state and local projects.
Toughening the Requirements to Be Designated Evidence Based
The federal government and the entire field of evidence-based policy should reflect on how model programs addressed to solving social problems in education, teen pregnancy, parenting, delinquency, and so forth earn the title of “evidence based.” Taken as a whole, the TPP evaluations are the biggest test to date of whether pro- grams judged by HHS to be evidence based will successfully replicate. The fact that in the TPP evaluations the previously untested innovative programs in Tier 2 were as likely or more likely to achieve positive impacts as the evidence-based programs in Tier 1 is prima facie evidence that the criteria now used for determining that a program is evidence based bear reexamination.8 Some of the model programs determined by HHS to be evidence based are a decade or two old and have not been replicated for many years—and many have never been replicated. In addition, only six9 of the 44 TPP programs now judged by HHS to be evidence based have ever been shown to produce an impact on use of birth control or pregnancy rates, the most important measures of program success, and most have had only modest impacts on behaviors related to sex.
Effect Sizes
To increase the rate of success of evidence-based programs, at least four improve- ments to TPP’s tiered-evidence strategy should be considered. In establishing the requirements for Tier 1 funding, the administration was wise to admit evidence based on high-quality evaluations (roughly, either RCTs or well-executed quasi- experimental designs). By contrast, allowing a program that produces a significant impact on only one of five broad domains (sexual activity, number of sexual part- ners, contraceptive use, sexually transmitted infections, and pregnancy) identified by HHS to be considered evidence based seems too weak. HHS has now identi- fied 4410 model TPP programs that meet the criteria for being evidence based. The
7 A recent letter to the Evidence-Based Policymaking Commission from the federal Interagency Council on Evaluation Policy presented five recommendations for improving the federal government’s ability to evaluate programs. The five recommendations were to make more federal administrative records available to evaluators, increase federal evaluation capacity in lagging agencies, ensure funding for evaluations, eliminate bureaucratic barriers that interfere with agency evaluations, and take steps to make it easier for agencies to obtain and use federal data on wages, taxes, health, and education. 8 It is useful to point out here that HHS contracted with Mathematica Policy Research to conduct a thorough literature review, and to update the review on an annual basis, in order to locate model teen pregnancy prevention programs that meet the HHS criteria for being evidence based. For an overview of the procedure, see Haskins and Margolis (2015, pp. 54–58). 9 We determined that only six of 44 programs reported results for birth control and pregnancy by examining publications of all 44 evaluations. 10 The most recent summary on the Teen Pregancy Prevention Evidence Review website identifies seven new evidence-based programs, on top of the 37 identified as of February 2015, for a total of 44 programs.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 297
determination of evidence based required only one statistically significant impact on one of the five sex-related behaviors in one study. Moreover, the evaluation study could have been conducted by the organization that developed the program. These are weak standards for determining that a program is evidence based. Philip Peters (2015) of the University of Missouri puts the matter succinctly:
The Department of Health and Human Services failed to complement its demanding re- search design requirements with equally tough requirements for the minimum outcomes needed to qualify for federal funding. (pp. 44–45)
As Professor Peters points out, there are currently no requirements about the mag- nitude of effects, replication, consistency of findings, or the durability of effects. One possibility for strengthening the standards for determining evidence based suggested by Peters would be to require that experimental-control differences on some vari- ables meet a minimum requirement for effect size (the difference between the mean value of the experimental and control groups on a given measure divided by the av- erage of the standard deviations of the two groups). An advantage of the effect-size measure is that, rather than an either/or determination of whether the difference be- tween the two groups could have occurred by chance alone with some probability,11
effect sizes provide a measure of magnitude of difference between the two groups. Federal agencies should consider establishing a minimum effect size that must be met before a model program could be determined to be evidence based. HHS offi- cials could consult with experts in the field to determine the minimum acceptable effect size, but something on the order of 0.25 seems appropriate for outcomes such as self-reported sexual activity.12 Bigger effect sizes might be appropriate for more important outcomes such as use of birth control and pregnancy.
Longer-Term Impacts
Another approach, which could be implemented in conjunction with the effect-size requirement, would be to specify the length of time after the end of the program the effect would have to persist. Fadeout of impacts is the bane of intervention programs (Bailey et al., forthcoming). But it may be reasonable to expect that an impact would last for, say, six months before a program could be considered evidence based. In addition to increasing the chances that evidence-based programs could be replicated, another favorable effect of expanding the duration requirement would be to institutionalize the practice of getting at least six months or a year of follow-up data in federally supported program evaluations.
Surrogate and Ultimate Outcome Measures
A more fundamental approach to strengthening the definition of evidence based would be to require an impact on the outcome that is the ultimate goal of the intervention; for example, reduced pregnancy rates, increased high school or post- secondary graduation rates, increased employment levels or wages, reduced delin- quency and crime. It is often easier to produce impacts on what might be called surrogate outcomes such as frequency of sex or use of birth control than to produce impacts on ultimate outcomes such as pregnancy or birth reduction. In the case
Of the additional seven, one showed an impact on pregnancy, none on sexually transmitted infections, five on contraceptive use, none on number of sexual partners, and four on sexual activity (Lugo-Gil et al., 2016). 11 Each evaluation is statistically adjusted for multiple hypothesis testing. 12 This is the minimum effect size requirement suggested by Philip Peters.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
298 / Presidential Address: Making Federal Social Programs Work
of the TPP programs, all the evaluations measured surrogate behaviors. These are good outcomes to measure, and research shows all are correlated with teen preg- nancy itself, but correlation does not guarantee reduced pregnancy or birth rates. The argument here is that measuring surrogate outcomes is second best to directly measuring the main outcome itself, in this case reduction in pregnancy or births.13
If a requirement to achieve the status of evidence based were an actual reduction in pregnancy rates in a rigorous evaluation, the odds of successful replication would in all likelihood be much higher.
A problem with the measurement of teen pregnancy reduction in the TPP and similar programs is that the pregnancy may occur long after the end of the inter- vention. Nearly 75 percent of the youngsters enrolled in the 41 TPP programs were age 14 at the time of enrollment. At that age, only about 10 percent of the kids would have even had sex, let alone gotten pregnant. In fact, even among 18-year- olds, only about 60 percent have had sex (Finer & Philbin, 2013). Thus, to measure the most important outcome, it would be necessary to follow kids participating in the programs for several years, something that none of the projects did. Still, if the field of TPP is to develop programs that actually reduce teen pregnancy, it will be necessary to conduct evaluations that follow adolescents for many years to identify the programs that truly work. One possibility that OAH could consider would be to follow kids from the programs that had an impact on use of birth control and pregnancy for another three to five years.
Replication
One more change in the criteria for determining whether programs are evidence based is perhaps the most important of all. In accord with the rules of science, one of the most important goals of both research studies and evaluation studies is to produce findings that other investigators can replicate if they follow the same procedures employed in the original study. Requiring model TPP programs to pro- duce impacts in at least two separate studies to earn the title evidence based would increase the probability that the program would replicate when expanded to new sites. The goal of the kind of work OAH is conducting with TPP—and which should be the goal of all federal evaluation studies—is to develop a body of evidence from high-quality evaluations about which model programs produce outcomes in new settings. Only replications can satisfy this goal.
Summary
If the goal of federal policy is to identify model social programs that can reliably reduce the nation’s social problems, I believe we must strengthen the criteria for defining evidence-based programs—especially because these are the programs that should be receiving the bulk of federal grant dollars. Some combination of increas- ing the criterion for determining whether observed experimental-control differences in evaluations reach statistical significance, increasing the length of time after the end of the intervention an impact must last, requiring impacts on the ultimate outcome measure of an intervention and not simply surrogate measures, and re- quiring at least two evaluations showing significant impacts would move the federal
13 The Coaltion for Evidence-Based Policy expressed concern that only two of the 28 model programs judged to be “evidence based” by HHS in their first round of reviewing teen pregnancy prevention programs had truly strong evidence of effectiveness while 26 of the 28 were backed by preliminary evidence (Coalition for Evidence-Based Policy, 2010).
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 299
criteria for awarding the term “evidence based” to a social intervention program seems necessary. Lest it be thought that these criteria are too tough, consider the criteria required by the California Evidence-Based Clearinghouse for Child Wel- fare (2016). In order to be deemed “well supported” (the Clearinghouse’s highest rating) as evidence based, a social intervention program must meet the following criteria:
� At least two rigorous RCTs in different usual care or practice settings have shown the practice or program to be superior to an appropriate comparison practice.
� In at least one of these RCTs, the practice or program has a sustained effect at least one year beyond the end of treatment, when compared to a control group.
� The RCTs have been reported in published, peer-reviewed literature (California Evidence-Based Clearinghouse, 2016).
Of course, implementing the changes discussed above in how the federal govern- ment defines evidence based will result in fewer programs winning the accolades and increased chances of federal funding for programs that accompany the des- ignation as evidence based. Even so, innovative programs, or programs with only preliminary evidence, could still qualify for smaller grants in Tier 2; if found effec- tive in those evaluations, the programs could then qualify as evidence based and join Tier 1 with its greater funding. Lengthening the process of earning the accolade of evidence based is a small price to pay for greatly increasing the chances that a given social intervention program will reliably have the impacts it advertises.
Development and Implementation of Evidence-Based Programs
As suggested by the possibility of reducing the number of evidence-based programs, a second approach to increasing the chances evidence-based programs will have impacts when replicated is to create better ways to develop and implement the programs (Goesling, personal communication, September 19, 2016). OAH has al- ready initiated an approach to helping innovative programs move in this direction. OAH requires Tier 2 grantees to work to ensure that their program, as an OAH staffer explained to us, is “implementation ready” by the end of their five-year grant. The process of making model programs implementation ready includes iden- tifying and clearly defining core components of the program, developing a theory of how the model program produces its impacts on the behavior of teens, and clearly specifying curriculum and training materials and developing monitoring tools that help new programs determine whether they are implementing the model with fidelity (Kappeler, 2013). These OAH procedures should be followed by all federal agencies that are working with local projects as part of the tiered-evidence strategy.
Balancing Fidelity and Adaptation During Program Implementation
The field of evidence-based policy needs to learn a lot more about how to balance fidelity to a program model with the necessity of adapting programs to local con- ditions. The need for program developers and implementers to identify the central features of a program discussed above can be useful in seeking this balance be- tween fidelity and adaptation. The central features of a program model are the ones that are most important to implement to achieve impacts. In ideal circumstances, the impacts of these features have been established by rigorous evaluations. But adaptations might be required because of characteristics of the setting, the staff, the parents, or other specific local conditions when the program is expanded to
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
300 / Presidential Address: Making Federal Social Programs Work
new sites. This process of balancing fidelity and adaptation is also important when, after a year or two of implementation, a model program is not producing impacts at an acceptable level. Again, OAH is leading the way here and has made it clear to grantees that their job is to carefully implement the model program they have selected, to conduct a rigorous evaluation of results, and to use their experience and the evaluation results to engage in continuous program improvement or even program replacement.14 Only by following these or similar procedures will federal agencies be able to help program operators accept the importance of evaluation, and use it not as evidence for a final up or down vote, but as a means of continually improving their programs.
CONCLUSION
The TPP program is the leading example of the administration’s efforts to focus grant dollars on results by using and building evidence. In particular, the TPP program uses a tiered-evidence design that should be more widely used across the federal government. But like any relatively new effort, it will need adjustments and improvements over time based on experience—including from the large trove of program evaluations recently released. Those evaluations show that few evidence- based models were able to successfully replicate in new conditions and produce positive impacts, particularly on reducing teen pregnancy.
If evidence-based policy is to fulfill its promise, the field cannot be satisfied with the level of impacts achieved by the TPP network. The vision of the field must be that programs supported by rigorous evidence will be developed to ad- dress and reduce the nation’s major social problems and that these programs will reliably produce impacts when implemented with fidelity in new settings. The field is a long way from achieving this most basic aim. A vital next step in the growth of evidence-based policy is developing strategies—such as those outlined here—to increase the likelihood that evidence-based programs will replicate when implemented in new settings. If that goal can be achieved over the next several years, a milestone in federal social policy will have been achieved and the path to using federal grant funds to effectively attack and reduce the nation’s social problems will be at hand—and we will enter the platinum age of evidence-based policy.
RON HASKINS is the Cabot Family Chair in Economic Studies and the Co-Director of the Center on Children and Families at the Brookings Institution, 1775 Massachusetts Avenue NW, Washington, DC 20036 (e-mail: [email protected]).
ACKNOWLEDGMENTS
The author thanks Nathan Joo for research assistance; Evelyn L. Kappeler, Amy Farb, and Amy L. Margolis of the Office of Adolescent Health for comments on the manuscript; and the Annie E. Casey Foundation and the Laura and John Arnold Foundation for financial support. Mistakes belong to the author.
14 OAH has used the lessons learned from the evaluations from the 2010 to 2014 cohort of grantees to inform their work with a new cohort of grantees that began in 2015. The new cohort will implement and test the effectiveness of several evidence-based programs brought to scale in multiple settings in communities so that adolescents will receive evidence-based programs several times during their devel- opment (OAH 2016c). OAH (2016b) provides multiple resources for helping grantees select, implement, and adapt teen pregnanacy prevention programs in a way that both maintains fidelity and encourages appropriate adaptation and innovation.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Presidential Address: Making Federal Social Programs Work / 301
REFERENCES
Bailey, D., Duncan, G. J., Odgers, C., & Yu, W. (2017). Persistence and fadeout in the impacts of child and adolescent interventions. Journal of Research on Educational Effectiveness, 10, 7–39.
Baron, J., & Sawhill, I. V. (2010, May 1). Federal programs for youth: More of the same won’t work. Washington, DC: Brookings. Retrieved October 4, 2016, from https://www.brookings.edu/opinions/federal-programs-for-youth-more-of-the-same- wont-work/.
Berlin, G. (2014). Impact on a large scale: The importance of evidence. New York, NY: MDRC.
California Evidence-Based Clearinghouse for Child Welfare. (2016). CEBC review and rating process. Retrieved October 4, 2016, from http://www.cebc4cw.org/home/how-are- programs-on-the-cebc-reviewed/.
Chetty, R., Hendren, R. N., & Katz, L. (2015). The long-term effects of exposure to better neighborhoods: New evidence from the moving to opportunity experiment. Working Paper. Boston, MA: Harvard University.
Chetty, R., Grusky, D., Hell, M., Hendren, N., Manduca, R., & Narang, J. (2016). The fading American dream: Trends in absolute income mobility since 1940. Working Paper No. 22910. National Bureau of Economic Research.
Coalition for Evidence-Based Policy. (2009). Suggestions for the new social invest- ment/entrepreneurship initiative. Retrieved January 2, 2017, from http://coalition4evi dence.org/wp-content/uploads/2009/06/ideas-for-social-entrepreneurship-initiative-12309. pdf.
Coalition for Evidence-Based Policy. (2010). HHS’s evidence-based teen pregnancy preven- tion program: Excellent first step, but only 2 of 28 approved models have strong evidence of effectiveness. Retrieved October 4, 2016, from http://coalition4evidence.org/wp- content/uploads/2010/05/Coalition-comments-HHS-Teen-Pregnancy-Prevention-May- 2010.pdf.
Deming, D. (2009). Early childhood intervention and life-cycle skill development: Evidence from Head Start. American Economic Journal: Applied Economics, 1, 111–134.
Farb, A. F., & Margolis A. L. (2016). The teen pregnancy prevention program (2010–2015): Synthesis of impact findings. American Journal of Public Health, 106, S9–S15.
Finer, L. B., & Philbin, J. M. (2013). Sexual initiation, contraceptive use, and pregnancy among young adolescents. Pediatrics, 131, 886–891.
Garces, E., Thomas, D., & Currie, J. (2002). Longer term effects of Head Start. American Economic Review, 92, 999–1012.
Haskins, R., & Margolis, G. (2015). Show me the evidence: Obama’s fight for rigor and results in social policy. Washington, DC: Brookings.
Kahneman, D. (2002). Daniel Kahneman—Biographical. Retrieved January 1, 2017, from http://www.nobelprize.org/nobel_prizes/economic-sciences/laureates/2002/kahneman- bio.html.
Kappeler, E. (2013, September 9). OAH grantee guidance, OAH2013–1: Packaging and dissemination expectations for OAH Teen Pregnancy Prevention (TPP) research and demonstration grantees. Retrieved October 4, 2016, from https://www.hhs.gov/ash/oah/oah- initiatives/for-grantees/program-guidance/Assets/tpp_packaging_guidance.pdf.
Lester, P. (2016, May 10). Pay for success efforts roll forward in Congress, admin- istration. Social Innovation Research Center. Retrieved December 31, 2016, from http://www.socialinnovationcenter.org/?p=2053.
Lewis, M. (2017). The undoing project: A friendship that changed our minds. New York, NY: Norton.
Lott, M. (2013, July 30). Gov’t knows best? White House creates “nudge squad” to shape behavior. Fox News. Retrieved December 28, 2016, from http://www.foxnews.com/ politics/2013/07/30/govt-knows-best-white-house-creates-nudge-squad-to-shape-behavior. html.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
302 / Presidential Address: Making Federal Social Programs Work
Lugo-Gil, J., Lee, A., Vohra, D., Adamek, K., Lacoe, J., & Goesling, B. (2016). Updated findings from the HHS teen pregnancy prevention evidence review: July 2014 through August 2015. Cambridge, MA: Mathematica Policy Research.
Manzi, J. (2012). Uncontrolled: The surprising payoff of trial-and-error for business, politics, and society. New York, NY: Basic Books.
MDRC. (n.d.). CABS: Center for applied behavioral science. New York, NY: MDRC.
Morabia, A. (Ed.). (2016). Building the evidence to prevent adolescent pregnancy: Office of Adolescent Health studies (2010–2015). American Journal of Public Health, 106.
Nonprofit Finance Fund. (2016). Pay for success 101. Retrieved December 31, 2016, from http://www.payforsuccess.org/learn-out-loud/pfs-101.
Office of Adolescent Health (OAH). (2016a, August 16). Summary of evaluated programs effective at changing behavior. Retrieved October 4, 2016, from http://www.hhs.gov/ ash/oah/oah-initiatives/evaluation/grantee-led-evaluation/summary.html.
Office of Adolescent Health (OAH). (2016b, November 1). Choosing and implement- ing evidence-based programs. Retrieved October 4, 2016, from https://www.hhs.gov/ash/ oah/oah-initiatives/teen_pregnancy/training/curriculum.html.
Office of Adolescent Health (OAH). (2016c, December 8). Grantees FY 2015–2019. Retrieved October 4, 2016, from https://www.hhs.gov/ash/oah/oah-initiatives/evaluation/grantee-led- evaluation/grantees-2015-2019.html.
Office of Adolescent Health (OAH). (2016d, December 27). TPP program grantees (FY2010–2014). Retrieved October 4, 2016, from http://www.hhs.gov/ash/oah/oah- initiatives/tpp_program/cohorts-fy-2010-2014.html.
Office of Management and Budget. (2016). Table 6.1. Retrieved December 30, 2016, from https://www/whitehouse.gov/sites/default/files/omb/budget/fy2017/assets/histo6z1.xls.
Overholser, G. (2015). Up for debate: Responses. Stanford Social Innovation Review, Fall, 37–38.
Peters, P. G. (2015). The federal experiment with evidence-based funding. Regulation, 38, 40–47.
Pew-MacArthur Results First Initiative. (2014, August). New Mexico’s evidence-based ap- proach to better governance. Retrieved December 31, 2016, from http://www.pewtrusts. org/�/media/assets/2014/08/nm_results_first_brief_web.pdf.
Puma, M., Bell, S., Cook, R., Heid, C., Broene, P., Jenkins, F., . . . Downer, J. (2012). Third grade follow-up to the Head Start impact study. Rockville, MD: Westat.
Rotz, D., Luca, D. L., Goesling, B., Cook, E., Murphy, K., & Stevens, J. (2016). Final impacts of the teen options to prevent pregnancy program: Impact report from the evaluation of ado- lescent pregnancy prevention approaches. Cambridge, MA: Mathematica Policy Research.
Social and Behavioral Sciences Team (SBST). (2015, September). Annual report. Washington, DC: Executive Office of the President.
Social and Behavioral Sciences Team (SBST). (2016, September). Annual report. Washington, DC: Executive Office of the President.
Thaler, R. H., & Sunstein, C. R. (2008). Nudge: Improving decisions about health, wealth, and happiness. New Haven, CT: Yale University Press.
Washington State Institute for Public Policy. (2016). Retrieved October 4, 2016, from http://www.wsipp.wa.gov/.
Journal of Policy Analysis and Management DOI: 10.1002/pam Published on behalf of the Association for Public Policy Analysis and Management
Copyright of Journal of Policy Analysis & Management is the property of John Wiley & Sons, Inc. and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use.