NEED IN 12 HOURS or LESS
External Validity and Policy Adaptation:
From Impact Evaluation to Policy Design
Martin J. Williams ∗
February, 2018
Abstract
With the growing number of impact evaluations worldwide, the question of how to ap- ply this evidence in policymaking processes has arguably become the main challenge for evidence-based policymaking. How can policymakers predict whether a policy will have the same impact in their context as it did elsewhere, and how should this influence the policy’s design and implementation? This paper suggests that failures of external validity (both in transporting and scaling up policy) can be understood as arising from an interaction between a policy’s theory of change and a dimension of the context in which it is being implemented. I survey existing approaches to analyzing external validity, and suggest that there has been more focus on the generalizability of impact evaluation results than on the applicability of evidence to specific contexts. To help fill this gap, I develop a method of “mechanism mapping” that maps a policy’s theory of change against salient contextual assumptions to identify external validity problems and suggest appropriate policy adaptations. In deciding whether and how to adapt a policy, there is a fundamental informational trade-off between the strength of evidence on the policy from other contexts and the policymaker’s information about the local context.
∗Associate Professor in Public Management, University of Oxford, Blavatnik School of Government. Email: [email protected]. I am grateful for conversations and comments from Jon Ahlberg, Noam Angrist, Alex Baron, Maria Barron Rodriguez, Eleanor Carter, Suvojit Chattopadhay, David Evans, Flavia Galvani, Frances Gardner, Julie Hennegan, David Humphreys, Robert Klitgaard, Julien Labonne, Adauto Modesto, Aoife O’Higgins, Daniel Rogger, and students at the Blavatnik School and the Escola Nacional de Administração Pública (Brazil). Any remaining errors are my own. A policy memo based on this paper with a five-step “how-to” guide is also available (Williams 2017).
External Validity and Policy Adaptation: From
Impact Evaluation to Policy Design
Abstract
With the growing number of impact evaluations worldwide, the question of how to ap-
ply this evidence in policymaking processes has arguably become the main challenge for
evidence-based policymaking. How can policymakers predict whether a policy will have the
same impact in their context as it did elsewhere, and how should this influence the policy’s
design and implementation? This paper suggests that failures of external validity (both in
transporting and scaling up policy) can be understood as arising from an interaction between
a policy’s theory of change and a dimension of the context in which it is being implemented.
I survey existing approaches to analyzing external validity, and suggest that there has been
more focus on the generalizability of impact evaluation results than on the applicability
of evidence to specific contexts. To help fill this gap, I develop a method of “mechanism
mapping” that maps a policy’s theory of change against salient contextual assumptions to
identify external validity problems and suggest appropriate policy adaptations. In deciding
whether and how to adapt a policy, there is a fundamental informational trade-off between
the strength of evidence on the policy from other contexts and the policymaker’s information
about the local context.
In 2015, Zimbabwe’s government rolled out a new HIV treatment nationwide. The decision
was evidence-based: a range of randomized control trials (RCTs) had shown the new treatment
to be an improvement over the previous drug cocktail, and the World Health Organization
(WHO) recommended that it be used as the standard treatment throughout sub-Saharan Africa.
Yet after the new treatment was rolled out, “reports soon followed about people quitting it in
droves” (Nordling 2017). It turned out that one of the drugs in the treatment, efavirenz, caused
significant neuropsychiatric adverse effects (e.g. hallucinations, suicide ideation) in individuals
with a particular genetic variant. This variant is rare worldwide, so these adverse effects were
not deemed a major problem by international researchers, but it happens to be quite common
in Zimbabwe (Masimirembwa et al 2016). Even though the Zimbabwean government and WHO
had based their decisions on extensive and rigorous empirical evidence, the policy decision to
include efavirenz had disastrous effects for a significant fraction of patients - an error that could
have been avoided, since Zimbabwean scientists had previously identified this genetic variant
and its interaction with efavirenz (Nyakutira et al 2008). Relying on empirical evidence from
elsewhere without also utilizing local information had led policymakers in Zimbabwe to make a
costly mistake.
Policymakers worldwide face similar challenges in trying to apply evidence to policy deci-
sions. With the recent boom in impact evaluations around the world, policymakers in many
sectors now have at their disposal an overwhelming amount of evidence about “what works” -
or at least what worked in a particular context. Following the example set by the medical sci-
ences, the premise of evidence-informed policymaking is that the design of policy can be based
on evidence of what has worked in other contexts, rather than each policymaker having to start
from scratch. Yet as impact evaluations have multiplied, it has become apparent that “the
same” policy can have very different effects in different populations (Deaton 2010, Pritchett
and Sandefur 2015, Vivalt 2017).1 Similarly, policies shown to be effective in small trials have
not always been as effective when implemented at scale, even in the same country (Bold et al
2016, Banerjee et al 2016a). This is the problem of the external validity of impact evaluations.
Although there is widespread agreement on the importance of evidence for informing policy de-
cisions, the limited external validity of impact evaluation evidence poses significant challenges
for policymakers: how can one know if a policy will have the same effect in this implementation
context as it did elsewhere? And to what extent should policymakers copy the design of policies
that have worked elsewhere, rather than use local information to try to adapt them to fit the
1
local context?
This paper begins by proposing a simple and flexible framework for thinking about these
questions, and about external validity more broadly. A policy can have a different impact in a
new context than it had in a previous context if part of a policy’s theory of change interacts with
a difference in contexts. A policy’s theory of change is a set of logical steps spanning inputs to
activities, outputs, intermediate outcomes, and final outcomes. Whether this mechanism works
as intended depends at each step on the validity of a set of contextual assumptions. While these
assumptions may have been true of the context in which a policy had previously been shown to
work, whether the policy will have the same effects in a new context depends on whether these
same contextual assumptions hold. Since context can include a wide range of factors - location,
target group, implementing organization, scale, time period, the existence of related policy
interventions, etc. - and the theory of change includes factors related to implementation as well
as impact, this parsimonious framework encompasses the range of typologies of external validity
failures discussed in existing literature (Deaton 2010, Cartwright and Hardie 2014, Muller 2015,
Banerjee et al 2016a) and applies equally to issues of scale-up as well as transporting successful
policies.
Within the topic of external validity, one can further distinguish between the generalizability
of evidence and the applicability of evidence. Questions of generalizability are about whether
an impact evaluation’s findings are likely to hold in general in other contexts (but without a
specific destination context in mind), whereas questions of applicability are about whether eval-
uation results from one or more other contexts will hold in a specific destination context. While
the generalizability of evidence is somewhat informative about whether a policy will work in
a specific context, the multi-dimensionality of both policies and contexts means that evidence
from elsewhere can never be fully determinative of a policy’s impact in a new context as even
apparently minor idiosyncratic details can have a major effect on policies’ effectiveness. The
problem is analogous to the “last mile” problem in public transportation and service delivery:
evidence from rigorous impact evaluations can get policymakers significantly closer to the “cor-
rect” policy choice, but there is still a need for structured guidance on how to bridge the gap
from “what works in general” to “what will work in my context”.
With this distinction in mind, this paper surveys the existing literature on external validity
across a range of disciplines. Overall, the academic literature on external validity has focused
mainly on questions of generalizability, whereas the question of applicability is the problem with
2
which most policymakers are confronted when trying to use evidence to design policy. This is
evident in the two main strands of work on external validity. First, one existing approach
is to estimate the average effect of an intervention across different contexts by aggregating the
results of multiple studies. This is, broadly speaking, the approach of replication, meta-analysis,
and systematic review (What Works Network 2014, Vivalt 2017). A second strand focuses
on what evaluators can do to increase the external validity of a particular study, including
a range of approaches such as formal theory and structural modeling (Deaton 2010), larger
evaluations (Muralidharan and Niehaus 2016), various econometric extrapolation techniques
(Angist and Fernandez-Val 2010, Gechter 2016, Kowalski 2016), and integrating “structured
speculation” on external validity into research papers (Banerjee et al 2016b). While both
strands are informative about the potential applicability of evidence to specific contexts, their
emphasis on generalizability offers policymakers little structured guidance on how to bridge the
gap between evidence from other contexts and the inevitable particularities of specific contexts.
The paper then introduces a method to help policymakers to bridge the gap between the best
evidence from other contexts - from impact evaluations, meta-analyses, model-based extrapo-
lations, etc. - and the effective application of this evidence to policy design and adaptation
in their own context. This method of mechanism mapping builds explicitly on the paper’s
proposed understanding of external validity as the interaction of mechanism with context by
juxtaposing 1) the policy’s theory of change with 2) the underlying contextual assumptions
needed for each step of this mechanism to operate, and 3) comparing these assumptions to the
actual characteristics of the policymaker’s context. If a necessary assumption does not hold
in the new context in the same way as it held in the old context, then the mechanism will be
interrupted and the policy’s impact will differ. The mechanism mapping process can also be
applied to questions of policy scale-up, since implementing a policy at scale involves different
contextual assumptions (e.g. implementation quality, resource requirements, general equilib-
rium effects, political economy) than a small pilot, even if the pilot was undertaken in the same
geographical location.
Evidence plays a crucial role within the mechanism mapping process. Undertaking mech-
anism mapping would ideally consist of a systematic process of seeking empirical evidence to
support contextual assumptions and understand actual contextual realities through descriptive
statistics, qualitative data, and evidence from relevant impact evaluations. At the most rigorous
extreme, one could undertake a series of “mechanism experiments” (Ludwig et al 2011) to val-
3
idate each step of the theory of change and its underlying contextual assumptions. Where time
or resource constraints make a more thorough process infeasible but a decision must nevertheless
be made, even conducting a relatively brief mechanism mapping may help policymakers struc-
ture their judgment and avoid sole reliance on intuitions or prejudices. As Ravallion (2009)
notes, learning from impact evaluations typically requires both theory and information from
outside the evaluation; this simple and intuitive diagnostic process gives policymakers a flexible
framework for marshalling all available empirical evidence from different sources and of different
levels of rigor in a structured way in support of policy decisions. Whereas the lack of quanti-
tative data has often hindered evidence-based policymaking in data-poor contexts, mechanism
mapping’s ability to integrate less formal types of evidence makes it particularly well suited to
such contexts. Section 3 discusses the application of mechanism mapping in more detail.
The process of mechanism mapping also feeds directly into policy adaptation, by identifying
specific aspects of the policy that are likely to work less well (or potentially better) than in
the policy’s original context. Policy adaptations thus flow directly from a diagnostic of the
relationship between the policy context and the policy’s theory of change, so that adaptations
are based on a combination of local, context-specific information and evaluation evidence from
other contexts. While this combination is a productive way to generate ideas for adaptation,
it also suggests a fundamental trade-off. Evaluation evidence on a policy’s effectiveness in
other contexts is likely to be more rigorous than available local information, but relying on
this evidence from elsewhere requires strict fidelity to the original policy design. On the other
hand, using mechanism mapping to identify potential adaptations makes efficient use of local
information, but making these adaptations decreases the relevance of evaluation evidence from
elsewhere. The optimal level of adaptation in each case will depend on the case-by-case balance
between 1) the strength and relevance of evaluation evidence on the policy from other contexts
and 2) the policymaker’s information about the local context. This optimal level will thus vary
not only by policy area and country, but also by the information set of the policymaker and the
nature of the policymaking process.
Of course, a key limitation of mechanism mapping as a tool is that it relies to an extent
on the judgment of policymakers, as with any other process of policy design or adaptation.
While this may provide opportunities for policymakers’ biases or preferences to influence pol-
icy, numerous such opportunities already exist in policy design and implementation, and the
structured process of comparing the theory of change, contextual assumptions, and contextual
4
realities can arguably reduce these by making them explicit and structuring deliberation. An-
other limitation of mechanism mapping it that it yields only directional predictions of policy
effectiveness rather than precise statistical point estimates of effect sizes and confidence inter-
vals. In some cases mechanism mapping may not generate an unambiguous overall prediction
for whether a policy is likely to be more or less effective than it had been in a previous con-
text, since multiple contextual differences may shift effectiveness in different directions (i.e. be
opposite-signed). Nonetheless, for many applied policy purposes - in particular for identifying
aspects of a policy that may benefit from adaptation - a directional prediction may suffice, and
the use of mechanism mapping does not preclude policymakers from also making use of more
precise quantitative tools. A practical approach for policymakers would be to use evidence
on from impact evaluations and meta-analyses on what policies or interventions are likely to
be most effective as a starting point from which to begin the process of mechanism mapping,
marshalling context-specific evidence, and adapting policy.
In its emphasis on understanding mechanism-context interactions, this paper is most sim-
ilar to recent work in public health (Moore et al 2015, Leviton 2017), economics (Bates and
Glennerster 2017), philosophy (Cartwright and Hardie 2014), and public management (Barzelay
2007), and to “realist” approaches to evaluation in sociology (Pawson and Tilley 1997). The
contribution of this paper is to 1) present a flexible and parsimonious conceptual approach to
understanding external validity, 2) survey the existing literature’s strengths and limitations in
helping policymakers analyze external validity issues in their own contexts, and 3) link this to a
simple, practical, and intuitive framework for identifying likely external validity failures which
4) feeds directly into the policy adaptation process. Finally, mechanism mapping is related to
adaptive policymaking (Pritchett et al 2012, World Bank 2015) in emphasizing the use of local
information to improve policy design; Section 4 discusses complementarities between the two
approaches.
In order to focus on issues of external validity and policy transportation arising from real
differences in context, this paper abstracts from the issues of the statistical or methodological
accuracy of published impact evaluations that have been the focus of much of the literature on
replication in the social sciences (Christensen and Miguel 2016). While these issues can also lead
to differences in estimated policy impacts across contexts, they have been discussed extensively
elsewhere and are conceptually distinct. For brevity, throughout this paper I therefore discuss
impact evaluations as if they represent true causal estimates of the policy’s impact in that
5
context, even though policymakers should obviously interpret published findings through a
critical lens.
The remainder of this paper proceeds as follows. Section 1 defines external validity and
elucidates the understanding of external validity failures as an interaction of context and theory
of change. Section 2 discusses the limitations of existing approaches to external validity, largely
arising from the high dimensionality of policies and contexts. Section 3 describes the process of
mechanism mapping in more detail and gives examples and practical recommendations, Section
4 discusses adaptation and the fundamental informational trade-off, and Section 5 concludes.
1 Understanding External Validity
1.1 Defining External Validity
An impact evaluation’s external validity refers to the extrapolation of its findings beyond the
study sample to another population. This contrasts with the internal validity of a study, which
is established by the identification of a causal effect via comparison with a valid counterfactual.
While academics may be concerned about establishing the extent to which a study has external
validity in general - across all other hypothetical contexts - the policymaker’s problem is whether
the findings of a study conducted elsewhere would continue hold in one specific context. In
Cartwright and Hardie’s (2014) framing, an impact evaluation answers the question “did it
work there?”, while policymakers are interested in the question “will it work here?” This
section distinguishes between two different types of external validity: scaling up a policy within
the same target population, and transporting a policy to a different target population.2
First, one might be interested in extrapolating the findings from the study sample (the indi-
viduals or units who actually participated in the evaluation) up to the broader target population
for whom the study’s results are intended to be applicable. Will a policy have the same impact
on the full target population as it did on a smaller pilot or trial group? This is the case of
policy scale-up, as represented by panel (a) of Figure 1, and comprises two distinct aspects.
One aspect of this can be achieved by having a study sample that is statistically representative
of the study population (usually through random sampling).3 However, in some cases this is not
possible. For example, many pharmaceutical trials are conducted on healthy individuals even
though these are systematically different than the target population on important dimensions.
Similarly, Henrich et al (2010) point out that most psychological studies are conducted on study
6
samples that are Western, Educated, Industrialized, Rich, and Democratic (WEIRD) - often
North American or European undergraduates - that are among the least representative sections
of humanity.
In addition to these concerns about the representativeness of the study sample when scaling
up a policy to the full target population, the second aspect of scale-up concerns the implemen-
tation and impact of the policy itself. Implementing at scale may mean implementing through
a government bureaucracy that also implements many other policies rather than a small and
closely supervised non-governmental organization or academic research team, which could un-
dermine effectiveness (Bold et al 2016, Cameron and Shah 2017). Treating a higher fraction of
the target population could also lead to higher or lower effects through spillovers (e.g. Miguel
and Kremer 2004.
Figure 1: Two Types of External Validity
Study sample
Target population
Overall population Overall population
Original target population
New target population
(a.) External validity for scale-up decisions (b.) External validity for policy transportation
?
?
??
Second, one might care about whether a policy or intervention would have the same effect in
a different target population than the original target population. This is the meaning of external
validity with which policy transportation is usually concerned. Panel (b) of Figure 1 illustrates
this meaning of external validity. External validity in this sense concerns the similarity of the
two target populations on key covariates, both observable and unobservable.
Despite their distinctions, both types of external validity can be thought of as specific cases
of the same underlying challenge: how to predict whether a policy will have the same effect(s)
7
in a new implementation context as it did in a previous context. As discussed further below,
the underlying factors that drive external validity failures of both types can be understood
with the same framework, and the mechanism mapping approach to diagnosing them is equally
applicable to both.
1.2 Failures of External Validity
Why might a policy be effective in one context but fail in another, or vice versa? This paper
proposes a framework that builds on two key concepts. First, a policy is defined by its theory of
change (also referred to as its mechanism, results chain, or logic model). This is a mapping of the
policy’s intended mechanism - how it is supposed to work. This begins with the specification
of the intended final outcomes or ultimate goals of a policy. In order to achieve these final
outcomes, a series of intermediate outcomes must occur, and a policy specifies outputs that
the government will deliver in order to trigger these intermediate outcomes. To deliver these
outputs, government plans to undertake a set of activities, which require certain inputs (e.g.
financial or human resources, information).4 The steps from the provision of inputs to the
delivery of outputs comprise the implementation of the policy, while the link from these outputs
to the policy’s final outcomes via intermediate outcomes represents the impact of these outputs
on society, or - in Bates and Glennerster’s (2017) terminology - the behavioral response to the
intervention.
Second, all policies are implemented in a particular context, and the characteristics of this
context may affect a policy’s effectiveness. Context here refers not just to location, but also to
the full range of population and other variables that could affect the policy’s implementation
and impact. While the range of potentially relevant characteristics is effectively limitless, some
particularly salient dimensions of context include:
• Location, polity, or society in which the policy is being implemented (e.g. Iceland or India),
together with all the social, cultural, economic, geographic, and political characteristics
that vary across locations;
• Target groups (e.g. working adults, single mothers, at-risk teens);
• Time the policy is being implemented, whether the year (e.g. 1965 vs. 2015), season, or
duration since the policy began;
8
• Existence of related policy interventions, including spillovers from implementation of the
policy in neighboring areas as well as availability of public services or infrastructure; and
• Implementing organization, including its competence, level of resources, and political con-
straints.
Combining these two concepts makes it clear how failures of external validity emerge. A
policy’s theory of change relies on each step actually occurring and leading to the next step
as intended, both in terms of implementation and impact. Will the correct level of inputs be
made available as intended? Will the activities needed to create outputs actually occur with
the requisite quality and sequence? Will society react to these outputs as hypothesized? The
answer to all of these questions may have been affirmative in the context in which a successful
impact evaluation was undertaken but transporting the policy to a new context requires making
assumptions about the answers to these questions in the new context. The implementation
and impact of a policy are thus a function of the combination of a policy’s theory of change
with the context in which it is being implemented. If we then observe that a policy had
one impact in one context but had a different impact in a different context, then it must be
the case that the differences in context undermined one or more critical links in the policy’s
theory of change. Identifying these interactions of context and theory of change is critical to
understanding external validity.
Two examples illustrate the diverse ways in which contextual differences can undermine
policy impact. The first example comes from Cartwright and Hardie’s (2014, 80-84) compari-
son of two World Bank-funded programmes: the Tamil Nadu Integrated Nutrition Programme
(TINP), and the Bangladesh Integrated Nutrition Programme (BINP). The TINP project was
implemented in the 1990’s and sought to improve child nutrition in rural Tamil Nadu by simul-
taneously delivering two interventions: supplementary food for pregnant or nursing mothers and
their children, and nutritional advice to mothers to correct a misperception that mothers should
reduce rather than increase their food intake during pregnancy. A rigorous impact evaluation
showed that the project was successful: mothers’ nutritional knowledge improved, mothers and
children consumed more food, and children malnutrition and stunting decreased significantly.
Figure 2 maps a simple version of this theory of change.
Following this evaluation, the program was copied and transported to rural Bangladesh,
where the same problems existed. Yet under BINP, while mothers’ nutritional knowledge im-
9
Figure 2: Theory of Change: Bangladesh Integrated Nutrition Programme
Inputs Activities Outputs Intermed. outcomes 1 Intermed.
outcomes 2 Final
outcomes
Gov’t provides supp. food and nutrition advice to mothers
Mothers’ nutrition aware- ness improves
Mothers decide to use extra food for selves and infants
Mother and infant nutrition improves
Procure food, hire and train workers
Funds, potential workers
Implementation Impact
Source: Author’s elaboration based on Save the Children (2003), White (2005), World Bank (2005a), World Bank (2005b), Cartwright and Hardie (2014).
proved, there was no impact on malnutrition. This was due to a key contextual difference:
whereas mothers were typically responsible for shopping and household food allocation deci-
sions in rural Tamil Nadu, in rural Bangladesh men usually conducted the shopping and their
mothers (the mothers-in-law of the pregnant or nursing women) controlled household food allo-
cations (White 2005; Cartwright and Hardie 2014). This difference in contexts interacted with
a key link in the theory of change: the hypothesis that greater nutritional knowledge in mothers
would lead them to decide to allocate the supplemental food to themselves and their children,
rather than distributing it to other members of the household.
A second example comes from the Tools of the Mind early childhood education program,
which aimed to improve executive function (e.g. resisting temptation, working memory). After
a small but widely publicized RCT in New Jersey showed strong positive impacts (Diamond et
al 2007), a federally funded scale-up in other states actually found negative impacts relative to
a control group. Evaluators explained that correctly implementing the curriculum - “the most
complex we have ever seen” - required two years of training, ongoing in-classroom teacher coach-
ing, and carefully sequenced implementation of the 60 activities that comprised the program
(Farran and Wilson 2014, 21). Although teachers actually implemented the formal components
of the program with relatively high fidelity (as measured by the number of activities imple-
mented) the closely specified structure of the program did not fit well into the school day which
- unlike the carefully controlled original RCT - also included many other non-program activities
and demands on teachers’ attention. While children undertook many of the structured parts of
10
the Tools of the Mind curriculum, there was little time for them to undertake the kind of free
play that would have allowed them to internalize the skills taught in the structured parts of the
program (Farran and Wilson 2014). In this case the interaction between context and theory
of change that undermined program effectiveness was quite subtle: implementing the program
in a “real-world” setting necessitated the compression of a program component that seemed
unimportant but turned out to be crucial.
This framework for understanding external validity also makes it clear that contextual dif-
ferences do not affect policy impact unless they interact with the policy’s theory of change.
Although contexts are characterized by an almost infinite number of dimensions and are thus
all unique, this does not imply that all policies must be designed with a particular context in
mind, since most of these contextual differences are irrelevant to the policy’s mechanisms. In
practice, of course, it can be difficult to identify salient contextual differences and judge their
relevance - the specific interactions that undermined both BINP and the Tools of the Mind
scale-up may not have been obvious ex ante. Section 3 below presents a structured approach to
helping policymakers identify which dimensions of context are likely to affect a policy’s impact.
Finally, although the examples presented here have been of negative interactions with con-
textual differences, these interactions could just as well be positive - leading a policy that was
not effective in its original context to be effective in a new context. As Section 4 discusses later,
policy adaptations can aim not just to mitigate threats from transportation to a new context
but also to improve their effectiveness.
2 Existing Approaches to External Validity
With this framework in mind, this section surveys a range of existing approaches to dealing
with external validity. On the whole, it finds that the academic literature on external validity
provides increasingly insightful answers to questions of the generalizability of impact evaluations
or bodies of evidence - whether evaluation results from a specific context will hold in unspecified
other contexts. However, it provides more limited insight into concerns about applicability of
evidence - whether evaluation results from various other contexts will hold in the specific context
in which a policymaker is working. Since researchers are typically highly engaged in the conduct
of impact evaluations and creation of evidence for a generalized audience of evidence consumers
rather than in trying to apply evidence to specific policy decisions in specific contexts, it is
11
perhaps natural that the literature on external validity has also focused more on questions of
generalizability than applicability. Although these existing approaches can be very powerful
and vary tremendously, from empirical to theoretical and formal to informal, their common
limitation is their inability to analyze the heterogeneity of policies’ impacts across more than
a handful of dimensions. This contrasts with the high dimensionality both of policies and of
contexts (Pritchett 2017), and limits the overall usefulness of these approaches to policymakers
rooted in specific contexts.
One empirically driven response to the variability of policy impacts across contexts is to
aggregate numerous studies of the same policy. In its simplest form, this could be a simple
replication in another context. As the policy is tried and evaluated in more contexts, it may
become possible to aggregate these results further, through a systematic review or a meta-
analysis. This empirically driven approach is perhaps most associated with the evidence-based
policy movement, drawing its inspiration largely from medicine.5 Aggregation in this way can
yield an average treatment effect across study samples (and if the samples are representative
of their target population, across these populations) in which the policy has been studied. But
this estimate is of an average treatment effect in the average context in which the policy has
been evaluated, which can differ from the policy’s effect in a specific new context in two ways.6
First, the populations in which the policy has previously been tried and/or evaluated may
differ systematically from the new context in important ways. For many social policy interven-
tions, for example, there exist numerous studies from OECD countries but little or no evidence
in developing countries, and Allcott (2015) has shown that policy experiments are often con-
ducted first in the most favorable locations, leading to a site selection bias effect. Policymakers
applying this evidence to their own contexts must therefore ask “is my context average?” Since
contexts have many dimensions, all contexts are unique in some ways, and it is unclear how
many and which of these dimensions of a context must be “average” in order for this average
treatment effect to pertain. This is not to say that systematic reviews are uninformative: under
a normal distribution one would expect most contexts to be closer to the average than the ex-
tremes across most dimensions, and so absent any further information about the new context,
an average treatment effect estimated from other contexts would be the best predictor of a pol-
icy’s impact. But while this makes systematic reviews a useful starting point for policymakers,
näıvely adopting a policy that has a positive “headline” average treatment effect in a systematic
review is likely to backfire in many contexts.7
12
Second, there can be significant heterogeneity in policy impact across contexts, so that a
policy that has a positive effect on average could have a negative effect in some contexts. The
main empirical approach to dealing with heterogeneous effects is to employ sub-group analysis,
which breaks down average treatment effects across important variables: age, gender, income,
region, implementing authority, and conceivably any other observable variable on which data
exists. Conducting sub-group analysis, either within a single study or in a meta-analysis, allows
evaluators to answer the more nuanced question “what works for whom?” This allows policy-
makers to compare their contexts to others on these covariates, and provides some guidance
about which dimensions of context might matter for a given policy.
While this information is useful, sub-group analysis is inherently limited in the number
of variables along which they can disaggregate results. Individual studies are limited in the
number of variables they can measure and collect, and sub-group analysis in meta-analysis
is even further restricted by the limited set of variables that are common to all (or at least
several) studies. Inevitably, there will be some contextual variables that mediate a policy’s
effectiveness - who controls household food allocations, the fit of a curriculum within the existing
school day, the prevalence of particular genetic variants - that are either difficult to measure
or that evaluators might not think to measure ex ante, and are thus unobserved. Even where
a given covariate is present across studies, one might question whether this variable interacts
with the policy in the same way across contexts.8 For instance, low income might undermine
the effectiveness of a skill upgrading intervention in rich countries because individuals do not
have time to attend the classes (e.g. if they are working multiple low-wage jobs, or cannot
arrange childcare), but in a poor country income may not be correlated with time poverty in the
same way, so the intervention might be more effective. The validity of proxies may also differ
across contexts: for example, medical researchers sometimes use genetic data from African-
American populations to extrapolate findings to African populations, even though African-
Americans are not genetically representative of Africa as a whole (Rajman et al 2017). While
aggregation of evidence across multiple studies (and use of sub-group analysis for disaggregation
within this) is indeed informative about the predicted success of an intervention in any given
context - as Vivalt (2017) shows empirically - even for the most highly researched interventions
there nonetheless remain numerous potential mechanism-context interactions that existing data
cannot fully predict.
A second approach focuses not on aggregating evidence across more contexts, but on mak-
13
ing out-of-sample extrapolations through structural modelling (Deaton and Cartwright 2016) or
other empirical methods that can, in some circumstances, be used to extrapolate results from
one study to other populations, by exploiting specific forms of selection and non-compliance
within RCTs or by adjusting estimated impacts based on heterogeneity over observed covari-
ates (Angrist and Fernandez-Val 2010, Gechter 2016, Kowalski 2016, Andrews and Oster 2018).
Such theoretical or empirical extrapolation methods take advantage of variation within the
study sample in order to better understand the underlying causal processes and extrapolate es-
timates from the actual study sample to other populations. These methods help researchers and
policymakers further improve the informativeness of the existing evidence about the predicted
impact of the intervention in a new context, but are also inherently limited in the number of
variables and types of scenarios across which they can extrapolate. As Low and Meghir (2017,
34) write:
“Structural economic models cannot possibly capture every aspect of reality, and any
effort to do so would make them unwieldy for either theoretical insight or applied
analysis. There will always be some economic choices left out of any particular model
- the key question is how to judge what aspects to leave out without rendering the
quantitative conclusions of the model irrelevant.”
Yet as the examples of BINP, Tools of the Mind, and Zimbabwe’s efavirenz rollout illustrate,
the range of contextual factors that can influence policy impact is immense. While structural
modeling can therefore provide powerful insights about the effect of some important contextual
factors, even the best-judged model will only be able to incorporate a small handful of the
numerous variables that policymakers must consider in policy design. As with the aggregation
approach discussed above, this is certainly informative for policymakers in specific contexts,
and is an improvement over simply having the results of an impact evaluation from another
context without such extrapolation, but still falls short of taking into account all the potential
mechanism-context interactions with which policymakers must concern themselves.
Similarly, the design of policy experiments has begun to deliberately vary aspects of the
policy that are important for understanding external validity, such as whether it is implemented
by an NGO or government (Bold et al 2016, Cameron and Shah 2017, Angrist 2017). Again,
the limitation is that trials can only feasibly vary one or two dimensions of a policy without
losing statistical power, while the number of dimensions of policy and context that could matter
14
- combined with their interactions - is effectively infinite. Likewise, larger experiments would
certainly improve external validity (Muralidharan and Niehaus 2016), but applying the results
will always require the consideration of contextual differences for any trial on a scale that is less
than global. Banerjee et al ’s (2016b) proposal for the inclusion of “structured speculation” on
external validity in reports of impact evaluation results is perhaps the closest in spirit to the
mechanism mapping approach developed in Section 3 of this paper, but is fundamentally a tool
for evaluators, not policy designers, since such speculation is necessarily undertaken without a
specific target context in mind.
The “realist evaluation” approach pioneered in sociology and social policy (Pawson and
Tilley 1997) shares with the recent external validity literature in economics an emphasis on
mechanisms and heterogeneity rather than simply establishing average treatment effects. In
asking “why a program works for whom and in what circumstances” and seeing the objective
of evaluation as the elaboration of “Context-Mechanism-Outcome configurations (CMOCs)”,
realist evaluation is also related to this paper’s mechanism mapping approach, albeit from
the perspective of the evaluator rather than the policymaker. However, the focus of realist
evaluation is typically on how best to evaluate a program rather than how to use existing
evidence to design a policy. This difference in target audiences has perhaps contributed to realist
evaluation being perceived as deeply philosophical as well as unwieldy and time-intensive in
practice (Marchal et al 2012). In contrast, this study’s mechanism mapping approach is intended
to be practical and simple enough for policymakers to incorporate into routine processes of policy
design.9
Finally, debates around external validity are perhaps most advanced in public health, where
discussions of the interaction between mechanism and context have become central to thinking
about the transportation and scaling of trial results (Moore et al 2015, Leviton 2017), the
complexity of interventions is widely acknowledged and is beginning to be explored empirically
(Hawe 2015), systematic reviews routinely take realist approaches to unpacking mechanisms and
heterogeneity (Jagosh et al 2015, Greenhalgh et al 2016), and a strong institutional architecture
is seeking to establish reporting conventions and other steps to embed these new approaches
in research (e.g. Wong et al 2016). While this is a model for other social sciences and policy
fields to follow in some respects, there is also a need for a simpler and less resource-intensive
analytical framework that is more accessible for other fields and policy contexts.
15
3 Mechanism Mapping
3.1 Basic Process and Example
Although these techniques for assessing and improving the external validity of impact evaluation
evidence are important to help guide policymakers towards policies and interventions that are
more likely to be successful, evidence-based policymakers are still faced with the challenging
task of diagnosing potential mechanism-context interactions that could influence the policy’s
effectiveness in their context. If external validity failures arise from interactions between a
policy’s theory of change and its context, then it follows that diagnosing such failures requires
a way to examine a policy’s theory of change alongside its context. Furthermore, if contexts
have numerous dimensions, for many of which hard data may not be available, then a useful
framework also needs to be able to integrate high-quality, rigorous evidence (“observables”) as
well as softer, potentially tacit or local information (“unobservables”). This section introduces
a mechanism mapping approach that fulfills both criteria. I introduce this method by first
presenting the approach itself, then giving a brief example.10 I then discuss some practical and
conceptual issues, including the role of empirical evidence, then suggest some ways in which this
tool can be integrated into processes of policymaking and evaluation in order to complement
other approaches to evidence-based policy.
The first step of mechanism mapping is to lay out the theory of change, or mechanism,
through which the policy had its measured impact in the previous context. As discussed in
Section 1, this can be thought of as a causal chain leading from a policy’s initial inputs to
its intended final outcomes, via activities, outputs, and intermediate outcomes. The second
step is to lay out the most important or salient contextual assumptions underpinning each step
of this chain. These are the characteristics of the context that are required for the policy to
actually function as the theory of change intends. If the policy in question had been shown
to be successful in another context, then presumably these assumptions would have been valid
in that context. The third and final step is to lay out the corresponding actual contextual
characteristics for each step of the chain, highlighting any differences between actual contextual
characteristics and the contextual assumptions necessary for the policy to function as intended.
These differences in context - whether negative or positive - are what policymakers can use to
predict whether the policy will have a similar, smaller, or larger impact on the final outcomes
than it did in its previous context, as well as pinpointing the stage at which the theory of change
16
is likely to be interrupted (and thus which aspects of the policy may need to be adapted, as
Section 4 discusses later).
To illustrate the approach, consider the Bangladesh Integrated Nutrition Programme (BINP)
discussed above, which had an identical design to the World Bank’s earlier TINP project in Tamil
Nadu.11 The intended final outcome of BINP was to improve mother and infant nutrition. To do
so, government was to provide two main outputs: nutritional advice delivered to pregnant and
nursing mothers, and the distribution of supplementary food to mothers to take home. These
outputs would lead to the final outcome via two sets of intermediate outcomes: first, mothers’
nutritional awareness would improve, alongside their receipt of the supplemental food; and
second, mothers would then decide to use the supplemental food for themselves and their infants
(as opposed to giving it to other family members, i.e. program “leakage”). In order to produce
these outputs, the government required inputs of adequate financial resources to purchase the
food and pay personnel, as well as a logistical system and potential pool of extension workers
to deliver the food and nutritional advice. Key activities for transforming inputs into outputs
could include procuring the food, hiring and training workers, and conducting outreach to
eligible mothers.
The contextual assumptions required for this theory of change to work are listed in the
second row of Figure 3. Sound implementation requires that government: dedicate adequate
financial and human resources to the project; procure and distribute food and hire workers
effectively, including quality assurance as well as prevention of excessive corruption, and train
workers adequately; and deliver these outputs to a pool of eligible mothers predictably and in
a timely fashion. Impact then requires that mothers are able to attend the sessions and trust
the advice they are being given; that mothers actually control household food allocation; and
that the supplementary food, if consumed, will actually lead to the desired improvement in
nutrition. In the Tamil Nadu context, these assumptions were presumably valid - hence the
impact evaluation finding that TINP significantly improved mother and infant nutrition (World
Bank 2005b).
The third row of Figure 3 contrasts these contextual assumptions to the actual contextual
characteristics of the new context, in this case rural Bangladesh. Recall that BINP succeeded
in distributing food and nutritional advice to the mothers, and that mothers’ nutritional aware-
ness did actually improve as a result, but that the program failed to improve mother and infant
nutrition because most of the supplementary food went to other family members. The key
17
Figure 3: Theory of Change: Bangladesh Integrated Nutrition Programme
Inputs Activities Outputs Intermed. outcomes 1 Intermed.
outcomes 2 Final
outcomes
T he
or y
of c
ha ng
e Gov’t provides supp. food and nutrition advice to mothers
Mothers’ nutrition aware- ness improves
Mothers decide to use extra food for selves and infants
Mother and infant nutrition improves
Procure food, hire and train workers
Funds, potential workers
Food and advice delivered to target mothers in timely fashion
Mothers able to attend sessions, trust advice
Mothers control house- hold food allocation
More food will improve nutrition
Adequate procure- ment, logistics, training processes
Funds will be dedicated and suitable workers available
C on
te xt
ua l a
ss um
pt io
ns
Food and advice delivered to target mothers in timely fashion
Mothers able to attend sessions, trust advice
Mothers do not control house- hold food allocation
More food will improve nutrition
Adequate procure- ment, logistics, training processes
Funds will be dedicated and suitable workers available
A ct
ua l c
on te
xt
Source: Author’s elaboration based on Save the Children (2003), White (2005), World Bank (2005a), World Bank (2005b), Cartwright and Hardie (2014).
contextual assumption that did not hold in Bangladesh was that mothers controlled house-
hold food allocation, and would thus be able to act on their improved nutritional awareness.
This broke the link between Intermediate Outcome 1 and Intermediate Outcome 2; since In-
termediate Outcome 2 was not achieved, neither was the Final Outcome. If the designers of
BINP had carried out a mechanism mapping when transporting the successful TINP program
to Bangladesh, perhaps they would have uncovered this crucial but implicit assumption.
3.2 Practical and Conceptual Issues
Each of these three steps - the theory of change, contextual assumptions, and actual charac-
teristics - are associated with important conceptual and practical challenges. The first is how
to decompose a policy or program into a theory of change. In many cases, policymakers in
new contexts can establish this based on the original impact evaluation or meta-analysis on
18
which they are basing the transportation or scale-up of the policy (although many published
impact evaluations do not fully specify their intended theory of change or evidence supporting
this mechanism, leaving it partially to the policymaker in the new context to specify this causal
chain). While this can be challenging, theories of change (and related tools such as logic mod-
els, results chains, and logframes) are an intuitive and commonly taught tool of policymaking
in many evaluation textbooks (Gertler et al 2016), government manuals (HM Treasury 2011),
and donor guidelines (DFID 2012). In practice, the aspect of laying out theories of change
that policymakers often find most challenging is related to classification of causal steps (e.g.
activities vs. outputs, etc.), which for the purpose of mechanism mapping matters less than
isolating each link in the causal chain.
Of course, some policies are more complicated than others. For clarity, mechanism mapping
is here illustrated using the type of simple, linear theory of change that adequately characterizes
many impact evaluations. Theories of change can of course be much more intricate, for example
by mapping out multiple components of a multi-faceted program. Supplementary Appendix A
presents an example of a policy whose theory of change contains 55 distinct steps across multiple
interlocking components. Since each of these 55 steps is numbered and has its precursor steps
clearly identified, the same basic procedure of juxtaposing the contextual assumptions against
actual contextual characteristics for each step can be applied (albeit perhaps in the form of a
table rather than underneath each step in the theory of change, for reasons of visual clarity).
Similarly, mechanism mapping can be adapted to policies that are intended to lead to mul-
tiple final outcomes (e.g. a cash transfer that is intended to increase consumption and improve
child school attendance); to examine the likelihood of negative outcomes or side effects of the
policy (Bonell et al 2014); and to shed light on questions of multiple competing mechanisms.
This involves laying out multiple parallel mechanism maps representing different potential path-
ways to the positive or negative outcomes, then analyzing each to identify which seem most
plausible and which key contextual assumptions or characteristics seem most important. Sup-
plementary Appendix B presents an example theory of change from a Botswanan NGO that
mapped out parallel theories of change against sets of contextual assumptions in order to ex-
amine competing mechanisms and scenarios leading to both positive and negative potential
outcomes as part of the intervention adaptation process prior to transporting an intervention
from Kenya to Botswana. Finally, some policies are not just complicated but also complex
in that their implementation involves a significant degree of uncertainty, simultaneity, and/or
19
feedback. While complex policies are inherently more difficult to create theories of change for
(as well as to conduct impact evaluations on), Davies (2004), Rogers (2008), and De Silva et al
(2014) provide examples and guidance for mapping theories of change for complex interventions,
which all share the common feature of breaking interventions down into a set of logical steps
and are thus amenable to mechanism mapping.
A second conceptual challenge confronting policymakers is how to identify which are the
most salient contextual assumptions to consider, since the high dimensionality of context makes
it unfeasible to consider all such assumptions. Although this is ultimately a matter of judgment,
four practical guidelines suggest themselves:
1. Results from sub-group analysis of impact evaluations or meta-analysis may give insights
into common determinants of program effectiveness and thus shed light onto key assump-
tions. For example, a meta-analysis of food supplementation programs showed that pro-
grams where supplements were delivered at feeding centers found lower average leakage
of food to other family members (15 percent) than when food was delivered at home (64
percent), suggesting that assumptions around intra-household food allocation are key to
the success of such policies (Kristjansson et al 2015).
2. Many dimensions of context are frequently salient and should be taken into consideration
for almost any policy: demographic and socioeconomic characteristics of the target pop-
ulation; resource availability; political support and resistance; social and cultural norms;
the effectiveness of implementing organizations; potential for corruption or resource diver-
sion; geographic accessibility and other logistical issues; etc. The UK government Magenta
Book (HM Treasury 2011) and Ravallion (2009) provide lists of aspects of context for pol-
icymakers to consider when transporting a program.
3. Important contextual factors specific to a particular policy are often implied by the pol-
icy’s theory of change (Moore et al 2015). For instance, laying out BINP’s theory of
change makes it clear that decisionmaking over household food allocation is a key contex-
tual assumption, and doing a mechanism map for the Tools of the Mind early childhood
education program scale-up would have made it clear that assumptions were necessary
about the fit of the program into the existing school day.
4. Participatory policymaking processes, where input from affected stakeholders such as im-
pacted populations (Parker et al 2008) or implementing staff (Leviton and Trujillo 2016)
20
is systematically sought during the policy design process, are especially well suited to
identifying salient contextual assumptions, since directly affected or involved individuals
are likely to be able to more accurately envision the practicalities of the policy’s impact
and implementation.
A third practical challenge - establishing actual contextual characteristics to compare to
these assumptions - is another aspect of mechanism mapping where empirical evidence is crucial.
In addition to compiling evidence on impact from existing impact evaluations or meta-analyses,
policymakers can also gather new data by (to return to the BINP example) examining budget
data and political context to shed light on resource availability, investigating the performance
of the implementing agency’s procurement processes, conducting a survey of eligible mothers’
trust of the state and baseline level of nutritional knowledge, undertaking (or reading existing)
qualitative research on household food allocation decisions in rural Bangladesh, and discussing
with public health experts the prevalence of diseases that might inhibit infants from absorbing
nutrients properly. Bates and Glennerster (2017) present several excellent examples of using
simple descriptive data, some gathered in just two weeks, to validate contextual assumptions.
Impact evaluation results from other contexts and systematic reviews can enter into mechanism
mapping via the contextual assumptions row, as policymakers can use the results of that eval-
uation to document the extent to which the contextual assumptions held in that context, and
possibly even to investigate how variation in these contextual factors was related to the policy’s
effectiveness. When the mechanism mapping is being conducted for a scale-up of a policy that
has already been trialed on a small scale in the same location, the mechanism mapper may
even have quite detailed evidence on these issues, and so the search for new empirical evidence
can focus on the aspects of context that are changing with the larger-scale implementation:
the effectiveness of the implementing agency, general equilibrium or spillover effects, political
economy issues, etc. One could even imagine policymakers conducting quick and cheap “mech-
anism experiments” (Ludwig et al 2011) to validate each step of the theory of change prior
to beginning full-scale implementation, or (less ambitiously) follow Rigterink and Schomerus
(2016) in compiling evidence from existing evaluations on the validity of some aspects of the
policy’s theory of change in cases where no impact evaluation of the full theory of change has
yet been conducted. In practice, of course, the available evidence on each of the contextual
assumptions, and hence each step of the theory of change, is likely to vary in terms of rigor
and depth. Mechanism mapping thus provides an integrative framework for policymakers to
21
apply all available evidence - from RCTs to administrative data to qualitative research to expert
judgment - to their policy decisions, and the same basic process can be scaled up or down in
terms of detail in order to fit within policymakers’ time, resource, and information constraints.
As a procedural matter, a practical way for policymakers to conduct mechanism mapping
is in a nested manner. The analyst begins by identifying only the most salient steps in the
policy’s theory of change along with accompanying contextual assumptions and characteristics,
following the guidance above. This presents a top-level picture of the overall fit of the policy’s
required assumptions with the context’s actual characteristics. At this stage, it is likely that
some steps of the theory of change will have a better fit than others (as in Figure 3). From this
top-level view, each of these links in the mechanism can then be broken down and analyzed
in more detail. Where the contextual assumptions seemed to fit well at the overall stage - for
instance, the activities or outputs steps of Figure 3 - breaking down the mechanism serves as a
further plausibility check. For instance, an ex ante mechanism map of BINP could have thought
in more detail about the steps involved in procuring and distributing food, in hiring and training
workers, and in coordinating these two program elements, and what resources and bureaucratic
processes and skills would be required to execute them. Where the contextual assumptions do
not seem to fit well at the top-level stage - for instance in Figure 3’s intermediate outcome
of mothers and infants consuming the extra food themselves - going into more detail can help
the analyst identify the root cause of the disjuncture. In the BINP case, this was that TINP’s
synergy between advice and food distribution to mothers would not apply in Bangladesh. Being
more precise in pinpointing the problem simplifies the problem of adaptation discussed in the
next section. Continuing this nested approach to mechanism mapping even further in detail
could be especially useful for bureaucratic planning processes, by linking a program’s theory of
change to a detailed set of tasks to be performed and coordinated.
While mechanism mapping is intended primarily as a tool for policymakers to use prospec-
tively to predict the impact of transporting or scaling up a policy that has been successful
elsewhere, mechanism mapping is also of potential value to evaluators and to policymakers de-
signing policy from scratch. First, it can be useful in the retrospective evaluation of policies
by helping evaluators to show clearly the intended and actual mechanism(s) through which a
policy had its impact (or non-impact). Showing intended versus actual mechanism maps in this
way can help evaluators clarify their own thinking and also make the evaluation more informa-
tive to readers and policymakers from other contexts. Second, prospective mechanism mapping
22
(e.g. during trial design or in a pre-analysis plan, as in the example discussed in Appendix B)
can help evaluators ensure that they collect the data necessary to assess each of the contextual
assumptions ex post, along with potential undesirable outcomes and the alternative mechanisms
that might bring them about. This process can even be useful in cases where the policy being
trialled is completely new, rather than transported from elsewhere or scaled up. Third, mech-
anism mapping can help evaluators make null results more informative, by specifying which
aspects of the policy (if any) worked as intended, where the causal chain broke down, and what
future trials should (or should not) considering adapting.
Of course, an important limitation of mechanism mapping for evaluation purposes is that
mechanism mapping is only intended to yield directional predictions about overall policy im-
pacts, unlike statistical methods that can yield point estimates and other more precise infor-
mation. However, directional predictions are still useful for many purposes - in particular for
optimizing policy design. Finally, the process of being explicit about a policy’s theory of change
and the fit between its contextual assumptions and actual characteristics is also important for
policymakers who are designing policies from scratch, without the aid of a successful trial from
elsewhere. While the lack of prior evidence makes this inherently more uncertain, particularly
in terms of accurately specifying the theory of change, the same basic structure and concepts
may be helpful nonetheless.
4 Policy Transportation and Adaptation
The external validity of impact evaluations is often framed as a question of “would the same
policy work in another context?” In practice, however, it is usually necessary to make at least
some adaptations to a policy in order for it to work in a new context. These can be superficial,
as in the translation of program materials into a different language, or more substantive, for
example by adapting the nutrition advice component of BINP to include not just mothers
but their husbands and mothers-in-law. Even where such adaptation is not strictly necessary,
appropriate adaptations may sometimes be able to optimize an already effective program. But
the number of adaptations that could be made to a policy or intervention is nearly infinite -
which aspects should be targeted for adaptation, and which left alone? And how much should a
policy that was successul in another context be adapted, since adaptations risk changing aspects
of the policy that make it effective?
23
Across social science disciplines, these questions of adaptation are even less well studied
than questions of external validity, with only ad hoc (if any) discussion of how external validity
concerns should be addressed through adaptations. For example, Bates and Glennerster (2017)
present a framework for analyzing external validity and several applied examples, but potential
adaptations resulting from their diagnostics are mentioned only briefly (both in the framework
and their discussion). The literature is most advanced in social policy and applied psychol-
ogy, where there is a well-worn debate on the trade-offs between fidelity to evidence-backed
interventions versus cultural adaptation of programs (Castro et al 2010), although even this
literature offers little guidance to policymakers beyond conducting focus groups and small-scale
pilots. A small literature examines this debate empirically, by comparing the effectiveness of
various policies or programs according to whether they were newly designed (“homegrown”),
transported but adapted, or transported without adaptation. The results are mixed: Hasson
et al (2014) compare 307 German and Swedish social interventions and find that novel and
adapted programs are more effective than programs that are transported without adaptation,
while Leijten et al (2016) find no difference on average between homegrown and transported
parenting interventions across a range of countries, and Gardner et al (2015) find that several
branded parenting interventions developed in the United States and Australia are at least as
effective in non-Western countries even with little adaptation. Of course, the challenge of trying
to use meta-analytic methods to ascertain the optimal level of adaptation is that it is unclear
exactly what changes in context the adaptations were responding to, or how appropriate the
adaptations were. As with policy choice, knowing the average effectiveness of adapted policies
is less useful for policymakers than guidance in identifying which adaptations are likely to be
necessary and effective in their specific context.
The diagnostic aspect of mechanism mapping helps fill this gap confronting policymakers.
Since mechanism mapping as a diagnostic tool focuses on the interaction between a policy’s
theory of change and differences in context, the diagnosis of whether a policy is likely to be
as effective in a new context as it was elsewhere inherently involves highlighting the aspects of
the policy that should be targeted for adaptation. In the case of BINP, for example, Figure 3
makes it obvious that the key aspect where adaptation was necessary was the nutritional advice
component, and specifically the individuals to whom this was targeted. The mechanism map
alone is not sufficient to determine exactly what the adaptation should be - whether it is possible
to simply include husbands and mothers-in-law through the existing delivery mechanism, for
24
example, requires additional context-specific knowledge and feasibility investigations, as in any
policy design process - but it does identify which aspect of the policy is problematic, and why.
Similarly, Figure 3 makes clear that the other steps in BINP’s theory of change fit well with
the contextual assumptions and previous context in which the program had been evaluated,
suggesting that there is little need for adaptation in these respects. The design of the resulting
adapted policy is thus informed both by evaluation evidence from other contexts - through the
aspects of the original policy that were maintained in the new context - as well as by local,
context-specific knowledge - through the aspects that were adapted.
The same framework can also be used to identify adaptations that might be necessary (or
potentially detrimental) in scaling up a policy that was successful in a small-scale trial. Most
obviously, contextual assumptions that held in the trial may not hold when implementing at
scale. For example, government agents may require different incentive or monitoring schemes
than non-governmental agents in order to elicit similar effort levels (Cameron and Shah 2017),
or incentive schemes that were effective for non-governmental implementers in small trials may
be unfeasible for political economy reasons when implemented at scale (Bold et al 2016), ne-
cessitating adaptation of the policy for scale-up. Similarly, adaptations might be imposed by
the scale-up process itself, as in the Tools of the Mind scale-up where implementation at scale
in schools without close experimental control led unstructured-but-crucial components of the
program to be crowded out by other demands on school time. Where such enforced adaptations
or risks can be foreseen in advance, mechanism mapping provides a framework for thinking
through their potential consequences and thus possible mitigating measures.
Of course, making even apparently superficial adaptations to an intervention also creates
new potential mismatches between the contextual assumptions for the adapted theory of change
and contextual realities. In practice, then, policymakers need to repeat the diagnostic process
outlined in the previous section on the adapted theory of change, then potentially re-adapt, and
so on in an iterative process. In this way, mechanism mapping complements another recent inno-
vation: adaptive policymaking, which views policy design and evaluation as an iterative process
of experimentation with tight feedback loops (Pritchett et al 2013, World Bank 2015). As a
policy diagnostic tool that focuses on mechanism, context, and their interaction, mechanism
mapping can be integrated with adaptive policymaking processes to help connect experimen-
tation to a more precise diagnosis of the barriers to the effectiveness of previous iterations of
a policy, thus adding precision to the experimental search process. Specifically, policymak-
25
ers can use the weakest assumption link in the theory of change (as identified by mechanism
mapping) as the starting point for adaptive experimentation. Similarly, since monitoring and
data collection strategies are typically based a policy’s theory of change or logframe, the data
that organizations generate during adaptive policymaking processes often closely aligns with
the evidence required to make mechanism mapping more empirically rigorous.
With respect to the second question, on the optimal extent of adaptation, mechanism map-
ping’s simultaneous use of evaluation evidence from other contexts and knowledge of the local
context highlights a fundamental trade-off. On one hand, evaluation evidence on a policy’s
effectiveness in other contexts is likely to be more rigorous (especially in causal identification)
than information about the local context. However, relying on this evidence requires strict
fidelity to the original policy design, implying that policymakers should refrain from making
adaptations as much as possible. On the other hand, using mechanism mapping to identify
potential adaptations can make efficient use of local information, which (even if less rigorous) is
specific to the context in question. However, using this local information to make adaptations
decreases the relevance of evaluation evidence from elsewhere. There is therefore a trade-off
between evidence from other contexts and local knowledge of the current context, making it
unclear how quick policymakers should be to make adaptations when they have identified a
difference in context. Should adaptations be made in response only to major differences, or also
to minor ones? And what constitutes a major or minor difference?
Figure 4: The Fidelity-Adaptation Spectrum
More adaptation
More fidelity
Homegrown policy
Substantive adaptations
Full fidelity
Superficial adaptations
Optimal policy with better evidence from other contexts
Optimal policy with better information about local context
There is no universally optimal solution to these questions, because the characteristics of
a specific context - however apparently minor or idiosyncratic they are - can undermine the
effectiveness of even the most evidence-backed policy, yet policymakers’ ability to foresee these
interactions is limited (hence the need to evaluate policies and use evidence in the first place).
26
That said, the respective roles of evidence and local information suggests that the optimal
extent of adaptation will vary from case to case, depending on several factors.
First, to the extent that available impact evaluation evidence on the policy is strong, con-
sistent, and from similar contexts, policymakers should make fewer adaptations (all else equal).
These factors reduce the uncertainty associated with a policy’s impact in its current form, thus
increasing the risk that adaptations could backfire. For example, Evans and Popova (2015),
Pritchett (2017), and Vivalt (2017) show that some types of development interventions exhibit
much greater variance across trials, suggesting that some interventions are more sensitive to
contextual differences than others and thus presumably have a greater need for adaptation.
Second, the greater the policymaker’s information about the local context - whether in the
form of formal evidence and data, or simply familiarity and tacit knowledge - the more a policy-
maker should be willing to adapt a policy, since this knowledge allows for better identification
of negative or positive context-mechanism interactions as well as suitable adaptations. This im-
plies that the optimal level of adaptation in a specific case will vary not only by policy area and
country, but also by the information set of the policymaker: ceteris paribus, a policymaker with
less familiarity, information, or ability to gather information about the target context should
generally make fewer adaptations to a transported policy than a policymaker who either has or
can obtain more detailed contextual information.
Third, the optimal level is also likely to vary according to the nature of the policy process.
An extensive participatory policy design process that engages key stakeholders and beneficiaries
will elicit more local information and is more likely to lead to useful adaptations than a quick
decision made by an individual policymaker (Parker et al 2008, Leviton and Trujillo 2017). For
instance, although the World Bank based the design of BINP on its successful program in Tamil
Nadu, the Bank’s own evaluation found that the view that “project design and implementation
should have sought to broaden the target audience for its nutritional messages. . . was expressed
by BINP fieldworkers and women themselves in project areas during field visits” (World Bank
2005b, 9). When such participatory processes are not practically or politically feasible, or when
policymakers do not have time to gather extensive data on actual contextual characteristics
and are thus forced to rely on their own knowledge, policymakers should weight evidence from
elsewhere relatively more heavily and thus usually make fewer adaptations.
Finally, making appropriate adaptations to a policy requires understanding not only of the
context but also of the policy’s mechanism, since what matters for impact is the interaction
27
between mechanism and context. While better understanding of the mechanism is unrelated to
the optimal level of adaptations to a policy (unlike better contextual information), one would
expect it to lead to more successful adaptations. This is an area in which bureaucratic expertise
and research - in particular high-quality evaluations or systematic reviews that are able to
trace mechanisms - can be especially useful. Effective policy adaptation thus stems not just
from contextual knowledge, but from its combination with rigorous evidence and professional
judgment.
5 Conclusion
As the harm caused by the neuropsychiatric interaction between the HIV drug efavirenz and
the rare genetic variant common in Zimbabwe’s population became evident, some of the same
scientists who predicted the problem designed a strategy to address it. “[T]he current ‘one
size fits all’ [efavirenz] dose strategy in sub-Saharan Africa needs to be carefully reevaluated by
considering integration of an individualized therapeutic approach” that combines individualized
testing, monitoring, and dosing adjustment (Masimirembwa et al 2016, 4). As one of the
scientists, Collen Masimirembwa, stated: “It’s not a bad drug. We just know it can be improved
in Africa” (Nordling 2017, 20).
Just as the spread of precision medicine promises to move medical treatment beyond one-
size-fits-all recommendations, so too is it necessary for impact evaluators and policymakers
to find ways to make evidence-based policymaking more responsive to the particularities of
specific contexts. While impact evaluations and systematic reviews provide excellent starting
points for doctors and policymakers alike, even before the advent of precision medicine actual
medical practice has always required doctors to combine rigorous research evidence with their
individual expertise and judgment for each case (Deaton 2010). While the rapidly growing
external validity literature has focused largely on the generalizability of a policy from a specific
context, the relevant question for policymakers is the applicability of evidence to their specific
context. This requires an understanding of the interactions between a policy’s theory of change,
as supported by contextual assumptions, and the actual characteristics of the context to which
a policy is being transported.
This article has introduced mechanism mapping as a flexible and conceptually simple diag-
nostic tool to help policymakers assess the fit of evidence-based policies with their own contexts.
28
Similarly, the same basic procedure can be applied to simple policies with a linear theory of
change and only one outcome of interest, or to complex policies with numerous potential out-
comes and feedback loops. This diagnostic process also feeds directly into the identification of
potential policy adaptations, by identifying the specific features of a policy that are likely to be
problematic in the new context.
The rapid growth of evidence-based policy has created a rich pool of rigorous impact evalua-
tions, and a growing literature on external validity has provided theoretical and empirical tools
for generalizing these results beyond their original context. However, this paper has highlighted
the paucity of guidance to policymakers on how to solve the “last mile” problems of policy
transportation and scale-up: the need to bridge the inevitable gap between the best evidence
available from other contexts and the particularities of their own context, and whether and
how to adapt evidence-based policies to better fit these contextual idiosyncrasies. While the
mechanism mapping framework is potentially useful in this regard, there is a need for more
and better research on key questions like how policymakers tend to make these contextual-fit
assessments and adaptations in practice, how they update their beliefs and make adaptations
based on this, and how research and institutional structures can improve their judgment in these
regards. Questions like these represent the basis for an intellectually rich and policy-relevant
agenda for researchers and practitioners alike.
Notes
1I use the terms policy, intervention, and program interchangeably throughout, since the distinctions between
them are not relevant for this paper’s purposes.
2For a more technically precise discussion of these issues, see Deaton and Cartwright (2016).
3Random sampling is of course only sufficient for representativeness with large enough sample sizes, but as
stated above, this article abstracts from this and other statistical issues.
4While there are numerous different disciplinary and institutional approaches and terminologies associated
with elaborating theories of change (e.g. DFID 2012, De Silva et al 2014), the aim of this paper is not to
adjudicate the debate between these various approaches, nor to suggest a best practice for doing so. Rather, this
paper takes a simple approach in order to focus on the core concepts. This approach can equally be applied to
theories of change written in different formats.
5The Cochrane and Campbell Collaborations and the UK government’s What Works Network are the most
prominent repositories of systematic reviews and meta-analyses.
6Meta-analysis is intended to capture real variation from differences in context as well as random statistical
variation from chance; as discussed previously, the latter is outside the scope of this paper.
29
7This point is not meant to caricature the views of authors of systematic reviews, most of whom have ap-
propriately nuanced views of how systematic reviews should be used by policymakers, but simply to clarify the
conceptual limitations of the “headline” average treatment effect that readers often focus on.
8More precisely, the correlation of observables with unobservables may differ across contexts. A further
limitation of sub-group analysis is of course the issue of limited power and the risk of false positives arising from
multiple testing (Petticrew et al 2012), but such statistical issues are beyond the scope of this paper.
9Pawson and Manzano-Santaella (2012) criticize the quality of many evaluations that claim to be realist,
particularly for their over-reliance on descriptive, qualitative data.
10The policy memo accompanying this paper (Williams 2017) presents a simplified “how-to” five-step guide to
mechanism mapping for policymakers.
11This example is based on Cartwright and Hardie’s excellent exposition (2014, 80-84), as well as on Save the
Children (2003), White (2005), World Bank (2005a), and World Bank (2005b). This article’s discussion of TINP
and BINP and their contexts is of course simplified for clarity and brevity.
References
Allcott, Hunt. 2015. “Site Selection Bias in Program Evaluation.” Quarterly Journal of
Economics : 1117-1165.
Andrews, Isaiah, and Emily Oster. 2018. “Weighting for External Validity.” NBER Working
Paper 23826.
Angrist, Noam. 2017. “An Application of the Jump from Internal to External Validity: Trans-
porting an HIV Prevention Intervention from Kenya to Botswana, and Testing Two Scale
Models Across Key Parameters of Heterogeneity.” Working Paper, 26 April.
Angrist, Joshua, and Ivan Fernandez-Val. 2010. “Extrapolate-ing: External Validity and Overi-
dentification in the LATE Framework.” NBER Working Paper 16566, December.
Banerjee, Abhijit, Rukmini Banerji, James Berry, Esther Duflo, Harini Kannan, Shobhini Muk-
erji, Marc Shotland, and Michael Walton. 2016a. “From Proof of Concept to Scalable Policies:
Challenges and Solutions, with an Application.” Mimeo, September.
Banerjee, Abhijit, Sylvain Chassang, and Erik Snowberg. 2016b. “Decision Theoretic Ap-
proaches to Experiment Design and External Validity.” Mimeo, July.
Barzelay, Michael. 2007. “Learning from Second-Hand Experience: Methodology for Extrapolation-
Oriented Case Research.” Governance 20(3): 521-43.
30
Bates, Mary Ann, and Rachel Glennerster. 2017. “The Generalizability Puzzle.” Stanford So-
cial Innovation Review, Summer, https://ssir.org/articles/entry/the generalizability puzzle,
accessed 22 June, 2017.
Bold, Tessa, Mwangi Kimenyi, Germano Mwabu, Alice Ng’ang’a, and Justin Sandefur. 2016.
“Experimental Evidence on Scaling Up Education Reforms in Kenya.” Mimeo, November.
Bonell, Chris, Farah Jamal, G.J. Melendez-Torres, and Steven Cummins. 2014. “‘Dark Logic’:
Theorising the Harmful Consequences of Public Health Interventions.” Journal of Epidemiol-
ogy and Community Health 69: 95-98.
Cameron, Lisa, and Manisha Shah. 2017. “Scaling Up Sanitation: Evidence from an RCT in
Indonesia.” IZA Discussion Paper 10619, March.
Cartwright, Nancy, and Jeremy Hardie. 2014. Evidence-Based Policy: A Practical Guide to
Doing It Better. Oxford: Oxford University Press.
Castro, Felipe González, Manual Barrera Jr., and Lori K. Holleran Steiker. 2010. “Issues
and Challenges in the Design of Culturally Adapted Evidence-Based Interventions.” Annual
Review of Clinical Psychology 6: 213-39.
Christensen, Garret, and Edward Miguel. 2016. “Transparency, Reproducibility, and the Cred-
ibility of Economics Research.” NBER Working Paper 22989.
Davies, Rick. 2004. “Scale, Complexity, and the Representation of Theories of Change.” Eval-
uation 10(1): 101-121.
De Silva, M.J., E. Breuer, L. Lee, L. Asher, N. Chowdhary, C. Lund, and V. Patel. 2014.
“Theory of Change: A Theory-Driven Approach to Enhance the Medical Research Council’s
Framework for Complex Interventions.” Trials 15(1): 267-78.
Deaton, Angus. 2010. “Instruments, Randomization, and Learning about Development.” Jour-
nal of Economic Literature 48: 424-55.
Deaton, Angus, and Nancy Cartwright. 2016. “Understanding and Misunderstanding Random-
ized Controlled Trials.” NBER Working Paper 22595, September.
Diamond, Adele, W. Steven Barnett, Jessica Thomas, and Sarah Munro. 2007. “Preschool
Program Improves Cognitive Control.” Science 318(5855): 1387-88.
31
DFID. 2012. “Review of the Use of ‘Theory of Change’ in International Development.” Review
Report, April.
Evans, David, and Anna Popova. 2015. “What Really Works to Improve Learning in Developing
Countries? An Analysis of Divergent Findings in Systematic Reviews.” World Bank Policy
Research Working Paper 7203.
Farran, Dale, and Sandra Jo Wilson. 2014. “Achievement and Self-Regulation in Pre-Kindergarten
Classrooms: Effects of the Tools of the Mind Curriculum.” Mimeo, Peabody Institute, July.
Gardner, Frances, and Paul Montgomery. 2015. “Transporting Evidence-Based Parenting Pro-
grams for Child Problem Behavior (Age 3-10) Between Countries: Systematic Review and
Meta-Analysis.” Journal of Clinical Child and Adolescent Psychology : 1-14.
Gechter, Michael. 2016. “Generalizing the Results from Social Experiments: Theory and
Evidence from Mexico and India.” Mimeo, December.
Gertler, Paul, Sebastian Martinez, Patrick Premand, Laura Rawlings, and Christel Vermeersch.
2016. “Impact Evaluation in Practice.” Second edition, Washington DC: World Bank.
Greenhalgh, Trisha, Fraser Macfarlane, Liz Steed, and Robert Walton. 2016. “What Works for
Whom in Pharmacist-Led Smoking Cessation Support: Realist Review.” BMC Medicine 14:
209-24.
Hasson, Henna, Knut Sundell, Andreas Beelmann, and Ulrica von Thiele Schwarz. 2014. “Novel
Programs, International Adoptions, or Contextual Adaptations? Meta-analytical Results from
German and Swedish Intervention Research.” BMC Health Services Research 14(Suppl. 2):
O32.
Hawe, Penelope. 2015. “Lessons from Complex Interventions to Improve Health.” Annual
Review of Public Health 36: 307-23.
Henrich, Joseph, Steven J. Heine, and Ara Norenzayan. 2010. “The Weirdest People in the
World?” Behavioral and Brain Sciences 33: 61-135.
HM Treasury. 2011. “The Magenta Book: Guidance for Evaluation.” April.
Jagosh, Justin, Paula L. Bush, Jon Salsberg, Ann C. Macaulay, Trish Greenhalgh, Geoff Wong,
Margaret Cargo, Lawrence W. Green, Carol P. Herbert, and Pierre Pluye. 2015. “A Realist
32
Evaluation of Community-Based Participatory Research: Partnership Synergy, Trust Building,
and Related Ripple Effects.” BMC Public Health 15: 725-35.
Kowalski, Amanda E. 2016. How to Examine External Validity Within an Experiment. Mimeo,
August.
Kristjansson, Elizabeth, Damian Francis, Selma Liberato, Maria Benkhalti Jandu, Vivian
Welch, Malek Batal, Trish Greenhalgh, Tamara Rader, Eamonn Noonan, Beverley Shea, Laura
Janzen, George Wells, and Mark Petticrew. 2015. “Food Supplementation for Improving the
Physical and Psychosocial Health of Socio-economically Disadvantaged Children Aged Three
Months to Five Years: A Systematic Review.” Campbell Systematic Reviews, 4 May.
Leijten, Patty, G.J. Melendez-Torres, Wendy Knerr, and Frances Gardner. 2016. “Transported
Versus Homegrown Parenting Interventions for Reducing Disruptive Child Behavior: A Mul-
tilevel Meta-Regression Study.” Journal of the American Academy of Child and Adolescent
Psychiatry 55(7): 610-17.
Leviton, Laura C. 2017. “Generalizing about Public Health Interventions: A Mixed-Methods
Approach to External Validity.” Annual Review of Public Health 38: 371-91.
Leviton, Laura, and Mathew Trujillo. 2017. “Interaction of Theory and Practice to Assess
External Validity.” Evaluation Review 41(5): 436-471.
Low, Hamish, and Costas Meghir. 2017. “The Use of Structural Models in Econometrics.”
Journal of Economic Perspectives 31(2): 33-58.
Ludwig, Jens, Jeffrey Kling, and Sendhil Mullainathan. 2011. “Mechanism Experiments and
Policy Evaluations.” Journal of Economic Perspectives 25(3): 17-38.
Marchal, Bruno, Sara van Belle, Josefien van Olmen, Tom Hoeré, and Guy Kegels. 2012. “Is
Realist Evaluation Keeping Its Promise? A Review of Published Empirical Studies in the
Field of Health Systems Research.” Evaluation 18(2): 192-212.
Masimirembwa, Collen, and Collet Dandara. 2016. “Rolling out Efavirenz for HIV Precision
Medicine in Africa: Are We Ready for Pharmacovigilance and Tackling Neuropsychiatric
Adverse Effects?” OMICS: A Journal of Integrative Biology 20(10): 1-6.
Miguel, Edward, and Michael Kremer. 2004. “Worms: Identifying Impacts on Education and
Health in the Presence of Treatment Externalities.” Econometrica 72(1): 159-217.
33
Moore, G.F., S. Audrey, M. Barker, L. Bond, C. Bonell, W. Hardeman, L. Moore, A. O’Cathain,
T. Tinati, D. Wight, and J. Baird. 2015. “Process Evaluationg of Complex Interventions:
Medical Research Council Guidance.” British Medical Journal : 350.
Muralidharan, Karthik, and Paul Niehaus. 2016. “Experimentation at Scale.” Mimeo, Septem-
ber.
Nordling, Linda. 2017. “Putting Genomes to Work in Africa.” Nature 544: 20-22.
Nyakutira, Christopher, Daniel Röshammar, Emmanuel Chigutsa, Prosper Chonzi, Michael
Ashton, Charles Nhachi, and Collen Masimirembwa. 2008. “High Prevlance of the CYP2B6
516G→T(*6) Variant and Effect on the Population Pharmacokinetics of Efavirenz in HIV/AIDS
Outpatients in Zimbabwe.” European Journal of Clinical Pharmacology 64: 357-65.
Parker, Edith, Barbara Israel, Thomas Robins, Graciela Mentz, Xihong Lin, Wilma Brakefield-
Caldwell, Erminia Ramirez, Katherine Edgren, Maria Salinas, and Toby Lewis. 2008. “Eval-
uation of Community Action Against Asthma: A Community Health Worker Intervetnion to
Improve Children’s Asthma-Related Health by Reducing Household Environmental Triggers
for Asthma.” Health Education and Behavior 35(3): 376-395.
Pawson, Ray, and Ana Manzano-Santaella. 2012. “A Realist Diagnostic Workshop.” Evaluation
18(2): 176-91.
Pawson, Ray, and Nicholas Tilley. 1997. Realistic Evaluation. London: SAGE Publications.
Petticrew, Mark, Peter Tugwell, Elizabeth Kristjansson, Sandy Oliver, Erin Ueffing, and Vivian
Welch. 2012. “Damned if You Do, Damned if You Don’t: Subgroup Analysis and Equity.”
Journal of Epidemiology and Community Health 66(1): 95-98.
Pritchett, Lant. 2017. “‘The Evidence’ About ‘What Works’ in Education: Graphs to Illustrate
External Validity and Construct Validity.” RISE Insight, June.
Pritchett, Lant, Salimah Samji, and Jeffrey Hammer. 2013. “It’s All About MeE: Using Struc-
tured Experiential Learning (“e”) to Crawl the Design Space.” Center for Global Development
Working Paper 322, April.
Pritchett, Lant, and Justin Sandefur. 2015. “Learning from Experiments when Context Mat-
ters.” American Economic Review: Papers and Proceedings 105(5): 471-75.
34
Rajman, Iris, Laura Knapp, Thomas Morgan, and Collen Masimirembwa. 2017. “African
Genetic Diversity: Implications for Cytochrome P450-mediated Drug Metabolism and Drug
Development.” EBioMedicine 17: 67-74.
Ravallion, Martin. 2009. “Evaluation in the Practice of Development.” World Bank Research
Observer 24: 29-53.
Rogers, Patricia. 2008. “Using Programme Theory to Evaluate Complicated and Complex
Aspects of Interventions.” Evaluation 14(1): 29-48.
Save the Children. 2003. “Thin on the Ground: Questioning the Evidence Behind World
Bank-funded Community Nutrition Projects in Bangladesh, Ethiopia, and Uganda.” Policy
Report.
SC4CCM. n.d. “SC4CCM Project Theory of Change.” Available at: http://sc4ccm.jsi.com/
emerging-lessons/theory-of-change/
Vivalt, Eva. 2017. “How Much Can We Generalize from Impact Evaluations?” Mimeo.
What Works Network. 2014. “What Works? Evidence for Decision Makers.” Presentation,
https://www.gov.uk/government/uploads/system/uploads/attachment data/file/378038/What
works evidence for decision makers.pdf, accessed 19 May, 2017.
White, Howard. 2005. “Comment on Contributions Regarding the Impact of the Bangladesh
Integrated Nutrition Project.” Health Policy and Planning 20(6): 408-11.
Williams, Martin J. 2017. “External Validity and Policy Adaptation: A Five-step Guide to
Mechanism Mapping.” Blavatnik School of Government Policy Memo, available at: http:
//www.martinjwilliams.com/research.
Wong, Geoff, Gill Westhorp, Ana Manzano, Joanne Greenhalgh, Justin Jagosh, and Trish
Greenhalgh. 2016. “RAMESES II Reporting Standards for Realist Evaluations.” BMC
Medicine 14: 96-113.
World Bank. 2005a. “The Bangladesh Integrated Nutrition Project: Effectiveness and Lessons.”
Bangladesh Development Series Paper No. 8, December.
World Bank. 2005b. “Project Performance Assessment Report: Bangladesh Integrated Nutri-
tion Project.” Report No. 32563, 13 June.
35
World Bank. 2015. “World Development Report 2015: Mind, Society, and Behavior.” Chapter
11.
Young 1ove. 2016. “Theory of Change: Relative Risk Information Campaign.” 25 April.
36
Appendix A Example of a Complicated Theory of Change
Figure A1: Theory of Change: Supply Chains 4 Community Case Management (SC4CCM)
55 . C
CM P
ol ic
y –
M O
H c
om m
it m
en t
to c
hi ld
s ur
vi va
l, or
ga ni
za ti
on al
s up
po rt
a nd
s tr
uc tu
re , s
ta ff
in g
of C
H W
s ex
is ts
53 . S
tr at
eg y
or P
la n
to a
ch ie
ve C
om m
od it
y Se
cu ri
ty fo
r CC
M p
ro du
ct s
ex is
ts
14 . S
up po
rt iv
e su
pe rv
is io
n of
CH
W s
w it
h SC
c om
po ne
nt is
pe
rf or
m ed
r eg
ul ar
ly
5. C
H W
s ar
e m
ot iv
at ed
t o
pe rf
or m
t he
ir r
ol es
in
t he
C CM
p ro
du ct
s up
pl y
ch ai
n
12 . F
ee db
ac k
is
co m
m un
ic at
ed
to C
H W
s
38 . S
up er
vi so
rs a
re
tr ai
ne d
in s
up po
rt iv
e su
pe rv
is io
n an
d th
e pr
oc ed
ur es
a nd
pr
oc es
se s
fo r
CC M
p ro
du ct
su
pp ly
c ha
in
37 .
Tr an
sp or
ta ti
on
an d
ot he
r re
so ur
ce s
av ai
la bl
e to
co
nd uc
t su
pe rv
is io
n 46 . S
tr ea
m lin
ed
pr oc
ed ur
es fo
r pr
ov id
in g
fe ed
ba ck
a nd
s up
er vi
si on
to
th e
CH W
s ex
is t a
nd a
re
do cu
m en
te d
39 . S
ta ff
re
sp on
si bl
e fo
r pr
ov id
in g
lo gi
st ic
s fe
ed ba
ck r
ep or
ts
ar e
t ra
in ed
a nd
pr
od uc
e re
po rt
s
52 . G
ov er
nm en
t de
si gn
at es
a b
ud ge
t lin
e /
o r
do no
r pr
ov id
es s
uf fic
ie nt
fu nd
in g
fo r
pr og
ra m
a nd
t ho
se fu
nd s
ar e
al lo
ca te
d an
d di
sb ur
se d
w he
n ne
ed ed
44 . S
tr ea
m lin
ed p
ro ce
du re
s fo
r or
de ri
ng , r
ep or
ti ng
, in
ve nt
or y
co nt
ro l,
st or
ag e
an d
di sp
os al
o f e
xp ir
ed /
da
m ag
ed h
ea lth
p ro
du ct
s ex
is t a
nd a
re d
oc um
en te
d
25 . C
H W
Su
pe rv
is or
s kn
ow o
f a nd
di
st ri
bu te
in
ce nt
iv es
w he
n ap
pr op
ri at
e
47 . D
ef in
ed C
H W
in
ce nt
iv e
sy st
em
th at
in cl
ud es
s up
pl y
ch ai
n pe
rf or
m an
ce
ex is
ts a
nd is
do
cu m
en te
d
13 . C
H W
s ro
ut in
el y
co lle
ct a
nd
re po
rt ti
m el
y,
ac cu
ra te
lo
gi st
ic s
da ta
33 . S
ta ff
a t
re su
pp ly
po
in t/
s ar
e tr
ai ne
d in
pr
oc ed
ur es
a nd
p ro
ce ss
es fo
r CC
M p
ro du
ct s
up pl
y ch
ai n
19 . S
ec ur
e an
d su
it ab
le s
to ra
ge
co nt
ai ne
rs o
r sh
el vi
ng fo
r CC
M
pr od
uc ts
a re
pr
oc ur
ed w
he re
ne
ed ed
34 . C
H W
s ar
e tr
ai ne
d in
pr
oc ed
ur es
a nd
pr
oc es
se s
fo r
CC M
p ro
du ct
s up
pl y
ch ai
n
17 . L
M IS
fo rm
s or
ot
he r
da ta
co
lle ct
io n
to ol
s ar
e av
ai la
bl e
fo r
CH W
s
CH W
s ha
ve u
sa bl
e an
d qu
al it
y m
ed ic
in es
a va
ila bl
e w
he n
ne ed
ed fo
r ap
pr op
ri at
e tr
ea tm
en t o
f c om
m on
c hi
ld ho
od
ill ne
ss es
*
Si ck
c hi
ld re
n re
ce iv
e ap
pr op
ri at
e tr
ea tm
en t
fo r
co m
m on
c hi
ld ho
od il
ln es
se s*
2. C
H W
s, or
p er
so n
re sp
on si
bl e
fo r
CH W
r es
up pl
y,
kn ow
h ow
, w he
re , w
ha t,
w he
n an
d ho
w m
uc h
of
ea ch
p ro
du ct
t o
re qu
is it
io n
or r
es up
pl y
an d
ac t a
s ne
ed ed
3. C
H W
s ha
ve a
de qu
at e
st or
ag e:
co rr
ec t
co nd
iti on
s, s
ec ur
it y
an d
ad eq
ua te
s pa
ce .
Sy st
em P
er fo
rm an
ce
To ol
s /
R es
ou rc
es
Kn ow
le dg
e /
Sk ill
s
Pr oc
es se
s
Fu nd
in g
Co m
m un
ic at
io n
/ In
fo rm
at io
n
Po lic
y /
St ra
te gy
G oa
ls a
nd O
ut co
m es
GF
D
C
1 51
. M an
uf ac
tu re
rs h
av e
in fo
rm at
io n
on e
st im
at ed
de
m an
d fo
r ac
ce pt
ab le
fo
rm ul
at io
ns o
f C CM
pr
od uc
ts , i
nc lu
di ng
v ia
bl e
pr ic
e po
in ts
A
49 . S
uf fic
ie nt
fu nd
in g
fo r
pr oc
ur in
g co
m m
od it
ie s
ar e
av ai
la bl
e
50 . S
tr ea
m lin
ed p
ro ce
du re
s fo
r ro
ut in
e qu
an ti
fic at
io n
an d
pi pe
lin e
m on
ito ri
ng
ex is
t an
d ar
e do
cu m
en te
d
48 . H
ig h-
qu al
it y,
c hi
ld -a
nd
su pp
ly c
ha in
- fr
ie nd
ly C
CM
pr od
uc ts
a re
a va
ila bl
e fr
om
gl ob
al , r
eg io
na l,
or lo
ca l
m ar
ke tp
la ce
30 . H
ig h-
qu al
ity , c
hi ld
- an
d su
pp ly
c ha
in fr
ie nd
ly
CC M
p ro
du ct
s ar
e on
N
EM L
27 . P
ro du
ct
sp ec
ifi ca
tio ns
a re
de
te rm
in ed
fo r
pr oc
ur em
en t
41 . C
CM p
ro du
ct
se le
ct io
n is
b as
ed
on s
ta nd
ar d
tr ea
tm en
t al
go ri
th m
s an
d su
pp ly
c ha
in
co ns
id er
at io
ns
40 . R
eg is
tr at
io n
of n
ew h
ig h-
qu al
ity , c
hi ld
- an
d su
pp ly
ch
ai n
fr ie
nd ly
CC
M p
ro du
ct s
oc cu
rs
16 . T
oo ls
a nd
r es
ou rc
es n
ee de
d to
im
pl em
en t
pr oc
ed ur
es a
re p
ro vi
de d
29 . R
ou ti
ne
qu an
ti fic
at io
n fo
r pr
oc ur
em en
t is
d on
e
28 . F
un ds
a re
a llo
ca te
d fo
r pr
oc ur
em en
t ba
se d
on q
ua nt
ifi ca
ti on
a nd
di
sb ur
se d
re gu
la rl
y
8. R
es up
pl y
po in
t/ s
ha
ve a
de qu
at e
st or
ag e:
co rr
ec t c
on di
tio ns
, se
cu ri
ty a
nd
ad eq
ua te
s pa
ce
9. S
ta ff
a t
al l r
es up
pl y
po in
ts a
re m
ot iv
at ed
an
d pe
rf or
m t
he ir
r ol
es
in t
he C
CM p
ro du
ct
su pp
ly c
ha in
a s
ex pe
ct ed
43 . S
ta ff
re
sp on
si bl
e fo
r qu
an ti
fic at
io n
a re
tr
ai ne
d
42 . V
is ib
ili ty
o f
pr og
ra m
, s up
pl y
an d
de m
an d
da ta
ex
is ts
32 . M
ec ha
ni sm
fo r
co m
m un
ic at
io n
be tw
ee n
st ak
eh ol
de rs
e xi
st s
31 . F
un di
ng a
nd
pr oc
ur em
en t
cy cl
es a
re
al ig
ne d
15 . T
im el
y pr
oc ur
em en
t of
q ua
lit y
CC M
p ro
du ct
s oc
cu rs
6. A
de qu
at e
qu an
ti ti
es o
f C CM
pr
od uc
ts a
re
av ai
la bl
e at
a ll
di st
ri bu
ti on
p oi
nt s
in c
ou nt
ry
1. N
ec es
sa ry
, u sa
bl e,
q ua
lit y
CC M
p ro
du ct
s ar
e av
ai la
bl e
at C
H W
re su
pp ly
p oi
nt /s
24 . S
up er
vi si
on
to ol
s ar
e av
ai la
bl e
54 . U
ni t/
s ex
is ts
t ha
t pl
an , m
an ag
e an
d co
or di
na te
re so
ur ce
s, s
er vi
ce s,
co
nt ra
ct s,
b ud
ge t,
p er
so nn
el e
tc .
E
B
26 . C
H W
s ha
ve
kn ow
le dg
e of
in ce
nt iv
es
10 . T
oo ls
a nd
re
so ur
ce s
ne ed
ed
to im
pl em
en t
pr oc
ed ur
es a
re
pr ov
id ed
18 . A
pp ro
pr ia
te
an d
se cu
re
st or
ag e
sp ac
e fo
r CC
M
pr od
uc ts
is
av ai
la bl
e
45 . S
tr ea
m lin
ed t
ra ns
po rt
at io
n pr
oc ed
ur es
f or
m ai
nt ai
ni ng
ve
hi cl
es a
nd d
is tr
ib ut
io n
/ co
lle ct
io n
of g
oo ds
e xi
st a
nd
ar e
do cu
m en
te d
36 . D
ri ve
rs a
re t
ra in
ed in
S C
tr an
sp or
t, a
pp ro
pr ia
te p
ro du
ct
ha nd
lin g
& m
ai nt
en an
ce
pr oc
ed ur
es a
nd t
ra ns
po rt
at io
n sc
he du
le
20 . T
ra ns
po rt
is
m ai
nt ai
ne d
on
a re
gu la
r ba
si s
fo r
pu rp
os e
of
CC M
p ro
du ct
di
st ri
bu tio
n
35 . S
pa re
pa
rt s
ar e
av ai
la bl
e fo
r m
ai nt
en an
ce
of t
ra ns
po rt
4. G
oo ds
a re
r ou
tin el
y tr
an sp
or te
d be
tw ee
n
re su
pp ly
p oi
nt s
an d
CH W
s
22 . A
de qu
at e
tr an
sp or
t
pr oc
ur ed
o r
hi re
d w
he n
ne ce
ss ar
y fo
r pu
rp os
e of
CC
M p
ro du
ct
di st
ri bu
ti on
11 . R
el ia
bl e,
t im
el y
an d
ap pr
op ri
at e
tr
an sp
or t
is a
va ila
bl e
to d
is tr
ib ut
e or
c ol
le ct
go
od s
be tw
ee n
re su
pp ly
p oi
nt a
nd C
H W
s
23 . S
ki lle
d dr
iv er
s ar
e av
ai la
bl e
w he
n an
d w
he re
ne
ed ed
21 . F
ue l
pr oc
ur ed
w he
n ne
ce ss
ar y
fo r
pu rp
os e
of
CC M
p ro
du ct
di
st ri
bu ti
on
7. T
ra ns
po rt
is
av ai
la bl
e to
di
st ri
bu te
o r
co lle
ct
CC M
p ro
du ct
s as
re
qu ir
ed to
re
su pp
ly p
oi nt
s
Le ad
o n
ac ti
vi tie
s
In vo
lv ed
b ut
n ot
le ad
In flu
en ce
t hr
ou gh
a dv
oc ac
y
Source: Reproduced from SC4CCM (n.d.).
37
Appendix B Example Theory of Change with Competing Mech-
anisms and Potential Negative Outcomes
One important consideration in the policy design process is the uncertainty of how a policy will
work in practice - i.e., the potential for multiple competing mechanisms. Some of these potential
impact pathways may lead to positive outcomes while others lead to null or even negative
outcomes. In its simplest form, the process of mechanism mapping can help policymakers
understand the plausibility of their intended mechanism. However, the same logic can be
applied to assess the plausibility of multiple mechanisms, some of which may lead to negative
rather than positive outcomes.
To illustrate this, consider the theory of change constructed in 2016 by the Botswanan NGO
Young 1ove for its planned transportation of the “sugar daddies” informational intervention on
teenage HIV infection risk in Kenya (Dupas 2011) in Figure A2. It considers the potential
for eight different competing mechanisms, half of which would lead to unintended outcomes in
terms of HIV rates, pregnancy incidence, and school attendance. Although this theory of change
does not exactly follow the suggested mechanism mapping structure, it does explicitly consider
the key assumptions underlying each step of the eight potential mechanisms. Many of these
assumptions are framed as empirical questions on which additional descriptive data could be
collected to help establish the plausibility of each of these mechanisms, and indeed the NGO did
collect descriptive data on many of these assumptions prior to beginning the trial (Noam Angrist,
personal communication, 19 February 2018). While Young 1ove’s example shows that the basic
concepts underlying mechanism mapping are intuitive and can be implemented even without
explicit reference to the tool, the structured guidance in undertaking this process presented in
this paper may nonetheless be useful for other organizations and policymakers.
This example also illustrates the feasibility of the mechanism mapping process for policy-
makers in the Global South and elsewhere. Young 1ove developed this theory of change through
a six-month collaborative process led by its Batswana staff, which reportedly helped increase
understanding of the program, surfaced key assumptions, and led to adaptations to the original
design of the policy in Kenya (Angrist, personal communication, 19 February 2018).
38
Figure A2: Theory of Change: Young 1ove
Inputs
Botswana Partners &
Ministry Buy-in
Curriculum
Facilitator Training
Funding
Support Staff
Outputs
Facilitators reach the in-school
youth and deliver the curriculum
Intermediate outcomes
Youth are able to negotiate condom use with an age mate, youth believe that condoms reduce the risk of
HIV/Pregnancy, youth have access to condoms, youth have comprehensive knowledge on correct & consistent condom
use - [**], Age mates advocate for protection as a function of perceived risk, Youth fear pregnancy with an age mate
who is unable to provide
Youth delay sexual debut, or refrain from sex
Condom use during sex increases
Age-mates are 9 times less risky (HIV), Age-mates are less likely to have acute HIV (40x), consistent and correct condom
use
School attendance increases
Anal sex – receptive anal sex is 18 times more risky than vaginal sex, young adults are 6 times less likely to have HIV, acuteness is ambiguous (MCP lower, young adults are more likely to default), anal sex will not result in pregnancy
Oral sex – Less likely to result in pregnancy/HIV
Ultimate outcomes
[**]Youth fear pregnancy with an older partner [socially unacceptable], Older partners have greater incentive to prevent
pregnancy [marital status, MCP, social stigma, consent laws]
RELATIVE RISK INFORMATION CAMPAIGN – Intended outcomes
Youth shift from older partners to age mates
School attendance increases
HIV: youth aren’t exposed to social stigma at school, youth have time and resources, youth’s health is not
compromised – [****] Pregnancy: All of the above + Schools are not
suspending youth from schools, youth don’t have competing preferences
Youth continue to have sex with older partners
Youth were getting income from older partners or young adults, youth lose the income they were getting in the shift to age-mates, youth perceive school as an add value to income generation, youth need or want
long-term income, school-going peer effect: girls want to spend more time at school with partners
School attendance increases
[**]youth & their young adult partners have (financial & social) access to condoms
HIV: [****]
Pregnancy: [****] + Schools are not suspending youth from schools, youth don’t have competing preferences
HIV: [****] Pregnancy: [****] + Schools are not suspending youth from schools, youth don’t have competing preferences
Youth have resources to attend school
Condom use during sex increases Pregnancy | HIV rates decrease
Youth shift from older partners to young adults
Youth have alternative sex [oral, anal]
Youth intended or had older sexual partners, Youth have agency over partner selection, youth internalize that higher HIV risk in older partners means own personal risk, youth’s priors on HIV rates for older sexual partners are lower than actual risk – [*], youth perceive the cost of HIV as higher
than the benefits of going with older partners,, youth perceive age-mates as the safest age group, youth perceive
agency over condom usage with age mates
Youth believe alternative sex is relatively less risky, alternative sex is more accessible and has higher agency - [***], youth fear pregnancy with an age mate who is unable to provide
Pregnancy rates decrease| HIV rates decrease [P > 75%]
Anal sex – Although age-mates are 9 times less likely to have HIV, receptive anal sex is 18 times more risky than vaginal sex, anal sex will not result in pregnancy, age-mates less likely to have acute HIV (40x)
Oral sex – Less likely to result in pregnancy/HIV
HIV: [****] Pregnancy: [****] + Schools are not suspending youth from schools, youth don’t have competing preferences
School attendance increases
Condom use during sex increases Pregnancy decreases| HIV rates decrease
[*]youth perceive the cost-benefit of young adults (the ‘sweet and sour spot’) as optimized over age-mates & older
partners, youth perceive agency over condom usage with young adults
School attendance increases
Youth receive “income sweet spot” – enough income to have resources to attend school but not enough to present a long-term source of income, youth need or want long-term income (threshold effect, longevity)
Pregnancy | HIV rates decrease
Consistent and correct condom use
School attendance increases
School attendance increases [P > 50%)Youth have alternative sex [oral, anal]
[***]Youth fear pregnancy with an older partner [socially unacceptable], youth believe alternative sex is relatively
less risky, Older partners have greater incentive to prevent pregnancy [marital status, MCP, social stigma, consent
laws]
Pregnancy rates decrease| HIV rates ambiguous
Anal sex: Receptive anal sex is 18 times more risky than vaginal sex, HIV acuteness is ambiguous (high MCP, more adherence), anal sex will not result in pregnancy
Oral sex: Less likely to result in pregnancy/HIV
HIV: see anal sex ambiguity Pregnancy: All of the above + Schools are not
suspending youth from schools, youth don’t have competing preferences
Youth intended or had sex with older partners, Youth have no agency over partner choice, Youth renegotiate incentives,
youth prioritize benefits conditional on minimizing risk, youth perceive agency over condom usage with older
partners
School attendance increases
School attendance increases
Pregnancy rates decrease| HIV rates decrease
The dominant mode of HIV transmission is sexual intercourse
Youth are attentive,
Facilitators are credible
& charismatic Facilitators cover the key points & activities
in the curriculum
Youth were getting income from older partners, youth lose the income they were getting in the shift to
celibacy, youth perceive school as an add value to income generation, youth need or want long-term
income, girls are not distracted by sexual partner (s)
School attendance increases [P > 50%]Youth have alternative sex [oral, anal]
[***]
Pregnancy rates decrease| HIV rates ambiguous
Anal sex – receptive anal sex is 18 times more risky than vaginal sex, young adults are 6 times less likely to have HIV, acuteness is ambiguous (MCP lower, young adults are more likely to default), anal sex will not result in pregnancy
Oral sex – Less likely to result in pregnancy/HIV
HIV: see anal sex ambiguity Pregnancy: All of the above + Schools are not
suspending youth from schools, youth don’t have competing preferences
School attendance increases
HIV: [****] Pregnancy: [****] + Schools are not suspending youth from schools, youth don’t have competing preferences
[*]Youth associate any sex with any type of partner as riskier than the benefits
Youth are able to identify men in
their 30s & 40s as having
high rates of HIV
Source: Reproduced from Young 1ove (2016).
39
Inputs
Botswana Partners &
Ministry Buy-in
Curriculum
Facilitator Training
Funding
Support Staff
Outputs
Facilitators reach the in-school
youth and deliver the curriculum
Intermediate outcomes
[**] Youth associate age-mates with ‘safety’, youth don’t mind pregnancy with an age-mate – socially acceptable
Condom use during sex decreases
Unprotected sex is more likely to result in pregnancy – [***], Age-mates are 9 times less risky (HIV), Age-mates are less
likely to have acute HIV (40x) -less likely to have be engaged in MCP, more likely to adhere to ART
School attendance increases [P>25%]
[***] - Young adults are 6 times less risky (HIV), HIV acuteness is ambiguous (MCP lower, young adults are more
likely to default)
Ultimate outcomes
[**] Youth don’t fear pregnancy with an older partner who can provide
RELATIVE RISK INFORMATION CAMPAIGN – Unintended outcomes
Youth shift from older partners to age mates
HIV: [see probability] Pregnancy: Youth are exposed to social stigma at
school, youth have time and resources, youth’s health is compromised, schools are suspending youth from
schools, youth have competing preferences
Youth were getting income from older partners, youth lose the income they were
getting in the shift to age mates, youth want or need income for school attendance and wellbeing, youth pursue other economic opportunities
for income, youth have limited time for attendance
School attendance is ambiguous
[**], youth associate young adults with some degree of ‘safety’-relative to older partners, youth are unable to
negotiate condom use with a young adult, youth don’t mind pregnancy with a young adult [able to provide, socially
acceptable]
HIV: [see HIV ambiguity] Pregnancy: youth are exposed to social stigma at
school, youth have limited time and resources, youth’s health is compromised, schools are not suspending
youth from schools, youth don’t have competing preferences
HIV: Youth are exposed to social stigma at school, youth have limited time and resources, youth’s health
is compromised Pregnancy: All of the above + Schools are suspending
youth from schools, youth have competing preferences
Condom use during sex decreases/remains the same Pregnancy rates increase | HIV rates decrease [P:50% -75%]
Youth shift from older partners to young adults
[*]Youth perceive age-mates as the safest age group
Condom use during sex decreases/remains the same Pregnancy increases| HIV rates decrease [P: 50 – 75%]
Youth intended or had older sexual partners, Youth have agency over partner selection, youth internalize that higher
HIV risk in older individuals means own personal risk, youth’s priors on HIV rates for older sexual partners are lower than actual risk, youth perceive the cost-benefit of young adults (the ‘sweet and sour spot’) as optimized over age-mates &
older partners
School attendance decreases
Youth have young adult partners who provide income, youth perceive school as an add
value to income generation, youth perceive attendance as unnecessary to their income
generation capacity, youth attend school less in order to spend time with young adult partners
Pregnancy | HIV rates increase
[***] - Older partners are 9 times more likely to have HIV, HIV acuteness is ambiguous (MCP rates are higher, although
older partners are more likely to adhere to ART)
School attendance decreases
School attendance remains the same
School attendance decreases
Pregnancy | HIV rates increase
Older partners are 9 times more likely to have HIV, HIV acuteness is ambiguous (MCP rates are higher in older
partner, although older partners are more likely to adhere)
Youth are attentive,
Facilitators are credible
& charismatic Facilitators cover the key points & activities
in the curriculum
Youth have older partners who provide income, youth perceive school as an add
value to income generation, youth perceive attendance as unnecessary to their income
generation capacity, youth attend school less in order to spend time with older partners
Youth intended or had sex with older partners, youth have agency over partner choice, Youth now know that older
partners are beneficial (sugar), youth seek out older partners in order to get incentives, youth perceive the costs of HIV as being lower than the benefits received from older
partners, youth’s priors on HIV rates for older sexual partners are higher than actual risk, youth have a defeatist
attitude towards HIV risk
School attendance decreases
HIV: youth are exposed to social stigma at school, youth have limited time and resources, youth’s health
is compromised Pregnancy: All of the above + Schools are suspending
youth from schools, youth have competing preferences
School attendance decreases
Condom use during sex stays the same
Youth don’t mind pregnancy with an older partner who can provide
Youth continue to have sex with older partners
Youth intended or had sex with older partners, Youth have no agency over partner choice, Youth renegotiate
incentives, youth perceive the cost of HIV as being lower than the benefits they are getting from older men, youth don’t internalize that higher HIV risk in older individuals
means own personal risk
Pregnancy | HIV rates remain the same
No shifts made.i.e. Youth are still having sex with older partners, condom use patterns are the same
Everything is constant
School attendance remains the same
Everything is constant
Condom use during sex decreases/remains the same
Youth are unable to negotiate condom use with an older partner, youth don’t believe that condoms reduce the risk of HIV/Pregnancy, youth don’t have comprehensive knowledge
on correct & consistent condom use, Youth don’t fear pregnancy with an older partner who can provide, youth
don’t fear contracting HIV (ART freely available)
Youth are able to identify
men in their 30s & 40s as having high rates
of HIV
More youth seek out older partners
Notes 1. Pregnancy fluctuations reflect pregnancy incidences 2. Under all anal sex scenarios pregnancy rates will always decrease
40