NO PLAGIARISM DUE MONDAY MAY 6, 2019. ATTACHED ARE CHAPTERS TO ASSIST WITH ASSIGNMENT

mztclass82
CRJ305CRIMEPREVENTIONCH4.docx

4 Evaluation Science

Learning Objectives

Upon finishing this chapter, students should be able to:

· Understand the purposes of evaluation

· Describe the different stages of the evaluation process

· Summarize the major types of evaluation analyses

· Understand how to evaluate the merits of evaluation research

· Identify common challenges of evaluation research.

Introduction

Now that you understand the historical and theoretical foundations of crime prevention, it is time for a brief review of how rigorous scientific evaluations are conducted. Briefly stated, an evaluation is a study to test whether an intervention works the way it was designed to work, whether it is effective in achieving what it intended to do. Good evaluations provide evidence of an intervention’s effectiveness by assessing the amount of change in criminal behavior that can be uniquely attributed to the program, practice, or policy being investigated. Since there is no theory‐free program design, evaluations are also tests of the validity of criminological theories (Shadish, Cook, and Leviton, 1991).

Although this chapter will provide a more detailed and sometimes technical description of different evaluation designs and methodologies, we will cover most issues fairly briefly. If you have already taken a Research Methods course, much of this information should be familiar. If you are intrigued by the material and want to learn about evaluation designs in more depth, we encourage you to consult the suggested readings listed at the end of this chapter. Our goal is to help you understand the most important issues related to evaluation so that you can judge for yourselves if claims about an intervention’s effectiveness are to be believed. This knowledge will be important not only in interpreting the information on effective interventions described in Section III of this text, but also more generally in your future careers in criminology, criminal justice, public health, social work, and other professions.

There are three basic purposes of such evaluations. First, evaluations are undertaken to provide feedback to a program developer about the particular intervention features and/or content which are working well and contributing to positive outcomes and the aspects that are not working well. This information guides the developer in making specific changes to the program to increase its ability to reduce crime. The second purpose is to test the validity of the causal mechanisms underlying the intervention. Recall from Chapter 3 that effective prevention programs, practices, and policies are based on theories of crime and attempt to change the underlying causes and processes that the theory states are related to criminal behavior. Evaluations play a critical role in validating theories of crime and extending our understanding about the causal mechanisms that increase or decrease the likelihood of offending. They can evaluate the claim that if a specific treatment, based on a particular criminological theory, is delivered to participants, it will cause them to be less involved in criminal behavior. If crime is reduced, support for the theory is demonstrated. The theory is further supported if the particular mechanisms identified by the theory as leading to crime are also changed. For example, if the intervention shows that participants’ relationships with others are strengthened and their criminal behavior is reduced, then social control theory is supported. Third, evaluations provide valuable information to decision‐makers who are responsible for taking action to reduce social problems like criminal behavior. By reviewing the results of evaluation studies, policy‐makers can decide whether or not they will encourage or even mandate the use or discontinuation of particular crime prevention strategies. All three reasons for conducting evaluation research are important, but our focus in this book will be on the latter two purposes: on how evaluations can identify effective interventions based on sound causal mechanisms and how this information can lead to informed choices about which programs, practices, and policies to implement in order to reduce rates of crime.

The processes used to evaluate the effectiveness of crime prevention strategies are similar to those used in medical research to determine the effectiveness of prescription drugs or medical procedures and protocols, like how to provide lifesaving CPR. In the USA, the Federal Drug Administration (FDA) is charged with reviewing evidence from evaluations to determine if a specific drug has the desired effect on the targeted disease or medical condition, identifying important negative side effects and judging if its benefits outweigh its potential harms. Only drugs shown to be effective and to have minimal side effects are certified or approved by the FDA and allowed to be sold to consumers. As we will describe further in Chapter 5, although determinations of the effectiveness of crime prevention programs, practices, and policies are modeled after the procedures used in medical research, no federal agency regulates these processes or certifies crime prevention programs as effective and ready to be marketed and implemented. Several governmental agencies and universities have compiled lists of interventions they consider to be effective in reducing crime, but the criteria and processes for making this determination vary considerably across institutions. As a result, the available lists of effective crime interventions differ substantially in what they consider to be effective, as you will see in Chapter 5. In this chapter, we describe how scientific evaluations should be conducted and which types of evaluations are likely to produce the most persuasive evidence about an intervention’s effectiveness.

The Stages of Evaluation

Evaluation is a process with distinct and sequenced stages. Each stage is designed to answer specific questions about the intervention being evaluated, as shown in Table 4.1. Evidence of effectiveness can be obtained from each stage, and conclusions about whether or not a program is effective should be based on the overall, cumulative evidence produced across all stages. The evaluation design differs according to the questions being asked, and methods typically become more complex as the evaluation sequence progresses.

Table 4.1 Questions addressed during the three stages of the evaluation process.

Evaluation stage

Question

1. Process evaluation

Is the intervention based on theory, practical to implement, and logical?

How difficult is the intervention to implement?

Can the intervention be integrated into existing agencies and social systems without significantly changing its content or methods of delivery?

2. Pre‐post outcome evaluation

Provides tentative answers to the questions:

Does the intervention have the desired effect on the targeted risk factors, protective factors, and outcomes?

How strong are the effects?

Does the intervention have high social, economic, and political importance?

3. Experimental outcome evaluation

Provides more definite answers to the questions:

Does the intervention have the desired effect on the targeted risk factors, protective factors, and outcomes?

How strong are the effects?

Does the intervention have high social, economic, and political importance?

Process evaluation

The evaluation sequence typically begins with a process evaluation. This type of evaluation is designed to answer the first three questions in Table 4.1. It examines the nature of the intervention and its chances of being successfully implemented. Until the foundations of the intervention are well understood, there is no reason to move to the next stage of evaluation, which examines the outcomes produced by an intervention.

Process evaluations rely on a review of the intervention’s theoretical model and descriptions of the theory(ies) underlying the program, the specific risk and/or protective factors the intervention will target for change based on this theory, and evidence from other empirical studies that show support for the theory and its stated causal mechanisms. As noted in Chapter 3, risk and protective factors differ in: (i) their impact on crime, with some showing a more powerful influence on offending, which means that they are also likely to have a stronger prevention effect in reducing crime; (ii) their malleability or ability to be changed by an intervention (e.g., it is harder to change poverty than academic performance); and (iii) the prevalence of these factors or conditions in the population or environment targeted to receive the intervention. For example, parent/child conflict is likely to be highest in adolescence compared to other stages of the life course. The theoretical model guides the development of the intervention by indicating which risk and protective factors should be targeted for change based on their impact on crime, their malleability, and their prevalence in the targeted population. The theory should also suggest the most appropriate developmental timing for the intervention. For example, social learning theory and tests of this perspective indicate that peer interactions will be most important in affecting criminal involvement during adolescence.

The process evaluation also requires a review of the program change model, which specifies the attitudes and behaviors to be changed by the intervention and the research evidence that these factors will change the prevalence of the targeted risk and protective factors. For example, research has shown that child abuse (a risk factor) can be reduced by teaching parents how to be more effective in disciplining and monitoring their children and that academic performance and high school graduation rates (protective factors) can be improved by creating close mentoring relationships between youth and adults. The program change model should also identify the participants who will be recruited or enrolled in the intervention; the setting or context in which the intervention will be delivered; the number and length of lessons, sessions, or treatments to be provided; the specific types of instruction to be used (e.g., lectures, discussions, or one‐on‐one therapy); and the credentials and experience required of implementers. As we discussed in Chapter 3, the theoretical rationale and the program change model together are referred to as the intervention logic model and provide a detailed description of why and how the intervention is expected to work.

In addition to considering the nature of the intervention, or what it looks like on paper, the process evaluation also examines how the intervention is delivered in practice. That is, it addresses the second question listed in Table 4.1: how difficult is it to implement the intervention? Can it be implemented as planned, particularly as specified in the change model, without too many problems? A process evaluation may, for example, investigate the degree to which an intervention has successfully recruited the types of individuals, groups, or communities targeted to receive it; provided the correct number of treatment sessions, contacts, or classes using the methods outlined in the change model; and ensured that implementers have the required skills and backgrounds needed to deliver the intervention.

The answers to these process evaluation questions require the careful collection and review of data describing all aspects of the implementation process. For example, a process evaluation of an individually focused crime prevention program should record each person’s eligibility for the intervention, their attendance and completion of program sessions, and the content actually delivered to them during each session. It may also collect information on participants’ levels of risk and protective factors and criminal involvement at the start of the intervention and after it has ended. To assess prevention practices and policies, data should be recorded to document the date they begin, the number of persons trained to deliver the practice or implement the policy, the length and content of any training required of implementers, how the policy was enforced, and crime rates before and after the practice or policy was delivered or enacted. All of these records are used to establish that the program is being implemented with fidelity, as closely as possible to the model described by the program designer, a concept we will describe much more fully in Chapter 10.

A process evaluation can also begin to evaluate the third question listed in Table 4.1: how well the intervention is integrated into the agencies and systems charged with delivering or supporting it. Such issues may be assessed by surveying agency administrators or systems leaders (e.g., school superintendents, the head of the state juvenile justice agency, etc.) about their use of and support for the intervention. These issues become very important when effective prevention programs are taken to scale and widely used across a state or country. We will continue the discussion of how to evaluate implementation fidelity in large‐scale replications in Section IV.

What does a typical process evaluation look like? To provide a real‐life example, consider the Life Skills Training Program (LST), a school‐based drug prevention program which has been subject to several process evaluations. One such evaluation, performed as part of the Blueprints for Healthy Youth Development initiative (Fagan and Mihalic, 2003), examined the theoretical and program change model during replications taking place in middle schools in the USA. The results supported that developer’s descriptions of the program (see: www.lifeskillstraining.com/and www.blueprintsprograms.com/LST), as follows:

Theoretical rationale. LST is based on social learning theory.1 As we described in Chapters 2 and 3, this theory states that individuals learn behaviors during social interactions by observing, imitating, and modeling behaviors of others. Especially during adolescence, behaviors like substance use and violence are likely to be learned during interactions with peers. Such behaviors help individuals achieve goals they believe they are unable to achieve in more conventional, law‐abiding ways.

Risk factors identified as important in social learning theory and targeted for change in LST include: favorable attitudes toward drug use, interactions with delinquent and/or substance‐using peers, and neighborhood laws and norms favorable to drug use or crime. Protective factors include: having clear standards for behavior, believing that drug use is risky, having strong problem‐solving skills, and being able to refuse offers of drugs.

Program Change Model. LST is designed to improve the general social and personal skills needed to successfully navigate developmental challenges faced when moving from childhood to adolescence. It also teaches young people ways to resist pro‐drug influences, refuse drug offers from peers, and identify and resist pro‐drug messages shown in movies, TV, and other media. The program is a universal, classroom‐based intervention for middle school or junior high school students. It involves 37 sessions taught by teachers over three years who use a variety of instructional techniques, including lecture, demonstration, feedback, reinforcement, and practice.

During the Blueprints process evaluation, data describing program implementation processes were collected in over 400 schools across the country (Fagan and Mihalic, 2003; Mihalic, Fagan, and Argamaso, 2008). A review of these data indicated that LST could be delivered as planned. All schools were successful in recruiting and training teachers to deliver the intervention to the targeted, eligible population of students in middle schools and junior high schools. Most teachers taught all of the lessons over a three‐year period and delivered all of the required material in each session. However, the process evaluation indicated that implementation procedures were not perfectly followed in all cases. A few teachers did not cover specific lessons, not every student attended all sessions in all three years and some students did not fully participate in every lesson. As we will discuss in Chapter 10, programs are not always implemented exactly as they were designed, but when deviations from the program logic model are relatively minor, the replication should still be able to produce its anticipated effects on crime.

Programs that show positive results from process evaluations, like LST, are then ready for the next stage of evaluation. If a program fails to demonstrate that it is grounded in theory, fails to alter relevant risk and protective factors, has no clear change strategy or cannot be implemented as planned (with fidelity), the program is not ready to move to the next stage of evaluation. Instead, the program should be re‐designed and re‐evaluated or completely abandoned.

Pre‐post outcome evaluation

The next stages of the evaluation process involve outcome or impact evaluations, which seek to answer Evaluation Questions #4–6 (see Table 4.1) relating to the effects of the intervention on criminal behavior. Outcome evaluations vary in their level of rigor and sophistication. More basic designs, like a non‐experimental pre‐post design associated with Stage 2 of the evaluation process, provide tentative or preliminary answers to the evaluation questions. Pre‐post evaluations involve the collection of data from individuals who receive a particular treatment at the start of the study and again when they complete the intervention or when they drop out of the study prior to its completion. As shown in Figure 4.1, all individuals participating in the study are expected to receive the intervention and the degree to which their criminal behaviors change is assessed using data collected at two time points: pre‐test and post‐test. If participation in the intervention has some effect on crime, the level of initiation or involvement in offending should be lower at post‐test compared to pre‐test. If no change is seen, there is reason to question the effectiveness of the intervention. If the average level of crime actually increases from pre‐test to post‐test, the intervention is potentially iatrogenic or harmful. The impact on crime is referred to as the main effect of the intervention and is the outcome of greatest interest to practitioners, funders, and policy‐makers.

Process flow diagrams of outcome evaluation designs: pre-post design (top) and experimental design (bottom).

Figure 4.1 Outcome evaluation designs.

The pre‐post evaluation can also compare participants’ levels of targeted risk and protective factors before and after the intervention. As with the main effect on crime, if the intervention is effective, then changes in these secondary effects should also be seen. That is, levels of risk should decline from pre‐test to post‐test and the levels of protection should increase. If this does not occur, the program’s effectiveness is less certain. It is also possible to test the degree to which a reduction in crime is directly related to a reduction in a risk factor and/or an increase in a protective factor. This type of analysis, shown in Figure 4.2, is called a mediating analysis, because the change in the risk or protective factor is thought to mediate or be responsible for the change in outcomes produced by the intervention.

Flow diagram of mediating effects starting from intervention to decrease in risk factor or increase in protective factor, and to reduction in crime, with paths A, B, and C indicated.

Figure 4.2 Mediating effects.

Mediating effects are often not examined, especially if no main effect on crime is produced, but they are important in helping to understand a program’s main effects and evaluate its logic model. For example, if the main effect indicates a reduction in crime (Path C in Figure 4.2), the intervention shows secondary effects on risk and protective factors (Path A), and the risk and protective factors are shown to reduce crime (Path B), we have more confidence in the theoretical rationale of the intervention, the claim that these particular risk and protective factors are causally linked to crime and that the intervention change strategy was effective. If the main effect shows reductions in crime (Path C) but there is no evidence that the risk and protective factors were changed by the intervention (Path A), this may lead us to question whether or not crime was actually reduced. The outcome could have been produced simply by chance (it happens!); or, if it was truly reduced, the intervention change strategy is likely to be faulty since the outcome was not due to changes in the targeted risk or protective factors.

The pre‐post impact evaluation provides an initial test of an intervention’s effectiveness. If an intervention appears to be ineffective or does not work the way it was expected to work, this design can help pinpoint whether or not the problem is with the theoretical rationale or the program change model. However, keep in mind that a failure to show a main effect may also be due to poor implementation of the program. If a process evaluation indicates that the intervention was not fully implemented or that major deviations from the program change model were made, then the expected outcomes are less likely to be seen. Thus, it is helpful to conduct a process evaluation along with the pre‐post evaluation, even if a process evaluation was conducted in the past.

Although pre‐post impact evaluations can provide preliminary evidence of program effectiveness, they have several limitations which, together, reduce their ability to show with confidence that an outcome like crime has been reduced. We review some of these problems below and will continue this discussion when describing more rigorous experimental outcome evaluations. What is most important to remember is that, compared to experimental research designs, pre‐post evaluations cannot demonstrate with certainty that a reduction in crime is uniquely the result of the intervention. They cannot rule out other alternative explanations for how a reduction in crime might or might not have occurred.

Consider, for example, an indicated intervention designed to reduce the frequency of offending among 17–20 year olds who have a history of frequent and serious offending. These offenders attend a “wilderness survival camp” lasting two weeks, followed by five hour‐long employment counseling sessions. A pre‐post evaluation finds that the average frequency of offending in the year following participation in this program is 20% less than it had been the year before entering the program, indicating a substantial pre‐post reduction in crime. Can we conclude that program participation caused this decrease in the rate of offending? Not necessarily. As noted in the discussion of the life‐course developmental paradigm in Chapter 3, studies consistently show that involvement in crime increases in early adolescence, peaks between ages 17 and 18 and then declines. This drop‐off in offending is called the maturation effect. For the majority of the population, criminal behavior lessens as one gets older, develops more mature ways of thinking and enters new social roles like work and marriage. This “natural” decrease in offending means that individuals are likely to reduce their involvement in crime during this age period whether they participated in the program or not. Maturation, or aging, cannot be ruled out as a possible explanation for the observed decline in offending, and any claim that the program caused this reduction is debatable.

History effects may also affect individual offending during the implementation period. A history effect is when an event that is not part of the intervention occurs during the same time period the intervention is being delivered – an event that might have some effect of its own on criminal behavior. Suppose a bullying prevention program is introduced into a school that was known to have high rates of bullying. A pre‐post assessment based on students’ self‐reported bullying perpetration revealed less bullying in the 3 months following the intervention compared to the 3 months prior to it. But a month after the program started, three senior boys known to be bullies were arrested by the police for breaking into cars and stealing CD players and were then expelled from school. Assuming the school had a relatively small student population, can we conclude the bullying prevention program caused the decline in bullying? Can we rule out the removal from school of three frequent bullies as an alternative explanation? Probably not.

Next, consider a community that implemented a drug prevention program in all three of its high schools. The pre‐post assessment of drug use, based on student self‐reports, indicated that while the levels of alcohol use declined slightly after the intervention, the rate of marijuana use actually increased substantially in all schools. This finding would point to an iatrogenic program effect. However, while the intervention was going on, the city district attorney announced that she would no longer prosecute anyone for the possession or use of marijuana. Can we conclude the program actually increased marijuana use? Can we rule out the possible history effect caused by the DA’s announcement?

As these examples show, what at first appears to be an intervention success based on a pre‐post evaluation may actually be due to some other explanation. Similarly, concluding that an intervention has failed based on the results of a pre‐post design may be premature; other possible explanations could explain why crime increased rather than decreased following the intervention. Given the inability to completely rule out alternative explanations of program success/failure, the findings of pre‐post evaluations must be considered tentative and taken with much caution. Why, then, do evaluators rely on this design? Why not just move to the next stage of the evaluation process and use a more rigorous methodology which can provide more definitive evidence of an intervention’s effectiveness?

There are three main reasons why pre‐post evaluations are conducted. First, such evaluations are not very expensive or difficult. They do not necessarily require experts in research methodology and so can be undertaken by local program developers or staff. Second, it is very difficult to obtain enough funding or professional expertise to conduct a more rigorous outcome evaluation, like an experimental research design, without preliminary evidence of intervention effectiveness such as that provided by a pre‐post evaluation. Policy‐makers and evaluators want to know in advance that more complex evaluations will be worth their time and money. Third, the pre‐post evaluation results can be helpful in making changes to an intervention to improve its effectiveness. If the primary goal, at least at this stage of program development, is to refine the intervention and maximize its ability to reduce crime, then pre‐post evaluations can serve a critical role.

Before describing experimental outcome evaluations, we must mention a few technical points relating to the evaluation of program outcomes. Evaluations that claim to find a pre‐post difference in the onset, prevalence or frequency of criminal behavior are usually referring to a change in crime from pre‐test to post‐test that is statistically significant. This means that the difference was unlikely to occur accidently or by chance. Usually, if the difference could occur more than 5 times out of every 100 comparisons, it is considered a non‐significant difference and no claim of intervention effectiveness is made. If the effect could occur 5 or fewer times per 100 comparisons, the difference is considered a significant or “real” difference. In this case, the risk of claiming program effectiveness when, in fact, the difference in outcomes occurred by chance is very low; the difference can be attributed to the intervention with much confidence. This “5% rule” for determining statistically significant differences in program outcomes is used in nearly all quantitative evaluations.

Does this rule mean that you have to conduct 100 evaluations to determine if your outcome is the result of chance? Thankfully, the answer is “no.” Using appropriate statistical procedures to evaluate pre‐post differences, you can assess program effectiveness using a single study. Although we could explain in more detail the mathematical calculations needed to identify the statistical probability of finding observed differences, we will leave that issue for you to investigate on your own (one good source to consult is Weisburd and Britt, 2014). What is important to remember is that when we state that an evaluation shows an effect on crime or a statistically significant difference in criminal behavior, we mean that this outcome has been found using the 5% rule.

A related statistical measure that is used to measure the size of a difference in outcomes is the effect size (ES). Although effect sizes are described in more detail later in this chapter, we mention the concept now to make the point that, while it is generally true that the bigger the observed difference, the more likely it is to be statistically significant, it is possible to find a statistically significant difference that is actually not very meaningful. That is, the difference could be relatively small, even if it is statistically significant, and have no practical value if our goal is to substantially reduce rates of crime. Statistical significance is directly related to the sample size of an evaluation, meaning that studies involving more participants have a greater likelihood of achieving statistical significance compared to studies with fewer participants even if the actual difference in criminal behavior is the same in both evaluations. For example, an evaluation of a probation program might indicate that it reduced the recidivism rate by 1%. If the number of participants in the evaluation study is small, this difference will not be statistically significant. If the number of participants is large, this 1% change may be statistically significant. Regardless, the size of the difference is small and does not reflect a very important reduction in crime. Furthermore, if the program is expensive, costing $20,000 per participant, then the change may be viewed as having very little practical value.

The goal in evaluation studies is to find differences that are both statistically significant and substantively important. The ES helps to evaluate the second issue since it identifies the size of an intervention effect without being overly influenced by sample size. The larger the ES, the greater the reduction of crime and the larger the potential social value of the intervention. Although not all program evaluations calculate effect sizes or include them in their published findings, we will note the ES of interventions described in Section III when they are available. We will also report on the return‐on‐investment (ROI), or financial return on investment of programs, a measure that takes into account both the effect size and the cost of the program as we will discuss at the end of this chapter. By examining these issues, evaluations can help address the fifth and sixth evaluation question listed in Table 4.1.

Experimental outcome evaluation

The third stage of evaluation involves an experimental research design, which can take several forms, including a causal comparative design, a quasi‐experimental design, or a true experimental design. Unlike pre‐post evaluations, these all involve a comparison or control group that does not receive the intervention, as well as a group that is provided the program or treatment. As such, these designs are better able to rule out alternative explanations for observed pre‐post intervention differences and provide more convincing evidence that the intervention caused a reduction in criminal behavior or some other outcome. They provide a more definitive answer to Evaluation Questions #4–6 in Table 4.1. We describe quasi‐experimental and true experimental evaluations in the next two sections of this chapter. Causal comparative evaluations, also called ex‐post‐facto studies, are rarely used in contemporary crime prevention evaluations and are not discussed here.

Quasi‐experimental design (QED) outcome evaluations

As shown in Figure 4.1, in a QED, a group of individuals (the study participants) is divided into a treatment group that receives the intervention to be evaluated and a “matched” comparison group that does not receive the intervention. Levels of risk and protection and targeted crime outcomes are assessed for both groups prior to the start of the intervention using the same measures and procedures. These results are examined in order to “match” the two groups, or make them as similar as possible before the study starts, particularly on their levels of risk and protection and involvement in crime. In practice, it is often not possible to achieve two groups that are completely matched in these outcomes. Sometimes, it is only possible to match the two groups in their demographic characteristics; for example, on the age or gender distribution of individuals in the two groups. However, if there is a close match on all of the intervention’s targeted risk and protective factors and outcomes at pre‐test, the evaluation is better able to rule out alternative explanations for program effects and make causal claims. After matching, the treatment group receives the intervention while the comparison group does not. At the end of the intervention, both groups are again assessed on the same measures of risk, protection, and criminal behaviors. Although Figure 4.1 shows only one post‐intervention assessment, multiple assessments can be made following the end of the intervention in order to determine longer term effects. As we will describe in Section III, some studies have conducted follow‐up assessments 10 to 20 years following the intervention!

True experimental outcome evaluations

True experiments also involve two groups of individuals and pre‐test and post‐test assessments of outcomes, but they require an additional and important feature: random assignment to the treatment and control groups. This type of experimental evaluation, also shown in Figure 4.1, is typically referred to as a randomized control trial (RCT).2 Whereas QEDs attempt to create two similar groups using a matching process, random assignment is much better able to produce treatment and control groups that are equivalent in key outcomes. Using this method, any participant or group (e.g., a family, classroom, or geographical area) in the study has the same chance of being placed in the treatment or control group. Who is assigned to each group is completely random. To ensure that group assignment is random, an evaluator may actually flip a coin and assign all participants who receive a “heads” to the treatment group and all those receiving a “tails” to the control group. Random assignment helps ensure that any differences between the two groups that may exist at the start of the study are due completely to chance, which will minimize bias when interpreting the results of the evaluation. To repeat, using randomization to create treatment and control groups is much better than the matching procedure employed in QEDs because it insures that the participants in the two groups are essentially the same, except for chance. It controls for all possible differences between those in the treatment and control groups, not just those characteristics that an evaluator thinks of ahead of time and includes in a matching process.

Random assignment allows us to rule out the possibility that the observed difference between those receiving or not receiving the intervention was due to some pre‐intervention difference. In fact, in a RCT, a pre‐test assessment is not required if the evaluator can be sure that the assignment to treatment and control groups was truly random. In practice, however, a pre‐test assessment is usually conducted and it can be helpful, particularly when the size of the treatment and control groups is relatively small. When less than 100 participants are involved in the study, the randomization process is less likely to produce equivalent groups and pre‐tests will be more necessary. Even with large sample sizes, pre‐tests can verify that the assignment was truly random and the two groups are equivalent. Having this information also increases the ability to rule out alternative explanations and calculate group differences more precisely.

Judging the Merit of Evaluation Designs

How can we judge the merits of these different types of evaluation strategies and determine which is best able to identify effective crime prevention programs, practices, and policies? Doing so relies on an assessment of each one’s validity, or its accuracy in determining that a change in criminal behavior is actually the result of participation in an intervention. When an evaluation has strong validity, we can place more confidence in its findings. Three types of validity must be considered.

First, evaluations should be judged according to their construct validity, or the degree to which their measures of targeted risk and protective factors, crime outcomes and any other processes or conditions related to the intervention accurately assess the constructs they are intended to measure. “Constructs” can be abstract ideas about some quality or trait of persons or contexts that cannot be seen, felt or heard by an outside observer; for example, an individual’s intelligence or self‐control, or the amount of bonding between a child and a parent. Since these types of constructs are abstract, they present a challenge to a scientific researcher who wants to accurately assess their presence or absence. Other constructs can be directly observed and counted and thus present a less significant measurement challenge. Such constructs, like “criminal behavior,” are typically measured using a direct count of the number of acts observed, recorded, or reported.

For crime prevention programs, we are usually interested in assessing the construct validity of criminal behavior. Recall from Chapter 1 that criminal acts can be defined and measured (i.e., counted) in different ways and that not all measures are equally “good.” In fact, this means that different types of measures have different levels of construct validity. We noted that arrests are less able to measure actual levels of criminal behavior because most criminal acts are undetected or are reported or observed but do not lead to an arrest. Using the number of arrests to evaluate the effectiveness of a youth employment program in reducing adolescent offending would therefore have weak construct validity. However, an arrest measure would have good construct validity if evaluating the effectiveness of a law enforcement policy designed to increase the use of arrests for domestic violence.

Compared to crime outcomes, it is more difficult to achieve high construct validity of measures of risk and protective factors because constructs like attitudes regarding violence or bonding to family or school represent abstract ideas or qualities which cannot be easily observed or counted. Knowing this, researchers have spent much time developing and testing measures of such factors. These usually rely on self‐reported surveys or observations to do so. They also tend to use multiple items or indicators to create scales which are better able to measure these more complex and abstract events or conditions. For example, a well‐known scale used to measure the individual risk factor of “low self‐control” relies on a set of 24 indicators or items assessing various aspects of this construct (Grasmick et al., 1993). Individuals are asked to rate their level of agreement with 24 statements including: “I don’t devote much thought and effort to preparing for the future;” “When things get complicated, I tend to quit or withdraw;” and “I lose my temper pretty easily.” Parental warmth, a protective factor, has been measured using home observations in which trained researchers count the number of times caregivers display different behaviors thought to indicate their love for and affection towards the child (Caldwell and Bradley, 1984). For example, observers rate whether or not or how often the parent praises; caresses, kisses, or hugs; or voices positive feelings to the child. Many established measures, already evaluated and shown to have good construct validity, are available for use in evaluation research. When there is no existing measure, the evaluator must pay careful attention to construct validity and create a new measure which accurately reflects the attitude, behavior, or condition of interest.

The second way to judge an evaluation is by its internal validity, the extent to which the evaluation has eliminated alternative explanations of an intervention’s effectiveness. We mentioned earlier that pre‐post evaluation studies have relatively weak internal validity, given that they cannot rule out alternative explanations like maturation and history effects. Other threats to the internal validity of an evaluation include testing effects, instrumentation effects, non‐equivalence, regression to the mean, and attrition (Shadish, Cook, and Campbell, 2002).

A testing effect refers to the situation whereby participants in an evaluation learn something from taking the pre‐test which artificially affects their responses to the post‐test. The more “practice” they have in completing assessments, especially when multiple post‐tests are given in a short period of time, the more problematic testing effects can be. Participants may become too familiar with the items used to measure particular outcomes, since the same items are supposed to be used during every assessment. Prior to the next assessment, participants may have thought about the questions to be asked and decided to answer them differently based on what they think the researcher is trying to study. Or, they may simply become “better” at test‐taking because they know what to expect and/or are less anxious about the testing process. Testing effects can result in higher or lower post‐test scores even if the intervention has had no real effect. If testing effects are known to have occurred, the evaluation will be less valid. One solution to this problem is to increase the length of the interval between pre‐test and post‐test. There is also a very sophisticated type of experimental design which actually controls for this possible effect called the Solomon Four Group Design (for more information on this type of design, see: Shadish, Cook, and Campbell, 2002).

An instrumentation effect indicates a change in the way a risk factor, protective factor, or crime outcome has been measured from pre‐test to post‐test. Any observed changes in the outcome may then be due to the change in the survey item(s) or in its meaning, rather than changes in the participants. To understand how this might occur, consider a multi‐year evaluation of a program implemented during the transition from adolescence to young adulthood. This period coincides with changes in the legal definition of what constitutes crime: some behavior that is illegal for adolescents is not illegal for adults, like purchasing alcohol, carrying a concealed weapon, or failing to attend school (i.e., truancy). If this difference is not taken into account, a simple count of the number of self‐reported crimes could show a pre‐post reduction that is not the result of the intervention, but rather a change in what is defined as a crime. This situation then introduces a potential alternative explanation for any observed pre‐post changes.

Another type of instrumentation effect is related to the possible bias introduced when a person delivering or involved with the intervention is the one completing the pre‐ and post‐test. For example, teachers delivering a violence prevention program may be asked to evaluate student participants’ levels of aggression in the classroom, or correctional officers delivering a vocational training program may be asked to assess inmates’ job skills. Knowing that a person is receiving an intervention can influence the rater’s perception of whether or not this person has changed, particularly if the rater believes the intervention is effective. In such cases, any improvement between pre‐test and post‐test may be due to rater bias, rather than a true change in risk, protection, or criminal behavior.

Regression effects can result when individuals assigned to a treatment group were selected or recruited because they were serious or frequent offenders at the peak of their criminal involvement. Given the age/crime curve described in Chapter 3, it can be assumed that without any intervention, rates of criminal behavior will decline as adolescents move into early adulthood. What may be interpreted as a decline in crime produced by an intervention is actually the result of the aging process. More generally, regression to the mean occurs as individuals with very high or very low levels of any outcome or measure move towards an average level of the outcome over time without any intervention. This “natural” change from an extreme to more typical level can skew evaluation findings and make it appear as if an intervention has either been successful or failed when, in fact, this is not the case.

We previously mentioned the problem of non‐equivalence, which occurs when participants in the treatment and control groups differ significantly in the outcomes of interest at the start of the study. One group might have higher levels of criminal behavior than the other or may be at greater risk of offending because they live in more disadvantaged neighborhoods or are exposed to more violence at home. Any control‐treatment group difference at pre‐test on a main outcome can provide an alternative explanation for an observed difference at post‐test. Matching participants in the control group to those in the treatment group to achieve groups that are similar in their outcomes is one way to try to achieve equivalence, but this can be difficult to do. Groups need to have similar levels not only of crime, but also of all risk and protective factors that could be related to crime. Unless all differences are removed, non‐equivalence remains an alternative explanation for differences at post‐test. As noted earlier, random assignment is a much stronger method than matching for achieving group equivalence at pre‐test.

Even the best evaluation study typically has participants drop out before the intervention is over. This loss is called attrition and poses another threat to the internal validity of an evaluation. If participants leaving the study have different characteristics than those remaining, especially different rates of criminal behavior, attrition could be alternative explanation for any observed differences at post‐test. For example, if males are more likely to drop out of a treatment group than a control group, greater reductions in crime for the treatment than the control group at post‐test could be due to this loss, given that males are more likely to commit crimes than females. Attrition is related to the problem of non‐equivalence, because treatment and control groups that were equal at pre‐test can become unequal at post‐test due to attrition. In the prior example, if both groups had equal numbers of males and females at pre‐test, at post‐test the control group will have more males than the treatment group. Even if the actual rate of attrition is relatively low, when attrition affects group equivalence, it can provide an alternative explanation for any observed post‐test differences.

The third aspect of an evaluation to consider is external validity. External validity is the degree to which the intervention is generalizable and can be transported to other settings and replicated with other participants with the same level of effectiveness. The greater the external validity or generalizability of an intervention, the broader its impact on crime is expected to be. Programs that have very strong external validity and are expected to work for diverse populations could potentially be replicated across a country or internationally and produce widespread reductions in crime.

There are three major types of external validity to consider: population, ecological and operations external validity. Population external validity evaluates whether or not the results of an evaluation shown for specific types of participants can be generalized to all persons with these same characteristics and to other types of persons. An intervention shown to be effective for Caucasians but not Asians, for example, would have low population external validity.

Ecological external validity is concerned with how the social or physical setting of the study influences outcomes and if effects can be replicated in other types of settings. For example, a program shown to reduce crime in urban, high‐poverty neighborhoods, wealthy neighborhoods, and rural communities would have high ecological external validity. Or consider an evaluation of a family‐based program seeking to increase parent/child bonding, reduce family conflict and prevent the initiation of delinquency. If children are asked to rate the quality of their interactions with parents and their illegal behaviors on surveys conducted in their homes, they may be reluctant to answer honestly for fear that their parents may see the results of their survey. If a second evaluation of the intervention relied on assessments conducted in a more neutral setting such as the child’s school, the youth’s responses may differ. The two evaluations may then produce different results due to poor ecological external validity, even if the change in behaviors was the same.

There are two ways to assess an evaluation’s population or ecological external validity. First, and ideally, the evaluation team would randomly select a sample of participants from the population of all persons who are eligible to receive the intervention. Specific study settings would also be randomly selected from the broad set of all possible contexts where the intervention was designed to take place. If these procedures are followed, the findings can be generalized from the smaller groups and settings to the larger population of all eligible persons and all possible study locations using statistical probability arguments. In practice, however, these procedures are rarely followed, at least in studies of crime prevention.3 A random process is almost never used to select participants or settings. As a result, there are few crime prevention programs that can claim to have very strong population or ecological external validity.

A second and more common strategy involves making a logical argument that the participants actually selected are typical of some larger group of intended and eligible persons and settings. To make this argument, the evaluation should describe in detail the characteristics of the participant sample, including their age, sex, socioeconomic status, race/ethnicity, level of education, and involvement in or general risk for criminal behavior. The intervention setting should also be described, including the types of families, schools, neighborhoods, or correctional facilities involved in the study. To the extent that these sample characteristic are similar to those intended to be targeted by the intervention, an informed judgment can be made regarding the external validity of the evaluation. Keep in mind, however, that this practice will provide much weaker evidence for external validity than the randomization strategy.

The third type, operational external validity, examines the degree to which findings from one evaluation are replicated in a second evaluation conducted by different investigators.4 This concern arises from the considerable evidence that evaluations completed by the developer of an intervention are more likely to find positive findings and stronger effects than when they are undertaken by other investigators (Eisner, 2009; Gandhi et al., 2007; Petrosino and Soydan, 2005). Typically, the developer of an intervention leads the first evaluation of the program and often has a personal interest, and sometimes a financial one, in its success. As a result, s/he may ensure that the intervention and the evaluation are implemented with a very high level of rigor and care that is difficult to duplicate in later tests by investigators who have no connection to the developer or intervention.

In general, the greater the number of replications, that is, of evaluations showing similar effects, the stronger the claim of external validity. Although some programs and policies have been subject to multiple replications, many have been tested only once. In these cases, investigators often avoid discussing the intervention’s external validity and simply note that the evaluation findings are limited to those individuals, groups, or settings actually participating in the evaluation. In our reviews of effective crime prevention programs, we will take care to describe the number of replications that have been conducted and to evaluate the logical claims that can be made for external validity.

Considering all three types of validity shown in Table 4.2, evaluations are most commonly judged by their construct and internal validity and less often for their external validity given the costs and complexity involved in doing so.5 Internal validity is actually of greatest concern and is very difficult to fully establish. Of the different designs we have discussed, RCTs are considered the “gold standard” and best method of achieving internal validity and producing the strongest claims of intervention effectiveness (National Science Foundation, 2013; Shadish, Cook, and Campbell, 2002). This design has the greatest potential for addressing the threats to internal validity and of eliminating the potential alternative explanations for an observed difference between control and treatment groups. The RCT design evaluation can, at least in theory, successfully deal with maturation, history, regression, and non‐equivalence threats.

Table 4.2 The three types of validity to be assessed when judging evaluation designs.

Type

Definition

Example

Construct validity

The degree to which a measure adequately captures or represents the construct it is intended to measure

Higher construct validity: using self‐reports of offending to measure criminal behavior Lower construct validity: using self‐reports of fear of crime to measure criminal victimization

Internal validity

The degree to which an evaluation can rule out other explanations of an intervention’s effectiveness

Higher internal validity: experimental evaluations, especially randomized control trials (RCTs) Lower internal validity: pre‐post evaluations

External validity

The degree to which an intervention’s effectiveness applies to a broad population of individuals, groups and/or contexts

Higher external validity: the Multisystemic Therapy (MST) program lowers recidivism for males and females, youth of different racial/ethnic groups and those with different levels and types of offending at pre‐test Lower external validity: the Project Northland program reduces alcohol use among Caucasian students from rural Minnesota but not among racially diverse students in urban Chicago

Well‐conducted QEDs can also address many of the threats to internal validity but are particularly vulnerable to the threat of non‐equivalence given that they rely on matching rather than randomization when trying to establish comparable treatment and comparison groups. Because the randomization requirement of RCTs is often difficult to achieve, and QEDs are often used, a fair amount of research has been conducted to compare results obtained using QEDs and RCTs. For example, studies have compared outcomes obtained using randomly assigned intervention and control groups and those obtained using the same intervention participants and a comparison group selected through methods other than randomization. Cook, Shadish, and Wong’s (2008) review of 12 such studies suggested that when comparison groups are obtained without very close matching, estimates of an intervention’s effects are likely to be inaccurate. This is true even when statistical techniques are used to adjust for observed differences between the two groups at pre‐test. Often QED studies match groups only on demographic variables, not on the outcomes of greatest interest to the evaluator, and these studies consistently fail to reproduce the results of RCTs. QED designs are more likely to produce valid results when there is very careful matching of the treatment and comparison groups at pre‐test, especially on measures of the outcome, risk/protective factors, and geographic location of the study.

RCTs and QEDs are more or less equally subject to testing, instrumentation and attrition threats. Since these evaluation designs cannot fully address these issues, evaluators must turn to statistical procedures to do so. We review some of these techniques and provide additional description of analysis strategies used to estimate the size and duration of program effects in the next section.

Evaluation Analyses

The main goal of a crime prevention evaluation is to establish a causal relationship between participation in a program or between the introduction of a new practice or policy and a reduction in criminal behavior. Evaluations involve a causal analysis, an analysis in which a cause(s) of crime is manipulated or controlled by a researcher via an intervention that is purposely delivered to some persons or in some contexts and not others. The demonstration of a causal relationship requires three conditions: (i) finding a statistically significant relationship between the cause and the effect, (ii) ensuring that the cause occurred prior to the effect, and (iii) ruling out possible alternative causes of the effect. It is relatively easy to satisfy the second condition, as the program, practice, or policy (the cause) always precedes the measure of change and the analysis of the effect (criminal behavior). As we have discussed, evaluations can also address the third condition by ruling out some of the threats to internal validity. To satisfy the first condition, statistical analyses must be used. In addition to identifying whether or not a cause produces an effect, these statistical procedures can also rule out some of the potential alternative explanations of changes in crime, like testing effects, instrumentation effects, and attrition effects.

The most basic analysis involves an assessment of the amount of change in the measures of risk, protection, and crime outcomes between pre‐test and post‐test. If the intervention is effective, the targeted risk and protective factors and crime measures will all have changed in a positive direction, demonstrating a statistically significant relationship between participation in the program (the cause), a reduction in risk and criminal behavior, and an increase in protection (the effects). When the evaluation includes both a treatment and a control group, the assessment of pre‐ to post‐intervention change is made for both groups, and the amount of change between groups is compared. If the program is effective, the amount of change will be in the expected direction and significantly greater for the treatment group than the control group. As noted earlier, the measured change in criminal behavior is the main effect. It is the primary goal of the intervention and often the only effect that is actually measured and reported. While the evaluation may assess secondary effects on risk or protective factors, or mediating effects examining whether or not changes in the risk and protective factors led to changes in outcomes, such analyses are not routinely performed at this time in the study of crime prevention.

Marginal and absolute effects

It is easy when reading results from experimental outcome studies to focus on changes that have occurred for the treatment group, but it is important to remember that in this type of evaluation, unlike the pre‐post design, outcomes are based on a comparison of the treatment group with a control group. It is also important to realize that there are two types of control groups: those who receive no treatment at all during the study period and those who receive an alternative intervention. The second group is often referred to as a “treatment as usual” control group. Providing control group participants with some type of service or treatment is common in crime prevention studies because it allows researchers to avoid the ethical dilemma faced when withholding services from those who may seem in need of assistance. In addition, a goal of many evaluations is to estimate whether a new intervention is more effective than what is usually offered. For example, probation is a standard type of intervention used by virtually all criminal and juvenile courts. If researchers or court officials sense that probation is not working very well to reduce crime – for example, they see that a growing number of offenders are violating the terms of their probation and being sent back to court – then a new intervention may be offered to some offenders (who would be in the intervention group) and its effectiveness compared to those who receive probation (the control group).

When the control group receives no intervention, the main effect is considered an absolute effect. It compares the effect of participating in the intervention to getting no services. When the control group participates in an alternative intervention, the marginal effect will indicate how much better or worse the new intervention is compared to the alternative intervention. The interpretation of main effects is quite different for these two types of evaluations. In the first scenario, if no differences in crime are found, the intervention would be considered ineffective. In the second case, finding no differences would indicate that the new treatment is no better than the alternative intervention or the treatment as usual. A decision‐maker could still decide to use the new treatment, especially if it was cheaper or considered easier to implement than existing services.

Attrition analyses

No matter what types of services a control group receives, main effects can be misleading if there is a high degree of participant attrition, a different attrition rate for treatment and control groups, or a loss of different types of participants in the two groups. As a first step in trying to minimize the threat of attrition, it is important to document how many and what types of participants drop out of the evaluation before it is over. This analysis can involve a simple count of the number of participants who complete the pre‐test and the post‐test and a calculation of the attrition rate. So, if 100 participants were enrolled at the start of the study and 90 completed the post‐test, the attrition rate is 10% (100‐90 divided by 100). If this rate is low, ideally less than 5% (Shultz and Grimes, 2002), or if the rate is about the same for the treatment and control groups, attrition is unlikely to threaten the internal validity of the study or to bias the estimate of the main effect.

A more problematic but somewhat common scenario is differential attrition, which indicates a difference between the treatment and control groups in the amount of attrition or in the type of participant that drops out. Even if the rates of attrition are the same in both groups, differential attrition can occur. For example, an evaluation may indicate a loss of 10% of participants in the treatment group and 10% in the control group. But, what if the majority of drop‐outs in the first group were high‐rate offenders, while most of those lost in the control group were low‐rate offenders? How might this situation affect the analysis of main effects? The evaluation may find that the treatment group showed greater reductions in offending compared to the control group. Differential attrition is problematic because it can introduce the threat of non‐equivalence, even in an RCT, if groups that were comparable at the start of a study are no longer equal at the end of the study and the difference is related to dropping out, not to the intervention itself.

Attrition makes it more difficult to conduct an intent‐to‐treat (ITT) analysis. In an ideal evaluation, all participants assigned to the treatment or control groups complete the pre‐test, are followed over time and complete the post‐test. This allows for an ITT analysis in which data from all participants is included in the analysis of main effects. When attrition occurs and data are missing from some participants, it is still possible to use an ITT approach. Because participant drop‐out is so common, several statistical methods have been developed to minimize the threat of attrition (Graham, 2012; Shafer, 1999). However, many evaluations with attrition problems have not used such statistical techniques to compensate for participant loss or have used them incorrectly.

What I advocate is treatment effectiveness research done well enough to produce both results we can trust and sufficient explanatory detail to understand why we got those results. Unfortunately, much contemporary treatment effectiveness research not only falls well short of this mark, but is, frankly, horrid.

Lipsey, 1988: 6

Some evaluators choose to abandon an ITT analysis and purposefully exclude information from those who failed to complete the treatment or who had no post‐test assessment. Those who endorse this practice claim that individuals who received none or only part of a treatment would have no reason to change their behaviors and thus should not be included in estimates of the intervention effect (Gupta, 2011). Although this rationale makes sense, the general consensus among prevention scientists (e.g., Moher, Schultz, and Altman, 2001) and our view is that failure to use an ITT approach can bias the estimate of the program’s effectiveness for several reasons. First, in any real world implementation some participants will drop out of the treatment condition. Subjects may move out of the intervention area; experience unexpected, negative side effects; feel threatened by the intervention or stigmatized by participation; or can no longer afford the cost of the intervention. Individuals may also switch from the control group into the treatment group or vice versa. Any of these changes can threaten the equivalence of the treatment and control groups and bias evaluation results. Second, when data are analyzed from a subset of individuals who had been randomly assigned to the treatment group at the start of the study, the results will be less generalizable to the program’s intended population of persons. In this case, the external validity of the evaluation is threatened. Third, the sample size of the subsample of individuals who complete the treatment will be smaller than the original group, which will reduce the evaluation’s ability to detect differences that are statistically significant. Given all of these potential problems, it is critical that evaluators make every effort to obtain post‐test assessments for all participants initially assigned to treatment and control groups. They should also become familiar with the statistical techniques than can help minimize attrition and missing data problems.

Mediation and moderation analyses

Although most evaluations focus on assessing main effects, additional analyses like mediation and moderation analyses can significantly add to the understanding of the intervention’s main effects and external validity. A mediation analysis, illustrated in Figure 4.2, helps explain exactly how the treatment produced a change in criminal behavior and tests the program’s logic model. The mediation analysis determines if the treatment changed the targeted risk and/or protective factors (Path A in the Figure) and if the change in risk/protection led to the change in criminal behavior (Path B in the Figure) as hypothesized in the theoretical rationale. If this type of analysis shows that the targeted risk/protective factors were changed by the treatment as expected, and this change led to the reduction in criminal behavior, a strong claim can be made for the effectiveness of the intervention and for the validity of the theoretical model underlying the intervention.

In some cases, however, a mediation analysis could show a significant main effect (Path C) and a significant change in the risk and protective factors (Path A), but fail to show that the change in the risk and protective factors was related to the main effect (Path B). This finding would weaken any conclusions about the effectiveness of the intervention and challenge the program’s theoretical rationale. The intervention would appear to work, but it would not be clear why or how it worked. If there were no main effect (Path C), but the treatment was effective in changing risk/protection (Path A), this also would challenge the intervention’s theoretical rationale. Finally, if the intervention failed to change the targeted risk/protective factors (Path A) but produced a significant main effect (Path C), this would challenge the program change model, since the change did not occur using the intended mechanisms. Even though these types of findings weaken causal claims about the intervention, they can still be useful in helping developers refine their theoretical and program change models in order to improve intervention effectiveness.

An analysis of moderator effects determines whether or not a third factor, usually related to participant characteristics like sex, race/ethnicity, or neighborhood poverty influence the direction or strength of the intervention’s effect. This type of effect is shown in Figure 4.3. Many evaluations have tested for and found evidence that a moderator variable influences an intervention’s main effect. For example, several of the evaluations of the Good Behavior Game (see Chapter 8) have shown that it is effective only with high‐risk participants (i.e., students who show aggressive behavior in kindergarten) and with males (Kellam et al., 2008). The intervention does not appear to reduce crime among lower risk students or females. For this program, both the “risk status” and sex of the participants would be considered moderating variables.

Figure 4.3 Moderating effects.

Diagram depicting the moderating effects. Intervention and Moderating factor both have arrows pointing to Reduction in crime.

In some cases, interventions may have a significant and positive main effect but show harmful outcomes for a sub‐set of participants.6 Moderating analyses can help identify these types of differential effects and assess the generalizability of an intervention: whether it appears to work equally well for all participants, is more effective for some groups than others, or has harmful effects on a specific subgroup. Like mediating analyses, moderating analyses can provide important feedback to program developers about how and for whom their program works, as well as inform potential users of a program of these differences. For example, developers may make changes to their program change model to make the intervention more effective with low‐risk participants. Or, they could ensure that the program is used only with the population for whom it has been shown to be effective. Knowing that certain conditions are necessary for program success can also help decision‐makers determine if an intervention is appropriate for the specific population they serve.

Given the value of the information provided by mediating and moderating analyses, it is surprising that such statistical procedures are not routinely included in crime prevention evaluations. A moderator analysis is more frequently performed than a mediator analysis, but typically, only a small number of potential moderating factors are considered, most often the sex or race/ethnicity of participants. Our descriptions of specific crime prevention programs and practices will identify, when the analyses have been conducted, the types of populations for whom the intervention has been shown to be effective.

Short‐ and long‐term effects

To this point we have referred to analyses that measure the effect of an intervention from pre‐test to post‐test, from before the intervention started to immediately after it ended. This type of analysis assesses the short‐term effect of an intervention. While important, such analyses do not provide the most complete or accurate test of an intervention’s effectiveness.

Consider an evaluation of a work program in a state correctional facility. Inmates interested in this program are randomly assigned to participate or not participate in the program. Pre‐test and post‐test assessments are made and official reports of prison infractions are used as the outcome measure. The measure of effectiveness is based on the change in rates of infraction during the six months before entering the program compared to the rate at the end of the intervention, a short‐term effect. Why is this time frame inadequate to assess the effectiveness of the program? For one reason, participants have been in custody throughout the entire intervention, including during the pre‐test assessment, have been closely supervised and have had limited opportunities to engage in criminal behavior. Secondly, the goal of the program was to provide work experience that would help inmates secure jobs after leaving prison. The theoretical rationale was that employment is a protective factor that would have its beneficial effect after the offenders had left the institution; it would keep them from re‐offending once they were released back to their communities. However, the short‐term analysis did not extend past the time of release, and given the restrictions imposed while participants were in custody, is this a fair test of the program’s effectiveness? It is likely that such an evaluation would find no effect on prison infractions, but does this mean that the program should be stopped on the grounds that it did not work?

A better evaluation would investigate the longer term effects of the program, using multiple post‐tests extending over a greater duration, at least until there is time for the intervention to have its intended effect. The specific amount of time needed between the pre‐test and final post‐test should be guided by the intervention’s logic model, as well as research relating to the developmental timing of particular behaviors. For example, an intervention providing academic enrichment services to preschool children in order to reduce delinquency should follow participants at least until the teenage years, since this is when crime is likely to occur. However, there will be pressure to have results from the study, and secondary effects on risk and protective factors like children’s reading skills and their attachment to teachers could still be examined in the short term, prior to adolescence.

In the case of the prison work program, a second post‐test one year after release from the correctional facility and a third post‐test two years after leaving would likely be adequate to capture the program’s main effects. A comparison of recidivism rates from the initial post‐test at the end of the program to the third post‐test is called a time‐to‐failure analysis. This analysis compares the treatment and control groups’ tendency to re‐offend right away or to delay offending for a longer period of time. If the treatment group remains crime free for a longer period of time (on average) than the control group, even if the actual rate of recidivism at the last post‐test is the same, this would be considered a positive outcome. In the shorter term, the evaluation could also measure secondary effects such as the acquisition of specific work skills, good work habits, and changes in attitudes toward work. Changes in these effects could then be incorporated into a mediating analysis to see if they led to reductions in recidivism.

All types of evaluations can incorporate multiple follow‐up periods and post‐tests, but these additional assessments will require more resources and effort compared to the administration of a single post‐test. Nonetheless, they are important to help determine the sustainability or durability of intervention effects. Outcomes could be brief or could last for many years, indicating that they are sustained over time. For example, a program operating in a correctional facility may demonstrate positive effects while prisoners are in custody, are closely guarded and receive reinforcements for good behavior. However, these effects could be quickly lost once prisoners are released into the same neighborhoods, homes, schools, or peer groups that helped shape their criminal behaviors and led to their incarceration. Longer term effects may be more difficult to achieve, but they are possible and we will provide many examples in Section III of interventions that have been shown to reduce criminal behavior many years following the end of an intervention. Although sustainable effects are optimal, an intervention that can delay the onset of criminal behavior, even if it does not ultimately prevent it, is still effective, particularly in light of the evidence discussed in Chapter 3 showing that early onset of crime is linked to more frequent and serious criminal behavior over the life course.

Science is always worth the wait.

The New York Times, 2011

Effect size

Effect sizes were discussed earlier in the chapter, but a few additional comments are relevant here. The effect size is a measure of the relative strength and substantive significance of the intervention main effect. It is intended to indicate how big a reduction in criminal behavior can be expected from participation in the intervention.

The effect size may be an unstandardized measure or a standardized measure. When the outcome is measured in units that have some intrinsic meaning, like criminal acts, arrests, or victimizations, an unstandardized effect size is typically used. For example, an evaluation may report that the effect size represents an average reduction of four crimes per person for those receiving the intervention. If the outcome is measured in units that have no intrinsic meaning, like 3 points on a 10‐point scale intended to assess the seriousness of an offense, a standardized measure is typically used. In this case, the effect size will vary from 0 to 1, with small effect sizes ranging from 0 to 0.20, moderate effects ranging from 0.21 to 0.40, and values greater than 0.40 representing large effects (Cohen, 1988), as shown in Figure 4.4. An added advantage of the standardized effect size measure is that the effects of different interventions can be compared even when the measure of the outcome is different.

Figure 4.4 Effect sizes indicating the amount of criminal behavior reduced by an intervention.

Three pie charts of the effect sizes indicating the amount of criminal behavior reduced by an intervention. They feature small effect size (left), moderate effect size (middle), and large effect size (right).

Another advantage of the effect size is that it is not influenced by the sample size of an evaluation, as is the test for statistical significance. Both small and large samples can be problematic when trying to identify the substantive or practical value of an intervention if you have to rely solely on tests of statistical significance. Small samples are said to have low statistical power. The smaller the sample, the greater the treatment‐control difference must be to pass the 5% rule and achieve the accepted level of statistical significance.7 But a calculation of the effect size can be performed in these cases in order to identify meaningful changes in crime. Conversely, in a large study (e.g., involving 1000 or more participants), it is much easier to find a statistically significant main effect even if the intervention has virtually no meaningful effect on criminal behavior, as would be indicated by a very small effect size. It is useful, then, to consider both the statistical significance and effect size of an intervention. If the main effect is not statistically significant, the effect size has little practical significance; if it is statistically significant, the effect size provides an estimate of how powerful the intervention is for reducing criminal behavior.

Meta‐analysis and systematic reviews

Meta‐analysis is also a useful strategy in studies with small samples sizes and low statistical power. A meta‐analysis provides an “analysis of analyses” and allows the results from multiple evaluation studies to be combined and analyzed together. This process eliminates the small sample size problem because results from participants involved in all studies are analyzed together. Meta‐analyses are useful when multiple evaluations of an intervention have shown conflicting evidence of effectiveness, such as differences in the strength of the main effect or both positive and negative outcomes. Meta‐analyses provide an estimate of the average effect size across all the studies included in the analysis and help make sense of contradictory evidence. This estimate is also considered more reliable than those from a single study as multiple evaluations contribute to the estimate.

In Section III, we will refer to findings from two types of meta‐analysis. A meta‐analysis of a prevention program will analyze the overall effects on crime of a single program, such as Scared Straight or Functional Family Therapy, which has been tested in multiple studies. A meta‐analysis of a prevention practice will be based on multiple evaluations of different interventions that use a common change strategy or target a common outcome, such as the law enforcement practice of increasing surveillance in hot spots, or areas in a police beat known to have high levels of crime. Most meta‐analyses in crime prevention are of the second type. Meta‐analyses of single programs and policies are relatively rare, primarily because most of these types of interventions have been evaluated only one or two times, which is not enough to justify a meta‐analysis.

A related approach to summarizing the evidence from multiple evaluations involves a systematic review. These reviews are based on having a set of procedures, specified in advance, which will guide the search, selection, evaluation, and synthesis of evaluation findings from multiple studies. The procedures should be clearly described so that other researchers could replicate the review. As we will describe more fully in Chapter 5, the Campbell Collaboration was one of the first organizations dedicated to conducting systematic reviews of interventions in crime and justice. As described on their website (www.campbellcollaboration.org/what_is_a_systematic‐review), a systematic review must have: clear criteria regarding the types of evaluations that will be included and excluded from the analysis, which should include published and unpublished reports, as well as international studies; a description of exactly how the search for evaluations will proceed; a detailed plan describing how evaluation findings will be reviewed and coded; and use of meta‐analysis where possible to estimate the average impact of a program, practice, or policy. In addition, they require that all procedures and findings be reviewed by other scientists prior to publication. This is a very rigorous review process that provides reliable estimates of effect sizes for individual programs, practices, and policies when there are multiple evaluations. As we will also describe in Chapter 5, the Office of Justice Programs’ CrimeSolutions.gov website (http://www.crimesolutions.gov/Programs.aspx#programs) uses findings published by the Campbell Collaboration to identify effective and ineffective crime prevention practices.

Return on investment (ROI)

Last, but certainly not least, we want to mention a statistical analysis strategy that allows estimation of the economic value, or return on investment (ROI), of a prevention intervention. This type of calculation is intended to help decision‐makers evaluate taxpayers’ and society’s financial benefits and determine the worthiness of investing in particular programs, practices, or policies. It places a monetary value on the intervention outcomes, based on the effect sizes for all analyzed outcomes, and compares these benefits to the costs of the intervention. Although we will not get into the technical details of how the ROI is estimated, recall from Chapter 1 that crime has significant costs to society, including expenses related to law enforcement, corrections and drug treatment services. Such figures are all considered in the ROI.

The ROI is helpful not only for evaluating the economic importance of a single intervention (recall this question from Table 4.1) but also in comparing different interventions. For example, according to the Washington State Institute for Public Policy (2015), providing vocational training while adult offenders are in prison has an ROI of $13.22. That is, this practice returns $13.22 in savings to taxpayers for every $1.00 invested in it. In comparison, the provision of vocational education and job assistance to offenders after they are released from prison is $47.79. Such information can provide useful guidance to correctional officials when evaluating potential treatment services, though it should not be the only consideration. When describing crime prevention strategies in Section III, we will provide the ROI where it has been calculated. As you will see, some interventions have a negative ROI, meaning that their costs exceed their benefits. For example, the Scared Straight intervention has an ROI of ‐$201; it costs taxpayers about $201 for every dollar put into it (Washington State Institute for Public Policy, 2015). As you will read in Chapter 9, this is because Scared Straight has no demonstrated impact on criminal behavior.

A Critical View of Current Evaluation Research

Although we have already pointed out some of the limitations of evaluation research, we summarize some of the most serious concerns here given their importance for conducting evaluations and informing decision‐makers’ views about which interventions to replicate. A first concern is over developer bias, a type of bias resulting from the involvement of program developers in evaluations of their own interventions that can distort estimates of the intervention’s effectiveness. To date, most high‐quality evaluations of crime prevention programs have been undertaken by the program developer (Gorman and Conde, 2007), and such evaluations have been shown to provide higher estimates of a program’s effect size compared to studies led by independent investigators (Gandhi et al., 2007; Petrosino and Soydan, 2005). This situation means that claims of program effectiveness may be exaggerated, and that if replicated on a wide‐scale, many interventions will not produce the same impact on crime as we would expect based on their developer‐based evaluation evidence.

Publication bias can also contribute to inflated estimates of a program’s effect size. It is commonly accepted among scientists that academic journals accept articles that report positive evaluation findings at a higher rate than articles reporting no significant findings or iatrogenic effects (Hunter and Schmidt, 1990; McCord, 2003; Wilson, 2009). Unpublished studies receive far less attention and are often not included in systematic reviews of the evaluation literature (although the Campbell Collaboration is an exception). As a result, their findings will not be able to balance out the more positive effects documented in published studies, thus contributing to biases in the estimate of an intervention’s average effect size. Although it is difficult to determine exactly how serious this problem is, it is likely that when agency decision‐makers evaluate different options for addressing crime, they will not have all relevant information.

Next, there is the problem of evaluations failing to assess the long‐term effectiveness of prevention programs, practices, and policies. Historically, the majority of evaluations have focused only on short‐term effects, largely because it is more costly and difficult to conduct multiple follow‐up assessments. Just consider the challenge in getting participants to complete two surveys, then think about how much more difficult it is to secure their agreement to participate in repeated assessments and in tracking them down year after year to do so. While easier to conduct, short‐term evaluations could produce misleading results of an intervention’s effectiveness if the treatment initially shows positive effects, but weaker or iatrogenic effects emerge in the long term. This has been known to occur (e.g., McCord, 2003; St. Pierre et al., 2005). Conversely, interventions may have “sleeper effects” which are not present early on but become evident over time (e.g., Hawkins et al., 1999). Such programs may be dismissed as ineffective if outcomes are evaluated only in the short term.

Findings can also be misinterpreted and interventions falsely perceived to be ineffective when evaluating marginal effects. The failure to find a statistically significant main effect when a new intervention is compared to “treatment as usual” should not be interpreted as evidence that the intervention was ineffective. This assessment may be correct if it is known that the treatment provided to the control group had no impact on crime. Otherwise, this finding means that the new intervention is no more effective than that the control group received. Stated differently, if the control group treatment reduced crime, and the new intervention’s effects were no different than those of the control group program, it also produced positive changes in offending.

An issue that we have not yet considered is the potential for evaluations to overlook problematic side effects caused by participation in the intervention. Although the developer’s review of research relating to the theoretical and change model should provide some guidance about possible negative effects, due to practical reasons such as having limited time to conduct surveys or the evaluator’s specialization in a particular field of study, evaluations often limit their analysis of effects to only one or two outcomes. In doing so, they may fail to identify unanticipated negative effects on unassessed outcomes. For example, the National Women’s Health Initiative Study examined the effects of a medical intervention intended to reduce heart disease in women. During the evaluation, the treatment provided to study participants was shown to increase women’s risk for stroke and breast cancer, outcomes which were not expected but which happened to be measured by the evaluators (Manson et al., 2013). These iatrogenic effects eventually led to the suspension of the intervention. In crime prevention, a long‐term follow‐up of the Milwaukee Domestic Violence Experiment revealed that domestic violence perpetrators who were randomly assigned to be arrested for their offenses were three times more likely to die of a homicide than those assigned to receive only a warning but not an arrest (Sherman and Harris, 2014). A somewhat higher death rate was also found for victims whose perpetrators had been arrested. Such outcomes were unanticipated and could have easily been overlooked if deaths had not been examined.

There is a moral imperative for … randomized experiments in crime and justice. That imperative develops from our professional obligation to provide valid answers to questions about the effectiveness of treatments, practices and programs.

David Weisburd, 2003: 339

A frequently cited ethical issue related to experimental evaluations, particularly RCTs, is that they recruit individuals who could all benefit from an intervention, but half are denied treatment: those assigned to the control group (Weisburd, 2003). In our view, this concern lacks validity because it assumes that the intervention is effective. But, the main goal of evaluation research is to answer the question: does this intervention work? If the treatment was known to reduce crime, it would be delivered to all participants. The ethical question could easily be turned around to ask: how can the justice system mandate individuals into interventions that have not been evaluated and which have no evidence of effectiveness? Such practices seem unethical given that some interventions, even those intended or expected to reduce crime, have been shown to be harmful, as you will see in Section III.

Another critique of RCTs is that they cost too much and that their high costs make them impractical to conduct. This concern has merit. For example, even a RCT considered “modest” in scope and complexity, such as an evaluation of a school‐based prevention program involving 20 schools and 2 post‐tests, can cost several million dollars and take 5 years to complete. Larger studies cost even more. The National Women’s Health Initiative mentioned earlier, which involved over 160,000 American women, cost 625 million dollars (Parker‐Pope, 2011). Costs can easily reach $10 to $15 million if an intervention is subject to an initial RCT and one or two replications to increase its external validity. Costs place a heavy burden on federal governments, as most experimental evaluations of crime prevention efforts are funded by national agencies like the Department of Justice and the National Institute on Drug Abuse in the USA. Local agencies are unlikely to be able to secure large grants from federal agencies or to have the required resources to conduct their own high‐quality RCT or even QED, leading to a shortage of information about the effectiveness of interventions commonly implemented in communities. Nonetheless, it is possible for local programs to move through the early stages of the evaluation process and conduct process and pre‐post evaluations. These studies can provide local providers with valuable information about their interventions and increase their chances of eventually obtaining funding for an experimental evaluation.

For all of these reasons, there has been some resistance to claims that experimental evaluations and especially RCTs are the “gold standard” in evaluation research and that only programs evaluated using these designs should be promoted as effective (Lab, 2014; Laycock, 2002). Critics also raise concerns about the limited external validity of RCTs, claiming they are too artificial and cannot adequately capture processes occurring in the real world. The heavy reliance on measuring outcomes using quantitative data based on official or self‐reported data has been critiqued and arguments made for the use of qualitative data. Such studies might, for example, rely more heavily on in‐depth interviews and/or observations of program participants, implementers and implementing agencies to understand the complex processes that affect criminal behavior. Pawson and Tilley (1997) have proposed the use of “realistic” or “realist” evaluation strategies which rely on a combination of qualitative and quantitative data to assess outcomes and change mechanisms of an intervention operating in different contexts.

In our view, qualitative studies have made significant contributions to the development of criminological theory, hypotheses about the causes of crime and identification of potential strategies to prevent or control it. However, this type of research is not equipped to validate these hypotheses or test the validity of proposed causal mechanisms. It is concerned with observing natural conditions, not manipulating causal variables. The latter is the specific purpose of an experimental study and the foundation for making a strong argument for a causal argument and/or intervention’s effectiveness. We will return to this debate in the next chapter, when we describe variation in standards across agencies in determining what works to reduce crime. As you might realize, we will be making the case that such claims require the use of well‐conducted RCTs and strong statistical analyses such as those described in this chapter.

The Blueprints for Healthy Youth Development Experience Reviewing Evaluations

As a final commentary on judging the quality of evaluation research, we conclude this chapter by sharing some of the first author’s experiences leading the Blueprints For Healthy Youth Development project ( www.blueprintsprograms.com /). Since 1996, Blueprints staff and members of its scientific advisory board have systematically compiled and reviewed findings from over 1300 prevention interventions, many of them aimed at reducing crime, delinquency, violence, and drug use (Mihalic and Elliott, 2015). The goal of the project is to identify interventions that have been demonstrated as effective in high‐quality research designs. The merit of each research evaluation is judged according to the criteria described in this chapter. (See Chapter 5 for more detail about the Blueprints project and criteria.)

To date, about 80% of the interventions reviewed by Blueprints have not been subject to a credible scientific evaluation. These findings suggest that most crime prevention programs currently in use in the USA and internationally have never been assessed using the evaluation process described earlier. Of those that have been evaluated using experimental designs, most of the experiments were judged to have low merit because most failed to adequately address threats to internal validity. By far the most frequent problem encountered has been the threat of non‐equivalence caused by contamination of the random assignment to treatment and control groups in RCTs, significant pre‐test differences in QEDs, failure to perform an intent‐to‐treat analysis, and differential attrition that is not addressed in the analysis (Mihalic and Elliott, 2015). When high‐quality evaluations have been performed, they have often indicated no significant effects on outcomes and a few studies have reported negative and harmful effects. These problems have resulted in a relatively short list of interventions identified as effective by Blueprints.

On the positive side, Blueprints has identified some interventions that have been evaluated using very high‐quality research designs and which have demonstrated positive effects and good ROI rates. Moreover, over time, the quality of crime prevention evaluations has improved. RCTs are now more frequently employed and new methods for addressing some of the threats to internal validity have been developed and utilized by evaluators (for further commentary on progress made in evaluation research, see: Farrington and Welsh, 2005; Flay and Collins, 2005).

Summary

This chapter has provided a complete but fairly general overview of the purpose, process, and nature of evaluation research. Evaluation has three purposes: (i) to provide feedback to developers to help them improve their interventions, (ii) to test the validity of the causal processes underlying the intervention, and (iii) to provide critical information to funders, policy‐makers and other decision‐makers so that they can make informed choices about what interventions to implement to reduce crime.

Evaluation is a process with stages that involve a progression from the use of more basic to more sophisticated research designs. This sequence begins with a process evaluation, followed by a pre‐post impact evaluation and then an experimental outcome evaluation. Each stage provides useful answers to particular research questions and information from all parts of the process should be considered when judging the overall effectiveness of an intervention. The merits of different types of evaluation are judged by their construct validity, internal validity, and external validity. In practice, few evaluations have provided good information on external validity and they are judged primarily on how well they address the threats to internal validity. Based on these criteria, RCTs are considered to be the “gold standard” when trying to determine the effectiveness of crime prevention programs, practices, and policies. When they are well conducted, RCTs do the best job of addressing the threats to internal validity and ruling out alternative explanations for the study findings.

We also reviewed different types of statistical analyses that can be used to help rule out alternative explanations of observed changes in criminal behaviors. Analytic procedures can also be used to confirm the underlying logic model and causal processes of the intervention, test for conditions that moderate the effects of the intervention and provide valuable feedback about the strengths and weaknesses of the intervention.

This chapter has identified the optimal practices to use when evaluating crime prevention interventions. Some of these methods have been critiqued and many are not followed in practice. Many experimental evaluations of crime interventions have weak levels of external or internal validity, have limited their design and analysis to short‐term rather than long‐term effects, may overstate their effect size, and are very expensive to undertake. For all of these reasons, Section III of this text, which describes interventions shown to be effective in reducing crime when evaluated in high‐quality research projects, may be shorter than you expect and may not include interventions that you may have heard about or seen used. In fact, many commonly used crime prevention strategies have not been evaluated using rigorous research methods and some have failed to show positive effects on offending. However, this situation is slowly changing and evaluations are increasingly relying on better methods and statistical procedures.