NO PLAGIARISM DUE MONDAY MAY 6, 2019. ATTACHED ARE CHAPTERS TO ASSIST WITH ASSIGNMENT

profilemztclass82
CRJ305CRIMEPREVENTIONCH5.docx

5 Establishing a Standard for Judging Intervention Effectiveness

Learning Objectives

Upon finishing this chapter, students should be able to:

· Understand why it is important to have standards, and consensus on these standards, for determining the effectiveness of crime prevention programs, practices, and policies

· Summarize the challenges in establishing rigorous, agreed‐upon standards for identifying effective crime prevention programs, practices, and policies

· Identify multiple web‐based databases that list effective crime prevention programs, policies, and practices

· Summarize the differences in standards and review processes used to determine program effectiveness across these databases

· Recommend a set of rigorous standards that should be used in future determinations of “what works” to prevent crime.

Introduction

Have you ever received an inoculation, vaccine, or other preventive medicine (e.g., a flu shot) in order to prevent an illness? How confident were you that this treatment would prevent you from getting sick? How worried were you that the medicine would make you feel worse?

When was the last time you were prescribed or bought over‐the‐counter medicine to cure a cold or other illness? How confident were you that this medicine would work to alleviate your symptoms? Were you worried about potential side effects, particularly whether or not taking the medicine would make you feel worse rather than better?

Although there are some exceptions, most people tend to place a lot of confidence in the medical profession, in doctors and in the medications offered to us when we are sick. When prescribed a medicine, we usually assume it will work and that harmful side effects will be minimal. Most of the time, we do not even give the topic much thought: we have faith that there are regulations governing the testing and approval of medications and that some capable group or organization is watching to make sure these policies are followed. Don’t we?

Can we have this same confidence in the work of prevention scientists? Are there any standards governing the creation, testing, and distribution of crime prevention programs, practices, and policies? Are there any regulations for how this work is done or any agencies or organizations charged with ensuring that such regulations are upheld and that harmful interventions are not implemented?

This chapter will consider these issues. We will describe the degree to which there are standards governing the work of prevention scientists and the challenges faced when trying to create, reach consensus on and apply these standards in crime prevention. We will identify the varied sets of standards currently used by different groups when identifying and recommending effective crime prevention strategies, and the resulting disparity in the interventions promoted on web‐based databases as “evidence‐based,” “best practices,” “model,” or “promising.” The chapter will conclude with our recommendations for a set of standards we would like to see used more consistently across agencies seeking to identify “what works” in crime prevention.

The Need for Standards about “What Works” to Prevent Crime

Today, many federal agencies emphasize the importance of using “what works” to improve public health and prevent or reduce crime, meaning that these problems should be addressed using evidence‐based or effective programs, practices, and policies. Yet, there are many different views about “what works”, what constitutes “evidence” and how to determine if a particular intervention is “effective” in achieving a particular outcome. These views vary across agencies, federal policy‐makers, scientists, and potential implementers of an intervention.

Although the increased focus on using evidence‐based interventions is encouraging, without clear definitions and agreed‐upon standards describing what this means, success is unlikely to be realized. Without such guidelines, practitioners have to decide for themselves what is effective or evidence‐based and what is not, and these decisions are very complicated. As we will describe, many organizations now promote evidence‐based prevention programs, but the number, type, and quality of interventions recommended vary widely across these agencies, resulting in a confusing array of options for practitioners. Some agencies judge the merits of interventions using strict and scientifically rigorous criteria, such as that described in Chapter 4 , when trying to identify what works, which usually results in a smaller list of recommended programs. Other agencies have lower standards, which generally increases the number of options but may also result in programs being nominated that have a weaker evidence base. There is a greater chance, then, that these interventions will be ineffective in reducing crime and a greater likelihood that taxpayers’ money will be wasted if such programs are replicated in communities.

By failing to provide practitioners with consistent, easy‐to‐access and valid information about what works, we risk their becoming frustrated with and skeptical about prevention science and scientists. They may ask themselves, for example: how can I trust a bunch of people who can’t even agree among themselves about what works? In the absence of good information, agencies and communities may also implement ineffective strategies, and when positive results are not achieved, they will become even more doubtful about the scientific process and prevention scientists.

How can we prevent this skepticism and ensure that effective crime prevention efforts are implemented and ineffective strategies avoided? The answer is straightforward, at least in theory: we must reach agreement on a set of standards or guidelines for what constitutes evidence that a program works and then ensure that these guidelines are consistently and widely used to identify effective interventions. If this is not accomplished, there is a real danger of promoting programs of questionable value and losing public confidence in prevention science.

How Politics, Science, and Practice Influence Decisions about What Works

If we want to have any chance of success in achieving consensus on guidelines for determining what works, we must first understand why such differences exist. Why is there such disparity across agencies and individuals in the standards used to identify effective crime prevention programs? What stands in the way of a “master list” of recommended programs and an agreed‐upon set of standards for identifying programs that all parties can agree to? In criminology, as in most disciplines, determinations about the level of evidence needed to determine program effectiveness, and which programs should be promoted based on this evidence, are influenced by politics, scientific ideals, ethical concerns, and practical considerations.

First, political ideologies strongly influence evaluation standards and the identification of effective and ineffective programs. One of the major differences across different lists of what works is whether or not a particular intervention must be evaluated using a randomized control trial (RCT). Although the RCT is widely promoted as the “gold standard” in evaluation research, meaning it is the best methodology for determining effectiveness (Cook and Campbell, 1979; Sherman, 2003; Weisburd, 2010), criminal justice agencies have not generally advocated for the use of RCTs. For example, the agencies responsible for funding most criminal justice research in the USA, the National Institute of Justice (NIJ) and the Office of Juvenile Justice and Delinquency Prevention (OJJDP), have historically not required that evaluations rely on RCTs. Instead, most of their funded research has supported non‐randomized trials. There have been some exceptions, usually coinciding with the arrival of a leader who is personally supportive of RCTs, but generally, these agencies have not widely encouraged the use of RCTs (Farrington, 2003; Weisburd, 2003). These views strongly influence scientific practices: when the political climate does not fund or support RCTs, scientists will be less likely to use and endorse RCTs.

There was a similar reluctance to support RCTs in the field of medicine. Until the early 1960s, the RCT was often viewed as too difficult, costly, and unethical to implement. But, in 1962, a law passed by President Kennedy dramatically changed this point of view in the USA. At that time, a sedative used to treat morning sickness in pregnant women, thalidomide, was shown to result in birth defects in children born in Europe, Canada, and other countries. The public demanded that the government do something to prevent such catastrophes in America, and the result was federal legislation, in the form of the US Kefauver‐Harris Drug Amendments Acts. These laws called for stricter guidelines and more safeguards to ensure that drugs sold in the USA went through rigorous testing prior to being released. They required that drug manufacturers prove scientifically, using RCTs, that medications were not only effective but also safe. In time, the National Institutes of Health, which provides billions of dollars to fund health‐related research each year, adopted similar standards and strongly emphasized the use of RCTs in their funded studies. In the fields of medicine and public health, thanks to political and public support, RCTs are now widely considered as the way to determine effectiveness.

Failure to discover whether a program is effective is unethical.

Boruch, 1975: 135

Politics do matter when it comes to emphasizing particular standards for determining effectiveness. It is also true, of course, that political forces can be difficult to overcome, and there is often political pressure to support particular programs even if scientific evidence indicates they are not effective. Federal and state agencies may be reluctant to identify as ineffective strategies which have been supported by public funds and endorsed by political allies. For example, mandating that some juvenile offenders be tried in the adult criminal justice system, Scared Straight interventions that involve visits to prisons for youth considered at high risk for engaging in crime, and the Drug Abuse and Resistance Education (D.A.R.E.) program taught in schools by police officers have received a lot of federal funding and have much political and popular support. This backing, and the considerable amount of personnel, materials, and training invested in these strategies, is difficult to counteract. Even though multiple studies have shown that each strategy is ineffective (see Section III for the facts), they are hard to remove from society. How can political and public pressures to use ineffective strategies be overturned? We will explore that question more fully in Section IV of this text, but the solution, we believe, lies in increasing support for prevention science, publicizing more broadly interventions that do work, requiring that criminal justice agencies be accountable for the outcomes they achieve with public funds, showing that effective prevention is cost effective and helping communities implement effective strategies.

Scientists, like politicians, play a key role in determining the standards to be used in judging program effectiveness. Unfortunately, they can be just as stubborn in their convictions about what works and what does not and in their views about what is needed to determine effectiveness. Returning to the example of the RCT, some criminologists support RCTs as the gold standard of evaluation research while others doubt their necessity, question their feasibility and emphasize their costs and difficulty in comparison to other methodologies (Hough, 2010; Sampson, 2010). Some argue that the methodological rigor of the RCT should not be the sole basis on which intervention effectiveness is determined, in large part because experiments are too artificial and have low external validity (Sampson, Winship, and Knight, 2013; Weisburd, 2003). That is, some critics contend that the experimental conditions manipulated by scientists to fit the demands of RCTs have little resemblance to everyday experiences and practices (Pawson and Tilley, 1994). As a result, an intervention deemed effective for participants in a research trial may not actually work in practice.

Sampson and colleagues (2013) contend that although RCTs can tell us what works, they are less able to tell us what will work in naturalistic conditions. They recommend that much more evidence must be accumulated prior to making recommendations for the use of particular interventions. Specifically, they would like to wait until programs document in detail the populations for whom they are most effective, exactly how and why outcomes are achieved and the larger contexts in which positive effects are most likely to be evidenced. Others assert that quantitative methodologies cannot capture all of the processes, outcomes and experiences potentially related to program effectiveness and that qualitative methods are better able to capture why and how behaviors change (Pawson and Tilley, 1994). Finally, some point out that local programs are usually not subject to randomized studies, because knowledge of how to conduct RCTs and sufficient funding to implement this methodology are often lacking in communities. This means that lists of what works often exclude local programs, even though these may be most appealing to local practitioners.

As discussed in Chapter 4, ethical concerns about experiments and RCTs have also been raised. Randomizing some individuals to a potentially effective treatment and withholding such services from other individuals is troubling to many. There may be concerns that decisions about who receives treatment and who does not will be biased and that, regardless, all individuals should be provided potentially beneficial treatments (Weisburd, 2003). Others point to the difficulty in convincing judges, parole officers, schools, families, or communities to agree to randomization. If participants are unlikely to agree to such a methodology, then more feasible approaches should be adopted. Finally, some suggest that randomized experiments cost more than other types of evaluations.

All of these concerns result in resistance to standards that limit the determination of what works to interventions that have been evaluated using RCTs. Most critics of the RCT advocate for more flexible standards that will allow for a variety of research evaluation methods and types of evidence. The advantage of loosening criteria is that doing so will likely increase the number of recommended programs that communities can adopt. Having more options can, in turn, increase the number of interventions in use. However, there is a corresponding increase in the risk of program failure, in that less rigorous standards generally correspond to more uncertainty that programs actually work. Then, as we have discussed, if programs are widely replicated and fail to produce results, confidence in science is undermined. Higher standards carry a lower risk of program failure, but also fewer choices for communities. When few options are available, practitioners may feel forced to implement a strategy that they do not agree with or which is difficult given their resources or capacity. Aware of this dilemma, public agencies may feel pressured to loosen standards of effectiveness.

This quickly begins to resemble a “chicken and egg” dilemma and leaves us with the question: how can we balance scientific standards with practical considerations in a way that will best ensure reductions in crime? We try to answer this question in the conclusion of this chapter. Before doing so, however, let’s review in more detail the standards that are currently used to determine program effectiveness and how these standards have come to be.

Historical Development of “What Works” Databases

The emphasis on identifying evidence‐based crime prevention strategies using high‐quality research evaluations emerged in the mid‐1990s in the USA. During this period, Martinson’s (1974) and others’ claims that “nothing works” (see Chapter 2) were still widely accepted. It was also recognized that many crime prevention efforts had not been evaluated for effectiveness and if they were, evaluations were generally of poor quality (Farrington, 2003). This was also a period in which crime rates, especially for youth‐perpetrated offenses, were surging, drawing public attention to the crime problem and increasing demands to do something to reverse these trends. What followed was the publication of various guides describing effective crime prevention strategies, though there was little description of the standards used to produce the lists of effective programs contained in these guides. Early guides to “what works” included the following:

1995: Guide for Implementing the Comprehensive Strategy for Serious, Violent, and Chronic Juvenile Offenders, Office of Juvenile Justice and Delinquency Prevention (OJJDP). The goal of this publication was to identify “Effective” and “Promising” programs evaluated in the USA. Given its sponsorship by OJJPD, all interventions in the Guide were aimed at preventing juvenile delinquency. The Guide also emphasized comprehensive approaches to prevention and included an array of universal, selected, and indicated programs – from interventions implemented early in life for non‐offenders to juvenile justice programs for incarcerated offenders. The Guide did not include information about criteria used to determine program effectiveness. It simply stated that Promising programs “did not have sufficiently strong research designs” compared to Effective programs (Office of Juvenile Justice and Delinquency Prevention, 1995: xi).

1996: Communities That Care Prevention Strategies Guide. This book was published as a resource for communities implementing the Communities That Care (CTC) prevention system (see Chapter 7 to learn more about CTC). It listed interventions considered to be “Effective” in reducing juvenile substance use and delinquency. Determinations about effectiveness were based on evidence showing that programs tried to reduce risk factors and increase protective factors and “showed positive effects in high‐quality” evaluations (Communities That Care: Prevention Strategies Guide to What Works, 1996: xvi).

1997: Preventing Drug Use Among Children and Adolescents: A Research‐Based Guide, National Institute on Drug Abuse (NIDA). This was the first attempt by NIDA to compile information on effective substance use prevention programs appropriate for children and adolescents. The Redbook (so named because of its red cover) reviewed information on risk and protective factors related to substance use and how interventions could target these factors to reduce substance use. Examples of “research‐tested” prevention programs were also given but criteria used to determine inclusion in the book were not provided.

1996–1997: Blueprints for Violence Prevention. In 1996, the Center for the Study and Prevention of Violence at the University of Colorado began the Blueprints Initiative. Its goal was to identify effective youth violence and substance use prevention programs based on systematic reviews of evaluation studies and rigorous, clearly identified criteria to judge effectiveness (see later for more details on these criteria). By 1997, 10 “Model” programs with strong evidence of effectiveness had been identified, were listed on the Blueprints website and were described in a series of short books (i.e., “blueprints”) designed for practitioners and agencies (Elliott, 1997).

1996–1997: Preventing Crime, What Works, What Doesn’t, What’s Promising. In 1996, Congress required that the US Attorney General provide a comprehensive evaluation of the effectiveness of crime prevention programs funded by the Department of Justice. The review, conducted by criminologists at the University of Maryland, included juvenile‐ and adult‐focused prevention programs as well as strategies that could be implemented by communities, law enforcement agencies, and correctional institutions (Sherman et al., 1997). In addition to listing effective and ineffective strategies, the review was important in creating a rating system known as the “Maryland Scale of Scientific Methods” (Sherman et al., 1998). It outlined specific criteria for differentiating between “Effective,” “Promising,” and “Ineffective” programs based on the methodological rigor of their evaluations.

2001: Youth Violence: A Report of the Surgeon General. This report was created by the US Department of Health and Human Services at the request of the US Surgeon General at that time, David Satcher, largely in response to concerns about increases in youth violence in the 1990s and the 1999 shootings at Columbine High School in Littleton, CO. This review of prevention programs was guided by life course developmental theories and focused on universal, selective, and indicated strategies aimed at youth. It did not review interventions for law enforcement or correctional agencies. The Report classified interventions as “Model,” “Promising,” and “Ineffective,” based on criteria including the Maryland Scale of Scientific Methods, the Blueprints ratings, and effect sizes produced by the interventions. Coincidently, its Senior Scientific Editor, Dr Del Elliott, is also the developer of the Blueprints Initiative and co‐author of this textbook.

Historical Development of “What Works” Databases

The emphasis on identifying evidence‐based crime prevention strategies using high‐quality research evaluations emerged in the mid‐1990s in the USA. During this period, Martinson’s (1974) and others’ claims that “nothing works” (see Chapter 2) were still widely accepted. It was also recognized that many crime prevention efforts had not been evaluated for effectiveness and if they were, evaluations were generally of poor quality (Farrington, 2003). This was also a period in which crime rates, especially for youth‐perpetrated offenses, were surging, drawing public attention to the crime problem and increasing demands to do something to reverse these trends. What followed was the publication of various guides describing effective crime prevention strategies, though there was little description of the standards used to produce the lists of effective programs contained in these guides. Early guides to “what works” included the following:

1995: Guide for Implementing the Comprehensive Strategy for Serious, Violent, and Chronic Juvenile Offenders, Office of Juvenile Justice and Delinquency Prevention (OJJDP). The goal of this publication was to identify “Effective” and “Promising” programs evaluated in the USA. Given its sponsorship by OJJPD, all interventions in the Guide were aimed at preventing juvenile delinquency. The Guide also emphasized comprehensive approaches to prevention and included an array of universal, selected, and indicated programs – from interventions implemented early in life for non‐offenders to juvenile justice programs for incarcerated offenders. The Guide did not include information about criteria used to determine program effectiveness. It simply stated that Promising programs “did not have sufficiently strong research designs” compared to Effective programs (Office of Juvenile Justice and Delinquency Prevention, 1995: xi).

1996: Communities That Care Prevention Strategies Guide. This book was published as a resource for communities implementing the Communities That Care (CTC) prevention system (see Chapter 7 to learn more about CTC). It listed interventions considered to be “Effective” in reducing juvenile substance use and delinquency. Determinations about effectiveness were based on evidence showing that programs tried to reduce risk factors and increase protective factors and “showed positive effects in high‐quality” evaluations (Communities That Care: Prevention Strategies Guide to What Works, 1996: xvi).

1997: Preventing Drug Use Among Children and Adolescents: A Research‐Based Guide, National Institute on Drug Abuse (NIDA). This was the first attempt by NIDA to compile information on effective substance use prevention programs appropriate for children and adolescents. The Redbook (so named because of its red cover) reviewed information on risk and protective factors related to substance use and how interventions could target these factors to reduce substance use. Examples of “research‐tested” prevention programs were also given but criteria used to determine inclusion in the book were not provided.

1996–1997: Blueprints for Violence Prevention. In 1996, the Center for the Study and Prevention of Violence at the University of Colorado began the Blueprints Initiative. Its goal was to identify effective youth violence and substance use prevention programs based on systematic reviews of evaluation studies and rigorous, clearly identified criteria to judge effectiveness (see later for more details on these criteria). By 1997, 10 “Model” programs with strong evidence of effectiveness had been identified, were listed on the Blueprints website and were described in a series of short books (i.e., “blueprints”) designed for practitioners and agencies (Elliott, 1997).

1996–1997: Preventing Crime, What Works, What Doesn’t, What’s Promising. In 1996, Congress required that the US Attorney General provide a comprehensive evaluation of the effectiveness of crime prevention programs funded by the Department of Justice. The review, conducted by criminologists at the University of Maryland, included juvenile‐ and adult‐focused prevention programs as well as strategies that could be implemented by communities, law enforcement agencies, and correctional institutions (Sherman et al., 1997). In addition to listing effective and ineffective strategies, the review was important in creating a rating system known as the “Maryland Scale of Scientific Methods” (Sherman et al., 1998). It outlined specific criteria for differentiating between “Effective,” “Promising,” and “Ineffective” programs based on the methodological rigor of their evaluations.

2001: Youth Violence: A Report of the Surgeon General. This report was created by the US Department of Health and Human Services at the request of the US Surgeon General at that time, David Satcher, largely in response to concerns about increases in youth violence in the 1990s and the 1999 shootings at Columbine High School in Littleton, CO. This review of prevention programs was guided by life course developmental theories and focused on universal, selective, and indicated strategies aimed at youth. It did not review interventions for law enforcement or correctional agencies. The Report classified interventions as “Model,” “Promising,” and “Ineffective,” based on criteria including the Maryland Scale of Scientific Methods, the Blueprints ratings, and effect sizes produced by the interventions. Coincidently, its Senior Scientific Editor, Dr Del Elliott, is also the developer of the Blueprints Initiative and co‐author of this textbook.

Recommendations for Achieving Consensus on Standards of Evidence

As all this information makes clear, at present, there is limited consensus regarding the standards of scientific evidence that should be used to designate or certify an individual program as effective or evidence‐based. The current lists use different selection processes and scientific standards to review and classify interventions. Not all use systematic search processes that take into account all relevant studies. Furthermore, they are not all regularly updated as new findings emerge. These limitations are problematic given recent increases in the number of program evaluations being conducted (Fagan and Eisenberg, 2012; Farrington and Welsh, 2005), and because findings from program replications are not always consistent with results from earlier studies (Valentine et al., 2011). Databases also vary in how well they describe their criteria for determining intervention effectiveness and in the level of rigor they apply when making these decisions. The result, as illustrated in Table 5.6, is that the same intervention can be considered very effective, somewhat effective, or not at all effective depending on the database one consults. This is confusing for the public, for policy‐makers, for funding agencies, for practitioners, and for scientists as well. The conflicting recommendations also threaten to undermine everyone’s confidence in the power of prevention science.

What can we do to improve what is currently a very messy process? How can we move towards a system – such as that used in medicine to regulate drug sales – that will certify programs as effective or ineffective, using consistent and transparent guidelines that will appeal to all stakeholders? Our recommendations are threefold. First, we call for more consistent use of the systematic review process to identify effective crime prevention programs, practices, and policies. Secondly, where there is enough evidence (e.g., at least three evaluations of the same or very similar programs), we recommend the use of meta‐analyses to calculate standardized effect sizes of programs and practices. Thirdly, where evidence has not accumulated sufficiently to allow for a meta‐analysis of the overall impact of an intervention, we put forth a set of rigorous but increasingly achievable criteria for determining program effectiveness.

As we have discussed, the systematic review process involves having a clearly described set of explicit procedures to find and evaluate the effectiveness of particular interventions. A comprehensive search of all studies that have evaluated a particular program, practice, or policy should be conducted and all such evaluations, not a subset of them, should be considered in the review (Wilson, 2009). In addition, the review should be as up‐to‐date as possible, with new findings incorporated as they emerge. When a sufficient number of evaluations are identified using a comprehensive search process, a meta‐analysis should be conducted to synthesize the findings and estimate an average effect size across all studies. This methodology allows for a more precise and reliable judgment of the degree to which a program, practice, or policy is likely to affect crime rates, as it takes all relevant research into account. Effect sizes are also fairly easy for practitioners and policy‐makers to understand, as a value of “zero” clearly indicates no effect on crime and scores between “zero” and “one” indicating a linear progression from having little effect to having a large impact on crime.

Unfortunately, relatively few programs currently have a sufficient number of high‐quality evaluations to allow for a reliable meta‐analysis of their impact. In these cases, we still need a standardized and rigorous method for identifying what works. Our recommendation is for databases to utilize the criteria developed by the Federal Collaboration on What Works (2005). In 2005, this group of federal agencies5 was directed by the White House Task Force on Disadvantaged Youth to set standards to be used across all federal agencies to identify effective prevention programs. After some no doubt very heated debate, the group agreed that determinations of program effectiveness should be based on ratings of: (i) the quality of the evaluation design, (ii) replication of the intervention, (iii) independence of the program developer and the program evaluator, and (iv) sustainability of effects after the intervention had ended (Working Group of the Federal Collaboration on What Works, 2005). Criterion #1, #2, and #4 should sound familiar, as they are already used by many groups and have been discussed in this chapter. Criterion #3, however, has not yet been discussed. This standard is based on the desire to keep science objective and uncontaminated from the potential bias that exists for program developers to show positive effects of their interventions (Gorman, 2005; Petrosino and Soydan, 2005). To guard against this bias, there is a preference that evaluations be conducted by researchers who are independent from program developers.

Under the guidelines of the Federal Collaboration on What Works, to be considered Effective,6 a program would need to have statistically significant, positive effects in a well‐conducted RCT, have at least one external or independent RCT replication, demonstrate evidence of sustained effects at least one year post‐intervention from at least one study and have no iatrogenic or harmful effects. The standard to be considered Promising required statistically significant positive effects from at least one well‐conducted RCT or QED (i.e., no replication required), sustained effects at least one year post‐intervention and no evidence of iatrogenic effects. The guidelines also allowed for identification of Ineffective, Inconclusive, and Harmful programs. Each designation had to meet all the requirements to be Promising, but instead of showing positive effects, Ineffective programs would demonstrate no statistically significant effects, Inconclusive programs would have mixed effects, and Harmful programs would show negative effects.

Although the goal of this group was to create uniform standards to be used by all federal agencies, including criminal justice agencies like NIJ and OJJDP, the criteria were never adopted by any organizations. We suspect there was reluctance to do so because the standards are very high and enforcement of these criteria would limit the number of interventions that could be recommended. As is evident in Table 5.7, none of the databases we reviewed in this chapter has fully adopted these standards. However, the criteria for consideration as a Blueprints Model program approach this level of rigor and, in Spring 2015, Blueprints decided to create a new classification: Model + programs. These interventions have to meet all the criteria to be certified as Model, as well as the independent evaluation criterion.

Table 5.7 Comparison of current databases with the federal collaboration on what works (2005) criteria.

Database

Study design for top rating

Replication

Independent replication

Sustained effects ≥1 year post‐intervention

Blueprints (Model)

RCT

Yes

Yes, for Model Plus programs

Yes

Coalition (Top Tier)

RCT

Yes

No

No

MPG/CS (Effective)

RCT or QED

No

No

No

NREPP

n/a, but requires a RCT or QED to be reviewed

No

No

No

C2

n/a

Yes

No

No

Despite the legitimate concern of limiting the number of programs recommended for use, we strongly advocate that the What Works standards be adopted and applied across agencies and databases. In our opinion, the government has a responsibility to spend public monies responsibly, to prevent public health problems, and to avoid harming individuals. As such, programs advocated by the government, which will result in considerable financial and human investment, should meet a high standard of effectiveness.

With this view in mind, we have some additional recommendations to ensure that standards have practical and scientific merit. First, given the importance of ensuring that interventions are based on credible theories and knowledge of the risk and protective factors associated with crime, we recommend that the Federal Collaboration on What Works criteria be expanded to include a rating of Intervention Specificity similar to that used in Blueprints. This would mean that, to be considered effective, interventions would have to clearly identify and describe their theoretical and empirical foundations, logic model, and the types of individuals or areas targeted for intervention. Even though we suspect that programs that fail to meet this criterion would be unlikely to achieve positive outcomes, having the standard draws program developers’ and users’ attention to the importance of these issues. Secondly, we advocate for the addition of the Blueprints Dissemination Readiness standard. Designating a program as effective and desirable for replication makes no sense if the program is not available for use. In addition, as we will discuss in Chapters 10 and 11, prevention programs can be challenging to implement, and if program developers lack the capacity to train and provide ongoing support to implementers, then successful replication is much less likely to occur.

We hesitate to add more criteria to our set of recommended standards at this point, for fear of limiting the number of options available to communities. However, as replications and evaluations continue to emerge, it may be possible to add two more standards: an analysis of mediating factors that produce behavioral changes and a cost–benefits analysis. As we discussed in Chapter 4, mediation analyses allow evaluation of the degree to which the risk and protective factors targeted for change are actually altered by the intervention and lead to anticipated reductions in crime. Demonstration of mediating pathways allows for a better understanding of why a program works and strengthens claims of its effectiveness. Documentation of the financial costs and benefits of an intervention has important practical implications. Potential users have limited funds and will want to know which program provides a “bigger bang for the buck.” Savvy consumers will want to compare the relative costs and effect sizes of different programs. It is also true that many practitioners and policy‐makers will be reluctant to spend funds on a new intervention even if it has scientific evidence of success. Being able to show that the program will generate financial returns as well as improve crime rates can help encourage adoption and widespread dissemination, a subject we will return to in Section IV. Although evaluations are increasingly incorporating analysis of mediating effects and costs versus benefits, this information is still not readily available for most programs and, as such, we will not yet recommend they be added to a list of required standards.

Summary

This chapter has reviewed the processes by which programs, practices, and policies are determined to be effective in preventing crime and the many challenges associated with making such recommendations. We have discussed the importance of setting high standards to differentiate effective and ineffective interventions, identified the range of standards that currently exist, and made recommendations for a set of rigorous standards we would like to see adopted by agencies in the future. We compared six web‐based databases that are currently in use and periodically updated in terms of the criteria and methods used to review intervention effectiveness and demonstrated how these differences result in different determinations across lists about what works.

As emphasized throughout this chapter, some lists have more rigorous criteria than others, which affects both the number of interventions that are designated as effective and our level of confidence that, if replicated, recommended programs will be likely to reduce crime. These two outcomes are inversely related: higher standards result in fewer nominated programs but a greater level of certainty in program effectiveness. In order to balance these two demands, our textbook will rely on multiple databases when describing interventions in Section III . The Blueprints and Coalition databases have the most rigorous standards for assessing the quality of prevention programs, and we will prioritize their recommendations. We rely more heavily on Blueprints, given that it is more comprehensive approach and has crime as a priority outcome of interest. Because Blueprints and the Coalition focus on programs intended to change youth outcomes, and because they do not review prevention practices, we will rely on the MPG/CrimeSolutions.gov and the C2 databases to identify adult‐focused interventions as well as law enforcement and situational crime prevention programs and practices. Since these databases often rely on less rigorous criteria, we will note their recommendations with some caution. We will not include ratings by NREPP given their reliance on less rigorous criteria.