U9D1-28 - Exploring Outcome-Based Evaluation Options - Please follow all instructions outlined below. Do not skip anything. Readings and Forms are attached.

profiledrcdopen82
Chapter11ImpactProgramEvaluationandHypothesisTesting.pdf

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 1/8

Chapter 11 Impact Program Evaluation and Hypothesis Testing

Chapter Overview The purpose of this chapter is to explain:

Impact program evaluation in more detail Theory success, theory failure, and program failure The logic model relationship between theory (cause), intervention (program), and result (effect) The two major reasons why programs fail The major types of impact program evaluations

The following topics are covered in this chapter:

Differentiating Impact Program Evaluation From Performance Measurement Impact Program Evaluation Impact Program Evaluation and Hypothesis Testing Research Designs for Impact Program Evaluation

Single Group Pretest/Posttest Design Nonequivalent Comparison Group Design Randomized Experimental Design

Summary Case Example Review Questions

Differentiating Impact Program Evaluation From Performance Measurement This chapter deals with the special case of what can be called impact program evaluation. Impact program evaluation differs from other types of evaluation in that the focus here is on changes in program participants and not on changes in organizations and communities. Impact program evaluation seeks to ask and answer the difficult question: Does this program actually have an impact and achieve positive outcomes (results) for participants? Before beginning a detailed discussion of impact program evaluation, it is useful to further distinguish it from performance measurement. Because impact program evaluation and performance measurement can, and frequently do, utilize the same types of data and information, confusion continues to exist over how these two assessment approaches differ (Magnabosco & Manderscheid, 2011; Nielsen & Ejler, 2008).

McDavid, Huse, and Hawthorn (2013, p. 324) identify a few key features that are important in distinguishing the purposes of impact program evaluation from those of performance measurement, as illustrated in Table 11.1.

Frequency—Impact program evaluation tends to be episodic, while performance measurement is an ongoing activity. Issue(s)—Impact program evaluation is frequently driven by specific stakeholder questions (e.g., Does the program work? How much did the program improve the lives of participants?). Performance measurement deals with more general performance issues that are relatively constant over time (e.g., Is the program accomplishing its output, quality, and outcome objectives?) Attribution of outcomes—Impact program evaluation is concerned with determining if the program actually caused the outcome (result) achieved or if the outcome (result) was caused by other external factors. In performance measurement, attribution is generally assumed.

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 2/8

Source: Adapted from McDavid, J., & Hawthorn, L. (2013). Program Evaluation & Performance Measurement (2nd ed.). Thousand Oaks, CA: SAGE Publications, Inc. Reprinted with permission.

Impact Program Evaluation Impact program evaluation attempts to demonstrate that an outcome (result) is attributable to the program and not some other variable. An impact program evaluation seeks to establish acause-and-effect relationship between a program and its outcome (result). An impact program evaluation can demonstrate that a human service program is (1) successful, or (2) unsuccessful due to theory failure, or (3) unsuccessful due to program failure. Figure 11.1 illustrates these three possibilities. All three of these findings are related to the intervention hypothesis that underpins a human service program and the implementation of the program design.

Using logic model terminology, a successful program is one where there is alignment between the program hypothesis (theory) and implementation of the program design (cause) that produces the desired outcome (result). We can call this theory success. An unsuccessful program usually occurs for one of two reasons:theory failure or program failure.

Theory failure takes place when there is alignment between the program hypothesis (theory) and implementation of program design, but the program does not produce the desired outcome (result). When this situation occurs, it is labeled theory failure because if the hypothesis was valid, then implementation of the program design should have produced the desired outcome (result). Because the desired outcome (result) was not produced, the hypothesis is not supported and therefore the theory is not supported. Program failure takes place when there is misalignment between the program hypothesis (cause) and implementation of the program design, which as a result fails to produce the desired outcome (effect). This situation should not come as a surprise. When a program is not implemented according to its design, failure is the likely result.

Figure 11.1 The Relationship Between Program Hypothesis, Program Design, and Desired Result

The important point here is that when a human service program fails to achieve its desired outcome (result) due to theory failure, the result is nevertheless still a valid test of the program. We can make inferences about the program because the test itself is valid. Consequently, we can learn from theory failure just as we learn from a successful program. However, when a program fails to achieve the desired outcome (result) due to program

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 3/8

failure, the program has not been properly tested. We cannot draw any inferences about the success or failure of a program when it is not implemented as planned.

Impact Program Evaluation and Hypothesis Testing In Chapter 6, we discussed program planning as a hypothesis-generating activity and impact program evaluation as a hypothesis-testing activity. A program hypothesis (1) makes explicit the assumptions underlying a human service program outcome (result); (2) establishes a framework to bring internal consistency to the implementation of a human service program; and (3) enables inputs, process, outputs (including quality outputs), and outcomes to be examined for internal consistency.

We offered a fairly detailed hypothesis in Chapter 6 related to the issue of domestic violence. We argued that women who are victims of domestic violence face a number of barriers to self-sufficiency (e.g., low self-esteem, social isolation, lack of financial resources, lack of basic education, lack of workplace skills). The suggestion was made that if we could identify the major barriers (hypothesis) andif we could successfully eliminate or reduce those barriers via a well-designed and well-implemented human service program, then we should see positive changes in the lives of the participants leading to an increase in self-sufficiency. Now let’s look again at the three tracks presented in Figure 11.1.

The first track, labeled “successful program,” depicts an ideal situation. Let’s assume that the hypothesis here is thatif a basic education and job training program is designed and implemented, then more women who are victims of domestic violence should be able to secure employment and become more self-sufficient. The program is implemented as designed (basic education and job training services are provided), and the desired outcome (result) is achieved (participants get jobs and become more self-sufficient). Given this finding, we conclude that the hypothesis is supported. But human service programs can also fail to achieve their desired outcomes (result) due to flaws in either the theory of the program or the implementation of the program.

The second track in Figure 11.1 is labeled “unsuccessful program (theory failure).” Theory failure describes a situation in which a human service program is implemented according to its design, but the anticipated outcome (result) is not achieved. Theory failure comes about because of a flaw in the hypothesis underlying the program: namely, that certain causal processes will lead directly to the desired outcome (result). The hypothesis is the same: If basic education and job training services are provided, then women who are victims of domestic violence will be able to secure employment and become more self-sufficient. The program is implemented according to the program design, but few participants get jobs and even fewer become more self-sufficient. Given this finding, we conclude that the hypothesis is not supported.

The third track in Figure 11.1 is labeled “unsuccessful program (program failure).” Program failure describes a situation where a program is not implemented according to its design. In this case, we can say nothing about the achievement of program outcome (result). Program outcome (result) may or may not be achieved, but neither success nor failure can be attributed to the program. For example, the hypothesis is the same: If basic education and job training services are provided, then women who are victims of domestic violence will be able to secure employment and become more self-sufficient. But let’s say that due to funding reductions, only basic education services are provided; no job training services are provided. In other words, the program is not implemented according to its design. In this case, the program hypothesis may or may not be correct. We will never know because the program hypothesis was not properly tested. Failure in this case is attributable not to flaws in the theory but to deficiencies in implementation of the program design.

Research Designs for Impact Program Evaluation The essence of impact program evaluation is comparison. The purpose of comparison is to determine, to the extent possible, what actually happens to clients as a result of participation in a human service program. Do program clients show improvement? How much improvement? And how do we know for sure that the

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 4/8

improvement is actually due to the program and not some other factor? These are difficult questions to pose and even more difficult to answer. Comparisons in impact program evaluation are usually made in one of two ways:

1. Two different groups are compared. 2. One group is compared to itself.

In the first instance, one of the two groups serves as the experimental group and the other serves as the comparison group. In the second instance, a single group serves as the experimental group, but also serves as its own comparison group. The purpose of creating a comparison group is to determine what is called the counterfactual. The counterfactual is evidence of what would have happened to participants if they were not in the program (McDavid et al., 2013). Through the use of the counterfactual, it is possible to estimate the impact of the human service program on clients.

A variety of impact evaluation designs are available for use depending on the particular situation, the resources available, and the expertise of human service agency administrators. These research designs vary in complexity, timeliness, cost, feasibility, and potential usefulness. Impact program evaluation can be a complicated and difficult undertaking that, to be successful, must be approached with realism, commitment, and considerable knowledge. Table 11.2 presents three types of impact program evaluation designs frequently utilized in human service programs: (1) the single-group pretest/posttest design, (2) the nonequivalent comparison group design, and (3) the randomized experimental design.

Some explanation of the symbols in Table 11.2 is in order. Each X represents a human service program provided to a defined client group. Each O refers to an observation, or measurement, of a defined client characteristic (attitude, behavior, status, knowledge, condition, etc.) that is intended to be changed by a program. Each R stands for random assignment. Random assignment is discussed later in the chapter. Temporal order in Table 11.2 is left to right.

Single-Group Pretest/Posttest Design

The first impact evaluation design to be discussed is the single-group pretest/posttest design. In this design, a single group serves as both the experimental group and its own comparison group. An initial measurement or observation (O1), called a pretest, takes place before clients begin the program. The measurement or observation is related to the client characteristic (attitude, behavior, status, knowledge, condition, etc.) the program is designed to change. The initial measurement or observation (O1) becomes the baseline that will be used later for comparison purposes. Next, clients participate in the program (X). After completion of the program, or the receipt of a full complement of services, a second measurement or observation (O2) called a posttest takes place. The pretest and posttest are exactly the same; they are simply administered at different times. Diagrammatically, the single-group pretest/posttest design looks like this:

O1 X O2

Let’s consider the example of a parenting skills training class to be taught to young mothers. From an impact evaluation standpoint we want to know this: Does participation in the program increase the parenting skills knowledge of the young mothers? Using the single-group pretest/posttest design, the participants take a written pretest (O1) designed to measure their existing knowledge of parenting skills. Then the young mothers participate in the program (X). When they complete the program, or receive a full complement of services, they then take the same written test, now called the posttest (O2). It should be noted that the pretest and posttest do

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 5/8

not have to be written tests; they can be any test, measurement, or observation, provided they are administered the same way before and after.

The impact program evaluation question can now be addressed: Did participation in the program increase the parenting skills knowledge of the young mothers? The answer is determined by comparing scores on the posttest (O2) with scores on the pretest (O1). The resulting change (impact), which hopefully is positive, is attributed to participation in the program. Diagrammatically, the process looks like this:

Change = O2 – O1

As an example, if a participant scored 75 points on the written posttest and 50 points on the written pretest, it can be concluded that there was an improvement of 25 points in her parenting skills knowledge. The counterfactual, the measure or estimate of what would have happened had she not participated in the program, is the difference between the scores (25 points). For confidentiality purposes, we would of course not discuss individual client scores. Instead, we would aggregate the client data. Then, we would compare the mean posttest scores (the arithmetic average designated by the symbol ) to the mean pretest scores. Diagrammatically, the process looks like this:

While the single-group pretest/posttest design provides for the creation of a counterfactual, we still cannot be certain that the observed change in the scores between the posttest and the pretest are the result of the program. The problem is that other potentially confounding factors may have affected the program or the clients. This is why we say that the observed change is attributable to the program. To be able to say with more assurance that the program caused the change, we must deal with what is called threats to internal validity. The Treasury Board of Canada (2010) identifies several major threats to internal validity: history, maturation, mortality, selection bias, diffusion or imitation of the program, testing, and instrumentation. Table 11.3 provides a definition of each of these threats.

Nonequivalent Comparison Group Design

The second impact program evaluation design to be discussed is the nonequivalent comparison group design, which begins to deal with the issue of threats to internal validity. In this design, the experimental group does not serve as its own comparison group. Instead two separate and distinct groups are created. The experimental group comprises individuals who will participate in the program. The second group, the comparison group, comprises individuals who are statistically similar to the clients participating in the program, but who will not actually

Change = ¯̄X̄ O2 − ¯̄X̄ O1

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 6/8

participate in the program. Statistically similar means that the characteristics of the individuals that constitute the comparison group are the same or similar to the characteristics of the individuals that constitute the experimental group on all relevant characteristics (e.g., age, ethnicity, gender). For example, the comparison group could comprise (a) people who are actually eligible to participate in the program but who cannot be served due to inadequate funding; (b) people who are eligible for the program but are unaware of its existence; or (c) people who are eligible for the program but are disqualified for some reason.

The comparison group usually receives some other type of program rather than no program at all. This practice avoids the ethical problem of withholding services from some people for the purposes of experimentation. With the nonequivalent comparison group design, two of the major threats to internal validity (selection bias and testing) are addressed. By attempting to make the individuals in the experimental group and the comparison group as statistically similar as possible, an attempt is made to control for selection bias. An attempt is made to control for the test effect (learning that can occur as a result of taking the pretest) in that both groups take the pretest, so its effect is already accounted for in both the experimental group and the comparison group. The impact of participation in the program is computed by subtracting the mean difference between the posttest and pretest scores for the experimental group compared with the difference between the posttest and pretest scores of the comparison group. Diagrammatically, the comparison looks like this:

Any difference between the scores of the experimental group and the comparison group, which hopefully is positive, is the measure of the program’s impact.

Randomized Experimental Design

The final impact program evaluation design to be discussed is the randomized experimental design. This design is the most valid of the three approaches discussed because it involves the random assignment of participants to the experimental group and the control group. Random assignment (designated by R in Table 11.2) is an important component of experimental impact program evaluation designs (Nathan, 2008). Random assignment is said to control for all the threats to internal validity except for the testing effect. Also, random assignment is more powerful in controlling for selection bias. Because of the strength of the randomized experimental design, any difference found to exist between the two groups on the client characteristics of interest (attitude, behavior, status, knowledge, condition, etc.) can be said with increased certainty to be attributable to the program and not to external factors.1

To understand how random assignment works, think of potential program participants standing in a line. The first person in line steps forward, and a coin is tossed in the air: Heads means the individual will be assigned to the experimental group, tails means the individual will be assigned to the comparison group. The coin comes up tails, and the individual is assigned to the comparison group. The second person in line steps forward. The coin is tossed and again it comes up tails. The second person is assigned to the comparison group. The third person in line steps forward. The coin comes up heads, and the individual is assigned to the experimental group. The key here is that assignment of one individual to either the experimental or comparison group does not in any way affect the probability of how another person will be assigned.

Once random assignment is accomplished, the process and analysis are the same as for the nonequivalent comparison group design. The randomized experimental design holds promise for producing data that most clearly demonstrate the real impact of a human service program. However, because of ethical issues (withholding participation in a program for experimental purposes) as well as legal requirements (it is illegal in many instances to withhold participation in a publicly funded human service program if a client meets eligibility requirements), human service administrators tend to rely on less rigorous designs when conducting impact program evaluation.

ExperimentalGroup = ¯̄X̄ O2 − ¯̄X̄ O1 minus Comparision Group = ¯̄X̄ O2 − ¯̄X̄ O3

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 7/8

Summary Conducting an impact program evaluation of a human service program is a difficult, but not impossible, undertaking. Impact program evaluation attempts to determine if a human service program actually works. In order to make this kind of a cause-and-effect statement, a counterfactual has to be created that identifies what would have happened to individuals had they not participated in the program. Additionally, the major threats to internal validity (history, maturation, mortality, selection bias, diffusion or imitation of the program, testing, and instrumentation) need to be controlled for. In this chapter, three impact program evaluation designs were introduced and discussed: single-group pretest/posttest, nonequivalent comparison group, and randomized experimental. The point was stressed that the strongest impact research design and the one that deals most strongly with threats to internal validity is the randomized experimental design.

Case Example The Winter Park Voluntary Pre-School Center was considering adopting a new curriculum. The center director, however, wanted to be sure that the new curriculum would be worth the cost and effort as well as the disruption in the daily routine that a change would create for the children. She decided to conduct an impact program evaluation. She randomly assigned 15 children to an experimental group and 15 children to a comparison group. Next, she administered a pretest to all 30 children designed to measure their existing knowledge of colors, numbers, and letters. She then exposed the children in the experimental group to the new curriculum, while continuing to use the existing curriculum with the children in the comparison group. At the end of 90 days, she administered a posttest to the same 30 children. When she compared the mean posttest/pretest scores of the children in the experimental group with the mean posttest/pretest scores of the children in the comparison group, the experimental group children scored 10 points higher. The executive director decided to implement the new curriculum for all children at the center.

Review Questions 1. Using the three factors (frequency, issues, and attribution of outcomes) from Table 11.1, how would you

define what the Winter Park Voluntary Pre-School Center director was attempting to accomplish? 2. If the experimental group children had scored the same as or worse than the comparison group, what

information would you need to determine if the program represented theory failure or program failure? 3. What figure represents the counterfactual in this example? 4. Which of the major threats to internal validity were controlled for in this impact program evaluation? 5. Do you think that a 10-point improvement in the children’s scores was worth the change to the new

curriculum?

Note 1. True experimental designs include random assignment of clients to ensure that no selection bias (e.g., screening for those clients most likely to benefit) influences the measurement of the program’s impact. Readers who would like more complete listings and critiques of experimental and quasi-experimental designs are referred to Gabor, Unrau, and Grinnel (1998) and Rossi, Lipsey, and Freeman (2004). Readers who would like more information on experimental designs with randomized assignment are referred to the Coalition for Evidence-Based Policy (2010).

References Coalition for Evidence-Based Policy. (2010). Checklist for reviewing a randomized controlled trial of a social

program or project, to assess whether it produced valid evidence. Retrieved from http://coalition4evidence.org/wp-content/uploads/uploads-dupes-safety/Checklist-For-Reviewing-a-RCT- Jan10.pdf

6/2/2018 Bookshelf Online: Designing and Managing Programs

https://online.vitalsource.com/#/books/9781483388328/cfi/6/54!/4/2/4/2@0:0 8/8

Gabor, P., Unrau, Y., & Grinnel, R. (1998). Evaluation for social workers. Boston, MA: Allyn & Bacon.

Magnabosco, J., & Manderscheid, R. (Eds.). (2011). Outcomes measurement in the human services (2nd ed.). Washington, DC: NASW Press.

McDavid, J., Huse, I., & Hawthorn, L. (2013). Program evaluation and performance measurement: An introduction to practice (2nd ed.). Thousand Oaks, CA: Sage.

Nathan, R. (2008). The role of random assignment in social policy research. Journal of Policy Analysis and Management, 27, 401–415.

Nielsen, S., & Ejler, N. (2008). Improving performance? Exploring the complementarities between evaluation and performance management. Evaluation, 14, 171–192.

Rossi, P., Lipsey, M., & Freeman, H. (2004). Evaluation: A systematic approach (7th ed.). Thousand Oaks, CA: Sage.

Treasury Board of Canada. (2010). Program evaluation methods: Measurement and attribution of program results (3rd ed.). Retrieved from https://www.tbs-sct.gc.ca/cee/pubs/meth/pem-mep-eng.pdf