SOCW wk 10 Discussion: Assessing Outcomes

profileSummerLove75
SOCW6070wk10resource4.pdf

Full Terms & Conditions of access and use can be found at https://www.tandfonline.com/action/journalInformation?journalCode=wasw21

Administration in Social Work

ISSN: 0364-3107 (Print) 1544-4376 (Online) Journal homepage: https://www.tandfonline.com/loi/wasw20

Moving from Outputs to Outcomes: A Review of the Evolution of Performance Measurement in the Human Service Nonprofit Sector

Kristen Lynch-Cerullo & Kate Cooney

To cite this article: Kristen Lynch-Cerullo & Kate Cooney (2011) Moving from Outputs to Outcomes: A Review of the Evolution of Performance Measurement in the Human Service Nonprofit Sector, Administration in Social Work, 35:4, 364-388, DOI: 10.1080/03643107.2011.599305

To link to this article: https://doi.org/10.1080/03643107.2011.599305

Published online: 26 Aug 2011.

Submit your article to this journal

Article views: 6781

View related articles

Citing articles: 39 View citing articles

Administration in Social Work, 35:364–388, 2011 Copyright © Taylor & Francis Group, LLC ISSN: 0364-3107 print/1544-4376 online DOI: 10.1080/03643107.2011.599305

Moving from Outputs to Outcomes: A Review of the Evolution of Performance Measurement

in the Human Service Nonprofit Sector

KRISTEN LYNCH-CERULLO Consultant, Boston, Massachusetts, USA

KATE COONEY School of Social Work, Boston University, Boston, Massachusetts, USA

Across the human service field, funders, executive directors, and program managers are faced with pressures to demonstrate effec- tiveness through measurable outcomes. Although performance measurement is often seen as an administrative burden imposed by funders to the detriment of direct service, it is increasingly accepted as crucial to achieving impact. Using a conceptual framework com- bining institutional theory and resource dependency theory, this article examines the field-level pressures facing human service orga- nizations and reviews the research on nonprofit-level responses to these pressures. After an examination of key innovations in social measurement, including the theory of change logic model, outcome standardization projects, and trends in calculating social value, as well as lessons learned from data-driven social innovation efforts, future directions in research and practice are proposed.

KEYWORDS evaluation, nonprofit accountability, performance management, outcome measurement, neo-institutional theory

INTRODUCTION

Today human service organizations (HSOs) face increasing pressure to demonstrate that their services add real social value by improving the lives of individuals and/or strengthening communities (Bliss, 2007; Cairns, Harris, Hutchison, & Tricker, 2005; Carman, 2010; Fisher, 2005; Olszak Management Consulting, 2003; Smith, 2010; Spilka, 2004; Weiss, 2004). Such accountability

Address correspondence to Kristen Lynch-Cerullo, 127 Chestnut St, Wakefield, MA 01880, USA. E-mail: [email protected]

364

Moving from Outputs to Outcomes 365

demands are not new to human services (Bliss, 2007; Zimmermann & Stevens, 2006); historically they have been addressed through periodic review by outside evaluators. However, the 1990s launched the “perfor- mance measurement era,” and with it instituted a nationwide, if not global, expectation that nonprofits develop the capacity to measure their own effec- tiveness and do so on an ongoing basis (Lampkin et al., 2006). This era is noteworthy for the emergence of: 1) process use versus findings use (Kramer & Pfitzer, 2007; Patton, 2004); 2) effectiveness versus efficiency perspectives (Martin & Kettner, 1996, 1997; but see Pallotta, 2009); and 3) demonstrated outcomes versus outputs (Spilka, 2004).

Evidence suggests that performance measurement, with its focus on demonstrating effectiveness, has become deeply embedded in how pol- icy makers and many funders and service providers think about programs designed to illicit changes in human beings. For example, in 2007 the federal government implemented the Program Assessment Rating Tool (PART), an outcome-focused survey that rates the effectiveness of government programs (U.S. Office of Management and Budget, 2007), and in 2009 established the “Social Innovation Fund,” a $200 million public-private partnership designed to support results-oriented nonprofits (Brest, 2010). At the state- level, Oregon law (ORS 182.515–525) now requires that by 2011, 75% of all funding (in criminal justice, youth authority, children and families, and addiction and mental health programs) be limited to proven, evidence-based practices. Lastly at the nonprofit sector-level, the Boston Foundation, the largest public charity in New England, announced in 2009 that it will begin giving larger grants with fewer restrictions to a smaller number of nonprofits based on demonstrated effectiveness (Ailworth, 2009).

This article reviews performance measurement among nonprofit HSOs according to a conceptual framework that uses neo-institutional theory to examine the diffusion, and perhaps institutionalization, of its practices. The framework focuses our review of the literature on performance measure- ment at multiple levels. At the field level, we examine the factors that encouraged performance measurement, as well as field leaders’ and inter- mediaries’ response to new measurement expectations. At the organizational level, we explore the depth and variability of current performance measure- ment practices. Lessons learned from research on measurement in practice and suggestions for future research are included.

Below, we provide a definition of performance measurement and then present a conceptual framework for examining the degree of its institutionalization in the United States.

PERFORMANCE MEASUREMENT DEFINED

Performance measurement, performance management, and evaluation are three distinct but related concepts (Mulvaney, Zwahr, & Baranowski, 2006). Performance measurement is one approach to assessing program

366 K. Lynch-Cerullo and K. Cooney

accountability (Bliss, 2007). It involves “an ongoing process of establish- ing performance objectives; transforming those objectives into measurable components; and collecting, analyzing, and reporting data on those mea- sures” (Mulvaney et al., 2006, p. 432). Process data (activities), output data (products and services), and outcome data (“changes in program participants stimulated by program activities”—Hendricks, Plantz, & Pritchard, 2008, p. 34) are all collected (Martin & Kettner, 1996; Mulvaney et al., 2006). By compar- ison, performance management involves using performance measurement data to make decisions through problem identification and strategic planning (Mulvaney et al., 2006).

Performance measurement cannot replace evaluation, but instead com- plements it, often as a precursor. While performance measurement looks at specific components of an organization, using programs as the primary unit of analysis, evaluation often looks at the effectiveness of the program or organiza- tion as a whole (Mulvaney et al., 2006). Evaluations are typically conducted by people outside the program, whereas performance measurement relies upon internal staff’s collection of data (Lampkin et al., 2006; Mulvaney et al., 2006). Unlike performance measurement, evaluation attempts to answer “why an initiative may or may not have been effective” (Mulvaney et al., 2006, p. 432), a question that performance measurement cannot answer. Indeed, performance measurement is limited to showing the changes that occur in people at specific points in time (i.e., before, during and after program participation; Gueron, 2005), while evaluation measures impacts, or the longstanding changes in people that are produced by a program that go “over and above what people would have accomplished on their own” (Gueron, 2005, p. 69). Whereas many evaluations are retrospective in nature, performance measurement allows for judgments “about the effectiveness of human service programs during imple- mentation, as well as after” (Martin & Kettner, 1996, p. 7). This is something that many researchers involved in evaluation have been advocating to do (Andrews, Stone Motes, Floyd, Crocker Flerx, & Fede, 2005; Mancini, Marek, Byrne, & Huebner, 2004).

For purposes of clarity, although performance measurement includes collection of activity and output data, this article and the movement we describe are focused primarily on outcome data collection, the core effec- tiveness piece of performance measurement. In addition, unless specifically noted, use of the term performance measurement in the article encompasses both performance measurement and management, consistent with its use in the literature.

THE INSTITUTIONALIZATION OF PERFORMANCE MEASUREMENT: A CONCEPTUAL FRAMEWORK

To examine the institutionalization of performance measurement in the United States, we start by proposing the use of neo-institutional theory

Moving from Outputs to Outcomes 367

to examine the shifts in broader institutional environment in which HSOs operate whereby, to maintain legitimacy, organizations are increasingly pres- sured to incorporate visible efforts to demonstrate effectiveness of their interventions. Neo-institutional theory (DiMaggio & Powell, 1983; Meyer & Rowan, 1977) posits that field-level institutionalization, by which we mean the extent to which this practice is an important signal of legitimacy, occurs through three primary mechanisms: coercive (through rules and regulations), normative (through professional networks), and mimetic (by imitation).

Combining neo-institutional theory with the insights from resource dependency theory (Oliver, 1991), which focuses on the active strategies organizations take in managing their interdependencies, allows us to con- centrate on organizational-level responses to these institutional pressures as we do in the second half of the paper. Christine Oliver (1991) hypothesizes that organizations have available to them a range of responses to institu- tional pressures, including acquiescence, compromise, avoidance, defiance, and manipulation (p. 160). Further, she advances a set of predictive factors (cause, constituents, content, control, and context) that serve to encourage one strategic response over another. For example, related to constituents, she hypothesizes that for organizations with multiple constituents exerting multiple and conflicting demands the strategic response to institutional pres- sure is more likely to be one of resistance than for organizations facing a more unified set of institutional pressures. Similarly, organizations with higher resource dependencies on institutional actors are more likely to acquiesce to institutional pressures than organizations with lower or more diversified resource dependencies (Oliver, 1991, p. 162).

In the next section, we utilize this framework to review the key factors leading to a rise in performance measurement and its institutionalization in the nonprofit human services sector.

SHIFTS IN THE INSTITUTIONAL ENVIRONMENT: FACTORS LEADING TO A RISE IN PERFORMANCE MEASUREMENT

The environment in which HSOs operate shifted dramatically over the past 20 years toward an intense focus on accountability, measurement, and the application of for-profit business practices to the nonprofit sector. The accountability pressures that emerged during the 1990s were the result of a confluence of factors including the commercialization of and increased competition within the nonprofit arena, (Lynn, 2002; Sasse & Trahan, 2006; Young & Salamon, 2002), the popularity of total quality management (TQM) approaches to management (Martin, 2000; Martin & Kettner, 1996), the increase in managed care penetration (Felty & Jones, 1998), greater pub- lic focus on philanthropic accountability (Behrens & Kelly, 2008; Hendricks et al., 2008), and legislative and regulatory changes enacted by the federal

368 K. Lynch-Cerullo and K. Cooney

government (Martin & Kettner, 1997). It was further propelled within human services by the emergence of evidence-based social work practice, as well as advances in computer technology that facilitate easier data collection (Lampkin et al., 2006; Rist, 2004). To understand the extent of this envi- ronmental shift, we review the Government Performance and Results Act (GPRA), a key source of institutional pressure to adopt performance mea- surement, and the trends that served to reinforce it and extend its reach into the nonprofit sector.

GPRA and the Nonprofit Sector

In 1992, based on a “what gets measured gets done” philosophy, the federal government passed several pieces of legislation calling for new per- formance measurement requirements. Of greatest significance, the GPRA, which was implemented in 1997, for the first time statutorily defined “out- put as a measure of ‘activity or effort’ and an outcome measure as the ‘results of the program’ compared to its intended purposes” (Martin, 2000, p. 32). Essentially, the GPRA created a “template” for performance reporting (Cozzens, 1997) by requiring quantitative versus qualitative information on “the extent to which the program services conformed to those that were planned, the definition of units of service, the cost of a unit of service, the settings in which the program was or was not effective or efficient, the char- acteristics of the clients who were positively affected by the program, and the extent to which the program met the needs of the community and targeted populations” (Kautz, Netting, Huber, Borders, & Davis, 1997, p. 368).

While performance measurement became increasingly important in the 1990s, previous efforts at measurement date back to the 1950s with two Hoover Commissions, followed by the National Center of Public Productivity, the Johnson administration’s planning-programming-budgeting systems (PPBS), Nixon’s management by objectives (MBO), and Carter’s zero-based budgeting (Kautz et al., 1997). However, while measurement of human service programs has a history, what is distinctive about the 1990 reforms is the degree to which nonprofits were affected by this change (Smith, 2010).

Despite being limited to federal agencies, the GPRA had far-reaching consequences for HSOs (Carman, 2009; Zimmermann & Stevens, 2006). First, programs receiving federal funding are expected to increasingly pro- vide similar types of outcome information (Carman, 2009; Martin & Kettner, 1996; Smith, 2010). Due to privatization trends rooted in the 1960s and strongly expedited during the 1980s, social services are now most often funded by government, but delivered by nonprofits (Felty & Jones, 1998; Smith, 2010). Indeed government, as opposed to private charity, now con- stitutes the single most important source of funding for nonprofits (Hetrick, 2004), with 52% of all nonprofit social service sector revenue coming from

Moving from Outputs to Outcomes 369

government (Independent Sector, 2009). Second, the federal government’s use of performance contracting, which often has an outcome-based focus, has grown in recent years, even within human service fields (Martin, 2000, p. 31; McBeath & Meezan, 2007; Smith, 2010).

Finally, organizations that in the past might have balanced their need to comply with government regulations by diversifying their resource depen- dencies increasingly face similar accountability demands from other types of funders across the sector in this era. Umbrella organizations such as the United Way and national organizations, including Big Brothers Big Sisters of America and Girl Scouts of America, as well as numerous foundations followed the lead of the GPRA and used many of its central concepts in establishing performance measurement expectations (Carman, 2009; Smith, 2010; Zimmermann & Stevens, 2006). Further, online websites such as Guidestar.org and Charity Navigator emerged in the late 1990s to assist indi- vidual donors in rating nonprofit organizations. Although this service was limited to efficiency versus efficacy measures, it illustrates the extension of this concern with measuring performance to the realm of the individual donor. Today even GuideStar is working to include effectiveness measures by rating nonprofits according to their use of field-specific best practices (Ailworth, 2009, pp. B7-B10).

While HSOs have faced accountability pressures in the past, the cur- rent level of institutionalization of performance measurement seems to be unique. Applying a neo-institutional theory analysis (DiMaggio & Powell, 1983; Meyer & Rowan, 1977) suggests that the coercive mechanism was initially dominant with enactment of the GPRA and the related regulations and requirements that extended to the nonprofit sector through contract- ing. However, the mimetic mechanism also appears to have played a key role as umbrella organizations and then foundations, to varying degrees, began replicating or imitating key GPRA principles. According to a neo- institutional/resource dependency lens, this scenario gives organizations less space to resist or defy these institutional pressures given their apparent uni- formity across constituents and the fact that these constituents all represent dependent resource relationships.

STRATEGIC FIELD-LEVEL RESPONSES TO INSTITUTIONAL PRESSURES: MANIPULATION, COMPROMISE AND ACQUIESCENCE

At the field level, the nonprofit sector responded in multiple ways to these institutional pressures. Initially, field-level actors in the sector moved to manipulate the institutional environment (Oliver, 1991) to preempt further accountability regulations. Indeed, in 2005 the Panel on the Nonprofit Sector, which represents both nonprofits and funders, recommended to the Senate Finance Committee that, “as a best practice, charitable organizations establish

370 K. Lynch-Cerullo and K. Cooney

procedures for measuring and evaluating their program accomplishments based on specific goals and objectives,” but according to their own stan- dards and specifications (Lampkin et al., 2006, p. 3). A second strategy involved a blend of acquiescence and compromise (Oliver, 1991) as field intermediaries, such as research institutes and umbrella funders, developed tools to assist nonprofits in meeting the new demands of their environment (Carman, 2010). As a result, there are now a number of tools available at the field level. Indeed, in 2009 the Foundation Center, seeking to central- ize these resources, compiled 150 different tools in Tools and Resources for Assessing Social Impact (TRASI), now available free of charge on their website (www.trasi.foundationcenter.org/browse_toolkit.php).

Below we briefly describe three key innovations useful to managers trying to better synthesize the myriad of approaches available. Logic mod- eling, standardization efforts, and methods for calculating social value are highlighted because of their focus on outcome measurement, as opposed to broad performance measurement tools such as dashboards and balanced scorecards, and because they illuminate the wide range, from basic to sophisticated, of the sector’s response.

As we view these tools, it is important to note that, although account- ability was the initial impetus for performance measurement and remains a key purpose, increasingly “program improvements, not external account- ability” is being stressed, certainly by the United Way and other large foundations, “as the primary reason for measuring outcomes” (Behrens & Kelly, 2008; Hendricks et al., 2008, p. 18). With this shift in mind, the tools below can be used to “learn about how to create change, rather than simply prove that change has occurred” (Behrens & Kelly, 2008, p. 43).

Logic Modeling

Logic models, used for years by evaluators, are now a key performance measurement tool (Mulvaney et al., 2006; Smith, 2010). Whether an organiza- tion collects data through the use of information technology or practitioners armed with paper and pencil, logic models are increasingly seen as either an essential “starting point” for sophisticated data collection systems (Wilson, 2009, p. 138), or the minimally acceptable measurement tool required of even the most cash-strapped nonprofits (Smith, 2010). There are different types of models, but funders generally now prefer outcome models (Martin, 2009), or one-page diagrams with a path from problem identified to inputs, outputs, and outcomes achieved (Bliss, 2007; Martin, 2009).

One type of outcome-based logic model increasingly in favor amongst funders is the theory of change logic model (Kellogg Foundation, 2004; Smith, 2010; Weiss, 2004). These models break down the assumptions inher- ent in all theories of change and answer why, based on research, a program’s key components are expected to achieve results (Carman, 2010; Learning

Moving from Outputs to Outcomes 371

for Sustainability, 2009; Weiss, 1995). For example, a civic education pro- gram designed to enhance personal responsibility in students (an outcome) assumes based on research that experiential learning and teacher training (two program components) are necessary to achieve the outcome. These two program components, which were originally assumptions articulated in developing the theory of change, are then operationalized to a point where they can be individually tracked in a logic model.

The benefits are considerable. The manager not only receives early feedback but also is better able to discern which aspects of the program are working and which are not, allowing for both adjustments in strategy and refinement of the theory of change. In addition to benefiting individual nonprofits, growth in field-level knowledge occurs when there is “diffusion, replication, critique, and modification” until proven theories of change in particular fields of practice emerge (Brest, 2010, p. 1). Further, funders look favorably upon theory of change logic models because they provide the ability to fund discreet program components (Weiss, 2004), perhaps those best aligned with their mission and interests.

Although a completed logic model is a powerful tool to possess, the process used to develop it has been found equally valuable (Hendricks et al., 2008; Patton, 2004). Sound logic models are achieved from intense discussions amongst key stakeholders that result in a high degree of clarity as to a program’s purpose and expected outcomes (Lampkin et al., 2006; Mulvaney et al., 2006; Poertner, Moore, & McDonald, 2008; Smith, 2010). To that end, a number of web-based tools are available to assist nonprof- its in developing sound outcome-based logic models, including theory of change logic models. Organizations with such tools include: the United Way, the Kellogg Foundation, the University of Wisconsin, the Harvard Family Research Project, and Learning for Sustainability.

Centralizing, Standardization, and Shared Measure Initiatives

In addition to logic modeling tools, key actors in the field have been developing standardized measurement tools. The purpose of standardiza- tion efforts generally is to help nonprofits reduce the time and cost of establishing outcomes and indicators, improve the overall quality of data collected by ensuring indicators had been carefully vetted by experts, assist in comparison of nonprofit results, and facilitate establishment of best practices (Kramer, Parkhurst, & Vaidyanathan, 2009; Lampkin et al., 2006; Snibbe, 2006). The desired impact of standardization is to “move beyond the current focus of one-off grants and capacity building for individual organizations” toward strengthening performance and collabo- ration within entire fields (Kramer et al., 2009, p. 3). Recently, Foundation Strategy Group Social Impact Advisors (FSG) categorized existing standard- ization tools conceptually: 1) shared measurement platforms, 2) comparative

372 K. Lynch-Cerullo and K. Cooney

performance systems, and 3) adaptive learning systems (Kramer et al., 2009, pp. 2–3).

SHARED MEASUREMENT

Although there are different approaches to shared measurement, all are specifically designed as an easy-to-access, low-cost resource for small non- profits. Although the United Way of America (UWA) has resisted the creation of a repository of standardized logic models, outcomes, and indicators (based on the belief that organic development increases the relevance to programs), they have created linkages to other such standardization efforts (Hendricks et al., 2008). For example, the UWA launched a website direc- tory (www.toolfind.org) that provides 46 highly screened tools created by field-specific experts that nonprofits can use to select evidence-based indica- tors within 11 education and youth development outcome areas (Retrieved March 15, 2010 from www.toolfind.org). In 2006, the Center for What Works and the Urban Institute went slightly further in its “shared measure- ment” approach. As opposed to creating a directory of existing tools, this initiative reviewed the array of outcomes and outcome indicators in use (culled from national nonprofit umbrella groups, national nonprofits with local affiliates, and accreditation agencies) to develop common, evidence- based outcomes and indicators for 14 different program areas arranged into “outcome sequence charts” that moved from outputs to mid-range then long- range outcomes (Lampkin et al., 2006, p. 6). These tools are now available online at no cost. The NeighborWorks America, a Congress-created non- profit umbrella organization in community development, implemented The Success Measures Data System in 2004 for its 230 affiliates (Retrieved March 10, 2010 from www.successmeasures,org) to provide, at minimal cost, not only 44 evidence-based indicators and 100 vetted tools for data collection (such as surveys, focus group questions etc), but also a data storage plat- form to track, tabulate, report, and manage data (Retrieved March 10 from www.successmeasures.org).

COMPARATIVE PERFORMANCE AND ADAPTIVE LEARNING SYSTEMS

Comparative performance and adaptive learning are two high-investment standardization approaches that, although early in development and not widely used, point the way to new trends in shared measurement. Comparative performance data systems allow nonprofits to track and report identical indicators, while adaptive learning systems facilitate engagement among “a large number of organizations working on different aspects of a single complex issue” (Kramer et al., 2009, p. 2). The former is focused on benchmarking and making organization-to-organization comparisons,

Moving from Outputs to Outcomes 373

while the latter is focused on collaboration, learning, and systemic impact via multiple organizations working on the same shared endeavor (Kramer et al., 2009).

Monetizing Social Value: Demonstrating a Return on Investment

Integrated cost initiatives, a term used by Tuan (2008, p. 1), are a series of approaches that are developing metrics that go beyond the demonstration of measurable outcomes, the general focus of this article, to monetize the social value, or benefits to stakeholders (i.e., employees, volunteers, recipients, donors. and the community), of those outcomes (Mook & Quarter, 2006). There are currently eight different integrated cost approaches being used within the social sector, and all attempt to demonstrate dollars saved (Tuan, 2008). For example, the Roberts Enterprise Development Fund (REDF) social return on investment (SROI) model, one of the first in the field, has demon- strated social value by monetizing the public savings that accrue due to their social purpose enterprise employees’ decreased reliance on government pro- grams, while the Robin Hood Foundation, which uses an SROI-like analysis with its grantees, demonstrates social value by monetizing the personal ben- efits (such as increased income) to clients (Brest, Harvey, & Low, 2009). Integrated cost approaches can produce powerful findings, including one study that showed “almost 4 to 1—for every dollar invested, $3.70 in social value was returned to the community” (Quarter & Richmond, 2001, p. 82).

Field leaders such as the Robin Hood Foundation, Acumen Fund, and the William and Flora Hewlett Foundation are beginning to use these approaches to not only monetize past and current impact, but predict future impact (Brest et al., 2009; Javits, 2008; Tuan, 2008). This work creates a “risk/return” profile that, similar to the private sector’s pre-investment anal- ysis, quantifies how much risk an investor, or in this case a donor, is willing to take in its investment and how much of a return per dollar they expect (Brest et al., 2009; Javits, 2008; Tuan, 2008).

With 10 years of experience, experts caution that integrated cost approaches “have not yet reached maturity” (Tuan, 2008, p. 6), lack basic infrastructure, and require further refinement (Gair, 2005; Javits, 2008; Tuan, 2008). Indeed, if applied to human services, this work would need to bet- ter address the reality that for certain populations, such as homeless youth, connection to social services is an important step toward reintegration in society, yet achieving this would result in a low SROI for those serving them (Gair, 2005). However, although integrated cost approaches are, at this stage of development, costly and flawed, they are noteworthy. Despite the GPRA’s focus on broadening measurement to programs, the vast majority of outcome metrics collected today focus on individual change. Integrated cost approaches go beyond examining impacts on the individual client to calculating the social impact of the program for the local community and

374 K. Lynch-Cerullo and K. Cooney

broader society; thus, this work holds the potential to shift perceptions of nonprofits from “users of resources” to “creators of value” (Mook & Quarter, 2006, p. 247). It is also calls attention to the fact that ultimate success for any nonprofit lies in showing both social and economic impact of work and directs managers to consider this in choosing outcomes and measures accordingly (Smith, 2010; Tuan, 2008; Wolk, Dholakia, & Kreitz, 2009).

As detailed above, there are now dozens of field-level actors investing in tools to support nonprofit performance measurement capabilities. Next we turn to what the literature tells us about the extent of organizational-level adoption of these practices, its challenges, and the lessons learned.

ORGANIZATIONAL RESPONSE TO INSTITUTIONAL PRESSURES

Neo-institutional theory suggests that organizations facing institutional pres- sures may respond by ceremonially enacting the practices required by the field to maintain legitimacy without allowing them to deeply penetrate the nature of the practice routines (Meyer & Rowan, 1977). A combined institutional/resource dependency perspective (Oliver, 1991) posits that an organization’s response to external institutional pressures depends, in part, on the way these environmental demands are refracted through the orga- nization’s resource dependencies. In this section we review the literature examining to what extent organizations are incorporating performance mea- surement into their service technologies and the lessons learned from those organizations moving to incorporate these practices deeply.

Adoption of Performance Measurement Practices

Although our review of the literature revealed very few recent empirical studies examining whether HSOs specifically are incorporating performance measurement into their service technologies and to what extent (but see Carman, 2007; Carman, 2009; 2010; Barman & MacIndoe, in press; Frumpkin, in press), progress at the organizational level has been characterized as slow (Carman, 2009, 2010; Carrillio, Packard, & Clapp, 2003; Spilka, 2004). Joanne G. Carman (2010, 2009, 2008, 2007) has done some of the only systematic survey research in this area to date, using samples of nonprofits in New York state (n = 178) and in Indiana (n = 189), as well as conducting in-depth interviews with nonprofit managers and funders.

Carman’s research suggests that, although many community-based HSOs appear to know “that they ‘should’ do evaluation and performance measurement” (2010, pp. 265–266), and many are trying to do so, often broad external accountability practices—such as producing reports for boards and funders, conducting program audits, experiencing site visits and program reviews, and using benchmarking toward strategic goals—are

Moving from Outputs to Outcomes 375

misconstrued for true evaluation and outcome measurement practices (Carman, 2007, p. 65). Specifically, Carman (2007) found that although 80%– 90% of the New York nonprofits surveyed report using the aforementioned practices, a much smaller number have a performance measurement system (41%), design program logic models (17%), collect comparison data (10%), or have data collection tools that extend beyond paper methods (3%). The Carman (2008) study of Indiana nonprofits resulted in very similar find- ings. In addition, although a non-representative sample not limited to HSOs, recent research of 284 nonprofits in North Carolina found that while 93% of nonprofits reported use of “external accountability” practices, “such as producing reports or site visits for external funders,” only 52% gathered information on best practices or benchmarks in the field, 34% used perfor- mance measurement systems, 23% used logic models, and 17% collected comparison data (Murphy & Mitchell, 2007, p. 3).

Carman interprets such findings to suggest that rigorous performance measurement practices have not been widely adopted by HSOs, at least in part due to the absence of standardized evaluation and performance mea- surement practices across the field, an unintended negative consequence of the proliferation of so many tools and approaches, and a lack of evaluation expertise on the part of nonprofit managers, most of whom she found do this work themselves (between 75% and 89%). The net result is managers who “do not understand or distinguish between reporting, monitoring, and management practices and evaluation. They tend to think about all of these activities together as part of the broader agenda and pressures to make community-based organizations more accountable” (Carman, 2007, p. 72). From a neo-institutional theory perspective, it also may suggest that a certain ceremonial enactment of practices exists (Meyer & Rowan, 1977), with man- agers engaged in activities to demonstrate a nod to accountability without ever truly endeavoring into practices related to core evaluation or outcome measurement.

With respect to resource dependency perspectives, Carman (2009) found that source of funding explained more of the variance in the degree of evaluation and performance measurement activity than did organizational characteristics such as age, budget size, and location in the service sector (11% vs. 6%), or licensing, certification, and accreditation status. However, her data suggest that the demands for evaluation and performance mea- surement are not manifested uniformly across funders. Nonprofit funding sources do significantly predict nonprofit engagement in specific types of evaluation activities, with federal funding and United Way funding both being significant predictors of whether a nonprofit is engaged in evaluation or performance measurement, while local and state government funding and foundation funding are not (Carman, 2009).

While this research provides important insight into the adoption of performance measurement, there is still much left unexplained about

376 K. Lynch-Cerullo and K. Cooney

organizational choices to incorporate these practices. As Carman (2007) points out, it appears that across organizations performance measurement is conceptualized and implemented in very different ways. Indeed, even within the United Way system, despite having a uniform, national approach (the kind of standardization Carman (2007) argues would benefit the field) within the 450 local United Ways engaged in outcome measurement, a wide variety of approaches exist (Hendricks et al., 2008). Below we provide an overview of the challenges to moving from output to outcome measurement systems and lessons learned from organizations that are successfully moving in this direction.

Challenges and Lessons Learned for Management

CHALLENGES

Adequacy of resources is often cited as the biggest obstacle to performance measurement. Indeed, a study of 138 South Carolina nonprofits (not lim- ited exclusively to HSOs like Carman) conducted by Zimmerman & Stevens (2006) found that nonprofits with large budgets are more likely to engage in performance measurement than those with small budgets. Specifically, about two-thirds of those with budgets of $300,000 or less reported engaging in performance measurement, compared to 96% of those with budgets of $1 million or more (Zimmermann & Stevens, 2006, p. 322).

Given that small nonprofits (frequently defined in the literature as those with budgets of less than $500,000) represent approximately 75% of all non- profits (Independent Sector, 2009), adequacy of resources may in fact be a significant factor in the slow diffusion of performance measurement tech- niques across the sector. Further, although funders have increased demands on organizations to provide outcome data, they have not provided the resources to do so. Carman (2009) finds the majority of NPOs use inter- nal funds to support evaluation/measurement activities, while one quarter assign no funds at all for these endeavors.

However, resources are not the only deterrent. Human service admin- istrators incorporating performance measurement into their practice tech- nologies are tasked with the challenge of choosing good outcome indicators and measuring them well. Client outcomes can be measured in terms of changes in client “affect, knowledge, behavior, environment, and status” (Poertner, 2000, p. 273). Ultimately funders desire outcome indicators within at least one of these categories that show “what happened to the program’s ‘customers’ that would not have happened without the program’s efforts” (Kautz et al., 1997, p. 370). Selecting such measurements in human services can be difficult. This is particularly true in fields, as is often the case, where policy does not dictate the specific outcomes to be achieved or where an evidence-based practice is not yet established (Poertner, 2009).

Moving from Outputs to Outcomes 377

First, many factors, such as changes in world or domestic economy, leg- islation, or judicial decisions, can influence the lives of program participants beyond actual program effects (Kautz et al., 1997). Further, human service programs often work in tandem with other programs, so attributing client outcomes to any one program can be difficult (Kautz et al., 1997), as well as accounting for people that simply improve on their own despite interven- tion (Gueron, 2005). Thus, as Kates, Marconi, and Mannle (2001) point out, providers must move beyond “what they control—that is their activities—to focus on what they merely influence—their results” (p. 147).

Second, some types of interventions are easier to develop outcome measures for than others. For example, prevention efforts are especially difficult to document (Kautz et al., 1997). In addition, social change can often occur very slowly and “a program that looks like a failure at year two or three may prove a raging success at year 20 or 30” (Snibbe, 2006, p. 42). Third, data collection, storage, and retrieval can require considerable staff time and resources, both of which may be diverted away from direct service (Poertner, 2000), a move historically criticized by both funders and the general public. Finally, human service programs are often funded by several sources, each of whose specific mission can result in different expectations with regard to the performance measures sought (Poertner, 2000; Snibbe, 2006).

For nonprofits that do collect performance data, an additional charge is to utilize the data to improve performance (Carman, 2007; Mulvaney et al., 2006). Zimmerman & Stevens (2006) found that 75% of nonprofits collecting performance data used it for decision making and/or to enhance manage- ment practices. However, other findings are not as positive. Drawing from the federal government’s experience, managers using performance measure- ment reported collecting an increased number of measures from 1997–2003, but not an increase in the usage of the data (Mulvaney et al., 2006, p. 438). Murphy and Mitchell (2007) found that although 82% of nonprofits reported that a purpose of evaluation and measurement is to make program changes, only 50% experienced program change as a benefit and 33% experienced positive changes in decision making and internal operations (p. 18–19). A review of United Ways, well versed in outcome measurement, found the use of outcome data specifically as lacking (Hendricks et al, 2008). The challenge of data usage is not limited to nonprofits, as the literature suggests that funders also lack the capacity to effectively use the data provided to them (Behrens & Kelly, 2008). Research shows that the end result is data collection that does not always reflect in increased mission-related accom- plishment (Christensen & Ebrahim, 2006) or direct benefits to consumers of services (Cairns et al., 2005).

The following broad reasons are suggested explanations for the inability to convert such measures into better management: 1) lack of staff/management expertise (Carman, 2010; Carrillio et al., 2003); 2) a man- agement focus on short-term activities versus long-term goals (Cozzens,

378 K. Lynch-Cerullo and K. Cooney

1997; Mulvaney et al., 2006); 3) disagreement on what measures to use and how to use them among stakeholders (Mulvaney et al., 2006; Snibbe, 2006); and 4) ideological opposition by human service professionals (Carrillio et al., 2003, p. 62; Lee, McMillen, Knudsen, & Woods, 2007; Snibbe, 2006). In addi- tion, a review of the UWA effort found that while there has been an intense focus on providing tools and trainings to collect data (what some call “front- end” work), there is a “dearth of back-end guidance” on how to use the data (Hendricks et al., 2008, p. 20). Our review of the literature suggests that this may be true of the field generally.

LESSONS LEARNED

Those practicing performance measurement have produced helpful insights that nonprofits grappling with these challenges can use in implementing and maintaining effective outcome measurement systems. First, the litera- ture suggests that organizational motivation, or in this case the underlying basis for establishing and maintaining a performance measurement system, may affect the ultimate success of these efforts. Specifically, it is suggested that performance measurement imposed from the outside is unlikely to be accepted and/or produce meaningful impacts to clients (Cairns et al., 2005), as is the collection of outcome data simply or primarily to appease funder demands (Carrillio et al., 2003; Fisher, 2005; Wolk et al., 2009). Thus, a “self-assessment” of the primary organizational motivation for con- ducting performance measurement may be an important step as the best systems appear to emanate from a desire to improve services versus a desire to appease funders. However, Zimmerman and Stevens (2006) found that although 87% of the nonprofits that used performance measurement were required to do so by a funder, 75% went on to use the data to improve services (Zimmermann & Stevens, 2006, p. 321). This limited bit of evidence suggests that even when funders impose performance measurement sys- tems, there remains the possibility of organizations institutionalizing these practices in meaningful ways. Indeed, in reviewing her own research in this area, one of Carman’s key recommendations is that funders should not only incentivize performance measurement but also performance manage- ment by “asking (and rewarding) for reports designed to demonstrate how they are using evaluation and performance data to improve service delivery” (Carman, 2007, p. 72).

Second, whether imposed externally or desired internally, experience has taught that initiating a quality performance measurement system often requires an overall organizational culture change (Carrilio et al., 2003; Fisher, 2005). Research shows that in an “environment where both leadership and employees are more concerned with avoiding blame than improving quality, progress toward the effective use of data for planning and quality man- agement is not likely” (Carrillio et al., 2003, p. 72). Further, research also

Moving from Outputs to Outcomes 379

suggests that even when “adequate intellectual and fiscal resources are avail- able, human service professionals may resist the use of evaluative data to guide services on ideological grounds” (Bliss, 2007; Carrillio et al., 2003, p. 62; Lee et al., 2007; but see Murphy & Mitchell, 2007).

Effective leadership is one of the most important factors in estab- lishing a culture conducive to measurement (Carrilio et al., 2003; Fisher, 2005; Hoole & Patterson, 2008; Mulvaney et al., 2006). To achieve this, experts advise use and/or knowledge of certain evaluative frameworks such as Patton’s utilization-focused evaluation and/or Fetterman’s empowerment evaluation, both of which facilitate development of an organizational culture in which measurement’s primary purpose is conceptualized as an opportu- nity for continual team learning and improvement (Snibbe, 2006). The goal is to create a learning organization typified by managers “willing to share their own mistakes, reward good ideas, encourage staff to hold each other respon- sible, and lend support to one another in their learning endeavors” (Hoole & Patterson, 2008, p. 96), and management and staff that can “think eval- uatively” or “weigh evidence, consider contradictions and inconsistencies, articulate values, and examine assumptions” (Patton, 2004, p. 4).

Central to both evaluative frameworks is a participatory process. Research shows that managers often error in limiting outcome and indica- tor development discussions to funders and/or administrative staff, despite evidence that practitioner involvement enhances both staff buy-in and the usefulness of the measures over time (Cairns et al., 2005). Research sug- gests that because practitioner and employee resistance tends to be centered around “how and if” outcome data will be used, regular usage of the data to improve services and make decisions can further promote staff buy-in (Fisher, 2005).

A key characteristic of such a “learning organization” is one in which “evaluation becomes part of the change effort” (Behrens & Kelly, 2008, p. 44). The Harlem Children’s Zone (HCZ), a nonprofit known for deeply integrating performance measurement into its practice approach, has taken the fairly uncommon approach of explicitly stating measurement as a fun- damental part of its theory of change (The Harlem Children’s Zone Project Model, 2009, p. 4). Although embedding measurement in an organization’s theory of change does not adequately address the complexities associated with an HSO’s data usage, it may be an important first step in engaging and encouraging staff to conceptualize measurement as an effort that extends beyond mere external accountability.

To promote data usage at the organization level, managers may con- sider use of existing software that is relatively inexpensive and has been found by experts to meet the needs of most nonprofits (Carman, 2008). Indeed, the majority of nonprofits today that have moved beyond paper methods appear to be collecting and storing data this way (Carman, 2008). Outcome data usage requires managers and staff to analyze the information

380 K. Lynch-Cerullo and K. Cooney

“to pinpoint where the program is having more and less success, interpret the implications, brainstorm possible ways to improve services, imple- ment trials, draw conclusions, and revise the program” (Hendricks et al., 2008, p. 20). The literature suggests that these steps are not necessarily apparent to nonprofits trying to engage in this work, but guidance (see Hatry, Cowan, & Hendricks, 2004; Morley & Lampkin, 2004) is increasingly available (Hendricks et al., 2008).

Third, programs have particular stages in their development (Brest, 2010; Emerson, 2009). Similarly, nonprofit measurement capacity also has stages of development that run along a broad continuum from very basic (i.e., transitioning from the collection of output data to outcome data) to highly sophisticated (i.e., use of comparative performance and adaptive learning shared measurement platforms). However, experts suggest that in applying measurement expectations and approaches, this great variance can be ignored (Carman, 2010; Emerson, 2009; Spilka, 2004). For exam- ple, Carman (2010) notes instances where nonprofits inappropriately set out to “‘prove causality’ or ‘estimate the counter-factual’ (i.e., prevention of a crime) for a funder,” both very sophisticated evaluation efforts, instead of establishing credible systems to collect simple baseline data and short-term outcomes well within their reach (p. 267).

The Edna McConnell Clark Foundation (EMCF) is tackling this issue by categorizing grantees and setting their measurement expectations accord- ing to the following stages: 1) apparent effectiveness (baseline data and outcome data are collected and compared); 2) demonstrated effective- ness (outcome data of participants is compared to a similar population of non-participants); and 3) proven effectiveness (impact on participants is sci- entifically confirmed through experimental research and randomized control groups) (Brest, 2010; Emerson, 2009, p. 2). Although EMCF funds nonprofits with moderate to sophisticated measurement capacity—indeed, according to EMCF most nonprofits today do not meet its most basic “apparent effective- ness standard” (Emerson, 2009)—its categorization of grantees calls attention to the broader problem of misapplication of measurement expectations, a common practice-level mistake (Carman, 2010; Emerson, 2009; Spilka, 2004).

The lesson for managers and funders is to be realistic in setting mea- surement expectations (Spilka, 2004) and to note that establishing effective systems takes significant time. For instance, according to the UWA, this work typically takes between two to four years (Hendricks et al., 2008). Realistic assessments, particularly for those new to outcome measurement, will likely lead to scaled back expectations, including initially establishing sound baseline data, important to later identification of meaningful targets (Hendricks et al., 2008), as well as the collection of intermediary versus long-term outcomes (Bliss, 2007). For example, a parenting education pro- gram (0–5) whose long-term outcome is school readiness might begin by measuring enhanced interactions between children and their parents and

Moving from Outputs to Outcomes 381

parental knowledge of positive discipline strategies. The goal for those new to outcome measurement is not necessarily to prove outcomes but to show continued progress in the efforts to capture this information (Hendricks et al., 2008).

In addition, “it is not uncommon to see intervention plans that are not strong or intense enough to produce hoped for outcomes” (Carman, 2010, p. 269). In assessing the applicability of outcome measurement to programs, managers need to consider the intensity of their program models to discern what kind of outcome collection is appropriate. For example, an organization operating 12-week workshops on self-defense and conflict resolution for girls in disadvantaged communities could utilize pre- and post- tests to determine a change in the girls’ knowledge and practice of self- defense and conflict resolution skills, but would have a more difficult time demonstrating a link between the program intervention and longer term outcomes, such as school attachment and academic performance, given the myriad of factors contributing to those outcomes.

Fourth, while addressing social issues requires dynamic programs and innovation, the use of existing field-specific research and knowledge can lend credibility to programs and save time and money. As opposed to developing new programs, many organizations, such as the Boys & Girls Clubs of America, use program models that have already been scientifically proven to work (Emerson, 2009; Snibbe, 2006). This has been made eas- ier by the growing number of organizations that compile field-specific best practices and make them available online (Snibbe, 2006). Examples include the What Works Clearinghouse (www.whatworks.ed.gov) for educational programs; the Campbell Collaboration (www.campbellcollaboration.org) for social welfare, education, and crime and justice programs; and the Colorado Center for the Study and Prevention and Violence (www.colorado. edu/cspv/blueprints) (Snibbe, 2006, p. 42).

For new programs and in instances where policy and best prac- tices have not yet been clearly established, the shared measurement tools described in this article can be utilized to develop research-based outcomes, indicators, and processes appropriate for the unique nature of the work done in the agency (Kramer et al., 2009; Lampkin et al., 2006; Wolk et al., 2009). Experts advise that the chief goal for managers is to use as much of the exist- ing research as possible, but to keep their outcomes and measures specific enough to reflect individual program goals (Wolk et al., 2009). In addi- tion, although there are generic tools to facilitate logic model development, often field-specific logic models are also available. For example, Bliss (2007) developed an outcome measurement system via a logic model for small substance abuse treatment programs, while Poertner (2008) provides field- specific outcome measurement advice for those in child welfare services.

To assist organizations in the process of adopting common indica- tor measurement tools, intermediaries and researchers from the “field” of

382 K. Lynch-Cerullo and K. Cooney

performance measurement itself have created some agnostic “best practices,” or guidelines, for managers. For example, nonprofit managers selecting measures for new programs are advised to seek a limited number of core indicators, those most central to a client’s success, that are: “specific (unique, unambiguous); observable (practical and cost-effective to collect, measur- able); understandable; relevant (measure important, significant dimensions to the client); time bound (cover a specific period of time); and valid (pro- vide reliable, accurate, unbiased data)” (Lampkin et al., 2006, p. 6). To achieve this and avoid the collection of useless data, experts advise non- profits to: 1) conduct an audit of any existing measures (Wolk et al., 2009); According to Poertner et al. (2008), 2) create a target number of measures; 3) keep only those measures that have actually been used to make decisions; 4) screen-out process measures from outcome measures; and 5) conduct a cost estimate for each measure added (p. 10).

Finally, as mentioned previously, Carman (2010) found that nonprofit managers, most of who have no formal evaluation training, often mistake accountability practices with true outcome measurement and evaluation. Carman (2010) urges the field to address this by strengthening evaluation offerings in graduate education and establishing “a set of simple, practi- cal evaluation and performance measurement expectations and standards” (p. 267). An easy, cost-effective strategy for managers put forth by Carman (2010) is to select at least one board member with knowledge of evaluation and measurement.

DISCUSSION AND CONCLUSION

The purpose of this article was to examine the rise of performance mea- surement in the institutional environment where nonprofit HSOs operate, explore the sector response, and share the lessons learned from the field. Although we initially set out to demonstrate the institutionalization of per- formance measurement, our assessment of the literature yields a portrait of a paradigm shift in progress. Demands from the federal government and United Way for measurable outcomes have altered the landscape in the nonprofit sector, producing a multitude of field-level tools to assist HSOs in capacity building. Yet despite institutional pressures and activity at the field level, the adoption of performance measurement practices at the organizational level appears varied and often superficial.

This may be explained in part through the use of our conceptual framework. By applying neo-institutional theory and resource dependency perspectives we see that, although the demands for outcome evaluation practices are an increasingly normative phenomenon, they appear to be unevenly distributed across funder types. Further exploration of the tepid organizational level adoption of core outcome measurement and evaluation

Moving from Outputs to Outcomes 383

practices may be assisted by Carman’s (2010) provocative assessment of the accountability movement itself via a theory of change logic model. According to Carman, the accountability movement has not achieved its desired result of improved efficacy because: 1) untrained HSO managers are overwhelmed by “unstandardized” performance measurement expectations and practices, an unintended consequence of the development of so many tools in the field; 2) an external orientation in reporting data that limits its usefulness to HSOs and thwarts organizational learning; and 3) an inadequate investment by funders that limits growth in HSO measurement capacity.

These are valid points and offer opportunity for field-level reflection and improvement. There are certainly numerous managers and practitioners providing effective services that, due to the factors cited by Carman (2010), are unsuccessful in capturing the results of their work. However, implied in this analysis and the literature generally is the notion that all programs are achieving results, and it is simply the lack of training and resources that prevent them from being captured. Using the logic model paradigm, this assumption should be challenged. In assessing the accountability movement in terms of sector-wide usage of core measurement practices, it is worth noting that, while improving service efficacy is the movement’s implicit goal, inherent in that goal has always been the notion that less efficacious programs would be “weeded out” (thereby clearing the way for greater resource delivery to effective interventions). It is plausible that a self-selection process is emerging in the field, such that HSOs with results-based programs are finding ways, even rudimentary ways, to invest in core measurement, while others that lack the intensity to produce outcomes do not or do so ceremonially.

From this vantage point, the accountability movement may be more or less successful, depending upon your viewpoint. Some may see this process as driving scarce resources to the most effective interventions, while oth- ers may see it as inadvertently driving out “less intensive programs” (i.e., a 12-week summer jobs program for disadvantaged youth) that make contribu- tions, albeit on a smaller scale than a neighborhood-based, cradle-to-college effort like HCZ, in a sweeping effort to eliminate truly ineffectual models. Indeed, although these less intensive models may not hold the potential in and of themselves to achieve impacts on long-term education or employ- ment trajectories, it is possible that engaging in a series of them over a lifetime (even unrelated programs), particularly in disadvantaged commu- nities, may make a difference. Future research could be useful in tracking whether or not such a selection process is emerging in the field and, if so, in clarifying its possible ramifications.

In keeping with Carman’s (2010) logic model application, future macro- level research might also benefit from framing improved service efficacy as the long-term outcome of performance measurement, and inserting “cre- ation of organizations that ‘think evaluatively”’ (Patton, 2004, p. 4) as the intermediary outcome. This research focus creates a more direct link, is

384 K. Lynch-Cerullo and K. Cooney

easier to measure and, we believe, holds the potential to yield findings of greater utility and immediacy to the field. A host of indicators could be generated, including the degree to which funders request not just data, but examples of how HSOs use the data to improve services, per Carman’s sug- gestion (2010). This lens could also further illuminate findings that suggest that, while field-level efforts did a good job of generating tools to capture outcome data, they have lagged in assisting HSOs with what to do with the data. The orientation may move a greater number of HSOs toward careful articulation of their theory of change, with consideration of such things as stage of development and program intensity in thoughtfully establishing tar- get outcomes and effective use of negative data to continually evolve their practice models. These are all things that research suggests the vast majority of HSOs are still not doing today.

In short, the longstanding (and perhaps at this point unhelpful) debate over performance measurement’s utility (between those that argue it is an unproven, funder-driven waste of scarce resources and a detriment to direct services and those that argue it is the only way to deliver effective ser- vices and achieve results) could be neutralized and instead the field could begin capitalizing on performance measurement’s untapped potential contri- bution to creating a greater number of organizations that “think evaluatively” (Patton, 2004, p. 4). It is hard to argue against the value of this. As Patton (2004) succinctly notes, “findings have a small window of relevance,” while “learning to think evaluatively can have ongoing impact” (p. 4).

In conclusion, while the research suggests that performance measure- ment has not deeply penetrated HSO practice uniformly, it is clear that for a growing number of funders and practitioners its central message of making a real, measurable difference is pervasive. As the performance measurement movement unfolds, social work administrators will have ample opportunities to work at the individual level in funder relationships and at the field level to co-construct meaningful approaches to measuring effectiveness. In addition, as the purpose of performance measurement continues to evolve from sim- ply “proving” to “improving,” (Behrens & Kelly, 2008), often in large part as a function of changes in organizational culture, social worker administrators, with their unique perspective and skills, have much to contribute to these discussions.

REFERENCES

Ailworth, E. (2009, September 16). Stressing results, charity retools grant-giving. Boston Globe, pp. A-1, A-7,

Andrews, A. B., Stone Motes, P., Floyd, A., Crocker Flerx, V., & Fede, L.-D. (2005). Building evaluation capacity in community-based organizations: Reflections of an empowerment evaluation team. Journal of Community Practice, 13(4), 85–104.

Moving from Outputs to Outcomes 385

Behrens, T. R., & Kelly, T. (2008). Paying the piper: Foundation evaluation capac- ity calls the tune. In J. G. Carman & K. A. Fredericks (Eds.), Nonprofits and evaluation: New directions for program evaluation. San Francisco, CA: Jossey-Bass.

Bliss, D. (2007). Implementing an outcomes measurement system in substance abuse treatment programs. Administration in Social Work, 31(4), 83–101.

Brest, P. (2010, Spring). The power of theories of change. Stanford Social Innovation Review. Retrieved from http://www.ssireview.org/site/printer/the_power_of_ theories_of_change/

Brest, P., Harvey, H., & Low, K. (2009, Winter). Calculated impact. Stanford Social Innovation Review. Retrieved from http://www.ssireview.org/site/ printer/calculated_impact/

Cairns, B., Harris, M., Hutchison, R., & Tricker, M. (2005). Improving perfor- mance? The adoption and implementation of quality systems in U.K. nonprofits. Nonprofit Management & Leadership, 16(2), 135–151.

Carman, J. G. (2007). Evaluation practice among community-based organizations: Research into the reality. American Journal of Evaluation, 28(1), 60–75.

Carman, J. G. (2008). Nonprofits, funders, and evaluation: Accountability in action. American Review of Public Administration, 39(4), 374–390.

Carman, J. G. (2009). Nonprofits, funders and evaluation: Accountability in action. The American Review of Public Administration, 39(4), 374–390.

Carman, J. G. (2010). The accountability movement: What’s wrong with this theory of change? Nonprofit and Voluntary Sector Quarterly, 39(2), 256–274.

Carrilio, T. E., Packard, T., & Clapp, J. D. (2003). Nothing in–nothing out: Barriers to the use of performance data in social service programs. Administration in Social Work, 27(4), 61–75.

Christensen, R., & Ebrahim, A. (2006). How does accountability affect mission? The case of a nonprofit serving immigrants and refugees. Nonprofit Management & Leadership, 17(2), 195–209.

Cozzens, S. (1997). The knowledge pool: Measurement challenges in evaluating fundamental research programs. Evaluation and Program Planning, 20(1), 77–89.

DiMaggio, P. J., & Powell, W., W. (1983). The iron cage revisited: Institutional isomor- phism and collective rationality in organizational fields. American Sociological Review, 48(2), 14–160.

Emerson, J. (2009, Winter). But does it work? How best to assess program per- formance. Stanford Social Innovation Review. Retrieved from http://www. ssireview.org/site/printer/but_does_it_work/

Felty, D., & Jones, M. (1998). Human services at risk. Social Service Review, 72(2), 155–171.

Fisher, E. A. (2005). Facing the challenges of outcomes measurement the role of transformational leadership. Administration in Social Work, 29(4), 35–49.

Gair, C. (2005). A report from the good ship SROI . Retrieved from http://www.redf. org/publications-sroi.htm#methodology

GPRA. (1993). Government Performance and Results Act of 1993, Public Law 103-62. Retrieved from http://www.usda.gov/news/gpra.htm.

386 K. Lynch-Cerullo and K. Cooney

Gueron, J. (2005). Throwing good money after bad: A common error mis- leads foundations and policymakers. Stanford Social Innovation Review, Fall, 69–71.

Hatry, H. P., Cowan, J., & Hendricks, M. (2004). Analyzing outcome information: Getting the most from data. Washington, DC: Urban Institute Press.

Hendricks, M., Plantz, M., & Pritchard, K. J. (2008). Measuring outcomes of United Way-funded programs: Expectations and reality. In J. G. Carman & K. A. Fredericks (Eds.), Nonprofits and evaluation: New directions for program evaluation (Vol. 119, pp. 13–35. San Francisco, CA: Jossey-Bass.

Hetrick, M. (2004). Performance of nonprofit human service organizations affiliated with the United Way. Dissertation abstracts international, A.: The humanities and social sciences, 65(6), 2, 356.

Hoole, E., & Patterson, T. E. (2008). Voices from the field: Evaluation as part of a learning culture. In J. G. Carman & K. A. Fredericks (Eds.), Nonprofits and evaluation. New Directions for Evaluation (Vol. 119, pp. 93–113).

Independent Sector. (2009, October 30). Facts and figures about charitable organi- zations. Retrieved from http://www.independentsector.org/programs/research/ Charitable_Fact_Sheet.pdf

Javits, C. (2008). REDF’s current approach to SROI . Retrieved from http://www.redf. org/download/sroi/REDFs-Current-Approach-to-SROI.pdf.

Kates, J., Marconi, K., & Mannle, T. (2001). Developing a performance management system for a federal public health program: the Ryan White Care Act titles I and II. Evaluation and Program Planning, 24, 145–155.

Kautz, J., Netting, E., Huber, R., Borders, K., & Davis, T. (1997). The Government Performance and Results Act of 1993: Implications for social work practice. Social Work, 42(4), 364–373.

Kellogg Foundation. (2004). Logic model development guide. Retrieved from http://www.wkkf.org/Pubs/Tools/Evaluation/Pub3669.pdf

Kramer, M., Parkhurst, M., & Vaidyanathan, L. (2009). Breakthroughs in shared mea- surement and social impact. FSG Social Impact Advisors. Retrieved from http:// www.fsg-impact.org/ideas/pdf/Breakthroughs_in_Shared_Measurement.pdf

Kramer, M., & Pfitzer, M. (2007, April). From insight to action: New direction in evaluation: Research synopsis for social sector funders. FSG Social Impact Advisors.

Lampkin, L., Kopczynski, M., Kerlin, J., Hatry, H., Natenshon, D., Saul, J., et al. (2006). Building a common outcome framework to measure nonprofit perfor- mance. Washington, DC: The Urban Institute and the Center for What Works. Retrieved from www.urban.org/publications/411404.html

Learning for Sustainability. (2009). Theory of change and logic models. Retrieved from http://learningforsustainability.net/evaluation/theoryofchange.php

Lee, B., McMillen, C., Knudsen, K., & Woods, C. (2007). Quality-directed activities and barriers to quality in social service organizations. Administration in Social Work, 31(2), 67–85.

Lynn, L. (2002). Social services and the state: The public appropriation of private charity. Social Service Review, March, 58–82.

Mancini, J., Marek, L., Byrne, R., & Huebner, A. (2004). Community-based program research: Context, program readiness, and evaluation usefulness. Journal of Community Practice, 12(1/2), 7–21.

Moving from Outputs to Outcomes 387

Martin, L. (2000). Performance contracting in the human services: An analysis of selected state practices. Administration in Social Work, 24(2), 29–43.

Martin, L. (2009). Program planning and management. In R. Patti (Ed.), The hand- book of human services management (pp. 339–350). Los Angeles, CA: Sage Publications.

Martin, L., & Kettner, P. (1996). Measuring the performance of human service programs. Thousand Oaks, CA: Sage Publications.

Martin, L., & Kettner, P. (1997). Performance measurement: The new accountability. Administration in Social Work, 21(1), 17–29.

McBeath, B., & Meezan, W. (2007). Nonprofit adaptation to performance-based, managed care contracting in Michigan’s foster care system. Administration in Social Work, 30(2), 39–70.

Meyer, J. W., & Rowan, B. (1977). Institutionalized organizations: Formal structure as myth and ceremony. American Journal of Sociology, 83(2), 340–363.

Mook, L., & Quarter, J. (2006). Accounting for the social economy: The socioeco- nomic impact statement. Annals of Public and Cooperative Economics, 77(2), 247–269.

Morley, E., & Lampkin, L. (2004). Using outcome information: Making data pay off . Washington, DC: The Urban Institute.

Mulvaney, R., Zwahr, M., & Baranowski, L. (2006). The trend toward accountability: What does it mean for HR managers? Human Resource Management Review, 16 , 431–442.

Murphy, D. M., & Mitchell, R. (2007). Building evaluation capacity in North Carolina’s nonprofit sector: A survey report. Raleigh, NC: Institute for Nonprofits, NC State University.

Oliver, C. (1991). Strategic responses to institutional processes. The Academy of Management Review, 16(1), 145–179.

Olszak Management Consulting. (2003). Community-based program information: The Roberts Enterprise Development Fund. A study of social enterprise train- ing and support models. Retrieved from http://www.olszak.com/nonprofit consulting/nonprofitresources/studyofsetrainingandsupportmodels.aspx

Pallotta, D. (2009). Charity navigator fixes its compass. Free the nonprofits. Retrieved from http://blogs.harvardbusiness.org/pallotta/2009/12/charity-navigator-fixes- its-compass.html

Patton, M. (2004). On evaluation use: Evaluative thinking and process use. The Evaluation Exchange, Winter, 4–5.

Poertner, J. (2000). Managing for service outcomes: The critical role of information. In R. Patti (Ed.), The handbook of social welfare management (pp. 267–281). Los Angeles, CA: Sage Publications.

Poertner, J. (2009). Managing for service outcomes: The critical role of information. In The handbook of social welfare management (pp. 165–181). Los Angeles, CA: Sage Publications.

Poertner, J., Moore, T., & McDonald, T. (2008). Managing for outcomes: The selec- tions of sets of outcome measures. Administration in Social Work, 32(4), 5–22.

Quarter, J., & Richmond, B. J. (2001). Accounting for social value in nonprofits and for-profits. Nonprofit Management & Leadership, 12(1), 75–85.

388 K. Lynch-Cerullo and K. Cooney

REDF. (2001). SROI methodology. Retrieved from http://www.redf.org Rist, R. (2004). On evaluation utilization: From studies to streams. The Evaluation

Exchange, IX(4), 1–8. Sasse, C., & Trahan, R. (2006). Rethinking the new corporate philanthropy. Business

Horizons, 50, 29–38. Smith, S. R. (2010). Nonprofits and public administration: Reconciling perfor-

mance management and citizen engagement. The American Review of Public Administration, 40(2), 129–152.

Snibbe, A. (2006). Drowning in data. Stanford Social Innovation Review, Fall, 38–45. Spilka, G. (2004). On community-based evaluation: Two trends. The Evaluation

Exchange, Winter, 6–7. The Harlem Children’s Zone Project Model. (2009). Executive summary. Retrieved

from http://www.tc.edu/i/a/document/9857_ExecutiveSummaryHCZ09.pdf Tuan, M. T. (2008). Measuring and/or estimating social value created: Insights

into eight integrated cost approaches. Retrieved from http://www.gatesfoun dation.org/learning/documents/wwl-report-measuring-estimating-social-value- creation.pdf.

Weiss, C. (1995). Nothing as good as a good theory: Exploring theory-based eval- uation for comprehensive community initiatives for children and families. In J. Connell, A. Kubisch, L. Schorr & C. Weiss (Eds.), New Approaches to Evaluation (pp. 65–73). Washington, DC: Aspen Institute.

Weiss, C. (2004). On theory-based evaluation: Winning friends and influencing people. The Evaluation Exchange, Winter, 2–3.

Wilson, S. (2009). Proactively managing for outcomes in statutory child protection- The development of a management model. Administration in Social Work, 33(2), 136–150.

Wolk, A., Dholakia, A., & Kreitz, K. (2009). Building a performance mea- surement system: Using data to accelerate social impact. Retrieved from www.rootcause.org/building-a-performance-measurement-system

Young, D., & Salamon, L. (2002). Commercialization, social ventures, and for-profit competition. In L. Salamon (Ed.), The state of nonprofit America (pp. 423–446). Washington, DC: Brookings Institute Press.

Zimmermann, J., & Stevens, B. (2006). The use of performance measurement in South Carolina nonprofits. Nonprofit Management & Leadership, 16(3), 315–327.