PADM704 Human Resources and the Legal and Regulatory Context of Public Administration 2

profileSweetness668
PublicAdministrationReview-2015-Gerrish-TheImpactofPerformanceManagementonPerformanceinPublicOrganizations.pdf

Research Synthesis

Ed Gerrish is assistant professor of public

administration at the University of South

Dakota. His research focuses on public sec-

tor performance, performance management,

and state and local taxation.

E-mail: [email protected]

48 Public Administration Review • January | February 2016

Public Administration Review,

Vol. 76, Iss. 1, pp. 48–66. © 2015 by

The American Society for Public Administration.

DOI: 10.1111/puar.12433.

Michael McGuire, Editor

Ed Gerrish University of South Dakota

Abstract: Performance-based management is pervasive in public organizations; countless governments have imple- mented performance management systems with the hope that they will improve organizational eff ectiveness. However, there has been little comprehensive review of their impact. Th is article conducts a meta-analysis on the impact of performance management on performance in public organizations. It contributes to the current literature in three ways. First, it examines the eff ect of the “average” performance management system. Second, it examines the infl u- ence of management: whether benefi cial performance management practices moderate the average eff ect. Th ird, it examines the eff ect of “time” on performance management. Using 2,188 eff ects from 49 studies, the analysis fi nds that performance management has a small average eff ect. However, the eff ect is substantially larger when indicators of best practices in high-quality studies are included, indicating that management practices have an important impact on the eff ectiveness of performance management systems. Evidence for the eff ect of time is mixed.

Practitioner Points • Th e act of measuring performance may not improve performance, but managing performance might. • Emphasize the use of benchmarking over time to provide a valid comparison and replicate success. • Performance management is present in a wide variety of policy areas. Ideas and best practices can be gleaned

from many experiences.

Collecting data from original studies that evaluate per- formance management, this meta-analysis combines data on performance with dummy variables represent- ing benefi cial performance management practices and indicators of study quality. In total, 2,188 eff ects were gathered from 49 original studies.

Th is analysis explores three related concepts. First, it examines the impact of PM on performance by com- bining all studies to estimate the eff ect of the average PM system, termed the mean eff ect size. Second, this analysis uses moderating variables to explore whether some indicators of benefi cial practices in PM (such as benchmarking and bottom-up implementation) infl uence the mean eff ect size, a test of the infl uence of management practices on performance.

Finally, this analysis explores the eff ect of “time” on the eff ectiveness of performance management systems. Time is operationalized in two ways. A “second-gen- eration” performance management system is defi ned as a system that has been in the same organization for at least two years and has been substantially changed (typically in response to perceived or actual failures). Th is defi nition is used in meta-regressions. Next, the eff ect of time is explored by examining the mean

Th e Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis

Considering how common performance management (PM) systems have become in public organizations, from policing to social

services, one might expect to fi nd a consensus that performance management systems are generally suc- cessful.1 Instead, one fi nds arguments that perfor- mance system values are misguided (Radin 2006), poorly applied (Frederickson 2003; Frederickson and Frederickson 2006; Radin 1998), or used for political ends (Lavertu and Moynihan 2012), as well as evi- dence that they do not substantially improve public performance (Gerrish 2014; Heckman, Heinrich, and Smith 1997; Hvidman and Andersen 2014; Rosenfeld, Fornango, and Baumer 2005) and that they induce behaviors that increase measured perfor- mance while adversely impacting actual performance (Courty and Marschke 2004, 2008; Heinrich and Marschke 2010).

Th ere are a number of important questions about PM, but the most fundamental is whether performance systems are associated with improved performance in public organizations. If there is little evidence that PM improves performance, then it seems senseless to con- sider performance management’s trade-off s with demo- cratic values (Radin 2006) or unintended consequences.

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 49

hypothesized how “management matters” to performance (Ingraham, Joyce, and Donahue 2003; Moynihan 2005) using policy case studies (Forsythe 2001) and self-reported performance surveys from manag- ers, which examine the eff ect of management generally (Moynihan 2005) and performance management specifi cally (Cavalluzzo and Ittner 2003; Julnes and Holzer 2001; Melkers and Willoughby 2005). In particular, they have been instrumental in generating and testing hypotheses. However, surveys that use self-reported performance have (at least) two drawbacks that limit their general applicability. First, they are subject to common method variance bias (Lindell and Whitney 2001); for example, when respondents are asked about the performance system and organizational eff ectiveness in the same instrument, common method variance bias tends to result in stronger correlations than multimethod instruments. Second, there may be a positive response bias if respondents are the performance offi cers tasked with both implementing the performance system and report- ing on perceived eff ectiveness for the survey instrument.

Nonetheless, performance management surveys have found a few consistent results. Support from managers for performance management is associated with both adoption and implementa- tion of performance management (Cavalluzzo and Ittner 2003). Use of performance information is both directly and indirectly related to the perception of performance management eff ective- ness (Yang and Hsieh 2007). Additionally, training and preparation for performance management implementation is associated with greater perceived eff ectiveness (Cavalluzzo and Ittner 2003; Julnes and Holzer 2001; Kroll and Moynihan 2015). Mission-orientation activities such as the establishment and reevaluation of mission goals are correlated with the implementation of performance manage- ment, but evidence for mission orientation’s impact on perceived eff ectiveness is lacking (Berman and Wang 2000; Wang and Berman 2001). Finally, voluntary performance management adoption may lead to “buy-in” and greater performance improvements (Julnes and Holzer 2001).

Th e second trend has been that public policy researchers have begun evaluating performance management systems within their respective fi elds, providing evidence on the impact of performance manage-

ment. However, these studies are isolated within policy subfi elds. Examples include policing using CompStat-like programs both in the United States and abroad (Chilvers and Weatherburn 2004; Jang, Hoover, and Joo 2010; Mazerolle, Rombouts, and McBroom 2007; Rosenfeld, Fornango, and Baumer 2005), waiting times in the National Health Service in England (Besley, Bevan, and Burchardi 2009; Propper et al. 2008,

2010), education accountability systems (Dee and Jacob 2011; Dee and Wyckoff 2013; Hanushek and Raymond 2005; Hvidman and Andersen 2014), child support enforcement (Gerrish 2014; Huang and Edwards 2009), and job training (Barnow 2000; Courty and Marschke 2008; Heckman, Heinrich, and Smith 2002; Heinrich 2002; Heinrich and Lynn 2001). Studies on job training and educa- tional accountability systems have been published more frequently than others and have also been linked to the incentives literature in economics (Baker, Jensen, and Murphy 1988; Holmstrom and Milgrom 1991).

impact of performance management using the data year, accumulat- ing the empirical evidence on performance management over time.

Th is analysis makes a signifi cant contribution to the literature by quantitatively examining the current state of performance manage- ment research. It combines studies from diverse fi elds and tests important theories about the impact of performance management, leveraging a large and sometimes contradictory body of existing empirical evidence. Th ese results have implications for a wide range of policy areas.

Th e following section discusses the recent foundations of perfor- mance-based management. It identifi es some theories tested by surveys about the moderating eff ect of managers on the relation- ship between performance management and performance. Next, it discusses the meta-analysis research method used, including the literature search process, coding of the original studies, and estima- tion of meta-regressions. After discussing results, this article off ers some suggestions for advancing the empirical research of perfor- mance management.

Managing for Performance: The Literature Public organizations have been managing for performance since at least the early 1990s (Williams 2003), although many key ideas started gaining traction in the 1970s (Moynihan 2008). Most scholars peg the modern incarnation of performance manage- ment to the late 1980s and early 1990s as part of the New Public Management (Hood 1995). Performance-based management in this era caught the attention of politicians of all stripes with a few key publications (Ammons 1995; Osborne and Gaebler 1992; Osborne and Plastrik 1997; Wholey and Hatry 1992), leading to the National Performance Review in the United States (Gore 1993). Performance-based management eff orts have been criticized as being fundamentally misguided because they supplanted demo- cratic values with technocratic ones (Radin 2006). Experiences with the Government Performance and Results Act (GPRA) at the federal level suggested that organizations may lack the capac- ity to implement sweeping performance reforms (Frederickson and Frederickson 2006; Kimm 1995; Mihm 1995), that the GPRA had a one-size-fi ts-all problem (Long and Franklin 2004; Radin 1998, 2000), and that perfor- mance measurement might be inappropriate, for example, within the U.S. Department of Health and Human Services, where programs fi nd it diffi cult to measure performance on rare diseases (Frederickson and Frederickson 2006). Despite these cautions, governments at every level have bet “the future of govern- ance on the use of performance information” (Moynihan 2008, 5). Th is continued during the George W. Bush administration under the GPRA’s successor, the Program Assessment Rating Tool (PART), and was modernized under President Barack Obama. Numerous surveys report that local governments, especially cities, use performance measurement widely, although less fre- quently for management (Melkers and Willoughby 2005; Wang and Berman 2001).

Th ere have been two parallel trends in performance management research during the last two decades. First, management scholars have

Public policy researchers have begun evaluating performance management systems within their respective fi elds, provid-

ing evidence on the impact of performance management.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

50 Public Administration Review • January | February 2016

Scholar search for “performance management” nets 306,000 results. Th erefore, it is important to establish a clear framework for includ- ing studies before beginning a systematic search. Studies that are acceptable meet all of the following criteria.

Study criteria. Th e research question for this synthesis, “What impact do performance management systems have on performance in public organizations?,” suggests four criteria for identifying a study as acceptable.

Th e fi rst criterion is an acceptable outcome variable of interest—typi- cally the dependent variable in a study that employs regression analy- sis. A broad defi nition of performance is employed here, although one that excludes self-reported measures of performance. As noted

earlier, self-reported performance is likely to be positively biased, for two reasons: managers in charge of performance systems may be more inclined to report success, and surveys that ask about performance systems and their perfor- mance may suff er from common method bias. Including survey responses of “customers” would be acceptable, but no studies examined employed such a survey. Examples of perfor-

mance used in this analysis include established child support orders, student test scores, future earnings, and reported crimes.

Th e second criterion is that the study must evaluate a performance management system. Th ere is no single acceptable defi nition of performance management systems, but there are some important elements. Moynihan defi nes performance management as “a system that generates performance information through strategic plan- ning and performance measurement routines and that connects this information to decision venues, where, ideally, the informa- tion infl uences a range of possible decisions” (2008, 5). Using Moynihan’s defi nition, along with additional literature in this area (Behn 2003; Hatry 2006; Kloot and Martin 2000; Melkers and Willoughby 2005; Moynihan and Pandey 2010; Wholey and Hatry 1992; Yang and Hsieh 2007), this analysis defi nes key elements of performance management systems as follows:

1. Setting performance goals or creating performance measures through fi at, negotiations, or models

2. Using incentives to achieve performance goals, including monetary rewards

3. Collecting performance information for use in strategic planning

4. Providing evidence that performance information is used in organizational decision making

5. Benchmarking current performance to previous perfor- mance or performance of other entities, inside and outside the organization; similarly, grading, categorizing, or recog- nizing performance from benchmarking

6. Linking agency, departmental, or organizational budgets or autonomy to achievement of performance goals

7. Publishing performance targets and results for managers, staff , stakeholders, and the public

To be included in this meta-analysis, two or more of the features described here ought to be evident in the original study. In some

Th is analysis leverages the fi ndings from the management litera- ture and data from policy research to address the three important questions about performance management described earlier. If the answer to all three questions points to a lack of association between performance management and performance, then it seems unneces- sary to consider value trade-off s or to use resources when the focus should be on alternatives to performance management, such as developing a public service motivation or ethic among managers (Perry and Wise 1990; Rainey 1982). Th ese questions are amenable to meta-analytic techniques.

Meta-Analysis: Data and Methods Meta-analysis, or analysis of analyses, combines quantitative fi nd- ings from a number of diff erent studies into a single study. It is more common in fi elds such as medicine, where much of the research comes from rand- omized trials of the same treatment. In policy analysis and management, meta-analysis is less common, for a few reasons. First, research is constantly shifting, meaning that we may not expect results to be replicated in new contexts or using diff erent methods. However, it is pos- sible to tease out context- and method-specifi c results using independent variables in meta-regressions. Second, policy analysis and management do not have a strong culture of study replication, a culture inherited from other social sciences. To a large extent, however, we discount the amount of parallel research that occurs in the social sciences. Not all studies make a methodo- logical or theoretical contribution on a particular subject and might therefore be omitted in a standard literature review.

Perhaps the largest advantage of quantitative meta-analytic tech- niques is that they allow analysts to accumulate fi ndings from the literature in a way that accounts for sample sizes (effi ciency) and strengths of the original research. Much like in the original studies, it is diffi cult to convey the eff ect of X on Y by examining individual observations, so eff ects are summarized using parameter estimates in a regression model.

Meta-analysis has its own terminology. Th e original analyses are called studies (or original studies), a term that encompasses manu- scripts, reports, books, and other publication outlets. Every study must have at least one statistical association between performance management (the X variable) and performance (Y). Each association is an eff ect. Every eff ect within every study is coded using the set of rules established later. Th e goal in meta-analysis is to estimate the size (direction and magnitude) of the average eff ect, the mean eff ect size. Th is mean eff ect size can be examined both unconditionally or conditioned on independent variables of interest.

Th e term “meta-analysis” describes the data collection process but, over time, has also encompassed statistical properties and a suite of tools (Borenstein et al. 2011; Card 2012; Ringquist 2013; Wolf 1986). Th e following sections provide more detail on data collec- tion, variable coding, and empirical techniques.

Data Collection Data collection in meta-analysis involves searching through the relevant literature to fi nd studies that are acceptable. A Google

Policy analysis and management do not have a strong culture

of study replication, a culture inherited from other social

sciences.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 51

working paper directories,8 and organizational websites, including those of government research bodies,9 nongovernmental researchers, and think tanks.10

Two additional processes ensure that as many acceptable studies as possible are found. Th e fi rst is an ancestry search, using both the references of acceptable studies as well as studies that have cited the original study (using Google Scholar). Th e second is to contact all of the authors who authored a study coded as acceptable, request- ing any other studies on the same subject. Both strategies yield additional studies that not found in the original search—contacting authors results in two additional dissertations that did not appear in searches of ProQuest’s Dissertations and Th eses Global database.

Literature searches result in just under 25,000 total hits (article titles that match the search terms) ending May 10, 2014. In total, 49 acceptable studies are included in this analysis and listed in the references. Th ey contain 2,188 total eff ects, the unit of analysis.11 Figure 1 presents a fl owchart of the literature search process. Studies cover a fairly wide swath of policy research but are dominated by studies in education—19 studies in education, compared with 10 in policing, 9 in job training (all from the JTPA/Workforce Investment Act), 6 in public health, and 5 in other areas, including

cases, acceptability is evident from the body of work on the perfor- mance management system rather than a single work (e.g., the Job Training Partnership Act [JTPA]). In other areas, it is necessary to fi nd additional background information (Dee and Wyckoff 2013; Fryer 2011, 2013). Th ese criteria typically exclude pay for perfor- mance and performance contracting studies.

Still, determining what is or is not a performance management system is as much art as science. For example, the National Institute for Excellence in Teaching (NIET) developed a teacher evalua- tion program, TAP, that has been evaluated by NIET researchers as well as independent evaluators. Th e program uses four elements of success: multiple career paths, ongoing applied professional growth, instructionally focused accountability, and performance- based compensation. While it shares two elements of perfor- mance management systems, there also appears to be an absence of organization-level use of teacher performance information for strategic management. Moreover, goals are set almost completely at the individual level. TAP also has a strong focus on professional development.

However, by the rules established here, TAP and some TAP-like teacher evaluation programs demand inclusion in this analysis. However, studies of teacher evaluation and incentives programs are diff erent enough from the other studies to conduct a robustness check excluding them.2

Th e third criterion for inclusion in this meta-analysis is that the original study data either must be completely composed of “public” organizations or must have a separate eff ect for public organiza- tions. Th e “publicness” of an organization is its own fi eld of inquiry (Bozeman and Bretschneider 1994).3 Th e criterion used here is that individuals or organizations must ultimately respond to an elected authority and must not have an explicit profi t motivation.4

Th e last criterion is that the study must have suffi cient information to transform the reported results into an eff ect that can be compared on an equal basis to other studies. Generally speaking, this only requires that the original study conduct a statistical analysis with a hypothesis test. Th is includes t, z, c 2, F, signifi cance stars, p-values, or any other fi gures signifying that a statistical test was performed against a null hypothesis.

Literature search process. The literature search fi nds studies that meet the four inclusion criteria. Because of the overwhelming number of hits for any search for “performance management,” the search is limited to public policy areas and a few specifi c performance systems such as CompStat and GPRA/PART. These terms are numerous, including “crime,” “policing,” “prisons,” “welfare,” “food assistance,” and “child support enforcement,” among others. A complete list of the exact policy/program-related search terms can be found in the notes.5 The foregoing terms are combined with performance management-related search terms, and the following fi ve exact phrases are used: “performance system,” “performance management,” “performance measurement,” “performance standard,” and “performance information.” This results in 100 total search permutations (20 policy search terms by fi ve performance-related terms). Each permutation is then searched in academic search engines,6 online conference proceedings,7

Table 1 Effects by Study

Authors Year Effects Authors Year Effects

Barnow 2000 4 Hvidman & Andersen 2013 3

Besley, Bevan, & Bur- chardi

2009 60 Jang 2008 17

Chilvers & Weatherburn 2004 8 Jang, Hoover, & Joo 2010 8

Chilvers & Weatherburn 2001 8 Lockwood & Porcelli 2013 38

Courty & Marschke 2008 12 Marsh et al. 2011 100

Cragg 1997 8 Mazerolle, McBroom, Rombouts

2011 7

Daley & Kim 2010 10 Mazerolle, Rombouts, McBroom

2007 14

Dee & Jacob 2009 231 Mazerolle, Rombouts, McBroom

2006 27

Dee & Jacob 2011 337 Nielsen 2013 23

Dee & Wyckoff 2013 43 Poister, Pasha, & Edwards 2013 2

Dickinson et al. 1988 19 Propper et al. 2010 51

Fryer 2013 78 Propper et al. 2008 45

Fryer 2011 124 Rosenfeld, Fornango, & Baumer

2005 2

Garicano & Heaton 2010 6 Schacter & Thum 2005 10

Gerrish 2014 56 Schacter et al. 2004 13

Glazerman & Seifullah 2012 257 Schochet & Burghardt 2008 4

Glazerman & Seifullah 2010 54 Schochet & Fortson 2014 105

Glazerman, McKie, & Carey

2009 76 Springer, Ballou, & Peng 2008 99

Hanushek & Raymond 2005 18 Springer et al. 2012 21

Heckman, Heinrich, & Smith

2002 32 Taylor & Tyler 2012 35

Heinrich 2002 10 Thibodeau et al. 2007 30

Heinrich & Lynn 2001 18 Thibodeau 2003 22

Horton 2010 15 Walker, Damanpour, Devece

2011 5

Huang & Edwards 2009 8 Wilborn et al. 2010 3 Hudson 2010 12 Total 2,188

Notes: Year indicates year published. Studies are listed in alphabetical order by fi rst author.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

52 Public Administration Review • January | February 2016

Notes: The fl owchart depicts how 49 studies with 2,188 effects were distilled from 24,737 hits through the literature search process. Studies were excluded for reasons enumerated in the text.

Figure 1 Flowchart of Study Selection and Inclusion

child support and general local government (e.g., English local gov- ernments). Twenty-eight of the 49 studies come from peer-reviewed journals; the others come from sources such as government reports, doctoral dissertations, and working papers. Table 1 lists the authors of each study used, the year it was published, and the number of eff ects coded within the study. Table 1 reveals that 949 eff ects, 43.4 percent of all eff ects, are coded from just four studies, one of which is the working paper of the journal article by Dee and Jacob (2011).12 Th ese four studies are removed in a robustness check to examine the sensitivity of results, fi nding the results to be robust to the removal of these studies.

Variable Coding After identifying an acceptable study, features about the study and each eff ect are coded. Except for the eff ect size, all of the other characteristics listed here are dummy variables.13 Many have only two categories, represented by a single dummy variable (for exam- ple, peer reviewed or not). In other cases, there are more than two categories; three categories are represented by two dummy variables, each refl ecting the diff erence between that category and the base.

Discussion of variables is grouped into fi ve sections. Th e fi rst section discusses the calculation of the dependent variable, the eff ect. Th e second section discusses coding a second-generation performance management system. Th e third section discusses coding performance management best practices. Th e fourth section discusses coding indi- cators of the quality of the original study. Th e last section is a discus- sion of other characteristics of the performance management systems, such as dummy variables for the policy area. Descriptive statistics of these variables, by group, are presented in table 2.

Calculating the effect size. Effects, both the dependent variable and unit of analysis, are calculated using a straightforward method but require some introduction for those versed in statistics but not meta-analysis. Because parameter estimates measure associations from different samples, it is necessary to convert parameter estimates from the original studies into a standardized measure. As a technique, meta-analysis combines such disparate measures of performance into a single variable, that is, an effect. The use of random-effects meta-regression explicitly assumes that these effects come from different but related measures of performance, adjusting

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 53

r is transformed to Fisher’s Z using the formula Z = 0.5 ln[(1 + r)/ (1 − r)] as a second step. Z has a known variance that does not depend on its value, Var(Z) = 1/(n − 3) and is unbounded. Th e diff erences in value between r and Z do not meaningfully impact the interpretation of coeffi cients; for Z values less than |0.40|, the diff erence between Z and r is less than 0.02. Th e diff erence in values approaches zero as either statistic approaches zero. More than 96 percent of the eff ects in this analysis have Z values less than |0.40|.

Second-generation performance management systems. One of the contributions of this article is to examine the impact of time on the performance of performance management. Time, in this analysis, is operationalized in two ways. The fi rst analyzes the mean effect size of performance management by year (of data), visually. This method is explained in greater detail in the Results section. The second way to operationalize time is to examine how second-generation performance management systems perform relative to their fi rst- generation counterparts. A second-generation system is distinguished here by two criteria. First, a performance system must be in place for at least two years. Second, extensive changes must have been made to the structure of the performance system, such as changing how the performance measures are defi ned or adding additional performance measures. This excludes the adoption of a performance system in a new setting. For example, a number of police departments outside of New York City modeled their systems on CompStat. They are not coded as second-generation systems.

Only a small number of second-generation systems have been examined by researchers, and almost none explicitly. For example, Barnow and Smith wrote that “[t]he JTPA performance standards system evolved considerably over its life from 1982 to 2000” and described a number of changes, including increasing the number of performance measures from 4 to 17 core measures (2004, 253). Similarly, the Child Support Performance and Incentives Act of 1998 changed child support enforcement performance measures from a single measure to fi ve performance measures, among other changes. In total, about 12 percent of eff ects used in this analysis are from second-generation systems.

Th e counterfactual for most second-generation systems is typi- cally fi rst-generation systems. Subsequently, studies do not identify the eff ect of a second-generation system from no system; reported results refl ect the impact of moving from a fi rst- to a second- generation system. To illustrate, imagine performance management in three stages: (A) an organization with no performance system, (B) an organization with a new performance system, and (C) an organization with a second-generation system. Ideally, research would identify the diff erence between C and A. However, all studies included in this meta-analysis compare C to B. Because meta-anal- ysis draws from many studies, this analysis is in the unique position to examine both the eff ect of moving from A to B and B to C.

Performance management best practices. The third set of variables indicates the quality of the performance management system as recommended by the performance management literature (Hatry 2006; Yang and Hsieh 2007) and is an attempt to peek into the black box of organizational performance. While these variables are referred to as “best practices,” they are, in fact, some practices that tend to be commonly recommended by both the theoretical and the “how-to”

(widening) confi dence intervals to account for different study settings. The method used here relies on the distributions of Pearson’s r and Fisher’s Z. Effects using these distributions are most common in social sciences and are called r-based effects.

Pearson’s r is fi rst calculated from the original studies. Pearson’s r takes values from –1 to +1, where |1| represents a perfect linear relationship and zero indicates no relationship. For example, the equation for Pearson’s r using the t-statistic from ordinary least squares is r = t2/(t2 + df ), where t is the t-statistic and df is degrees of freedom. Because distributions of z, t, F, c 2, and others are related, similar calculations can convert other associations into r using the test statistic and the degrees of freedom (Ringquist 2013).

Conservative estimates of r are used when original studies only report p-values or signifi cance stars. For instance, a reported p-value < .10 is converted to the test statistic value associated with p = .10 given the degrees of freedom. Results without signifi cance stars are converted to an eff ect size of zero. Both of these assump- tions bias eff ect sizes toward zero.

Despite being easy to calculate, there are two known limitations to r. Th e fi rst is that r is heteroskedastic; its variance depends on its value. Second, it is bounded by –1/1, that is, censored. As a result,

Ta ble 2 Descriptive Statistics

Expected Sign

Full Data Set Without TAP

N % N %

Second-generation PM system (+) 260 11.9% 182 12.3% Best practices Benchmarking is … (base is multiple forms)

Limited (–) 204 9.3% 124 8.4% Absent (–) 1,067 48.8% 478 32.4%

Top-down adoption (–) 1,967 89.9% 1,300 88.1% Output measure (+) 222 10.1% 196 13.3% Study quality Research design is … (base is randomized trial)

Quasi-experimental research design, pre-post with comparison group (+/–) 677 30.9% 284 19.2%

Weaker (+/–) 1,426 65.2% 1,171 79.3% Endogeneity strategy is … (base is randomized or regression discontinuity)

DiD, FE, matching, etc. (+/–) 1,381 63.1% 1,072 72.6% Absent (+/–) 289 13.2% 279 18.9%

Self-selection (+) 287 13.1% 207 14.0% Not peer-reviewed (+/–) 1,282 58.6% 701 47.5% Other characteristics Study domain is non-U.S. (+/–) 292 13.3% 292 19.8% Study domain is state

and local Policy arena is. . . (base is education) (+/–) 933 42.6% 244 16.5% Job training (+/–) 212 9.7% 212 14.4% Crime/policing (+/–) 112 5.1% 112 7.6% Public health (+/–) 211 9.6% 211 14.3% Child support (+/–) 64 2.9% 64 4.3% Other (+/–) 41 1.9% 41 2.8%

Total observations 2,188 1,476

Notes: Column N represents the number of observations from the sample coded 1. All variables except the dependent variable are dummy variables. Without TAP indicates that effects from evaluations of the TAP teacher evaluation program and other teacher evaluation programs are not included in the sample. This represents 712 effects, or about one-third of the sample.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

54 Public Administration Review • January | February 2016

the performance system. Legislation or executive action is coded as a top-down approach, as is the implementation of performance management by the researcher. For example, the adoption of the No Child Left Behind Act examined by Dee and Jacob (2011) would be a top-down approach, even though the act was modeled after state experience. Th e base category is implementation by management (bottom-up).

It is also important to consider how performance is being measured. Th e program evaluation logic model divides these into fi ve categories: inputs, activities, outputs, outcomes, and impacts. In included stud- ies, performance is typically measured using outputs, outcomes, or a measure of effi ciency (Julnes and Holzer 2001). In this research, one dummy variable is used to represent output performance measures. Th e base case represents outcomes or impacts. Effi ciency measures, an outcome or output over an input, is classifi ed by its numerator. For example, child support collected and distributed (outcome) divided by administrative expenditures (input) is classifi ed as an outcome. Less than 4 percent of all eff ects are measures of effi ciency. Th e expected sign on the dummy variable for use of output performance measures is positive, as outputs are more directly under the control of organiza- tions than either outcomes or impacts. For example, CompStat-type programs likely have a stronger eff ect on broken windows arrests (output) than reported broken windows crimes (outcome).

Study quality. The next set of dummy variables code for the quality of the original studies. These variables include strength of the research design, fi xes for endogeneity, and self-selection into a performance management system. They are also entered in reverse so that high- quality studies are the base case and lower-quality effects are coded with a 1. Study quality often varies within studies; one model may contain fi xed effects, while another does not. Aside from self-selection, which ought to have a positive sign, there are no ex ante expectations on the signs of these variables; higher-quality studies may or may not be associated with a stronger relationship.

Th e fi rst two variables code for the strength of the research design. Th e base case is a randomized fi eld experiment. Th e fi rst dummy variable indicates that the design contains observations both before and after implementation of the performance management system, and it includes a nonequivalent comparison group (in tables, abbre- viated as pre-post w/comparison). All other research designs are coded in the second variable indicating a weaker design. While there are a large number of such designs, and some recover causal estimates more reliably than others, only one dummy variable is used repre- senting the average eff ect of all such designs.

Similarly, the next group of dummy vari- ables explains how original studies handle threats to the exogeneity of the treatment. Because the choice of adopting a perfor- mance system is possibly endogenous, many studies directly address threats to unbiased parameter estimates. In short, an exogenous treatment is one that is not theoretically cor- related with an unobserved error term and would recover unbiased results—this is a key

reason experimental designs are considered the “gold standard” of causal inference. Even if an experimental design is not used, there

literature. The selected set consists of variables that can be consistently identifi ed and coded, making it shorter than a complete list of best practices. For example, the use of performance leadership teams (and regular performance meetings) as suggested by the PerformanceStat movement (Behn 2014; Smith and Bratton 2001) cannot be included because an omission of leadership teams does not necessarily mean that they are not present. Additionally, other concepts that are commonly marked as important practices to the success of performance management cannot be included because of a lack of consistent information. These include goal orientation, resources (Julnes and Holzer 2001), implementation training, (Julnes and Holzer 2001; Kroll and Moynihan 2015), and support/opposition of client groups, the public, or the media (Moynihan and Pandey 2010). As a result, this list should be considered an incomplete list of plausibly codable good practices that are benefi cial to the practice of performance management. From a statistical perspective, they may also be incomplete proxies for other benefi cial practices.

Variables in this section include evidence of performance bench- marking, an indicator of “bottom-up” versus “top-down” adoption of performance management, and the use of an outcome or impact performance measure. Th ese variables are entered into the model in reverse. In other words, a dummy variable for output measures takes a value of 1, while outcome measures take a value of zero. Th is results in an intercept that captures many ideal features of a perfor- mance management system.

Th e fi rst two variables in this set indicate whether there is no evidence of performance benchmarking or whether benchmark- ing is limited in scope. Th e base case includes at least two forms of benchmarking, benchmarking to past performance and to other units within the organization or organizations in another jurisdic- tion. Limited benchmarking means that organizations or individuals only compare their performance to past performance. Both dummy variables have a negative expected sign, and performance systems that utilize multiple forms of benchmarking ought to have stronger associations with performance compared with those that do not. Benchmarking is a common practice supported by a number of instructional texts. Hatry (2006) devoted a chapter to identifying appropriate benchmarks.

Also included is a variable that defi nes whether the performance management system appears to be adopted through a “bottom- up” approach versus a “top-down” approach. As Julnes and Holzer note, “if we take an internal policy requiring the organization to have performance measures as a proxy for voluntary participation, we can speculate that it will have a strong eff ect on adoption. An internal policy may represent ‘buy in,’ or more of a commitment to make performance measurement work” (2001, 696). Similarly, if top-level managers and line staff are involved in the adoption and implementation of performance management, it is more likely that it will be implemented for primarily instrumental rather than symbolic reasons, an important theoretical distinction made by Moynihan (2008). Th is variable is coded by reading the introduction or background of the original studies, which typically contain information about the adoption of

If top-level managers and line staff are involved in the adop- tion and implementation of

performance management, it is more likely that it will be imple- mented for primarily instrumen- tal rather than symbolic reasons.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 55

r-based eff ect, Fisher’s Z, on the independent variables described earlier. Each observation is weighted by its standard error so that greater emphasis is placed on studies with more observations.

Th e strength of meta-regression is that it accounts for both the effi ciency (from sample size) and magnitude of eff ects, whereas a count of positive versus negative results cannot and may therefore be misleading. Th e trade-off , however, is that meta-regression must also meet more stringent assumptions of parametric models. For example, because there are multiple eff ects within each study, these eff ects are not independent. Th ere are two options to deal with this problem. Th e fi rst is to collapse eff ects within studies into a single eff ect, creating a mean eff ect size within each study. However, because independent variables vary within studies (e.g., fi xed eff ects in one model but not another), this method would discard eff ect- level variation. Th e second option, employed here, is to use cluster- robust variance estimators, recognizing the structure of the error dependence. Th is strategy preserves variation in the independent variables and bases inferential statistics on the number of studies (49) rather than the number of eff ects, sacrifi cing precision for con- sistency.16 Another consideration for meta-analysis generally (and meta-regression specifi cally) is whether eff ects come from a fi xed- eff ects framework or a random-eff ects framework. While this shares the nomenclature of panel data, the concept has more in common with hierarchical linear models. Th e standard assumption for meta- analyses in social sciences is random eff ects (Ringquist 2013). A Q-test of fi xed versus random eff ects confi rms that random eff ects is the appropriate framework, meaning that eff ects come from a distribution of eff ects that are diff erent by more than sampling error alone (p < .001).17

Publication bias. Publication bias has the potential to bias the estimated mean effect size upward or downward (although typically not toward zero). Publication bias arises during data collection but impacts the estimation of the mean effect size. It is similar to truncation in that it causes an unknown fraction of the sample to be unobserved. Publication bias is caused by two related phenomena. The fi rst is journals rejecting studies that fi nd either no signifi cant result or a result contrary to expectations (positive publication bias). The second is researchers not submitting their studies for publication because the results are not statistically signifi cant or run counter to the expectations of the literature (the fi le drawer problem). These sources of publication bias are often confl ated (rejected studies may end up in a fi le drawer), but the difference is meaningful for meta-analysis. The meta-analytic researcher can fi nd rejected studies and include them in the analysis. However, the researcher cannot include results that were never written up. A 2014 review of publication bias in the journal Science found strong publication bias in the social sciences. Specifi cally, among a sample of peer-reviewed National Science Foundation proposals, the researchers found that “strong results are 40 percentage points more likely to be published than null results, and 60 percentage points more likely to be written up” (Franco, Malhotra, and Simonovits 2014, 1502). Both problems are more likely to occur in small-N studies, and both lead to biased estimates in meta-analysis.

Th ere are a few methods to detect publication bias among the peer- reviewed subset of studies. Th e fi rst is visual using the confunnel

are numerous methods to treat endogeneity, the threat to exog- enous treatment. Th e base category includes randomized treatment or regression discontinuity design (RDD). Although RDD is a quasi-experimental design, it is notable for its ability to recover local causal eff ects (Imbens and Lemieux 2008). Th e fi rst dummy variable includes the use of fi xed eff ects, diff erence-in-diff erences, matching, and instrumental variables. Th ese strategies work better in some contexts than others, but rather than subjectively assessing each attempt, these strategies are grouped into one category. Th e second dummy variable indicates that there does not appear to be a strategy for addressing possible endogeneity.

Regardless of the study’s ability to address threats, a dummy is included for whether organizations or individuals self-selected into the performance management system. For example, in one analysis, schools choose whether to participate in a teacher evaluation pro- gram (Daley and Kim 2010). Th is would create a clear self-selection problem that researchers may or may not have been able to address. Th is occurs less often than one might think in performance manage- ment. A fi nal dummy variable represents whether the original study was not published in a peer-reviewed journal.

Other characteristics. The fi nal set of dummy variables helps explain the variation in the mean effect size, but including this set of variables means that the intercept no longer represents an ideal performance management system. This includes the study context (foreign or domestic, local or national) and the area of public policy.

Th e fi rst dummy variable in this set indicates whether the data from the original study are located outside the United States. Th e second variable indicates that study data are contained entirely within a state or local government, with the base category indicating that the study is either at the interstate or federal level.14 Th e anticipated sign on this dummy variable is ambiguous. Some authors sug- gest that lower levels of government may not have the capacity to monitor performance (Radin 1998), although it might also be true that lower levels of government are more capable of understanding (therefore managing) the drivers of organizational performance.

Five dummy variables are also included that defi ne the broad area of public policy from which the study is derived. Th e base case is edu- cation. Dummy variables represent job training, crime and policing, public health, child support enforcement, and “other.” Two of the studies in the other category examine local governments in England, and the fi nal examines public transportation. Note that this is the policy area of the original study, not the publication outlet or the training of the researchers.

All of the dummy variables described in the previous sections, their group, the base case, their expected sign in meta-regressions, and descriptive statistics are presented in table 2.15 Descriptive statistics are presented for the full data set and a subset that excludes teacher evaluation programs. Th e next section describes the procedure for estimating the eff ect of the independent variables on the eff ect of performance management on performance—meta-regression.

Meta-Regression Meta-regression is the primary tool used in this meta-analysis. It employs a framework similar to weighted least squares to regress the

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

56 Public Administration Review • January | February 2016

study generates a large number of the positive statistical outliers in the lower-right quadrant of the top panel (Dee and Jacob 2011), and the working paper by the same author creates outliers in the bottom panel, but otherwise there appears to be a large number of eff ects in the lower-middle and lower-left quadrants in both published and unpublished papers, indicating (at least visually) that publication bias is not a signifi cant problem.

A second method to test for publication bias is a linear parametric test with eff ect sizes regressed on eff ect standard errors (Egger’s test). Th e intuition of this test is that an eff ect’s standard error should not be correlated with its size. If, for example, the slope estimate is negative and signifi cant, this would indicate that, on average, studies with a larger standard errors (smaller N) also have larger eff ects, an indication of publication bias. A modifi ed form of Egger’s test is conducted, taking into account the nonindependence of the observations. Results from this test indicate that there does not appear to be an important amount of publication bias among the 28 peer-reviewed studies (t = 1.39). Th is fi nding is not wholly surprising given the relative disagreement among scholars about performance management—null results in this area are as impor- tant and informative as positive ones. Finding no need to correct for publication bias, meta-regression results are presented in the next section.18

Results Th is section begins with an examination of the unconditional mean eff ect size using a forest plot—a visual representation of the mean eff ect size by study. Next, it discusses the results of meta-regressions that examine the impact of performance management on perfor- mance in public organizations. Th e last part of the section discusses two robustness checks. One excludes teacher evaluation programs, while the other excludes three studies that contain the largest num- ber of total eff ects.

Figure 3 contains a forest plot of all 49 studies. Studies are sorted by one of six policy areas, then the year of publication, with earlier studies appearing near the top of each grouping. Each study has a mean eff ect size indicated by a black dot, and the 95 percent confi dence interval is indicated by a black horizontal line. Th e right column of text reports the relative study weight—how much the grand mean relies on each individual study (and group). Th e diamond at the bottom of the fi gure estimates the unconditional mean eff ect size of all studies at 0.031 with 95 percent confi dence intervals of 0.021 to 0.042. Th e 95 percent prediction interval is –0.02 to 0.09. Th e vertical dotted line helps examine each study relative to the global mean. Th e magnitude of the grand mean eff ect size is discussed in greater detail below when it is compared with other meta-regression results.19 Each policy area also has its own diamond, representing the confi dence interval for that that group of studies. However, these groupings do not control for some important characteristics of performance management, so it may be misleading to conclude that performance management is more successful in some areas than others without further investiga- tion. Th e forest plot also makes apparent that there is signifi cant heterogeneity between studies—while many are near or around the average, there are signifi cant outliers on both sides, suggesting that study context is important and no single study is representative of the research in this area.

plots in fi gure 2. On the top panel, the black plus signs represent 907 eff ects from peer-reviewed studies. Th e x-axis indicates their eff ect size. Th e y-axis indicates the standard error of each eff ect. Th e funnel eff ect is created by pseudo-confi dence interval bars: as the sample size increases, the standard error decreases and eff ects move toward the mean. Th e black line represents the mean eff ect size of the peer-reviewed studies. Th e bottom panel contains the 907 pub- lished eff ects and overlays the 1,281 unpublished eff ects, indicated with a gray x. Again, the black (right) line represents the mean eff ect size of the peer reviewed studies, and the gray (left) line represents studies not peer reviewed. Most scholars assume that large studies are published regardless of the result because they are both expensive and precisely estimated.

However, researchers tend to shelve small-N studies with results close to zero or against the expectations in the fi eld. In fi gure 2, one

Notes: The top fi gure contains N = 907 peer-reviewed effects. Each plus sign represents one effect. Shaded areas are pseudo-confi dence interval bars. The bottom fi gure contains all 2,188 effects; each x represents an effect from a study not peer reviewed placed over the peer reviewed studies. Effects are sorted from the largest studies (top) to the smallest. Publication bias is typically refl ected by missing effects from the lower quadrants of the top fi gure, potentially represented by studies not peer-reviewed.

Figure 2 Confunnel Plot for Publication Bias

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 57

Notes: Left y-axis lists studies by policy area, then year of publication. The x-axis contains a dot for the mean effect size within study, and horizontal lines refl ect the 95% confi dence interval. Mean effect size is indicated by the middle diamonds, with the diamond edges refl ecting the 95% confi dence interval. The outside lines on the diamonds indicate the 95% prediction interval. t = .0268.

Figure 3 Forest Plot by Policy Area

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

58 Public Administration Review • January | February 2016

Ta ble 3. Meta-Regression: Impact of Performance Management on Performance in Public Organizations

Model 1 Model 2 Model 3 Model 4 Model 5 Model 6 Model 7

Mean effect size 0.030** 0.030* 0.102*** 0.019 0.117*** 0.049** 0.098*** (0.010) (0.012) (0.024) (0.013) (0.032) (0.016) (0.023)

Second-generation PM system 0.001 –0.037* –0.077* (0.017) (0.017) (0.029)

Best practices Benchmarking is … (base is multiple forms)

Limited –0.048* –0.077* –0.079* (0.022) (0.029) (0.032)

Absent –0.070*** –0.095*** –0.122*** (0.013) (0.023) (0.030)

Top-down –0.034 –0.012 –0.029 (0.022) (0.026) (0.024)

Output measure 0.037 (0.059)

0.029 (0.059)

–0.002 (0.030)

Study quality Research design is … (base is randomized)

Pre-post w/comparison –0.016 0.003 –0.008 (0.017) (0.012) (0.009)

Weaker –0.004 –0.018 –0.011 (0.007) (0.009) (0.007)

Endogeneity strategy is. . . (base is randomized or RDD) DiD, FE, matching, etc. 0.047*** –0.009 0.014

(0.013) (0.009) (0.009) Absent 0.040 0.022 0.054*

(0.022) (0.017) (0.022)

Self-selection 0.001 0.038* 0.059* (0.023) (0.017) (0.029)

Not peer-reviewed –0.020 –0.009 –0.018 (0.014) (0.012) (0.011)

Other characteristics Study domain is non-U.S. –0.053 –0.031

(0.040) (0.032) Study domain is state and local –

–0.043* 0.070*

(0.031) Policy arena is … (base is education)

Job training –0.055** 0.057 (0.017) (0.041)

Crime/policing 0.044 –0.087* (0.033) (0.035)

Public health 0.113** 0.061* (0.042) (0.027)

Child support –0.009 0.051 (0.042) (0.042)

Other 0.095* 0.113** (0.045) (0.039)

R2 0.000 0.000 0.111 0.041 0.131 0.113 0.176 Adjusted R2 0.000 –0.000 0.109 0.039 0.126 0.110 0.169 t 2 0.006 0.006 0.006 0.005 0.005 0.005 0.003 I2 0.837 0.835 0.818 0.829 0.829 0.824 0.794

Notes: N = 2,188 from 49 studies. Robust standard errors clustered by study in parentheses. Observations are weighted by inverse variance plus between study variance estimator, t 2. Mean effect sizes in models 3–5 represent best practices and/or ideal study qualities and can be interpreted as such. However, the mean effect sizes in models 6 and 7 are conditional on other characteristics. *p < .05; **p < .01; ***p < .001.

Meta-regression results are reported in table 3, which is organized as follows: model 1 includes the mean eff ect size (the intercept), and model 2 adds a dummy for second-generation performance management systems. Model 3 includes best practice indicators, and model 4 includes study quality indicators. Model 5 combines the variables in models 2, 3, and 4. In models 3–5, the mean eff ect size represented by the intercept can be interpreted as an incrementally better performance management system from the perspective of research, practice, or both. Th e intercept in models 6 and 7 is also conditioned on other characteristics and can be interpreted as set in education, in the United States, and contained within one state or

local government. Th us, the mean eff ect size is not as easily inter- pretable as in models 1–5.

Th is analysis sets out to examine three main ideas. Th e fi rst is the direction and strength of the average eff ect of performance manage- ment on performance—the unconditional mean eff ect size in model 1. Th e second is to examine the eff ect of management practices on the mean eff ect size, hypothesizing that management matters in per- formance management. Th e third is to examine the eff ect of time on performance management—do performance systems show stronger associations with performance over time?

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 59

models 1–5. Yet in combination with performance management best practices in model 5, the intercept is the largest of the fi ve. Th e small coeffi cient on studies that are not peer-reviewed is additional evidence that publication bias is not problematic in this context. Inclusion of other study characteristics in models 6 and 7 does help explain variation in estimated eff ect (R2 increases slightly) but also muddles the interpretation of the mean eff ect size. However, it suggests that there is no evidence that the impact of performance management is uniquely American: studies from Australia, Great Britain, Germany, and other countries have a negative sign in both models 6 and 7, but not strongly so.

Evidence for the impact of performance management in state and local settings has confl icting signs. Without controls for best prac- tices and study quality, the coeffi cient is negative. With controls, the coeffi cient is positive. Both are statistically diff erent from zero. Th ese results make it hard to evaluate the whether state and local govern- ments have the capacity to use performance information eff ectively.

Performance management may have some diff erential impact between policy areas, although it is important to use caution in reading too much into this. Performance management systems in public health (and “other”) are most consistently associated with performance compared with the base case of education, particularly when controlling for best practices and study quality. While it is interesting to speculate why some policy areas appear to be more eff ective than others, it should be noted that in three of the fi ve policy areas examined (job training, crime and policing, and child support), the sign of the coeffi cient changes when study quality and best practice indicators are included. Th is sign switching is not nearly as common in the indicators of best practices. It does not appear that the impact of performance management in particular policy areas is nearly as systematic as how they are implemented.

A last fi nding from table 3 is that only a relatively small amount of the variation in performance management is explained by the factors used in this meta-regression. Th e R2 statistic ranges from about 4 percent to 18 percent. Th e I2 statistic is a measure of the heterogeneity of eff ects between studies and starts at 83.7 percent and declines to 79.4 percent in model 7. I2 is a statistic created by Higgins and Th ompson (2002) that estimates the proportion of total variation in eff ects that can be attributed to heterogeneity. In other words, I2 is the percent of variation that is due to diff erences in study settings and characteristics rather than by random variation alone. Th e fact that a meta-regression with 18 covariates (model 7) only slightly decreases this statistic suggests that there is still signifi - cant unexplained variation in the impact of performance manage- ment on performance in the public sector.

Th e robustness check in table 4 tells largely the same story as table 3, particularly with respect to the intercepts (the conditional or unconditional mean eff ect sizes). Table 4 contains the same model specifi cations as table 3; however, all studies that exam- ined teacher evaluation programs have been excluded, resulting in 1,476 eff ects from 38 studies. One clear diff erence is that the mean eff ect size in models 1 and 2 are almost double the estimated eff ect in table 3. Th is would indicate that the teacher evaluation programs included in this analysis have a negative mean eff ect size, on average.21 However, better practices in models 3 and 5 are still

First, the unconditional mean eff ect size of performance manage- ment systems in table 3, model 1, is 0.030 (p < .01) on a scale from –1 to +1.20 Th e interpretation of this eff ect is typically considered to be negligible. However, there are two items to note about this fi nding. First, this eff ect is stronger than other recent meta-analyses in public policy and management using the same methodology. It is stronger than the eff ect of education vouchers on student perfor- mance (0.009) (Anderson, Guzman, and Ringquist 2013), the eff ect of public service motivation on self-reported organizational performance (0.025) (Warren and Chen 2013), and the eff ect of poverty deconcentration policies on economic well-being (–0.01) and negative behaviors (0.003) (Bolinger and Xu 2013). Th ese stud- ies suggest that the mean eff ect size of most policy interventions will be small. Second, the unconditional mean eff ect size is interpreted as the average performance management system on average per- formance in the average public organization—none of these exist. However, it a useful placeholder to compare against the conditional mean eff ect sizes found in models 2–7.

Performance management best practices signifi cantly increase the strength of the mean eff ect size. In short, management matters. In model 3, including indicators of best practices alone more than dou- bles the estimated mean eff ect size to 0.081 ( p < .001). When com- bined with study quality indicators, it increases more than threefold to 0.103 (p < .001). While still considered negligible, this eff ect represents an incremental increase in organizational performance.

Performance management systems with limited or no benchmarking have weaker associations with performance, and these variables have the expected sign (negative) in all models. While adoption through a top-down approach has the expected sign (negative), the coef- fi cient is close to zero. Th e use of an output rather than an outcome measure of performance has the expected sign in two of three mod- els but is not statistically diff erent from zero. Together, however, all three variables importantly impact the mean eff ect size.

Second-generation performance management systems in models 2, 5, and 7 do not appear to be associated with stronger performance compared with their fi rst-generation counterparts, and the coef- fi cient is negative and signifi cant in two models. Recall, however, that the interpretation is the intercept plus the indicator, meaning that second-generation systems are not associated with negative performance; there is simply less performance gain than their fi rst- generation counterparts. Th ese results suggest that “time,” at least in this context, does not appear to be associated with improved performance of performance management. Th is result could exist because fi rst-generation systems capture low-hanging fruit—subse- quent systems are reformed to maintain performance gains or avoid gaming behavior, and such systems might be expected to be less successful over time within organizations. However, because these results come from a small number of observations, further investiga- tion is warranted.

Th ere are a number of secondary results worth noting. Study quality dummies do not have a common sign, and they are not robust—lower-quality studies do not tend to fi nd more positive results. Including these variables in model 4 does alter the mean eff ect size, however. Including study quality indicators alone in model 4, the conditional mean eff ect size is the smallest among

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

60 Public Administration Review • January | February 2016

Ta ble 4 Impact of Performance Management on Performance in Public Organizations, Excluding Studies of Teacher Evaluation Programs

Model 1 Model 2 Model 3 Model 4 Model 5 Model 6 Model 7

Mean effect size 0.055** 0.057** 0.081* 0.004*** 0.103* 0.051** –0.178** (0.016) (0.019) (0.035) (0.000) (0.041) (0.015) (0.054)

Second-generation PM system –0.021 –0.033 –0.028 (0.028) (0.034) (0.036)

Best practices Benchmarking is … (base is multiple forms)

Limited –0.070 –0.108** –0.144** (0.038) (0.039) (0.048)

Absent –0.068*** –0.104*** –0.180*** (0.013) (0.018) (0.047)

Top-down –0.011 0.005 0.025 (0.035) (0.032) (0.037)

Output measure 0.109 (0.115)

0.093 (0.085)

–0.031 (0.068)

Study quality Research design is … (base is randomized)

Pre-post w/comparison 0.042 0.033 0.042 (0.037) (0.027) (0.024)

Weaker –0.021 0.007 0.008 (0.045) (0.034) (0.024)

Endogeneity strategy is … (base is randomized or regression discontinuity) DiD, FE, matching, etc. 0.070* –0.034 0.218***

(0.031) (0.034) (0.041) Absent 0.038 0.003 0.219***

(0.031) (0.031) (0.031) Self-selection –0.022 0.047 0.122*

(0.036) (0.036) (0.047) Not peer-reviewed 0.013 –0.019 –0.013

(0.044) (0.033) (0.023) Other characteristics Study domain is non-U.S. –0.037 –0.016

(0.058) (0.041) Study domain is state and local –0.034 0.336***

(0.026) (0.070) Policy arena is … (base is education)

Job training –0.059*** 0.102* (0.015) (0.050)

Crime/policing 0.015 –0.251*** (0.050) (0.052)

Public health 0.146* 0.120** (0.061) (0.039)

Child support –0.006 –0.021 (0.046) (0.036)

Other 0.093 0.041 (0.065) (0.065)

R2 0.000 0.002 0.098 0.042 0.114 0.102 0.143 Adjusted R2 0.000 0.001 0.096 0.039 0.107 0.098 0.132 t 2 0.018 0.018 0.018 0.017 0.017 0.015 0.013 I2 0.860 0.860 0.846 0.849 0.849 0.849 0.823

Notes: N = 1,476 from 38 studies. Robust standard errors clustered by study in parentheses. Observations are weighted by inverse variance plus the between study vari- ance estimator, t 2. Sample excludes effects from evaluations of TAP and TAP-like teacher evaluation programs. Mean effect sizes in models 3–5 represent best practices and/or ideal study qualities and can be interpreted as such. Teacher evaluation programs include Schacter et al. (2004), Schacter and Thum (2005), Glazerman, McKie, and Carey (2009), Daley and Kim (2010), Glazerman and Seifullah (2010), Hudson (2010), Glazerman and Seifullah (2012), Fryer (2011), Taylor and Tyler (2012), Fryer (2013), and Dee and Wyckoff (2013). *p < .05; **p < .01; ***p <.001.

larger than model 1. Management still matters in this robustness check. Th e other major diff erence is that the intercept (mean eff ect size) in model 7 is now strongly negative. Th is intercept represents the impact of performance management on performance in state and local U.S. education. Additionally, in model 7, the lack of a strategy to address endogeneity is associated with more positive results.

Similarly, table 5 indicates that the results are not driven by studies that contain a large number of eff ects. Th is table presents results of the same model specifi cations as table 3 but excludes the three

largest studies by number of eff ects (and the working paper of one published study), resulting in 1,239 eff ects from 45 studies. In this robustness check, the intercepts in models that include best practices are the largest in comparison to the unconditional mean eff ect size in model 1, suggesting the largest studies do not drive the results found previously and may even dampen them.

Th e main limitation to the results in tables 3, 4, and 5 is that they do not explicitly account for the gamut of behaviors that may increase measured, although not actual, performance. Some of the studies explicitly examine these behaviors and fi nd them (e.g.,

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 61

Tab le 5 Impact of Performance Management on Performance in Public Organizations, Excluding Three Largest Studies by Number of Effects

Model 1 Model 2 Model 3 Model 4 Model 5 Model 6 Model 7

Mean effect size 0.036* 0.037 0.102** 0.004 0.125*** 0.018 0.036 (0.015) (0.019) (0.031) (0.035) (0.033) (0.018) (0.049)

Second–generation PM system –0.005 –0.055 –0.076* (0.024) (0.034) (0.029)

Best practices Benchmarking is … (base is multiple forms)

Limited –0.050 –0.102** –0.073 (0.032) (0.037) (0.039)

Absent –0.076* –0.124** –0.107* (0.032) (0.038) (0.040)

Top-down –0.031 0.003 –0.019 (0.024) (0.027) (0.027)

Output measure 0.067 0.043 0.003 (0.076) (0.068) (0.043)

Study quality Research design is … (base is randomized)

Pre-post w/comparison 0.016 0.009 0.013 (0.019) (0.014) (0.020)

Weaker –0.009 –0.017 –0.000 (0.023) (0.024) (0.022)

Endogeneity strategy is … (base is randomized or RDD) 0.051 –0.005 0.008 (0.033) (0.016) (0.018)

Absent 0.043 0.029 0.048 (0.027) (0.023) (0.028)

Self-selection –0.004 0.051* 0.087* (0.037) (0.024) (0.034)

Not peer-reviewed –0.002 –0.012 –0.020 (0.028) (0.021) (0.024)

Other characteristics Study domain is non-U.S. –0.031 –0.015

(0.054) (0.045) Study domain is state and local –0.004 0.093**

(0.021) (0.030) Policy arena is … (base is education)

Job training –0.025 0.081 (0.018) (0.047)

Crime/policing 0.017 –0.086* (0.043) (0.037)

Public health 0.156** 0.119* (0.057) (0.058)

Child support 0.022 0.094 (0.044) (0.050)

Other 0.113* 0.140* (0.055) (0.054)

R2 0.000 0.000 0.108 0.033 0.140 0.128 0.176 Adjusted R2 0.000 –0.001 0.105 0.028 0.132 0.123 0.164 t 2 0.012 0.012 0.012 0.011 0.011 0.009 0.007 I2 0.887 0.885 0.876 0.880 0.880 0.877 0.859

Notes: N = 1,239 from 45 studies. Robust standard errors clustered by study in parentheses. Observations are weighted by inverse variance plus the between study variance estimator, t 2. Sample excludes effects from the three largest studies by number of effects. Mean effect sizes in models 3–5 represent best practices and/or ideal study qualities and can be interpreted as such. However, the mean effect sizes in models 6 and 7 are conditional on other characteristics. *p < .05; **p <.01; ***p <.001.

Courty and Marschke 2004, 2008; Gerrish 2014), and others look but do not fi nd them (Hanushek and Raymond 2005), but most studies used in this analysis do not test for dysfunctional/gaming behaviors at all. Th us, it is diffi cult to determine what fraction of performance gains found here are illusory.

Cumulative mean effect size. As discussed previously, another way to operationalize the effect of time on performance management is to examine the mean effect size over time by year. After calculating the mean effect size using the method described earlier, effects are sorted by the last year that the performance management system was in use, typically the last year of the original author’s data. Each effect

is then multiplied by its weight and cumulatively summed by year. The average effect is calculated by dividing the total weighted effect size by the total weight.22

Figure 4 reports the cumulative mean eff ect size and confi dence interval by the last year of the original authors’ data. Th e left y-axis shows the last year that the performance management system was in use. Th e right y-axis calculates the mean eff ect size and confi dence interval, accumulating each previous year. Th is is also shown visu- ally along the x-axis. All 49 studies are included in this fi gure, and the number of eff ects added each year is reported in the rightmost column.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

62 Public Administration Review • January | February 2016

Notes: The left y-axis reports the most recent year of data in the analysis. The right y-axis reports the average effect size (and confi dence interval) in that year. Mean effect size is indicated by the middle diamond, and horizontal lines refl ect the 95% confi dence interval. Numbers to the right are the number of new effects added to the cumulative total in that year. See the Cumulative Mean Effect Size section for more information on how effects are combined by year.

Figure 4 Cumulative Pooled Estimates, Mean Effect Size by Year

Th ere are two interesting results from fi gure 4. Th e fi rst is that prior to 2005, the standard error of the mean eff ect size is quite large and not diff erent from zero. After 2005, the cumulative standard error begins to shrink and the mean eff ect size is positive, and there is suffi cient evidence that the confi dence interval does not include zero. Part of this story is that more data have been brought to bear on this question since the 2000s, explaining the shrinking standard errors. As would be expected from more data, any particular outlier does not impact the mean as strongly as earlier studies. Although confi dence intervals before 2005 contain positive results, evidence from fi gure 4 (together with the meta-regression results on second-generation systems) suggests that perhaps it is not time within an organization that matters. Rather, per- formance management appears to have been adopted more successfully over time in public organizations. Second, fi gure 4 suggests that, as an empirical question, the eff ect of performance management on perfor- mance has become amenable to meta-analytic techniques only recently. Current research drives the more precisely estimated mean eff ect size found in the meta-regressions. In other words, while the debate about performance management has persisted for some time, a systematic review and meta-analysis would have only been able to distinguish the average eff ect from zero after 2006.

Discussion: Quantifying How Management Matters Th e results in the previous section help quantify something we think we know—the use of per- formance measurement does not improve per- formance (much) in and of itself (Behn 2003). Performance management must be moderated by the management of performance information. Th ere are a number of reasons these management infl uences may be associated with better performance. Benchmarking, for instance, can generate a list of valid counterfactuals against which organizations can measure themselves. If

they fi nd that their performance lags behind their comparison group, they have a list of candidate organizations to learn from. Alternatively, it is possible that benchmarking and support from managers are merely indicators of a culture that promotes a number of latent constructs that could not be operationalized in this analysis because of limitations in the original studies. In either case, it is apparent that these infl uences matter by a quantifi able amount—performance management systems that use best practices are two to three times more eff ective than the “average” performance management system.

Understanding how time impacts performance management is more diffi cult. On one hand, performance management systems that exist in a single organization over a longer period of time do not look much diff erent from new systems. On the other hand, the impact of time on performance management in real time (from the 1980s to today) appears to be large. Th ere are (at least) two ways to interpret these results. Th e fi rst is that performance management captures low-hang- ing fruit, making quick performance improvements without altering the secular trend in performance. However, because of changing organ- izational goals or gaming behavior, performance management systems must evolve over time to keep pace. Th e second interpretation is that learning among organizations is more important than within organiza- tions. As a result, new systems can avoid old problems.

Conclusion Th is meta-analysis examines the impact of performance manage- ment systems on performance in public organizations. Specifi cally, it examines three related ideas: fi rst, whether performance manage- ment is positively associated with performance in public organiza- tions, unconditionally; second, whether management practices moderate this eff ect; and third, whether experience and/or time alters the mean eff ect.

Constructing criteria for study inclusion and performing the search results in 49 original coded studies containing 2,188 eff ects. Th e dependent variable and unit of analysis, an eff ect, represents the statistical association between performance management and perfor- mance in public organizations. Th e primary empirical method uses random-eff ects meta-regression, a form of weighted least squares with cluster-robust standard errors.

Meta-regressions fi nd that performance management systems tend to have a small but positive average impact on performance in public organizations, with an unconditional mean eff ect size of 0.03 using the full data set. Th ese results can be interpreted similarly to Pearson’s r, bounded between –1 and +1. When combined with

performance management best practices and indicators for high-quality studies, the mean eff ect size is much larger: two to three times as large (0.10 and 0.12, respectively). While the estimated eff ect size, either conditional or unconditional, is small using typical interpre- tations of this statistic, it is larger than other recent examinations of policy interventions (Anderson, Guzman, and Ringquist 2013;

Bolinger and Xu 2013; Warren and Chen 2013).

Th ese results provide a few recommendations for practitioners man- aging performance. First, understand that measuring performance

Performance management systems tend to have a small but positive average impact on performance in public

organizations.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 63

is not an end unto itself; this research suggests that performance systems using best practice techniques are two to three times more eff ective than average. Benchmarking, in particular, appears to be an eff ective method for learning who is performing well. Second, performance management is used in a wide spectrum of policy areas. No single area appears to dominate the results, suggesting that performance improvements can be made by learning from other areas.

As mentioned earlier, this analysis has two main limitations. First, quantitative meta-analysis cannot balance the small positive eff ect of performance management found here against other values or trade-off s. For example, it cannot address the eff ect of performance management on intrinsic or public service motivation. Second, the extent to which gaming behaviors create illusory performance gains cannot be accounted for in the analysis and is left for future work.

Advancing Future Research Reading (and coding) the quantitative research literature provides a unique vantage point to survey the body of empirical research. Th e performance management literature’s current strength is that it is receiving signifi cant attention from researchers using increasingly sophisticated casual identifi cation strategies: 12 of the studies used in this analysis were written since 2012. Unfortunately, the gold standard of techniques, the randomized fi eld trial, is rare outside of education. Th is may be partially attributable to the fact that performance man- agement is usually considered a process rather than a new program. Th eoretically, elements of performance management could be rand- omized among horizontal entities in a multi-arm trial, which would help answer whether performance management works and why.

Second, researchers need to anticipate gaming responses from the targets of performance schemes. As mentioned earlier, there is a large literature that hypothesizes, models, and fi nds dysfunctional responses in performance systems. However, a number of recent studies do not address gaming, making it diffi cult to systematically account for here. Even when not tied to fi nancial rewards, performance systems can induce gaming behaviors and ought to be considered.

Th ird, performance management requires resources. However, few researchers examine the relationship between cost and performance: effi ciency. Only 3.6 percent of the eff ects examined in this article were measures of effi ciency. Ironically, if performance management systems have displaced democratic values with technocratic values such as effi ciency, they are doing a good job of hiding it.

Finally, research on performance management needs to be cross- pollinated in other policy areas. Th ere is much to be gained from understanding how performance systems are being designed and managed in other contexts, how gaming behaviors are being created, and evidence of failure or success. As it appears that public organiza- tions will be committed to some form of performance management for the foreseeable future, drawing lessons from other contexts to avoid common mistakes should be an important goal of practical policy research.

Acknowledgments Th is article is dedicated to the memory of Evan J. Ringquist, a men- tor. Th anks to Liz Baldwin, Tom Rabovsky, Dave Warren, Shannon

Watkins, Zach Wendling, and Shuang Zhao for helpful comments and suggestions.

Author’s Note Th e author has no confl icts of interest or funding to disclose.

Notes 1. It is unclear how common performance management really is, although most

evidence suggests it is very common. Federally, there is no punchy four-letter acronym for performance management in the Obama administration aside from the modernization of the Government Performance and Results Act (GPRA). However, the Offi ce of Management and Budget notes that elements of perfor- mance measurement are still in use in many federal agencies (Joyce 2011). How common performance management systems are among local governments is also fuzzy, although survey evidence from a variety of sources suggests that it is quite common. A 2005 survey reported that “almost half of the respondents (47.8 percent) noted that all departments within their city or county use performance measures, and another 20 percent noted that at least half of their departments do so” (Melkers and Willoughby 2005, 183).

2. Th ese studies include Schacter et al. (2004), Schacter and Th um (2005), Glazerman, McKie, and Carey (2009), Daley and Kim (2010), Glazerman and Seifullah (2010), Hudson (2010), Glazerman and Seifullah (2012), Fryer (2011), Taylor and Tyler (2012), Fryer (2013), and Dee and Wyckoff (2013).

3. Th e Job Partnership Training Act, for example, created some interesting dilemmas with respect to coding the publicness of the organizations. Many of the studies in this area (Barnow 2000; Barnow and Smith 2004; Cragg 1997; Heckman, Heinrich, and Smith 1997, 2002) note that there were both public and private job training services. In some cases, the estimated eff ect refl ects contracting with only private sector organizations. Whenever it was possible, the eff ects from the private providers were excluded from this analysis. Th is was pos- sible in most analyses in which it was understood that the eff ect of performance incentives might diff er between the two groups.

4. In contracting, the organization can be profi t motivated while still being accountable to an elected authority. Lack of accountability to public offi cials also means that this analysis excludes nonprofi t organizations unless they are working directly on behalf of public organizations.

5. Th e exact search phrases were “child support enforcement,” “food stamps,” “food assistance,” “nutrition assistance,” “Medicaid,” “Medicare,” “state children’s health insurance program” (abbreviated and unabbreviated), “CompStat,” “CitiStat,” “police AND crime,” “prison,” “education,” “No Child Left Behind,” “transporta- tion,” “construction,” “welfare,” “Job Training Partnership Act,” “Government Performance and Results Act,” and “Program Assessment Rating Tool.”

6. Academic Search Premier, British Library Document Supply Service, Business Source Premier, ERIC, Google Scholar, ISI Web of Knowledge, JSTOR, ProQuest, ProQuest Dissertations and Th eses Global, PsycINFO, and WorldCat.

7. Association for Public Policy Analysis and Management, American Society for Public Administration, and Political Research Online.

8. EconLit, National Bureau of Economic Research, and SSRN. 9. Government Accountability Offi ce, Offi ce of Management and Budget, and

Congressional Budget Offi ce. 10. Abt Associates, American Enterprise Institute, Brookings Institution, Heritage

Foundation, Mathematica Policy Research, MDRC, National Academy of Public Administration, National Center on Performance Incentives, Performance Institute, RAND Corporation, Urban Institute, and Westat.

11. No acceptable studies found are excluded from this analysis. Any acceptable studies not found and included are an error on the part of the author.

12. To avoid duplication of eff ects from Dee and Jacob (2009, 2011), all of the eff ects that occur in both papers have been removed from coding in the working paper.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

64 Public Administration Review • January | February 2016

13. Data extraction was conducted using standard coding documents, and they are available upon request. All variables were coded by the author; variable coding by a single author is a noted weakness in meta-analysis because of a lack of construct agreement statistics.

14. Because not all research takes place in the United States, the diff erence between interstate and federal may not be as sharp in some cases as in others.

15. Overall study quality is moderate to low; 65.2 percent of studies have a design that is weaker than pre-post with a comparison group. Th is rises to 79.3 percent when teacher evaluation programs are excluded, owing to the fact that some randomized control experiments were conducted evaluating TAP. However, most studies (all but 13 percent) had some strategy for addressing possible endogene- ity, even if it was just fi xed-eff ects dummy variables.

16. Cameron, Gelbach, and Miller (2008) fi nd that cluster-robust variance estima- tors can still have infl ated alpha rejection rates (standard errors are biased downward) when the number of groups is fewer than 30 and the number of observations within groups varies. However, simulation results suggest that as the number of groups increase, this problem diminishes. Th ey test up to 30 groups, fi nding that the bias diminishes with size. Th is analysis contains 49 studies.

17. A Q-test for fi xed versus random eff ects examines the diff ering variation between studies expected by chance versus the observed variation. Use of random eff ects also means that an estimate of the between-study variance of study eff ects (known as t 2) is multiplied to the inverse of the regression weights. See Stata’s metareg command for information on calculating t 2.

18. A further test can be conducted by examining the eff ect of the coeffi cient for peer- reviewed studies in meta-regression. Th is test largely follows the logic of Egger’s test while also conditioning eff ects on other factors that might be correlated with publication, such as the strength of the research design. Th is covariate provides further evidence that publication bias does not appear to be problematic.

19. Th e average eff ect size by study was calculated by multiplying each eff ect size by its weight (adjusted by t using all 2,188 eff ects), divided by the sum of the weights within each study. Th is arrives at an average eff ect size that ignores dif- ferences in the models or data used. Th e grand mean, 0.031, will diff er slightly more than the mean found in meta-regressions because of rounding diff erences caused by fi rst aggregating eff ects by study.

20. As noted previously, the scale of Fisher’s Z is actually unbounded, but converting Z to r does not change the estimated eff ect size.

21. Th e exact eff ect size for teacher evaluation programs is –0.0004. Th is is not the same as merely subtracting the two mean eff ect sizes because including these studies alters the study weights.

22. Mathematically, , where Z is the transformed eff ect, Fisher’s Z;

Z‒t is the mean eff ect size by year; i indexes the eff ect; w is the weight; and N – 3. N is the number of observations in the original study. Zt is calculated cumula- tively each year.

References *Indicates study coded for meta-analysis. Amm ons, David N. 1995. Overcoming the Inadequacies of Performance

Measurement in Local Government: Th e Case of Libraries and Leisure Services. Public Administration Review 55(1): 37–47.

And erson, Mary R., Tatyana Guzman, and Evan J. Ringquist. 2013. Evaluating the Eff ectiveness of Educational Vouchers. In Meta-Analysis for Public Management and Policy, edited by Evan J. Ringquist and Mary R. Anderson, 202–37. San Francisco: Jossey-Bass.

Bak er, George P., Michael C. Jensen, and Kevin J. Murphy. 1988. Compensation and Incentives: Practice vs. Th eory. Journal of Finance 43(3): 593–616.

Bar now, Burt S. 2000. Exploring the Relationship between Performance Management and Program Impact: A Case Study of the Job Training Partnership Act. Journal of Policy Analysis and Management 19(1): 118–41.*

Bar now, Burt S., and Jeff rey A. Smith. 2004. Performance Management of U.S. Job Training Programs: Lessons from the Job Training Partnership Act. Public Finance and Management 4(3): 247–87.

Beh n, Robert D. 2003. Why Measure Performance? Diff erent Purposes Require Diff erent Measures. Public Administration Review 63(5): 586–606.

——— . 2014. Th e PerformanceStat Potential: A Leadership Strategy for Producing Results. Washington, DC: Brookings Institution Press.

Ber man, Evan, and XiaoHu Wang. 2000. Performance Measurement in U.S. Counties: Capacity for Reform. Public Administration Review 60(5): 409–20.

Be sley, Timothy J., Gwyn Bevan, and Konrad B. Burchardi. 2009. Naming and Shaming: Th e Impacts of Diff erent Regimes on Hospital Waiting Times in England and Wales. Discussion Paper no. 7306, Center for Economic and Policy Research, Washington, DC.*

Bol inger, Joe, and Lanlan Xu. 2013. Th e Eff ects of Federal Assisted Deconcentration Eff orts on Economic Self-Suffi ciency and Problematic Behaviors. In Meta- Analysis for Public Management and Policy, edited by Evan J. Ringquist and Mary R. Anderson, 276–308. San Francisco: Jossey-Bass.

Bor enstein, Michael, Larry V. Hedges, Julian P. T. Higgins, and Hannah R. Rothstein. 2011. Introduction to Meta-Analysis. Chichester, UK: Wiley.

Boz eman, Barry, and Stuart Bretschneider. 1994. Th e “Publicness Puzzle” in Organization Th eory: A Test of Alternative Explanations of Diff erences between Public and Private Organizations. Journal of Public Administration Research and Th eory 4(2): 197–224.

Cam eron, A. Colin, Jonah B. Gelbach, and Douglas L. Miller. 2008. Bootstrap- Based Improvements for Inference with Clustered Errors. Review of Economics and Statistics 90(3): 414–27.

Car d, Noel A. 2012. Applied Meta-Analysis for Social Science Research. New York: Guilford Press.

Cav alluzzo, Ken S., and Christopher D. Ittner. 2003. Implementing Performance Measurement Innovations: Evidence from Government. Accounting, Organizations and Society 29(3–4): 243–67.

Chilvers, Marilyn, and Don Weatherburn. 2001. Do Targeted Arrests Reduce Crime? Crime and Justice Bulletin, no. 63.*

——— . 2004. Th e New South Wales “CompStat” Process: Its Impact on Crime. Australian and New Zealand Journal of Criminology 37(1): 22–48.*

Cou rty, Pascal, and Gerald Marschke. 2004. An Empirical Investigation of Gaming Responses to Explicit Performance Incentives. Journal of Labor Economics 22(1): 23–56.

——— . 2008. A General Test for Distortions in Performance Measures. Review of Economics and Statistics 90(3): 428–41.*

Cra gg, Michael. 1997. Performance Incentives in the Public Sector: Evidence from the Job Training Partnership Act. Journal of Law, Economics, and Organization 13(1): 147–68.*

Dal ey, Glenn, and Lydia Kim. 2010. A Teacher Evaluation System Th at Works. Working paper, National Institute for Excellence in Teaching, Santa Monica, CA. http://fi les.eric.ed.gov/fulltext/ED533380.pdf [accessed July 7, 2015].*

Dee , Th omas, and Brian Jacob. 2009. Th e Impact of No Child Left Behind on Student Achievement. Working Paper no. 15531, National Bureau of Economic Research, Cambridge, MA. http://www.nber.org/papers/w15531 [accessed July 7, 2015].*

——— . 2011. Th e Impact of No Child Left Behind on Student Achievement. Journal of Policy Analysis and Management 30(3): 418–46.*

Dee , Th omas, and James Wyckoff . 2013. Incentives, Selection, and Teacher Performance: Evidence from IMPACT. Working Paper no. 19529, National Bureau of Economic Research, Cambridge, MA. http://www.nber.org/papers/ w19529 [accessed July 7, 2015].*

Dickinson, Katherine P., Richard W. West, Deborah J. Kogan, David A. Drury, Marlene S. Franks, Laura Schlichtmann, and Mary Vencil. 1988. Evaluation of the Eff ects of JTPA Performance Standards on Clients, Services, and Costs.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

The Impact of Performance Management on Performance in Public Organizations: A Meta-Analysis 65

Research Report no. 88-16., National Commission for Employment Policy, Menlo Park, CA.*

For sythe, Dall W. 2001. Quicker, Better, Cheaper? Managing Performance in American Government. Albany: State University of New York Press.

Fra nco, Annie, Neil Malhotra, and Gabor Simonovits. 2014. Publication Bias in the Social Sciences: Unlocking the File Drawer. Science 345(6203): 1502–5.

Fre derickson, David G. 2003. Performance Measurement and Th ird-Party Government: A Study of the Implementation of the Government Performance and Results Act in Five Health Agencies. PhD diss., Indiana University.

Fre derickson, David G., and H. George Frederickson. 2006. Measuring the Performance of the Hollow State. Washington, DC: Georgetown University Press.

Fry er, Roland G. 2011. Teacher Incentives and Student Achievement: Evidence from New York City Public Schools. Working Paper no. 16850, National Bureau of Economic Rese arch, Cambridge, MA. http://www.nber.org/papers/w16850 [accessed July 7, 2015].*

———. 2013. Teacher Incentives and Student Achievement: Evidence from New York City Public Schools. Journal of Labor Economics 31(2): 373–407.*

Garicano, Luis, and Paul Heaton. 2010. Information Technology, Organization, and Productivity in the Public Sector: Evidence from Police Departments. Journal of Labor Economics 28(1): 167–201.*

Ger rish, Ed. 2014. Th e Eff ect of the Child Support Performance and Incentive Act of 1998 on Rewarded and Unrewarded Performance Goals. Working paper, Indiana University. http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2476033 [accessed July 7, 2015].*

Gla zerman, Steven, Allison McKie, and Nancy Carey. 2009. An Evaluation of the Teacher Advancement Program (TAP) in Chicago: Year One Impact Report. Washington, DC: Mathematica Policy Research.*

Glazerman, Steven, and Allison Seifullah. 2010. An Evaluation of the Teacher Advancement Program (TAP) in Chicago: Year Two Impact Report. Washington, DC: Mathematica Policy Research.*

——— . 2012. An Evaluation of the Chicago Teacher Advancement Program (Chicago TAP) after Four Years: Final Report. Washington, DC: Mathematica Policy Research.*

Gor e, Albert. 1993. From Red Tape to Results: Creating a Government Th at Works Better and Costs Less: Report of the National Performance Review. Washington, DC: U.S. Government Printing Offi ce.

Han ushek, Eric A., and Margaret E. Raymond. 2005. Does School Accountability Lead to Improved Student Performance? Journal of Policy Analysis and Management 24(2): 297–327.*

Hat ry, Harry P. 2006. Performance Measurement: Getting Results. Washington, DC: Urban Institute Press.

Hec kman, James, Carolyn Heinrich, and Jeff rey Smith. 1997. Assessing the Performance of Performance Standards in Public Bureaucracies. American Economic Review 87(2): 389–95.

——— . 2002. Th e Performance of Performance Standards. Journal of Human Resources 37(4): 778–811.*

Hei nrich, Carolyn J. 2002. Outcomes-Based Performance Management in the Public Sector: Implications for Government Accountability and Eff ectiveness. Public Administration Review 62(6): 712–25.*

Hei nrich, Carolyn J., and Laurence E. Lynn, Jr. 2001. Means and Ends: A Comparative Study of Empirical Methods for Investigating Governance and Performance. Journal of Public Administration Research and Th eory 11(1): 109–38.*

Hei nrich, Carolyn J., and Gerald Marschke. 2010. Incentives and Th eir Dynamics in Public Sector Performance Management Systems. Journal of Policy Analysis and Management 29(1): 183–208.

Hig gins, Julian P. T., and Simon G. Th ompson. 2002. Quantifying Heterogeneity in a Meta-Analysis. Statistics in Medicine 21(11): 1539–58.

Hol mstrom, Bengt, and Paul Milgrom. 1991. Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design. Special issue, Journal of Law, Economics and Organization 7: 24–52.

Hoo d, Christopher. 1995. Th e “New Public Management” in the 1980s: Variations on a Th eme. Accounting, Organizations and Society 20(2–3): 93–109.

Horton, James. 2010. An Examination of the Applicability of the CitiStat Performance Management System to Municipal Fire Departments. PhD diss., University of Texas at Arlington.*

Hua ng, Chien Chung, and Richard L. Edwards. 2009. Th e Relationship between State Eff orts and Child Support Performance. Children and Youth Services Review 31(2): 243–48.*

Hud son, Sally. 2010. Th e Eff ects of Performance-Based Teacher Pay on Student Achievement. Undergraduate thesis, Stanford University.*

Hvi dman, Ulrik, and Simon Calmar Andersen. 2014. Impact of Performance Management in Public and Private Organizations. Journal of Public Administration Research and Th eory 24(1): 35–58.*

Imb ens, Guido W., and Th omas Lemieux. 2008. Regression Discontinuity Designs: A Guide to Practice. Journal of Econometrics 142(2): 615–35.

Ing raham, Patricia W., Philip G. Joyce, and Amy Kneedler Donahue. 2003. Government Performance: Why Management Matters. Baltimore: Johns Hopkins University Press.

Jang, Hyunseok. 2008. Evaluation of CompStat Policing Strategy Using Broken Windows Enforcement in Two Texas Police Departments: A Time Series Analysis. PhD diss., Sam Houston State University.*

Jan g, Hyunseok, Larry T. Hoover, and Hee-Jong Joo. 2010. An Evaluation of CompStat’s Eff ect on Crime: Th e Fort Worth Experience. Police Quarterly 13(4): 387–412.*

Joy ce, Philip G. 2011. Th e Obama Administration and PBB: Building On the Legacy of Federal Performance-Informed Budgeting? Public Administration Review 71(3): 356–67.

Jul nes, Patria De Lancer, and Marc Holzer. 2001. Promoting the Utilization of Performance Measures in Public Organizations: An Empirical Study of Factors Aff ecting Adoption and Implementation. Public Administration Review 61(6): 693–708.

Kimm, Victor J. 1995. GPRA: Early Implementation. Public Manager 24(1): 11–14. Klo ot, Louise, and John Martin. 2000. Strategic Performance Management: A

Balanced Approach to Performance Management Issues in Local Government. Management Accounting Research 11(2): 231–51.

Kro ll, Alexander, and Donald P. Moynihan. 2015. Does Training Matter? Evidence from Performance Management Reforms. Public Administration Review 75(3): 411–20.

Lav ertu, Stéphane, and Donald P. Moynihan. 2012. Agency Political Ideology and Reform Implementation: Performance Management in the Bush Administration. Journal of Public Administration Research and Th eory 23(3): 521–49.

Lin dell, Michael K., and David J. Whitney. 2001. Accounting for Common Method Variance in Cross-Sectional Research Designs. Journal of Applied Psychology 86(1): 114–21.

Lockwood, Ben, and Francesco Porcelli. 2013. Incentive Schemes for Local Government: Th eory and Evidence from Comprehensive Performance Assessment in England. American Economic Journal: Economic Policy 5(3): 254–86.*

Lon g, Edward, and Aimee L. Franklin. 2004. Th e Paradox of Implementing the Government Performance and Results Act: Top-Down Direction for Bottom-Up Implementation. Public Administration Review 64(3): 309–19.

Marsh, Julie A., Matthew G. Springer, Daniel F. McCaff rey, Kun Yuan, Scott Epstein, Julia Koppich, Nidhi Kalra, Catherine DiMartino, and Art (Xiao) Peng. 2011. A Big Apple for Educators. Santa Monica, CA: RAND Corporation.*

Mazerolle, Lorraine, James McBroom, and Sacha Rombouts. 2011. CompStat in Australia: An Analysis of the Spatial and Temporal Impact. Journal of Criminal Justice 39(2): 128–36.*

Mazerolle, Lorraine, Sacha Rombouts, and James McBroom. 2006. Th e Impact of Operational Performance Reviews on Reported Crime in Queensland. Trends and Issues in Crime and Criminal Justice, no. 313.*

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense

66 Public Administration Review • January | February 2016

——— . 2007. Th e Impact of COMPSTAT on Reported Crime in Queensland. Policing: An International Journal of Police Strategies and Management 30(2): 237–56.*

Mel kers, Julia, and Katherine Willoughby. 2005. Models of Performance- Measurement Use in Local Governments: Understanding Budgeting, Communication, and Lasting Eff ects. Public Administration Review 65(2): 180–190.

Mih m, J. Christopher. 1995. GPRA and the New Dialogue. Public Manager 24(4): 15–18.

Moy nihan, Donald P. 2005. Testing How Management Matters in an Era of Government by Performance Management. Journal of Public Administration Research and Th eory 15(3): 421–39.

——— . 2008. Th e Dynamics of Performance Management: Constructing Information and Reform. Washington, DC: Georgetown University Press.

Moy nihan, Donald P., and Sanjay K. Pandey. 2010. Th e Big Question for Performance Management: Why Do Managers Use Performance Information? Journal of Public Administration Research and Th eory 20(4): 849–66.

Nielsen, Poul A. 2013. Performance Management, Managerial Authority, and Public Service Performance. Journal of Public Administration Research and Th eory 24(2): 431–58.*

Osb orne, David, and Ted Gaebler. 1992. Reinventing Government: How the Entrepreneurial Spirit Is Transforming the Public Sector. Reading, MA: Addison-Wesley.

Osb orne, David, and Peter Plastrik. 1997. Banishing Bureaucracy: Th e Five Strategies for Reinventing Government. Reading, MA: Addison-Wesley.

Per ry, James L., and Lois Recascino Wise. 1990. Th e Motivational Bases of Public Service. Public Administration Review 50(3): 367–73.

Poister, Th eodore H., Obed Q. Pasha, and Lauren Hamilton Edwards. 2013. Does Performance Management Lead to Better Outcomes? Evidence from the U.S. Public Transit Industry. Public Administration Review 73(4): 625–36.*

Pro pper, Carol, Matt Sutton, Carolyn Whitnall, and Frank Windmeijer. 2008. Did “Targets and Terror” Reduce Waiting Times in England for Hospital Care? Health Care Economics and Policy 8(2): 1935–82.*

——— . 2010. Incentives and Targets in Hospital Care: Evidence from a Natural Experiment. Journal of Public Economics 94(3–4): 318–35.*

R adin, Beryl A. 1998. Th e Government Performance and Results Act (GPRA): Hydra-Headed Monster or Flexible Management Tool? Public Administration Review 58(4): 307–16.

——— . 2000. Th e Government Performance and Results Act and the Tradition of Federal Management Reform: Square Pegs in Round Holes? Journal of Public Administration Research and Th eory 10(1): 111–35.

———. 2006. Challenging the Performance Movement: Accountability, Complexity, and Democratic Values. Washington, DC: Georgetown University Press.

Rai ney, Hal G. 1982. Reward Preferences among Public and Private Managers: In Search of the Service Ethic. American Review of Public Administration 16(4): 288–302.

Rin gquist, Evan. 2013. Meta-Analysis for Public Management and Policy. Edited by Mary R. Anderson. San Francisco: Jossey-Bass.

Ros enfeld, Richard, Robert Fornango, and Eric Baumer. 2005. Did Ceasefi re, CompStat, and Exile Reduce Homicide? Criminology and Public Policy 4(3): 419–49.*

Sch acter, John, and Yeow Meng Th um. 2005. TAPping into High Quality Teachers: Preliminary Results from the Teacher Advancement Program Comprehensive School Reform. School Eff ectiveness and School Improvement 16(3): 327–53.*

Sch acter, John, Yeow Meng Th um, Daren Reifsneider, and Tamara Schiff . 2004. Th e Teacher Advancement Program Report Two: Year Th ree Results from Arizona and Year One Results from South Carolina TAP Schools. Working paper, National Institute of Education and Training, Cupertino, CA.*

Schochet, Peter Z., and John A. Burghardt. 2008. Do Job Corps Performance Measures Track Program Impacts? Journal of Policy Analysis and Management 27(3): 556–76.*

Schochet, Peter Z., and Jane Fortson. 2014. When Do Regression-Adjusted Performance Measures Track Longer-Term Program Impacts? A Case Study for Job Corps. Journal of Policy Analysis and Management 33(2): 495–525.*

Smi th, Dennis C., and William J. Bratton. 2001. Performance Management in New York City: CompStat and the Revolution in Police Management. In Quicker, Better, Cheaper? Managing Performance in American Government, edited by Dall W. Forsythe, 453–82. Albany, NY: Rockefeller Institute Press.

Springer, Matthew, Dale Ballou, and Art (Xiao) Peng. 2008. Impact of the Teacher Advancement Program on Student Test Score Gains: Findings from an Independent Appraisal. Working paper, National Center on Performance Incentives, Nashville, TN.*

Springer, Matthew G., John F. Pane, Vi-Nhuan Le, Daniel F. McCaff rey, Susan Freeman Burns, Laura S. Hamilton, and Brian Stecher. 2012. Team Pay for Performance Experimental Evidence from the Round Rock Pilot Project on Team Incentives. Educational Evaluation and Policy Analysis 34(4): 367–90.*

Tay lor, Eric S., and John H. Tyler. 2012. Th e Eff ect of Evaluation on Teacher Performance. American Economic Review 102(7): 3628–51.*

Th ibodeau, Nicole. 2003. Improving the Organizational Architecture of Public Enterprise: An Investigation of the Eff ect of the Federal Government’s Latest Eff ort through the Veterans Health Administration. PhD diss., University of Pittsburgh.*

Th ibodeau, Nicole, John Harry Evans III, Nandu J. Nagarajan, and Jeff Whittle. 2007. Value Creation in Public Enterprises: An Empirical Analysis of Coordinated Organizational Changes in the Veterans Health Administration. Accounting Review 82(2): 483–520.*

Walker, Richard M., Fariborz Damanpour, and Carlos A. Devece. 2011. Management Innovation and Organizational Performance: Th e Mediating Eff ect of Performance Management. Journal of Public Administration Research and Th eory 21(2): 367–86.*

Wan g, XiaoHu, and Evan Berman. 2001. Hypotheses about Performance Measurement in Counties: Findings from a Survey. Journal of Public Administration Research and Th eory 11(3): 403–28.

War ren, David C., and Li-Ting Chen. 2013. Th e Relationship between Public Service Motivation and Performance. In Meta-Analysis for Public Management and Policy, edited by Evan J. Ringquist and Mary R. Anderson, 309–31. San Francisco: Jossey-Bass.

Who ley, Joseph S., and Harry P. Hatry. 1992. Th e Case for Performance Monitoring. Public Administration Review 52(6): 604–10.

Wilborn, Doris, Ulrike Grittner, Th eo Dassen, and Jan Kottner. 2010. Th e National Expert Standard Pressure Ulcer Prevention in Nursing and Pressure Ulcer Prevalence in German Health Care Facilities: A Multilevel Analysis. Journal of Clinical Nursing 19(23–24): 3364–71.*

Wil liams, Daniel W. 2003. Measuring Government in the Early Twentieth Century. Public Administration Review 63(6): 643–59.

Wol f, Fredric M. 1986. Meta-Analysis: Quantitative Methods for Research Synthesis. Newbury Park: Sage Publications.

Yan g, Kaifeng, and Jun Yi Hsieh. 2007. Managerial Eff ectiveness of Government Performance Measurement: Testing a Middle-Range Model. Public Administration Review 67(5): 861–79.

Supporting Information Additional supporting information may be found in the online version of this article at http://onlinelibrary.wiley.com/journal/10.1111/ (ISSN)1540-6210.

15406210, 2016, 1, D ow

nloaded from https://onlinelibrary.w

iley.com /doi/10.1111/puar.12433 by L

iberty U niversity, W

iley O nline L

ibrary on [31/08/2024]. See the T erm

s and C onditions (https://onlinelibrary.w

iley.com /term

s-and-conditions) on W iley O

nline L ibrary for rules of use; O

A articles are governed by the applicable C

reative C om

m ons L

icense