Week 6 - Assignment: Prepare Job Descriptions and Calls for Applicants and Week 7 - Assignment: Appraise Strategies to Assist in Effective Job Evaluations
https://doi.org/10.1177/0734371X17704891
Review of Public Personnel Administration 2017, Vol. 37(3) 275 –294
© The Author(s) 2017 Reprints and permissions:
sagepub.com/journalsPermissions.nav DOI: 10.1177/0734371X17704891
journals.sagepub.com/home/rop
Article
Cognitive Biases in Performance Appraisal: Experimental Evidence on Anchoring and Halo Effects With Public Sector Managers and Employees
Nicola Belle1, Paola Cantarelli2, and Paolo Belardinelli2
Abstract A systematic literature review of performance appraisal in a selection of public administration journals revealed a lack of investigations on the cognitive biases that affect raters’ evaluation of ratees’ performance. To address this gap, we conducted two artefactual field experiments on a sample of 600 public sector managers and employees. Results show that anchoring and halo effects systematically biased performance ratings. For the former, average scores were higher when subjects were exposed to a high rather than a low anchor. For the latter, higher ability on one performance dimension led participants to provide a higher average score on another performance dimension. Halo effect was moderated by rater’s gender. We conclude by discussing the study limitations and providing suggestions for future work in this area.
Keywords performance appraisal, anchoring effect, halo effect, systematic literature review, artefactual field experiments
1Scuola Superiore Sant’Anna MHL Laboratory, Pisa, Italy 2Bocconi University, Milan, Italy
Corresponding Author: Nicola Belle, Scuola Superiore Sant’Anna MHL Laboratory, Piazza Martiri della Libertà 33, Pisa 56127, Italy. Email: [email protected]
704891ROPXXX10.1177/0734371X17704891Review of Public Personnel AdministrationBelle et al. research-article2017
276 Review of Public Personnel Administration 37(3)
Introduction
Performance appraisal of public employees was originally introduced in 1978 in the United States as a key provision of the Civil Service Reform Act, and it was later included as a foundational element in the New Public Management reform waves. It has since been adopted by governments and public sector organizations around the world (e.g., Christensen, Dong, Painter, & Walker, 2012; Lah & Perry, 2008; Liebert, 2014; Organization for Economic Cooperation and Development [OECD], 2012). As a result, an abundant amount of attention of public administration research has been dedicated to illuminating our understanding of individual performance evaluation of government workers. Scholars’ and practitioners’ work in this area is unlikely to decline as “effectively managing performance appraisal in the public sector is increas- ingly important given the drive toward greater accountability for results” (Battaglio, 2015, p. 207).
The first aim of our work is to contribute to this body of knowledge by identifying the main research topics on performance appraisal that researchers in our field have explored so far and the primary research designs that they have employed. We do so by conducting a systematic literature review of a selection of public administration journals included in the 2015 Institute for Scientific Information (ISI; 2015) Journal Citation Reports (©Thomson Reuters). We find that early research (e.g., J. S. Bowman, 1999; Martin & Bartol, 1986) argued that public employees need to be trained if they are to be competent in evaluating others’ performance, and need to be made aware of cognitive biases characteristic of human nature if their performance scores are to be free of systematic errors. Recently, Battaglio (2015) convincingly stated that “one of the primary concerns of performance appraisal is the error and bias of raters. Given that all performance appraisals are subject to human involvement, error and bias are constant threats to effective evaluation” (p. 203). Regardless of such warnings, how- ever, our systematic literature review did not discover any empirical study on cogni- tive biases in performance appraisal. Unlike in our field, experimental research in disciplines such as applied psychology (e.g., Thorsteinson, Breier, Atwell, Hamilton, & Privette, 2008) and behavioral economics (e.g., Furnham & Boo, 2011; Kahneman, 2011) have long suggested that cognitive biases may systematically affect raters’ eval- uation of ratees’ performance.
In particular, anchoring and halo effects consistently have been shown to affect the performance scores that raters assign to ratees. “Anchoring is a pervasive and robust effect in human decisions regardless of factors such as types of anchors, relevance of anchor cues, expertise, motivation and cognitive load” (Furnham & Boo, 2011, p. 41). Similarly, halo errors are often included among the most common rating errors (e.g., Balzer & Sulsky, 1992; Battaglio, 2015; Bechger, Maris, & Hsiao, 2010; J. S. Bowman, 1999; Martin & Bartol, 1986). The second contribution of our work lies in providing much needed novel experimental evidence of the anchoring and halo effects in perfor- mance appraisal using real public sector managers and employees as raters. In so doing, we follow up on recent calls in our field to encourage cross-fertilization among disci- plines (e.g., Grimmelikhuijsen, Jilke, Olsen, & Tummers, 2017). Also, we further
Belle et al. 277
nurture the efforts of other scholars in our field aimed at studying cognitive biases in public administration and management (e.g., Andersen & Hjortskov, 2015; Marvel, 2015, 2016).
Performance Appraisal in Public Administration
Cross-national research as of 2008 shows that about 93% of the countries belonging to the OECD have adopted performance appraisal systems for their civil servants (Lah & Perry, 2008). Recent observational research also provides evidence of the adoption of individual performance evaluation in government and public sector organizations in non-Western countries (e.g., Bissessar, 2000; Christensen et al., 2012; Liebert, 2014; Liu & Dong, 2012).
To capitalize on existing knowledge about public employees’ performance appraisal and better highlight the contributions of our study, we conducted a systematic litera- ture review in a selection of public administration journals. In particular, we searched for articles that contained “performance appraisal” or “performance evaluation” in the title or abstract and were published in the Journal of Public Administration Research and Theory, Public Administration Review, Review of Public Personnel Administration, or Public Personnel Management. We selected the first two journals because they are top-tier in our field (ISI Journal Citation Reports ©Thomson Reuters 2015) and the last two journals because they are specifically dedicated to public human resources management. We also included the International Public Management Journal for its international as well as managerial scopes. Due to restricted access to this last source, and based on the need to keep the search manageable, in this last case, we could search only for articles that contained “performance appraisal” anywhere in the text and had been published since 2005.
These searches, which we performed in August 2016, returned 293 articles. Upon eliminating the studies that did not deal with public sector workers’ individual perfor- mance appraisal, we retained 181 articles. We performed an in-depth analysis of these articles with a double aim. On one side, we were interested in identifying the studies’ main research question to group them into main research themes. On the other side, we aimed to describe these manuscripts based on their research methods or approach. Like in many other systematic literature reviews (e.g., R. Bowman & Grimmelikhuijsen, 2016; Cantarelli, Belardinelli, & Belle, 2016; Tummers, Bekkers, Vink, & Musheno, 2015), we had to make judgmental calls in identifying research themes in an effort to balance between synthesizing available knowledge and not losing too much of the studies’ individual qualities.
Table 1 shows the results of our systematic literature review on the performance appraisal of public employees. The 181 studies are categorized based on their primary research theme and primary research method or approach. With respect to the former, about 32% of the articles illuminate our understanding of the overall design of perfor- mance appraisal systems. For instance, work in this area addressed questions such as what conditions were expected to be more conducive for the real implementation of performance evaluation (e.g., Brown, 1982; Feild & Holley, 1975; Hyde, 1982) or the
278
T a b
le 1
. N
um be
r an
d G
ra nd
T o ta
l P er
ce nt
ag e
o f th
e St
ud ie
s o n
In di
vi du
al P
er fo
rm an
ce A
pp ra
is al
in t
he P
ub lic
S ec
to r
by T
he ir
P ri
m ar
y R
es ea
rc h
T he
m e
an d
Pr im
ar y
R es
ea rc
h D
es ig
n/ A
pp ro
ac h
(R o w
s an
d C
o lu
m ns
A re
O rd
er ed
B as
ed o
n T
he ir
R es
pe ct
iv e
T o ta
ls ).
Su rv
ey N
o rm
at iv
e ap
pr o ac
h C
as e
st ud
y Li
te ra
tu re
re
vi ew
Pr e–
po st
te
st T
o ta
l
D es
ig n
o f th
e pe
rf o rm
an ce
a pp
ra is
al s
ys te
m 13
7% 30
17 %
11 6%
4 2%
58 32
% A
ss o ci
at io
n be
tw ee
n pe
rf o rm
an ce
a pp
ra is
al
an d
em pl
o ye
es ’ p
er ce
pt io
ns /m
o ti va
ti o n/
ch ar
ac te
ri st
ic s
44 24
% 2
1% 1
1% 1
1% 48
27 %
Sh o rt
co m
in gs
o f pe
rf o rm
an ce
a pp
ra is
al 4
2% 11
6% 7
4% 22
12 %
Pr ac
ti ce
s o f pe
rf o rm
an ce
a pp
ra is
al 7
4% 1
1% 12
7% 20
11 %
Pe rf
o rm
an ce
a pp
ra is
al a
nd m
er it p
ay 2
1% 6
3% 4
2% 12
7% Pe
rf o rm
an ce
a pp
ra is
al a
nd le
ga l h
ur dl
es /
lit ig
at io
ns 1
1% 5
3% 3
2% 1
1% 10
6%
Pe rf
o rm
an ce
a pp
ra is
al in
c o nt
ex ts
o th
er t
ha n
th e
U ni
te d
St at
es 1
1% 2
1% 2
1% 1
1% 6
3%
Ef fe
ct s
o f pe
rf o rm
an ce
a pp
ra is
al o
n em
pl o ye
es ’
pe rc
ep ti o ns
5 3%
5 3%
T o ta
l 72
40 %
57 31
% 40
22 %
7 4%
5 3%
18 1
68 %
Belle et al. 279
differences and similarities of performance appraisal as compared with tools such as total quality management (e.g., J. S. Bowman, 1994; Cederblom & Pemerl, 2002) and management by objectives (e.g., Brumback, 1978; Brumback & McFee, 1982). Other research included in this category recommended the adoption and use of practices that are able to identify employees’ areas of improvement and professional growth (e.g., Herbert & Doverspike, 1990; Nanry, 1988).
Accounting for about 27% of the grand total, the second major area of inquiry in terms of research theme expanded our knowledge about the relationship between per- formance appraisal systems and employees’ perceptions, motivational forces, and/or demographic characteristics. Studies in this area, for example, analyzed workers’ per- ceptions of fairness and effectiveness of the performance appraisal process (e.g., Harrington & Lee, 2015; Kim, 2002; Kim & Rubianty, 2011; H. Lee, Cayer, & Lan, 2006), employees’ job satisfaction in the presence of performance evaluation practices (e.g., Ellickson, 2002; Kim, 2002; Yang & Kassekert, 2010), and links with different work motivation forces (e.g., Christensen et al., 2013; Oh & Lewis, 2009; Park, 2014).
Studies that identified the major shortcomings of the performance appraisal sys- tems (e.g., Clifford, 1999; Grøn, 2011; Hassan & Rohrbaugh, 2009; Roberts, 1998) made up about 12% of the 181 articles included in our literature review. Similarly, work describing practices of individual performance appraisal in public sector organi- zations (e.g., Riccucci & Lurie, 2001; Thornton & Morris, 2001; Van Hein, Kramer, & Hein, 2007) constituted another 11% of the manuscripts that we analyzed in depth. The remainder of articles in our systematic literature review was composed as follows: 7% investigated the link between performance appraisal and merit pay (e.g., among the most recent work, McKinney, Mulvaney, & Grodsky, 2013), 6% unveiled potential and actual legal hurdles and litigation due to performance appraisal systems and pro- cedures (e.g., among the most recent work Daley, 2007; J. A. Lee, Havighurst, & Rassel, 2004), 3% explored performance appraisal in governments and public sector organizations outside of the United States (e.g., Lah & Perry, 2008; Liebert, 2014; Liu & Dong, 2012), and 3% tested the effects of the introduction of performance appraisal systems on employees’ perceptions through pre–post comparisons (e.g., De Leon & Ewen, 1997; Gabris & Ihrke, 2000; Moussavi & Ashbaugh, 1995).
Regarding the studies’ primary research design or approach, Table 1 shows that the analysis of survey data was the most widely adopted design: it accounted for about 40% of the grand total. One group of manuscripts in this category analyzed data from large-scale surveys, such as those administered by the Merit System Protection Board (e.g., Cho & Lewis, 2012; Daley, 2008; Kim & Rubianty, 2011; Kim & Holzer, 2016) or the Federal Employees Viewpoint Survey administered by the Office of Personnel Management (e.g., H. Lee et al., 2006; Yang & Kassekert, 2010). A second group of manuscripts in this area, instead, analyzed data collected through ad hoc surveys administered either in one organization/government (e.g., Daley, Vasu, & Weinstein, 2002; Daley, 1996; Ellickson 2002; Gabris & Ihrke, 2001) or across comparable orga- nizations/governments (e.g., Christensen et al., 2012; Frazier & Swiss, 2008; Hassan & Rohrbaugh, 2009).
280 Review of Public Personnel Administration 37(3)
About 31% of the articles that we analyzed in depth, then, adopted a normative approach. Manuscripts in this category spanned all of the primary research themes that we identified. Case studies were the third most commonly used research design and accounted for about 31% of the 181 articles that we analyzed in depth. Those included the description of performance appraisal systems, practices, or correlates in selected organizations such as enforcement agencies (e.g., Cederblom & Pemerl, 2002; Cozzetto, 1990; Glen, 1990; LaVan, Katz, & Carley, 1993), local governments (e.g., Di Mascio & Natalini, 2013; Mulvaney, McKinney, & Grodsky, 2012), state govern- ments (e.g., Cozzetto, 1991; Daley, 1983), and supranational organizations (e.g., Grøn, 2011). Finally, through our search strategy, we identified seven (about 4% of the grand total) unsystematic literature reviews on specific themes about individual performance appraisal and five studies (3%) that used a pre–posttest design.
Our systematic search for studies on performance appraisal in a selection of public administration journals, thus, reveals a gap in our understanding of whether and how cognitive biases affect the performance ratings that raters assign to ratees in govern- ment and public sector organizations. Martin and Bartol (1986) argued that, as long as organizations want their performance appraisal systems to achieve the intended goals, “it is vitally important that raters be trained to fulfill their appraisal function ade- quately” (p. 108) and that such training includes discussions of the typical rating errors and strategies to limit them. Similarly, J. S. Bowman (1999) noted that while “the use of ratings assumes that evaluators are reasonably objective and precise . . . well known errors occur in the process based on cognitive limitations” (p. 566). Both Martin and Bartol (1986) as well as J. S. Bowman (1999) discussed halo effects as a common cognitive bias in performance appraisal. Unlike in our field, abundant empirical evi- dence from disciplines such as applied psychology (e.g., Thorsteinson et al., 2008) and behavioral economics (e.g., Furnham & Boo, 2011; Kahneman, 2011) suggests that performance scores can be biased because of cognitive limitations.
As a final note, it is worth emphasizing that our search strategies returned several recent articles that we did not include in our classification in Table 1 because they dealt with organizational rather than individual performance but, nevertheless, informed our research. In fact, they investigated whether and how cognitive biases influence citizens’ evaluation of government performance. For instance, novel experimental evidence sug- gests that parents’ satisfaction with their kids’ public schools, often used as an indicator of organizational performance in the public sector, is not entirely consistent with ratio- nal models but, rather, is systematically biased (e.g., Andersen & Hjortskov, 2015). Similarly, Barrows, Henderson, Peterson, and West (2016) conducted a survey experi- ment and found that citizens’ assessment of the quality of local services depends on the information about how they rank compared with other services. Further randomized trials show that citizens’ evaluation of government performance is influenced by uncon- scious and implicit negative associations between public administration and good per- formance (e.g., Marvel, 2015, 2016). Overall, thus, our work shares with this nascent and very promising research area the curiosity for detecting errors in performance appraisal that are systematic rather than random. Unlike these works, however, our focus is on raters’ biases in scoring subordinates’ job achievements.
Belle et al. 281
Cognitive Biases in Performance Appraisal
Anchoring Effect
Anchoring is the cognitive tendency to estimate unknown quantities by making adjust- ments from an initial value. Experimental evidence has consistently shown that “dif- ferent starting points yield different estimates, which are biased toward the initial values” (Tversky & Kahneman, 1974, p. 1128). Thus, according to the anchoring effect, people asked to estimate an unknown quantity and to consider a certain number for their estimation tend to yield a final answer that is an insufficient adjustment of the initial value (Kahneman, 2011; Tversky & Kahneman, 1974). Classical studies dem- onstrate that incomplete computations, numbers generated randomly in the presence of the decision maker, and numbers provided with a priming role function as anchors and lead to anchoring effects.
In early experimental tests of the anchoring effect, high school students provided a greater median answer when asked to compute 8 × 7 × 6 × 5 × 4 × 3 × 2 × 1 rather than 1 × 2 × 3 × 4 × 5 × 6 × 7 × 8: 2,250 and 520, respectively. Admittedly, these two opera- tions are identical and have the same result, which, by the way, is 40,320 (Tversky & Kahneman, 1974). In another classical experiment, the median estimate of the percent- age of African Nations of the United Nations was higher when students were exposed to the spin of a wheel of fortune that landed on the number 65 as compared with the number 10. The wheel of fortune could not possibly provide any useful information about anything. Before providing their answers, participants were asked to think whether their estimate was going to be higher or lower compared with the reference value that was presented to them (Tversky & Kahneman, 1974). In short, “anchors that are obviously random can be just as effective as potentially informative anchors” (Kahneman, 2011, p. 125).
Extant experimental research uncovers anchoring effects in domains such as gen- eral knowledge (e.g., Simmons, LeBoeuf, & Nelson, 2010), legal judgments made by professionals with different degrees of experience and expertise (e.g., Bennett, 2014), and negotiations (e.g., Orr & Guthrie, 2006). Recent research on anchoring also has been conducted in economic valuations on the field (e.g., Alevy, Landry, & List, 2015). Furnham and Boo (2011) provided a literature review on the domains in which anchor- ing has been explored, explanations that have been provided for the effect, and indi- viduals’ factors that may contribute to the bias.
Abundant work finds systematic biases in performance ratings due to anchors pro- vided to the raters. The presence of anchoring effects in performance appraisals seems to hold across types of raters, rates, and anchors. For instance, Thorsteinson et al. (2008) showed that the overall performance rating that students assigned to a fictional sales representative and to their instructor was consistently higher for the group exposed to a high anchor then for the groups exposed to either a low anchor or no anchor. Results held across experimental manipulations of the anchors and across applicability of the anchor to the final judgment. Similarly, Chen and Kemp (2015) found evidence of anchoring effects in the average rating score assigned to a fictional
282 Review of Public Personnel Administration 37(3)
lecturer applying for promotion across raters (students and faculty members at differ- ent stages in their professional careers), types of role (department head or independent evaluator), types of anchors (provided by the experimenter or generated by the sub- jects), and types of promotion (good or excellent). Experimental work has also identi- fied prior performance scores as anchors that have an impact on raters’ scores for ratees’ subsequent performance (e.g., Kravitz & Balzer, 1992; Murphy, Balzer, Lockhart, & Eisenman, 1985; Smither, Reilly, & Buda, 1988; Sumer & Knight, 1996).
Halo Effect
Extant scholarship locates early evidence of the halo effect in the work of Thorndike (1920). Required to consider independently four performance dimensions and provide a score for each, superiors in the army assigned ratings to officers whose “correlations are too high and too even” (Thorndike, 1920, p. 27). A similar pattern of findings emerged from the analysis of the correlations between rating of general abilities and rating of specific abilities of army officers and teachers evaluated by their superiors: “a halo of general merit is extended to influence the rating for the special ability, or vice versa” (Thorndike, 1920, p. 27). About a century later, Kahneman (2011) described the halo effect as “the tendency to like (or dislike) everything about a per- son—including things you have not observed” (p. 81) and explained that we tend to exaggerate the consistency of judgments to maintain simple and coherent explanatory narratives.
Meanwhile, literature on the halo effect has developed in both normative and empirical terms. From the normative perspective, extant scholarship provides several definitions and operational measures of the halo effect (e.g., Balzer & Sulsky, 1992; Murphy, Jako, & Anhalt, 1993). One element tends to be consistent across definitions: the bias toward making ratings consistent across different dimensions, regardless of, or even contrary to, available information. Accordingly, in public administration research (e.g., Battaglio, 2015; J. S. Bowman, 1999; Martin & Bartol, 1986), the halo effect has been described as occurring “when raters allow a rating on one criterion (e.g., arriving late to work) to influence ratings on subsequent criteria (e.g., processing paperwork, communicating with customer)” (Battaglio, 2015, p. 204).
From the empirical perspective, there is evidence that raters’ general impressions of ratees influence the ratings of specific competencies (e.g., Lance, LaPointe, & Stewart, 1994) and that scores on one dimension spread to a second dimension (e.g., Bechger et al., 2010; Cooper, 1981; Dennis, 2007; Nisbett & Wilson, 1977; Solomonson & Lance, 1997; Thorndike, 1920). Halo effects in performance appraisal have been found in settings as different as post commanders and sergeants judging the perfor- mance of their nonprobationary troopers in a state police agency (e.g., King, Hunter, & Schmidt, 1980); workers of a manufacturing company assessing subordinates, themselves, and peers (e.g., Holzbach, 1978); and students’ evaluating faculty mem- bers (e.g., Jacobs & Kozlowski, 1985).
In essence, raters affected by halo error tend to transfer their impression (general or specific) of each ratee from one domain to another by providing consistently high (or
Belle et al. 283
low, or average) ratings across performance dimensions, even though, in fact, ratees are likely to exhibit significant relative strengths and weaknesses on different perfor- mance dimensions (Borman, 1975).
Method
Experiment 1: Participants, Design, and Measures
Experiment 1 tested the anchoring effect. Participants in Experiment 1 were 600 Italian public sector managers and employees. They were recruited in June 2016 by the Qualtrics Software Company that also administered the survey experiment online. Of the respondents, 64% were managers (i.e., have formal responsibilities to manage people in their organizations) whereas 36% were public sector employees without managerial responsibilities. About 56% were female. About 16% of the respondents were employed in the health care industry, about 49% in education, about 21% in administration, and the remaining 14% in other public sector industries. About 38% of participants held a scientific degree, 38% a humanistic degree, and the remaining 24% did not have a degree. Average age was 44 years (SD = 10 years).
Building on Tversky and Kahneman (1974) and Kahneman (2011), we designed two scenarios that were identical in all regards except for our manipulation of the anchor. Public managers and employees in our study were randomly assigned to respond to only one of the two scenarios. Both scenarios asked participants to imagine themselves as the supervisor of the employee described in the vignette. The scenarios informed that the subordinate met the majority of goals, had good interpersonal skills, and showed moderate creativity. Both scenarios, then, reported the performance score that respondents (in the rater’ shoes) assigned to that subordinate the previous year.
Participants in the low anchor group (n = 291) read that last year they assigned to that ratee a performance rating of 51/100 and were asked to decide whether they were going to assign a rating lower or higher than 51/100 this year.
Participants in the high anchor group (n = 309) read that last year they assigned to that ratee a performance rating of 91/100 and were asked to decide whether they were going to assign a rating lower or higher than 91/100 this year.
All participants, then, indicated the performance score that they wanted to assign to their subordinate for this year on a 0-100 scale. This performance score was our depen- dent variable.
To guard against omitted variables bias, we followed recent experimental research in public administration using samples of real civil servants (e.g., Belle, 2013, 2014, 2015) and measured two sets of control variables. On one side, we measured the fol- lowing demographic variables: managerial status, gender, industry of employment, educational background, and age. On the other side, we used Donnellan, Oswald, Baird, and Lucas (2006) scale to observe personality traits: extraversion, agreeable- ness, conscientiousness, neuroticism, and intellect/imagination. Items were measured on a 1 (strongly disagree) to 7 (strongly agree) Likert-type scale. Cronbach’s alphas were .75 for extraversion, .69 for agreeableness, .64 for conscientiousness, .63 for
284 Review of Public Personnel Administration 37(3)
neuroticism, and .68 for intellect/imagination. Given that they were not the primary focus of our research, we took an exploratory approach in analyzing eventual interac- tions of demographic variables and personality traits with our manipulated factor.
Experiment 2: Participants, Design, and Measures
Participants in Experiment 2 were the same as participants in Experiment 1. Building on Thorndike (1920) and Kahneman (2011), we endeavored to understand halo effects between dimensions of performance. As before, we designed two scenarios that were identical in all regards except for our experimental manipulation. As in Experiment 1, participants in Experiment 2 were exposed to only one randomly selected scenario. In both versions, the scenarios asked participants to imagine being a rater who had to evaluate a ratee along two dimensions: accuracy in carrying out job duties and inter- personal skills. The former dimension was our independent variable, which we manip- ulated at two levels: low versus high.
Participants in the low accuracy group (n = 287) learned that their ratee’s error rates in processing document had always been among the highest within the organization.
Participants in the high accuracy group (n = 313), instead, read that their ratee’s error rate in processing documents had always been among the lowest in the organization.
Both groups, then, read that the ratees had average interpersonal skills. All study participants were lastly asked to indicate on a 0-100 points scale their evaluation of the ratee along the two performance dimensions: job accuracy and interpersonal skills. The former measure served as our manipulation check, which we included to make sure that our experimental treatment produced the intended effect. The latter, that is, the score on interpersonal skills, was our dependent variable.
The control variables in Experiment 2 were the same as the control variables in Experiment 1.
Results
Experiment 1
Chi-square tests indicated that the participants who were randomly assigned to the low anchor group, compared with the participants who were assigned to the high anchor group, did not differ at the .05 significance level in terms of proportion of managers, females, and distribution of degrees. The two groups differed marginally in terms of frequencies of fields of employment within the public sector (Fisher’s exact test p value = .054). A two-sample mean comparison test showed that average age did not differ at the .05 level between the two groups. Therefore, the random assignment of participants to treatments based on demographic variables worked correctly.
Figure 1 reports the mean performance score that participants assigned to the ratee in the two treatment groups. On average, raters’ assessment of the ratee’s performance was higher in the high anchor group (M = 88.47, SD = 15.63) than in the low anchor
Belle et al. 285
group (M = 71.07, SD = 13.01), p < .001. This is in line with our theoretical expecta- tions about the presence of the anchoring effect in performance appraisal in the public sector. Indeed, our findings suggest that civil servants’ judgment of ratees’ perfor- mance tends to be systematically biased toward anchors that are provided to them before they choose their final performance score.
This result held true across our demographics variables and personality traits. A series of ANOVAs showed no interaction between the experimental treatment of the anchor and our observed control variables.
Experiment 2
The percentage of managers, proportion of females, and distributions of respon- dents by industry of employment and degrees were not statistically different in the low accuracy group and high accuracy group. A test of the equality of means between the two samples showed that average age was marginally different in the low and high accuracy groups (M = 43.15, SD = 10.21; M = 45.29, SD = 10.10, respectively, p = .10).
Figure 1. Mean performance rating that raters assigned to the ratee this year by anchor (Experiment 1). Note. CI = confidence intervals.
286 Review of Public Personnel Administration 37(3)
A two-sample mean comparison test showed that the average rating assigned by the raters to the ratee’s accuracy was lower in the low accuracy group (M = 39.04, SD = 29.95) as compared with the high accuracy group (M = 84.25, SD = 17.32), p < .001. Therefore, our experimental manipulation was effective in producing the intended effects.
Figure 2 reports the mean rating that participants assigned to the ratee’s interper- sonal skills, by ratee’s accuracy on the job. On average, raters in the low accuracy group scored ratee’s interpersonal skills lower (M = 60.30, SD = 22.40) than raters in the high accuracy group did (M = 67.74, SD = 19.46), p < .001. As expected, our find- ings suggest that public sector workers can be systematically prone to halo effects in assessing different dimensions of ratees’ performance. Indeed, the mean score assigned to interpersonal abilities was contingent upon the manipulation of job accuracy skills.
A series of ANOVAs suggested that these findings were robust across our control variables except for the gender of the rater. The results of a 2 (accuracy: low, high) × 2 (raters’ gender: female, male) ANOVA showed a significant interaction, F(1, 599) = 5.35, p = .021 (Table 2). Findings from an ordinary least square regression revealed a positive interaction between high accuracy and being female. In other words, the effect
Figure 2. Mean rating that raters assigned to ratee’s interpersonal skills by ratee’s accuracy on the job (Experiment 2). Note. CI = confidence intervals.
Belle et al. 287
of the ratee’s high accuracy score on the score assigned to the ratee’s interpersonal skills was larger when raters were female.
Figure 3 shows the interaction between our experimental manipulation and respon- dents’ gender graphically, using estimated marginal means. In the low accuracy group, female raters scored ratee’s interpersonal skills lower (M = 57.57, SD = 22.98) than male raters did (M = 63.83, SD = 21.19), p = .019. To the contrary, ratings of interper- sonal skills did not differ at the .05 significance level between female and male raters in the high accuracy group (p = .452). While female raters in the low anchor group provided a lower score to ratee’s interpersonal skills than female raters in the high anchor group did (M = 68.45, SD = 20.17), p < .001, male raters did not, p = .219. In short, our findings reveal that only female public sector managers and employees were prone to halo effects in assessing different performance dimensions of a fictional subordinate.
Discussion and Conclusion
This article synthesizes existing scholarship on performance appraisal in public administration and illuminates our understanding of the effects of cognitive biases on performance scores in public sector settings. Our systematic literature review on a selection of public administration journals showed that about one third of the articles addressed questions related to the overall design of performance appraisal systems in government and public organizations. An additional one third of the articles explored the associations between performance evaluation systems and public employees’ per- ceptions, work motivation drivers, and demographic characteristics. As far as the research method or approach is concerned, instead, the greatest majority of manu- scripts (about 40%) analyzed data collected through surveys. Those were followed by 31% of articles that adopted normative lenses to argue, for instance, how performance
Table 2. Two-Way ANOVA for the Interaction of Ratee’s Accuracy With Rater’s Gender on the Score Assigned to Ratee’s Interpersonal Skills (Experiment 2).
Source Partial SS df MS F p
Model 11,282.64 3 3,760.88 8.66 .000 Accuracy 7,168.21 1 7,168.21 16.51 .000 Female 781.20 1 781.20 1.80 .180 Accuracy × Female 2,323.76 1 2,323.76 5.35 .021 Residual 258,689.92 596 434.04 Total 269,972.56 599 450.71 Root MSE = 20.83 N = 600 R2 = .0418 Adjusted R2 = .0370
Note. SS = sum of squares; MS = mean square; MSE = mean square error.
288 Review of Public Personnel Administration 37(3)
appraisal should be designed and what the likely shortcomings were, and 22% of arti- cles that presented insights from case studies. Our systematic literature review unveiled a lack of empirical studies on the cognitive limitations of raters that can systematically bias their assessment of ratees’ performance.
Results of Experiments 1 and 2 revealed that public employees required to rate the work achievements of a fictional subordinate were biased by both anchoring and halo effects. In Experiment 1, the average overall performance score assigned to a subordi- nate was higher when raters were exposed to a high anchor related to the previous year’s performance than when subjects were exposed to a low anchor. In Experiment 2, the average score on one performance dimension (interpersonal skills) was higher when raters were led to assign a higher score to the subordinate on another perfor- mance dimension (accuracy). Interestingly, our results found evidence of the halo effect only for female respondents.
The first contribution of our work lies in providing a synthesis of the scholarship on performance appraisal in governmental organizations and classifying studies along
Figure 3. Two-way interaction of ratee’s accuracy with rater’s gender on the score assigned to ratee’s interpersonal skills (Experiment 2). aEstimated marginal means.
Belle et al. 289
two dimensions: primary research topic and primary research method/approach. We hope this can facilitate the work of scholars and practitioners who are interested in advancing knowledge and practice that take stock of existing evidence.
Our study further contributes to extant scholarship on performance appraisal in public administration by providing the first experimental test of anchoring and halo effects on public sector managers and employees. In joining recent efforts to explore cognitive biases in public administration and management (e.g., Andersen & Hjortskov, 2015; Marvel, 2015, 2016), we broaden the substantive scope of available studies. Furthermore, we complement evidence of cognitive biases in performance appraisal available in other disciplines, like applied psychology and behavioral economics.
The findings of our empirical work should be interpreted in light of the limitations that are typical of artefactual field experiments. More precisely, our design seems to be well equipped to meet internal validity requirements and, therefore, establish a causal link between our experimental manipulations and the effects on average performance ratings. However, our design may be prone to external validity concerns. In other words, we cannot be certain that our results would replicate under more natural set- tings and circumstances. For this reason, we encourage scholars in our field to test our experimental manipulation with different samples of public sector managers in differ- ent contexts.
Also, our study may fall short in terms of construct validity. In particular, we are unable to conclude that we would have observed the same anchoring and halo effects had we used different treatment levels and different measures of the performance scores. Scholars in our field can address these limitations by varying such elements and capitalizing on the experimental designs adopted in other disciplines.
Finally, we call for observational research on anchoring and halo effects in perfor- mance evaluation in public administration. This would serve a double purpose. On one side, it would triangulate findings with existing experimental work. On the other side, it would expand our knowledge of the micromechanisms behind the cognitive biases that systematically affect human judgment and decision making on the job.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
References
A complete list of the references included in the systematic literature review is available upon request to the authors.
Alevy, J. E., Landry, C. E., & List, J. A. (2015). Field experiments on the anchoring of eco- nomic valuations. Economic Inquiry, 53, 1522-1538.
290 Review of Public Personnel Administration 37(3)
Andersen, S. C., & Hjortskov, M. (2015). Cognitive biases in performance evaluations. Journal of Public Administration Research and Theory, 26, 647-662.
Balzer, W. K., & Sulsky, L. M. (1992). Halo and performance appraisal research: A critical examination. Journal of Applied Psychology, 77, 975-985.
Barrows, S., Henderson, M., Peterson, P. E., & West, M. R. (2016). Relative performance infor- mation and perceptions of public service quality: Evidence from American school districts. Journal of Public Administration Research and Theory, 26, 571-583.
Battaglio, P. R., Jr. (2015.). Public Human Resource Management: Strategies and practices in the 21st century. Thousand Oaks, CA: CQ Press.
Bechger, T. M., Maris, G., & Hsiao, Y. P. (2010). Detecting halo effects in performance-based examinations. Applied Psychological Measurement, 34, 607-619.
Belle, N. (2013). Experimental evidence on the relationship between public service motivation and job performance. Public Administration Review, 73, 143-153.
Belle, N. (2014). Leading to make a difference: A field experiment on the performance effects of transformational leadership, perceived social impact, and public service motivation. Journal of Public Administration Research and Theory, 24, 109-136.
Belle, N. (2015). Performance-related pay and the crowding out of motivation in the public sec- tor: A randomized field experiment. Public Administration Review, 75, 230-241.
Bennett, M. W. (2014). Confronting cognitive “anchoring effect” and “blind spot” biases in federal sentencing: A modest solution for reforming a fundamental flaw. The Journal of Criminal Law and Criminology, 104, 489-534.
Bissessar, A. M. (2000). The introduction of new appraisals systems in the public services of the Commonwealth Caribbean. Public Personnel Management, 29, 277-292.
Borman, W. C. (1975). Effects of instructions to avoid halo error on reliability and validity of performance evaluation ratings. Journal of applied psychology, 60, 556-560.
Bowman, J. S. (1994). At last, an alternative to performance appraisal: Total quality manage- ment. Public Administration Review, 54, 129-136.
Bowman, J. S. (1999). Performance appraisal: Verisimilitude trumps veracity. Public Personnel Management, 28, 557-576.
Bowman, R., & Grimmelikhuijsen, S. (2016). Experimental public administration from 1992 to 2014: A systematic literature review and ways forward. International Journal of Public Sector Management, 29, 110-131.
Brown, R. W. (1982). Performance appraisal: A policy implementation analysis. Review of Public Personnel Administration, 2, 69-85.
Brumback, G. B. (1978). Toward a new theory and system of performance evaluation-a standard- ized, mbo (management by objectives)-oriented approach. Public Personnel Management, 7, 205-211.
Brumback, G. B., & McFee, T. S. (1982). From MBO to MBR. Public Administration Review, 42, 363-371.
Cantarelli, P., Belardinelli, P., & Belle, N. (2016). A Meta-analysis of job satisfaction correlates in the public administration literature. Review of Public Personnel Administration, 36, 115- 144.
Cederblom, D., & Pemerl, D. E. (2002). From performance appraisal to performance manage- ment: One agency’s experience. Public Personnel Management, 31, 131-140.
Chen, Z., & Kemp, S. (2015). Anchoring effects in simulated academic promotion decisions: How the promotion criterion affects ratings and the decision to support an application. Journal of Behavioral Decision Making, 28, 137-148.
Belle et al. 291
Cho, Y. J., & Lewis, G. B. (2012). Turnover intention and turnover behavior implications for retaining federal employees. Review of Public Personnel Administration, 32, 4-23.
Christensen, R. K., Whiting, S. W., Im, T., Rho, E., Stritch, J. M., & Park, J. (2013). Public service motivation, task, and non-task behavior: A performance appraisal experiment with Korean MPA and MBA students. International Public Management Journal, 16, 28-52.
Christensen, T., Dong, L., Painter, M., & Walker, R. M. (2012). Imitating the West? Evidence on administrative reform from the upper echelons of Chinese provincial government. Public Administration Review, 72, 798-806.
Clifford, J. P. (1999). The collective wisdom of the workforce: Conversations with employees regarding performance evaluation. Public Personnel Management, 28, 119-155.
Cooper, W. H. (1981). Conceptual similarity as a source of illusory halo in job performance ratings. Journal of Applied Psychology, 66, 302-307.
Cozzetto, D. (1990). The officer fitness report as a performance appraisal tool. Public Personnel Management, 19, 235-244.
Cozzetto, D. (1991). Public sector grievances: The case of North Dakota. Review of Public Personnel Administration, 12, 5-13.
Daley, D., Vasu, M. L., & Weinstein, M. B. (2002). Strategic human resource management: Perceptions among North Carolina county social service professionals. Public Personnel Management, 31, 359-375.
Daley, D. M. (1983). Performance appraisal as a guide for training and development: A research note on the Iowa performance evaluation system. Public Personnel Management, 12, 159- 166.
Daley, D. M. (1996). Humanistic management and organizational success: The effect of job and work environment characteristics on organizational effectiveness, public responsiveness, and job satisfaction. Public Personnel Management, 15, 131-142.
Daley, D. M. (2007). If a tree falls in the forest: The effect of grievances on employee percep- tions of performance appraisal, efficacy, and job satisfaction. Review of Public Personnel Administration, 27, 281-296.
Daley, D. M. (2008). The burden of dealing with poor performers: Wear and tear on supervisory organizational engagement. Review of Public Personnel Administration, 28, 44-59.
De Leon, L., & Ewen, A. J. (1997). Multi-source performance appraisals: Employee perceptions of fairness. Review of Public Personnel Administration, 17, 22-36.
Dennis, I. (2007). Halo effects in grading student projects. Journal of Applied Psychology, 92, 1169-1176.
Di Mascio, F., & Natalini, A. (2013). Context and mechanisms in administrative reform pro- cesses: Performance management within Italian local government. International Public Management Journal, 16, 141-166.
Donnellan, M. B., Oswald, F. L., Baird, B. M., & Lucas, R. E. (2006). The Mini-IPIP scales: Tiny-yet-effective measures of the Big Five factors of personality. Psychological Assessment, 18, 192-203.
Ellickson, M. C. (2002). Determinants of job satisfaction of municipal government employees. Public Personnel Management, 31, 343-358.
Feild, H. S., & Holley, W. H. (1975). Traits in performance ratings: Their importance in public employment. Public Personnel Management, 9, 327-330.
Frazier, M. A., & Swiss, J. E. (2008). Contrasting views of results-based management tools from different organizational levels. International Public Management Journal, 11, 214-234.
292 Review of Public Personnel Administration 37(3)
Furnham, A., & Boo, H. C. (2011). A literature review of the anchoring effect. The Journal of Socio-Economics, 40, 35-42.
Gabris, G. T., & Ihrke, D. M. (2000). Improving employee acceptance toward performance appraisal and merit pay systems: The role of leadership credibility. Review of Public Personnel Administration, 20, 41-53.
Gabris, G. T., & Ihrke, D. M. (2001). Does performance appraisal contribute to heightened levels of employee burnout? The results of one study. Public Personnel Management, 30, 157-172.
Glen, R. M. (1990). Performance appraisal: An unnerving yet useful process. Public Personnel Management, 19, 1-10.
Grøn, C. H. (2011). Having it both ways? Balancing knights and knaves when motivat- ing European commission officials. Review of Public Personnel Administration, 31, 369-385.
Grimmelikhuijsen, S. G., Jilke, S., Olsen, A. L., & Tummers, L. G. (2017). Behavioral pub- lic administration: Combining insights from public administration and psychology. Public Administration Review, 77, 45-56.
Harrington, J. R., & Lee, J. H. (2015). What drives perceived fairness of performance appraisal? Exploring the effects of psychological contract fulfillment on employees’ perceived fair- ness of performance appraisal in U.S. federal agencies. Public Personnel Management, 44, 214-238.
Hassan, S., & Rohrbaugh, J. (2009). Incongruity in 360-degree feedback ratings and com- peting managerial values: Evidence from a public agency setting. International Public Management Journal, 12, 421-449.
Herbert, G. R., & Doverspike, D. (1990). Performance appraisal in the training needs analysis process: A review and critique. Public Personnel Management, 19, 253-270.
Holzbach, R. L. (1978). Rater bias in performance ratings: Superior, self-, and peer ratings. Journal of Applied Psychology, 63, 579-588.
Hyde, A. C. (1982). Performance appraisal in the post reform era. Public Personnel Management, 11, 294-305.
Institute for Scientific Information. (2015). Journal citation reports (Science edition). ©Thomson Reuters. Retrieved from http://thomsonreuters.com/en/products-services/schol- arly-scientific-research/research-management-and-evaluation/journal-citation-reports.html
Jacobs, R., & Kozlowski, S. W. (1985). A closer look at halo error in performance ratings. Academy of Management Journal, 28, 201-212.
Kahneman, D. (2011). Thinking, fast and slow. New York City, NY: Macmillan. Kim, S. (2002). Organizational support of career development and job satisfaction: A case study
of the Nevada Operations Office of the Department of Energy. Review of Public Personnel Administration, 22, 276-294.
Kim, S. E., & Rubianty, D. (2011). Perceived fairness of performance appraisals in the federal government does it matter? Review of Public Personnel Administration, 31, 329-348.
Kim, T., & Holzer, M. (2016). Public employees and performance appraisal: A study of anteced- ents to employees’ perception of the process. Review of Public Personnel Administration, 36, 31-56.
King, L. M., Hunter, J. E., & Schmidt, F. L. (1980). Halo in a multidimensional forced-choice performance evaluation scale. Journal of Applied Psychology, 65, 507-516.
Kravitz, D., & Balzer, W. (1992). Context effects in performance appraisal: A methodological critique and empirical study. Journal of Applied Psychology, 77, 24-31.
Belle et al. 293
Lah, T. J., & Perry, J. L. (2008). The diffusion of the civil service reform act of 1978 in OECD countries: A tale of two paths to reform. Review of Public Personnel Administration, 28, 282-299.
Lance, C. E., LaPointe, J. A., & Stewart, A. M. (1994). A test of the context dependency of three causal models of halo rater error. Journal of Applied Psychology, 79, 332-340.
LaVan, H., Katz, M., & Carley, C. (1993). The arbitration of grievances of police officers and fire fighters. Public Personnel Management, 22, 433-444.
Lee, H., Cayer, N. J., & Lan, G. Z. (2006). Changing federal government employee attitudes since the Civil Service Reform Act of 1978. Review of Public Personnel Administration, 26, 21-51.
Lee, J. A., Havighurst, L. C., & Rassel, G. (2004). Factors related to court references to perfor- mance appraisal fairness and validity. Public Personnel Management, 33, 61-77.
Liebert, S. (2014). Challenges of reforming the civil service in the post-Soviet era: The case of Kyrgyzstan. Review of Public Personnel Administration, 34, 403-420.
Liu, X., & Dong, K. (2012). Development of the civil servants’ performance appraisal system in China: Challenge and improvement. Review of Public Personnel Administration, 32, 149- 168. doi:10.1177/073437112438244
Martin, D. C., & Bartol, K. M. (1986). Training the raters: A key to effective performance appraisal. Public Personnel Management, 15, 101-109.
Marvel, J. D. (2015). Public opinion and public sector performance: Are individuals’ beliefs about performance evidence-based or the product of anti–public sector bias? International Public Management Journal, 18, 209-227.
Marvel, J. D. (2016). Unconscious bias in citizens’ evaluations of public sector performance. Journal of Public Administration Research and Theory, 26, 143-158.
McKinney, W. R., Mulvaney, M. A., & Grodsky, R. (2013). The development of a model for the distribution of merit pay increase monies for municipal agencies: A case study. Public Personnel Management, 42, 471-492.
Moussavi, F., & Ashbaugh, D. L. (1995). Perceptual effects of participative, goal-oriented per- formance appraisal: A field study in public agencies. Journal of Public Administration Research and Theory, 5, 331-343.
Mulvaney, M. A., McKinney, W. R., & Grodsky, R. (2012). The development of a pay-for- performance appraisal system for municipal agencies: A case study. Public Personnel Management, 41, 505-533.
Murphy, K. R., Balzer, W. K., Lockhart, M. C., & Eisenman, E. J. (1985). Effects of previous performance on evaluations of present performance. Journal of Applied Psychology, 70, 72-84.
Murphy, K. R., Jako, R. A., & Anhalt, R. L. (1993). Nature and consequences of halo error: A critical analysis. Journal of Applied Psychology, 78, 218-225.
Nanry, C. (1988). Performance linked training. Public Personnel Management, 17, 457-463. Nisbett, R. E., & Wilson, T. D. (1977). The halo effect: Evidence for unconscious alteration of
judgments. Journal of Personality and Social Psychology, 35, 250-256. Oh, S. S., & Lewis, G. B. (2009). Can performance appraisal systems inspire intrinsically moti-
vated employees? Review of Public Personnel Administration, 29, 158-167. Organization for Economic Cooperation and Development. (2012). Rewarding performance in
the public sector and performance related pay in OECD countries. Paris, France: Author. Orr, D., & Guthrie, C. (2006). Anchoring, information, expertise, and negotiation: New insights
from meta-analysis. Ohio State Journal on Dispute Resolution, 21, 597-628.
294 Review of Public Personnel Administration 37(3)
Park, S. (2014). Motivation of public managers as raters in performance appraisal: Developing a model of rater motivation. Public Personnel Management, 43, 387-414.
Riccucci, N. M., & Lurie, I. (2001). Employee performance evaluation in social welfare offices. Review of Public Personnel Administration, 21, 27-37.
Roberts, G. E. (1998). Perspectives on enduring and emerging issues in performance appraisal. Public Personnel Management, 27, 301-320.
Simmons, J. P., LeBoeuf, R. A., & Nelson, L. D. (2010). The effect of accuracy motivation on anchoring and adjustment: Do people adjust from provided anchors? Journal of Personality and Social Psychology, 99, 917-932.
Smither, J. W., Reilly, R. R., & Buda, R. (1988). Effect of prior performance information on ratings of present performance: Contrast versus assimilation revisited. Journal of Applied Psychology, 73, 487-496.
Solomonson, A. L., & Lance, C. E. (1997). Examination of the relationship between true halo and halo error in performance ratings. Journal of Applied Psychology, 82, 665-674.
Sumer, H. C., & Knight, P. A. (1996). Assimilation and contrast effects in performance ratings: Effects of rating the previous performance on rating subsequent performance. Journal of Applied Psychology, 81, 436-442.
Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4, 25-29.
Thornton, G. C., & Morris, D. M. (2001). The application of assessment center technology to the evaluation of personnel records. Public Personnel Management, 30, 55-66.
Thorsteinson, T. J., Breier, J., Atwell, A., Hamilton, C., & Privette, M. (2008). Anchoring effects on performance judgments. Organizational Behavior and Human Decision Processes, 107, 29-40.
Tummers, L. L., Bekkers, V., Vink, E., & Musheno, M. (2015). Coping during public service delivery: A conceptualization and systematic review of the literature. Journal of Public Administration Research and Theory, 25, 1099-1126.
Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185, 1124-1130.
Van Hein, J. L., Kramer, J. J., & Hein, M. (2007). The validity of the Reid Report for selection of corrections staff. Public Personnel Management, 36, 269-280.
Yang, K., & Kassekert, A. (2010). Linking management reform with employee job satisfaction: Evidence from federal agencies. Journal of Public Administration Research and Theory, 20, 413-436.
Author Biographies
Nicola Belle is an assistant professor at the Scuola Superiore Sant’Anna (Pisa, Italy). His research focuses on behavioral management.
Paola Cantarelli is a postdoctoral scholar at Bocconi University (Milan, Italy). She received her PhD in Public Affairs from the University of Texas at Dallas. Her current research focuses on work motivation and managerial decision making under uncertainty.
Paolo Belardinelli is a PhD candidate at Bocconi University (Milan, Italy). His research focuses on behavioral public administration.