Article Component Papers (Leadership Measurement)

profileDeonpalmgrove
Fleenoretal.2010.LQ.21.1005-1034.pdf

The Leadership Quarterly 21 (2010) 1005–1034

Contents lists available at ScienceDirect

The Leadership Quarterly

journal homepage: www.elsevier.com/locate/leaqua

Self–other rating agreement in leadership: A review

John W. Fleenor a,⁎, James W. Smither b, Leanne E. Atwater c, Phillip W. Braddy a, Rachel E. Sturm c

a Center for Creative Leadership, Greensboro, NC, United States b School of Business, La Salle University, United States c Department of Management, University of Houston, United States

a r t i c l e i n f o

⁎ Corresponding author. Center for Creative Leaders E-mail address: [email protected] (J.W. Fleenor).

1048-9843/$ – see front matter © 2010 Elsevier Inc. doi:10.1016/j.leaqua.2010.10.006

a b s t r a c t

Keywords:

This paper reviews the theoretical and empirical literature on self–other rating agreement (SOA) related to leadership in the workplace, focusing primarily on research published between 1997 (the year of Atwater & Yammarino's seminal paper on SOA) and the present. Much of the current interest in SOA derives from its purported relationships with self-awareness and leader effectiveness. The literature, however, has used a variety of metrics to assess SOA, resulting in discrepancies between findings across studies. As multi-rater (360-degree; multisource) feedback instruments continue to be widely used as a measure of leadership in organizations, it is important that we more clearly understand the relationships between SOA and its predictors and outcomes. To this end, in this article, we review (a) models of agreement, (b) factors affecting self-ratings and the congruence between self–others' ratings, (c) factors affecting others' ratings, (d) correlates of agreement, and (e) measurement issues and data analytic techniques. We conclude with discussions of practitioner issues and directions for future research.

© 2010 Elsevier Inc. All rights reserved.

Self–other rating agreement Leadership Literature review Rating congruence 360-degree feedback Multi-rater feedback Multisource feedback

In leadership research, self–other rating agreement (SOA) is typically defined as the degree of agreement or congruence between a leader's self-ratings and the ratings of others, usually coworkers such as superiors, peers, and subordinates (Atwater, Wang, Smither, & Fleenor, 2009; Fleenor, McCauley, & Brutus, 1996; Ostroff, Atwater, & Feinberg, 2004; Yammarino & Atwater, 1993). These ratings are usually gathered using multisource (360-degree) feedback instruments, which solicit ratings from various sources regarding the focal leader's effectiveness (cf. Bracken, Timmreck, & Church, 2001; Fleenor, Taylor, & Chappelow, 2008). Traditionally, differences between rating sources have been thought of as measurement error that should be reduced or eliminated; however, with multisource ratings, the lack of agreement between different perspectives is itself of major interest. These differences have been characterized as useful information for both research and applied purposes (Tornow, 1993). Much of the current interest in SOA research derives from two primary factors: (a) it is posited to be an indicator of self-awareness, and (b) it appears to be related to several outcomes of interest, including leader effectiveness and derailment.

According to Tsui and Ohlott (1988), individuals do not appear to be good judges of how they are seen or rated by bosses, direct reports, or peers, regardless of whether the ratings are based on skills, leadership behaviors, performance, or personality. While a number of studies have suggested that individuals inflate their own ratings relative to their raters (e.g., Mabe & West, 1982; Podsakoff & Organ, 1986), other research indicates that some individuals provide self-ratings that agree with others' ratings, and some self-raters provide ratings that are lower than others' ratings (Atwater & Yammarino, 1992; Fleenor et al., 1996; Van Velsor, Taylor, & Leslie, 1993).

Although it appears that the use of self-ratings of leadership alone is problematic (Harris & Schaubroeck, 1988), ratings provided by others (e.g., bosses, direct reports, and peers) should not necessarily be considered as the “true scores” of leader effectiveness (Atwater & Yammarino, 1997). According to Dunnette (1993), self-ratings have accurate components; therefore,

hip, One Leadership Place, Greensboro, NC 27410, United States.

All rights reserved.

1006 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

others' ratings should not always be presumed to be more accurate. Moreover, when compared to the ratings of others, self-ratings may provide an indication of a leader's level of self-awareness. Specifically, leaders whose self-ratings agree with the ratings of others may be seen as more self-aware than leaders whose self-ratings are not congruent with others' ratings. The ratings of others have important implications for leaders regardless of their perceived accuracy—leaders must take others' ratings seriously even if they disagree with them. In SOA research, it is the level of congruence between self-ratings and others' ratings that is of primary concern.

In this review, we will not address SOA in domains such as mentoring, vocational interests, education (e.g., student groups), clinical psychology, psychiatry, health care, or aggressive behavior. While there is a large literature in these areas (e.g., De Los Reyes, & Kazdin, 2004; Szarota, Zawadzki, & Strelau, 2002), we will confine this review to SOA as it pertains to leader and managerial performance in the workplace. This review will focus primarily on theoretical and empirical literature published between 1997 and the present. We selected 1997 because it was the year of publication of a seminal and influential paper on SOA in leadership (Atwater & Yammarino, 1997). In the current article, we will examine the following topics related to SOA: (1) models of agreement, (2) factors affecting self-ratings and the congruence between self- and others' ratings, (3) factors affecting others' ratings, (4) correlates of agreement, (5) measurement issues and data analytic techniques, (6) practitioner issues, and (7) directions for future research.

According to Atwater and Yammarino (1992), investigations of the correlates of SOA are of interest to both researchers and practitioners. The investigation of differences between self- and others' ratings appears to have important implications for both research and practice, especially in the area of leadership development. The relationships between SOA and various outcomes of interest (e.g., leader effectiveness) justify further investigation of this topic; therefore, the objective of the current paper is to review and synthesize the relevant research on SOA.

1. Models of self–other rating agreement

Self-ratings alone, in general, are not considered to be accurate predictors of leadership outcomes because they are likely inflated by leniency bias (Halverson, Tonidandel, Barlow, & Dipboye, 2005; Hough, Keyes, & Dunnette, 1983; Podsakoff & Organ, 1986). Other researchers have documented that self-ratings are unreliable, invalid, and inaccurate when compared to the ratings of others or to objective criteria (Ashford, 1989; Harris & Schaubroeck, 1988; Mabe & West, 1982; Yammarino & Atwater, 1993).

In a meta-analytic study, Mabe and West (1982) found that raters who provided more accurate ratings (and were therefore more self-aware) tended to score higher on characteristics such as intelligence, achievement status, and internal locus of control. Ashford (1989) developed a model of self-evaluation in which she described the theoretical antecedents of incongruent self- ratings and their effects on individuals. In an effort to better understand why self-ratings tend to be inaccurate, Yammarino and Atwater (1993) developed a conceptual model of self-perception accuracy based on several studies that were designed to increase the understanding of SOA. Their model of agreement encompasses the entire range of rating sources, including subordinates, peers, and superiors. Yammarino and Atwater posit that individual and organizational outcomes are: (a) enhanced when self- perception is accurate, (b) diminished when self-perception is inflated (over-estimators), and (c) mixed when self-perception is deflated (under-estimators).

Based on this earlier research (Ashford, 1989; Mabe & West, 1982; Yammarino & Atwater, 1993), Atwater and Yammarino (1997) developed a comprehensive model of SOA that included explanations of why self-ratings are often found to differ from others' ratings. Following Ashford (1989), Atwater and Yammarino proposed that factors such as cognitive processes, job experience, biographical information, personality characteristics, and contextual factors affect self-perception and, as a result, self- ratings. Additionally, they proposed that factors such as tenure and emotional stability affect the ratings of others.

In their model, Atwater and Yammarino (1997) posit that leaders whose self-ratings agree with others' ratings as to their high levels of effectiveness are more likely to be linked to positive individual and organizational outcomes. These in-agreement/good estimators are good performers with views of themselves that are similar to those of other raters. On the other hand, leaders who agree with others about their low level of effectiveness are more likely to be linked to negative outcomes. These in-agreement/ poor estimators are poor performers who recognize their low level of performance, but are unwilling or unable to change.

In the literature review that follows, we will describe correlates of self-ratings, correlates of others' ratings, and correlates (predictors or consequences) of SOA. Unfortunately, the literature on SOA has used a variety of metrics to assess this phenomenon (as discussed in the section on measurement issues) and these metrics are not equivalent. Hence, discrepancies between findings across studies may be the result, at least in part, of how SOA was operationalized in different studies. Also, the results of any single study may or may not be replicated if SOA had been operationalized differently in that particular study. The inconsistent operationalizations of SOA constitute a major stumbling block to drawing broad conclusions about the causes and consequences of SOA; therefore, we strongly recommend that future research use rigorous operationalizations of SOA when feasible. For example, when SOA is being treated as a predictor, researchers should use widely accepted methods such as polynomial regression (Edwards, 1994) and within and between analysis (WABA; Yammarino, 1998) to analyze SOA, rather than alternative methodologies.

2. Factors affecting self-ratings and congruence between self- and others' ratings

Self-assessment is used for a wide variety of purposes in organizations including as part of the performance appraisal process, in 360-degree feedback programs, or as input into selection decisions, to name a few. Because of a variety of individual and

1007J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

contextual factors, however, self-ratings do not always reflect reality or the ratings given by others. In this section of the paper, we review what has been learned about predictors of, or factors affecting, self-ratings, primarily since the publication of Atwater and Yammarino's (1997) paper. Also included, where relevant, are factors affecting SOA. Table 1 presents a summary of the factors that affect self-ratings and the congruence between self- and others' ratings.

2.1. Biographical characteristics

2.1.1. Gender In a sample of students and faculty, Visser, Ashton, and Vernon (2008) found that men gave higher self-estimates of their

spatial, mathematical, and kinesthetic abilities than did women, and that these higher estimates were not explained by sex differences in measured abilities in these areas. They concluded from this that women were more accurate in their estimates because their degree of over-estimation was smaller. Moshavi, Brown, and Dodd (2003) found that male managers gave themselves higher ratings of transformational leadership than did their female counterparts although there was no relationship between gender and follower ratings. Brutus, Fleenor, and McCauley (1999) also found that male managers were more likely to overestimate their effectiveness, while female managers had self-ratings that were aligned with the ratings from their peers and subordinates. Data from a 360-degree feedback program with managers revealed that the tendency to overestimate one's own effectiveness as a leader was greater for men than for women (Vecchio & Anderson, 2009).

Lindeman, Sundvik, and Rouhiainen (1995) found that over twice as many males as female job applicants over-rated their abilities in the areas of sales and marketing. Nearly three-fourths of males over-rated while only one-third of females over-rated their abilities. Jones and Fletcher (2002) found a tendency for men to inflate self-ratings compared to women in a selection context. This finding appears to be pervasive as Patiar and Mia (2008) found the same leniency tendency in self-ratings of performance relative to boss ratings among male managers in the hotel industry.

Fletcher (1999) examined these differences in ratings relative to gender and postulated that some of the contributing factors are due to personality differences and feedback-seeking propensities. Specifically, women were more likely to correctly select their own assessment reports than men, and this skill was partly explained by their personality factor O (self-assured versus apprehensive as measured by the 16PF). Factor O is defined as those who are more prone to worry, feel anxious, and feel less accepted by others. The authors suggested that being more anxious and self-critical could contribute to a tendency to seek more feedback and to monitor one's effect on others. Women are also more self-disclosing than men (Dindia & Allen, 1992), which could be a sign of greater self-awareness.

2.1.2. Age According to Brutus, Fleenor, and McCauley (1999), older managers, as compared to younger managers, tended to over-rate

their performance in relation to ratings provided by their supervisors. Vecchio and Anderson (2009) also noticed the tendency for older managers to overestimate their effectiveness on a 360-feedback instrument.

Similarly, Ostroff et al. (2004) found that older managers gave themselves higher ratings of leadership than did younger managers, yet they were rated lower than younger managers by their subordinates and peers. Moshavi et al. (2003) also found higher self-ratings of transformational leadership among older managers but follower ratings were also higher for older managers. However, the Moshavi et al. study was from a single organization and was not as robust and reliable as the findings from Ostroff et al., which included managers from 527 organizations.

Table 1 Factors affecting self-ratings and congruence between self- and others' ratings.

Topics Summary

Biographical characteristics In general, with regard to age and gender, males and older individuals tend to over-rate their leadership, abilities, and effectiveness relative to other raters. This over-rating contributes to greater discrepancies between self- and others' ratings. Personality differences and feedback-seeking propensities among males and females were postulated to contribute to females' less inflated ratings. Those in higher positions tend to over-rate, which may result from a lack of appropriate feedback from others. Non-Whites and those with less education also were found to be more likely to provide inflated self-ratings.

Personality and individual characteristics Extraversion, openness to experience, agreeableness, conscientiousness, and dominance were positively related to self-ratings of leadership while neuroticism was negatively related. However, these relationships with personality seem only to hold for self-ratings, not for others' ratings. Empathy was the only individual characteristic related to congruence between self- and others' ratings. Narcissists and those with high self-esteem tend to over-rate their performance. Intelligence and meta-cognitive ability are associated with more agreement between self- and others' ratings. Those with an internal locus of control provide higher self-ratings. Interestingly, depressed individuals tend to provide self-ratings that are in greater agreement with others, perhaps stemming from their focus on cues from others.

Job relevant experiences When raters receive feedback over time, their self-ratings become more congruent with ratings from others. However, this congruence may be a result of changes in others' ratings rather than increased self-awareness on the part of the self-rater.

1008 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

2.1.3. Position Brutus, Fleenor, and McCauley (1999) found that organizational level was related to discrepancies among self–others' ratings

and, in particular, they suggested that managers in higher positions (such as senior executives) in the organization were more inclined to provide inflated self-ratings when compared to others' ratings of their performance. Sala (2003) took this notion a step further by examining the relationship of position and ratings on a 360 instrument administered to 1214 individuals from a variety of organizations. The results of this study indicated that higher-level employees had greater disagreement between self- and others' ratings than lower-level individuals.

Gentry, Hannum, Ekelund, and de Jong's (2007) results replicated previous research (Ostroff et al., 2004; Sala, 2003) by showing that as managerial level increases, so does the gap between self- and others' ratings. In addition, their study extends the literature by indicating that this incongruity among ratings may be caused by inflated ratings at higher managerial levels. They suggested that these incongruities may occur because managers in higher positions do not receive appropriate (or any) feedback or because they are too over-confident.

2.1.4. Race Ostroff et al. (2004), in a study of over 3000 managers across 527 organizations, found the race of managers to be correlated

with self-ratings of leadership. Specifically, non-White managers rated their own leadership higher than did their White counterparts. Race of the manager was unrelated to subordinates', peers', or supervisors' ratings of the managers' leadership behaviors. The explanation given for these race effects was that non-White managers may feel the need to strongly assert themselves especially in positions traditionally held by Whites. Non-Whites may also be more reluctant to accept feedback, particularly if it comes from Whites. A third possibility suggested by these authors was that non-Whites who achieve a high-status managerial position may see this as a statement about their worth and value.

2.1.5. Education Ostroff et al. (2004) found a positive correlation between self-ratings of leadership and level of education. Interestingly, they

also found that those with more education were in more agreement with others when self- and others' ratings were compared. This may have occurred in part because those with more education are likely to also have higher levels of analytic and cognitive abilities (cf. Kingston, Hubbard, Lapp, Schroeder, & Wilson, 2003), which may help them better process self-relevant information.

2.2. Personality and other individual characteristics

2.2.1. Big Five personality factors There is evidence that personality is more strongly related to self-estimates of intelligence than to actual measures of

intelligence (Furnham, Moutafi, & Chamorro-Premuzic, 2005). For example, Furnham et al. found that neuroticism and agreeableness were both negatively related to self-estimates of intelligence but unrelated to measured intelligence. Judge, LePine, and Rich (2006) found that extraversion, openness to experience, agreeableness, and conscientiousness were positively related to self-ratings of leadership, while this relationship for neuroticism was negative. Visser et al. (2008) also found that extraversion, conscientiousness, and openness to experience were positively related to self-estimated ability. Similar results were obtained by Bell and Arthur (2008) who found a positive relationship between participants' self-ratings of assessment center performance and their level of extraversion. This relationship, however, was not observed for assessor ratings and participant extraversion. Extraverts seemed to believe they were better performers but this belief was not shared by their assessors.

2.2.2. Dominance Jackson, Stillman, Burke, and Englert (2007) found three measures of dominance (confidence in social settings, confidence in

one's intellectual abilities, and determination and forcefulness) to be positively associated with self-ratings of performance in assessment center tasks. However, independent assessors did not share perceptions of their “better” performance with the dominant self-raters.

2.2.3. Empathy Brutus, Fleenor, and McCauley (1999) found an interesting relationship between empathy and SOA. Empathy was the only

personality trait in their study that predicted congruence across all others' ratings with self-ratings. For example, they found that managers who rated themselves high in empathy also received higher ratings from other rating sources.

2.2.4. Self-esteem Levy (1993) found self-esteem to be related to self-ratings on a hypothetical test of managerial potential. Goffin and Anderson

(2007) measured job performance of 204 managers through the utilization of multisource ratings. Regression analysis indicated that managers high in self-esteem tended to overestimate their performance. They also found that managers with high anxiety under-rated their performance relative to their supervisors' ratings.

2.2.5. Narcissism Narcissists (those with an inflated sense of self-importance) also tend to over-rate. Robins and John (1997) asked narcissists

and non-narcissists to rate their individual performance on a group task. While there were no true performance differences,

1009J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

narcissists rated themselves more favorably than did non-narcissists. Interestingly, after narcissists and non-narcissists viewed their performance on videotape, the non-narcissists' self-ratings decreased (to nearly perfect accuracy) while the narcissists' self- ratings increased.

Judge et al. (2006) conducted two studies assessing the relationship between narcissistic personality and self- and others' ratings of leadership. They found that narcissism was related to higher self-ratings even when the Big Five personality traits were controlled. In addition, while narcissism was positively related to self-ratings of leadership, it was negatively related to others' ratings of leadership.

From the previous discussion, we can conclude that a variety of personality characteristics are associated with self-rating tendencies.

2.2.6. Meta-cognitive ability/intelligence Kruger and Dunning (1999) found that an individual's true ability was most poorly recognized by those who scored lowest on

ability. For example, the bottom quartile of individuals on a test of logical reasoning scored on average at the 12th percentile while they estimated their score at the 62nd percentile. Kruger and Dunning attributed this lack of awareness to a deficit in meta- cognitive skill. As they concluded in their paper, “unskilled individuals suffer a dual burden: Not only do they perform poorly but they fail to realize it” (p. 1131). Moreover, London (1995) suggested that individuals with higher intelligence are more likely to provide more accurate self-assessments because their greater intelligence allows them to collect, process, and remember more self-relevant information.

Cognitive, analytical, and numerical abilities as measured with standardized tests were all found to be negatively related to employees' self-ratings of performance (Beehr, Ivanitskaya, Hansen, Erofeev, & Gudanowski, 2001). Again, this suggests that those with more ability are less likely to inflate their self-ratings.

2.2.7. Private and public self-consciousness Sosik and Megerian (1999) found no relationships between self-ratings of leadership and leaders' private or public self-

consciousness. However, higher ratings of public self-consciousness were more likely among those who overestimated their transformational leadership when compared to subordinate ratings.

2.2.8. Self-monitoring Self-monitoring, or the ability to adapt one's behavior in the presence of social cues (Snyder & Copeland, 1989), may be related

to SOA in that high self-monitors would be expected to use their self-monitoring skills to better interpret how they were perceived by others. However, somewhat counter-intuitively, while high self-monitoring was related to higher self-ratings among a young student sample, low self-monitors showed greater agreement between their self-ratings of performance and ratings from others (Miller & Cardy, 2000). These results were explained in terms of the consistency of presentation by low self-monitors to other raters. That is, they do not adjust their presentations of self to others' expectations and therefore are not evaluated differently by different others.

Blakely, Andrews, and Fuller (2003) advocated this position as well by indicating that low self-monitors are less concerned with their own impact on others and are guided more by their internal feelings and attitudes than by situational cues. Thus, low self-monitors may be more aware of their own behavior and will rate themselves more accurately in terms of how others perceive them (hence, high rating congruence). Miller and Cardy's (2000) study was also interesting because in the young student sample, high self-monitoring was associated with higher self-appraisals while in a field setting this was not the case. The authors explain this as follows: “High self-monitors' hyper vigilant approach to the social milieu as young adults entering the workforce may evolve into a tendency to be more inwardly focused and self-critical as years and experience accumulate” (Miller & Cardy, 2000, p. 621). Self-monitoring was also unrelated to self-ratings of transformational leadership in a sample of 63 technology managers (Sosik & Megerian, 1999). It seems that either age/maturity/experience is relevant to the relationship between self-monitoring and self-ratings or the results from the student sample were spurious.

2.2.9. Efficacy and locus of control Personal efficacy, the belief that one can make things happen, and social self-confidence (i.e., one's comfort in groups) were

both significantly positively related to self-ratings of transformational leadership (Sosik & Megerian, 1999). In addition, Levy (1993) found locus of control to be significantly related to self-ratings of performance on a hypothetical test of managerial potential. Those with an internal locus of control provided higher self-ratings of managerial potential, which Levy suggested fit within an attribution framework. That is, those who believe they have more responsibility for their positive outcomes are more likely to view their behavior more positively. However, because no objective measure of managerial potential was available, this could not be clearly concluded. It is also possible that those with an internal locus of control deserve higher ratings because they are more capable as managers (Whetten & Cameron, 2007).

2.2.10. Depression Bagby et al. (1998) found high congruence between depressed individuals' self-ratings of four of the Big Five dimensions of

personality (agreeableness, neuroticism, openness to experience, and conscientiousness) and others' ratings of personality of these same individuals. Although depressed individuals tended to under-rate their extraversion, these results indicate overall that depressed individuals are more accurate in assessing their personality relative to others' assessments. The possible explanation for

1010 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

this is that depressed individuals provide lower self-ratings, which are more likely to be in line with ratings from others. Or perhaps, depressed individuals are more attuned to cues from others.

2.3. Context

2.3.1. Culture While Atwater and Yammarino (1997) did not discuss culture in relation to self-ratings or SOA, research has suggested that

culture may come into play, particularly in how it affects self-ratings. For example, Farh, Dobbins, and Cheng (1991) proposed a cultural-relativity hypothesis suggesting that individualism/collectivism influences self-rating behavior. Specifically, this hypothesis argues that raters with a collectivist orientation (simplistically defined as a preference for being treated as a group member as opposed to an individual) will show less leniency bias in their self-ratings. A number of studies comparing individuals who live in more or less collectivist countries have supported this hypothesis (cf. Farh et al.; Farh & Cheng, 1997; Yik, Bond, & Paulhus, 1998). Xie, Roy, and Chen (2006) expanded this investigation by measuring not only country-level culture but individual perceptions of their individualist/collectivist orientations. Using a sample of university students from Canada, China, Taiwan, Hong Kong, and Japan, they found that individualists were more likely to inflate ratings of cognitive ability and cognitive performance when actual levels of cognitive ability were controlled.

In a related study, Korsgaard, Meglino, and Lester (2004) measured “other orientation” among 450 new hires in a U.S. healthcare center. Other orientation was defined as concern for others as opposed to self-orientation. They found a significant relationship between self-ratings of performance and the individual's other orientation (higher self-ratings for those with greater self-orientation). Those with other orientation also showed more agreement between self- and supervisor ratings. They speculated that greater SOA among other-oriented individuals may be due, in part, to their tendency to engage in “hansei” (greater self-reflection). Thus, they may be more likely to accept negative feedback from their supervisors. Both greater self-reflection and greater acceptance of negative feedback by those with more other orientation could help reduce inflated self-ratings. Gentry et al. (2007) also found SOA to be greater for American managers relative to European managers in certain respects (i.e., problems with interpersonal relationships).

Atwater et al. (2009) found that leaders across 21 countries were more likely to give themselves higher ratings of leadership in cultures that were higher on individualism (as characterized by the GLOBE study; see House, Hanges, Javidan, Dorfman, & Gupta, 1999). This tendency occurred when compared to both peer and subordinate ratings. It is not clear whether this suggests a leniency bias in individualist cultures or a modesty bias in collectivist cultures, or both.

In a study of managerial derailment, Gentry, Braddy, Fleenor, and Howard (2008) compared SOA of Hispanic leaders to leaders from three other ethnic groups (Caucasians, African-Americans, and Asians). With a sample of over 1000 leaders, the authors found no differences in SOA between Hispanics and Caucasians. Hispanics, however, were found to demonstrate greater rating discrepancies than African-Americans, but smaller discrepancies than Asians. Self-boss and self-direct report rating discrepancies appeared to have resulted from inflated self-ratings.

Gentry, Yip, and Hannum (2010) analyzed multisource ratings for managers from Southern Asia (n=261) and Confucian Asia (n=599) for cultural differences in SOA. Multivariate regressions revealed that self–other rating discrepancies were larger for managers from Southern Asia compared to Confucian Asia. The discrepancies resulted from the managers' self-ratings being different across cultures (i.e., they were not due to others' ratings).

2.3.2. Controllability Rothermund, Bak, and Brandtstädter (2005) investigated whether self-rating inflation would be more pronounced for

characteristics of the self that were deemed more or less controllable. Controllable characteristics, for example, were things such as class attendance or types of classes selected, whereas uncontrollable characteristics included attachment to one's parents or interest in science. The authors found that when uncontrollable characteristics were classified as highly related to success, these characteristics were far more likely to receive higher self-ratings. Those characteristics deemed positive and controllable were under-rated. The authors concluded that self-enhancement bias is more salient for attributes that raters believe are not open to improvement.

2.3.3. Political purpose Heidemeier and Moser (2009) included employee development in their study of SOA and unexpectedly found that managers

tended to over-rate their ability in developmental performance ratings versus research-based settings. In contrast, when managers expected a validation of ratings, they tended to rate themselves lower (i.e., more modestly). This illustrates that the context in which the rating occurs can have an impact on how managers rate themselves. Thus, research-based settings wherein managers may feel pressure to self-report accurately led to self–other congruence, whereas developmental settings seemed to have created inflated self-ratings.

2.3.4. Similarity Ostroff et al. (2004) found that the similarity between the age or gender of managers and the age or gender of other raters (e.g.,

their subordinates, peers, and supervisors) was unrelated to the manager's self-ratings.

1011J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Caligiuri and Day (2000) examined self-monitoring and rater–ratee national similarity across three varying performance dimensions. They found that high self-monitors were rated more favorably by supervisors of the same nationality, indicating that cultural similarity may play a role in SOA.

2.4. Job relevant experiences

2.4.1. Feedback There is evidence that managers' perceptions of their competence increase and that their developmental needs decrease over

time when the appropriate feedback measures are utilized. In a study of management competence over a two-year period, Bailey and Fletcher (2002) found significant increases in agreement between self-supervisor ratings and self-subordinate ratings. This congruence was linked to a further understanding of the managers' self-awareness of their competence, which was initially overestimated relative to subordinates' ratings and under-estimated relative to supervisors' ratings. However, this congruence among ratings, as shown by the descriptive data, may be the result of subordinates changing their ratings and not the result of managers becoming more self-aware and thus revising their self-ratings to fit feedback.

3. Factors affecting others' ratings

Vance, Winne, and Wright (1983) examined performance ratings provided by 90 supervisors who rated 350 subordinates on five occasions over three and one-half years. The nature of their data set allowed them to hold either raters or ratees constant, thereby enabling inferences regarding the sources of reliable variance due to raters or ratees. They demonstrated that, although reliable variance in mean ratings was partly attributable to ratees, it was mainly attributable to raters. A comprehensive presentation of all the factors that might influence others' ratings is beyond the scope of this paper. In this section, we review the major factors that are likely to shape others' ratings. These include (a) the rater's cognitive processes, (b) characteristics of the rater (e.g., ability, personality, and beliefs), (c) rater motivation, (d) contextual factors, and (e) rater–ratee interactions and expectations. Table 2 presents a summary of the factors that affect others' ratings.

3.1. The rater's cognitive processes

Cognitive models of performance appraisal (DeCotiis & Petit, 1978; DeNisi, 1996; DeNisi, Cafferty, & Meglino, 1984; Feldman, 1981; Ilgen & Feldman, 1983) have focused on raters' cognitive processes. These models, borrowing heavily from research on person perception and social cognition in social psychology, look at how raters recognize, attend to, and observe employee behavior (or other information related to employee performance), represent, organize, and store this information in memory, retrieve the information from memory, and integrate the information to form a judgment of the employee. This research illustrated the important role that categories and schemas play in automatic processing (where little cognitive energy is expended) and showed that controlled processing (which requires more deliberate cognitive effort) occurs only when information is acquired about an employee that is quite inconsistent with the category or schema to which the employee has already been assigned by the rater. It also showed how categories and schemas shape subsequent information processing (e.g., what raters attend to and recall) and ratings of performance. That is, raters initially use observations of the employee's behavior or other information about the employee to place him/her into a preexisting category or schema (e.g., typical engineer, typical salesperson, irresponsible employee, or high- potential employee). The rater's subsequent appraisal of the employee's performance reflects not only the employee's performance but also the rater's beliefs about how a typical individual in that category performs.

Several authors (Lord, Foti, & de Vader, 1984; Rush, Phillips, & Lord, 1981) noted that ratings of leadership reflect not only the leader's actual behavior but also the rater's implicit leadership theory (ILT) that describes the behaviors, abilities, and traits needed for effective leadership. ILTs are cognitive schemas that can affect encoding (e.g., filtering new information) and retrieval (by relying on one's implicit theories to fill gaps in memory due to ambiguous, lost, or unavailable information about the leader). Thus, raters evaluate leaders in terms of prototypicality or the extent to which the leader's attributes are consistent with the rater's prototype of a leader (i.e., ILT). Consistent with ILT, holding actual leader behavior constant, leaders of purportedly high- performing teams receive more favorable behavioral descriptions than do leaders of purportedly low-performing teams. Gioia and Sims (1985) found that ratings made using more behaviorally specific scales were less affected by raters' ILTs and better reflected actual leader behaviors than ratings made using a more general scale (i.e., the Leader Behavior Description Questionnaire).

3.2. Characteristics of the rater

3.2.1. Rater ability Long ago, Taft (1955) noted that intelligence was positively correlated with the ability to accurately judge the personality

characteristics of others. Around the same time, Schneider and Bayroff (1953) and Bayroff, Haggerty, and Rundquist (1954) found that cognitive ability was positively associated with the validity of ratings. Borman (1979) found that 17% of the variance in rating accuracy was explained by 12 individual difference variables with the highest associations of accuracy being with rater intelligence, attention to detail, and personal adjustment.

More recently, several laboratory studies (Hauenstein & Alexander, 1991; Smither & Reilly, 1987) have found that rater intelligence is associated with more accurate ratings. Cardy and Kehoe (1984) examined the relationship between field

Table 2 Factors affecting others' ratings.

Topics Summary

The rater's cognitive processes Raters' categories and schemas shape what they attend to and recall about ratee performance. Raters initially use observations of the employee's behavior or other information to place the employee into a preexisting category or schema, and the rater's subsequent appraisal of the employee's performance reflects not only the employee's performance but also the rater's beliefs about how a typical individual in that category performs. Raters evaluate leaders in terms of prototypicality—the extent to which the leader's attributes are consistent with the rater's prototype of a leader.

Characteristics of the rater Rater ability (intelligence) is associated with more accurate ratings. Several personality-related characteristics of raters appear to be associated with rating quality, including rater cynicism, agreeableness, self-monitoring, self-efficacy, and conscientiousness. Raters in a positive mood evaluate others more favorably. A rater's implicit person theory about the malleability of personal attributes (e.g., personality and ability) affects whether the rater recognizes and acknowledges change in employee behavior. Raters who have a high level of appraisal discomfort provide more lenient ratings.

Rater motivation Rater motivation can affect all stages of the performance appraisal process including observing employee performance, storing the observed information in long-term memory, retrieving information, integrating information, rating the employee, and providing feedback. Motivated raters are more likely to use deliberate or controlled information processing strategies rather than quick, heuristic-based, or automatic information processing strategies at each of these stages. The rater's goal need not necessarily be to provide an accurate rating of the employee's performance. Raters can pursue a variety of goals that include increasing the employee's (ratee's) performance, building and maintaining good working relationships with employees, and enhancing the rater's status in the organization (e.g., giving favorable ratings to convey a positive image of the rater's skills as a manager). Ratings are especially likely to be shaped by political considerations. Situational factors can enhance rater motivation. These include accountability (having to justify one's ratings to others), interdependence (when the rater's outcomes and rewards are highly dependent on the employee's performance), trust in the appraisal system (believing that others will provide fair, rather than lenient, ratings of their employees), and ease of use (appraisal forms that are not too difficult or time-consuming to complete).

Contextual factors Frame-of-reference training for raters is an effective approach to increase rating accuracy. Appraisals obtained for administrative purposes are generally more favorable than those obtained for employee development or research purposes. Norms can develop within a group about what constitutes acceptable performance ratings. Raters who believe that other raters are purposely inflating ratings to get a better deal for their direct reports are themselves more likely to provide inflated ratings. The strength of the situation will determine the rater's opportunity to observe the ratee's underlying traits and competencies. In strong situations, behavioral expectations are straightforward, so there is little variability in observed behavior (thereby masking any between-individual differences in underlying traits and competencies). There is more variability in how people respond to weak situations (thereby revealing between-individual differences in underlying traits and competencies). National culture can influence choice of raters (e.g., in high power distance cultures, those higher in status might be unreceptive to feedback from those lower in status). Raters in high collectivism cultures are likely to emphasize relationships, loyalty, and cooperativeness in their evaluations. Although culture may be important, there is a risk of taking such generalizations too far (leading to unwarranted stereotypes) due to the wide range of behavior that can often be observed within any one culture. The performance of a ratee's peers can affect evaluations of the ratee (e.g., the greater the proportion of noncompliant workers in a unit, the more favorable the supervisor's judgments of compliant workers).

Rater–ratee interactions and expectations

The relationship between rater and ratee affects performance ratings (e.g., in-group members receive more favorable evaluations). Managers and subordinates provide the most favorable evaluations of each other when both perceive themselves as similar. Similarly, raters' interpersonal affect (liking) toward a ratee is positively related to the favorableness of performance ratings. Knowledge of a ratee's prior performance affects ratings of the ratee's current performance sometimes resulting in assimilation effects (where current ratings are biased in the direction of prior performance) and other times resulting in contrast effects. Raters who have initially promoted an employee provide more favorable ratings of the employee than raters who had not initially promoted the employee. A leader's positive expectations of another person (target) have an influence on the target such that the target ultimately changes his or her behavior or performance in a way that confirms the leader's positive expectations. In such cases, the higher ratings received by targets are justified because the targets have actually performed better. Opportunity to observe (e.g., amount of time the rater has been the ratee's supervisor) is associated with more valid ratings and higher inter-rater reliability of job performance ratings. Supervisors' ratings are higher when they receive a favorable (rather than unfavorable) self-appraisal from the employee.

1012 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

independence and rating accuracy. Field independent persons can separately attend to the various dimensions of a multidimensional stimulus, whereas field dependent persons have difficulty doing so. Cardy and Kehoe argued that, to the extent that individuals are influenced by irrelevant factors in the stimulus field, the accuracy of their judgments (e.g., ratings) should be adversely affected. They found that field independent raters provided more accurate ratings than did field dependent raters, but Härtel (1993) found that this occurred only when scale formats were holistic (i.e., where raters provide only a single

1013J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

rating for each dimension). In a related study, Cellar, Durr, Halsell, and Doverspike (1989) found that field independent raters provided more accurate job evaluation ratings. However, other studies have reported that rater intelligence and field independence did not affect rating quality (Bravo & Kravitz, 1996). Also, Bernardin, Cardy, and Carlyle (1982) found that raters' cognitive complexity was unrelated to accuracy.

3.2.2. Rater job experience and performance Cascio and Valenzi (1977) found a small, positive relationship between raters' length of job experience (and level of education)

and ratings of others. Kirchner and Reisberg (1962) found that better supervisors were more discriminating in their ratings of subordinates. Also, whereas poorer supervisors paid more attention to loyalty and cooperation, better supervisors regarded independent, forward looking action by subordinates as important.

3.2.3. Rater personality Early research found little or no relationship between rater personality and halo in ratings (inter-correlations among ratings on

various dimensions) or agreement with consensus ratings (Dubin, Burke, Katz, & Chesler, 1954a,b). However, more recent research has identified several personality-related characteristics of raters that appear to be associated with rating quality. For example, employees who express cynicism towards upper management and the upward feedback process have lower intentions to provide honest upward feedback ratings (Smith & Fortunato, 2008). Raters high in agreeableness provide more lenient ratings when they expect to have a face-to-face meeting with the employee, but this effect is attenuated when using a behavior checklist rather than a graphic-rating scale (Yun, Donahue, Dudley, & McFarland, 2005).

Jawahar (2001) found that high self-monitors provided less accurate ratings than low self-monitors and that attitudes toward accurate appraisal (e.g., believing that personal biases should not influence ratings) were positively related to accuracy of ratings for low self-monitors but not for high self-monitors. Finally, Tziner, Murphy, and Cleveland (2005) suggested that raters with low self-efficacy might lack sufficient motivation to provide well-documented and accurate evaluations, whereas raters with high self- efficacy are likely to approach the appraising task more conscientiously. They also argue that raters high on conscientiousness might complete performance appraisals with greater diligence, thereby leading to better discrimination among ratees and to less inflated ratings. Also, Tziner et al. predicted that raters who are high self-monitors (i.e., sensitive to environmental cues) are especially likely to be affected by the rating context.

3.2.4. Rater mood Fried, Levi, Ben-David, Tiegs, and Avital (2000) noted that laboratory research shows that raters who experience positive mood

evaluate others more favorably, whereas raters who experience negative mood evaluate others less favorably. Using a realistic organizational simulation, they found that negative mood predisposition was negatively related to performance ratings, whereas positive mood predisposition was unrelated to performance ratings.

3.2.5. Rater beliefs about human nature Heslin and his colleagues (Heslin, Latham, & VandeWalle, 2005; Heslin & VandeWalle, 2008) have shown that a rater's implicit

person theory (IPT) about the malleability of personal attributes (e.g., personality and ability) can affect whether raters acknowledge change in employee behavior. Specifically, some people hold an “entity” theory that human attributes are innate, stable over time, and unalterable, whereas other people hold an “incremental” theory that personal attributes can be developed. They found that managers who hold an “entity” theory tend to inadequately recognize actual changes in employee performance (and are also less inclined to coach employees about how to improve their performance), whereas managers who hold an “incremental” theory are more likely to acknowledge an improvement in employee performance and hence provide more accurate performance appraisals (and helpful employee coaching).

Wexley and Youtz (1985) found that raters with positive beliefs about human nature (e.g., people in general are altruistic and trustworthy) provided more lenient ratings. Rater accuracy was negatively correlated with beliefs about the variability of human nature.

3.2.6. Rater attitudes Forsyth, Heiney, and Wright (1997) found that participants who are conservative in their attitudes toward the role of women in

contemporary society rated a task-oriented female leader more favorably than a relationship-oriented female leader. Perhaps this occurred because she more closely matched their stereotypically masculine leader prototypes (whereas ratings by participants with liberal attitudes revealed the opposite pattern).

3.2.7. Rater's organizational commitment Tziner et al. (2005) noted that the process of appraising others' performance can require raters to invest a good deal of effort

and, in some instances, to take risks (because unfavorable ratings might damage interpersonal relationships or lead to resentment and complaints). They suggested that, to some extent, raters' willingness to engage in the rating process will reflect their attitudes about the organization such that raters who identify with the organization will exert more effort to develop the skills required to evaluate others effectively. Tziner, Murphy, Cleveland, Beaudin, and Marchand (1998) and Tziner and Murphy (1999) reported that organizational commitment was positively correlated with overall level of ratings. However, Jawahar (2001) found no relationship between instrumental or affective organizational commitment and ratings.

1014 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

3.2.8. Discomfort with appraisal Raters who have a high level of appraisal discomfort provide more lenient ratings than raters who have a low level of appraisal

discomfort (Villanova, Bernardin, Dahmus, & Sims, 1993).

3.2.9. Rater's organization level relative to the ratee A meta-analysis by Conway and Huffcutt (1997) found that correlations between rating sources are generally low

(subordinate-peer, .22; subordinate-supervisor, .22; subordinate-self, .14; supervisor-self, .22; peer-self, .19; and supervisor-peer, .34). Based on comparisons of between-source correlations with within-source reliabilities, they concluded that different sources (and thus people at different organizational levels) have different perspectives on performance.

3.3. Rater motivation

Over half a century ago, Taft (1955) noted the central role of motivation in making accurate judgments of others. Harris (1994) noted that much of the research concerning performance appraisal had focused on the rater's ability to provide accurate ratings but little attention had been directed to the rater's motivation in the appraisal process. He argued that there are three determinants of rater motivation: rewards (e.g., whether providing accurate ratings or feedback to employees will be rewarded by the organization or will indirectly affect the rater's rewards by leading to improved employee performance), negative consequences (e.g., whether ratings or feedback might demoralize the employee or damage the manager–employee relationship), and impression management (e.g., adhering to organizational norms or wanting to give favorable ratings so that the manager is perceived by others as having an effective work group). Harris argued that several situational factors will enhance rater motivation. These include accountability (having to justify one's ratings to others), interdependence (when the rater's outcomes and rewards are highly dependent on the employee's performance), trust in the appraisal system (believing that others will provide fair, rather than lenient, ratings of their employees), and ease of use (appraisal forms that are not too difficult or time-consuming to complete). Harris also emphasized that rater motivation can affect all stages of the performance appraisal process including observing employee performance, storing the observed information in long-term memory, retrieving complete information, integrating this information, rating the employee, and providing feedback. Moreover, motivated raters are more likely to use deliberate or controlled information processing strategies rather than quick, heuristic-based, or automatic, information processing strategies at each of these stages.

3.3.1. Rater goals Murphy and Cleveland (1995) noted that raters can use performance ratings to pursue a variety of goals that include increasing

the employee's (ratee's) performance, building and maintaining good working relationships with employees, and enhancing the rater's status in the organization (e.g., giving favorable ratings to convey a positive image of the rater's skills as a manager). Moreover, the rater's goal need not necessarily be to provide an accurate rating of the employee's performance. For example, Murphy and Cleveland (1991, 1995) noted that judgments (the rater's private view) and ratings (the rater's public statement) are not identical. Because ratings do not necessarily reflect the rater's judgments, a supervisor (or other raters) might hold an accurate view of an employee's performance but deliberately provide an inaccurate rating.

3.3.2. Politics Ratings (as contrasted with judgments) are especially likely to be shaped by political considerations, and several authors have

described the role that politics can play in performance ratings (Ferris, Fedor, Chachere, & Pondy, 1989; Kozlowski, Chao, & Morrison, 1998; Longenecker, Sims, & Gioia, 1987; Poon, 2004). Indeed, based on in-depth interviews with executives, Longenecker and colleagues (Gioia & Longenecker, 1994) Longenecker et al. found that politics often play a role in performance ratings. They noted that senior executives have considerable latitude in evaluating their employees and that appraisals can serve as political tools to control people and resources. A scale to measure perceptions of the extent to which performance appraisals are affected by organizational politics has been developed by Tziner, Latham, Price, and Haccoun (1996).

Numerous examples of how politics influence ratings have been observed. For example, because a poor rating might damage the supervisor's relationship with an employee, the supervisor might provide an inaccurate (more favorable) rating and thereby create a climate where the two can work together more comfortably going forward. Managers might also inflate their ratings of employees to maximize the merit increase available to the employee (especially if the merit ceiling is low) or to protect or encourage an employee whose performance is suffering due to personal problems. Or a supervisor might provide a poor rating (that is lower than the supervisor's judgment of the employee's performance) to teach a disruptive or uncooperative employee a lesson, to send a message that the employee should consider leaving the organization, or to create a documented record of poor performance that can speed up the termination process.

3.3.3. Rater accountability Mero and Motowidlo (1995) showed that raters who are made to feel accountable by having to justify their ratings in writing

provided more accurate ratings. Similarly, raters accountable to others with authority or higher status provide more accurate ratings compared with raters who are accountable to a lower status audience and raters who do not have to justify their ratings (Mero, Guidice, & Brownlee, 2007).

1015J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

3.3.4. Rater incentives Murphy and Cleveland (1995) have noted that, in applied settings, there are generally no extrinsic rewards (or punishments)

associated with providing accurate (or inaccurate) performance ratings. In a laboratory study, Salvemini, Reilly, and Smither (1993) found that bias resulting from knowledge of ratees' prior performance levels was eliminated when raters were offered incentives for accurate ratings. Also, incentives improved the accuracy of ratings.

3.4. Contextual factors

3.4.1. Situational strength Mischel (1973, 1977) has shown that the strength of a situation affects between-individual variability in behavior. In strong

situations, behavioral expectations and demands are unambiguous and straightforward. As a result, most people respond to the situation in a similar manner (i.e., there is little variability in observed behavior), thereby masking any between-individual differences in underlying traits and competencies. In weak situations, people differ in terms of how they perceive the situation and in their anticipations concerning the consequences of various responses. As a result, there is more variability in how people respond to such situations, thereby revealing between-individual differences in underlying traits and competencies. Moreover, when the same individual is observed across several strong, albeit different, situations, the cross-situational consistency of the individual's behavior is likely to be small or nonexistent. Trait-activation potential refers to the opportunity to observe between- individual differences in trait-relevant behaviors within a given situation (Haaland & Christiansen, 2002; Tett & Guterman, 2000). The more likely it is that between-individual differences in trait-relevant behavior can be observed in a situation, the higher that situation's trait-activation potential is said to be.

Also important is the relevance of the situation to the trait of interest. That is, situations vary in terms of the extent to which they provide cues for trait-relevant behavior (Murtha, Kanfer, & Ackerman, 1996). To illustrate, consider the example of attending a professional or trade conference. In such a setting, we would expect broad consensus that verbally (or physically) aggressive behavior would be unacceptable. Also, it is unlikely that the situation would provide any cues to trigger such behavior. Therefore, the situation has low trait-activation potential for the trait of aggression. In contrast, the same situation would likely have high trait-activation potential for the trait of sociability (e.g., because one can readily observe between-individual differences in how people interact, develop relationships, and share information with others who are attending the conference). In sum, the strength of the situations to activate particular traits in which the rater observes the ratee will determine the rater's opportunity to observe the ratee's underlying traits and competencies.

3.4.2. Culture Several authors (e.g., Davis, 1998; Day & Greguras, 2009; Fletcher & Perry, 2001) have described how national culture can affect

performance ratings. Davis suggested that culture can affect the way employees represent “ideal” performance and categorize performance information. This, in turn, can influence the work behaviors that are observed, remembered, and recorded when it is time to complete performance ratings. Culture can also influence choice of raters such that in high power distance cultures, those higher in status might be unreceptive to feedback from those lower in status (Davis), and workers in collectivist (rather than individualist) cultures might be reluctant to evaluate their peers (Fletcher & Perry, 2001). Day and Greguras (2009) suggested that employees from individualistic cultures are likely to provide more differentiated ratings of others, whereas employees from collectivistic cultures are more likely to provide less differentiated and perhaps more favorable ratings of others (as well as being more likely to provide group-level rather than individual-level feedback) in an effort to maintain harmony and facilitate teamwork (Aycan & Kanungo, 2001; Fletcher & Perry). They also suggest that raters from high humane-orientation cultures (where members encourage and reward friendliness, generosity, and support) will show greater tolerance for mistakes than those from low humane-orientation cultures. Several authors (Aycan & Kanungo; Fletcher & Perry) have proposed that raters in high collectivistic cultures are more likely to emphasize relationships, loyalty, and cooperativeness in their evaluations. Although culture may be important, there is a risk of taking such generalizations too far (leading to unwarranted stereotypes) due to the wide range of behavior that can often be observed within any one culture.

3.4.3. Difficulty of rating task The difficulty of the rating task can also affect the quality of ratings. For example, in a laboratory study, Smither and Reilly

(1987) found that ratings were more accurate when the true intercorrelation among job components was high rather than low. When the true intercorrelation among job components is high, raters can easily form an accurate overall impression of the ratee (who is consistently good or consistently poor across dimensions). In contrast, when the true intercorrelation among job components is low, the rating task is more difficult because accurate ratings can only be obtained if raters attend to and recall the ratee's performance on each dimension (because the ratee's performance is not consistently good or consistently poor across dimensions).

3.4.4. Observation context versus rating context Uleman (1991) noted that the context in which the rater initially observes an employee's performance and the context in

which the rater later recalls the employee's performance can affect recall and ratings of that performance. For example, consider a rater who initially observed an employee's relatively average performance in the context of several other excellent employees but when the rater subsequently (weeks or months later) recalls the employee's performance, that context is no longer salient.

1016 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Higgins and Lurie (1983) examined the aforementioned process by asking participants to read about the sentencing decisions of a target judge in the context of other judges who consistently gave either longer or shorter sentences than the target judge. Not surprisingly, participants tended to categorize the target judge as lenient in the former condition and harsh in the latter condition. One week later, the participants read about the sentencing decisions of additional judges (who were either harsh, moderate, or lenient, thereby creating a harsh, moderate, or lenient norm) and then recalled the sentencing decisions of the target judge who they read about one week earlier. Participants who were initially exposed to the same target judge in the same context, and initially categorized the target judge in the same way, nevertheless remembered the target judge's behavior differently if the norm were different at the time of recall.

3.4.5. Negative versus positive performance incidents It is likely that instances of poor performance will receive more weight in overall evaluations than will instances of good

performance. Baumeister, Bratslavsky, Finkenauer, and Vohs (2001) noted that people tend to pay more attention to negative than to positive phenomena, bad information is processed more thoroughly than good information, and bad impressions and bad stereotypes are more quickly formed and are more resistant to disconfirmation than good ones.

3.4.6. Primacy and recency effects In an examination of primacy and recency effects, Steiner and Rain (1989) showed that the serial position of a single poor or

good performance among three average performances by a ratee led to a recency effect for ratings of overall performance when the good performance occurred last and all four performances were viewed in a single session. However, when the four performances were viewed over a period of four days, a recency effect for ratings of overall performance occurred when poor performance was viewed last.

3.4.7. Anchoring effects Murphy and Cleveland (1995) noted that research on anchoring effects in judgment suggests that raters who evaluate the same

person on multiple occasions might be insensitive to changes in the person's behavior over time. For example, an employee's previous performance (or the rater's first impression of the employee) creates an expectation or anchor; when the employee's subsequent performance changes, the rater's evaluation of that subsequent performance will not change (increase or decrease) as much as it should, resulting in an assimilation effect.

3.4.8. Proportion of women or minorities in the work group Sackett, DuBois, and Noe (1991) found that women received lower ratings when the proportion of women in the work group

was small; however, this effect did not generalize to race. This is consistent with the finding of Ostroff et al. (2004) mentioned earlier, which suggested that groups consisting of more men are considered of higher value.

3.4.9. Performance of the ratee's peers Woehr and Roch (1996) showed that ratings of an average performer were affected by the performance level of previously

observed performers. Mitchell and Liden (1982) found that a poor performer was rated more positively and good performers more negatively when the poor performer was high (rather than low) in social or leadership skills. In a field study, Grey and Kipnis (1976) found that the greater the proportion of noncompliant workers in a unit, the more favorable the supervisor's judgments of his or her compliant workers. Linville and Jones (1980) showed that when an out-group ratee's (Black or opposite-sex) credentials were positive, the out-group member was evaluated more favorably than an identically qualified in-group member (White or same-sex); however, when the out-group ratee's credentials were weak, the out-group member was evaluated more negatively than an identically qualified in-group member. This effect occurred because people have more complex schemas concerning in- group members than out-group members, which in turn leads to moderation in evaluations.

3.4.10. Task interdependence Liden and Mitchell (1983) showed that the task interdependence of group members affected ratings such that good performers

were rated higher (and poor performers were rated lower) when working independently, whereas good performers were rated lower (and poor performers were rated higher) when group members were interdependent. Perhaps raters have difficulty evaluating the contribution of individual group members when group members are very interdependent.

3.4.11. Rater training Efforts to create rater training programs have focused on increasing the accuracy of ratings (Bernardin & Buckley, 1981;

Bernardin & Pence, 1980; Hauenstein, 1998). Such training (often referred to as frame-of-reference training) generally involves familiarizing raters with the definitions and behavioral indicators of each performance dimension, providing opportunities to complete practice ratings (using either written vignettes or videos to present the performance examples), and delivering feedback concerning the accuracy of the practice ratings (by comparing them with target ratings that represent the organization's estimate of the effectiveness levels demonstrated in the performance examples). A cumulative research review by Woehr and Huffcutt (1994) showed that frame-of-reference training is an effective approach to increase rating accuracy. Gorman and Rentsch (2009) recently showed that such training leads raters to possess schemas of performance that are similar to an expert-like schema, which in turn enhances accuracy.

1017J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

3.4.12. Rating purpose A meta-analysis by Jawahar and Williams (1997) found that the purpose of the appraisal influences leniency in ratings such

that appraisals obtained for administrative purposes (e.g., to influence pay raises or promotions) were about one-third of a standard deviation higher than those obtained for employee development or research purposes, especially when the ratings were made by practicing managers in real-world settings.

3.4.13. Trust in the appraisal process Raters who believe that other raters are purposely inflating ratings to get a better deal for their direct reports are themselves

more likely to provide inflated ratings (Bernardin & Cardy, 1982; Bernardin & Orban, 1990; Tziner et al., 1998).

3.4.14. Norms Murphy and Cleveland (1995) suggested that norms can develop within a group regarding what constitutes acceptable

performance ratings. For example, a norm might develop that influences raters to believe that average or below average ratings are unacceptable, even when such ratings might be accurate.

3.5. Rater–ratee interactions and expectations

3.5.1. Leader–member exchange It has long been assumed that the relationship between rater and ratee affects performance ratings (Kallejian, Brown, &

Weschler, 1953). Consistent with this assumption, Murphy and Cleveland (1995) summarized research on leader–member exchange theory (LMX) showing that in-group members receive more favorable evaluations, out-group members typically receive more extreme (usually, but not always, negative) evaluations than do in-group members, and out-group members are perceived as relatively homogeneous and hence might receive similar evaluations (see also Vecchio & Gobdel, 1984). A meta-analysis by Gerstner and Day (1997) found that leader-reported LMX was positively associated with leaders' ratings of subordinates' performance (r=.41); member-reported LMX was also positively associated with leaders' ratings of subordinates' performance (r=.28). Finally, LMX was positively related to measures of objective performance but to a lesser degree (r=.10).

3.5.2. Ratee's past performance Several studies have found that knowledge of a ratee's prior performance affects ratings of the ratee's current performance

(Kravitz & Balzer, 1992; Smither, Reilly, & Buda, 1988), sometimes resulting in assimilation effects (where current ratings are biased in the direction of prior performance) and other times resulting in contrast effects (when ratees are compared to each another rather than to a standard). Individual raters told that a group had been judged to be very good (before observing the group) subsequently recalled more effective behaviors (including behaviors that had not occurred) and fewer ineffective behaviors than raters told that the (same) group had been judged to be very poor (Martell & Leavitt, 2002). That is, knowledge of the target's performance serves as a cue that leads raters to recall cue-consistent attributes (effective or ineffective behaviors) as having occurred (even if they had not).

In a related study, Smither, Collins, and Buda (1989) examined whether knowledge of a ratee's satisfaction affected performance ratings. They found that students who were told that an instructor had high job satisfaction rated his performance more favorably than students who were told that he was dissatisfied. In a separate experiment, they asked participants to complete an in-basket task and a satisfaction questionnaire; participants provided with bogus feedback indicating that their task satisfaction was high evaluated their own performance more favorably than participants provided with dissatisfaction feedback.

3.5.3. Prior commitment to the ratee Bazerman, Beekun, and Schoorman (1982) provided raters with negative performance data concerning an employee who had

either been promoted by the rater or by someone else. Compared to raters who had not initially promoted the employee, raters who had initially promoted the employee provided more favorable ratings of the employee, provided larger rewards, and were more optimistic about the employee's future performance, suggesting an escalation of commitment effect.

3.5.4. Rater expectations about the ratee The Pygmalion effect is a type of self-fulfilling prophecy whereby a person's positive expectations of another person (target)

have an influence on the target such that the target ultimately changes his or her behavior or performance in a way that confirms the other person's positive expectations. For example, a manager's high expectations concerning an employee lead the manager to treat the employee differently, resulting in higher employee self-efficacy that, in turn, results in higher employee motivation and performance. A related concept, the Golem effect, is essentially the reverse of the Pygmalion effect (e.g., the target confirms the other person's low expectations). Two meta-analyses have examined such effects in organizational settings (Kierein & Gold, 2000; McNatt, 2000). Their results showed that Pygmalion interventions can have a sizeable effect on performance, with the effect being greater in military than in business settings. Note that, in such cases, the higher ratings received by targets are justified because the targets have actually performed better.

1018 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

3.5.5. Familiarity with ratees Landy and Farr (1980) noted that implicit personality theories are most likely to affect ratings in situations in which the rater

has little familiarity with the ratee. Smith and Fortunato (2008) found that opportunity to observe their supervisors was a key predictor of employee intentions to provide honest upward feedback ratings. Guion (1998) noted that the job relevance of a rater's contact with the ratee is more important than the mere frequency of contact.

Consistent with these arguments, Moser, Schuler, and Funke (1999) found that opportunity to observe was associated with more valid ratings. Another study found that ratings made by acquainted raters of briefings on research projects were more accurate than those made by unacquainted raters (Van Scotter, Moustafa, Burnett, & Michael, 2007). Freeberg (1969) found that performance ratings were more valid when raters possessed task-relevant acquaintance with the ratee than when they possessed merely task-irrelevant acquaintance (e.g., socializing).

Kingstrom and Mainstone (1985) found that ratees with whom supervisors had established relatively high task acquaintance (familiarity with the employee's job behavior) and personal acquaintance received more favorable performance ratings than other ratees. These ratees also exhibited higher objective sales productivity, thereby suggesting that differences in ratings were at least partially due to true differences among ratees rather than to rater bias.

Inter-rater reliability and agreement are also affected by opportunity to observe the ratee. For example, Rothstein (1990) found that the inter-rater reliability of job performance increased with increasing opportunity to observe the ratee. Heger (2007) found that observability (i.e., how easy or difficult a competency is to see, assuming that it is being engaged in by the ratee) was positively related to inter-rater agreement. Additionally, Haaland and Christiansen (2002) showed that the poor convergence of assessment center ratings is, in part, a result of correlating ratings from exercises that differ in the extent that behavior relevant to the target trait can be observed.

3.5.6. Ratee impression management tactics Several studies have illustrated that impression management tactics used by ratees can influence ratings (Ferris, Judge,

Rowland, & Fitzgibbons, 1994; Wayne & Ferris, 1990; Wayne & Kacmar, 1991). For example, Wayne and Ferris concluded that a subordinate's impression management tactics led to supervisors liking the subordinate more and rating the subordinate's performance more favorably.

3.5.7. Ratee self-appraisals In a lab experiment, Shore, Adams, and Tashchian (1998) showed that self-appraisals can influence supervisors' appraisals of

employees such that the supervisors' ratings are higher when they receive a favorable (rather than unfavorable) self-appraisal from the employee.

3.5.8. Similarity to the ratee and rater affect toward the ratee Research on the similarity-attraction paradigm (Berscheid & Walster, 1969; Byrne, 1971) has shown that there is a strong

correlation between similarity and interpersonal attraction. Alternatively, Rosenbaum (1986) has argued and provided evidence that attitudinal similarity does not lead to liking but that dissimilarity leads to repulsion. In either event, in the context of employment interviews, research has found a “similar-to-me” effect such that rater–applicant similarity in terms of demographic, attitudinal, and personality (i.e., conscientiousness) variables leads to more favorable judgments by raters (Baskett, 1973; Lin, Dobbins, & Farh, 1992; Rand & Wexley, 1975; Sears & Rowe, 2003).

In the context of performance ratings, Tsui and O'Reilly (1989) found that increasing dissimilarity in supervisor-subordinate demographic characteristics was associated with supervisors rating subordinates as being less effective, less personal attraction of supervisors for subordinates, and subordinates experiencing more role ambiguity. Pulakos and Wexley (1983) also found that managers and subordinates provided the most favorable evaluations of each other when both perceived themselves as similar. Antonioni and Park (2001a) found that rater–ratee similarity in conscientiousness was positively associated with peer ratings even after controlling for interpersonal affect.

A meta-analysis by Kraiger and Ford (1985) found that white raters gave more favorable ratings to white ratees than to black ratees, whereas black raters gave more favorable ratings to black ratees than to white ratees. However, Sackett and DuBois (1991), after presenting additional data and analyzing the Kraiger and Ford study more closely, challenged the conclusion that raters generally give more favorable ratings to members of their own race. In a study in the Israeli military, Fox, Ben-Nahum, and Yinon (1989) found that the accuracy of peer ratings was positively related to the rater's perceived similarity to those peers.

Conway (1998) showed that interpersonal affect accounts for significant method variance in performance appraisal ratings. For example, Tsui and Barry (1986) found that raters' interpersonal affect (liking) toward a ratee was positively related to the leniency (favorableness) of performance ratings. Also, raters with negative feelings toward the ratee displayed more halo (across dimensions) in their ratings. Antonioni (1999) found that, when effects of the managers' performance results were controlled, subordinates who expressed positive affect toward their managers gave more favorable upward appraisal ratings than subordinates who did not indicate strong positive liking. Antonioni and Park (2001b) found that the influence of rater interpersonal affect (liking) on the leniency of ratings was greater in upward and peer feedback than in downward feedback.

1019J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

4. Correlates of self–other rating agreement

According to Brutus, Fleenor, and Tisak (1999), discrepancies between self-ratings and others' ratings allow for a rare insight into a leader's interpersonal world. A number of theoretical arguments regarding the contribution of self-awareness, self-insight, and self-perception to individual and organizational outcomes have appeared in the literature (e.g., Atwater & Yammarino, 1997; London, 1995; Tsui & Ashford, 1994). Through the use of ratings generated by multi-rater instruments, the degree of agreement between self-perceptions and the perceptions of others can be employed to test such arguments. For example, using ratings from such instruments, Nilsen and Campbell (1993) found that SOA appears to be stable over time. They reported that discrepancies on skill-based, multi-rater instruments were related to discrepancies on personality-based, multi-rater instruments, suggesting that self-perception accuracy is a stable individual difference.

Atwater and Yammarino (1992) proposed that the degree and direction of SOA are related to various outcome measures that have important implications for both individuals and organizations. They assessed both predictors and outcomes of SOA and found that agreement was a moderator of leadership–performance relationships. Specifically, they found that leadership ratings of over- estimators, under-estimators, and in-agreement raters were predicted by different variables. Atwater and Yammarino concluded that SOA, as an individual difference variable, could moderate predictor–outcome relationships. Table 3 presents a summary of the findings on the correlates (predictors or outcomes) of SOA that are discussed below.

4.1. Correlates of self–other rating agreement at the individual level

Atwater and Yammarino (1997) proposed that leaders with discrepant ratings may misdiagnose their strengths and weaknesses, which can adversely influence their effectiveness. Leaders who rate themselves higher than others may set unrealistic goals for themselves and for their employees, resulting in negative outcomes for both the individual and the organization. Additionally, because these over-estimators believe their level of performance is already high, they may ignore developmental feedback and fail to improve their performance (Bass & Yammarino, 1991).

Much of the early research on SOA focused on the relationship between agreement and leader effectiveness (Atwater & Yammarino, 1992; Bass & Yammarino, 1991; Fleenor et al., 1996; Roush & Atwater, 1992; Van Velsor et al., 1993). In theory, this relationship could take many forms—a straightforward hypothesis is that leaders with congruent ratings are more effective than those whose ratings are incongruent. It appears, however, that the relationship between SOA and leader effectiveness is much more complex than initially conceptualized (Brutus et al., 1999).

4.1.1. Leader performance Several early studies (e.g., Atwater & Yammarino, 1992; Bass & Yammarino, 1991) found that SOA was related to higher

performance by leaders relative to those whose ratings disagreed with others. Van Velsor et al. (1993) found that over-estimators (those with self-ratings higher than others' ratings) received the lowest direct-report ratings of effectiveness compared to under- estimators (those with self-ratings lower than others' ratings) and in-agreement raters (those with self-ratings similar to others'

Table 3 Correlates of self–other rating agreement (SOA).

Topics Summary

Individual level

Atwater and Yammarino (1992) reported that SOA was related to higher performance by leaders versus those whose ratings disagreed with others. Van Velsor et al. (1993) found that over-estimators received lower ratings from subordinates compared to under-estimators and in-agreement raters. Furnham and Stringfield (1994) replicated Bass and Yammarino's (1991) findings that self-subordinate agreement was related to leader performance. Atwater and Yammarino (1997) posited that leaders with discrepant ratings misdiagnose their strengths and weaknesses, adversely affecting their effectiveness. Atwater et al. (1998) reported that SOA was related to performance outcomes and uncovered an underlying three-dimensional relationship between self- and others' ratings and effectiveness; however, Bazigos (1999) found that not all over-estimators were seen as less effective leaders. Atkins and Wood (2002) reported that high self-ratings were associated with poor performance in an assessment center, especially when others' ratings were low; however, a high level of SOA was not found to predict high performance, suggesting that the relationship between SOA and performance is not linear. Ostroff et al. (2004) found that SOA was related to performance, compensation, and organizational level, although rating patterns differed. Atwater et al. (2005) replicated earlier studies demonstrating that simultaneous consideration of self- and others' ratings is necessary for predicting performance. They found that the effect of self- and others' ratings in the prediction of performance differed between U.S. and European leaders. McCall and Lombardo (1983) found that inflated self-ratings were related to managerial derailment, whereas Bass and Yammarino (1991) reported that leaders whose self-ratings agreed with the ratings of their subordinates were more promotable. Atwater and Yammarino (1992) found that promotion recommendations from superiors were negatively related to effectiveness for over-estimators, positively related for in-agreement raters, and unrelated for under-estimators. London and Smither (1995) reported that accurate self-ratings were important for setting realistic goals. In a study of influence tactics, Berson and Sosik (2007) found that in-agreement/good leaders outperformed over-estimators and in-agreement/poor raters in championing a climate of innovation and quality. Jeffcoat (2000) reported that SOA was related to feedback fairness, while Kwan et al. (2008) suggested that SOA was a predictor of psychological adjustment.

Organizational level

Szell and Henderson (1997) investigated the effects of SOA on subordinate job satisfaction, organizational commitment, and variables posited to account for rating discrepancies. They found that agreement was positively related to subordinate job satisfaction and organizational commitment, specifically with subordinates' identification with the organization. Halverson et al. (2005) suggested that the effects of SOA may be cumulative over time and have important implications for organizations; therefore, higher levels of SOA will likely result in positive outcomes for organizations.

1020 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

ratings). Both Van Velsor et al. and Atwater, Roush, and Fischthal (1995), however, reported that leaders who rated themselves lower in relation to others' ratings were rated higher by direct reports.

In another early study, Furnham and Stringfield (1994) examined self- and subordinate ratings for almost 400 Chinese and Caucasian leaders. These ratings, as well as a measure of self-subordinate agreement, were compared to superior ratings. The authors found several relationships between self- and subordinate ratings, few relationships between self- and superior ratings, and numerous relationships between subordinate and superior ratings. However, there were few cross-cultural differences. Regression analyses indicated that the best predictors of superior ratings were subordinate ratings of innovation and the measure of rating agreement. This study replicated the findings of Bass and Yammarino (1991) that self-subordinate agreement is related to leader performance.

In an examination of leaders who overestimate their performance, Bazigos (1999) focused on leaders whose self-ratings exceeded their followers' ratings of the leaders' performance. With a sample of 309 leaders and 1394 direct reports from several organizations and industries, Bazigos reported that a subset of high-performing leaders reside in the over-rater category. This subset was large enough to significantly raise the group's overall performance. Greater variability in direct-report ratings was found for the over-estimator group, supporting the hypothesis that not all over-estimators are seen as less effective leaders.

A few studies have suggested that SOA is unrelated to leader performance (e.g., Brutus et al., 1999; Fleenor et al., 1996). The authors of these studies argue that the confounding of agreement group with performance level results in a spurious relationship between these two factors. Using a six-group categorization model, Fleenor et al. found that when the performance level of focal leaders was accounted for, SOA was not related to leader performance. Using polynomial regression, Brutus et al. found that leader performance was predicted by an additive function of peer and self-ratings, with peer ratings being the most influential; however, agreement between self- and peer ratings had no relationship with performance. They concluded that simultaneous consideration of both self- and others' ratings is of little importance; the ratings of others are the most important factor in explaining leadership outcomes.

On the other hand, several studies, including more recent investigations, have found positive relationships between SOA and leader performance. Using polynomial regression with a sample of over 1400 leaders who were rated by themselves, direct reports, peers, and supervisors, Atwater et al. (1998) found that both self- and others' ratings were positively related to performance outcomes. They also uncovered an underlying three-dimensional relationship between self-ratings, others' ratings, and effectiveness. In a related study, Ostroff et al. (2004) investigated the relationship between SOA and several variables, including performance. Using a sample of 3217 leaders from 527 organizations, Ostroff et al. employed polynomial regression to examine the relationships between SOA and important outcomes. Results indicated that SOA was positively related to performance, compensation, and organizational level, although rating patterns differed. The results of the Atwater et al. and Ostroff et al. studies demonstrate that the association between self-ratings, others' ratings, and performance is somewhat more complex than previous conceptualizations of this relationship.

In an investigation of the relevance of SOA in countries other than the United States, Atwater et al. (2005) replicated earlier studies demonstrating that simultaneous consideration of self- and others' ratings is necessary for predicting leadership performance. They investigated the impact of SOA on performance for 2732 leaders from five European countries and found that the effect of self- and others' ratings in the prediction of performance differed between U.S. and European leaders. The simultaneous inclusion of both self- and others' ratings was found to be generally less useful in other countries than in the U.S. Furthermore, the effects of SOA were found to vary among the European countries.

4.1.2. Assessment center performance Using an independent performance measure, Atkins and Wood (2002) examined the relationship between SOA and

performance in an assessment center. They found that self-ratings from a multi-rater instrument were negatively related to assessment center ratings—those who rated themselves the highest were the poorest performers in the assessment center. Leaders who placed themselves in the mid-range of the rating scale were more likely to be high performers in the assessment center than those rating themselves on the upper or lower ends of the scale. While it was not the case that others' ratings alone were predictive of performance (cf. Fleenor et al., 1996), a high level of SOA was not found to predict high performance (cf. Atwater et al., 1998). The results of polynomial regression indicated that a new pattern of relationships had emerged in which high self-ratings were associated with the poorest performance, especially when others' ratings were low. Superior ratings were found to predict the performance of over-estimators, suggesting that supervisors were able to see through inflated self-ratings. Superior ratings, however, were less successful at predicting the performance of under-estimators, indicating that the performance of more modest participants was under-estimated by their superiors. Additionally, peers were found to overestimate the performance of poor performers. The results of Atkins and Wood suggest that the relationship between SOA and performance is not linear.

4.1.3. Performance improvement Smither et al. (1995) found that leaders who provided low self-ratings did not improve their performance after receiving low

ratings from their direct reports. As indicated by self-consistency theory (Korman, 1976), it appears that leaders are satisfied with feedback that is consistent with their self-perceptions, even if these self-perceptions are negative.

In an investigation of whether agreement among raters influences performance improvement, Johnson and Ferstl (1999) examined how self-ratings change after feedback. Self-ratings and direct report ratings were collected from 1888 managers at two points in time one year apart. Using polynomial regression, the authors found that leaders who over-rated themselves relative to how others rated them tended to improve their performance from one year to the next, while under-estimators tended to decline.

1021J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Self-ratings tended to decrease for over-estimators and increase for under-estimators, but this effect was not constant throughout the range of self-ratings. According to Johnson and Ferstl, these findings are consistent with the predictions of self-consistency theory (Korman, 1976); however, self-enhancement theory may be a viable alternative explanation (Dipboye, 1977).

4.1.4. Promotion According to Atwater and Yammarino (1997), in-agreement raters, who accurately diagnose their strengths and weaknesses,

are able to make more effective decisions related to their careers. For example, Bass and Yammarino (1991) found that naval officers whose self-ratings agreed with the ratings of their subordinates attained higher ranks and were rated as more promotable by their superior officers. Atwater and Yammarino (1992) reported that promotion recommendations from superiors were negatively related to independent leadership measures for over-estimators, positively related for in-agreement raters, and unrelated for under-estimators

Using a sample of over 2000 military officers, Halverson et al. (2005) found that greater insight into one's own performance was positively related to promotion rate. The authors also reported that agreement between self and subordinates was more important in predicting promotion rates than was agreement between self and superior and self and peers.

4.1.5. Derailment McCall and Lombardo (1983) found that inflated self-ratings were related to managerial derailment (i.e., likelihood a leader

will plateau, be demoted, or be fired). Using ratings of managerial derailment for a sample of over 1700 European leaders and over 1900 American leaders, Gentry et al. (2007) found differences between self-ratings and others' ratings on the likelihood that leaders would derail. Superior ratings were used to measure likelihood of derailment, whereas direct report and peer ratings were used to determine self–other discrepancies. These discrepancies appeared to widen as managerial level increased, primarily as the result of inflated self-ratings. The authors reported differences in SOA between European and American managers on three derailment factors. Self–other discrepancies were found to be greater for the American managers than for the European managers on all three derailment factors.

4.1.6. Goal setting According to Ashford (1989), individual goal setting may be affected by the level of agreement between self- and others'

ratings. Individual leaders set goals and then evaluate their success in meeting these goals. As noted by London and Smither (1995), SOA is key to leaders' perceptions of their goal-performance discrepancies. Accurate self-ratings are important to help individuals decide how to allocate their efforts toward meeting these goals. For example, over-estimators may fail to set developmental goals because they believe their strengths are not recognized by other raters, and they fail to recognize their weaknesses that are perceived by others. In-agreement raters who receive low ratings from others may not be motivated to achieve their goals, whereas in-agreement raters who receive high ratings from others should set realistic developmental goals (Atwater & Yammarino, 1997).

4.1.7. Influence tactics Berson and Sosik (2007) examined the extent to which others' ratings of influence tactics were related to the self-awareness of

leaders. Self-awareness was operationalized by categorizing leaders as over-estimators, under-estimators, in-agreement/poor or in agreement/good raters based on the differences between self-ratings and direct report ratings of charismatic leadership. Influence tactics were categorized as soft (i.e., consultation, ingratiation and inspirational appeals), hard (i.e., legitimating, pressure and exchange), or rational persuasion. Results indicated that under-estimators tended to use more rational persuasion than over-estimators and in-agreement/poor leaders. Over-estimators tended to use fewer soft tactics than under-estimators and in-agreement/good managers. In-agreement/good leaders tended to use more exchange tactics and outperformed over- estimators and in-agreement/poor managers in championing a climate of innovation and quality.

4.2. Feedback fairness

To examine the relationship between self–other rating discrepancy and feedback fairness, rater fairness, and feedback usefulness, Jeffcoat (2000) asked 182 participants to complete an in-basket exercise and to rate their performance and self- efficacy. Each participant ostensibly received a feedback report from either one superior, one peer, three peers, or one superior and two peers, resulting in four feedback groups. In actuality, three trained raters evaluated in-basket performance and provided the feedback for the participants. After receiving feedback, the participants rated source credibility, feedback fairness, source fairness, and feedback usefulness. The ratings on feedback fairness, rater fairness, and feedback usefulness were then regressed on rater credibility, SOA, self-efficacy, and feedback group. Source credibility was found to predict feedback fairness, source fairness, and feedback usefulness, and rating discrepancy was found to be inversely related to feedback fairness.

4.2.1. Psychological adjustment Kwan et al. (2008) conducted an investigation of the relationship between self-enhancement (overly positive self-perceptions)

and psychological adjustment. The longstanding view has been that accurate self-perceptions are keys to effective functioning. Using a component-based approach to assessing self-enhancement, the authors found that self-enhancing individuals were low in psychological adjustment, suggesting that SOA may be a predictor of psychological adjustment.

1022 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

4.3. Correlates of self–other rating agreement at the organizational level

4.3.1. Job satisfaction Szell and Henderson (1997) investigated the effects of SOA on subordinates' job satisfaction, organizational commitment, and

on variables that were posited to account for rating discrepancies. Hypotheses were derived from research on role theory, impression management, and self-awareness. The sample consisted of 83 leader/follower dyads who rated each other's performance and motivation. Subordinates also completed standard measures of job satisfaction and organizational commitment. Results indicated that agreement was positively related to subordinate job satisfaction, specifically satisfaction with supervision, social, and growth aspects of their job. Organizational commitment was also positively correlated with rating agreement, specifically with subordinates' identification with the organization.

4.4. Measurement issues and data analytic techniques in self–other rating agreement

In this section of the article, we review measurement issues and data analytic techniques generally encountered when conducting research on SOA. Consistent with previous work on measurement issues in the multisource rating area (e.g., see Yammarino, 2003), we review topics on operationalizing SOA, assessing rating similarity and agreement, and examining measurement equivalence of ratings across rating sources. We do not cover traditional methods for assessing reliability and validity of measures, but we assume that these issues will be addressed by the researcher. In addition, we review the majority of data analytic techniques (or methods) that have been employed in past studies of SOA. These include difference scores, polynomial regression, multivariate regression, categories of agreement, WABA, and hierarchical linear modeling (HLM). Tables 4 and 5 provide summaries of the measurement issues and data analytic techniques discussed in this section.

4.5. Measurement issues in self–other rating agreement research

4.5.1. Operationalizing SOA Researchers have generally assessed individuals' self-perceptions using a social comparison or self-insight approach. The

rationale for the social comparison approach was derived from Festinger's (1954) social comparison theory and requires comparing the self-ratings of individuals to their ratings of others. Conversely, the self-insight method is based on Allport's (1937) notion of self-insight and involves comparing the self-ratings of individuals to others' ratings of them (Kwan et al., 2004, 2008). While both methods have been used in the self-enhancement/SOA literature, the self-insight method has primarily been used to operationalize SOA in multisource rating studies (e.g., Atwater et al., 1998; Fleenor et al., 1996).

In recent reviews, however, Kwan et al. (2004, 2008) have called attention to problems associated with the aforementioned approaches. Their primary criticisms were that these approaches fail to take into account, or ignore, certain factors of interpersonal perception. Namely, the social comparison index does not account for how others rate a target individual, while the self-insight index does not account for the way the target typically rates others. To illustrate these problems, Kwan et al. (2004) formulated a hypothetical example using information from Charles Darwin's autobiography. Specifically, they proposed that Darwin would have rated himself average on ability (a 7 on an 11-point scale), rated others less able than himself (a rating of 6), and experts

Table 4 Measurement Issues in self–other rating agreement (SOA) research.

Topics Summary

SOA operationalizations In the SOA and self-enhancement literature, SOA has been operationalized using a social comparison method (compare the self-ratings of individuals to their ratings of others), a self-insight method (compare the self-ratings of individuals to others' ratings of them), and a componential approach (compute this index by subtracting or removing perceiver and target effects from one's self-rating). Although the self-insight method has most frequently been used in SOA studies, Kwan et al. (2004, 2008) recently pointed out criticisms of both the self-insight and social comparison approaches. They suggest that the componential approach is superior.

Rater similarity and agreement

Past research has supported the discrepancy hypothesis—the notion that rating similarity is high within rating groups (e.g., peers) but is low between rating groups (e.g., peers versus bosses). However, recent research by LeBreton et al. (2003) challenged this view and proposed an alternative called the restriction of variance hypothesis. This hypothesis purports that rating similarity between rater groups has erroneously been deemed to be low in past research because authors have investigated rating similarity using correlation-based statistics on range restricted performance ratings. They demonstrated that when using correlation-based statistics, rating similarity is low both within and between rating sources, whereas when using non-correlation-based statistics such as rWG, rating similarity is high both within and between rating groups; these results supported their restriction of variance hypothesis. LeBreton et al.'s findings call into question the meaningfulness of the SOA tradition of comparing self-ratings to the ratings of direct reports, peers, and bosses separately rather than comparing self-ratings to “others” (i.e., direct reports, peers, and bosses combined).

ME/I ME/I of ratings should be established prior to testing substantive hypotheses. ME/I can be investigated using item response theory techniques or logistic regression but is most commonly assessed using multiple-group confirmatory factor analysis (CFA) in organizational research. Although results of past studies are somewhat mixed, they generally indicated that ME/I can be found in multisource performance ratings. In their 2005 study, Woehr et al. illustrated the limitations of using traditional multiple-group CFA approach to assess ME/I. As an alternative, they proposed that future research use multi-trait multi-method (or rater) CFA models to assess ME/I.

Table 5 Data analytic techniques in self–other rating agreement (SOA) research.

Topics Summary

Difference scores Difference score indices were used in the early years of SOA research, but their popularity decreased quite substantially after this approach was critiqued by Edwards (1993, 1994, & 2002) and Edwards and Parry (1993). Example criticisms include that difference scores cannot be interpreted unambiguously, and they generally have lower reliability than the separate component measures used to compute them.

Polynomial regression Edwards (1993, 1994, & 2002) and Edwards and Parry (1993) recommended using polynomial regression instead of difference scores when investigating SOA as an independent or predictor variable. This technique overcomes the limitations of difference scores by keeping component measures (e.g., self- and others' ratings) separate and by incorporating higher-order terms (self-ratings squared, others' ratings squared, and the cross product of self- and others' ratings) to examine the relationship between SOA and outcomes in three dimensions.

Multivariate regression Edwards (1995) proposed using multivariate regression instead of using difference scores when investigating SOA as a dependent variable. In this analysis, self- and others' ratings are kept separate and are simultaneously regressed on continuous predictors and/or a mix of continuous and categorical predictors. If the multivariate tests are significant, relationships between the predictors and each dependent variable (self-ratings, peer ratings, etc.) can be examined to determine the nature of SOA effects (see Ostroff et al., 2004).

Categories of agreement Atwater and Yammarino (1992) initially proposed classifying individuals into three rating categories (under-estimators, over-estimators, and in-agreement raters) to assist with analyzing SOA data. In 1997, they expanded this model to include four groups by distinguishing between in-agreement/good and in-agreement/poor raters. Around the same time, Fleenor et al. (1996) expanded the four-group model to include six groups: under-estimators/good, under-estimators/poor, over-estimators/good, over-estimators/poor, in-agreement/good, and in-agreement/poor. Although this approach has intuitive appeal, it inherits many of the limitations of difference scores and artificially categorizes continuous ratings. As such, this approach is generally not used in isolation.

WABA Like polynomial and multivariate regression, WABA overcomes the limitations of difference scores and categories of agreement. This technique assesses SOA by analyzing where the source of variation/covariation in scores primarily lies. In other words, is variation/covariation within, between, or both/neither within and between self–other dyads? WABA can yield four patterns of results for self–other dyads: patterned agreement, patterned disagreement, nonpatterned disagreement, or a lack of variation/ covariation.

HLM HLM can be used to test hypotheses regarding self- and other scores in nested data. For example, Atwater et al. (2009) used HLM to examine the nature of the relationships between self- and others' ratings of leadership and whether the magnitude of these relationships varied as a function of cultural context variables, which served as cross-level moderators.

1023J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

would have rated him extremely capable (a rating of 11). If one computes a social comparison index in this example, Darwin would be classified as a self-enhancer. Using the self-insight method, however, Darwin would be labeled a self-effacer. This simple example illustrates that the social comparison index did not consider how others perceived Darwin's intelligence (a rating of 11), whereas the self-insight method did not account for how Darwin perceived the intelligence of others (a rating of 6).

To overcome these limitations, Kwan et al. (2004, 2008) proposed using a componential approach to operationalize SOA. As they described, this approach builds on the Social Relations Model (SRM; Kenny, 1994; Kenny & La Voie, 1984) by defining self- perception as a form of interpersonal perception in which the self is both the perceiver and the target perceived by others. Specifically, the componential approach implies that self-perception (Xss) can be decomposed into three variable components (a perceiver effect or Ps, target effect or Ts, and relationship with self or Rss) and a constant (Cs; the mean of all self-ratings). The first variable component, the perceiver effect, refers to whether one evaluates others harshly or leniently. The second component, the target effect, refers to whether one is perceived negatively or positively by others. Finally, the relationship with self can be computed by subtracting the perceiver and target effects (and a constant) from one's observed self-perception or self-rating.

Interestingly, individuals can perceive themselves positively because (a) they tend to perceive others positively (Ps), (b) they are perceived positively by others (Ts), and (c) they have an excessively positive view of themselves (Rss). In this context, given that Rss is computed by subtracting (or removing) the perceiver and target effects from one's self-rating, Kwan et al. (2004, 2008) noted that Rss is the only relevant indicator of self-perception bias as it captures one's idiosyncratic self-view. Unlike the social comparison and self-insight indices, Kwan et al. contended that Rss can be viewed as an unconfounded index of SOA. If Rss is negative, it is indicative of self-effacement (under-estimation), whereas positive values are indicative of self-enhancement (over- estimation).

In virtually each experiment they conducted, Kwan et al. (2004, 2008) concluded that the componential index was a more appropriate measure of self-enhancement compared to the other two indices. For example, Kwan et al. (2004) noted that, similar to the Darwin example that they provided, the social comparison index incorrectly classified some participants as self-enhancers despite the fact that they performed better on tasks than other group members in their experiment. The self-insight method incorrectly classified some participants as self-enhancers when they actually displayed a general person-positivity effect (i.e., perceiving all others including oneself positively).

In sum, Kwan et al. (2004, 2008) proposed using a componential approach to operationalizing SOA. They suggested that this approach represents a new direction in the self-enhancement or SOA literature. However, while the componential approach may overcome some of the limitations of the social comparison and self-insight indices of SOA, Kwan et al. used a difference score approach to operationalize SOA in their studies and thus collapsed component measures into a single index (Rss). As such, the componential approach as used in this manner appears to be subject to some of the difference score limitations outlined by Edwards (1993, 1994, 2002) and Edwards and Parry (1993).

1024 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

4.5.2. Rater similarity and agreement In many past studies on SOA (e.g., Atwater et al., 1998, 2005), researchers have compared target leaders' self-ratings to the

ratings of different rating groups. This has primarily been done because of the belief that individuals from different rating groups provide unique perspectives on a leader's performance. In other words, it is assumed that while rating similarity within rating source is high, rating similarity between rating sources is low. This is commonly called the discrepancy hypothesis and is purported to occur because leaders do not act similarly around all rating groups, different rating groups have different opportunities to observe leaders, rating groups attend to different aspects of leader performance, and rating groups place different levels of importance on identical leader behaviors (Borman, 1997; LeBreton et al., 2003; Murphy & Cleveland, 1995; Tornow, 1993). As summarized in LeBreton et al., past research has provided some empirical support (e.g., Conway & Huffcutt, 1997; Harris & Schaubroeck, 1988) for this hypothesis.

In contrast, LeBreton et al. (2003) proposed the restriction of variance hypothesis. This hypothesis is based on the premise that the low level of rating similarity observed to exist between rating groups is an artifact of the analytic techniques and data used in previous studies. That is, past studies have examined rating similarity using correlation-based statistics on range restricted performance ratings. LeBreton et al. proposed that performance ratings are generally restricted in range due to rating biases (e.g., central tendency and leniency), effective organizational interventions (e.g., selection, performance, and training) that reduce between target leader variability, and the fact that managers engage in similar sets of behaviors across time and situations, which may reduce both within and between target leader variability. Given that restricted variability in ratings in turn attenuates correlations, LeBreton et al. argued that past studies have erroneously concluded that rating similarity between rating groups is low. In fact, their restriction of variance hypothesis stated that rating similarity indexed via Pearson correlations and intraclass correlations (ICCs) would be low within and between rating groups, while inter-rater agreement estimates indexed via rWG (James, Demaree, & Wolf, 1984, 1993), which is not correlation based, would be large in magnitude both within and between rating sources.

These two competing hypotheses (discrepancy and restriction of variance) were tested in LeBreton et al.'s (2003) study by applying correlations, ICCs, and rWG (James et al., 1984, 1993) to multisource ratings. As the authors predicted, their correlations and ICCs were generally small in magnitude, indicating fairly low levels of rating similarity within and between rating sources. Also, as expected, the rWG index revealed high levels of agreement (with many values close to or above .70) for both within-group and between-group ratings. In short, these findings indicated that rating agreement is much higher within and between rating sources than previously thought, providing strong support for the restriction of variance hypothesis.

Based on these findings, one may question the meaningfulness of the SOA tradition of separately comparing self-ratings to the ratings of peers, subordinates, and bosses instead of comparing self-ratings to all rating groups combined (i.e., “others”). However, this issue is not clear cut and has continued to be a topic of debate. For example, recent studies using multi-trait–multi-rater (MTMR) methodology have strongly advocated for the ecological perspective (e.g., see Lance, Hoffman, Gentry, & Baranik, 2008). Like the discrepancy hypothesis, this perspective implies that rating groups can often have different, yet equally valid views of a leader's performance. If one believes in this perspective, comparing self-ratings to each rating group may be justified.

Whether self-ratings are compared to multiple rating groups or to all rating groups combined (i.e., self versus all others), one needs to demonstrate sufficient within-group inter-rater agreement in the chosen comparison group(s) prior to aggregating ratings within rating group(s) and testing SOA hypotheses. To review indices available for computing agreement, we recommend reading LeBreton and Senter's (2008) article. They provided a nice overview of these indices, answered common questions about inter-rater agreement, and provided SPSS syntax and tutorials for computing agreement indices.

4.5.3. Measurement equivalence/invariance (ME/I) ME/I can be presumed to exist in multisource ratings when measures display very similar psychometric properties across

rating groups (Barr & Raju, 2003). When this occurs, one can conclude that differences in ratings between rating sources are not due to, or created by, the multisource instrument itself. This in turn gives researchers more confidence that differences in ratings can be meaningfully and accurately interpreted as true differences in how rating groups perceive target leaders (Woehr, Sheehan, & Bennett, 2005). Although ME/I is sometimes assessed via item response theory (IRT) techniques (Barr & Raju; Facteau & Craig, 2001; Meade & Lautenschlager, 2004) and logistic regression (Penny, 2003), in organizational research, it is most commonly assessed using multiple-group confirmatory factor analysis (CFA; Vandenberg & Lance, 2000; Woehr et al.).

Using multiple-group CFA procedures, numerous studies have examined the ME/I of multisource performance ratings. While the majority of these studies (e.g., Cheung, 1999; Facteau & Craig, 2001; Maurer, Raju, & Collins, 1998) revealed that performance measures were generally equivalent across rating sources, a study conducted by Lance and Bennett (1997) yielded mixed results. Lance and Bennett factor analyzed multisource ratings on a global performance dimension using eight U.S. Air Force samples. In five of their samples, Lance and Bennett found that ME/I did not exist across rating groups. Similarly, several studies using IRT- based methods (e.g., Barr & Raju, 2003) found evidence of a lack of ME/I for multiple items comprising multisource assessments. In short, these studies provided somewhat mixed results, although they implied that ME/I can often be found in multisource ratings.

In a more recent article, Woehr et al. (2005) critiqued the CFA approach used in previous studies. In particular, they noted that traditional multiple-group CFA procedures model ratings as a function of the underlying performance dimension(s) (and unique variance) and test whether rating source moderates the relationships between items and the latent performance factor(s). Unfortunately, as they indicated, this approach fails to consider the direct effects of rating source on performance ratings because rater factors are not directly modeled. Stated differently, while traditional multiple-group CFA allows an examination of the direct effect of performance dimensions and Performance Dimension×Rating Source interactions, it does not permit the examination of

1025J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

the main effect of rating source. Additionally, because the main effect of rating source is omitted from the model, estimates of performance dimension and Performance Dimension×Rating Source effects may be biased.

As an alternative to the traditional multiple-group CFA approach, Woehr et al. (2005) proposed using a multi-trait–multi-rater (MTMR) CFA framework. In their study, they compared a CFA performance dimension-only model to a MTMR CFA model using ratings on eight performance dimensions in a U.S. Air Force sample. The CFA performance dimension-only model was specified with only performance factors (and unique effects) directly included. In contrast, in the MTMR CFA model, ratings were modeled as a function of their performance dimension, rating source, and unique effects. Specifically, each of the eight performance factors was indicated by self, peer, and supervisor ratings of that dimension, whereas the three rating source factors were indicated via their respective ratings on all eight dimensions. Thus, while performance dimensions represented traits, rating sources represented “method” or rating factors. In this model, inter-correlations among traits and inter-correlations among rating source factors were permitted. However, trait and rating source factors were not allowed to correlate with each other. This model served as a test of configural (factor structure) invariance. Woehr et al. also evaluated two additional hierarchically nested models that were variations of the aforementioned MTMR CFA model. These two models provided tests of metric (item factor loadings) and error variance invariance.

Findings from their study revealed that the MTMR CFA model fit the data much better than did the traditional multiple-group CFA model. Therefore, Woehr et al. (2005) suggested that MTMR CFA should be used in future ME/I investigations of multisource ratings. They also found support for configural and metric invariance, meaning that performance ratings reflected the underlying performance dimensions equivalently across rating sources. Using an improved CFA approach, their results therefore provided initial evidence that meaningful comparisons on performance can be made across rating groups. Nonetheless, we encourage researchers to investigate the ME/I of multisource ratings (using MTMR CFA models) in their own studies prior to testing substantive SOA hypotheses.

4.6. Data analytic techniques in SOA research

4.6.1. Difference scores As described in Atwater and Yammarino (1997), Atwater et al. (1998), and Edwards (2002), difference score indices were

commonly used in the initial years of SOA research to assess degree of agreement as well as to test how rating agreement related to importantoutcomes. Despite its initial popularity, however, the use of differencescores in SOA research decreased after Edwards (1993, 1994, 2002) and Edwards and Parry (1993) provided critiques of this approach. As they pointed out, whether using profile similarity indices or algebraic, absolute, or squared difference indices, the difference score approach is statistically flawed for multiple reasons.

Some of the criticisms of difference scores offered by Edwards (1993, 1994, 2002) and Edwards and Parry (1993) were that they cannot be interpreted unambiguously, they confound the effects of their components (which conceals the contributions of the components to the relationships between the difference score indices and outcomes), they generally have lower reliability than the separate component measures used to compute them, and they reduce inherently three-dimensional relationships among component measures and outcomes to two dimensions. Given the limitations of this approach, Edwards and Edwards and Parry recommended using two alternative procedures (polynomial and multivariate regression) for testing rating congruence hypotheses. Both of these data analytic techniques are described next.

4.6.2. Polynomial regression Edwards (1993, 1994, 2002) and Edwards and Parry (1993) recommended using polynomial regression when treating SOA as

an independent (or predictor) variable. For example, it would be appropriate to use this data analytic technique when examining if rating congruence on managerial skills predicted managerial performance ratings. Importantly, unlike the difference score approach, polynomial regression enables researchers to keep component measures such as self- and others' ratings separate. This procedure also incorporates higher-order terms such as squared self- and others' ratings that are necessary for appropriately examining the relationships between self- and others' ratings and outcomes in three dimensions.

In the SOA literature, polynomial regression has typically been conducted in two stages. In the first stage, researchers generally conduct a hierarchical regression analysis in which a continuous outcome is regressed on self-ratings and others' ratings in step 1, and squared self-ratings, squared others' ratings, and the interaction of self-ratings and others' ratings are added in step 2. If the ΔR2 between steps 1 and 2 is significant, it indicates that nonlinear effects may be present. Next, in stage 2, researchers generally use the obtained regression coefficients to conduct response surface tests (Atwater et al., 1998; Edwards, 1993, 1994; Edwards & Parry, 1993; Halverson et al., 2005; Ostroff et al., 2004).

Response surface tests permit the examination of salient features of the surfaces corresponding to the polynomial regression equations and help clarify the nature of the relationships indicated in the regression model (Edwards, 2002). As described in Atwater et al. (1998) and Ostroff et al. (2004), these tests can be used in examining interesting SOA hypotheses. For instance, SOA researchers have traditionally investigated the slope and curvature of the line of perfect agreement (where self-rating=other rating; S=O) to determine if outcomes (e.g., performance ratings) increase or decrease as both self-ratings and others' ratings simultaneously increase (and thus are in agreement). These tests also reveal whether the slope of the S=O line is linear or nonlinear. In addition, researchers have traditionally examined the slope and curvature of the line perpendicular to the line of perfect agreement, or S=−O (e.g., self-ratings=5 and others' ratings=1), to determine how under-estimation or over- estimation of self-ratings relates to an outcome (e.g., do over-estimators have higher performance than under-estimators; see Halverson et al., 2005; Ostroff et al.).

1026 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

4.6.3. Multivariate regression This analytic technique is used when examining SOA as a dependent (or outcome) variable. For example, it may be used when

examining if gender and age predict rating congruence in managerial performance ratings. Importantly, multivariate regression is superior to difference scores because it does not combine self-ratings and others' ratings into a single index. Instead, with this analysis, self-ratings and others' ratings are kept separate and are examined simultaneously (Edwards, 1995). In the SOA literature (e.g., see Ostroff et al., 2004), this procedure has been applied by jointly regressing each set of self- and others' ratings (e.g., self- and peer ratings) on continuous predictor variables or on a mix of continuous and categorical predictors. This analysis yields several tests of interest for interpreting SOA effects.

The first test to interpret is the omnibus multivariate test, which is based on an overall Wilks' Lambda. This is a test of the multivariate association between the set of predictors and the set of self- and others' ratings and represents a joint test of the equations for self- and others' ratings. If this test is significant, the canonical correlation between the set of predictor and criterion variables is generally examined to determine the magnitude of their relationship (see Ostroff et al., 2004). Subsequently, multivariate tests based on Wilks' Lambda are conducted to assess the effects of each predictor on the set of self- and others' ratings. These tests are generally used to make two assessments: (1) are equal but opposite effects occurring?, and (2) does a predictor have a significant effect on self- and others' ratings when these ratings are considered jointly? Equal but opposite effects would be plausible if a non-significant Wilks' Lambda were found but a predictor's regression coefficient was significantly different from zero. If, however, the Wilks' Lambda is statistically significant, the predictor is related to the self- and others' ratings when considered jointly (see Edwards, 1995; Ostroff et al., 2004).

Next, for each significant predictor, additional regression analyses are conducted to examine the relationship between the predictor and corresponding self- and others' ratings (the dependent variables) separately. Regression coefficients and graphs are subsequently examined to determine the source (self-ratings, others' ratings, or both) of any rating discrepancies that exist. For example, the predictor may be positively related to self-ratings but not to others' ratings. In this case, over-rating would be occurring due to higher self-ratings (see Ostroff et al., 2004).

4.6.4. Categories of agreement In 1992, Atwater and Yammarino introduced the idea of using rating agreement categories to assist with analyzing SOA data.

This approach requires computing difference scores between self- and others' ratings, calculating the mean and standard deviation of the difference scores, and subsequently classifying individuals into groups based on the magnitude of their self–other differences. Atwater and Yammarino initially recommended using three rating agreement categories with this approach. They included: (1) under-estimators (individuals whose difference scores were more than one-half standard deviation below the mean self–other difference), (2) over-estimators (individuals whose difference scores were more than one-half standard deviation above the mean self–other difference), and (3) in-agreement raters (individuals with difference scores within one-half standard deviation of the mean self–other difference).

In 1997, Atwater and Yammarino expanded their 1992 model to include four categories by distinguishing between in- agreement/good raters (i.e., individuals with self- and others' ratings in agreement that they were performing well) and in- agreement/poor raters (i.e., individuals with self- and others' ratings in agreement that they were performing poorly). Atwater and Yammarino recommended making this distinction because they hypothesized that it was relevant to understanding individual and organizational outcomes. Around the same time, Fleenor et al. (1996) expanded the four-group categorization model to six groups (Fleenor et al. cited the Atwater & Yammarino paper as in press, which explains the earlier publication date of their article). Following Atwater and Yammarino's treatment of the in-agreement groups, Fleenor et al. distinguished between good and poor performers for under-estimators, over-estimators, and in-agreement raters, resulting in six categories of agreement. They included under-estimators/good, under-estimators/poor, over-estimators/good, over-estimators/poor, in- agreement/good, and in-agreement/poor.

Agreement categories formed using the aforementioned models have frequently been used as independent variables in SOA research. In these cases, statistics such as univariate analysis of variance (ANOVA) and multivariate analysis of variance (MANOVA) have been employed to test for category differences in continuous individual or organizational outcomes (e.g., Fleenor et al., 1996; Halverson et al., 2005). While the categorization approach has intuitive appeal, it unfortunately inherits the problems associated with difference scores and artificially categorizes continuous ratings (Brutus et al., 1999). As such, agreement categories are generally no longer used as a primary method for testing SOA hypotheses.

4.6.5. WABA Another technique that overcomes the limitations of difference scores and categories of agreement and is useful for analyzing

SOA data is WABA. As its name implies, this technique assesses SOA by analyzing where the source of variation/covariation in scores primarily lies. In other words, is variation/covariation within, between, or both/neither within and between self–other dyads? More specifically, Atwater and Yammarino (1997) described how WABA may be performed in several steps to analyze self- and other scores. In the first step, self- and other scores are assessed at the dyad-level for each variable of interest to ascertain whether the variation in scores is within, between, or within and between self–other dyads. Next, self- and other scores are assessed at the dyad-level for each pair of relationships between variables of interest to determine if the covariation among variables is within, between, within and between, or neither within nor between the self–other dyads. In the final step, the WABA equation is used to combine results from the two previous steps, and within and between dyad components are then examined to formulate overall SOA conclusions.

1027J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

According to Atwater and Yammarino (1997), WABA analyses can yield four types of results in terms of the source of variation/ covariation for self–other dyads. First, the variation/covariation could largely be between self–other dyads such that self- and others' ratings are largely congruent for each manager, but differences exist in self–other dyad ratings across managers (Atwater & Yammarino). This is called patterned agreement (Yammarino, 2003). Second, the variation/covariation could largely be within self–other dyads. As an example, either under-estimation (self-ratings are lower than others' ratings) or over-estimation (self- ratings are higher than others' ratings) could be systematically occurring across multiple dyads. This is called patterned disagreement. Third, the variation/covariation could be within and between the dyads such that there are no systematic relationships between self- and other scores (Atwater & Yammarino). This represents a lack of agreement or nonpatterned disagreement (Yammarino). Finally, there be could no variation/covariation within or between the dyads. For example, each dyad could have the same self- and other scores (Atwater & Yammarino). If one obtains results similar to those described in the latter two conditions, Atwater and Yammarino indicated that other levels of analysis should be investigated as the self–other dyads do not meaningfully aid one's understanding.

In short, WABA is useful for analyzing the level of analysis at which agreement exists (for example, group or dyad), the patterns of agreement that are occurring, and the magnitude of such agreement. As previously explained, agreement can be assessed for one variable or for the relationship between two or more variables. Additionally, agreement can be assessed within rating groups (e.g., peers), between self–other rating groups (e.g., self versus peers), or between other–other rating groups (e.g., peers versus subordinates; Yammarino, 2003). For those interested in learning more about WABA, we recommend reading Dansereau and Yammarino (2000), Yammarino (1998, 2003), and Yammarino and Markham (1992) as a starting point.

4.6.6. HLM One final technique to consider is HLM. Although describing this procedure in detail is beyond the scope of this article, we

highlight a recent study that illustrates how HLM might be used to test hypotheses regarding the nature of the relationships that exist between self- and others' ratings in nested data. Specifically, Atwater et al. (2009) used HLM to examine relationships between self- and others' ratings of leadership to ascertain if these relationships varied as a function of cultural context. Most of Atwater et al.'s hypotheses specified that the relationships between self- and others' ratings of leadership would be positive and that these relationships would be stronger in cultures considered higher on assertiveness, individualism, and power distance.

The Atwater et al. (2009) hypotheses were tested via examining cultural context variables as cross-level moderators. Specifically, they investigated if their three cultural context variables predicted the relationships between others' ratings and self- ratings of leadership. They used others' ratings as the Level-1 predictors of self-ratings and the three cultural context variables as Level-2 predictors of the random intercept and random slope from the Level-1 regression. To determine if their hypotheses were supported, Atwater et al. examined whether the regression coefficients for the cultural context variables were significant. If so, they implied that the cultural context variables were related to the other rating/self-rating relationship slope. In other words, this implied a cross-level interaction such that the other rating/self-rating relationship was stronger in some cultures (e.g., those higher on assertiveness) than in others. To better understand each of their cross-level interactions, they plotted the relationships between self- and others’ ratings at conditional levels (plus and minus one standard deviation) of the cultural context variables. In short, we can envision this multi-level procedure being used to examine similar hypotheses in other nested data sets.

5. Discussion

In this article, we have reviewed the literature on SOA as it is related to leadership in the workplace. In this section, we discuss the implications of our findings for practitioners and for future research, and we summarize the conclusions that we have drawn from our review of the relevant literature.

5.1. Practitioner issues

As suggested by Yammarino and Atwater (1993), SOA can be an important factor in increasing the self-perception accuracy of participants in leadership development programs that use multi-rater assessments. In programs of this type, leaders are typically asked to compare their self-ratings to others' ratings at the scale level to determine their level of rating agreement. A key issue in this area is how to operationalize SOA so that it can be more easily used in practice. While sophisticated statistical techniques such as multivariate and polynomial regression are powerful tools for testing hypotheses in empirical research, in practice, these techniques are not very useful for indexing SOA or self-awareness. Thus, when giving 360-degree feedback to leaders, we recommend using simple indices such as comparisons of self-ratings to the mean ratings across rater groups (as discussed previously, inter-rater agreement within rating sources should be assessed prior to computing mean ratings). In these situations, an overall index of rating agreement would be a useful indicator of whether an individual has a general tendency, for example, to under- or overestimate his or her performance. Brutus et al. (1999) suggest using standardized ratings from self- and other raters to compute agreement categories, similar to the early categorization schemes (e.g., Atwater & Yammarino, 1997; Fleenor et al., 1996). Before creating such categories, however, researchers should first use techniques such as polynomial regression to establish what inferences can be made (e.g., in-agreement raters are more effective than over-estimators). If the results support a model that can be explained by agreement categories, then these categories may be used in feedback sessions for leadership development purposes.

1028 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

To improve the accuracy of performance ratings, there are several practical steps that organizations can take. For example, organizations could implement formal, structured performance appraisal systems in which raters record observed behaviors throughout the performance period and make their ratings using behaviorally anchored scales. Additionally, organizations could demonstrate their commitment to the rating process by providing rater training and incentives to raters (cf. Murphy, 2008). Finally, organizations could increase the rating accuracy in performance appraisals by holding supervisors accountable for the ratings that they provide of their employees. For example, supervisors could be required to justify their ratings in writing or to justify their ratings to individuals with authority at higher levels in the organization.

It is our experience that practitioners often ask which rating source is most predictive of leadership and other important organizational outcomes. Several SOA studies (e.g., Halverson et al., 2005) have found stronger effects for direct-report ratings, indicating that subordinates may have the best perspective for assessing leadership outcomes. Additionally, both Furnham and Stringfield (1994) and Bass and Yammarino (1991) have found that self-subordinate agreement is related to leader effectiveness. Given that direct reports directly experience leader behavior, it follows that subordinate ratings may be the best indicator of leader effectiveness. However, in a study of managerial derailment, Gentry, Braddy, Weber, and Thompson (2008) found that peer ratings on derailment variables accounted for more variance in leader performance ratings than did direct report ratings or self-ratings of derailment. Brutus et al. (1999) also found that peer ratings were the most predictive of leader effectiveness. Conway and Huffcutt (1997) reported that correlations between rating sources were generally low; thus, different rating sources may have different perspectives on leader performance. Self- ratings, however, appear to be the least predictive (e.g., Harris & Schaubroeck, 1988). Before firm conclusions can be drawn as to which rating source is the best predictor of leadership-related outcomes, additional research is needed.

5.2. Directions for future research

Research has demonstrated that there are a variety of contextual influences on self-ratings that can affect their accuracy as well as the extent of their agreement with others' ratings. As suggested by Murphy (2008), there are situational constraints in organizations (e.g., the ability of raters to observe performance, the purpose of the ratings, and the rating scales themselves) that affect the accuracy of both self-ratings and others' ratings. Cultural factors, such as unrealistic performance expectations by raters, may also affect the accuracy of ratings. To facilitate more accurate ratings, organizations should foster environments in which raters have incentives, tools, and opportunities to accurately observe and recall ratees' performance. More research is needed to develop tools and techniques that organizations can use to improve the accuracy of performance ratings, and consequently, the degree of congruence among raters.

Another area for future research concerns the relationship between self-awareness and SOA. Some research (e.g., Van Velsor et al., 1993) has suggested that in-agreement raters are more effective leaders than are individuals who underestimate or overestimate their ratings. Based on such findings, it is often assumed that individuals with in-agreement ratings are also more self-aware. As such, SOA is sometimes used as a proxy for self-awareness (e.g., Berson & Sosik, 2007).

On the other hand, several studies have found that both over-estimators and/or under-estimators may be effective leaders (e.g., Bazigos, 1999; Fleenor et al., 1996). Our experience with multi-rater feedback in applied settings suggests that differences in SOA, particularly those greater than one point on a five-point scale, are indicators of low self-awareness, especially for over-estimators. In much of the research in this area, self-awareness is measured with the same instrument used to determine rating agreement. In other words, the multi-rater instruments employed to determine SOA often contain a scale that is used to measure self-awareness. In order to appropriately test the relationship between self-awareness and SOA, valid and independent measures of self- awareness need to be developed. With more valid and independent measures, it will be possible to conduct more thorough investigations of the relationship between SOA and self-awareness.

As discussed in the section on measurement issues, the SOA literature has used a variety of metrics that are not equivalent to assess this phenomenon. Discrepancies between findings across studies may be the result, therefore, of how SOA was operationalized in the different studies. Moreover, the results of any single study may or may not be replicated if SOA had been operationalized differently in that study. These multiple and inconsistent operationalizations of SOA constitute a major obstacle to drawing broad conclusions about the antecedents and consequences of SOA. We strongly recommend that when SOA is being treated as a predictor, researchers use a widely accepted analytic technique, such as polynomial regression (Edwards, 1993) or WABA (Yammarino, 1998), rather than alternative methodologies (e.g., difference scores or agreement categorization) to analyze SOA.

As noted by an anonymous reviewer, another key question for future research is: What does SOA actually measure? Is SOA a single construct or is it a constellation of personal and interpersonal constructs? To fully investigate this question, relationships between SOA and several additional variables should be explored. Examples of such variables include: private and public self- consciousness, self-regulation, self-monitoring, felt authenticity, self-deception, self-construals, and values. Future research should be informed by the literatures on self-identity, positive psychology, and authentic leadership. For example, Peterson and Seligman's (2004) research on the virtue of temperance and measurement of the character strength of self-regulation/control may be relevant.

5.3. Conclusion

In 1997, Atwater and Yammarino proposed that SOA was related to various outcome measures that have important implications for both individuals and organizations. Among their predictions were that leaders whose self-ratings agreed with

1029J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

others' ratings as to their high levels of effectiveness are more likely to be linked to positive individual and organizational outcomes. Alternatively, leaders who agreed with others about their low level of effectiveness are more likely to be linked to negative outcomes. As described in this literature review, many of Atwater and Yammarino's (1997) hypotheses have been upheld by the extant research.

SOA research continues to demonstrate that there are a variety of individual and contextual influences on self-ratings that can affect their “accuracy” as well as their congruence with others' ratings. There are also techniques that can be used to increase the accuracy of self-ratings as well as their degree of congruence with others.

Likewise, others' ratings are influenced by a variety of factors including the rater's cognitive processes, ability, and motivation. Each of these factors can enhance or diminish the accuracy of others' ratings. In addition, contextual factors can affect others' ratings in a variety of ways. Many of the factors reviewed in this paper may result in others' ratings departing from self-ratings, even when those self-ratings are accurate. To increase rating agreement, organizations must create an environment where “raters have: (a) incentives, tools and opportunities to observe and recall ratees' job performance, (b) incentives to provide ratings that faithfully reflect the rater's evaluation of each ratee's performance, and (c) protection against the negative consequences of giving honest ratings” (Murphy, 2008, p. 157).

As multi-rater feedback continues to be widely used in organizations, it is important to better understand the relationships between SOA and its predictors and outcomes. For example, there is some evidence that in-agreement raters place more importance on job satisfaction and organizational commitment (e.g., Szell & Henderson, 1997) and that they may be more psychologically well-adjusted than over-estimators (Kwan et al., 2008). Improving our knowledge of the multi-rater process will be critical to providing more specific and helpful feedback to the leaders involved. For example, understanding that older, more experienced leaders tend to over-rate themselves relative to others suggests that feedback coaches should pay particular attention to this trend when providing feedback to these leaders (Ostroff et al., 2004).

Studies that used early methods for determining SOA (e.g., categorization; Fleenor et al., 1996) failed to demonstrate that SOA predicts outcomes such as leader effectiveness. Recent research using more sophisticated analyses (e.g., polynomial regression and multivariate regression; Atwater et al., 1998; Ostroff et al., 2004), however, has found that SOA can predict outcomes such as effectiveness and derailment. Moreover, such research indicates that the relationship between SOA and leadership outcomes is more complex than a simple linear association between agreement and these outcomes. For example, some studies have found stronger effects for direct-report ratings, indicating that subordinates may have the best perspective for assessing leadership outcomes (e.g., Halverson et al., 2005). Other studies found stronger effects for rating agreement for American leaders compared to European leaders (e.g., Atwater et al., 2005; Gentry et al., 2007).

In this review, we discussed several important measurement issues facing SOA researchers. In addition to establishing reliability and validity of measures used, researchers must operationalize SOA, decide whether to compare self-ratings to all rater groups separately or to all groups combined, assess inter-rater agreement for within-source rating groups (e.g., peers or subordinates) prior to testing substantive hypotheses, and establish measurement equivalence of the ratings from different rater groups. Moreover, researchers must choose the appropriate statistical technique(s) for analyzing their data. As recommended in this review, difference score indices and categories of agreement should not be used as primary methods for testing SOA hypotheses in scholarly research. Depending upon the specific research question being investigated, researchers could use polynomial regression and response surface methodology, multivariate regression, WABA, or HLM.

According to Lance, Baxter, and Mahan (2006), the literature on multisource ratings (e.g., London & Smither, 1995) has consistently found that levels of SOA tend to be, at best, moderate. Although the amount of variance accounted for by SOA is typically small, these effects may be cumulative and may have important consequences for individuals and organizations (Halverson et al., 2005). For example, using multisource feedback to determine SOA may be critical to enhancing the self- perception accuracy of participants in leadership development programs (Yammarino & Atwater, 2001). There is some indication that higher rating agreement is related to higher levels of self-awareness, which should lead to in-agreement raters setting more realistic expectations and goals. This will likely result in positive outcomes for employees and organizations (Atwater & Yammarino, 1997; Halverson et al., 2005).

Acknowledgements

We thank Francis Yammarino, William A. Gentry, and an anonymous reviewer for their helpful comments on this article.

References

Allport, G. W. (1937). Personality: A psychological interpretation. New York: Holt. Antonioni, D. (1999). Predictors of upward appraisal ratings. Journal of Managerial Issues, 11, 26−36. Antonioni, D., & Park, H. (2001a). The effects of personality similarity on peer ratings of contextual work behaviors. Personnel Psychology, 54, 331−360. Antonioni, D., & Park, H. (2001b). The relationship between rater affect and three sources of 360-degree feedback ratings. Journal of Management, 27, 479−495. Ashford, S. J. (1989). Self-assessments in organizations: A literature review and integrative model. In L. L. Cummings & B.W. Staw (Eds.), Research in organizational

behavior, 11. (pp. 133−174). Greenwich, CT: JAI Press. Atkins, P. W., & Wood, R. E. (2002). Self-versus others' ratings as predictors of assessment center ratings: Validation evidence for 360-degree feedback programs.

Personnel Psychology, 55, 871−904. Atwater, L. E., Ostroff, C., Yammarino, F. J., & Fleenor, J. W. (1998). Self–other agreement: Does it really matter? Personnel Psychology, 51, 577−598. Atwater, L. E., Roush, P., & Fischthal, A. (1995). The influence of upward feedback on self- and follower ratings of leadership. Personnel Psychology, 48, 35−59.

1030 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Atwater, L. E., Waldman, D., Ostroff, C., Robie, C., & Johnson, K. M. (2005). Self–other agreement: Comparing its relationship with performance in the U.S. and Europe. International Journal of Selection and Assessment, 13, 25−40.

Atwater, L. E., Wang, M., Smither, J. W., & Fleenor, J. W. (2009). Are cultural characteristics associated with the relationship between self and others' ratings of leadership? Journal of Applied Psychology, 94, 876−886.

Atwater, L. E., & Yammarino, F. J. (1992). Does self–other agreement on leadership perceptions moderate the validity of leadership and performance predictions? Personnel Psychology, 45, 141−164.

Atwater, L. E., & Yammarino, F. J. (1997). Self–other rating agreement. In G. R. Ferris (Ed.), Research in personnel and human resources management, 15. (pp. 121−174). JAI Press.

Aycan, Z., & Kanungo, R. N. (2001). Cross-cultural industrial and organizational psychology: A critical appraisal of the field and future directions. In N. Anderson, D. S. Ones, H. K. Sinangil, & C. Viswesvaran (Eds.), Handbook of industrial, work, and organizational psychology. Personnel psychology, 1. (pp. 385−408). London: Sage.

Bagby, R. M., Rector, N. A., Bindseil, K., Dickens, S. E., Levitan, R. D., & Kennedy, S. H. (1998). Self-report ratings and informants' ratings of personalities of depressed outpatients. American Journal of Psychiatry, 155, 437−438.

Bailey, C., & Fletcher, C. (2002). The impact of multiple source feedback on management development: Findings from a longitudinal study. Journal of Organizational Behavior, 23, 853−867.

Barr, M. A., & Raju, N. S. (2003). IRT-based assessments of rater effects in multiple-source feedback instruments. Organizational Research Methods, 6, 15−43. Baskett, G. D. (1973). Interview decisions as determined by competency and attitude similarity. Journal of Applied Psychology, 57, 343−345. Bass, B., & Yammarino, F. J. (1991). Congruence of self and others' leadership ratings of naval officers for understanding successful performance. Applied Psychology:

An International Review, 40, 437−454. Baumeister, R. F., Bratslavsky, E., Finkenauer, C., & Vohs, K. D. (2001). Bad is stronger than good. Review of General Psychology, 5, 323−370. Bayroff, A. G., Haggerty, H. R., & Rundquist, E. A. (1954). Validity of ratings as related to rating techniques and conditions. Personnel Psychology, 7, 93−112. Bazerman, M. H., Beekun, R. I., & Schoorman, F. D. (1982). Performance evaluation in a dynamic context: A laboratory study of the impact of a prior commitment to

the ratee. Journal of Applied Psychology, 67, 873−876. Bazigos, M. N. (1999). The relationship of upward feedback disparities to leader performance: Understanding “overestimation”. Dissertation Abstracts International:

Section B: The Sciences and Engineering, 60(5-B). (pp. 2395). Beehr, T. A., Ivanitskaya, L., Hansen, C. P., Erofeev, D., & Gudanowski, D. M. (2001). Evaluation of 360-degree feedback ratings: Relationships with each other and

with performance and selection predictors. Journal of Organizational Behavior, 22, 775−788. Bell, S. T., & Arthur, W., Jr. (2008). Feedback acceptance in developmental assessment centers: The role of feedback message, participant personality, and affective

response to the feedback session. Journal of Organizational Behavior, 29, 681−703. Bernardin, H. J., & Buckley, M. R. (1981). Strategies in rater training. Academy of Management Review, 6, 205−212. Bernardin, H. J., & Cardy, R. L. (1982). Appraisal accuracy: The ability and motivation to remember the past. Public Personnel Management, 11, 352−357. Bernardin, H. J., Cardy, R. L., & Carlyle, J. J. (1982). Cognitive complexity and appraisal effectiveness: Back to the drawing board? Journal of Applied Psychology, 67,

151−160. Bernardin, H. J., & Orban, J. A. (1990). Leniency effect as a function of rating format, purpose of appraisal, and rater individual differences. Journal of Business and

Psychology, 5, 197−211. Bernardin, H. J., & Pence, E. C. (1980). Effects of rater error training: Creating new response sets and decreasing accuracy. Journal of Applied Psychology, 65, 60−66. Berscheid, E., & Walster, E. (1969). Interpersonal attraction. Reading, MA: Addison-Wesley. Berson, Y., & Sosik, J. J. (2007). The relationship between self–other rating agreement and influence tactics and organizational processes. Group & Organization

Management, 32, 675−698. Blakely, G. L., Andrews, M. C., & Fuller, J. (2003). Are chameleons good citizens? A longitudinal study of the relationship between self-monitoring and

organizational citizenship behavior. Journal of Business and Psychology, 18, 131−144. Borman, W. C. (1979). Individual differences correlates of accuracy in evaluating others' performance effectiveness. Applied Psychological Measurement, 3,

103−115. Borman, W. C. (1997). 360 ratings: An analysis of assumptions and a research agenda for evaluating their validity. Human Resource Management Review, 7, 299−315. Bracken, D., Timmreck, C., & Church, A. (2001). The handbook of multisource feedback. San Francisco: Jossey-Bass. Bravo, I. M., & Kravitz, D. A. (1996). Context effects in performance appraisals: Influence of target value, context polarity, and individual differences. Journal of

Applied Social Psychology, 26, 1681−1701. Brutus, S., Fleenor, J. W., & McCauley, C. D. (1999). Demographic and personality predictors of congruence in multi-source ratings. The Journal of Management

Development, 18, 417−435. Brutus, S., Fleenor, J. W., & Tisak, J. (1999). Exploring the link between rating congruence and managerial effectiveness. Canadian Journal of Administrative Sciences,

16, 308−322. Byrne, D. (1971). The attraction paradigm. New York: Academic Press. Caligiuri, P. M., & Day, D. V. (2000). Effects of self-monitoring on technical, contextual, and assignment-specific performance. Group and Organization Management,

25, 154−174. Cardy, R. L., & Kehoe, J. F. (1984). Rater selective attention ability and appraisal effectiveness: The effect of a cognitive style on the accuracy of differentiation among

ratees. Journal of Applied Psychology, 69, 589−594. Cascio, W. F., & Valenzi, E. R. (1977). Behaviorally anchored rating scales: Effects of education and job experience of raters and ratees. Journal of Applied Psychology,

62, 278−282. Cellar, D. F., Durr, M. L., Halsell, S., & Doverspike, D. (1989). The effect of field independence, job analysis format, and sex of rater on the accuracy of job evaluation

ratings. Journal of Applied Social Psychology, 19, 363−376. Cheung, G. W. (1999). Multifaceted conceptions of self–others' ratings disagreement. Personnel Psychology, 52, 1−36. Conway, J. M. (1998). Understanding method variance in multitrait–multirater performance appraisal matrices: Examples using general impressions and

interpersonal affect as measured method factors. Human Performance, 11, 29−55. Conway, J. M., & Huffcutt, A. I. (1997). Psychometric properties of multisource performance ratings: A meta-analysis of subordinate, supervisor, peer, and self-

ratings. Human Performance, 10, 331−360. Dansereau, F., & Yammarino, F. J. (2000). Within and between analysis: The variant paradigm as an underlying approach to theory building and testing. In K. J. Klein

& S.W.J. Kozlowski (Eds.), Multilevel theory, research, and methods in organizations: Foundations, extensions, and new directions (pp. 425−466). San Francisco: Jossey-Bass.

Davis, D. D. (1998). International performance measurement and management. In J. W. Smither (Ed.), Performance appraisal: State-of-the-art in practice (pp. 95−131). San Francisco: Jossey-Bass.

Day, D. D., & Greguras, G. J. (2009). Performance management in multi-national companies. In J. W. Smither & M. London (Eds.), Performance management: Putting research into practice. San Francisco: Jossey-Bass.

De Los Reyes, A., & Kazdin, A. E. (2004). Measuring informant discrepancies in clinical child research. Psychological Assessment, 16, 330−334. DeCotiis, T. A., & Petit, A. (1978). The performance appraisal process: A model and some testable hypotheses. Academy of Management Review, 21, 635−646. DeNisi, A. S. (1996). Cognitive approach to performance appraisal: A program of research. New York: Routledge. DeNisi, A. S., Cafferty, T. P., & Meglino, B. M. (1984). A cognitive view of the performance appraisal process: A model and research propositions. Organizational

Behavior and Human Performance, 33, 360−396. Dindia, K., & Allen, M. (1992). Sex differences in self-disclosure: A meta-analysis. Psychological Bulletin, 112, 106−124. Dipboye, R. L. (1977). A critical review of Korman's self-consistency theory of work motivation and occupational choice. Organizational Behavior & Human

Performance, 18, 108−126.

1031J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Dubin, S. S., Burke, L. K., Katz, A., & Chesler, D. J. (1954a). Characteristics of raters whose ratings reflect halo. US Army Personnel Research Branch Note, 37, 5. Dubin, S. S., Burke, L. K., Katz, A., & Chesler, D. J. (1954b). Characteristics of raters whose ratings agree with consensus ratings. US Army Personnel Research Branch

Note, 35, 8. Dunnette, M. (1993). My hammer or your hammer? Human Resource Management, 32, 373−384. Edwards, J. R. (1993). Problems with the use of profile similarity indices in the study of congruence in organizational research. Personnel Psychology, 46, 641−665. Edwards, J. R. (1994). The study of congruence in organizational behavior research: Critique and a proposed alternative. Organizational Behavior and Human

Processes, 58, 51−100. Edwards, J. R. (1995). Alternatives to difference scores as dependent variables in the study of congruence in organizational research. Organizational Behavior and

Human Decision Processes, 64, 307−324. Edwards, J. R. (2002). Alternatives to difference scores: Polynomial regression analysis and response surface methodology. In F. Drasgow & N. Schmitt (Eds.),

Measuring and analyzing behavior in organizations: Advances in measurement and data analysis (pp. 350−400). San Francisco, CA: Jossey-Bass, Inc. Edwards, J. R., & Parry, M. E. (1993). On the use of polynomial regression equations as an alternative to difference scores in organizational research. Academy of

Management Journal, 6, 1577−1633. Facteau, J. D., & Craig, S. B. (2001). Are performance appraisal ratings from different rating sources comparable? Journal of Applied Psychology, 86, 215−227. Farh, J. -L., & Cheng, B. -S. (1997). Modesty bias in self-rating in Taiwan: Impact of item wording, modesty value, and self-esteem. Chinese Journal of Psychology, 39,

103−118. Farh, J. -L., Dobbins, G. H., & Cheng, B. -S. (1991). Cultural relativity in action: A comparison of self-rating made by Chinese and U.S. workers. Personnel Psychology,

44, 129−147. Feldman, J. M. (1981). Beyond attribution theory: Cognitive processes in performance appraisal. Journal of Applied Psychology, 66, 127−148. Ferris, G. R., Fedor, D. B., Chachere, J. G., & Pondy, L. R. (1989). Myths and politics in organizational contexts. Group & Organization Studies, 14, 83−103. Ferris, G. R., Judge, T. A., Rowland, K. M., & Fitzgibbons, D. E. (1994). Subordinate influence and the performance evaluation process: Test of a model. Organizational

Behavior and Human Decision Processes, 58, 101−135. Festinger, L. (1954). A theory of social comparison processes. Human Relations, 7, 117−140. Fleenor, J. W., McCauley, C. D., & Brutus, S. (1996). Self–other rating agreement and leader effectiveness. Leadership Quarterly, 7, 487−506. Fleenor, J. W., Taylor, S., & Chappelow, C. (2008). Leveraging the impact of 360-degree feedback. San Francisco: Pfeiffer. Fletcher, C. (1999). The implication of research on gender differences in self-assessment and 360 degree appraisal. Human Resource Management Journal, 9, 39−46. Fletcher, C., & Perry, E. L. (2001). Performance appraisal and feedback: A consideration of national culture and a review of contemporary research and future

trends. In N. Anderson, D. S. Ones, H. K. Sinangil, & C. Viswesvaran (Eds.), Handbook of industrial, work, and organizational psychology. Personnel psychology, 1. (pp. 127−144). London: Sage.

Forsyth, D. R., Heiney, M. M., & Wright, S. S. (1997). Biases in appraisals of women leaders. Group Dynamics: Theory, Research, and Practice, 1, 98−103. Fox, S., Ben-Nahum, Z., & Yinon, Y. (1989). Perceived similarity and accuracy of peer ratings. Journal of Applied Psychology, 74, 781−786. Freeberg, N. E. (1969). Relevance of rater–ratee acquaintance in the validity and reliability of ratings. Journal of Applied Psychology, 53, 518−524. Fried, Y., Levi, A. S., Ben-David, H. A., Tiegs, R. B., & Avital, N. (2000). Rater positive and negative mood predispositions as predictors of performance ratings of ratees

in simulated and real organizational settings: Evidence from US and Israeli samples. Journal of Occupational and Organizational Psychology, 73, 373−378. Furnham, A., Moutafi, J., & Chamorro-Premuzic, T. (2005). Personality and intelligence: Gender, the Big Five, self-estimated and psychometric intelligence.

International Journal of Selection and Assessment, 13, 11−24. Furnham, A., & Stringfield, P. (1994). Congruence of self and subordinate ratings of managerial practices as a correlate of supervisor evaluation. Journal of

Occupational and Organizational Psychology, 67, 57−67. Gentry, W. A., Braddy, P. W., Fleenor, J. W., & Howard, P. J. (2008). Self-observer rating discrepancies on the derailment behaviors of Hispanic managers. The

Business Journal of Hispanic Research, 2, 76−87. Gentry, W. A., Braddy, P. W., Weber, T., & Thompson, L. F. (2008, Apr). Predictive utility of peer versus direct-report ratings of derailment tendencies. Paper presented at

the meeting of the Society for Industrial and Organizational Psychology, San Francisco. Gentry, W. A., Hannum, K. M., Ekelund, B. Z., & de Jong, A. (2007). A study of the discrepancy between self- and observer-ratings on managerial derailment

characteristics of European managers. European Journal of Work and Organizational Psychology, 16, 295−325. Gentry, W. A., Yip, J., & Hannum, K. M. (2010). Self-observer rating discrepancies of managers in Asia: A study of derailment characteristics and behaviors in

Southern and Confucian Asia. International Journal of Selection and Assessment, 18, 237−250. Gerstner, C. R., & Day, D. V. (1997). Meta-analytic review of leader–member exchange theory: Correlates and construct issues. Journal of Applied Psychology, 82,

827−844. Gioia, D. A., & Longenecker, C. O. (1994). Delving into the dark side: The politics of executive appraisal. Organizational Dynamics, 22, 47−58. Gioia, D. A., & Sims, H. P. (1985). On avoiding the influence of implicit leadership theories in leader behavior descriptions. Educational and Psychological

Measurement, 45, 217−232. Goffin, R. D., & Anderson, D. W. (2007). The self-rater's personality and self–other disagreement in multi-source performance ratings: Is disagreement healthy?

Journal of Managerial Psychology, 22, 271−289. Gorman, C. A., & Rentsch, J. R. (2009). Evaluating frame-of-reference rater training effectiveness using performance schema accuracy. Journal of Applied Psychology,

94, 1336−1344. Grey, R. J., & Kipnis, D. (1976). Untangling the performance appraisal dilemma: The influence of perceived organizational context on evaluative processes. Journal

of Applied Psychology, 61, 329−335. Guion, R. M. (1998). Assessment, measurement, and prediction for personnel decisions. Mahwah, NJ: Lawrence Erlbaum. Haaland, S., & Christiansen, N. D. (2002). Implications of trait-activation theory for evaluating the construct validity of assessment center ratings. Personnel

Psychology, 55, 137−163. Halverson, S. K., Tonidandel, S., Barlow, C. B., & Dipboye, R. L. (2005). Self–other agreement on a 360-degree leadership evaluation. In S. Reddy (Ed.), Multi-source

performance assessment: Perspective and Insights (pp. 125−144). Hyderabad, India: ICFAI University Press. Harris, M. M. (1994). Rater motivation in the performance appraisal context: A theoretical framework. Journal of Management, 20, 737−756. Harris, M. M., & Schaubroeck, J. (1988). A meta-analysis of self-supervisor, self-peer, and peer-supervisor ratings. Personnel Psychology, 41, 43−62. Härtel, C. E. (1993). Rating format research revisited: Format effectiveness and acceptability depend on rater characteristics. Journal of Applied Psychology, 78,

212−217. Hauenstein, N. M. A. (1998). Training raters to increase the accuracy of appraisals and the usefulness of feedback. In J. W. Smither (Ed.), Performance appraisal: State

of the art in practice (pp. 404−442). San Francisco: Jossey-Bass. Hauenstein, N. M., & Alexander, R. A. (1991). Rating ability in performance judgments: The joint influence of implicit theories and intelligence. Organizational

Behavior and Human Decision Processes, 50, 300−323. Heger, D. K. (2007). Factors affecting agreement in multi-rater performance appraisal: Observability and opportunity-to-observe. Dissertation Abstracts

International: Section B: The Sciences and Engineering, 68(5-B), 3436. Heidemeier, H., & Moser, K. (2009). Self–other agreement in job performance ratings: A meta-analytic test of a process model. Journal of Applied Psychology, 94,

353−370. Heslin, P. A., Latham, G. P., & VandeWalle, D. (2005). The effect of implicit person theory on performance appraisals. Journal of Applied Psychology, 90, 842−856. Heslin, P. A., & VandeWalle, D. (2008). Managers' implicit assumptions about personnel. Current Directions in Psychological Science, 17, 219−223. Higgins, E. T., & Lurie, L. (1983). Context, categorization, and recall: The “change-of-standard” effect. Cognitive Psychology, 15, 525−547. Hough, L., Keyes, M., & Dunnette, M. (1983). An evaluation of three “alternative” selection procedures. Personnel Psychology, 36, 261−275. House, R. J., Hanges, P. J., Javidan, M., Dorfman, P., & Gupta, V. (1999). Culture, leadership and organizations. Thousand Oaks, CA: Sage. Ilgen, D. R., & Feldman, J. M. (1983). Performance appraisal: A process focus. Research in Organizational Behavior, 5, 141−197.

1032 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Jackson, D. J. R., Stillman, J. A., Burke, S., & Englert, P. (2007). Self versus assessor ratings and their classification in assessment centers: Profiling the self-rater. New Zealand Journal of Psychology, 36, 93−99.

James, L. R., Demaree, R. G., & Wolf, G. (1984). Estimating within-group reliability with and without response bias. Journal of Applied Psychology, 69, 85−98. James, L. J., Demaree, R. G., & Wolf, G. (1993). rWG: An assessment of within-group interrater agreement. Journal of Applied Psychology, 78, 306−309. Jawahar, I. M. (2001). Attitudes, self-monitoring, and appraisal behaviors. Journal of Applied Psychology, 86, 875−883. Jawahar, I. M., & Williams, C. R. (1997). Where all the children are above average: The performance appraisal purpose effect. Personnel Psychology, 50, 905−925. Jeffcoat, K. A. (2000). Multi-source feedback, source credibility, and performance discrepancy: The dynamics of fairness and usefulness perceptions. Dissertation

Abstracts International: Section B: The Sciences and Engineering, 61(4-B), 2256. Johnson, J. W., & Ferstl, K. L. (1999). The effects of interrater and self–other agreement on performance improvement following upward feedback. Personnel

Psychology, 52, 272−303. Jones, L., & Fletcher, C. (2002). Self-assessment in a selection situation: An evaluation of different measurement approaches. Journal of Occupational &

Organizational Psychology, 75, 145−161. Judge, T. A., LePine, J. A., & Rich, B. L. (2006). Loving yourself abundantly: Relationship of the narcissistic personality to self- and other perceptions of workplace

deviance, leadership, and task and contextual performance. Journal of Applied Psychology, 91, 762−776. Kallejian, V., Brown, P., & Weschler, I. R. (1953). The impact of interpersonal relations on ratings of performance. Public Personnel Review, 14, 166−170. Kenny, D. A. (1994). Interpersonal perception: A social relations analysis. New York: Guilford Press. Kenny, D. A., & La Voie, L. (1984). The social relations model. In L. Berkowitz (Ed.), Advances in experimental social psychology, Vol. 18. (pp. 142−182). Orlando, FL:

Academic Press. Kierein, N. M., & Gold, M. A. (2000). Pygmalion in work organizations: A meta-analysis. Journal of Organizational Behavior, 21, 913−928. Kingston, P. W., Hubbard, R., Lapp, B., Schroeder, P., & Wilson, J. (2003). Why education matters. Sociology of Education, 76, 53−70. Kingstrom, P. O., & Mainstone, L. E. (1985). An investigation of the rater–ratee acquaintance and rater bias. Academy of Management Journal, 28, 641−653. Kirchner, W. K., & Reisberg, D. J. (1962). Differences between better and less-effective supervisors in appraisal of subordinates. Personnel Psychology, 15, 295−302. Korman, A. K. (1976). Hypothesis of work behavior revisited and an extension. Academy of Management Review, 1, 50−63. Korsgaard, M. A., Meglino, B. M., & Lester, S. W. (2004). The effect of other orientation on self-supervisor rating agreement. Journal of Organizational Behavior, 25,

873−891. Kozlowski, S. W. J., Chao, G. T., & Morrison, R. F. (1998). Games raters play: Politics, strategies, and impression management in performance appraisal. In J. W.

Smither (Ed.), Performance appraisal: State-of-the-art in practice. San Francisco: Jossey-Bass. Kraiger, K., & Ford, J. K. (1985). A meta-analysis of ratee race effects in performance ratings. Journal of Applied Psychology, 70, 56−65. Kravitz, D. A., & Balzer, W. K. (1992). Context effects in performance appraisal: A methodological critique and empirical study. Journal of Applied Psychology, 77, 24−31. Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it: How difficulties in recognizing one's own incompetence lead to inflated self-assessments. Journal of

Personality and Social Psychology, 77, 1121−1134. Kwan, V. S. Y., John, O. P., Kenny, D. A., Bond, M. H., & Robins, R. W. (2004). Reconceptualizing individual differences in self-enhancement bias: An interpersonal

approach. Psychological Review, 111, 94−110. Kwan, V. S. Y., John, O. P., Robins, R. W., & Kuang, L. L. (2008). Conceptualizing and assessing self-enhancement bias: A componential approach. Journal of Personality

and Social Psychology, 94, 1062−1077. Lance, C. E., Baxter, D., & Mahan, R. (2006). Evaluation of alternative perspectives on source effects in multisource performance measures. In W. Bennett, C. E. Lance,

& D. J. Woehr (Eds.), Performance measurement: Current perspectives and future challenges (pp. 49−76). Mahwah, NJ: Lawrence Erlbaum. Lance, C. E., & Bennett, W., Jr. (1997, April). Rater source differences in cognitive representations of performance dimensions. Paper presented at the meeting of the

Society for Industrial and Organizational Psychology, St. Louis, MO. Lance, C. E., Hoffman, B. J., Gentry, W. A., & Baranik, L. E. (2008). Rater source factors represent important subcomponents of the criterion construct space, not rater

bias. Human Resource Management Review, 18, 223−232. Landy, F. J., & Farr, J. L. (1980). Performance rating. Psychological Bulletin, 87, 72−107. LeBreton, J. M., Burgess, J. R., Kaiser, R. B., Atchley, R. B., & James, L. R. (2003). The restriction of variance hypothesis and interrater reliability and agreement. Are

ratings from multiple sources really dissimilar? Organizational Research Methods, 6, 80−128. LeBreton, J. M., & Senter, J. L. (2008). Answers to 20 questions about interrater reliability and agreement. Organizational Research Methods, 11, 815−852. Levy, P. E. (1993). Self-appraisal and attributions: A test of a model. Journal of Management, 19, 51−62. Liden, R. C., & Mitchell, T. R. (1983). The effects of group interdependence on supervisor performance evaluations. Personnel Psychology, 36, 289−299. Lin, T., Dobbins, G. H., & Farh, J. (1992). A field study of race and age similarity effects on interview ratings in conventional and situational interviews. Journal of

Applied Psychology, 77, 363−371. Lindeman, M., Sundvik, L., & Rouhiainen, P. (1995). Under- or overestimation of self? Person variables and self-assessment accuracy in work settings. Journal of

Social Behavior & Personality, 10, 123−134. Linville, P. W., & Jones, E. E. (1980). Polarized appraisals of out-group members. Journal of Personality and Social Psychology, 38, 689−703. London, M. (1995). Self and interpersonal insight: How people learn about themselves and others in organizations. New York: Oxford University Press. London, M., & Smither, J. W. (1995). Can multi-source feedback change perceptions of goal accomplishment, self-evaluations, and performance-related outcomes?

Theory-based applications and directions for research. Personnel Psychology, 48, 803−839. Longenecker, C. O., Sims, H. P., & Gioia, D. A. (1987). Behind the mask: The politics of employee appraisal. Academy of Management Executive, 1, 183−193. Lord, R. G., Foti, R. J., & de Vader, C. L. (1984). A test of leadership categorization theory: Internal structure, information processing, and leadership perceptions.

Organizational Behavior & Human Performance, 34, 343−378. Mabe, P., & West, J. (1982). Validity of self-evaluation of ability: A review and meta-analysis. Journal of Applied Psychology, 67, 280−296. Martell, R. F., & Leavitt, K. N. (2002). Reducing the performance-cue bias in work behavior ratings: Can groups help? Journal of Applied Psychology, 87, 1032−1041. Maurer, T., Raju, N. S., & Collins, W. C. (1998). Peer and subordinate performance appraisal measurement equivalence. Journal of Applied Psychology, 83, 693−703. McCall, M., & Lombardo, M. (1983). Off the track: Why and how successful executives get derailed. Greensboro, NC: Center for Creative Leadership. McNatt, D. B. (2000). Ancient Pygmalion joins contemporary management: A meta-analysis of the result. Journal of Applied Psychology, 85, 314−322. Meade, A. W., & Lautenschlager, G. J. (2004). A comparison of item response theory and confirmatory factor analytic methodologies for establishing measurement

equivalence/invariance. Organizational Research Methods, 7, 361−388. Mero, N. P., Guidice, R. M., & Brownlee, A. L. (2007). Accountability in a performance appraisal context: The effect of audience and form of accounting on rater

response and behavior. Journal of Management, 33, 223−252. Mero, N. P., & Motowidlo, S. J. (1995). Effects of rater accountability on the accuracy and the favorability of performance ratings. Journal of Applied Psychology, 80,

517−524. Miller, J. S., & Cardy, R. L. (2000). Self-monitoring and performance appraisal: Rating outcomes in project teams. Journal of Organizational Behavior, 21(6), 609−626. Mischel, W. (1973). Toward a cognitive social learning reconceptualization of personality. Psychological Review, 80, 252−253. Mischel, W. (1977). The interaction of person and situation. In D. Magnussen & N. Endler (Eds.), Personality at the crossroads: Current issues in interactional

psychology. Hillsdale, NJ: Erlbaum. Mitchell, T. R., & Liden, R. C. (1982). The effects of the social context on performance evaluations. Organizational Behavior & Human Performance, 29, 241−256. Moser, K., Schuler, H., & Funke, U. (1999). The moderating effect of raters' opportunities to observe ratees' job performance on the validity of an assessment centre.

International Journal of Selection and Assessment, 7, 133−141. Moshavi, D., Brown, F. W., & Dodd, N. G. (2003). Leader self-awareness and its relationship to subordinate attitudes and performance. Leadership and Organizational

Development Journal, 24, 407−418. Murphy, K. R. (2008). Explaining the weak relationship between job performance and ratings of job performance. Industrial and Organizational Psychology:

Perspectives on Science and Practice, 1, 148−160.

1033J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Murphy, K. R., & Cleveland, J. N. (1991). Performance appraisal: An organizational perspective. Boston: Allyn & Bacon. Murphy, K. R., & Cleveland, J. N. (1995). Understanding performance appraisal: Social, organizational, and goal-based perspectives. Thousand Oaks, CA: Sage. Murtha, T. C., Kanfer, R., & Ackerman, P. L. (1996). Toward an interactionist taxonomy of personality and situations: An integrative situational–dispositional

representation of personality traits. Journal of Personality and Social Psychology, 71, 193−207. Nilsen, D., & Campbell, D. P. (1993). Self-observer rating discrepancies: Once an overrater, always and overrater? Human Resource Management, 32, 265−282. Ostroff, C., Atwater, L. E., & Feinberg, B. J. (2004). Understanding self–other agreement: A look at rater and ratee characteristics, context, and outcomes. Personnel

Psychology, 57, 333−375. Patiar, A., & Mia, L. (2008). The effect of subordinates' gender on the difference between self-ratings, and superiors' ratings, of subordinates' performance in hotels.

International Journal of Hospitality Management, 27, 53−64. Penny, J. A. (2003). Exploring differential item functioning in a 360-degree assessment: Rater source and method of delivery. Organizational Research Methods, 6,

61−79. Peterson, C., & Seligman, M. E. P. (2004). Character strengths and virtues: A handbook and classification. New York: Oxford/American Psychological Association. Podsakoff, P., & Organ, D. (1986). Self-reports in organizational research: Problems and prospects. Journal of Management, 12, 531−544. Poon, J. M. L. (2004). Effects of performance appraisal politics on job satisfaction and turnover intention. Personnel Review, 33, 322−334. Pulakos, E. D., & Wexley, K. N. (1983). The relationship among perceptual similarity, sex, and performance ratings in manager-subordinate dyads. Academy of

Management Journal, 26, 129−139. Rand, T. M., & Wexley, K. N. (1975). Demonstration of the effect, “similar to me”, in simulated employment interviews. Psychological Reports, 36, 535−544. Robins, R. W., & John, O. P. (1997). Effects of visual perspective and narcissism on self-perception: Is seeing believing? Psychological Science, 8, 37−42. Rosenbaum, M. E. (1986). The repulsion hypothesis: On the nondevelopment of relationships. Journal of Personality and Social Psychology, 51, 1156−1166. Rothermund, K., Bak, P. M., & Brandtstädter, J. (2005). Biases in self-evaluation: Moderating effects of attribute controllability. European Journal of Social Psychology,

35, 281−290. Rothstein, H. R. (1990). Interrater reliability of job performance ratings: Growth to asymptote level with increasing opportunity to observe. Journal of Applied

Psychology, 75, 322−327. Roush, P., & Atwater, L. E. (1992). Using the MBTI to understand transformational leadership and self-perception accuracy. Military Psychology, 4, 17−34. Rush, M. C., Phillips, J. S., & Lord, R. G. (1981). Effects of a temporal delay in rating on leader behavior descriptions: A laboratory investigation. Journal of Applied

Psychology, 66, 442−450. Sackett, P. R., & DuBois, C. L. (1991). Rater–ratee race effects on performance evaluation: Challenging meta-analytic conclusions. Journal of Applied Psychology, 76,

873−877. Sackett, P. R., DuBois, C. L., & Noe, A. W. (1991). Tokenism in performance evaluation: The effects of work group representation on male–female and White–Black

differences in performance ratings. Journal of Applied Psychology, 76, 263−267. Sala, F. (2003). Executive blind spots: Discrepancies between self- and others' ratings. Consulting Psychology Journal: Practice and Research, 55, 222−229. Salvemini, N. J., Reilly, R. R., & Smither, J. W. (1993). The influence of rater motivation on assimilation effects and accuracy in performance ratings. Organizational

Behavior and Human Decision Processes, 55, 41−60. Schneider, D. E., & Bayroff, A. G. (1953). The relationship between rater characteristics and validity of ratings. Journal of Applied Psychology, 37, 278−280. Sears, G. J., & Rowe, P. M. (2003). A personality-based similar-to-me effect in the employment interview: Conscientiousness, affect-versus competence-mediated

interpretations, and the role of job relevance. Canadian Journal of Behavioural Science/Revue Canadienne Des Sciences Du Comportement, 35, 13−24. Shore, T. H., Adams, J. S., & Tashchian, A. (1998). Effects of self-appraisal information, appraisal purpose, and feedback target on performance appraisal ratings.

Journal of Business and Psychology, 12, 283−298. Smith, A. F. R., & Fortunato, V. J. (2008). Factors influencing employee intentions to provide honest upward feedback ratings. Journal of Business and Psychology, 22,

191−207. Smither, J. W., Collins, H., & Buda, R. (1989). When ratee satisfaction influences performance evaluations: A case of illusory correlation. Journal of Applied

Psychology, 74, 599−605. Smither, J. W., London, M., Vasilopoulos, N., Reilly, R., Millsap, R., & Salvemini, N. (1995). An examination of the effects of an upward feedback program over time.

Personnel Psychology, 48, 1−33. Smither, J. W., & Reilly, R. R. (1987). True intercorrelation among job components, time delay in rating, and rater intelligence as determinants of accuracy in

performance ratings. Organizational Behavior and Human Decision Processes, 40, 369−391. Smither, J. W., Reilly, R. R., & Buda, R. (1988). Effect of prior performance information on ratings of present performance: Contrast versus assimilation revisited.

Journal of Applied Psychology, 73, 487−496. Snyder, M., & Copeland, J. (1989). Self-monitoring processes in organizational settings. Hillsdale, NJ: Lawrence Erlbaum & Associates. Sosik, J. J., & Megerian, L. E. (1999). Understanding leader emotional intelligence and performance: The role of self–other agreement on transformational

leadership perceptions. Group & Organization Management, 24, 367−390. Steiner, D. D., & Rain, J. S. (1989). Immediate and delayed primacy and recency effects in performance evaluation. Journal of Applied Psychology, 74, 136−142. Szarota, P., Zawadzki, B., & Strelau, J. (2002). Big five domain and gender as determinants of rater agreement: A comparison based on self- and peer-rating on the

Polish Adjective List. Personality and Individual Differences, 33, 1265−1277. Szell, S., & Henderson, R. (1997). The impact of self-supervisor/subordinate performance rating agreement on subordinates' job satisfaction and organisational

commitment. Journal of Applied Social Behaviour, 3, 25−37. Taft, R. (1955). The ability to judge people. Psychological Bulletin, 52, 1−23. Tett, R. P., & Guterman, H. A. (2000). Situation trait relevance, trait expression, and cross-situational consistency: Testing a principle of trait-activation. Journal of

Research in Personality, 34, 397−423. Tornow, W. (1993). Perceptions or reality: Is multi-perspective measurement a means or an end? Human Resource Management, 32, 221−230. Tsui, A. S., & Ashford, S. J. (1994). Adaptive self-regulation: A process view of managerial effectiveness. Journal of Management, 20, 93−121. Tsui, A. S., & Barry, B. (1986). Interpersonal affect and rating errors. Academy of Management Journal, 29, 586−599. Tsui, A. S., & Ohlott, P. (1988). Multiple assessment of managerial effectiveness: Interrater agreement and consensus in effectiveness models. Personnel Psychology,

41, 799−803. Tsui, A. S., & O'Reilly, C. A. (1989). Beyond simple demographic effects: The importance of relational demography in superior-subordinate dyads. Academy of

Management Journal, 32, 402−423. Tziner, A., Latham, G. P., Price, B. S., & Haccoun, R. (1996). Development and validation of a questionnaire for measuring perceived political considerations in

performance appraisal. Journal of Organizational Behavior, 17, 179−190. Tziner, A., & Murphy, K. R. (1999). Additional evidence of attitudinal influences in performance appraisal. Journal of Business and Psychology, 13, 407−419. Tziner, A., Murphy, K. R., & Cleveland, J. N. (2005). Contextual and rater factors affecting rating behavior. Group & Organization Management, 30, 89−98. Tziner, A., Murphy, K. R., Cleveland, J., Beaudin, G., & Marchand, S. (1998). Impact of rater beliefs regarding performance appraisal and its organizational contexts

on appraisal quality. Journal of Business and Psychology, 12, 457−467. Uleman, J. S. (1991). Leadership ratings: Toward focusing more on specific behaviors. Leadership Quarterly, 2, 175−187. Van Scotter, J. R., Moustafa, K., Burnett, J. R., & Michael, P. G. (2007). Influence of prior acquaintance with the ratee on rater accuracy and halo. Journal of

Management Development, 26, 790−803. Van Velsor, E., Taylor, S., & Leslie, J. (1993). An examination of the relationships among self-perception accuracy, self-awareness, gender, and leader effectiveness.

Human Resource Management, 32, 249−264. Vance, R. J., Winne, P. S., & Wright, E. S. (1983). A longitudinal examination of rater and ratee effects in performance ratings. Personnel Psychology, 36, 609−620. Vandenberg, R. J., & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for

organizational research. Organizational Research Methods, 3, 4−69.

1034 J.W. Fleenor et al. / The Leadership Quarterly 21 (2010) 1005–1034

Vecchio, R. P., & Anderson, R. J. (2009). Agreement in self–others' ratings of leader effectiveness: The role of demographics and personality. International Journal of Selection and Assessment, 17, 165−179.

Vecchio, R. P., & Gobdel, B. C. (1984). The vertical dyad linkage model of leadership: Problems and prospects. Organizational Behavior & Human Performance, 34, 5−20.

Villanova, P., Bernardin, J., Dahmus, S. A., & Sims, R. L. (1993). Rater leniency and performance appraisal discomfort. Educational and Psychological Measurement, 53, 789−799.

Visser, B. A., Ashton, M. C., & Vernon, P. A. (2008). What makes you think you're so smart? Measured abilities, personality, and sex differences in relation to self- estimates of multiple intelligences. Journal of Individual Differences, 29, 35−44.

Wayne, S. J., & Ferris, G. R. (1990). Influence tactics, affect, and exchange quality in supervisor–subordinate interactions: A laboratory experiment and field study. Journal of Applied Psychology, 75, 487−499.

Wayne, S. J., & Kacmar, K. M. (1991). The effects of impression management on the performance appraisal process. Organizational Behavior and Human Decision Processes, 48, 70−88.

Wexley, K. N., & Youtz, M. A. (1985). Rater beliefs about others: Their effects on rating errors and rater accuracy. Journal of Occupational Psychology, 58, 265−275. Whetten, D. A., & Cameron, K. S. (2007). Developing management skills. Upper Saddle River, NJ: Prentice Hall. Woehr, D. J., & Huffcutt, A. I. (1994). Rater training for performance appraisal: A quantitative review. Journal of Occupational and Organizational Psychology, 67,

189−205. Woehr, D. J., & Roch, S. G. (1996). Context effects in performance evaluation: The impact of ratee sex and performance level on performance ratings and behavioral

recall. Organizational Behavior and Human Decision Processes, 66, 31−41. Woehr, D. J., Sheehan, M. K., & Bennett, W., Jr. (2005). Assessing measurement equivalence across rating sources: a multitrait–multirater approach. Journal of

Applied Psychology, 3, 592−600. Xie, J. L., Roy, J. -P., & Chen, Z. (2006). Cultural and individual differences in self-rating behavior: An extension and refinement of the cultural relativity hypothesis.

Journal of Organizational Behavior, 27, 341−364. Yammarino, F. J. (1998). Multivariate aspects of the variant/WABA approach: A discussion and illustration. Leadership Quarterly, 9, 203−227. Yammarino, F. J. (2003). Modern data analytic techniques for multisource feedback. Organizational Research Methods, 6, 6−14. Yammarino, F. J., & Atwater, L. E. (1993). Understanding self-perception accuracy: Implications for human resource management. Human Resource Management,

32, 231−247. Yammarino, F. J., & Atwater, L. E. (2001). Understanding agreement in multisource feedback. In D. Bracken, C. Timmreck, & A. Church (Eds.), The handbook of

multisource feedback (pp. 204−220). San Francisco: Jossey-Bass. Yammarino, F. J., & Markham, S. E. (1992). On the application of within and between analysis: Are absence and affect really group-based phenomena? Journal of

Applied Psychology, 77, 168−176. Yik, M. S. M., Bond, M. H., & Paulhus, D. L. (1998). Do Chinese self-enhance or self-efface? It's a matter of domain. Personality and Social Psychology Bulletin, 24,

399−406. Yun, G. J., Donahue, L. M., Dudley, N. M., & McFarland, L. A. (2005). Rater personality, rating format, and social context: Implications for performance appraisal

ratings. International Journal of Selection and Assessment, 13, 97−107.

  • Self–other rating agreement in leadership: A review
    • Models of self–other rating agreement
    • Factors affecting self-ratings and congruence between self- and others' ratings
      • Biographical characteristics
        • Gender
        • Age
        • Position
        • Race
        • Education
      • Personality and other individual characteristics
        • Big Five personality factors
        • Dominance
        • Empathy
        • Self-esteem
        • Narcissism
        • Meta-cognitive ability/intelligence
        • Private and public self-consciousness
        • Self-monitoring
        • Efficacy and locus of control
        • Depression
      • Context
        • Culture
        • Controllability
        • Political purpose
        • Similarity
      • Job relevant experiences
        • Feedback
    • Factors affecting others' ratings
      • The rater's cognitive processes
      • Characteristics of the rater
        • Rater ability
        • Rater job experience and performance
        • Rater personality
        • Rater mood
        • Rater beliefs about human nature
        • Rater attitudes
        • Rater's organizational commitment
        • Discomfort with appraisal
        • Rater's organization level relative to the ratee
      • Rater motivation
        • Rater goals
        • Politics
        • Rater accountability
        • Rater incentives
      • Contextual factors
        • Situational strength
        • Culture
        • Difficulty of rating task
        • Observation context versus rating context
        • Negative versus positive performance incidents
        • Primacy and recency effects
        • Anchoring effects
        • Proportion of women or minorities in the work group
        • Performance of the ratee's peers
        • Task interdependence
        • Rater training
        • Rating purpose
        • Trust in the appraisal process
        • Norms
      • Rater–ratee interactions and expectations
        • Leader–member exchange
        • Ratee's past performance
        • Prior commitment to the ratee
        • Rater expectations about the ratee
        • Familiarity with ratees
        • Ratee impression management tactics
        • Ratee self-appraisals
        • Similarity to the ratee and rater affect toward the ratee
    • Correlates of self–other rating agreement
      • Correlates of self–other rating agreement at the individual level
        • Leader performance
        • Assessment center performance
        • Performance improvement
        • Promotion
        • Derailment
        • Goal setting
        • Influence tactics
      • Feedback fairness
        • Psychological adjustment
      • Correlates of self–other rating agreement at the organizational level
        • Job satisfaction
      • Measurement issues and data analytic techniques in self–other rating agreement
      • Measurement issues in self–other rating agreement research
        • Operationalizing SOA
        • Rater similarity and agreement
        • Measurement equivalence/invariance (ME/I)
      • Data analytic techniques in SOA research
        • Difference scores
        • Polynomial regression
        • Multivariate regression
        • Categories of agreement
        • WABA
        • HLM
    • Discussion
      • Practitioner issues
      • Directions for future research
      • Conclusion
    • Acknowledgements
    • References