Wk 3, IOP 490: Creating a Change Strategy
303
C H A P T E R 1 0
PERFORMANCE MEASUREMENT AT WORK: A MULTILEVEL
PERSPECTIVE Jessica L. Wildman, Wendy L. Bedwell, Eduardo Salas, and Kimberly A. Smith-Jentsch
For decades, one of the primary goals of organiza- tional research has been the improvement and man- agement of organizational performance. Inherent to the goal of improving performance is the concept of performance measurement (PM). PM is the mecha- nism that allows managers and researchers to gain an understanding of individual, team, and overall orga- nizational performance. Without the ability to accu- rately measure a construct such as performance, it is impossible to truly understand, control, or improve it. As Sink and Tuttle (1989) asserted, one cannot manage what one cannot measure. Ultimately, the effective training and management of employees, teams, and organizations in any context is contingent on the quality of PM. Accordingly, much effort has been devoted over the past several decades to explor- ing theories, methods, and practices associated with PM (e.g., Bititci, Turner, & Begemann, 2000; Campbell, McCloy, Oppler, & Sager, 1993; Folan & Browne, 2005; Gershoni & Rudy, 1981; Kendall & Salas, 2004; Pun & White, 2005).
The PM literature can generally be categorized into three distinct perspectives: individual-level PM, team- level PM, and organizational-level PM. Very little research has simultaneously examined multiple levels. This is problematic given that actual performance in organizations takes place at all three levels simultane- ously, and perhaps more important, all three levels of performance are intertwined. Teams are becoming the predominant method for achieving organizational goals. These teams are made up of individual employ- ees, who actually engage in behaviors that lead to per-
formance. Thus, there is a need to integrate these three streams of PM research into one comprehen- sive understanding of PM and its implications.
To address this need, this chapter presents a multilevel perspective on the field of PM. First, we discuss the criterion problem, which represents a broad issue underscoring the importance of PM. Next, we briefly describe five critical considerations when choosing or designing any PM system. Then, after the core underlying issues are clear, we dive into PM as described from the individual, team, and orga- nizational perspectives. This includes the general def- inition of performance, key theories, and common measurement strategies used in each stream of litera- ture. Once each perspective is discussed separately, we discuss a multilevel approach to PM. The chapter concludes with a review of current trends requiring future research and some concluding remarks. (See also Vol. 2, chap. 9, this handbook.)
THE CRITERION PROBLEM
The development and measurement of appropriate performance criteria is of importance to both researchers and managers alike, as they are both focused on influencing performance. The practical significance of measurement on the basis of sound criteria has long been accepted (e.g., Scott, 1917); however, rigorous research on the “necessary concep- tual, taxonomic, and methodological prerequisites for the pursuit of understanding criteria” ( J. T. Austin & Villanova, 1992, p. 836) did not become a prominent area of concern until the early 1990s. Researchers
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 303
http://dx.doi.org/10.1037/12169-010 APA Handbook of Industrial and Organizational Psychology, Vol 1: Building and Developing the Organization, edited by S. Zedeck Copyright © 2011 American Psychological Association. All rights reserved.
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
have noted the necessity of well-developed criteria for measuring individual- and team-level performance as well as evaluation of organizational programs and training initiatives (Schmitt & Klimoski, 1991).
Performance criteria are initially conceptual in nature and are thus defined on the basis of subjec- tive statements of what is considered successful performance. Simply stated, performance criteria represent whatever aspects of performance a certain set of stakeholders have identified as critical. Thus, the selected dimensions of any given criterion mea- sure are based largely on the defined conceptual criteria (Nagle, 1953; Toops, 1944). For example, if a set of stakeholders are conceptually interested in assessing the productivity of a professor, this could be assessed using measures of effectiveness such as number of publications, number of graduate stu- dents sponsored, and number of conference pre- sentations. The important aspect of selecting criteria measures is to make sure these measures are ratio- nally linked to the conceptual criteria and are suffi- ciently covering the criterion space. In other words, do the measures of performance effectiveness include all of the things that stakeholders deem important to performance in a particular job?
One of the most troubling issues in performance research has been the lack of focus on the conscious choice and development of criteria measures. Unfortunately, organizations and researchers often select criteria on the basis of availability or how easy the criteria are to collect. This is problematic because the choice of a performance measure influences how well selected predictors can actually forecast future performance. The choice of outcome criteria is just as important as the choice of predictors, if not more so. The best selection test or training system in the world could be developed; however, without sound criteria to serve as a measure of effectiveness, it is difficult to provide evidence for its validity. If the performance measure chosen as the criterion is not conceptually related to the outcome of interest, or if the measure is contaminated, the selection process may be excellent or the training may be well designed; however, there will never be a demonstrated connec- tion to performance.
Therefore, there is a need for an increased focus on developing sound performance criteria and sys-
tematically linking those criteria to other constructs of interest. Additionally, performance criteria should be measured without the influence of halo and other sources of error to capture true performance. These issues fall under the term criterion problem (e.g., Flanagan, 1956; Smith, 1976). Essentially, this refers to problems associated with developing and measuring the multidimensional nature of per- formance criteria given the constraints of the mea- surement purpose and situational factors ( J. T. Austin & Villanova, 1992). For example, Viswesvaran, Schmidt, and Ones (2005) found that less than 10% of the variance in a set of job performance ratings or rankings could be attributed to valid performance-related information. The rest was attributed to such things as halo error, rater leniency error, and random error, among others, demonstrating that these problems are important to consider when measuring performance.
These errors can be divided into three common categories: distributional errors, illusory halo, and other types of errors. Distributional errors relate to incorrect representations of performance distribu- tions across employees being evaluated (Borman, 1991). These errors can occur in both the rating means (e.g., severity or leniency) and variance (e.g., range restriction and central tendency). If a rater provides ratings that are lower (severity) or higher (leniency) than actually warranted by the performance because of inaccurate norms, then ratings will be erroneously deflated (severity) or inflated (leniency). If a rater fails to sufficiently dif- ferentiate between two or more ratees on the same dimension, then restriction of range has occurred. This is similar to the error of central tendency; however, with central tendency, ratings tend to be clustered around the midpoint of any given scale (Tsui & Barry, 1986). The second category of errors is illusory halo, which results in correlations between ratings of two different dimensions being higher (or lower) than the correlation between the actual behaviors reflecting those dimensions. Essentially, raters are either overestimating (higher correlations) or underestimating (lower correlations) the relation- ship between dimensions (Borman, 1991; Fisicaro, 1988). The final category of other errors includes such perceptual errors as the similar-to-me error and
Wildman et al.
304
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 304
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
the first-impression error. Similar-to-me error occurs when the rater projects his or her own personal char- acteristics onto the employee (Latham, Wexley, & Pursell, 1975). If the rater is heavily influenced by early experiences with the ratee, then first-impression error has occurred (Latham et al., 1975). This can cause biased ratings that are lower or higher than the actual performance warrants, depending on whether first impressions are negative or positive.
There are additional issues associated with the development of high-quality performance criteria. Because criteria essentially focus on the results or outcomes of performance, there are several steps required to directly link criteria measurement to asso- ciated predictors (J. T. Austin & Villanova, 1992). This is problematic in that other variables, such as sit- uational factors, may constrain this translation from predictors to behaviors to results (Binning & Barrett, 1989). It is important to clearly define the constructs within this context. Borman (1991), in his seminal chapter on job performance, defined behavior (what people do), performance (individual contributions toward organizational goals), and effectiveness (outcomes such as promotion rate or salary level). Campbell et al. (1993) suggested that performance is the actual behavior and therefore measuring the behavior constitutes measuring performance. Regardless of the adopted definition, it is fairly easy to measure behaviors, as they are generally observable and can easily be recorded. It is also a relatively sim- ple process to measure results using quantity, qual- ity, or customer satisfaction. The difficulty lies in (a) tying specific behaviors to specific results in the context of performance and (b) measuring the cogni- tive aspects associated with behaviors.
Others have suggested that criteria dimensions are also problematic because they are context sensitive (Bailey, 1983). Given this assertion, measures appro- priate for use in one situation would be inappropriate within a different context. Also, as noted previously, the selected dimensions of a criterion construct are based on how the conceptual criteria are defined. An additional issue contributing to the criterion problem is the lack of description often provided as to why certain dimensions were selected and other seemingly important dimensions were ignored (J. T. Austin & Villanova, 1992).
SETTING THE STAGE: BASIC CONSIDERATIONS IN PERFORMANCE MEASUREMENT
This section outlines five critical issues to consider when choosing or designing a PM system: (a) the purpose of the measurement, (b) the content of the measurement, (c) the timing of measurement, (d) the fidelity of the measurement setting, and (e) the technique or tools used for measurement. We refer to these issues simply as the why, what, when, where, and how of PM (see Figure 10.1). These considerations are important to keep in mind when examining existing PM strategies, because each strategy presents advantages and disadvantages regarding these considerations. Given that the pri- mary focus of this chapter is on the methods for measuring performance at the individual, team, and organizational levels, the consideration of how (i.e., how to measure performance) is divided into these three perspectives for discussion and repre- sents a large portion of the content in this chapter.
Why: Purpose of Measurement The first critical issue to consider when choosing or designing a PM system is the purpose for the mea- surement, as the purpose will drive the entire mea- surement process (Salas, Burke, & Fowlkes, 2006). The purpose determines whether multiple criteria measures (e.g., Bartram, 2005) or a single composite criterion measure (e.g., Viswesvaran et al., 2005) is used. There are numerous uses for PM data, ranging from basic research to a variety of applied purposes such as training development and strategic plan- ning. The most common purposes for PM include research, feedback development, training develop- ment, performance evaluation, and organization planning. Multiple measures are appropriate if the purpose is to diagnose performance issues, as this allows for a more accurate picture of areas needing improvement and aids in planning for training and employee development. Composite measures, on the other hand, are better for comparing across units who may not do the same type of work. This is the basis of the Productivity Measurement and Enhancement System (ProMES), a PM system that is discussed later in the chapter.
Performance Measurement at Work
305
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 305
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Basic research. Accurate PM is absolutely critical to all research endeavors. As stated by Tannenbaum (2006), “measurement lies at the heart of scientific study” (p. 297). Without the ability to accurately and reliably measure performance and other constructs of interest to researchers, it would be impossible to gain any scientific knowledge. Brannick and Prince (1997) also pointed out that “measurement is cen- tral to the evaluation and elaboration of theories” (p. 5). Theories would not be validated, or basic relationships tested, without proper measurement. Measurement is the most basic ingredient in any research, for any purpose.
Feedback development. There is a large base of lit- erature connecting feedback to improved performance both for individuals and groups (e.g., Pritchard, Youngcourt, Philo, McMonagle, & David, 2007). PM plays a critical role in the development of feedback. Specifically, performance must be measured to assess how an individual or team is performing, including what they are doing right, what they are doing wrong, and where improvements in performance can be made. These performance data can then be used to develop focused feedback, centered on identified strengths, weaknesses, and areas for improvement. Therefore, accurate and thorough PM is the first step in any feedback system. By accurately measuring and describing the performance of an individual, feedback can be used as specific instructions for performance improvement.
Training development and evaluation. The use of PM data for feedback development is quite similar to the use of PM data for training development. Another intervention designed to improve perfor- mance, training, aims to develop an individual’s or team’s knowledge or skills by providing information and opportunities for practice. It is important that training systems are designed to address specific deficiencies in employee performance, as they can often be costly and time consuming both to design and implement. PM data are a necessary first step in the development of training. These data are used to identify deficiencies and pinpoint knowledge, skills, and abilities in need of improvement. PM also plays a role in the assessment of the training system effec- tiveness. Specifically, performance must be mea- sured at the conclusion of the training program and linked to relevant outcomes to assess whether the training is imparting the desired knowledge or skills (i.e., whether there was learning) and ultimately contributing to the performance of the employees and organization as a whole (i.e., whether there was training transfer). Training is a cycle of providing instruction, assessing learning and outcomes, and adjusting instruction on the basis of that assessment. PM facilitates this cycle.
In addition to remedial efforts, training can also be used to provide new knowledge, skills, or attitudes. For example, nearly 88% of organizations with rev- enues exceeding $10 billion have executive develop- ment programs aimed at providing executives with
Wildman et al.
306
FIGURE 10.1. Critical considerations in performance measurement.
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 306
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
skills or knowledge that are not necessarily directly related to their current positions (Czarnowsky, 2008). PM can aid in the development of these pro- grams as well. By measuring the performance of top executives, an organization can establish criteria of what a successful executive looks like by focusing on the strengths of each individual that positively contri- bute to personal and organizational effectiveness. These criteria can then be used to develop training targeting future executive development efforts.
Performance evaluations. PM data can also be more simply used for evaluation purposes. Employee evaluations are yearly or quarterly assessments used to determine the comparative success of indi- viduals within an organization. Data from these evaluations, usually in the form of subjective ratings performed by supervisors, can then be used to deter- mine various human resources decisions such as promotions, salary changes, or bonuses. PM data for evaluation purposes often serve as the justification behind these types of decisions that must be made in all organizations. Given that performance evaluation data can often impact individual employees in salient and life-changing ways (e.g., firing, promotion), it is absolutely critical that performance data used for this purpose are accurate and nonbiased. Team-level PM can also be used for evaluative purposes, similarly to individual-level data.
Organizational planning. Up to this point, every purpose for PM discussed has focused on the indi- vidual or team level. However, PM is also critical and necessary at the overall organizational level as well. All organizational-level decision making and planning relies on accurate measurement of perfor- mance at the individual, team, and organizational level. For example, if an organization puts a new policy or program in place, they will undoubtedly need to measure performance at some point after implementation to assess whether that program is working as intended and to decide whether the pro- gram should be modified, expanded, or eliminated (Tannenbaum, 2006). Assessing big picture perfor- mance also allows for an organization to keep track of their organizational health, which can lead to high-level decisions such as mergers or acquisitions. Without PM at the organizational level, decisions
such as these would be made blindly. Finally, orga- nizations are interested in overall measures to inform decisions regarding human resources (e.g., recruitment, training).
What: Content of Measurement Once the purpose of the measurement has been identified, it is necessary to determine what corre- sponding content should be captured. Depending on the reason behind the PM, and how performance is being defined, there are numerous behavioral aspects that could be measured. For example, if the purpose of the measurement is to develop taskwork training for pilots, then measuring task-related per- formance would likely be the best choice of content. However, if the purpose is to look at how aircrews function together as a cohesive unit, teamwork- related behaviors should be the focus of measure- ment. As is described later in the chapter, there are many different types of performance that can be measured that focus on task performance, interper- sonal performance, or the outcomes of performance. Which type of performance is measured should be decided on carefully to best match the purpose of the PM system. This consideration relates heavily to the criterion problem and the importance of choos- ing performance measures that represent the con- ceptual criteria of interest.
Criteria: A deeper look. The conceptual criterion can be described as a verbal statement of the impor- tant outcomes related to a particular problem (Borman, 1991). Conceptual criteria are abstract statements of what is important to the stakeholder and represent the starting point that drives the devel- opment of performance measures. Essentially, con- ceptual criteria are the gold standard of what a highly successful employee, team, or organization would look like if performing at the highest level. Consequently, conceptual criteria are very subjective in nature. Subject matter experts (SMEs) can provide insight, but the bottom line is that the conceptual criteria should conceptually relate back to the orga- nizational mission and goals. Measures of effective- ness (outcomes) are developed on the basis of the conceptual criteria. These should be developed ratio- nally to ensure they map onto the conceptual crite-
Performance Measurement at Work
307
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 307
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
ria. Measures of effectiveness can be considered the operational definition of the conceptual criteria, with an evaluative component.
Muchinsky (2009) provided an example of con- ceptual criteria by using the example of a successful college student. He suggested that one important dimension is intellectual growth, noting that highly successful college students experience more intellec- tual growth than unsuccessful, or less successful, colleagues. Additionally, he pointed to emotional growth as a second dimension, positing that a col- lege education should allow successful students to clarify values and beliefs, aiding in their develop- ment and stability. A final dimension to be consid- ered might be citizenship, whereby successful college students desire to engage in civic activities and positively contribute to their surrounding com- munity. Muchinksy suggested that these three fac- tors are the defining criteria, or conceptual criteria, of what constitutes a successful college student. Yet, these are theoretical; therefore, the challenge is to convert these theoretical ideas of desired behaviors into something that can be quantifiably measured.
Performance measures: General characteristics. Generally, performance measures should provide information regarding products, services, or the tasks that individuals or teams complete to produce those products or services. Performance measures are essentially tools that let decision makers see how well individuals, groups, teams, or organizations are doing. Additionally, they provide insight into whether goals are being met, whether customers are satisfied, whether processes are indeed working as desired, and where improvements are needed.
Performance measures can also be multidimen- sional. There are numerous examples of this dimen- sionality. For example, number of accidents or injuries per million hours worked is one indicator of a company’s safety program. However, the cost of injuries provides additional information regarding safety program effectiveness. This type of measure provides more detailed information than just the first example, which is a single dimensional mea- sure. Essentially, whatever is measured must be expressed in measurement units that are meaningful given the entire purpose of measurement.
When: Timing of Measurement Another consideration focuses on when the con- struct will be measured. Performance is dynamic and changes over time. Processes that are happening at the beginning of a performance cycle may change or even be replaced with different processes at the end of the performance cycle. Therefore, the point during performance at which a construct is measured and the amount of times it is measured (i.e., once or repeated measures) may have a significant impact on what information is captured. For example, many performance measures can be considered “lagging measures” in that they capture performance out- comes long after the behavior that led to those outcomes occurred. End-of-the-year performance reviews and archival data are two good examples of this type of measurement. These sources are practical and useful measures of past performance, and they can be used to link specific performance behaviors to more distal organizational outcomes such as financial success. However, they may provide an inaccurate, or outdated, understanding of current performance, especially if the performance in question is likely to change quickly or often over time (i.e., is cyclical in nature).
Measuring performance throughout a perfor- mance period is advantageous because it provides a real-time understanding of what behaviors are actu- ally occurring that lead to the performance outcome. Several of the measures commonly used in the team literature, such as event-based measurement and communication analysis, take this approach. How- ever, the limitation with midperformance measure- ment is that usually it is a more intrusive method of measurement. If individuals are aware they are being observed or evaluated, there may be an issue with eliciting maximum versus typical performance, which is discussed in more detail in the following section.
One advantage of a repeated measures design for PM is the ability to determine the magnitude of any gains in performance. For example, assume perfor- mance is being measured to assess the effectiveness of a training program or some other intervention. A pre- and postmeasurement approach allows for a comparison of levels prior to training and levels posttraining. It is also a useful method for capturing the dynamic aspect of performance. Specifically, dif-
Wildman et al.
308
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 308
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
ferent performance processes may become more or less critical at different points during a performance cycle, and by measuring performance repeatedly, these changes in process can be captured. However, there are also limitations with this design. Without a control group, it is difficult to conclude with cer- tainty that any noted improvements in performance were specifically due to the intervention and not some outside influence.
Where: Fidelity of the Measurement Setting Another issue to consider when measuring perfor- mance is how characteristics of the setting will impact the process of PM. This issue pertains mostly to observational or rating-type measures, as knowl- edge tests and financial data are generally indepen- dent of the setting. Observational methods, however, are directly influenced by the realism of the mea- surement setting. In both laboratory and on-the-job settings, the level of fidelity influences the process of PM. Hays and Singer (1989) defined fidelity as “the similarity between the . . . situation and the opera- tional situation which is simulated” (p. 50). In the case of PM, this refers to how closely the measure- ment setting replicates the actual performance situa- tion it is intended to represent. Fidelity can be further defined in terms of two dimensions: (a) the physical characteristics of the measurement environment (i.e., the look and feel of the equipment and envi- ronment) and (b) the functional characteristics of the measurement environment (i.e., the functional aspects of the task and equipment).
Depending on the purpose and nature of the per- formance measures, different levels of fidelity will be more or less appropriate. For example, if the measure- ment is intended to capture day-to-day performance of employees on the job, and the job in question is highly dependent on various changes that occur in a fast-paced dynamic environment, a strictly controlled laboratory setting with a low level of fidelity may result in misleading findings. Imagine trying to mea- sure the performance of a team of emergency medical technicians (EMTs) while they are responding to a severe vehicle collision. In this situation, PM may be more accurate if gathered from the natural job envi- ronment or in a high-fidelity laboratory setting (i.e.,
simulation) designed to closely mimic the complex dynamic environment faced by the EMTs. If the sim- ulated environment does not accurately represent the potentially complex environmental factors that are inherent in situations commonly faced by EMTs (e.g., quickly changing medical status of victims, severe weather, vehicle fires or explosions), the mea- sures may not capture performance that is indicative of day-to-day actions. On the other end of the spec- trum, some tasks can be appropriately measured using lower fidelity situations. For example, an assembly line worker most likely could perform his or her task in an artificially contrived task simulation in relatively the same manner as he or she performs in the actual work environment. Overall, the level of fidelity of the setting, nature of the task, and purpose of the measurement must be considered in tandem when choosing the setting in which to conduct PM.
Another important issue to consider when exam- ining the fidelity of a measurement environment is the problem of maximum versus typical performance. Maximum performance can be defined as the highest level of performance possible to achieve under opti- mal conditions, whereas typical performance is the average performance on a day-to-day basis (Mangos & Arnold, 2008). Sackett, Zedeck, and Fogli (1988) provided an example to illustrate the differences between the two constructs, using grocery store register clerks. Typical performance was opera- tionalized as the average number of items scanned less the number of voids per shift, whereas maxi- mum performance was operationalized as the speed and accuracy of scanning items averaged across sev- eral timed observations. They found that the mea- sures were not statistically related, which suggests that typical and maximum performance are distinct constructs. Additionally, research has shown that each construct has different antecedents. For exam- ple, intelligence is more highly related to maximum performance and personality is more predictive of typical performance (Dubois, Sackett, Zedeck, & Fogli, 1993). There are cultural implications to maximum and typical performance as well. Dubois et al. (1993) found that Caucasians outperformed African Americans on typical performance; how- ever, the differences were minimal when looking at maximum performance.
Performance Measurement at Work
309
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 309
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Certain environmental cues can elicit maximum performance conditions, which can be problematic if the measurement is intended to capture typical, day-to-day performance. When individuals are acutely aware that they are being observed and eval- uated, they will likely try to perform to the best of their ability (Sacket et al., 1988). Consequently, maximum performance may be unintentionally elicited and measured. This same phenomenon can occur when observers are present in the natural work environment and the individuals being observed are aware of their presence. This is a signif- icant problem, as research has shown that both max- imum and typical performance are predicted by different variables (Campbell et al., 1993; Lim & Ployhart, 2004). Therefore, if the goal of measure- ment is to represent typical performance, the knowl- edge of being observed may trigger maximum performance instead and may distort findings regarding the relationships between performance criteria and predictor variables.
Fidelity is often associated with a trade-off in terms of the level of experimental control in a mea- surement setting. For example, although measure- ment in the operational work environment is as realistic (i.e., high fidelity levels) as possible, this usually makes it more difficult to isolate and identify the causes underlying performance because so many uncontrolled variables are freely influencing perfor- mance. In other words, because the experimenter does not design and control the setting of the mea- surement, there is the potential for any number of environmental factors to influence or, more impor- tant, confound, results. Therefore, it is often more difficult to assess the effectiveness of training or other performance interventions in the field because effects can be hidden by various outside influences. However, results found while measuring perfor- mance in the operational setting will most often be more externally valid than results found in more artificial or lower fidelity settings (i.e., lab setting), because data are collected directly in the environ- ment to which they are intended to generalize.
How: Measurement Techniques One final consideration when choosing or designing a PM system is the technique or approach for mea-
surement. Rather than first presenting the theoretical background for each level, and then the measure- ment techniques separately, we group them together within the realm of individual, team, and organiza- tional perspectives to ensure that each measurement technique is considered within the appropriate theo- retical context. Our hope is that this delineation will provide insight into the most commonly used mea- surement techniques at the different levels in addi- tion to providing the necessary context for why consideration of multilevels with regard to measure- ment warrants attention. Therefore, in the next sec- tions, we discuss each level (individual, team, and organizational) and describe common techniques frequently used for PM at that level.
It is absolutely critical to note that the measure- ment strategies discussed in each perspective are in no way used exclusively within that perspective. Many of the measurement tools mentioned are clearly applicable to, and consequently have been used across, multiple levels of PM. Additionally, the list of measurement strategies we provide is by no means exhaustive. However, as our primary goal is to compare and integrate three distinct streams of literature, we discuss each selected measurement strategy within the theoretical perspective in which it is discussed or most frequently used. We bring all three perspectives together at the conclusion of the chapter.
PERFORMANCE MEASUREMENT FROM THE INDIVIDUAL PERSPECTIVE
In the following section, we discuss PM approaches focused on capturing individual-level phenomenon. First, we define individual performance. Second, we describe several key theories of individual perfor- mance that have been developed over the years. Finally, we summarize and describe the most com- mon performance measurement approaches used at the individual level.
Defining Individual Performance A majority of the PM literature has been devoted to measurement at the individual level. The most basic resource in an organization is the individual employee, and therefore the first place to start man-
Wildman et al.
310
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 310
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
aging performance is at the individual level. For as long as organizational science has existed, the term performance has been misused and vaguely defined. Campbell et al. (1993) attempted to rectify this mis- use of the term by defining performance as synony- mous with behavior. In this view, performance must be the actions of the individual in question. Job per- formance, specifically, “includes only those actions or behaviors that are relevant to the organization’s goals and that can be scaled in terms of each indi- vidual’s proficiency” (Campbell et al., 1993, p. 40). Job performance at the individual level, simply stated, is what employees are hired to do.
Therefore, performance is not the outcome or the consequence of behavior; it is the behavior itself (Campbell et al., 1993). This distinction defines the difference between performance, effectiveness, and productivity. Performance is the actions taken by the individual, effectiveness is the “evaluation of the results of performance” and productivity is “the ratio of effectiveness to the cost of achieving that level of effectiveness” (Campbell et al., 1993, p. 41). Each of these terms is an independent construct. One can measure performance without evaluating that performance or without comparing that evalua- tion with cost. However, when trying to use PM data for any practical purpose, evaluating that perfor- mance is a critical step.
Theories of Individual Performance Several theories of individual performance have been developed looking at various aspects of perfor- mance such as job performance, organizational citi- zenship behavior, contextual performance, adaptive performance, integrated work role performance, and counterproductive work behavior. Each of these theories is described in more detail in this section.
Job performance behaviors. Along with their defi- nition of performance as synonymous with behavior, Campbell et al. (1993) also broke job performance down into eight major behavioral components: (a) job-specific task proficiency, (b) non–job-specific task proficiency, (c) written and oral communica- tion task proficiency, (d) demonstrating effort, (e) maintaining personal discipline, (f ) facilitating peer and team performance, (g) supervision or lead- ership, and (h) management or administration. Job-
specific task proficiency reflects the “degree to which the individual can perform the core substantive or technical tasks that are central to the job” (p. 46). Non–job-specific task proficiency is the degree to which the individual can perform tasks in the work- place that are not specific to a particular job (i.e., teamwork skills). Written and oral communication task proficiency is the proficiency with which a job incumbent can write or speak. Demonstrating effort is a reflection of the consistency, frequency, and willingness of an individual to demonstrate effort. Maintaining personal discipline involves the extent to which negative behaviors (i.e., alcohol and substance abuse, excessive absenteeism) are avoided at work. Facilitating peer and team performance is the extent to which an individual supports their peers. Supervision or leadership is the degree to which an individual engages in behaviors directed at influencing the performance of subordinates. Last, management or administration includes performance behaviors directed at management tasks such as articulating goals or monitoring progress. One critical contribu- tion of the Campbell et al. theory of performance is the conceptualization of performance as a multi- dimensional construct. By breaking job performance down into multiple components, they acknowledged that performance is not just one behavior that can be captured by one simple measure.
Organizational citizenship behavior. One limita- tion of the Campbell et al. (1993) model of job per- formance is that it focuses solely on task performance as defined by the job description. It does not account for behaviors that are not technically part of the job yet contribute to job performance. In response to this gap in the literature, several new performance con- cepts were developed, such as citizenship behavior (e.g., Borman et al., 2001). In a review of organiza- tional citizenship behavior (OCB), Podsakoff, MacKenzie, Paine, and Bachrach (2000) summarized the literature into seven core types of citizenship behaviors: (a) helping behaviors, (b) sportsmanship, (c) organizational loyalty, (d) organizational compli- ance, (e) individual initiative, (f) civic virtue, and (g) self-development. (See also Vol. 2, chap. 10, this handbook.)
Helping behaviors include helping others with work-related problems as well as actively preventing
Performance Measurement at Work
311
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 311
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
problems for others. Sportsmanship includes behav- iors such as not complaining even when inconve- nienced and generally maintaining a positive attitude in the face of difficulty. Organizational loyalty is composed of behaviors such as protecting, endorsing, and defending the organization and its objectives. Organizational compliance is another organizationally focused type of citizenship behavior that includes internalized and accepting the organization’s rules, regulations, and procedures even when not being directly observed. Individual initiative OCBs are behaviors in which the individual goes above and beyond minimum requirements in task-related situa- tions. Civic virtue describes a high level commitment to the organization as a whole. This commitment is displayed through behaviors such as attending volun- tary meetings and monitoring environmental changes that could impact the organization. Finally, self- development refers to voluntary behaviors aimed at improving one’s own knowledge, skills, and abilities.
Overall, OCBs pose an interesting dilemma for PM in that by definition they are not explicit require- ments of the job, and thus including them as part of a formal review or performance evaluation may be unethical. Simply stated, it may be inappropriate to evaluate an employee in terms of behaviors that are not explicitly stated as part of their job role, espe- cially if the evaluation is then used as the basis for pay and promotion decisions. If OCBs are included in formal performance reviews, this in essence makes them part of the job description, and therefore they are no longer extrarole. If an organization chooses to include OCBs as part of their official job descrip- tions, then this approach is appropriate. However, the defining feature of OCBs is that they are per- formed without being required (i.e., extrarole behav- iors; Organ, 1997); therefore, formally measuring them for evaluative purposes could potentially change the nature of these behaviors. In fact, there is an ongoing debate in the literature regarding whether measuring and evaluating OCBs changes the fundamental nature of the behaviors.
Contextual performance. The concept of contex- tual performance is very similar to citizenship behavior in that they both describe on-the-job behavior that is not directly recognized as part of the
job (i.e., it is not a job requirement) yet still con- tributes to job effectiveness. Contextual performance is behavior that contributes to organizational effec- tiveness through its impact on the psychological, social, and organizational context (Motowidlo, 2003). Borman et al. (2001) presented a refined model of contextual performance that categorizes behaviors as personal support, organizational sup- port, and conscientious initiative. Personal support includes behaviors such as helping others with tasks and showing courtesy and tact when interacting with others. Organizational support includes actions such as defending and promoting the organization. Conscientious initiative focuses on behaviors such as devoting extra effort to the job or taking advantage of opportunities for self-development. There is a noticeable amount of overlap between the conceptu- alizations of contextual behavior and OCB as described previously. The same issues regarding measurement of OCBs applies to measuring contex- tual behavior. Given that contextual performance is composed of behaviors that are not formally recog- nized as part of the job, it is unethical to evaluate individuals on the basis of contextual performance without their explicit knowledge. Accordingly, if contextual performance is required as part of a job, it is by definition no longer contextual.
Adaptive performance. The Campbell et al. (1993) model of job performance also does not account for work behaviors that contribute to effectiveness in dynamic, complex, uncertain, and interdependent settings (Griffin, Neal, & Parker, 2007). Pulakos, Arad, Donovan, and Plamondon (2000) developed a theoretically and empirically based model of perfor- mance focused on the concept of adaptivity, with eight dimensions of adaptive performance. This model is intended to assess how well individuals adjust or adapt to new conditions or unexpected job requirements. The eight dimensions of adaptive per- formance are (a) handling emergencies or crisis situa- tions; (b) handling work stress; (c) solving problems creatively; (d) dealing with uncertain and unpre- dictable work situations; (e) learning work tasks, technologies, and procedures; (f ) demonstrating interpersonal adaptability; (g) demonstrating cultural adaptability; and (h) demonstrating physically ori- ented adaptability.
Wildman et al.
312
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 312
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
The first dimension, handling emergencies or cri- sis situations, involves reacting appropriately in life- threatening or dangerous situations. The dimension of handling work stress includes remaining calm when faced with difficulties and effectively manag- ing frustration. Solving problems creatively refers to behaviors such as finding innovative ideas to com- plex problems and considering a wide range of pos- sibilities when solving a problem. Dealing with uncertain and unpredictable work situations is similar to handling emergency and crisis situations and handling work stress in that it involved reacting appropriately to a cue, but this dimension focuses on changing plans, goals, and strategies in response to unexpected events or situations rather than just remaining calm. Learning work tasks, technologies, and procedures is the most task-relevant dimension of adaptive performance and includes keeping up to date with changing procedures and technology nec- essary for the job. The dimension of demonstrating interpersonal adaptability includes being open- minded and considerate when dealing with other people and maintaining effective relationships. Demonstrating cultural adaptability specifically focuses on interacting with people from other cul- tures and adjusting behavior to make these inter- actions effective. Finally, demonstrating physically
oriented adaptability refers to adjusting to physical environmental conditions such as temperature or training to become more physically proficient.
Integrated work role performance. Bringing together several of the previous understandings of work performance, Griffin et al. (2007) recently developed an integrated model of work role perfor- mance (see Table 10.1). They proposed that context plays a major role in the behaviors that will be viewed as valuable performance in an organization. Specifically, they proposed that uncertainty in the environment influences to what extent roles can be formalized and that interdependence with the envi- ronment influences how embedded work roles are in the larger system. Their model attempts to address the difficulty of capturing the total set of perfor- mance dimensions in a job by cross-classifying the three levels at which work behaviors can contribute to effectiveness (individual, team, and organization) with the three different forms of work behavior (proficiency, adaptivity, and proactivity). This cross- classification resulted in nine subdimensions of work role performance: (a) individual task proficiency, (b) individual task adaptivity, (c) individual task proactivity, (d) team member proficiency, (e) team member adaptivity, (f ) team member proactivity,
Performance Measurement at Work
313
TABLE 10.1
Model of Positive Work Role Behaviors
Proficiency: Fulfills the Adaptivity: Copes with, Proactivity: Initiates prescribed or predictable responds to, and supports change, is self-starting
Individual work role behaviors requirements of the role change and future directed
Individual task behaviors: Behavior contributes to individual effectiveness
Team member behaviors: Behavior contributes to team effectiveness rather than individual effectiveness
Organization member behaviors: Behavior contributes to organization effectiveness rather than individual or team effectiveness
Note. From “A New Model of Work Role Performance: Positive Behavior in Uncertain and Interdependent Contexts,” by M. A. Griffin, A. Neal, and S. K. Parker, 2007, Academy of Management Journal, 50, p. 330. Copyright 2007 by the Academy of Management. Adapted with permission.
Individual task proactivity
Team member proactivity
Organization member proactivity
Individual task adaptivity
Team member adaptivity
Organization member adaptivity
Individual task proficiency
Team member proficiency
Organization member proficiency
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 313
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
(g) organization member proficiency, (h) organiza- tion member adaptivity, and (i) organization mem- ber proactivity.
Individual task proficiency represents the formal task performance that contributes to individual effec- tiveness. Team member proficiency moves one level beyond individual task proficiency and includes task- related behaviors that an individual engages in that contribute to team effectiveness. This includes behav- iors such as helping other team members or monitor- ing the work of other team members. Organization member proficiency describes task-related behaviors that contribute to organizational effectiveness, such as defending the organization’s reputation. This pat- tern of behavior remains consistent throughout the rest of the model. Specifically, individual task adaptiv- ity, team member adaptivity, and organization member adaptivity all refer to behaviors such as appropriately responding to changes in the environment that con- tribute to individual, team, and organizational effec- tiveness. Similarity, individual task proactivity, team member proactivity, and organization member proactiv- ity represent the three levels of self-starting, future- directed behavior. This model of performance is more robust than previously mentioned models in that it integrates the broad concepts of role behavior, adap- tive behavior, and proactive behavior and applies these concepts across multiple levels of analysis; however, it is still lacking in comprehensiveness (e.g., counterproductive work behaviors [CWBs]).
Counterproductive work behavior. Thus far, all theories of individual performance have focused on the positive behaviors job incumbents can engage in. However, humans are also capable of negative, or dysfunctional, work behaviors, and research has labeled this CWB (Sackett, 2002). CWB refers to any type of intentional employee behavior that is contrary to the organization’s interests. CWBs include various deviant acts such as theft, destruc- tion of property, drug abuse, and poor attendance. Some have argued that CWB is not a distinct con- struct but is rather a representation of the negative end of the citizenship behavior continuum. Recently, however, Sackett, Berry, Wiemann, and Laczo (2006) empirically supported that CWB is a sepa- rate and distinct construct from OCB.
Strategies for Measuring Individual Performance The most common approaches for measuring indi- vidual performance include performance appraisals, multiple-source ratings, objective measures, job knowledge tests, and work sample tests. It is impor- tant to note that these measurement strategies have not been used exclusively for capturing individual performance but are most often seen in the individ- ual realm. Further detail is provided for each strategy in this section.
Performance appraisals. Performance appraisals (PAs) are one of the most commonly used methods of PM in organizations. Traditionally, the term per- formance appraisal referred to a process involving a supervisor completing an annual report on an employee’s performance and discussing it with the employee in an interview (Fletcher, 2001). In a tradi- tional PA system, the supervisor prepares a written evaluation of the employee on the basis of informa- tion gathered from coworkers, customers, and any pertinent documentation regarding the employee (Aldakhilallah & Parente, 2002). Then the super- visor schedules a meeting with the employee to review his or her job description, his or her perfor- mance against this description, and the organiza- tion’s goals. The supervisor also addressees the employee’s career progress and identifies opportuni- ties for further development and improvement. This information is forwarded to higher management to be used in promotion and salary decisions. PA sys- tems are often used to make promotional decisions based on past performance, to identify skill deficien- cies and need for training, to make salary decisions, and to provide employees with feedback.
As with most PM approaches, PAs have received both praise and criticism. Some voice the merits of PA systems, including feedback, goal setting, career management, objective assessment, and legal protec- tion (Nickols, 2007). Also, some have claimed quite simply, “having a traditional PA system is better than having no system at all” (Aldakhilallah & Parente, 2002, p. 44), although this claim should be qualified to include only accurate and effective PA systems. It is likely that a bad PA system could actually be worse than no system in that it might force employees to
Wildman et al.
314
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 314
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
focus on the “wrong” behaviors—ones that are not really important to effective performance. Another advantage that has been mentioned previously relates to salary—that it is a motivational tool to improve employee performance, and therefore, salary deci- sions are often tied to the PA system (Rynes, Gerhart, & Parks, 2005). Finally, the outcomes of PAs provide the necessary justifications for decisions such as ter- mination, promotion, transfer, or a change in salary.
The more supported vein of thought, however, posits that supervisor-based PA systems are costly in terms of both time and money, and usually only result in negative emotions with very little demon- strable value (Nickols, 2007). Because the system is almost entirely controlled by the supervisor, there is a high chance that the employee’s appraisal will be based on the supervisors’ opinion alone. Depending on the ethical nature and personality of the super- visor in question, this could make PAs unfair or inaccurate. Data from an Internet survey revealed several other issues associated with PAs (Nickols, 2007). Often, PAs are accompanied by periods of reduced productivity; heightened negative emo- tional states such as anxiety, depression, and stress; and lowered morale and motivation. When they are linked to short-term rewards or consequences, they also tend to foster a focus on short-term goals at the sacrifice of long-term goals, which can result in neg- ative consequences in the long run. Additionally, people are often hesitant to convey negative feed- back, and therefore supervisors may engage in avoidance, delay, or distortion of PA feedback. Benedict and Levine (1988) found that, in particu- lar, female raters may delay appraisals, delay sched- uling feedback sessions, and more positively distort their ratings, especially when rating low performers. Because of these inherent issues, much of the research on PA has focused on making more objec- tive and accurate ratings (Fletcher, 2001).
Multiple-source ratings. Multiple-source ratings, also known as “360-degree feedback,” have been defined as “evaluations gathered about a target par- ticipant from two or more rating sources, including self, supervisor, peers, direct reports, internal cus- tomers, external customers, vendors, or suppliers” (Dalessio, 1998, p. 278). This PM method extends the PA concept by retaining subjective evaluations
as the main form of measurement, but this time including ratings from multiple sources, rather than from the supervisor alone. 360-degree feedback sys- tems were originally developed for purely develop- mental purposes, with no intention of evaluative use. Very quickly, however, organizations began integrating this method into their PA systems, mak- ing it evaluative rather than only developmental.
The use of 360-degree feedback at first seems like a logical choice for evaluative purposes (Waldman, Atwater, & Antonioni, 1998). This measurement system, unlike a traditional PA system, provides per- formance feedback from not only the supervisor but also from subordinates, peers and coworkers, clients (if applicable), and the self, reducing the chances for bias or unfair evaluations (Waldman et al., 1998). Specifically, by having peers rate the performance of other workers at their level, they are in a position to have a deep understanding of the job requirements and conditions, and therefore should provide more accurate ratings. They also generally have more opportunities to observe and monitor the work of the ratee because they work directly with them. Supervisors often are too far removed to have this level of understanding. On this basis, peer-report ratings seem to be a good alternative to self-report ratings because the tendency for inflating perfor- mance is reduced.
Additionally, the traditional PA system flows only downward from supervisors to subordinates; supervisors and higher management do not receive any feedback or evaluation. Having feedback flowing in all directions allows for the development of management and leadership, and subordinates and customers are in a good position to evaluate manage- rial performance (Morgeson, Mumford, & Campion, 2005). 360-degree feedback is also intended to be given anonymously, which theoretically should result in more honest (and therefore accurate) feed- back (Ghorpade, 2000). A review of the literature by Morgeson et al. (2005) delineated several other advantages of 360-degree feedback, such as an increase in information and formal feedback between employees, an increase in management learning, encouragement of goal setting and skill development, a change in corporate culture, and improved managerial effectiveness.
Performance Measurement at Work
315
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 315
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
However, 360-degree feedback has limitations as a performance evaluation approach. Waldman et al. (1998) suggested that the bidirectional nature of the system could encourage employees to deliberately try to sabotage the system, with supervisors striking deals with their subordinates to give high ratings in exchange for high ratings. They also contended that this type of sabotage is much less likely if feedback is purely developmental, as there is no immediate or direct tangible outcome associated with good or bad evaluations. Also, peer-report ratings come with problems of their own. Specifically, Viswevaran et al. (2005) found that peer ratings had more halo error than supervisor ratings.
Other issues arise regarding the improper imple- mentation of 360-degree feedback systems. Often, organizations implementing the process do not clearly define the mission and scope of the program beforehand or do not provide clear rules for infor- mation sharing. Consequently, feedback can end up unrelated to actual employee performance, and employees may be unable to use the results to develop goals or plans (Ghorpade, 2000). Toegel and Conger (2003) argued for the separation of developmental and evaluative tools because of the contradictions and competing goals between using 360-degree feedback as developmental versus an evaluative technique.
Objective measures. Another type of measure often used to assess individual performance is objective data. Specifically, indices such as absences, produc- tion rates, sales, or number of disciplinary cases can be used as a measure of an individual’s performance (Borman, 1991). Objective measures are appealing to many organizations because they are easy to gather and interpret, and are not as vulnerable to rater error or subjectivity as subjective measures. However, there are disadvantages to objective measures of individual effectiveness as well. To begin, the quality of an objective measure depends on what the stakeholder in question considers important. If the measurement system is designed to reduce absenteeism, then mea- suring absences is an appropriate choice. However, if the measurement system is intended to improve cus- tomer satisfaction, individual absences would be a very deficient measure. Another important factor to
consider with regard to objective measures is the notion of controllability. Controllability of measures can be conceptualized as the extent to which individ- uals (or teams) control the indicator of performance by varying the amount of effort they allocate to the measured tasks. Controllability—both the actual level and the perception of control—can impact motiva- tion, performance, and ultimately organizational effectiveness. For example, if a measure of perfor- mance is bed utilization in a hospital, doctors and nurses have very little control over this, as it is largely determined by the nature of the illnesses that affect the patients. The only way to affect this measure is for patients to remain longer than needed to “use a bed.” This is clearly not an effective practice. Therefore, it is critical that objective measures of individual perfor- mance are used only when they appropriately repre- sent the criteria of interest to the stakeholder and they are under the control of the individual being measured.
Job knowledge tests. Job knowledge tests are usu- ally written or computer-based tests that assess the extent of an individual’s knowledge regarding the content and procedures necessary for the job (Borman, 1991). Prior to developing a job knowledge test, a thorough job analysis should be conducted to gain a clear understanding of the knowledge, skills, and abilities necessary to perform the job. From this information, a set of items can be developed. As in any other written test, items can take many forms, such as multiple choice, true–false, or essay. Job knowledge tests are inherently best suited for positions that require high levels of declarative and procedural knowledge. It is important to note that job knowledge tests assess only the extent to which an individual can recall the appropriate information or procedure for a job but not his or her skill in applying that knowledge or performing that procedure. Job knowledge tests can therefore be appropriate performance measures when combined with other indicators (i.e., work-sam- ple tests; see below) or for jobs that are highly depen- dent on declarative knowledge (e.g., tour guides).
Work-sample tests. Work-sample tests are the practical, organizationally based equivalent of a laboratory-based measurement system (Cascio & Phillips, 1979). Because hands-on work-sample tests
Wildman et al.
316
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 316
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
require employees to engage in a simulated version of their normal taskwork specifically for measure- ment purposes, they are best suited for jobs with tasks that can be easily replicated in artificial set- tings. When broadly defined to include assessment centers, this type of measurement is suited for many different jobs. Example jobs that commonly use work-sample tests include baggage screeners, assem- bly line workers, aircraft pilots, or business execu- tives. These jobs include tasks that can be easily simulated or tasks with behaviors that can be easily recorded and scored. Jobs that would not be as well- suited for work-sample tests are more complex or ambiguously defined strategic planning or research positions, as these jobs require tasks that occur over a much longer period of time and tend to include aspects of performance that cannot be outwardly observed as scored during a short, simulated session.
Summary Several of the most commonly used measures of individual performance were described, including PAs, multiple-source ratings, objective measures, job knowledge tests, and work-sample tests (see Table 10.2 for a summary). These measures capture a wide range of performance behaviors and outcomes from a variety of perspectives. However, to reiterate,
the measures discussed do not represent an exhaus- tive list of the techniques for assessing individual performance. They represent only a sampling of the most commonly used and studies techniques.
PERFORMANCE MEASUREMENT FROM THE TEAM PERSPECTIVE
The science of teams has developed rapidly in recent years, and developments in the measurement of team performance have increased as well. Before we describe the measurement strategies most com- monly used to capture team performance, we define team performance and provide a summary of the key theories regarding team performance. (See also chap. 19, this volume.)
Defining Team Performance “Teams do not behave, individuals do” (Zalesny, Salas, & Prince, 1995, p. 99). This assertion creates a challenge for teams researchers in that it assumes teams do not engage in measurable behaviors. However, it can be argued that although teams do not behave, the behavioral interactions between team members and the behavioral processes the team engages in as a whole can be measured at the team level or aggregated to the team level from the individ-
Performance Measurement at Work
317
TABLE 10.2
Individual-Level Measurement Techniques
Technique Description References
Performance appraisals
Multiple-source ratings
Objective measures
Job knowledge tests
Work-sample tests
Traditionally, a measurement process involving a supervisor completing a written evaluation of an employee’s performance
Evaluations about a target participant collected from two or more source ratings such as self, supervisors, peers, customers; also referred to as “360-degree feedback”
Objective indices such as absences, production rates, sales, or number of disciplinary cases
Written or computer-based tests that assess the extent of an individual’s knowledge regarding the content and procedures necessary for the job
Hands-on simulated versions of everyday taskwork used for measurement purposes
Aldakhilallah and Parente (2002); Arvey and Murphy (1998); Nickols (2007)
Beehr et al. (2001); Ghorpade (2000); Morgeson et al. (2005); Toegel and Conger (2003)
Borman (1991); Ghalayini and Noble (1996); Pandey (2005); Paranjape et al. (2006)
Osborn and Campbell (1976)
Cascio and Phillips (1979)
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 317
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
ual level. Various models of team performance have identified critical behaviors that facilitate team per- formance, such as communication, coordination, mutual performance monitoring, and backup behav- ior (e.g., Marks, Mathieu, & Zaccaro, 2001; Salas, Sims, & Burke, 2005). Volumes have been devoted to the development of behavioral, team-level measure- ment systems (e.g., Brannick, Salas, & Prince, 1997).
When measuring team performance, it is critical to first start with a definition. Salas, Stagl, Burke, and Goodwin (2007) noted that team performance is often considered either a behavioral or cognitive act, and it is the resulting outputs that are considered per- formance outcomes. They suggested that team perfor- mance is multilevel, characterized by both taskwork (e.g., writing a paper) and teamwork (e.g., backup behavior) competencies exhibited by one or more team members, as well as team-level action (e.g., adaptation). This is not a new conceptualization. Early research has pointed to the multilevel nature of team performance as well (e.g., Kozlowski & Klein, 2000). Salas et al. (2007) postulated that team perfor- mance is a bottom-up emergent process, beginning with individuals and progressing toward teams. However, it is important not to ignore the top-down process inherent in team performance. Higher level (i.e., organizational) factors can directly impact, or have a moderating effect on, team performance.
Additionally, it is important to acknowledge that process measures are distinct from, although related to, outcome measures. Process measures capture and “describe the strategies, steps, or procedures used to accomplish a task” (Smith-Jentsch, Cannon-Bowers, Tannenbaum, & Salas, 1998, p. 62). Outcome mea- sures “evaluate the quantity or quality of the end result of those processes” (Smith-Jentsch et al., 1998, p. 62). Although outcome measures are ultimately what the researcher or organization is interested in predicting or managing, it is necessary to measure both process and outcome. This is due to the fact that outcome measures are influenced by much more than just individual or team performance. Outcomes can be affected by environmental factors, situational factors, or even luck.
Cannon-Bowers and Salas (1997) suggested that team PM should capture both processes and out- comes at the individual and team levels. They posited
that outcome measures are not very diagnostic because they often do not indicate the underlying causes of that outcome. This is where process mea- sures provide additional information by capturing the underlying behavioral mechanisms of performance. Therefore, it is critical to measure processes as well as outcomes for a robust and comprehensive under- standing of team performance.
Theories of Team Performance Salas et al. (2007) conducted a thorough review of the literature to determine what is known regarding team performance. They found over 130 models of team performance or effectiveness that addressed at least three constructs believed to be relevant in the nomological network of team performance. The inclusion criteria prohibited inclusion of the thou- sands of additional “models” that only considered one or two constructs. They reviewed 11 of these models (e.g., Gersick, 1988; Hackman, 1987; Nieva, Fleishman, & Reick, 1978), noting that their selec- tion in no way invalidated the significance of the remaining models. On the basis of this review, Salas et al. (2007) created an integrated model of team effectiveness. This model attempted to provide a comprehensive snapshot of the variables that impact team effectiveness.
Below, we use a similar approach in summarizing the team performance literature. Space precludes us from thoroughly examining all potential models of team performance; therefore, we highlight just a selection of the plethora of team performance mod- els. However, through the above discussion, we wish to illustrate that there is no universally accepted model of team performance to use as a departure point for measurement purposes.
Team processes. Capturing performance at multi- ple levels is also a critical part of the framework presented by Cannon-Bowers and Salas (1997). Specifically, they believed measures of both individual- and team-level competencies are necessary for team PM. Teams are made up of indi- viduals, and teamwork is composed of individual behaviors, so individual-level measurement is criti- cal. However, individual-level measurement is not sufficient, as team performance is more than the
Wildman et al.
318
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 318
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
sum of its parts. There must also be team-level mea- surement to capture the processes and outcomes emerging from team interactions.
Marks et al. (2001) recently delineated 10 core team processes that occur throughout the perfor- mance cycle of a team, such as mission analysis, coor- dination, and conflict management (see Table 10.3). Completing the circle, the outcomes of team perfor- mance are also multidimensional. Prior research has examined performance outcomes such as SME ratings of operational readiness in experienced mili- tary battalions (Lim & Klein, 2006), performance of
undergraduate students in a low-fidelity flight simula- tion (Mathieu, Heffner, Goodwin, Salas, & Cannon- Bowers, 2000), and safety and efficiency in air-traffic controllers (Smith-Jentsch, Mathieu, & Kraiger, 2005). The definition, and measurement, of team per- formance varies widely across different populations, settings, and purposes. With so many interrelated inputs, processes, and outcomes, it is impossible for one measurement tool to capture every aspect of team performance simultaneously.
Team performance is also a dynamic phenome- non, which means that static measurement systems
Performance Measurement at Work
319
TABLE 10.3
Team Processes
Process Definition Transition processes
Mission analysis formulation and planning
Goal specification
Strategy formulation
Action processes
Monitoring progress toward goals
Systems monitoring
Team monitoring and backup behavior
Coordination
Interpersonal processes
Conflict management
Motivation and confidence building
Affect management
Note. From “A Temporally Based Framework and Taxonomy of Team Processes,” by M. A. Marks, J. E. Mathieu, and S. J. Zaccaro, 2001, Academy of Management Review, 26, p. 363. Copyright 2001 by the Academy of Management. Adapted with permission.
Interpreting and evaluating the team’s mission, including identifying its main tasks as well as the operative environmental conditions and team resources available for mission execution
Identifying and prioritizing goals and subgoals for mission accomplishment
Developing alternative courses of action for mission accomplishment
Tracking task and progress toward mission accomplishment, interpreting system information in terms of what needs to be accomplished for goal attainment, and transmitting progress to team members
Tracking team resources and environmental conditions as they relate to mission accomplishment, which involves (a) internal systems monitoring (tracking team resources such as personnel, equipment, and other information that is generated and contained within the team) and (b) envi- ronmental monitoring (tracking the environmental conditions relevant to the team)
Assisting team members to perform their tasks; assistance may occur by (a) providing a teammate verbal feedback or coaching, (b) helping a teammate behaviorally in carrying out actions, or (c) assuming and completing a task for a teammate
The process of orchestrating the sequence and timing of interdependent actions
Preemptive conflict management involves establishing conditions to prevent, control, or guide team conflict before it occurs; reactive conflict management involves working through tasks and interpersonal disagreements among team members
Generating and preserving a sense of collective confidence, motivation, and task-based cohesion with regard to mission accomplishment
Regulating member emotions during mission accomplishment, including (but not limited to) social cohesion, frustration, and excitement
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 319
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
capture only a cross-section rather than the entire performance episode as it progresses over time. Marks et al. (2001) specifically focused on this temporal aspect of team performance when develop- ing the previously mentioned 10 core teamwork processes. They posited that different team processes are critical during different phases of task execution. Specifically, teams cycle through action and transi- tion phases, and depending on which phase the team is currently in, the processes occurring will dif- fer. Therefore, measuring team performance at one point in time may paint a completely different pic- ture than measuring it at a different point in time or even at multiple points in time. Researchers have long been working to overcome the measurement challenges inherent in the complex, multidimen- sional, dynamic nature of team performance (e.g., Salas, Priest, & Burke, 2005).
Guzzo and Dickson (1996) reviewed the literature on team performance and noted empirical support for several important team process and outcome variables. Research has found positive associations between cohesiveness and performance (Evans & Dion, 1991; Guzzo & Shea, 1992). For example, Evans and Dion (1991) found that 18% of variance in performance was accounted for by cohesion after cor- recting for measurement error. Campion, Medsker, and Higgs (1993) looked at several variables, includ- ing composition. They found a positive relationship between team size and effectiveness (r = .23, p < .05) and a composite measure of composition including heterogeneity, flexibility, size, and preference for group work (r = .21, p < .05). Although Campion et al. found no significant effect, or a negative one (r = −.05, ns), on heterogeneity and performance, Watson, Kumar, and Michaelsen (1993) found time to be influential, specifically that heterogeneous teams (operationalized by cultural diversity) who worked together long enough overcame the perfor- mance deficits relative to homogeneous teams. Others have found positive relationships between perfor- mance and (a) familiarity (e.g., Goodman & Leyden, 1991; finding increases from 1.8% to 11% in produc- tivity), (b) leader expectations (e.g., Eden, 1990; ω2 = .19 and .17 for performance operationalized as theo- retical specialty and practical specialty, respectively), (c) leader mood (e.g., George & Bettenhausen, 1990;
r = .43, p < .01, for prosocial behavior, which was significantly correlated with sales performance), (d) motivation as defined by team efficacy and group potency (e.g., Gully, Incalcaterra, Joshi, & Beaubien, 2002; ρ = .41 and .37, respectively), (e) self-efficacy (e.g., Earley, 1994; effort and self-efficacy accounted for a .66 change in R2 over demographic variables alone, with all variables together accounting for 69% of the variance in performance), (f) team goals (e.g., Weingart & Weldon, 1991; r = .45, p < .05) and (g) quality of feedback (e.g., Pritchard, Harrell, DiazGranados, & Guzman, 2008; r = .45, p < .01). The variability in these findings suggests the presence of moderating variables. Considering the negative effects that Campion et al. found, these moderating variables could change not only the magnitude of the relationships but also the direction. This is important to note for measurement. Because team performance is complex, if relevant impacting variables are not measured, it may lead to incorrect conclusions regarding team performance.
Salas, Sims, and Burke (2005) took a different approach to team performance with the advancement of a theory called the “Big Five in Teamwork.” Their goal was to present a parsimonious framework, high- lighting the essence of teamwork. They theorized that teamwork is composed of five core processes: (a) team leadership, (b) team orientation, (c) mutual performance monitoring, (d) backup behavior, and (e) adaptability. They noted the importance of three additional variables of interest: (a) shared mental models (SMMs), (b) closed-loop communication, and (c) mutual trust. This model incorporates not only core team skill-based competencies but also impor- tant affective competencies essential for effective team performance. Current research is focusing on the development of performance measures based on this model (e.g., Wiese et al., 2006).
Team adaptability. Burke, Stagl, Salas, Pierce, and Kendall (2006) proposed that a critical skill of any team, and by extension, a critical skill to be mea- sured, is the ability of a team to adapt. Team adapta- tion is defined as a change in team performance resulting from an identified cue or cue pattern, lead- ing to effective and efficient outcomes for the entire team (Burke et al., 2006). Team adaptability is cru-
Wildman et al.
320
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 320
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
cial to organizations, as teams are called on to make decisions and solve problems in complex, dynamic environments with increasing frequency. Adaptation of team performance processes is of great importance to those working to understand effective team per- formance and how teams should alter in response to rapidly changing task demands.
Burke et al. (2006) developed a cross-level mixed determinants model that integrates several perspec- tives of organizational theory. This multidisciplinary, multilevel, and multiphasic model is one of the most comprehensive conceptualizations of team adapta- tion available in the literature. The adaptive cycle comprises four process-oriented phases: (a) situation assessment, (b) plan formulation, (c) plan execution, and (d) team learning, as well as emergent cognitive states that function as both proximal outcomes and inputs throughout the cycle. Operating on the assumption that different team processes are critical at different phases of the team cycle (Marks et al., 2001), Burke et al. posited that plan formulation consists of transition processes and that plan execu- tion consists of action processes, with interpersonal processes occurring in both.
Team cognition. The study of individual mental models quickly made the leap to the team level, resulting in the concept of SMMs. SMMs are mea- sured by capturing the amount of “sharedness,” or overlap, between a set of mental models, usually through an aggregation technique. Other team-level cognition constructs that can be measured as a dimension of team performance include transactive
memory systems, team situation awareness, and metacognition (e.g., J. R. Austin, 2003; Hinsz, 2004; Prince, Ellis, Brannick, & Salas, 2007; see Table 10.4).
Kraiger and Wenzel (1997) suggested that SMMs can be measured on the basis of three key compo- nents: knowledge, behaviors, and attitudes, including perceptions, reactions, and structures. They provided numerous methodologies for measuring these three components of SMM, such as card sorts, structural assessments, and attitude perception surveys. They provided several hypotheses with regard to frame- work of antecedents, outcomes, and components of SMMs. Kraiger and Wenzel also noted another important issue with regard to measuring SMMs: measure weighting. They postulated that what mea- sures to use and how to weight them are very situation specific yet cautioned that any selected measures should be sensitive to factors such as organizational culture.
Strategies for Measuring Team Performance There is no one universally accepted measure of team performance. Guzzo and Dickson (1996) defined team performance effectiveness on the basis of earlier work by Hackman (1987) and Sundstrom, De Meuse, and Futrell (1990). They suggested that team perfor- mance effectiveness is characterized by (a) team out- puts (e.g., quantity or quality, customer satisfaction), (b) consequences for members, or (c) an increase in ability of a team to effectively perform at a later time. The approaches described below tend to fall under one or more of these three areas.
Performance Measurement at Work
321
TABLE 10.4
Components of Team Cognition
Component Definition Source
Team mental models
Transactive memory system
Metacognition
Team situation awareness
Team-level stable mental representations, including key knowledge about undertaking team tasks related to both teamwork and taskwork
A cognitive system teams use to encode, store, and retrieve information
What group members know about the way groups process information
The mental representation associated with a dynamic understanding of the current situation that is developed by team members moment by moment
Rico et al. (2008, p. 167)
Lewis (2003)
Hinsz (2004, p. 35)
Rico et al. (2008, p. 167)
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 321
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Observational approaches. Observational methods of PM, at first glance, are enticing because they offer the objectivity of an outside observer. However, evidence has shown that techniques that use human observers to rate performance have very low inter- rater reliability (Fowlkes, Lane, Salas, Franz, & Oser, 1994). This could be because of several reasons, such as inadequate rater training, leniency effect, or halo effect. Because of this reliability issue, some observa- tional methods have been developed that focus on capturing more objective data such as frequency or existence of specific performance dimensions rather than subjective ratings of behavior as posi- tive or negative. Examples of both rating-based and checklist-based observational methods are described in the following sections. In both categories, mea- surement requires a third party observer to capture performance.
Behavioral observation scales. The behavioral observation scale (BOS) approach to PM involves the use of observers who provide subjective ratings of the frequency of team performance. In this method, an observer physically watching a team uses a Likert- type scale to rate the amount of times that the team engages in a certain specified process. For example, an observer may be recording the frequency of infor- mation exchange in a team. Behaviors representative of information exchange would be rated on a scale from 1 to 5, representing intervals between none and always. Specifically, the item may ask “How often does this team share information about the environ- ment with each other?” The observer would then rate the behavior as happening never, seldom, some- times, frequently, or always on a 1-to-5 scale.
The distinguishing characteristic of this measure- ment method is that it assesses the typical behavior of a team over time. However, this characteristic also makes ratings more susceptible to recency effects. Over time, observers may ignore earlier behavior in favor of the more recently observed behaviors. For example, a team may begin their task by sharing information 50 times an hour but may slow down to 20 times an hour as they become more familiar with the task and each other. Because the most recent behavior viewed by the observer was a subjectively “low” 20 utterances an hour, they may rate this team as “seldom” communicating, when on average they
were actually communicating 35 times an hour. Another problem with the BOS method is the subjec- tivity of the rating scales. Specifically, frequency is a very subjective concept. Some people may consider the same objective amount of communication (e.g., 50 utterances) as never communicating, whereas others may consider that as always communicating. Finally, there is a temporal aspect that bears discus- sion. Consider the information exchange example above. Perhaps in the beginning, teams should have been communicating 50 times, but as team members become more familiar with each other, they should only communicate 20 times. If the measurement scale is not adapted to reflect optimal levels of per- formance, teams may inadvertently receive low rat- ings (i.e., 20 times per hour as low), when in reality that amount of information exchange is quite appro- priate for effective performance.
Behaviorally anchored rating scales. The behav- iorally anchored rating scales (BARS; Smith & Kendall, 1963) method is very similar to the BOS approach in that it requires an outside expert to observe, classify, and rate behavior (Kendall & Salas, 2004). Originally, this measurement technique was created to evaluate individuals; however, BARS has been readily adopted for use within the team perfor- mance arena. Just as in BOS methods, the observer rates the occurrence of team behavior on a numerical scale. However, rather than simply rating the fre- quency of behavior within the team, behavior is rated in regard to quality. There are specific examples of high-quality and low-quality behavior attached, or anchored, to each rating point in the scale. For exam- ple, the scale may range from 1 to 5, with 1 being poor behavior and 5 being excellent behavior. The behavioral anchors are usually generated by SMEs who provide performance episodes that represent both exceptional and unacceptable performance in the specified situation. These behavioral examples, in the form of short written descriptions, are intended to facilitate more accurate ratings by observers by mak- ing sure that similar behaviors are rated as the same number across raters. The point of this measurement technique, as used with teams, is to capture the rela- tive frequency of a behavior as well as providing con- text as to where that frequency lies on a good to bad continuum. Essentially, just noting that a behavior
Wildman et al.
322
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 322
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
occurred 20 times does not provide evaluative infor- mation as to what level of quality that frequency rep- resents. Therefore, BARS provide more usable information than BOS in most cases.
There are some issues inherent in the BARS method, however. Because behavioral anchors focus on specific types of behaviors, observers tend to watch for those types of behaviors and rate perfor- mance on the basis of them regardless of the possibil- ity that the team is using an equally successful, but different, behavior and regardless of the overall per- formance (Kendall & Salas, 2004). This issue actually relates back to the issue of criterion dimensionality, meaning several different approaches can be used to successfully complete a given task. Therefore, BARS are best suited for PM in work situations that requires very specific behavioral responses and in which cri- terion dimensionality is not a likely possibility.
Event-based performance measurement. Event- based measurement techniques, often used in train- ing exercises, are “event-based” because they involve systematically scripting events into a relevant exer- cise, task, or scenario to trigger specific behavioral responses from the team in question (Fowlkes et al., 1994). In this measurement approach, similarly to BOS and BARS, specific behaviors are identified by SMEs as critical, and these behaviors are recorded as they are observed. It is the high level of control over the appearance of relevant behaviors that makes this a unique measurement technique (Salas, Burke, Fowlkes, & Priest, 2003). The intention of scripting scenarios and identifying behaviors a priori is to increase the level of reliability in measurement. Additionally, event-based measurement is distinct from BOS and BARS techniques because behaviors are not necessarily rated on the basis of quality. One specific method for event-based team PM is known as the targeted acceptable responses to generated events or tasks (TARGETs) methodology (Fowlkes et al., 1994).
Targeted acceptable responses to generated events or tasks. In the TARGETs method of PM, a checklist of specific, observable behaviors is generated on the basis of the purpose of the observer and the task characteristics of the team being observed. For example, TARGETs for a helicopter aircrew may
include statements such as “Pilots question unsafe navigation procedure” or “Pilots acknowledge com- munications” (Fowlkes et al., 1994, p. 52). These are specific, easily observable behaviors that an observer could capture during team task execution. Therefore, a trained observer could simply watch the team in action and check off boxes as each behavior occurs. One important characteristic of the TARGETs method is that the events within the sce- nario being measured are controlled to elicit the behaviors of interest, because routine scenarios gen- erally will not result in a wide enough range of observable behaviors. This makes the TARGETs methodology of PM, or any other event-based approach, much better suited for laboratory-based research than for field or practical use. Additionally, it may be noted that this particular technique (among others used within team PM) can be equally applied to individual PM. It is important to remember that team performance is frequently measured at the indi- vidual level and then aggregated to the team level. Therefore, it is important to use measures that fully capture individual performance yet can account for the complexities of team performance.
Team dimensional training. Another form of event-based measurement, specifically designed for training teamwork skills, is team dimensional train- ing (TDT; Smith-Jentsch et al., 1997). This training system focuses on measuring and improving four core teamwork behaviors: information exchange, initiative or leadership, supporting behavior, and communication. As in all event-based approaches, a training scenario is carefully designed to provide ample opportunities for the trainees to engage in the four teamwork behaviors. As the trainees go through the scenario, an instructor observes and records examples of strong and weak execution of the team- work behaviors using a coding sheet. The coding sheet includes a column with the times when each scripted event should occur and a column where the instructor writes a description of how team performs in reaction to those events.
Communication analysis. Another common method for examining teamwork performance has been through the analysis of communication tran- scripts (Salas et al., 2003). The analysis of communi-
Performance Measurement at Work
323
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 323
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
cation data generally takes two forms: content analy- sis and flow analysis. Content analysis focuses on analyzing the linguistic content of the communica- tion data, such as the topics of discussions, the fre- quency of specific words, or the frequency of questions asked. Typically in content analysis, com- munication utterances are categorized into groups that represent behavioral constructs such as infor- mation exchange or backup behavior, and the fre- quency of those utterances represent the amount of that behavior occurring. Flow analysis differs from content analysis in that it ignores the meaning or content of the communications and focuses only on the pattern of the team’s interactions. For example, flow analysis might investigate whether there is a consistent pattern in the types of communications occurring, such as responses occurring immediately after questions.
This is a unique method for during-performance PM in that it requires no participation during the team task except for the presence of audio or video recording equipment. The actual evaluation of per- formance is performed post hoc. Archival measure- ment such as communication analysis is appealing because the source of the data is permanent, easy to access, and objective in nature. The problem is that measures capture performance well after it has occurred and therefore may draw an outdated pic- ture. Additionally, analysis of communication tran- scripts requires the prior development of a coding or classification scheme to guide coders in interpreting communication data in the context of teamwork. As this technique requires outside raters to qualita- tively assess behavior, it is critical that the raters are properly trained in the classification scheme. Classification schemes for communication analysis can take many forms. Communication transcripts could be coded for the frequency of communication, the content of communication, the pattern commu- nication, or a combination thereof. For example, the number of request for information could be counted, as well as whether these requests for infor- mation were regarding the taskwork of others or the environment. These communication instances could be considered as representations of team and sys- tems monitoring, which are two of the critical team processes identified by Marks et al. (2001).
Automated measurement. Automated measure- ment techniques are one of the more recently dis- cussed methods in team performance literature. This is not a set of specific measurement tools but is rather a specific strategy for implementing various perfor- mance metrics. Specifically, automated computer systems can be used to continuously monitor team processes such as communication utterances or body movements (Kendall & Salas, 2004). The recorded behaviors are compared against an expert standard, and the system can provide automated feedback to the team. Automated measures are less obtrusive than other measures (Salas, Priest, & Burke, 2005), but they can be used to measure only overt behaviors and are often quite expensive to implement and maintain.
Summary The team-level measurement techniques discussed in this section included observational methods such as BOS and BARS, event-based methods such as TARGETS and TDT, as well as communi- cation analysis and automated measurement (see Table 10.5). There are clearly some parallels between team-level and individual-level measure- ment in that BOS and BARS are both rating methods just as PAs or multiple-source feedback are. However, the team PM literature tends to be heavily focused on the use of measurement for feedback and training purposes, and places a heavy emphasis on the mea- surement of performance at multiple levels. As we move on to the organization PM literature, the emphasis shifts from multiple levels of analysis to multiple simultaneous dimensions of performance. This point is more fully developed throughout the remainder of this chapter.
PERFORMANCE MEASUREMENT FROM THE ORGANIZATIONAL PERSPECTIVE
The third level of performance measurement focuses on the processes and outcomes of the organization as a whole. In the following sections, we define orga- nizational performance, summarize the key theories of organizational performance, and present the most common measurement strategies for capturing organizational-level phenomena.
Wildman et al.
324
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 324
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Defining Organizational Performance The larger system under which organizational teams are embedded provides the context for team perfor- mance (Guzzo & Dickson, 1996). Researchers have long argued that there has been a lack of emphasis on tying team performance to the overall organizational performance (Levine & Moreland, 1990). McGrath (1991) argued that teams are partially nested and loosely coupled to the broader organization. This refers to the fact that individuals are often part of more than one team and that teams can be part of more than one organization. Team performance, thus, impacts organizational performance.
The literature pertaining to organizational PM is incredibly vast and ranges from theories and reviews of performance (e.g., Folan & Browne, 2005; Herman & Renz, 2008) to case studies and descriptions of specialized PM systems developed for individual companies (e.g., Bhasin, 2007; Khan & Wibisono, 2008). A comprehensive review of this literature is far beyond the scope of this chapter; thus we focus on several of the more commonly studied theories and strategies for organizational PM, as well as some of the newest approaches.
Traditionally, scholars of organizational perfor- mance have focused on measuring financial outcomes (Ghalayini & Noble, 1996). Financial outcomes such as return on investment, productivity, or sales per employee provide a concrete, easily accessible mea- sure of overall organizational performance. Yet in terms of assessing the entire organizational perfor- mance domain, these measures may be deficient. Therefore, the focus on organizational-level PM has shifted from financial outcomes to more integrated measures of performance that include financial per- formance along with other dimensions of perfor- mance such as customer service and organizational learning. Some define organizational performance from a systems management approach, and therefore collect measures of both internal and external perfor- mance information (Jensen & Sage, 2000). Only very recently has organizational performance literature begun to conceptualize performance as a process and explore the measurement of processes in relation to strategic goals and company policy (Nenadal, 2008).
Theories of Organizational Performance The measure of organizational-level performance used depends on the model of organizational effectiveness
Performance Measurement at Work
325
TABLE 10.5
Team-Level Measurement Techniques
Technique Description Source
Behavioral observation scales
Behaviorally anchored rating scales
Targeted acceptable responses to generated events or tasks
Team dimensional training
Communication analysis
Automated measurement
A subjective evaluation of team performance completed by an uninvolved third-party observer, usually based on a Likert-scale
A subjective evaluation of team performance completed by an uninvolved third-party observer that includes illustrative exam- ples of good and bad behavior on which to anchor the ratings
A measurement system that involves scripting a scenario to elicit behaviors of interest and then completing a checklist of those behaviors
A form of event-based measurement that focuses on measuring and providing feedback regarding information exchange, initia- tive or leadership, supporting behavior, and communication
The analysis of communication data or transcripts for indicators of performance behaviors
Any measurement system that uses automated computer systems to continuously monitor and record team processes such as communication or body movements
Dominick et al. (1997); Kendall and Salas (2004); Taggar and Brown (2001)
Kendall and Salas (2004); Motowidlo and Borman (1977)
Dwyer et al. (1997); Fowlkes et al. (1994)
Smith-Jentsch et al. (1997, 2008)
Bowers et al. (1998); Dong (2005); Landauer et al. (1998); Muniz et al. (1996)
Kendall and Salas (2004); Salas, Priest, and Burke (2005)
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 325
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
serving as the theoretical basis. Ahmed (1999) reviewed several of the most widely studied models of organizational effectiveness, including the goal model, system model, internal process model, human relations model, and the political approach. These models differ in their emphasis and prioritiza- tion of different dimensions of performance.
The goal model (Georgopoulos & Tannenbaum, 1971) defines organizational effectiveness in terms of an organization’s achievements of its stated official goals. This is one of the earliest and most dominant models of effectiveness, given it directly attempts to align PM with the organizational strategy. The sys- tem model of effectiveness views organizations as open systems that work in a close relationship with the external environment. In this model, effective- ness is seen as the organization’s ability to survive and adapt in a changing environment. The internal process model conceptualizes effectiveness as the extent to which the internal business processes of an organization are smooth, orderly, continuous, predictable, and with minimal conflict. The human relations approach has a person-centered focus and maintains employee needs (i.e., employee satisfac- tion, self-efficacy) as the most important aspect of organizational effectiveness. The political approach is a very unusual model of organizational effective- ness, and it uses criteria such as “responsiveness, accountability, representativeness, and adherence to democratic values” (Ahmed, 1999, p. 544). Clearly, each of these theories focuses on very different aspects of organizational performance and would require very different methods of PM.
Total quality management. From a business per- spective, total quality management (TQM) repre- sents a management philosophy of satisfying the customer’s requirements continually, at a low cost, by involving everyone’s daily commitment (Kanji, 1990). TQM defines quality as a process that must be managed. The objective of TQM is to train all lev- els of an organization to accept a new set of rules, methods, and lifestyles that focus on continuous improvement (Chung, Tien, Hsieh, & Tsai, 2008). There are four stages to the TQM process: (a) identi- fying and collecting information about the areas that need improvement in the organization, (b) making
sure that management understands and accepts the TQM philosophy, (c) identifying and resolving issues by involving all of management and supervi- sion in a proper scheme of training and communica- tion, and (d) starting new initiatives with new targets and spreading the improvement process to all aspects of the organization including supplier and customer links. TQM has a definite multilevel per- spective in that it advocates a philosophy of continu- ous PM and improvement in all systems within a business. A recent study examined the value of TQM in actual organizations and found that 15 organiza- tions that used TQM had above-average financial ability (Chung et al., 2008).
Knowledge management. The knowledge manage- ment (KM) approach to PM budded out of the orga- nizational management paradigm shift from focusing on managing physical goods to focusing on manag- ing intangible assets (Nielsen, 2005). In this manage- ment philosophy, knowledge is viewed as a critical resource in an organization, and management of this resource can lead to competitive advantage. Of course, to manage knowledge in an organization, it must be measured. The literature dealing with KM can generally be grouped into content and process perspectives (Nielsen, 2005).
The content perspective looks at the “what” of knowledge in organizations, specifically examining the categorization and transferability of different types of knowledge in organizations (Nielsen, 2005). Organizational knowledge can be categorized into explicit and tacit knowledge. Explicit knowledge is the knowledge that is exchanged using formal, systematic language and includes things such as explicit facts (Nielsen, 2005). Often, this knowledge can be found in instantiated formats such as books, databases, and computer programs (Small & Sage, 2005/2006). Tacit knowledge, on the other hand, is more intuitive, non- verbalized knowledge that is not directly articulated but still exists (Nielsen, 2005). This type of knowl- edge is hard to put into words and usually is rooted in contextual experiences (Small & Sage, 2005/2006). Other content-based KM research has looked at the holistic concept of organizational knowledge and the distinctions between data, information, and knowl- edge (Small & Sage, 2005/2006).
Wildman et al.
326
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 326
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
The process perspective of KM focuses on the “how” of organizational knowledge and therefore looks at the different stages of information process- ing in which knowledge is “embodied, embedded, embrained, encultured, and encoded” within mem- bers of the organization (Nielsen, 2005, p. 4). One goal of process perspective KM research is to uncover how the processes for accumulating and internalizing knowledge can be enhanced or improved (Nielsen, 2005).
Strategies for Measuring Organizational Performance The approaches used to measure organizational performance have changed over the years. Earlier measures included simpler outcomes, such as finan- cial data and balanced scorecards (BSCs). More complex, integrated measures have emerged, such as Six Sigma and ProMES. Each of these approaches in described in more detail in this section.
Financial data. As mentioned previously, tradi- tional organizational PM literature focused on finan- cial outcome measures. Financial measures are easy to obtain, easy to interpret, address the bottom line, and often are already recorded in organizational reports, thereby removing the need for a separate PM system. This makes them an enticing option for organizations looking to measure their overall levels of performance. Productivity, earlier defined as the ratio of effectiveness to the cost of achieving that level of effectiveness, has been one of the most widely used indicators of financial performance in organizations (Ghalayini & Noble, 1996).
Although financial performance measures appear to be a simple, easy-to-use option for measuring organizational performance, the enticing simplicity of financial data brings with it many inherent prob- lems. One of the problems attributed with financial measures such as return on investment is the poten- tial for managers to engage in short-sighted decision making to maximize the short-term return, and con- sequently, sacrificing the long-term well-being of the firm (Pandey, 2005), a potential contributor to the current economic crisis. Along the same lines, some have argued that financial measures “tell the story of past events” and therefore are inadequate
for information age companies moving at a fast pace (Paranjape, Rossiter, & Pantano, 2006, p. 6). Additionally, they may not be the best choice of organizational performance indicators for nonprofit organizations because profit is not the primary focus of those groups.
Ghalayini and Noble (1996) outlined several general limitations of traditional financial mea- sures. One limitation is the missing link between measuring performance in financial terms and the improvement efforts needed to improve that finan- cial performance. It is often difficult to diagnose the causes of performance problems at the individual or team level if all one has is organizational-level financial data. It is this problem of translation between the process of performance and its out- comes that makes financial performance measures less than ideal.
Additionally, traditional financial measures do not take into account the strategy or goals of the organization as an entity, unless those goals are purely financial to begin with. For example, an orga- nization may have high-level qualitative goals such as becoming well known in a particular target popu- lation or developing a reputation for timely and accu- rate service. The attainment of these goals cannot be measured using financial data. This is problematic, especially if the goals are central to the organization’s identity. This is where the measurement of individual- and team-level behaviors, cognitions, and attitudes can improve a PM system.
Balanced scorecard. One of the first PM tools created in response to the limitations of traditional financial measures is known as the BSC. The BSC is “a system of combining financial and nonfinancial measures of performance in one single scorecard” (Pandey, 2005, p. 51). BSC is not a strategy but is rather a management tool focusing on the financial and nonfinancial goals of an organization (Pandey, 2005). It was developed in response to the realiza- tion that quite often, nonfinancial goals are the drivers of business success, and therefore purely financial measures of performance are not sufficient alone. This integrated PM system complements financial measures with critical nonfinancial per- spectives (Paranjape et al., 2006). A recent estimate
Performance Measurement at Work
327
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 327
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
states that over 60% of organizations use a scorecard (Kaplan & Norton, 2005).
BSC suggests the inclusion of measures of four perspectives—(a) financial, (b) customer, (c) internal business processes, and (d) learning and growth— although it does not have to be restricted to these four dimensions of performance (Pandey, 2005). All four of these perspectives are necessary to get an accurate picture of organizational performance. One of the unique aspects of this measurement approach is that BSC implies a complex causality between the four dimensions, with learning and growth influenc- ing the internal business processes, and the processes have an impact on the financial outcomes, either directly or through the customer perspective (Dror, 2008).
Financial measures such as growth, profit mar- gin, return on investment, and so on, are backward- looking indicators but are necessary to see whether process improvements are translating into financial success. Customer-focused measures include cus- tomer satisfaction, customer retention, market share, and customer profitability. For example, poor cus- tomer satisfaction is a leading indicator of future performance decline (Pandey, 2005). The internal process perspective looks at the quality of the busi- ness processes within the organization, and the key objectives are process improvement and suppliers’ relations. Finally, the learning and growth perspec- tive focuses on innovation, creativity, and capability. Measures of learning and growth include employee satisfaction, employee retention, and employee pro- ductivity. In the BSC approach, each of the four per- spectives includes (a) objectives to attain (e.g., higher customer satisfaction, higher return on investment), (b) measures of those objectives (e.g., financial data, customer survey data), (c) target values of those measures, and (d) initiatives needed to achieve those targets. The success of BSC depends on the clear identification of nonfinancial and financial variables, their accurate measurement, and linking performance to rewards and penalties (Pandey, 2005).
As literature on BSC has grown, there have been two distinct and conflicting viewpoints emerging. One division of the literature advocates the successes of BSC. The main advantage of the BSC system is the ability to translate an organization’s vision and cen-
tral mission into tangible, achievable objectives and measures (Kanji, 2002; Pandey, 2005). BSC is unique from other measurement systems in that it contains both the outcome measures (i.e., financial measures) and the performance drivers of those outcomes (i.e., customer, innovation, internal process; Kanji, 2002). Kanji (2002) delineated several additional strengths of the BSC approach. While focusing on multiple dimensions of performance to get a more holistic view of organizational effectiveness, it simultane- ously limits the number of measures so as to avoid information overload. It is also a relatively flexible PM approach and can be individualized for each organization. Finally, it places strong emphasis on customers and the market, which are often ignored by traditional measures.
Alternatively, literature also criticizes the approach and calls attention to the lack of scientific evidence linking BSC to improved organizational performance (Paranjape et al., 2006). Other criticisms of BSC claim that it lacks a mechanism for building and maintaining relevance of measures once they are defined (Dror, 2008). In response, Kaplan and Norton (2004) introduced the concept of a strategy map, which provides such a mechanism for connect- ing strategic objectives with each other. Kanji (2002) described several other weaknesses of BSC. It was designed as a conceptual model and, therefore, is hard to convert into measurement. It is not a com- prehensive system approach and leaves out impor- tant stakeholders in addition to the role of individual employees and suppliers. The effectiveness of the strategy map has yet to be empirically investigated.
Six Sigma. Six Sigma is another performance management technique that has gained increasing popularity. Technically, the name Six Sigma can be translated into “3.4 defects per million opportunities” (Banuelas & Antony, 2003). This technique was developed at Motorola as a response to a large loss in productivity due to a lack of quality (Raisinghani, Ette, Pierce, Cannon, & Daripaly, 2005). Although the application of Six Sigma is spreading into increasingly different business sectors, historically, Six Sigma has been predominantly manufacturing- based (McAdam, Hazlett, & Henderson, 2005). This method of performance management “relies on
Wildman et al.
328
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 328
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
planned change, team-based collaboration, a focus on performance improvement, a systems perspec- tive, and reliance on the scientific method and sta- tistical methodologies” ( Jeffery, 2005, p. 21). Six Sigma is a very process-focused method and there- fore attempts to reduce errors and mistakes by experimenting with processes until more stable and robust processes are achieved.
When Six Sigma was first developed, it was pri- marily seen as a statistical method for tracking and reducing variability in processes (McAdam et al., 2005). However, as the popularity and use of Six Sigma have increased, the approach has grown into a much more complex operations improvement and problem-solving methodology. One of the most inter- esting aspects of the Six Sigma approach is the change in organizational culture that is often associated with its use. Many organizations using Six Sigma view it not as a statistical method for reducing error, but as a complete business philosophy requiring high levels of organizational and management support. This buy-in is crucial to the success of the strategy.
The Six Sigma method can be implemented using two different approaches—a continuous improve- ment design (reactive strategy) or a design–redesign approach known as Design for Six Sigma (DFSS; Banuelas & Antony, 2003). The continuous improve- ment design follows a sequence of five phases (a) define, (b) measure, (c) analyze, (d) improve, and (e) control (DMAIC). In the define phase, the problem to be addressed is identified and defined, and the critical stakeholders are identified. In the measure phase, the measurement capability of the organization is assured, current levels of performance are identified (most often using existing corporate performance measures), and goals for improvement are then set. In the analyze phase, the causes of prob- lems and the key variables that may be linked to defects are identified. In the improve phase, experi- mentation is used to quantify the impact of variables on the process, and the process is subsequently improved. Last, in the control phase, continuous monitoring and adjustments are used to maintain the improved process (Goh & Xie, 2004). This improve- ment approach assumes the initial process is essen- tially correct and thus only needs to be adjusted for optimal performance (Banuelas & Antony, 2003).
Raisinghani et al. (2005) described five general steps in the DMAIC approach. These include (a) determin- ing the appropriateness of the existing PM equipment through analysis, (b) looking for process deviations that require improvement, (c) using the “design of experiments” technique, (d) conducting a failure mode and effects analysis, and (e) taking one final measure of quality to ensure that Six Sigma levels have been achieved.
DFSS follows a slightly different sequence of phases: (a) define, (b) measure, (c) analyze, (d) design, and (e) verify (DMADV). It differs only in that rather than improving an existing process, and controlling that change, a brand new process is designed and then verified. DFSS is a more proactive approach in that it involves designing processes capable of reaching Six Sigma levels, effectively pre- venting problems before they arise rather than after (Banuelas & Antony, 2003). This approach does not assume that the initial process is correct; instead, it aims to replace the process with a better, more effec- tively designed process. Therefore, both efficiency and effectiveness of the process can be improved using the DMADV approach.
As with the implementation of any PM system, the context surrounding the Six Sigma approach is critical. McAdam et al. (2005) delineated a list of key steps in the Six Sigma processes on the basis of a review of the literature. The organization must gain senior manager support and involvement to create the desired deep level. A team of specialists must be trained to lead the process, requiring an initial investment of time devoted to training and develop- ing a set of experts in the process. Much of the lim- ited empirical evaluation of Six Sigma focuses on case studies (e.g., Antony, Kumar, & Madu, 2005).
Productivity Measurement and Enhancement System. ProMES is an intervention designed to improve performance by actively measuring it and using those measures to provide detailed performance-based feedback (Pritchard, 1990; Pritchard et al., 2008; Pritchard, Holling, Lammers, & Clark, 2002). A typical ProMES intervention involves a design team (composed of people respon- sible for doing the work) who identifies the main objectives of the unit and develops quantitative
Performance Measurement at Work
329
11819-10_Ch10_rev4.qxd 4/6/10 12:00 PM Page 329
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
measures of these objectives. These measures are designed to be valid measures of all the important aspects of work required from the unit. The system is then reviewed and approved by higher manage- ment. On the basis of the measures, a regularly occurring feedback report is provided to the work units to guide desired improvements.
The conceptual framework underlying ProMES is the NPI theory (Naylor, Pritchard, & Ilgen, 1980). NPI theory is a refinement of expectancy theory (Arvey, 1972; Campbell & Pritchard, 1976; Vroom, 1964) with a novel approach to motivation. Motivation within NPI is considered a process of allocating energy and time to various activities (Pritchard et al., 2002). The NPI motivational process involves acts (individual behavior), focusing on both direction (why certain acts are chosen over others) and amplitude (the amount of effort devoted to completion of the act). These acts lead to the gen- eration of products (individual outputs). These prod- ucts are then evaluated, which results in evaluations. These evaluations come from supervisors, peers, self, subordinates, family, and so on. Outcomes are tied to those evaluations and can either be intrinsic (i.e., increased satisfaction with one’s work) or extrinsic (i.e., pay raise, promotion). Outcomes are motivating because of their tie to needs satisfaction, which are very individual desires regarding every- thing from status to health and more. Acts, products, evaluations, outcomes, and needs satisfaction are combined into motivational force, which refers to an individual belief that changes in devoted time and energy (effort) toward different acts (tasks) will lead to changes in the amount of needs that are satisfied (Pritchard et al., 2002).
A key characteristic of NPI is contingencies. Between each of the motivational components are component links (contingencies), signifying the interconnectivity among components (Pritchard et al., 2002). For example, actions generate products, thus the creation of products is contingent on the amount of effort directed toward acts or behaviors that directly relate to generating the product. Each contingency is based on individually perceived rela- tionships. Again looking to the actions to products link, acts-to-products contingencies depict the relation- ship between the amount of effort or time devoted
toward an act and the expected amount of generated product.
ProMES consists of seven steps (Pritchard, 1990; Pritchard et al., 2002). For a more detailed review, see Pritchard (1990). The first step is creating a design team of seven to eight people who are ulti- mately responsible for creating the new system. On the basis of the notion of participation in decision making, the team is composed of employees who do the work under evaluation. The second step focuses on the identification of overall unit objectives, devel- oped through group discussion to consensus. The objectives should describe the purpose of the work group. All dimensions of the work must be reflected in the overall objectives. The third step involves iden- tifying indicators or quantifiable measures of effec- tiveness to determine how well the objectives are being met. The objectives and indicators are pre- sented to upper management for approval, and any disagreements are discussed to consensus. The fourth step moves toward defining contingencies. Contingencies are basically a function of the amount of the indicator as compared with its value to the organization. Indicators are charts, plotting effective- ness on the y-axis and the objective on the x-axis. The fifth step is to design the feedback system. Data are collected from each indicator, and indicator effectiveness scores are calculated. These individual scores are aggregated into an overall effectiveness score. A feedback report is then prepared with these data and is compared with historical data to deter- mine improvements and well as areas that require further attention. The sixth step is the actual feed- back session between the supervisor and employee, who discuss causes for improvements and any noted decreases, and devise a plan for continued improve- ment or changes. The seventh step involves monitor- ing the project overtime. If it is determined after several feedback sessions that some aspect of the measurement system needs to be changed, then the entire process should start again to come to agree- ment on how it should be fixed. If there are signifi- cant changes in the work or policy changes, then the system should also be reviewed.
This system is highly effective. Not only does it lead to large productivity improvements in many dif- ferent types of settings, these effects have also been
Wildman et al.
330
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 330
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
shown to last over time (Pritchard et al., 2008). The mean weighted effect size is 1.44, which is equivalent to the 93rd percentile of performance under baseline. Pritchard et al. (2008) also investigated a number of potential moderators. They found that the degree to which the study followed the ProMES methodology, the quality of the feedback given, whether changes were made to the new system, the degree of inter- dependence of the work group, and the degree of centralization of the organizations moderated the effectiveness of a ProMES intervention.
Summary Financial data, BSC, Six Sigma, and ProMES repre- sent some of the most widely studied and discussed approaches for measuring organizational perfor- mance (see Table 10.6). As the science of organiza- tional PM has grown, there has been a definite push toward more integrated, multidimensional PM sys- tems. Financial data alone are no longer considered an appropriate gauge of organizational effectiveness. There has also been a push toward PM systems that are more diagnostic of the causes underlying perfor- mance. We contend that taking a multilevel approach to PM is one of the most effective ways to diagnose organizational performance. This approach is described in detail in the following section.
MULTILEVEL PERFORMANCE MEASUREMENT: THE LINKAGES BETWEEN LEVELS
We contend that PM should serve three fundamental purposes: describing, evaluating, and diagnosing. The most basic purpose for PM is the accurate description of performance. Simply stated, PM is “the process of quantifying action” (Neely, Gregory, & Platts, 1995, p. 80). Often, what are referred to as PM systems in modern organizations are truly performance evalua- tion systems: that is, they are systems used to deter- mine whether a particular measure of performance is satisfactory. However, the first basic step in PM is the accurate description of performance without any judgment.
Once performance has been described, the next step is to evaluate that described action. As previously mentioned, PM can occur with or without a subse- quent evaluation of that performance. However, the link between these two systems is a critical one if PM data are to serve any practical purpose. If the inter- ested party cannot gauge whether the performance level captured by a PM system is “good” or “bad,” then they are unable to manage and improve that performance. However, the reverse is also true—if a well-designed evaluation system is applied to an inaccurate or unreliable measure, the product will
Performance Measurement at Work
331
TABLE 10.6
Organizational-Level Measurement Techniques
Technique Description Sources
Financial data
Balanced scorecard
Six Sigma
Productivity Measurement and Enhancement System
Outcomes measures focused on financial success such as return on investment or productivity
A performance measurement system that combines financial and nonfinancial measures into a single scorecard for organizational effectiveness
A performance measurement system that focuses on improv- ing process design and reducing process error and waste
A performance intervention designed to improve performance by carefully measuring it and providing detailed feedback
Ghalayini and Noble, (1996); Pandey (2005); Paranjape et al. (2006)
Ahn (2005); Kanji (2002); Kaplan and Norton (1996); Norreklit (2000); Pandey (2005); Paranjape et al. (2006); Schwartz (2005)
Antony (2006); Banuelas and Antony (2003); Goeke and Offodile (2005); Goh and Xie (2004); Raisinghani et al. (2005)
Pritchard et al. (1989, 2002, 2008)
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 331
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
again be meaningless. Therefore, it is imperative that organizations have both sound PM strategies as well as sound evaluation strategies for interpreting perfor- mance data.
The third step after describing and evaluating performance is to diagnose the causes of effective and ineffective performance. This step is perhaps the most important in the PM process because without an understanding of the causes underlying perfor- mance, the development of feedback and training becomes difficult and suboptimal, if not impossible. Identifying deficiencies in performance is not very useful unless one can also identify the underlying causal mechanisms to change to rectify those defi- ciencies. Diagnosing the underlying causes of effec- tive and ineffective performance is the only way PM can be used to manage and improve performance. This concept of linking specific performance out- comes to specific causal mechanisms to achieve effective PM has been called to our attention by other organizational researchers (e.g., Ittner & Larcker, 2003).
Therefore, we suggest that for any PM system to provide the most accurate description, evaluation, and diagnosis of performance, it should take an integrative, multilevel approach to measurement (see Figure 10.2). This position has been also been advocated quite strongly in the team PM literature (e.g., Salas et al., 2003). Individuals are the units that make up teams, and therefore team PM should include both team-level and individual-level indices. Following the same pattern, organizations are made up of teams, and teams are made up of individuals, so organizational PM should include all three levels of measurement if it is intended to capture the most beneficial and applicable information. The team PM literature focuses heavily on the distinctions between process and outcome, and how measure- ment of both is necessary if one is to accurately diagnose the underlying mechanisms influencing performance outcomes. Applying this concept to the organizational-, individual- and team-level processes are often the underlying causal mecha- nisms behind organizational outcomes. Therefore, for organizational PM to be diagnostic, it must cap- ture and connect performance at all active levels of analysis.
Although the paradigm shift in organizational PM from financial measures to more integrated multidimensional measurement systems illustrates a much needed change, organizational performance measures still tend to focus heavily on tangible out- comes at a very high level. Therefore, there is still some level of disconnect between effective perfor- mance at the organizational level and the processes that enable that performance at the individual and team levels. This disconnect makes it difficult to translate organizational-level measures into action- able goals for improvement. For example, a corpo- ration may measure profit per unit production to assess overall organizational-level performance and find they are making $5 in profit for each unit pro- duced, and decide they would ultimately like an increase of $2 in profit. Yet, how does this translate into terms of improvement? One can not tell employees to work “$2 harder.” How much extra effort is necessary for two more dollars of profit? What changes in the production process are neces- sary? What factors influence level of profit? This is why organizational-level PM is important, but not sufficient, for improved organizational performance.
To truly diagnose and improve organizational performance, multiple levels of performance must be captured to link overall financial outcomes to underlying individual- and team-level processes. Effective PM systems should focus on making con-
Wildman et al.
332
FIGURE 10.2. Linkages between levels of performance.
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 332
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
nections between the different levels of performance and should take advantage of the various techniques available from each perspective. This will allow the organization to gain an understanding of what indi- vidual-level behaviors and processes impact team processes and team outcomes, what team-level processes impact organizational outcomes, and so on. By doing this, changes in organizational-level financial performance can be stimulated by manag- ing performance at the employee and team levels. Ultimately, a multilevel understanding of perfor- mance will enable the organization to pinpoint the causes behind performance problems and effectively address them.
To illustrate the potential use of multilevel PM, a hypothetical example is described. Assume an orga- nization is planning to implement a new PM system throughout all levels of its personnel. Perhaps the company employs factory workers that work individ- ually to make products on an assembly line, as well as new product development teams that create new ideas for the products, among many other positions. For the sake of parsimony, we discuss only these two specific job positions. Now assume the organization chooses to use an automated measure of perfor- mance for the assembly line workers after carefully considering the purpose, content, timing, and setting desired. This automated system records the number of units produced, number of errors made, and several other measures of productivity. This choice of measurement allows for continuous on-the-job, real-time capturing of data. The organization then decides, after careful consideration, that the new product development will be individually rated using 360-degree feedback from themselves, their team- mates, and the team leader on several dimensions including innovation, cooperation, and communica- tion. The team will also be rated as a unit on the same dimensions. These subjective rating measures are better suited for measurement of highly creative and innovative jobs. Finally, organizational perfor- mance will be measured using overall customer satis- faction surveys as well as archival data reporting the annual profit of the company.
Now, because the organization collected PM data from all three levels of analysis, the organization is capable of linking outcomes to the underlying causal
processes. Specifically, perhaps analysis of the data suggests that higher ratings on communication and innovation at the team level within the new product development teams are related to higher levels of customer satisfaction, whereas the number of units produced by assembly line workers is quite unrelated to customer satisfaction scores. This information would allow the organization to pinpoint the process that needs remediation (i.e., communication and innovation within teams) if an increase in customer satisfaction is their top priority. Similarly, they may find that the number of errors made on the assembly line is highly related to overall profit, and therefore an increase in profit would more likely occur because of a reduction in these errors. The main point of this example is that multilevel PM leads to a richer under- standing of performance and the links between differ- ent processes and different outcomes, and therefore enables the organization to more effectively and effi- ciently manage their performance.
The connections between different levels of analysis occur not only in a bottom-up direction but also from the top down. Because individuals and teams operate within the particular context of the organization, factors from the organizational level may impact both team and individual functioning (Salas et al., 2003). For example, a lack of resources provided by the organization or a set of organiza- tional procedures may limit the ability of a team or individual to perform in a certain way. Thus it is crit- ical to measure contextual variables at the organiza- tional level to take those into consideration when trying to assess individual or team performance.
EMERGING ISSUES FOR FUTURE RESEARCH IN PERFORMANCE MEASUREMENT
An astounding amount of research has examined PM at the individual, team, and organizational lev- els. So much so that only a fraction of that research was touched on in this chapter. However, despite the abundance of research in the field of PM, there is always room for growth and further exploration. In particular, as the world changes and technologies advance, several areas have become paramount to the understanding of performance at any level: dis-
Performance Measurement at Work
333
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 333
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
tribution, diversity, dysfunction, and culture. We posit that these concepts represent the future of PM research.
Performance Measurement in Distributed Environments Globalization has impacted the way organizations work. Corporate offices are distributed throughout the world. Teams are distributed throughout various offices. Yet, another important phenomenon is also occurring—an increase in the number of workers who are working outside of the traditional office envi- ronment. A survey of senior-level executives world- wide indicated that two thirds of the global workforce was involved in distributed work arrange- ments (AT&T, 2004). In the United States, telecom- muting, the most common form of distributed work, has increased exponentially over the past sev- eral years. In 1997, 11.6 million full-time employees worked remotely at least 1 day per month (WorldatWork, 2006). This number more than quadrupled to 45 million by 2006, the majority of whom are women, ages 35–44 (Gajendran & Harrison, 2007; Roitz, Allenby, & Atkyns, 2002).
The increased prevalence creates a measurement issue—how does one measure the performance of employees when one never observes them actually “performing”? Many theories have put forth the notion that technology-mediated interaction leads to decrements (e.g., Media Richness Theory; Daft & Lengel, 1986). This is seemingly evidence that per- haps distributed performance should be measured differently than performance that is directly observed, to account for these issues. However, others speculate that good measurement is good measurement regard- less of the setting and that, for example, what con- stitutes good PM for teams that are collocated is the same for PM for teams that are distributed— specifically, that similar principles should apply to the development of a system for collocated or distrib- uted employees (Blackburn, Furst, & Rosen, 2003). Blackburn et al. (2003) suggested that the what, how, why and when of PM are closely related in either face-to-face or distributed environments. They out- lined several characteristics that should apply to either situation (e.g., performance measures should always be tied to desired organizational goals, the
purpose of the measure should be clear). However, there are some noted differences. In face-to-face envi- ronments, performance ratings are based not only on the outcomes but also on the perceived effort on the part of the employee. However, this is not possible when measuring virtual performance and thus should be taken into consideration. Additionally, it is impor- tant to consider the multiple ways at which to arrive at the same outcome (criterion dimensionality). These issues and many more need to be systemati- cally addressed by future research. Researchers need to understand the impact of distribution on perfor- mance ratings; the value added by distributed team or organizational members; and if performance should be measured differently, then how? Some ini- tial work has been done in this area (e.g., Hofmann, Klar, Mohr, Quick, & Siegle, 1994); however, more research is needed, given the changing nature of technology and the increased prevalence of this type of work. The implications of this research extend beyond teleworkers to any employees who are work- ing away from their home office (e.g., expatriates).
Performance Measurement in Diverse Settings Today’s workforce is more diverse than ever before. This poses a unique challenge for designing process- oriented PM systems that evaluate fairly across demo- graphic groups. Specifically, criterion dimensionality issues can play a big role in diverse settings. As men- tioned previously, criterion dimensionality is the idea that two people on the same job may be equally effective but engage in very different behaviors to reach that level of effectiveness (Borman, 1991). This can become a problem if there are gender or ethnicity differences in the strategies used to complete a task and the measurement system assumes only one cor- rect strategy. For example, men and women may dif- fer systematically on leadership styles because of their differences (i.e., women tend to exhibit more democ- ratic or participative styles whereas men tend to por- tray autocratic or directive styles; Eagly & Johnson, 1990). If a measure of performance assumes one style of leadership to be more effective than another with- out actually validating this effect, the measurement system could lead to unfair test bias and adverse impact issues. Therefore, it is critical that researchers
Wildman et al.
334
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 334
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
examine issues of demographic difference in process and how this influences subsequent effectiveness. If findings indeed suggest that multiple strategies can lead to the same level of overall effectiveness, this dimensionality should be reflected in the measure- ment system.
Performance Measurement and Dysfunction The study of performance has been pervasive in the industrial and organizational psychology liter- ature for decades. Yet only recently has the field begun to consider unintended consequences of measuring performance. As noted by Grizzle (2002), “we expect that measuring efficiency leads to greater efficiency and measuring outcomes leads to better outcomes, but we don’t always get the results we expect” (p. 363). She noted several examples such as distorting financial data as an unintended consequence of measuring quarterly earnings per share. Many companies, such as Enron, resorted to an incomplete disclosure of financial data in an attempt to manipulate this measure. In terms of process measures, she noted that measuring quantity can lead to a diminished focus on quality. The measure may be indicating positive findings (i.e., high levels of quantity); however, the actual product may be of very poor quality. Others have noted the unintended conse- quences of measuring performance on the employees. For example, performance monitoring in call centers has been shown to cause employees distress, which impacts their well-being and satisfaction (Holman, Chissick, & Totterdell, 2002).
More research is needed that focuses on these unintended consequences: What are they? Under what circumstances do they occur? How much and in what ways do they affect performance measures? Finally, what types of interventions alleviate the potentially harmful effects of these unintended con- sequences of PM?
Performance Measurement and Culture One emerging issue in PM research is the impact of cultural differences on PM. As the globalization of business increases, organizations are beginning to measure performance across national and cultural
boundaries like never before. Cultural background influences the qualities and traits people value, and therefore changes the way people rate performance dimensions. For example, Nonaka and Takeuchi (1995) argued that Westerners and Japanese view knowledge differently, which could influence how KM is viewed in multinational organizations. For example, Japanese view knowledge as being primar- ily tacit, whereas Western cultures tend to focus on explicit knowledge. Therefore, in a PM system that focuses on capturing levels of explicit knowledge, there may be cultural differences between Western and Eastern employees. These cultural differences in how people perceive performance dimensions could indicate that an employee being rated by a manager from one culture may receive a completely different rating from a manager of another culture. The chal- lenge in this situation is either selecting dimensions of performance that are comparable across cultures or building a common understanding of the perfor- mance dimensions already being measured. Otherwise, ratings will not be equivalent.
For example, Gillespie (2005) compared 360-degree feedback ratings across Great Britain, Hong Kong, Japan, and the United States and found that there were differences between countries in responses to the survey. Specifically, Gillespie found that respondents from different countries had differ- ent understandings of the constructs included on the survey and how these constructs related to each other. Therefore, they responded to the survey differ- ently, meaning that the results of 360-degree feed- back across cultures were not comparable. This calls into question the appropriateness of using 360- feedback in multicultural contexts because the rat- ings given by members of different cultures may not be compatible for aggregation. This is a very impor- tant issue for multinational corporations to consider, and further research is needed to explore this issue and possible ways to account for these differences.
CONCLUDING REMARKS
The purpose of this chapter was to highlight PM issues as they relate to various levels (individuals, teams, and organizations). Given the scope of the literature on PM developed over the past several
Performance Measurement at Work
335
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 335
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
decades, we have presented a brief overview of this topic. By noting issues relevant to individuals, teams, and organizations as well as presenting several mea- surement strategies common to each level, we hope to synthesize existing literature on this topic in a sys- tematic manner. Although the literature has taught researchers a great deal about PM, there is still much the field does not know. Future research suggestions are designed to spur further investigation in areas that are currently underdeveloped in terms of both scientific and practical understanding. It is clear that progress has been made with regard to improved measurement techniques, the consideration of multi- ple levels of performance, and a more concerted effort to consider conceptual criteria. Yet for a more broad understanding of PM, more research is still required.
References Ahmed, S. (1999). The emerging measure of effectiveness
for human resource management: An exploratory study with performance appraisal. Journal of Management Development, 18, 543–556.
Ahn, H. (2005). Insights from research: How to individu- alise your balanced scorecard. Measuring Business Excellence, 9, 5–12.
Aldakhilallah, K. A., & Parente, D. H. (2002). Redesigning a square peg: Total quality management performance appraisals. Total Quality Management, 13, 39–51.
Antony, J. (2006). Six Sigma for service processes. Business Process Management Journal, 12, 234–248.
Antony, J., Kumar, M., & Madu, C. (2005). Six Sigma in small- and medium-sized UK manufacturing enter- prises: Some empirical observations. International Journal of Quality & Reliability Management, 22, 860–874.
Arvey, R. D. (1972). Task performance as a function of perceived effort performance and performance reward contingencies. Organizational Behavior & Human Performance, 8, 423–433.
Arvey, R. D., & Murphy, K. R. (1998). Performance evalu- ations in work settings. Annual Review of Psychology, 49, 141–168.
AT&T. (2004). The remote working revolution. Retrieved August 21, 2008, from http://www.corp.att.com/ emea/docs/remote_working_2004.pdf
Austin, J. R. (2003). Transactive memory in organizational groups: The effects of content, consensus, specializa- tion, and accuracy, on group performance. Journal of Applied Psychology, 88, 866–878.
Austin, J. T., & Villanova, P. (1992). The criterion problem: 1917–1992. Journal of Applied Psychology, 77, 836–874.
Bailey, C. T. (1983). The measurement of job performance. Aldershot, England: Gower Press.
Banuelas, R., & Antony, J. (2003). Going from Six Sigma to Design for Six Sigma: An exploratory study using analytic hierarchy process. The TQM Magazine, 15, 334–344.
Bartram, D. (2005). The great eight competencies: A criterion-centric approach to validation. Journal of Applied Psychology, 90, 1185–1203.
Beehr, T. A., Ivanitskaya, L., Hansen, C. P., Erofeev, D., & Gudanowski, D. M. (2001). Evaluation of 360-degree feedback ratings: Relationships with each other and with performance and selection predictors. Journal of Organizational Behavior, 22, 775–788.
Benedict, M. E., & Levine, E. L. (1988). Delay and distor- tion: Tacit influences on performance appraisal effec- tiveness. Journal of Applied Psychology, 73, 507–514.
Bhasin, S. (2007). Lean and performance measurement. Journal of Manufacturing Technology Management, 19, 670–684.
Binning, J., & Barrett, G. Y (1989). Validity of personnel decisions: A review of the inferential and evidential bases. Journal of Applied Psychology, 74, 478–494.
Bititci, U. S., Turner, T., & Begemann, C. (2000). Dynamics of performance measurement systems. International Journal of Operations & Production Management, 20, 692–704.
Blackburn, R. S., Furst, S. A., & Rosen, B. (2003). Building a winning virtual team. In C. B. Gibson & S. G. Cohen (Eds.), Virtual teams that work: Creating the conditions for virtual team effectiveness (pp. 95–120). San Francisco: Jossey-Bass.
Borman, W. C. (1991). Job behavior, performance and effectiveness. In M. D. Dunnette & L. M. Hough (Eds.), Handbook of industrial and organizational psychology (Vol. 2., pp. 271–326). Palo Alto, CA: Consulting Psychologists Press.
Borman, W. C., Buck, D. E., Hanson, M. A., Motowidlo, S. J., Starks, S., & Drasgow, F. (2001). An examina- tion of the comparative reliability, validity, and accu- racy of performance ratings made using computerized adaptive rating scales. Journal of Applied Psychology, 86, 965–973.
Bowers, C. A., Jentsch, F., Salas, E., & Braun, C. C. (1998) Analyzing communication sequences for team training needs assessment. Human Factors, 40, 672–679.
Brannick, M. T., & Prince, C. (1997). An overview of team performance measurement. In M. T. Brannick, E. Salas, & C. Prince (Eds.), Team performance assessment and measurement: Theory, methods, and applications (pp. 3–16). Mahwah, NJ: Erlbaum.
Brannick, M. T., Salas, E., & Prince, C. (Eds.). (1997). Team performance assessment and measurement: Theory, methods, and applications. Mahwah, NJ: Erlbaum.
Wildman et al.
336
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 336
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Burke, C. S., Stagl, K. C., Salas, E., Pierce, L., & Kendall, D. L. (2006). Understanding team adaptation: A con- ceptual analysis and model. Journal of Applied Psychology, 91, 1189–1207.
Campbell, J. P., McCloy, R. A., Oppler, S. H., & Sager, C. E. (1993). A theory of performance. In N. Schmitt & W. C. Borman (Eds.), Personnel selection in organi- zations (pp. 35–70). San Francisco, CA: Jossey-Bass.
Campbell, J. P., & Pritchard, R. D. (1976). Motivation theory in industrial and organizational psychology. In M. D. Dunnette (Ed.), Handbook of industrial and organizational psychology (pp. 63–130). Chicago: Rand McNally.
Campion, M. A., Medsker, G. J., & Higgs, A. C. (1993). Relations between work group characteristics and effectiveness: Implications for designing effective work groups. Personnel Psychology, 46, 823–850.
Cannon-Bowers, J. A., & Salas E. (1997). A framework for developing team performance measures in train- ing. In M. T. Brannick, E. Salas, & C. Prince (Eds.), Team performance assessment and measurement: Theory, methods, and applications (pp. 45–62). Mahwah, NJ: Erlbaum.
Cascio, W. F., & Phillips, N. F. (1979). Performance test- ing: A rose among thorns? Personnel Psychology, 32, 751–766.
Chung, Y. C., Tien, S. W., Hsieh, C. H., & Tsai, C. H. (2008). A study of the business value of total quality management. Total Quality Management & Business Excellence, 19, 367–379
Czarnowsky, M. (2008, September). Executive develop- ment: An ASTD research report. Training and Development, 62, 44–45.
Daft, R. L., & Lengel, R. H. (1986). Organizational infor- mation requirements, media richness and structural design. Management Science, 32, 554–571.
Dalessio, A. T. (1998). Using multisource feedback for employee development and personnel decisions. In J. W. Smither (Ed.), Performance appraisal: State-of- the-art in practice (pp. 278–330). San Francisco: Jossey-Bass.
Dominick, P. G., Reilly, R. R., & McGourty, J. W. (1997). The effects of peer feedback on team member behav- ior. Group & Organization Management, 22, 508–520.
Dong, A. (2005). The latent semantic approach to study- ing design team communication. Design Studies, 26, 445–461.
Dror, S. (2008). The balanced scorecard versus quality award models as strategic frameworks. Total Quality Management & Business Excellence, 19, 583–593.
Dubois, C. L. Z., Sackett, P. R., Zedeck, S., & Fogli, L. (1993). Further exploration of typical and maximum performance criteria: Definitional issues, prediction,
and White–Black differences. Journal of Applied Psychology, 78, 205–211.
Dwyer, D. J., Fowlkes, J. E., Oser, R. L., Salas, E., & Lane, N. E. (1997). Team performance measurement in distributed environments: The TARGETs methodol- ogy. In M. T. Brannick, E. Salas, & C. Prince (Eds.), Team performance assessment and measurement: Theory, methods, and applications (pp. 137–153). Mahwah, NJ: Erlbaum.
Eagly, A. H., & Johnson, B. T. (1990). Gender and leader- ship style: A meta-analysis. Psychological Bulletin, 108, 233–256.
Earley, P. (1994). Self or group? Cultural effects of train- ing on self-efficacy and performance. Administrative Science Quarterly, 39, 89–117.
Eden, D. (1990). Pygmalion without interpersonal con- trast effects: Whole groups gain from raising man- ager expectations. Journal of Applied Psychology, 75, 394–398.
Evans, C. R., & Dion, K. L. (1991). Group cohesion and group performance: A meta-analysis. Small Group Research, 22, 175–186.
Fisicaro, S. A. (1988). A re-examination of the relation between halo error and accuracy. Journal of Applied Psychology, 73, 239–244.
Flanagan, J. C. (1956). The evaluation of methods in applied psychology and the problem of criteria. Occupational Psychology, 30, 1–9.
Fletcher, C. (2001). Performance appraisal and manage- ment: The developing research agenda. Journal of Occupational and Organizational Psychology, 74, 473–487.
Folan, P., & Browne, J. (2005). A review of performance measurement: Towards performance management. Computers in Industry, 56, 663–680.
Fowlkes, J. E., Lane, N. E., Salas, E., Franz, T., & Oser, R. (1994). Improving the Measurement of Team Performance: The TARGETs Methodology. Military Psychology, 6, 47–61.
Gajendran, R. S., & Harrison, D. A. (2007). The good, the bad, and the unknown about telecommuting: Meta- analysis of psychological mediators and individual consequences. Journal of Applied Psychology, 92, 1524–1541.
George, J. M., & Bettenhausen, K. (1990). Understanding prosocial behavior, sales performance, and turnover: A group-level analysis in a service context. Journal of Applied Psychology, 75, 698–709.
Georgopoulos, B. S., & Tannenbaum, A. S. (1971). A study of organisational effectiveness. In J. Ghorpade (Ed.), Assessment of organizational effectiveness: Issues, analysis, and readings (pp. 177–188). Pacific Palisades, CA: Goodyear.
Performance Measurement at Work
337
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 337
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Gershoni, H., & Rudy, N. (1981). An analysis of the total variation of work measurement techniques. International Journal of Production Research, 19, 303–316.
Gersick, C. G. (1988). Time and transition in work teams: Toward a new model of group development. Academy of Management Journal, 31, 9–41.
Ghalayini, A. M., & Noble, J. S. (1996). The changing basis of performance measurement. International Journal of Operations & Production Management, 16, 63–80.
Ghorpade, J. (2000). Managing five paradoxes of 360- degree feedback. Academy of Management Executive, 14, 140–150.
Gillespie, T. L. (2005). Internationalizing 360-degree feedback: Are subordinate ratings comparable? Journal of Business and Psychology, 19, 361–380.
Goeke, R. J., & Offodile, O. F. (2005). Forecasting man- agement philosophy life cycles: A comparative study of Six Sigma and TQM. The Quality Management Journal, 12, 34–46.
Goh, T. N., & Xie, M. (2004). Improving on the Six Sigma paradigm. The TQM Magazine, 16, 235–240.
Goodman, P. S., & Leyden, D. P. (1991). Familiarity and group productivity. Journal of Applied Psychology, 76, 578–586.
Griffin, M. A., Neal, A., & Parker, S. K. (2007). A new model of work role performance: Positive behavior in uncertain and interdependent contexts. Academy of Management Journal, 50, 327–347
Grizzle, G. A. (2002). Performance measurement and dys- function: The dark side of quantifying work. Public Performance and Management Review, 25, 363–369.
Gully, S., Incalcaterra, K., Joshi, A., & Beaubien, J. (2002). A meta-analysis of team efficacy, potency, and performance: Interdependence and level of analysis as moderators of observed relationships. Journal of Applied Psychology, 87, 819–832.
Guzzo, R. A., & Dickson, M. W. (1996). Teams in organi- zations: Recent research on performance and effec- tiveness. Annual Review of Psychology, 47, 307–338.
Guzzo, R. A., & Shea, G. P. (1992). Group performance and intergroup relations in organizations. In M. D. Dunnette & L. M. Hough (Eds.), Handbook of indus- trial and organizational psychology (Vol. 3, 2nd ed., pp. 269–313). Palo Alto, CA: Consulting Psychologists Press.
Hackman, J. R. (1987). The design of work teams. In J. Lorsch (Ed.), Handbook of organizational behavior (315–342). Englewood Cliffs, NJ: Prentice-Hall.
Hays, R. T., & Singer, M. J. (1989). Simulation fidelity in training system design: Bridging the gap between reality and training. New York: Springer-Verlag.
Herman, R. D., & Renz, D. O. (2008). Advancing non- profit organizational effectiveness research and the- ory: Nine theses. Nonprofit Management and Leadership, 18, 399–415.
Hinsz, V. B. (2004). Metacognition and mental models in groups: An illustration with metamemory of group recognition memory. In E. Salas & S. M. Fiore (Eds.), Team cognition: Understanding the factors that drive process and performance (pp. 33–58). Washington, DC: American Psychological Association.
Hofmann, R., Klar, R., Mohr, B., Quick, A., & Siegle, M. (1994). Distributed performance monitoring: Methods, tools, and applications. Parallel and Distributed Systems, 5, 585–598.
Holman, D., Chissick, C., & Totterdell, P. (2002). The effects of performance monitoring on emotional labor and well-being in call centers. Motivation and Emotion, 26, 57–81.
Ittner, C. D., & Larcker, D. F. (2003, November). Coming up short on nonfinancial performance mea- surement. Harvard Business Review, 88–95.
Jeffery, A. B. (2005). Integrating organization develop- ment and Six Sigma: Six Sigma as a process improve- ment intervention in action research. Organizational Development Journal, 23(4), 20–31.
Jensen, A. J., & Sage, A. P. (2000). A system management approach for improvement of organizational perfor- mance measurement systems. Information, Knowledge, Systems Management, 2, 33–61.
Kanji, G. K. (1990). Total quality management: The sec- ond industrial revolution. Total Quality Management, 1, 3–12.
Kanji, G. K. (2002). Performance measurement system. Total Quality Management, 13, 715–728.
Kaplan, R. S., & Norton, D. P. (1996, January–February). Using the balanced scorecard as a strategic manage- ment system. Harvard Business Review, 150–161.
Kaplan, R. S., & Norton, D. P. (2004). The strategy map: Guide to aligning intangible assets. Strategy & Leadership, 32, 10–17.
Kaplan, R. S., & Norton, D. P. (2005). The office of strat- egy management. Strategic Finance, 87, 8–60.
Kendall, D. L., & Salas, E. (2004). Measuring team perfor- mance: Review of current methods and consideration of future needs. In J. W. Ness, V. Tepe, & D. Ritzer (Eds.), The science and simulation of human perfor- mance (Vol. 5, pp. 307–326). New York: Elsevier.
Khan, M. K., & Wibisono, D. (2008). A hybrid knowl- edge-based performance measurement system. Business Process Management, 14, 129–146.
Kozlowski, S. W. J., & Klein, K. (2000). A multilevel approach to theory and research in organizations: Contextual, temporal, and emergent processes. In
Wildman et al.
338
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 338
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
K. J. Klein (Ed.), Multilevel theory, research, and methods in organizations: Foundations, extensions, and new directions (pp. 3–90). San Francisco: Jossey-Bass.
Kraiger, K., & Wenzel, L. H. (1997). Conceptual develop- ment and empirical evaluation of measures of shared mental models as indicators of team effectiveness. In M. T. Brannick, E. Salas, & C. Prince (Eds.), Team performance assessment and measurement: Theory, methods, and applications (pp. 63–84). Mahwah, NJ: Erlbaum.
Landauer, T. K., Foltz, P. W., & Darrel, L. (1998). An introduction to latent semantic analysis. Discourse Processes, 25, 259–284.
Latham, G. P., Wexley, K. N., & Pursell, E. D. (1975). Training managers to minimize rating errors in the observation of behavior. Journal of Applied Psychology, 60, 550–555.
Levine, J. M., & Moreland, R. L. (1990). Progress in small group research. Annual Review of Psychology, 41, 585–634.
Lewis, K. (2003). Measuring transactive memory systems in the field: Scale development and validation. Journal of Applied Psychology, 88, 587–604.
Lim, B. C., & Klein, K. J. (2006). Team mental models and team performance: A field study of the effects of team mental model similarity and accuracy. Journal of Organizational Behavior, 27, 403–418.
Lim, B., & Ployhart, R. E. (2004). Transformational leader- ship: Relations to the five-factor model and team per- formance. Journal of Applied Psychology, 89, 610–621.
Mangos, P. M., & Arnold, R. D. (2008). Enhancing mili- tary training through the application of maximum and typical performance measurement principles. Performance Improvement, 47, 29–35.
Marks, M. A., Mathieu, J. E., & Zaccaro, S. J. (2001). A temporally based framework and taxonomy of team processes. Academy of Management Review, 26, 356–376.
Mathieu, J. E., Heffner, T. S., Goodwin, G. F., Salas, E., & Cannon-Bowers, J. A. (2000). The influence of shared mental models on team processes and perfor- mance. Journal of Applied Psychology, 85, 273–283.
McAdam, R., Hazlett, S.-A., & Henderson, J. (2005). A critical review of Six Sigma: Exploring the dichotomies. International Journal of Organizational Analysis, 13, 151–174.
McGrath, J. E. (1991). Time, interaction, and perfor- mance (TIP): A theory of groups. Small Group Research, 22, 147–174.
Morgeson, F. P., Mumford, T. V., & Campion, M. A. (2005). Coming full circle: Using research and prac- tice to address 27 questions about 360-degree feed- back programs. Consulting Psychology Journal: Practice and Research, 57, 196–209.
Motowidlo, S. J. (2003). Job performance. In W. C. Borman, D. R. Ilgen, & R. J. Klimoski (Eds.), Handbook of psychology: Vol. 12. Industrial and orga- nizational psychology (pp. 39–53). New York: Wiley.
Motowidlo, S. J., & Borman, W. C. (1977). Behavioral anchoring scales for measuring morale in military units. Journal of Applied Psychology, 62, 177–183.
Muchinsky, P. M. (2009). Psychology applied to work: An introduction to industrial and organizational psychol- ogy. Summerfield, NC: Hypergraphic Press.
Muniz, E., Stout, R. J., & Salas, E. (1996, March). Communications as an indicator of team situation awareness. Paper presented at the 42nd Annual Meeting of the Southeastern Psychological Association, Norfolk, VA.
Nagle, B. (1953). Criterion development. Personnel Psychology, 6, 211–289.
Naylor, J. C., Pritchard, R. D., & Ilgen, D. R. (1980). A theory of behavior in organizations. New York: Academic Press.
Nenadal, J. (2008). Process performance measurement in manufacturing organizations. International Journal of Productivity and Performance Management, 57, 460–467.
Nickols, F. (2007). Performance appraisal: Weighed and found wanting in the balance. The Journal for Quality and Participation, 30, 13–16.
Nielsen, B. B. (2005). Strategic knowledge management research: Tracing the co-evolution of strategic man- agement and knowledge management perspectives. Competitiveness Review, 15, 1–13.
Nieva, V., Fleishman, E. A., & Reick, A. (1978). Team dimensions: Their identity, their measurement, and their relationships (Contract No. DAHC19-78-C-0001). Washington DC: Response Analysis Corp.
Nonaka, I., & Takeuchi, H. (1995). The knowledge-creating company. New York: Oxford University Press.
Norreklit, H. (2000). The balance on the balanced score- card: A critical analysis of some of its assumptions. Management Accounting Research, 11, 65–88.
Organ, D. W. (1997). Organizational citizenship behav- ior: It’s construct clean-up time. Human Performance, 10, 85–97.
Osborn, W. C., & Campbell, R. C. (1976, October). Developing skills qualifications tests. Paper presented at the annual convention of the Military Testing Association, Gulf State Park, AL.
Pandey, I. M. (2005). Balanced scorecard: Myth and real- ity. The Journal for Decision Makers, 30, 51–66.
Paranjape, B., Rossiter, M., & Pantano, V. (2006). Insights from the balanced scorecard performance measure- ment systems: Successes, failures, and future— A review. Measuring Business Excellence, 10, 4–14.
Performance Measurement at Work
339
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 339
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
Podsakoff, P. M., MacKenzie, S. B., Paine, J. B., & Bachrach, D. G. (2000). Organizational citizenship behaviors: A critical review of the theoretical and empirical literature and suggestions for future research. Journal of Management, 26, 513–563.
Prince, C., Ellis, E., Brannick, M. T., & Salas, E. (2007). Measurement of team situation awareness in low experience level aviators. The International Journal of Aviation Psychology, 17, 41–57.
Pritchard, R. D. (1990). Enhancing work motivation through productivity measurement and feedback. In U. Kleinbeck, H. Quast, H. Thierry, & H. Hacker (Eds.), Work motivation: Past, present, and future (pp. 119–132). Hillsdale, NJ: Erlbaum.
Pritchard, R. D., Harrell, M. M., DiazGranados, D., & Guzman, M. J. (2008). The productivity measure- ment and enhancement system: A meta-analysis. Journal of Applied Psychology, 93, 540–567.
Pritchard, R. D., Holling, H., Lammers, F., & Clark, B. D. (Eds.). (2002). Improving organizational performance with the productivity measurement and enhancement system: An international collaboration. Huntington, NY: NOVA Science Publishers.
Pritchard, R. D., Jones, S. D., Roth, P. L., Stuebing, K. K., & Ekeberg, S. E. (1989). The evaluation of an inte- grated approach to measuring organizational produc- tivity. Personnel Psychology, 42, 69–115.
Pulakos, E. D., Arad, S., Donovan, M. A., & Plamondon, K. E. (2000). Adaptability in the workplace: Development of a taxonomy of adaptive perfor- mance. Journal of Applied Psychology, 85, 612–624.
Pun, K. F., & White, A. S. (2005). A performance measure- ment paradigm for integrating strategy formulation: A review of systems and frameworks. International Journal of Management Reviews, 7, 49–71.
Raisinghani, M. S., Ette, H., Pierce, R., Cannon, G., & Daripaly, P. (2005). Six Sigma: Concepts, tools, and applications. Industrial Management & Data Systems, 105, 491–505.
Rico, R., Sánchez-Manzanares, M., Gil, F., & Gibson, C. (2008). Team implicit coordination processes: A team knowledge-based approach. The Academy of Management Review, 33, 163–184.
Roitz, J., Allenby, B., & Atkyns, R. (2002). 2001/2002 employee survey results—Telework, business benefit, and the decentralized enterprise. Retrieved June 4, 2003, from http://www.att.com/telework/ article_library/survey_results_2002.html
Rynes, S. L., Gerhart, B., & Parks, L. (2005). Personnel psychology: Performance evaluation and pay for performance. Annual Review of Psychology, 56, 571–600.
Sackett, P. R. (2002). The structure of counterproductive work behaviors: Dimensionality and relationships
with facets of job performance. International Journal of Selection and Assessment, 10, 5–11.
Sackett, P. R., Berry, C. M., Wiemann, S. A., & Laczo, R. M. (2006). Citizenship and counterproductive behavior: Clarifying relations between the two domains. Human Performance, 19, 441–464.
Sackett, P. R., Zedeck, S., & Fogli, L. (1988). Relations between measures of typical and maximum job per- formance. Journal of Applied Psychology, 73, 482–486.
Salas, E., Burke, C. S., & Fowlkes, J. E. (2006). Measuring team performance “in the wild”: Challenges and tips. In W. Bennett, C. Lance, & D. Woehr (Eds.), Performance measurement: Current perspectives and future challenges (pp. 245–272). Mahwah, NJ: Erlbaum.
Salas, E., Burke, C. S., Fowlkes, J. E., & Priest, H. A. (2003). On measuring teamwork skills. In J. C. Thomas & M. Hersen (Eds.), Comprehensive hand- book of psychological assessment (pp. 427–442). Indianapolis, IN: Wiley.
Salas, E., Priest, H. A., & Burke, C. S. (2005). Teamwork and team performance measurement. In J. R. Wilson & N. Corlett (Eds.), Evaluation of human work (3rd ed., pp. 793–808). New York: Taylor & Francis.
Salas, E., Sims, D. E., & Burke, C. S. (2005). Is there “big five” in teamwork? Small Group Research, 36, 555–599.
Salas, E., Stagl, K. C., Burke, C. S., & Goodwin, G. F. (2007). Fostering team effectiveness in organizations: Toward an integrative theoretical framework of team performance. In J. W. Shuart, W. Spaulding, & J. Poland (Eds.), Nebraska Symposium on Motivation: Vol. 51. Modeling complex systems: Motivation, cogni- tion and social processes (pp. 185–243) Lincoln: University of Nebraska Press.
Scott, W. D. (1917). A fourth method of checking results in vocational selection. Journal of Applied Psychology, 1, 61–67.
Schmitt, N. W., & Klimoski, R. J. (1991). Research methods in human resources management. Cincinnati, OH: Southwestern Publishing.
Schwartz, J. (2005). The balanced scorecard versus total quality management: Which is better for your orga- nization? Military Medicine, 170, 855–858.
Sink, D. S., & Tuttle, T. C. (1989). Planning and measure- ment in your organization of the future. Norcross, GA: Industrial Engineering and Management Press.
Small, C. T., & Sage, A. P. (2005/2006). Knowledge management and knowledge sharing: A review. Information Knowledge Systems Management, 5, 153–169.
Smith, P. C. (1976). Behavior, results, and organiza- tional effectiveness: The problem of criteria. In M. D. Dunnette (Ed.), Handbook of industrial and
Wildman et al.
340
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 340
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .
organizational psychology (pp. 745–775). Chicago: Rand McNally.
Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47, 149–155.
Smith-Jentsch, K. A., Cannon-Bowers, J. A., Tannenbaum, S. I., & Salas, E. (2008). Guided team self-correction: Impacts on team mental models, processes, and effec- tiveness. Small Group Research, 39, 303–327.
Smith-Jentsch, K. A., Mathieu, J. E., & Kraiger, K. (2005). Investigating linear and interactive effects of shared mental models on safety and efficiency in a field set- ting. Journal of Applied Psychology, 90, 523–535.
Sundstrom, E., De Meuse, K. P., & Futrell, D. (1990). Work teams: Applications and effectiveness. American Psychologist, 45, 120–133.
Tannenbaum, S. I. (2006). Applied performance measure- ment: Practical issues and challenges. In W. Bennett, C. E. Lance, & D. J. Woehr (Eds.), Performance mea- surement: Current perspectives and future challenges (pp. 297–318). Mahwah, NJ: Erlbaum.
Toegel, G., & Conger, J. A. (2003). 360-degree assessment: Time for reinvention. Academy of Management Learning and Education, 2, 297–311.
Toops, H. A. (1944). The criterion. Educational and Psychological Measurement, 4, 271–297.
Tsui, A. S., & Barry, B. (1986). Interpersonal affect and rating errors. Academy of Management Journal, 29, 586–599.
Viswesvaran, C., Schmidt, F. L., & Ones, D. S. (2005). Is there a general factor in ratings of job performance?
A meta-analytic framework for disentangling sub- stantive and error influences. Journal of Applied Psychology, 90, 108–113.
Vroom, V. H. (1964). Work and motivation. New York: Wiley.
Waldman, D. A., Atwater, L. E., & Antonioni, D. (1998). Has 360-degree feedback gone amok? Academy of Management Executive, 12, 86–94.
Watson, W. E., Kumar, K., & Michaelsen, L. K. (1993). Cultural diversity’s impact on interaction process and performance: Comparing homogeneous and diverse task groups. Academy of Management Journal, 36, 590–602.
Weingart, L., & Weldon, E. (1991). Processes that medi- ate the relationship between a group goal and group member performance. Human Performance, 4(1), 33–54.
Wiese, E. E., Merket, D., Stacy, W., Nelson-Walwanis, M., Freeman, J., & Aten, T. (2006, December). Enhancing distributed debriefs with performance measurement objects. Proceedings of the 2006 Interservice/Industry Training, Simulation, and Education Conference (pp. 1–11). Orlando, FL.
WorldatWork. (2006). Telework trendlines for 2006. Retrieved August 21, 2008, from http://www.work- ingfromanywhere.org/news/Trendlines_2006.pdf
Zalesny, M. D., Salas, E., & Prince, C. (1995). Conceptual and measurement issues in coordination: Implications for team behavior and performance. In G. R. Ferris (Ed.), Research in personnel and human resources man- agement (Vol. 13, pp. 81–115). Greenwich, CT: JAI Press.
Performance Measurement at Work
341
11819-10_Ch10_rev3.qxd 3/30/10 11:52 AM Page 341
Co py
ri gh
t Am
er ic
an P sy
ch ol og ic al A ss oc ia ti on . No t fo r fu
rt he
r di
st ri
bu ti
on .