For TUTOR GRACEE only
This is Chapter 3 from the following book:
Johnson, C. M., Mawhinney, T. C. & Redmon, W. K. (2001). Handbook of Organizational Performance: Behavior Analysis and Management (1st ed.). New York: Haworth Press.
Chapter 3
Developing Performance Appraisals: Criteria for What and How Performance Is Measured
Despite the lip service paid to the fair definition and accurate appraisal of target behaviors, these two steps are often glossed over when setting up motivational programs. Yet, both the industrial/organization (I/O) and applied behavior analysis (ABA) communities emphasize the importance of what and how performance is measured. In I/O psychology, performance appraisal is one of five major areas (Campbell, 1990; Cardy and Dobbins, 1994; Dunnette, 1963; James, 1973; Latham, Skarlicki, Irvine, and Siegel, 1993; Latham and Latham, 2000; Murphy and Cleveland, 1995; Shaw, Schneier, Beatty, and Baird, 1995; Smither, 1998). I/O psychologists constantly grapple with their failure to develop decent indices of performance, what they refer to as the “criterion problem” (Blum and Naylor, 1968). Steers, Porter, and Bigley (1996) admit that even in “the bestdesigned reward systems …, the evaluation or appraisal of performance [is] perhaps the most basic concern” (p. 500). ABA researchers agree that the way in which targets are defined and measured profoundly influence the ultimate goal of enhancing desired performance (Bellack and Hersen, 1988; Ciminero, Calhoun, and Adams, 1986; Goldfried and Kent, 1972; Johnston and Pennypacker, 1993). Weist, Ollendick, and Finney (1991) frown on such dubious practices as basing definitions on expediency and choosing erroneous or irrelevant indices. Foster and Cone (1986) question whether the expectations of the raters will bias the results.
These criticisms continue, however. This faultfinding is not restricted to academics. Employees complain as well. Some have successfully sought compensation in court, affirming their charges that raters were influenced by irrelevant characteristics such as their race, gender, or age rather than employees’ sustained performance on the job (Ashe and McRae, 1985; Cascio and Bernardin, 1981; Feild and Holley, 1982; Werner and Bolino, 1997). Two major problems have been identified with traditional performance appraisal systems: the vague way in which performance is defined and the bias of the raters. An analysis of law court cases since 1973 indicated that “in six of ten cases decided against the organization, the plaintiffs were able to show that subjective standards had been applied unevenly to minority and majority employees” (Barrett and Kernan, 1987, p. 501).
This chapter begins by illustrating some prevalent problems. To do this, we draw upon articles in the research and professional literature, as well as court cases. Although the examples presented here take place in work settings, they can be found in virtually any applied setting. To counteract these criticisms, five criteria for appraisal systems are proposed, referred to as the SURF and C (Komaki, 1998a). Examples are from diverse settings with different work groups doing a variety of jobs.
PROBLEMS WITH WHAT TO APPRAISE
One of the major problems is the substance, or content, of what is measured, whether it be employees’ personality traits (e.g., cooperation, dependability), the outcomes of the work (e.g., the number of injuries per million hours worked, the accuracy of preventive maintenance checklists), or the work behaviors themselves (e.g., providing customer service, being safe).
When Judgments of Performance Are Based on Aspects of the Job That Do Not Adequately Represent the Job
What is measured does not always reflect what is critical to the job or include all of the many essential aspects of the job. For example, a sales agent at a California carrier was dismayed to find that an important aspect of her job—providing quality service—was apparently disregarded in appraising her performance (Bravo, 1991). One day the agent received a call from a distraught customer whose relative had died. The agent not only made a complicated set of arrangements, booking flights from one remote area to another, but also managed to keep the customer calm. Later, she was appalled when her supervisor reprimanded her for spending too much time on the call. Evidently, she was judged on the amount of time she had spent with one customer, and she had exceeded the prescribed amount. While efficiency was, no doubt, an integral aspect of the job, another vital aspect—the quality of the service—was downplayed. Thus, she questioned whether her supervisor's appraisal took into consideration both of these important aspects of her job.
This same concern about substance surfaces in the complaints aired about the Internal Revenue Service (Rosenbaum, 1998, p. A15). A major criticism was its heavy emphasis on “production quotas” to the detriment of the taxpayer (“After Critical Inquiry,” 1992, p. A18). A consultant notes that “in a numbers-driven organization, … if they say, ‘You've got to collect X amount of dollars…, well, all of a sudden the taxpayer becomes subordinate to that goal'” (p. A18). If IRS agents are judged primarily in terms of the amount of revenue they bring in and the high “producers” are promoted, then other aspects of the job, such as following prescribed regulations and properly justifying tax assessments, will fall by the wayside. Although perhaps inadvertently, those who “take advantage of taxpayers” (p. A18) may be reinforced.
Questions have also been raised about what is being monitored electronically (Griffith, 1993; Schrage, 1992). The aspects of the job that are typically assessed are those things that are amenable to being counted; for example, the number of keystrokes, the length of phone calls, and “down” time. Because of the ease of obtaining information on these factors, more and more organizations (department stores, airlines, insurance agencies, and telephone companies) are using them to appraise their employees. Unfortunately, other critical aspects, such as the accuracy of the information entered, may be neglected. Assessments of insurance clerks, for example, are typically made on the quantity of claim forms processed, regardless of accuracy. In each of these examples, the issue is the substance of the appraisal and whether it adequately reflects what employees do; whether the critical aspects are included and the noncritical excluded.
When Judgments of Performance Are Based on Aspects of the Job That Persons Cannot Sufficiently Influence
Another content-related issue concerns the control that employees can realistically exert over the indices of their performance.
A woman's mother, who had suffered a serious heart attack, lay dying in a hospital in need of an operation that no doctor would perform (Byer, 1992). The cardiologist believed that the woman's mother would die within a few days if the operation was not performed, but the surgeons whom the family contacted either refused to operate or wanted to wait a week. Why did the surgeons refuse to operate? In this case, the surgery was too risky and the surgeons contacted had too many black marks on their names. One doctor who spoke frankly with the woman and her family was quoted as saying, “Don't you think that the chief of surgery would love to do this operation? … He's a great surgeon but he's taken too many highrisk cases lately …” (Byer, 1992, p. A23). At that time (it has since changed), the New York State Department of Health judged surgeons solely by what is known as a “surgical scorecard”: a record of the number of patient deaths (Altman, 1992). What the scorecards failed to reflect, however, was the complexity of the operation and the patient's health condition before surgery. Patients may die not because of the physician's skill in performing the surgery, but because of these “uncontrollable factors.” Hence, to judge surgeons’ competence solely on the basis of fatalities would not fairly reflect their performance in the operating room.
A similar control issue is raised in relation to an index sometimes used to reflect employee performance—stock prices. Although some economists would argue that the best measure of a company's performance (and by inference, the employees in the company) is the stock price (e.g., Baker, Gibbons, and Murphy, 1994), stock prices reflect many factors over which employees can exert relatively little control. A memorable illustration of this can be seen in a graph depicting stock prices for United Airlines from January through September 1990 (Berg, 1990). The carrier's stock price tumbled dramatically in July from $160 to less than $100 per share. If one assumes that stock prices reflect employee performance, then by inference one would have to conclude that their performance had plummeted at the same time. Astute observers of the airlines, however, attributed the precipitous drop in stock prices to a variety of uncontrollable factors: takeover bids confounded by the departure of three chief executives in four years, rising fuel prices, and a need to cut costs. Furthermore, countervailing evidence suggested that the employees’ performance was actually on the upswing during the same period—the airline had improved its service, the number of times that the ground crews had mishandled baggage was down, and on-time arrivals and departures were up.
Relying on stock prices or surgical scorecards as the sole evaluation makes it more likely that the “uncontrollables” get overestimated and employees’ performance is underestimated. Latham and colleagues (1993) refer to this problem as the omission of “the organizational context in which the appraisal process is embedded” (p. 122).
In summary, problems can occur with the substance of what is measured. When the appraisals are primarily based on indices over which employees can exert relatively little influence and when measures are incomplete, neglecting all of the critical aspects of the job, misleading estimates of employees’ performance may be produced.
THE “HOW” ISSUES IN PERFORMANCE APPRAISAL
Besides these two content-oriented problems, another challenge concerns the way in which performance is measured: the relatively low (or nonexistent) levels of interrater reliability, the indirect methods used in evaluating performance, and the rarity of appraisals.
When Judgments of Performance Are Based on Information That Has Been Collected Semiannually or Annually
A common practice in many organizations is to evaluate employees once, or at most twice, a year. About one-third of a sample of about 700 survey respondents, who were exempt employees in a manufacturing organization, reported on an anonymous questionnaire that their performance was not even evaluated at least once a year (Landy, Barnes, and Murphy, 1978). (Note: Significantly enough, the frequency of appraisals was found to be one of the factors related to employee perceptions of the fairness and accuracy of their evaluations.)
In discussing the difficulties of appraising employees, one executive admitted: “Many of us have trouble rating for the entire year. If one of my people has a stellar three months prior to the review … you don't want to do anything that impedes that person's momentum and progress” (Longenecker, Sims, and Gioia, 1987, p. 188). The implication here is that employee ratings occur once a year and reflect performance over less than the entire year.
Because of a reliance on retrospective and rarely made accounts, questions are raised about the accuracy of the judgments in relation to performance.
When Performance Judgments Are Based on Information That Is Collected Indirectly
Many appraisals do not go directly to the source, but instead rely on secondhand information. For instance, a store manager evaluates supervisors by the letters received from distraught customers; a city manager judges the performance of the superintendent of the wastewater treatment plant by relying on the opinion of the assistant superintendent; a department chair judges the quality of professors’ teaching by relying on student complaints. Among the problems with these secondary sources is that the information may be generated from biased sources, and hence sifted through a nonneutral filter. The assistant superintendent may aspire to the job of the superintendent. Customers and students who complain may not be representative.
Even when the sources are neutral, the fidelity of the information may suffer when it is not directly sampled. The case of a bond trader, Jett, at a Wall Street firm provides an involved but nonetheless realistic example of what happens when one relies on secondhand information (Nasar, 1994). To judge Jett's performance, his supervisor depended on the firm's accounting system. However, discrepancies turned up between what was recorded on the system as trades (and profit) and the trades that were actually consummated. This particular discrepancy turned out to be major, resulting in a $350 million inflation of Kidder's profits for the year. In discussing how such a discrepancy could have occurred, Jett's supervisor vehemently disagreed, pointing out that it was “unrealistic and irresponsible to suggest one person could personally supervise a dozen departments without relying on a firm's internal auditing system” (p. D 14). An analyst at another company pointed out that, even though Jett's supervisor didn't have responsibility for the auditing system, he did have “responsibility for independently verifying what's going on” (p. D 14). Another Wall Street manager agreed that it was critical for a general manger to be aware of what employees are doing with respect to their trading practices. “He should never have had to rely entirely on secondary sources” (p. D 14). Only when Jett's supervisor did an “initial looksee” did he uncover the loophole of trades that counted as trades even though they were never consummated.
When Judgments of Performance Are Based on Information That Is Not Confirmed by More Than One Rater
Rarely does more than one supervisor rate the performance of employees; a lone supervisor usually conducts the appraisal. Not surprisingly, few checks are made to determine if one supervisor would rate an employee the same way as another. That is often a problem. In a comprehensive and thoughtful analysis of observational measures, Foster and Cone (1986) express skepticism about having a single rater do an evaluation alone. They point to rater drift in which a rater's definitions shift over time, and rater bias in which a rater's beliefs and perceptions about who and what they are observing can taint their findings. For the latter, it has been shown that both conscious and inadvertent biases or distortions can creep in and reduce the accuracy of the information obtained. In a study comparing direct observations and self-ratings, biases were found even when college students estimated something as neutral as how often they interacted with given individuals in their dorm. The tendency was to discount interactions with persons with “… whom they were relatively low-frequency cointeractants, in favor of those with whom they were reciprocally high” (Hammer, 1985, p. 201). Discrepancies have even been found between the data recorded by managers themselves and their estimations of what they had done (Lewis and Dahl, 1976). The discrepancies were not necessarily random. Managers in another study were found to substantially overestimate some activities (e.g., the time spent on production) and underestimate others (e.g., personnel) (Burns, 1954), leading to biased estimates of performance.
Rater bias continues to be raised by employees as well: women at the Voice of America (Kilborn, 2000), Asian-American scientists at the Los Alamos weapon laboratory (Glanz, 2000) and African-American agents at the Federal Bureau of Investigation (Johnston, 1998). One court case involved Texaco's employment practices particulary regarding its performance appraisal system (Roberts v. Texaco, 1994). The class action suit filed on behalf of 1,400 minority employees asserted that Texaco systematically discriminates against minority employees in promotions. Besides generating front-page headlines with taped conversations of Texaco executives disparaging African-American employees and identifying them as being “glued to the bottom of the bag” (Eichenwald, 1996, p. D4), the lawsuit characterized the company's performance appraisal system as “entirely arbitrary, and … used as a pretext for denying qualified minority employees promotions to which they are otherwise entitled” (Roberts v. Texaco, 1994, p. 9). As evidence, plaintiff Bari-Ellen Roberts pointed out that in Texaco's own Diversity Assessment Survey throughout the company, all groups—women and men, minority and majority—saw “criteria other than performance as being barriers to promotions, with most employees feeling that promotions are based on who you know, rather than performance” (p. 11).
Complicating matters is the fact that the traditional performance appraisal systems used in most any Fortune 1000 firm are faulty. Even the most well meaning raters would have difficulty rendering fair and accurate judgments, given the forms they are asked to fill out with few if any definitions of performance, the sparseness of their training, and, perhaps most important, the utter lack of follow-through to ensure that employees are rated irrespective of their gender, age, race, or sexual orientation on only what matters—their sustained record of performance on the job. Supervisors are typically asked to rate employees on a variety of generic dimensions such as quantity and quality of work, job or trade knowledge, ability to learn, cooperation, dependability, industry, and attendance. Generic dimensions are often used so that supervisors in a wide variety of departments can utilize preprinted forms to evaluate all of their employees. Problems have occurred with these generic dimensions, however. In another court case, for example, referred to as James v. Stockham Valves and Fittings Co. (1977), an employee named James brought a case against his employer, Stockham Valves and Fittings Co. James’ contention was that the seven dimensions used to appraise him were so general as to be open to a wide variety of individual definitions, interpretations, and possible biases. For example, in assessing dependability, Supervisor A might think of it as the person being on the job every day, regardless of circumstances such as illness or bad weather, whereas Supervisor B might see it as being able to count on the person to do what he said he would do in a timely manner. This could lead to a situation in which James would be judged highly by Supervisor A and poorly by Supervisor B. The company could have countered by providing data showing that interrater reliability checks had been made between supervisors collecting data on the same employees and that agreements between the different raters were uniformly high. If such evidence had been presented, a case could have been made that the definitions were clear, the supervisors were trained, and that factors unrelated to performance did not bias the ratings employees were given.
Besides the lack of interrater reliability checks, the infrequency of the appraisals and the indirectness of the sampling were raised as methodological issues. As can be seen, both the method in which evaluations are made and the content of the appraisals have been problematic.
WHAT CAN BE DONE TO IMPROVE CONTENT AND METHOD?
Given the gamut of complaints that exist in relation to what and how to assess performance, we have surveyed the literature in I/O psychology that focuses on appraising employee performance. In some cases, we found that problems were satisfactorily resolved. At an air freight company (Doyle and Shapiro, 1980), it was found that numerous errors and delays occurred in tracking shipment counts: “it took from three to five months for feedback on sales to reach” sales representatives, and “it was often impossible to determine whether they or the salespeople on the other end should get credit for the sale” (p. 139). This lack of consistency in pinpointing who was responsible for closing the sale was identified as causing problems in motivating the sales personnel. In this case, the appraisal system was redesigned to ensure accurate and timely sales information. After this change, the three test offices “moved to among the top producers in their respective regions, increasing shipments an average of 34.7 percent” (p. 140).
Revisions to a measure of preventive maintenance (PM) were also successfully made in the U.S. Marines (Komaki, 1998a). During Year 1, a measure was developed including the time Marines utilized, and a program was introduced that included the exchange of time-off with pay. Data were collected for a year. Surprisingly, no changes were forthcoming. The failure forced Komaki (1998a) to see how she had fallen prey to expediency. Among the reasons for the failure was her choice of the target. Time utilization could be easily and reliably defined in four words—”manipulating tools or equipment.” It entailed no specialized knowledge. Interrater reliability was obtained quickly. Unfortunately, it did not meet the criterion of being under workers’ control. The marines complained that they could not utilize their time well if they could not control when they were sent to do the work. In Year 2, she redesigned the measure of PM so that it was primarily under the marines’ control, thus meeting all of the SURF & C criteria. Even though the program was less potent, utilizing only feedback, significant improvements were found. Supervisory personnel rated the intervention as “very” to “extremely” effective. All parties agreed that they had a better idea of the maintenance effort. One unit supervisor remarked that the targets were “probably as objective as any evaluation could be” (p. 271).
Traditional Performance Appraisal Literature Is Primarily Descriptive
Disappointingly, the previous examples are rarities. Much of the literature does not go beyond a description of the different types of appraisals. One listing divides them into (1) “objective” indices such as production and sales data, as well as personnel information about absences, tardiness, promotions, tenure, and accidents; and (2) judgmental indices in which supervisors give their opinion of employee performance (DeVries et al., 1980). The latter includes: (1) global essays (a narrative about the worker's performance); (2) graphic rating scales (“rate this person on dependability where 1 = excellent and 5 = poor”); (3) ranking (“list employees in order in terms of dependability”), and four techniques combining various ways of producing the target behaviors and presenting the information to the rater; (4) the critical incident technique (“describe a situation in which employees, were either very dependable or very undependable”) (Flanagan, 1954); (5) mixed standard scales (Blanz and Ghiselli, 1972); (6) behaviorally anchored rating scales (Smith and Kendall, 1963); and (7) behavioral observation scales (Latham and Wexley, 1981). (For more detailed information, please refer to the citations.)
Occasionally, comparisons are made among the different types of appraisals. DeVries and colleagues (1980) evaluated the above types of appraisals according to various criteria: content validity (or job relatedness), criterion/construct validity, reliability, discriminability for selecting employees, usefulness for administrative decisions such as promotions, and feedback for motivational purposes—the focus of the present chapter. When feedback is the aim, behaviorally anchored rating scales (BARS) and the objective approach were judged to be the “strongest” (p. 50). Although these comparisons go beyond describing the types of appraisals, few studies have been conducted to determine whether one type of appraisal is superior to another (c.f., Gomez-Mejia, 1988).
Prescriptions Are Limited
The earliest research to focus on appraisals concerned the format, or layout, of graphic rating scales (Landy and Farr, 1980). Formats were varied with changes in the types of anchors, the position of the high end of the scale, the spatial orientation of the scale, the segmentation of the scale line, the numbering of scale levels, and the number of response categories. For example, when the research consistently showed that an excessive number of response categories had detrimental effects on reliability, no more than nine scale anchors were recommended.
Unfortunately, this format-oriented research is limited to rating scales. Moreover, many of the results were inconclusive. Though raters had preferences for some formats (e.g., physical arrangements of high and low anchors and numbering systems), Landy and Farr (1980) conclude that even when they were granted their preferences, the new formats had relatively little effect on the quality of the ratings.
Recently, attention has shifted to the raters themselves and the cognitive processes involved when appraising others. A variety of models have been proposed (e.g., DeNisi, 1996; Feldman, 1981; Ilgen and Feldman, 1983; Landy and Farr, 1980). All view the rater as an active collector of information, and all are based on a process whereby raters acquire, store, recall, and combine information to make judgments. The major difference among the models is that each focuses on a different part of the process. The DeNisi (1996) model, for instance, concentrates on the information acquisition activities of the rater and also emphasizes the importance of information storage in memory. A basic assumption in this model is that the pattern in which information is acquired has a significant impact on later information processing. The authors postulate that the observation of behavior is affected by the rater’s preconceived notions or impressions of the person he or she is rating, the purpose for which the appraisal is conducted, time pressures on the rater, and the nature of the rating instrument (including the different dimensions). Raters are not thought to use all information available to them, but rather to form global impressions of employees that obscure behavioral detail.
The primary recommendation that has come from this line of research is that raters should be trained, usually by learning the standards by which they are to judge performance or in the cognitive processes they go through in observing, encoding, storing, recalling, and integrating performance information (Borman, 1991). To make more accurate ratings, De-Nisi, Cafferty, and Meglino (1984) suggest the use of diary keeping and familiarity with rating scales. These techniques are thought to reduce distortion by providing a way in which raters can structure and organize information to be stored in memory (to facilitate processing), and also to rely less upon memories of performance. If this is the case, then this recommendation would have direct implications for improving the interrater reliability of the raters. Even if raters could be trained until they are reliable, the training deals with only one methodological aspect of the process and does not address any of the content issues.
Few Specifics Related to Content
Although few investigators have touched upon the content and method problems identified earlier in this chapter, suggestions about the content of appraisals have not been entirely absent. Many of these discussions take place over the selection or hiring of employees. When hiring a worker, various tests (e.g., of cognitive ability or motor skills) are sometimes given to prospective employees to predict the person’s performance on the job, referred to as the “criterion.” Many problems have occurred when defining what constitutes an adequate assessment of a person’s performance. To avoid this criterion problem (Blum and Naylor, 1968), many suggestions have been made. The traditional advice is to make sure that the criterion is valid and related to the job, and that it represents important aspects of the job, is free from contamination (or the influence of factors other than important aspects of the job), and is free from deficiency (or the incompleteness and omission of important aspects) (Smith, 1976).
Three problems exist, however, with this general guidance. One, many of the definitions are not clear. How relevant does a criterion need to be before it is “relevant”? Exactly what is considered “deficient”? Two, few, if any, constructive suggestions are made about what to measure. For example, as long as measurements of any of three basic categories of content (individual personality traits, behaviors, and outcomes/results) identified by DeVries and colleagues (1980) are part of a job, they are considered appropriate for use. If we take the case of IRS agents, several questions can be posed. Should agents be judged on outcomes/results—the amount of revenue they amass? Should other factors be included? If so, which ones? Furthermore, almost no mention is made of how the information is collected other than to highlight the importance of reliability.
In short, few recommendations have indicated how to constructively improve what is appraised in order to ensure that the target behaviors include critical aspects of the job and are responsive to workers’ efforts. At the same time, relatively little research deals with how employees are evaluated. These issues are not addressed by the comparisons among the different types of performance. The models generated about the cognitive processes of raters and the research conducted on the graphic rating scale have implications for only one methodological issue, that of the reliability of the raters.
NEW CRITERIA FOR CRITERIA
SURF&C
In response to the concerns raised, five criteria (or standards) are recommended (Komaki, 1998a). These criteria, referred to by the mnemonic SURF & C, include:
|
S: |
the target being directly sampled rather than relying on filtered or secondary sources; |
|
U: |
the target being primarily under the control of workers, responsive to their efforts, and minimally affected by extraneous factors; |
|
R: |
independent observers consistently agreeing on their recordings and obtaining interrater reliability scores of 80 percent to (ideally) 90 percent or better during the formal data collection period; |
|
F: |
the target being assessed frequently and on a regular reoccurring basis—at least twenty and ideally thirty times—during the period of the intervention; and |
|
C: |
evidence indicating that the target is critical for the successful completion of the task. Data must be provided showing a significant relationship between the target and the intended out-come. |
When two of the criteria—being under control (U) and being critical (C)—are met, the target is considered to be appropriate in content for appraising performance. To ensure the method used in obtaining the information is appropriate, the appraisal should meet the criteria of being directly sampled (S), passing the test of interrater reliability (R), and being frequently (F) assessed.
The term target is used here in the same sense as operational definition or criterion. It is not restricted to behaviors, but it can include the outcomes of these behaviors as well. Hence, making errors as well as the errors them-selves could be considered targets.
Application
These SURF & C criteria are designed to provide standards for measuring behaviors needing improvement. As such, they are explicitly limited to these motivational situations. When the aim is performance improvement, however, they are highly recommended.
The SURF & C criteria lend themselves to providing consequences that are frequent, positive, and contingent, the hallmarks of a well-designed motivational program based on operant conditioning principles (Kazdin, 2000; Komaki, Coombs, Redding, and Schepman, 2000; Malott, Whaley, and Malott, 1997; Sulzer-Azaroff and Mayer, 1991). Furthermore, they reflect the operant model of effective supervision inspired by the theory (Komaki, 1998b). The frequent (F) assessment of targets is the foundation for providing consequences on a regularly reoccurring basis, as required in the frequency of reinforcement principle (Miller, 1997). The frequency criterion counters the common practice in organizations to appraise performance semiannually or annually. When targets are responsive to workers’ efforts, workers are more apt to engage in the desired targets. The reliability of the assessments (R), and the frequent (F) and direct sampling (S) of the target, enhance obtaining more representative and accurate information. In turn, the higher quality information enhances the closeness of the relationship between behavior and consequences. Thus, the groundwork is laid for providing contingent consequences. The criteria of being critical (C) and under control (U) also lend themselves to providing positive consequences. Management is more likely to recognize workers for target behaviors tapping critical aspects of the task. Furthermore, the process of gathering evidence to identify which behaviors are critical aids in the specification of the desired behaviors. Ensuring that the target behaviors are responsive to workers’ efforts and minimally affected by extraneous influences makes it more likely that workers will be motivated to improve and maintain their behavior.
Examples
To show how the criteria can be applied, examples are presented of different work groups doing a variety of jobs. Summarized in Table 3.1 , the examples are drawn from four areas, identified as safety, service, human services and education, and professional/complex.
To set the context for each of the areas, we will briefly identify the research questions, subjects and settings, and some of the dependent variables or targets. A frequent aim of safety experiments is to improve workers’ practices, with the ultimate aim of reducing accidents (Fellner and Sulzer-Azaroff, 1984; Hopkins, Conrad, and Smith, 1986; Komaki, Collins, and Penn, 1982; Ludwig and Geller, 1991; Reber and Wallin, 1983). The targets have consisted of various behaviors (e.g., wearing hard hats, ear protection, and goggles; working in areas with exhaust ventilation; spraying chemicals away from other workers; using seat belts and turn signals) and outcomes (e.g., housekeeping conditions or the status of materials in the organization). These studies took place in a paper mill, a plastic products manufacturing plant, and a pizzeria, with laborers, machine operators, and deliverers.
The second area deals with the service provided by bank tellers, foreign exchange clerks, and retail salespeople (Elizur, 1987; George, 1991; Komaki, Collins, and Temlock, 1987; Luthans, Paul, and Baker, 1981). Different aspects of service were assessed, including both functional—contact with customers (e.g., words of courtesy, smiling, approaching), service rendered (e.g., assessing customer’s needs, relating merchandise to needs), merchandise handling (e.g., tagging, arranging, replenishing, or unpacking merchandise—and dysfunctional behaviors (e.g., socializing, standing idly, and leaving the work area).
In the third area, staff worked in various human service and educational organizations—an institution, a center for psychological services, and a pediatric clinic (Callahan and Redmon, 1987; Frederiksen et al., 1982; Ingham and Greer, 1992; Iwata et al., 1976; Johnson and Frederiksen, 1983). The dependent variables included patient care (e.g., dental care and patient contact), the rate and accuracy of teaching behaviors, being on or off the unit, use of time (e.g., number of patients seen), and recording the progress of patients (e.g., errors made on client charts).
TABLE 3.1. Examples of Appraisals Meeting One or More SURF & C Criteria
The fourth area—designated professional/complex—was exemplified by a community board, a barnstorming baseball team, resource room teachers, graduate students working on their theses, and Marine Corps personnel within a heavy artillery battalion (Briscoe, Hoffman, and Bailey, 1975; Dillon, Kent, and Malott, 1980; Heward, 1978; Komaki, 1998a; Maher, 1982). The behaviors assessed included problem solving, teaching and discussion, team members’ overall contribution to the team’s run production, thesis-completion activities, and the preventive maintenance of heavy equipment.
IDENTIFYING WHAT SHOULD BE APPRAISED
To specify what should be evaluated, the criteria of U and C should be applied.
Make Sure the Target Is Under (U) the Control of the Worker
The importance of making sure that behaviors are under the control of workers is well illustrated in the area of safety. A typical way of judging employees’ safety practices is to look at the number of injuries that have occurred and then to infer that workers are performing unsafely when the injury rate is high. This assumption is not always warranted. Some injuries occur not because the worker has performed unsafely, but because of some extraneous factors (to that particular worker) such as the condition of the work environment and equipment. If coal miners have accidents after a roof caves in or a seat mechanism fails, it might not be appropriate to assume that their injuries were a function of unsafe acts on their part. When assessing how safely pizza deliverers drove, safety was defined not in terms of accidents occurring on the job, but in terms of driving practices—the use of safety belts and turn signals (Ludwig and Geller, 1991), each of which were under the control of the deliverers.
Along the same lines, behaviors representing service—words of courtesy, eye contact, and smiling—were chosen to represent the quality of customer service rendered in an Israeli bank (Elizur, 1987). The traditional measures, which include the number of transactions completed, customers’ letters of commendation or complaint, and voluntary responses to customer surveys, are more likely to be affected by forces beyond the control of employees, such as customer traffic, the season of the year, the economic climate, the merchandise mix, and customer perceptions of service.
An assumption sometimes made in the human services field is that better performing staff have patients who make more progress. The problem with this inference is that the progress of patients is typically a function of many factors, only one of which is staff performance. Patients may improve in spite of the quality of the care they receive, for instance. At the same time, patients may decline in functioning for reasons having little to do with the quality of care. Hence, strictly patientoriented appraisals are not recommended as the best (or sole) measure of job performance. Instead, the focus should be on performance that the staff can more readily control. Iwata and his colleagues (1976), for example, defined staff performance in terms of nine behaviors ranging from indirect and direct custodial work with residents to area supervision, each of which was minimally affected by extraneous factors.
The same case can be made for assessing students’ progress in completing their theses and dissertations. Sometimes the lack of student progress may not be solely attributable to a lack of student effort, but rather to factors not wholly under their control. As every student who has attempted to complete a thesis will attest, difficulties often arise in trying to gain acceptance from a diverse set of committee members. To assess students’ progress toward completion of their theses, an index was developed (Dillon, Kent, and Malott, 1980). From one to six major tasks were examined depending on the stage of the work: reading, reviewing, and discussing one article per week; presenting new data each week; attending a weekly halfhour meeting with supervisors; keeping a log of ideas, procedures, and changes gained from meetings, courses, and faculty; reporting the total number of hours worked on thesisresearch activities; formally writing 750 new words or rewriting; editing; and preparing a research proposal. Each of the tasks was more or less under the student’s control, and together they provided a more sensitive barometer of students’ progress toward completing their theses.
As shown in Table 3.1 , the indices under the control heading are more responsive to worker efforts than the measures that are typically used.
Empirically Verify the Critical (C) Aspects of the Task
The criteria of criticalness also should be applied when deciding what target behavior to assess. To do this, empirical verification of the relationship between the target and the ultimate criterion is required. When such a relationship exists, then and only then is the target behavior considered critical. Expert opinion, no matter how exalted, does not count as empirical evidence. Convening a corporate personnel manager, a store manager, and several highly ranked store associates to generate definitions of customer service (Komaki, Collins, and Temlock, 1987), for instance, would be considered only expert opinion and not as meeting the criterion of criticalness. The target behavior is judged as critical only when data are gathered on site about the target behavior and a relationship is found between the target behavior and the outcome of interest. To obtain such evidence, a Southwestern retailer gathered data on various customer service behaviors—as rated by superiors and the employees themselves—and sales, defined as sales per hour and standardized within departments. When a correlation of .20, p < .01, was found between the service behaviors and sales, these behaviors were thought to be critical.
Similar evidence was obtained in the area of safety. Although previous experimenters had designed observational measures of safety (e.g., Fellner and Sulzer-Azaroff, 1984; Hopkins, Conrad, and Smith, 1986; Komaki, Bar-wick, and Scott, 1978), they merely assumed that these measures were related to accidents. No attempt was made to collect similarly constructed observational data in a number of different departments and then examine the injury rates in these departments. Reber and Wallin (1983) took the commendable step of actually doing that. They assessed the relationship between an observational measure of safety (like that of Fellner and Sulzer-Azaroff, Hopkins, Conrad, and Smith, and Komaki, Barwick, and Scott) and both overall and lost-time injury rates over a three-year period. When correlations were obtained ranging from -.65 to -.76 between the safety behaviors and accidents, the observational measures of safety they were using as well as other similarly constructed measures were considered critical.
Confirming evidence of the target behavior was found in both of the previous examples. In some cases, confirmation is not obtained despite extensive data collection efforts. Johnson and Frederiksen (1983), for example, tried to obtain evidence that their target was critical. They defined the target behavior in their study as the total number of daily nursing staff contacts with patients in group sessions of a “reality orientation program.” The intended outcome—improved patient orientation to person, place, and time—was measured by the number of correct patient responses to questions posed by staff members assessing orientation to person, place, time, and past and recent events. Both the target behavior and the outcome were assessed for twenty-two weeks, and periodic interrater reliability checks were performed during data collection to check for accuracy. An improvement was shown in the target behavior of staff contact, but no relationship was found between the target behavior and the intended outcome of patient orientation. Although the results were not confirmatory, the investigators were a step ahead of many in the field: They knew that they should abandon the target behaviors and search for alternatives.
How does one generate ideas about target behaviors? Our recommendation is to look for opposites. Compare known extreme groups; that is, contrast a group known to possess a certain characteristic with a group lacking it. In developing targets in a poultry processing plant, for example, neophyte and seasoned employees were compared while performing the same operation (Komaki, Collins, and Perm, 1982). The differences in timing and motion that distinguished between the groups were used as the safety targets. In a much more extensive effort, Crawley and colleagues (1982) observed the top-pro-ducing sales people “during four months for 1,000 hours, as they worked with actual customers in stores and in homes” (p. 187) to see what these top producers actually did. Using the same strategy, Johnson and Frederiksen (1983) compared staff that have been successful in improving patient orientation with staff who have not been successful. The identified differences between the two groups were used as the basis for a new set of target behaviors.
Another way of sparking ideas about new target behaviors is to make comparisons between successful and unsuccessful situations. Contrasts can be made between weeks judged by superiors to be successful and those judged to be unsuccessful. To generate a new target in the area of preventive maintenance, Komaki (1998a) went on site and observed what occurred when preventive maintenance was and was not done successfully. During the successful weeks, she found that deficiencies were detected in the equipment; when these deficiencies were identified, the marines could order replacement parts and continue the maintenance chain. In contrast, during the unsuccessful weeks, few deficiencies were detected or reported or both, therefore reducing the number of successful follow-throughs. In this way, the new target of detected deficiencies was generated. To verify the criticalness of the targets, the relationship between the targets and the ultimate criterion (in this case, the actual deficiencies in the equipment) was found to be positively correlated. As the targets improved, the actual deficiencies declined over time. These two pieces of evidence lent credence to the criticalness of the targets.
Of the five criteria of criticalness fewer examples exist in which this criterion can be illustrated. However, as shown in Table 3.1 , a careful search of the literature unearthed at least one in each of the areas. In each case, no armchair opinions, no matter how expert or exalted, were taken as evidence. Data had to be gathered and the data had to show a significant relationship between the target behavior and the ultimate criterion.
Designating How Performance Should Be Appraised
To identify how the target behavior should be assessed, the criteria of S (direct sampling), R (interrater reliability), and F (frequent data collection) should be used.
Use the Test of Interrater Reliability (R)
A series of checks and balances are recommended, referred to as the test of interrater reliability (IRR). In this IRR test, two raters independently sample the work of an employee, using the same standards, and then check to see each time they agree or disagree. When the agreement score reaches 80 or 90 percent, then one has evidence that the appraisal actually reflects their performance.
This test helps ensure the clarity of definitions, the rigor of observer training, and the lack of bias in the formal data collection, and has been successfully used in a variety of cases, as shown in Table 3.1 . Reliability was calculated during data collection in a study of a time management program for resource room teachers (Maher, 1982). The formula used for occurrence reliability (of the targeted behaviors) was the total number of agreements that behaviors had occurred divided by the sum of the agreements and disagreements. Scores averaged 96 percent for pretraining sessions and 97 percent for posttraining sessions (the range was 84 to 100 percent).
To ensure that observers in a safety experiment were trained adequately, they completed qualifying practice observations until they obtained agreement scores of at least 90 percent (Hopkins, Conrad, and Smith, 1986). Reliability checks were also conducted during formal data collection. The checks were randomly performed on 17 percent of the items. If observers’ scores dropped below 90 percent, they needed to requalify before they could collect any more data. The median score reported for safety behaviors was 97 percent, with 94 percent reported for housekeeping conditions.
The test of interrater reliability was used at three points during the data collection process when assessing the ephemeral area of customer service (Komaki, Collins, and Temlock, 1987). First, it was used during the development of the measure as a benchmark to determine whether the definitions were clear and objectively defined. The definitions continued to be revised until the observers could reliably agree with one another on virtually all of the definitions. The observers frequently disagreed when they initially scored the quality of service provided. After the researchers made several abortive attempts to change the definition and to improve the reliability scores, the global “quality” item was dropped and a list of specific types of comment categories was added. Only then did the observers pass the tests of interrater reliability. Second, the test of interrater reliability was used in training the observers; observers were not considered to be trained until they obtained scores of at least 90 percent on three consecutive, representative occasions. Lastly, to ensure that customer service performance was being accurately measured, reliability was assessed during the formal data collection phase. The reliability percentage score was calculated as the number of agreements divided by the number of agreements plus disagreements. Interrater reliability was assessed during 8 percent of the formal observations. Overall, the average reliability was 97.8 percent.
Interrater reliability was calculated in two different ways in a two-phase study conducted in a mental health setting (Ingham and Greer, 1992). In the first part, the test was used to train the primary observer, or to “calibrate” the individual to a standard. This was necessary because the teachers, the subjects of the study, would not allow videotaping or more than one observer to collect data. Therefore, the supervisor, who was the primary observer, had to learn to be a reliable and accurate coder before the data were collected. To do this, she coded “previously validated videotapes” of the same types of teachers in a similar setting. Indices of agreement were calculated by dividing the number of agreements by the number of agreements plus disagreements. Percentages ranged from 85 to 97 percent, with a mean of 91 percent. In the second phase, conducted in the same setting one year later, data were collected via the supervisor, the teachers themselves, and videotape. A second observer coded the videotape, and interobserver agreement scores were calculated using the same formula. (Note: Ideally, the data in each of these two phases should be collected during a random and a less restricted set of observations.) Scores were calculated for each behavior measured. Scores between the supervisor and the observer ranged from 75 to 100 percent, with a mean of about 96 percent. The average percentage agreement between the supervisor and teachers was 94 percent, ranging between 85 and 100 percent.
Collect Data Frequently (F)
To meet the frequency criterion, information should be collected at least twenty to thirty times during the intervention period. This recommendation is based on: (a) theoretical considerations (Miller, 1997)—one of the most straightforward ways to increase the potency of an intervention is to increase its frequency; (b) statistical concerns regarding the representativeness of the information obtained (Cronbach et al, 1972) and the risks involved in drawing conclusions with too few data points in a time series (Gottman, 1981), with one statistician (R. Millsap, personal communication, April 10, 1997) stipulating that at least twenty to thirty data points are necessary to discern trends reliably; and (c) the reactions of target subjects. Employees who thought their evaluations were more fair and accurate identified their appraisals as being conducted more often (Landy, Barnes, and Murphy, 1978).
The assessment of safety practices during the forty-six-week period of one study was judged to meet the criterion of frequency (Komaki, Collins, and Penn, 1982). The field experiment consisted of four groups observed over two conditions. Across the four groups, thirty-two, thirty, twenty-nine, and thirty-two observations of behaviors were reported during the first experimental condition; seventy-seven, sixty-four, fifty, and thirty-nine were reported for the second condition.
During the experimental phase of another study, service behaviors during the intervention phase were reported on a graph eighteen times in the experimental group and twenty-one times in the control group (Luthans, Paul, and Baker, 1981).
Data on three variables—the amount of time a patient spent in the clinic, staff use of time, and patient satisfaction—were collected for approximately five months in a pediatric outpatient clinic (Callahan and Redmon, 1987). The baseline phase of one type of patient scheduling system consisted of twenty-eight data points, while two intervention phases of another type of scheduling system contained nineteen data points each. Data were also collected eight times in the return-!to-baseline phase separating the interventions and thirteen times during a follow-up period.
In an experiment by Komaki (1998a) conducted in the Marine Corps to improve preventive maintenance of equipment, data were collected in one group on two target behaviors twenty-one to twenty-three times during the thirty-five-week intervention period, and in the other group fifteen to sixteen times during the twenty-five-week intervention period. This was an average data collection of .6 times per week. At the end of the intervention period, which occurred after a year-long experiment, it was determined if the intervention had been effective.
Directly Sample (S) the Target
Rather than secondhand and/or filtered reports, assessments should be firsthand. Workers should be appraised as they are doing the work, or the products of their work should be examined. To assess safety, for example, observers recorded the condition of materials and equipment or observed the safe and unsafe behaviors of employees while they were operating a machine or conducting a task (Fellner and Sulzer-Azaroff, 1984). Similarly, observers watched Israeli bank employees while standing in line be-tween customers (Elizur, 1987). This was done during “high load” hours when many customers were waiting. Along the same line, client charts, a common product in the area of mental health, were sampled for errors (Frederiksen et al., 1982). In the professional/complex area, community board meetings were videotaped and then coded for target behaviors (Briscoe, Hoffman, and Bailey, 1975). An interesting and unique measure of performance for barnstorming baseball players also utilized direct sampling (Heward, 1978). The measure consisted of a numerical description of the individual’s contribution to the team’s run production. This was calculated by dividing the total number of times the player went to the plate into the total number of hits, runs, runs batted in, walks, sacrifices, and hits by a pitch that the player accumulated during times at bat.
In all of the previous examples, investigators showed how they had successfully applied one or more of the SURF & C criteria. In some cases, the researchers could make use of existing indices; Heward (1978) did this when he tallied individual baseball players’ actions such as the number of times a player goes to the plate. In other cases, Briscoe, Hoffman, and Bailey (1975) had to get permission to set up videotaping equipment and then code the data to ensure that community board members’ actions would be directly sampled.
By presenting examples occurring in ongoing work settings, we hope that readers will be better able to evaluate their current and planned appraisals. To return to one of the previous examples of the IRS: An audit of the agency substantiated claims that taxpayers were being treated shoddily (Rosenbaum, 1998). A major criticism was that the appraisals were primarily based on “statistical enforcement goals and not on the services they provided taxpayers” (p. A15). A proposed plan was to simply deemphasize the financial goals by “ending the practice or ranking … districts on the basis of how closely they met their goal of tax collections” (p. A15). Just because districts are no longer ranked, however, does not mean that employees will necessarily be judged on the judicious treatment of taxpayers. Another suggestion—to bolster the treatment of taxpayers by setting up “a board made up largely of private citizens … to make it easier for taxpayers to prevail in tax court” (p. A15)—is no substitute for a fair and accurate measure. Besides questions of content, such a board could not meet the frequency, reliability, or direct sampling criteria.
At the same time, we hope that these actual examples will help future investigators to better see how they too can enhance the appraisals in their own organizations and perhaps be inspired to initiate these efforts. The criteria can be beneficially applied to the problematic but unfortunately widespread rating scale used at both Texaco and Stockham Valves (Komaki, in preparation). To clarify the definitions of such vague dimensions as dependability and work quality, the test of interrater reliability could be profitably used in the development of a new performance appraisal form. When raters disagree with one another, for example, their disagreements should provoke discussions about how to change the definitions so that the disagreements diminish. After rewriting the definitions and obtaining improvements in the reliability scores, it should be clearer as to what constitutes performance, leading hopefully to more uniform evaluations on the next round. Similarly, the IRR test should be used, in the training of raters as a standard to show they are qualified, and during the formal data collection to show that the raters’ standards are not shifting over time. Both of these will also enable the screening out of biased evaluators. Finally, the supervisors should collect data directly and frequently.
Likewise, the assessment of service, whether it be for airlines, hotels, or restaurants, could be improved by meeting the criterion of criticalness and seeking out empirical rationale such as that provided by Bitner, Booms, and Tetreault (1990). They collected hundreds of critical incidents, providing evidence from the customers’ point of view of satisfactory service encounters. The incidents went beyond smiling and greeting to highlight the ways in which personnel handled failures and how they responded to customers with special needs. Based on criticality data like these, developers of appraisal instruments could empirically verify what matters to the customer in providing constitutes quality service.
In short, when developing and evaluating performance appraisals, the SURF & C criteria are highly recommended. The criteria of being under the workers’ control (U) and critical to the task at hand (C) address the content of the target. The criteria of sampling directly (S), collecting information frequently (F), and conducting interrater reliability (R) checks identify how the appraisals should be done. Meeting all these criteria helps to minimize such perennial and pernicious problems as rater bias and to promote the fair and accurate appraisal of workers’ performance on the job, the foundation for any effective and sustained improvement effort.
Johnson, C. M., Mawhinney, T. C. & Redmon, W. K. (2001). Handbook of Organizational Performance: Behavior Analysis and Management (1st ed.). New York: Haworth Press.
REFERENCES
After critical inquiry, I.R.S. turns to ethics expert. (1992). The New York Times, April 7, p. A18.
Altman, L. K. (1992). Surgical scorecards: Can doctors be rated just like ballplayers? The New York Times, January 14, p. C3.
Ashe, R. L. and McRae, G. S. (1985). Performance evaluations go to court in the 1980s. Mercer Law Review, 36, 887-905.
Baker, G., Gibbons, R., and Murphy, K. (1994). Subjective performance measures in optimal incentive contracts. Quarterly Journal of Economics, 109, 1125-1156.
Barrett, G. V. and Kernan, M. C. (1987). Performance appraisal and terminations: A review of court decisions since Brito v. Zia with implications for personnel practices. Personnel Psychology, 40, 489-503.
Bellack, A. S. and Hersen, M. (Eds.). (1988). Behavioral assessment: A practical
Berg, E. N. (1990). United thrives amid turmoil. The New York Times, October 2, p. C1.
Bitner, M. J., Booms, B. H., and Tetreault, M. S. (1990). The service encounter: diagnosing favorable and unfavorable incidents. Journal of Marketing, 54, 71-84.
Blanz, R. and Ghiselli, E. E. (1972). The mixed standard scale: A new rating system. Journal of Applied Psychology, 63, 677-688.
Blum, M. L. and Naylor, J. C. (1968). Industrial psychology, its theoretical and social foundations (Revised edition). New York: Harper & Row.
Borman, W. C. (1991). Job behavior, performance, and effectiveness. In M. Dunnette and L. Hough (Eds.), Handbook of industrial and organizationalpsychology (Vol. 2, pp. 271-326). Palo Alto, CA: Consulting Psychologists Press.
Bravo, E. (1991). Mistrust and manipulation: Electronic monitoring of the American workforce. USA Today, 119 May, 46-48.
Briscoe, R. V., Hoffman, D. B., and Bailey, J. S. (1975). Behavioral community psychology: Training a community board to problem solve. Journal of Applied Behavior Analysis, 8, 157-168.
Burns, T. (1954). The directions of activity and communication in a departmental executive group. Human Relations, 7, 73-97.
Byer, M. J. (1992). Faint hearts. The New York Times, March 21, p. A23.
Callahan, N. M. and Redmon, W. K. (1987). Effects of problem-based scheduling on patient waiting and staff utilization of time in a pediatric clinic. Journal of Applied Behavior Analysis, 20, 193-199.
Campbell, J. P. (1990). Modeling the performance prediction problem in industrial and organizational psychology. In M. D. Dunnette and L. M. Hough (Eds.), Handbook of industrial and organizational psychology (pp. 687-732). Palo Alto, CA: Consulting Psychologists Press.
Cardy, R. L. and Dobbins, G. H. (1994). Performance appraisal: Alternative perspectives. Florence, KY: South Western.
Cascio, W. F. and Bernardin, H. J. (1981). Implications of performance appraisal litigation for personnel decisions. Personnel Psychology, 34, 211-226.
Ciminero, A. R., Calhoun, K. S., and Adams, H. E. (Eds.). (1986). Handbook of behavioral assessment. New York: Wiley-Interscience.
Crawley, W. J., Adler, B. S, O’Brien, R. M., and Duffy, E. M. (1982). Making salesman: Behavioral assessment and intervention. In R. M. O’Brien, A. M. Dickinson, and M. P. Rosow (Eds.), Industrial behavior modification: A management handbook (pp. 184-199). New York: Pergamon Press.
Cronbach, L. J., Gleser, G. C, Nanda, H., and Rajarathnam, N. (1972). The dependability of behavioral measures. New York: Wiley.
DeNisi, A. S. (1996). A cognitive approach to performance appraisal: A program of research. London: Routledge.
DeNisi, A. S., Cafferty, T. P., and Meglino, B. M. (1984). A cognitive view of the appraisal process: A model and research propositions. Organizational Behavior and Human Performance, 33, 360-396.
DeVries, D. L., Morrison, A. M., Shullman, S. L., and Gerlach, M. L. (1980). Performance appraisal on the line. Greensboro, NC: Center for Creative Leadership.
Dillon, M. J., Kent, H. M., and Malott, R. W. (1980). A supervisory system for accomplishing long-range projects: An application to master’s thesis research. Journal of Organizational Behavior Management, 2, 213-227.
Doyle, S. X. and Shapiro, B. P. (1980). What counts most in motivating your sales force. Harvard Business Review, May-June, 133-140.
Dunnette, M. D. (1963). A note on the criterion. Journal of Applied Psychology, 47, 251-254.
Eichenwald, K. (1996). Texaco executives, on tape, discussed impeding a bias suit. The New York Times, November 4, pp. A1, D4.
Elizur, D. (1987). Effect of feedback on verbal and nonverbal courtesy in bank setting. Applied Psychology: An International Review, 36, 147-156.
Feild, H. S. and Holley, W. (1982). The relationship of performance appraisal system characteristics to verdicts in selected employee discrimination cases. Academy of Management Journal, 25, 392-406.
Feldman, J. M. (1981). Beyond attribution theory: Cognitive processes in performance appraisal. Journal of Applied Psychology, 66, 127-148.
Fellner, D. J. and Sulzer-Azaroff, B. (1984). Increasing industrial safety practices and conditions through posted feedback. Journal of Safety and Research, 15, 7-21.
Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51, 327-355.
Foster, S. L. and Cone, J. D. (1986). Design and use of direct observation procedures. In A. R. Ciminero, K. S. Calhoun, and H. E. Adams (Eds.), Handbook of behavioral assessment (Second edition, pp. 253-324). New York: Wiley-Inter-science.
Frederiksen, L. W., Richter, W. T, Johnson, R. P., and Solomon, L. J. (1982). Specificity of performance feedback in a professional service delivery setting. Journal of Organizational Behavior Management, 5(4), 41-53.
George, J. M. (1991). State or trait: Effects of positive mood on prosocial behaviors at work. Journal of Applied Psychology, 76, 299-307.
Glanz, J. (2000). Amid race profiling claims, Asian-Americans avoid labs. The New York Times, July 7, p. Al.
Goldfried, M. R. and Kent, R. N. (1972). Traditional versus behavioral personality assessment: A comparison of methodological and theoretical assumptions. Psychological Bulletin, 77(6), 409-420.
Gomez-Mejia, L. R. (1988). Evaluating employee performance: Does the appraisal instrument make a difference? Journal of Organizational Behavior Management, 9(2), 155-172.
Gottman, J. M. (1981). Time series analysis: A comprehensive introduction for social scientists. Cambridge, UK: Cambridge University Press.
Griffith, T. L. (1993). Teaching big brother to be a team player: Computer monitoring and quality. Academy of Management Executive, 7, 73-80.
Hammer, M. (1985). Implications of behavioral and cognitive reciprocity in social network data. Social Networks, 7,189-201.
Heward, W. L. (1978). Operant conditioning of a .300 hitter? The effects of reinforcement on the offensive efficiency of a barnstorming baseball team. Behavior Modification, 2, 25-40.
Hopkins, B. L., Conrad, R. J., and Smith, M. J. (1986). Effective and reliable behavioral control technology. American Industrial Hygiene Association Journal, 47(12), 785-791.
Ilgen, D. R. and Feldman, J. M. (1983). Performance appraisal: A process focus. In L. Cummings and B. Staw (Eds.), Research in organizational behavior (Vol. 5, pp. 141-197). Greenwich, CT: JAI.
Ingham, P. and Greer, R. D. (1992). Changes in student and teacher responses in observed and generalized settings as a function of supervisor observations. Journal of Applied Behavior Analysis, 25, 153-164.
Iwata, B. A., Bailey, J. S., Brown, K. M., Foshee, T. J., and Alpern, M. (1976). A performance-based lottery to improve residential care and training by institutional staff. Journal of Applied Behavior Analysis, 9, 417-431.
James v. Stockham Valves and Fittings Co., 559 F.2d 310 (U.S. Ct. App. 5th Circuit 1977).
James, L. R. (1973). Criterion models and construct validity for criteria. Psychological Bulletin, 80, 75-83.
Johnson, R. P. and Frederiksen, L. W. (1983). Process vs. outcome feedback and goal setting in a human service organization. Journal of Organizational Behavior Management, 5(3/4), 37-56.
Johnston, D. (1998). Black F.B.I. agents renew bias complaint. The New York Times, October 15, p. A24.
Johnston, J. M. and Penny packer, H. S. (1993). Strategies and tactics of human behavioral research (Second edition). Hillsdale, NJ: Lawrence Erlbaum.
Kazdin, A. E. (2000). Behavior modification in applied settings (Sixth edition). Belmont, CA: Wadsworth Thomson Learning.
Kilborn, P. T. (2000). For women in bias case, the wounds remain. The New York Times, March 24, p. A14.
Komaki, J. L. (1998a). When performance improvement is the goal: A new set of criteria for criteria. Journal of Applied Behavior Analysis, 31(2), 263-280.
Komaki, J. L. (1998b). Leadership from an operant perspective. London: Routledge.
Komaki, J. L. (in preparation). Daring to dream: How organizations can come closer to fulfilling Martin Luther King’s (1963) dream. Applied Psychology: International Review.
Komaki, J. L., Barwick, K. D., and Scott, L. R. (1978). A behavioral approach to occupational safety: Pinpointing and reinforcing safe performance in a food manufacturing plant. Journal of Applied Psychology, 63, 434-445.
Komaki, J. L., Collins, R. L., and Penn, P. (1982). The role of performance antecedents and consequences in work motivation. Journal of Applied Psychology, 67, 334-340.
Komaki, J. L., Collins, R. L., and Temlock, S. (1987). An alternative performance measurement approach: Applied operant measurement in the service sector. Applied Psychology: An International Review, 36, 71-89.
Komaki, J. L., Coombs, T., Redding Jr., T. P., and Schepman, S. (2000). A rich and rigorous examination of applied behavior analysis research in the world of work. In C. L. Cooper and I. T. Robertson (Eds.), International Review of Industrial and Organizaitonal Psychology. Sussex, England: John Wiley.
Landy, F. J., Barnes, J. L., and Murphy, K. R. (1978). Correlates of perceived fairness and accuracy of performance evaluation. Journal of Applied Psychology, 63, 751-754.
Landy, F. J. and Farr, J. L. (1980). Performance rating. Psychological Bulletin, 87, 72-107.
Latham, G. P. and Latham, S. D. (2000). Overlooking theory and research in performance appraisal at one’s peril: Much done, more to do. In C. Cooper and E. A. Locke (Eds.), International Review of Industrial-Organizational Psychology. Chichester, England: Wiley.
Latham, G. P., Skarlicki, D., Irvine, D., and Siegel, J. P. (1993). The increasing importance of performance appraisals to employee effectiveness in organizational settings in North America. In C. L. Cooper and I. T. Robertson (Eds.), International review of industrial and organizational psychology (Vol. 8, pp. 87-131). Chichester, NY: John Wiley & Sons.
Latham, G. P. and Wexley, K. N. (1981). Increasing productivity through performance appraisal. Reading, MA: Addison-Wesley Publication Co.
Lewis, D. R. and Dahl, T. (1976). Time management in higher education administration: A case study. Higher Education, 5, 49-66.
Longenecker, C. O., Sims, H. P., and Gioia, D. A. (1987). Behind the mask: The politics of employee appraisal. The Academy of Management Executive, I, 183-193.
Ludwig, T. S. and Geller, E. S. (1991). Improving the driving practices of pizza deliverers: Response generalization and moderating effects of driving history. Journal of Applied Behavior Analysis, 24, 31-44.
Luthans, F, Paul, R., and Baker, D. (1981). An experimental analysis of the impact of contingent reinforcement on salespersons’ performance behavior. Journal of Applied Psychology, 3, 314-323.
Maher, C. A. (1982). Improving teacher instructional behavior: Evaluation of a time management training program. Journal of Organizational Behavior Management, 4(3/4), 27-36.
Malott, R. W, Whaley, D. L., and Malott, M. E. (1997). Elementary principles of behavior (Third edition). NJ: Prentice-Hall.
Miller, L. K. (1997). Principles of everyday behavior analysis (Third edition). Pacific Grove, CA: Brooks/Cole.
Murphy, K. R. and Cleveland, J. (1995). Understanding performance appraisal: Social, organizational, and goalbased perspectives. Thousand Oaks, CA: Sage.
Nasar, S. (1994). Jett’s supervisor at Kidder breaks silence. The New York Times, July 26, p. D1, D14.
Reber, R. A. and Wallin, J. A. (1983). Validation of a behavioral measure of occupational safety. Journal of Organizational Behavior Management, 5(2), 69-77.
Roberts v. Texaco, 94 Civ. 2015 (CLB, 1994).
Rosenbaum, D. E. (1998). Internal audit confirms abusive I.R.S. practices. The New York Times, January 14, p. A15.
Schrage, M. (1992). When technology heightens office tensions. The New York Times, October 5, p. A12.
Shaw, D. G., Schneier, C. E., Beatty, R. W., and Baird, L. S. (1995). The performance measurement, management, and appraisal sourcebook. Amherst, MA: Human Resource Development Press.
Smith, P. C. (1976). Behaviors, results, and organizational effectiveness: The problem of criteria. In M. Dunnette (Ed.), Handbook of industrial and organizational psychology (pp. 745-775). New York: John Wiley & Sons.
Smith, P. C. and Kendall L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47, 149-155.
Smither, J. W. (Ed.). (1998). Performance appraisal. San Francisco: Jossey-Bass.
Steers, R. M, Porter, L. W. and Bigley, G. A. (1996). Motivation and leadership at work. New York: McGraw-Hill.
Sulzer-Azaroff, B. and Mayer, G. R. (1991). Behavior analysis for lasting change. Fort Worth: Holt, Rinehart and Winston.
Weist, M. D., Ollendick, T. H., and Finney, J. W. (1991). Toward the empirical validation of treatment targets in children. Clinical Psychology Review, 11, 515-538.
Werner, J. M. and Bolino, M. C. (1997). Explaining U.S. courts of appeals decisionsz involving performance appraisal: Accuracy, fairness, and evaluation. Personnel Psychology, 50(1), 1-24.
(Redmon)
Redmon, William K. Handbook of Organizational Performance. Routledge, 20130403. VitalBook file.
The citation provided is a guideline. Please check each citation for accuracy before use.