HUMAN RESOURCE ASSIGNMENT

profilejaybird
topic_6_art_1.pdf

A Meta-Analysis of the Relationship Between Individual Assessments and Job Performance

Scott B. Morris, Rebecca L. Daisley, Megan Wheeler, and Peggy Boyer Illinois Institute of Technology

Though individual assessments are widely used in selection settings, very little research exists to support their criterion-related validity. A random-effects meta-analysis was conducted of 39 individual assess- ment validation studies. For the current research, individual assessments were defined as any employee selection procedure that involved (a) multiple assessment methods, (b) administered to an individual examinee, and (c) relying on assessor judgment to integrate the information into an overall evaluation of the candidate’s suitability for a job. Assessor recommendations were found to be useful predictors of job performance, although the level of validity varied considerably across studies. Validity tended to be higher for managerial than nonmanagerial occupations and for assessments that included a cognitive ability test. Validity was not moderated by the degree of standardization of the assessment content or by use of multiple assessors for each candidate. However, higher validities were found when the same assessor was used across all candidates than when different assessors evaluated different candidates. These results should be interpreted with caution, given a small number of studies for many of the moderator subgroups as well as considerable evidence of publication bias. These limitations of the available research base highlight the need for additional empirical work to inform individual assessment practices.

Keywords: executive selection, meta-analysis, validation study, personnel selection, individual psychological assessment

Supplemental materials: http://dx.doi.org/10.1037/a0036938.supp

Much of the research on personnel selection has focused on standardized assessment methods, such as cognitive ability tests, self-report personality inventories, and structured interviews. These assessment methods are designed to be administered and scored efficiently for large groups of job applicants. Another tradition in assessment takes a very different approach, emphasiz- ing assessments that are tailored to the individual and the organi- zation. This latter approach, called individual psychological as- sessment, emphasizes the measurement of the person as a whole (Highhouse, 2002) and the use of expert judgment to interpret the pattern of responses across a variety of assessment methods (Silzer & Jeanneret, 2011). This job- and person-specific approach is particularly attractive in settings such as executive selection, where there are few candidates for a single job opening and where job requirements are highly individualized (Hollenbeck, 2009).

Individual psychological assessment is defined as a process of gathering information regarding a person’s knowledge, skills, ap- titude, and temperament (Jeanneret & Silzer, 1998) through the use

of individually administered selection tools, including tests and interviews, and integrating this information to make an infer- ence regarding the individual’s appropriateness for a particular position (Prien, Schippman & Prien, 2003). Prien et al. (2003) stated that the distinguishing feature of individual assessment is that it combines integrative and interpretive treatments of as- sessment data at the individual-case level. The uniqueness of individual assessment rests on the idea that through the use of various tools, professionals use the individual psychological assessment to evaluate the applicant as a whole and make a prediction about the applicant’s appropriateness for a position within an organization holistically. In addition, individual as- sessment differs from other forms of testing in that the assessor must draw a conclusion about a candidate’s fit by utilizing his or her own subjective interpretation of the candidate’s perfor- mance on the tests administered, the interview, and any other data collected (Jeanneret & Silzer, 1998).

Individual assessment is widely used in employee selection, particularly for higher level positions (Thornton, Hollenbeck & Johnson, 2010). Despite this, there has been relatively little re- search published on the validity of individual assessments. The most extensive review of the validity evidence was provided by Prein et al. (2003), who identified approximately 20 validation studies. For the most part, the evidence supported the validity of individual assessments across a variety of occupations, including mangers (Albrecht, Glaser & Marks, 1964), first-line supervisors (Dicken & Black, 1965; Dunnette & Kirchner, 1958; Handyside & Duncan, 1954), and management consultants (Miner, 1970), among several others. However, there was considerable variability

This article was published Online First May 26, 2014. Scott B. Morris, Rebecca L. Daisley, Megan Wheeler, and Peggy Boyer,

Department of Psychology, Illinois Institute of Technology. The authors wish to thank Paige Olson for her assistance with the coding

of studies. We also thank the many people who helped us obtain informa- tion on unpublished studies, including Joy Hazucha, Maynard Goff, Bob Barnett, and Dave Sowinski.

Correspondence concerning this article should be addressed to Scott B. Morris, Department of Psychology, Illinois Institute of Technology, 3105 South Dearborn, Chicago, IL 60616. E-mail: [email protected]

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

Journal of Applied Psychology © 2014 American Psychological Association 2015, Vol. 100, No. 1, 5–20 0021-9010/15/$12.00 http://dx.doi.org/10.1037/a0036938

5

in the results, with validity coefficients ranging from –.14 (Miner, 1970) to as high as .70 (Phelan, 1962).

Hunter and Schmidt (2004) have noted that observed differences across studies can often be attributed to statistical artifacts, rather than true differences in validity. Meta-analysis provides a means to disentangle true and artificial sources of variability and therefore provides a stronger basis for estimating the true validity and identifying moderators. In this study, we addressed the inconsis- tencies that have been shown in previous reviews and investigated whether individual assessments are valid across situations. In addition, surveys of the field have reported considerable variability in assessment practices (Ryan & Sackett, 1987, 1992). By model- ing differences in validity across studies, we sought to identify best practices in individual assessment.

Reliability and Validity of Assessor Judgments

Most individual assessments include tests (e.g., cognitive abil- ity, personality, biodata) that have been shown to predict work performance. Still, it is an open question whether this relationship remains when holistic integration is used to make predictions concerning performance. A central feature of the individual as- sessment is the reliance on the assessors’ expert judgment to draw inferences from patterns of behavior (Silzer & Jeanneret, 2011) and fit the multiple pieces of information together into a coherent picture of the candidate as a whole (Highhouse, 2002). Therefore, it is important to assess not only the validity of the component test scores but also the validity of assessor judgments. To the extent that assessors can go beyond individual test scores to provide a holistic evaluation of candidates, assessor judgments would be expected to show greater validity than individual assessment tools. On the other hand, if assessors are biased or inconsistent in their use of test information, assessor judgments may show lower va- lidity than obtained from individual test scores.

Ryan and Sackett (1998) pointed out that lack of reliability of ratings and recommendations is a major issue with individual assessments. Several studies have examined the interrater agree- ment of assessors who made ratings based on assessment reports (DeNelsky & McKee, 1969; Dicken & Black, 1965) and those who had access to test scores and interview notes (Hilton, Bolin, Parker, Taylor, & Walker, 1955), as well as individuals who made pre- dictions based on the entire assessment that they conducted them- selves (Ryan & Sackett, 1989). These studies suggest that there can be considerable disagreement among assessors in ratings of candidate attributes and overall person-job fit, as well as inconsis- tencies in the assessment process itself (Ryan & Sackett, 1989).

Impact of Assessment Practices on Validity

The individual assessment process generally consists of three stages: information input, information evaluation, and information output (Weiner, 2003). During the information input stage, the individual assessor uses a variety of tools to collect information about the job candidate. These tools may include any combination of an interview, personality tests, cognitive ability tests, and other instruments. In the second stage, the collected data are interpreted and integrated. In the final stage, the information output stage, the assessor provides a report and recommendation regarding the suitability of the individual job candidate for a particular position.

There is significant variability in the methods of conducting individual assessments (Jeanneret & Silzer, 1998; Ryan & Sackett, 1987, 1992; Silzer & Jeanneret, 2011), and the choice of strategies for information input and integration are likely to influence the validity of assessor recommendations. Research on employment interviews, which are similar to individual assessments in their reliance on the interpretative skills of the interviewer, suggests that validity can differ as a function of both the content and structure of the interview. In particular, most reviews show higher reliability and validity for more structured interviews (Arvey & Campion, 1982;Campion, Palmer, & Campion, 1997; Harris, 1989; Huffcutt & Arthur, 1994; Schmitt, 1976; Ulrich & Trumbo, 1965). There- fore, the literature on structure provides a useful framework for examining differences in individual assessment practices. The following sections discuss components of information input and information integration that may affect the validity of individual assessments.

Information Input

In the information input stage, the collection of data is the primary interest. One of the defining characteristics of the indi- vidual assessment is the use of multiple assessment tools. Al- though there may be some overlap in the constructs that are assessed by each of the assessment tools, research has shown that combinations of different assessment tools can add incrementally to the validity of the selection battery (Schmidt & Hunter, 1998). According to a survey of individual assessors, most reported using a combination of personal history forms (bio-data), ability tests, personality inventories, and interviews (Ryan & Sackett, 1987). A substantial body of research supports the use of each of these selection methods. Personal history forms show validities of .30 – .37 (Mumford & Stokes, 1992; Reilly & Chao, 1982). The validity of general cognitive ability tests ranges from .23 for the least complex jobs to .58 for professional jobs (Schmidt & Hunter, 1998). The personality trait of conscientiousness has been found to be a consistent predictor of job performance, with validities rang- ing from .20 to .23 across a variety of occupational groups (Barrick & Mount, 1991).

Most individual assessments also include an interview (Ryan & Sackett, 1987). Personal contact between the assessor and the candidate is a central feature that distinguishes individual assess- ments from group administration of test batteries. Silzer and Jean- neret (2011) maintained that this personal contact allows the assessor to observe patterns of candidate behavior, yielding unique insights about factors such as interpersonal and communication skills that are not well assessed using standardized measures. Further, the interview allows the assessor to probe for additional information in order to test hypotheses about candidate character- istics. Research on employment interviews yields validity coeffi- cients ranging from .20 to .62 (Conway, Jako, & Goodman, 1995; McDaniel, Whetzel, Schmidt, & Mauer, 1994), with higher valid- ities for more structured interviews.

Although all of these assessment tools are useful predictors of job performance, it is important to note differences between the research literature and how these tools are used in the context of individual assessment. For example, although 80% of individual assessors report using personal history data (Ryan & Sackett, 1987), there is little information on exactly how personal history

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

6 MORRIS, DAISLEY, WHEELER, AND BOYER

data are collected and scored. Research on biodata has emphasized systematic methods for developing, scoring, and weighting per- sonal history items. When biodata items are not developed from empirical evidence or are not weighted properly, it is questionable how valid and reliable they may be (Stokes & Cooper, 2004).

Similarly, research on the validity of personality assessment has focused on personality trait measures, particularly those based on the Big Five theory (Barrick & Mount, 1991). Although this approach is used in many individual assessments, others use per- sonality inventories designed to identify abnormal personality pat- terns, such as the Minnesota Multiphasic Personality Inventory (Butcher, 1994) or projective techniques such as the Thematic Apperception Test (Reilly & Chao, 1982). Although there is some support for the validity of the MMPI (Aamodt, 2004) and projec- tive tests (Reilly & Chao, 1982) as predictors of job performance, research on the use of these techniques for personnel selection is limited and has been criticized for methodological shortcomings (Kinslinger, 1966).

Research on the employment interview has consistently shown greater reliability and validity for interviews that have greater structure, that is, interviews that are based on a job analysis, that employ well-trained interviewers, and that use behaviorally an- chored rating scales (Campion, Pursell, & Brown, 1988; Huffcutt & Arthur, 1994; Wiesner & Cronshaw, 1988). Individual assess- ments, due to their individualized nature, are likely to impose less structure on the interview.

In sum, individual assessments tend to involve administration of a variety of selection tools that are known to predict job perfor- mance, and the use of multiple assessments should further enhance the validity the process. At the same time, the use of less structured approaches may limit validity to some extent. Nevertheless, the research on the components of individual assessment suggests that each should provide some incremental validity to the assessor’s recommendation.

Hypothesis 1a: Assessor recommendations from assessments that include a test of general cognitive ability will be better predictors of performance than those from assessments that do not include general cognitive ability.

Hypothesis 1b: Assessor recommendations from assessments that include a personality questionnaire will be better predic- tors of performance than those from assessments that do not include a personality questionnaire.

Hypothesis 1c: Assessor recommendations from assessments that include a biodata measure will be better predictors of performance than those from assessments that do not include a biodata measure.

Hypothesis 1d: Assessor recommendations from assessments that include an interview will be better predictors of perfor- mance than those from assessments that do not include an interview.

Information input can also be characterized by the degree of structure. At the highest level of structure, the questions that are asked of the candidates are identical in every way (Campion et al., 1997). As the interview for each candidate becomes more unique, it becomes increasingly difficult to assess candidates on the same

criteria, which can affect the validity of the recommendations. Given the individualized nature of individual assessments, candi- dates are often asked to complete slightly different assessment components. A survey by Ryan and Sackett (1987) found that only 15% of individual assessment practitioners reported using a highly standardized format, 73% reported a using a loose structure, and 12% reported an unstructured format. Many assessors utilize an adaptive interviewing approach in which questions are designed to test hypotheses or seek clarification on information from the assessment battery (McPhail & Jeanneret, 2011; Ryan & Sackett, 1987). The standardization of the assessment content is important because it ensures that the collected information is comparable across candidates. As such, the standardization of the content allows the assessors to make meaningful comparisons across ap- plicants, which should lead to higher validities.

Hypothesis 2: Assessor recommendations from assessment batteries that are identical for each candidate will be more predictive of performance than those from batteries that differ across participants.

Information Integration

After the assessment tools have been administered, the infor- mation must be integrated by the individual assessor, who com- bines all of the collected information into a final recommendation for the hiring organization. Prien et al. (2003) noted that the information integration stage is really the heart of the individual assessment process. It is through this process that the assessor uses both quantitative and qualitative information to make a holistic assessment of the job candidate. Because this is a largely subjec- tive process, the individual assessors themselves are central to the information integration stage.

The integration of information from an assessment battery can take place either statistically or through expert judgment. Statisti- cal integration of information involves techniques that combine the scores from each assessment tool mechanically into a composite score. Judgmental integration relies on the assessor to combine information in a nonstatistical, clinical manner. Holt (1958) de- scribed clinical judgment in terms of making decisions and reach- ing conclusions by thinking over the facts and theories that are already available to the decision maker. Individual assessments tend to use this judgmental integration approach, and this will be the focus of the current study.

In a definitive piece, Meehl (1954) argued that there is no theoretical or empirical basis supporting the idea that people can combine information in their heads as efficiently as they can using statistical techniques. Several subsequent articles have supported this argument, showing that statistical decision-making techniques consistently outperform, or at least perform equally as well as clinical decision making (e.g., Ægisdóttir et al., 2006. Dawes, Faust, & Meehl, 1989; Grove & Meehl, 1996; Grove, Zald, Lebow, Snitz, & Nelson, 2000; Kuncel, Klieger, Connelly, & Ones, 2013). So despite the popularity of individual assessments as a selection tool, the utility of holistic assessment has been contro- versial for at least half a century, and the question remains whether holistic assessment is actually better than traditional statistical techniques.

However, the belief that statistical prediction is superior to clinical predictions is not universal. In response to Meehl’s argu-

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

7INDIVIDUAL ASSESSMENT META-ANALYSIS

ment, Holt (1958) differentiated between two types of clinical predictions, “naïve clinical” and “sophisticated clinical” (p. 4). He posited that naïve clinical prediction is characterized by the use of data that are primarily qualitative with no attempt to use objective criteria in the decision-making process. Predictions are made in an entirely intuitive manner without emphasis on objectivity. Sophis- ticated clinical prediction, on the other hand, includes the use of qualitative data from sources such as interviews, biodata, and projective techniques that are used in addition to objective facts and scores. It differs from the naïve clinical prediction techniques because it includes objectivity, organization, and scientific meth- odology in the planning of the assessment, the gathering of data, and in the analysis. At the same time, though objectivity is em- ployed in the process, the clinician is retained as one of the primary instruments of prediction, with an effort to be as reliable and valid a data-processor as possible. According to Holt (1958), sophisti- cated clinical methods, or those that involve a scientific approach to holistic assessment, might yield even better predictions than statistical techniques.

Building on Holt’s contentions, one way to improve on the scientific approach and the structure of the clinical assessment may be to involve more than one assessor to collectively make predic- tions about each candidate. In the interview literature, it has been suggested that the use of multiple interviewers can contribute to the structure of the assessment by reducing the impact of idiosyn- cratic biases among interviewers (Campion et al., 1988; Hakel, 1982), limiting the number of irrelevant inferences that are not job related (Arvey & Campion, 1982), and increasing accuracy as a result of the range of information and judgment that is obtained from different perspectives (Dipboye, 1992). Research confirms that averaging ratings across multiple assessors improves the reli- ability and validity of unstructured interviews (Schmidt & Zim- merman, 2004). Given this, individual assessments are expected to be most valid when several assessors are used.

Hypothesis 3: Recommendations from individual assessments that use multiple assessors to assess each candidate will be better predictors of performance than those that utilize a single assessor.

Another aspect of structure concerns whether the same assessor, or same panel of assessors, is used across all candidates (Campion et al. 1997; Huffcutt & Woehr, 1999). In individual assessment situations, the integration of information probably differs across assessors, which may lead to predictions that are not necessarily comparable across the job candidates. Ryan and Sackett (1998) noted that disagreement is common among individual assessors. They suggested that this disagreement may be attributed to differ- ences in assessor competence, different ideologies, or differences in how assessors organize and use the same information. If differ- ent candidates are evaluated by different assessors, then these assessor idiosyncrasies may add variability to ratings that is unre- lated to the competencies being assessed, thereby lowering validity (O’Brien & Rothstein, 2011). The use of a single assessor across candidates should reduce error due to assessor differences and therefore may result in higher validity.

Hypothesis 4: Recommendations from individual assessments that employ the same assessor across all candidates will show

higher validity than those that use a variety of assessors across candidates.

Occupation

Individual assessments are most commonly used to fill high- level management positions (Ryan & Sackett, 1987). Given the cost, individual assessments are most likely to be used in hiring situations where the stakes are high, such as upper level manage- ment. The assessment and hiring of senior-level executives have been identified as distinct from other types of positions, due to the constantly changing performance expectations and the latitude each incumbent has to shape the nature of the work (Silzer & Jeanneret, 2011; Thornton et al., 2010). The complexities of as- sessing managerial potential call for a highly flexible approach, such as that provided by individual assessment.

Given the popularity of individual assessment for management positions, it is particularly important to examine whether assess- ment practices are effective in this context. Previous meta-analytic research found that the validity of many predictors depends on the type of job. Cognitive ability tests have been found to be more predictive of performance for jobs with greater complexity (Hunter & Hunter, 1984). In contrast, the validity of situational interviews has been found to be lower for high complexity jobs (Huffcutt, Conway, Roth, & Klehe, 2004). The personality trait of extrover- sion has been found to better predict performance for managerial jobs (Barrick & Mount, 1991). These findings suggest that validity might differ for managerial and nonmanagerial positions, but the direction of this difference is uncertain.

Research Question 1: Does the validity of assessor recommen- dations differ for managerial and nonmanagerial jobs?

Methodological Factors

All of the considerations that have been discussed thus far deal with the individual assessment procedure and how different as- pects of the procedure itself may contribute to the validity of the individual assessment. Several research study design issues can also influence the results of individual assessment validation stud- ies, which are discussed in the following sections.

Source of Recommendation

When designing a study, researchers make numerous choices that may impact the results. In the case of individual assessments, one important issue is the fidelity of the research design, that is, the extent to which the study design reflects individual assessments as they occur in applied settings. An important component of fidelity in the individual assessment literature is the source of the predictor score used to compute the validity coefficient. Given the goal of evaluating the validity of assessor recommendations, the optimal research design would obtain numerical ratings directly from the individuals who conducted the assessment. This is consistent with the standard practice of individual assessment, where the assessor who has access to all of the assessment data as well as personal interaction with the examinee is the same person who writes the assessment report.

A difficulty arises in conducting validation research because some individual assessments do not use numeric ratings, providing

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

8 MORRIS, DAISLEY, WHEELER, AND BOYER

instead a qualitative report of the candidate’s strengths, weak- nesses, and fit with the hiring organization. Validating these as- sessments requires an additional step whereby the assessment report is translated into a numerical rating for the purpose of the validation research. In several studies, one person conducted the assessment and wrote the qualitative assessment report; then a second assessor later reviewed this report and provided the nu- meric ratings needed for the validation study. In such cases, the second individual assessor, whose prediction was used to validate the individual assessment, did not have direct access to the job candidate.

The source of the predictor data are likely to impact the results of the validation study. The individual who conducted the original assessment had the greatest opportunity to develop a deep under- stand of the job candidate through interaction. In contrast, when ratings made by secondary sources from narrative assessment reports, some information is inevitably lost and the validity of ratings may be lower.

Hypothesis 5: Studies in which recommendations were ob- tained directly from the individual who conducted the assess- ment will have higher validity than studies where recommen- dations were made by secondary sources after reading assessment reports.

Another important element of study design is the type of vari- able used to measure the performance criterion. Employee success can be operationalized in a variety of ways. In order to aggregate results across studies, it is extremely important that the criteria included in the meta-analysis are comparable to one another. Many studies used supervisor ratings of performance as the criterion variable. In other situations, administrative data, such as bonuses or organizational advancement, were the only criteria available. Administrative decisions might be influenced by factors other than individual work productivity and therefore reflect a more distal measure of performance. Because these outcomes reflect substan- tially different types of criterion measures (Bommer, Johnson, Rich, Podsakoff, & MacKenzie, 1995), we conducted separate analyses for the validity of individual assessments predicting sub- jective performance ratings and administrative decisions.

Research Question 2: Are recommendations from individual assessments more predictive of subjective performance ratings or administrative decisions?

Method

Selection of Studies

A literature search was conducted to identify published and unpublished criterion-related validity studies of individual assess- ments used for selection purposes. First, a computer search was done of PsycINFO and Dissertation Abstracts through 2011 in order to find all references to individual assessment in employment selection using the following search terms: individual assessment, individualized assessment, individual psychological assessment, clinical assessment, clinical judgment, holistic assessment. Sec- ond, a manual search was conducted that consisted of checking the sources cited in the reference section of literature reviews, articles, and books on the topic of individual assessment (e.g., Highhouse,

2002; Prien et al., 2003; Ryan & Sackett, 1998). Additionally, manual searches of the reference sections of included studies were conducted to identify additional individual assessment studies. We also sought unpublished validation studies through a variety of methods, including contacting consulting firms that conduct as- sessments, contacting researchers who have published on the topic, and posting a request for studies on a listserv for human resource professionals (HRNET). Only English-language research reports were sought.

For the current research, individual assessments were defined as any employee selection procedure that involves (a) multiple as- sessment methods (b) administered to an individual examinee and (c) relying on assessor judgment to integrate the information into an overall evaluation of the candidate’s suitability for a job. Therefore, studies that employed a statistical or mechanical inte- gration of collected data to arrive at a prediction were not included in the current study. In addition, the validity coefficient must have reflected an assessment of an individual, using individual assess- ment techniques as opposed to group assessments. Studies that included activities such as role plays or simulations were included only if they were conducted on an individual basis. For example, Silzer’s (1984) findings were excluded because the holistic assess- ment included group exercises, and no recommendation was made that excluded the results of the group exercise. In contrast, a study conducted by Handyside and Duncan (1954) was retained for the current study because it reported separate validity coefficients based on holistic evaluations of the candidates with and without a group exercise.

It is also important to point out that Ryan and Sackett’s (1987) definition of individual assessment reflects the concept of “one psychologist making an assessment decision for a personnel- related purpose about one individual” (p. 456). However, for this study, the definition of individual assessment was not restricted to include only one psychologist assessor. Several of the studies included panel decisions about individuals as well as multiple assessors in the study design. In addition, the current study in- cluded assessments that were made by both psychologists and non-psychologists. Specifically, for 16 of the samples, the asses- sors were psychologists; in one sample, a non-psychologist was used as an assessor; and in 10 samples, both human resources representatives and psychologists were utilized as assessors (the assessor background was unknown for the remaining samples).

With these criteria, the search yielded 42 research reports, of which 24 were acceptable for inclusion in this analysis. The studies included in the analysis are summarized in Appendix A in the supplemental online material. The remaining 18 studies were ex- cluded for the following reasons: two studies did not involve multiple assessment methods, three studies did not use assessor judgment to integrate information, five studies involved group exercises, and eight studies did not provide a validity coefficient for predicting job performance or enough information to derive a validity coefficient from the data provided. The excluded studies are listed in Appendix B in the supplemental online material.

A total of 39 samples were obtained from the 24 research reports as several articles reported results for more than one sample. Four of the published articles reported results based on two samples, and one study reported results for four separate samples. Of the un- published reports, two provided two samples each, and one re- ported on seven samples. Nine samples were reported in the 1950s,

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

9INDIVIDUAL ASSESSMENT META-ANALYSIS

eight in the 1960s, four in the 1970s, two in the 1900s, and 16 after 2000. Twenty-one samples were from journal articles, two were from book chapters, three were from unpublished doctoral disser- tations, and 13 were from conference articles or unpublished technical reports.

Effect Size

The validity of each study was computed as the correlation between an overall assessor rating and a measure of job perfor- mance. The predictor consisted of an overall rating of job suitabil- ity, fit, or potential, or a composite of ratings on multiple dimen- sions. Two types of outcome measures were included. Subjective ratings consisted of ratings or rankings made by either a supervisor or an administrator at a higher level in the organization than the participant. Administrative decisions included organizational de- cisions such as promotions, salary changes, and receipt of bonuses. We only included measures of change in job level or compensa- tion. We did not use current job level or salary, because both variables can be influenced by the initial job offer and therefore could be contaminated by the assessment rating that was used to make the employment decision.

Separate analyses were conducted for supervisor ratings and administrative data. Seven studies included both supervisor ratings and administrative outcomes and were included in both analyses, whereas 30 included supervisor ratings only and two included administrative data only.

Coding of Variables

Coding of each study was conducted independently by two trained coders using a structured coding sheet. Whenever there was a discrepancy in coding, the discrepancy was discussed until consensus was reached.

Content of assessments. The assessment batteries were coded for inclusion of several commonly used assessment practices, including cognitive ability tests, personality measures, biodata, and an interview. It should be noted that some research reports did not include a complete description of the assessment battery, choosing instead to report only examples of the assessment methods used. An assessment method was treated as absent if it was not men- tioned in the research report.

Source of recommendation. The studies were categorized into two types of research design, based on whether the individual who conducted the assessment was the same person who provided the ratings used in the validation study. The study was classified as an assessor recommendation when there was direct interaction between the participant and the assessor who made the prediction that was used in the validation study. These assessors had primary access to all standardized test data and interview information upon which a judgment was made for each candidate. There were some studies in which a single assessor had primary access to all information, having been the only interviewer, but a panel decision was made by several psychologists who did not have direct contact with the participant. A decision was made in these cases to code such studies as assessor recommendations, as long as at least one person on the assessment panel had engaged in direct contact with the individual in question.

Secondary source recommendations included studies in which numerical ratings were made based on a written report provided to

the hiring organization about each individual candidate. In these studies, the assessment data were collected, and an assessment report was written by the original assessor; however, because the report did not include a numerical rating of the candidate, another person read the report and assigned the predictor score that was used for the validation. As such, the validation study assessor had only secondary access to the assessment report and did not have any direct contact with study participants. Five samples were identified as secondary studies based on the information in the original research report. Four additional samples reported in Miner (1970) were classified as secondary based on the description of these studies in a review by Prien et al. (2003).

It should be noted that the distinction between assessor recom- mendations and secondary sources was relevant only for assess- ment practices that involved some form of interaction between the assessor and the assessee that went beyond the test battery. Studies of assessment practices that did not include an interview were excluded from the analysis of source of recommendation.

The initial coding of this variable resulted in low interrater agreement (52% agreement, � � .27) due to discrepancies in the coding of assessor panels and assessments with no interaction between the assessor and the candidate. The adoption of the decision rules for these situations, as noted earlier, resolved the discrepancies.

Standardization of assessment battery. This variable indi- cated whether the same set of assessment methods was adminis- tered to all candidates (� � .80; 65% agreement). Differences in procedure included the use of different tests for the same construct or assessment of different constructs across candidates. For exam- ple, in Hilton et al. (1955), slightly more than half of the candidates were given a test of practical judgment. In some instances, the author mentioned that the procedure varied without providing specific information about the nature of the differences (e.g., DeNelsky & McKee, 1969).

Single vs. multiple assessors. This variable reflected whether a single assessor was given access to the candidates’ assessment battery or multiple assessors were used to assess each of the candidates. In cases where multiple assessors were used, more than one assessor was given access to the results of the applicant’s inventories and interview performance, and a prediction was made for each candidate using all of the assessors’ input. In secondary source designs, the number of assessors was determined based on the original assessors who had direct contact with candidates.

Agreement on the coding of this variable was fairly low (� � .61; 72% agreement). The discrepancies were primarily a side- effect of studies using a secondary source for recommendations. Differences in coding studies as assessor versus secondary source designs led to differences in identifying the primary assessors, which in turn affected judgments about the number of primary assessors.

Same vs. different assessors across candidates. Whether each of the candidates within a study was assessed by the same assessor or a variety of assessors were used to assess the candi- dates within a study was also coded. In secondary research de- signs, this was determined based on the original assessors who had direct contact with candidates. Interrater agreement on this vari- able was fairly low (� � .47; 68% agreement). Many of the disagreements were a side-effect of problems with classifying studies into assessor versus secondary source designs and were

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

10 MORRIS, DAISLEY, WHEELER, AND BOYER

resolved once consensus was reached on who was the primary assessor.

Occupation. The final variable of interest was the occupation of the position to be filled (� � .89; percent agreement � 92%). Occupations were classified into managerial and nonmanagerial work. Most of the managerial samples were first-line supervisors in manufacturing firms and other managerial positions included department managers and general division managers. Nonmana- gerial occupations included psychiatrists, engineers, consultants, police, fire fighters, sales personnel, students in graduate pro- grams, pilots, and CIA agents. Two samples were composed of multiple occupations and were classified as nonmanagerial.

Criterion type. The performance criterion was coded as either a subjective rating or an administrative decision. Because some studies included both types of outcomes, we coded the presence of each type of criterion measure separately. Subjective ratings con- sisted of ratings or rankings made by either a supervisor or an administrator at a higher level in the organization than the partic- ipant (� � .57; percentage of agreement � 84%). Administrative decisions included organizational decisions such as promotions, salary changes, and receipt of bonuses (� � .92; percentage of agreement � 96%).

Purpose of performance ratings. Subjective performance ratings were classified into those that were developed for the validation study (research based) and those that were conducted for administrative purposes.

Meta-Analytic Procedure

A random-effects meta-analysis was conducted using the pro- cedures described in Hedges and Olkin (1985). This approach weights each effect size by the reciprocal of its variance, which is the sum of the sampling variance and the between study variance. For the estimation of the sampling variance, the population corre- lation was set at a constant equal to the sample size weighted average correlation (Hunter & Schmidt, 2004). A restricted maximum-likelihood estimator was used for the between-study variance (Viechtbauer, 2005). The meta-analysis was conducted in R using the metafor package (Viechtbauer, 2010). Heterogeneity of validity, that is, the extent to which results varied across studies more than would be expected due to sampling error, was evaluated for statistical significance using the Q-within test and for practical significance using the I2 index (Higgins & Thompson, 2002).

Moderator tests were conducted by comparing the average validity for subgroups, using the Q-between test (Hedges & Olkin, 1985).

Validity was estimated using uncorrected correlation coeffi- cients, and then the average correlation was corrected for attenu- ation due to criterion reliability. Because reliability information was available for only eight of the samples (total N � 331), the corrections were conducted at the aggregate level, using the aver- age interrater reliability of the criterion measures (M � .82, SD � .06). Adjustments to the variance of validities due to criterion unreliability were found to be trivial and are not reported here. Few studies provided enough information for us to investigate the effect of range restriction, and therefore correction for range restriction was not applied.

Supplemental analyses were conducted to assess the sensitivity of our results to publication bias. Publication bias occurs when the sample of studies available for the analysis differs systematically from the full body of research on a topic. For example, validation studies that produce nonsignificant findings may be less likely to be published and organizations that conduct validation studies on their own assessments may be less likely to share unfavorable technical reports. We conducted two types of publication bias analyses. First, we used a visual inspection of funnel plots to identify patterns of asymmetry consistent with publication bias (Sterne, Becker & Egger, 2005). Second, we used a trim-and-fill analysis to estimate what the average validity would be if the hypothetical missing studies were included in the meta-analysis (Duval & Tweedie, 2000).

The number of studies included in many of the subgroup anal- yses was fairly small, increasing the chance that a single extreme study could substantially influence the results. Therefore, outlier analyses were conducted on both the overall analysis and each of the moderator tests. A study was considered an influential data point if the absolute value of the Studentized residual was greater than 2.5 and either the Cook’s D or the standardized DFBETA statistic was greater than 1.0 (Viechtbauer & Cheung, 2010).

Results

Overall Meta-Analysis

Table 1 shows the results of the meta-analyses conducted across all available studies. Separate analyses were conducted for subjec-

Table 1 Meta-Analysis Results for Individual Assessments Predicting Subjective Ratings and Administrative Decisions as Criteria

Criterion N K

Mean r

SDr SD� I 2 Q w/in

T&F

N-weighted Random effect (SE) Corrected �K r

Subjective criteria 3922 37 .24 .27 (.03) .30 .16 .12 62 91.3�� 9a .21 Administrative decisions 600 9 .19 .17 (.06) — .17 .11 45 14.5 —b —

Note. N � combined sample size; K � number of samples; Mean r: N-weighted � average sample-size weighted validity, Random effect � average validity under from the random effects meta-analysis, Corrected � average validity corrected for criterion unreliability; SDr � standard deviation of the observed validities; SD� � standard deviation of the validity corrected for sampling error; I2 � percentage of variance beyond sampling error; Q w/in � chi-square test for homogeneity of observed validities of studies; T&F �K � number of effect sizes imputed by trim-and-fill analysis; T&F r � trim-and-fill estimate of average correlation. a Imputed on the left side of the distribution. b Trim-and-fill analysis was not conducted for administrative decisions due to the small number of studies. �� p � .01.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

11INDIVIDUAL ASSESSMENT META-ANALYSIS

tive and administrative performance outcomes, with several sam- ples appearing in both analyses because they reported both types of criteria.

The first meta-analysis estimated the validity of all individual assessments for predicting ratings of job performance. The esti- mate of the mean-corrected validity was .30 (SD� � .12) across 37 studies and a total sample size of 3,922. A 95% confidence interval of [.24, .35] indicates that we can be quite confident the true average validity of individual assessments exceeds zero. The Q-within test was significant, and the I2 statistic of 62% indicated most of the observed variability in validity was due to true differ- ence and not sampling error. The 95% credibility interval of [.07, .50] suggests that although individual assessments were predictive of subjective performance ratings in most studies, the strength of the relationship varied considerably. Thus, moderator analyses are in order.

The second meta-analysis was conducted to assess the mean validity of individual assessments predicting administrative deci- sions, such as change in job level and salary. This analysis in- cluded 9 correlations with a total sample size of 600. Because reliability information was not available for the administrative decisions, no correction for measurement error was applied. The mean validity for individual assessments was estimated at .17. The 95% confidence interval of [.06, .28] indicates that we can be confident the true average validity of individual assessments pre- dicting administrative decisions is positive. The Q-within test indicated no significant variance between studies. Outlier analyses identified a single influential study (Gaudet, 1957), and when this study was removed, the validity dropped to .13. Due to the lack of significant variance across studies, as well as the limited number of studies, moderator analyses were not conducted for administrative criteria.

Moderator Analyses for Subjective Criteria

As can be seen in Table 2, the moderator variables were some- what intercorrelated. Individual assessments that reported using cognitive ability tests were also likely to report use of a personality scale. Assessments practices that had a standardized battery were also more likely to use multiple assessors for each candidate. Assessments that used the same assessor across all candidates were less likely to report using an interview. Assessments used for

managerial candidates were more likely to use the same procedure across candidates and were less likely to use a secondary source for the recommendation. The negative correlations between pub- lication type and assessment tools reflect the fact that unpublished studies were fairly consistent in using all three assessment tools (cognitive ability, personality, and interview), whereas published studies were more variable in type of tools used. Unpublished studies were also more likely to use the same battery across candidates and multiple assessors and were less likely to use the same assessor for all candidates.

The results of the moderator analyses are reported in Table 3. The moderator variables can be organized into four categories: the content of the assessment, the degree of structure, occupation, and methodological factors.

Assessment content. An initial descriptive analysis was con- ducted to characterize the types of assessment methods used in the individual assessments. For this analysis, an assessment tool was coded as being present if it was mentioned in the research report and absent if it was not mentioned. However, because many of the reports provided only a partial description of the assessment bat- tery, the results reported here may underestimate the use of these methods. Most studies reported the use of a measure of cognitive ability (84%) and to a lesser extent a personality scale (68%). Use of personal history or biodata was reported less often (22%). Most assessment practices included an interview (78%).

The first four moderator analyses examined whether the validity differed depending on whether a specific assessment tool was used in the assessment process. Consistent with Hypothesis 1a, individ- ual assessments that included a cognitive ability test had higher corrected validity (.32) than those without a cognitive ability test (.14), Q(1, K � 37) � 4.3, p � .05. It should be noted, however, that the sample of studies with no cognitive ability test was quite limited (six studies, N � 385). Further, an outlier analysis identi- fied Russell (2001) as an influential data point. This study had a substantially higher validity (.48) than other studies without a cognitive ability test. Removing this study produced a lower va- lidity for assessments without cognitive ability (.05) and did not change the significance of the moderator test.

Hypotheses 1b–1d were not supported. No significant differ- ences were found for individual assessments that included person- ality or biodata. There was an interesting trend such that individual

Table 2 Correlations Among Moderator Variables

Moderator 1 2 3 4 5 6 7 8 9 10 11

1. Battery included cognitive ability 1.00 2. Battery included personality .64�� 1.00 3. Battery included biodata �.13 �.20 1.00 4. Assessment included interview .13 .20 .28 1.00 5. Standardized assessment battery �.12 �.09 .25 �.23 1.00 6. Multiple assessors per candidate �.10 �.21 �.02 .28 .37� 1.00 7. Same assessor across candidates �.11 �.18 .17 �.44� .06 �.05 1.00 8. Managerial occupation .23 .13 .17 .10 .50�� .29 .09 1.00 9. Assessor was source of recommendation .46�� .18 .32 .31 .28 .19 .27 .35� 1.00

10. Performance rating collected for research .38� .08 �.06 �.10 �.26 �.44� .10 �.20 .00 1.00 11. Journal publication �.41� �.52�� �.04 �.35� �.48�� �.35� .45�� �.32 �.31 .13 1.00

Note. K � 26 –37. � p � .05. �� p � .01.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

12 MORRIS, DAISLEY, WHEELER, AND BOYER

assessments that were based on the assessors’ interpretation of a test battery with no interview showed higher validity (corrected r � .42, k � 8) than those that included an interview (corrected r � .27, k � 29); however, the difference was not statistically significant, Q (1, K � 37) � 2.0, ns. An outlier analysis identified two influential studies (Miner, 1970, Study 1 and Study 4). These two samples showed substantially lower validity than other studies with no interview (–.04 and –.14 for Study 1 and Study 4, respec- tively). Removing these two studies increased the mean validity for noninterview studies to .64, and yielded a significant moderator test. However, it should be noted that the number of studies without an interview was quite small (eight studies, N � 304), and therefore this trend should be interpreted with caution.

Degree of structure. Three moderator analyses examined whether more structured individual assessment practices were associated with higher validity. Hypothesis 2 predicted that assessment protocols that administered the same assessment

battery to all candidates would produce higher validity than those for which the assessment content differed across candi- dates. The nonsignificant moderator test indicated that the validities for these moderator groups were not different from one another, Q(1, K � 32) � 0.0, ns. As such, there was no support for Hypothesis 2.

Hypothesis 3 predicted that practices with multiple assessors of each candidate would show higher validity than those with a single assessor per candidate. The difference was not significant, Q(1, K � 31) � 0.9, ns. Thus, Hypothesis 3 was not supported.

Hypothesis 4 predicted that use of the same assessor across all candidates would lead to higher validity than use of different assessors for different candidates. A significant moderator test indicated support for this hypothesis, Q(1, K � 32) � 4.4, p � .05. The mean-corrected validity for studies using the same assessors across all candidates was .44, which was substantially higher than the validity of .27 for the studies that used different assessors across candidates.

Table 3 Moderators of Validity for Subjective Performance Ratings

Moderator N K

Mean r

SDr SD� I 2

Q T&F

N-weighted Random effect (SE) Corrected Mod W/in �K r

Battery included cognitive ability 4.3�

Yes 3537 31 .35 .29 (.03) .32 .15 .10 57 69.4�� 13 .21 No 385 6 .16 .13 (.09) .14 .23 .19 71 18.8�� — —

Battery included personality 0.1 Yes 2981 25 .23 .25 (.02) .28 .12 .07 39 38.3� 10 .20 No 941 12 .29 .28 (.08) .31 .26 .23 81 50.3�� 0 —

Battery Included biodata 0.3 Yes 1122 8 .19 .26 (.07) .28 .19 .16 79 29.2�� — — No 2800 29 .26 .28 (.03) .30 .16 .10 52 57.2�� 7 .23

Assessment included interview 2.0 Yes 3618 29 .23 .25 (.02) .27 .12 .07 42 48.9�� 9 .20 No 304 8 .37 .38 (.12) .42 .34 .29 76 36.6�� — —

Standardization of assessment battery 0.0 Same procedure 3134 26 .25 .30 (.03) .33 .17 .13 69 74.7 10 .21 Different procedure 495 6 .28 .28 (.04) .31 .10 .00 0 2.3 — —

Single vs. multiple assessors 0.9 Single assessor 1570 15 .20 .24 (.04) .27 .14 .09 46 26.1� 6 .18 Multiple assessors 1998 16 .26 .28 (.03) .31 .11 .06 35 24.5 5 .23

Same vs. different assessors across candidates 4.4�

Same assessor 421 9 .38 .40 (.07) .44 .21 .14 47 14.4 — — Different assessors 2948 23 .24 .24 (.03) .27 .15 .12 67 61.7�� 0 —

Occupation 5.6�

Managers 2346 22 .28 .32 (.03) .35 .15 .11 58 48.1�� 8 .24 Nonmanagers 1576 15 .18 .19 (.04) .21 .15 .10 53 30.8�� 1 .19

Source of recommendation 1.3 Assessor 2910 23 .23 .25 (.03) .28 .12 .08 47 43.6�� 7 .20 Secondary source 750 8 .19 .18 (.06) .20 .17 .13 59 15.7� — —

Purpose of performance rating 0.3 Research 2317 18 .22 .27 (.04) .29 .16 .12 68 50.3�� 1 .26 Administrative 1436 14 .24 .23 (.04) .26 .15 .11 56 28.7�� 0 —

Publication type 1.1 Journal articles 1383 20 .29 .30 (.05) .33 .21 .17 68 55.7�� 0 -- Other sources 2539 17 .21 .23 (.03) .26 .11 .07 44 28.7� 6 .19

Note. All trim-and-fill estimates were imputed on the left side of the distribution; Trim-and-fill analyses were not conducted for subgroups with fewer than 10 studies. N � combined sample size; K � number of samples; N-weighted � average sample-size weighted validity; random effect � average validity from the random effects meta-analysis; corrected � average validity corrected for criterion unreliability; SDr � standard deviation of the observed validities; SD� � standard deviation of the validity corrected for sampling error; I

2 � percentage of variance beyond sampling error; Q mod � chi-square moderator test; Q w/in � chi-square test for homogeneity of observed validities; T&F �K � number of effect sizes imputed by trim-and-fill analysis; T&F r � trim and fill estimate of average correlation. � p � .05. �� p � .01.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

13INDIVIDUAL ASSESSMENT META-ANALYSIS

Occupation. Another possible moderator of individual assess- ment validity involves the occupation of the study participants, which we classified into managerial and nonmanagerial occupa- tions. The significant moderator test supports the presence of differences across occupational groups, Q(1, K � 37) � 5.6, p � .05. For managers, the mean corrected validity was .35, whereas for nonmanagers, the mean corrected validity was .21. Consider- able variability remained within each of the occupational sub- groups (SD� was about .10).

These findings clearly indicate that individual assessments are most effectively used to predict managerial performance. The results also show a positive relationship with performance ratings for nonmanagerial jobs, only to a lesser degree than managerial jobs.

Methodological Factors

We investigated two methodological features of the validation studies as moderators of validity. Hypothesis 5 suggested that validity would be higher in designs in which the recommendation came directly from the assessor, rather than from a secondary source who provided a rating after reading a narrative assessment report. The moderator test was not significant, Q(1, K � 31) � 1.3, ns, and validities were quite similar for the two types of analysis. Hypothesis 5 was not supported.

An additional analysis was conducted to compare validities for performance ratings collected for research purposes to those based on administrative performance ratings. We did not find a signifi- cant difference between administrative and research-based perfor- mance ratings, Q(1, K � 32) � 0.3, ns.

Meta-Regression Analysis

The interpretation of the moderator analyses is potentially complicated by the correlations among the moderator variables. However, none of the significant moderators were found to be significantly correlated with one another. Nevertheless, a meta- regression analysis was conducted including two of the moderators simultaneously (occupation and use of the same assessor across candidates). The third significant moderator, whether the battery included a cognitive ability assessment, was excluded from this analysis due to the small number of studies with no cognitive ability test. In the meta-regression analysis, both moderators showed an effect similar to that found in the separate moderator analyses. Validities were .12 higher for managerial jobs than for nonmanagerial jobs, p � .05. Similarly, individual assessments that used the same assessor across candidates had validities .14 higher than those with different assessors, p � .05.

Publication Bias Analysis

Publication bias occurs when the sample of studies available for a meta-analysis is unrepresentative of the full body of research on the topic. The editorial practices of journals can make it less likely for studies with nonsignificant findings to be published. Similarly, unpublished research with nonsignificant findings may be less accessible because researchers may be less motivated to write up or share unfavorable results. Both processes can tend to produce a distribution of effect sizes where studies that have lower (closer to

zero) validity and smaller sample size are less likely to be in- cluded, leading to a potential upward bias in the mean effect size.

Several methods can be used to examine the presence of pub- lication bias. First, we conducted a moderator analysis comparing the results from journal articles to those from other sources (see Table 3). The results do not suggest a difference between pub- lished and unpublished studies, Q(1, K � 37) � 1.1, ns. Validities from journal articles (corrected r � .33) were similar to those from unpublished sources (corrected r � .26). However, the comparison of published to unpublished findings is a limited approach to detect publication bias, because it fails to account for the degree of publication bias within each subgroup (Kepes, Banks, McDaniel & Whetzel, 2012).

Another method to detect publication bias is through visual inspection of funnel plot asymmetry (Sterne et al., 2005). A funnel plot graphs the uncorrected validity against precision (the inverse of the standard error). If there is no publication bias, one would expect the funnel plot to be symmetrical, with a wide range of validities at the bottom and narrowing toward the top. To the extent that publication bias causes exclusion of nonsignificant results, the funnel plot will be asymmetric, with missing validities in the lower left part of the graph. The funnel plot for the subjec- tive criteria analysis is consistent with this pattern (see Figure 1), suggesting that publication bias is a potential concern with this data.

In a third approach to examining publication bias, the trim-and- fill method (Duval & Tweedie, 2000), researchers impute addi- tional validities in order to make the funnel plot symmetric and then examine the impact of these studies on the mean validity. The difference between the original validity and the trim-and-fill esti- mate indicates the sensitivity of the results to publication bias.

Consistent with the funnel plot, the trim-and-fill analysis im- puted nine studies on the left side of the distribution of effect sizes for subjective performance ratings. The results indicated that the original uncorrected validity estimate of .27 may be inflated by as

Observed Outcome

S ta

nd ar

d E

rr or

0. 31

4 0.

23 5

0. 15

7 0.

07 8

0. 00

0

-0.50 0.00 0.50 1.00

Figure 1. Funnel plot for validities predicting subjective criteria. The x-axis indicates the correlation coefficient.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

14 MORRIS, DAISLEY, WHEELER, AND BOYER

much as 29% in comparison to the trim-and-fill adjusted validity estimate of .21. Still, even after correction for publication bias, a positive correlation was found between assessor recommendations and subjective performance ratings, albeit a weaker relationship than indicated in the initial analysis.

Sterne et al. (2005) noted that trim-and-fill analysis may be inaccurate when there is substantial between study variability. When moderators are present, it is better to conduct the publication bias analysis on moderator subgroups. Following the recommen- dation of Kepes, Banks, McDaniel and Whetzel (2012), we report trim-and-fill analyses only for moderator subgroups with at least 10 studies.

Many of the subgroup analyses showed asymmetry consistent with publication bias. As shown in Table 3, the trim-and-fill analysis for most of the moderator subgroups imputed validities on the left, and in several cases the number of imputed studies was substantial relative to the number of actual studies.

In many subgroups, the trim-and-fill imputed validity was over 20% smaller than the original estimate. For example, the trim-and- fill estimate for managerial occupations was .24, which is substan- tially lower than the original uncorrected validity of .32. None of these reductions was large enough to change the conclusion about the usefulness of the individual assessments. The trim-and-fill estimates still represented useful levels of validity, and the pattern of differences across moderator subgroups was similar. Neverthe- less, the results suggest that many of the estimates from this analysis may be slightly inflated.

Discussion

Despite the popularity of individual assessment for employee selection, it has received much less scrutiny in the research liter- ature than other selection approaches. Few published reports of criterion-related validity are available, and much of this research was conducted about half a century ago. The current meta-analysis compiled the limited evidence from the published literature, along with a number of more recent unpublished validation reports. While necessarily limited by the sparse research base, the current study represents the most comprehensive summary to date of the research on individual assessment validity.

Overall, individual assessments were found to have moderate levels of validity for predicting job performance. The average corrected validity for subjective performance ratings was .30, and validity was higher (.35) for managerial jobs. These values repre- sent useful levels of predictive validity and are comparable to many other widely use employee selection methods. Validity was lower for administrative criteria (e.g., pay raise and promotion rates), with an uncorrected validity of .17.

Because individual assessments typically include a battery of tests and an interview, it is useful to compare the resulting validity to what would be expected for the components of an assessment battery if they were interpreted without the aid of assessor judg- ment. The data available for the meta-analysis did not permit an assessment of the component test validity in these particular stud- ies; however, the existing literature provides considerable infor- mation on typical validities for these methods.

Most of the individual assessment batteries we examined in- cluded a cognitive ability test and an interview, and many also included a personality scale and a biodata measure. Based on a

summary of the available meta-analytic evidence, Bobko, Roth, and Potosky (1999) reported average validities of .51 for cognitive ability tests, .48 for structured interviews, .22 for conscientious- ness, and .32 for biodata. An optimally weighted composite of these four predictors would be expected to have a validity of .63 (De Corte, Lievens, & Sackett, 2008).

When placed in the context of the validity of the components, the support for individual assessment was only modest. That is, although there is evidence for the predictive validity of individual assessments, these validities do not exceed what could be obtained from use of a cognitive ability test or structured interview alone, and the overall estimate of .30 is similar to what could be obtained from a self-report biodata measure.

A defining feature of the individual assessment is the expert assessor who administers, interprets, and integrates the results of the battery. The literature on individual assessment has pointed to several potential benefits of the assessment process that go beyond the content of the assessment battery. First, through personal interaction with the candidate, an individual assessor may be able to obtain information not available through standardized tests. The assessor is able to form and test hypotheses and probe for addi- tional information to clarify discrepancies (Silzer & Jeanneret, 2011). Further, individual assessments often include lengthy and intensive interviews, which may reduce the impact of impression management strategies (Tsai, Chen & Chiu, 2005).

Our results suggest that the benefits derived from the assessor’s interaction with the candidate may be limited. Individual assess- ments that included an interview, providing an opportunity for the assessor to interact personally with the candidate, did not show higher validity than individual assessments based only on a battery of tests. In fact, individual assessments conducted without an interview showed slightly higher validity than those with an inter- view, although this difference was not statistically significant. Although intriguing, these results should be interpreted with cau- tion, given the limited size of the research base for assessments where there was no interview (i.e., there were only eight studies with a total sample size of 304). Nevertheless, these results do not support the idea that assessors add unique insights to the evalua- tion of candidates by observing subtle behavioral cues or through the application of hypothesis testing during the interview (Silzer & Jeanneret, 2011).

The ability of expert assessors to interpret complex patterns of behavior and apply configural decision rules is at the same time a hallmark of individual assessment (Prien et al., 2003) and a source of considerable debate (Highhouse, 2008). According to Silzer and Jeanneret (2011), effective assessors are able to interpret individ- ual data points in the context of broader patterns of behavior or psychological constructs. For example, a deeper understanding of the competencies developed through work experience can be ob- tained by looking at how work experiences build on each other, rather than in isolation (Dragoni, Oh, Vankatwyk, & Tesluk, 2011) Further, the assessor can take situational factors and context into account and consider so-called “broken leg” cues when interpret- ing specific pieces of information (McPhail & Jeanneret, 2011).

Unfortunately, there is a lack of evidence that experts are able to reliably and validly apply configural rules and form accurate judgments based on complex patterns of behavior. Contrary to the idea that assessors provide more sophisticated interpretation of the data, research suggests that expert judgments tend to focus on only

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

15INDIVIDUAL ASSESSMENT META-ANALYSIS

a few salient cues (Hastie & Dawes, 2001). Decades of research has consistently shown that trained experts cannot outperform simple linear models in prediction tasks and may actually be less accurate than a mechanical combination of the available informa- tion (Ægisdóttir et al., 2006; Grove & Meehl, 1996). Although the incremental validity of assessor judgments was not tested in the current study, our results are consistent with the literature on statistical versus mechanical prediction, in that the validity of assessor recommendations was not greater that would be expected from the component tests.

Even if assessors do not add to predictive validity, individual assessments may still be useful in situations where it would not be practical to develop a mechanical decision process. Developing an optimal linear model requires conducting a large-sample criterion- related validity study to obtain optimal predictor weights. This approach is often infeasible for upper-level management positions, where there may only be one individual in a particular job and only a handful of applicants (Guion, 1998; Hollenbeck, 2009; Ryan & Sackett, 1998). An advantage of individual assessment is the ability of the assessor to provide recommendations tailored to the organization’s needs and culture, even in situations where criterion-related approaches are not feasible.

Surveys have shown that practitioners of individual assessment vary considerably in their approaches (Ryan & Sackett, 1987, 1992). Our results suggest that these differences may have impor- tant consequences. The validity of assessor ratings was found to vary considerably across studies. A 95% credibility interval sug- gests that the validity for predicting job performance ratings ranged from essentially zero to as high as .50. Similarly, moderator analyses yielded mean corrected validity for subgroups ranging from low (� � .14) to moderate (� � .44). Thus, choices in how assessments are conducted may have led to substantially higher or lower validity.

The selection literature suggests several features of individual psychological assessment that may enhance validity, including the content of the assessments, the standardization of data collection, and the process used to integrate information into the final assess- ment report. In terms of content, a substantial body of research exists on the predictor constructs that are likely to be useful components of the assessment battery. This research consistently shows strong validity for tests of general cognitive ability (Schmidt & Hunter, 1998). Other predictor constructs such as conscientiousness, integrity, and prior experience have more mod- est validity but nevertheless add incremental validity beyond cog- nitive ability (Schmidt & Hunter, 1998). Our results partly support these recommendations, in that higher validity was found for assessments involving a cognitive ability test compared with those that did not measure cognitive ability. The reader should note, however, that data on assessments without a cognitive ability test was quite limited (i.e., six studies with a total sample size of 385), and therefore the comparison should be interpreted with caution.

In contrast, the inclusion of personality or biodata measures was not associated with higher validity. It may be that the multimethod nature of individual assessments lessened the importance of in- cluding specific measures for these constructs. Given the wide- spread use of personality and background information in evaluat- ing candidates, it is unlikely that these constructs were completely absent from any of the assessments. Notably, a survey of individ- ual assessment practice studies revealed that most collect personal

history information (Ryan & Sackett, 1987), and yet only 22% of the studies in the current meta-analysis reported using a personal history form. For situations in which a standardized measure was not administered, the assessor may have gathered information about prior experience through the interview. Thus, the assessors in most situations would have access to information about prior training, education, and work experience, regardless of whether a separate personal history form was administered. Similarly, most assessors likely drew some inferences about personality constructs as part of the interview, even when no standardized personality questionnaire was administered. In both cases, the ability to obtain relevant information from multiple methods may have compen- sated for the absence of specific measures.

Another feature that differentiates among assessment practices is the degree of standardization of the assessment process. The literature on employment interviews has consistently shown greater validity for more structured interviews (Arvey & Campion, 1982; Campion et al., 1997; Harris, 1989; Huffcutt & Arthur, 1994). However, our analysis found mixed results regarding the impact on structure of individual assessment validity.

The only aspect of structure found to impact validity was the use of the same assessor across candidates. Predictions were most valid when the same assessor or assessment panel was used to make predictions about the entire pool of candidates. Although the comparison was significant, it is notable that limited data were available on studies using the same assessor across candidates (i.e., only nine studies with a total sample size of 421).

The use of a common assessor has been identified in the inter- view literature as a way to enhance reliability and validity (Cam- pion et al., 1997; Huffcutt & Woehr, 1999). Having a single assessor across candidates provides a consistent frame of reference and removes differences due to idiosyncratic judgments by differ- ent assessors. Given evidence for low agreement among practitio- ners of individual assessment (Ryan & Sackett, 1989), the use of different assessors for different candidates is likely to add construct-irrelevant variance to suitability judgments (O’Brien & Rothstein, 2011), thereby reducing validity.

At the same time, the use of a single assessor for all candidates is useful only if the individual selected to conduct the assessments is known to provide accurate judgments. Research on employment interviews suggests that some individuals may make more valid judgments than others (Dreher, Ash, & Hancock, 1988), although research on this topic is mixed (Pulakos, Schimitt, Whitney, & Smith, 1996; Van Iddekinge, Sager, Burnfield, & Heffner, 2006). Therefore, the selection of effective assessors is critical for ensuring validity (Silzer & Jeanneret, 2011). Further, when different assessors are to be used for different candidates, training to provide a common frame of reference may enhance validity (Lievens, 2001).

Another strategy to reduce the impact of idiosyncratic rater effects would be to have each candidate evaluated by multiple assessors. This practice has been recommended in both the interview (Campion et al., 1997) and assessment center (International Task Force on Assessment Center Guidelines, 2009) literatures. When data are combined across multiple assessors, the judgment errors and biases associated with any one assessor are minimized, resulting in more reliable recommenda- tions and reducing construct-irrelevant variance.

The benefits of using multiple assessors, however, were not supported by the current meta-analysis. Practices with multiple assessors were not found to be more valid than those based on a

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

16 MORRIS, DAISLEY, WHEELER, AND BOYER

single assessor. Despite the apparent advantages of using multiple assessors, prior research in the context of employment interviews has been mixed. In fact, a meta-analysis by Huffcutt and Woehr (1999) found lower validity for panel interviews than for those with a single interviewer. It may be that use of multiple assessors has negative consequences that counteract the benefits of multiple raters. For example, interacting with multiple assessors may place increased demands on the candidate and increase the stressfulness of the assessment processes (Campion et al., 1997).

The usefulness of having multiple assessors may depend on how information from the multiple sources is combined to obtain the final recommendation. The benefits of multiple perspectives may be limited if a single assessor is responsible for integrating data from the multiple sources. Research on assessment centers has found that both mechanical approaches (e.g., average ratings across assessors) and decisions based on consensus discussion can produce valid recommendations (Pynes & Bernardin, 1992). Ide- ally, a consensus-based decision gives equal weight to all asses- sors; however, in practice the assessors may differ in their influ- ence (Sackett & Wilson, 1982), and this may undermine the benefits of a multiple assessor approach.

Contrary to expectations, standardization of the assessment pro- tocol did not moderate the validity of individual assessments. It was expected that assessment protocols with a standardized battery would be more predictive of performance than those that differed somewhat in the process, but the results showed that the validities of these approaches were quite similar.

Two aspects of the individual assessments may allow reliable and accurate assessments with a less structured process. Most individual assessors form judgments by integrating information from multiple tests, simulations, and an interview. The redundancy inherent in this approach may make the consistent application of any one component less critical. Further, the training and experi- ence of the assessor may allow the assessor to compensate for lack of perfectly consistent sources of information.

As such, the consistency of the assessment battery may be less important than the structure of the judgment and information integra- tion process. If assessors are not using information in a consistent fashion, it does not matter if the input is the same. McPhail and Jeanneret (2011) discussed three important ways in which structure could be built into the judgment process. First, the assessment dimen- sions should be clearly linked to a systematic analysis of the job. Second, there should be a clear mapping of assessment tools to dimension ratings. Third, assessors should use well-defined behav- ioral rating scales for recording observations from simulation exer- cises. The use of rating forms could also be beneficial for recording impressions during interviews (Campion et al., 1997).

In addition to these assessment design features, occupation was also found to moderate validity. Individual assessments were more valid for predicting supervisory and managerial performance than they were for other occupations. This finding is important given the widespread use of individual assessments for managerial and executive selection (Thornton et al., 2010). Selection for high-level managerial positions creates unique challenges due to the small number of candidates considered for a position and the highly individualized nature of job requirements (Hollenbeck, 2009). The flexibility of individual assessments makes them well suited for this context. Although the strong validity results for managerial occupations are supportive of this practice, it is important to note

that the managerial jobs included in our meta-analysis were mostly lower and mid-level. Additional research is needed to determine whether these findings generalize to executive-level positions.

To some extent, the higher validity for managerial occupations may be a result of occupational differences in the validity of the assessment tools used in the assessment. Cognitive ability tests, which were included in most of the individual assessments we studied, tend to show higher validity for high-complexity jobs, such as management. Similarly, the personality trait of extrover- sion tends to show higher validity for managerial jobs (Barrick & Mount, 1991). Consequently, an assessment battery with a cogni- tive ability test and a Big Five personality scale would be expected to perform well for managerial jobs.

Limitations and Future Research

As with all research, this study is not without limitations. The most obvious limitation of this study is the small number of samples available for inclusion. The current meta-analysis in- cluded only 39 studies, which made interpretation of results some- what difficult. This, combined with the heterogeneity of validity across studies, means that the power of moderator tests was low. It is also important to note that over half of the included studies were published before 1980, so the results may reflect outdated individ- ual assessment practices.

A related issue is the considerable evidence of publication bias. It is possible that the results of this study may reflect a biased subset of individual assessment validation studies. While publica- tion bias typically refers to a tendency for published articles to overrepresent significant findings, our data revealed a different pattern. Evidence of publication bias was more pronounced among unpublished rather than published research reports. Although ef- forts were made to identify both published and unpublished vali- dation reports, it is possible that additional unpublished studies were conducted that were never written up or that vendors of individual assessments chose not to share research reports with unfavorable findings (McDaniel, Rothstein, & Whetzel, 2006).

The trim-and-fill analysis (Duval & Tweedie 2000) indicated that publication bias may have inflated the average validity in several of the moderator subgroups. Consequently, the true valid- ity may be slightly lower than indicated in our results. However, even after accounting for this effect, the individual assessment still demonstrated useful levels of validity.

Another obvious limitation of this study concerns the coding of the studies. Though the authors attempted to code the studies for all possible moderators, many of the studies did not provide sufficient information to properly categorize the samples. As a result, many of the studies were placed into “unknown” categories, which made the interpretation of results somewhat ambiguous.

Along the same lines, limited information was provided con- cerning the range restriction information for the included studies. It has been suggested that candidates in individual assessment situations are likely to have been through several screening hurdles prior to the assessment, so psychologists rarely see the entire applicant pool (Ryan & Sackett, 1998). However, because the primary researchers rarely reported range restriction information, we were unable to develop an artifact distribution for range re- striction to correct for this artifact. As a result, it is possible that the

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

17INDIVIDUAL ASSESSMENT META-ANALYSIS

corrected mean validities derived in the current studies are under- estimates of the true validity of individual assessments.

Above anything else, the small number of studies available for inclusion in the current study underscores the need for additional validation research in the area of individual assessment. Though individual assessments have become a major source of income for many practitioners, very few validation studies have been con- ducted on the subject. Although admittedly it is difficult to conduct validation studies for assessments that are typically done with small sample sizes, the literature provides suggestions concerning how practitioners can conceptualize validation studies to address this limitation (Prien et al., 2003; Ryan & Sackett, 1998).

Several questions remain that can be addressed in future studies. While recent literature has given increased attention to individual assessment (McPhail & Jeanneret, 2011; Silzer & Jeanneret, 2011), additional research is needed to identify the practices that are most effective. Although our findings did not strongly support benefits of increased structure, additional research on this topic may be informa- tive. Practices that have been found to enhance assessor judgments in assessment centers (e.g., Lievens, 2001) may prove useful if applied to the design of individual assessments as well. Similarly, future research should examine whether training of assessors can improve the consistency of information interpretation and integration.

Like interviews or assessment centers, individual assessment represents a data collection method that can be used to assess a variety of constructs (Arthur & Villado, 2008). Although assessors commonly form judgments about multiple dimensions, the current study only examined overall recommendations. Future research should explore the specific constructs assessed in individual as- sessment and examine both the construct and predictive validity of more specific assessor judgments.

Conclusion

Overall, the results indicate that individual assessments are useful predictors of job performance, especially for managerial positions. However, the wide range of validities suggests that not all assessment practices are equally valid. Additional research is needed to identify the features of individual assessments that are most effective. We hope that this work will prompt more researchers and practitioners to conduct validation studies on individual assessments and to publish their results. Building the empirical literature on individual assess- ment validity will provide a foundation for improving the effective- ness of this important area of practice.

References

References marked with an asterisk indicate studies included in the meta-analysis.

Aamodt, M. G. (2004). Special issue on using MMPI–2 Scale configura- tions in law enforcement selection: Introduction and meta-analysis. Applied H. R. M. Research, 9, 41–52.

Ægisdóttir, S., White, M. J., Spengler, P. M., Maugherman, A. S., Ander- son, L. A., Cook, R. S., . . . Rush, J. D. (2006). The meta-analysis of Clinical Judgment Project: Fifty-six years of accumulated research on clinical versus statistical prediction. The Counseling Psychologist, 34, 341–382. doi:10.1177/0011000005285875

�Albrecht, P. A., Glaser, E. M., & Marks, J. (1964). Validation of a multiple-assessment procedure for managerial personnel. Journal of Applied Psychology, 48, 351–360. doi:10.1037/h0042422

Arthur, W., & Villado, A. J. (2008). The importance of distinguishing between constructs and methods when comparing predictors in person- nel selection research and practice. Journal of Applied Psychology, 93, 435– 442. doi:10.1037/0021-9010.93.2.435

Arvey, R. D., & Campion, J. E. (1982). The employment interview: A summary and review of recent research. Personnel Psychology, 35, 281–322. doi:10.1111/j.1744-6570.1982.tb02197.x

�Barnett, R., & Beatty, A. (2013, April). The validity of assessor judgment in individual psychological assessment. Paper presented at the 28th annual conference of the Society for Industrial and Organizational Psychology, Houston, TX.

Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimen- sion and job performance: A meta-analysis. Personnel Psychology, 44, 1–26. doi:10.1111/j.1744-6570.1991.tb00688.x

Bobko, P., Roth, P. L., & Potosky, D. (1999). Derivation and implications of a meta-analytic matrix incorporating cognitive ability, alternative predictors, and job performance. Personnel Psychology, 52, 561–589. doi:10.1111/j.1744-6570.1999.tb00172.x

Bommer, W. H., Johnson, J. L., Rich, G. A., Podsakoff, P. M., & MacK- enzie, S. B. (1995). On the interchangeability of objective and subjective measures of employee performance: A meta-analysis. Personnel Psy- chology, 48, 587– 605. doi:10.1111/j.1744-6570.1995.tb01772.x

Butcher, J. N. (1994). Psychological assessment of airline pilot applicants with the MMPI-2. Journal of Personality Assessment, 62, 31– 44. doi: 10.1207/s15327752jpa6201_4

�Campbell, J. T., Otis, J. L., Liske, R. E., & Prien, E. P. (1962). Assess- ments of higher level personnel: II. Validity of the over-all assessment process. Personnel Psychology, 15, 63–74. doi:10.1111/j.1744-6570 .1962.tb01847.x

Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50, 655–702. doi:10.1111/j.1744-6570.1997.tb00709.x

Campion, M. A., Pursell, E. D., & Brown, B. K. (1988). Structured interviewing: Raising the psychometric properties of the employment interview. Personnel Psychology, 41, 25– 42. doi:10.1111/j.1744-6570 .1988.tb00630.x

Conway, J. M., Jako, R. A., & Goodman, D. F. (1995). A meta-analysis of interrater and internal consistency reliability of selection interviews. Journal of Applied Psychology, 80, 565–579. doi:10.1037/0021-9010.80 .5.565

Dawes, R. M., Faust, D., & Meehl, P. E. (1989, March 31). Clinical versus actuarial judgment. Science, 243, 1668 –1674. doi:10.1126/science .2648573

De Corte, W., Lievens, F., & Sackett, P. R. (2008). Validity and adverse impact potential of predictor composite formation. International Journal of Selection and Assessment, 16, 183–194. doi:10.1111/j.1468-2389 .2008.00423.x

�DeNelsky, G. Y., & McKee, M. G. (1969). Prediction of job performance from assessment reports: Use of a modified Q-sort technique to expand predictor and criterion variance. Journal of Applied Psychology, 53, 439 – 445. doi:10.1037/h0028654

�Dicken, C., & Black, J. (1965). Predictive validity of psychometric evaluation of supervisors. Journal of Applied Psychology, 49, 34 – 47. doi:10.1037/h0021695

Dipboye, R. L. (1992). Selection interviews: Process perspectives. Cincin- nati, OH: Southwestern.

Dragoni, L., Oh, I.-S., Vankatwyk, P., & Tesluk, P. E. (2011). Developing executive leaders: The relative contribution of cognitive ability, person- ality, and the accumulation of work experience in predicting strategic thinking competency. Personnel Psychology, 64, 829 – 864. doi:10.1111/ j.1744-6570.2011.01229.x

Dreher, G. F., Ash, R. A., & Hancock, P. (1988). The role of the traditional research design in underestimating the validity of the employment

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

18 MORRIS, DAISLEY, WHEELER, AND BOYER

interview. Personnel Psychology, 41, 315–325. doi:10.1111/j.1744-6570 .1988.tb02387.x

�Dunnette, M. D., & Kirchner, W. K. (1958). Validation of psychological tests in industry. Personnel Administration, 21, 20 –27.

Duval, S. J., & Tweedie, R. L. (2000). A nonparametric “trim and fill” method of accounting for publication bias in meta-analysis. Journal of the American Statistical Association, 95, 89 –98.

�Francoeur, K. A., Schnur, A. C., Bell, D., & Kinney, T. B. (2010, April). Using individual assessment to predict executive “promotability”. Paper presented at the 25th annual conference of the Society for Industrial and Organizational Psychology, Atlanta, GA.

�Gaudet, F. J. (1957). A study of psychological tests as instruments for management evaluation. In American Management Association (Ed.), Executive selection, development, and inventory (Personnel Series No. 171, pp. 17–27). New York, NY: American Management Association.

�Goudy, K., & Sowinski, D. (2010, April). Individual assessment valida- tion: Using assessment to identify high performance and potential. Paper presented at the 25th annual conference of the Society for Industrial and Organizational Psychology, Atlanta, GA.

Grove, W. M., & Meehl, P. E. (1996). Comparative efficiency of informal (subjective, impressionistic) and formal (mechanical, algorithmic) pre- diction procedures: The clinical-statistical controversy. Psychology, Public Policy, and Law, 2, 293–323. doi:10.1037/1076-8971.2.2.293

Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12, 19 –30. doi:10.1037/1040-3590.12.1.19

Guion, R. M. (1998). Assessment, measurement, and prediction for per- sonnel decisions. Mahwah, NJ: Erlbaum.

Hakel, M. D. (1982). Employment interviewing. In K. Rowland & G. Ferris (Eds.), Personnel manager (pp. 129 –155). Boston, MA: Allyn & Bacon.

�Handyside, J. D., & Duncan, D. C. (1954). Four years later: A follow-up of an experiment in selecting supervisors. Occupational Psychology, 28, 9 –23.

Harris, M. M. (1989). Reconsidering the employment interview: A review of recent literature and suggestions for future research. Personnel Psy- chology, 42, 691–726. doi:10.1111/j.1744-6570.1989.tb00673.x

Hastie, R., & Dawes, R. M. (2001). Rational choice in an uncertain world: The psychology of judgment and decision making. Thousand Oaks, CA: Sage.

�Hausknecht, J. P., Langevin, A. M., & Schruhl, J. (2011). Working paper: Predictors of executive success. Unpublished manuscript, Cornell Uni- versity, Ithaca, NY.

Hedges, L. V., & Olkin, I. (1985). Statistical methods for meta-analysis. Orlando, FL: Academic Press.

Higgins, J. P., & Thompson, S. G. (2002). Quantifying heterogeneity in meta-analysis. Statistics in Medicine, 21, 1539 –1558. doi:10.1002/sim .1186

Highhouse, S. (2002). Assessing the candidate as a whole: A historical and critical analysis of individual assessment for personnel decision making. Personnel Psychology, 55, 363–396. doi:10.1111/j.1744-6570.2002 .tb00114.x

Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology: Per- spectives on Science and Practice, 1, 333–342. doi:10.1111/j.1754-9434 .2008.00058.x

�Hilton, A. C., Bolin, S. F., Parker, J. W., Taylor, E. K., & Walker, W. B. (1955). The validity of personnel assessments by professional psychol- ogists. Journal of Applied Psychology, 39, 287–293. doi:10.1037/ h0042236

Hollenbeck, G. P. (2009). Executive selection—What’s right and what’s wrong. Industrial and Organizational Psychology: Perspectives on Sci- ence and Practice, 2, 130 –143. doi:10.1111/j.1754-9434.2009.01122.x

�Holt, R. R. (1958). Clinical and statistical prediction: A reformulation and some new data. Journal of Abnormal and Social Psychology, 56, 1–12. doi:10.1037/h0041045

Huffcutt, A. I., & Arthur, W. (1994). Hunter and Hunter (1984) revisited: Interview validity for entry-level jobs. Journal of Applied Psychology, 79, 184 –190. doi:10.1037/0021-9010.79.2.184

Huffcutt, A. I., Conway, J. M., Roth, P. L., & Klehe, U. (2004). The impact of job complexity and study design on situational and behavior descrip- tion interview validity. International Journal of Selection and Assess- ment, 12, 262–273. doi:10.1111/j.0965-075X.2004.280_1.x

Huffcutt, A. I., & Woehr, D. J. (1999). Further analysis of employment interview validity: A quantitative evaluation of interviewer-related struc- turing methods. Journal of Organizational Behavior, 20, 549 –560. doi:10.1002/(SICI)1099-1379(199907)20:4�549::AID-JOB921�3.0 .CO;2-Q

Hunter, J. E., & Hunter, R. F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin, 96, 72–98. doi: 10.1037/0033-2909.96.1.72

Hunter, J. E., & Schmidt, F. L. (2004). Methods of meta-analysis: Cor- recting error and bias in research findings (2nd ed.). Newbury Park, CA: Sage.

�Huse, E. F. (1962). Assessments of higher level personnel: IV. The validity of assessment techniques based on systematically varied infor- mation. Personnel Psychology, 15, 195–205. doi:10.1111/j.1744-6570 .1962.tb01861.x

International Task Force on Assessment Center Guidelines. (2009). Guide- lines and ethical considerations for assessment center operations. Inter- national Journal of Selection and Assessment, 17, 243–253. doi: 10.1111/j.1468-2389.2009.00467.x

Jeanneret, R., & Silzer, R. (1998). An overview of individual psychological assessment. In R. Jeanneret & R. Silzer (Eds.), Individual psychological assessment: Predicting behavior in organizational settings (pp. 3–26). San Francisco, CA: Jossey–Bass.

Kepes, S., Banks, G. C., McDaniel, M. A., & Whetzel, D. L. (2012). Publication bias in the organizational sciences. Organizational Research Methods, 15, 624 – 662. doi:10.1177/1094428112452760

Kinslinger, H. J. (1966). Application of projective techniques in personnel psychology since 1940. Psychological Bulletin, 66, 134 –149. doi: 10.1037/h0023609

�Korn/Ferry International. (2011). [The value of Korn/Ferry leadership assessments]. Unpublished technical report, Korn/Ferry International, Minneapolis, Minnesota.

Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98, 1060 – 1072. doi:10.1037/a0034156

�Kwaske, I. H. (2006). An exploratory, multi-level study to validate indi- vidual psychological assessments for entry-level police officer and fire fighter positions (Unpublished doctoral dissertation). Chicago, IL: Illi- nois Institute of Technology.

�LaGanke, J. S. (2008). Validation of an individual assessment process. (Unpublished doctorial dissertation). Wayne State University, Detroit, MI.

�Lees-Hotton, C., Kinney, T. B., & Kung, M. (2010, April). Using exec- utive assessment in predicting success on the job. Paper presented at the 25th annual conference of the Society for Industrial and Organizational Psychology, Atlanta, GA.

Lievens, F. (2001). Assessor training strategies and their effects on accu- racy, interrater reliability, and discriminant validity. Journal of Applied Psychology, 86, 255–264. doi:10.1037/0021-9010.86.2.255

McDaniel, M. A., Rothstein, H. R., & Whetzel, D. L. (2006). Publication bias: A case study of four test vendors. Personnel Psychology, 59, 927–953. doi:10.1111/j.1744-6570.2006.00059.x

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

19INDIVIDUAL ASSESSMENT META-ANALYSIS

McDaniel, M. A., Whetzel, D. L., Schmidt, F. L., & Mauer, S. D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology, 79, 599 – 616. doi: 10.1037/0021-9010.79.4.599

McPhail, S. M., & Jeanneret, P. R. (2011). Individual psychological assess- ment In N. Schmitt (Ed.), Oxford handbook of personnel assessment and selection (pp. 411–442). New York, NY: Oxford University Press.

Meehl, P. E. (1954). Clinical versus statistical prediction: A theoretical analysis and a review of the evidence. Minneapolis: University of Minnesota Press. doi:10.1037/11281-000

�Meyer, H. H. (1956). An evaluation of a supervisory selection program. Personnel Psychology, 9, 499–513. doi:10.1111/j.1744-6570.1956.tb01082.x

�Miner, J. B. (1970). Psychological evaluations as predictors of consulting success. Personnel Psychology, 23, 393– 405. doi:10.1111/j.1744-6570 .1970.tb01665.x

Mumford, M. D., & Stokes, G. S. (1992). Developmental determinants of individual action: Theory and practice in applying background measures. In M. Dunnette & L. Hough (Eds.), Handbook of industrial and organizational psychology (pp. 61–138). Palo Alto, CA: Consulting Psychologists Press.

O’Brien, J., & Rothstein, M. G. (2011). Leniency: Hidden threat to large- scale, interview-based selection systems. Military Psychology, 23, 601– 615. doi:10.1080/08995605.2011.616791

�Phelan, J. G. (1962). Projective techniques in the selection of management personnel. Journal of Projective Techniques, 26, 102–104. doi:10.1080/ 08853126.1962.10381083

Prein, E. P., Schippman, J. S., & Prien, K. O. (2003). Individual assessment as practiced in industry and consulting. Mahwah, NJ: Erlbaum.

Pulakos, E. D., Schimitt, N., Whitney, D., & Smith, M. (1996). Individual differences in interviewer ratings: The impact of standardization, consensus discussion, and sampling error on the validity of a structured interview. Per- sonnel Psychology, 49, 85–102. doi:10.1111/j.1744-6570.1996.tb01792.x

Pynes, J., & Bernardin, H. J. (1992). Mechanical vs. consensus-derived assessment center ratings: A comparison of job performance validities. Public Personnel Management, 21, 17–28.

Reilly, R. R., & Chao, G. T. (1982). Validity and fairness of some alternative employee selection procedures. Personnel Psychology, 35, 1– 62. doi:10.1111/j.1744-6570.1982.tb02184.x

�Russell, C. J. (1990). Selecting top corporate leaders: An example of biographical information. Journal of Management, 16, 73– 86. doi: 10.1177/014920639001600106

�Russell, C. J. (2001). A longitudinal study of top-level executive perfor- mance. Journal of Applied Psychology, 86, 560 –573. doi:10.1037/0021- 9010.86.4.560

Ryan, A. M., & Sackett, P. R. (1987). A survey of individual assessment practices by I/O psychologists. Personnel Psychology, 40, 455– 488. doi:10.1111/j.1744-6570.1987.tb00610.x

Ryan, A. M., & Sackett, P. R. (1989). Exploratory study of individual assessment practices: Interrater reliability and judgments of assessor effectiveness. Journal of Applied Psychology, 74, 568 –579. doi: 10.1037/0021-9010.74.4.568

Ryan, A. M., & Sackett, P. R. (1992). Relationships between graduate training, professional affiliation, and individual psychological assess- ments for personnel decisions. Personnel Psychology, 45, 363–387. doi:10.1111/j.1744-6570.1992.tb00854.x

Ryan, A. M., & Sackett, P. R. (1998). Individual assessment: The research base. In R. Jeanneret & R. Silzer (Eds.), Individual psychological as- sessment: Predicting behavior in organizational settings (pp. 3–26). San Francisco, CA: Jossey–Bass.

Sackett, P. R., & Wilson, M. A. (1982). Factors affecting the consensus judgment process in managerial assessment centers. Journal of Applied Psychology, 67, 10 –17. doi:10.1037/0021-9010.67.1.10

Schmidt, F., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124, 262–274. doi:10.1037/0033-2909.124.2.262

Schmidt, F. L., & Zimmerman, R. D. (2004). A counterintuitive hypothesis about employment interview validity and some supporting evidence. Journal of Applied Psychology, 89, 553–561. doi:10.1037/0021-9010.89.3.553

Schmitt, N. (1976). Social and situational determinants of interview deci- sions: Implications for the employment interview. Personnel Psychol- ogy, 29, 79 –101. doi:10.1111/j.1744-6570.1976.tb00404.x

Silzer, R. F. (1984). Clinical and statistical prediction in a management assessment center (Unpublished doctoral dissertation). University of Minnesota, Minneapolis, MN.

Silzer, R. J., & Jeanneret, R. (2011). Individual psychological assessment: A practice and science in search of common ground. Industrial and Organizational Psychology: Perspectives on Science and Practice, 4, 270 –296. doi:10.1111/j.1754-9434.2011.01341.x

�Stern, G. G., Stein, M. I., & Bloom, B. S. (1956). Methods in personality assessment. Glencoe, NY: Free Press.

Sterne, J. A., Becker, B. J., & Egger, M. (2005). The funnel plot. In H. R. Rothstein, A. J. Sutton, & M. Borenstein (Eds.), Publication bias in meta-analysis: Prevention, assessment, and adjustments (pp. 75–98). West Sussex, United Kingdom: Wiley.

Stokes, G. S., & Cooper, L. A. (2004). Biodata. In J. C. Thomas & M. Hersen (Eds.), Comprehensive handbook of psychological assessment: Vol. 4. Industrial and organizational assessment (pp. 243–268). Hobo- ken, NJ: Wiley.

Thornton, G. C., Hollenbeck, G. P., & Johnson, S. K. (2010). Selecting leaders: Executives and high potentials. In J. L. Farr & N. T. Tippins (Eds.), Handbook of employee selection (pp. 823–840). New York, NY: Routledge.

Tsai, W.-C., Chen, C.-C., & Chiu, S.-F. (2005). Exploring boundaries of the effects of applicant impression management tactics in job interviews. Journal of Management, 31, 108 –125. doi:10.1177/0149206304271384

Ulrich, L., & Trumbo, D. (1965). The selection interview since 1949. Psychological Bulletin, 63, 100 –116. doi:10.1037/h0021696

Van Iddekinge, C. H., Sager, C. E., Burnfield, J. L., & Heffner, T. S. (2006). The variability of criterion-related validity estimates among interviewers and interview panels. International Journal of Selection and Assessment, 14, 193–205. doi:10.1111/j.1468-2389.2006.00352.x

Viechtbauer, W. (2005). Bias and efficiency of meta-analytic variance estimators in the random-effects model. Journal of Educational and Behavioral Statistics, 30, 261–293. doi:10.3102/10769986030003261

Viechtbauer, W. (2010). Conducting meta-analysis in R with the metafor package. Journal of Statistical Software, 36, 1– 48.

Viechtbauer, W., & Cheung, M. W.-L. (2010). Outlier and influence diagnostics for meta-analysis. Research Synthesis Methods, 1, 112–125. doi:10.1002/jrsm.11

Weiner, I. B. (2003). The assessment process. In J. R. Graham & J. A. Naglieri (Eds.), Handbook of psychology: Vol. 10. Assessment psychol- ogy (pp. 3–25). Hoboken, NJ: Wiley.

Wiesner, W. H., & Cronshaw, S. F. (1988). A meta-analytic investigation of the impact of interview format and degree of structure on the validity of the employment interview. Journal of Occupational Psychology, 61, 275–290. doi:10.1111/j.2044-8325.1988.tb00467.x

Received October 21, 2009 Revision received January 30, 2014

Accepted March 31, 2014 �

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

20 MORRIS, DAISLEY, WHEELER, AND BOYER

  • A Meta-Analysis of the Relationship Between Individual Assessments and Job Performance
    • Reliability and Validity of Assessor Judgments
    • Impact of Assessment Practices on Validity
      • Information Input
      • Information Integration
      • Occupation
    • Methodological Factors
      • Source of Recommendation
    • Method
      • Selection of Studies
      • Effect Size
      • Coding of Variables
        • Content of assessments
        • Source of recommendation
        • Standardization of assessment battery
        • Single vs. multiple assessors
        • Same vs. different assessors across candidates
        • Occupation
        • Criterion type
        • Purpose of performance ratings
      • Meta-Analytic Procedure
    • Results
      • Overall Meta-Analysis
      • Moderator Analyses for Subjective Criteria
        • Assessment content
        • Degree of structure
        • Occupation
      • Methodological Factors
      • Meta-Regression Analysis
      • Publication Bias Analysis
    • Discussion
      • Limitations and Future Research
      • Conclusion
    • References