essay
PERSONNEL PSYCHOLOGY 2015, 68, 899–927
COHEN’S d CORRECTED FOR CASE IV RANGE RESTRICTION: A MORE ACCURATE PROCEDURE FOR EVALUATING SUBGROUP DIFFERENCES IN ORGANIZATIONAL RESEARCH
JOHNSON CHING-HONG LI University of Manitoba
Organizational and staffing researchers are often interested in evaluat- ing whether subgroup differences exist (e.g., between Caucasian and African-American individuals) on predictors of job performance. To investigate subgroup differences, researchers often will collect data from current employees to make inferences about subgroup differ- ences among job applicants. However, the magnitude of subgroup dif- ferences (i.e., Cohen’s d) within incumbent samples may be different (i.e., smaller) than the magnitude of subgroup differences in applicant samples because selection of applicants typically reduces the variance of scores on the predictors (i.e., because lower scoring applicants are not selected). If researchers seek to generalize a d value in an incum- bent sample to the applicant population, they may use Bobko, Roth, and Bobko’s (correcting the effect size of d for range restriction and unrelia- bility, 2001) Case II or III correction. By extension, Hunter, Schmidt, and Le (implications of direct and indirect range restriction for meta-analysis methods and findings, 2006) have proposed a Case IV correction, which is more realistic than Bobko et al.’s approach. Therefore, this paper develops a Case IV correction for d (i.e., dc4). The simulation results showed that the dc4 was generally accurate across 6,000 simulation con- ditions. Moreover, 2 published datasets were reanalyzed to show the influence of the Case IV correction on d. In addition, implications and future directions of the dc4 are discussed.
Research on personnel selection has tended to focus on identifying constructs (e.g., cognitive ability, personality) and methods (e.g., situ- ational judgment tests, employment interviews) that predict applicants’ future job performance or other valued outcomes (e.g., turnover; Hunter & Schmidt, 2004). This emphasis not only improves the efficiency of an organization in seeking the most appropriate applicants but also sup- ports the use of selection procedures for hiring and promotion purposes
I thank Ying Cui, Wai Chan, and the anonymous reviewers for their critiques and comments on earlier drafts.
Correspondence and requests for reprints should be addressed to Johnson Ching-Hong Li, Department of Psychology, University of Manitoba, P517B, Duff Roblin Building, Winnipeg, MB, R3T 2N2, Canada; [email protected].
C© 2014 Wiley Periodicals, Inc. doi: 10.1111/peps.12096
899
900 PERSONNEL PSYCHOLOGY
(Bobko & Roth, 2013). Selection validity is commonly evaluated based on the level of Pearson’s correlation r between a selection predictor (i.e., X) and a performance criterion (i.e., Y). Although the algorithm for r is computationally straightforward, its evaluation is often more complicated than one may believe, given that the range of observations in predictor X is often restricted as a result of a selection process (Hunter, Schmidt, & Le, 2006). This is because applicants who score below a certain cut- off on predictor X (e.g., cognitive ability test) may not have a chance to work in an organization, and therefore, there are no scores for crite- rion Y (e.g., job performance). This scenario is known as direct range restriction (RR), in which selection occurs on the basis of predictor X (Chan & Chan, 2004). In addition, if applicants are selected based on a third variable Z (e.g., interview performance) that is correlated with pre- dictor X (e.g., cognitive ability), the situation is known as indirect RR, which causes range (or variance) restriction in both X and Y (Li, Chan, & Cui, 2011).
No matter whether a selection process is direct or indirect RR, the observed validity of a screening test X in predicting performance Y based on an incumbent sample is often weaker (or smaller) than that based on an applicant sample (Schmidt, Oh, & Le, 2006). Organizational and staffing researchers who are interested in examining the validity of predictors of job performance are expected to be aware of the population that they seek to generalize to (Sackett & Yang, 2000). That is, if they attempt to gener- alize the validity for the incumbent population, they can simply interpret the observed validity obtained in an incumbent sample. On the other hand, if they seek to generalize the validity obtained in an (restricted) incumbent sample to the (unrestricted) applicant population, they can use Thorndike’s (1949) well-known Case II (or III) correction, in which the ratio of the restricted to unrestricted SD of X (or Z) is used to adjust for the downward bias arising from direct (or indirect) RR. Correction for RR has become a widely discussed issue in various disciplines including psychology (e.g., Mendoza & Mumford, 1987; Thorndike, 1949), management (e.g., Burke, Normand, & Doran, 1989; Yang, Sackett, & Nho, 2004), and education (e.g., Andre & Hegland, 1998; Fife, Mendoza, & Terry, 2012).
An Expanded Classification System for RR Scenarios
A common practical problem, however, lies in whether or not the information required for using Thorndike’s Case II or III correction is available. In light of this, Sackett and Yang (2000) proposed an expanded classification system, which conceptualizes a number of additional RR scenarios—(a) whether selection occurs based on X, Y, or Z; (b) whether the unrestricted SD for the relevant variable is known; and (c) whether Z,
JOHNSON CHING-HONG LI 901
if involved in a selection process, is measured or unmeasured—thereby producing a total of 11 conditions (for details, refer to Sackett & Yang, p. 114). Of the 11 conditions, the scenario in which Z is unmeasured and the unrestricted SD of X is known (i.e., 2d) has been the focus in the literature on personnel psychology (Sackett & Yang, 2000). In particular, much attention has been given to Hunter et al.’s (2006) developed Case IV correction (e.g., Banks, Batchelor, & McDaniel, 2010; Christian, Bradley, Wallace, & Burke, 2009; Fife, Mendoza, & Terry, 2013). This is in part due to a more realistic selection or decision-making process assumed in this model and to the fact that fewer parameters are required to adjust for the bias.
Subgroup Differences in Predictors of Job Performance
In addition to evaluating the validity of predictors in predicting job performance, organizational and staffing researchers are often interested in examining subgroup differences (e.g., Caucasian/African American), if any, that exist in these predictors (e.g., Bobko & Roth, 2013; Roth, Van Iddekinge, Huffcutt, Eidson, & Bobko, 2002). Indeed, evaluating subgroup differences is an important legal issue (U.S. Equal Opportunity Employment Commission et al., 1978, section 3) because organizations are expected to develop selection tests that predict job performance and do not adversely impact the selection of protected groups of applicants. Sub- group differences in personnel selection are commonly assessed through Cohen’s (1988) standardized mean difference d, which is defined as the mean difference between two groups of X observations divided by the pooled standard deviation (SD; Sp). That is,
d = ( X 1 − X 2)/Sp (1)
where X 1 and X 2 are the mean scores in Groups 1 and 2, respectively, and Sp is the pooled standard deviation of Groups 1 and 2, that is, Sp = {[(N1 − 1)S21 + (N2 − 1)S22 ]/(N1 + N2 − 2)}1/2, where Ni and Si are the sample size and SD of observations in Group i = 1, 2 respectively. For example, a Caucasian/African-American subgroup difference d of 1.0 in cognitive ability suggests that the Caucasian applicants outscored the African-American applicants, on average, by 1.0 SD. Note that d is functionally related to the point-biserial r, given that both statistics as- sess the association between a nominal variable G (e.g., ethnicity) and interval-scale variable X (e.g., cognitive ability).
In light of the importance of d in evaluating subgroup differences and its relation to the point-biserial r, it is surprising that many studies have ignored the impact of RR on d, as mentioned in studies such as Bobko and
902 PERSONNEL PSYCHOLOGY
Roth (2013) and Roth et al. (2002). For example, a researcher evaluates gender differences on a cognitive ability test from an incumbent sample and attempts to generalize the observed d to the applicant population. The range of the cognitive ability (i.e., Sp) may be reduced by preselection vari- ables such as interview performance, college achievement, and so forth.
To better understand the potential threat of RR on d, Bobko and Roth (2013) conducted a comprehensive review of Caucasian/African- American differences in frequently used predictors of job performance such as cognitive ability tests, situational judgment tests, and biodata. Bobko and Roth found that the majority of ds reported in the literature are based on incumbent samples rather than on applicant samples. For ex- ample, many primary and meta-analytic studies in personnel psychology found that the Caucasian/African-American d for conscientiousness (one of the Big Five personality traits) is close to .00 (Bobko & Roth, 2013); however, the majority of these studies used incumbent samples to obtain this value. An exception can be found in Weekley, Ployhart, and Harold (2004), which examined the d in two large samples of incumbents (N = 2,989) and applicants (N = 7,259) from five retail companies in the United States. Although the Caucasian/African-American d for conscientiousness was –.01 in the incumbent sample, it was –.33 in the applicant sample.
Corrections for d Under Direct and Indirect RR
Correcting estimates of subgroup differences (i.e., d) for RR has not received nearly as much attention as has correcting validity coefficients (i.e., r) for RR. However, Bobko, Roth, and Bobko (2001) developed two bias-correction procedures for the range-restricted d, which are based on Thorndike’s (1949) conventional Case II and III corrections for r. Specifically, direct RR for d (with a grouping variable G and interval- scale variable X) assumes that applicants are selected based on their X scores (e.g., cognitive ability test) top down, which causes RR in X and G. Indirect RR for d assumes that there is a third variable Z (e.g., inter- view performance) that is correlated with variable X, and hence selection that occurs on Z causes RR in both X and G. Given these RR scenarios, Bobko et al. used the mathematical linkage between r and d to develop the Cases II and III corrections for d. Building on Bobko et al.’s recom- mendation, a number of studies have examined the impact of RR on d. For example, Roth et al. (2002) found White/Black ds of .36 and .56 for two forms of a behavioral interview. However, Bobko et al.’s (2001) Cases II and III corrections are not sufficient for researchers to use in practice: For Case II, researchers may know the unrestricted variance of X, but X is not the single variable that determines selection; for Case III, researchers may not know which variable is Z that causes selection, and
JOHNSON CHING-HONG LI 903
hence, its unrestricted variance is unknown. Rather, a more common and realistic scenario is that Z is often an unmeasured construct and that only the unrestricted variance of X is available.
Hence, this paper aims to develop a comparable correction for d based on Hunter et al.’s (2006) Case IV correction for r. The details of the Case IV model for d and the correction procedures are provided in the following sections. Moreover, this paper presents a Monte Carlo study, which evaluates the performance of the new correction approach. In addi- tion, this paper reanalyzes data from two previously published studies to demonstrate how the proposed Case IV bias-corrected procedure can be used in practice and how correcting ds for RR can influence conclusions.
Case III Correction Procedure for d Under Indirect RR
On the basis of Thorndike’s (1949) RR framework, Bobko et al. (2001) developed the Case III correction for d under indirect RR. For example, an organization uses a test Z (e.g., college GPA) to select applicants, and those with a Z score below a certain cutoff are not selected. Later, if the organization examines the standardized mean difference (X) between Caucasian and African-American applicants on a cognitive ability test (G), the observed d based on this incumbent sample is downwardly biased. To adjust for the bias, Bobko et al. substituted the equation between r and d into Thorndike’s correction and developed the Case III correction for d when it is subject to indirect RR. This procedure is (Bobko et al., 2001, Equation 6)1
dc3 = (
dXi√ r X Xi
) M +
( u−2Z − 1
) √
P Q
√ 1 pq
+ d 2 Zi
u2Z + ( u−2Z − 1
)( r X Zi√ r X Xi
√ r Z Zi
)2 [ 1 pq
+ (
dZi√ r Z Zi
)2] (2)(
r X Zi√ r X Xi
√ r Z Zi
)( dZi√ r Z Zi
) − (
dXi√ r X Xi
)2 M 2 − 2
( u−2Z − 1
) M (
r X Zi√ r X Xi
√ r Z Zi
)( dXi√ r X Xi
)( dZi√ r Z Zi
)
1According to Bobko et al. (2001), the mathematical linkage between r and d is r = d/
√ ( P Q)−1 + d 2 , where P is the proportion of participants in Group 1, and Q is
the proportion of participants in Group 2. The Case III correction procedure for Pearson’s correlation is rc3 = [r X Yi + r X Zi rY Zi (u−2Z − 1)]/
√ [1+r 2X Zi
(u−2Z −1)][1+r 2 Y Zi
(u−2Z −1)], where r X Yi , r X Zi , and rY Zi are the restricted (or incumbent) correlations between X–Y, X–Z, and Y–Z, respec- tively, and u Z = sZ /SZ is the ratio of the restricted to unrestricted SD of Z.
904 PERSONNEL PSYCHOLOGY
S
Xt G
X
tSX
XGdtX X
tX Gd
Figure 1: A Conceptual Model of Case IV Indirect Range Restriction (RR) for Cohen’s d Based on Hunter et al. (2006), and Le and Schmidt (2006).
Note. The direction of RR shows the relationship between variables. S = suitability construct that creates RR in Xt; G = group variable; Xt = true score of variable X; ρX t S = correlation between Xt and S; ρX t X = correlation between Xt and X, and its square is the reliability of X, ρX X ; dX t G = Cohen’s d between Xt and G; dX G = Cohen’s d between X and G.
where dc3 is the d corrected for Case III indirect RR and unreliability; r X Zi is the restricted r between X and Z; dX i is the restricted d based on X; dZi is the restricted d based on Z; uZ is the ratio of the restricted to unrestricted SD of Z; p and q (or P and Q) are the proportions of the restricted (or unrestricted) participants in Groups 1 and 2: respectively; r X X i is the restricted reliability of X; and r Z Zi is the restricted reliability of Z.
M = {[
1 + pq (
dZi√ r Z Zi
)2]/[ 1 + pq
( dX i√ r X X i
)2]}1/2
Case IV Correction Procedure for d Under Indirect RR
In addition, Schmidt, On, and Le (2006) mentioned that the Case III correction for indirect RR is seldom used in practice because the parameters (i.e., uZ, r X Zi , dX i , and dZi ) required for the correction are rarely available, and selection is rarely based on a single and observable variable Z. Rather, a group of job applicants is usually selected through a composite of several variables (e.g., GPA, interview), which is called the suitability construct (i.e., S). On the basis of Hunter et al.’s (2006) Case IV model for r under indirect RR, this paper proposes and develops a comparable Case IV model for d under indirect RR. As shown in Figure 1,
JOHNSON CHING-HONG LI 905
this model contains four variables: the S, the true score of X (i.e., Xt), the observed score of X, and the grouping variable G.
First, this model assumes that organizations judge the suitability (i.e., S) of applicants based on a combination of variables, which may be quan- tified (e.g., GPA) or not quantified (e.g., impressions based on applicants’ resumes or interview performance). Applicants are then selected top down based on their S, which causes RR and reduces the variance of Xt. Ac- cording to Hunter et al. (2006), it is unrealistic that an organization can predict the errors of measurement associated with X.
Second, the model contains an arrow from S to Xt, and from Xt to G. Thus, there is an indirect effect of RR from S on G. No other arrow is assumed to connect S and G, meaning that the selection process does not evaluate any characteristic that would directly affect the grouping variable G for factors other than those measured by the construct Xt (Hunter et al., 2006). In other words, all effects of RR that are caused by S to G are fully mediated by Xt. An example of this situation would be if an organization judges applicants’ suitability (S) based on their resumés and performance during an interview, and judgments of suitability are not based on any variables that would directly affect G (e.g., ethnic subgroup differences) other than by the predictor under examination (e.g., incumbent scores on a cognitive ability test). This assumption is generally tenable, given that organizations should select applicants with highest S level regardless of their membership in G (e.g., ethnicity). Hence, there is no direct effect of RR from S to G except through Xt.
When this assumption is met, the model is identical to Bobko et al.’s (2001) Case III. When this assumption is violated, the most accurate cor- rection should be Bobko et al.’s Case III; however, the parameters required for correction in Equation 2 (i.e., uZ, r X Zi , dX i , and dZi ) are rarely avail- able in practice. Alternatively, one may choose to report the uncorrected, Case II or Case IV adjusted ds. For the case of r, simulation results have shown that the best alternative should be the Case IV correction (e.g., Le & Schmidt, 2006; Li et al., 2011; Li, Cui, & Chan, 2013). Hunter et al.’s model did not only draw attention in simulation studies, but it also led to a number of recent meta-analyses in personnel psychology that reported this correction (e.g., Banks et al., 2010; Christian et al., 2009).
In light of this, this study aims to develop a comparable Case IV correction for d. First, if the RR effect of S on G is fully mediated by Xt as assumed in the Case IV model, the ratio of the restricted to unrestricted SD of Xt (i.e., ut) can reflect the level of restriction that exists in S. Thus, one can use ut to adjust for the bias, and it can be estimated using Equation 3
906 PERSONNEL PSYCHOLOGY
from Schmidt et al. (2006)
ut = √
[u2X − (1 − r X X a )]/r X X a (3)
where r X X a is the unrestricted (or applicant) reliability of X, and u X = sX /SX is the ratio of the restricted to the unrestricted SD of X. Note that Equation 3 contains two unrestricted parameters. The first parameter, r X X a , can be estimated from the restricted reliability (i.e., r X X i ) in a restricted (or incumbent) sample through (Li et al., 2011, Equation 5).
r X X a = 1 − u2X (1 − r X X i ) (4)
The second parameter, SX (i.e., unrestricted SD of X), is unknown in a restricted (or incumbent) sample. Researchers can trace the sources of SX from a technical manual or journal article that is based on an unrestricted sample2 (Li et al., 2011).
Second, one can estimate the restricted d corrected for unreliability (i.e., dc) through (Bobko et al., 2001, Equation 8).
dc = dr / √
r X X i (5)
where dr is the restricted d, and r X X i is the restricted (or incumbent) reliability of X.
Third, one can convert dc in Equation 5 to the Case IV adjusted r (i.e., rc4; Schmidt et al., 2006, Equation 4) through
3
rc4 = ⎛ ⎝ dc
ut √
( P Q)−1 + d 2c
⎞ ⎠ /√√√√√
[( 1
u2t
) − 1
]⎛⎝ dc√ ( P Q)−1 + d 2c
⎞ ⎠
2
+ 1 (6)
where P is the proportion of the unrestricted applicants who belong to Group 1, and Q = 1 – P is the proportion of the unrestricted applicants who belong to Group 2.
2Although P and Q in Equation 6 are also estimated from an unrestricted group of applicants, these values are typically observable because organizations often record the proportion of the unselected job applicants in each subgroup.
3According to Schmidt et al. (2006; Equation 4), the correlation corrected for Case IV indirect RR (rc 4) can be estimated based on the correlation r X Yc corrected for unreliability,
rc4 = ( 1ut )r X Yc / √
[( 1 u2t
)−1]r 2X Yc +1 Given that r and d can be linked mathematically by (Bobko
et al., 2001; Equation 1, r = d/ √
( P Q)−1 + d 2, rc4 can be estimated based on the restricted dr, rc4 = ( dc
ut
√ ( P Q)−1 +d 2c
)/ √
[( 1 u2t
) − 1]( dc√ ( P Q)−1 +d 2c
)2 + 1
JOHNSON CHING-HONG LI 907
Finally, one can transform rc4 in Equation 6 back to the d metric in order to produce a final estimate of the Case IV adjusted d; that is,
dc4 = rc4/ √
P Q ( 1 − r 2c4
) (7)
To illustrate how the proposed correction works, consider the following example. A researcher investigates whether gender difference exists in cognitive ability (X) based on a (restricted) incumbent sample and obtains a d of .20, an SD of .50, and a reliability of .80. If the researcher seeks to generalize the restricted d to the applicant population, the researcher may be able to find an estimate of the unrestricted SD of X (e.g., 1.00), such as from a test manual, to obtain the unrestricted re- liability in Equation 4 (i.e., r X X a = 1 − (.50/1.0)2(1 − .80) = .95), which can be plugged into Equation 3 to obtain ut, that is, ut =
√ [(.50/1.0)2 − (1 − .95)]/.95 = .4588. Next, the researcher
can correct the restricted d for unreliability using Equation 5, that is, dc = .20/
√ .80 = .2236. Given r X X a = .95, ut = .4588, and dc = .2236,
and P and Q are assumed to be equal (i.e., 50% of men and women in ap- plicant population), the researcher can convert dc to rc4 via Equation 6, that
is, rc4 = ( .2236 .4588
√ (.50 · .50)−1+.22362
)/ √
[( 1 .45882
) − 1]( .2236√ (.50 · .50)−1+.22362
)2 + 1 = .2368. The last step involves the back transformation from r to d via Equation 7, that is, dc4 = .2368/
√ .50 · .50 · (1 − .23682) = .4875.
Hence, the researcher can obtain a Case IV corrected d (i.e., dc4) of .49 for the applicant population.
In sum, researchers and practitioners can use Equations 3–7 to ob- tain the Case IV adjusted d. The following section presents a Monte Carlo study designed to test the effectiveness of the Case IV corrected d for RR.
Monte Carlo Study
Design
Six factors that could affect the accuracy of the uncorrected and bias- corrected ds were manipulated in this study: the restricted sample size, selection ratio, magnitude of d, proportion of applicants in each subgroup, correlation between S and Xt, and reliability of X.
Factor 1: Restricted sample size (nr; four levels). Restricted sample size nr was manipulated at values of 30, 50, 100, and 300, respectively. These values reflect from small- to medium-large scale companies (Enterprise and Industry Publications, 2005). It is clear that sample size affects the precision of the restricted and adjusted ds. In previous simulations (e.g.,
908 PERSONNEL PSYCHOLOGY
Li et al, 2011; 2013), the minimum restricted sample size ranged from 30 to 50 for an adequate Case IV correction for r, and hence, the sizes of 30 and 50 were evaluated. In addition, larger sample sizes of 100 and 300 were examined.
Factor 2: selection ratio (π ; five levels). Selection ratio is the ratio of the restricted to the total sample size (i.e., π = nr /N ), which was simulated at .10, .30, .50, .70, and .90.
Factor 3: Unrestricted population effect size d (δ; five levels). Five lev- els of the unrestricted effect sizes δ—.20, .50, .80, 1.20, and 1.50—were manipulated. The first three levels indicate small, moderate, and large effect sizes (Cohen, 1988), whereas the last two levels represent two ex- ceptionally large effect sizes. For simplicity, Group 1 observations were simulated to have higher X scores, on average, than Group 2 observa- tions, based on the manipulated level of the unrestricted population effect size δ.
Factor 4: Proportion of unrestricted applicants in Groups 1 and 2 (P and Q; five levels). Five proportions of P—.10, .30, .50, .70, and .90— were evaluated. In many applied situations, the proportions of the unre- stricted applicants tend to be equal between the two groups (e.g., 50% women and 50% men). The remaining four levels are comprehensive enough to cover the unbalanced proportions in practice. Note that Q = 1 – P.
Factor 5: Correlation between S and Xt (ρS X t ; three levels). Three levels of ρS X t , .20, .50, and .80, were evaluated. The first two values refer to the small-to-medium and large effect size levels (Cohen, 1988), whereas the exceptionally large correlation of .80 was simulated in order to evaluate the performance of the corrected d in this adverse selection scenario.
Factor 6: Unrestricted reliability level of X (ρX X a ; four levels). In this study, four distributions of ρX X a with the means—.60, .70, .80, and .90— were evaluated. These values were selected to reflect the reliability levels that range from relatively low to relatively high values. The associated SDs of these ρX X a values were set to .16, .12, .08, and .04, respectively, given that these values have been used in many previous meta-analyses and simulation studies (Le & Schmidt, 2006). The unrestricted sample re- liability r X X a for each replication was generated from a normal distribution with the specified mean and SD.
To summarize, the six factors were combined to produce a design with 4 × 5 × 5 × 5 × 3 × 4 = 6, 000 conditions. Each condition was repli- cated 2,000 times, given that the minimum number of replications for a Monte Carlo study should be 1,000 (Mooney, 1997).
JOHNSON CHING-HONG LI 909
Simulation Procedure
For each condition, 2,000 (i.e., number of replications) random sam- ples of size N1 and N2 were generated for Xt, respectively, based on a nor- mal distribution, where N1 = N P = (nr /π ) P , and N2 = N Q = (nr /π )Q with the values of nr, π , P, and Q determined according to the specified simulation conditions. The mean and SD of the Xt scores in Group 2 (i.e., μ2 and σ 2) were fixed at 0 and 1.0, respectively. Similarly, the mean and SD of the Xt scores in Group 1 were fixed at δ (.2. .5, .8, 1.2, or 1.5) and 1.0, respectively, thereby manipulating the unrestricted population effect sizes at the specified levels. Second, the generated Xt scores were converted into the observed X scores through
X = ρX X t X t + eX = √
r X X a X t + eX , (8) where eX is the measurement error of X, which was generated from a nor- mal distribution with mean 0 and S D =
√ 1/r X X a − 1.4 In this way, the
generated X scores contained measurement errors, and the associated reli- ability was manipulated at a generated value of r X X a . Third, the selection variable S (or called Z in Case III) scores of the unrestricted applicants were generated based on5
S = ρS X t X t + eS (9) where eS is the measurement error of S, which was generated from a normal distribution with mean 0 and S D =√1−ρ2S Xt . Hence the correlation between the generated S and Xt scores was fixed at ρS X t according to the specified simulation conditions. Given the generated S, X, and G, the simulated examinees were rank-ordered by the S scores, producing a data matrix ⎡
⎢⎣ Sr Su
∣∣∣∣∣ Xr ...
∣∣∣∣∣ Gr ...
⎤ ⎥⎦ (10)
where (Sr, Xr, Gr) is the restricted sample of size nr, and Su is the un- selected sample of size nu (i.e., N – nr), where N is the sample size of
4Note that the observed X created by Equation 8 will have the standard deviation of√ 1/
√ r X X , so the effect size of X in the unrestricted population will be smaller than the true
effect size of Xt by a factor of √
r X X . 5The notation S is used to indicate the suitability construct S for Case IV indirect RR.
For simplicity, the same notation S is used to label the third variable Z for Case III indirect RR.
910 PERSONNEL PSYCHOLOGY
the unrestricted applicants. The remaining procedures followed Equa- tions 2 and 3–7,6 resulting in estimates of the Cases III and IV corrected ds. As noted, the unrestricted SX (or SZ for Case III) are sometimes re- ported in previous studies or in test manuals. In the present simulation, the unrestricted SX (or SZ) was estimated based on a generated unrestricted sample of X (or Z).
Evaluation Criteria
Two evaluation criteria were used. First, percentage bias was used to evaluate the performance of the restricted and bias-adjusted ds in each manipulated condition. That is, bias = [(∀ − δ)/δ] × 100%, where ∀ is the mean of the Cases III, IV, and uncorrected ds across 2,000 replications, respectively. As in Li et al. (2011), a parameter estimate was considered reasonable if a bias was within ±10%. Second, to summarize the biases across T simulation conditions, Flores (1986) proposed the mean absolute percentage error (MAPE): MAPE = ∑Ti =1 |bias(i )|/T . A MAPE within 10% was regarded as an appropriate fit (Brockwell & Davis, 2002).
Results
The simulation program stopped frequently due to the lack of simu- lated applicants in conditions in which the proportion of applicants was highly unbalanced and or in which the restricted sample size was small. Hence, the results are discussed in two sections. The first presents the re- sults of 3,000 simulation conditions when nr = 100 and 300. The second shows the results of 1,800 simulation conditions when nr = 30 and 50, but the extreme proportions (i.e., P = .1, and .9) are dropped.
When nr = 100 and 300
As shown in Figure 2, the restricted d (i.e., dr) values were downwardly biased across the 3,000 simulation conditions. The biases ranged from –65.7% to .3%, with a mean of –28.3%. Of the 3,000 conditions, only 298 (or 9.9%) yielded a bias within the nominal level of ±10%.
The Case III corrected d (dc3) values were only adequate. The biases ranged from –62.6% to 80.4%, with a mean of –3.7%. Of the 3,000 conditions, 1,750 (or 58.3%) fell inside the nominal range. Overall, the MAPE was 12.6%, which was slightly larger than the criterion of 10%.
6In practice, the restricted reliability r X Xi can be obtained either using the real data or converting an unrestricted reliability r X Xa to a restricted reliability r X Xi through Equation 4. Given that the present simulation study generated a single variable X only, it used the latter approach in order to convert the unrestricted to restricted reliability, r X Xi .
JOHNSON CHING-HONG LI 911
-.80
-.60
-.40
-.20
.00
.20
.40
.60
.80 Restricted
-.80
-.60
-.40
-.20
.00
.20
.40
.60
.80 Case III
-.80
-.60
-.40
-.20
.00
.20
.40
.60
.80 Case IV
nr = 100 nr = 300
Figure 2: Biases Obtained by the Restricted, Case III, and Case IV Corrected ds Across the 3,000 Simulation Conditions When nr = 100 or 300.
Comparatively, the Case IV corrected d (dc4) was more accurate than the dc3. The biases ranged from –37.5% to 45.4%, with a mean of –3.5%. Of the 3,000 conditions, 2,537 (or 84.6%) yielded a bias within the nominal range of ±10%. The MAPE, as an indicator of the overall accuracy, was 5.7%, which was within the nominal value of 10%.
In sum, the Case IV corrected dc4 often yielded a better estimate of the true unrestricted δ than the uncorrected dr and the Case III adjusted dc3. The following section discusses the effects of each of the manipulated factors on the uncorrected and adjusted ds.
912 PERSONNEL PSYCHOLOGY
Effects of the Manipulated Factors on the Restricted d (Table 1)
Given that the total number of simulation conditions was large, the simulation results that come from the two extremes of each manipulated factor are presented in Table 1. First, when the selection process became more stringent, that is, with a smaller selection ratio (π ) and a stronger correlation between S and Xt (ρS X t ), the restricted d became smaller. This is reasonable as the variability of the restricted X scores was further restricted given its strong relationship with the selection construct, and only a small proportion of simulated applicants could be selected. Second, better reliability slightly improved the accuracy of the restricted d because the observed effect size is less attenuated by measurement error in the instrument. Third, the sample size, proportion of participants in each group, and magnitude of d did not show any obvious impact on the restricted d.
Effects of the Manipulated Factors on the Case III Adjusted d (Table 2)
First, the factor that appeared to have the strongest influence on the Case III corrected d was the magnitude of the true unrestricted d. When the magnitude was .2, 96 (or 80%) of the 120 conditions produced an acceptable adjusted d. However, when it was 1.5, only 57 (or 47.5%) of the 120 conditions produced a reasonable d. This finding was likely due to the two key parameters (i.e., dZi and dX i ) that require the nonlin- ear r-to-d transformations and back transformations in using the Case III correction. Specifically, the original Case III correction for r requires three restricted bivariate correlations (i.e., r X Yi , r X Zi , and rY Zi ; Bobko et al., 2001, Equation 5 for correction. The present Case III correction for d needed two nonlinear r-to-d transformations (i.e., r X Yi to dX i and rY Zi to dZi ) and also two back transformations; this tended to distort the accuracy of the correction substantially. An anonymous reviewer noted that the formula underlying this r-to-d transformation is based on the assumption of bivariate normality of the variables (i.e., Xt and G and S and G in this case). Yet, in the current simulation, this assumption is violated because G is a categorical variable and Xt is a bimodal vari- able, which explains the inaccuracy of Case III. Second, this distortion became more severe when the proportion of participants was more un- balanced, the selection ratio was more stringent, and/or the correlation between S and Xt was stronger. Third, the accuracy of the corrected d slightly improved when the unrestricted reliability increased. Fourth, the two sample sizes (100 and 300) did not show obvious effects on the accuracy.
JOHNSON CHING-HONG LI 913
T A
B L
E 1
M ea
n s
a n
d B
ia se
s o
f th
e R
es tr
ic te
d d
in S
el ec
te d
S im
u la
ti o
n C
o n
d it
io n
s W
h en
R es
tr ic
te d
S iz
es W
er e
1 0
0 a
n d
3 0
0
δ n r
ρ X
X a
ρ S
X t
π P
= .1
.3 .5
.7 .9
.2 10
0 .6
.2 .1
.1 44
(− .2
8) .1
31 (−
.3 4)
.1 44
(− .2
8) .1
48 (−
.2 6)
.1 57
(− .2
2) .9
.1 52
(− .2
4) .1
40 (−
.3 0)
.1 43
(− .2
8) .1
56 (−
.2 2)
.1 39
(− .3
0) .5
.1 .1
32 (−
.3 4)
.1 10
(− .4
5) .1
20 (−
.4 0)
.1 14
(− .4
3) .1
19 (−
.4 1)
.9 .1
44 (−
.2 8)
.1 39
(− .3
1) .1
39 (−
.3 1)
.1 48
(− .2
6) .1
51 (−
.2 5)
.8 .1
.0 78
(− .6
1) .0
84 (−
.5 8)
.0 80
(− .6
0) .0
92 (−
.5 4)
.0 65
(− .6
7) .9
.1 26
(− .3
7) .1
25 (−
.3 8)
.1 26
(− .3
7) .1
36 (−
.3 2)
.1 03
(− .4
9) .9
.2 .1
.1 81
(− .1
0) .1
76 (−
.1 2)
.1 99
(− .0
1) .1
93 (−
.0 3)
.1 95
(− .0
2) .9
.1 93
(− .0
3) .1
85 (−
.0 8)
.1 78
(− .1
1) .1
89 (−
.0 6)
.1 99
(. 00
) .5
.1 .1
90 (−
.0 5)
.1 69
(− .1
6) .1
62 (−
.1 9)
.1 70
(− .1
5) .1
55 (−
.2 2)
.9 .1
63 (−
.1 8)
.1 91
(− .0
5) .1
84 (−
.0 8)
.1 68
(− .1
6) .1
63 (−
.1 8)
.8 .1
.1 17
(− .4
2) .1
10 (−
.4 5)
.1 20
(− .4
0) .1
19 (−
.4 1)
.1 24
(− .3
8) .9
.1 63
(− .1
9) .1
65 (−
.1 8)
.1 69
(− .1
6) .1
68 (−
.1 6)
.1 79
(− .1
0) 30
0 .6
.2 .1
.1 47
(− .2
6) .1
44 (−
.2 8)
.1 31
(− .3
4) .1
36 (−
.3 2)
.1 34
(− .3
3) .9
.1 39
(− .3
0) .1
41 (−
.2 9)
.1 49
(− .2
5) .1
46 (−
.2 7)
.1 57
(− .2
1) .5
.1 .1
15 (−
.4 2)
.1 20
(− .4
0) .1
26 (−
.3 7)
.1 15
(− .4
3) .1
28 (−
.3 6)
.9 .1
42 (−
.2 9)
.1 42
(− .2
9) .1
50 (−
.2 5)
.1 34
(− .3
3) .1
39 (−
.3 0)
.8 .1
.0 88
(− .5
6) .0
83 (−
.5 9)
.0 82
(− .5
9) .0
80 ( −
.6 0)
.0 80
(− .6
0) .9
.1 24
(− .3
8) .1
30 (−
.3 5)
.1 26
(− .3
7) .1
38 (−
.3 1)
.1 23
(− .3
9) .9
.2 .1
.1 72
(− .1
4) .1
91 (−
.0 4)
.1 74
(− .1
3) .1
78 (−
.1 1)
.2 10
(. 05
) .9
.2 00
(. 00
) .2
05 (.
03 )
.1 96
(− .0
2) .1
93 (−
.0 4)
.1 65
(− .1
7) .5
.1 .1
58 (−
.2 1)
.1 60
(− .2
0) .1
61 (−
.1 9)
.1 62
(− .1
9) .1
70 (−
.1 5)
.9 .1
90 (−
.0 5)
.1 90
(− .0
5) .1
84 (−
.0 8)
.1 81
(− .0
9) .1
75 (−
.1 3)
.8 .1
.1 31
(− .3
5) .1
25 (−
.3 7)
.1 23
(− .3
9) .1
23 (−
.3 9)
.1 24
(− .3
8) .9
.1 64
(− .1
8) .1
77 (−
.1 2)
.1 69
(− .1
6) .1
63 (−
.1 8)
.1 84
(− .0
8) 1.
5 10
0 .6
.2 .1
1. 07
8 (−
.2 8)
1. 06
8 (−
.2 9)
1. 07
7 (−
.2 8)
1. 08
4 (−
.2 8)
1. 08
7 (−
.2 8)
.9 1.
10 0
(− .2
7) 1.
10 2
(− .2
7) 1.
11 1
(− .2
6) 1.
10 4
(− .2
6) 1.
10 2
(− .2
7)
co n ti
n u ed
914 PERSONNEL PSYCHOLOGY T
A B
L E
1 (c
on ti
nu ed
)
δ n r
ρ X
X a
ρ S
X t
π P
= .1
.3 .5
.7 .9
.5 .1
.9 43
(− .3
7) .9
42 (−
.3 7)
.9 22
(− .3
9) .9
25 (−
.3 8)
.9 48
(− .3
7) .9
1. 06
2 (−
.2 9)
1. 06
0 (−
.2 9)
1. 04
8 (−
.3 0)
1. 03
5 (−
.3 1)
1. 04
5 (−
.3 0)
.8 .1
.6 60
(− .5
6) .6
21 (−
.5 9)
.6 17
(− .5
9) .5
96 (−
.6 0)
.5 70
(− .6
2) .9
1. 05
0 (−
.3 0)
1. 02
3 (−
.3 2)
.9 94
(− .3
4) .9
16 (−
.3 9)
.8 44
(− .4
4) .9
.2 .1
1. 39
9 (−
.0 7)
1. 39
1 (−
.0 7)
1. 39
0 (−
.0 7)
1. 38
8 (−
.0 7)
1. 41
6 (−
.0 6)
.9 1.
41 8
(− .0
5) 1.
42 6
(− .0
5) 1.
41 4
(− .0
6) 1.
41 9
(− .0
5) 1.
40 6
(− .0
6) .5
.1 1.
26 2
(− .1
6) 1.
24 3
(− .1
7) 1.
24 8
(− .1
7) 1.
21 8
(− .1
9) 1.
22 0
(− .1
9) .9
1. 38
9 (−
.0 7)
1. 37
2 (−
.0 9)
1. 37
1 (−
.0 9)
1. 35
6 (−
.1 0)
1. 31
7 (−
.1 2)
.8 .1
.9 72
(− .3
5) .9
14 (−
.3 9)
.8 90
(− .4
1) .8
67 (−
.4 2)
.8 70
(− .4
2) .9
1. 39
0 (−
.0 7)
1. 34
4 (−
.1 0)
1. 29
3 (−
.1 4)
1. 22
0 (−
.1 9)
1. 13
6 (−
.2 4)
30 0
.6 .2
.1 1.
06 5
(− .2
9) 1.
07 7
(− .2
8) 1.
05 0
(− .3
0) 1.
05 5
(− .3
0) 1.
05 4
(− .3
0) .9
1. 08
2 (−
.2 8)
1. 09
0 (−
.2 7)
1. 09
6 (−
.2 7)
1. 08
9 (−
.2 7)
1. 09
6 (−
.2 7)
.5 .1
.9 42
(− .3
7) .9
22 (−
.3 9)
.9 25
(− .3
8) .9
02 (−
.4 0)
.9 31
(− .3
8) .9
1. 07
2 (−
.2 9)
1. 05
7 (−
.3 0)
1. 04
1 (−
.3 1)
1. 03
7 (−
.3 1)
.9 91
(− .3
4) .8
.1 .6
66 (−
.5 6)
.6 12
(− .5
9) .5
74 (−
.6 2)
.5 97
(− .6
0) .5
63 (−
.6 2)
.9 1.
05 3
(− .3
0) 1.
02 8
(− .3
1) .9
79 (−
.3 5)
.9 32
(− .3
8) .8
48 (−
.4 3)
.9 .2
.1 1.
38 9
(− .0
7) 1.
37 9
(− .0
8) 1.
39 0
(− .0
7) 1.
38 6
(− .0
8) 1.
39 1
(− .0
7) .9
1. 40
7 ( −
.0 6)
1. 39
9 (−
.0 7)
1. 40
8 (−
.0 6)
1. 41
4 (−
.0 6)
1. 39
4 (−
.0 7)
.5 .1
1. 24
9 (−
.1 7)
1. 24
2 (−
.1 7)
1. 25
2 (−
.1 7)
1. 22
1 (−
.1 9)
1. 24
6 (−
.1 7)
.9 1.
37 4
(− .0
8) 1.
37 8
(− .0
8) 1.
35 0
(− .1
0) 1.
35 0
(− .1
0) 1.
34 0
(− .1
1) .8
.1 .9
54 (−
.3 6)
.8 86
(− .4
1) .8
67 (−
.4 2)
.8 59
(− .4
3) .8
47 (−
.4 4)
.9 1.
38 5
(− .0
8) 1.
33 9
(− .1
1) 1.
30 0
(− .1
3) 1.
22 4
(− .1
8) 1.
13 2
(− .2
5)
N o
te .
δ is
th e
tr ue
un re
st ri
ct ed
C oh
en ’s
d ,n
r is
th e
re st
ri ct
ed sa
m pl
e si
ze ,ρ
X X
a is
th e
un re
st ri
ct ed
re li
ab il
it y,
ρ S
X t
is th
e un
re st
ri ct
ed co
rr el
at io
n be
tw ee
n S
an d
X t,
π is
th e
se le
ct io
n ra
ti o,
an d
P is
th e
pr op
or ti
on of
un re
st ri
ct ed
pa rt
ic ip
an ts
in G
ro up
1. B
ia se
s ar
e pr
es en
te d
in pa
re nt
he se
s. B
ia se
s ou
ts id
e th
e no
m in
al ra
ng e
of ±1
0% ar
e pr
es en
te d
in bo
ld .
JOHNSON CHING-HONG LI 915
T A
B L
E 2
M ea
n s
a n
d B
ia se
s o
f th
e C
a se
II I
C o
rr ec
te d
d in
S el
ec te
d S
im u
la ti
o n
C o
n d
it io
n s
W h
en R
es tr
ic te
d S
iz es
W er
e 1
0 0
a n
d 3
0 0
δ n
r ρ
X X
a ρ
S X
t π
P =
.1 .3
.5 .7
.9
.2 10
0 .6
.2 .1
.1 92
(− .0
4) .1
72 (−
.1 4)
.1 86
(− .0
7) .1
87 (−
.0 6)
.1 83
(− .0
8) .9
.1 98
(− .0
1) .1
79 (−
.1 0)
.1 81
(− .0
9) .1
94 (−
.0 3)
.1 74
(− .1
3) .5
.1 .2
12 (.
06 )
.1 67
(− .1
6) .1
85 (−
.0 8)
.1 70
(− .1
5) .1
64 (−
.1 8)
.9 .1
97 (−
.0 1)
.1 87
(− .0
7) .1
82 (−
.0 9)
.1 96
(− .0
2) .1
95 (−
.0 3)
.8 .1
.1 95
(− .0
3) .1
98 (−
.0 1)
.1 86
(− .0
7) .1
69 (−
.1 6)
.1 40
(− .3
0) .9
.1 88
(− .0
6) .1
84 (−
.0 8)
.1 83
(− .0
9) .1
99 (−
.0 1)
.1 50
(− .2
5) .9
.2 .1
.1 88
(− .0
6) .1
96 (−
.0 2)
.2 15
(. 07
) .2
08 (.
04 )
.2 00
(. 00
) .9
.2 13
(. 07
) .1
97 (−
.0 2)
.1 89
(− .0
5) .1
98 (−
.0 1)
.2 08
(. 04
) .5
.1 .2
22 (.
11 )
.2 10
(. 05
) .1
94 (−
.0 3)
.2 02
(. 01
) .1
84 (−
.0 8)
.9 .1
85 (−
.0 7)
.2 12
(. 06
) .2
04 (.
02 )
.1 85
(− .0
8) .1
77 (−
.1 1)
.8 .1
.2 20
(. 10
) .1
99 (.
00 )
.2 04
(. 02
) .1
82 (−
.0 9)
.1 88
(− .0
6) .9
.2 03
(. 02
) .1
96 (−
.0 2)
.2 01
(. 01
) .1
98 (−
.0 1)
.2 05
(. 03
) 30
0 .6
.2 .1
.2 00
(. 00
) .1
87 (−
.0 6)
.1 69
(− .1
5) .1
71 (−
.1 4)
.1 64
(− .1
8) .9
.1 77
(− .1
2) .1
78 (−
.1 1)
.1 86
(− .0
7) .1
84 (−
.0 8)
.1 97
(− .0
2) .5
.1 .1
73 (−
.1 3)
.1 83
(− .0
8) .1
90 (−
.0 5)
.1 57
(− .2
1) .1
79 (−
.1 0)
.9 .1
89 (−
.0 5)
.1 89
(− .0
6) .1
99 (.
00 )
.1 74
(− .1
3) .1
81 (−
.1 0)
.8 .1
.2 17
(. 08
) .2
11 (.
06 )
.2 05
(. 03
) .1
63 (−
.1 8)
.1 75
(− .1
3) .9
.1 82
(− .0
9) .1
90 (−
.0 5)
.1 82
(− .0
9) .1
99 (−
.0 1)
.1 72
(− .1
4) .9
.2 .1
.1 88
(− .0
6) .2
06 (.
03 )
.1 90
(− .0
5) .1
90 (−
.0 5)
.2 19
(. 09
) .9
.2 16
(. 08
) .2
19 (.
09 )
.2 09
(. 04
) .2
05 (.
03 )
.1 74
(− .1
3) .5
.1 .2
13 (.
07 )
.2 07
(. 04
) .1
89 (−
.0 5)
.1 85
(− .0
7) .2
04 (.
02 )
.9 .2
12 (.
06 )
.2 11
(. 05
) .2
02 (.
01 )
.2 00
(. 00
) .1
87 (−
.0 7)
.8 .1
.2 48
(. 24
) .2
09 (.
05 )
.2 12
(. 06
) .1
86 (−
.0 7)
.1 96
( − .0
2) .9
.2 00
(. 00
) .2
13 (.
06 )
.1 99
(− .0
1) .1
92 (−
.0 4)
.2 14
(. 07
) 1.
5 10
0 .6
.2 .1
1. 71
2 (.
14 )
1. 60
4 (.
07 )
1. 47
8 (−
.0 1)
1. 29
5 (−
.1 4)
1. 12
4 (−
.2 5)
.9 1.
50 7
(. 00
) 1.
53 1
(. 02
) 1.
53 9
(. 03
) 1.
47 8
(− .0
1) 1.
38 2
(− .0
8)
co n ti
n u ed
916 PERSONNEL PSYCHOLOGY T
A B
L E
2 (c
on ti
nu ed
)
δ n
r ρ
X X
a ρ
S X
t π
P =
.1 .3
.5 .7
.9
.5 .1
2. 17
5 (.
45 )
1. 67
2 (.
11 )
1. 21
2 (−
.1 9)
.9 30
(− .3
8) .7
84 (−
.4 8)
.9 1.
57 7
(. 05
) 1.
56 3
(. 04
) 1.
54 7
(. 03
) 1.
41 5
(− .0
6) 1.
27 0
(− .1
5) .8
.1 2.
52 9
(. 69
) 1.
39 0
(− .0
7) .8
28 (−
.4 5)
.5 62
(− .6
3) .6
06 (−
.6 0)
.9 1.
70 9
(. 14
) 1.
68 3
(. 12
) 1.
59 9
(. 07
) 1.
35 9
(− .0
9) 1.
03 4
(− .3
1) .9
.2 .1
1. 83
2 (.
22 )
1. 65
8 (.
11 )
1. 49
5 (.
00 )
1. 33
3 (−
.1 1)
1. 19
5 (−
.2 0)
.9 1.
60 0
(. 07
) 1.
58 1
(. 05
) 1.
55 3
(. 04
) 1.
51 8
(. 01
) 1.
44 6
(− .0
4) .5
.1 2.
34 7
(. 56
) 1.
66 5
(. 11
) 1.
23 6
(− .1
8) .9
19 (−
.3 9)
.8 37
(− .4
4) .9
1. 67
8 (.
12 )
1. 61
0 (.
07 )
1. 57
7 (.
05 )
1. 47
6 (−
.0 2)
1. 31
4 (−
.1 2)
.8 .1
2. 70
6 (.
80 )
1. 37
1 (−
.0 9)
.8 26
(− .4
5) .5
73 (−
.6 2)
.6 76
(− .5
5) .9
1. 82
3 (.
22 )
1. 70
0 (.
13 )
1. 59
5 (.
06 )
1. 39
9 (−
.0 7)
1. 10
5 (−
.2 6)
30 0
.6 .2
.1 1.
70 6
(. 14
) 1.
60 5
(. 07
) 1.
44 0
(− .0
4) 1.
27 0
(− .1
5) 1.
08 0
(− .2
8) .9
1. 44
5 (−
.0 4)
1. 49
5 (.
00 )
1. 49
8 (.
00 )
1. 45
9 (−
.0 3)
1. 37
5 (−
.0 8)
.5 .1
2. 18
7 (.
46 )
1. 60
9 (.
07 )
1. 22
0 (−
.1 9)
.9 02
(− .4
0) .6
72 (−
.5 5)
.9 1.
54 1
(. 03
) 1.
55 3
(. 04
) 1.
51 3
(. 01
) 1.
41 8
(− .0
5) 1.
21 1
(− .1
9) .8
.1 2.
54 2
(. 69
) 1.
37 3
(− .0
8) .7
79 (−
.4 8)
.5 27
(− .6
5) .4
10 (−
.7 3)
.9 1.
66 2
(. 11
) 1.
64 9
(. 10
) 1.
55 6
(. 04
) 1.
37 8
(− .0
8) 1.
05 1
(− .3
0) .9
.2 .1
1. 82
4 (.
22 )
1. 64
5 (.
10 )
1. 49
6 (.
00 )
1. 32
9 (−
.1 1)
1. 17
7 (−
.2 2)
.9 1.
55 7
(. 04
) 1.
54 7
(. 03
) 1.
54 3
(. 03
) 1.
52 4
(. 02
) 1.
44 2
(− .0
4) .5
.1 2.
32 1
(. 55
) 1.
69 3
(. 13
) 1.
25 1
(− .1
7) .9
31 (−
.3 8)
.7 25
(− .5
2) .9
1. 62
3 (.
08 )
1. 61
9 (.
08 )
1. 54
6 (.
03 )
1. 47
8 (−
.0 1)
1. 34
8 (−
.1 0)
.8 .1
2. 69
4 (.
80 )
1. 35
2 (−
.1 0)
.7 98
(− .4
7) .5
22 (−
.6 5)
.4 17
(− .7
2) .9
1. 76
4 (.
18 )
1. 68
3 (.
12 )
1. 59
5 (.
06 )
1. 41
0 (−
.0 6)
1. 13
0 (−
.2 5)
N o
te .
δ is
th e
tr ue
un re
st ri
ct ed
C oh
en ’s
d ,n
r is
th e
re st
ri ct
ed sa
m pl
e si
ze ,ρ
X X
a is
th e
un re
st ri
ct ed
re li
ab il
it y,
ρ S
X t
is th
e un
re st
ri ct
ed co
rr el
at io
n be
tw ee
n S
an d
X t,
π is
th e
se le
ct io
n ra
ti o,
an d
P is
th e
pr op
or ti
on of
un re
st ri
ct ed
pa rt
ic ip
an ts
in G
ro up
1. B
ia se
s ar
e pr
es en
te d
in pa
re nt
he se
s. B
ia se
s ou
ts id
e th
e no
m in
al ra
ng e
of ±1
0% ar
e pr
es en
te d
in bo
ld .
JOHNSON CHING-HONG LI 917
Effects of the Manipulated Factors on the Case IV Corrected d (Table 3)
First, the unrestricted reliability appeared to be an influential factor. When ρX X a = .6, only 71 (or 59.2%) of the 120 conditions obtained a reasonable adjustment; however, of the remaining 120 conditions with a reliability of .9, 108 (or 90%) resulted in a good adjusted d. Second, as ρS X t increased, the accuracy appeared to further decrease. When ρS X t = .8 and ρX X a = .6, only 14 (or 35%) of the 40 conditions yielded an accept- able correction. Third, the accuracy of the corrected d increased when the selection ratio was less stringent. Fourth, other factors, including the proportion of participants, restricted sample size, and true d did not show obvious effects on the accuracy of the corrected d.
When nr = 30 and 50
Because no simulated examinee could be selected from the less priv- ileged group when the unrestricted proportion P was too extreme, this section discusses the results based on P = .30, .50, and .70, with a total of 1,800 simulation conditions.7 As in the aforementioned section, the uncorrected ds were inaccurate. The biases ranged from –65.2% to 5.1%, with a mean of –27.1%. Of the 1,800 conditions, only 208 (or 11.6%) were within the nominal range of ±10%. The overall performance across the 1,800 conditions was poor, given that the MAPE was only 27.1% outside the criterion of 10%.
The two biased-corrected d estimates performed appropriately. The Case III corrected ds produced biases ranging from –55.0% to 32.3%, with a mean of –1.3%. Of the 1,800 conditions, 1,266 (or 70.3%) were within the nominal range of ±10%. The overall performance was reasonable, as reflected by the MAPE (i.e., 8.6%), which was within the criterion of 10%. Moreover, the Case IV corrected ds produced biases, which ranged from –33.7% to 50.8%, with a mean of 2.0%. This range or fluctuation was larger than that reported in the previous section (i.e., nr � 100), given that a smaller restricted sample size led to a wider variability. Specifically, larger variability was found when nr = 30. That is, the biases ranged from –33.7% to 50.8%, with a mean of 4.3%. When nr = 50, the biases ranged from –25.6% to 26.8%, with a mean of –.3%. Moreover, of the 1,800 conditions, 1,498 (or 83.2%) were within the criterion of ±10%. To summarize the overall performance across the 1,800 conditions, the MAPE was 6.2% within the criterion, meaning that the Case IV corrected
7The pattern of the results is similar to that shown in Figure 2. Details of the pattern are available upon request.
918 PERSONNEL PSYCHOLOGY T
A B
L E
3 M
ea n
s a
n d
B ia
se s
o f
th e
C a
se IV
C o
rr ec
te d
d in
S el
ec te
d S
im u
la ti
o n
C o
n d
it io
n s
W h
en R
es tr
ic te
d S
iz es
W er
e 1
0 0
a n
d 3
0 0
δ n
r ρ
X X
a ρ
S X
t π
P =
.1 .3
.5 .7
.9
.2 10
0 .6
.2 .1
.1 84
(− .0
8) .1
70 (−
.1 5)
.1 84
(− .0
8) .1
91 (−
.0 4)
.2 03
(. 01
) .9
.1 90
(− .0
5) .1
76 (−
.1 2)
.1 80
(− .1
0) .1
94 (−
.0 3)
.1 73
(− .1
3) .5
.1 .1
91 (−
.0 5)
.1 58
(− .2
1) .1
74 (−
.1 3)
.1 69
(− .1
6) .1
73 (−
.1 4)
.9 .1
86 (−
.0 7)
.1 81
(− .1
0) .1
80 (−
.1 0)
.1 93
(− .0
3) .1
96 (−
.0 2)
.8 .1
.1 57
(− .2
2) .1
74 (−
.1 3)
.1 59
(− .2
0) .1
89 (−
.0 6)
.1 25
(− .3
7) .9
.1 78
(− .1
1) .1
75 (−
.1 2)
.1 76
(− .1
2) .1
93 (−
.0 3)
.1 43
(− .2
8) .9
.2 .1
.1 95
(− .0
2) .1
88 (−
.0 6)
.2 17
(. 08
) .2
09 (.
05 )
.2 10
(. 05
) .9
.2 06
(. 03
) .1
96 (−
.0 2)
.1 89
(− .0
5) .2
00 (.
00 )
.2 12
(. 06
) .5
.1 .2
29 (.
15 )
.2 03
(. 01
) .1
96 (−
.0 2)
.2 05
(. 02
) .1
88 (−
.0 6)
.9 .1
80 (−
.1 0)
.2 10
(. 05
) .2
02 (.
01 )
.1 85
(− .0
8) .1
80 (−
.1 0)
.8 .1
.1 90
( − .0
5) .1
82 (−
.0 9)
.1 96
(− .0
2) .1
93 (−
.0 3)
.2 04
(. 02
) .9
.1 94
(− .0
3) .1
95 (−
.0 2)
.2 00
(. 00
) .2
01 (.
00 )
.2 13
(. 07
) 30
0 .6
.2 .1
.1 88
(− .0
6) .1
86 (−
.0 7)
.1 70
(− .1
5) .1
74 (−
.1 3)
.1 71
(− .1
4) .9
.1 74
(− .1
3) .1
76 (−
.1 2)
.1 86
(− .0
7) .1
83 (−
.0 8)
.1 98
(− .0
1) .5
.1 .1
63 (−
.1 9)
.1 71
(− .1
5) .1
81 (−
.1 0)
.1 66
(− .1
7) .1
85 (−
.0 8)
.9 .1
83 (−
.0 8)
.1 84
(− .0
8) .1
95 (−
.0 3)
.1 72
(− .1
4) .1
80 (−
.1 0)
.8 .1
.1 72
(− .1
4) .1
64 (−
.1 8)
.1 66
(− .1
7) .1
55 (−
.2 3)
.1 59
(− .2
0) .9
.1 73
(− .1
3) .1
83 (−
.0 9)
.1 75
(− .1
2) .1
93 (−
.0 3)
.1 72
(− .1
4) .9
.2 .1
.1 85
(− .0
7) .2
06 (.
03 )
.1 87
(− .0
6) .1
92 (−
.0 4)
.2 27
(. 13
) .9
.2 12
(. 06
) .2
18 (.
09 )
.2 08
(. 04
) .2
05 (.
03 )
.1 76
(− .1
2) .5
.1 .1
89 (−
.0 6)
.1 91
(− .0
4) .1
93 (−
.0 4)
.1 94
(− .0
3) .2
03 (.
01 )
.9 .2
09 (.
04 )
.2 08
(. 04
) .2
02 (.
01 )
.1 99
(. 00
) .1
92 (−
.0 4)
.8 .1
.2 14
(. 07
) .2
03 (.
02 )
.1 99
(. 00
) .1
99 (.
00 )
.2 04
(. 02
) .9
.1 95
(− .0
3) .2
09 (.
05 )
.1 98
(− .0
1) .1
92 (−
.0 4)
.2 16
(. 08
) 1.
5 10
0 .6
.2 .1
1. 33
4 (−
.1 1)
1. 36
0 (−
.0 9)
1. 42
9 (−
.0 5)
1. 48
8 (−
.0 1)
1. 47
2 (−
.0 2)
.9 1.
36 3
(− .0
9) 1.
37 8
(− .0
8) 1.
39 1
(− .0
7) 1.
39 7
(− .0
7) 1.
39 6
(− .0
7)
C o
n ti
n u
ed
JOHNSON CHING-HONG LI 919
T A
B L
E 3
(C on
ti nu
ed )
δ n
r ρ
X X
a ρ
S X
t π
P =
.1 .3
.5 .7
.9
.5 .1
1. 26
9 (−
.1 5)
1. 44
3 (−
.0 4)
1. 60
1 (.
07 )
1. 67
3 (.
12 )
1. 51
3 (.
01 )
.9 1.
38 3
(− .0
8) 1.
37 7
(− .0
8) 1.
39 7
(− .0
7) 1.
39 1
(− .0
7) 1.
39 4
(− .0
7) .8
.1 1.
28 7
(− .1
4) 1.
69 7
(. 13
) 2.
10 1
(. 40
) 1.
81 0
(. 21
) 1.
28 7
(− .1
4) .9
1. 45
2 (−
.0 3)
1. 45
6 (−
.0 3)
1. 43
8 (−
.0 4)
1. 37
7 (−
.0 8)
1. 25
8 (−
.1 6)
.9 .2
.1 1.
45 3
(− .0
3) 1.
47 6
(− .0
2) 1.
53 4
(. 02
) 1.
58 0
(. 05
) 1.
58 6
(. 06
) .9
1. 49
9 (.
00 )
1. 51
1 (.
01 )
1. 50
6 (.
00 )
1. 52
0 (.
01 )
1. 50
9 (.
01 )
.5 .1
1. 43
2 (−
.0 5)
1. 53
1 (.
02 )
1. 69
8 (.
13 )
1. 71
7 (.
14 )
1. 59
0 (.
06 )
.9 1.
52 3
(. 02
) 1.
51 1
(. 01
) 1.
52 7
(. 02
) 1.
53 9
(. 03
) 1.
48 6
(− .0
1) .8
.1 1.
50 2
(. 00
) 1.
70 9
(. 14
) 1.
87 3
(. 25
) 1.
80 9
(. 21
) 1.
56 2
(. 04
) .9
1. 63
6 (.
09 )
1. 58
0 (.
05 )
1. 55
4 (.
04 )
1. 51
3 (.
01 )
1. 40
8 (−
.0 6)
30 0
.6 .2
.1 1.
31 3
(− .1
2) 1.
35 1
(− .1
0) 1.
39 4
(− .0
7) 1.
43 8
(− .0
4) 1.
40 3
(− .0
6) .9
1. 35
5 (−
.1 0)
1. 36
0 (−
.0 9)
1. 36
9 (−
.0 9)
1. 37
7 (−
.0 8)
1. 38
5 (−
.0 8)
.5 .1
1. 26
0 (−
.1 6)
1. 37
2 (−
.0 9)
1. 56
9 (.
05 )
1. 59
2 (.
06 )
1. 46
9 (−
.0 2)
.9 1.
38 7
(− .0
8) 1.
37 5
(− .0
8) 1.
37 9
(− .0
8) 1.
38 9
(− .0
7) 1.
31 9
(− .1
2) .8
.1 1.
25 6
(− .1
6) 1.
63 6
(. 09
) 1.
87 2
(. 25
) 1.
76 8
(. 18
) 1.
29 9
(− .1
3) .9
1. 45
7 (−
.0 3)
1. 44
5 (−
.0 4)
1. 41
7 (−
.0 6)
1. 39
0 (−
.0 7)
1. 25
1 (−
.1 7)
.9 .2
.1 1.
43 9
(− .0
4) 1.
45 5
(− .0
3) 1.
52 9
(. 02
) 1.
57 2
(. 05
) 1.
55 5
(. 04
) .9
1. 49
0 (−
.0 1)
1. 48
4 (−
.0 1)
1. 50
1 (.
00 )
1. 51
2 (.
01 )
1. 49
3 (.
00 )
.5 .1
1. 41
0 (−
.0 6)
1. 52
1 (.
01 )
1. 69
1 (.
13 )
1. 70
2 (.
13 )
1. 62
8 (.
09 )
.9 1.
50 8
(. 01
) 1.
51 8
(. 01
) 1.
50 4
(. 00
) 1.
52 6
(. 02
) 1.
50 8
(. 01
) .8
.1 1.
45 9
(− .0
3) 1.
62 9
(. 09
) 1.
80 1
(. 20
) 1.
77 4
(. 18
) 1.
54 4
(. 03
) .9
1. 61
7 (.
08 )
1. 57
2 (.
05 )
1. 55
8 (.
04 )
1. 50
7 (.
00 )
1. 39
2 (−
.0 7)
N o
te .
δ is
th e
tr ue
un re
st ri
ct ed
C oh
en ’s
d ,n
r is
th e
re st
ri ct
ed sa
m pl
e si
ze ,ρ
X X
a is
th e
un re
st ri
ct ed
re li
ab il
it y,
ρ S
X t
is th
e un
re st
ri ct
ed co
rr el
at io
n be
tw ee
n S
an d
X t,
π is
th e
se le
ct io
n ra
ti o,
an d
P is
th e
pr op
or ti
on of
un re
st ri
ct ed
pa rt
ic ip
an ts
in G
ro up
1. B
ia se
s ar
e pr
es en
te d
in pa
re nt
he se
s. B
ia se
s ou
ts id
e th
e no
m in
al ra
ng e
of ±1
0% ar
e pr
es en
te d
in bo
ld .
920 PERSONNEL PSYCHOLOGY
d is appropriate even when the total restricted size is relatively small (i.e., 30).
In sum, both the Cases III and IV corrected ds performed appropri- ately across the simulation conditions. The Case IV corrected d slightly outperformed the Case III corrected d, given that the nonlinear r-to-d transformations required in Case III appeared to contaminate the accu- racy of the correction. When the restricted sample size was 100 or above, the Case IV corrected d achieved good estimates of its true value, even when the selection scenario was very stringent (e.g., 10% selection ratio, 9:1 proportion of unrestricted applicants in two groups, .80 correlation between S and Xt, .60 unrestricted reliability). A smaller restricted sample size (e.g., 30) still produced reasonable Case IV adjusted d values, when the proportion of unrestricted applicants was not severely unbalanced (i.e., 7:3), while other factors were held constant.
Real-World Example 1: Ethnic Group Differences on a Cognitive Ability Test
This section illustrates how the Case IV correction can be used based on the information given in two published studies and shows the influence of the Case IV correction in evaluating subgroup differences. The Case III adjusted d is not presented, given that the parameters associated with the third variable Z (i.e., u Z , dZi , and r X Zi in Equation 2) are unknown in these studies. Note that the interpretation of the Case IV corrected d is based on the assumption that the effect of RR from S to G is fully mediated by the true score of variable X (i.e., cognitive ability in Example 1 and hardiness in Example 2).
Berry, Cullen, and Meyer (2014) evaluated whether there were ethnic subgroup differences on the General Aptitude Test Battery (GATB). The researchers obtained a GATB dataset from a large manufactory organiza- tion in the United States. Of the 8,837 applicants for the sales job, 1,417 received a job offer after a selection process assessing their interview per- formance, cognitive ability, and so forth. The organization provided the test scores for both the applicant and incumbent samples in Berry et al., and therefore, this study can compare the performance of the Case IV adjusted d with that of the unrestricted d obtained in the applicant pool.
According to Berry et al. (2014, table 5, p. 28), the Caucasian/African- American d for the GATB was .39 for an incumbent sample. When one seeks to generalize these results to the applicant population, one can ad- just for the bias through the Case IV correction. First, the unrestricted reliability of GATB was found to be .81 based on 515 validation studies across 30 years (Hartigan & Wigdor, 1989), and the SD ratio was reported as u X = .46/.51 = .90 in Berry et al. Hence ut becomes .88 through Equation 3. Second, given the SD ratio (.90) and the unrestricted
JOHNSON CHING-HONG LI 921
reliability (.81), the restricted reliability is found to be .77 through Equa- tion 4; this value, together with the restricted d (.39), can be plugged into Equation 5 to obtain the d corrected for unreliability, that is, dc = .39/
√ .77 = .45. Third, given that the proportion of Caucasian ap-
plicants was .96 (i.e., P = .96 and Q = .04) in Berry et al., and the values of ut and dc were known (i.e., .88 and .45), the r adjusted for Case IV can be estimated through Equation 6, and it becomes .10. Fourth, this r can be converted back to the d metric through Equation 7, thereby producing dc4 = .51, which represents a moderate effect size. Note that Berry et al. had the data for the applicant pool, and the d was reported as .49, which became .54 when corrected for unreliability (i.e., .49/
√ .81), which
is highly comparable to dc4 (i.e., .51). Following the same procedure, we can also obtain the d corrected
for Case IV for the Caucasian/Hispanic subgroup difference. Berry et al. (2014) reported that the d was .24 for the incumbent sample. Given that the proposed Case IV correction, dc4, was found to be .32 (details of the calculation are available upon request), which is comparable with the d = .28 for the applicant sample in Berry et al.’s study, and the value became .31 when corrected for unreliability (i.e., .28/
√ .81).
To summarize, the Case IV adjusted ds for the Caucasian/African- American and Caucasian/Hispanic subgroup differences in cognitive abil- ity are found to be highly comparable with the d values obtained in the applicant sample in Berry et al. (2014), meaning that the proposed Case IV correction is a trustworthy method that adjusts for the downward bias arising from RR.
Real-World Example 2: Military Graduate/Nongraduate Differences in Hardiness
In the literature on military training, an important personality di- mension related to performance is psychological hardiness (Hystad, Eid, Laberg, & Bartone, 2011). Hardiness is a personality style associated with resilience and success under a range of stressful conditions (Bartone, 1999; Kobasa, 1979), and it has been commonly used to evaluate per- formance of military specialties across the globe (e.g., Bartone, Roland, Picano, & Williams, 2008; Hystad et al., 2011). People with a strong sense of hardiness are more committed to life and work even under highly challenging and stressful situations.
Bartone et al. (2008) examined candidates who had applied to join the United States Special Forces Army, which is regarded as the most elite and challenging department. They were required to go through a num- ber of competitive selection processes (e.g., a highly demanding 4-week selection and assessment course). In Bartone et al., the Army recruited
922 PERSONNEL PSYCHOLOGY
1,138 candidates and found that 637 (i.e., 56.0%) passed the course, whereas 501 (i.e., 44.0%) did not. Bartone et al. reported a d of .24 for graduates versus nongraduates on a measure of hardiness. The d increased to .28 when corrected for unreliability (i.e., .24/
√ .73).
However, the participants in Bartone (2008) had previously been se- lected into the Army based on a series of tests (e.g., an interview, a personality measure) used to assess applicants’ suitability for the Army; this forms a judgment variable for evaluating their suitability to succeed in the job. In fact, Bartone (1991) presented the psychometric properties of the hardiness scale in 16 studies, which consisted of both civilian and military samples with the total sample size of 21,938 participants. These samples are sufficiently comprehensive in representing the characteristics of an unrestricted normative sample. In Bartone the pooled normative SD for the hardiness scale across these studies was 9.95, and the pooled nor- mative coefficient alpha (as a measure of reliability) was found to be .78. By contrast, the restricted SD was only 6.10, and the restricted coefficient alpha was only .73.
If a researcher is interested in evaluating whether hardiness generally predicts training performance, correcting for prior selection on hardiness appears to be appropriate. First, given u X = 6.10
/ 9.95 = .61 and r X X a =
.78, ut becomes .45 in Equation 3. Second, given r X X i = .73, dc becomes
.28, in Equation 5. Third, given that dc = .28, ut = .45, and P and Q are assumed to be .50,8 we can use Equation 6 to obtain an estimate of rc4, which equals .30 in this example. Fourth, we can convert rc4 back to the d metric using Equation 7, and hence the final estimate of dc4 is .63, which is substantially larger than the restricted dr (i.e., .24) and the restricted dr corrected for unreliability (i.e., .28).
The dc4 statistic (i.e., .63) has implications for understanding the utility of self-report psychological tests in military training. This is because many previous studies have not succeeded in observing obvious differences be- tween successful and unsuccessful military graduates in their self-report skills, as evidenced by the small coefficients (i.e., ds ranged from .10 to .30). Indeed, some studies (e.g., Hystad et al., 2011) have questioned whether self-report psychological tests are useful predictors of military training performance. The present findings provide a solid foundation for
8Ideally, we would know the unrestricted proportion of graduates versus nongraduates in the applicant population (i.e., all Army soldiers who apply for the Special Forces Unit). However, it was not possible to determine the unrestricted proportions because the training is administered after soldiers have been preliminarily accepted into the Special Forces. Given that approximately 50% of new Special Forces soldiers pass the training (Bartone et al., 2008), it was assumed that if all soldiers had an opportunity to participate in the training, the passing rate would be similar (i.e., 50%).
JOHNSON CHING-HONG LI 923
pursuing research into and developing self-report psychological instru- ments in military training.
Discussion and Conclusions
This paper has developed a bias-correction procedure for d when a sample is subject to the Case IV indirect RR, which is regarded as a more realistic and practical restriction model in management and psychology. The simulation results showed that the Case IV adjusted d outperformed the conventional Case III and uncorrected ds, when the model assumption is met (i.e., all effects of RR that are caused by S to G are fully mediated by Xt). Two real-world examples were also included to present the impact of RR on d and the Case IV adjusted d.
The proposed Case IV correction for d is important because evaluating the d values based on incumbent samples may be misleading. For example, selection researchers and practitioners typically are interested in whether selection procedures (e.g., cognitive ability tests) will produce subgroup differences among job applicants. In the first real-world example, the incumbent-based Caucasian/African-American and Caucasian/Hispanic ds were .39 and .24, respectively, prior to correction, and .49 and .32 after correction. Hence, the magnitude of ethnic subgroup differences is somewhat larger when the population of interest (i.e., job applicants) is considered. As Bobko et al. (2001) noted, the lack of emphasis on the applicant or population level of subgroup analysis may retard theory development and inhibit equal employment opportunity.
The issue of RR also is relevant beyond the selection context. For example, organizational behavior researchers may encounter job or or- ganizational factors that restrict the range of scores (e.g., on measures of autonomy, decision-making) relative to the general population of em- ployed individuals. In the second real-world example, human resources or training researchers may wonder whether hardiness generally predicts training performance or job performance. In this case, correcting for prior selection on hardiness (or selection on predictors that correlate with har- diness) appears to be appropriate. The results revealed a d of .28 between graduates and nongraduates on a measure of hardiness based on a restricted trainee sample in Bartone et al. (2008). If the researchers are interested in examining whether hardiness generally predicts training performance, the d statistic becomes .63 when corrected for Case IV RR, thereby showing a great difference in understanding the predictive validity of hardiness.
It is important to note that the decision of whether or not to correct es- timates for RR as well as which correction procedure to use depend on the population that researchers seek to generalize to or in which they are interested. For example, some researchers are interested in whether
924 PERSONNEL PSYCHOLOGY
current employees from different subgroups (e.g., ethnicity, gender) pos- sess similar levels of certain attributes (e.g., personality). In this case, there is no need to adjust for the d value obtained in an incumbent sample. In the second real-world example, if researchers are interested in the magnitude of subgroup differences that will manifest when hardiness is used to de- termine whether newly hired employees or trainees will make it through training, the researchers do not need to correct the observed d for RR. In this situation, it may be inappropriate to correct the observed d back to the general or applicant population if the organization plans to continue to select employees on predictors that may restrict the variance in hardi- ness among new hires entering training. Generally, when researchers are concerned with how job incumbents think, feel, and behave, and may not necessarily want or need to generalize their incumbent-based ds to the population of individuals who apply to work in an organization, they do not need to perform the correction.
Limitations and Directions for Future Research
A first area of future research involves reexamining the incumbent- based ds in published primary or meta-analytic studies. Bobko and Roth (2013) found that the majority of the ds reported in the existing literature were confusing because they were often estimated based on an incumbent sample rather than an applicant sample. This study only reanalyzed the d values reported in two published studies, and hence, additional research is required to provide a more complete picture of the subgroup differences in predictors of job performance. It is hoped that the effect of this study will be comparable to Hunter et al. (2006) for Case IV r, that is, changing the norm among researchers and practitioners to report the restricted, Case II, and Case IV adjusted ds.
A second area lies in investigating the sampling distribution of the Case IV adjusted d. This study evaluated the point estimate; by extension, the sampling distribution, as well as the confidence intervals, can provide further information about the precision of the sample estimate. Given that the bootstrap procedure has recently been used to estimate the confidence intervals of r corrected for RR (e.g., Chan & Chan, 2004; Li et al., 2011), additional research can evaluate the accuracy of the bootstrap-based con- fidence intervals for d.
Third, to use the Case IV correction, researchers must have a reason- able estimate of the unrestricted SD of scores on the focal variable (i.e., SX). However, these values are not always available and, as an anonymous reviewer noted, using an inaccurate estimate of SX will transfer additional bias to the corrected d. Thus, future research might explore the impact of an inaccurate estimate of SX on the correction.
JOHNSON CHING-HONG LI 925
A fourth area of future research lies in evaluating the performance of the Case IV adjusted d when the model assumption is violated. Although simulation findings (e.g., Le & Schmidt, 2006; Li et al., 2013) showed that the Case IV adjusted r still outperformed the uncorrected r and Case II adjusted r when the model assumption was violated in practice, additional studies can provide further empirical evidence of the accuracy of the Case IV adjusted d given different degrees of violation. For example, one may separate the total RR effect from S to G via two paths: direct (i.e., S to G) and indirect (i.e., S to G via Xt). The Case IV model assumes that 100% of the total RR effect should go through the indirect path. Further research can examine whether the Case IV adjusted d performs reasonably if other proportions (e.g., 50% to 90%) of the total RR effect is allowed to go through the indirect path, which implies that 50% to 10% of the total RR effect will go through the direct path (Le & Schmidt, 2006).
REFERENCES
Andre T, Hegland S. (1998). Range restriction, outliers, and the use of the graduate record examination to predict graduate school performance. American Psychologist, 53, 574–575. doi:10.1037/0003-066X.53.5.574
Banks GC, Batchelor JH, McDaniel MA. (2010). Smarter people are (a bit) more sym- metrical: A meta-analysis of the relationship between intelligence and fluctuating asymmetry. Intelligence, 38, 393–401. doi:10.1016/j.intell.2010.04.003
Bartone PT. (1991, August). Development and validation of a short hardiness measure. Paper presented at the third annual convention of the American Psychological As- sociation, Washington, DC.
Bartone PT. (1999). Hardiness protects against war-related stress in Army Reserve Forces. Consulting Psychology Journal, 51, 72–82. doi:10.1037/1061-4087.51.2.72
Bartone PT, Roland RR, Picano JJ, Williams TJ. (2008). Psychological hardiness predicts success in US Army Special Forces candidates. International Journal of Selection and Assessment, 16, 78–81. doi:10.1111/j.1468-2389.2008.00412.x
Berry CM, Cullen MJ, Meyer JM. (2014). Racial/ethnic subgroup differences in cognitive ability test range restriction: Implications for differential validity. Journal of Applied Psychology, 99, 21–37. doi:10.1037/a0034376
Bobko P, Roth PL. (2013), Reviewing, categorizing, and analyzing the literature on black- white mean differences for predictors of job performance: Verifying some per- ceptions and updating/correcting others. PERSONNEL PSYCHOLOGY, 66, 91–126. doi:10.1111/peps.12007
Bobko P, Roth PL, Bobko C. (2001). Correcting the effect size of d for range restriction and unreliability. Organizational Research Methods, 4, 46–61. doi:10.1177/109442810141003
Brockwell PJ, Davis RA. (2002). Introduction to time series and forecasting (2nd ed.). New York, NY: Springer Science + Business Media, Inc.
Burke, MJ, Normand J, Doran LI. (1989). Estimating unrestricted population parameters from restricted sample data in employment testing. Applied Psychological Measure- ment, 13, 161–166. doi:10.1177/014662168901300206
926 PERSONNEL PSYCHOLOGY
Chan W, Chan, DW-L. (2004). Bootstrap standard error and confidence intervals for the cor- relation corrected for range restriction: A simulation study. Psychological Methods, 9, 369–385. doi:10.1037/1082-989X.9.3.369
Christian MS, Bradley JC, Wallace JC, Burke MJ. (2009). Workplace safety: A meta- analysis of the roles of person and situation factors. Journal of Applied Psychology, 94, 1103–1127. doi:10.1037/a0016172
Cohen J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Erlbaum.
Enterprise and Industry Publications. (2005). The new SME definition. User guide and model declaration. Retrieved from http://ec.europa.eu/enterprise/policies/ sme/files/sme_definition/sme_user_guide_en.pdf
Fife DA, Mendoza JL, Terry R. (2012). The assessment of reliability under range restriction: A comparison of α, ω, and test-retest reliability for dichotomous data. Educational and Psychological Measurement, 72, 862–888. doi:10.1177/0013164411430225
Fife DA, Mendoza JL, Terry R. (2013). Revisiting Case IV: A reassessment of bias and standard errors of Case IV under range restriction. British Journal of Mathematical and Statistical Psychology, 66, 521–542. doi: 10.1111/j.2044-8317.2012.02060.x
Flores BE. (1986). A pragmatic view of accuracy measurement in forecasting. Omega- International Journal of Management Science, 14, 93–98. doi:10.1016/0305- 0483(86)90013-7
Hartigan JA, Wigdor AK. (1989). Fairness in employment testing: Validity generalization, minority issues, and the general aptitude test battery. Washington D.C.: National Academy Press.
Hunter JE, Schmidt FL. (2004). Methods of meta-analysis: Correcting error and bias in research findings (2nd ed.). Thousand Oaks, CA: Sage.
Hunter JE, Schmidt FL, Le H. (2006). Implications of direct and indirect range restriction for meta-analysis methods and findings. Journal of Applied Psychology, 91, 594–612. doi:10.1037/0021-9010.91.3.594
Hystad SW, Eid JL, Jon C, Bartone PT. (2011). Psychological hardiness predicts admis- sion into Norwegian military officer schools. Military Psychology, 23, 381–389. doi:10.1080/08995605.2011.589333
Kobasa SC. (1979). Stressful life events, personality, and health: Inquiry into hardi- ness. Journal of Personality and Social Psychology, 37, 1–11. doi:10.1037/0022- 3514.37.1.1
Le H, Schmidt FL. (2006). Correcting indirect range restriction in meta-analysis: Testing a new analytic procedure. Psychological Methods, 11, 416–438. doi:10.1037/1082- 989X.11.4.416
Li JC-H, Chan W, Cui Y. (2011). Bootstrap standard error and confidence intervals for the correlations corrected for indirect range restriction. British Journal of Mathematical and Statistical Psychology, 64, 367–387. doi:10.1348/2044-8317.002007
Li JC-H, Cui Y, Chan W. (2013). Bootstrap confidence intervals for the mean correla- tion corrected for Case IV range restriction: A more adequate procedure for meta- analysis. Journal of Applied Psychology, 98, 183–193. doi:10.1037/a0029946
Mendoza JL, Mumford M. (1987). Corrections for attenuation and range re- striction on the predictor. Journal of Educational Statistics, 12, 282–293. doi:10.3102/10769986012003282
Mooney CZ. (1997). Monte Carlo simulation. Sage University Paper series on Quantitative Applications in the Social Sciences, Series No. 07–116. Thousand Oaks, CA: Sage.
Roth PL, Van Iddekinge CH, Huffcutt AI, Eidson CE., Jr., Bobko P. (2002). Corrections for range restriction in structured interview ethnic group differences: The values may
JOHNSON CHING-HONG LI 927
be larger than researchers thought. Journal of Applied Psychology, 87, 369–376. doi:10.1037/0021-9010.87.2.369
Sackett PR, Yang H. (2000). Correction for range restriction: An expanded typology. Journal of Applied Psychology, 85, 112–118. doi:10.1037//0021-9010.85.1.112
Schmidt FL, Oh I-S, Le H. (2006). Increasing the accuracy of corrections for range re- striction: Implications for selection procedure validities and other research results. PERSONNEL PSYCHOLOGY, 59, 281–305. doi:10.1111/j.1744-6570.2006.00037.x
Thorndike RL. (1949). Personnel selection. New York, NY: Wiley. U.S. Equal Opportunity Employment Commission, U.S. Civil Service Commission, U.S.
Department of Labor, U.S. Department of Justice (1978). Uniform Guidelines on employee selection procedures. Federal Register, 43(4), 38295–38309.
Weekley JA, Ployhart RE., Harold CM. (2004). Personality and situational judg- ment tests across applicant and incumbent contexts: An examination of valid- ity, measurement, and subgroup differences. Human Performance, 17, 433–461. doi:10.1080/08959280802137820
Yang H, Sackett PR, Nho Y. (2004). Developing a procedure to correct for range restriction that involves both institutional selection and applicants’ rejection of job offers. Organizational Research Methods, 7, 442–455. doi:10.1177/1094428104269054