Assessment Guide

profileEverleigh
issuesofculturaldiversityasitappliestopsychologicalassessment..pdf

Methodological and Statistical Advances in the Consideration of Cultural Diversity in Assessment: A Critical Review of Group Classification and

Measurement Invariance Testing

Kyunghee Han, Stephen M. Colarelli, and Nathan C. Weed Central Michigan University

One of the most important considerations in psychological and educational assessment is the extent to which a test is free of bias and fair for groups with diverse backgrounds. Establishing measurement invariance (MI) of a test or items is a prerequisite for meaningful comparisons across groups as it ensures that test items do not function differently across groups. Demonstration of MI is particularly important in assessment settings where test scores are used in decision making. In this review, we begin with an overview of test bias and fairness, followed by a discussion of issues involving group classification, focusing on categorizations of race/ethnicity and sex/gender. We then describe procedures used to establish MI, detailing steps in the implementation of multigroup confirmatory factor analysis, and discussing recent developments in alternative procedures for establishing MI, such as the alignment method and moderated nonlinear factor analysis, which accommodate reconceptualization of group categorizations. Lastly, we discuss a variety of important statistical and conceptual issues to be considered in conducting multigroup confirmatory factor analysis and related methods and conclude with some recommendations for applications of these procedures.

Public Significance Statement This article highlights some important conceptual and statistical and issues that researchers should consider in research involving MI to maximize the meaningfulness of their results. Additionally, it offers recommendations for conducting MI research with multigroup confirmatory factor analysis and related procedures.

Keywords: test bias and fairness, categorizations of race/ethnicity and sex/gender, measurement invariance, multigroup CFA

Supplemental materials: http://dx.doi.org/10.1037/pas0000731.supp

When psychological tests are used in diverse populations, it is assumed that a given test score represents the same level of the underlying construct across groups and predicts the same outcome score. Suppose that two hypothetical examinees, a middle aged Mexican immigrant woman and a Jewish European American male college student, each produced the same score on a measure of depression. We would like to conclude that the examinees exhibit the same severity and breadth of depression symptoms and that their therapists would rate them similarly on

relevant behavioral and symptom measures. If empirical evi- dence indicates otherwise, and such conclusions are not justi- fied, scores on the measure are said to be biased.

Although it has been defined variously, a representative definition refers to psychometric bias as “systematic error in estimation of a value”). A biased test “is one that systematically overestimates or underestimates the value of the variable it is intended to assess” due to group membership, such as ethnicity or gender (Reynolds & Suzuki, 2013, p. 83). The “value of the variable it is intended to assess” can either be a “true score” (see S1 in the online supplemental materials) on the latent construct or a score on a specified criterion measure. The former appli- cation concerns what is sometimes termed measurement bias, in which the relationship between test scores and the latent attri- bute that these test scores measure varies for different groups (Borsboom, Romejin, & Wicherts, 2008; Millsap, 1997), whereas the latter application concerns what is referred to as predictive bias, which entails systematic inaccuracies in the prediction of a criterion from a test depending upon group membership (Clearly, 1968; Millsap, 1997).

Kyunghee Han, Stephen M. Colarelli, and Nathan C. Weed, Department of Psychology, Central Michigan University.

This article has not been published elsewhere, nor has it been submitted simultaneously for publication elsewhere. The author(s) de- clared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article. The author(s) received no funding for this study.

Correspondence concerning this article should be addressed to Kyung- hee Han, Department of Psychology, Central Michigan University, Mount Pleasant, MI 48859. E-mail: [email protected]

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

Psychological Assessment © 2019 American Psychological Association 2019, Vol. 31, No. 12, 1481–1496 1040-3590/19/$12.00 http://dx.doi.org/10.1037/pas0000731

1481

Test bias should not be confused with test fairness. Although the two concepts have been used interchangeably at times (e.g., Hunter & Schmidt, 1976), test fairness entails a broader and more sub- jective evaluation of assessment outcomes from perspectives of social justice (Kline, 2013), whereas test bias is an empirical property of test scores, estimated statistically (Jensen, 1980). Ap- praisals of test fairness include multifaceted aspects of the assess- ment process, lack of test bias being only one facet (American Educational Research Association, American Psychological Asso- ciation [APA], & National Council on Measurement in Education, 2014; Society for Industrial Organizational Psychology, 2018; see S2 in the online supplemental materials).

In the example above, the measure of depression may be unfair for the Mexican female client if an English language version of the measure was used without evaluating her English proficiency, if her score was derived using American norms only, if computerized administration was used, or if use of the test leads her to be less likely than members of other groups to be hired for a job. Although test bias is not a necessary condition for test unfairness to exist, it may be a sufficient condition (Kline, 2013). Accordingly, it is especially important to evaluate whether test scores are biased against vulnerable groups.

The evaluation of test bias and test fairness each entails a comparison of one group of people with another. While asking the question, “Is a test biased?” we are also implicitly asking “against or for which group?” Similarly, if we are concerned about using a test fairly, we must ask: are the outcomes based on the results of the test apportioned fairly to groups of people who have taken the test? Thus, the categorization of people into distinct groups is a sine qua non of many aspects of psychological assessment re- search. Racial/ethnic and sex/gender categories are prominent fea- tures of the social, cultural, and political landscapes in the United States (e.g., Helms, 2006; Hyde, Bigler, Joel, Tate, & van Anders, 2019; Jensen, 1980; Newman, Hanges, & Outtz, 2007), and have therefore been the most commonly studied group variables in bias research (e.g., Warne, Yoon, & Price, 2014). Most of the initial research on and debates about test bias and fairness in the United States stemmed from political movements addressing race and sex discrimination (e.g., Sackett & Wilk, 1994). In service of pressing research on questions of discrimination and economic inequality, it thus became commonplace among psychologists and social scien- tists to categorize people crudely into groups (based primarily on race, ethnicity, and sex/gender) without much thought to the mean- ing and validity of those categorizations (e.g., Hyde et al., 2019; Yee, 1983; Yee, Fairchild, Weizmann, & Wyatt, 1993). This has changed somewhat over the past two decades as scholarship by psychologists and others has increasingly focused on nuances of identity, multiculturalism, intersectionality, and multiple position- alities (Cole, 2009; Song, 2017). This scholarship has emphasized that racial, ethnic, and gender classifications can be complex, ambiguous, and debatable—and that identities are often self- constructed and can be fluid (Helms, 2006; Hyde et al., 2019). The first goal of this review, therefore, is to overview contemporary issues involving race/ethnicity and sex/gender classifications in bias research and to describe alternative approaches to the mea- surement of these variables.

The psychometric methods used to examine test bias usually depend on the definition of test bias operating for a given appli- cation. Evaluating predictive bias (i.e., establishing predictive in-

variance) often involves regressing total scores from a criterion measure onto total scores on the measure of interest, and compar- ing regression slopes and intercepts across groups (Clearly, 1968). Evaluating measurement bias (i.e., establishing measurement in- variance [MI]) often necessitates more advanced quantitative meth- ods, such as confirmatory factor analysis (CFA) or methods deriv- ing from item response theory, to compare the properties of item scores and scores on latent variables across different groups. Multigroup confirmatory factor analysis (MGCFA) has been one of the most commonly used techniques to examine MI (Davidov, Meuleman, Cieciuch, Schmidt, & Billiet, 2014) because it pro- vides a comprehensive framework for evaluating different forms of MI. The second goal of this review is to provide a broad overview of MGCFA and related procedures and their relevance to psychological assessment.

Although MGCFA is a well-established procedure in the eval- uation of MI, it has limitations. MGCFA is not an optimal method for conducting MI tests when many groups are involved. More- over, the grouping variable in MGCFA must be categorical, and therefore does not permit MI testing with continuous grouping variables (e.g., age). As modern research questions may require MI testing across many groups, and with continuous reconceptualiza- tions of some of the grouping variables (e.g., gender), more flex- ible techniques are needed. Our third goal, therefore, is to describe two recent alternative methods for MI testing, the alignment method and moderated nonlinear factor analysis, that aim to over- come these limitations. We conclude the review with a discussion of some important statistical and conceptual issues to be consid- ered when evaluating MI, and include a list of recommended practices.

Group Classifications Used in Bias Research

Racial and Ethnic Classifications

Race and ethnicity (see S3 in the online supplemental materials) are conceptually vague and empirically complex social constructs that have been examined by numerous researchers across many disciplines (Betancourt & López, 1993; Helms, Jernigan, & Mascher, 2005; Yee et al., 1993). Consider race. As a biological concept, it is essentially meaningless. In most cases, there is more genetic variation within so-called racial groups than between racial groups (Witherspoon et al., 2007). Even if we allow race to be defined by a combination of specific morphological features and ancestry, few “racial” populations are pure (Gibbons, 2017). Most are mixed—like real numbers, with infinite gradations. For exam- ple, although many African Americans trace their ancestry to West Africa, about 20% to 30% of their genetic heritage is from Euro- pean and American Indian ancestors (Parra et al., 1998), and racial admixture continues as the frequency of interracial marriages increases (Rosenfeld, 2006; U.S. Census Bureau, 2008). Even if one were to accept race as a combination of biological features and cultural and social identities (shared cultural heritage, hardships, and discrimination), there is the problem of degree. For example, while many Black Americans share social and cultural identities based on roots in American slavery and racial discrimination, not all do, such as recent Black immigrants from the Caribbean. Racial and ethnic classifications are often conflated. In psychological research, “Asian” is commonly used both as a cultural (Nisbett,

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1482 HAN, COLARELLI, AND WEED

Peng, Choi, & Norenzayan, 2001) and racial category (Rushton, 1994). Yet it is a catch-all term based primarily on geography. It typically refers to people from (or whose ancestors are from) South, Southeast, and Eastern Asia. The term Hispanic often conflates linguistic, cultural, and sometimes even morphological features (Humes, Jones, & Ramirez, 2010).

In public policy, mixtures of racial (or ethnic) background has only recently begun to be addressed. The U.S. Census, for exam- ple, did not include a multiracial category until 2000 (Nobles, 2000). We are only beginning to see assessment studies that parse people from traditional broad groupings into smaller, more mean- ingful and homogeneous groups. In one of the few studies that identified different types of Asians, Appel, Huang, Ai, and Lin (2011) found significant (and sometimes major) differences in physical, behavioral, and mental health problems among Chinese, Vietnamese, and Filipina women in the U.S. More recently, Tal- helm et al. (2014) found important differences in culture and thought patterns within only one Asian country, China. People in northern China were significantly more individualistic than those in southern China, who were more collectivistic. With current and historical farming practices as their theoretical centerpiece, they examined farming practices as causal factors. In northern China wheat has been farmed as a staple crop for millennia, whereas in southern China rice has been (and is) the staple crop. Talhelm et al. argued that the farming practices required by these two crops required different types of social organization that, over time, influenced cultural values and cognition. The work by Talhelm and colleagues is important because it is one of the first studies to show—along with a powerful theoretical rationale—that there are important cultural differences between people from what has typ- ically been thought of as a relatively homogeneous racial and cultural group.

In another seminal article, Gelfand and colleagues (2011) ex- amined the looseness-tightness dimension of cultures in 33 coun- tries. This dimension reflects the strength of norms and the toler- ance of deviant behavior. Loose cultures have weaker norms and are more tolerant of deviant behavior. While there was substantial variation between countries, there was still considerable variation among countries typically considered “Asian.” Hong Kong was the loosest (6.3), while Malaysia was the tightest (11.8), with the People’s Republic of China (7.9), Japan (8.6), South Korea (10), and Singapore (10.4) in between. To say that all Asian countries are culturally similar is untenable when, for example, Malaysian culture is 88% tighter than Hong Kong culture.

Gender Classifications

Binary categories. Gender (see S4 in the online supplemental materials) differences or similarities on psychological constructs have been a widely researched topic (e.g., Feingold, 1994; Hyde, 2005) since the 1970s (Eagly & Riger, 2014), with many studies assuming the existence of clear qualitative and quantitative differ- ences between genders (Brizendine, 2006; Ruigrok et al., 2014). These include numerous studies examining bias or MI across gender (e.g., Baker & Mason, 2010; Linn & Kessel, 2010) in which researchers have tended to employ a binary categorization of gender. However, the binary gender categorization and the presumption of qualitative gender difference have recently been

challenged (APA, 2015, 2017; Richards et al., 2016; Richards, Bouman, & Barker, 2017; Schellenberg & Kaiser, 2018).

In a recent article, Hyde and colleagues (2019) reviewed em- pirical findings from five disciplines and challenged the legitimacy of binary gender classification in each: (a) neuroscience (sexual dimorphism of the human brain), (b) behavioral neuroendrocrinol- ogy (the argument for genetically fixed, nonoverlapping, sexually dimorphic hormonal systems), (c) research on psychological vari- ables (the inference of clear gender differences on psychological constructs), (d) research with transgender and nonbinary individ- uals (the assumption that gender identities and experiences are consistent with gender assigned at birth), and (e) research from developmental psychology (arguing that binary gender categories are culturally universal and unmalleable).

Nonbinary identities. There is a wide variety of nonbinary gen- der identities (APA, 2015; Hyde et al., 2019; Richards et al., 2016, 2017). Intersex individuals have physical characteristics outside the typical binary male-female categories, although most still identify their gender within the binary system (Richards & Barker, 2013). More common are people who are not physiologically intersex but who have nonbinary gender identities. Although terms are often subsumed within the umbrella terms of nonbinary or genderqueer identities, various more specific labels have been used: androgynous, mixed gender, or pangender to indicate incor- porating aspects of both male and female, but having a fixed identity; bigender or gender fluid to indicate movement between gender in a fluid way; trigender to indicate moving between more than two genders; third gender or other gender to identify a specific additional gender; gender queer to challenge the binary gender system; agender, gender neutral, genderless, nongendered, or neuter to indicate no gender (Richards et al., 2016). The frequency of people with nonbinary gender identities, while small compared to the population at large, is not trivial. For example, one Dutch study found that 4.6% of people assigned male identities at birth and 3.2% assigned female identities at birth have ambivalent gender identities (Kuyper & Wijsen, 2014). Others—people who regard themselves as asexual— have what might be called no gender identity (Carrugan, 2015).

The trend toward acceptance of diversity in gender classification has been reflected in various professional and societal contexts. Within the context of mental health care, the Diagnostic and Statistical Manual, 5th edition (American Psychiatric Association, 2013) removed the Diagnostic and Statistical Manual–IV–TR (American Psychiatric Association, 2000) diagnosis of gender identity disorder and recognized nonbinary genders within the diagnostic taxonomy. People with nonbinary gender identities have become more politically active, with visible results in recent years both in official documentation and in media coverage (Scelfo, 2015). For example, New Zealand passport holders can claim one of three gender identities: male, female, or X (other). New York, New York, has recently joined Oregon, California, Wash- ington, and New Jersey in offering a nonbinary gender marker (X) on birth certificates for residents who do not identify as male or female (Hafner, 2019). In an effort to foster an environment of inclusiveness and supporting students’ preferred form of self-identification, many universities in the United States allow students to choose preferred gender pronouns (https://www.npr.org/2015/11/08/455202525/more- universities-move-to-include-gender-neutral-pronouns). This move- ment has also been reflected in product development and marketing;

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1483METHODOLOGICAL AND STATISTICAL ADVANCES

numerous vendors specialize in gender neutral clothing or other items (e.g., https://www.lgbtqnation.com/2017/07/target-hits-back-binary- gender-neutral-clothing-line/; http://www.foxnews.com/us/2015/08/ 13/target-going-gender-neutral-in-some-sections.html).

Emerging Recommendations Regarding Race/Ethnicity and Sex/Gender Categorizations

Results of studies of test bias rest to a considerable degree on how grouping variables such as race/ethnicity and sex/gender are operationalized. Cultural researchers have recently proposed a number of suggestions for reconceptualizing and operationalizing these grouping variables. Hyde et al. (2019) and other researchers (e.g., Bittner & Goodyear-Grant, 2017; Schellenberg & Kaiser, 2018; Tate, Ledbetter, & Youssef, 2013) warn about the costs of binary sex/gender categorization in research and provide helpful suggestions for reconceptualizing and measuring sex/gender: (a) providing multiple categories of gender identity within available response options (“female,” “male,” “transgender female,” “trans- gender male,” “genderqueer” [click for more options: “agender,” “bigender,” etc.], “intersex,” and other [specify]); (b) asking about both birth-assigned and self-assigned gender/sex identities; (c) using an open-ended response format (e.g., “what is your gen- der?”); (d) treating gender/sex constructs as multidimensional (use of multiple measures of gender/sex identities, stereotypes, or be- haviors), dynamic, and continuous (e.g., “In the past 12 months, have you thought of yourself as a man?” “In the past 12 months, have you thought of yourself as a woman?” “How would you rate yourself on the continuum from 0 to 100 regarding “maleness?” “How would you rate yourself on the continuum from 0 to 100 regarding “femaleness?”); and (e) asking about gender identity at the end of a study so that it does not influence responding.

Numerous researchers (e.g., Cole, 2009; Helms & colleagues, 2005; Yee, 1983) have proposed ways to reconceptualize race/ ethnicity in research. When researchers separate people into groups and compare them, careful attention must be given to (a) what the categories mean, (b) attempt to ensure their homogeneity, (c) how the categories are theoretically related to the substantive constructs under study, (d) treating race/ethnicity constructs as multidimensional (use of multiple measures of race/ethnicity iden- tities, stereotypes, socioeconomic status [SES], or racism), and (d) intersectional views of the categories (see below). Racial catego- rization should be avoided in research without clear conceptual reasons.

It is clear that people fall into multiple grouping categories— each person is simultaneously a member of a gender, race, ethnic- ity, class, and sexual orientation (APA, 2017). The concept of “intersectionality” has been developed by feminist and critical race theorists (Cole, 2009; Eagly & Riger, 2014) to encourage research- ers to understand individuals from “a vast array of cultural, struc- tural, sociobiological, economic, and social contexts by which individuals are shaped and with which they identify” (APA, 2017, p. 19). How might we conceptualize and examine test bias across combinations of categories? Studies of psychometric bias are typically conducted from the perspective of one group membership at a time mainly due to methodological complications in imple- menting multiple group categories into testable models. Research teams such as Corral and Landrine (2010) and Else-Quest and Hyde (2016) have provided recommendations for incorporating

intersectional approaches to numerous facets of research (i.e., theory, design, sampling techniques, measurement, data analytic strategies, and interpretation and framing). More work building on these recommendations will be needed to meet the research chal- lenges of intersectionality in the context of bias testing.

Testing Bias

As mentioned earlier, the statistical methods used to examine test bias usually depend on the definition of test bias operating for a given application (see S5 in the online supplemental materials). If test scores can be used to predict some future outcome (a criterion), they are said to demonstrate predictive validity (Clearly, 1968). In general, a test is considered not biased (demonstrating “predictive invariance”) if its scores predict future outcomes equally well for individuals from different groups. However, if test scores are better at predicting outcomes for some groups than for others, the test scores are said to reflect differential predictive validity or slope bias (Camilli, 2006). If test scores systematically overpredict or underpredict criterion scores for one group relative to another, the scores are said to reflect intercept bias (see S6 in the online supplemental materials).

Relative to predictive invariance, MI has been investigated more frequently and more rigorously in recent years due to advances in statistical procedures, including the development of software that allows researchers to carry out MI testing more easily. Researchers have examined how predictive invariance is related to MI and argued that evidence for one form of invariance is not evidence in support of the other, but may in some cases serve as evidence against the other (Borsboom et al., 2008; see Millsap, 1995, 1997, 2007, for a mathematical proof of this argument). Investigating both forms of invariance simultaneously in the same study is ideal, but seldom achieved (Millsap, 2007), with some exceptions (Cul- hane, Morera, Watson, & Millsap, 2009; Wright, Kutschenko, Bush, Hannum, & Braddy, 2015). Researchers tend to prefer to evaluate one form over the other depending upon the context of assessment. For example, in assessment research related to per- sonnel section, evaluating predictive invariance is favored over evaluating MI (Borsboom, 2006; Society for Industrial Organiza- tional Psychology, 2018). However, even in this context, item- level MI analyses are recommended when conducting cross- cultural research involving linguistically different populations.

The ease of evaluating predictive bias varies greatly depend- ing upon the psychometric integrity of the criterion variable in question. It is possible for differential predictive validity to reflect statistical artifact due to measurement error in scores on the criterion and predictor (Kane & Mroch, 2010; Warne et al., 2014). Moreover, other confounding variables (e.g., time gap between obtaining test and criterion data) need to be controlled across groups when examining predictive invariance. These complications (see Borsboom et al., 2008, for a comprehensive discussion) are especially problematic in cross-cultural research involving translated measures. Therefore, Borsboom and col- league (2008) argued that psychologists should favor MI testing over predictive invariance testing in defining and measuring bias. For the rest of this article, we focus on procedures asso- ciated with MI.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1484 HAN, COLARELLI, AND WEED

Testing Measurement Invariance: Multigroup Confirmatory Factor Analysis

The popularity of research involving MI has increased exponen- tially in recent years. An informal PsycINFO database search of titles, keywords, and abstracts associated with peer reviewed jour- nal articles from 1980 to January 31, 2018, on group and mea- surement and invariance or equivalence located 55 unique articles published from 1980 to 1989, 91 articles published from 1990 to 1999, 459 articles published from 2000 to 2009, and 1,504 articles published from 2010 to 2017. The increased interest in this topic may reflect globalization of the social sciences, increased empha- sis on fair assessment for diverse populations, and statistical ad- vances permitting rigorous testing of MI (Davidov et al., 2014; Sass, 2011; Vandenberg, 2002).

Since its development by Jöreskog (1971) and Sörbom (1974) for use with continuous indicators, MI testing via MGCFA has substantially evolved, and has been applied in an increasingly wide variety of research contexts (Sass, Schmitt, & Marsh, 2014). Greiff and Scherer (2018) showed that MGCFA was the method of choice in more than 70% of the MI studies published in the European Journal of Personality Assessment from January 1995 to August 2017. We overview MGCFA procedures here.

Levels of Invariance

MI is evaluated using MGCFA by first identifying a well-fitting “baseline” factor model, and then applying a series of hierarchi- cally ordered cross-sample tests of the equivalence of (a) the factor structure or form (i.e., the pattern of the causal links between indicator and factors; this is sometimes referred to as a test of configural invariance), (b) factor loadings (i.e., the magnitude and direction of the causal relationships between indicators and fac- tors; metric invariance), (c) intercepts/thresholds (i.e., the point of origin; scalar invariance), (d) unique or residual invariance (i.e., precision of each scale), (e) factor variances (i.e., variability of each dimension’s true scores), (f) factor covariances (i.e., relation- ship among true scores), and (g) factor means (i.e., means of true scores; Vandenberg, 2002). Although the number and order of the MI tests vary across studies, configural, metric, and scalar invari- ances are the most frequently examined and reported (Vandenberg & Lance, 2000). These three types of invariance, along with residual invariance, involve the measurement aspects of the model, and are considered types of MI or measurement equivalence. The other three types of invariance (variance, covariance, and mean) address characteristics of the latent variables and their interrela- tions, and are therefore said to comprise structural invariance or structural equivalence (van de Vijver, 2011). In this article, we focus principally on MI testing: configural, metric, scalar, and residual invariance.

Configural invariance. Configural invariance (Meredith, 1993; Vandenberg & Lance, 2000), the first and lowest level in the MI testing sequence, requires a demonstration that the same items or indicators load on the same factor(s) across groups. It requires mini- mal constraints to ensure that the pattern of fixed and free factor loadings are the same across groups; that is, the same items are forced to load on the same factors, but the factor loadings are allowed to vary across groups. The fit of a configural invariance test is approximately the average of the CFA fits for each group considered separately. For

example, if a specific four-factor model shows a good fit separately for two ethnic group samples, configural invariance is likely to be established, suggesting that both groups conceptualize the construct in similar way. A configurally invariant model can be used as a baseline model for further invariance testing (Vandenberg & Lance, 2000).

Metric invariance. Metric invariance, the second level in the MI testing sequence, requires that factor loadings for each item are similar across groups. Demonstration of metric invariance is some- times referred to as indicating “weak invariance” (Meredith, 1993) because it represents a minimum level of evidence at which researchers feel comfortable ascribing the property of invariance to a measure. Equal factor loadings across groups indicate that the strength of the causal relationships between items and their under- lying dimensions are the same across groups. Consequently, metric invariance suggests invariance in measurement unit, which implies that equal increases in indicator scores across groups represent equal increases in corresponding latent variable scores across groups. Metric noninvariance is observed when a latent construct is similar across groups, but some items reflect the latent construct better for members of one group than another (Chen, 2008). For example, it was shown that a Center for Epidemiological Studies– Depression item referencing “crying spells” had a much higher factor loading on a Depressive Affect factor for females than for males (e.g., Verhoeven, Sawyer, & Spence, 2013). In a cross- cultural study of substance abuse (Wang et al., 2014), items about marijuana use showed much higher factor loadings in a U.S. sample than in a Korean sample, likely due to the relative unavail- ability of marijuana in Korea. Metric noninvariance can also be observed when groups employ different response styles (Chen, 2008; Cheung & Rensvold, 2000). For example, Lee (2012) found that Koreans used fewer extreme points on a measure of aggres- sion than did Americans, attributing the differences to a relative cultural preference for modesty among Koreans. Gilman et al. (2008) found similar results in a sample of Korean adolescents on a life satisfaction scale. Avoiding extreme responses would result in a restricted range of item responses, potentially lowering cor- relations among indicators and factor loadings (Chen, 2008).

Scalar invariance. Once metric invariance has been demon- strated, scalar invariance (also known as “strong invariance”) can be examined by constraining to be equal across groups item intercepts (i.e., the predicted score of an item when its latent factor score is zero; the point of origin for continuous indicators) or item thresholds (i.e., the point where response categories shift, e.g., from strongly disagree to disagree or from false to true; item difficulties for ordered categorical indicators). Scalar invariance is said to exist when both factor loadings and intercepts/thresholds are not significantly different across groups (Vandenberg & Lance, 2000). When established, the scale is considered to possess both the same measurement units (factor loadings) and origin (intercept/ thresholds) across groups. Consequently, scalar invariance— equality of factor loadings and item intercepts across groups—is prerequi- site for the comparison of latent factor means because it permits the conclusion that group differences in latent factor means reflect true group mean differences on latent constructs (Van de Velde, Levecque, & Bracke, 2009). Group differences attributable to response styles, however, such as social desirability, acquiescence, or the use of differential reference frameworks in making judg- ments, contribute to scalar noninvariance (Chen, 2008; Cheung & Rensvold, 2000). Such response patterns might lead to systematically

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1485METHODOLOGICAL AND STATISTICAL ADVANCES

higher or lower item responses in one group compared to another, affecting observed means. For example, Locke and Baik (2009) discovered that Koreans were more acquiescent than Americans. Shou, Sellbom, and Han (2017) found that Chinese students were less likely than Americans to endorse extreme responses on items of a psychopathy measure, a cultural difference in response pattern that led to scalar noninvariance.

Residual invariance. Residual variance refers to the variance of an indicator that is not attributable to the variance of the associated latent variable (Cheung & Rensvold, 2002). Residual invariance implies that in addition to invariant factor loadings and intercepts, item uniqueness, precision, or measurement error is also invariant across groups. Residual invariance is sometimes called “strict invariance” because residual invariance is a highly con- strained model that cannot be easily be met in practice. Differences in familiarity with response format, item sentence structure, use of idiom, or cultural experiences may contribute to residual nonin- variance (Cheung & Rensvold, 2002). For example, it was shown that on average, Asian respondents scored much higher (d � 1.5) than American respondents on a scale (Variable Response Incon- sistency) that consisted of item pairs with similar content (Cheung, Song, & Zhang, 1996; Ketterer, Han, Hur, & Moon, 2010). It was also shown that the mean absolute correlations of Variable Re- sponse Inconsistency item pairs were much stronger for the Amer- ican respondents than for their Asian counterparts (Ketterer et al., 2010). Together these results indicated that Asian respondents displayed a higher level of inconsistency than did Americans in their responses to item pairs with similar content. The authors posited that these observed differences were either due to genuine cultural differences or to subtle changes in the meaning of items that were introduced during translation. In studies in which bilin- gual respondents complete two language versions of the same measure (Chung, Weed, & Han, 2006; Pudumjee, Weed, Pant, Ahluwalia, & Chakranarayan, 2016), it has been shown that items with double negatives, complicated sentence structure, or transla- tions of American idioms can produce relatively small cross- language item correlations. As establishing residual invariance permits confidence that group mean differences on the scale scores represent genuine group differences on the construct rather than on other factors (Meredith, 1993), residual invariance needs to be demonstrated prior to observed mean comparisons. However, sca- lar invariance is considered sufficient to permit meaningful com- parison of latent mean differences, as residuals are not part of the latent factors (Chen, 2008; Cheung & Lau, 2012).

Evaluation of Invariance

Establishing criteria for determining invariance is a complex task, as numerous characteristics of a model need to be considered. Characteristics of the indicators (continuous, ordered categorical, or binary) and of distributions (degree of multivariate normality) may require different types of estimation methods, which also determine different approaches for evaluating model fit. At pres- ent, an approach commonly used with models with continuous or ordered categorical indicators that do not have serious departures from multivariate normality is the maximum likelihood estimation method. Approaches used with models with categorical data will be discussed later. For determining model fit for both single group and multigroup CFA, see Sellbom and Tellegen (2019, this issue).

MI at the metric or scalar levels is said to be “full” or “com- plete” when all the relevant parameters jointly are equal across groups. As full MI is difficult to achieve, some researchers argue that when full invariance cannot be demonstrated, analyses at a subsequent level of MI testing should be allowed if a subset of indicators functions invariantly across groups, a phenomenon termed partial invariance (Byrne, Shavelson, & Muthén, 1989; Steenkamp & Baumgartner, 1998). Partial metric invariance is said to have been established when the parameters of at least two indicators per dimension (at least one invariant item aside from the referent, which is assumed to be invariant) are invariant across groups (Byrne et al., 1989; Steenkamp & Baumgartner, 1998). Vandenberg and Lance (2000) argued that a factor can be consid- ered partially invariant if the majority of items on the factor are invariant.

Problems With MGCFA and Alternative Methods

Despite the popularity of MGCFA in MI testing, it has limita- tions. Several new statistical techniques for MI testing have re- cently been developed to make MI testing less stringent by per- mitting weaker model constraints, to make the process of determining partial invariance less laborious than in MGCFA, or to make MI testing more flexible by allowing continuous and/or categorical group variables (the alignment method, Asparouhov & Muthén, 2014 and alignment within CFA, Marsh et al., 2018; moderated nonlinear factor analysis, Bauer, 2017; exploratory structural equa- tion modeling, Asparouhov & Muthén, 2009; Marsh et al., 2009 [see Sellbom & Tellegen, 2019]; multigroup Bayesian structural equation modeling, Muthén & Asparouhov, 2013; multilevel fac- tor analysis, Muthén & Asparouhov, 2017). A comprehensive comparison of some of these new techniques can be found in Hildebrandt, Lüdtke, Robitzsch, Sommer, and Wilhelm (2016); Jang et al. (2017); and Kim, Cao, Wang, and Nguyen (2017). We present brief summaries of two of these new developments: the alignment method (Asparouhov & Muthén, 2014) and moderated nonlinear factor analysis (Bauer, 2017), as these methods are most relevant to the reconceptualization of group classification dis- cussed above.

Alignment method. When researchers conduct MGCFA with many groups (e.g., examinees from 20 countries), they encounter unique challenges, as MI testing is usually conducted with two groups at a time (Kim et al., 2017). As noted above, however, there is often good reason to group examinees into smaller, more mean- ingful and homogenous samples (e.g., instead of combining Ko- reans, Chinese, Japanese, Malaysians, and Vietnamese into a sin- gle “Asian” group). As the number of groups studied increases, the number of pairwise comparisons between groups increases expo- nentially, resulting in a dramatic increase in the probability of falsely detecting noninvariance (Kim et al., 2017; Rutkowski & Svetina, 2014). Accordingly, with a large number of groups, invariance criteria suggested for two-group MGCFA are often too stringent (Rutkowski & Svetina, 2014).

Asparouhov and Muthén (2014) have recently described an alternative to MGCFA that they refer to as the “alignment method” or “alignment optimization” for use in MI testing with a large number of groups. Typically, when many groups are compared, it is difficult to establish full scalar invariance, especially for lengthy measures. In such cases, partial invariance testing is conducted by

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1486 HAN, COLARELLI, AND WEED

consulting modification indices to free the equality constraints on the measurement model one parameter at a time, identifying the model with the fewest noninvariant parameters (Byrne et al., 1989). In studies with a large number of groups and indicators, this stepwise process becomes cumbersome and is likely to capitalize on chance, producing a solution that is not replicable (Marsh et al., 2018). The goal of the alignment method is to identify the model that minimizes the number of noninvariance parameters and amount of noninvariance. Once such a model is identified, re- searchers can compare factor means across groups with minimal noninvariance.

MI testing using the alignment method begins with the assump- tion that researchers have identified the same well-fitting baseline model for all groups. As a first step the configural model is estimated in each group; all factor loadings and intercepts are freely estimated while factor means and variances are fixed at zero and one, respectively, to avoid identification problems (Asp- arouhov & Muthén, 2014). This is the best fitting model because it has no across-group parameter restrictions. Next, the factor means and variances are freed in all groups, and an alignment model is estimated using alignment optimization. Unlike MGCFA, which tests MI models sequentially (i.e., configural ¡ metric ¡ scalar), alignment optimization examines the invariance of factor loadings and intercepts in tandem. Factor means and variances for each group are computed with the approximate invariance assump- tion that noninvariance of factor loadings and intercepts should be minimized across all groups. In other words, factor means and variances for each group are determined where the sum of nonin- variance in the factor loadings and intercepts across all possible pairs of groups is minimized. This less stringent approach is in contrast to an approach that applies the exact invariance assump- tion that factor loadings and intercepts are identical across all groups. The alignment model is set to provide the same fit as the initial configural model (i.e., the best fitting model), a procedure that is described as corresponding to rotation in exploratory factor analysis models, whereby a structure is simplified by reapportion- ing loadings without compromising overall model fit. This align- ment technique also generates information regarding which param- eters are approximately invariant and which are not. If the degree of noninvariance is not very severe (i.e., at least 75% of the measurement parameters in the final model are approximately invariant; Muthén & Asparouhov, 2014), factor means of groups are said to be meaningfully comparable. Although comparisons of factor means across groups are a main goal of the alignment method, results regarding the degree of (non)invariance of the measurement parameters can offer insights into differential man- ifestation of the construct across groups.

The primary strengths of the alignment method are its ability to estimate group-specific factor means and variances under mini- mum conditions of noninvariance and its provision of detailed information about noninvariance and factor mean differences. Its major weaknesses are that (a) models with cross-loadings are not accommodated; (b) covariates cannot be included in a model, as they are not estimated; and (c) no guidelines are available regard- ing how much noninvariance is permissible for interpretation of factor mean differences. The suggested “75% rule” (Muthén & Asparouhov, 2014) needs to be validated, as it does not take into account degree and location of noninvariance (Kim et al., 2017). The alignment method appears to be best suited for testing MI

across a large number of groups and using a factor model with a relatively simple structure. Very recently, an extended, more flex- ible version of the alignment method, alignment within CFA, was introduced (Marsh et al., 2018). It is expected that alignment within CFA will be used extensively for comparisons involving many groups (e.g., see Jang et al., 2017).

Moderated nonlinear factor analysis (MNLFA). The strength of MGCFA is that it provides a well-established procedure to evaluate the invariance of all model parameters (e.g., factor load- ings, intercepts/thresholds, residuals, etc.). A major limitation, however, is that parameters are compared across one categorical grouping variable only (a “moderator,” e.g., sex [male vs. female] or race [European American vs. African American]). When the moderator is a continuous variable (e.g., age, SES), researchers sometimes categorize it (e.g., old vs. young group) for MGCFA, a practice that should be avoided due to its arbitrary borders between groups (Hildebrandt et al., 2006), loss of information about indi- vidual differences within groups, and inability to account for nonlinear relationships (MacCallum, Zhang, Preacher, & Rucker, 2002). The multiple-indicator multiple-cause (see S7 in the online supplemental materials) model allows both categorical and con- tinuous moderators in the assessment of MI, but MI is evaluated only for the factor means and item intercepts. Bauer (2017) de- veloped MNLFA, which combines the strengths of MGCFA and multiple-indicator multiple-cause models, to serve as a method of examining MI flexible enough to permit full and simultaneous assessment of MI across multiple categorical and/or continuous moderators. To the extent that grouping variables such as gender identity is reconceptualized as continuous within the context of MI testing and if the need to test multiple moderators (e.g., gender and SES) for examining intersectionalities is present, MNLFA seems an especially promising procedure.

Bauer illustrated the use of MNLFA with MI testing of mea- sures of violent and nonviolent delinquent behaviors across sex (operationalized categorically) and age (12–18 years). The proce- dure involves three steps. In Step 1, the basic factor model is determined based on previous studies, theoretical or rational judg- ment, and/or preliminary analyses using the full sample or specific subsamples. Bauer proposed a two-factor model for the delinquent measures: violent delinquent behaviors and nonviolent delinquent behaviors. In Step 2, each factor is separated in isolation of the other factors and one-factor MNLFAs are fit to identify optimal specification for each factor. To accomplish this procedure, mod- eration functions for factor means and variances should be iden- tified based on informed knowledge and data exploration (Step 2a). Bauer hypothesized that delinquent behaviors would be higher for boys than for girls and would show a curvilinear trend across the age span. Five terms representing main and interaction effects (sex, age, age2, Sex � Age, and Sex � Age2) were included on the factor mean and variance via the functions given in equations (see Bauer, 2017, pp. 513–516). The next step (Step 2b) involves identifying items exhibiting differential item functioning (DIF; a term Bauer used to describe lack of MI). Bauer used an iterative method, first starting with the assumption of all items being invariant, then testing DIF associated with the set of five moder- ator terms sequentially, then identifying items with and without DIF, and then removing nonsignificant DIF terms (except lower- order terms involved in higher order effects; e.g., the age term would not be removed if the Sex � Age term is significant). The

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1487METHODOLOGICAL AND STATISTICAL ADVANCES

last step (Step 3) involves fitting a multidimensional model by recombining the optimal unidimensional models obtained from Step 2 and analyzing the full item set simultaneously.

Bauer acknowledged that specifying the MNLFA model can be quite complex due to the mixture of linear or log-linear moderation functions implemented, depending on the model parameters (e.g., linear functions for factor loadings, and log-linear functions for factor variances). Interpretation of the results is challenging as well; the author recommends scaling the grouping variables to meaningful units (e.g., mean-centering the age variable) to make baseline estimates interpretable. We refer the readers to Bauer (2017) for more detailed descriptions of implementation, interpre- tation, limitations, and suggestions for future research on MNLFA.

Statistical and Conceptual Issues Surrounding Evaluation of Measurement Invariance

As many interesting theoretical tests involving group compari- sons are discouraged until MI of an instrument has been estab- lished, researchers sometimes treat MI as an obstacle that must be overcome (Meade & Bauer, 2007). Consequently, evaluation of MI is not always conducted with the care and thoughtfulness scholars bring to their primary research interests, and results are not as informative as they could be. In the following section, we highlight some important statistical and conceptual issues (a) that investigators should consider in the context of their research in- volving MI to maximize the meaningfulness of their results and (b) that require theoretical and empirical development to advance the study of diversity in psychological assessment.

Criteria for Making Decisions About Invariance

MI tests typically involve a decision regarding whether the fit of a restrictive model (e.g., a metric model) is worse than that of an adjacent, less restrictive model (e.g., a configural model). For decision making criteria, researchers commonly rely on statistical significance tests and/or empirically derived cutoffs. For example, a chi-square change (��2) value is typically calculated to compare two nested models, with its associated p value used as the criterion. It is well known, however, that significance tests for �2 and ��2

are affected by sample size (Cheung & Lau, 2012). In large samples, even a minimal difference between two �2 values could be flagged as statistically significant, implying measurement non- invariance. Values of ��2 are also affected by model complexity, degree of departure from multivariate normality, and model mis- specification (French & Finch, 2008; Sass et al., 2014). To cir- cumvent these issues, empirically derived fit criteria have been proposed. Cheung and Rensvold’s (2002) simulation study led to a recommendation of comparative fit index change (�CFI) � �.01 as a criterion for identifying model noninvariance (i.e., the CFI should not decrease by more than .01 for invariance to be inferred). Chen’s (2007) simulation resulted in a proposal of more diverse cutoff criteria that varied according to sample size, equality of sample size across groups, and patterns of noninvariance. For example, a stricter criterion (�CFI � �.005) was suggested when sample size is small (total N � 300), sample sizes are unequal, and pattern of noninvariance is uniform. Meade, Johnson, and Brad- dy’s (2008) simulation study led to an even stricter recommended criterion (�CFI � �.002), as they found that more liberal criteria did not detect some forms of noninvariance.

While the above empirical cutoff criteria may serve as useful guidelines, they are best applied to models with continuous indi- cators that are estimated using the maximum likelihood method (Chen, 2007; Cheung & Rensvold, 2002; Meade et al., 2008; Sass et al., 2014). A recent study (Sass et al., 2014) showed that these traditional empirical criteria should not be used with the weighted least square means and variance adjusted estimator, which is the optimal method for handling binary or ordered categorical indica- tors. For binary indicators with a weighted least square means and variance adjusted estimator, the DIFFTEST option in Mplus has been recommended to test ��2 (Muthén & Muthén, 1998 –2015).

In short, there is no single standard that should be used to determine invariance in all situations (Greiff & Scherer, 2018; Meade et al., 2008). More relaxed standards for invariance may be called for in cross-cultural research in which lack of invariance is commonly observed due to multiple sources of potential bias (translation adequacy, linguistic idioms, cultural experiences, etc.). More conservative criteria for invariance should be applied if results could have a negative impact on disadvantaged groups in a selection situation. Unfortunately, no statistical standards are available for testing MI specifically in cross-cultural research (Svetina & Rutkowski, 2017). MI tests comparing Chinese Americans whose native language is English with European Americans using an English version of a measure can be expected to contain fewer errors than MI tests comparing Chinese persons residing in China with European Americans on their respective language versions of the same measure. Teasing out various sources of bias (translation adequacy, linguistic idioms, cultural experiences, etc.) using quan- titative methods is extremely difficult (Rutkowski & Svetina, 2014). For this reason, it is important for cross-cultural researchers to rely on measures that have been produced via rigorous transla- tion procedures. It is also important that cultural groups should be equivalent in terms of SES, age, education, and other relevant demographics, so that researchers can be confident that noninvari- ance is attributable to cultural differences.

Referent Indicators (RI) and Methods Identifying Invariant RI

When researchers propose CFA models, they must assign a unit of measurement to the latent factor so that the latent factors will have a scale and so that the model is identified (Cheung & Rensvold, 1999; Jöreskog & Sörrbom, 1989). Two procedures have been commonly used. One approach is to select an item (a marker or referent) for each factor and fix their factor loadings to 1. The other approach is to set the variance of each latent factor to 1, which standardizes the latent factors. For a single group CFA, the choice between procedures is not consequential as the two approaches produce almost the same model fit. For a MGCFA, however, the choice should be made more cautiously, as the two procedures entail different assumptions with respect to MI. In the first approach, setting the factor loading of the referent indicator (RI) to 1 across groups assumes that the true factor loadings of the referent are equal across groups. To the extent this assumption is incorrect, inaccurate estimates of other model parameters will ensue, and comparisons of factor loadings across groups may be biased. In the second approach, if the true variances of the latent factors are unequal across groups then tests of factorial invariance may be biased because the factor loadings for each group are using

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1488 HAN, COLARELLI, AND WEED

different metrics (Cheung & Rensvold, 1999). Because the latter is usually judged to be a more serious problem, the first approach has been more commonly used in MI research (e.g., Millsap & Yun- Tein, 2004). However, identifying candidates to serve as invariant RI can be complicated, especially when there are many indicators (Cheung & Rensvold, 1999; Jung & Yoon, 2017). Researchers rarely know which items are invariant prior to conducting any MI tests (Cheung, & Lau, 2012).

To assist in identifying invariant RI, Cheung and Rensvold (1999; Rensvold & Cheung, 1998) proposed a factor ratio test that examines each indicator as the RI in a set of models with each other variable constrained to be invariant across groups. If there

are p items to be tested, then a total of p�p � 1�

2 tests are needed. For example, if there are four items (X1–X4) associated with one factor, six invariance tests are required; when X1 � RI, there are three invariant tests (for X2, X3, and X4); when X2 � RI, there are two invariant tests (for X3 and X4); and when X3 � RI, there is one invariant test (for X4; see p. 11 in Cheung & Rensvold, 1999, for a detailed description of this method). Using chi-square differ- ence tests, invariant and nonvariant sets of items are identified. While this procedure can be efficacious in correctly identifying noninvariant items, it becomes labor intensive as the number of indicators and factors increases (e.g., 21 tests are needed for a seven-item single factor; French & Finch, 2008).

To simplify the search for an invariant RI using the factor-ratio test, Cheung and Lau (2012) proposed the use of bias-corrected bootstrapping confidence intervals, an extension of the estimation of confidence intervals for factor loading differences proposed by Meade and Bauer (2007). A major strength of this procedure is that it permits examination of all item-level MI tests using a single model estimation. It cannot, however, be applied to tests across three or more groups, or to scale-level MI evaluations. Readers interested in this topic are referred to Jung and Yoon (2017) for an explication of alternative methods for identifying RI (Jung & Yoon, 2017; Raykov, Marcoulides, & Millsap, 2013; Yoon & Millsap, 2007).

Study Characteristics and Patterns of Noninvariance

It is generally expected that as parameter differences (e.g., factor loadings) between groups increase, the observed increase in change in fit indices (e.g., ��2, �CFI) represents corresponding levels of test bias across groups (Chen, 2007). However, a number of simulation studies have shown that changes in fit indices are affected by other factors. Meade and Bauer (2007) demonstrated that certain manifestations of a study’s psychometric soundness can lead to a greater chance of detecting factor loading differences between groups. Larger sample size, factor overdetermination (i.e., more indicators per factor), and higher item communality (i.e., the proportion of variance in the item explained by factors) were associated with more accurate estimates of factor loading differ- ences and thereby a higher probability of detecting factor loading differences between groups. Ironically, as a failure to reject the null hypothesis is taken as evidence of invariance, some undesir- able characteristics of the data or the model (small sample size, few indicators per factor, and low item communality) can lead to spurious inference of MI.

Chen (2007) examined the sensitivity of change fit indices to measurement noninvaraince at various levels of invariance (e.g.,

factor loading, intercept, residual) using simulation methods. It was found that changes in fit indices were influenced by the interaction between degree of invariance and patterns of invariance when invariance testing was at the factor loading and intercept level. When noninvariance was reflected in mixed form (i.e., some of the factor loadings were higher for one group, some were higher for the other group), the relationship between degree of noninvari- ance and change in fit statistics was monotonic, as expected. That is, under these conditions, changes in fit indices were largest when invariance was highest, and changes were smallest when invari- ance was the lowest. When noninvariance was uniform (i.e., factor loadings were consistently higher for one group than for the other), however, the relationship between the degree of noninvariance and change in fit statistics was nonmonotonic. Surprisingly, under these conditions, change in fit indices was smallest when noninvariance was highest, whereas change in fit indices was largest when noninvariance was only moderate. Sample sizes also affected change indices when invariance tests were at the factor loading, intercept, and residual levels. For all three levels of invariance, equal sample sizes resulted in larger values in change indices than did unequal sample sizes, which implies that invariance tests are more likely to fail to detect nonin- variance when sample sizes are unequal.

A subsequent simulation study (Chen, 2008) found a similar interaction effect of degree and pattern of invariance on predictive validity (see S8 in the online supplemental materials) and on mean comparisons. Other simulation studies have found that when factor loading noninvariance between groups is large and in a uniform pattern, the bias in estimates of factor mean differences is large (Wang, Whittaker, & Beretvas, 2012; Xu & Green, 2016). These studies should caution researchers to supplement interpretation of fit indices with an understanding of both the study characteristics and observed pattern of noninvariance when making sense of MI test results.

Effect Size Measures

Again, while statistical significance tests and empirically de- rived criteria can be useful in practice, they must be used with caution. In the last few decades, behavioral scientists have begun placing less emphasis on dichotomous reject/do not reject deci- sions in favor of estimating magnitude of effect (Cohen, 1990, 1994; Wilkinson, 1999). Guidelines for reporting psychological research now commonly recommend or require reporting effect size estimates along with, or even instead of, statistical signifi- cance testing. Approximate fit indices (e.g., CFI) and their change values (e.g., �CFI) cannot be considered effect size measures; these values do not provide meaningful information about the nature of fit or invariance that can be interpreted outside of the context of the individual analysis. Moreover, just as there is nothing special about the threshold between p � .05 and p � .06, the distinction between �CFI � �.01 and �CFI � �.02 is a matter of minor degree rather than a categorical answer to a dichotomous question. Meade and Bauer (2007) recommend that researchers not view MI as an either– or proposition and calculate effect sizes and confidence intervals for factor loading differences when metric invariance is lacking.

Despite the importance of effect size measures to psychological research and practice, they have not been well developed in the MI literature. To date, Nye and Drasgow (2011) have made most

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1489METHODOLOGICAL AND STATISTICAL ADVANCES

systematic attempt to operationalize effect size within MI testing. They introduced a procedure for estimating degree of noninvari- ance called dMACS (where MACS is mean and covariance struc- ture). For each group, an item score (y) is regressed on a latent trait score (x), and the two regression lines are superimposed. dMACS quantifies DIF by indexing the area between the two regression lines. In providing a heuristic for the magnitude of effect, the authors turned to Cohen’s (1988) general guidelines for d (0.2, 0.5, and 0.8 for small, medium, and large effects, respectively). The authors provided an executable program that computes not only dMACS but also the effects of DIF on the mean (�M) and variance (�var) of a scale. At present, this method is limited to models with a simple structure (i.e., no secondary loadings) with at least one unbiased referent item (i.e., a marker item that is used to scale a latent factor). As dMACS represents the area between the two regression lines that are a function of (non)invariance of both factor loadings and intercept, this method does not distinguish between metric and scale invariance. We recommend that more researchers consider reporting such effect size measures along with statistical significance tests and fit index scores. Future work is greatly needed to develop and refine effect size measures and to incorporate them within MI software packages to promote wider use. Ultimately, then, degree of noninvariance can begin to be scaled to practical impact on examinees and groups of examinees, just as the increased use of effect size measures in other areas of psychology has led to greater insight about the practical impact of measures and their effects.

Relatedly, there have been some efforts to construct confidence bands around important MI parameters. In response to criticisms surrounding the use of �CFI (i.e., cutoff values being arbitrary, and a significance test of �CFI not being allowed due to unknown sampling distribution), Cheung and Lau (2012) proposed a direct- model comparison approach to MI testing using bias-corrected bootstrapping confidence intervals. In this approach, instead of comparing nested models, different forms of invariance are exam- ined in the same model systematically by performing a bootstrap analysis of differences between groups on specific parameters (Cheung & Lau, 2012). The procedure creates 1,000 bootstrap samples, calculates the parameter (e.g., differences between groups on factor loadings or intercepts) for each bootstrap sample, and makes adjustments to the bootstrap distribution of the parameter to form upper and lower confidence intervals. If zero falls within the conference intervals, then the null hypothesis of invariance is retained, as the parameter is considered to be invariant between groups. Even if zero falls outside of the confidence intervals, researchers might still pursue further research questions (e.g., group comparisons on latent means) as long as the confidence intervals are small and close to zero (Meade & Bauer, 2007). While the major strength of this method is its ability to determine where noninvariance exists within a model, it can presently only be applied to tests between two groups, or to item-level MI evalua- tions (Cheung & Lau, 2012). A few recent studies (e.g., Lui, 2019; McDermott et al., 2017) demonstrated how this method could be incorporated into other model indices in MI testing.

Appropriateness of CFA

The topics covered above are premised on the assumption the measure under study has a theoretically sound structural model

that shows good fit for all groups. Producing a single well-fitting initial model for all groups, however, is no easy task, depending on the characteristics of the measure and of the data matrix. The numerous tests used in clinical psychological assessment are ex- tremely diverse in terms of length, item response format, and test construction strategy, and certain categories of measures may simply be poorly suited to structural analysis using MGCFA. Consider two tests with very different characteristics. The first test is de- signed to measure a narrow facet of a normal personality trait known to show ample variability across a wide range of settings. It has a manageable number of items (e.g., 15–20) assembled using a deductive strategy valuing high internal consistency (Burisch, 1984). An item-level factor model (e.g., a three-factor model with five indicators per factor) can be proposed based on the theory that guided item selection. The model is tested comparing groups of about 200 cases each, and data meet multivariate normality as- sumption. Such a measure is an excellent candidate for MGCFA testing of MI. The proposed model is very likely to fit both groups well, and therefore serve as a baseline model for tests of metric or scalar invariance.

The second test is a multiscale inventory designed to measure broad and somewhat overlapping domains of personality and psy- chopathology. It has many items (e.g., over 75) and scales were assembled through a combination of construction approaches, in- cluding deductive, inductive, and empirical (Burisch, 1984). Re- flecting the psychodiagnostic constructs at its core, the measure’s items cover a wide range of content reflecting thoughts, emotions, attitudes, and behaviors, including symptom reports that vary widely in prevalence across settings. Such a test is a poor candidate for MGCFA testing on many grounds. Due to its construction methodology and the complexity of its constructs, such a measure is not likely to have a clean and simple internal structure, even within individual scales, even if scores on the test have shown substantial criterion-related validity. Sample sizes to test MI would need to be extremely large, and composition of the samples would be critical for adequate variability in the constructs measured. A researcher might attempt to find a good-fitting initial model for an individual scale by proposing a one-factor model and allowing some correlated error terms based on modification indices for all groups. Alternatively, a researcher might rely on EFA to identify a common factor model across groups and then apply it to MI testing. These approaches might still fail, however, as empirically derived factor models tend not to replicate well, leading to poor fit in independent samples.

Group Classifications

As discussed earlier, conventional classifications of racial/ethnic and sex/gender groupings commonly used in psychological re- search, including bias research, have recently been challenged. With respect to reconceptualizations of sex/gender, there are few precedents to guide current MI research. Replacing binary catego- ries with single or multiple ordinal or continuous variables (e.g., rating of maleness or femaleness) seems an option for MI testing, but the first such studies will break new ground. The same is true for the option of treating sex/gender categorically but expanding response options to include nonbinary groupings. An obvious related concern is the impact of number of groupings on sample size; including nonbinary groupings might produce some catego-

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1490 HAN, COLARELLI, AND WEED

ries with samples too small to include in formal MI testing. With Open Science initiatives gaining popularity, however, aggregation of data across open repositories may facilitate research addressing novel questions and allowing underrepre- sented gender variability to be better reflected in psychological assessment research.

Identifying meaningful and homogeneous racial/ethnic catego- ries is similarly complicated. For international or cross-cultural research, group homogeneity can sometimes be justified quite easily in terms of linguistic, geographical, and cultural common- alities. However, lines cannot always be drawn so easily. It might be a simple matter to decide to separate a Japanese sample from a Korean sample on the grounds that they do not share a primary language. But what of hypothetical samples from North and South Korea, whose language and geography are very similar, and who

shared a common culture for centuries until quite recently in historical terms? Does the cultural divergence of North and South Koreans since the Korean War justify treating them as separate cultural groups in an MI context? Conversely, might a growing global Internet culture someday permit classifying South Korean and American adolescents within the same cultural group in an MI study? It does not seem especially useful to propose universal heuristics for determining the correct number of groups in an MI study, or for describing the qualitative boundaries that ought to separate groups. The answers to questions about the propriety of group membership must always be found in the research context, guided either by theories about the phenomena under study, or by previous empirical literature justifying the separation or aggrega- tion of groups, or by the practical applications for which the measure is intended.

Table 1 Summary of Recommendations for Research Involving Evaluation of Measurement Invariance (MI) via Multigroup Confirmatory Factor Analysis and Related Statistical Methods

Study/paper component Recommendations

Introduction: Group classifications

Justify the group classifications used in your study, describe the possible implications of test bias and test unfairness, and propose specific hypotheses about invariance of your measure.

What do the categories mean? How are the categories theoretically related to the main constructs studied? Are all essential group categories included? Provide exhaustive response options to cover full range of categories. Try to avoid binary gender categories. If data on nonbinary gender categories are not included in the study and removed due to small N, these limitations should be discussed. Are your groups free of lumping errors? Ensure that heterogeneous groups (e.g., Asian Americans with diverse heritages and cultures) are not treated as a homogenous group. Explore the possibility of examining intersectional views of the categories.

Measures If different language versions are used in MI testing, verify the integrity of translation procedures and translation accuracy. Item-level MI testing, rather than scale-level MI testing, is recommended, as the results of the former may help identify poorly translated items.

Choice of model Justify your MI models and statistical methods. Consider using alternate statistical methods or multiple methods, if appropriate: Alignment method: model with many groups; one categorical grouping variable Moderated nonlinear factor analysis: model with single or multiple categorical or continuous grouping variables Exploratory SEM: complex modeling; one categorical grouping variable Multigroup Bayesian SEM: model with one categorical grouping variable; a method that allows researchers to incorporate prior knowledge of parameters.

Examine predictive invariance, if possible, along with MI testing.

Results Examine and report factor loadings and intercepts (or thresholds) for each group. Discuss patterns of these parameters. Ensure that the marker indicators (referents) are invariant across groups.

Use with caution criteria recommended for concluding measurement (non)invariance. Adjust criteria as appropriate to the context or consequence of MI, and provide justification. Report effect size estimates if possible.

Discussion Provide post hoc explanations regarding (non)invariant items as hypotheses to be tested in future studies. What did you learn from your study of MI testing regarding group differences and similarities? If the measure is used for important decision making purposes (e.g., selection, educational or employment enhancement, diagnosis), what implications do the results have regarding (in)equality and test (un)fairness? Provide recommendations regarding noninvariant items. Remove, interpret with caution, or retranslate.

Note. As MI methods are specialized forms of confirmatory factor analysis (CFA) and structural equation modeling (SEM), many of the guidelines suggested for SEM and CFA (see especially Table 7 in Appelbaum et al., 2018) are strongly recommended for MI testing. The list here is designed to address issues specific to measurement invariance research that are not discussed as frequently. Many other resources on SEM and CFA (e.g., Kline, 2016; Mueller & Hancock, 2011; Schmitt & Kuljanin, 2008; Thompson, 2000); on MI (e.g., Putnick & Bornstein, 2016); and on race/ethnicity, sex/gender, and intersectionality (Corral & Landrine, 2010; Else-Quest & Hyde, 2016; Helms, Jernigan, & Mascher, 2005; Hyde, Bigler, Joel, Tate, & van Anders, 2018; Schellenberg & Kaiser, 2018) offer excellent conceptual and/or methodological guidelines and recommendations for various stages in conducting measurement invariance studies.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1491METHODOLOGICAL AND STATISTICAL ADVANCES

Conclusions and Recommendations

Human organizations have become increasingly inclusive and demographically complex as a consequence of social and political changes. Responding to these changes in the context of assessing bias in testing requires that we (a) seek a more nuanced under- standing of the grouping variables at the center of our research, (b) be more explicit about the meaning and relevance of our group categories, (c) be more sensitive to and inclusive of differences in personal identity, (d) consider ways of defining and assessing intersectionalities (Colarelli, Han, & Yang, 2010; Cole, 2009; Pauker, Meyers, Sanchez, Gaither, & Young, 2018), and (e) sup- plement or replace traditional categorical operationalizations with alternatives when appropriate. While the broad, simple classifica- tions at the center of past research have some advantages in parsimony and statistical power, their scientific utility has clear limits.

MI testing was introduced to the psychological literature more than 5 decades years ago (e.g., Jöreskog, 1971), and statistical techniques for testing invariance have become increasingly acces- sible to researchers due to marked increases in computing capacity, the development of user-friendly software, and the availability of helpful, step-by-step guides to standard MI testing procedures. The increased accessibility has unfortunately also led to an increase in superficial or perfunctory practices: Group categorizations are not well-justified; specific hypotheses regarding the outcome of MI testing are not offered; modification indices are overused to claim good model fit; factor loadings, intercepts, and thresholds are unreported and undiscussed; and specific cultural implications of MI results are neglected.

Instead, researchers need to use MI testing in a way that opti- mizes advancement of knowledge of the constructs and diverse groups under study, intentionally and thoughtfully drawing infer- ences relevant to their samples and constructs, and keeping in mind the conceptual and statistical challenges inherent in the procedures. Limitations in MGCFA should lead researchers to consider alter- nate techniques such as the alignment method when a large num- ber of homogenous groups are involved in MI testing, or MNLFA when testing MI using categorical and/or continuous group vari- ables. Negative results should not go unreported, but be used as an opportunity to learn about the manifestations of construct expres- sion across groups (Greiff & Scherer, 2018). Researchers should bear in mind that characteristics of the measure itself—its devel- opment, breadth, size, complexity of constructs, and sensitivity to sample— have critical implications for the appropriateness of CFA or MGCFA. As many researchers have noted before (e.g., Hop- wood & Donnellan, 2010), internal structural analyses represent just one method of establishing construct validity. Depending on the intended applications of the test, and especially in clinical psychological assessment, criterion validation approaches may not only be more appropriate, they may also be more clinically rele- vant.

It is our hope that knitting substantive assessment issues within a presentation of the advances in this methodology will foster more thoughtful MI research submissions to Psychological Assessment and elsewhere. Table 1 summarizes our recommendations for con- ducting MI (also see Table S1 in the online supplemental materials for an overview of key publications on CFA, MGCFA, and related methods). It is our hope that greater familiarity with these procedures

and the conceptual and statistic issues that surround them will guide assessment researchers toward the development and revision of mea- sures that are maximally useful in cross-cultural or multicultural contexts.

References

American Educational Research Association, American Psychological As- sociation (APA), & National Council on Measurement in Education, Joint Committee on Standards for Educational and Psychological Test- ing. (2014). Standards for educational and psychological testing. Wash- ington, DC: American Educational Research Association.

American Psychiatric Association (APA). (2000). Diagnostic and statisti- cal manual of mental Disorders (4th ed., text rev.). Washington, DC: Author.

American Psychiatric Association (APA). (2013). Diagnostic and statisti- cal manual of mental disorders (5th ed.). Washington, DC: Author.

American Psychological Association (APA). (2012). Guidelines for psy- chological practice with lesbian, gay, and bisexual clients. American Psychologist, 67, 10 – 42. http://dx.doi.org/10.1037/a0024659

American Psychological Association (APA). (2015). Guidelines for psycho- logical practice with transgender and gender nonconforming people. Amer- ican Psychologist, 70, 832– 864. http://dx.doi.org/10.1037/a0039906

American Psychological Association (APA). (2017). Multicultural guide- lines: An ecological approach to context, identity, and intersectionality. Washington, DC: Author. Retrieved from https://www.apa.org/about/ policy/multicultural-guidelines.aspx

Appel, H. B., Huang, B., Ai, A. L., & Lin, C. J. (2011). Physical, behavioral, and mental health issues in Asian American women: Results from the National Latino Asian American Study. Journal of Women’s Health, 20, 1703–1711. http://dx.doi.org/10.1089/jwh.2010.2726

Appelbaum, M., Cooper, H., Kline, R. B., Mayo-Wilson, E., Nezu, A. M., & Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73, 3–25. http://dx.doi .org/10.1037/amp0000191

Arbisi, P. A., Ben-Porath, Y. S., & McNulty, J. (2002). A comparison of MMPI-2 validity in African American and Caucasian psychiatric inpa- tients. Psychological Assessment, 14, 3–15. http://dx.doi.org/10.1037/ 1040-3590.14.1.3

Asparouhov, T., & Muthén, B. (2009). Exploratory structural equation mod- eling. Structural Equation Modeling, 16, 397– 438. http://dx.doi.org/10 .1080/10705510903008204

Asparouhov, T., & Muthén, B. (2014). Multiple-group factor analysis align- ment. Structural Equation Modeling, 21, 495–508. http://dx.doi.org/10 .1080/10705511.2014.919210

Baker, N. L., & Mason, J. L. (2010). Gender issues in psychological testing of personality and abilities. In J. C. Chrisler & D. R. McCreary (Eds.), Handbook of gender research in psychology: Vol. 2: Gender research in social and applied psychology. New York, NY: Springer. http://dx.doi .org/10.1007/978-1-4419-1467-5_4

Bauer, D. J. (2017). A more general model for testing measurement invariance and differential item functioning. Psychological Methods, 22, 507–526. http://dx.doi.org/10.1037/met0000077

Berry, C. M., Clark, M. A., & McClure, T. K. (2011). Racial/ethnic differences in the criterion-related validity of cognitive ability tests: A qualitative and quantitative review. Journal of Applied Psychology, 96, 881–906. http://dx.doi.org/10.1037/a0023222

Betancourt, H., & López, S. R. (1993). The study of culture, ethnicity, and race in American psychology. American Psychologist, 48, 629 – 637. http://dx.doi.org/10.1037/0003-066X.48.6.629

Bittner, A., & Goodyear-Grant, E. (2017). Sex isn’t gender: Reforming concepts and measurements in the study of public opinion. Political Behavior, 39, 1019 –1041. http://dx.doi.org/10.1007/s11109-017-9391-y

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1492 HAN, COLARELLI, AND WEED

Borsboom, D. (2006). The attack of the psychometricians. Psychometrika, 71, 425– 440. http://dx.doi.org/10.1007/s11336-006-1447-6

Borsboom, D., Romeijn, J. W., & Wicherts, J. M. (2008). Measurement invariance versus selection invariance: Is fair selection possible? Psy- chological Methods, 13, 75–98. http://dx.doi.org/10.1037/1082-989X.13 .2.75

Brizendine, L. (2006). The female brain. New York, NY: Morgan Road Books.

Burisch, M. (1984). Approaches to personality inventory construction: A comparison of merits. American Psychologist, 39, 214 –227. http://dx .doi.org/10.1037/0003-066X.39.3.214

Butcher, J. N., Dahlstrom, W. G., Graham, J. R., Tellegen, A., & Kaem- mer, B. (1989). The Minnesota Multiphasic Personality Inventory-2 (MMPI-2): Manual for administration and scoring. Minneapolis: Uni- versity of Minnesota Press.

Byrne, B. M., Shavelson, R. J., & Muthén, B. (1989). Testing for the equivalence of factor covariance and mean structures: The issue of partial measurement invariance. Psychological Bulletin, 105, 456 – 466. http://dx.doi.org/10.1037/0033-2909.105.3.456

Byrne, B. M., & Stewart, S. (2006). Teacher’s corner: The MACS ap- proach to testing for multigroup invariance of a second-order structure: A walk through the process. Structural Equation Modeling, 13, 287– 321. http://dx.doi.org/10.1207/s15328007sem1302_7

Camilli, G. (2006). Test fairness. In R. L. Brennan (Ed.), Educational measurement (4th ed., pp. 221–256). Westport, CT: Praeger.

Carrugan, M. (2015). Asexuality. In C. Richards & M. J. Barker (Eds.), The Palgrave handbook of the psychology of sexuality and gender (pp. 1–23). London, United Kingdom: Palgrave Macmillan. http://dx.doi.org/ 10.1057/9781137345899_2

Chen, F. F. (2007). Sensitivity of goodness of fit indexes to lack of measurement invariance. Structural Equation Modeling, 14, 464 –504. http://dx.doi.org/10.1080/10705510701301834

Chen, F. F. (2008). What happens if we compare chopsticks with forks? The impact of making inappropriate comparisons in cross-cultural re- search. Journal of Personality and Social Psychology, 95, 1005–1018. http://dx.doi.org/10.1037/a0013193

Chen, F. F., Sousa, K. H., & West, S. G. (2005). Teacher’s corner: Testing measurement invariance of second-order factor models. Structural Equation Modeling, 12, 471– 492. http://dx.doi.org/10.1207/s153280 07sem1203_7

Cheung, F. M., Song, W. Z., & Zhang, J. X. (1996). The Chinese MMPI-2: Research and applications in Hong Kong and the People’s Republic of China. In J. N. Butcher (Ed.), International adaptations of the MMPI-2: A handbook of research and applications (pp. 137–161). Minneapolis: University of Minnesota Press.

Cheung, G. W. (2008). Testing equivalence in the structure, means, and variances of higher-order constructs with structural equation modeling. Organizational Research Methods, 11, 593– 613. http://dx.doi.org/10 .1177/1094428106298973

Cheung, G. W., & Lau, R. S. (2012). A direct comparison approach for testing measurement invariance. Organizational Research Methods, 15, 167–198. http://dx.doi.org/10.1177/1094428111421987

Cheung, G. W., & Rensvold, R. B. (1999). Testing factorial invariance across groups: A reconceptualization and proposed new method. Journal of Man- agement, 25, 1–27. http://dx.doi.org/10.1177/014920639902500101

Cheung, G. W., & Rensvold, R. B. (2000). Assessing extreme and acqui- escence response sets in cross-cultural research using structural equation modeling. Journal of Cross-Cultural Psychology, 31, 187–212. http:// dx.doi.org/10.1177/0022022100031002003

Cheung, G. W., & Rensvold, R. B. (2002). Evaluating goodness-of-fit indexes for testing measurement invariance. Structural Equation Mod- eling, 9, 233–255. http://dx.doi.org/10.1207/S15328007SEM0902_5

Chun, S., Stark, S., Kim, E. S., & Chernyshenko, O. S. (2016). MIMIC methods for detecting DIF among multiple groups: Exploring a new

sequential-free baseline procedure. Applied Psychological Measurement, 40, 486 – 499. http://dx.doi.org/10.1177/0146621616659738

Chung, J. J., Weed, N. C., & Han, K. (2006). Evaluating cross-cultural equivalence of the Korean MMPI-2 via bilingual test-retest. Interna- tional Journal of Intercultural Relations, 30, 531–543. http://dx.doi.org/ 10.1016/j.ijintrel.2005.08.009

Clarizio, H. F. (1979). In defense of the IQ test. School Psychology Review, 8, 79 – 88.

Clearly, T. A. (1968). Test bias: Prediction of grades of Negro and White students in integrated colleges. Journal of Educational Measurement, 5, 115–124. http://dx.doi.org/10.1111/j.1745-3984.1968.tb00613.x

Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Erlbaum.

Cohen, J. (1990). Things I have learned (so far). American Psychologist, 45, 1304 –1312. http://dx.doi.org/10.1037/0003-066X.45.12.1304

Cohen, J. (1994). The earth is round (p � .05). American Psychologist, 49, 997–1003. http://dx.doi.org/10.1037/0003-066X.49.12.997

Colarelli, S. M., Han, K., & Yang, C. (2010). Biased against whom? The problems of “group” definition and membership in test bias analyses. Industrial and Organizational Psychology: Perspectives on Science and Practice, 3, 228 –231. http://dx.doi.org/10.1111/j.1754-9434.2010.01229.x

Cole, E. R. (2009). Intersectionality and research in psychology. American Psychologist, 64, 170 –180. http://dx.doi.org/10.1037/a0014564

Corral, I., & Landrine, H. (2010). Methodological and statistical issues in research with diverse samples: The problem of measurement equiva- lence. In H. Landrine & N. F. Russo (Eds.), Handbook of diversity in feminist psychology (pp. 83–134). New York, NY: Springer.

Culhane, S. E., Morera, O. F., Watson, P. J., & Millsap, R. E. (2009). Assessing measurement and predictive invariance of the Toronto Alex- ithymia Scale-20 in U.S. Anglo and U.S. Hispanic student samples. Journal of Personality Assessment, 91, 387–395. http://dx.doi.org/10 .1080/00223890902936264

Davidov, E., Meuleman, B., Cieciuch, J., Schmidt, P., & Billiet, J. (2014). Measurement equivalence in cross-national research. Annual Review of Sociology, 40, 55–75. http://dx.doi.org/10.1146/annurev-soc-071913- 043137

Eagly, A. H., & Riger, S. (2014). Feminism and psychology: Critiques of methods and epistemology. American Psychologist, 69, 685–702. http:// dx.doi.org/10.1037/a0037372

Else-Quest, N. M., & Hyde, J. S. (2016). Intersectionality in quantitative psychological research: I. Theoretical and epistemological issues. Psy- chology of Women Quarterly, 40, 155–170. http://dx.doi.org/10.1177/ 0361684316629797

Feingold, A. (1994). Gender differences in personality: A meta-analysis. Psychological Bulletin, 116, 429 – 456. http://dx.doi.org/10.1037/0033- 2909.116.3.429

French, B. F., & Finch, W. H. (2008). Multigroup confirmatory factor analysis: Locating the invariant referent sets. Structural Equation Mod- eling, 15, 96 –113. http://dx.doi.org/10.1080/10705510701758349

Gelfand, M. J., Raver, J. L., Nishii, L., Leslie, L. M., Lun, J., Lim, B. C., . . . Yamaguchi, S. (2011). Differences between tight and loose cultures: A 33-nation study. Science, 332, 1100 –1104. http://dx.doi.org/10.1126/ science.1197754

Gibbons, A. (2017). Busting myths of origin. Science, 356, 678 – 681. Gilman, R., Scott, H. E., Tian, L., Park, N., O’Byrne, J., Schiff, M., . . .

Langknecht, H. (2008). Cross-national adolescent multidimensional life satisfaction reports: Analyses of mean scores and response style differ- ences. Journal of Youth and Adolescence, 37, 142–154. http://dx.doi.org/ 10.1007/s10964-007-9172-8

Greiff, S., & Scherer, R. (2018). Still comparing apples with oranges? Some thoughts on the principles and practices of measurement invari- ance testing. European Journal of Psychological Assessment, 34, 141– 144. http://dx.doi.org/10.1027/1015-5759/a000487

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1493METHODOLOGICAL AND STATISTICAL ADVANCES

Hafner, J. (2019, March 3). Gender “X”: New York City adds gender- neutral option to birth certificates. USA Today. Retrieved from https:// www.usatoday.com/story/news/nation/2019/01/03/new-york-city-birth- certificates-now-feature-third-gender-option-x/2472189002/

Helms, J. E. (2006). Fairness is not validity or cultural bias in racial-group assessment: A quantitative perspective. American Psychologist, 61, 845– 859. http://dx.doi.org/10.1037/0003-066X.61.8.845

Helms, J. E., Jernigan, M., & Mascher, J. (2005). The meaning of race in psychology and how to change it: A methodological perspective. Amer- ican Psychologist, 60, 27–36. http://dx.doi.org/10.1037/0003-066X.60 .1.27

Hildebrandt, A., Lüdtke, O., Robitzsch, A., Sommer, C., & Wilhelm, O. (2016). Exploring factor model parameters across continuous variables with local structural equation models. Multivariate Behavioral Re- search, 51, 257–258. http://dx.doi.org/10.1080/00273171.2016.1142856

Hirschfeld, G., & von Brachel, R. (2014). Improving multiple-group confirmatory factor analysis in R-A tutorial in measurement invari- ance with continuous and ordinal indicators. Practical Assessment, Research & Evaluation, 19, 1–12. Retrieved from http://pareonline .net/getvn.asp?v�19&n�7

Hopwood, C. J., & Donnellan, M. B. (2010). How should the internal structure of personality inventories be evaluated? Personality and Social Psychology Review, 14, 332–346. http://dx.doi.org/10.1177/1088868 310361240

Humes, K. R., Jones, N. A., & Ramirez, R. R. (2010). Overview of race and Hispanic origin: 2010 (Census briefs). Retrieved from https://www .census.gov/prod/cen2010/briefs/c2010br-02.pdf

Hunter, J. E., & Schmidt, F. L. (1976). Critical analysis of the statistical and ethical implications of various definitions of test bias. Psychological Bulletin, 83, 1053–1071. http://dx.doi.org/10.1037/0033-2909.83.6.1053

Hyde, J. S. (2005). The gender similarities hypothesis. American Psychol- ogist, 60, 581–592. http://dx.doi.org/10.1037/0003-066X.60.6.581

Hyde, J. S., Bigler, R. S., Joel, D., Tate, C. C., & van Anders, S. M. (2019). The future of sex and gender in psychology: Five challenges to the gender binary. American Psychologist, 74, 171–193.

Jang, S., Kim, E., Cao, C., Allen, T. D., Cooper, C. L., Lapierre, L. M., . . . Woo, J. (2017). Measurement invariance of the Satisfaction with Life Scale across 26 countries. Journal of Cross-Cultural Psychology, 48, 560 –576. http://dx.doi.org/10.1177/0022022117697844

Jensen, A. R. (1980). Bias in mental testing. New York, NY: Free Press. Jöreskog, K. G. (1971). Simultaneous factor analysis in several popula-

tions. Psychometrika, 36, 409 – 426. http://dx.doi.org/10.1007/BF0 2291366

Jöreskog, K. G., & Sörrbom, D. (1989). L1SREL 7: A guide to the program and applications. Chicago, IL: SPSS.

Jung, E., & Yoon, M. (2017). Two-step approach to partial factorial invariance: Selecting a reference variable and identifying the source of noninvariance. Structural Equation Modeling, 24, 65–79. http://dx.doi .org/10.1080/10705511.2016.1251845

Kane, M. T., & Mroch, A. A. (2010). Modeling group differences in OLS and orthogonal regression: Implications for differential validity studies. Applied Measurement in Education, 23, 215–241. http://dx.doi.org/10 .1080/08957347.2010.485990

Kaplan, R. M., & Saccuzzo, D. P. (2009). Psychological testing: Princi- ples, applications, and issues (7th ed.). Belmont, CA: Wadsworth.

Ketterer, H. L., Han, K., Hur, J., & Moon, K. (2010). Development and validation of Variable Response Inconsistency (VRIN) and True Re- sponse Inconsistency (TRIN) validity scales for use with the Korean MMPI-2. Psychological Assessment, 16, 379 –385.

Khojasteh, J., & Lo, W. J. (2015). Investigating the sensitivity of goodness- of-fit indices to detect measurement invariance in a bifactor model. Structural Equation Modeling, 22, 531–541. http://dx.doi.org/10.1080/ 10705511.2014.937791

Kim, E. S., Cao, C., Wang, Y., & Nguyen, D. T. (2017). Measurement invariance testing with many groups: A comparison of five approaches. Structural Equation Modeling, 24, 524 –544. http://dx.doi.org/10.1080/ 10705511.2017.1304822

Kline, R. B. (2013). Assessing statistical aspects of test fairness with structural equation modelling. Educational Research and Evaluation: An International Journal on Theory and Practice, 19, 204 –222. http:// dx.doi.org/10.1080/13803611.2013.767624

Kline, R. B. (2016). Principles and practice of structural equation mod- eling (4th ed.). New York, NY: Guilford Press.

Kuyper, L., & Wijsen, C. (2014). Gender identities and gender dysphoria in the Netherlands. Archives of Sexual Behavior, 43, 377–385. http://dx .doi.org/10.1007/s10508-013-0140-y

Lai, M. H. C., & Yoon, M. (2015). A modified comparative fit index for factorial invariance studies. Structural Equation Modeling, 22, 236 –248. http://dx.doi.org/10.1080/10705511.2014.935928

Lee, H. (2012). Equivalence and faking issues of the aggression question- naire and the conditional reasoning test for aggression in Korean and American samples (Doctoral dissertation). Retrieved from PsycINFO Database (Accession No. 2014 –99100-226).

Linn, M. C., & Kessel, C. (2010). Assessment and gender. In J. C. Chrisler & D. R. McCreary (Eds.), Handbook of gender research in psychology: Vol. 2: Gender research in social and applied psychology (pp. 40 –50). New York, NY: Springer.

Locke, K. D., & Baik, K. (2009). Does an acquiescent response style explain why Koreans are less consistent than Americans? Journal of Cross-Cultural Psychology, 40, 319 –323. http://dx.doi.org/10.1177/ 0022022108328915

Lui, P. P. (2019). College alcohol beliefs: Measurement invariance, mean differences, and correlations with alcohol use outcomes across sociode- mographic groups. Journal of Counseling Psychology. Advance online publication. http://dx.doi.org/10.1037/cou0000338

MacCallum, R. C., Zhang, S., Preacher, K. J., & Rucker, D. D. (2002). On the practice of dichotomization of quantitative variables. Psychological Methods, 7, 19 – 40. http://dx.doi.org/10.1037/1082-989X.7.1.19

Markus, H. R. (2008). Pride, prejudice, and ambivalence: Toward a unified theory of race and ethnicity. American Psychologist, 63, 651– 670. http://dx.doi.org/10.1037/0003-066X.63.8.651

Marsh, H. W., Guo, J., Parker, P. D., Nagengast, B., Asparouhov, T., Muthén, B., & Dicke, T. (2018). What to do when scalar invariance fails: The extended alignment method for multi-group factor analysis com- parison of latent means across many groups. Psychological Methods, 23, 524 –545.

Marsh, H. W., Morin, A. J. S., Parker, P. D., & Kaur, G. (2014). Exploratory structural equation modeling: An integration of the best features of explor- atory and confirmatory factor analysis. Annual Review of Clinical Psychol- ogy, 10, 85–110. http://dx.doi.org/10.1146/annurev-clinpsy-032813-153700

Marsh, H. W., Muthén, B., Asparouhov, T., Lüdtke, O., Robitzsch, A., Morin, A. J. S., & Trautwein, U. (2009). Exploratory structural equation modeling, integrating CFA and EFA: Application to students’ evalua- tions of university teaching. Structural Equation Modeling, 16, 439 – 476. http://dx.doi.org/10.1080/10705510903008220

Mattern, K. D., Patterson, B. F., Shaw, E. J., Kobrin, J. L., & Barbuti, S. M. (2008). Differential validity and prediction of the SAT (College Board Research Report No. 2008-4). New York, NY: College Board.

McDermott, R. C., Levant, R. F., Hammer, J. H., Hall, R. J., McKelvey, D. K., & Jones, Z. (2017). Further examination of the factor structure of the Male Role Norms Inventory-Short Form (MRNI-SF): Measurement consider- ations for women, men of color, and gay men. Journal of Counseling Psychology, 64, 724 –738. http://dx.doi.org/10.1037/cou0000225

Meade, A. W., & Bauer, D. J. (2007). Power and precision in confirmatory factor analytic tests of measurement invariance. Structural Equation Modeling, 14, 611– 635. http://dx.doi.org/10.1080/10705510701575461

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1494 HAN, COLARELLI, AND WEED

Meade, A. W., Johnson, E. C., & Braddy, P. W. (2008). Power and sensitivity of alternative fit indices in tests of measurement invariance. Journal of Applied Psychology, 93, 568 –592. http://dx.doi.org/10.1037/ 0021-9010.93.3.568

Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58, 525–543. http://dx.doi.org/10.1007/ BF02294825

Millsap, R. E. (1995). Measurement invariance, predictive invariance, and the duality paradox. Multivariate Behavioral Research, 30, 577– 605. http://dx.doi.org/10.1207/s15327906mbr3004_6

Millsap, R. E. (1997). Invariance in measurement and prediction: Their relationship in the single-factor case. Psychological Methods, 2, 248 –260. http://dx.doi.org/10.1037/1082-989X.2.3.248

Millsap, R. E. (2007). Invariance in measurement and prediction revisited. Psychometrika, 72, 461– 473. http://dx.doi.org/10.1007/s11336-007- 9039-7

Millsap, R. E. (2011). Statistical approaches to measurement invariance. New York, NY: Routledge.

Millsap, R. E., & Yun-Tein, J. (2004). Assessing factorial invariance in ordered-categorical measures. Multivariate Behavioral Research, 39, 479 –515. http://dx.doi.org/10.1207/S15327906MBR3903_4

Mueller, R. O., & Hancock, G. R. (2011). Best practices in structural equation modeling. In J. Osborne (Ed.), Best practices in quantitative methods (pp. 488 –508). Thousand Oaks, CA: Sage.

Muthén, B., & Asparouhov, T. (2013). BSEM measurement invariance analysis. Retrieved from https://www.statmodel.com/examples/ webnotes/webnote17.pdf

Muthén, B., & Asparouhov, T. (2014). IRT studies of many groups: The alignment method. Frontiers in Psychology, 5, 978.

Muthén, B., & Asparouhov, T. (2017). Recent methods for the study of measurement invariance with many groups: Alignment and random effects. Sociological Methods & Research, 47, 637– 664. http://dx.doi .org/10.1177/0049124117701488

Muthén, L. K., & Muthén, B. (1998 –2015). Mplus user’s guide (Version 7.11). Los Angeles, CA: Author.

Newman, D. A., Hanges, P. J., & Outtz, J. L. (2007). Racial groups and test fairness, considering history and construct validity. American Psychol- ogist, 62, 1082–1083. http://dx.doi.org/10.1037/0003-066X.62.9.1082

Nisbett, R. E., Peng, K., Choi, I., & Norenzayan, A. (2001). Culture and systems of thought: Holistic versus analytic cognition. Psychological Review, 108, 291–310. http://dx.doi.org/10.1037/0033-295X.108.2.291

Nobles, M. (2000). Shades of citizenship: Race and the census in modern politics. Stanford, CA: Stanford University Press.

Nye, C. D., & Drasgow, F. (2011). Effect size indices for analyses of measurement equivalence: Understanding the practical importance of differences between groups. Journal of Applied Psychology, 96, 966 – 980. http://dx.doi.org/10.1037/a0022955

Parra, E. J., Marcini, A., Akey, J., Martinson, J., Batzer, M. A., Cooper, R., . . . Shriver, M. D. (1998). Estimating African American admixture proportions by use of population-specific alleles. American Journal of Human Genetics, 63, 1839 –1851. http://dx.doi.org/10.1086/302148

Pauker, K., Meyers, C., Sanchez, D. T., Gaither, S. E., & Young, D. M. (2018). A review of multiracial malleability: Identity, categorization, and shifting racial attitudes. Social and Personality Psychology Com- pass, 12, e12392. http://dx.doi.org/10.1111/spc3.12392

Pendergast, L. L., von der Embse, N., Kilgus, S. P., & Eklund, K. R. (2017). Measurement equivalence: A non-technical primer on categori- cal multi-group confirmatory factor analysis in school psychology. Jour- nal of School Psychology, 60, 65– 82. http://dx.doi.org/10.1016/j.jsp .2016.11.002

Pudumjee, S. B., Weed, N. C., Pant, H., Ahluwalia, H., & Chakranarayan, C. (2016, May). Initial evaluation of the Hindi MMPI-2 via bilingual test-retest. Paper presented at the 51st Annual MMPI Symposium and Workshops, Fort Lauderdale, FL.

Putnick, D. L., & Bornstein, M. H. (2016). Measurement invariance conventions and reporting: The state of the art and future directions for psychological research. Developmental Review, 41, 71–90. http://dx.doi .org/10.1016/j.dr.2016.06.004

Raykov, T., Marcoulides, G. A., & Millsap, R. E. (2013). Factorial invari- ance in multiple populations: A multiple testing procedure. Educational and Psychological Measurement, 73, 713–727. http://dx.doi.org/10 .1177/0013164412451978

Rensvold, R. B., & Cheung, G. W. (1998). Testing measurement model for factorial invariance: A systematic approach. Educational and Psycho- logical Measurement, 58, 1017–1034. http://dx.doi.org/10.1177/001 3164498058006010

Reynolds, C. R., & Suzuki, L. A. (2013). Bias in psychological assessment: An empirical review and recommendations. In J. R. Graham, J. A. Naglieri, & I. B. Weiner (Eds.), Handbook of psychology: Assessment psychology (pp. 82–113). Hoboken, NJ: Wiley.

Richards, C., & Barker, M. (2013). Sexuality and gender for mental health professionals: A practical guide. Thousand Oaks, CA: Sage. http://dx .doi.org/10.4135/9781473957817

Richards, C., Bouman, W. P., & Barker, M. (Eds.). (2017). Genderqueer and non-binary genders. London, United Kingdom: Palgrave Mac- millan. http://dx.doi.org/10.1057/978-1-137-51053-2

Richards, C., Bouman, W. P., Seal, L., Barker, M. J., Nieder, T. O., & T’Sjoen, G. (2016). Non-binary or genderqueer genders. International Review of Psychiatry, 28, 95–102. http://dx.doi.org/10.3109/09540261 .2015.1106446

Rosenfeld, M. J. (2006). Young adulthood as a factor in social change in the United States. Population and Development Review, 32, 27–51. http://dx.doi.org/10.1111/j.1728-4457.2006.00104.x

Roth, P. L., Bevier, C. A., Bobko, P., Switzer, F. S., III, & Tyler, P. (2001). Ethnic group differences in cognitive ability in employment and educa- tional settings: A meta-analysis. Personnel Psychology, 54, 297–330. http://dx.doi.org/10.1111/j.1744-6570.2001.tb00094.x

Ruigrok, A. N., Salimi-Khorshidi, G., Lai, M. C., Baron-Cohen, S., Lom- bardo, M. V., Tait, R. J., & Suckling, J. (2014). A meta-analysis of sex differences in human brain structure. Neuroscience and Biobehavioral Reviews, 39, 34 –50. http://dx.doi.org/10.1016/j.neubiorev.2013.12.004

Rushton, J. P. (1994). Race, evolution, and behavior: A life history per- spective. New Brunswick, NJ: Transaction Publishers.

Rutkowski, L., & Svetina, D. (2014). Assessing the hypothesis of mea- surement invariance in the context of large-scale international surveys. Educational and Psychological Measurement, 74, 31–57. http://dx.doi .org/10.1177/0013164413498257

Sackett, P. R., & Wilk, S. L. (1994). Within-group norming and other forms of score adjustment in preemployment testing. American Psychol- ogist, 49, 929 –954. http://dx.doi.org/10.1037/0003-066X.49.11.929

Sass, D. A. (2011). Testing measurement invariance and comparing latent factor means within a confirmatory factor analysis framework. Journal of Psychoeducational Assessment, 29, 347–363. http://dx.doi.org/10 .1177/0734282911406661

Sass, D. A., Schmitt, T. A., & Marsh, H. W. (2014). Evaluating model fit with ordered categorical data within a measurement invariance frame- work: A comparison of estimators. Structural Equation Modeling, 21, 167–180. http://dx.doi.org/10.1080/10705511.2014.882658

Scelfo, J. (2015, Feb. 5). A university recognizes a third gender: Neutral. New York Times. Retrieved from https://www.nytimes.com/2015/02/08/ education/edlife/a-university-recognizes-a-third-gender-neutral .html?_r�0

Schellenberg, D., & Kaiser, A. (2018). The sex/gender distinction: Beyond F and M. In by C. B. Travis, J. W. White, A. Rutherford, W. S. Williams, S. L. Cook, & K. F. Wyche (Eds.), APA handbook of the psychology of women: History, theory, and battlegrounds (Vol. 1, pp. 165–187). Washington, DC: American Psychological Association.

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1495METHODOLOGICAL AND STATISTICAL ADVANCES

Schmitt, N., & Kuljanin, G. (2008). Measurement invariance: Review of practice and implications. Human Resource Management Review, 18, 210 –222. http://dx.doi.org/10.1016/j.hrmr.2008.03.003

Sellbom, M., & Tellegen, A. (2019). Factor analysis in psychological assessment research: Common pitfalls and recommendations. Psycho- logical Assessment, 31, 1428 –1441. http://dx.doi.org/10.1037/ pas0000623

Shou, Y., Sellbom, M., & Han, J. (2017). Evaluating the construct validity of the Levenson Self-Report Psychopathy Scale in China. Assessment, 24, 1008 –1023. http://dx.doi.org/10.1177/1073191116637421

Society for Industrial Organizational Psychology. (2018). Principles for the application and use of personnel selection procedures (5th ed.). Bowling Green, OH: Author.

Song, S. (2017, Spring). Multiculturalism. The Stanford Encyclopedia of Philosophy. Retrieved from https://plato.stanford.edu/archives/spr2017/ entries/multiculturalism/

Sörbom, D. (1974). A general method for studying differences in factor means and factor structure between groups. British Journal of Mathe- matical & Statistical Psychology, 27, 229 –239. http://dx.doi.org/10 .1111/j.2044-8317.1974.tb00543.x

Steenkamp, J. E. M., & Baumgartner, H. (1998). Assessing measurement invariance in cross-national consumer research. The Journal of Con- sumer Research, 25, 78 –90. http://dx.doi.org/10.1086/209528

Svetina, D., & Rutkowski, L. (2017). Multidimensional measurement invariance in an international context: Fit measure performance with many groups. Journal of Cross-Cultural Psychology, 48, 991–1008. http://dx.doi.org/10.1177/0022022117717028

Talhelm, T., Zhang, X., Oishi, S., Shimin, C., Duan, D., Lan, X., & Kitayama, S. (2014). Large-scale psychological differences within China explained by rice versus wheat agriculture. Science, 344, 603– 608. http://dx.doi.org/10.1126/science.1246850

Tate, C. C., Ledbetter, J. N., & Youssef, C. P. (2013). A two-question method for assessing gender categories in the social and medical sci- ences. Journal of Sex Research, 50, 767–776.

Thompson, B. (2000). Ten commandments of structural equation model- ing. In L. G. Grimm & P. R. Yarnold (Eds.), Reading and understanding MORE multivariate statistics (pp. 261–283). Washington, DC: Ameri- can Psychological Association.

U.S. Census Bureau. (2008, March 30). Table 1: Race of wife by race of husband: 1960, 1970, 1980, 1991, and 1992. Retrieved March 30, 2008, from http://www.census.gov/population/socdemo/race/interractab1.txt

Vandenberg, R. J. (2002). Toward a further understanding of and improve- ment in measurement invariance methods and procedures. Organiza- tional Research Methods, 5, 139 –158. http://dx.doi.org/10.1177/109 4428102005002001

Vandenberg, R. J., & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recom- mendations for organizational research. Organizational Research Meth- ods, 3, 4 –70. http://dx.doi.org/10.1177/109442810031002

Van de Velde, S., Levecque, K., & Bracke, P. (2009). Measurement equivalence of the CES-D 8 in the general population in Belgium: A gender perspective. Archives of Public Health, 67, 15–29. http://dx.doi .org/10.1186/0778-7367-67-1-15

Van de Vijver, F. J. R. (2011). Capturing bias in structural equation modeling. In E. Davidov, P. Schmidt, & J. Billiet (Eds.), Cross-cultural analysis: Methods and applications (pp. 3–34). New York, NY: Rout- ledge.

Verhoeven, M., Sawyer, M. G., & Spence, S. H. (2013). The factorial invariance of the CES-D during adolescence: Are symptom profiles for depression stable across gender and time? Journal of Adolescence, 36, 181–190. http://dx.doi.org/10.1016/j.adolescence.2012.10.007

Wang, D., Whittaker, T. A., & Beretvas, S. N. (2012). The impact of violating factor scaling method assumptions on latent mean difference testing in structured means models. Journal of Modern Applied Statis- tical Methods, 111, 3. http://dx.doi.org/10.22237/jmasm/1335844920

Wang, J., Han, K., Weed, N. C., McCabe, B. J., Dykhouse, A. S., McLaughlan, J. K., & Gilson, A. L. (2014, April). Establishing mea- surement invariance of the MMPI-2 RC4 (Antisocial Behavior) using American and Korean clinical samples. Poster presented at the 49th Annual Workshops and Symposia on Recent Developments in the Use of the MMPI-2/MMPI-2-RF/MMPI-A, Scottsdale, AZ.

Warne, R. T., Yoon, M., & Price, C. J. (2014). Exploring the various interpretations of “test bias.” Cultural Diversity & Ethnic Minority Psychology, 20, 570 –582. http://dx.doi.org/10.1037/a0036503

Wilkinson, L. (1999). Statistical methods in psychology journals: Guide- lines and explanations. American Psychologist, 54, 594 – 604. http://dx .doi.org/10.1037/0003-066X.54.8.594

Witherspoon, D. J., Wooding, S., Rogers, A. R., Marchani, E. E., Watkins, W. S., Batzer, M. A., & Jorde, L. B. (2007). Genetic similarities within and between human populations. Genetics, 176, 351–359. http://dx.doi .org/10.1534/genetics.106.067355

Wright, N. A., Kutschenko, K., Bush, B. A., Hannum, K. M., & Braddy, P. W. (2015). Measurement and predictive invariance of a work–life boundary measure across gender. International Journal of Selection and Assessment, 23, 131–148. http://dx.doi.org/10.1111/ijsa.12102

Xu, Y., & Green, S. B. (2016). The impact of varying the number of measurement invariance constraints on the assessment of between-group differences of latent means. Structural Equation Modeling, 23, 290 –301. http://dx.doi.org/10.1080/10705511.2015.1047932

Yee, A. H. (1983). Ethnicity and race: Psychological perspectives. Edu- cational Psychologist, 18, 14 –24. http://dx.doi.org/10.1080/004615 28309529257

Yee, A. H., Fairchild, H. H., Weizmann, F., & Wyatt, G. E. (1993). Addressing psychology’s problems with race. American Psychologist, 48, 1132–1140. http://dx.doi.org/10.1037/0003-066X.48.11.1132

Yoon, M., & Millsap, R. E. (2007). Detecting violations of factorial invariance using data-based specification searches: A Monte Carlo study. Structural Equation Modeling, 14, 435– 463. http://dx.doi.org/10 .1080/10705510701301677

Yuan, K. H., & Chan, W. (2016). Measurement invariance via multigroup SEM: Issues and solutions with chi-square-difference tests. Psycholog- ical Methods, 21, 405– 426. http://dx.doi.org/10.1037/met0000080

Received June 7, 2018 Revision received April 8, 2019

Accepted April 9, 2019 �

T hi

s do

cu m

en t

is co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

ti on

or on

e of

it s

al li

ed pu

bl is

he rs

. T

hi s

ar ti

cl e

is in

te nd

ed so

le ly

fo r

th e

pe rs

on al

us e

of th

e in

di vi

du al

us er

an d

is no

t to

be di

ss em

in at

ed br

oa dl

y.

1496 HAN, COLARELLI, AND WEED

  • Methodological and Statistical Advances in the Consideration of Cultural Diversity in Assessment ...
    • Group Classifications Used in Bias Research
      • Racial and Ethnic Classifications
      • Gender Classifications
        • Binary categories
        • Nonbinary identities
      • Emerging Recommendations Regarding Race/Ethnicity and Sex/Gender Categorizations
    • Testing Bias
    • Testing Measurement Invariance: Multigroup Confirmatory Factor Analysis
      • Levels of Invariance
        • Configural invariance
        • Metric invariance
        • Scalar invariance
        • Residual invariance
      • Evaluation of Invariance
      • Problems With MGCFA and Alternative Methods
        • Alignment method
        • Moderated nonlinear factor analysis (MNLFA)
    • Statistical and Conceptual Issues Surrounding Evaluation of Measurement Invariance
      • Criteria for Making Decisions About Invariance
      • Referent Indicators (RI) and Methods Identifying Invariant RI
      • Study Characteristics and Patterns of Noninvariance
      • Effect Size Measures
      • Appropriateness of CFA
      • Group Classifications
    • Conclusions and Recommendations
    • References