Ethical and Professional Issues in Psychology Testing
CHAPTER 4 Validity and Test
Development TOPIC 4A Basic Concepts of Validity 4.1 Validity: A Definition 4.2 Content Validity 4.3 Criterion-Related Validity 4.4 Construct Validity 4.5 Approaches to Construct Validity 4.6 Extravalidity Concerns and the Widening Scope of Test Validity
As most every student of psychology knows, the merit of a psychological test is determined first by its reliability but then ultimately by its validity. In the preceding chapter we pointed out that reliability can be appraised by many seemingly diverse methods ranging from the conceptually straightforward test–retest approach to the theoretically more complex
methodologies of internal consistency. Yet, regardless of the method used, the assessment of reliability invariably boils down to a simple summary statistic, the reliability coefficient. In this chapter, the more difficult and complex issue of validity—what a test score means—is investigated. The concept of validity is still evolving and, therefore, stirs up a great deal more controversy than its staid and established cousin, reliability (AERA, APA, & NCME, 1999). In Topic 4A, Basic Concepts of Validity, we introduce essential concepts of validity, including the standard tripartite division into content, criterion-related, and construct validity. We also discuss extravalidity concerns, which include side effects and unintended consequences of testing. Extravalidity concerns have fostered a wider definition of test validity that extends beyond the technical notions of content, criteria, and constructs. In Topic 4B, Test Construction, we stress that validity must be built into the test from the outset rather than being limited to the final stages of test development.
Put simply, the validity of a test is the extent to which it measures what it claims to measure. Psychometricians have long acknowledged that validity is the most fundamental and important characteristic of a test. After all, validity defines the meaning of test scores. Reliability is important, too, but only insofar as it constrains validity. To the extent that a test is unreliable, it cannot be valid. We can express this point from an alternative perspective: Reliability is a necessary but not a sufficient precursor of validity. Test developers have a responsibility to demonstrate that new instruments fulfill the purposes for which they are designed. However, unlike test reliability, test validity is not a simple issue that is easily resolved on the basis of a few rudimentary studies. Test validation is a developmental process that begins with test construction and continues indefinitely: After a test is released for operational use, the
interpretive meaning of its scores may continue to be sharpened, refined, and enriched through the gradual accumulation
of clinical observations and through special research projects. . . . Test validity is a living thing; it is not dead and embalmed when the test is released. (Anastasi, 1986)
Test validity hinges upon the accumulation of research findings. In the sections that follow, we examine the kinds of evidence sought in the validation of a psychological test.
4.1 VALIDITY: A DEFINITION We begin with a definition of validity, paraphrased from the influential Standards for Educational and Psychological Testing (AERA, APA, & NCME, 1999): A test is valid to the extent that inferences made
from it are appropriate, meaningful, and useful.
Notice that a test score per se is meaningless until the examiner draws inferences from it based on the test manual or other research findings. For example, knowing that an examinee has obtained a slightly elevated score on the MMPI-2 Depression scale is not
particularly helpful. This result becomes valuable only when the examiner infers behavioral characteristics from it. Based on existing research, the examiner might conclude, “The elevated Depression score suggests that the examinee has little energy and has a pessimistic outlook on life.” The MMPI-2 Depression scale possesses psychometric validity to the extent that such inferences are appropriate, meaningful, and useful. Unfortunately, it is seldom possible to summarize the validity of a test in terms of a single, tidy statistic. Determining whether inferences are appropriate, meaningful, and useful typically requires numerous studies of the relationships between test performance and other independently observed behaviors. Validity reflects an evolutionary, research-based judgment of how adequately a test measures the attribute it was designed to measure. Consequently, the validity of tests is not easily captured by neat statistical summaries but is instead characterized on a continuum ranging from weak to acceptable to strong.
Traditionally, the different ways of accumulating validity evidence have been grouped into three categories: • Content validity • Criterion-related validity • Construct validity
We will expand on this tripartite view of validity shortly, but first a few cautions. The use of these convenient labels does not imply that there are distinct types of validity or that a specific validation procedure is best for one test use and not another: An ideal validation includes several types of
evidence, which span all three of the traditional categories. Other things being equal, more sources of evidence are better than fewer. However, the quality of the evidence is of primary importance, and a single line of solid evidence is preferable to numerous lines of evidence of questionable quality. Professional judgment should guide the decisions regarding the forms of evidence that are most necessary and feasible in light of the intended uses of the
test and any likely alternatives to testing. (AERA, APA, & NCME, 1985)
We may summarize these points by stressing that validity is a unitary concept determined by the extent to which a test measures what it purports to measure. The inferences drawn from a valid test are appropriate, meaningful, and useful. In this light, it should be apparent that virtually any empirical study that relates test scores to other findings is a potential source of validity information (Anastasi, 1986; Messick, 1995).
4.2 CONTENT VALIDITY Content validity is determined by the degree to which the questions, tasks, or items on a test are representative of the universe of behavior the test was designed to sample. In theory, content validity is really nothing more than a sampling issue (Bausell, 1986). The items of a test can be visualized as a sample drawn from a larger population of potential items that define what the researcher really wishes to measure. If the
sample (specific items on the test) is representative of the population (all possible items), then the test possesses content validity. Content validity is a useful concept when a great deal is known about the variable that the researcher wishes to measure. With achievement tests in particular, it is often possible to specify the relevant universe of behaviors in advance. For example, when developing an achievement test of spelling, a researcher could identify nearly all possible words that third graders should know. The content validity of a third- grade spelling achievement test would be assured, in part, if words of varying difficulty level were randomly sampled from this preexisting list. However, test developers must take care to specify the relevant universe of responses as well. All too often, a multiple-choice format is taken for granted: If the constructor thinks about his aims with an
open mind he will often decide that the task should call for a response constructed by the student—written open-end responses
or, if inhibitions are to be minimized, oral responses. Nor are the directions to the subject and the social setting of the test to be neglected in defining the task. (Cronbach, 1971)
In reference to spelling achievement, it cannot be assumed that a multiple-choice test will measure the same spelling skills as an oral test or a frequency count of misspellings in written compositions. Thus, when evaluating content validity, response specification is also an integral part of defining the relevant universe of behaviors. Content validity is more difficult to assure when the test measures an ill-defined trait. How could a test developer possibly hope to specify the universe of potential items for a measure of anxiety? In these cases in which the measured trait is less tangible, no test developer in his or her right mind would try to construct the literal universe of potential test items. Instead, what usually passes for content validity is the considered opinion of expert judges. In effect, the test developer asserts that “a panel of experts
reviewed the domain specification carefully and judged the following test questions to possess content validity.” Figure 4.1 reproduces a sample judge’s item rating form for determining the content validity of test questions.
FIGURE 4.1 Sample Judges Item-Rating Form for Determining Content Validity Source: Based on Martuza (1977), Hambleton (1984), Bausell (1986). Quantification of Content Validity Martuza (1977) and others have discussed statistical methods for determining the overall content validity of a test from the judgments of experts. These methods tend to be very specialized and have not been widely accepted. Nonetheless, their approaches can serve as a model for a commonsense viewpoint on
interrater agreement as a basis for content validity. When two expert judges evaluate individual items of a test on the four-point scale proposed in Figure 4.1, the ratings of each judge on each item can be dichotomized into weak relevance (ratings of 1 or 2) versus strong relevance (ratings of 3 or 4). For each item, then, the conjoint ratings of the two judges can be entered into the two-by-two agreement table depicted in Figure 4.2. For example, if both judges believed an item was quite relevant (strong relevance), it would be placed in cell D. If the first judge believed an item was very relevant (strong relevance) but the second judge deemed it be only slightly relevant (weak relevance), the item would be placed in cell B. Notice that cell D is the only cell that reflects valid agreement between judges. The other cells involve disagreement (cells B and C) or agreement that an item doesn’t belong on the test (cell A). We have reproduced hypothetical results for a 100-item test in Figure 4.3. A
coefficient of content validity can be derived from the following formula:
FIGURE 4.2 Interrater Agreement Model for Content Validity For example, on our 100-item test both judges concurred that 87 items were strongly relevant (cell D), so the coefficient of content validity would be 87/(4 + 4 + 5 + 87) or .87. If more than two judges are used, this computational procedure could be completed with all possible
pair-wise combinations of judges, and the average coefficient reported. An important note: A coefficient of content validity is just one piece of evidence in the evaluation of a test. Such a coefficient does not by itself establish the validity of a test. The commonsense approach to content validity advocated here serves well as a flagging mechanism to help cull out existing items that are deemed inappropriate by expert raters. However, it cannot identify nonexistent items that should be added to a test to help make the pool of questions more representative of the intended domain. A test could possess a robust coefficient of content validity and still fall short in subtle ways. Quantification of content validity is no substitute for careful selection of items. Face Validity
FIGURE 4.3 Hypothetical Example of Agreement Model of Content Validity for a 100-Item Test We digress briefly here to mention face validity, which is not really a form of validity at all. Nonetheless, the concept is encountered in testing and, therefore, needs brief explanation. A test has face validity if it looks valid to test users, examiners, and especially the examinees. Face validity is really a matter of social acceptability and not a technical form of validity in the same category as content, criterion-
related, or construct validity (Nevo, 1985). From a public relations standpoint, it is crucial that tests possess face validity—otherwise those who take the tests may be dissatisfied and doubt the value of psychological testing. However, face validity should not be confused with objective validity, which is determined by the relationship of test scores to other sources of information. In fact, a test could possess extremely strong face validity—the items might look highly relevant to what is presumably measured by the instrument—yet produce totally meaningless scores with no predictive utility whatever.
4.3 CRITERION-RELATED VALIDITY Criterion-related validity is demonstrated when a test is shown to be effective in estimating an examinee’s performance on some outcome measure. In this context, the variable of primary interest is the outcome measure, called a criterion. The test score is useful only
insofar as it provides a basis for accurate prediction of the criterion. For example, a college entrance exam that is reasonably accurate in predicting the subsequent grade point average of examinees would possess criterion-related validity. Two different approaches to validity evidence are subsumed under the heading of criterion- related validity. In concurrent validity, the criterion measures are obtained at approximately the same time as the test scores. For example, the current psychiatric diagnosis of patients would be an appropriate criterion measure to provide validation evidence for a paper-and-pencil psychodiagnostic test. In predictive validity, the criterion measures are obtained in the future, usually months or years after the test scores are obtained, as with the college grades predicted from an entrance exam. Each of these two approaches is best suited to different testing situations, discussed in the following sections. However, before we review the nature of concurrent and predictive validity,
let us examine a more fundamental question: What are the characteristics of a good criterion? Characteristics of a Good Criterion As noted, a criterion is any outcome measure against which a test is validated. In practical terms, a criterion can be most anything. Some examples will help to illustrate the diversity of potential criteria. A simulator-based driver skill test might be validated against a criterion of “number of traffic citations received in the last 12 months.” A scale measuring social readjustment might be validated against a criterion of “number of days spent in a psychiatric hospital in the last three years.” A test of sales potential might be validated against a criterion of “dollar amount of goods sold in the preceding year.” The choice of criteria is circumscribed, in part, by the ingenuity of the test developer. However, criteria must be more than just imaginative; they must also be reliable, appropriate, and free of contamination from the test itself.
The criterion must itself be reliable if it is to be a useful index of what the test measures. If you recall the meaning of reliability—consistency of scores—the need for a reliable criterion measure is intuitively obvious. After all, unreliable means unpredictable. An unreliable criterion will be inherently unpredictable, regardless of the merits of the test. Consider the case in which scores on a college entrance exam (the test) are used to predict subsequent grade point average (the criterion). The validity of the entrance exam could be studied by computing the correlation (rxy) between entrance exam scores and grade point averages for a representative sample of students. For purposes of a validity study, it would be ideal if the students were granted open or unscreened enrollment so as to prevent a restriction of range on the criterion variable. In any case, the resulting correlation coefficient is called a validity coefficient.1 The theoretical upper limit of the validity coefficient is constrained by the reliability of both the test and the criterion:
The validity coefficient is always less than or equal to the square root of the test reliability multiplied by the criterion reliability. In other words, to the extent that the reliability of either the test or the criterion (or both) is low, the validity coefficient is also diminished. Returning to our example of an entrance exam used to predict college grade point average, we must conclude that the validity coefficient for such a test will always fall far short of +1.00, owing in part to the unreliability of college grades and also in part to the unreliability of the test itself. A criterion measure must also be appropriate for the test under investigation. The Standards for Educational and Psychological Testing sourcebook (AERA, APA, & NCME, 1985) incorporates this important point as a separate standard: All criterion measures should be described
accurately, and the rationale for choosing
them as relevant criteria should be made explicit.
For example, in the case of interest tests, it is sometimes unclear whether the criterion measure should indicate satisfaction, success, or continuance in the activities under question. The choice between these subtle variants in the criterion must be made carefully, based on an analysis of what the interest test purports to measure. A criterion must also be free of contamination from the test itself. Lehman (1978) has illustrated this point in a criterion-related validity study of a life change measure. The Schedule of Recent Events, or the SRE (Holmes & Rahe, 1967) is a widely used instrument that provides a quantitative index of the accumulation of stressful life events (e.g., divorce, job promotion, traffic tickets). Scores on the SRE correlate modestly with such criterion measures as physical illness and psychological disturbance. However, many seemingly appropriate criterion measures incorporate items that are similar or identical to
SRE items. For example, screening tests of psychiatric symptoms often check for changes in eating, sleeping, or social activities. Unfortunately, the SRE incorporates questions that check for the following: • Change in eating habits • Change in sleeping habits • Change in social activities
If the screening test contains the same items as the SRE, then the correlation between these two measures will be artificially inflated. This potential source of error in test validation is referred to as criterion contamination, since the criterion is “contaminated” by its artificial commonality with the test. Criterion contamination is also possible when the criterion consists of ratings from experts. If the experts also possess knowledge of the examinees’ test scores, this information may (consciously or unconsciously) influence their ratings. When validating a test against a criterion of expert ratings, the test scores must be held in strictest confidence until the ratings have been collected.
Now that the reader knows the general characteristics of a good criterion, we will review the application of this knowledge in the analysis of concurrent and predictive validity. Concurrent Validity In a concurrent validation study, test scores and criterion information are obtained simultaneously. Concurrent evidence of test validity is usually desirable for achievement tests, tests used for licensing or certification, and diagnostic clinical tests. An evaluation of concurrent validity indicates the extent to which test scores accurately estimate an individual’s present position on the relevant criterion. For example, an arithmetic achievement test would possess concurrent validity if its scores could be used to predict, with reasonable accuracy, the current standing of students in a mathematics course. A personality inventory would possess concurrent validity if diagnostic classifications derived from it roughly matched the opinions of psychiatrists or clinical psychologists.
A test with demonstrated concurrent validity provides a shortcut for obtaining information that might otherwise require the extended investment of professional time. For example, the case assignment procedure in a mental health clinic can be expedited if a test with demonstrated concurrent validity is used for initial screening decisions. In this manner, severely disturbed patients requiring immediate clinical workup and intensive treatment can be quickly identified by paper-and-pencil test. Of course, tests are not intended to replace mental health specialists, but they can save time in the initial phases of diagnosis. Correlations between a new test and existing tests are often cited as evidence of concurrent validity. This has a catch-22 quality to it—old tests validating a new test—but is nonetheless appropriate if two conditions are met. First, the criterion (existing) tests must have been validated through correlations with appropriate nontest behavioral data. In other words, the network of interlocking relationships must touch ground with real-world behavior at some point.
Second, the instrument being validated must measure the same construct as the criterion tests. Thus, it is entirely appropriate that developers of a new intelligence test report correlations between it and established mainstays such as the Stanford-Binet and Wechsler scales. Predictive Validity In a predictive validation study, test scores are used to estimate outcome measures obtained at a later date. Predictive validity is particularly relevant for entrance examinations and employment tests. Such tests share a common function—determining who is likely to succeed at a future endeavor. A relevant criterion for a college entrance exam would be first-year- student grade point average, while an employment test might be validated against supervisor ratings after six months on the job. In the ideal situation, such tests are validated during periods of open enrollment (or open hiring) so that a full range of results is possible on the outcome measures. In this manner, future use of the test as a selection device for
excluding low-scoring applicants will rest on a solid foundation of validational data. When tests are used for purposes of prediction, it is necessary to develop a regression equation. A regression equation describes the best-fitting straight line for estimating the criterion from the test. We will not discuss the statistical approach to fitting the straight line, except to mention that it minimizes the sum of the squared deviations from the line (Ghiselli, Campbell, & Zedeck, 1981). For current purposes, it is more important to understand the nature and function of regression equations. Ghiselli and associates (1981) provide a simple example of regression in the service of prediction, summarized here. Suppose we are trying to predict success on a job Y (evaluated by the supervisor on a 7-point scale ranging from poor to excellent performance) from scores on a preemployment test X (with scores that range from a low of 0 to a high of 100). The regression equation Y = .07X + .2
might describe the best-fitting straight line and therefore produce the most accurate predictions. For an individual who scored 55 on the test, the predicted performance level would be 4.05; that is, .07(55) + .2. A test score of 33 yields a predicted performance level of 2.51, that is, .07(33) + .2. Additional predictions are made likewise. Validity Coefficient and the Standard Error of the Estimate The relationship between test scores and criterion measures can be expressed in several different ways. Perhaps the most popular approach is to compute the correlation between test and criterion (rxy). In this context, the resulting correlation is known as a validity coefficient. The higher the validity coefficient rxy, the more accurate is the test in predicting the criterion. In the hypothetical case where rxy is 1.00, the test would possess perfect validity and allow for flawless prediction. Of course, no such test exists, and validity coefficients are more commonly in the low- to midrange of
correlations and rarely exceed .80. But how high should a validity coefficient be? There is no general answer to this question. However, we can approach the question indirectly by investigating the relationship between the validity coefficient and the corresponding error of estimate. The standard error of estimate (SEest) is the margin of error to be expected in the predicted criterion score. The error of estimate is derived from the following formula:
In this formula, is the square of the validity coefficient and SDy is the standard deviation of the criterion scores. Perhaps the reader has noticed the similarities between this index and the standard error of measurement (SEM). In fact, both indices help gauge margins of error. The SEM indicates the margin of measurement error caused by unreliability of the test, whereas SEest indicates the margin of prediction error caused by the imperfect validity of the test.
The SEest helps answer the fundamental question: “How accurately can criterion performance be predicted from test scores?” (AERA, APA, & NCME, 1985). Consider the common practice of attempting to predict college grade point average from high school scores on a scholastic aptitude test. For a specific aptitude test, suppose we determine that the SEest for predicted grade point average is .2 (on the usual 0.0 to 4.0 grade point scale). What does this mean for the examinee whose college grade point is predicted to be 3.1? As is the case with all standard deviations, the standard error of the estimate can be used to bracket predicted outcomes in a probabilistic sense. Assuming that the frequency distribution of grades is normal, we know that the chances are about 68 in 100 that the examinee’s predicted grade point will fall between 2.9 and 3.3 (plus or minus one SEest). In like manner, we know that the chances are about 95 in 100 that the examinee’s predicted grade point will fall between 2.7 and 3.5 (plus or minus two SEest).
What is an acceptable standard of predictive accuracy? There is no simple answer to this question. As the reader will discern from the discussion that follows, standards of predictive accuracy are, in part, value judgments. To explain why this is so, we need to introduce the basic elements of decision theory (Taylor & Russell, 1939; Cronbach & Gleser, 1965). Decision Theory Applied to Psychological Tests Proponents of decision theory stress that the purpose of psychological testing is not measurement per se but measurement in the service of decision making. The personnel manager wishes to know whom to hire; the admissions officer must choose whom to admit; the parole board desires to know which felons are good risks for early release; and the psychiatrist needs to determine which patients require hospitalization. The link between testing and decision making is nowhere more obvious than in the context of predictive validation studies. Many of these
studies use test results to determine who will likely succeed or fail on the criterion task so that, in the future, examinees with poor scores on the predictor test can be screened from admission, employment, or other privilege. This is the rationale by which admissions officers or employers require applicants to obtain a certain minimum score on an appropriate entrance or employment exam—previous studies of predictive validity can be cited to show that candidates scoring below a certain cutoff face steep odds in their educational or employment pursuits. Psychological tests frequently play a major role in these kinds of institutional decision making. In a typical institutional decision, a committee —or sometimes a single person—makes a large number of comparable decisions based on a cutoff score on one or more selection tests. In order to present the key concepts of decision theory, let us oversimplify somewhat and assume that only a single test is involved. Even though most tests produce a range of scores along a continuum, it is usually possible
to identify a cutoff or pass/fail score that divides the sample into those predicted to succeed versus those predicted to fail on the criterion of interest. Let us assume that persons predicted to succeed are also selected for hiring or admission. In this case, the proportion of persons in the “predicted-to-succeed” group is referred to as the selection ratio. The selection ratio can vary from 0 to 1.0, depending on the proportion of persons who are considered good bets to succeed on the criterion measure. If the results of a selection test allow for the simple dichotomy of “predicted to succeed” versus “predicted to fail,” then the subsequent outcome on the criterion measure likewise can be split into two categories, namely, “did succeed” and “did fail.” From this perspective, every study of predictive validity produces a two-by-two matrix, as portrayed in Figure 4.4. Certain combinations of predicted and actual outcomes are more likely than others. If a test has good predictive validity, then most persons predicted to succeed will succeed and most persons predicted to fail will fail. These are
examples of correct predictions and serve to bolster the validity of a selection instrument. Outcomes in these two cells are referred to as hits because the test has made a correct prediction.
FIGURE 4.4 Possible Outcomes When a Selection Test Is Used to Predict Performance on a Criterion Measure But no selection test is a perfect predictor, so two other types of outcomes are also possible. Some persons predicted to succeed will, in fact,
fail. These cases are referred to as false positives. And some persons predicted to fail would, if given the chance, succeed. These cases are referred to as false negatives. False positives and false negatives are collectively known as misses, because in both cases the test has made an inaccurate prediction. Finally, the hit rate is the proportion of cases in which the test accurately predicts success or failure, that is, hit rate = (hits)/(hits + misses). False positives and false negatives are unavoidable in the real-world use of selection tests. The only way to eliminate such selection errors would be to develop a perfect test, an instrument which has a validity coefficient of +1.00, signifying a perfect correlation with the criterion measure. A perfect test is theoretically possible, but none has yet been observed on this planet. Nonetheless, it is still important to develop selection tests with very high predictive validity, so as to minimize decision errors. Proponents of decision theory make two fundamental assumptions about the use of selection tests:
1. The value of various outcomes to the institution can be expressed in terms of a common utility scale. One such scale—but by no means the only one—is profit and loss. For example, when using an interest inventory to select salespersons, a corporation can anticipate profit from applicants correctly identified as successful but will lose money when, inevitably, some of those selected do not sell enough even to support their own salary (false positives). The cost of the selection procedure must also be factored in to the utility scale as well.
2. In institutional selection decisions, the most generally useful strategy is one that maximizes the average gain on the utility scale (or minimizes average loss) over many similar decisions. For example, which selection ratio produces the largest average gain on the utility scale? Maximization is, thus, the fundamental decision principle.
The application of decision theory is much more complicated than illustrated here, mainly
because of the difficulty of finding a common utility scale for different outcomes. Consider the plight of the admissions officer at any large university. If the selection ratio is quite strict, then most of the admitted students will also succeed. But some students not admitted might have succeeded, too, and their financial support to the university (tuition, fees) is, therefore, lost. However, if the selection ratio is too lenient, then the percentage of false positives (students admitted who subsequently fail) skyrockets. How is the cost of a false positive to be calculated? The financial cost can be estimated —for example, advisers dedicate a certain number of hours at a known pay rate counseling these students. But no single utility scale can encompass the other diverse consequences such as the need for additional remedial services (which require money), the increase in faculty cynicism (an issue of morale), and the dashed hopes of misled students (whose heartbreak affects public perception of the university and may even influence future state funding!). Clearly, the neat statistical notions of decision
theory oversimplify the complex influences that determine utility in the real world. Nonetheless, in large institutional settings where a common utility scale can be identified, principles of decision theory can be applied to selection problems with thought-provoking results. For example, Schmidt, Hunter, McKenzie, and Muldrow (1979) analyzed the potential impact of using the Programmer Aptitude Test (PAT, Hughes & McNamara, 1959) in the selection of computer programmers by the federal government. They based their analysis on the following facts and assumptions: 1. PAT scores and measures of later on-the-job
programming performance correlate quite substantially; the validity coefficient of the PAT is .76 (fact).
2. The government hires 600 new programmers each year (fact).
3. The cost of testing is about $10 per examinee (fact).
4. Programmers stay on the job for about nine years and receive pay raises according to a known pay scale (fact).
5. The yearly productivity in dollars of low- performing, average, and superior programmers can be accurately estimated by supervisors (assumption).
Based on these facts and assumptions, Schmidt et al. (1979) then compared the hypothetical use of the PAT against other selection procedures of lesser validity. Since the usefulness of a test is partly determined by the percentage of applicants who are selected for employment, the researchers also looked at the impact of different selection ratios on overall productivity. In each case, they estimated the yearly increase in dollar-amount productivity from using the PAT instead of an alternative and less efficacious procedure. In general, the use of the PAT was estimated to increase productivity by tens of millions of dollars. The specific estimated increase depended on the selection ratio and the validity coefficient of hypothetical alternative procedures. For example, if 80 percent of the applicants were hired (selection ratio of .80), using the PAT would increase the productivity of the federal government by at
least $5.6 million (if the alternative procedure had a validity coefficient of .50) and possibly as much as $16.5 million (if the alternative procedure had no validity at all). If the selection ratio were quite small, the use of the PAT for selection boosted productivity even more— possibly as much as nearly $100 million. Schmidt et al. (1979) concluded that “the impact of valid selection procedures on work-force productivity is considerably greater than most personnel psychologists have believed.” 1We have purposefully refrained from referring to such a statistic as the validity coefficient. Remember that validity is a unitary concept determined by multiple sources of information that may include the correlation between test and criterion.
4.4 CONSTRUCT VALIDITY The final type of validity discussed in this unit is construct validity, and it is undoubtedly the most difficult and elusive of the bunch. A construct is a theoretical, intangible quality or trait in which individuals differ (Messick, 1995). Examples of constructs include leadership ability, overcontrolled hostility, depression, and
intelligence. Notice in each of these examples that constructs are inferred from behavior but are more than the behavior itself. In general, constructs are theorized to have some form of independent existence and to exert broad but to some extent predictable influences on human behavior. A test designed to measure a construct must estimate the existence of an inferred, underlying characteristic (e.g., leadership ability) based on a limited sample of behavior. Construct validity refers to the appropriateness of these inferences about the underlying construct. All psychological constructs possess two characteristics in common: 1. There is no single external referent sufficient
to validate the existence of the construct; that is, the construct cannot be operationally defined (Cronbach & Meehl, 1955).
2. Nonetheless, a network of interlocking suppositions can be derived from existing theory about the construct (AERA, APA, & NCME, 1985).
We will illustrate these points by reference to the construct of psychopathy (Cleckley, 1976), a personality constellation characterized by antisocial behavior (lying, stealing, and occasionally violence), a lack of guilt and shame, and impulsivity.2 Psychopathy is surely a construct, in that there is no single behavioral characteristic or outcome sufficient to determine who is strongly psychopathic and who is not. On average we might expect psychopaths to be frequently incarcerated, but so are many common criminals. Furthermore, many successful psychopaths somehow avoid apprehension altogether (Cleckley, 1976). Psychopathy cannot be gauged only by scrapes with the law. Nonetheless, a network of interlocking suppositions can be derived from existing theory about psychopathy. The fundamental problem in psychopathy is presumed to be a deficiency in the ability to feel emotional arousal—whether empathy, guilt, fear of punishment, or anxiety under stress (Cleckley, 1976). A number of predictions follow from this appraisal. For
example, psychopaths should lie convincingly, have a greater tolerance for physical pain, show less autonomic arousal in the resting state, and get into trouble because of their lack of behavioral inhibition. Thus, to validate a measure of psychopathy, we would need to check out a number of different expectations based on our theory of psychopathy. Construct validity pertains to psychological tests that claim to measure complex, multifaceted, and theory-bound psychological attributes such as psychopathy, intelligence, leadership ability, and the like. The crucial point to understand about construct validity is that “no criterion or universe of content is accepted as entirely adequate to define the quality to be measured” (Cronbach & Meehl, 1955). Thus, the demonstration of construct validity always rests on a program of research using diverse procedures outlined in the following sections. To evaluate the construct validity of a test, we must amass a variety of evidence from numerous sources.
Many psychometric theorists regard construct validity as the unifying concept for all types of validity evidence (Cronbach, 1988; Messick, 1995). According to this viewpoint, individual studies of content, concurrent, and predictive validity are regarded merely as supportive evidence in the cumulative quest for construct validation. 2The construct of psychopathy is very similar to what is now designated as antisocial personality disorder (American Psychiatric Association, 1994).
4.5 APPROACHES TO CONSTRUCT VALIDITY How does a test developer determine whether a new instrument possesses construct validity? As previously hinted, no single procedure will suffice for this difficult task. Evidence of construct validity can be found in practically any empirical study that examines test scores from appropriate groups of subjects. Most studies of construct validity fall into one of the following categories:
• Analysis to determine whether the test items or subtests are homogeneous and therefore measure a single construct
• Study of developmental changes to determine whether they are consistent with the theory of the construct
• Research to ascertain whether group differences on test scores are theory- consistent
• Analysis to determine whether intervention effects on test scores are theory-consistent
• Correlation of the test with other related and unrelated tests and measures
• Factor analysis of test scores in relation to other sources of information
• Analysis to determine whether test scores allow for the correct classification of examinees
We examine these sources of construct validity evidence in more detail in the following. Test Homogeneity If a test measures a single construct, then its component items (or subtests) likely will be
homogeneous (also referred to as internally consistent). In most cases, homogeneity is built into the test during the development process discussed in more detail in the next unit. The aim of test development is to select items that form a homogeneous scale. The most commonly used method for achieving this goal is to correlate each potential item with the total score and select items that show high correlations with the total score. A related procedure is to correlate subtests with the total score in the early phases of test development. In this manner, wayward scales that do not correlate to some minimum degree with the total test score can be revised before the instrument is released for general use. Homogeneity is an important first step in certifying the construct validity of a new test, but standing alone it is weak evidence. Kline (1986) has pointed out the circularity of the procedure: If all our items in the item pool were wide of the
mark and did not measure what we hoped, they would be selecting items by the
criterion of their correlation with the total score, which can never work. It is to be noted that the same argument applies to the factoring of the item pool. A general factor of poor items is still possible. This objection is sound and has to be refuted empirically. Having found by item analysis a set of homogeneous items, we must still present evidence concerning their validity. Thus to construct a homogeneous test is not sufficient, validity studies must be carried out.
In addition to demonstrating the homogeneity of items, a test developer must provide multiple other sources of construct validity, discussed subsequently. Appropriate Developmental Changes Many constructs can be assumed to show regular age-graded changes from early childhood into mature adulthood and perhaps beyond. Consider the construct of vocabulary knowledge as an example. It has been known since the inception of intelligence tests at the
turn of the century that knowledge of vocabulary increases exponentially from early childhood into late childhood. More recent research demonstrates that vocabulary continues to grow, albeit at a slower pace, into old age (Gregory & Gernert, 1990). For any new test of vocabulary, then, an important piece of construct validity evidence would be that older subjects score better than younger subjects, assuming that education and health factors are held constant. Of course, not all constructs lend themselves to predictions about developmental changes. For example, it is not clear whether a scale measuring “assertiveness” should show a pattern of increasing, decreasing, or stable scores with advancing age. Developmental changes would be irrelevant to the construct validity of such a scale. We should also mention that appropriate developmental changes are but one piece in the construct validity puzzle. This approach does not provide information about how the construct relates to other constructs.
Theory-Consistent Group Differences One way to bolster the validity of a new instrument is to show that, on average, persons with different backgrounds and characteristics obtain theory-consistent scores on the test. Specifically, persons thought to be high on the construct measured by the test should obtain high scores, whereas persons with presumably low amounts of the construct should obtain low scores. Crandall (1981) developed a social interest scale that illustrates the use of theory-consistent group differences in the process of construct validation. Borrowing from Alfred Adler, Crandall (1984) defined social interest as an “interest in and concern for others.” To measure this construct, he devised a brief and simple instrument consisting of 15 forced-choice items. For each item, one of the two alternatives includes a trait closely related to the Adlerian concept of social interest (e.g., helpful), whereas the other choice consists of an equally attractive but nonsocial trait (e.g., quick-witted). The
subject is instructed to “choose the trait which you value more highly.” Each of the 15 items is scored 1 if the social interest trait is picked, 0 otherwise; thus, total scores on the Social Interest Scale (SIS) can range from 0 to 15. Table 4.1 presents average scores on the SIS for 13 well-defined groups of subjects. The reader will notice that individuals likely to be high in social interest (e.g., nuns) obtain the highest average scores on the SIS, whereas the lowest scores are earned by presumably self-centered persons (e.g., models) and those who are outright antisocial (felons). These findings are theory-consistent and support the construct validity of this interesting instrument. TABLE 4.1 Mean Scores on the Social Interest Scale for Selected Groups Group N Mean
Score Ursuline sisters 6 13.3 Adult church members 147 11.2 Charity volunteers 9 10.8 High school students nominated for high social interest
23 10.2
Source: Adapted with permission from Crandall, J. (1981). Theory and measurement of social interest: Empirical tests of Alfred Adler’s concept. New York: Columbia University Press. Theory-Consistent Intervention Effects Another approach to construct validation is to show that test scores change in appropriate direction and amount in reaction to planned or unplanned interventions. For example, the scores of elderly persons on a spatial orientation test battery should increase after these subjects receive cognitive training specifically designed to enhance their spatial orientation abilities.
University students nominated for high social interest
21 9.5
University employees 327 8.9 University students 1,784 8.2 University students nominated for low social interest
35 7.4
Professional models 54 7.1 High school students nominated for low social interest
22 6.9
Adult atheists and agnostics 30 6.7 Convicted felons 30 6.4
More precisely, if the test battery possesses construct validity, we can predict that spatial orientation scores should show a greater increase from pretest to posttest than found on unrelated abilities not targeted for special training (e.g., inductive reasoning, perceptual speed, numerical reasoning, or verbal reasoning). Willis and Schaie (1986) found just such a pattern of test results in a cognitive training study with elderly subjects, supporting the construct validity of their spatial orientation measure. Convergent and Discriminant Validation Convergent validity is demonstrated when a test correlates highly with other variables or tests with which it shares an overlap of constructs. For example, two tests designed to measure different types of intelligence should, nonetheless, share enough of the general factor in intelligence to produce a hefty correlation (say, .5 or above) when jointly administered to a heterogeneous sample of subjects. In fact, any
new test of intelligence that did not correlate at least modestly with existing measures would be highly suspect, on the grounds that it did not possess convergent validity. Discriminant validity is demonstrated when a test does not correlate with variables or tests from which it should differ. For example, social interest and intelligence are theoretically unrelated, and tests of these two constructs should correlate negligibly, if at all. In a classic paper often quoted but seldom emulated, Campbell and Fiske (1959) proposed a systematic experimental design for simultaneously confirming the convergent and discriminant validities of a psychological test. Their design is called the multitrait-multimethod matrix, and it calls for the assessment of two or more traits by two or more methods. Table 4.2 provides a hypothetical example of this approach. In this example, three traits (A, B, and C) are measured by three methods (1, 2, and 3). For example, traits A, B, and C might be social interest, creativity, and dominance. Methods 1, 2, and 3 might be self-report inventory, peer
ratings, and projective test. Thus, A1 would represent a self-report inventory of social interest, B2 a peer rating of creativity, C3 a dominance measure derived from projective test, and so on. TABLE 4.2 Hypothetical Multitrait- Multimethod Matrix
Notice in this example that nine tests are studied (three traits are each measured by three methods). When each of these tests is administered twice to the same group of subjects and scores on all pairs of tests are correlated, the result is a multitrait- multimethod matrix (Table 4.2). This matrix is
a rich source of data on reliability, convergent validity, and discriminant validity. For example, the correlations along the main diagonal (in parentheses) are reliability coefficients for each test. The higher these values, the better, and preferably we like to see values in the .80s or .90s here. The correlations along the three shorter diagonals (in boldface) supply evidence of convergent validity—the same trait measured by different methods. These correlations should be strong and positive, as shown here. Notice that the table also includes correlations between different traits measured by the same method (in solid triangles) and different traits measured by different methods (in dotted triangles). These correlations should be the lowest of all in the matrix, insofar as they supply evidence of discriminant validity. The Campbell and Fiske (1959) methodology is an important contribution to our understanding of the test validation process. However, the full implementation of this procedure typically requires too monumental a commitment from researchers. It is more common for test
developers to collect convergent and discriminant validity data in bits and pieces, rather than producing an entire matrix of intercorrelations. Meier (1984) provides one of the few real-world implementations of the multitrait-multimethod matrix in an examination of the validity of the “burnout” construct. Factor Analysis Factor analysis is a specialized statistical technique that is particularly useful for investigating construct validity. We discuss factor analysis in substantial detail in Topic 5A, Intelligence Tests and Factor Analysis; here, we provide a quick preview so that the reader can appreciate the role of factor analysis in the study of construct validity. The purpose of factor analysis is to identify the minimum number of determiners (factors) required to account for the intercorrelations among a battery of tests. The goal in factor analysis is to find a smaller set of dimensions, called factors, that can account for the observed array of intercorrelations among individual tests. A typical approach in factor
analysis is to administer a battery of tests to several hundred subjects and then calculate a correlation matrix from the scores on all possible pairs of tests. For example, if 15 tests have been administered to a sample of psychiatric and neurological patients, the first step in factor analysis is to compute the correlations between scores on the 105 possible pairs of tests.3 Although it may be feasible to see certain clusterings of tests that measure common traits, it is more typical that the mass of data found in a correlation matrix is simply too complex for the unaided human eye to analyze effectively. Fortunately, the computer- implemented procedures of factor analysis search this pattern of intercorrelations, identify a small number of factors, and then produce a table of factor loadings. A factor loading is actually a correlation between an individual test and a single factor. Thus, factor loadings can vary between −1.0 and +1.0. The final outcome of a factor analysis is a table depicting the correlation of each test with each factor.
We can illustrate the use of factor analysis in the study of construct validity by referring to a specific instrument, the Wechsler Adult Intelligence Scale-IV (WAIS-IV, Wechsler, 2008), discussed in more detail in the next chapter. The 10 core subtests of the WAIS-IV yield not only a Full Scale IQ, but also four Index scores designed to provide a meaningful and theoretically sound partition of intelligence into subcomponents. These Index scores are Verbal Comprehension (3 subtests), Perceptual Reasoning (3 subtests), Working Memory (2 subtests), and Processing Speed (2 subtests). When factor analysis is applied to WAIS-IV subtest scores for large samples of adults, four factors are found, just as predicted by the structure of the test (Ryan, Sattler, & Tree, 2009). Further, each core subtest usually demonstrates its highest factor loading on the appropriate factor. For example, the Vocabulary subtest shows its highest factor loading on Verbal Comprehension, and the Matrix Reasoning subtest reveals its highest factor loading on Perceptual Reasoning. Findings like
this bolster the construct validity of the WAIS- IV. Classification Accuracy Many tests are used for screening purposes to identify examinees who meet (or don’t meet) certain diagnostic criteria. For these instruments, accurate classification is an essential index of validity. As a basis for illustrating this approach to validation, we consider the Mini-Mental State Examination (MMSE), a short screening test of cognitive functioning. The MMSE consists of a number of simple questions (e.g., What day is this?) and easy tasks (e.g., remembering three words). The test yields a score from 0 (no items correct) to 30 (all items correct). Although used for many purposes, a major application of the MMSE is to identify elderly individuals who might be experiencing dementia. Dementia is a general term that refers to significant cognitive decline and memory loss caused by a disease process such as Alzheimer’s disease or the accumulation of small strokes. Both the MMSE and various
forms of dementia are described in more detail in Chapter 10, Neuropsychological Assessment and Screening. The MMSE is one of the most widely researched screening tests in existence. Much is known about its measurement qualities, such as the accuracy of the tool in detecting individuals with dementia. In exploring its utility, researchers have paid special attention to two psychometric features that bear upon validity: sensitivity and specificity. Sensitivity has to do with accurate identification of patients who have a syndrome—in this case, dementia. Specificity has to do with accurate identification of normal patients. These ideas are clarified later. Understanding these concepts is pertinent to the validity of every screening test used in mental health and medicine. Thus, we provide modest coverage here, using the MMSE as an exemplar of a more general principle. Our discussion loosely follows the presentation found in Gregory (1999). The concepts of sensitivity and specificity are chiefly helpful in dichotomous diagnostic
situations in which individuals are presumed either to manifest a syndrome or not. For example, in medicine a patient either has prostate cancer or he does not. In this case, the criterion of truth, against which a screening test is measured, would be a tissue biopsy. Similarly, in research studies on the sensitivity and specificity of the MMSE, patients are known from independent, comprehensive medical and psychological workups either to meet the criteria for dementia or not. This is the “gold standard” against which the screening instrument is validated. The rationale for the screening test is pragmatic: It is unrealistic to refer every patient with suspected dementia for comprehensive evaluations that would include, for example, many hours of professional time (psychologist, neurologist, geriatric specialist, etc.) and expensive brain scans. The purpose of the MMSE—or any screening test—is to determine the need for additional assessment. Screening tests typically provide a cutoff score used to identify possible cases of the syndrome in question. With the MMSE, a common cutting
score is 23/24 out of the 30 points possible. Thus, a score of 23 points and below indicates the likelihood of dementia, whereas 24 points and above is considered normal. In this context, the sensitivity of the MMSE is the percentage of patients known to have dementia who score 23 points or lower. For example, if 100 patients are known from independent, comprehensive evaluations to exhibit dementia, and 79 of them score 23 or below, then the sensitivity of the test is 79 percent. The specificity of the MMSE is the other side of the coin, the percentage of patients known to be normal who score 24 points or higher. For example, if 83 of 100 normal patients score 24 points or higher, then the specificity of the test is 83 percent. In general, the validity of a screening test is bolstered to the extent that it possesses both high sensitivity and high specificity. There are no exact cutoffs, but for many purposes a test will need sensitivity and specificity that exceed 80 or 90 percent in order to justify its use. As we will see later, the standards for sensitivity and specificity are unique to each situation and
depend on the costs—both financial and otherwise—of different kinds of errors in classification. An ideal screening test, of course, would yield 100 percent sensitivity and 100 percent specificity. No such test exists in the real world. The reality of assessment is that the examiner must choose a cutoff score that provides a balance between sensitivity and specificity. What makes this problematic is that sensitivity and specificity are inversely related. Choosing a cutoff score that increases sensitivity invariably will reduce specificity, and vice versa. The inverse relationship between sensitivity and specificity is not only an empirical fact, but it is also a logical necessity—if one improves, the other must decline—no exceptions are possible. Practitioners need to select a cutoff score that produces a livable balance between sensitivity and specificity. But exactly where is that point of equilibrium? In the case of the MMSE, the answer depends not just on the age and education of the client but also on the relative advantages and drawbacks of correct or
incorrect decisions. Robust levels of sensitivity and specificity provide corroborating evidence of test validity, and test developers should strive to achieve the highest possible levels of both. 3The general formula for the number of pairings among N tests is N(N − 1)/2. Thus, if 15 tests are administered, there will be 15 × 14/2 or 105 possible pairings of individual tests.
4.6 EXTRAVALIDITY CONCERNS AND THE WIDENING SCOPE OF TEST VALIDITY We begin this section with a review of extravalidity concerns, which include side effects and unintended consequences of testing. By acknowledging the importance of the extravalidity domain, psychologists confirm that the decision to use a test involves social, legal, and political considerations that extend far beyond the traditional questions of technical validity. In a related development, we will also review how the interest in extravalidity concerns has spurred several theorists to broaden the
concept of test validity. As the reader will discover, value implications and social consequences are now encompassed within the widening scope of test validity. Even if a test is valid, unbiased, and fair, the decision to use it may be governed by additional considerations. Cole and Moss (1998) outline the following factors: • What is the purpose for which the test is
used? • To what extent are the purposes
accomplished by the actions taken? • What are the possible side effects or
unintended consequences of using the test? • What possible alternatives to the test might
serve the same purpose? We survey only the most prominent extravalidity concerns here and show how they have served to widen the scope of test validity. Unintended Side Effects of Testing The intended outcome of using a psychological test is not necessarily the only consequence. Various side effects also are possible, indeed,
they are probable. The examiner must determine whether the benefits of giving the test outweigh the costs of the potential side effects. Furthermore, by anticipating unintended side effects, the examiner might be able to deflect or diminish them. Cole and Moss (1998) cite the example of using psychological tests to determine eligibility for special education. Although the intended outcome is to help students learn, the process of identifying students eligible for special education may produce numerous negative side effects: • The identified children may feel unusual or
dumb. • Other children may call the children names. • Teachers may view these children as
unworthy of attention. • The process may produce classes segregated
by race or social class. A consideration of side effects should influence an examiner’s decision to use a particular test for a specified purpose. The examiner might appropriately choose not to use a test for a
worthy purpose if the likely costs from side effects outweigh the expected benefits. Consider the common practice in years past of using the Minnesota Multiphasic Personality Inventory (MMPI) to help screen candidates for peace officer positions such as police officer or sheriff’s deputy. Although the MMPI was originally designed as an aid in psychiatric diagnosis, subsequent research indicated that it is also useful in the identification of persons unsuited to a career in law enforcement (Hiatt & Hargrave, 1988). In particular, peace officers who produce MMPI profiles with mild elevations (e.g., T score 65 to 69) on Scales F (Frequency), Masculinity-Femininity, Paranoia, and Hypomania tend to be involved in serious disciplinary actions; peace officers who produce more “defensive” MMPI profiles with fewer clinical scale elevations tend not to be involved in such actions. Thus, the test possessed modest validity for the worthy purpose of screening law enforcement candidates. But no test, not even the highly respected MMPI, is perfectly valid. Some good applicants will be passed over
because their MMPI results are marginal. Perhaps their Paranoia Scale is at a T score of 66, or the Hypomania Scale is at a T score of 68. On the MMPI, a T score of 70 is often considered the upper limit of the “normal” range. One unintended side effect of using the MMPI for evaluation of peace officer applicants is that job candidates who are unsuccessful with one agency may be tagged with a pathological label such as psychopathic, schizophrenic, or paranoid. The label may arise in spite of the best efforts of the consulting psychologist, who may never have used any pejorative terms in the assessment report on the candidate. Typically, the label is conceived when administrators at the referring department look at the MMPI profile and see that the candidate obtained his or her highest score on a scale with a horrendous title such as Psychopathic Deviate, Schizophrenia, Hypochondriasis, or Paranoia. Unfortunately, the law enforcement community can be a very closed fraternity. Police chiefs and sheriffs commonly exchange verbal reports about their
job applicants, so a pejorative label may follow the candidate from one setting to another, permanently barring the applicant from entry into the law enforcement profession. The repercussions are not only unfair to the candidate, but they also raise the specter of lawsuits against the agency and the consulting psychologist. All things considered, the consulting psychologist may find it preferable to use a technically less valid test for the same purpose, particularly if the alternative instrument does not produce these unintended side effects. The renewed sensitivity to extravalidity issues has caused several test theorists to widen their definition of test validity. We review these recent developments in the following section, cautioning the reader that a final consensus about the nature of test validity is yet to emerge. The Widening Scope of Test Validity By now the reader is familiar with the narrow, traditionalist perspective on test use, which states that a test is valid if it measures “what it
purports to measure.” The implicit implication of this perspective is that technical validity is the most essential basis for recommending test use. After all, valid tests provide accurate information about examinees—and what could be wrong with that? Recently, several psychometric theoreticians have introduced a wider, functionalist definition of validity that asserts that a test is valid if it serves the purpose for which it is used (Cronbach, 1988; Mes-sick, 1995). For example, a reading achievement test might be used to identify students for assignment to a remedial section. According to the functionalist perspective, the test would be valid—and its use, therefore, appropriate—if the students selected for remediation actually received some academic benefit from this application of the test. The functionalist perspective explicitly recognizes that the test validator has an obligation to determine whether a practice has constructive consequences for individuals and institutions and especially to guard against
adverse outcomes (Messick, 1980). Test validity, then, is an overall evaluative judgment of the adequacy and appropriateness of inferences and actions that flow from test scores. Messick (1980, 1995) argues that the new, wider conception of validity rests on four bases. These are (1) traditional evidence of construct validity, for example, appropriate convergent and discriminant validity, (2) an analysis of the value implications of the test interpretation, (3) evidence for the usefulness of test interpretations in particular applications, and (4) an appraisal of the potential and actual social consequences, including side effects, from test use. A valid test is one that answers well to all four facets of test validity. This wider conception of test validity is admittedly controversial, and some theorists prefer the traditional view that consequences and values are important but nonetheless separate from the technical issues of test validity. Everyone can agree on one point: Psychological measurement is not a neutral
endeavor, it is an applied science that occurs in a social and political context. Utility: The Last Horizon of Test Validity Finally, we introduce the concept of test utility, which is widely neglected in the research literature on psychological testing (Hunsley & Bailey, 1999). As noted by Wood, Garb, and Nezworski (2007), test utility can be summed up by the question “Does use of this test result in better patient outcomes or more efficient delivery of services?” For example, we might envision an experiment in which individual psychotherapy clients were randomly assigned to two groups. One group is tested with the Beck Depression Inventory-2 (Beck, Steer, & Brown, 1996) and the results provided to their therapists, while the other group is not tested but instead proceeds directly for treatment. If the tested group showed more improvement or required fewer sessions to achieve the same level of improvement, we would conclude that utility has been demonstrated for the test.
Unfortunately, there is very little research on the utility of psychological tests, and the research that does exist is indirect. For example, Finn and Tonsager (1992) have shown that a highly structured method for giving feedback on personality test findings to college students awaiting psychotherapy has initial therapeutic effects in its own right. However, this does not answer the question whether the ultimate client outcome is better as a result of the test usage. For some tests such as the Rorschach inkblot technique, discussed later in the text, the question of utility is especially pertinent because of the time required by a psychologist to administer, score, interpret, and document the results. The total time easily can run to many hours. It is lamentable that the utility of this instrument and many other tests has not been systematically investigated. TOPIC 4B Test Construction 4.7 Defining the Test 4.8 Selecting a Scaling Method 4.9 Representative Scaling Methods
4.10 Constructing the Items 4.11 Testing the Items 4.12 Revising the Test 4.13 Publishing the Test
Creating a new test involves both science and art. A test developer must choose strategies and materials and then make day-to-day research decisions that will affect the quality of his or her emerging instrument. The purpose of this section is to discuss the process by which psychometricians create valid tests. Although we will discuss many separate topics, they are united by a common theme: Valid tests do not just materialize on the scene in full maturity— they emerge slowly from an evolutionary, developmental process that builds in validity from the very beginning. We will emphasize the basics of test development here; readers who desire a more advanced presentation should consult Kline (1986), McDonald (1999), and Bernstein and Nunnally (1994).
Test construction consists of six intertwined stages:
By way of preview, we can summarize these steps as follows: defining the test consists of delimiting its scope and purpose, which must be known before the developer can proceed to test construction. Selecting a scaling method is a process of setting the rules by which numbers are assigned to test results. Constructing the items is as much art as science, and it is here that the creativity of the test developer may be required. Once a preliminary version of the test is available, the developer usually administers it to a modest-sized sample of subjects in order to collect initial data about test item characteristics. Testing the items entails a variety of statistical procedures referred to collectively as item analysis. The purpose of item analysis is to determine which items should be retained, which revised, and which thrown out. Based on item analysis and other sources of
Defining the test Testing the items Selecting a scaling method Revising the test Constructing the items Publishing the test
information, the test is then revised. If the revisions are substantial, new items and additional pretesting with new subjects may be required. Thus, test construction involves a feedback loop whereby second, third, and fourth drafts of an instrument might be produced (Figure 4.5). Publishing the test is the final step. In addition to releasing the test materials, the developer must produce a user-friendly test manual. Let us examine each of these steps in more detail.
4.7 DEFINING THE TEST
FIGURE 4.5 The Test Construction Process In order to construct a new test, the developer must have a clear idea of what the test is to measure and how it is to differ from existing instruments. Insofar as psychological testing is now entering its second one hundred years, and insofar as thousands of tests have already been published, the burden of proof clearly rests on the test developer to show that a proposed instrument is different from, and better than, existing measures. Consider the daunting task faced by a test developer who proposes yet another measure of general intelligence. With dozens of such instruments already in existence, how could a new test possibly make a useful contribution to the field? The answer is that contemporary research continually adds to our understanding of intelligence and impels us to seek new and more useful ways to measure this multifaceted construct. Kaufman and Kaufman (1983) provide a good model of the test definition process. In proposing the Kaufman Assessment Battery for
Children (K-ABC), a new test of general intelligence in children, the authors listed six primary goals that define the purpose of the test and distinguish it from existing measures: 1. Measure intelligence from a strong
theoretical and research basis 2. Separate acquired factual knowledge from
the ability to solve unfamiliar problems 3. Yield scores that translate to educational
intervention 4. Include novel tasks 5. Be easy to administer and objective to score 6. Be sensitive to the diverse needs of
preschool, minority group, and exceptional children (Kaufman & Kaufman, 1983)
The K-ABC represents an interesting departure from traditional intelligence tests. For now, the important point is that the developers of this instrument, now in its second edition (K-ABC- II), explained its purpose explicitly and proposed a fresh focus for measuring intelligence, long before they started constructing test items.
4.8 SELECTING A SCALING METHOD The immediate purpose of psychological testing is to assign numbers to responses on a test so that the examinee can be judged to have more or less of the characteristic measured. The rules by which numbers are assigned to responses define the scaling method. Test developers select a scaling method that is optimally suited to the manner in which they have conceptualized the trait(s) measured by their test. No single scaling method is uniformly better than the others. For some traits, ordinal ranking of expert judges might be the best measurement approach; for other traits, complex scaling of self-report data might yield the most valid measurements. There are so many distinctive scaling methods available to psychometricians that we will be satisfied to provide only a representative sample here. Readers who wish a more thorough and detailed review should consult Gulliksen (1950), Nunnally (1978), or Kline (1986). However, before reviewing selecting scaling methods, we
need to introduce a related concept, levels of measurement, so that the reader can better appreciate the differences between scaling methods. Levels of Measurement According to Stevens (1946), all numbers derived from measurement instruments of any kind can be placed into one of four hierarchical categories: nominal, ordinal, interval, or ratio. Each category defines a level of measurement; the order listed is from least to most informative. In a nominal scale, the numbers serve only as category names. For example, when collecting data for a demographic study, a researcher might code males as “1” and females as “2.” Notice that the numbers are arbitrary and do not designate “more” or “less” of anything. In nominal scales the numbers are just a simplified form of naming. An ordinal scale constitutes a form of ordering or ranking. If college professors were asked to rank order four cars as to which they would
prefer to own, the preferred order might be “1” Cadillac, “2” Chevrolet, “3” Volkswagen, “4” Hyundai. Notice here that the numbers are not interchangeable. A ranking of “1” is “more” than a ranking of “2,” and so on. The “more” refers to the order of preference. However, ordinal scales fail to provide information about the relative strength of rankings. In this hypothetical example, we do not know whether college professors strongly prefer Cadillacs over Chevrolets or just marginally so. An interval scale provides information about ranking, but also supplies a metric for gauging the differences between rankings. To construct an interval scale, we might ask our college professors to rate on a scale from 1 to 100 how much they would like to own the four cars previously listed. Suppose the average ratings work out as follows: Cadillac, 90; Chevrolet, 70; Volkswagen, 60; Hyundai, 50. From this information we could infer that the preference for a Cadillac is much stronger than for a Chevrolet, which, in turn, is mildly stronger than the preference for a Volkswagen. More
important, we can also make the assumption that the intervals between the points on this scale are approximately the same: The difference between professors’ preference for a Chevrolet and Volkswagen (10 points) is about the same as that between a Volkswagen and a Hyundai (also 10 points). In short, interval scales are based on the assumption of equal-sized units or intervals for the underlying scale. A ratio scale has all the characteristics of an interval scale but also possesses a conceptually meaningful zero point in which there is a total absence of the characteristic being measured. The essential characteristics of the four levels of measurement are summarized in Figure 4.6. Ratio scales are rare in psychological measurement. Consider whether there is any meaningful sense in which a person can be thought to have zero intelligence. Not really. The same is true for most constructs in psychology: Meaningful zero points just do not exist. However, a few physical measures used by psychologists qualify as ratio scales. For example, height and weight qualify, and perhaps
some physiological measures such as electrodermal response qualify, too. But by and large the best a psychologist can hope for is interval-level measurement.
FIGURE 4.6 Essential Characteristics of Four Levels of Measurement Levels of measurement are relevant to test construction because the more powerful and useful parametric statistical procedures (e.g., Pearson r, analysis of variance, multiple regression) should be used only for scores derived from measures that meet the criteria of
interval or ratio scales. For scales that are only nominal or ordinal, less-powerful non- parametric statistical procedures (e.g., chi- square, rank order correlation, median tests) must be employed. In practice, most major psychological testing instruments (especially intelligence tests and personality scales) are assumed to employ approximately interval-level measurement even though, strictly speaking, it is very difficult to demonstrate absolute equality of intervals for such instruments (Bausell, 1986). Now that the reader is familiar with levels of measurement, we introduce a representative sample of scaling methods, noting in advance that different scaling methods yield different levels of measurement.
4.9 REPRESENTATIVE SCALING METHODS Expert Rankings Suppose we wanted to measure the depth of coma in patients who had suffered a recent head injury that rendered them unconscious. A depth
of coma scale could be very important in predicting the course of improvement, because it is well known that a lengthy period of unconsciousness offers a poor prognosis for ultimate recovery. In addition, rehabilitation personnel have a practical need to know whether a patient is deeply comatose or in a partially communicative state of twilight consciousness. One approach to scaling the depth of coma would be to rely on the behavioral rankings of experts. For example, we could ask a panel of neurologists to list patient behaviors associated with different levels of consciousness. After the experts had submitted a large list of diagnostic behaviors, the test developers—preferably experts on head injuries—could rank the indicator behaviors along a continuum of consciousness ranging from deep coma to basic orientation. Using precisely this approach, Teasdale and Jennett (1974) produced the Glasgow Coma Scale. Instruments similar to this scale are widely used in hospitals for the
assessment of traumatic brain injury (Figure 4.7).
FIGURE 4.7 Example of the Use of the Glasgow Coma Scale for Recording Depth of Coma Source: Reprinted with permission from Jennett, B., Teasdale, G. M., & Knill-Jones, R. P. (1975). Predicting outcome after head injury. Journal of the Royal College of Physicians of London, 9, 231–237. The Glasgow Coma Scale is scored by observing the patient and assigning the highest level of functioning on each of three subscales.
On each sub-scale, it is assumed that the patient displays all levels of behavior below the rated level. Thus, from a psychometric standpoint, this scale consists of three sub-scales (eyes, verbal response, and motor response) each yielding an ordinal ranking of behavior. In addition to the rankings, it is possible to compute a single overall score that is something more than an ordinal scale, although probably less than true interval-level measurement. If numbers are attached to the rankings (e.g., for eyes open a coding of “none” = 1, “to pain” = 2, and so on), then the numbers for the rated level for each subscale can be added, yielding a maximum possible score of 14. The total score on the Glasgow Coma Scale predicts later recovery with a very high degree of accuracy (Jennett, Teasdale, & Knill-Jones, 1975). We see, then, that quite plain psychological tests derived from the very simplest scaling methods can, nonetheless, provide valid and useful information. Method of Equal-Appearing Intervals
Early in the twentieth century, L. L. Thurstone (1929) proposed a method for constructing interval-level scales from attitude statements. His method of equal-appearing intervals is still used today, marking him as one of the giants of psychometric theory. The actual methodology of constructing equal-appearing intervals is somewhat complex and statistically laden, but the underlying logic is easy to explain (Ghiselli, Campbell, & Zedeck, 1981). We illustrate the method by summarizing the steps involved in constructing a scale of attitudes toward physical exercise. First, a large number of true–false statements reflecting a range of positive and negative attitudes toward physical exercise would be compiled. Two extreme examples might be: • “I feel that physical exercise is generally
boring and tedious.” • “Physical exercise should be a significant
part of everyone’s daily life.” Of course, many items of moderate attitudinal valence would be written as well. The idea at this point-of-scale development is to produce an
excess of items with the expectation that unsuitable items later will be dropped. Next, these attitude statements would be presented to a group of judges (up to a dozen individuals) who would sort each statement into 1 of 11 categories that range from “extremely favorable” to “extremely unfavorable.” Then, the average favorability for each item (−1.0 to +1.0) would be calculated, along with the standard deviation. Items with larger standard deviations would be dropped, because they produce unreliable ratings. Finally, about 20 to 30 items would be chosen to cover the range of the dimension (favorable to unfavorable). The items on the final scale are assumed to meet the criteria of an interval scale. The score for persons who take the attitude scale is the average scale value of those items endorsed as true (or false, in the case of negatively worded items). Ghiselli et al. (1981) note that the preceding scaling method merely produces the attitude scale. Reliability and validity analyses of the
scale are still needed to determine its appropriateness and usefulness. A study by Russo (1994) illustrates a modern application of the Thurstone method. She used a Thurstone scaling approach to evaluate 216 items from three prominent self-report depression inventories. The judges included 527 undergraduates and 37 clinical faculty members at a medical school. The 216 items were randomized and rated with respect to depressive severity from 1 representing no depression to 11 representing extreme depression. She discovered that all three self-report inventories lacked items and response options typical of mild depression. The distribution of the 216 items was bimodal with many items bunched near the bottom (no depression) and many items bunched near the middle (moderate depression). A characteristic finding for one set of items from a prominent depression scale was as follows:
The reader will notice that the original scoring on these items deviates substantially from the depression ratings provided by the panel of students and clinical faculty. It is also evident that the actual scale values are discontinuous, jumping from 1.0 to 3.4 and higher. A similar pattern was observed for many items on all three inventories, leading Russo (1994) to conclude: The present results suggest that if the original
scoring is used for the three scales examined here, then the distinctions between well-being and absence of depression as well as between moderate
Rated Depressi
on Original Scoring Item Content
1.0 1 I never feel downhearted or sad.
3.4 2 I sometimes feel downhearted or sad.
4.1 3 I feel downhearted or sad a good part of the time.
4.4 4 I feel downhearted or sad most of the time.
and severe will be difficult to make. Such imprecision will make it difficult to assess the efficacy of treatments for depression, because a lack thereof must be a function of added measurement error due to ordinal measures. Such error could also wreak havoc in longitudinal studies, especially in those in which memory is involved.
We see in this example that Thurstone’s approach to item scaling has powerful applications in test development. Based on these findings, researchers are now in a position to develop improved self-report scales that assess the full range of symptomatology in depression. Method of Absolute Scaling Thurstone (1925) also developed the method of absolute scaling, a procedure for obtaining a measure of absolute item difficulty based on results for different age groups of test takers. The methodology for determining individual item difficulty on an absolute scale is quite complex, although the underlying rationale is not too difficult to understand. Essentially, a set
of common test items is administered to two or more age groups. The relative difficulty of these items between any two age groups serves as the basis for making a series of interlocking comparisons for all items and all age groups. One age group serves as the anchor group. Item difficulty is measured in common units such as standard deviation units of ability for the anchor group. The method of absolute scaling is widely used in group achievement and aptitude testing (Donlon, 1984). Thurstone (1925) illustrated the method of absolute scaling with data from the testing of 3,000 schoolchildren on the 65 questions from the original Binet test. Using the mean of Binet test intelligence of 3½-year-old children as the zero point and the standard deviation of their intelligence as the unit of measurement, he constructed a scale that ranged from −2 to +10 and then located each of the 65 questions on that scale. Thurstone (1925) found that the scale “brings out rather strikingly the fact that the questions are unduly bunched at certain ranges [of difficulty] and rather scarce at other ranges.”
A modern test developer would use this kind of analysis as a basis for dropping redundant test items (redundant in the sense that they measure at the same difficulty level) and adding other items that test the higher (and lower) ranges of difficulty. Likert Scales Likert (1932) proposed a simple and straightforward method for scaling attitudes that is widely used today. A Likert scale presents the examinee with five responses ordered on an agree/disagree or approve/disapprove continuum. For example, one item on a scale to assess attitudes toward church membership might read: Church services give me inspiration and help me to live up to my best during the following week. Do you:
|| || || || || Strongly
Agree Agr ee
Undeci ded
Disagr ee
Strongly Disagree
Depending on the wording of an individual item, an extreme answer of “strongly agree” or “strongly disagree” will indicate the most favorable response on the underlying attitude measured by the questionnaire. Likert (1932) assigned a score of 5 to this extreme response, 1 to the opposite extreme, and 2, 3, and 4 to intermediate replies. The total scale score is obtained by adding the scores from individual items. For this reason, a Likert scale is also referred to as a summative scale. Guttman Scales On a Guttman scale, respondents who endorse one statement also agree with milder statements pertinent to the same underlying continuum (Guttman, 1947). Thus, if the examiner knows an examinee’s most extreme endorsement on the continuum, it is possible to reconstruct the intermediate responses as well. Guttman scales are produced by selecting items that fall into an ordered sequence of examinee endorsement. A perfect Guttman scale is seldom achieved because of errors of measurement, but is
nonetheless a fitting goal for certain types of tests. Although the Guttman approach was originally devised to determine whether a set of attitude statements is unidimensional, the technique has been used in many different kinds of tests. For example, Beck used Guttman-type scaling to produce the individual items of the Beck Depression Inventory (BDI, Beck, Steer, & Garbin, 1988). Items from the BDI resemble the following: • ( ) I occasionally feel sad or blue. • ( ) I often feel sad or blue. • ( ) I feel sad or blue most of the time. • ( ) I always feel sad and I can’t stand it.
Clients are asked to “check the statement from each group that you feel is most true about you.” A client who endorses an extreme alternative (e.g., “I always feel sad and I can’t stand it”) almost certainly agrees with the milder statements as well. Method of Empirical Keying
The reader may have noticed that most of the scaling methods discussed in the preceding section rely upon the authoritative judgment of experts in the selection and ordering of items. It is also possible to construct measurement scales based entirely on empirical considerations devoid of theory or expert judgment. In the method of empirical keying, test items are selected for a scale based entirely on how well they contrast a criterion group from a normative sample. For example, a Depression scale could be derived from a pool of true-false personality inventory questions in the following manner: 1. A carefully selected and homogeneous
group of persons experiencing major depression is gathered to answer the pool of true–false questions.
2. For each item, the endorsement frequency of the depression group is compared to the endorsement frequency of the normative sample.
3. Items which show a large difference in endorsement frequency between the depression and normative samples are
selected for the Depression scale, keyed in the direction favored by depression subjects (true or false, as appropriate).
4. Raw score on the Depression scale is then simply the number of items answered in the keyed direction.
The method of empirical keying can produce some interesting surprises. A common finding is that some items selected for a scale may show no obvious relationship to the construct measured. For example, an item such as “I drink a lot of water” (keyed true) might end up on a Depression scale. The momentary rationale for including this item is simply that it works. Of course, the challenge posed to researchers is to determine why the item works. However, from the practical standpoint of empirical scale construction, theoretical considerations are of secondary importance. We discuss the method of empirical keying further in Topic 8B, Self- Report and Behavioral Assessment of Psychopathology.
Rational Scale Construction (Internal Consistency) The rational approach to scale construction is a popular method for the development of self- report personality inventories. The name rational is somewhat of a misnomer, insofar as certain statistical methods are essential to this approach. Also, the name implies that other approaches are nonrational or irrational, which is untrue. The heart of the method of rational scaling is that all scale items correlate positively with each other and also with the total score for the scale. An alternative and more appropriate name for this approach is internal consistency, which emphasizes what is actually done. Gough and Bradley (1992) explain how the rational approach earned its descriptive title: The idea of rationality enters the scene in that
the central theme or unifying dimension around which the items cluster is one that was conceptually articulated beforehand by the developer of the measure and from which the scoring of each item is
determined in a logical and understandable way.
We will follow their presentation to illustrate the features of the rational approach. Suppose a test developer desires to develop a new self-report scale for leadership potential. Based on a review of relevant literature, the researcher might conclude that leadership potential is characterized by self-confidence, resilience under pressure, high intelligence, persuasiveness, assertiveness, and the ability to sense what others are thinking and feeling (Gough & Bradley, 1992). These notions suggest that the following true–false items might be useful in the assessment of leadership potential: • Most of the time I am pretty confident and
sure of myself (T) • When others disagree with me, I usually let
things go. (F) • I know that I am smarter than most people.
(T) • I am not very good at understanding how
others react. (F)
• My friends would describe me as a dominant person. (T)
The T and F after each statement would indicate the rationally keyed direction for leadership potential. Of course, additional items with similar intentions also would be proposed. The test developer might begin with 100 items that appear—on a rational basis—to assess leadership potential. These preliminary items would be administered to a large sample of individuals similar to the target population for whom the scale is intended. For instance, if the scale is designed to identify college students with leadership potential, then it should be administered to a cross-section of several hundred college students. For scale development, very large samples are desirable. In this hypothetical case, let us assume that we obtain results for 500 college students. The next step in rational scale construction is to correlate scores on each of the preliminary items with the total score on the test for the 500 subjects in the tryout sample. Because scores on
the items are dichotomous (1 is arbitrarily assigned to an answer corresponding to the scoring key, 0 to the alternative), a biserial correlation coefficient rbis is needed. Once the correlations are obtained, the researcher scans the list in search of weak correlations and reversals (negative correlations). These items are discarded because they do not contribute to the measurement of leadership potential. Up to half of the initial items might be discarded. If a large proportion of items is initially discarded, the researcher might recalculate the item-total correlations based upon the reduced item pool to verify the homogeneity of the remaining items. The items that survive this iterative procedure constitute the leadership potential scale. The reader should keep in mind that the rational approach to scale construction merely produces a homogeneous scale thought to measure a specified construct. Additional studies with new subject samples would be needed to determine the reliability and validity of the new scale.
4.10 CONSTRUCTING THE ITEMS Constructing test items is a painful and laborious procedure that taxes the creativity of test developers. The item writer is confronted with a profusion of initial questions: • Should item content be homogeneous or
varied? • What range of difficulty should the items
cover? • How many initial items should be
constructed? • Which cognitive processes and item
domains should be tapped? • What kind of test item should be used?
We will address the first three questions briefly before turning to a more detailed discussion of the last two topics, which are commonly referred to under the rubrics of table of specifications and item formats. Initial Questions in Test Construction The first question pertains to the homogeneity versus heterogeneity of test item content. In
large measure, whether item content is homogeneous or varied is dictated by the manner in which the test developer has defined the new instrument. Consider a culture-reduced test of general intelligence. Such an instrument might incorporate varied items, so long as the questions do not presume specific schooling. The test developer might seek to incorporate novel problems equally unfamiliar to all examinees. On the other hand, with a theory- based test of spatial thinking, subscales with homogeneous item content would be required. The range of item difficulty must be sufficient to allow for meaningful differentiation of examinees at both extremes. The most useful tests, then, are those that include a graded series of very easy items passed by almost everyone as well as a group of incrementally more difficult items passed by virtually no one. A ceiling effect is observed when significant numbers of examinees obtain perfect or near-perfect scores. The problem with a ceiling effect is that distinctions between high-scoring examinees are not possible, even though these examinees
might differ substantially on the underlying trait measured by the test. A floor effect is observed when significant numbers of examinees obtain scores at or near the bottom of the scale. For example, the WAIS-R possessed a serious floor effect in that it failed to discriminate between moderate, severe, and profound levels of mental retardation—all persons with significant developmental disabilities would fail to answer virtually every question. Test developers expect that some initial items will prove to make ineffectual contributions to the overall measurement goal of their instrument. For this reason, it is common practice to construct a first draft that contains excess items, perhaps double the number of questions desired on the final draft. For example, the 550-item MMPI originally consisted of more than 1,000 true–false personality statements (Hathaway & McKinley, 1940). Table of Specifications
Professional developers of achievement and ability tests often use one or more item-writing schemes to help ensure that their instrument taps a desired mixture of cognitive processes and content domains. For example, a very simple item-writing scheme might designate that an achievement test on the Civil War should consist of 10 multiple-choice items and 10 fill-in-the- blank questions, half of each on factual matters (e.g., dates, major battles) and the other half on conceptual issues (e.g., differing views on slavery). Before development of a test begins, item writers usually receive a table of specifications. A table of specifications enumerates the information and cognitive tasks on which examinees are to be assessed. Perhaps the most common specification table is the content-by- process matrix, which lists the exact number of items in relevant content areas and details the precise composite of items that must exemplify different cognitive processes (Millman & Greene, 1989).
Consider a science achievement test suitable for high school students. Such a test must cover many different content areas and should require a mixture of cognitive processes ranging from simple recall to inferential reasoning. By providing a table of specifications prior to the item-writing stage, the test developer can guarantee that the resulting instrument contains a proper balance of topical coverage and taps a desired range of cognitive skills. A hypothetical but realistic table of specifications is portrayed in Table 4.3. TABLE 4.3 Example of a Content-by-Process Table of Specifications for a Hypothetical 100-Item Science Achievement Test
Conten t Area
Process Factual Knowledge a
Information Competenceb
Inferential Reasoningc
Astron omy
8 3 3
Botany 6 7 2 Chemis try
10 5 4
aFactual Knowledge: Items can be answered based on simple recognition of basic facts. bInformation Competence: Items require usage of information provided in written text. cInferential Reasoning: Items can be answered by making deductions or drawing conclusions. Item Formats When it comes to the method by which psychological attributes are to be assessed, the test developer is confronted with dozens of choices. Indeed, it would be easy to write an entire chapter on this topic alone. For reviews of item formats, the interested reader should consult Bausell (1986), Jensen (1980), and Wesman (1971). In this section, we will quickly survey the advantages and pitfalls of the more common varieties of test items.
Geolog y
10 5 2
Physics 8 5 6 Zoolog y
8 5 3
Totals 50 30 20
For group-administered tests of intellect or achievement, the technique of choice is the multiple-choice question. For example, an item on an American history achievement test might include this combination of stem and options: The president of the United States during the Civil War was • Washington • Lincoln • Hamilton • Wilson
Proponents of multiple-choice methodology argue that properly constructed items can measure conceptual as well as factual knowledge. Multiple-choice tests also permit quick and objective machine scoring. Furthermore, the fairness of multiple-choice questions can be proved (or occasionally disproved!) with very simple item analysis procedures discussed subsequently. The major shortcomings of multiple-choice questions are, first, the difficulty of writing good distractor options and, second, the possibility that the presence of the response may cue a half-
knowledgeable respondent to the correct answer. Guidelines for writing good multiple-choice items are listed in Table 4.4. Matching questions are popular in classroom testing, but suffer serious psychometric shortcomings. An example of a matching question: Using the letters on the left, match the name to the accomplishment:
TABLE 4.4 Guidelines for Writing Multiple- Choice Items
A. Binet ____ translated a major intelligence test
B. Woodwor th
____ no correlation between grades and mental tests
C. Cattell ____ developed true/false personality inventory
Choose words that have precise meanings. Avoid complex or awkward word arrangements. Include all information needed for response selection. Put as much of the question as possible in the stem. Do not take stems verbatim from textbooks. Use options of equal length and parallel phrasing. Use “none of the above” and “all of the above” rarely. Minimize the use of negatives such as not. Avoid the use of nonfunctional words. Avoid unessential specificity in the stem. Avoid unnecessary clues to the correct response. Submit items to others for editorial scrutiny.
The most serious problem with matching questions is that responses are not independent —missing one match usually compels the examinee to miss another. Another problem is that the options in a matching question must be very closely related or the question will be too easy. For individually administered tests, the procedure of choice is the short-answer objective item. Indeed, the simplest and most straightforward types of questions often possess the best reliability and validity. A case in point is the Vocabulary subtest from the WAIS-IV, which consists merely of asking the examinee to define words. This subtest has very high reliability (.96) and is often considered the single best measure of overall intelligence on the test.
D. McKinley
____ battery of sensorimotor tests
E. Wissler ____ developed first useful intelligence test
F. Goddard
____ screening test for emotional disturbance
Personality tests often use true–false questions because they are easy for subjects to understand. Most people find it simple to answer true or false to items such as:
Critics of this approach have pointed out that answers to such questions may reflect social desirability rather than personality traits (Edwards, 1961). An alternative format designed to counteract this problem is the forced-choice methodology in which the examinee must choose between two equally desirable (or undesirable) options: Which would you rather do: _____ Mop a gallon of syrup from the floor. _____ Volunteer for a half day at a nursing home. Although the forced-choice approach has many desirable psychometric properties, personality test developers have not rushed to embrace this interesting methodology.
T F ____ ____ I like sports magazines.
4.11 TESTING THE ITEMS Psychometricians expect that numerous test items from the original tryout pool will be discarded or revised as test development proceeds. For this reason, test developers initially produce many, many excess items, perhaps double the number of items they intend to use. So, how is the final sample of test questions selected from the initial item pool? Test developers use item analysis, a family of statistical procedures, to identify the best items. In general, the purpose of item analysis is to determine which items should be retained, which revised, and which thrown out. In conducting a thorough item analysis, the test developer might make use of item-difficulty index, item-reliability index, item-validity index, item-characteristic curve, and an index of item discrimination. We turn now to a brief review of these statistical approaches to item analysis. Readers who wish an in-depth discussion and critique of these topics should consult Hambleton (1989) and Nunnally (1978).
Item-Difficulty Index The item difficulty for a single test item is defined as the proportion of examinees in a large tryout sample who get that item correct. For any individual item i, the index of item difficulty is pi, which varies from 0.0 to 1.0. An item with difficulty of .2 is more difficult than an item with difficulty of .7, because fewer examinees answered it correctly. The item-difficulty index is a useful tool for identifying items that should be altered or discarded. Suppose an item has a difficulty index near 0.0, meaning that nearly everyone has answered it incorrectly. Unfortunately, this item is psychometrically unproductive because it does not provide information about differences between examinees. For most applications, the item should be rewritten or thrown out. The same can be said for an item with a difficulty index near 1.0, where virtually all subjects provide a correct answer. What is the optimal level of item difficulty? Generally, item difficulties that hover around .5,
ranging between .3 and .7, maximize the information the test provides about differences between examinees. However, this rule of thumb is subject to one important qualification and one very significant exception. For true–false or multiple-choice items, the optimal level of item difficulty needs to be adjusted for the effects of guessing. For a true– false test, a difficulty level of .5 can result when examinees merely guess. Thus, the optimal item difficulty for such items would be .75 (halfway between .5 and 1.0). In general, the optimal level of item difficulty can be computed from the formula (1.0 + g)/2, where g is the chance success level. Thus, for a four-option multiple- choice item, the chance success level is .25, and the optimal level of item difficulty would be (1.0 + .25)/2, or about .63. If a test is to be used for selection of an extreme group by means of a cutting score, it may be desirable to select items with difficulty levels outside the .3 to .7 range. For example, a test used to select graduate students for a university that admits only a select few of its many
applicants should contain many very difficult items. A test used to designate children for a remedial-education program should contain many extremely easy items. In both cases, there will be useful discrimination among examinees near the cutting score—a very high score for the graduate admissions and a very low score for students eligible for remediation—but little discrimination among the remaining examinees (Allen & Yen, 1979). Item-Reliability Index A test developer may desire an instrument with a high level of internal consistency in which the items are reasonably homogeneous. A simple way to determine whether an individual item “hangs together” with the remaining test items is to correlate scores on that item with scores on the total test. However, individual items are typically right or wrong (often scored 1 or 0), whereas total scores constitute a continuous variable. In order to correlate these two different kinds of scores it is necessary to use a special type of statistic called the point-biserial
correlation coefficient. The computational formula for this correlation coefficient is equivalent to the Pearson r discussed earlier, and the point-biserial coefficient conveys much the same kind of information regarding the relationship between two variables (one of which happens to be dichotomous and scored 0 or 1). In general, the higher the point-biserial correlation riT between an individual item and the total score, the more useful is the item from the standpoint of internal consistency. The usefulness of an individual dichotomous test item is also determined by the extent to which scores on it are distributed between the two outcomes of 0 and 1. Although it sounds incongruous, it is possible to compute the standard deviation for dichotomous items; as with a continuously scored variable, the standard deviation of a dichotomous item indicates the extent of dispersion of the scores. If an individual item has a standard deviation of zero, everyone is obtaining the same score (all right or all wrong). The more closely the item approaches a 50–50 split of right and wrong
scores, the greater is its standard deviation. In general, the greater the standard deviation of an item, the more useful is the item to the overall scale. Although we will not provide the derivation, it can be shown that the item-score standard deviation si for a dichotomously scored item can be computed from
We may summarize the discussion up to this point as follows: The potential value of a dichotomously scored test item depends jointly on its internal consistency as indexed by the correlation with the total score (riT) and also its variability as indexed by the standard deviation (si). If we compute the product of these two indices, we obtain siriT, which is the item- reliability index. Consider the characteristics of an item that possesses a relatively large item- reliability index. Such an item must exhibit strong internal consistency and produce a good dispersion of scores between its two alternatives. The value of this index in test construction is simply this: By computing the
item-reliability index for every item in the preliminary test, we can eliminate the “outlier” items that have the lowest value on this index. Such items would possess poor internal consistency or weak dispersion of scores and therefore not contribute to the goals of measurement. Item-Validity Index For many applications, it is important that a test possess the highest possible concurrent or predictive validity. In these cases, one overriding question governs test construction: How well does each preliminary test item contribute to accurate prediction of the criterion? The item-validity index is a useful tool in the psychometrician’s quest to identify predictively useful test items. By computing the item-validity index for every item in the preliminary test, the test developer can identify ineffectual items, eliminate or rewrite them, and produce a revised instrument with greater practical utility.
The first step in figuring an item-validity index is to compute the point-biserial correlation between the item score and the score on the criterion variable. In general, the higher the point-biserial correlation riC between scores on an individual item and the criterion score, the more useful is the item from the standpoint of predictive validity. As previously noted, the utility of an item also depends upon its standard deviation si. Thus, the item-validity index consists of the product of the standard deviation and the point-biserial correlation: siriC. Item-Characteristic Curves Also known as an item response function, an item-characteristic curve (ICC) is a graphical display of the relationship between the probability of a correct response and the examinee’s position on the underlying trait measured by the test. However, we do not have direct access to underlying traits, so observed test scores must be used to estimate trait quantities.
A separate ICC is graphed for each item, based upon a plot of the total test scores on the horizontal axis versus the proportion of examinees passing the item on the vertical axis (Figure 4.8). An ICC is actually a mathematical idealization of the relationship between the probability of a correct response and the amount of the trait possessed by test respondents. Different ICC models use different mathematical functions based on initial assumptions. The simplest ICC model is the Rasch Model, based upon the item-response theory of the Danish mathematician Georg Rasch (1966). The Rasch Model is the simplest model because it makes just two assumptions: (1) test items are unidimensional and measure one common trait, and (2) test items vary on a continuum of difficulty level. In general, a good item has a positive ICC slope. If the ability to solve a particular item is normally distributed, the ICC will resemble a normal ogive (curve a in Figure 4.8). The normal ogive is simply the normal distribution graphed in cumulative form.
The desired shape of the ICC depends on the purpose of the test. Psychometric purists would prefer that test item ICCs approximate the normal ogive, because this curve is convenient for making mathematical deductions about the underlying trait (Lord & Novick, 1968). However, for selection decisions based on cutoff scores, a step function is preferred. For example, when combined with other similar items, the item that produced curve b in Figure 4.8 would be the best for selecting examinees with high levels of the measured trait.
FIGURE 4.8 Some Sample Item- Characteristic Curves ICCs are especially useful for identifying items that perform differently for subgroups of examinees (Allen & Yen, 1979). For example, a test developer may discover that an item performs differently for men and women. A sex- biased question involving football facts comes to mind here. For men, the ICC for this item might have the desired positive slope, whereas for women the ICC might be quite flat (such as curve c in Figure 4.8). Items with ICCs that differ among subgroups of examinees can be revised or eliminated. The underlying theory of ICC is also known as item response theory and latent trait theory. The usefulness of this approach has been questioned by Nunnally (1978), who points out that the assumption of test unidimensionality (implied in the ICC curve, which plots percentage passing against the unidimensional horizontal axis of trait value) is violated when many psychological tests are considered. If there were no serious technical and practical problems involved, “one
wonders why ICC theory was not adopted long ago for the actual construction and scoring of tests” (Nunnally, 1978). The merits of the ICC approach are still debated. ICC theory seems particularly appropriate for certain forms of computerized adaptive testing (CAT) in which each test taker responds to an individualized and unique set of items that are then scored on an underlying uniform scale (Weiss, 1983). The CAT approach to assessment would not be possible in the absence of an ICC approach to measurement. CAT is discussed in Topic 12B, Computerized Assessment and the Future of Testing. Readers who wish a more detailed discussion of ICC and other latent trait models should consult Hambleton (1989) and Embretson and Reise (2000). Item-Discrimination Index It should be clear from the discussion of ICCs that an effective test item is one that discriminates between high scorers and low scorers on the entire test. An ideal test item is
one that most of the high scorers pass and most of the low scorers fail (see curve a in Figure 4.8). Simple visual inspection of the ICC provides a coarse basis for gauging the discriminability of a test item: If the slope of the curve is positive and the curve is preferably ogive-shaped, the item is doing a good job of separating high and low scorers. But visual inspection is not a completely objective procedure; what is needed is a statistical tool that summarizes the discrimination power of individual test items. An item-discrimination index is a statistical index of how efficiently an item discriminates between persons who obtain high and low scores on the entire test. There are many indices of item discrimination, including such indirect measures as riT, the point-biserial correlation between scores on an individual item and the total test score. However, we will restrict our discussion here to a direct measure, the item- discrimination index, symbolized by the lowercase, italicized letter d. On an item-by- item basis, this index compares the performance
of subjects in the upper and lower regions of total test score. The upper and lower ranges are generally defined as the upper- and lower- scoring 10 percent to 33 percent of the sample. If the total test scores are normally distributed, the optimal comparison is the highest-scoring 27 percent versus the lowest-scoring 27 percent of the examinees. If the distribution of total test scores is flatter than the normal curve, the optimal percentage is larger, approaching 33 percent. For most applications, any percentage between 25 and 33 will yield similar estimates of d (Allen & Yen, 1979). The item-discrimination index for a test item is calculated from the formula: d = (U − L)/N where U is the number of examinees in the upper range who answered the item correctly, L is the number of examinees in the lower range who answered the item correctly, and N is the total number of examinees in the upper or lower range. Let us illustrate the computation and use of d with a hypothetical example. Suppose that a test
developer has constructed the preliminary version of a multiple-choice achievement test and has administered the exam to a tryout sample of 400 high school students. After computing total scores for each subject, the test developer then identifies the high-scoring 25 percent and low-scoring 25 percent of the sample. Since there are 100 students in each group (25 percent of 400), N in the preceding formula will be 100. Next, for each item, the developer determines the number of students in the upper range and the lower range who answered it correctly. To compute d for each item is a simple matter of plugging these values into the formula (U − L)/N. For example, suppose on the first item that 49 students in the upper range answered it correctly, whereas 23 students in the lower range answered it correctly. For this item, d is equal to (49 − 23)/ 100 or .26. It is evident from the formula for d that this index can vary from −1.0 to +1.0. Notice, too, that a negative value for d is a warning signal that a test item needs revision or replacement.
After all, such an outcome indicates that more of the low-scoring subjects answered the item correctly than did the high-scoring subjects. If d is zero, exactly equal numbers of low- and high- scoring subjects answered the item correctly; since the item is not discriminating between low- and high-scoring subjects at all, it should be revised or eliminated. A positive value for d is preferred, and the closer to +1.0 the better. Table 4.5 illustrates item-discrimination indices for six items from the hypothetical test proposed here. A test developer can supplement the item- discrimination approach by inspecting the number of examinees in the upper- and lower- scoring groups who choose each of the incorrect alternatives. If a multiple-choice item is well written, the incorrect alternatives should be equally attractive to subjects who do not know the correct answer. Of course, we expect that high-scoring examinees will choose the correct alternative more often than low-scoring examinees—that is the purpose in computing item-discrimination indices. But, in addition, a
good item should show proportional dispersion of incorrect choices for both high- and low- scoring subjects. Assume that we investigate the choices of 100 high-scoring and 100 low-scoring subjects on a hypothetical multiple-choice test. Correct choices are indicated by an asterisk (*). Item 1 demonstrates the desired pattern of answers, with incorrect choices about equally dispersed.
On item 2, we notice that no examinees picked alternative d. This alternative should be replaced with a more appealing distractor:
Item 3 is probably a poor item in spite of the fact that it discriminates effectively between high- and low-scoring subjects. The obvious problem is that high-scoring examinees prefer alternative a to the correct alternative, d:
Alternatives Item 1 a b c* d e High Scorers 5 6 80 5 4 Low Scorers 15 14 40 16 15
Item 2 a b* c d e High Scorers 5 75 10 0 10 Low Scorers 21 34 20 0 25
TABLE 4.5 Item-Discrimination Indices for Six Hypothetical Items
Perhaps by rewriting alternative a, this item could be rescued. In any case, the main point here is that test developers should pry into every corner of every test item by every means possible, including visual inspection of the pattern of answers.
Item 3 a b c* d e High Scorers 43 6 5 37 9 Low Scorers 20 19 22 10 29
Ite m
U L (U − L)/N
Comment
1 49 23 0.26 Very good item with high difficulty
2 79 19 0.60 Excellent item but rarely achieved
3 52 52 0.00 Poor item that should be revised
4 100 0 1.00 Ideal item but never achieved 5 20 80 -0.60 Terrible item that should be
eliminated 6 0 100 -1.00 Theoretically worst possible
item
Reprise: The Best Items From all the methods of item analysis previously portrayed, which ones should the test developer use to identify the best items for a test? The answer to this question is neither simple nor straightforward. After all, the choice of “best” items depends on the objectives of the test developer. For example, a theoretically inclined research psychologist might desire a measurement instrument with the highest possible internal consistency; item-reliability indices are crucial to this goal. A practically minded college administrator might wish for an instrument with the highest possible criterion validity; item-validity indices would be useful for this purpose. A remediation-oriented mental retardation specialist might desire an intelligence test with minimal floor effect; item- difficulty indices would be helpful in this regard. In sum, there is no single preferred method for item selection ideally suited to every context of assessment and test development.
4.12 REVISING THE TEST The purpose of item analysis, discussed previously, is to identify unproductive items in the preliminary test so that they can be revised, eliminated, or replaced. Very few tests emerge from this process unscathed. It is common in the evolutionary process of test development that many items are dropped, others refined, and new items added. The initial repercussion is that a new and slightly different test emerges. This revised test likely contains more discriminating items with higher reliability and greater predictive accuracy—but these improvements are known to be true only for the first tryout sample. The next step in test development is to collect new data from a second tryout sample. Of course, these examinees should be similar to those for whom the test is ultimately intended. The purpose of collecting additional test data is to repeat the item analysis procedures anew. If further changes are of the minor fine-tuning variety, the test developer may decide the test is
satisfactory and ready for cross-validational study, discussed in the following section. If major changes are needed, it is desirable to collect data from a third and even perhaps a fourth tryout sample. But at some point, psychometric tinkering must end; the developer must propose a finalized instrument and proceed to the next step, cross validation. Cross Validation When a tryout sample is used to ascertain that a test possesses criterion-related validity, the evidence is quite preliminary and tentative. It is prudent practice in test development to seek fresh and independent confirmation of test validity before proceeding to publication. The term cross validation refers to the practice of using the original regression equation in a new sample to determine whether the test predicts the criterion as well as it did in the original sample. Ghiselli, Campbell, and Zedeck (1981) outline the rationale for cross validation: Whether items are chosen on the basis of
empirical keying or whether they are
corrected or weighted, the obtained results should, unless additional data are collected, be viewed as specific to the sample used for the statistical analyses. This is necessary because the obtained results have likely capitalized on chance factors operating in that group and therefore are applicable only to the sample studied.
Validity Shrinkage A common discovery in cross-validation research is that a test predicts the relevant criterion less accurately with the new sample of examinees than with the original tryout sample. The term validity shrinkage is applied to this phenomenon. For example, a biographically based predictor of sales potential might perform quite well for the sample of subjects used to develop the instrument but demonstrate less validity when applied to a new group of examinees. Mitchell and Klimoski (1986) studied validity shrinkage of an instrument designed to foretell which students will succeed in real estate, as measured by the real-world
criterion of obtaining a real estate license two years later. In one analysis based on the sample used to derive the test, the biographically based predictor test correlated .6 with the criterion. But when this same test was tried out on a new sample of real estate students, the correlation with the criterion was lower, about .4, demonstrating typical validity shrinkage. Validity shrinkage is an inevitable part of test development and underscores the need for cross validation. In most cases, shrinkage is slight and the instrument withstands the challenge of cross validation. However, shrinkage of test validity can be a major problem when derivation and cross-validation samples are small, the number of potential test items is large, and items are chosen on a purely empirical basis without theoretical rationale. A classic paper by Cureton (1950) demonstrates a worst-case scenario: using a very small sample to select empirically keyed items from a large item pool, then validating the test on the same sample. The criterion in his study was grade point average, artificially dichotomized into
grades of B or better and grades below B. His “test” items consisted of 85 tags, numbered on one side. For each of 29 students, the tags were shaken in a container and dropped on the table. All tags that fell with numbers up were recorded as indicating the presence of that “item” for the student. Next, Cureton conducted an item analysis, using the dichotomized grades as the criterion. Based on this analysis, 24 items were found to be maximally predictive of students’ grades. Nine items occurred more often among students with the higher grades, and these items were weighted +1. Fifteen items occurred more often among students with the lower grades, and these items were weighted −1. The score on this test (facetiously named the “B-Projective Psychokinesis Test”) consisted of the sum of these 24 item weights. In spite of the nonsensical nature of his test, Cureton (1950) found that test scores correlated .82 with grades. Of course, the strength of this correlation was due entirely to capitalization upon chance. If we were to conduct a series of cross-validation studies using new samples of
students, the correlation between the B- Projective Psychokinesis Test and grades would likely hover right around zero, because this test is completely devoid of predictive validity. There is an important lesson here that applies to serious tests as well: Demonstrate validity through cross validation, do not assume it based merely on the solemn intentions of a new instrument. Feedback from Examinees In test revision, feedback from examinees is a potentially valuable source of information that is normally overlooked by test developers. We can illustrate this approach with research by Nevo (1992). He developed the Examinee Feedback Questionnaire (EFeQ) to study the Inter- University Psychometric Entrance Examination, a major requirement for admission to the six universities in Israel. The Inter-University entrance exam is a group test consisting of five multiple-choice subtests: General Knowledge, Figural Reasoning, Comprehension, Mathematical Reasoning, and English. The
EFeQ was designed as an anonymous posttest administered immediately after the Inter- University entrance exam. The EFeQ is a short and simple questionnaire designed to elicit candid opinions from examinees as to these features of the test– examiner–respondent matrix: • Behavior of examiners • Testing conditions • Clarity of exam instructions • Convenience in using the answer sheet • Perceived suitability of the test • Perceived cultural fairness of the test • Perceived sufficiency of time • Perceived difficulty of the test • Emotional response to the test • Level of guessing • Cheating by the examinee or others
The final question on the EFeQ is an open- ended essay: “We are interested in any remarks or suggestions you might have for improving the exam.” Nevo (1992) determined that the EFeQ questionnaire possesses modest reliability, with
a test–retest reliability of about .70. Regardless of the psychometric properties of his scale, the tradition of asking examinees for feedback about tests has proved invaluable. The Inter- University entrance exam was modified in numerous ways in response to feedback: The answer sheet format was modified in ways suggested by examinees; the time limit was increased for specific tests reported to be too speeded; certain items perceived as culturally biased or unfair were deleted. In addition, security measures were revised and tightened in order to minimize cheating, which was much more prevalent than examiners had anticipated. Nevo (1992) also cites a hidden advantage to feedback questionnaires: They convey the message that someone cares enough to listen, which reduces postexamination stress. Examinee feedback questionnaires should become a routine practice in group standardized testing.
4.13 PUBLISHING THE TEST
The test construction process does not end with the collection of cross-validation data. The test developer also must oversee the production of the testing materials, publish a technical manual, and produce a user’s manual. A number of relevant guidelines can be offered for each of these final steps, as outlined in the following sections. Finally, we close this chapter with a provocative comment on the conservatism of modern test publishers. Production of Testing Materials Testing materials must be user friendly if they are to receive wide acceptance by psychologists and educators. Thus, a first guideline for test production is that the physical packaging of test materials must allow for quick and smooth administration. Consider the challenge posed by some performance tests, in which the examiner must wrestle with pencil, clipboard, test form, stopwatch, test manual, item shield, item box, and a disassembled cardboard object, all the while maintaining conversation with the examinee. If it is possible for the test developer
to simplify the duties of the examiner while leaving examinee task demands unchanged, the resulting instrument will have much greater acceptability to potential users. For example, if the administration instructions can be summarized on the test form, the examiner can put the test manual aside while setting out the task for the examinee. Another welcome addition to psychological test packaging is the stand-up ring binder that shows the test question on the side facing the examinee and provides instructions for administration on the reverse side facing the examiner. Technical Manual and User’s Manual Technical data about a new instrument are usually summarized with appropriate references in a technical manual. Here, the prospective user can find information about item analyses, scale reliabilities, cross-validation studies, and the like. In some cases, this information is incorporated in the user’s manual, which gives instructions for administration and also provides guidelines for test interpretation.
Test manuals should communicate information to many different groups ranging in background and training from measurement specialist to classroom teacher. Test manuals serve many purposes, as outlined in the Standards for Educational and Psychological Testing (AERA, APA, & NCME, 1985, 1999). The influential Standards manual suggests that test manuals accomplish the following goals: • Describe the rationale and recommended
uses for the test • Provide specific cautions against anticipated
misuses of a test • Cite representative studies regarding general
and specific test uses • Identify special qualifications needed to
administer and interpret the test • Provide revisions, ammendations, and
supplements as needed • Use promotional material that is accurate
and research based • Cite quantitative relationships between test
scores and criteria
• Report on the degree to which alternative modes of response (e.g., booklet versus an answer sheet) are interchangeable
• Provide appropriate interpretive aids to the test taker
• Furnish evidence of the validity of any automated test interpretations
Finally, test manuals should provide the essential data on reliability and validity rather than referring the user to other sources—an unfortunate practice encountered in some test manuals. Testing Is Big Business By now the reader should appreciate the intimidating task faced by anyone who sets out to develop and publish a new test. Aside from the gargantuan proportions of the endeavor, test development is extraordinarily expensive, which means that publishers are inherently conservative about introducing new tests. Jensen (1980) provides the following provocative view on this topic:
To produce a new general intelligence test that would be a really significant improvement over existing instruments would be a multimillion-dollar project requiring a large staff of test construction experts working for several years. Today we possess the necessary psychometric technology for producing considerably better tests than are now in popular use. The principal hindrances are copyright laws, vested interests of test publishers in the established tests in which they have already made enormous investments, and the market economy for tests. Significant improvement of tests is not an attractive commercial venture initially and would probably have to depend on large-scale and long-term subsidies from government agencies and private foundations.