Ethical and Professional Issues in Psychology Testing
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 1/70
CHAPTER 4 Validity and Test Development
TOPIC 4A Basic Concepts of Validity
4.1 Validity: A De�inition (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec1#ch04lev1sec1)
4.2 Content Validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec2#ch04lev1sec2)
4.3 Criterion-Related Validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec3#ch04lev1sec3)
4.4 Construct Validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec4#ch04lev1sec4)
4.5 Approaches to Construct Validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec5#ch04lev1sec5)
4.6 Extravalidity Concerns and the Widening Scope of Test Validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec6#ch04lev1sec6)
As most every student of psychology knows, the merit of a psychological test is determined �irst by its reliability but then ultimately by its validity. In the preceding chapter we pointed out that reliability can be appraised by many seemingly diverse methods ranging from the conceptually straightforward test– retest approach to the theoretically more complex methodologies of internal consistency. Yet, regardless of the method used, the assessment of reliability invariably boils down to a simple summary statistic, the reliability coef�icient. In this chapter, the more dif�icult and complex issue of validity—what a test score means—is investigated. The concept of validity is still evolving and, therefore, stirs up a great deal more controversy than its staid and established cousin, reliability (AERA, APA, & NCME, 1999 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib30) ). In Topic 4A (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04#ch04box1) , Basic Concepts of Validity, we introduce essential concepts of validity, including the standard tripartite division into content, criterion-related, and construct validity. We also discuss extravalidity concerns, which include side effects and unintended consequences of testing. Extravalidity concerns have fostered a wider de�inition of test validity that extends beyond the technical notions of content, criteria, and constructs. In Topic 4B (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec6#ch04box2) , Test Construction, we stress that validity must be built into the test from the outset rather than being limited to the �inal stages of test development.
Put simply, the validity of a test is the extent to which it measures what it claims to measure. Psychometricians have long acknowledged that validity is the most fundamental and important characteristic of a test. After all, validity de�ines the meaning of test scores. Reliability is important, too, but only insofar as it constrains validity. To the extent that a test is unreliable, it cannot be valid. We can express this point from an alternative perspective: Reliability is a necessary but not a suf�icient precursor of validity.
Test developers have a responsibility to demonstrate that new instruments ful�ill the purposes for which they are designed. However, unlike test reliability, test validity is not a simple issue that is easily resolved on the basis of a few rudimentary studies. Test validation is a developmental process that begins with test construction and continues inde�initely:
After a test is released for operational use, the interpretive meaning of its scores may continue to be sharpened, re�ined, and enriched through the gradual accumulation of clinical observations and through special research projects. . . . Test validity is a living thing; it is not dead and embalmed when the test is released. (Anastasi, 1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib47) )
Test validity hinges upon the accumulation of research �indings. In the sections that follow, we examine the kinds of evidence sought in the validation of a psychological test.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 2/70
4.1 VALIDITY: A DEFINITION We begin with a de�inition of validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss341) , paraphrased from the in�luential Standards for Educational and Psychological Testing (AERA, APA, & NCME, 1999 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib30) ):
A test is valid to the extent that inferences made from it are appropriate, meaningful, and useful.
Notice that a test score per se is meaningless until the examiner draws inferences from it based on the test manual or other research �indings. For example, knowing that an examinee has obtained a slightly elevated score on the MMPI-2 Depression scale is not particularly helpful. This result becomes valuable only when the examiner infers behavioral characteristics from it. Based on existing research, the examiner might conclude, “The elevated Depression score suggests that the examinee has little energy and has a pessimistic outlook on life.” The MMPI-2 Depression scale possesses psychometric validity to the extent that such inferences are appropriate, meaningful, and useful.
Unfortunately, it is seldom possible to summarize the validity of a test in terms of a single, tidy statistic. Determining whether inferences are appropriate, meaningful, and useful typically requires numerous studies of the relationships between test performance and other independently observed behaviors. Validity re�lects an evolutionary, research-based judgment of how adequately a test measures the attribute it was designed to measure. Consequently, the validity of tests is not easily captured by neat statistical summaries but is instead characterized on a continuum ranging from weak to acceptable to strong.
Traditionally, the different ways of accumulating validity evidence have been grouped into three categories:
Content validity Criterion-related validity Construct validity
We will expand on this tripartite view of validity shortly, but �irst a few cautions. The use of these convenient labels does not imply that there are distinct types of validity or that a speci�ic validation procedure is best for one test use and not another:
An ideal validation includes several types of evidence, which span all three of the traditional categories. Other things being equal, more sources of evidence are better than fewer. However, the quality of the evidence is of primary importance, and a single line of solid evidence is preferable to numerous lines of evidence of questionable quality. Professional judgment should guide the decisions regarding the forms of evidence that are most necessary and feasible in light of the intended uses of the test and any likely alternatives to testing. (AERA, APA, & NCME, 1985 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib29) )
We may summarize these points by stressing that validity is a unitary concept determined by the extent to which a test measures what it purports to measure. The inferences drawn from a valid test are appropriate, meaningful, and useful. In this light, it should be apparent that virtually any empirical study that relates test scores to other �indings is a potential source of validity information (Anastasi, 1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib47) ; Messick, 1995 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1135) ).
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 3/70
4.2 CONTENT VALIDITY Content validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss67) is determined by the degree to which the questions, tasks, or items on a test are representative of the universe of behavior the test was designed to sample. In theory, content validity is really nothing more than a sampling issue (Bausell, 1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib110) ). The items of a test can be visualized as a sample drawn from a larger population of potential items that de�ine what the researcher really wishes to measure. If the sample (speci�ic items on the test) is representative of the population (all possible items), then the test possesses content validity.
Content validity is a useful concept when a great deal is known about the variable that the researcher wishes to measure. With achievement tests in particular, it is often possible to specify the relevant universe of behaviors in advance. For example, when developing an achievement test of spelling, a researcher could identify nearly all possible words that third graders should know. The content validity of a third-grade spelling achievement test would be assured, in part, if words of varying dif�iculty level were randomly sampled from this preexisting list.
However, test developers must take care to specify the relevant universe of responses as well. All too often, a multiple-choice format is taken for granted:
If the constructor thinks about his aims with an open mind he will often decide that the task should call for a response constructed by the student—written open-end responses or, if inhibitions are to be minimized, oral responses. Nor are the directions to the subject and the social setting of the test to be neglected in de�ining the task. (Cronbach, 1971 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib374) )
In reference to spelling achievement, it cannot be assumed that a multiple-choice test will measure the same spelling skills as an oral test or a frequency count of misspellings in written compositions. Thus, when evaluating content validity, response speci�ication is also an integral part of de�ining the relevant universe of behaviors.
Content validity is more dif�icult to assure when the test measures an ill-de�ined trait. How could a test developer possibly hope to specify the universe of potential items for a measure of anxiety? In these cases in which the measured trait is less tangible, no test developer in his or her right mind would try to construct the literal universe of potential test items. Instead, what usually passes for content validity is the considered opinion of expert judges. In effect, the test developer asserts that “a panel of experts reviewed the domain speci�ication carefully and judged the following test questions to possess content validity.” Figure 4.1 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec2#ch04�ig1) reproduces a sample judge’s item rating form for determining the content validity of test questions.
FIGURE 4.1 Sample Judges Item-Rating Form for Determining Content Validity Source: Based on Martuza (1977 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1055) ), Hambleton (1984 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib687) ), Bausell (1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib110) ).
Quanti�ication of Content Validity Martuza (1977 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1055) ) and others have discussed statistical methods for determining the overall content validity of a test from the judgments of experts. These methods tend to be very specialized and have not been widely accepted. Nonetheless, their approaches can serve as a model for a commonsense viewpoint on interrater agreement as a basis for content validity.
When two expert judges evaluate individual items of a test on the four-point scale proposed in Figure 4.1 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec2#ch04�ig1) , the ratings of each judge on each item can be dichotomized into weak relevance (ratings of 1 or 2) versus strong relevance (ratings of 3 or 4). For each item, then, the conjoint ratings of the two judges can be entered into the two-by- two agreement table depicted in Figure 4.2 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec2#ch04�ig2) . For example, if both judges believed an item was quite relevant (strong relevance), it would be placed in cell D. If the �irst judge believed an item was very relevant (strong relevance) but the second judge deemed it be only slightly relevant (weak relevance), the item would be placed in cell B.
Notice that cell D is the only cell that re�lects valid agreement between judges. The other cells involve disagreement (cells B and C) or agreement that an item doesn’t belong on the test (cell A). We have reproduced hypothetical results for a 100-item test in Figure 4.3 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec2#ch04�ig3) . A coef�icient of content validity can be derived from the following formula:
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 4/70
FIGURE 4.2 Interrater Agreement Model for Content Validity
For example, on our 100-item test both judges concurred that 87 items were strongly relevant (cell D), so the coef�icient of content validity would be 87/(4 + 4 + 5 + 87) or .87. If more than two judges are used, this computational procedure could be completed with all possible pair-wise combinations of judges, and the average coef�icient reported. An important note: A coef�icient of content validity is just one piece of evidence in the evaluation of a test. Such a coef�icient does not by itself establish the validity of a test.
The commonsense approach to content validity advocated here serves well as a �lagging mechanism to help cull out existing items that are deemed inappropriate by expert raters. However, it cannot identify nonexistent items that should be added to a test to help make the pool of questions more representative of the intended domain. A test could possess a robust coef�icient of content validity and still fall short in subtle ways. Quanti�ication of content validity is no substitute for careful selection of items.
Face Validity
FIGURE 4.3 Hypothetical Example of Agreement Model of Content Validity for a 100-Item Test
We digress brie�ly here to mention face validity, which is not really a form of validity at all. Nonetheless, the concept is encountered in testing and, therefore, needs brief explanation. A test has face validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss111) if it looks valid to test users, examiners, and especially the examinees. Face validity is really a matter of social acceptability and not a technical form of validity in the same category as content, criterion-related, or construct validity (Nevo, 1985 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1229) ). From a public relations standpoint, it is crucial that tests possess face validity—otherwise those who take the tests may be dissatis�ied and doubt the value of psychological testing. However, face validity should not be confused with objective validity, which is determined by the relationship of test scores to other sources of information. In fact, a test could possess extremely strong face validity—the items might look highly relevant to what is presumably measured by the instrument—yet produce totally meaningless scores with no predictive utility whatever.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 5/70
4.3 CRITERION-RELATED VALIDITY Criterion-related validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss81) is demonstrated when a test is shown to be effective in estimating an examinee’s performance on some outcome measure. In this context, the variable of primary interest is the outcome measure, called a criterion. The test score is useful only insofar as it provides a basis for accurate prediction of the criterion. For example, a college entrance exam that is reasonably accurate in predicting the subsequent grade point average of examinees would possess criterion-related validity.
Two different approaches to validity evidence are subsumed under the heading of criterion-related validity. In concurrent validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss60) , the criterion measures are obtained at approximately the same time as the test scores. For example, the current psychiatric diagnosis of patients would be an appropriate criterion measure to provide validation evidence for a paper- and-pencil psychodiagnostic test. In predictive validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss251) , the criterion measures are obtained in the future, usually months or years after the test scores are obtained, as with the college grades predicted from an entrance exam. Each of these two approaches is best suited to different testing situations, discussed in the following sections. However, before we review the nature of concurrent and predictive validity, let us examine a more fundamental question: What are the characteristics of a good criterion?
Characteristics of a Good Criterion As noted, a criterion is any outcome measure against which a test is validated. In practical terms, a criterion can be most anything. Some examples will help to illustrate the diversity of potential criteria. A simulator-based driver skill test might be validated against a criterion of “number of traf�ic citations received in the last 12 months.” A scale measuring social readjustment might be validated against a criterion of “number of days spent in a psychiatric hospital in the last three years.” A test of sales potential might be validated against a criterion of “dollar amount of goods sold in the preceding year.” The choice of criteria is circumscribed, in part, by the ingenuity of the test developer. However, criteria must be more than just imaginative; they must also be reliable, appropriate, and free of contamination from the test itself.
The criterion must itself be reliable if it is to be a useful index of what the test measures. If you recall the meaning of reliability—consistency of scores—the need for a reliable criterion measure is intuitively obvious. After all, unreliable means unpredictable. An unreliable criterion will be inherently unpredictable, regardless of the merits of the test.
Consider the case in which scores on a college entrance exam (the test) are used to predict subsequent grade point average (the criterion). The validity of the entrance exam could be studied by computing the correlation (rxy) between entrance exam scores and grade point averages for a representative sample of students. For purposes of a validity study, it would be ideal if the students were granted open or unscreened enrollment so as to prevent a restriction of range on the criterion variable. In any case, the resulting correlation coef�icient is called a validity coef�icient.1 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec3#ch04fn01)
The theoretical upper limit of the validity coef�icient is constrained by the reliability of both the test and the criterion:
The validity coef�icient is always less than or equal to the square root of the test reliability multiplied by the criterion reliability. In other words, to the extent that the reliability of either the test or the criterion (or both) is low, the validity coef�icient is also diminished. Returning to our example of an entrance exam used to predict college grade point average, we must conclude that the validity coef�icient for such a test will always fall far short of +1.00, owing in part to the unreliability of college grades and also in part to the unreliability of the test itself.
A criterion measure must also be appropriate for the test under investigation. The Standards for Educational and Psychological Testing sourcebook (AERA, APA, & NCME, 1985 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib29) ) incorporates this important point as a separate standard:
All criterion measures should be described accurately, and the rationale for choosing them as relevant criteria should be made explicit.
For example, in the case of interest tests, it is sometimes unclear whether the criterion measure should indicate satisfaction, success, or continuance in the activities under question. The choice between these subtle variants in the criterion must be made carefully, based on an analysis of what the interest test purports to measure.
A criterion must also be free of contamination from the test itself. Lehman (1978 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib965) ) has illustrated this point in a criterion-related validity study of a life change measure. The Schedule of Recent Events, or the SRE (Holmes & Rahe, 1967 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib778) ) is a widely used instrument that provides a quantitative index of the accumulation of stressful life events (e.g., divorce, job promotion, traf�ic tickets). Scores on the SRE correlate modestly with such criterion measures as physical illness and psychological disturbance. However, many seemingly appropriate criterion measures incorporate items that are similar or identical to SRE items. For example, screening tests of psychiatric symptoms often check for changes in eating, sleeping, or social activities. Unfortunately, the SRE incorporates questions that check for the following:
Change in eating habits Change in sleeping habits Change in social activities
If the screening test contains the same items as the SRE, then the correlation between these two measures will be arti�icially in�lated. This potential source of error in test validation is referred to as criterion contamination, since the criterion is “contaminated” by its arti�icial commonality with the test.
Criterion contamination is also possible when the criterion consists of ratings from experts. If the experts also possess knowledge of the examinees’ test scores, this information may (consciously or unconsciously) in�luence their ratings. When validating a test against a criterion of expert ratings, the test scores must be held in strictest con�idence until the ratings have been collected.
Now that the reader knows the general characteristics of a good criterion, we will review the application of this knowledge in the analysis of concurrent and predictive validity.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 6/70
Concurrent Validity In a concurrent validation study, test scores and criterion information are obtained simultaneously. Concurrent evidence of test validity is usually desirable for achievement tests, tests used for licensing or certi�ication, and diagnostic clinical tests. An evaluation of concurrent validity indicates the extent to which test scores accurately estimate an individual’s present position on the relevant criterion. For example, an arithmetic achievement test would possess concurrent validity if its scores could be used to predict, with reasonable accuracy, the current standing of students in a mathematics course. A personality inventory would possess concurrent validity if diagnostic classi�ications derived from it roughly matched the opinions of psychiatrists or clinical psychologists.
A test with demonstrated concurrent validity provides a shortcut for obtaining information that might otherwise require the extended investment of professional time. For example, the case assignment procedure in a mental health clinic can be expedited if a test with demonstrated concurrent validity is used for initial screening decisions. In this manner, severely disturbed patients requiring immediate clinical workup and intensive treatment can be quickly identi�ied by paper-and-pencil test. Of course, tests are not intended to replace mental health specialists, but they can save time in the initial phases of diagnosis.
Correlations between a new test and existing tests are often cited as evidence of concurrent validity. This has a catch-22 quality to it—old tests validating a new test—but is nonetheless appropriate if two conditions are met. First, the criterion (existing) tests must have been validated through correlations with appropriate nontest behavioral data. In other words, the network of interlocking relationships must touch ground with real-world behavior at some point. Second, the instrument being validated must measure the same construct as the criterion tests. Thus, it is entirely appropriate that developers of a new intelligence test report correlations between it and established mainstays such as the Stanford-Binet and Wechsler scales.
Predictive Validity In a predictive validation study, test scores are used to estimate outcome measures obtained at a later date. Predictive validity is particularly relevant for entrance examinations and employment tests. Such tests share a common function—determining who is likely to succeed at a future endeavor. A relevant criterion for a college entrance exam would be �irst-year-student grade point average, while an employment test might be validated against supervisor ratings after six months on the job. In the ideal situation, such tests are validated during periods of open enrollment (or open hiring) so that a full range of results is possible on the outcome measures. In this manner, future use of the test as a selection device for excluding low-scoring applicants will rest on a solid foundation of validational data.
When tests are used for purposes of prediction, it is necessary to develop a regression equation. A regression equation (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss273) describes the best-�itting straight line for estimating the criterion from the test. We will not discuss the statistical approach to �itting the straight line, except to mention that it minimizes the sum of the squared deviations from the line (Ghiselli, Campbell, & Zedeck, 1981 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib585) ). For current purposes, it is more important to understand the nature and function of regression equations.
Ghiselli and associates (1981 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib585) ) provide a simple example of regression in the service of prediction, summarized here. Suppose we are trying to predict success on a job Y (evaluated by the supervisor on a 7-point scale ranging from poor to excellent performance) from scores on a preemployment test X (with scores that range from a low of 0 to a high of 100). The regression equation
Y = .07X + .2
might describe the best-�itting straight line and therefore produce the most accurate predictions. For an individual who scored 55 on the test, the predicted performance level would be 4.05; that is, .07(55) + .2. A test score of 33 yields a predicted performance level of 2.51, that is, .07(33) + .2. Additional predictions are made likewise.
Validity Coef�icient and the Standard Error of the Estimate The relationship between test scores and criterion measures can be expressed in several different ways. Perhaps the most popular approach is to compute the correlation between test and criterion (rxy). In this context, the resulting correlation is known as a validity coef�icient (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss342) . The higher the validity coef�icient rxy, the more accurate is the test in predicting the criterion. In the hypothetical case where rxy is 1.00, the test would possess perfect validity and allow for �lawless prediction. Of course, no such test exists, and validity coef�icients are more commonly in the low- to midrange of correlations and rarely exceed .80. But how high should a validity coef�icient be? There is no general answer to this question. However, we can approach the question indirectly by investigating the relationship between the validity coef�icient and the corresponding error of estimate.
The standard error of estimate (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss306) (SEest) is the margin of error to be expected in the predicted criterion score. The error of estimate is derived from the following formula:
In this formula, is the square of the validity coef�icient and SDy is the standard deviation of the criterion scores. Perhaps the reader has noticed the similarities between this index and the standard error of measurement (SEM). In fact, both indices help gauge margins of error. The SEM indicates the margin of measurement error caused by unreliability of the test, whereas SEest indicates the margin of prediction error caused by the imperfect validity of the test.
The SEest helps answer the fundamental question: “How accurately can criterion performance be predicted from test scores?” (AERA, APA, & NCME, 1985 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib29) ). Consider the common practice of attempting to predict college grade point average from high school scores on a scholastic aptitude test. For a speci�ic aptitude test, suppose we determine that the SEest for predicted grade point average is .2 (on the usual 0.0 to 4.0 grade point scale). What does this mean for the examinee whose college grade point is predicted to be 3.1? As is the case with all standard deviations, the standard error of the estimate can be used to bracket predicted outcomes in a probabilistic sense. Assuming that the frequency distribution of grades is normal, we know that the chances are about 68 in 100 that the examinee’s predicted grade point will fall between 2.9 and 3.3 (plus or minus one SEest). In like manner, we know that the chances are about 95 in 100 that the examinee’s predicted grade point will fall between 2.7 and 3.5 (plus or minus two SEest).
What is an acceptable standard of predictive accuracy? There is no simple answer to this question. As the reader will discern from the discussion that follows, standards of predictive accuracy are, in part, value judgments. To explain why this is so, we need to introduce the basic elements of decision theory (Taylor & Russell, 1939; Cronbach & Gleser, 1965).
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 7/70
Decision Theory Applied to Psychological Tests Proponents of decision theory (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss89) stress that the purpose of psychological testing is not measurement per se but measurement in the service of decision making. The personnel manager wishes to know whom to hire; the admissions of�icer must choose whom to admit; the parole board desires to know which felons are good risks for early release; and the psychiatrist needs to determine which patients require hospitalization.
The link between testing and decision making is nowhere more obvious than in the context of predictive validation studies. Many of these studies use test results to determine who will likely succeed or fail on the criterion task so that, in the future, examinees with poor scores on the predictor test can be screened from admission, employment, or other privilege. This is the rationale by which admissions of�icers or employers require applicants to obtain a certain minimum score on an appropriate entrance or employment exam—previous studies of predictive validity can be cited to show that candidates scoring below a certain cutoff face steep odds in their educational or employment pursuits.
Psychological tests frequently play a major role in these kinds of institutional decision making. In a typical institutional decision, a committee—or sometimes a single person—makes a large number of comparable decisions based on a cutoff score on one or more selection tests. In order to present the key concepts of decision theory, let us oversimplify somewhat and assume that only a single test is involved.
Even though most tests produce a range of scores along a continuum, it is usually possible to identify a cutoff or pass/fail score that divides the sample into those predicted to succeed versus those predicted to fail on the criterion of interest. Let us assume that persons predicted to succeed are also selected for hiring or admission. In this case, the proportion of persons in the “predicted-to-succeed” group is referred to as the selection ratio. The selection ratio can vary from 0 to 1.0, depending on the proportion of persons who are considered good bets to succeed on the criterion measure.
If the results of a selection test allow for the simple dichotomy of “predicted to succeed” versus “predicted to fail,” then the subsequent outcome on the criterion measure likewise can be split into two categories, namely, “did succeed” and “did fail.” From this perspective, every study of predictive validity produces a two- by-two matrix, as portrayed in Figure 4.4 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec3#ch04�ig4) .
Certain combinations of predicted and actual outcomes are more likely than others. If a test has good predictive validity, then most persons predicted to succeed will succeed and most persons predicted to fail will fail. These are examples of correct predictions and serve to bolster the validity of a selection instrument. Outcomes in these two cells are referred to as hits because the test has made a correct prediction.
FIGURE 4.4 Possible Outcomes When a Selection Test Is Used to Predict Performance on a Criterion Measure
But no selection test is a perfect predictor, so two other types of outcomes are also possible. Some persons predicted to succeed will, in fact, fail. These cases are referred to as false positives (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss117) . And some persons predicted to fail would, if given the chance, succeed. These cases are referred to as false negatives (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss116) . False positives and false negatives are collectively known as misses, because in both cases the test has made an inaccurate prediction. Finally, the hit rate is the proportion of cases in which the test accurately predicts success or failure, that is, hit rate = (hits)/(hits + misses).
False positives and false negatives are unavoidable in the real-world use of selection tests. The only way to eliminate such selection errors would be to develop a perfect test, an instrument which has a validity coef�icient of +1.00, signifying a perfect correlation with the criterion measure. A perfect test is theoretically possible, but none has yet been observed on this planet. Nonetheless, it is still important to develop selection tests with very high predictive validity, so as to minimize decision errors.
Proponents of decision theory make two fundamental assumptions about the use of selection tests:
The value of various outcomes to the institution can be expressed in terms of a common utility scale. One such scale—but by no means the only one—is pro�it and loss. For example, when using an interest inventory to select salespersons, a corporation can anticipate pro�it from applicants correctly identi�ied as successful but will lose money when, inevitably, some of those selected do not sell enough even to support their own salary (false positives). The cost of the selection procedure must also be factored in to the utility scale as well. In institutional selection decisions, the most generally useful strategy is one that maximizes the average gain on the utility scale (or minimizes average loss) over many similar decisions. For example, which selection ratio produces the largest average gain on the utility scale? Maximization is, thus, the fundamental decision principle.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 8/70
The application of decision theory is much more complicated than illustrated here, mainly because of the dif�iculty of �inding a common utility scale for different outcomes. Consider the plight of the admissions of�icer at any large university. If the selection ratio is quite strict, then most of the admitted students will also succeed. But some students not admitted might have succeeded, too, and their �inancial support to the university (tuition, fees) is, therefore, lost. However, if the selection ratio is too lenient, then the percentage of false positives (students admitted who subsequently fail) skyrockets. How is the cost of a false positive to be calculated? The �inancial cost can be estimated—for example, advisers dedicate a certain number of hours at a known pay rate counseling these students. But no single utility scale can encompass the other diverse consequences such as the need for additional remedial services (which require money), the increase in faculty cynicism (an issue of morale), and the dashed hopes of misled students (whose heartbreak affects public perception of the university and may even in�luence future state funding!). Clearly, the neat statistical notions of decision theory oversimplify the complex in�luences that determine utility in the real world.
Nonetheless, in large institutional settings where a common utility scale can be identi�ied, principles of decision theory can be applied to selection problems with thought-provoking results. For example, Schmidt, Hunter, McKenzie, and Muldrow (1979 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1465) ) analyzed the potential impact of using the Programmer Aptitude Test (PAT, Hughes & McNamara, 1959 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib794) ) in the selection of computer programmers by the federal government. They based their analysis on the following facts and assumptions:
PAT scores and measures of later on-the-job programming performance correlate quite substantially; the validity coef�icient of the PAT is .76 (fact). The government hires 600 new programmers each year (fact). The cost of testing is about $10 per examinee (fact). Programmers stay on the job for about nine years and receive pay raises according to a known pay scale (fact). The yearly productivity in dollars of low-performing, average, and superior programmers can be accurately estimated by supervisors (assumption).
Based on these facts and assumptions, Schmidt et al. (1979 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1465) ) then compared the hypothetical use of the PAT against other selection procedures of lesser validity. Since the usefulness of a test is partly determined by the percentage of applicants who are selected for employment, the researchers also looked at the impact of different selection ratios on overall productivity. In each case, they estimated the yearly increase in dollar-amount productivity from using the PAT instead of an alternative and less ef�icacious procedure. In general, the use of the PAT was estimated to increase productivity by tens of millions of dollars. The speci�ic estimated increase depended on the selection ratio and the validity coef�icient of hypothetical alternative procedures. For example, if 80 percent of the applicants were hired (selection ratio of .80), using the PAT would increase the productivity of the federal government by at least $5.6 million (if the alternative procedure had a validity coef�icient of .50) and possibly as much as $16.5 million (if the alternative procedure had no validity at all). If the selection ratio were quite small, the use of the PAT for selection boosted productivity even more—possibly as much as nearly $100 million. Schmidt et al. (1979 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1465) ) concluded that “the impact of valid selection procedures on work-force productivity is considerably greater than most personnel psychologists have believed.”
1We have purposefully refrained from referring to such a statistic as the validity coef�icient. Remember that validity is a unitary concept determined by multiple sources of information that may include the correlation between test and criterion.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev1… 9/70
4.4 CONSTRUCT VALIDITY The �inal type of validity discussed in this unit is construct validity, and it is undoubtedly the most dif�icult and elusive of the bunch. A construct (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss63) is a theoretical, intangible quality or trait in which individuals differ (Messick, 1995 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1135) ). Examples of constructs include leadership ability, overcontrolled hostility, depression, and intelligence. Notice in each of these examples that constructs are inferred from behavior but are more than the behavior itself. In general, constructs are theorized to have some form of independent existence and to exert broad but to some extent predictable in�luences on human behavior. A test designed to measure a construct must estimate the existence of an inferred, underlying characteristic (e.g., leadership ability) based on a limited sample of behavior. Construct validity refers to the appropriateness of these inferences about the underlying construct.
All psychological constructs possess two characteristics in common:
There is no single external referent suf�icient to validate the existence of the construct; that is, the construct cannot be operationally de�ined (Cronbach & Meehl, 1955 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib376) ). Nonetheless, a network of interlocking suppositions can be derived from existing theory about the construct (AERA, APA, & NCME, 1985 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib29) ).
We will illustrate these points by reference to the construct of psychopathy (Cleckley, 1976 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib313) ), a personality constellation characterized by antisocial behavior (lying, stealing, and occasionally violence), a lack of guilt and shame, and impulsivity.2 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec4#ch04fn02) Psychopathy is surely a construct, in that there is no single behavioral characteristic or outcome suf�icient to determine who is strongly psychopathic and who is not. On average we might expect psychopaths to be frequently incarcerated, but so are many common criminals. Furthermore, many successful psychopaths somehow avoid apprehension altogether (Cleckley, 1976 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib313) ). Psychopathy cannot be gauged only by scrapes with the law.
Nonetheless, a network of interlocking suppositions can be derived from existing theory about psychopathy. The fundamental problem in psychopathy is presumed to be a de�iciency in the ability to feel emotional arousal—whether empathy, guilt, fear of punishment, or anxiety under stress (Cleckley, 1976 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib313) ). A number of predictions follow from this appraisal. For example, psychopaths should lie convincingly, have a greater tolerance for physical pain, show less autonomic arousal in the resting state, and get into trouble because of their lack of behavioral inhibition. Thus, to validate a measure of psychopathy, we would need to check out a number of different expectations based on our theory of psychopathy.
Construct validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss64) pertains to psychological tests that claim to measure complex, multifaceted, and theory-bound psychological attributes such as psychopathy, intelligence, leadership ability, and the like. The crucial point to understand about construct validity is that “no criterion or universe of content is accepted as entirely adequate to de�ine the quality to be measured” (Cronbach & Meehl, 1955 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib376) ). Thus, the demonstration of construct validity always rests on a program of research using diverse procedures outlined in the following sections. To evaluate the construct validity of a test, we must amass a variety of evidence from numerous sources.
Many psychometric theorists regard construct validity as the unifying concept for all types of validity evidence (Cronbach, 1988 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib375) ; Messick, 1995 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1135) ). According to this viewpoint, individual studies of content, concurrent, and predictive validity are regarded merely as supportive evidence in the cumulative quest for construct validation.
2The construct of psychopathy is very similar to what is now designated as antisocial personality disorder (American Psychiatric Association, 1994 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib32) ).
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 10/70
4.5 APPROACHES TO CONSTRUCT VALIDITY How does a test developer determine whether a new instrument possesses construct validity? As previously hinted, no single procedure will suf�ice for this dif�icult task. Evidence of construct validity can be found in practically any empirical study that examines test scores from appropriate groups of subjects. Most studies of construct validity fall into one of the following categories:
Analysis to determine whether the test items or subtests are homogeneous and therefore measure a single construct Study of developmental changes to determine whether they are consistent with the theory of the construct Research to ascertain whether group differences on test scores are theory-consistent Analysis to determine whether intervention effects on test scores are theory-consistent Correlation of the test with other related and unrelated tests and measures Factor analysis of test scores in relation to other sources of information Analysis to determine whether test scores allow for the correct classi�ication of examinees
We examine these sources of construct validity evidence in more detail in the following.
Test Homogeneity If a test measures a single construct, then its component items (or subtests) likely will be homogeneous (also referred to as internally consistent). In most cases, homogeneity is built into the test during the development process discussed in more detail in the next unit. The aim of test development is to select items that form a homogeneous scale (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss147) . The most commonly used method for achieving this goal is to correlate each potential item with the total score and select items that show high correlations with the total score. A related procedure is to correlate subtests with the total score in the early phases of test development. In this manner, wayward scales that do not correlate to some minimum degree with the total test score can be revised before the instrument is released for general use.
Homogeneity is an important �irst step in certifying the construct validity of a new test, but standing alone it is weak evidence. Kline (1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib900) ) has pointed out the circularity of the procedure:
If all our items in the item pool were wide of the mark and did not measure what we hoped, they would be selecting items by the criterion of their correlation with the total score, which can never work. It is to be noted that the same argument applies to the factoring of the item pool. A general factor of poor items is still possible. This objection is sound and has to be refuted empirically. Having found by item analysis a set of homogeneous items, we must still present evidence concerning their validity. Thus to construct a homogeneous test is not suf�icient, validity studies must be carried out.
In addition to demonstrating the homogeneity of items, a test developer must provide multiple other sources of construct validity, discussed subsequently.
Appropriate Developmental Changes Many constructs can be assumed to show regular age-graded changes from early childhood into mature adulthood and perhaps beyond. Consider the construct of vocabulary knowledge as an example. It has been known since the inception of intelligence tests at the turn of the century that knowledge of vocabulary increases exponentially from early childhood into late childhood. More recent research demonstrates that vocabulary continues to grow, albeit at a slower pace, into old age (Gregory & Gernert, 1990 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib652) ). For any new test of vocabulary, then, an important piece of construct validity evidence would be that older subjects score better than younger subjects, assuming that education and health factors are held constant.
Of course, not all constructs lend themselves to predictions about developmental changes. For example, it is not clear whether a scale measuring “assertiveness” should show a pattern of increasing, decreasing, or stable scores with advancing age. Developmental changes would be irrelevant to the construct validity of such a scale. We should also mention that appropriate developmental changes are but one piece in the construct validity puzzle. This approach does not provide information about how the construct relates to other constructs.
Theory-Consistent Group Differences One way to bolster the validity of a new instrument is to show that, on average, persons with different backgrounds and characteristics obtain theory-consistent scores on the test. Speci�ically, persons thought to be high on the construct measured by the test should obtain high scores, whereas persons with presumably low amounts of the construct should obtain low scores.
Crandall (1981 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib368) ) developed a social interest scale that illustrates the use of theory-consistent group differences in the process of construct validation. Borrowing from Alfred Adler, Crandall (1984) de�ined social interest as an “interest in and concern for others.” To measure this construct, he devised a brief and simple instrument consisting of 15 forced-choice items. For each item, one of the two alternatives includes a trait closely related to the Adlerian concept of social interest (e.g., helpful), whereas the other choice consists of an equally attractive but nonsocial trait (e.g., quick-witted). The subject is instructed to “choose the trait which you value more highly.” Each of the 15 items is scored 1 if the social interest trait is picked, 0 otherwise; thus, total scores on the Social Interest Scale (SIS) can range from 0 to 15.
Table 4.1 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec5#ch04tab1) presents average scores on the SIS for 13 well-de�ined groups of subjects. The reader will notice that individuals likely to be high in social interest (e.g., nuns) obtain the highest average scores on the SIS, whereas the lowest scores are earned by presumably self-centered persons (e.g., models) and those who are outright antisocial (felons). These �indings are theory-consistent and support the construct validity of this interesting instrument.
TABLE 4.1 Mean Scores on the Social Interest Scale for Selected Groups
Group N Mean Score Ursuline sisters 6 13.3 Adult church members 147 11.2 Charity volunteers 9 10.8 High school students nominated for high social interest 23 10.2 University students nominated for high social interest 21 9.5
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 11/70
Group N Mean Score University employees 327 8.9 University students 1,784 8.2 University students nominated for low social interest 35 7.4 Professional models 54 7.1 High school students nominated for low social interest 22 6.9 Adult atheists and agnostics 30 6.7 Convicted felons 30 6.4
Source: Adapted with permission from Crandall, J. (1981). Theory and measurement of social interest: Empirical tests of Alfred Adler’s concept. New York: Columbia University Press.
Theory-Consistent Intervention Effects Another approach to construct validation is to show that test scores change in appropriate direction and amount in reaction to planned or unplanned interventions. For example, the scores of elderly persons on a spatial orientation test battery should increase after these subjects receive cognitive training speci�ically designed to enhance their spatial orientation abilities. More precisely, if the test battery possesses construct validity, we can predict that spatial orientation scores should show a greater increase from pretest to posttest than found on unrelated abilities not targeted for special training (e.g., inductive reasoning, perceptual speed, numerical reasoning, or verbal reasoning). Willis and Schaie (1986) found just such a pattern of test results in a cognitive training study with elderly subjects, supporting the construct validity of their spatial orientation measure.
Convergent and Discriminant Validation Convergent validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss70) is demonstrated when a test correlates highly with other variables or tests with which it shares an overlap of constructs. For example, two tests designed to measure different types of intelligence should, nonetheless, share enough of the general factor in intelligence to produce a hefty correlation (say, .5 or above) when jointly administered to a heterogeneous sample of subjects. In fact, any new test of intelligence that did not correlate at least modestly with existing measures would be highly suspect, on the grounds that it did not possess convergent validity.
Discriminant validity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss92) is demonstrated when a test does not correlate with variables or tests from which it should differ. For example, social interest and intelligence are theoretically unrelated, and tests of these two constructs should correlate negligibly, if at all.
In a classic paper often quoted but seldom emulated, Campbell and Fiske (1959 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib259) ) proposed a systematic experimental design for simultaneously con�irming the convergent and discriminant validities of a psychological test. Their design is called the multitrait-multimethod matrix, and it calls for the assessment of two or more traits by two or more methods. Table 4.2 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec5#ch04tab2) provides a hypothetical example of this approach. In this example, three traits (A, B, and C) are measured by three methods (1, 2, and 3). For example, traits A, B, and C might be social interest, creativity, and dominance. Methods 1, 2, and 3 might be self-report inventory, peer ratings, and projective test. Thus, A1 would represent a self-report inventory of social interest, B2 a peer rating of creativity, C3 a dominance measure derived from projective test, and so on.
TABLE 4.2 Hypothetical Multitrait-Multimethod Matrix
Notice in this example that nine tests are studied (three traits are each measured by three methods). When each of these tests is administered twice to the same group of subjects and scores on all pairs of tests are correlated, the result is a multitrait-multimethod matrix (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss213) (Table 4.2
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 12/70
(http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec5#ch04tab2) ). This matrix is a rich source of data on reliability, convergent validity, and discriminant validity.
For example, the correlations along the main diagonal (in parentheses) are reliability coef�icients for each test. The higher these values, the better, and preferably we like to see values in the .80s or .90s here. The correlations along the three shorter diagonals (in boldface) supply evidence of convergent validity—the same trait measured by different methods. These correlations should be strong and positive, as shown here. Notice that the table also includes correlations between different traits measured by the same method (in solid triangles) and different traits measured by different methods (in dotted triangles). These correlations should be the lowest of all in the matrix, insofar as they supply evidence of discriminant validity.
The Campbell and Fiske (1959 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib259) ) methodology is an important contribution to our understanding of the test validation process. However, the full implementation of this procedure typically requires too monumental a commitment from researchers. It is more common for test developers to collect convergent and discriminant validity data in bits and pieces, rather than producing an entire matrix of intercorrelations. Meier (1984 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1121) ) provides one of the few real-world implementations of the multitrait-multimethod matrix in an examination of the validity of the “burnout” construct.
Factor Analysis Factor analysis is a specialized statistical technique that is particularly useful for investigating construct validity. We discuss factor analysis in substantial detail in Topic 5A (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch05#ch05box1) , Intelligence Tests and Factor Analysis; here, we provide a quick preview so that the reader can appreciate the role of factor analysis in the study of construct validity. The purpose of factor analysis (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss113) is to identify the minimum number of determiners (factors) required to account for the intercorrelations among a battery of tests. The goal in factor analysis is to �ind a smaller set of dimensions, called factors, that can account for the observed array of intercorrelations among individual tests. A typical approach in factor analysis is to administer a battery of tests to several hundred subjects and then calculate a correlation matrix from the scores on all possible pairs of tests. For example, if 15 tests have been administered to a sample of psychiatric and neurological patients, the �irst step in factor analysis is to compute the correlations between scores on the 105 possible pairs of tests.3 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec5#ch04fn03) Although it may be feasible to see certain clusterings of tests that measure common traits, it is more typical that the mass of data found in a correlation matrix is simply too complex for the unaided human eye to analyze effectively. Fortunately, the computer-implemented procedures of factor analysis search this pattern of intercorrelations, identify a small number of factors, and then produce a table of factor loadings. A factor loading (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss114) is actually a correlation between an individual test and a single factor. Thus, factor loadings can vary between −1.0 and +1.0. The �inal outcome of a factor analysis is a table depicting the correlation of each test with each factor.
We can illustrate the use of factor analysis in the study of construct validity by referring to a speci�ic instrument, the Wechsler Adult Intelligence Scale-IV (WAIS- IV, Wechsler, 2008 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1737) ), discussed in more detail in the next chapter. The 10 core subtests of the WAIS-IV yield not only a Full Scale IQ, but also four Index scores designed to provide a meaningful and theoretically sound partition of intelligence into subcomponents. These Index scores are Verbal Comprehension (3 subtests), Perceptual Reasoning (3 subtests), Working Memory (2 subtests), and Processing Speed (2 subtests). When factor analysis is applied to WAIS-IV subtest scores for large samples of adults, four factors are found, just as predicted by the structure of the test (Ryan, Sattler, & Tree, 2009 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1423) ). Further, each core subtest usually demonstrates its highest factor loading on the appropriate factor. For example, the Vocabulary subtest shows its highest factor loading on Verbal Comprehension, and the Matrix Reasoning subtest reveals its highest factor loading on Perceptual Reasoning. Findings like this bolster the construct validity of the WAIS-IV.
Classi�ication Accuracy Many tests are used for screening purposes to identify examinees who meet (or don’t meet) certain diagnostic criteria. For these instruments, accurate classi�ication is an essential index of validity. As a basis for illustrating this approach to validation, we consider the Mini-Mental State Examination (MMSE), a short screening test of cognitive functioning. The MMSE consists of a number of simple questions (e.g., What day is this?) and easy tasks (e.g., remembering three words). The test yields a score from 0 (no items correct) to 30 (all items correct). Although used for many purposes, a major application of the MMSE is to identify elderly individuals who might be experiencing dementia. Dementia is a general term that refers to signi�icant cognitive decline and memory loss caused by a disease process such as Alzheimer’s disease or the accumulation of small strokes. Both the MMSE and various forms of dementia are described in more detail in Chapter 10 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch10#ch10) , Neuropsychological Assessment and Screening.
The MMSE is one of the most widely researched screening tests in existence. Much is known about its measurement qualities, such as the accuracy of the tool in detecting individuals with dementia. In exploring its utility, researchers have paid special attention to two psychometric features that bear upon validity: sensitivity and speci�icity. Sensitivity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss291) has to do with accurate identi�ication of patients who have a syndrome—in this case, dementia. Speci�icity (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss300) has to do with accurate identi�ication of normal patients. These ideas are clari�ied later. Understanding these concepts is pertinent to the validity of every screening test used in mental health and medicine. Thus, we provide modest coverage here, using the MMSE as an exemplar of a more general principle. Our discussion loosely follows the presentation found in Gregory (1999 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib650) ).
The concepts of sensitivity and speci�icity are chie�ly helpful in dichotomous diagnostic situations in which individuals are presumed either to manifest a syndrome or not. For example, in medicine a patient either has prostate cancer or he does not. In this case, the criterion of truth, against which a screening test is measured, would be a tissue biopsy. Similarly, in research studies on the sensitivity and speci�icity of the MMSE, patients are known from independent, comprehensive medical and psychological workups either to meet the criteria for dementia or not. This is the “gold standard” against which the screening instrument is validated. The rationale for the screening test is pragmatic: It is unrealistic to refer every patient with suspected dementia for comprehensive evaluations that would include, for example, many hours of professional time (psychologist, neurologist, geriatric specialist, etc.) and expensive brain scans. The purpose of the MMSE—or any screening test—is to determine the need for additional assessment.
Screening tests typically provide a cutoff score used to identify possible cases of the syndrome in question. With the MMSE, a common cutting score is 23/24 out of the 30 points possible. Thus, a score of 23 points and below indicates the likelihood of dementia, whereas 24 points and above is considered normal. In this context, the sensitivity of the MMSE is the percentage of patients known to have dementia who score 23 points or lower. For example, if 100 patients are known from independent, comprehensive evaluations to exhibit dementia, and 79 of them score 23 or below, then the sensitivity of the test is 79 percent. The speci�icity of the MMSE is the other side of the coin, the percentage of patients known to be normal who score 24 points or higher. For example, if 83 of 100 normal patients score 24 points or higher, then the speci�icity of the test is 83 percent.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 13/70
In general, the validity of a screening test is bolstered to the extent that it possesses both high sensitivity and high speci�icity. There are no exact cutoffs, but for many purposes a test will need sensitivity and speci�icity that exceed 80 or 90 percent in order to justify its use. As we will see later, the standards for sensitivity and speci�icity are unique to each situation and depend on the costs—both �inancial and otherwise—of different kinds of errors in classi�ication.
An ideal screening test, of course, would yield 100 percent sensitivity and 100 percent speci�icity. No such test exists in the real world. The reality of assessment is that the examiner must choose a cutoff score that provides a balance between sensitivity and speci�icity. What makes this problematic is that sensitivity and speci�icity are inversely related. Choosing a cutoff score that increases sensitivity invariably will reduce speci�icity, and vice versa. The inverse relationship between sensitivity and speci�icity is not only an empirical fact, but it is also a logical necessity—if one improves, the other must decline—no exceptions are possible. Practitioners need to select a cutoff score that produces a livable balance between sensitivity and speci�icity. But exactly where is that point of equilibrium? In the case of the MMSE, the answer depends not just on the age and education of the client but also on the relative advantages and drawbacks of correct or incorrect decisions. Robust levels of sensitivity and speci�icity provide corroborating evidence of test validity, and test developers should strive to achieve the highest possible levels of both.
3The general formula for the number of pairings among N tests is N(N − 1)/2. Thus, if 15 tests are administered, there will be 15 × 14/2 or 105 possible pairings of individual tests.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 14/70
4.6 EXTRAVALIDITY CONCERNS AND THE WIDENING SCOPE OF TEST VALIDITY We begin this section with a review of extravalidity concerns (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss108) , which include side effects and unintended consequences of testing. By acknowledging the importance of the extravalidity domain, psychologists con�irm that the decision to use a test involves social, legal, and political considerations that extend far beyond the traditional questions of technical validity. In a related development, we will also review how the interest in extravalidity concerns has spurred several theorists to broaden the concept of test validity. As the reader will discover, value implications and social consequences are now encompassed within the widening scope of test validity.
Even if a test is valid, unbiased, and fair, the decision to use it may be governed by additional considerations. Cole and Moss (1998) outline the following factors:
What is the purpose for which the test is used? To what extent are the purposes accomplished by the actions taken? What are the possible side effects or unintended consequences of using the test? What possible alternatives to the test might serve the same purpose?
We survey only the most prominent extravalidity concerns here and show how they have served to widen the scope of test validity.
Unintended Side Effects of Testing The intended outcome of using a psychological test is not necessarily the only consequence. Various side effects also are possible, indeed, they are probable. The examiner must determine whether the bene�its of giving the test outweigh the costs of the potential side effects. Furthermore, by anticipating unintended side effects, the examiner might be able to de�lect or diminish them.
Cole and Moss (1998) cite the example of using psychological tests to determine eligibility for special education. Although the intended outcome is to help students learn, the process of identifying students eligible for special education may produce numerous negative side effects:
The identi�ied children may feel unusual or dumb. Other children may call the children names. Teachers may view these children as unworthy of attention. The process may produce classes segregated by race or social class.
A consideration of side effects should in�luence an examiner’s decision to use a particular test for a speci�ied purpose. The examiner might appropriately choose not to use a test for a worthy purpose if the likely costs from side effects outweigh the expected bene�its.
Consider the common practice in years past of using the Minnesota Multiphasic Personality Inventory (MMPI) to help screen candidates for peace of�icer positions such as police of�icer or sheriff ’s deputy. Although the MMPI was originally designed as an aid in psychiatric diagnosis, subsequent research indicated that it is also useful in the identi�ication of persons unsuited to a career in law enforcement (Hiatt & Hargrave, 1988 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib742) ). In particular, peace of�icers who produce MMPI pro�iles with mild elevations (e.g., T score 65 to 69) on Scales F (Frequency), Masculinity-Femininity, Paranoia, and Hypomania tend to be involved in serious disciplinary actions; peace of�icers who produce more “defensive” MMPI pro�iles with fewer clinical scale elevations tend not to be involved in such actions. Thus, the test possessed modest validity for the worthy purpose of screening law enforcement candidates. But no test, not even the highly respected MMPI, is perfectly valid. Some good applicants will be passed over because their MMPI results are marginal. Perhaps their Paranoia Scale is at a T score of 66, or the Hypomania Scale is at a T score of 68. On the MMPI, a T score of 70 is often considered the upper limit of the “normal” range.
One unintended side effect of using the MMPI for evaluation of peace of�icer applicants is that job candidates who are unsuccessful with one agency may be tagged with a pathological label such as psychopathic, schizophrenic, or paranoid. The label may arise in spite of the best efforts of the consulting psychologist, who may never have used any pejorative terms in the assessment report on the candidate. Typically, the label is conceived when administrators at the referring department look at the MMPI pro�ile and see that the candidate obtained his or her highest score on a scale with a horrendous title such as Psychopathic Deviate, Schizophrenia, Hypochondriasis, or Paranoia. Unfortunately, the law enforcement community can be a very closed fraternity. Police chiefs and sheriffs commonly exchange verbal reports about their job applicants, so a pejorative label may follow the candidate from one setting to another, permanently barring the applicant from entry into the law enforcement profession. The repercussions are not only unfair to the candidate, but they also raise the specter of lawsuits against the agency and the consulting psychologist. All things considered, the consulting psychologist may �ind it preferable to use a technically less valid test for the same purpose, particularly if the alternative instrument does not produce these unintended side effects.
The renewed sensitivity to extravalidity issues has caused several test theorists to widen their de�inition of test validity. We review these recent developments in the following section, cautioning the reader that a �inal consensus about the nature of test validity is yet to emerge.
The Widening Scope of Test Validity By now the reader is familiar with the narrow, traditionalist perspective on test use, which states that a test is valid if it measures “what it purports to measure.” The implicit implication of this perspective is that technical validity is the most essential basis for recommending test use. After all, valid tests provide accurate information about examinees—and what could be wrong with that?
Recently, several psychometric theoreticians have introduced a wider, functionalist de�inition of validity that asserts that a test is valid if it serves the purpose for which it is used (Cronbach, 1988 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib375) ; Mes-sick, 1995 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1135) ). For example, a reading achievement test might be used to identify students for assignment to a remedial section. According to the functionalist perspective, the test would be valid—and its use, therefore, appropriate—if the students selected for remediation actually received some academic bene�it from this application of the test.
The functionalist perspective explicitly recognizes that the test validator has an obligation to determine whether a practice has constructive consequences for individuals and institutions and especially to guard against adverse outcomes (Messick, 1980 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1134) ). Test validity, then, is an overall evaluative judgment of the adequacy and appropriateness of inferences and actions that �low from test scores.
Messick (1980 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1134) , 1995 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1135) ) argues that the new, wider conception of validity rests on four bases.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 15/70
These are (1) traditional evidence of construct validity, for example, appropriate convergent and discriminant validity, (2) an analysis of the value implications of the test interpretation, (3) evidence for the usefulness of test interpretations in particular applications, and (4) an appraisal of the potential and actual social consequences, including side effects, from test use. A valid test is one that answers well to all four facets of test validity.
This wider conception of test validity is admittedly controversial, and some theorists prefer the traditional view that consequences and values are important but nonetheless separate from the technical issues of test validity. Everyone can agree on one point: Psychological measurement is not a neutral endeavor, it is an applied science that occurs in a social and political context.
Utility: The Last Horizon of Test Validity Finally, we introduce the concept of test utility, which is widely neglected in the research literature on psychological testing (Hunsley & Bailey, 1999 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib798) ). As noted by Wood, Garb, and Nezworski (2007 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1785) ), test utility can be summed up by the question “Does use of this test result in better patient outcomes or more ef�icient delivery of services?” For example, we might envision an experiment in which individual psychotherapy clients were randomly assigned to two groups. One group is tested with the Beck Depression Inventory-2 (Beck, Steer, & Brown, 1996 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib118) ) and the results provided to their therapists, while the other group is not tested but instead proceeds directly for treatment. If the tested group showed more improvement or required fewer sessions to achieve the same level of improvement, we would conclude that utility has been demonstrated for the test.
Unfortunately, there is very little research on the utility of psychological tests, and the research that does exist is indirect. For example, Finn and Tonsager (1992 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib500) ) have shown that a highly structured method for giving feedback on personality test �indings to college students awaiting psychotherapy has initial therapeutic effects in its own right. However, this does not answer the question whether the ultimate client outcome is better as a result of the test usage. For some tests such as the Rorschach inkblot technique, discussed later in the text, the question of utility is especially pertinent because of the time required by a psychologist to administer, score, interpret, and document the results. The total time easily can run to many hours. It is lamentable that the utility of this instrument and many other tests has not been systematically investigated.
TOPIC 4B Test Construction
4.7 De�ining the Test (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec7#ch04lev1sec7)
4.8 Selecting a Scaling Method (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec8#ch04lev1sec8)
4.9 Representative Scaling Methods (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec9#ch04lev1sec9)
4.10 Constructing the Items (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec10#ch04lev1sec10)
4.11 Testing the Items (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04lev1sec11)
4.12 Revising the Test (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec12#ch04lev1sec12)
4.13 Publishing the Test (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec13#ch04lev1sec13)
Creating a new test involves both science and art. A test developer must choose strategies and materials and then make day-to-day research decisions that will affect the quality of his or her emerging instrument. The purpose of this section is to discuss the process by which psychometricians create valid tests. Although we will discuss many separate topics, they are united by a common theme: Valid tests do not just materialize on the scene in full maturity—they emerge slowly from an evolutionary, developmental process that builds in validity from the very beginning. We will emphasize the basics of test development here; readers who desire a more advanced presentation should consult Kline (1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib900) ), McDonald (1999 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1090) ), and Bernstein and Nunnally (1994 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib151) ).
Test construction consists of six intertwined stages:
De�ining the test Testing the items Selecting a scaling method Revising the test Constructing the items Publishing the test
By way of preview, we can summarize these steps as follows: de�ining the test consists of delimiting its scope and purpose, which must be known before the developer can proceed to test construction. Selecting a scaling method is a process of setting the rules by which numbers are assigned to test results. Constructing the items is as much art as science, and it is here that the creativity of the test developer may be required. Once a preliminary version of the test is available, the developer usually administers it to a modest-sized sample of subjects in order to collect initial data about test item characteristics. Testing the items entails a variety of statistical procedures referred to collectively as item analysis. The purpose of item analysis is to determine which items should be retained, which revised, and which thrown out. Based on item analysis and other sources of information, the test is then revised. If the revisions are substantial, new items and additional pretesting with new subjects may be required. Thus, test construction involves a feedback loop whereby second, third, and fourth drafts of an instrument might be produced (Figure 4.5 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec7#ch04�ig5) ). Publishing the test is the �inal step. In addition to releasing the test materials, the developer must produce a user-friendly test manual. Let us examine each of these steps in more detail.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 16/70
4.7 DEFINING THE TEST
FIGURE 4.5 The Test Construction Process
In order to construct a new test, the developer must have a clear idea of what the test is to measure and how it is to differ from existing instruments. Insofar as psychological testing is now entering its second one hundred years, and insofar as thousands of tests have already been published, the burden of proof clearly rests on the test developer to show that a proposed instrument is different from, and better than, existing measures.
Consider the daunting task faced by a test developer who proposes yet another measure of general intelligence. With dozens of such instruments already in existence, how could a new test possibly make a useful contribution to the �ield? The answer is that contemporary research continually adds to our understanding of intelligence and impels us to seek new and more useful ways to measure this multifaceted construct.
Kaufman and Kaufman (1983 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib861) ) provide a good model of the test de�inition process. In proposing the Kaufman Assessment Battery for Children (K-ABC), a new test of general intelligence in children, the authors listed six primary goals that de�ine the purpose of the test and distinguish it from existing measures:
Measure intelligence from a strong theoretical and research basis Separate acquired factual knowledge from the ability to solve unfamiliar problems Yield scores that translate to educational intervention Include novel tasks Be easy to administer and objective to score Be sensitive to the diverse needs of preschool, minority group, and exceptional children (Kaufman & Kaufman, 1983 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib861) )
The K-ABC represents an interesting departure from traditional intelligence tests. For now, the important point is that the developers of this instrument, now in its second edition (K-ABC-II), explained its purpose explicitly and proposed a fresh focus for measuring intelligence, long before they started constructing test items.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 17/70
4.8 SELECTING A SCALING METHOD The immediate purpose of psychological testing is to assign numbers to responses on a test so that the examinee can be judged to have more or less of the characteristic measured. The rules by which numbers are assigned to responses de�ine the scaling method. Test developers select a scaling method that is optimally suited to the manner in which they have conceptualized the trait(s) measured by their test. No single scaling method is uniformly better than the others. For some traits, ordinal ranking of expert judges might be the best measurement approach; for other traits, complex scaling of self-report data might yield the most valid measurements.
There are so many distinctive scaling methods available to psychometricians that we will be satis�ied to provide only a representative sample here. Readers who wish a more thorough and detailed review should consult Gulliksen (1950 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib672) ), Nunnally (1978 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1243) ), or Kline (1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib900) ). However, before reviewing selecting scaling methods, we need to introduce a related concept, levels of measurement, so that the reader can better appreciate the differences between scaling methods.
Levels of Measurement According to Stevens (1946 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1575) ), all numbers derived from measurement instruments of any kind can be placed into one of four hierarchical categories: nominal, ordinal, interval, or ratio. Each category de�ines a level of measurement; the order listed is from least to most informative.
In a nominal scale (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss216) , the numbers serve only as category names. For example, when collecting data for a demographic study, a researcher might code males as “1” and females as “2.” Notice that the numbers are arbitrary and do not designate “more” or “less” of anything. In nominal scales the numbers are just a simpli�ied form of naming.
An ordinal scale (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss230) constitutes a form of ordering or ranking. If college professors were asked to rank order four cars as to which they would prefer to own, the preferred order might be “1” Cadillac, “2” Chevrolet, “3” Volkswagen, “4” Hyundai. Notice here that the numbers are not interchangeable. A ranking of “1” is “more” than a ranking of “2,” and so on. The “more” refers to the order of preference. However, ordinal scales fail to provide information about the relative strength of rankings. In this hypothetical example, we do not know whether college professors strongly prefer Cadillacs over Chevrolets or just marginally so.
An interval scale (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss165) provides information about ranking, but also supplies a metric for gauging the differences between rankings. To construct an interval scale, we might ask our college professors to rate on a scale from 1 to 100 how much they would like to own the four cars previously listed. Suppose the average ratings work out as follows: Cadillac, 90; Chevrolet, 70; Volkswagen, 60; Hyundai, 50. From this information we could infer that the preference for a Cadillac is much stronger than for a Chevrolet, which, in turn, is mildly stronger than the preference for a Volkswagen. More important, we can also make the assumption that the intervals between the points on this scale are approximately the same: The difference between professors’ preference for a Chevrolet and Volkswagen (10 points) is about the same as that between a Volkswagen and a Hyundai (also 10 points). In short, interval scales are based on the assumption of equal-sized units or intervals for the underlying scale.
A ratio scale (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss269) has all the characteristics of an interval scale but also possesses a conceptually meaningful zero point in which there is a total absence of the characteristic being measured. The essential characteristics of the four levels of measurement are summarized in Figure 4.6 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec8#ch04�ig6) .
Ratio scales are rare in psychological measurement. Consider whether there is any meaningful sense in which a person can be thought to have zero intelligence. Not really. The same is true for most constructs in psychology: Meaningful zero points just do not exist. However, a few physical measures used by psychologists qualify as ratio scales. For example, height and weight qualify, and perhaps some physiological measures such as electrodermal response qualify, too. But by and large the best a psychologist can hope for is interval-level measurement.
FIGURE 4.6 Essential Characteristics of Four Levels of Measurement
Levels of measurement are relevant to test construction because the more powerful and useful parametric statistical procedures (e.g., Pearson r, analysis of variance, multiple regression) should be used only for scores derived from measures that meet the criteria of interval or ratio scales. For scales that are only nominal or ordinal, less-powerful non-parametric statistical procedures (e.g., chi-square, rank order correlation, median tests) must be employed. In practice, most major psychological testing instruments (especially intelligence tests and personality scales) are assumed to employ approximately interval-level measurement even though, strictly speaking, it is very dif�icult to demonstrate absolute equality of intervals for such instruments (Bausell, 1986
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 18/70
(http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib110) ). Now that the reader is familiar with levels of measurement, we introduce a representative sample of scaling methods, noting in advance that different scaling methods yield different levels of measurement.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 19/70
4.9 REPRESENTATIVE SCALING METHODS
Expert Rankings Suppose we wanted to measure the depth of coma in patients who had suffered a recent head injury that rendered them unconscious. A depth of coma scale could be very important in predicting the course of improvement, because it is well known that a lengthy period of unconsciousness offers a poor prognosis for ultimate recovery. In addition, rehabilitation personnel have a practical need to know whether a patient is deeply comatose or in a partially communicative state of twilight consciousness.
One approach to scaling the depth of coma would be to rely on the behavioral rankings of experts. For example, we could ask a panel of neurologists to list patient behaviors associated with different levels of consciousness. After the experts had submitted a large list of diagnostic behaviors, the test developers— preferably experts on head injuries—could rank the indicator behaviors along a continuum of consciousness ranging from deep coma to basic orientation. Using precisely this approach, Teasdale and Jennett (1974 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1620) ) produced the Glasgow Coma Scale. Instruments similar to this scale are widely used in hospitals for the assessment of traumatic brain injury (Figure 4.7 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec9#ch04�ig7) ).
FIGURE 4.7 Example of the Use of the Glasgow Coma Scale for Recording Depth of Coma Source: Reprinted with permission from Jennett, B., Teasdale, G. M., & Knill-Jones, R. P. (1975). Predicting outcome after head injury. Journal of the Royal College of Physicians of London, 9, 231–237.
The Glasgow Coma Scale is scored by observing the patient and assigning the highest level of functioning on each of three subscales. On each sub-scale, it is assumed that the patient displays all levels of behavior below the rated level. Thus, from a psychometric standpoint, this scale consists of three sub-scales (eyes, verbal response, and motor response) each yielding an ordinal ranking of behavior.
In addition to the rankings, it is possible to compute a single overall score that is something more than an ordinal scale, although probably less than true interval-level measurement. If numbers are attached to the rankings (e.g., for eyes open a coding of “none” = 1, “to pain” = 2, and so on), then the numbers for the rated level for each subscale can be added, yielding a maximum possible score of 14. The total score on the Glasgow Coma Scale predicts later recovery with a very high degree of accuracy (Jennett, Teasdale, & Knill-Jones, 1975 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib827) ). We see, then, that quite plain psychological tests derived from the very simplest scaling methods can, nonetheless, provide valid and useful information.
Method of Equal-Appearing Intervals Early in the twentieth century, L. L. Thurstone (1929 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1645) ) proposed a method for constructing interval-level scales from attitude statements. His method of equal-appearing intervals is still used today, marking him as one of the giants of psychometric theory. The actual methodology of constructing equal-appearing intervals is somewhat complex and statistically laden, but the underlying logic is easy to explain (Ghiselli, Campbell, & Zedeck, 1981 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib585) ). We illustrate the method by summarizing the steps involved in constructing a scale of attitudes toward physical exercise.
First, a large number of true–false statements re�lecting a range of positive and negative attitudes toward physical exercise would be compiled. Two extreme examples might be:
“I feel that physical exercise is generally boring and tedious.” “Physical exercise should be a signi�icant part of everyone’s daily life.”
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 20/70
Of course, many items of moderate attitudinal valence would be written as well. The idea at this point-of-scale development is to produce an excess of items with the expectation that unsuitable items later will be dropped.
Next, these attitude statements would be presented to a group of judges (up to a dozen individuals) who would sort each statement into 1 of 11 categories that range from “extremely favorable” to “extremely unfavorable.” Then, the average favorability for each item (−1.0 to +1.0) would be calculated, along with the standard deviation. Items with larger standard deviations would be dropped, because they produce unreliable ratings. Finally, about 20 to 30 items would be chosen to cover the range of the dimension (favorable to unfavorable). The items on the �inal scale are assumed to meet the criteria of an interval scale. The score for persons who take the attitude scale is the average scale value of those items endorsed as true (or false, in the case of negatively worded items).
Ghiselli et al. (1981 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib585) ) note that the preceding scaling method merely produces the attitude scale. Reliability and validity analyses of the scale are still needed to determine its appropriateness and usefulness.
A study by Russo (1994 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1420) ) illustrates a modern application of the Thurstone method. She used a Thurstone scaling approach to evaluate 216 items from three prominent self-report depression inventories. The judges included 527 undergraduates and 37 clinical faculty members at a medical school. The 216 items were randomized and rated with respect to depressive severity from 1 representing no depression to 11 representing extreme depression. She discovered that all three self-report inventories lacked items and response options typical of mild depression. The distribution of the 216 items was bimodal with many items bunched near the bottom (no depression) and many items bunched near the middle (moderate depression). A characteristic �inding for one set of items from a prominent depression scale was as follows:
Rated Depression Original Scoring Item Content 1.0 1 I never feel downhearted or sad. 3.4 2 I sometimes feel downhearted or sad. 4.1 3 I feel downhearted or sad a good part of the time. 4.4 4 I feel downhearted or sad most of the time.
The reader will notice that the original scoring on these items deviates substantially from the depression ratings provided by the panel of students and clinical faculty. It is also evident that the actual scale values are discontinuous, jumping from 1.0 to 3.4 and higher. A similar pattern was observed for many items on all three inventories, leading Russo (1994 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1420) ) to conclude:
The present results suggest that if the original scoring is used for the three scales examined here, then the distinctions between well-being and absence of depression as well as between moderate and severe will be dif�icult to make. Such imprecision will make it dif�icult to assess the ef�icacy of treatments for depression, because a lack thereof must be a function of added measurement error due to ordinal measures. Such error could also wreak havoc in longitudinal studies, especially in those in which memory is involved.
We see in this example that Thurstone’s approach to item scaling has powerful applications in test development. Based on these �indings, researchers are now in a position to develop improved self-report scales that assess the full range of symptomatology in depression.
Method of Absolute Scaling Thurstone (1925 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1644) ) also developed the method of absolute scaling (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss200) , a procedure for obtaining a measure of absolute item dif�iculty based on results for different age groups of test takers. The methodology for determining individual item dif�iculty on an absolute scale is quite complex, although the underlying rationale is not too dif�icult to understand. Essentially, a set of common test items is administered to two or more age groups. The relative dif�iculty of these items between any two age groups serves as the basis for making a series of interlocking comparisons for all items and all age groups. One age group serves as the anchor group. Item dif�iculty is measured in common units such as standard deviation units of ability for the anchor group. The method of absolute scaling is widely used in group achievement and aptitude testing (Donlon, 1984 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib431) ).
Thurstone (1925 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1644) ) illustrated the method of absolute scaling with data from the testing of 3,000 schoolchildren on the 65 questions from the original Binet test. Using the mean of Binet test intelligence of 3½-year-old children as the zero point and the standard deviation of their intelligence as the unit of measurement, he constructed a scale that ranged from −2 to +10 and then located each of the 65 questions on that scale. Thurstone (1925 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1644) ) found that the scale “brings out rather strikingly the fact that the questions are unduly bunched at certain ranges [of dif�iculty] and rather scarce at other ranges.” A modern test developer would use this kind of analysis as a basis for dropping redundant test items (redundant in the sense that they measure at the same dif�iculty level) and adding other items that test the higher (and lower) ranges of dif�iculty.
Likert Scales Likert (1932 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib988) ) proposed a simple and straightforward method for scaling attitudes that is widely used today. A Likert scale (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss186) presents the examinee with �ive responses ordered on an agree/disagree or approve/disapprove continuum. For example, one item on a scale to assess attitudes toward church membership might read:
Church services give me inspiration and help me to live up to my best during the following week.
Do you:
|| || || || || Strongly Agree Agree Undecided Disagree Strongly Disagree
Depending on the wording of an individual item, an extreme answer of “strongly agree” or “strongly disagree” will indicate the most favorable response on the underlying attitude measured by the questionnaire. Likert (1932 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib988) ) assigned a score of 5 to this extreme response, 1 to the opposite extreme, and 2, 3, and 4 to intermediate replies. The total scale score is obtained by adding the scores from individual items. For this reason, a Likert scale is also referred to as a summative scale.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 21/70
Guttman Scales On a Guttman scale, respondents who endorse one statement also agree with milder statements pertinent to the same underlying continuum (Guttman, 1947 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib676) ). Thus, if the examiner knows an examinee’s most extreme endorsement on the continuum, it is possible to reconstruct the intermediate responses as well. Guttman scales are produced by selecting items that fall into an ordered sequence of examinee endorsement. A perfect Guttman scale is seldom achieved because of errors of measurement, but is nonetheless a �itting goal for certain types of tests.
Although the Guttman approach was originally devised to determine whether a set of attitude statements is unidimensional, the technique has been used in many different kinds of tests. For example, Beck used Guttman-type scaling to produce the individual items of the Beck Depression Inventory (BDI, Beck, Steer, & Garbin, 1988 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib119) ). Items from the BDI resemble the following:
( ) I occasionally feel sad or blue. ( ) I often feel sad or blue. ( ) I feel sad or blue most of the time. ( ) I always feel sad and I can’t stand it.
Clients are asked to “check the statement from each group that you feel is most true about you.” A client who endorses an extreme alternative (e.g., “I always feel sad and I can’t stand it”) almost certainly agrees with the milder statements as well.
Method of Empirical Keying The reader may have noticed that most of the scaling methods discussed in the preceding section rely upon the authoritative judgment of experts in the selection and ordering of items. It is also possible to construct measurement scales based entirely on empirical considerations devoid of theory or expert judgment. In the method of empirical keying (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss201) , test items are selected for a scale based entirely on how well they contrast a criterion group from a normative sample. For example, a Depression scale could be derived from a pool of true-false personality inventory questions in the following manner:
A carefully selected and homogeneous group of persons experiencing major depression is gathered to answer the pool of true–false questions. For each item, the endorsement frequency of the depression group is compared to the endorsement frequency of the normative sample. Items which show a large difference in endorsement frequency between the depression and normative samples are selected for the Depression scale, keyed in the direction favored by depression subjects (true or false, as appropriate). Raw score on the Depression scale is then simply the number of items answered in the keyed direction.
The method of empirical keying can produce some interesting surprises. A common �inding is that some items selected for a scale may show no obvious relationship to the construct measured. For example, an item such as “I drink a lot of water” (keyed true) might end up on a Depression scale. The momentary rationale for including this item is simply that it works. Of course, the challenge posed to researchers is to determine why the item works. However, from the practical standpoint of empirical scale construction, theoretical considerations are of secondary importance. We discuss the method of empirical keying further in Topic 8B (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch08lev1sec11#ch08box2) , Self-Report and Behavioral Assessment of Psychopathology.
Rational Scale Construction (Internal Consistency) The rational approach to scale construction is a popular method for the development of self-report personality inventories. The name rational is somewhat of a misnomer, insofar as certain statistical methods are essential to this approach. Also, the name implies that other approaches are nonrational or irrational, which is untrue. The heart of the method of rational scaling (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss203) is that all scale items correlate positively with each other and also with the total score for the scale. An alternative and more appropriate name for this approach is internal consistency, which emphasizes what is actually done. Gough and Bradley (1992) explain how the rational approach earned its descriptive title:
The idea of rationality enters the scene in that the central theme or unifying dimension around which the items cluster is one that was conceptually articulated beforehand by the developer of the measure and from which the scoring of each item is determined in a logical and understandable way.
We will follow their presentation to illustrate the features of the rational approach.
Suppose a test developer desires to develop a new self-report scale for leadership potential. Based on a review of relevant literature, the researcher might conclude that leadership potential is characterized by self-con�idence, resilience under pressure, high intelligence, persuasiveness, assertiveness, and the ability to sense what others are thinking and feeling (Gough & Bradley, 1992). These notions suggest that the following true–false items might be useful in the assessment of leadership potential:
Most of the time I am pretty con�ident and sure of myself (T) When others disagree with me, I usually let things go. (F) I know that I am smarter than most people. (T) I am not very good at understanding how others react. (F) My friends would describe me as a dominant person. (T)
The T and F after each statement would indicate the rationally keyed direction for leadership potential.
Of course, additional items with similar intentions also would be proposed. The test developer might begin with 100 items that appear—on a rational basis—to assess leadership potential. These preliminary items would be administered to a large sample of individuals similar to the target population for whom the scale is intended. For instance, if the scale is designed to identify college students with leadership potential, then it should be administered to a cross-section of several hundred college students. For scale development, very large samples are desirable. In this hypothetical case, let us assume that we obtain results for 500 college students.
The next step in rational scale construction is to correlate scores on each of the preliminary items with the total score on the test for the 500 subjects in the tryout sample. Because scores on the items are dichotomous (1 is arbitrarily assigned to an answer corresponding to the scoring key, 0 to the alternative), a biserial correlation coef�icient rbis is needed. Once the correlations are obtained, the researcher scans the list in search of weak correlations and reversals (negative correlations). These items are discarded because they do not contribute to the measurement of leadership potential. Up to half of the initial items might be discarded. If a large proportion of items is initially discarded, the researcher might recalculate the item-total correlations based upon the reduced item
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 22/70
pool to verify the homogeneity of the remaining items. The items that survive this iterative procedure constitute the leadership potential scale. The reader should keep in mind that the rational approach to scale construction merely produces a homogeneous scale thought to measure a speci�ied construct. Additional studies with new subject samples would be needed to determine the reliability and validity of the new scale.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 23/70
4.10 CONSTRUCTING THE ITEMS Constructing test items is a painful and laborious procedure that taxes the creativity of test developers. The item writer is confronted with a profusion of initial questions:
Should item content be homogeneous or varied? What range of dif�iculty should the items cover? How many initial items should be constructed? Which cognitive processes and item domains should be tapped? What kind of test item should be used?
We will address the �irst three questions brie�ly before turning to a more detailed discussion of the last two topics, which are commonly referred to under the rubrics of table of speci�ications and item formats.
Initial Questions in Test Construction The �irst question pertains to the homogeneity versus heterogeneity of test item content. In large measure, whether item content is homogeneous or varied is dictated by the manner in which the test developer has de�ined the new instrument. Consider a culture-reduced test of general intelligence. Such an instrument might incorporate varied items, so long as the questions do not presume speci�ic schooling. The test developer might seek to incorporate novel problems equally unfamiliar to all examinees. On the other hand, with a theory-based test of spatial thinking, subscales with homogeneous item content would be required.
The range of item dif�iculty must be suf�icient to allow for meaningful differentiation of examinees at both extremes. The most useful tests, then, are those that include a graded series of very easy items passed by almost everyone as well as a group of incrementally more dif�icult items passed by virtually no one. A ceiling effect is observed when signi�icant numbers of examinees obtain perfect or near-perfect scores. The problem with a ceiling effect is that distinctions between high-scoring examinees are not possible, even though these examinees might differ substantially on the underlying trait measured by the test. A �loor effect is observed when signi�icant numbers of examinees obtain scores at or near the bottom of the scale. For example, the WAIS-R possessed a serious �loor effect in that it failed to discriminate between moderate, severe, and profound levels of mental retardation—all persons with signi�icant developmental disabilities would fail to answer virtually every question.
Test developers expect that some initial items will prove to make ineffectual contributions to the overall measurement goal of their instrument. For this reason, it is common practice to construct a �irst draft that contains excess items, perhaps double the number of questions desired on the �inal draft. For example, the 550-item MMPI originally consisted of more than 1,000 true–false personality statements (Hathaway & McKinley, 1940 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib715) ).
Table of Speci�ications Professional developers of achievement and ability tests often use one or more item-writing schemes to help ensure that their instrument taps a desired mixture of cognitive processes and content domains. For example, a very simple item-writing scheme might designate that an achievement test on the Civil War should consist of 10 multiple-choice items and 10 �ill-in-the-blank questions, half of each on factual matters (e.g., dates, major battles) and the other half on conceptual issues (e.g., differing views on slavery).
Before development of a test begins, item writers usually receive a table of speci�ications. A table of speci�ications (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss322) enumerates the information and cognitive tasks on which examinees are to be assessed. Perhaps the most common speci�ication table is the content-by-process matrix, which lists the exact number of items in relevant content areas and details the precise composite of items that must exemplify different cognitive processes (Millman & Greene, 1989 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1152) ).
Consider a science achievement test suitable for high school students. Such a test must cover many different content areas and should require a mixture of cognitive processes ranging from simple recall to inferential reasoning. By providing a table of speci�ications prior to the item-writing stage, the test developer can guarantee that the resulting instrument contains a proper balance of topical coverage and taps a desired range of cognitive skills. A hypothetical but realistic table of speci�ications is portrayed in Table 4.3 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec10#ch04tab3) .
TABLE 4.3 Example of a Content-by-Process Table of Speci�ications for a Hypothetical 100-Item Science Achievement Test
Content Area
Process
Factual Knowledgea (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec10#ch04fn4)
Information Competenceb (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec10#ch04fn5)(http://c
Astronomy 8 3 Botany 6 7 Chemistry 10 5 Geology 10 5 Physics 8 5 Zoology 8 5 Totals 50 30
aFactual Knowledge: Items can be answered based on simple recognition of basic facts. bInformation Competence: Items require usage of information provided in written text. cInferential Reasoning: Items can be answered by making deductions or drawing conclusions.
Item Formats When it comes to the method by which psychological attributes are to be assessed, the test developer is confronted with dozens of choices. Indeed, it would be easy to write an entire chapter on this topic alone. For reviews of item formats, the interested reader should consult Bausell (1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib110) ), Jensen (1980
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 24/70
(http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib832) ), and Wesman (1971 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1750) ). In this section, we will quickly survey the advantages and pitfalls of the more common varieties of test items.
For group-administered tests of intellect or achievement, the technique of choice is the multiple-choice question. For example, an item on an American history achievement test might include this combination of stem and options:
The president of the United States during the Civil War was
Washington Lincoln Hamilton Wilson
Proponents of multiple-choice methodology argue that properly constructed items can measure conceptual as well as factual knowledge. Multiple-choice tests also permit quick and objective machine scoring. Furthermore, the fairness of multiple-choice questions can be proved (or occasionally disproved!) with very simple item analysis procedures discussed subsequently. The major shortcomings of multiple-choice questions are, �irst, the dif�iculty of writing good distractor options and, second, the possibility that the presence of the response may cue a half-knowledgeable respondent to the correct answer. Guidelines for writing good multiple-choice items are listed in Table 4.4 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec10#ch04tab4) .
Matching questions are popular in classroom testing, but suffer serious psychometric shortcomings. An example of a matching question:
Using the letters on the left, match the name to the accomplishment:
A. Binet ____ translated a major intelligence test B. Woodworth ____ no correlation between grades and mental tests C. Cattell ____ developed true/false personality inventory
TABLE 4.4 Guidelines for Writing Multiple-Choice Items
Choose words that have precise meanings. Avoid complex or awkward word arrangements. Include all information needed for response selection. Put as much of the question as possible in the stem. Do not take stems verbatim from textbooks. Use options of equal length and parallel phrasing. Use “none of the above” and “all of the above” rarely. Minimize the use of negatives such as not. Avoid the use of nonfunctional words. Avoid unessential speci�icity in the stem. Avoid unnecessary clues to the correct response. Submit items to others for editorial scrutiny.
D. McKinley ____ battery of sensorimotor tests E. Wissler ____ developed �irst useful intelligence test F. Goddard ____ screening test for emotional disturbance
The most serious problem with matching questions is that responses are not independent—missing one match usually compels the examinee to miss another. Another problem is that the options in a matching question must be very closely related or the question will be too easy.
For individually administered tests, the procedure of choice is the short-answer objective item. Indeed, the simplest and most straightforward types of questions often possess the best reliability and validity. A case in point is the Vocabulary subtest from the WAIS-IV, which consists merely of asking the examinee to de�ine words. This subtest has very high reliability (.96) and is often considered the single best measure of overall intelligence on the test.
Personality tests often use true–false questions because they are easy for subjects to understand. Most people �ind it simple to answer true or false to items such as:
T F ____ ____ I like sports magazines.
Critics of this approach have pointed out that answers to such questions may re�lect social desirability rather than personality traits (Edwards, 1961). An alternative format designed to counteract this problem is the forced-choice methodology (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss122) in which the examinee must choose between two equally desirable (or undesirable) options:
Which would you rather do:
_____ Mop a gallon of syrup from the �loor.
_____ Volunteer for a half day at a nursing home.
Although the forced-choice approach has many desirable psychometric properties, personality test developers have not rushed to embrace this interesting methodology.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 25/70
4.11 TESTING THE ITEMS Psychometricians expect that numerous test items from the original tryout pool will be discarded or revised as test development proceeds. For this reason, test developers initially produce many, many excess items, perhaps double the number of items they intend to use. So, how is the �inal sample of test questions selected from the initial item pool? Test developers use item analysis, a family of statistical procedures, to identify the best items. In general, the purpose of item analysis is to determine which items should be retained, which revised, and which thrown out. In conducting a thorough item analysis, the test developer might make use of item-dif�iculty index, item-reliability index, item-validity index, item-characteristic curve, and an index of item discrimination. We turn now to a brief review of these statistical approaches to item analysis. Readers who wish an in-depth discussion and critique of these topics should consult Hambleton (1989 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib688) ) and Nunnally (1978 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1243) ).
Item-Dif�iculty Index The item dif�iculty for a single test item is de�ined as the proportion of examinees in a large tryout sample who get that item correct. For any individual item i, the index of item dif�iculty is pi, which varies from 0.0 to 1.0. An item with dif�iculty of .2 is more dif�icult than an item with dif�iculty of .7, because fewer examinees answered it correctly.
The item-dif�iculty index (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss171) is a useful tool for identifying items that should be altered or discarded. Suppose an item has a dif�iculty index near 0.0, meaning that nearly everyone has answered it incorrectly. Unfortunately, this item is psychometrically unproductive because it does not provide information about differences between examinees. For most applications, the item should be rewritten or thrown out. The same can be said for an item with a dif�iculty index near 1.0, where virtually all subjects provide a correct answer.
What is the optimal level of item dif�iculty? Generally, item dif�iculties that hover around .5, ranging between .3 and .7, maximize the information the test provides about differences between examinees. However, this rule of thumb is subject to one important quali�ication and one very signi�icant exception.
For true–false or multiple-choice items, the optimal level of item dif�iculty needs to be adjusted for the effects of guessing. For a true–false test, a dif�iculty level of .5 can result when examinees merely guess. Thus, the optimal item dif�iculty for such items would be .75 (halfway between .5 and 1.0). In general, the optimal level of item dif�iculty can be computed from the formula (1.0 + g)/2, where g is the chance success level. Thus, for a four-option multiple-choice item, the chance success level is .25, and the optimal level of item dif�iculty would be (1.0 + .25)/2, or about .63.
If a test is to be used for selection of an extreme group by means of a cutting score, it may be desirable to select items with dif�iculty levels outside the .3 to .7 range. For example, a test used to select graduate students for a university that admits only a select few of its many applicants should contain many very dif�icult items. A test used to designate children for a remedial-education program should contain many extremely easy items. In both cases, there will be useful discrimination among examinees near the cutting score—a very high score for the graduate admissions and a very low score for students eligible for remediation—but little discrimination among the remaining examinees (Allen & Yen, 1979 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib18) ).
Item-Reliability Index A test developer may desire an instrument with a high level of internal consistency in which the items are reasonably homogeneous. A simple way to determine whether an individual item “hangs together” with the remaining test items is to correlate scores on that item with scores on the total test. However, individual items are typically right or wrong (often scored 1 or 0), whereas total scores constitute a continuous variable. In order to correlate these two different kinds of scores it is necessary to use a special type of statistic called the point-biserial correlation coef�icient. The computational formula for this correlation coef�icient is equivalent to the Pearson r discussed earlier, and the point-biserial coef�icient conveys much the same kind of information regarding the relationship between two variables (one of which happens to be dichotomous and scored 0 or 1). In general, the higher the point-biserial correlation riT between an individual item and the total score, the more useful is the item from the standpoint of internal consistency.
The usefulness of an individual dichotomous test item is also determined by the extent to which scores on it are distributed between the two outcomes of 0 and 1. Although it sounds incongruous, it is possible to compute the standard deviation for dichotomous items; as with a continuously scored variable, the standard deviation of a dichotomous item indicates the extent of dispersion of the scores. If an individual item has a standard deviation of zero, everyone is obtaining the same score (all right or all wrong). The more closely the item approaches a 50–50 split of right and wrong scores, the greater is its standard deviation. In general, the greater the standard deviation of an item, the more useful is the item to the overall scale. Although we will not provide the derivation, it can be shown that the item-score standard deviation si for a dichotomously scored item can be computed from
We may summarize the discussion up to this point as follows: The potential value of a dichotomously scored test item depends jointly on its internal consistency as indexed by the correlation with the total score (riT) and also its variability as indexed by the standard deviation (si). If we compute the product of these two indices, we obtain siriT, which is the item-reliability index (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss174) . Consider the characteristics of an item that possesses a relatively large item-reliability index. Such an item must exhibit strong internal consistency and produce a good dispersion of scores between its two alternatives. The value of this index in test construction is simply this: By computing the item-reliability index for every item in the preliminary test, we can eliminate the “outlier” items that have the lowest value on this index. Such items would possess poor internal consistency or weak dispersion of scores and therefore not contribute to the goals of measurement.
Item-Validity Index For many applications, it is important that a test possess the highest possible concurrent or predictive validity. In these cases, one overriding question governs test construction: How well does each preliminary test item contribute to accurate prediction of the criterion? The item-validity index (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss177) is a useful tool in the psychometrician’s quest to identify predictively useful test items. By computing the item-validity index for every item in the preliminary test, the test developer can identify ineffectual items, eliminate or rewrite them, and produce a revised instrument with greater practical utility.
The �irst step in �iguring an item-validity index is to compute the point-biserial correlation between the item score and the score on the criterion variable. In general, the higher the point-biserial correlation riC between scores on an individual item and the criterion score, the more useful is the item from the standpoint
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 26/70
of predictive validity. As previously noted, the utility of an item also depends upon its standard deviation si. Thus, the item-validity index consists of the product of the standard deviation and the point-biserial correlation: siriC.
Item-Characteristic Curves Also known as an item response function, an item-characteristic curve (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss170) (ICC) is a graphical display of the relationship between the probability of a correct response and the examinee’s position on the underlying trait measured by the test. However, we do not have direct access to underlying traits, so observed test scores must be used to estimate trait quantities.
A separate ICC is graphed for each item, based upon a plot of the total test scores on the horizontal axis versus the proportion of examinees passing the item on the vertical axis (Figure 4.8 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04�ig8) ). An ICC is actually a mathematical idealization of the relationship between the probability of a correct response and the amount of the trait possessed by test respondents. Different ICC models use different mathematical functions based on initial assumptions. The simplest ICC model is the Rasch Model, based upon the item-response theory of the Danish mathematician Georg Rasch (1966). The Rasch Model is the simplest model because it makes just two assumptions: (1) test items are unidimensional and measure one common trait, and (2) test items vary on a continuum of dif�iculty level.
In general, a good item has a positive ICC slope. If the ability to solve a particular item is normally distributed, the ICC will resemble a normal ogive (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss221) (curve a in Figure 4.8 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04�ig8) ). The normal ogive is simply the normal distribution graphed in cumulative form.
The desired shape of the ICC depends on the purpose of the test. Psychometric purists would prefer that test item ICCs approximate the normal ogive, because this curve is convenient for making mathematical deductions about the underlying trait (Lord & Novick, 1968 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1006) ). However, for selection decisions based on cutoff scores, a step function is preferred. For example, when combined with other similar items, the item that produced curve b in Figure 4.8 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04�ig8) would be the best for selecting examinees with high levels of the measured trait.
FIGURE 4.8 Some Sample Item-Characteristic Curves
ICCs are especially useful for identifying items that perform differently for subgroups of examinees (Allen & Yen, 1979 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib18) ). For example, a test developer may discover that an item performs differently for men and women. A sex-biased question involving football facts comes to mind here. For men, the ICC for this item might have the desired positive slope, whereas for women the ICC might be quite �lat (such as curve c in Figure 4.8 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04�ig8) ). Items with ICCs that differ among subgroups of examinees can be revised or eliminated.
The underlying theory of ICC is also known as item response theory and latent trait theory. The usefulness of this approach has been questioned by Nunnally (1978 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1243) ), who points out that the assumption of test unidimensionality (implied in the ICC curve, which plots percentage passing against the unidimensional horizontal axis of trait value) is violated when many psychological tests are considered. If there were no serious technical and practical problems involved, “one wonders why ICC theory was not adopted long ago for the actual construction and scoring of tests” (Nunnally, 1978 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1243) ).
The merits of the ICC approach are still debated. ICC theory seems particularly appropriate for certain forms of computerized adaptive testing (CAT) in which each test taker responds to an individualized and unique set of items that are then scored on an underlying uniform scale (Weiss, 1983 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1744) ). The CAT approach to assessment would not be possible in the absence of an ICC approach to measurement. CAT is discussed in Topic 12B (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch12lev1sec5#ch12box3) , Computerized Assessment and the Future of Testing. Readers who wish a more detailed discussion of ICC and other latent trait models should consult Hambleton (1989 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib688) ) and Embretson and Reise (2000 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib459) ).
Item-Discrimination Index It should be clear from the discussion of ICCs that an effective test item is one that discriminates between high scorers and low scorers on the entire test. An ideal test item is one that most of the high scorers pass and most of the low scorers fail (see curve a in Figure 4.8 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04�ig8) ). Simple visual inspection of the ICC provides a coarse basis for gauging the discriminability of a test item: If the slope of the curve is positive and the curve is preferably ogive-shaped, the item is doing a good job of separating
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 27/70
high and low scorers. But visual inspection is not a completely objective procedure; what is needed is a statistical tool that summarizes the discrimination power of individual test items.
An item-discrimination index (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss172) is a statistical index of how ef�iciently an item discriminates between persons who obtain high and low scores on the entire test. There are many indices of item discrimination, including such indirect measures as riT, the point-biserial correlation between scores on an individual item and the total test score. However, we will restrict our discussion here to a direct measure, the item-discrimination index, symbolized by the lowercase, italicized letter d. On an item-by-item basis, this index compares the performance of subjects in the upper and lower regions of total test score. The upper and lower ranges are generally de�ined as the upper- and lower-scoring 10 percent to 33 percent of the sample. If the total test scores are normally distributed, the optimal comparison is the highest-scoring 27 percent versus the lowest-scoring 27 percent of the examinees. If the distribution of total test scores is �latter than the normal curve, the optimal percentage is larger, approaching 33 percent. For most applications, any percentage between 25 and 33 will yield similar estimates of d (Allen & Yen, 1979 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib18) ).
The item-discrimination index for a test item is calculated from the formula:
d = (U − L)/N
where U is the number of examinees in the upper range who answered the item correctly, L is the number of examinees in the lower range who answered the item correctly, and N is the total number of examinees in the upper or lower range.
Let us illustrate the computation and use of d with a hypothetical example. Suppose that a test developer has constructed the preliminary version of a multiple- choice achievement test and has administered the exam to a tryout sample of 400 high school students. After computing total scores for each subject, the test developer then identi�ies the high-scoring 25 percent and low-scoring 25 percent of the sample. Since there are 100 students in each group (25 percent of 400), N in the preceding formula will be 100. Next, for each item, the developer determines the number of students in the upper range and the lower range who answered it correctly. To compute d for each item is a simple matter of plugging these values into the formula (U − L)/N. For example, suppose on the �irst item that 49 students in the upper range answered it correctly, whereas 23 students in the lower range answered it correctly. For this item, d is equal to (49 − 23)/100 or .26.
It is evident from the formula for d that this index can vary from −1.0 to +1.0. Notice, too, that a negative value for d is a warning signal that a test item needs revision or replacement. After all, such an outcome indicates that more of the low-scoring subjects answered the item correctly than did the high-scoring subjects. If d is zero, exactly equal numbers of low- and high-scoring subjects answered the item correctly; since the item is not discriminating between low- and high-scoring subjects at all, it should be revised or eliminated. A positive value for d is preferred, and the closer to +1.0 the better. Table 4.5 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/ch04lev1sec11#ch04tab5) illustrates item-discrimination indices for six items from the hypothetical test proposed here.
A test developer can supplement the item-discrimination approach by inspecting the number of examinees in the upper- and lower-scoring groups who choose each of the incorrect alternatives. If a multiple-choice item is well written, the incorrect alternatives should be equally attractive to subjects who do not know the correct answer. Of course, we expect that high-scoring examinees will choose the correct alternative more often than low-scoring examinees—that is the purpose in computing item-discrimination indices. But, in addition, a good item should show proportional dispersion of incorrect choices for both high- and low- scoring subjects.
Assume that we investigate the choices of 100 high-scoring and 100 low-scoring subjects on a hypothetical multiple-choice test. Correct choices are indicated by an asterisk (*). Item 1 demonstrates the desired pattern of answers, with incorrect choices about equally dispersed.
Alternatives Item 1 a b c* d e High Scorers 5 6 80 5 4 Low Scorers 15 14 40 16 15
On item 2, we notice that no examinees picked alternative d. This alternative should be replaced with a more appealing distractor:
Item 2 a b* c d e High Scorers 5 75 10 0 10 Low Scorers 21 34 20 0 25
Item 3 is probably a poor item in spite of the fact that it discriminates effectively between high- and low-scoring subjects. The obvious problem is that high- scoring examinees prefer alternative a to the correct alternative, d:
Item 3 a b c* d e High Scorers 43 6 5 37 9 Low Scorers 20 19 22 10 29
TABLE 4.5 Item-Discrimination Indices for Six Hypothetical Items
Item U L (U − L)/N Comment 1 49 23 .26 Very good item with high dif�iculty 2 79 19 .60 Excellent item but rarely achieved 3 52 52 .00 Poor item that should be revised 4 100 0 1.00 Ideal item but never achieved 5 20 80 −.60 Terrible item that should be eliminated 6 0 100 −1.00 Theoretically worst possible item
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 28/70
Perhaps by rewriting alternative a, this item could be rescued. In any case, the main point here is that test developers should pry into every corner of every test item by every means possible, including visual inspection of the pattern of answers.
Reprise: The Best Items From all the methods of item analysis previously portrayed, which ones should the test developer use to identify the best items for a test? The answer to this question is neither simple nor straightforward. After all, the choice of “best” items depends on the objectives of the test developer. For example, a theoretically inclined research psychologist might desire a measurement instrument with the highest possible internal consistency; item-reliability indices are crucial to this goal. A practically minded college administrator might wish for an instrument with the highest possible criterion validity; item-validity indices would be useful for this purpose. A remediation-oriented mental retardation specialist might desire an intelligence test with minimal �loor effect; item-dif�iculty indices would be helpful in this regard. In sum, there is no single preferred method for item selection ideally suited to every context of assessment and test development.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 29/70
4.12 REVISING THE TEST The purpose of item analysis, discussed previously, is to identify unproductive items in the preliminary test so that they can be revised, eliminated, or replaced. Very few tests emerge from this process unscathed. It is common in the evolutionary process of test development that many items are dropped, others re�ined, and new items added. The initial repercussion is that a new and slightly different test emerges. This revised test likely contains more discriminating items with higher reliability and greater predictive accuracy—but these improvements are known to be true only for the �irst tryout sample.
The next step in test development is to collect new data from a second tryout sample. Of course, these examinees should be similar to those for whom the test is ultimately intended. The purpose of collecting additional test data is to repeat the item analysis procedures anew. If further changes are of the minor �ine-tuning variety, the test developer may decide the test is satisfactory and ready for cross-validational study, discussed in the following section. If major changes are needed, it is desirable to collect data from a third and even perhaps a fourth tryout sample. But at some point, psychometric tinkering must end; the developer must propose a �inalized instrument and proceed to the next step, cross validation.
Cross Validation When a tryout sample is used to ascertain that a test possesses criterion-related validity, the evidence is quite preliminary and tentative. It is prudent practice in test development to seek fresh and independent con�irmation of test validity before proceeding to publication. The term cross validation (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss85) refers to the practice of using the original regression equation in a new sample to determine whether the test predicts the criterion as well as it did in the original sample. Ghiselli, Campbell, and Zedeck (1981 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib585) ) outline the rationale for cross validation:
Whether items are chosen on the basis of empirical keying or whether they are corrected or weighted, the obtained results should, unless additional data are collected, be viewed as speci�ic to the sample used for the statistical analyses. This is necessary because the obtained results have likely capitalized on chance factors operating in that group and therefore are applicable only to the sample studied.
Validity Shrinkage A common discovery in cross-validation research is that a test predicts the relevant criterion less accurately with the new sample of examinees than with the original tryout sample. The term validity shrinkage (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss343) is applied to this phenomenon. For example, a biographically based predictor of sales potential might perform quite well for the sample of subjects used to develop the instrument but demonstrate less validity when applied to a new group of examinees. Mitchell and Klimoski (1986 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1166) ) studied validity shrinkage of an instrument designed to foretell which students will succeed in real estate, as measured by the real-world criterion of obtaining a real estate license two years later. In one analysis based on the sample used to derive the test, the biographically based predictor test correlated .6 with the criterion. But when this same test was tried out on a new sample of real estate students, the correlation with the criterion was lower, about .4, demonstrating typical validity shrinkage.
Validity shrinkage is an inevitable part of test development and underscores the need for cross validation. In most cases, shrinkage is slight and the instrument withstands the challenge of cross validation. However, shrinkage of test validity can be a major problem when derivation and cross-validation samples are small, the number of potential test items is large, and items are chosen on a purely empirical basis without theoretical rationale.
A classic paper by Cureton (1950 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib382) ) demonstrates a worst-case scenario: using a very small sample to select empirically keyed items from a large item pool, then validating the test on the same sample. The criterion in his study was grade point average, arti�icially dichotomized into grades of B or better and grades below B. His “test” items consisted of 85 tags, numbered on one side. For each of 29 students, the tags were shaken in a container and dropped on the table. All tags that fell with numbers up were recorded as indicating the presence of that “item” for the student. Next, Cureton conducted an item analysis, using the dichotomized grades as the criterion. Based on this analysis, 24 items were found to be maximally predictive of students’ grades. Nine items occurred more often among students with the higher grades, and these items were weighted +1. Fifteen items occurred more often among students with the lower grades, and these items were weighted −1. The score on this test (facetiously named the “B-Projective Psychokinesis Test”) consisted of the sum of these 24 item weights.
In spite of the nonsensical nature of his test, Cureton (1950 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib382) ) found that test scores correlated .82 with grades. Of course, the strength of this correlation was due entirely to capitalization upon chance. If we were to conduct a series of cross-validation studies using new samples of students, the correlation between the B-Projective Psychokinesis Test and grades would likely hover right around zero, because this test is completely devoid of predictive validity. There is an important lesson here that applies to serious tests as well: Demonstrate validity through cross validation, do not assume it based merely on the solemn intentions of a new instrument.
Feedback from Examinees In test revision, feedback from examinees is a potentially valuable source of information that is normally overlooked by test developers. We can illustrate this approach with research by Nevo (1992 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1230) ). He developed the Examinee Feedback Questionnaire (EFeQ) to study the Inter-University Psychometric Entrance Examination, a major requirement for admission to the six universities in Israel. The Inter-University entrance exam is a group test consisting of �ive multiple-choice subtests: General Knowledge, Figural Reasoning, Comprehension, Mathematical Reasoning, and English. The EFeQ was designed as an anonymous posttest administered immediately after the Inter-University entrance exam.
The EFeQ is a short and simple questionnaire designed to elicit candid opinions from examinees as to these features of the test–examiner–respondent matrix:
Behavior of examiners Testing conditions Clarity of exam instructions Convenience in using the answer sheet Perceived suitability of the test Perceived cultural fairness of the test Perceived suf�iciency of time Perceived dif�iculty of the test Emotional response to the test Level of guessing Cheating by the examinee or others
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 30/70
The �inal question on the EFeQ is an open-ended essay: “We are interested in any remarks or suggestions you might have for improving the exam.”
Nevo (1992 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1230) ) determined that the EFeQ questionnaire possesses modest reliability, with a test–retest reliability of about .70. Regardless of the psychometric properties of his scale, the tradition of asking examinees for feedback about tests has proved invaluable. The Inter-University entrance exam was modi�ied in numerous ways in response to feedback: The answer sheet format was modi�ied in ways suggested by examinees; the time limit was increased for speci�ic tests reported to be too speeded; certain items perceived as culturally biased or unfair were deleted. In addition, security measures were revised and tightened in order to minimize cheating, which was much more prevalent than examiners had anticipated. Nevo (1992 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib1230) ) also cites a hidden advantage to feedback questionnaires: They convey the message that someone cares enough to listen, which reduces postexamination stress. Examinee feedback questionnaires should become a routine practice in group standardized testing.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 31/70
4.13 PUBLISHING THE TEST The test construction process does not end with the collection of cross-validation data. The test developer also must oversee the production of the testing materials, publish a technical manual, and produce a user’s manual. A number of relevant guidelines can be offered for each of these �inal steps, as outlined in the following sections. Finally, we close this chapter with a provocative comment on the conservatism of modern test publishers.
Production of Testing Materials Testing materials must be user friendly if they are to receive wide acceptance by psychologists and educators. Thus, a �irst guideline for test production is that the physical packaging of test materials must allow for quick and smooth administration. Consider the challenge posed by some performance tests, in which the examiner must wrestle with pencil, clipboard, test form, stopwatch, test manual, item shield, item box, and a disassembled cardboard object, all the while maintaining conversation with the examinee. If it is possible for the test developer to simplify the duties of the examiner while leaving examinee task demands unchanged, the resulting instrument will have much greater acceptability to potential users. For example, if the administration instructions can be summarized on the test form, the examiner can put the test manual aside while setting out the task for the examinee. Another welcome addition to psychological test packaging is the stand-up ring binder that shows the test question on the side facing the examinee and provides instructions for administration on the reverse side facing the examiner.
Technical Manual and User’s Manual Technical data about a new instrument are usually summarized with appropriate references in a technical manual (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss323) . Here, the prospective user can �ind information about item analyses, scale reliabilities, cross-validation studies, and the like. In some cases, this information is incorporated in the user’s manual (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm01#bm01gloss340) , which gives instructions for administration and also provides guidelines for test interpretation.
Test manuals should communicate information to many different groups ranging in background and training from measurement specialist to classroom teacher. Test manuals serve many purposes, as outlined in the Standards for Educational and Psychological Testing (AERA, APA, & NCME, 1985 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib29) , 1999 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib30) ). The in�luential Standards manual suggests that test manuals accomplish the following goals:
Describe the rationale and recommended uses for the test Provide speci�ic cautions against anticipated misuses of a test Cite representative studies regarding general and speci�ic test uses Identify special quali�ications needed to administer and interpret the test Provide revisions, ammendations, and supplements as needed Use promotional material that is accurate and research based Cite quantitative relationships between test scores and criteria Report on the degree to which alternative modes of response (e.g., booklet versus an answer sheet) are interchangeable Provide appropriate interpretive aids to the test taker Furnish evidence of the validity of any automated test interpretations
Finally, test manuals should provide the essential data on reliability and validity rather than referring the user to other sources—an unfortunate practice encountered in some test manuals.
Testing Is Big Business By now the reader should appreciate the intimidating task faced by anyone who sets out to develop and publish a new test. Aside from the gargantuan proportions of the endeavor, test development is extraordinarily expensive, which means that publishers are inherently conservative about introducing new tests. Jensen (1980 (http://content.thuzelearning.com/books/Gregory.8055.17.1/sections/bm02#bm02bib832) ) provides the following provocative view on this topic:
To produce a new general intelligence test that would be a really signi�icant improvement over existing instruments would be a multimillion-dollar project requiring a large staff of test construction experts working for several years. Today we possess the necessary psychometric technology for producing considerably better tests than are now in popular use. The principal hindrances are copyright laws, vested interests of test publishers in the established tests in which they have already made enormous investments, and the market economy for tests. Signi�icant improvement of tests is not an attractive commercial venture initially and would probably have to depend on large-scale and long-term subsidies from government agencies and private foundations.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 32/70
REFERENCES
Aamodt, M. G., Keller, R., Crawford, K., & Kimbrough, W. (1981). A critical incident job analysis of the university housing resident assistant position. Psychological Reports, 49, 983–986.
Abel, E. L. (1995). An update on incidence of FAS: FAS is not an equal opportunity birth defect. Neurobehavioral Toxicology, 17, 437–443. Abel, E. L. (2009). Fetal alcohol syndrome: Same old, same old. Addiction, 104, 1274–1275. Abell, S. C., Briesen, P., & Watz, L. (1996). Intellectual evaluations of children using human �igure drawings: An empirical investigation of two methods. Journal of
Clinical Psychology, 52, 67–74. Achenbach, T. M. (1991). Manual for the Teacher’s Report Form and 1991 Pro�ile. Burlington: University of Vermont, Department of Psychiatry. Achenbach, T. M. (1992). Manual for the Child Behavior Checklist/2–3 and 1992 Pro�ile. Burlington: University of Vermont, Department of Psychiatry. Achenbach, T. M., & Rescorla, L. A. (2000). Manual for the ASEBA preschool forms and pro�iles. Burlington: University of Vermont, Research Center for Children,
Youth, and Families. Adams, G. A., Elacqua, T., & Colarelli, S. (1994). The employment interview as a sociometric selection technique. Journal of Group Psychotherapy, Psychodrama,
and Sociometry, 47 (Fall), 99–113. Adams, K. M., & Heaton, R. K. (1985). Automated interpretation of neuropsychological test data. Journal of Consulting and Clinical Psychology, 53, 790–802. Agbenyega, S., & Jiggetts, J. (1999). Minority children and their over-representation in special education. Education, 119, 619–633. Aguinis, A., Culpepper, S. A., & Pierce, C. A. (2010). Revival of test bias research in preemployment testing. Journal of Applied Psychology, 95, 648–680. Ahrens, J., Evans, R., & Barnett, R. (1990). Factors related to dropping out of school in an incarcerated population. Educational and Psychological Measurement,
50, 611–617. Aiken, L. R. (1989). Assessment of personality. Boston: Allyn and Bacon. Ainsworth, M., & Bowlby, J. (1965). Child care and the growth of love. London: Penguin Books. Albers, C., & Grieve, A. (2007). Test Review: Bayley, N. (2006). Bayley Scales of Infant and Toddler Development—Third Edition. San Antonio, TX—Harcourt
Assessment. Journal of Psychoeducational Assessment, 25, 180–190. Albert, S., Fox, H. M., & Kahn, M. W. (1980). Faking psychosis on the Rorschach: Can expert judges detect malingering? Journal of Personality Assessment, 44, 115–
119. Alkhadher, O., Clarke, D., & Anderson, N. (1998). Equivalence and predictive validity of paper-and-pencil and computerized adaptive formats of the Differential
Aptitude Tests. Journal of Occupational and Organizational Psychology, 71, 205–217. Allen, M. J., & Yen, W. M. (1979). Introduction to measurement theory. Monterey, CA: Brooks/Cole. Allport, G. W. (1937). Personality: A psychological interpretation. New York: Holt, Rinehart and Winston. Allport, G. W. (1950). The individual and his religion. New York: Macmillan. Allport, G. W., & Odbert, H. (1936). Trait names, a psycholexical study. Psychological Monographs, 47 (Whole No. 211). Allport, G. W., & Ross, J. (1967). Personal religious orientation and prejudice. Journal of Personality and Social Psychology, 5, 432–443. Altepeter, T. S. (1989). The PPVT-R as a measure of psycholinguistic functioning: A caution. Journal of Clinical Psychology, 45, 935–941. Altepeter, T. S., & Johnson, K. A. (1989). Use of the PPVT-R for intellectual screening with adults: A caution. Journal of Psychoeducational Assessment, 7, 39–45. Alzheimer’s Disease and Related Disorders Association. (2000). General statistics/demographics. Chicago: Author. Amabile, T. M. (1983). The social psychology of creativity. New York: Springer-Verlag. Ambrosini, P. J. (2000). Historical development and present status of the schedule for affect disorders and schizophrenia for school-age children (K-SADS).
Journal of the American Academy of Child and Adolescent Psychiatry, 39, 49–58. American Association for Counseling and Development. (1988). Ethical standards. Washington, DC: Author. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (1985). Standards for
educational and psychological testing. Washington, DC: American Psychological Association. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (1999). Standards for
educational and psychological testing (2nd ed.). Washington, DC: American Psychological Association. American Federation of Teachers, National Council on Measurement in Education, & National Education Association. (1990). Standards for teacher competence in
educational assessment of students. Washington, DC: Author. American Psychiatric Association. (1994). Diagnostic and statistical manual of mental disorders (4th ed.). Washington, DC: Author. American Psychiatric Association. (2000). Diagnostic and statistical manual of mental disorders (4th ed., text revision). Washington, DC: Author. American Psychological Association. (1953). Ethical standards of psychologists. Washington, DC: Author. American Psychological Association. (1986). Guidelines for computer-based tests and interpretations. Washington, DC: Author. American Psychological Association. (1988). In the Supreme Court of the United States: Clara Watson v. Fort Worth Bank & Trust. American Psychologist, 43,
1019–1028. American Psychological Association. (1992a). Ethical principles of psychologists and code of conduct. American Psychologist, 47, 1597–1611. American Psychological Association. (1992b). Psychological testing of language minority and culturally different children. Washington, DC: Author. American Psychological Association. (1993). Guidelines for providers of psychological services to ethnic, linguistic, and culturally diverse populations. American
Psychologist, 48, 45–48. American Psychological Association. (1994). Report of the ethics committee, 1993. American Psychologist, 49, 659–666. American Psychological Association. (2002). Ethical principles of psychologists and code of conduct. American Psychologist, 57, 1060–1073. American Psychological Association. (2012). Specialty guidelines for forensic psychology. American Psychologist, 68, 7–19. American Speech-Language-Hearing Association. (1991). Code of ethics of the American Speech-Language Hearing Association. Rockville, MD: Author. Ammer, C. (2003). The American Heritage dictionary of idioms. New York: Houghton Mif�lin Harcourt. Anastasi, A. (1975). Review of the Goodenough-Harris Drawing Test. The seventh mental measurements yearbook. Lincoln: University of Nebraska Press. Anastasi, A. (1985). Psychological testing (6th ed.). New York: Macmillan.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 33/70
Anastasi, A. (1986). Emerging concepts of test validation. Annual Review of Psychology, 37, 1–15. Anastasi, A. (1988). Psychological testing (6th ed.). New York: Macmillan. Andersen, P., & Vandehey, M. A. (2011). Career counseling and development in a global economy (2nd ed.). Belmont, CA: Cengage Learning. Andersson, H. W. (1996). The Fagan Test of Infant Intelligence: Predictive validity in a random sample. Psychological Reports, 78, 1015–1026. Andreasen, N. (2001). Brave new brain: Conquering mental illness in the era of the genome. New York: Oxford University Press. Andreasen, N. C., & Black, D. (1995). Introductory textbook of psychiatry (2nd ed.). Washington, DC: American Psychiatric Press. Andrew, D. M., Peterson, D. G., & Longstaff, H. P. (1979). Minnesota Clerical Test Manual. San Antonio, TX: The Psychological Corporation. Andrews, F. M. (1975). Social and psychological factors which in�luence the creative process. In I. A. Taylor & J. W. Getzels (Eds.), Perspectives in creativity.
Chicago: Aldine. Ansorge, C. J. (1985). Review of the Cognitive Abilities Test. Ninth mental measurements yearbook. Lincoln: University of Nebraska Press. Anstey, K. J., Jorm, A. F., Reglade-Méslin, C., & others. (2007). Weekly alcohol consumption, brain atrophy, and white matter hyperintensities in a community-
based sample aged 60 to 64 years. Psychosomatic Medicine, 68, 778–785. Anthony, J. C., LeResche, L., Niaz, U., Von Korff, M., & Folstein, M. (1982). Limits of the Mini-Mental State as a screening test for dementia and delirium among
hospital patients. Psychological Medicine, 12, 397–408. Anthony, J., & Assel, M. (2007). A �irst look at the validity of the DIAL-3 Spanish version. Journal of Psychoeducational Assessment, 25, 165–179. APA Task Force. (2006). Evidence-based practice in psychology. American Psychologist, 61, 271–285. Arizona Senate Research Staff. (2008, August 27). Arizona State Senate issue brief: AIMS (Arizona Instrument to Measure Standards). Phoeniz, AZ: Author. Arnau, R. C., Meagher, M. W., Norris, M. P., & Bramson, R. (2001). Psychometric evaluation of the Beck Depression Inventory-II with primary care medical patients.
Health Psychology, 20, 112–119. Arvey, R. D., & Campion, J. E. (1982). The employment interview: A summary and review of recent research. Personnel Psychology, 35, 281–332. Arvey, R. D., & Faley, R. H. (1988). Fairness in selecting employees. Reading, MA: Addison-Wesley. Arvey, R. D., & Murphy, K. R. (1998). Performance evaluation in work settings. Annual Review of Psychology, 49, 141–168. Asher, J. J., & Sciarrino, J. A. (1974). Realistic work samples: A review. Personnel Psychology, 27, 519–533. Assel, M., & Anthony, J. (2009). Factor structure of the DIAL-3: A test of a theory-driven conceptualization versus an empirically driven conceptualization in a
nationally representative sample. Journal of Psychoeducational Assessment, 27, 113–124. Atkins v. Virginia. (2002). U.S. Supreme Court Cases, 536, 304–354. Retrieved from http://docs.justia.com/cases/supreme/536/304.pdf
(http://docs.justia.com/cases/supreme/536/304.pdf)
Atkinson, L., Bevc, I., Dickens, S., & Blackwell, J. (1992). Concurrent validities of the Stanford-Binet (Fourth Edition), Leiter, and Vineland with developmentally delayed children. Journal of School Psychology, 30, 165–173.
Austin, J. T., & Villanova, P. (1992). The criterion problem: 1917–1992. Journal of Applied Psychology, 77, 836–874. Axelrod, B. N., Greve, K., & Goldman, R. (1994). Comparison of four Wisconsin Card Sorting Test Scoring guides with novice raters. Assessment, 1, 115–121. Aylward, G. P., & Carson, A. (2005, April 1). Use of the Test Observation Checklist with the Stanford-Binet Intelligence Scales for Early Childhood, Fifth Edition (Early
SB5). Paper presented at the National Association of School Psychologists, Atlanta, GA. Bach, P. J., Harowski, K., Kirby, K., Peterson, P., & Schulein, M. (1981). The interrater reliability of the Luria-Nebraska Neuropsychological Battery. Clinical
Neuropsychology, 3, 19–21. Baddeley, A. (1986). Working memory. Oxford: Clarendon Press/Oxford University Press. Baer, D. M., Harrison, R., Fradenburg, L., Petersen, D., & Milla, S. (2005). Some pragmatics in the valid and reliable recording of directly observed behavior.
Research on Social Work Practice, 15, 440–451. Bagby, R. M., Rogers, R., Buis, T., & Kalemba, V. (1994). Malingered and defensive response styles on the MMPI-2: An examination of validity scales. Assessment, 1,
31–38. Bailey, D., Larson, L., Borgen, F., & Gasser, C. (2008). Changing of the guard: Interpretive continuity of the 2005 Strong Interest Inventory. Journal of Career
Assessment, 16, 135–155. Baker, C., Koenig, A., & Sowell, V. (1995). Relationship of the Blind Learning Aptitude Test to Braille reading skills. Journal of Visual Impairment & Blindness, 89,
440–447. Baker, F. B. (2001). The basics of item response theory (2nd ed.). College Park, MD: ERIC Clearing House on Assessment and Evaluation. Balboni, G., Pedrabissi, L., Molteni, M., & Villa, S. (2001). Discriminant validity of the Vineland Scales: Score pro�iles of individuals with mental retardation and a
speci�ic disorder. American Journal of Mental Retardation, 106, 162–172. Ballard, J., & Zettel, J. (1977). Public Law 94-142 and Sec. 504: What they say about rights and protections. Exceptional Children, 44, 177–185. Baltes, P. B., Reese, H., & Nesselroade, J. (1977). Life-span developmental psychology: Introduction to research methods. Belmont, CA: Wadsworth. Bandura, A. (1965). Vicarious processes: A case of no-trial learning. In L. Berkowitz (Ed.), Advances in experimental social psychology (vol. 2). New York:
Academic Press. Bandura, A. (1971). Social learning theory. Morristown, NJ: General Learning Press. Bandura, A. (1977). Social learning. Englewood Cliffs, NJ: Prentice Hall. Bandura, A. (1982). Self-ef�icacy mechanism in human agency. American Psychologist, 37, 122–147. Bandura, A. (1997). Self-ef�icacy: The exercise of control. New York: Freeman. Bandura, A. (2006). Guide for constructing self-ef�icacy scales. In T. Urdan & F. Pajares (Eds.), Self-ef�icacy beliefs of adolescents (pp. 307–337). Greenwich, CT:
Information Age Publishing. Bandura, A., & Walters, R. H. (1963). Social learning and personality development. New York: Holt, Rinehart and Winston. Barber, M., & Stott, D. (2004). Validity of the Telephone Interview for Cognitive Status (TICS) in post-stroke subjects. International Journal of Geriatric Psychiatry,
19, 75–79. Barkley, R. A. (1996). Attention-de�icit/hyperactivity disorder. In E. J. Mash & R. A. Barkley (Eds.), Child psychopathology (pp. 63–112). New York: Guilford. Barlow, D. (2005). What’s new about evidence-based assessment? Psychological Assessment, 17, 308–311. Barnett, W. S., & Camilli, G. (2002). Compensatory pre-school education, cognitive development, and “race.” In J. Fish (Ed.), Race and intelligence: Separating
science from myth. Mahwah, NJ: Erlbaum. Bar-On, R. (1997). Bar-On Emotional Quotient Inventory: Technical manual (EQ-i). Toronto, Canada: Multi-Health Systems.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 34/70
Bar-On, R. (2000). Emotional and social intelligence: Insights from the Emotional Quotient Inventory (EQ-i). In R. Bar-On & J. Parker (Eds.), Handbook of emotional intelligence (pp. 363–388). San Francisco: Jossey-Bass.
Bar-On, R., & Parker, J. D. (2000). Bar-On Emotional Quotient Inventory: Youth version. North Tonawanda, NY: Multi-Health Systems Incorporated. Barrett, L. F. (2009). The future of psychology: Connecting mind to brain. Perspectives on Psychological Science, 4, 326–339. Barrett, P. K. (2000). Validation of the Test of Nonverbal Intelligence-Third Edition (TONI-3) for Jamaican students. Unpublished Doctoral Dissertation, Auburn
University, Auburn, AL. Barrick, M. R., Swider, B. W., & Stewart, G. L. (2010). Initial evaluations in the interview: Relationships with subsequent interviewer evaluations and employment
offers. Journal of Applied Psychology, 95(6), 1163–1172. Barron, F. (1953). An ego-strength scale which predicts response to psychotherapy. Journal of Consulting Psychology, 17, 327–333. Barron, F. (1955). The disposition toward originality. Journal of Abnormal and Social Psychology, 51, 478–485. Barron, F. (1968). Creativity and personal freedom. Princeton, NJ: Van Nostrand. Barron, F., & Harrington, D. M. (1981). Creativity, intelligence, and personality. Annual Review of Psychology, 32, 439–476. Barry, A. E. (2005). How attrition impacts the internal and external validity of longitudinal research. Journal of School Health, 75, 267–270. Bartol, C., & Bartol, A. (2004). Introduction to forensic psychology: Research and application. Thousand Oaks, CA: Sage. Bartsch, A. J., Homola, G., Biller, A., & others. (2007). Manifestations of early brain recovery associated with abstinence from alcoholism. Brain, 130, 36–47. Bate, A., Mathias, J., & Crawford, J. (2001). Performance on the Test of Everyday Attention and standard tests of attention following severe traumatic brain injury.
Clinical Neuropsychologist, 15, 405–422. Batey, M. (2007). A psychometric investigation of everyday creativity. Unpublished doctoral dissertation, University College, London. Batey, M., & Furnham, A. (2006). Creativity, intelligence, and personality: A critical review of the scattered literature. Genetic, Social, and General Psychology
Monographs, 132, 355–429. Batson, C. D., Schoenrade, P., & Ventis, W. (1993). Religion and the individual: A social-psychological perspective. New York: Oxford University Press. Bausell, R. B. (1986). A practical guide to conducting empirical research. New York: Harper & Row. Bayless, J. D., Varney, N. R., & Roberts, R. J. (1989). Tinker Toy Test performance and vocational outcome in patients with closed-head injuries. Journal of Clinical
and Experimental Neuropsychology, 11, 913–917. Bayley, N. (1969). Bayley Scales of Infant Development. San Antonio, TX: The Psychological Corporation. Bayley, N. (2006). Bayley Scales of Infant and Toddler Development—Third Edition. San Antonio, TX: Harcourt Assessment. Beck, A. T. (1976). Cognitive therapy and the emotional disorders. New York: New American Library. Beck, A. T. (1983). Negative cognitions. In E. Levitt, B. Lubin, & J. Brooks (Eds.), Depression: Concepts, controversies, and some new facts (2nd ed.). Hillsdale, NJ:
Erlbaum. Beck, A. T. (1987). Cognitive models of depression. Journal of Cognitive Psychotherapy: An International Quarterly, 1, 5–37. Beck, A. T., & Steer, R. A. (1987). Manual for the revised Beck Depression Inventory. San Antonio, TX: The Psychological Corporation. Beck, A. T., Steer, R. A., & Brown, G. K. (1996). Manual for the Beck Depression Inventory-II. San Antonio, TX: The Psychological Corporation. Beck, A. T., Steer, R. A., & Garbin, M. G. (1988). Psychometric properties of the Beck Depression Inventory: Twenty-�ive years of evaluation. Clinical Psychology
Review, 8, 77–100. Beck, A. T., Ward, C. H., Mendelsohn, M., Mock, J., & Erbaugh, J. (1961). An inventory for measuring depression. Archives of General Psychiatry, 4, 561–571. Behling, O. (1998). Employee selection: Will intelligence and conscientiousness do the job? Academy of Management Executive, 12, 77–86. Beirne-Smith, M., Ittenbach, R. F., & Patton, J. R. (2002). Mental retardation (6th ed.). Upper Saddle River, NJ: Merrill (Prentice Hall). Belcher, M. J. (1992). Review of the Wonderlic Personnel Test. The eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Bell, L., & Casebourne, J. (2008). Increasing employment for ethnic minorities: A survey of research �indings. London: Center for Economic and Social Inclusion. Bell, N., Lassiter, K., Matthews, T., & Hutchinson, M. (2001). Comparison of the Peabody Picture Vocabulary Test-Third Edition and Wechsler Adult Intelligence
Scale-Third Edition with university students. Journal of Clinical Psychology, 57, 417–422. Bell, N., Matthews, T., Lassister, K., & Leverett, J. (2002). Validity of the Wonderlic Personnel Test as a measure of �luid or crystallized intelligence: Implications for
career assessment. North American Journal of Psychology, 4, 113–120. Bellak, L. (1992). The Thematic Apperception Test, the Children’s Apperception Test, and the Senior Apperception Technique in clinical use (5th ed.). Orlando, FL:
Grune & Stratton. Bellak, L., & Bellak, S. S. (1991). Children’s Apperception Test Manual (CAT) (8th rev. ed.). Larchmont, NY: C. P. S. Bellak, L., & Bellak, S. S. (1994). Children’s Apperception Test Human Figures (CAT-H) (11th ed.). Larchmont, NY: C. P. S. Belsky, J., & Pluess, M. (2009). The nature (and nurture?) of plasticity in early human development. Perspectives on Psychological Science, 4, 345–351. Bem, D., & Funder, D. (1978). Predicting more of the people more of the time: Assessing the personality of situations. Psychological Review, 85, 485–501. Bender, L. (1938). A visual motor gestalt test and its clinical use. New York: American Orthopsychiatric Association. Bennett, G. K., Seashore, H. G., & Wesman, A. G. (1974). Fifth edition manual for the Differential Aptitude Tests, Forms S and T. San Antonio, TX: The Psychological
Corporation. Bennett, G. K., Seashore, H. G., & Wesman, A. G. (1982). Differential Aptitude Tests: Administrator’s handbook. San Antonio, TX: The Psychological Corporation. Bennett, G. K., Seashore, H. G., & Wesman, A. G. (1984). Differential Aptitude Tests: Technical Supplement. San Antonio, TX: The Psychological Corporation. Bennett, T. (1988). Use of the Halstead-Reitan Neuropsychological Test Battery in the assessment of head injury. Cognitive Rehabilitation, 6, 18–25. Ben-Porath, Y. S., & Butcher, J. N. (1989). Psychometric stability of rewritten MMPI items. Journal of Personality Assessment, 53, 645–653. Ben-Porath, Y. S., & Tellegen, A. (2008). MMPI-2-RF (Minnesota Multiphasic Personality Inventory-2 Restructured Form): Manual for administration, scoring, and
interpretation. Minneapolis: University of Minnesota Press. Benson, D. F. (1988). Disorders of visual gnosis. In J. W. Brown (Ed.), The neuropsychology of visual perception. Hillsdale, NJ: Erlbaum. Benson, D. F. (1994). The neurology of thinking. New York: Oxford University Press. Benson, P. G. (1985). Minnesota Importance Questionnaire. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques (vol. 2). Kansas City, MO: Test Corporation of
America. Benson, P., Donahue, M., & Erickson, J. (1993). The Faith Maturity Scale: Conceptualization, measurement, and empirical validation. In M. L. Lynn & D. O. Moberg
(Eds.), Research in the social scienti�ic study of religion (vol. 5). Greenwich, CN: JAI Press. Benton, A., Hamsher, K., Rey, G., & Sivan, A. (1994). Multilingual Aphasia Examination (3rd ed.). Iowa City, IA: AJA Associates.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 35/70
Benton, A., Sivan, A., Hamsher, K., Varney, N., & Spreen, O. (1994). Contributions to neuropsychological assessment (2nd ed.). New York: Oxford University Press. Beran, T. (2007). Differential Ability Scales (2nd ed.). Canadian Journal of School Psychology, 22, 128–132. Berg, E. A. (1948). A simple objective test for measuring �lexibility in thinking. Journal of General Psychology, 39, 15–22. Berger, S. G., Chibnall, J., & Gfeller, J. (1994). The Category Test: A comparison of computerized and standard versions. Assessment, 3, 255–258. Berk, R. A. (Ed.). (1984). A guide to criterion-referenced test construction. Baltimore: Johns Hopkins University Press. Bernreuter, R. G. (1931). The personality inventory. Stanford, CA: Stanford University Press. Bernstein, D. M., & Loftus, E. F. (2009). How to tell if a particular memory is true or false. Perspectives on Psychological Science, 4, 370–374. Bernstein, I., & Nunnally, J. (1994). Psychometric theory. New York: McGraw-Hill. Berry, C., Sackett, P., & Wiemann, S. (2007). A review of recent developments in integrity test research. Personnel Psychology, 60, 271–301. Berry, D. J., Bridges, L. J., & Zaslow, M. J. (2004). Early childhood measures pro�iles. Washington, DC: Child Trends. Bersoff, D. N. (1988). Should subjective employment devices be scrutinized? Its elementary, my dear Ms. Watson. American Psychologist, 43, 1016–1018. Bertrand, J., Floyd, R., Weber, K., & others. (2004). National task force on fetal alcohol syndrome and fetal alcohol effect. Fetal alcohol syndrome: Guidelines for
referral and diagnosis. Atlanta, GA: Centers for Disease Control and Prevention. Bialik, C. (2010, September 4). Seven careers in a lifetime? Think twice, researchers say. Wall Street Journal. Bianchini, K., Etherton, J., Greve, K., Heinly, M., & Meyers, J. (2008). Classi�ication accuracy of MMPI-2 validity scales in the detection of pain-related malingering:
A known-groups study. Assessment, 15, 435–449. Bickley, P. G., Keith, T. Z., & Wolfe, L. M. (1995). The three-stratum theory of cognitive abilities: Test of the structure of intelligence across the life span.
Intelligence, 20, 309–328. Bilker, W. B., Hansen, J. A., Brensinger, C. M., & others. (2012). Development of abbreviated nine-item forms of the Raven’s Standard Progressive Matrices Test.
Assessment, 19, 354–369. Binet, A., & Simon, T. (1905). Methodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. Annee Psychologique, 11, 191–244. Blake, R. J., Potter, E., III, & Sliwak, R. (1993). Validation of the structural scales of the CPI for predicting the performance of junior of�icers in the U.S. Coast Guard.
Journal of Business Psychology, 7, 431–448. Blin, Dr. (1902). Les debilites mentales. Revue de Psychiatrie, 8, 337–345. Bloch, A. (2002). Refugees’ opportunities and barriers in employment and training, Research Report 179. Leeds, UK: Department for Work and Pensions. Block, J. (1961). The Q-sort method in personality assessment and psychiatric research. Spring�ield, IL: Charles C. Thomas. Block, J. (2008). The Q-Sort in character appraisal: Encoding subjective impressions of persons quantitatively. Washington, DC: American Psychological Association. Blum, G. (1950). The Blacky Pictures. New York: The Psychological Corporation. Blumenthal, J. A. (1985). Review of Jenkins Activity Survey. In J. V. Mitchell, Jr. (Ed.). The ninth mental measurements yearbook (vol. 1). Lincoln: Buros Institute of
Mental Measurements of the University of Nebraska-Lincoln. Blustein, D. L. (2006). The psychology of working: A new perspective. New York: Routledge. Blustein, D. L., Kenna, A., Gill, N., & DeVoy, J. (2008). The psychology of working: A new framework for counseling practice and public policy. The Career
Development Quarterly, 56, 294–308. Boake, C. (2002). From the Binet-Simon to the Wechsler-Bellevue: Tracing the history of intelligence testing. Journal of Clinical and Experimental
Neuropsychology, 24, 383–405. Board of Trustees of the Society for Personality Assessment. (2005). The status of the Rorschach in clinical and forensic practice: An of�icial statement by the
Board of Trustees of the Society for Personality Assessment. Journal of Personality Assessment, 85, 219–237. Boccaccini, M., Turner, D., & Murrie, D. (2008). Do some evaluators report consistently higher or lower PCL-R scores than others? Findings from a statewide
sample of sexually violent predator evaluations. Psychology, Public Policy, and Law, 14, 262–283. Boden, M. (2004). The creative mind: Myths and mechanisms (2nd ed.). London: Routledge. Boggs, D. H., & Simon, J. R. (1968). Differential effect of noise on tasks of varying complexity. Journal of Applied Psychology, 52, 148–153. Boggs, K. (1999). Campbell Interest and Skill Survey: Review and critique. Measurement and Evaluation in Counseling and Development, 32, 168–182. Bond, L. (1996). Norm- and criterion-referenced testing. Practical Assessment, Research and Evaluation, [Online journal], 5. Available: ericae.net (http://ericae.net)
. Bonner, C. M. (1988). Utilization of spiritual resources by patients experiencing a recent cancer diagnosis. Unpublished master’s thesis, University of Pittsburgh. Bonner, M. F., Ash, S., & Grossman, M. (2010). The new classi�ication of primary progressive aphasia into semantic, logopenic, or non�luent/agrammatic variants.
Current Neurology and Neuroscience Reports, 10, 484–490. Boring, E. G. (1923, June). Intelligence as the tests test it. New Republic, 35–37. Boring, E. G. (1950). A history of experimental psychology (2nd ed.). New York: Appleton-Century-Crofts. Borkowski, J. (1985). Signs of intelligence: Strategy generalization and metacognition. In S. R. Yussen (Ed.), The growth of re�lection in children. Orlando:
Academic Press. Borman, W., Ilgen, D., Klimoski, R., & Weiner, I. (2003). Handbook of psychology, industrial and organizational psychology. San Francisco: Jossey-Bass. Bornstein, M. H. (1994). Infancy. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Bornstein, R. F., & Masling, J. M. (2005). Scoring the Rorschach: Seven validated systems. Mahwah, NJ: Erlbaum. Boter, R., & Hoekstra-Vrolijk, S. (1994). ITVIC, an intelligence test for visually impaired children. In A. Kooijman & P. Looijestijn (Eds.), Low vision: Research and
new developments in rehabilitation (pp. 135–138). Amsterdam: IOS Press. Bouchard, T. J., Jr. (1994). Twin studies. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Bouchard, T. J., Jr., Lykken, D., McGue, M., Segal, N., & Tellegen, A. (1990). Sources of human psychological differences: The Minnesota Study of Twins Reared
Apart. Science, 250, 223–228. Bowden, E., & Jung-Beeman, M. (2003). Normative data for 144 compound remote associate problems. Behavior Research Methods, Instruments & Computers, 35,
634–639. Bowers, T., & Pantle, M. (1998). Shipley Institute for Living Scale and the Kaufman Brief Intelligence Test as screening instruments for intelligence. Assessment, 5,
187–195. Bowling, A. (1997). Measuring health: A review of quality of life measurement scales (2nd ed.). Buckingham, UK: Open University Press. Bowling, A. (2001). Measuring disease: A review of disease-speci�ic quality of life measurement scales (2nd ed.). Buckingham, UK: Open University Press.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 36/70
Bowman, M. (1989). Testing individual differences in ancient China. American Psychologist, 44, 576–578. Boyd, T. M., & Sauter, S. (1993). Route-�inding: A measure of everyday executive functioning in the head-injured adult. Applied Cognitive Psychology, 7, 171–181. Bracken, B. A., & Fagan, T. K. (1990). Guest editors’ introduction to the conference “Intelligence: Theories and Practice.” Journal of Psychoeducational Assessment,
8, 221–222. Brackett, M., & Mayer, J. (2003). Convergent, discriminant, and incremental validity of competing measures of emotional intelligence. Personality and Social
Psychology Bulletin, 29, 1147–1158. Braden, J. (1992). Intellectual assessment of deaf and hard of hearing people: A quantitative and qualitative research synthesis. School Psychology Review, 21, 82–
94. Braden, J., & Hannah, J. (1998). Assessment of hearing impaired and deaf children with the WISC-III. In D. Saklofske & A. Pri�itera (Eds.), Use of the WISC-III in
clinical practice. New York: Houghton Mif�lin. Bradley, K., Boyd-Wickizer, J., Powell, S., & Burman, M. (1998). Review: Some alcohol screening tests have acceptable test properties for use in general clinical
populations of U.S. women. Journal of the American Medical Association, 280, 166–171. Bradley, R., Corwyn, R., Pipes McAdoo, H., & Garcia Coll, C. (2001). The home environments of children in the United States Part I: Variations by age, ethnicity, and
poverty status. Child Development, 72, 1844–1867. Bradley, R. H., & Caldwell, B. M. (1984). 174 children: A study of the relationship between home environment and cognitive development during the �irst 5 years.
In A. W. Gottfried (Ed.), Home environment and early cognitive development: Longitudinal research. Orlando, FL: Academic Press. Bradley, R. H., & Rock, S. L. (1985). The HOME Inventory: Its relation to school failure and development of an elementary-age version. In W. K. Frankenburg, R. N.
Emde, & J. W. Sullivan (Eds.), Early identi�ication of children at risk. New York: Plenum. Bradley, R. H., Mundfrom, D., Whiteside, L., Case, P., & Barrett, K. (1994). A factor analytic study of the Infant-Toddler and Early Childhood versions of the HOME
Inventory administered to white, Black, and Hispanic American parents of children born preterm. Child Development, 65, 880–888. Bradley, R. H., Rock, S. L., Caldwell, B. M., & Brisby, J. A. (1989). Use of the HOME Inventory for families with handicapped children. American Journal on Mental
Retardation, 94, 313–330. Bradley-Johnson, S. (2001). Cognitive assessment for the youngest children: A critical review of tests. Journal of Psychoeducational Assessment, 19, 19–44. Bradshaw, J. L., & Mattingley, J. B. (1995). Clinical neuropsychology: Behavioral and brain science. San Diego, CA: Academic Press. Bradway, K. P. (1944). IQ constancy on the Revised Stanford-Binet from the preschool to the junior high school level. Journal of Genetic Psychology, 65, 197–217. Braithwaite, V., & Law, H. (1985). Structure of human values: Testing the adequacy of the Rokeach Value Survey. Journal of Personality and Social Psychology, 49,
250–263. Brannick, M. T., Michaels, C. E., & Baker, D. P. (1989). Construct validity of in-basket scores. Journal of Applied Psychology, 74, 957–963. Brannigan, G. G., & Decker, S. L. (2003). Bender Visual-Motor Gestalt Test (2nd ed.). Itasca, IL: Riverside Publishing. Brass, D. J., & Oldham, G. R. (1976). Validating an in-basket test using an alternative set of leadership scoring dimensions. Journal of Applied Psychology, 61, 652–
657. Brauer, B., Braden, J., Pollard, R., & Hardy-Braz, S. (1998). Deaf and hard of hearing people. In J. Sandoval, C. Frisby, K. Geisinger, J. Scheuneman, & J. Grenier (Eds.),
Test interpretation and diversity. Washington, DC: American Psychological Association. Brazelton, T. B., & Nugent, J. (1995). Neonatal Behavioral Assessment Scale (3rd ed.). London: Cambridge University Press. Breaugh, J. A. (2009). The use of biodata for employee selection: Past research and future directions. Human Resource Management Review, 19, 219–231. Bremner, J. D. (2005). Brain imaging handbook. New York: Norton. Breslau, N. (1994). A gradient relationship between low birth weight and IQ at age 6 years. Archives of Pediatric and Adolescent Medicine, 148, 377–383. Breslau, N., Chilcoat, H., Susser, E., & others. (2001). Stability and change in children’s Intelligence Quotient scores: A comparison of two socioeconomically
disparate communities. American Journal of Epidemiology, 154, 711–717. Breuer, J., & Freud, S. (1893–1895). Studies on hysteria. In J. Strachey (Ed., in collaboration with A. Freud). The standard edition of the complete psychological
works of Sigmund Freud (vol. 2). London: Hogarth, 1955. Brief, D. E., & Comrey, A. L. (1993). A pro�ile of personality for a Russian sample: As indicated by the Comrey Personality Scales. Journal of Personality Assessment,
60, 267–284. Britt, G., & Myers, B. (1994). The effects of Brazelton intervention: A review. Infant Mental Health Journal, 15, 278–292. Brodal, A. (1981). Neurological anatomy (3rd ed.). New York: Oxford University Press. Brody, E. B., & Brody, N. (1976). Intelligence: Nature, determinants and consequences. New York: Academic Press. Bromberg, W. (1959). The mind of man: A history of psychotherapy and psychoanalysis. New York: Harper & Row. Brooks, B., Holdnack, J. A., & Iverson, G. L. (2011). Advanced clinical interpretation of the WAIS-IV and WMS-IV: Prevalence of low scores varies by level of
intelligence and years of education. Assessment, 18, 156–167. Brooks, B., Iverson, G., Holdnack, J., & Feldman, H. (2008). Potential for misclassi�ication of mild cognitive impairment: A study of memory scores on the Wechsler
Memory Scale-III in healthy older adults. Journal of the International Neuropsychological Society, 14, 463–478. Brooks-Gunn, J., Klebanov, P., & Duncan, G. (1996). Ethnic differences in children’s intelligence test scores: Role of economic deprivation, home environment, and
maternal characteristics. Child Development, 67, 396–408. Brown, I. T., Chen, T., Gehlert, N. C., & Piedmont, R. L. (2012, October 8). Age and gender effects on the Assessment of Spirituality and Religious Sentiments
(ASPIRES) Scale: A cross-sectional analysis. Psychology of Religion and Spirituality [online publication]. Bruininks, R. H., Woodcock, R. W., Weatherman, R. F., & Hill, B. K. (1996). Scales of Independent Behavior-Revised, Interviewer’s Manual. Allen, TX: DLM Teaching
Resources. Bruyere, S. M., & O’Keeffe, J. (Eds.). (1994). Implications of the Americans with Disabilities Act for psychology. New York: Springer. Buck, J. (1948). The H-T-P technique, a qualitative and quantitative scoring method. Journal of Clinical Psychology Monograph Supplement, 5, 1–120. Buck, J. (1981). The House-Tree-Person technique: A revised manual. Los Angeles: Western Psychological Services. Bufford, R., & Parker, T., Jr. (1985). Religion and well-being: Concurrent validation of the Spiritual Well-Being Scale. Paper presented at the annual meeting of the
American Psychological Association, Los Angeles. Bufford, R., Paloutzian, R., & Ellison, C. (1991). Norms for the Spiritual Well-Being Scale. Journal of Psychology and Theology, 19, 56–70. Bullock, E., & Reardon, R. (2008). Interest pro�ile elevation, Big Five personality traits, and secondary constructs on the Self-Directed Search: A replication and
extension. Journal of Career Assessment, 16, 326–338. Burke, H. R. (1958). Raven’s Progressive Matrices: A review and critical evaluation. Journal of Genetic Psychology, 93, 199–228.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 37/70
Buschke, H., & Fuld, P. A. (1974). Evaluating storage, retention, and retrieval in disordered memory and learning. Neurology, 24, 1019–1025. Buss, A. (1997). Evolutionary perspectives on personality traits. In R. Hogan, J. Johnson, & S. Briggs (Eds.), Handbook of personality psychology. San Diego, CA:
Academic Press. Buss, D. M. (2009). How can evolutionary psychology successfully explain personality and individual differences? Perspectives on Psychological Science, 4, 359–
366. Butcher, J. N. (1985). Introduction to the special series. Journal of Consulting and Clinical Psychology, 53, 746–747. Butcher, J. N. (1993). The Minnesota Report user’s guide. Minneapolis, MN: National Computer System. Butcher, J. N. (2005). MMPI-2: A practitioner’s guide. Washington, DC: American Psychological Association. Butcher, J. N. (2011). A beginner’s guide to the MMPI-2 (3rd ed.). Washington, DC: American Psychological Association. Butcher, J. N. (Ed.). (1987). Computerized psychological assessment: A practitioner’s guide. New York: Basic Books. Butcher, J. N. (Ed.). (2000). Basic sources on the MMPI-2. Minneapolis, MN: University of Minnesota Press. Butcher, J. N., & Williams, C. L. (1992). Essentials of MMPI-2 and MMPI-A interpretation. Minneapolis: University of Minnesota Press. Butcher, J. N., & Williams, C. L. (2000). Essentials of MMPI-2 and MMPI-A interpretation. Minneapolis: University of Minnesota Press. Butcher, J. N., Dahlstrom, W. G., Graham, J. R., Tellegen, A., & Kaemmer, B. (1989). Minnesota Multiphasic Personality Inventory-2: Manual for administration and
scoring. Minneapolis: University of Minnesota Press. Butcher, J. N., Graham, J. R., Williams, C. L., & Ben-Porath, Y. S. (1990). Development and use of the MMPI-2 content scales. Minneapolis: University of Minnesota
Press. Butcher, J., Perry, J., & Atlis, M. (2000). Validity and utility of computer-based test interpretation. Psychological Assessment, 12, 6–18. Buxbaum, L. J., Dawson, A. M., & Linsley, D. (2012). Reliability and Validity of the Virtual Reality Lateralized Attention Test in Assessing Hemispatial Neglect in
Right-Hemisphere Stroke. Neuropsychology, 26, 430–441. Caldwell, B. M., & Bradley, R. H. (1984). Home observation for measurement of the environment. Little Rock: University of Arkansas at Little Rock. Caldwell, B. M., & Bradley, R. H. (1994). Environmental issues in developmental follow-up research. In S. L. Friedman & H. C. Haywood (Eds.), Developmental
follow-up: Concepts, domains, and methods. San Diego, CA: Academic Press. Caldwell, B. M., & Richmond, J. (1967). Social class level and the stimulation potential of the home. In J. Hellmuth (Ed.), The exceptional infant (vol. 1). Seattle, WA:
Special Child Publications. Callahan, L. A., McGreevy, M., Cirincione, C., & Stead-man, H. (1992). Measuring the effects of the Guilty But Mentally Ill (GBMI) verdict. Law and Human Behavior,
16, 447–462. Campbell, C. D. (1988). Coping with hemodialysis: Cognitive appraisals, coping behaviors, spiritual well-being, assertiveness, and family adaptability and
cohesion as correlates of adjustment (Doctoral dissertation, Western Conservative Baptist Seminary, 1983). Dissertation Abstracts International, 49, 538B. Campbell, D. (2002). The history and development of the Campbell Interest and Skill Survey. Journal of Career Assessment, 10, 150–168. Campbell, D. P. (1971). Handbook for the Strong Vocational Interest Blank. Stanford, CA: Stanford University Press. Campbell, D. P. (1974). Manual for the Strong-Campbell Vocational Interest Blank. Stanford, CA: Stanford University Press. Campbell, D. P., Hyne, S., & Nilsen, D. (1992). Manual for the Campbell Interest and Skill Survey. Minneapolis, MN: National Computer Systems. Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56, 81–105. Campbell, J. P., Gasser, M., & Oswald, F. (1996). The substantive nature of job performance variability. In K. R. Murphy (Ed.), Individual differences and behavior in
organizations. San Francisco: Jossey-Bass. Campbell, J., & McCord, D. (1996). The WAIS-R Comprehension and Picture Arrangement Subtests as measures of social intelligence: Testing traditional
interpretations. Journal of Psychoeducational Assessment, 14, 240–249. Campbell, J., Bell, S., & Keith, L. (2001). Concurrent validity of the Peabody Picture Vocabulary Test-Third Edition as an intelligence and achievement screener for
low SES African American children. Assessment, 8, 85–94. Campion, J. E. (1972). Work sampling for personnel selection. Journal of Applied Psychology, 56, 40–44. Campion, M. A., Pursell, E. D., & Brown, B. K. (1988). Structured interviewing: Raising the psychometric properties of the employment interview. Personnel
Psychology, 41, 25–42. Campione, J., & Brown, A. (1978). Toward a theory of intelligence: Contributions from research with retarded children. Intelligence, 2, 279–304. Can�ield, A. A. (1951). The “sten” scale—A modi�ied C-scale. Educational and Psychological Measurement, 11, 295–297. Cannell, J. J. (1988). Nationally normed elementary achievement testing in America’s public schools: How all 50 states are above the national average.
Educational Measurement: Issues and Practice, 7, 5–9. Capraro, R., & Capraro, M. (2002). Myers-Briggs Typica Indicator score reliability across studies: A meta-analytic reliability generalization study. Educational and
Psychological Measurement, 62, 590–602. Carless, S. (2000). The validity of scores on the Multidimensional Aptitude Battery. Educational and Psychological Measurement, 60, 592–603. Carlson, C. F., Kula, M., & St. Laurent, C. (1997). Rorschach revised DEPI and CDI with inpatient major depressives and borderline personality disorder with major
depression: Validity issues. Journal of Clinical Psychology, 53, 51–58. Carpenter, M. B. (1991). Core text of neuroanatomy (4th ed.). Baltimore: Williams & Wilkins. Carroll, D. (1988). How accurate is polygraph lie detection? In A. Gale (Ed.), The polygraph test: Lies, truth and science. London: Sage. Carroll, J. B. (1993). Human cognitive abilities. New York: Cambridge University Press. Carson, S., Peterson, J. B., & Higgins, D. M. (2005). Reliability, validity and factor structure of the Creative Achievement Questionnaire. Creativity Research Journal,
17, 37–50. Carter, C., Mintun, M., Nichols, T., & Cohen, J. (1997). Anterior cingulate gyrus dysfunction and selection attention de�icits in schizophrenia. American Journal of
Psychiatry, 154, 1670–1675. Carver, C., & Scheier, M. (2002). Optimism. In C. R. Snyder & S. Lopez (Eds.), The handbook of positive psychology (pp. 434–445). New York: Oxford University
Press. Carver, C., & Scheier, M. (2003). Optimism. In S. Lopez & C. R. Snyder (Eds.), Positive psychological assessment: A handbook of models and measures. Washington,
DC: American Psychological Association. Cascio, W. F. (1976). Turnover, biographical data, and fair employment practice. Journal of Applied Psychology, 61, 576–580. Cascio, W. F. (1987). Applied psychology in personnel management (3rd ed.). Englewood Cliffs, NJ: Prentice Hall.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 38/70
Cathers-Schiffman, T., & Thompson, M. (2007). Assessment of English- and Spanish-speaking students with the WISC-III and Leiter-R. Journal of Psychoeducational Assessment, 25, 41–52.
Cattell, H. E. P., & Mead, A. D. (2008). The Sixteen Personality Factor Questionnaire (16PF). In G. J. Boyle, G. Matthews, & D. H. Saklofske (Eds.), The SAGE handbook of personality theory and assessment (vol. 2, pp. 135–159). Thousand Oaks, CA: SAGE Publishers.
Cattell, J. McK. (1890). Mental tests and measurements. Mind, 15, 373–380. Cattell, R. (1950). Personality: A systematic theoretical and factual study. New York: McGraw-Hill. Cattell, R. B. (1941). Some theoretical issues in adult intelligence testing. Psychological Bulletin, 38, 592 (abstract). Cattell, R. B. (1971). Abilities: Their structure, growth, and action. Boston: Houghton Mif�lin. Cattell, R. B. (1973). Personality pinned down. Psychology Today, 7, 40–46. Cautela, J. R. (1977). Behavioral analysis forms for clinical intervention. Champaign, IL: Research Press. Ceci, S. (1996). On intelligence: A bio-ecological treatise on intellectual development. (Expanded ed.). Cambridge, MA: Harvard University Press. Ceci, S. J. (1994). Bioecological theory of intellectual development. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Centers for Disease Control and Prevention. (2012). Alcohol use and binge drinking among women of child-bearing age—United States, 2006–2010. Morbidity
and Mortality Weekly Report, 61, 534–538. Chaffee, J. W. (1985). The thorny gates of learning in Sung China: A social history of examinations. Cambridge: Cambridge University Press. Chalmers, T. (1833). On the power, wisdom, and goodness of God as manifested in the adaptation of external nature to the moral and intellectual constitution of
man. London: William Pickering. Chamberlin, J. (2009). How do you spot raw legal talent? Take this test. Monitor on Psychology, 40(6), 12. Chan, R. (2000). Attentional de�icits in patients with closed head injury: A further study to the discriminative validity of the Test of Everyday Attention. Brain
Injury, 14, 227–236. Chan, R., & Lai, M. (2006). Latent structure of the Test of Everyday Attention: Convergent evidence from patients with traumatic brain injury. Brain Injury, 20,
653–659. Chan, R., Lai, M., & Robertson, I. (2006). Latent structure of the Test of Everyday Attention in a non-clinical Chinese sample. Archives of Clinical Neuropsychology,
21, 477–485. Chapell, M. S., Blanding, Z. B., Silverstein, M. E., & others. (2005). Test Anxiety and Academic Performance in Undergraduate and Graduate Students. Journal of
Educational Psychology, 97, 268–274. Chapman, L. J., & Chapman, J. P. (1967). Genesis of popular but erroneous psychodiagnostic observations. Journal of Abnormal Psychology, 74, 271–280. Chase, C. I. (1985). Review of the Torrance Tests of Creative Thinking. Ninth mental measurements yearbook. Lincoln, NB: University of Nebraska Press. Cherpitel, C. (2002). Screening for alcohol problems in the U.S. general population: Comparison of the CAGE, RAPS4, and RAPS4-QF by gender, ethnicity, and
service utilitzation. Alcoholism: Clinical and Experimental Research, 26, 1686–1691. Chiaravalloti, N. D., & DeLuca, J. (2003). Assessing the behavioral consequences of multiple sclerosis: An application of the Frontal Systems Behavior Scale
(FrSBe). Cognitive and Behavioral Neurology, 16, 54–67. Chibnall, J., & Detrick, P. (2003). The NEO-PI-R, Inwald Personality Inventory, and MMPI-2 in the prediction of police academy performance: A case for
incremental validity. American Journal of Criminal Justice, 27, 233–248. Chin, C., Ledesma, H., Cirino, P., & others. (2001). Relation between Kaufman Brief Intelligence Test and WISC-III scores of children with RD. Journal of Learning
Disabilities, 34, 2–8. Choi, H., & Proctor, T. (1994). Error-prone subtests and error types in the administration of the Stanford-Binet Intelligence Scale: Fourth Edition. Journal of
Psychoeducational Assessment, 12, 165–171. Chung, J. (2009). Clinical validity of Fuld Object Memory Evaluation to screen for dementia in Chinese society. International Journal of Geriatric Psychiatry, 24,
156–162. Chung, J., & Ho, W. (2009). Validity of Fuld Object Memory Evaluation for the detection of dementia in nursing home residents. Aging and Mental Health, 13, 274–
279. Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Erlbaum. Clark, D. A. (1988). The validity of measures of cognition: A review of the literature. Cognitive Therapy and Research, 12, 1–20. Clarkin, J. F., Hull, J., Cantor, J., & Sanderson, C. (1993). Borderline personality disorder and personality traits: A comparison of SCID-II BPD and NEO-PI.
Psychological Assessment, 5, 472–476. Clarren, S., Randels, S., Sanderson, M., & Fineman, R. (2001). Screening for fetal alcohol syndrome in primary schools: A feasibility study. Teratology, 63, 3–10. Cleary, T. A., Humphreys, L. G., Kendrick, S. A., & Wesman, A. (1975). Educational uses of tests with disadvantaged students. American Psychologist, 30, 15–41. Cleckley, H. (1941). The mask of sanity. St. Louis, MO: C. V. Mosby. Cleckley, H. (1976). The mask of sanity (5th ed.). St. Louis, MO: Mosby. Clemans, W. V. (1971). Test administration. In R. L. Thorndike (Ed.), Educational measurement (2nd ed.). Washington, DC: American Council on Education. Cleveland, J. N., Murphy, K. R., & Williams, R. E. (1989). Multiple uses of performance appraisal: Prevalence and correlates. Journal of Applied Psychology, 74, 130–
135. Cohen, J. (1960). A coef�icient of agreement for nominal scales. Educational and Psychological Measurement, 20, 37–46. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Erlbaum. Cohen, M. (1997). Children’s Memory Scale. San Antonio, TX: Psychological Corporation. Cohen, S., & Janicki-Deverts, D. (2009). Can we improve our physical health by altering our social networks? Perspectives on Psychological Science, 4, 375–378. Colby, A., & Kohlberg, L. (1987). The measurement of moral judgment (vol. I). Cambridge: Cambridge University Press. Colby, A., Kohlberg, L., Gibbs, J. C., & others. (1978). Measuring moral judgment: Standardized scoring manual. Cambridge, MA: Harvard University, Moral
Education Research Foundation. Colby, A., Kohlberg, L., Gibbs, J., & Lieberman, M. (1983). A longitudinal study of moral judgment. Monographs for the Society for Research in Child Development,
48, 1, 2. Cole, N. S., & Moss, P. A. (1989). Bias in test use. In R. L. Linn (Ed.), Educational measurement (3rd ed.). New York: ACE/Macmillan. College Board. (2005). Retrieved from www.collegeboard.com/student/testing/sat/ (http://www.collegeboard.com/student/testing/sat/) on September 21, 2005. Collins, J. M., & Schmidt, F. L. (1993). Personality, integrity, and white collar crime: A construct validity study. Personnel Psychology, 46, 295–311.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 39/70
Colom, R., Quiroga, M., & Juan-Espinosa, M. (1999). Are cognitive sex differences disappearing? Evidence from Spanish populations. Personality and Individual Differences, 27, 1189–1195.
Committee on Ethical Guidelines for Forensic Psychologists. (1991). Specialty guidelines for forensic psychologists. Law and Human Behavior, 15, 655–665. Community Research Partners. (2007). School readiness assessment: A review of the literature. Columbus, OH: Author. Comrey, A. (1995). Career assessment and the Comrey Personality Scales. Journal of Career Assessment, 3, 140–156. Comrey, A. L. (1970). Manual for the Comrey Personality Scales. San Diego, CA: EdITS. Comrey, A. L. (1973). A �irst course in factor analysis. New York: Academic Press. Comrey, A. L. (1980). Handbook of interpretations for the Comrey Personality Scales. San Diego, CA: EdITS. Comrey, A. L. (2008). The Comrey Personality Scales. In G. J. Boyle, G. Matthews, & D. H. Saklofske (Eds.), The SAGE handbook of personality theory and assessment,
vol 2: Personality measurement and testing (pp. 113–134). Thousand Oaks, CA: Sage Publications. Comrey, A. L., & Backer, T. (1970). Construct validation of the Comrey Personality Scales. Multivariate Behavior Research, 5, 469–477. Comrey, A. L., & Schiebel, D. (1983). Personality test correlates of psychiatric outpatient status. Journal of Consulting and Clinical Psychology, 51, 756–762. Comrey, A. L., & Schiebel, D. (1985). Personality test correlates of psychiatric case history data. Journal of Consulting and Clinical Psychology, 53, 470–479. Conn, H. O. (2011). Normal pressure hydrocephalus (NPH): More about NPH by a physician who is the patient. Clinical Medicine, 11(2), 162–165. Conners, C. K. (1990). Conners’ Rating Scales. Los Angeles: Western Psychological Services. Conners, C. K. (1991). Conners’ Teacher Rating Scales- 39. North Tonawanda, NY: Multi-Health Systems, Inc. Conners, C. K. (1995). Conners’ Continuous Performance Test II (CPT II). North Tonawanda, NY: Multi-Health Systems, Inc. Conners, C. K. (1997). Conners’ Rating Scales-Revised. North Tonawanda, NY: Multi-Health Systems. Conoley, C. W. (1992). Review of Beck Depression Inventory. The eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Conoley, C. W., Plake, B., & Kemmerer, B. (1991). Issues in computer-based test interpretive systems. Computers in Human Behavior, 7, 97–101. Constantino, G., & Malgady, R. (1996). Development of TEMAS, a multicultural thematic apperception test: Psychometric properties and clinical utility. In G. R.
Sodowsky & J. C. Impara (Eds.), Multicultural assessment in counseling and clinical psychology. Lincoln, NE: The Buros Institute of Mental Measurements. Constantino, G., & Malgady, R. (2000). Multicultural and cross-cultural utility of the TEMAS (Tell-Me-A-Story) Test. In R. Dana (Ed.), Handbook of cross-cultural
and multicultural personality assessment. Mahwah, NJ: Erlbaum. Constantino, G., Malgady, R., & Rogler, L. (1988). Tell-Me-A-Story (TEMAS): Manual. Los Angeles: Western Psychological Services. Conte, J. (2005). A review and critique of emotional intelligence measures. Journal of Organizational Behavior, 26, 433–440. Conway, J. M., Jako, R., & Goodman, D. (1995). A meta-analysis of interrater and internal consistency reliability of selection interviews. Journal of Applied
Psychology, 80, 565–579. Cooper, D., & Shepard, K. (1992). Review of DIAL-R. Learning Disabilities Research & Practice, 7, 171–174. Corkin, S. (1968). Acquisition of motor skill after bilateral medial temporal-lobe excision. Neuropsychologia, 6, 255–265. Cornelius, S. W., & Caspi, A. (1987). Everyday problem solving in adulthood and old age. Psychology and Aging, 2, 144–153. Cosden, M. (1992). Review of the Draw A Person: A Quantitative Scoring System. The eleventh mental measurements yearbook. Lincoln: University of Nebraska
Press. Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and the NEO Five Factor Inventory (NEO-FFI) professional manual. Odessa, FL:
Psychological Assessment Resources. Costa, P. T., Herbst, J. H., McCrae, R. R., & Siegler, I. C. (2000). Personality at midlife: Stability, intrinsic maturation, and response to life events. Assessment, 7, 365–
378. Costa, P. T., Jr. (1991). Clinical use of the �ive-factor model. Journal of Personality Assessment, 57, 393–398. Costa, P. T., Jr., & McCrae, R. (1989). NEO Five-Factor Inventory test manual. Port Huron, MI: Sigma Assessment Systems. Costa, P. T., Jr., & McCrae, R. (1992). NEO PI-R test manual. Port Huron, MI: Sigma Assessment Systems. Costa, P. T., Jr., McCrae, R. R., & Holland, J. L. (1984). Personality and vocational interests in an adult sample. Journal of Applied Psychology, 69, 390–400. Costa, P., McCrae, R., & Martin, T. (2005). The NEO-PI-3: A more readable revised NEO Personality Inventory. Journal of Personality Assessment, 84, 261–270. Costa, P., McCrae, R., & Martin, T. (2008). Incipient adult personality: The NEO-PI-3 in middle-school-aged children. British Journal of Developmental Psychology,
26, 71–89. Costenbader, V., & Ngari, S. (2001). A Kenya standardization of the Raven’s Coloured Progressive Matrices. School Psychology International, 22, 258–268. Cote, L., & Crutcher, M. D. (1991). The basal ganglia. In E. R. Kandel, J. H. Schwartz, & T. M. Jessell (Eds.), Principles of neural science (3rd ed.). New York: Elsevier. Courvoisier, D. S., Eid, M., & Lischetzke, T. (2012). Compliance to a cell phone-based ecological momentary assessment study: The effect of time and personality
characteristics. Psychological Assessment, 24, 713–720. Coveny, T. E. (1972). A new test for the visually handicapped: Preliminary analysis of reliability and validity of the Perkins-Binet. Education of the Handicapped, 4,
97–101. Cowdery, K. M. (1926–27). Measurement of professional attitudes: Differences between lawyers, physicians, and engineers. Journal of Personnel Research, 5, 131–
141. Craig, R. J. (Ed.). (1993). The Millon Clinical Multiaxial Inventory: A clinical research information synthesis. Hillsdale, NJ: Erlbaum. Cramond, B., Matthews-Morgan, J., Bandalos, D., & Zuo, L. (2005). A report on the 40-year follow-up of the Torrance Tests of Creative Thinking: Alive and well in
the new millennium. Gifted Child Quarterly, 49, 283–291. Crandall, J. E. (1981). Theory and measurement of social interest: Empirical tests of Alfred Adler’s concept. New York: Columbia University Press. Crawford, J. R., Sommerville, J., & Robertson, I. (1997). Assessing the reliability and abnormality of subtest differences on the Test of Everyday Attention. British
Journal of Clinical Psychology, 36, 609–617. Creed, P., Patton, W., & Bartrum, D. (2002). Multidimensional properties of the LOT-R: Effects of optimism and pessimism on career and well-being related
variables in adolescents. Journal of Career Assessment, 10, 42–61. Cripe, L. (1996). The ecological validity of executive function testing. In R. J. Sbordone & C. J. Long (Eds.), Ecological validity of neuropsychological testing. Delray
Beach, FL: GR Press/St. Lucie Press. Critchley, M. (1953). The parietal lobes. London: Edward Arnold. Cronbach, L. J. (1951). Coef�icient alpha and the internal structure of tests. Psychometrika, 16, 297–334. Cronbach, L. J. (1971). Test validation. In R. L. Thorndike (Ed.), Educational measurement (2nd ed.). Washington, DC: American Council on Education.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 40/70
Cronbach, L. J. (1988). Five perspectives on the validity argument. In H. Wainer & H. I. Braun (Eds.), Test validity. Hillsdale, NJ: Lawrence Erlbaum. Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52, 281–302. Culbertson, J., & Edmonds, A. (1996). Learning Disabilities. In R. Adams, O. Parsons, J. Culbertson, & S. Nixon (Eds.). Neuropsychology for clinical practice: Etiology,
assessment, and treatment of common neurological disorders. Washington, DC: American Psychological Association. Cullen, M., & Sackett, P. (2004). Integrity testing in the workplace. In J. Thomas (Ed.), Comprehensive handbook of psychological assessment, Vol. 4: Industrial and
organizational assessment. Hoboken, NJ: John Wiley. Cummings, N. A. (2007). Treatment and assessment take place in an economic context, always. In S. O. Lilienfeld & W. T. O’Donohue (Eds.), The great ideas of
clinical science: 17 principles that every mental health professional should understand (pp. 163–184). New York: Routledge. Cummings, R., Maddux, C., Harlow, S., & Dyas, L. (2002). Academic misconduct in undergraduate teacher education students and its relationship to their
principled moral reasoning. Journal of Instructional Psychology, 29, 286–296. Cunningham, M., Wong, D., & Barbee, A. (1994). Self-presentation dynamics on overt integrity tests: Experimental studies with the Reid Report. Journal of Applied
Psychology, 79, 643–658. Cureton, E. E. (1950). Validity, reliability, and baloney. Educational and Psychological Measurement, 10, 94–96. Cutler, B. L., & Kovera, M. B. (2011). Expert psychological testimony. Current Directions in Psychological Science, 20, 53–57. da Costa Armentano, C. G., Porto, C. S., Brucki, S., & Nitrini, R. (2009). Study on the Behavioural Assessment of the Dysexecutive Syndrome (BADS) performance in
healthy individuals, mild cognitive impairment and Alzheimer’s disease: A preliminary study. Dementia & Neuropsychologia, 3, 101–107. Dahlstrom, W. G., Welsh, G. S., & Dahlstrom, L. E. (1975). An MMPI handbook: Vol. II. Research applications. Minneapolis: University of Minnesota Press. Daley, T., Whaley, S., Sigman, M., Espinosa, M., & Neumann, C. (2003). IQ on the rise: The Flynn effect in rural Kenyan children. Psychological Science, 14, 215–219. Dana, R. H. (1959). Proposal for objective scoring of the TAT. Perceptual and Motor Skills, 10, 27–43. Das J. P., Naglieri, J., & Kirby, J. (1994). Assessment of cognitive processes: The PASS theory of intelligence. Boston: Allyn and Bacon. Das, J. P. (1994). Serial and parallel processing. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Das, J. P., & Naglieri, J. A. (1993). Cognitive assessment system: Standardization version. Chicago: Riverside. Das, J. P., Kirby, J. R., & Jarman, R. F. (1979). Simultaneous and successive cognitive processes. New York: Academic Press. Das, J. P., Kirby, J. R., & Jarman, R. F. (1979). Simultaneous and successive cognitive processes. Orlando, FL: Academic Press. Davis, A. S., Johnson, J. A., & D’Amato, R. C. (2005). Evaluating and using long-standing school neuro-psychological batteries: The Halstead-Reitan and the Luria-
Nebraska neuropsychological batteries. In R. C. D’Amato, E. Fletcher-Jansen, & C. R. Reynolds (Eds.), Handbook of school neuropsychology (pp. 236–263). Hoboken, NJ: Wiley.
Davis, C. (1980). Perkins-Binet Tests of Intelligence for the blind. Watertown, MA: Perkins School for the Blind. Davis, E., Glynn, L., Schetter, C., & others. (2007). Prenatal exposure to maternal depression and cortisol in�luences infant temperament. Journal of the American
Academy of Child and Adolescent Psychiatry, 46, 737–746. Davison, M., Gasser, M., & Ding, S. (1996). Identifying major pro�ile patterns in a population: An exploratory study of WAIS and GATB patterns. Psychological
Assessment, 1, 26–31. Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus actuarial judgment. Science, 243, 1668–1674. Dawis, R. V. (1996). The theory of work adjustment and person-environment correspondence counseling. In D. Brown & L. Brooks (Eds.), Career choice and
development (3rd ed., pp. 75–120). San Francisco: Jossey-Bass. Dawis, R. V. (2002). Person-Environment Correspondence theory. In D. Brown & Associates (Eds.), Career choice and development (4th ed., pp. 427–464). San
Francisco: Jossey-Bass. Dawis, R. V., & Lofquist, L. H. (1984). A psychological theory of work adjustment. Minneapolis: University of Minnesota Press. Dawson, A., Buxbaum, L. J., & Rizzo, A. A. (2008). The Virtual Reality Lateralized Attention Test: Sensitivity and validity of a new clinical tool for assessing
hemispatial neglect. Virtual Rehabilitation, 2008, 77–82. Dayan, K., Fox, S., & Kasten, R. (2008). The preliminary employment interview as a predictor of assessment center outcomes. International Journal of Selection
and Assessment, 16, 102–111. de Bildt, A., Kraijere, D., Sytema, S., & Minderaa, R. (2005). The psychometric properties of the Vineland Adaptive Behavior Scales in children and adolescents
with mental retardation. Journal of Autism and Developmental Disorders, 35, 53–62. de Raad, B., & Perugini, M. (Eds.). Big �ive assessment. Ashland, OH: Hogrefe and Huber Publishers. Decker, S. L. (2008). Measuring growth and decline in visual-motor processes with the Bender-Gestalt second edition. Journal of Psychoeducational Assessment,
26, 3–15. DeCrans, M. (1990). Spiritual well-being in the rural elderly. Unpublished manuscript, Marquette University, Milwaukee, WI. Delis, D. C., & Kaplan, E. (1982). Assessment of aphasia with the Luria-Nebraska Neuropsychological Battery: A conceptual critique. Journal of Consulting and
Clinical Psychology, 50, 32–39. Delis, D. C., Kramer, J., Kaplan, E., & Ober, B. (2000). California Verbal Learning Test—Second Edition. San Antonio, TX: The Psychological Corporation. Dellas, M., & Gaier, E. L. (1970). Identi�ication of creativity: The individual. Psychological Bulletin, 73, 55–73. Deri, S. (1949). Introduction to the Szondi Test. New York: Grune & Stratton. Dey, A. N., Schiller, J. S., & Tai, D. A. (2004). Summary health statistics for U.S. children: National Health Interview Survey, 2002. Washington, DC: National Center for
Health Statistics. Diamond, S. (1980). Wundt before Leipzig. In R. W. Rieber (Ed.), Wilhelm Wundt and the making of a scienti�ic psychology. New York: Plenum Press. Dickens, W., & Flynn, J. (2006). Black Americans reduce the racial IQ gap. Psychological Science, 17, 913–920. Dickens, W., & Flynn, J. R. (2006). Black Americans reduce the racial IQ gap: Evidence from standardization samples. Psychological Science, 17, 913–920. Diener, E. (2009). Editor’s introduction: Special issue on the next big questions in psychology. Perspectives on Psychological Science, 4, 325. Diessner, R., & Lewis, G. (2007). Further validation of the Gratitude, Resentment, and Appreciation Test. Journal of Social Psychology, 147, 445–447. Digman, J. (1990). Personality structure: Emergence of the �ive-factor model. Annual Review of Psychology, 41, 417–440. Dikmen, S., Machamer, J., Winn, H., & Temkin, N. (1995). Neuropsychological outcome at 1-year post head injury. Neuropsychology, 9, 80–90. DiLalla, L. F., Thompson, L. A., Plomin, R., & others. (1990). Infant predictors of preschool and adult IQ: A study of infant twins and their parents. Developmental
Psychology, 26, 759–769. Dillon, R. F., Pohlmann, J. T., & Lohman, D. F. (1981). A factor analysis of Raven’s Advanced Progressive Matrices freed of dif�iculty factors. Educational and
Psychological Measurement, 41, 1295–1302.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 41/70
Dixon, C. E. (2011, May 16). Traumatic brain injury produced by exposure to blasts, a critical problem in current wars: Biomarkers, clinical studies, and animal modes. Proceedings of the Society of Photo-Optical Instrumentation Engineers, 80290M.
Dodge, K. A. (2009). Mechanisms of gene-environment interaction effects in the development of Conduct disorder. Perspectives on Psychological Science, 4, 408– 414.
Dodrill, C. B. (1979). Sex differences on the Halstead-Reitan Neuropsychological Battery and on other neuropsychological measures. Journal of Clinical Psychology, 35, 236–241.
Dodrill, C. B. (1981). An economical method of measuring general intelligence in adults. Journal of Consulting and Clinical Psychology, 49, 668–673. Dodrill, C. B., & Warner, M. H. (1988). Further studies of the Wonderlic Personnel Test as a brief measure of intelligence. Journal of Consulting and Clinical
Psychology, 56, 145–147. Doll, E. A. (1935). The Vineland Social Maturity Scale. Training School Bulletin, 32, 1–7, 25–32, 48–55, 68–74. Dolliver, R. H., Irvin, J. A., & Bigley, S. E. (1972). Twelve-year follow-up of the Strong Vocational Interest Blank. Journal of Counseling Psychology, 19, 212–217. Donders, J. (1995). Validity of the Kaufman Brief Intelligence Test (K-BIT) in children with traumatic brain injury. Assessment, 2, 219–224. Donders, J., & Levitt, T. (2012). Criterion validity of the neuropsychological assessment battery after traumatic brain injury. Archives of Clinical Neuropsychology,
27, 440–445. Donders, J., Tulsky, D., & Zhu, J. (2001). Criterion validity of new WAIS-III subtest scores after traumatic brain injury. Journal of the International
Neuropsychological Society, 7, 892–898. Donlon, T. F. (Ed.). (1984). The College Board technical handbook for the Scholastic Aptitude Test and Achievement Tests. New York: College Entrance Examination
Board. Donnay, D., & Borgen, F. (1996). Validity, structure, and content of the 1994 Strong Interest Inventory. Journal of Counseling Psychology, 43, 275–291. Donnay, D., Thompson, R., Morris, M., & Schaubhut, N. (2004). Technical brief for the newly revised Strong Interest Inventory assessment: Content, reliability and
validity. Mountain View, CA: Consulting Psychologists Press. Drakeley, R. J., Herriot, P., & Jones, A. (1988). Biographical data, training success and turnover. Journal of Occupational Psychology, 61, 145–152. Drasgow, F., Olson-Buchanan, J., & Moberg, P. (1999). Development of an interactive assessment: Trials and tribulations. In F. Drasgow & J. Olson-Buchanan (Eds.),
Innovations in computerized assessment. Mahwah, NJ: Erlbaum. Drebing, C., Van Gorp, W., Stuck, A., Mitrushina, M., & Beck, J. (1994). Early detection of cognitive decline in higher cognitively functioning older adults: Sensitivity
and speci�icity of a neuropsychological screening battery. Neuropsychology, 8, 31–37. DuBois, P. E. (1939). A test standardized on Pueblo Indian children. Psychological Bulletin, 36, 523. DuBois, P. H. (1970). A history of psychological testing. Boston: Allyn and Bacon. Dumenci, L. (1995). Construct validity of the Self-Directed Search using hierarchically nested structural models. Journal of Vocational Behavior, 47, 21–34. Dumont, R., Cruse, C., Price, L., & Whelley, P. (1996). The relationship between the Differential Ability Scales (DAS) and the Wechsler Intelligence Scale for
Children-Third Edition (WISC-III). Psychology in the Schools, 33, 203–209. Dunai, F., & Porter, R. (2001). Armed Services Vocational Aptitude Battery predictors of entry-level radiography students’ success. Military Medicine, 166, 422–
426. Dunn, L. M., & Dunn, D. M. (2007). Examiner’s manual: Peabody Picture Vocabulary Test—Fourth Edition. New York: Pearson. Dunn, L. M., & Dunn, L. M. (1981). Peabody Picture Vocabulary Test-Revised. Circle Pines, MN: American Guidance Service. Dunn, L. M., & Dunn, L. M. (1998). Examiner’s Manual: Peabody Picture Vocabulary Test-III. Circle Pines, MN: American Guidance Service. Dyce, J. A. (1996). Factor structure of the Beck Hopelessness Scale. Journal of Clinical Psychology, 52, 555–558. Ebbinghaus, H. (1885/1913). Memory: A contribution to experimental psychology. Translated by Henry A. Ruger & Clara E. Bussenius. New York: Teachers College
Press. Ebbinghaus, H. (1897). Ueber eine neue Methode zur Pruefung geistiger Faehigkeiten und ihre Anwendung bei Schulkindern. Zeitschrift fuer Angewandte
Psychologie, 13, 401–459. Eccles, J. C. (1973). The understanding of the brain. New York: McGraw-Hill. Educational Testing Service. (1989). Guidelines for proper use of GRE scores. Princeton, NJ: Author. Eggerth, D. D. (2008). From theory of work adjustment to person-environment correspondence counseling: Vocational psychology as positive psychology. Journal
of Career Assessment, 16, 60–74. Eisenstein, N., & Engelhart, C. (1997). Comparison of the KBIT with short forms of the WAIS-R in a neuropsychological population. Psychological Assessment, 9,
57–62. Elder, G. (1974). Children of the great depression: Social change in life experience. Boulder, CO: Westview Press. Elliott, C. D. (1990). The Differential Ability Scales: Introductory and technical handbook. San Antonio, TX: The Psychological Corporation. Elliott, C. D. (2007). Differential Ability Scales—Second Edition: Introductory and technical manual. San Antonio, TX: Harcourt Assessment. Ellis, A. (1962). Reason and emotion in psychotherapy. New York: Lyle Stuart. Ellison, C. W. (1983). Spiritual well-being: Conceptualization and measurement. Journal of Psychology and Theology, 11, 330–340. Ellison, C. W., & Smith, J. (1991). Toward an integrative measure of health and well-being. Journal of Psychology and Theology, 19, 35–48. Embretson, S. E. (1996). The new rules of measurement. Psychological Assessment, 8, 341–349. Embretson, S. E., & Reise, S. (2000). Item response theory for psychologists. Mahwah, NJ: Erlbaum. Emmons, R., McCullough, M., & Tsang, J. (2003). The assessment of gratitude. In S. Lopez & C. R. Snyder (Eds.), Positive psychological assessment (pp. 345–360).
Washington, DC: American Psychological Association. Eonta, S. E., Carr, W., McArdle, J. J., & others. (2011). Automated neuropsychological assessment metrics: Repeated assessments with two military samples.
Aviation, Space, and Environmental Medicine, 82, 34–39. Erard, R. E. (2012). Expert testimony using the Rorschach Performance Assessment System in psychological injury cases. Psychological Injury and the Law, 5,
122–134. Erdberg, P. (1985). The Rorschach. In C. S. Newmark (Ed.), Major psychological assessment instruments. Boston: Allyn and Bacon. Esquirol, J. E. D. (1845/1838). Mental maladies. (trans. E. K. Hunt). Philadelphia: Lea & Blanchard. Estes, W. K. (1974). Learning theory and intelligence. American Psychologist, 29, 740–749. Evans, D. A., Funkenstein, H., Albert, M., & others. (1989). Prevalence of Alzheimer’s Disease in a community population of older persons. Journal of the American
Medical Association, 262, 2551–2556.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 42/70
Ewing, J. A. (1984). Detecting alcoholism: The CAGE questionnaire. Journal of the American Medical Association, 252, 1905–1907. Exner, J. E., Jr. (1991). The Rorschach: A comprehensive system, Volume 2. Current research and advanced interpretation (2nd ed.). New York: Wiley. Exner, J. E., Jr. (1993). The Rorschach: A comprehensive system, Volume 1. Basic foundations (3rd ed.). New York: Wiley. Exner, J. E., Jr. (1995). Issues and methods in Rorschach research. Mahwah, NJ: Erlbaum. Exner, J. E., Jr., & Weiner, I. B. (1994). The Rorschach: A comprehensive system, Volume 3. Assessment of children and adolescents (2nd ed.). New York: Wiley. Eyde, L. D., & Primhoff, E. S. (1992). Responsible test use. In M. Zeidner and R. Most (Eds.), Psychological testing: An inside view. Palo Alto, CA: Consulting
Psychologists Press. Eyde, L. D., Robertson, G. J., & Krug, S. (2009). Responsible test use: Case studies for assessing human behavior. Washington, DC: American Psychological
Association. Eyde, L. D., Robertson, G. J., Krug, S., & others. (1993). Responsible test use: Case studies for assessing human behavior. Washington, DC: American Psychological
Association. Eysenck, H. J. (1986). Is intelligence? In R. J. Sternberg & D. K. Detterman (Eds.), What is intelligence? Contemporary viewpoints on its nature and de�inition.
Norwood, NJ: Ablex. Eysenck, H. J. (1986). Toward a new model of intelligence. Personality and Individual Differences, 7, 731–736. Eysenck, H. J., & Eysenck, M. W. (1975). Manual of the Eysenck Personality Questionnaire. San Diego: Educational and Industrial Testing Service. Eysenck, H. J., & Eysenck, M. W. (1985). Personality and individual differences: A natural science approach. New York: Plenum Press. Factor, S., & Weiner, W. (2008). Parkinson’s disease: Diagnosis and clinical management (2nd ed.). New York: Demos Medical Publishing. Fagan, J. F. III, & Haiken-Vasen, J. (1997). Selective attention to novelty as a measure of information processing across the lifespan. In J. Burack & J. Enns (Eds.),
Attention, development, and psychopathology. New York: Guilford. Fagan, J. F. III, & McGrath, S. K. (1981). Infant recognition memory and later intelligence. Intelligence, 5, 121–130. Fagan, J. F. III, & Shepherd, P. A. (1986). The Fagan Test of Infant Intelligence: Training manual. Cleveland, OH: Infantest Corporation. Fagan, J. F. III, Singer, L., Montie, J., & Shepherd, P. (1986). Selective screening device for the early detection of normal or delayed cognitive development in infants
at risk for later mental retardation. Pediatrics, 78, 1021–1026. Fagan, J. F. III. (1984). Infant memory. In M. Moscovitch (Ed.), Infant memory. New York: Plenum Press. Fagan, J., & Holland, C. (2006). Racial equality in intelligence: Predictions from a theory of intelligence as processing. Intelligence, 15, 319–334. Fancher, R. E. (1985). The intelligence men: Makers of the IQ controversy. New York: Norton. Farrell, M., & Phelps, L. (2000). A comparison of the Leiter-R and the Universal Nonverbal Intelligence Test (UNIT) with children classi�ied as language impaired.
Journal of Psychoeducational Assessment, 18, 268–274. Faul, M., Xu, L., Wald, M. M., & Coronado, V. G. (2010). Traumatic brain injury in the United States: Emergency department visits, hospitalizations, and deaths.
Atlanta, GA: Centers for Disease Control and Prevention. Federal Rules of Evidence for United States Courts and Magistrates. (1975). St. Paul, MN: West Publishing Company. Fehring, R., Brennan, P., & Keller, M. (1987). Psychological and spiritual well-being in college students. Research in Nursing and Health, 10, 391–398. Feist, G. (1999). Autonomy and independence. Encyclopedia of creativity (vol. 1, pp. 157–163). San Diego, CA: Academic Press. Feist, G., & Barron, F. (2003). Predicting creativity from early to late adulthood: Intellect, potential, and personality. Journal of Research in Personality, 37, 62–88. Feldman, R. D. (1982). Whatever happened to the quiz kids? Chicago: Chicago Review Press. Feldstein, S., & Miller, W. (2007). Does subtle screening for substance abuse work? A review of the Substance Abuse Subtle Screening Inventory (SASSI).
Addiction, 102, 41–50. Feldt, L. S., & Brennan, R. L. (1989). In R. L. Linn (Ed.), Educational measurement (3rd ed.). New York: American Council on Education/Macmillan. Ferris, G., Judge, T., Rowland, K., & Fitzgibbons, D. (1994). Subordinate in�luence and the performance evaluation process: Test of a model. Organizational
Behavior and Human Decision Processes, 58, 101–135. Ferris, S. H. (1992). Diagnosis by specialists: Psychological testing. Acta Neurologica Scandinavica, 85, 32–35. Finholt, T., & Olson, G. (1997). From laboratories to col-laboratories: A new organizational form for scienti�ic collaboration. Psychological Science, 8, 28–36. Finn, S. E. (1996). A manual for using the MMPI-2 as a therapeutic intervention. Minneapolis: University of Minnesota Press. Finn, S. E., & Tonsager, M. E. (1992). Therapeutic effects of providing MMPI-2 feedback to college students awaiting therapy. Psychological Assessment, 9, 374–
385. Finn, S. E., & Tonsager, M. E. (1992). Therapeutic effects of providing MMPI-2 test feedback to college students awaiting therapy. Psychological Assessment, 4,
278–287. Finn, S. E., & Tonsager, M. E. (1997). Information-gathering and therapeutic models of assessment: Complementary paradigms. Psychological Assessment, 9, 374–
385. Fiorello, C. A., & Primerano, D. (2005). Research into practice: Cattell-Horn-Carroll cognitive assessment in practice: Eligibility and program development issues.
Psychology in the Schools, 42, 525–536. First, M., & Gibbon, M. (2004). The structured clinical interview for DSM-IV axis I disorders (SCID-I) and the structured clinical interview for DSM-IV axis II
disorders (SCID-II). In M. Hilsenroth & D. Segal (Eds.), Comprehensive handbook of psychological assessment, Vol 2: Personality Assessment (pp. 134–143). Hoboken, NJ: John Wiley.
Fish, J. M. (Ed.). (2002). Race and intelligence: Separating science from myth. Mahwah, NJ: Erlbaum. Fisher, S., & Greenberg, R. P. (1984). The scienti�ic credibility of Freud’s theories and therapy. New York: Columbia University Press. Fiske, D. W. (1986). The trait concept and the personality questionnaire. In A. Angleitner & J. S. Wiggins (Eds.), Personality assessment via questionnaires: Current
issues in theory and measurement. Berlin: Springer-Verlag. Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51, 327–358. Flanagan, J. C. (1956). The evaluation of methods in applied psychology and the problem of criteria. Occupational Psychology, 30, 1–9. Flanagan, R., & di Guiseppe, R. (1999). Critical review of the TEMAS: A step within the development of thematic apperception instruments. Psychology in the
Schools, 36, 21–30. Flavell, J. (1976). Metacognitive aspects of problem-solving. In L. Resnick (Ed.), The nature of intelligence. Hillsdale, NJ: Erlbaum. Fletcher, J., & Vaughn, S. (2009). Response to intervention: Preventing and remediating academic dif�iculties. Child Development Perspectives, 3, 30–37. Floyd, R. G., Evans, J. J., & McGrew, K. S. (2003). Relations between measures of Cattell-Horn-Carroll (CHC) cognitive abilities and mathematics achievement
across the school-age years. Psychology in the Schools, 40, 155–171.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 43/70
Flynn, J. R. (1984). The mean IQ of Americans: Massive gains 1932 to 1978. Psychological Bulletin, 95, 29–51. Flynn, J. R. (1987). Massive IQ gains in 14 nations: What IQ tests really measure. Psychological Bulletin, 101, 171–191. Flynn, J. R. (1994). IQ gains over time. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Flynn, J. R. (2007a). What is intelligence: Beyond the Flynn effect. Cambridge: Cambridge University Press. Flynn, J. R. (2007b). Solving the IQ puzzle. Scienti�ic American Mind, 18, 25–31. Flynn, J. R., & Rossi-Casé, L. (2012). IQ gains in Argentina between 1964 and 1998. Intelligence, 40, 145–150. Folstein, M., Folstein, S., & McHugh, P. (1975). Mini-Mental State: A practical method for grading the cognitive state of patients for the clinician. Journal of
Psychiatric Research, 12, 189–198. Fonseca, R., Scherer, L., de Oliveira, C., & others. (2009). Hemisphere specialization for communicative processing: Neuroimaging data on the role of the right
hemisphere. Psychology and Neuroscience, 2, 25–33. Forbey, J., & Ben-Porath, Y. (2002). Use of the MMPI-2 in the treatment of offenders. International Journal of Offender Therapy and Comparative Criminology, 46,
308–318. Forbey, J., & Ben-Porath, Y. (2007). Computerized adaptive personality testing: A review and illustration with the MMPI-2 computerized adaptive version.
Psychological Assessment, 19, 14–24. Forrest, D. W. (1974). Francis Galton: The life and work of a Victorian genius. New York: Taplinger Publishing. Forster, A. A. (1994). Learning Disability. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Fowler, R. D. (1985). Landmarks in computer-assisted psychological assessment. Journal of Consulting and Clinical Psychology, 53, 748–759. Frank, G. (1983). The Wechsler enterprise: An assessment of the development, structure, and use of the Wechsler tests of intelligence. New York: Pergamon Press. Frank, G. (1990). Research on the clinical usefulness of the Rorschach: 1. The diagnosis of schizophrenia. Perceptual and Motor Skills, 71, 573–578. Frank, L. K. (1939). Projective methods for the study of personality. Journal of Psychology, 8, 389–413. Frank, L. K. (1948). Projective methods. Spring�ield, IL: Thomas. Franke, W. (1963). The reform and abolition of the traditional Chinese examination system. Cambridge, MA: Harvard University Press. Frankenburg, W. K. (1985). The Denver approach to early case �inding: A review of the Denver Developmental Screening Test and a brief training program in
developmental diagnosis. In W. K. Frankenburg, R. M. Emde, & J. W. Sullivan (Eds.), Identi�ication of children at risk: An international perspective. New York: Plenum Press.
Frankenburg, W. K., & Dodds, J. B. (1967). The Denver developmental screening tests. Journal of Pediatrics, 71, 181–191. Frankenburg, W. K., Dodds, J., Archer, P., & others. (1990). Denver II: Technical manual. Denver, CO: Denver Developmental Materials. Frankl, V. (1963). Man’s search for meaning: An introduction to logotherapy. New York: Washington Square Press. Frauenheim, J. G., & Heckerl, J. R. (1983). A longitudinal study of psychological and achievement test performance in severe dyslexic adults. Journal of Learning
Disabilities, 16, 339–347. Frechtling, J. A. (1989). Administrative uses of school testing programs. In R. L. Linn (Ed.), Educational measurement (3rd ed.). New York: American Council on
Education/Macmillan. Frederickson, L. C. (1985). Goodenough-Harris Drawing Test. In D. J. Keyser & R. C. Sweetland (Eds.). Test critiques (vol. 2). Kansas City, MO: Test Corporation of
America. Frederiksen, N. (1962). Factors in In-basket Performance. Psychological Monographs, 76, Whole No. 541. Freud, A. (1946). The ego and the mechanisms of defense. New York: International Universities Press. Freud, S. (1900). The interpretation of dreams. In J. Strachey, (Ed., in collaboration with A. Freud). The standard edition of the complete psychological works of
Sigmund Freud. London: Hogarth, 1955, vols. 4 and 5. Freud, S. (1927/1961). The future of an illusion (J. Strachey, trans.). New York: Basic Books. (Originally published 1900). Freud, S. (1933). New introductory lectures on psychoanalysis. New York: Norton. Frey, M. C., & Detterman, D. K. (2004). Scholastic Assessment or g? The relationship between the scholastic assessment test and general cognitive ability.
Psychological Science, 15, 373–378. Fridlund, A. J., Ekman, P., & Oster, H. (1987). Facial expressions of emotion: Review of literature 1970–1983. In A. W. Siegman & S. Feldstein (Eds.), Nonverbal
behavior and communication (2nd ed.). Hillsdale, NJ: Erlbaum. Friedman, A. F. (1987). Eysenck Personality Questionnaire. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques compendium. Kansas City, MO: Test Corporation
of America. Friedman, T. L. (2009). The world is �lat 3.0: A brief history of the twenty-�irst century. New York: Picador. Fuchs, D., & Fuchs, L. (2005). Responsiveness-to-intervention: A blueprint for practitioners, policymakers, and parents. Teaching Exceptional Children, 38, 57–61. Fuess, C. M. (1950). The College Board: Its �irst �ifty years. New York: Columbia University Press. Fuld, P. A. (1977). Fuld Object-Memory Evaluation. Chicago: Stoelting Co. Fuld, P. A., Masur, D. M., Blau, A. D., Crystal, H., & Aronson, M. K. (1990). Object-Memory Evaluation for prospective detection of dementia in normal functioning
elderly: Predictive and normative data. Journal of Clinical and Experimental Neuropsychology, 12, 520–528. Fuller, G. B., Parmelee, W. M., & Carroll, J. L. (1982). Performance of delinquent and nondelinquent highschool boys on the Rotter Incomplete Sentences Blank.
Journal of Personality Assessment, 46, 506–510. Funder, D. C. (2009). Naıv̈e and obvious questions. Perspectives on Psychological Science, 4, 340–344. Fuqua, D. R., & Newman, J. L. (1994). An evaluation of the Career Beliefs Inventory. Journal of Counseling and Development, 72, 429–430. Furnham, A., Batey, M., Anand, K., & Man�ield, J. (2008). Personality, hypomania, intelligence and creativity. Personality and Individual Differences, 44, 1060–1069. Furnham, A., Moutari, J., & Crump, J. (2003). The relationship between the revised NEO-Personality Inventory and the Myers-Briggs Type Indicator. Social
Behavior and Personality, 31, 577–584. Furnham, A., Toop, A., Lewis, C., & Fisher, A. (1995). P-E �it and job satisfaction: A failure to support Holland’s theory in three British samples. Personality and
Individual Differences, 19, 677–690. Galton, F. (1879). Psychometric experiments. Brain, 2, 149–162. Galton, F. (1883). Inquiries into human faculty and its development. London: Macmillan. Galton, F. (1888). Natural inheritance. London: Macmillan. Garb, H. N. (1994). Judgment research: Implications for clinical practice and testimony in court. Applied and Preventive Psychology, 3, 173–183.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 44/70
Garb, H. N., Florio, C., & Grove, W. (1998). The validity of the Rorschach and the MMPI: Results from meta-analyses. Psychological Science, 9, 402–404. Gardner, H. (1983). Frames of mind: The theory of multiple intelligence. New York: Basic Books. Gardner, H. (1986). The waning of intelligence tests. In R. J. Sternberg & D. K. Detterman (Eds.), What is intelligence? Contemporary viewpoints on its nature and
de�inition. Norwood, NJ: Ablex. Gardner, H. (1992). Assessment in context: The alternative to standardized testing. In B. R. Gifford & M. C. O’Connor (Eds.), Alternative views of aptitude,
achievement, and instruction. Boston: Klummer. Gardner, H. (1993). Multiple intelligences: The theory in practice. New York: Basic Books. Gardner, H. (1998). Are there additional intelligences? The case for naturalistic, spiritual, and existential intelligences. In J. Kane (Ed.), Education, information,
and transformation. Englewood Cliffs, NJ: Prentice Hall. Gardner, J. (1967). The adjustment of drug addicts as measured by the sentence completion test. Journal of Projective Techniques and Personality Assessment, 31,
28–29. Gardner, R. A. (1981). Digits forward and digits backward as two separate tests: Normative data on 1567 school children. Journal of Clinical Child Psychology, 10,
131–135. Gast, J., & Hart, K. J. (2010). The performance of juvenile offenders on the Test of Memory Malingering. Journal of Forensic Psychology Practice, 10, 53–68. Gaudry, E., Vagg, P., & Spielberger, C. D. (1975). Validation of the state-trait distinction in anxiety research. Multivariate Behavioral Research, 10, 331–341. Gavett, B. E., Lou, K. R., Daneshvar, D. H., & others. (2012). Diagnostic accuracy statistics for seven Neuro-psychological Assessment Battery (NAB) test variables
in the diagnosis of Alzheimer’s disease. Applied Neuro-psychology, 19, 108–115. Gazzaniga, M. S. (1970). The bisected brain. New York: Appleton-Century-Crofts. Gazzaniga, M. S., & LeDoux, J. E. (1978). The integrated mind. New York: Plenum Press. Geary, D. C., & Whitworth, R. H. (1988). Is the factor structure of the WISC-R different for Anglo- and Mexican-American children? Journal of Psychoeducational
Assessment, 6, 253–260. GED Testing Service (1991). Examiner’s manual: Test of General Educational Development. Washington, DC: GED Testing Service of the American Council on
Education. Gelb, S. (1986). Henry H. Goddard and the immigrants, 1910–1917: The studies and their social context. Journal of the History of the Behavioral Sciences, 22, 324–
332. George, C., & Solomon, J. (1999). Attachment and care-giving: The caregiving behavioral system. In J. Cassidy & P. Shaver (Eds.), Handbook of attachment: Theory,
research and clinical application (pp. 649–670). New York: Guilford Press. Gerard, A. B. (1993). Manual for Parent–Child Relationship Inventory. Los Angeles: Western Psychological Services. Geschwind, N. (1972). Language and the brain. Scienti�ic American, 226, 76–83. Geschwind, N., & Galaburda, A. M. (1987). Cerebral lateralization: Biological mechanisms, associations, and pathology. Cambridge, MA: MIT Press. Getz, I. R. (1984). Moral judgment and religion: A review of the literature. Counseling and Values, 28, 94–116. Ghez, C. (1991). The cerebellum. In E. R. Kandel, J. H. Schwartz, & T. M. Jessell (Eds.), Principles of neural science (3rd ed.). New York: Elsevier. Ghiselli, E. E. (1966). The validity of occupational aptitude tests. New York: Wiley. Ghiselli, E. E., Campbell, J. P., & Zedeck, S. (1981). Measurement theory for the behavioral sciences. San Francisco: W. H. Freeman. Gibbons, R., Weiss, D., Kupfer, D., Frank, E., Fagiolini, A., & others. (2008). Using computerized adaptive testing to reduce the burden of mental health assessment.
Psychiatric Services, 59, 361–368. Gifford, R. (1991). Applied psychology: Variety and opportunity. Boston: Allyn and Bacon. Gignac, G. (2006). A con�irmatory examination of the factor structure of the Multidimensional Aptitude Battery: Contrasting oblique, higher order, and nested
factor models. Educational and Psychological Measurement, 66, 136–145. Gilberstadt, H., & Duker, J. (1965). A handbook for clinical and actuarial MMPI interpretation. Philadelphia: W. B. Saunders. Glascoe, F. P. (1991). Developmental screening: Rationale, methods and application. Infants and Young Children, 4, 1–10. Glascoe, F. P. (2005). Commonly used screening tests. Retrieved from www.dbpeds.org/articles (http://www.dbpeds.org/articles) on September 2, 2005. Glascoe, F. P., & Byrne, K. E. (1993). The accuracy of three developmental screening tests. Journal of Early Intervention, 17, 368–379. Glascoe, F. P., & Shapiro, H. (2005). Introduction to developmental and behavioral screening. Retrieved from www.dbpeds.org/articles
(http://www.dbpeds.org/articles) on September 2, 2005. Goddard, H. H. (1910a). A measuring scale for intelligence. The Training School, 6, 146–155. Goddard, H. H. (1910b). Four hundred feebleminded children classi�ied by the Binet method. Pedagogical Seminary, 17, 387–397. Goddard, H. H. (1911). Two thousand normal children measured by the Binet measuring scale of intelligence. Pedagogical Seminary, 18, 232–259. Goddard, H. H. (1912). Feeble-mindedness and immigration. Training School Bulletin, 9, 91. Goddard, H. H. (1917). The mental level of a group of immigrants. Psychological Bulletin, 14, 68–69. Goddard, H. H. (1919). Psychology of the normal and sub-normal. New York: Dodd, Mead, and Co. Goddard, H. H. (1928). Feeblemindedness: A question of de�inition. Journal of Psycho-Asthenics, 33, 219–227. Gof�in, R. D., Rothstein, M., & Johnston, N. (1996). Personality testing and the assessment center: Incremental validity for managerial selection. Journal of Applied
Psychology, 81, 746–756. Gof�in, R., Rothstein, M., & Johnston, N. (2000). Predicting job performance using personality constructs: Are personality tests created equal? In R. Gof�in & E.
Helmes (Eds.), Problems and solutions in human assessment: Honoring Douglas N. Jackson at seventy. New York: Kluwer Academic/Plenum Publishers. Goldberg, L. R. (1965). Diagnosticians vs. diagnostic signs: The diagnosis of psychosis vs. neurosis from the MMPI. Psychological Monographs, 79 (9, Whole No.
602). Goldberg, L. R. (1981a). Developing a taxonomy of trait-descriptive terms. In D. Fiske (Ed.), New directions for methodology of social and behavioral science:
Problems with language imprecision (no. 9). San Francisco: Jossey-Bass. Goldberg, L. R. (1981b). Language and individual differences: The search for universals in personality lexicons. In L. Wheeler (Ed.), Review of personality and
social psychology. Beverly Hills, CA: Sage. Goldberg, L. R. (1990). An alternative “description of personality”: The big-�ive factor structure. Journal of Personality and Social Psychology, 59, 1216–1229. Golden, C. (2004). The Adult Luria-Nebraska Neuro-psychological Battery. In G. Goldstein, S. Beers, & M. Hersen (Eds.), Intellectual and neuropsychological
assessment (pp. 133–146). Hoboken, NJ: Wiley. Golden, C. J., Purish, A. D., & Hammeke, T. A. (1980). Luria-Nebraska Neuropsychological Battery: Manual. Los Angeles: Western Psychological Services.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 45/70
Golden, C. J., Purish, A. D., & Hammeke, T. A. (1986). Luria-Nebraska Neuropsychological Battery: Forms I and II. Los Angeles: Western Psychological Services. Goldfried, M. R., & Zax, M. (1965). The stimulus value of the TAT. Journal of Projective Techniques, 29, 46–57. Golding, S. (1993). Interdisciplinary Fitness Interview—Revised: A training manual. State of Utah Division of Mental Health. Goldstein, I. L. (1991). Training in work organizations. In M. D. Dunnette & L. M. Hough (Eds.), Handbook of industrial and organizational psychology (vol. 2). Palo
Alto, CA: Consulting Psychologists Press. Goldstein, I. L. (1992). Training (3rd ed.). Monterey, CA: Brooks/Cole. Goldstein, K. (1944). The mental changes due to frontal lobe damage. Journal of Psychology, 17, 187–208. Goleman, D. (1995). Emotional intelligence: Why it can matter more than IQ. New York: Bantam. Goodenough, F. L. (1926). Measurement of intelligence by drawings. New York: Harcourt, Brace & World. Goodenough, F. L. (1949). Mental testing: Its history, principles, and applications. New York: Rinehart. Goodglass, H., Kaplan, E., & Barresi, B. (2000). Boston Diagnostic Aphasia Examination (3rd ed.). Austin, TX: The Psychological Corporation. Goodman, J. (1990). Infant intelligence: Do we, can we, should we assess it? In C. C. Reynolds & R. W. Kamphaus (Eds.), Handbook of psychological and educational
assessment of children: Intelligence and achievement. New York: Guilford. Gordon, G., & Charanian, T. (1964). Measuring the creativity of research scientists and engineers. Working paper cited in I. A. Taylor and J. W. Getzels (Eds.),
Perspectives in creativity. Chicago: Aldine. Gordon, M., & Keiser, S. (1998). Accommodations in higher education under the Americans with Disabilities Act (ADA). DeWitt, NY: GSI Publications. Gordon, M., & Mettelman, B. B. (1988). The assessment of attention: I. Standardization and reliability of a behavior-based measure. Journal of Clinical Psychology,
44, 688–690. Goslin, D. A. (1963). The search for ability: Standardized testing in social perspective. New York: Russell Sage Foundation. Gothard, S., Viglione, D., Meloy, J. R., & Sherman, M. (1996). Detection of malingering in competency to stand trial evaluations. Law and Human Behavior, 19, 493–
505. Gottfredson, G. D., & Holland, J. L. (1975). Vocational choices of men and women: A comparison of predictors from the Self-Directed Search. Journal of Counseling
Psychology, 22, 28–34. Gottfredson, G. D., & Holland, J. L. (1989). Dictionary of Holland Occupational Codes (2nd ed.). Odessa, FL: Psychological Assessment Resources. Gottfredson, L. S. (2005). Using Gottfredson’s theory of circumscription and compromise in career guidance and counseling. In S. D. Brown & R. W. Lendt (Eds.),
Career development and counseling: Putting theory and research to work (pp. 71–100). New York: John Wiley & Sons. Gough, H. (1995). Career assessment and the California Psychological Inventory. Journal of Career Assessment, 3, 101–122. Gough, H. G. (1984). A managerial potential scale for the California Psychological Inventory. Journal of Applied Psychology, 69, 233–244. Gough, H. G. (1987). California Psychological Inventory manual. Palo Alto, CA: Consulting Psychologists Press. Gough, H. G., & Bradley, P. (1992a). Comparing two strategies for developing personality scales. In M. Zeidner & R. Most (Eds.), Psychological testing: An inside
view. Palo Alto, CA: Consulting Psychologists Press. Gough, H. G., & Bradley, P. (1992b). Delinquent and criminal behavior as assessed by the Revised California Psychological Inventory. Journal of Clinical Psychology,
48, 298–307. Gough, H. G., & Bradley, P. (1996). CPI manual (3rd ed.). Mountain View, CA: Consulting Psychologists Press. Gould, S. J. (1981). The mismeasure of man. New York: Norton. Gow, A. J., Johnson, W., Pattie, A., & others. (2011). Stability and change in intelligence from age 11 to ages 70, 79, and 87: The Lothian Birth Cohorts of 1921 and
1936. Psychology and Aging, 26, 232–240. Grace, J., & Malloy, P. F. (2001). Frontal Systems Behavior Scale professional manual. Lutz, FL: Psychological Assessment Resources. Graham, J. (1961). Lavater’s physiognomy in England. Journal of the History of Ideas, 22, 561–572. Graham, J. R. (1987). The MMPI: A practical guide (2nd ed.). New York: Oxford University Press. Graham, J. R. (1993). MMPI-2: Assessing personality and psychopathology. New York: Oxford. Graham, J. R. (2000). MMPI-2: Assessing personality and psychopathology (3rd ed.). New York: Oxford University Press. Granstrom, S. L. (1987). A comparative study of loneliness, Buberian religiosity and spiritual well-being in cancer patients. Paper presented at the conference of the
National Hospice Organization. Gray, B. (2001). A factor analytic study of the Substance Abuse Subtle Screening Inventory (SASSI). Educational and Psychological Measurement, 61, 102–118. Green, D., & Rosenfeld, B. (2011). Evaluating the gold standard: A review and meta-analysis of the Structured Interview of Reported Symptoms. Psychological
Assessment, 23, 95–107. Greenough, W. T., Black, J. E., & Wallace, C. S. (1987). Experience and brain development. Child Development, 58, 539–559. Gregory, R. J. (1987). Adult intellectual assessment. Boston: Allyn and Bacon. Gregory, R. J. (1994a). Aptitude tests. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Gregory, R. J. (1994b). Pro�ile interpretation. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Gregory, R. J. (1994c). Classi�ication of intelligence. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Gregory, R. J. (1998). Testing in clinical psychology. In S. Cullari (Ed.), Foundations of clinical psychology. Boston: Allyn and Bacon. Gregory, R. J. (1999). Foundations of intellectual assessment: The WAIS-III and other tests in clinical practice. Boston: Allyn and Bacon. Gregory, R. J. (2009). Testing bias. In I. Weiner & E. Craighead (Eds.), Corsini’s encyclopedia of psychology. New York: Wiley. Gregory, R. J., & Gernert, C. H. (1990). Age trends for �luid and crystallized intelligence in an able subpopulation. Unpublished manuscript. Gresham, F. M. (1993). “What’s wrong in this picture?”: Response to Motta et al.’s review of human �igure drawings. School Psychology Quarterly, 8, 182–186. Greve, K., Love, J., Sherwin, E., & others. (2002). Temporal stability of the Wisconsin Card Sorting Test in a chronic traumatic brain injury sample. Assessment, 9,
271–277. Grös, D. F., Antony, M. M., Simms, L. J., & McCabe, R. E. (2007). Psychometric properties of the State-Trait Inventory for Cognitive and Somatic Anxiety (STICSA):
Comparison to the State-Trait Anxiety Inventory (STAI). Psychological Assessment, 19, 369–381. Grossman, S. A., Richards, C., Anglin, D., & Hutson, H. (2000). Caring for the patient with mental retardation in the ED. Annals of Emergency Medicine, 35, 69–76. Groth-Marnat, G. (1997). Handbook of psychological assessment (2nd ed.). New York: Wiley. Groth-Marnat, G. (2003). Handbook of psychological assessment (4th ed.). New York: Wiley.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 46/70
Grove, W., & Barden, C. (1999). Protecting the integrity of the legal system: The admissibility of testimony from mental health experts under Daubert/Kumho analyses. Psychology, Public Policy, and Law, 5, 224–242.
Grove, W., Barden, C., Garb, H., & Lilienfeld, S. (2002). Failure of Rorschach-Comprehensive-System-Based testimony to be admissible under the Daubert-Joiner- Kumho standard. Psychology, Public Policy, and Law, 8, 216–234.
Grove, W., Zald, D., Lebow, B., Snitz, B., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12, 19–30. Guaiana, G., Tyson, P., & Mortimer, A. (2004). The Rivermead Behavioural Memory Test can predict social functioning among schizophrenic patients treated with
clozapine. International Journal of Psychiatry in Clinical Practice, 8, 245–249. Gudjonsson, G. H. (1995). The Standard Progressive Matrices: Methodological problems associated with the administration of the 1992 adult standardisation
sample. Personality and Individual Differences, 18, 441–442. Guilford, J. P. (1954). Psychometric methods. New York: McGraw-Hill. Guilford, J. P. (1959). Personality. New York: McGraw-Hill. Guilford, J. P. (1967). The nature of human intelligence. New York: McGraw-Hill. Guilford, J. P. (1985). The Structure-of-Intellect model. In B. B. Wolman (Ed.), Handbook of intelligence: Theories, measurements, and applications. New York: Wiley. Guilford, J. P., & Fruchter, B. (1978). Fundamental statistics in psychology and education (6th ed.). New York: McGraw-Hill. Guilford, J. P., & Guilford, J. S. (1980). Christensen-Guilford Fluency Tests. Orange, CA: Sheridan Psychological Services. Guilford, J. P., & Hoepfner, R. (1971). The analysis of intelligence. New York: McGraw-Hill. Guion, R. (1998). Assessment, measurement, and prediction for personnel decisions. Mahwah, NJ: Erlbaum. Gulliksen, H. (1950). Theory of mental tests. New York: Wiley. Gunning, M. D., Denison, F. C., Stockley, C. J., & others. (2010). Assessing maternal anxiety in pregnancy with the State-Trait Anxiety Inventory: Issues of validity,
location and participation. Journal of Reproductive and Infant Psychology, 28, 266–273. Gutkin, R. B., & Reynolds, C. R. (1981). Factorial similarity of the WISC-R for white and black children from the standardization sample. Journal of Educational
Psychology, 73, 227–231. Guttman, L. (1944). A basis for scaling qualitative data. American Sociological Review, 9, 139–150. Guttman, L. (1947). The Cornell technique for scale and intensity analysis. Educational and Psychological Measurement, 7, 247–280. Gynther, M. D., & Gynther, R. A. (1976). Personality inventories. In I. B. Weiner (Ed.), Clinical methods in psychology. New York: Wiley. Haaland, K. Y., & Delaney, H. D. (1981). Motor de�icits after left or right hemisphere damage due to stroke or tumor. Neuropsychologia, 19, 17–27. Haber, A., & Fichtenberg, N. (2006). Replication of the Test of Memory Malingering (TOMM) in a traumatic brain injury and head trauma sample. The Clinical
Neuropsychologist, 20, 524–532. Hachinski, V. C., Iliff, L., Zilha, E., & others. (1975). Cerebral blood �low in dementia. Archives of Neurology, 32, 632–637. Hack, M., Taylor, G., Drotar, D., & others. (2005). Poor predictive validity of the Bayley Scales of Infant Development for cognitive function of extremely low birth
weight children at school age. Pediatrics, 116, 333–341. Haedt-Matt, A. A., & Keel, P. K. (2011). Revisiting the affect regulation model of binge eating: A meta-analysis of studies using ecological momentary assessment.
Psychological Bulletin, 37, 660–681. Hain, J. (1964). The Bender-Gestalt Test: A scoring method for identifying brain damage. Journal of Consulting and Clinical Psychology, 28, 34–40. Haladyna, T. M. (1992). Review of the Millon Clinical Multiaxial Inventory-II. Eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Hale, J., Fiorello, C., Dumont, R., & others. (2008). “Differential Ability Scales-Second Edition”: (Neuro) Psychological predictors of math performance for typical
children and children with math disorders. Psychology in the Schools, 45, 838–858. Hall, P., & Hall, D. (1983). The handshake as interaction. Semiotica, 45, 249–264. Hambleton, R. K. (1984). Validating the test scores. In R. A. Berk (Ed.), A guide to criterion-referenced test construction. Baltimore: Johns Hopkins University Press. Hambleton, R. K. (1989). Principles and selected applications of item response theory. In R. L. Linn (Ed.), Educational measurement (3rd ed.). New York:
American Council on Education/Macmillan. Hambleton, R. K., & Zenisky, A. (2003). Advances in criterion-referenced testing methods and practices. In C. R. Reynolds & R. W. Kamphaus (Eds.), Handbook of
psychological and educational assessment of children: Intelligence, aptitude, and achievement (2nd ed., pp. 377–404). New York: Guilford Press. Hammill, D. D. (1999). Detroit Tests of Learning Aptitude-4 (DTLA-4). Austin, TX: PRO-ED. Handler, L., & Clemence, A. (2005). The Rorschach Prognostic Rating Scale. In R. F. Bornstein & J. M. Masling (Eds.), Scoring the Rorschach: Seven validated
systems. Mahwah, NJ: Erlbaum. Hansen, J. (2007). Evidence of validity for the skill scale scores of the Campbell Interest and Skill Survey. Journal of Vocational Behavior, 71, 23–44. Hansen, J. C. (1992). Strong user’s guide, Revised edition. Palo Alto, CA: Consulting Psychologists Press. Hansen, J. C., & Campbell, D. P. (1985). Manual for the Strong Interest Inventory Form T325 of the Strong Vocational Interest Blanks, Fourth Edition. Stanford, CA:
Stanford University Press. Hansen, J.-I., & Neuman, J. (1999). Evidence of concurrent prediction of the Campbell Interest and Skill Survey (CISS) for college major selection. Journal of Career
Assessment, 7, 239–247. Hanson, G. A. (1991). To catch a thief: The legal and policy implications of honesty testing in the workplace. Law and Inequality, 9, 497–531. Hanzel, E. P. (2003). Assessment of cognitive abilities in high-functioning children with autistic disorder: A comparison of the WISC-III and the Leiter-R.
Dissertation Abstracts International: Section B: The Sciences and Engineering, 64(3-B), 1492. Hare, R. D. (1996). Psychopathy: A clinical construct whose time has come. Criminal Justice and Behavior, 23, 25–54. Hare, R. D. (2003). The Hare Psychopathy Checklist-Revised (PCL-R) (2nd ed.). Toronto: Multi-Health Systems. Hare, R. D., & McPherson, L. (1984). Violent and aggressive behavior by criminal psychopaths. International Journal of Law and Psychiatry, 7, 35–50. Hare, R. D., Harpur, T., & Hakstian, R., & others. (1990). The Revised Psychopathy Checklist: Descriptive statistics, reliability, and factor structure. Psychological
Assessment: A Journal of Consulting and Clinical Psychology, 1, 6–17. Hare, R., & Neuman, C. (2006). The PCL-R assessment of psychopathy: Development, structural properties, and new directions. In C. Patrick (Ed.), Handbook of
psychopathy (pp. 58–88). New York: Guilford. Hargrave, G., & Hiatt, D. (1989). Use of the California Psychological Inventory in law enforcement of�icer selection. Journal of Personality Assessment, 53, 267–277. Hargrave, G., Hiatt, D., Ogard, E., & Karr, C. (1994). Comparison of the MMPI and the MMPI-2 for a sample of peace of�icers. Psychological Assessment, 6, 27–32. Harmon, L. W. (1989). Counseling. In R. L. Linn (Ed.), Educational measurement (3rd ed.). New York: American Council on Education/Macmillan.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 47/70
Harmon, L. W., Hansen, J. C., Borgen, F., & Hammer, A. (1994). Strong Interest Inventory applications and technical guide. Palo Alto, CA: Consulting Psychologists Press.
Harrell, T. H., Honaker, L., & Parnell, T. (1992). Equivalence of the MMPI-2 with the MMPI in psychiatric patients. Psychological Assessment, 4, 460–465. Harrington, D. M. (1975). Effect of explicit instructions to “be creative” on the psychological meaning of divergent thinking test scores. Journal of Personality, 43,
434–454. Harris, D. B. (1963). Children’s drawings as measures of intellectual maturity. New York: Harcourt, Brace & World. Harris, M. M., & Schaubroeck, J. (1988). A meta-analysis of self-supervisor, self-peer, and peer-supervisor ratings. Personnel Psychology, 38, 43–62. Harrison, D. A., & Hulin, C. L. (1989). Investigations of absenteeism: Using event history models to study the absence-taking process. Journal of Applied
Psychology, 74, 300–316. Harrison, D. A., & Shaffer, M. (1994). Comparative examinations of self-reports and perceived absenteeism norms: Wading through Lake Wobegon. Journal of
Applied Psychology, 79, 240–251. Harrison, P. L., & Schock, H. H. (1994). Draw-A-Figure test. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Hartung, P., Borges, N., & Jones, B. (2005). Using person matching to predict career specialty choice. Journal of Vocational Behavior, 67, 102–117. Hathaway, S. R., & McKinley, J. C. (1940). A multiphasic personality schedule (Minnesota): I. Construction of the schedule. Journal of Psychology, 10, 249–254. Hathaway, S. R., & McKinley, J. C. (1942). A multiphasic personality schedule (Minnesota): III. The measurement of symptomatic depression. Journal of
Psychology, 14, 73–84. Hathaway, S. R., & McKinley, J. C. (1943). The Minnesota Multiphasic Personality Inventory (rev. ed.). Minneapolis: University of Minnesota Press. Hawkins, D. B. (1988). Interpersonal behavior traits, spiritual well-being, and their relationships to blood pressure (doctoral dissertation, Western Conservative
Baptist Seminary, 1986). Dissertation Abstracts International, 48, 3680B. Hawkins, D. B., & Larson, R. (1984). The relationship between measures of health and spiritual well-being. Unpublished manuscript, Western Conservative Baptist
Seminary, Portland, OR. Hawkins, K. A., Faraone, S. V., Pepple, J. R., Seidman, L. J., & Tsuang, M. T. (1990). WAIS-R validation of the Wonderlic Personnel Test as a brief intelligence measure
in a psychiatric sample. Psychological Assessment: A Journal of Consulting and Clinical Psychology, 2, 198–201. Hawkins, K., Dean, D., & Pearlson, G. (2004). Alternative forms of the Rey Auditory Verbal Learning Test: A review. Behavioral Neurology, 15, 99–107. Hawthorne, J. (2009). Promoting development of the early parent-infant relationship using the Neonatal Behavioural Assessment Scale. In J. Barlow & P.
Svanberg (Eds.), Keeping the baby in mind: Infant mental health in practice. New York: Routledge/Taylor & Francis Group. Hayes, P. A. (2008). Addressing cultural complexities in practice: Assessment, diagnosis, and therapy (2nd ed.). Washington, DC: American Psychological
Association. Haynes, S. G., Feinleib, M., & Eaker, E. (1983). Type A behavior and the ten-year incidence of coronary heart disease in the Framingham heart study. In R. H.
Rosenman (Ed.), Psychosomatic risk factors and coronary heart disease. Bern, Switzerland: Huber. Haynes, S. N. (2001). Introduction to the special section on clinical applications of analogue behavioral observation. Psychological Assessment, 13, 3–4. Hayslip, B., & Panek, P. E. (1989). Adult development and aging. New York: Harper & Row. Heaton, R. K., Chelune, G., Talley, J., & others. (1993). Wisconsin Card Sorting Test manual: Revised and expanded. Odessa, FL: Psychological Assessment Resources. Heaton, R. K., Smith, H. H., Jr., Lehman, R. A. W., & Vogt, A. T. (1978). Prospects for faking believable de�icits on neuropsychological testing. Journal of Consulting
and Clinical Psychology, 46, 892–900. Hebb, D. O. (1939). Intelligence in man after large removals of cerebral tissue: Report of four left frontal lobe cases. Journal of General Psychology, 21, 73–87. Heilbrun, A. B., Jr., & Georges, M. (1990). The measurement of principled morality by the Kohlberg Moral Dilemma Questionnaire. Journal of Personality
Assessment, 55, 183–194. Heilbrun, K. (1992). The role of psychological testing in forensic assessment. Law and Human Behavior, 16, 257–272. Helms, J. E. (1992). Why is there no study of cultural equivalence in standardized cognitive ability testing? American Psychologist, 47, 1083–1101. Helson, R., & Soto, C. J. (2005). Up and down in middle age: Monotonic and nonmonotonic changes in roles, status, and personality. Journal of Personality and
Social Psychology, 89, 194–204. Helson, R., Kwan, V., John, O. P., & Jones, C. (2002). The growing evidence for personality change in adulthood: Findings from research with personality
inventories. Journal of Research in Personality, 36, 287–306. Hendriks, A., Hofstee, W., & De Raad, B. (1999). The Five-Factor Personality Inventory. Personality and Individual Differences, 27, 307–325. Herman, D. O. (1988). Blind Learning Aptitude Test. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques (vol. 5). Kansas City, MO: Test Corporation of America. Hernandez-Reif, M., Field, T., Diego, M., & Ruddock, M. (2006). Greater arousal and less attentiveness to face/voice stimuli by neonates of depressed mothers on
the Brazelton neonatal Behavioral Assessment Scale. Infant Behavior and Development, 29, 594–598. Herrnstein, R. J., & Murray, C. (1994). The bell curve: Intelligence and class structure in American life. New York: Free Press. Hersen, M., & Bellack, A. S. (Eds.). (1988). Dictionary of behavioral assessment techniques. New York: Pergamon. Herzberg, P., Glaesmer, H., & Hoyer, J. (2006). Separating optimism and pessimism: A robust psychometric analysis of the Revised Life Orientation Test (LOT-R).
Psychological Assessment, 18, 433–438. Heyman, R. (2001). Observation of couple con�licts: Clinical assessment applications, stubborn truths, and shaky foundations. Psychological Assessment, 13, 5–35. Hiatt, D., & Hargrave, G. E. (1988). MMPI pro�iles of problem peace of�icers. Journal of Personality Assessment, 52, 722–731. Higgs, M. (2001). Is there a relationship between the Myers-Briggs Type Indicator and emotional intelligence? Journal of Managerial Psychology, 16, 509–533. Highhouse, S., & Nolan, K. P. (in press). One history of the assessment center. In D. J. R. Jackson, C. E. Lance, & B. J. Hoffman (Eds.), The psychology of assessment
centers (pp. 25–44). New York: Routledge/Taylor & Francis Group. Hill, B. (2005). ICAP User’s Group Home Page. Retrieved from www.cpinternet.com/bhill/icap (http://www.cpinternet.com/bhill/icap) on September 13, 2005. Hill, P. C., & Hood, R. W. (Eds.). (1999). Measures of religiosity. Birmingham, AL: Religious Education Press. Hill, P. C., & Pargament, K. I. (2008). Advances in the conceptualization and measurement of religion and spirituality: Implications for physical and mental health
research. Psychology of Religion and Spirituality, S(1), 3–17. Hilliard, A. G. (1984). IQ testing as the emperor’s new clothes: A critique of Jensen’s Bias in Mental Testing. In C. R. Reynolds & R. T. Brown (Eds.), Perspectives on
bias in mental testing. New York: Plenum Press. Hintze, J., Volpe, R., & Shapiro, E. (2002). Best practices in the systematic direct observation of student behavior. In A. Thomas & J. Grimes (Eds.), Best practices in
school psychology IV. Washington, DC: National Association of School Psychologists. Hiskey, M. S. (1966). Manual for the Hiskey-Nebraska Test of Learning Aptitude. Lincoln, NE: Union College Press.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 48/70
Hofer, S., Sliwinski, M., & Flaherty, B. (2002). Understanding ageing: Further commentary on the limitations of cross-sectional designs for ageing research. Gerontology, 48, 22–29.
Hoffart, A., Friis, S., Strand, J., & Olsen, B. (1994). Symptoms and cognitions during situational and hyperventilatory exposure in agoraphobic patients with and without panic. Journal of Psychopathology and Behavioral Assessment, 16, 15–32.
Hoffman, F. J., Sheldon, K. L., Minskoff, E. H., & others. (1987). Needs of learning disabled adults. Journal of Learning Disabilities, 20, 43–52. Hofmann, S. G., & Reinecke, M. A. (2010). Cognitive-behavioral therapy with adults. A guide to empirically-informed assessment and intervention. New York:
Cambridge University Press. Hogan, A. E., Scott, K. G., & Bauer, C. R. (1992). The Adaptive Social Behavior Inventory (ASBI): A new assessment of social competence in high-risk three-year-
olds. Journal of Psychoeducational Assessment, 10, 230–239. Hogan, J., & Hogan, R. (1986). Manual for the Hogan Personnel Selection System. Minneapolis, MN: National Computer Systems. Hogan, R. (2002). The Hogan Personality Inventory. In B. de Raad & M. Perugini (Eds.), Big �ive assessment (pp. 329–346). Ashland, OH: Hogrefe and Huber. Hogan, R. T. (1986). Manual for the Hogan Personality Inventory. Minneapolis, MN: National Computer Systems. Hoge, C. W., McGurk, D., Thomas, J. L., & others. (2008). Mild traumatic brain injury in U.S. soldiers returning from Iraq. New England Journal of Medicine, 358(5),
453–463. Hoge, D. R. (1996). Religion in America: The demographics of belief and af�iliation. In E. P. Shafranske (Ed.), Religion and the clinical practice of psychology.
Washington, DC: American Psychological Association. Hoge, S., Bonnie, R., Poythress, N., & Monahan, J. (1999). The MacArthur Competence Assessment Tool—Criminal Adjudication. Odessa, FL: Psychological
Assessment Resources. Holland, J. L. (1959). A theory of vocational choice. Journal of Counseling Psychology, 6, 35–44. Holland, J. L. (1966). The psychology of vocational choice. Waltham, MA: Blaisdell. Holland, J. L. (1978). The occupations �inder. Palo Alto, CA: Consulting Psychologists Press. Holland, J. L. (1985). Vocational Preference Inventory (VPI) manual—1985 edition. Odessa, FL: Psychological Assessment Resources. Holland, J. L. (1985a). Making vocational choices: A theory of vocational personalities and work environments (2nd ed.). Englewood Cliffs, NJ: Prentice Hall. Holland, J. L. (1985b). Self-Directed Search: Professional manual—1985 edition. Odessa, FL: Psychological Assessment Resources. Holland, J. L. (1985c). Vocational Preference Inventory (VPI) manual—1985 edition. Odessa, FL: Psychological Assessment Resources. Holland, J. L. (1987). 1987 manual supplement for the Self-Directed Search. Odessa, FL: Psychological Assessment Resources. Holland, J. L., Johnston, J., Asama, N. F., & Polys, S. (1993). Validating and using the Career Beliefs Inventory. Journal of Career Development, 19, 233–244. Hollander, E., Kolevzon, A., & Coyle, J. (2011). Textbook of autism spectrum disorders. Washington, DC: American Psychiatric Publishing. Hollingshead, A., & Redlich, F. (1958). Social class and mental illness. New York: Wiley. Hollingworth, H.L. (1943). Leta Stetter Hollingworth. Lincoln, NE: University of Nebraska Press. Hollingworth, L. (1914). Variability as related to sex differences in achievement: A critique. American Journal of Sociology, 19, 510–530. Hollingworth, L. (1928). Children clustering at 165 IQ and children clustering at 146 IQ compared for three years in achievement. In G. Whipple (Ed.), The
twenty-seventh yearbook of the National Society for the Study of Education: Nature and nurture, Part II—Their in�luence upon achievement. Bloomington, IL: Public School Publishing.
Hollingworth, L. (1935). The comparative beauty of the faces of highly intelligent adolescents. Journal of Genetic Psychology, 47, 268–281. Hollingworth, L., & Monahan, J. (1926). Tapping-rate of children who test above 135 IQ (Stanford-Binet). Journal of Educational Psychology, 17, 505–518. Holmes, T., & Rahe, R. (1967). The Social Readjustment Rating Scale. Journal of Psychosomatic Research, 11, 213–218. Holzinger, K. J., & Swineford, F. (1939). A study in factor analysis: The stability of a bi-factor solution. University of Chicago, Supplementary Educational
Monographs, No. 48. Holzinger, K., & Harman, H. (1941). Factor analysis: A synthesis of factorial methods. Chicago: University of Chicago Press. Holzman, P., Levy, D., & Johnston, M. H. (2005). The use of the Rorschach technique for assessing formal thought disorder. In R. F. Bornstein & J. M. Masling (Eds.),
Scoring the Rorschach: Seven validated systems. Mahwah, NJ: Erlbaum. Honzik, M. (1957). Developmental studies of parent-child resemblance in intelligence. Child Development, 28, 215–228. Hooper, S., Hatton, D., Baranek, G., Roberts, J., & Bailey, D. (2000). Nonverbal assessment of IQ, attention, and memory abilities in children with fragile-X
syndrome using the Leiter-R. Journal of Psychoeducational Assessment, 18, 255–267. Horn, J. L. (1968). Organization of abilities and the development of intelligence. Psychological Review, 75, 242–259. Horn, J. L. (1985). Remodeling old models of intelligence. In B. B. Wolman (Ed.), Handbook of intelligence: Theories, measurements, and applications. New York:
Wiley. Horn, J. L. (1994). Theory of �luid and crystallized intelligence. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Horn, J. L., & Cattell, R. B. (1966). Re�inement and test of the theory of �luid and crystallized general intelligences. Journal of Educational Psychology, 57, 253–270. Horn, J. L., & Masunaga, H. (2000). New directions for research into aging and intelligence: The development of expertise. In T. J. Perfect & E. A. Maylor (Eds.),
Models of cognitive aging (pp. 125–159). Oxford, England: Oxford University Press. Horton, A. (2008). The Halstead-Reitan Neuropsychological Test Battery: Past, present, and future. In A. Horton & D. Wedding (Eds.), The neuropsychology
handbook (3rd ed.) (pp. 251–278). New York: Springer. Hough, L. M., Eaton, N., Dunnette, M., Kamp, J., & McCloy, R. (1990). Criterion-related validities of personality constructs and the effect of response distortion on
those validities [Monograph]. Journal of Applied Psychology, 75, 581–595. Howell, R. J., & Richards, L. (1989). Review of the Rogers Criminal Responsibility Assessment Scales. The tenth mental measurements yearbook. Lincoln:
University of Nebraska Press. Huffcutt, A. (2007). Employment interviews. In D. Whetzel & G. Wheaton (Eds.), Applied measurement: Industrial psychology in human resources management (pp.
181–199). New York: Taylor & Francis/Erlbaum. Huffcutt, A. I., & Roth, P. (1998). Racial group differences in employment interview evaluations. Journal of Applied Psychology, 83, 179–189. Hughes, J. L., & McNamara, W. J. (1959). Manual for the revised Programmer Aptitude Test. New York: The Psychological Corporation. Human Rights Watch. (2001). Beyond reason: The death penalty and offenders with mental retardation. Human Rights Watch Publications, 13, 1–50. Humphreys, L. G. (1971). Theory of intelligence. In R. Cancro (Ed.), Intelligence: genetic and environmental in�luences. New York: Grune & Stratton. Hunsberger, B. (1995). Religion and prejudice: The role of religious fundamentalism, quest, and right-wing authoritarianism. Journal of Social Issues, 51, 113–
129.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 49/70
Hunsley, J., & Bailey, J. (1999). The clinical utility of the Rorschach: Unful�illed promises and an uncertain future. Psychological Assessment, 11, 266–277. Hunsley, J., & Mash, E. J. (2005). Introduction to the special section on developing guidelines for the evidence-based assessment (EBA) of adult disorders.
Psychological Assessment, 17, 251–255. Hunter, J. E. (1989). The Wonderlic Personnel Test as a predictor of training success and job performance. North�ield, IL: E. F. Wonderlic Personnel Test. Hunter, J. E. (1994). General Aptitude Test Battery. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Hunter, J. E., & Schmidt, F. L. (1976). Critical analysis of the statistical and ethical implications of various de�initions of test bias. Psychological Bulletin, 83, 1053–
1071. Hurtz, G., & Donovan, J. (2000). Personality and job performance: The Big Five revisited. Journal of Applied Psychology, 83, 869–879. Hutt, M. L. (1977). The Hutt Adaptation of the Bender-Gestalt Test. New York: Grune & Stratton. Hutt, M. L., & Briskin, G. J. (1960). The clinical use of the revised Bender-Gestalt Test. New York: Grune & Stratton. Institute of Medicine. (2001). Crossing the quality chasm: A new health system for the 21st century. Washington, DC: National Academy Press. International Psychogeriatric Association. (2002). Behavioral and Psychological Symptoms of Dementia (BPSD) educational pack. Skokie, IL: Author. Inwald, R. (2008). The Inwald Personality Inventory (IPI) and Hilson Research Inventories: Development and rationale. Aggression and Violent Behavior, 13, 298–
327. Inwald, R. E. (1988). Five-year follow-up of departmental terminations as predicted by 16 preemployment psychological indicators. Journal of Applied
Psychology, 73, 703–710. Irwin, P. M. (1992). Elementary and Secondary Education Act of 1965: FY 1993 Guide to Programs. Congressional Research Service. Washington, DC: Government
Printing Of�ice. Itard, J. M. G. (1932/1801). The wild boy of Aveyron. Trans. by G. & M. Humphrey. New York: Appleton-Century-Crofts. Ivcevic, Z., & Mayer, J. D. (2009). Mapping dimensions of creativity in the life-space. Creativity Research Journal, 21, 152–165. Iversen, G., Williamson, D., Ropacki, M., & Reilly, K. (2007). Frequency of abnormal scores on the Neuro-psychological Assessment Battery Screening Module (S-
NAB) in a mixed neurological sample. Applied Neuropsychology, 14, 178–182. Jaberg, P. E., Dixon, D. J., & Weis, G. M. (2009). Replication evidence in support of the psychometric properties of the Devereux Early Childhood Assessment.
Canadian Journal of School Psychology, 24, 158–166. Jackson, A., Brooks-Gunn, J., Huang, C., & Glassman, M. (2000). Single mothers in low-wage jobs: Financial strain, parenting, and preschoolers’ outcomes. Child
Development, 71, 1409–1423. Jackson, D. N. (1970). A sequential system for personality scale development. In C. D. Spielberger (Ed.), Current topics in clinical and community psychology (vol.
2). Orlando, FL: Academic Press. Jackson, D. N. (1984a). Manual for the Multidimensional Aptitude Battery. Port Huron, MI: Research Psychologists Press. Jackson, D. N. (1984b). Personality Research Form manual. Port Huron, MI: Research Psychologists Press. Jackson, D. N. (1991). Jackson Vocational Interest Survey manual (3rd ed.). Port Huron, MI: Research Psychologists Press. Jackson, D. N. (1998). Manual for the Multidimensional Aptitude Battery, Second Edition. Port Huron, MI: Research Psychologists Press. Jackson, D. N. (2000). Career Directions Inventory manual. Port Huron, MI: Sigma Assessment Systems. Jackson, D. N., & Messick, S. (1968). Creativity. In P. London & D. Rosenhan (Eds.). Foundations of abnormal psychology. New York: Holt. Jackson, J., Mulick, J., & Rojahn, J. (Eds.). (2007). Handbook of intellectual and developmental disabilities. New York: Springer. James, W. (1902). The varieties of religious experience. New York: Longman. Jankowski, D. (2002). A beginner’s guide to the MCMI-III. Washington, DC: American Psychological Association. Jennett, B., & Teasdale, G. (1981). Management of head injuries. Philadelphia: F. A. Davis. Jennett, B., Teasdale, G. M., & Knill-Jones, R. P. (1975). Predicting outcome after head injury. Journal of Royal College of Physicians of London, 9, 231–237. Jensen, A. (1998). The g factor: The science of mental ability. Westport, CT: Praeger. Jensen, A. R. (1969). How much can we boost IQ and scholastic achievement? Harvard Educational Review, 39, 1–123. Jensen, A. R. (1977). Cumulative de�icit in IQ of blacks in the rural south. Developmental Psychology, 13, 184–191. Jensen, A. R. (1979). g: outmoded theory or unconquered frontier? Creative Science and Technology, 2, 16–29. Jensen, A. R. (1980). Bias in mental testing. New York: Free Press. Jensen, A. R. (1981). Raising the IQ: The Ramey and Haskins Study. Intelligence, 5, 29–40. Jensen, A. R. (1984). Test bias: Concepts and criticisms. In C. R. Reynolds & R. T. Brown (Eds.), Perspectives on bias in mental testing. New York: Plenum Press. Jensen, A. R. (1998). The g factor: The science of mental ability. Westport, CT: Praeger. Jensen, A. R. (2006). Clocking the mind: Mental chronometer individual differences. Amsterdam: Elsevier. Jensen, A. R. (2011). The theory of intelligence and its measurement. Intelligence, 39, 171–177. Jensen, A. R., & Osborne, R. T. (1979). Forward and backward digit span interaction with race and IQ: A longitudinal developmental comparison. Berkeley:
University of California. (ERIC Document Reproduction Service No. ED 173 384). John, O. P., Donahue, E. M., & Kentle, R. L. (1991). The Big Five Inventory: Versions 4a and 54. Berkeley, CA: University of California, Berkeley, Institute of
Personality and Social Research. John, O. P., Naumann, L. P., & Soto, C. J. (2008). Paradigm shift to the integrative Big-Five trait taxonomy: History, measurement, and conceptual issues. In O. P.
John, R. W. Robins, & L. A. Pervin (Eds.), Handbook of personality: Theory and research (pp. 114–158). New York: Guilford Press. Johnson, J. H., & Williams, T. A. (1975). The use of online computer technology in a mental health admitting system. American Psychologist, 30, 388–390. Johnson, R. C., McClearn, G. E., Yuen, S., Nagoshi, C. T., Ahern, F. M., & Cole, R. E. (1985). Galton’s data a century later. American Psychologist, 40, 875–892. Johnson, S. T. (1994). Scholastic Assessment Tests. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Johnston, D. W. (1986). Behavior therapy. In R. Harre & R. Lamb (Eds.), The dictionary of physiological and clinical psychology. Cambridge, MA: MIT Press. Johnston, W. T., & Bolen, R. M. (1984). A comparison of the factor structures of the WISC-R for Blacks and Whites. Psychology in the Schools, 21, 42–44. Joint Committee on Testing Practices. (1988). Code of fair testing practices in education. Washington, DC: Author. Jones, K. L., Smith, D. W., Ulleland, C. N., & Streissguth, A. P. (1973). Patterns of malformation in offspring of chronic alcoholic mothers. Lancet, 1, 1267–1271. Jones, K., & Barber, J. (2012). Help for unemployed Americans. APA Monitor, 43(1), 18–19. Julian, E. (2005). Validity of the Medical College Admission Test for predicting medical school performance. Academic Medicine, 80, 910–917. Jung, C. G. (1910). The association method. American Journal of Psychology, 21, 219–269.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 50/70
Kaiser, H. F., & Michael, W. B. (1975). Domain validity and generalizability. Educational and Psychological Measurement, 35, 31–35. Kalat, J. (2012). Biological psychology (11th ed.). Belmont, CA: Wadsworth. Kamin, L. J., & Goldberger, A. S. (2001). Twin studies in behavioral research: A skeptical view. Unpublished manuscript. Kamphaus, R. W. (1993). Clinical assessment of children’s intelligence. Boston: Allyn and Bacon. Kanaya, T., Scullin, M., & Ceci, S. (2003). The Flynn effect and U.S. Policies: The impact of rising IQ scores on American society via mental retardation diagnoses.
American Psychologist, 58, 778–790. Kandel, E. R. (1991). Perception of motion, depth, and form. In E. R. Kandel, J. H. Schwartz, & T. M. Jessell (Eds.), Principles of neural science (3rd ed.). New York:
Elsevier. Kandel, E. R., Schwartz, J. H., & Jessell, T. M. (1995). Essentials of neural science and behavior. Norwalk, CT: Appleton & Lange. Kandel, E. R., Schwartz, J. H., Jessel, T. M., Siegelbaum, S. A., & Hudspeth, A. J. (2013). Principles of neural science (5th ed. rev.). New York: McGraw-Hill Medical. Kane, R. L. (1991). Standardized and �lexible batteries in neuropsychology: An assessment update. Neuropsychology Review, 2, 281–339. Kapuscinski, A. N., & Masters, K. S. (2010). The current status of measures of spirituality: A critical review of scale development. Psychology of Religion and
Spirituality, 2, 191–205. Kaufman, A. S. (1983). Test review: WAIS-R. Journal of Psychoeducational Assessment, 1, 309–319. Kaufman, A. S. (1990). Assessing adolescent and adult intelligence. Boston: Allyn and Bacon. Kaufman, A. S., & Kaufman, N. L. (1983). K-ABC administration and scoring manual. Circle Pines, MN: American Guidance Service. Kaufman, A. S., & Kaufman, N. L. (2004a). Kaufman Brief Intelligence Test (2nd ed.). Circle Pines, MN: American Guidance Service. Kaufman, A. S., & Kaufman, N. L. (2004b). Kaufman Test of Educational Achievement (2nd ed.). Circle Pines, MN: American Guidance System Publishing. Kaufman, A. S., & Lichtenberger, E. O. (2002). Assessing adolescent and adult intelligence (2nd ed.). Boston: Allyn & Bacon. Kaufman, A. S., McLean, J. E., & Reynolds, C. R. (1988). Sex, race, residence, region, and education differences on the 11 WAIS-R subtests. Journal of Clinical
Psychology, 44, 231–248. Kaufman, J. C., & Baer, J. (2004). Hawking’s Haiku, Madonna’s math: Why it is hard to be creative in every room of the house. In R. J. Sternberg, E. L. Grigorenko, &
J. L. Singer (Eds.), Creativity: From potential to realization (pp. 3–19). Washington, DC: American Psychological Association. Kaufman, J. C., Cole, J. C., & Baer, J. (2009). The construct of creativity: A structural model for self-reported creativity ratings. Journal of Creative Behavior, 43, 119–
134. Kaufman, J. D. (2012). Development of the Kaufman Domains of Creativity Scale (K-DOCS). Psychology of Aesthetics, Creativity, and the Arts, 6, 298–308. Kausler, D. (1991). Experimental psychology, cognition, and human aging (2nd ed.). New York: Springer-Verlag. Kazdin, A. E. (1990). Evaluation of the Automatic Thoughts Questionnaire: Negative cognitive processes and depression among children. Psychological
Assessment: A Journal of Consulting and Clinical Psychology, 2, 73–79. Keith, T. Z. (1999). Effects of general and speci�ic abilities on student achievement: Similarities and differences across ethnic groups. School Psychology Quarterly,
14, 239–262. Kelley, T. L. (1928). Crossroads in the mind of man: A study of differentiable mental abilities. Stanford, CA: Stanford University Press. Kelly, E. L., & Fiske, D. W. (1951). The prediction of performance in clinical psychology. Ann Arbor: University of Michigan Press. Kendall, P. C., & Hollon, S. D. (1989). Anxious self-talk: Development of the Anxious Self-Statements Questionnaire (ASSQ). Cognitive Therapy and Research, 13,
81–93. Kennedy, C., & Moore, J. (Eds.). (2010). Military neuro-psychology. New York: Springer Publishing. Kennedy, W. A., Van de Riet, V., & White, J. C., Jr. (1963). A normative sample of intelligence and achievement of negro elementary school children in the southeast
United States. Monographs of the Society for Research in Child Development, 28 [No. 90], 68. Kent, G. H., & Rosanoff, A. J. (1910). A study of association in insanity. American Journal of Insanity, 67, 37–96; 317–390. Kerr, B., & Gagliardi, C. (2003). Measuring creativity in research and practice. In S. Lopez & C. R. Snyder (Eds.), Positive psychological assessment: A handbook of
models and measures. Washington, DC: American Psychological Association. Kertesz, A. (1982). Aphasia and associated disorders: Taxonomy, localization, and recovery. New York: Grune & Stratton. Kertesz, A. (2006). Western Aphasia Battery-Revised. San Antonio, TX: Harcourt. Keyser, D. J., & Sweetland, R. C. (Eds.). (1984–1994). Test Critiques (volumes I–X). Kansas City, MO: Test Corporation of America. Khaleefa, O., & Lynn, R. (2008). Normative data for Raven’s Coloured Progressive Matrices Scale in Yemen. Psychological Reports, 103, 170–172. Kiecolt-Glaser, J. K. (2009). Psychoneuroimmunology: Psychology’s gateway to the biomedical future. Perspectives on Psychological Science, 4, 367–369. Kiernan, R., Mueller, J., & Langston, J. W. (2009). Cognistat manual. Fairfax, CA: Cognistat, Inc. Kifer, E. (1985). Review of ACT Assessment Program. Ninth mental measurements yearbook. Lincoln: University of Nebraska Press. Killian, G. A. (1987). House-Tree-Person technique. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques compendium. Kansas City, MO: Test Corporation of
America. Kim, K. H. (2006). Can we trust creativity tests? A review of the Torrance Tests of Creative Thinking (TTCT). Creativity Research Journal, 18, 3–14. Kim, W. J., Kim, L. I., & Rue, D. S. (1997). Korean American children. In G. Johnson-Powell, J. Yamamoto, G. E. Wyatt, & W. Arroyo (Eds.), Transcultural child
development: Psychological assessment and treatment (pp. 183–207). Hoboken, NJ: John Wiley & Sons. Kim, Y., Pilkonis, P. A., Frank, E., Thase, M. E., & Reynolds, C. F. (2002). Differential functioning of the Beck Depression Inventory in late-life patients: Use of item
response theory. Psychology and Aging, 17, 379–391. King, K. (2001). A critique of behavioral observational coding systems of couples’ interaction: CISS and RCISS. Journal of Social and Clinical Psychology, 20, 1–23. Kinnear, P. R., & Gray, C. D. (1997). SPSS for Windows made simple (2nd ed.). Trowbridge, UK: Psychology Press. Kinsbourne, M. (1994). Neuropsychology of attention. In D. W. Zaidel (Ed.), Neuropsychology. San Diego, CA: Academic Press. Kirk, J. W., Harris, B., Hutaff-Lee, C. F., & others. (2010). Performance on the Test of Memory Malingering (TOMM) among a large clinic-referred pediatric sample.
Child Neuropsychology, 17, 242–254. Kirkpatrick, L., & Hood, R. (1990). Intrinsic-Extrinsic Religious Orientation: The boon or bane of contemporary psychology of religion? Journal for the Scienti�ic
Study of Religion, 29, 442–462. Kleiman, L., & Faley, R. (1985). The implications of professional and legal guidelines for court decisions involving criterion-related validity: A review and analysis.
Personnel Psychology, 38, 303–833. Klieger, D. M., & Franklin, M. E. (1993). Validity of the fear survey schedule in phobia research: A laboratory test. Journal of Psychopathology and Behavioral
Assessment, 15, 207–218.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 51/70
Klimoski, R., & Palmer, S. (1994). The ADA and the hiring process in organizations. In S. M. Bruyere & J. O’Keeffe (Eds.), Implications of the Americans with Disabilities Act for psychology. New York: Springer.
Kline, P. (1986). A handbook of test construction. New York: Methuen. Kline, P. (1993). The handbook of psychological testing. London: Routledge. Kline, P. (1999). The handbook of psychological testing (2nd ed.). London: Routledge. Klingler, D. E., Miller, D., Johnson, J., & Williams, T. (1977). Process evaluation of an online computer-assisted unit for intake assessment of mental health patients.
Behavior Research Methods and Instrumentation, 9, 110–116. Klove, H. (1963). Clinical neuropsychology. In F. M. Forster (Ed.), The medical clinics of North America. New York: Saunders. Kohlberg, L. (1958). The development of modes of moral thinking and choice in the years ten to sixteen. Unpublished doctoral dissertation, University of Chicago. Kohlberg, L. (1981). Essays on moral development: Vol. 1. The philosophy of moral development. San Francisco: Harper & Row. Kohlberg, L. (1984). Essays on moral development: Vol. 2. The psychology of moral development. San Francisco: Harper & Row. Kohlberg, L., & Elfenbein, D. (1975). The development of moral judgments concerning capital punishment. American Journal of Orthopsychiatry, 45, 614–639. Kohlberg, L., & Kramer, R. (1969). Continuities and discontinuities in children and adult moral development. Human Development, 12, 225–252. Kolb, B., & Milner, B. (1981). Performance of complex arm and facial movements after focal brain lesions. Neuropsychologia, 19, 491–503. Kolb, B., & Whishaw, I. Q. (2002). Fundamentals of human neuropsychology (5th ed.). New York: Worth/Freeman. Kolb, B., & Whishaw, I. Q. (2011). An introduction to brain and behavior (3rd ed.). New York: Worth Publishers. Kolb, B., Milner, B., & Taylor, L. (1983). Perception of faces by patients with localized cortical excisions. Canadian Journal of Psychology, 37, 8–18. Koppitz, E. (1963). The Bender Gestalt Test for young children. New York: Grune and Stratton. Koppitz, E. (1975). The Bender Gestalt Test for young children, Volume II: Research and application, 1963–1975. New York: Grune and Stratton. Koss, E., Patterson, M., Mack, J., Smyth, K., & Whitehouse, P. (1998). Reliability and validity of the Tinkertoy Test in evaluating individuals with Alzheimer’s
disease. Clinical Neuropsychologist, 12, 325–329. Kostrubala, C., & Braden, J. (1998). The American Sign Language translation of the WAIS-III. San Antonio, TX: The Psychological Corporation. Kraus, J. F., & MacArthur, D. L. (1996). Epidemiologic aspects of brain injury. Neurologic Clinics, 14(2): 435–450. Krikorian, R., & Bartok, J. (1998). Developmental data for the Porteus Maze Test. Clinical Neuropsychologist, 12, 305–310. Krokoff, L. J., Gottman, J., & Hass, S. (1989). Validation of a global rapid couples interaction scoring system. Behavioral Assessment, 11, 65–79. Krugman, M. (1970). H-T-P: House, Tree, and Person. In O. K. Buros (Ed.), Personality tests and reviews. Highland Park, NJ: Gryphon Press. Krumboltz, J. (1999). Career Beliefs Inventory: Applications and technical guide. Palo Alto, CA: Consulting Psychologists Press. Krumboltz, J. D. (1993). Integrating career and personal counseling. Career Development Quarterly, 42, 143–148. Krumboltz, J. D. (1996). A learning theory of career counseling. In M. L. Savickas & W. B. Walsh (Eds.), Handbook of career counseling theory and practice (pp. 55–
80). Palo Alto, CA: Davies-Black. Krumboltz, J. D. (2009). The happenstance learning theory. Journal of Career Assessment, 17, 135–154. Krumboltz, J. D., & Vosvick, M. A. (1996). Career assessment and the Career Beliefs Inventory. Journal of Career Assessment, 4, 345–361. Kuder, G. F. (1934). Kuder preference record. Chicago: Science Research Associates. Kuder, G. F., & Richardson, M. W. (1937). The theory of estimation of test reliability. Psychometrika, 2, 151–160. Kuncel, N. R., & Sackett, P. R. (2007). Selective citation mars conclusions about test validity and predictive bias. American Psychologist, 62, 145–146. Kuncel, N., Campbell, J., & Ones, D. (1998). Validity of the Graduate Record Examination: Estimated or tacitly known? American Psychologist, 53, 567–568. Kuncel, N., Hezlett, S., & Ones, D. (2001). A comprehensive meta-analysis of the predictive validity of the Graduate Record Examinations: Implications for
graduate student selection and performance. Psychological Bulletin, 127, 162–181. Kupfermann, I. (1991). Hypothalamus and limbic system: Peptidergic neurons, homeostasis, and emotional behavior. In E. R. Kandel, J. H. Schwartz, & T. M. Jessell
(Eds.), Principles of neural science (3rd ed.). New York: Elsevier. Kurtines, W., & Greif, E. B. (1974). The development of moral thought: Review and evaluation of Kohlberg’s approach. Psychological Bulletin, 81, 453–470. Kvaal, K., Ulstein, I., Nordhus, I. H., & Engedal, K. (2005). The Spielberger State-Trait Anxiety Inventory (STAI): The state scale in detecting mental disorders in
geriatric patients. International Journal of Geriatric Psychiatry, 20, 629–634. Kwate, N. (2001). Intelligence or misorientation? Eurocentrism in the WISC-III. Journal of Black Psychology, 27, 221–238. La Rue, A. (1992). Aging and neuropsychological assessment. New York: Plenum. Laatsch, L., & Choca, J. (1994). Cluster-branching methodology for adaptive testing and the development of the Adaptive Category Test. Psychological Assessment,
6, 345–351. LaBarbera, D. (2005). Physician assistant Self-Directed Search Holland Codes. Journal of Career Assessment, 13, 337–346. Lachar, D. (1974). The MMPI: Clinical assessment and automated interpretation. Los Angeles: Western Psychological Services. Lachar, D. (1987). Automated assessment of child and adolescent personality. In J. N. Butcher (Ed.), Computerized psychological assessment: A practitioner’s guide.
New York: Basic Books. Lachar, D., & Gdowski, C. L. (1979). Actuarial assessment of child and adolescent personality: An interpretive guide for the Personality Inventory for Children pro�ile.
Los Angeles: Western Psychological Services. Lachar, D., & Gruber, C. (2001). Manual: Personality Inventory for Children-2. Los Angeles: Western Psychological Services. Lacks, P. (1999). Bender-Gestalt screening for brain dysfunction (2nd ed.). New York: Wiley. Lah, M. I. (1989). New validity, normative, and scoring data for the Rotter Incomplete Sentences Blank. Journal of Personality Assessment, 53, 607–620. Lah, M. I., & Rotter, J. B. (1981). Changing college student norms on the Rotter Incomplete Sentences Blank. Journal of Consulting and Clinical Psychology, 49, 985. Lamp, R., & Krohn, E. (2001). A longitudinal predictive validity investigation of the SB:FE and K-ABC with at-risk children. Journal of Psychoeducational
Assessment, 19, 334–349. Landy, F. (1996). The psychology of work behavior (5th ed.) Monterey, CA: Brooks/Cole. Landy, F. J., & Farr, J. L. (1983). The measurement of work performance: Methods, theory and applications. New York: Academic Press. Lane, S. (1992). Review of the Iowa Tests of Basic Skills. Eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. LaPiana, W. P. (1998). A history of the Law School Admission Council and the LSAT. Keynote Address to the 1998 LSAC Annual Meeting. Larrabee, G. (2008). Flexible vs. �ixed batteries in forensic neuropsychological assessment: Reply to Bigler and Hom. Archives of Clinical Neuropsychology, 23(7–
8), 763–776.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 52/70
Larson, G. E. (1994). Armed Services Vocational Aptitude Battery. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Larson, G. E., & Wolfe, J. (1995). Validity results for g from an expanded test base. Intelligence, 20, 15–25. Lassiter, K., & Bardos, A. (1995). The relationship between young children’s academic achievement and measures of intelligence. Psychology in the Schools, 32,
170–177. Latham, G. P., & Skarlicki, D. (1995). Criterion-related validity of the situational and patterned behavior description interviews with organizational citizenship
behavior. Human Performance, 8, 67–80. Lau, B. C., Collins, M. W., & Lovell, M. R. (2011). Sensitivity and speci�icity of subacute computerized neurocognitive testing and symptom evaluation in predicting
outcomes after sports-related concussion. American Journal of Sports Medicine, 39(6), 1209–1216. Laux, J., Salyers, K., & Kotova, E. (2005). A psychometric evaluation of the SASSI-3 in a college sample. Journal of College Counseling, 8, 41–51. LaVoie, A. L. (1987). The Blacky Pictures. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques compendium. Kansas City, MO: Test Corporation of America. LeBuffe, P. A., & Naglieri, J. A. (1999a). Devereux Early Childhood Assessment (DECA): A measure of within-child protective factors in preschool children. NHSA
Dialog, 3, 75–80. LeBuffe, P. A., & Naglieri, J. A. (1999b). Devereux Early Childhood Assessment Program: Technical manual. Lewisville, NC: Kaplan Press. LeBuffe, P. A., & Naglieri, J. A. (2003). The Devereux Early Childhood Assessment Clinical Form (DECA-C): A measure of behaviors related to risk and resilience in
preschool children. Lewisville, NC: Kaplan Press. Ledbetter, M., Smith, L., Vosler-Hunter, W., & Fischer, J. (1991). An evaluation of the research and clinical usefulness of the Spiritual Well-Being Scale. Journal of
Psychology and Theology, 19, 49–55. Lee, M. S., Wallbrown, F., & Blaha, J. (1990). Note on the construct validity of the Multidimensional Aptitude Battery. Psychological Reports, 67, 1219–1222. Lefebvre, M. F. (1981). Cognitive distortion and cognitive errors in depressed psychiatric and low back pain patients. Journal of Consulting and Clinical
Psychology, 49, 517–525. Lehman, R. E. (1978). Symptom contamination of the Schedule of Recent Events. Journal of Consulting and Clinical Psychology, 46, 1564–1565. Leiter, R. G. (1948). Leiter International Performance Scale. Chicago: Stoelting Co. Leiter, R. G. (1979). Leiter International Performance Scale: Instruction manual. Chicago: Stoelting Co. Leli, D. A., & Filskov, S. B. (1984). Clinical detection of intellectual deterioration associated with brain damage. Journal of Clinical Psychology, 40, 1435–1441. Lent, R. W., Brown, S. D., & Hackett, G. (2000). Contextual supports and barriers to career choice: A social cognitive analysis. Journal of Counseling Psychology, 47,
36–49. Lester, B. M. (1984). Data analysis and prediction. In T. B. Brazelton (Ed.), Neonatal Behavioral Scale (2nd Ed.). London: Spastics International Medical
Publications. Levashina, J., Morgeson, F. P., & Campion, M. A. (2012). Tell me some more: Exploring how verbal ability and item veri�iability in�luence responses to biodata
questions in a high-stakes selection context. Personnel Psychology, 65, 359–383. Levin, H., Song, J., Ewing-Cobbs, L., & Roberson, G. (2001). Porteus Maze performance following traumatic brain injury in children. Neuropsychology, 15, 557–567. Levinson, E. M. (1990). Vocational assessment involvement and use of the Self-Directed Search by school psychologists. Psychology in the Schools, 27, 217–228. Lewinsohn, P. M. (1965). Psychological correlates of overall quality of �igure drawings. Journal of Consulting Psychology, 29, 504–512. Lewinsohn, P. M., Munoz, R. F., Youngren, M. A., & Zeiss, A. M. (1986). Control your depression: Reducing depression through learning self-control techniques,
relaxation training, pleasant activities, social skills, constructed thinking, planning ahead, and more (rev. ed.). New York: Prentice Hall. Lewinsohn, P., & Talkington, J. (1979). Studies on the measurement of unpleasant events and relations with depression. Applied Psychological Measurement, 3,
83–101. Lewis, M., & Brooks-Gunn, J. (1981). Visual attention at three months as a predictor of cognitive functioning at two years of age. Intelligence, 5, 131–140. Lewis, M., & Sullivan, M. W. (1985). Infant intelligence and its assessment. In B. B. Wolman (Ed.), Handbook of intelligence: Theories, measurements, and
applications. New York: Wiley. Lezak, M. (1982). The problem of assessing executive functions. International Journal of Psychology, 17, 281–297. Lezak, M. (1983). Neuropsychological assessment (2nd ed.). New York: Oxford University Press. Lezak, M. (1995). Neuropsychological assessment (3rd ed.). New York: Oxford University Press. Lezak, M. D., & O’Brien, K. P. (1990). Chronic emotional, social, and physical changes after traumatic brain injury. In E. D. Bigler (Ed.), Traumatic brain injury:
Mechanisms of damage, assessment, intervention, and outcome. Austin, TX: PRO-ED. Lezak, M. D., Howieson, D. B., Bigler, E. D., & Tranel, D. (2012). Neuropsychological assessment (5th ed.). New York: Oxford University Press. Lezak, M., Howieson, D., & Loring, D. (2004). Neuropsychological assessment (4th ed.). New York: Oxford University Press. Lichtenberg, P., Manning, Vangel, S., & Ross. T. (1995). Normative and ecological validity data in older urban medical patients: A program of neuropsychological
research. Advances in Medical Psychotherapy, 8, 121–136. Lichtenberger, E., & Kaufman, A. (2009). Essentials of WAIS-IV assessment. New York: Wiley. Lien, M. T., & Carlson, J. S. (2009). Psychometric properties of the Devereux Early Childhood Assessment in a Head Start sample. Journal of Psychoeducational
Assessment, 27, 386–396. Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 140. Lilienfeld, S. O., Ammirati, R., & Land�ield, K. (2009). Giving debiasing away: Can psychological research on correcting cognitive errors promote human welfare?
Perspectives on Psychological Science, 4, 390–398. Lilienfeld, S., Wood, J., & Garb, H. (2000). The scienti�ic status of projective techniques. Psychological Science in the Public Interest, 2, 27–66. Lilienfeld, S., Wood, J., & Garb, H. (2001, May). What’s wrong with this picture? Scienti�ic American, 81–87. Lindal, E., & Stefansson, J. (1993). Mini-Mental State Examination scores: Gender and lifetime psychiatric disorders. Psychological Reports, 72, 631–641. Lindenberger, U., & Baltes, P. (1994). Aging and intelligence. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Lindvall, C. M. (1967). Measuring pupil achievement and aptitude. New York: Harcourt, Brace & World. Lindzey, G. (1959). On the classi�ication of projective techniques. Psychological Bulletin, 56, 158–168. Linn, R. L. (1989). Review of the Iowa Tests of Basic Skills. Tenth Mental Measurements Yearbook. Lincoln: University of Nebraska Press. Lipsitt, P. D. (1970). Competency Screening Test. Boston: Competency to Stand Trial and Mental Illness Project. Lipsitz, J. D., Dworkin, R., & Erlenmeyer-Kimling, L. (1993). Wechsler Comprehension and Picture Arrangement subtests and social adjustment. Psychological
Assessment, 5, 430–437. Lishman, W. A. (1997). Organic psychiatry: The psychological consequences of cerebral disorder (3rd ed.). Oxford: Blackwell Scienti�ic Publications.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 53/70
Liskow, B., Campbell, J., Nickel, E., & Powell, B. (1995). Validity of the CAGE questionnaire in screening for alcohol dependence in a walk-in (triage) clinic. Journal of Studies on Alcohol, 56, 277–281.
Loe, S. A., Kadlubek, R. M., & Williams, W. J. (2007). Administration and scoring errors on the WISC-IV among graduate student examiners. Journal of Psychoeducational Assessment, 25, 237–247.
Lofquist, L. H., & Dawis, R. V. (1991). Essentials of person-environment correspondence counseling. Minneapolis: University of Minnesota Press. Lohman, D., & Hagen, E. (2001). Cognitive Abilities Test, Form 6; Examiner’s manual. Boston: Houghton Mif�lin. Lopez, S. J., & Snyder, C. R. (Eds.). (2003). Positive psychological assessment: A handbook of models and measures. Washington, DC: American Psychological
Association. Lopez, S., & Snyder, C. R. (Eds.). (2003). Positive psychological assessment: A handbook of models and measures. Washington, DC: American Psychological
Association. Lord, F. M., & Novick, M. R. (1968). Statistical theories of mental test scores. Menlo Park, CA: Addison-Wesley. Lord, F., & Novick, M. (1968). Statistical theories of mental tests. New York: Addison-Wesley. Lovell, M. R. (2006). The ImPACT neuropsychological test battery. In R. J. Echemendia (Ed.), Sports neuro-psychology: Assessment and management of traumatic
brain injury (pp. 193–215). New York: Guilford Press. Lovell, M. R., Iverson, G. L., Podell, M. W., & others. (2006). Measurement of symptoms following sports-related concussion: Reliability and normative data for the
post-concussion scale. Applied Neuropsychology, 13(3), 166–174. Lowe, P. A., Lee, S. W., Witteborg, K. M., & others. (2008). The Test Anxiety Inventory for Children and Adolescents (TAICA): Examination of the psychometric
properties of a new multidimensional measure of test anxiety among elementary and secondary school students. Journal of Psychoeducational Assessment, 26, 215–230.
Lubinski, D., Benbow, C., & Ryan, J. (1995). Stability of vocational interests among the intellectually gifted from adolescence to adulthood: A 15-year longitudinal study. Journal of Applied Psychology, 80, 196–200.
Lüdtke, O., Roberts, B. W., Trautwein, U., & Nagy, G. (2011). A random walk down university avenue: Life paths, life events, and personality trait change at the transition to university life. Journal of Personality and Social Psychology, 101, 620–637.
Lukasik, C. (2004). The physiognomy of biometrics. Retrieved from www.common-place.org (http://www.common-place.org) , 5, 1–4. Lukin, M. E., Dowd, E. T., Plake, B., & Kraft, R. (1985). Comparing computerized vs. traditional psychological assessment. Computers in Human Behavior, 1, 49–58. Lunz, M., & Bergstrom, B. (1994). Computer adaptive testing: A national pilot study. In M. Wilson (Ed.), Objective measurement: Theory into practice (vol. 2).
Norwood, NJ: Ablex. Lunz, M., Bergstrom, B., & Wright, B. (1994). Reliability of alternate computer-adaptive tests. In M. Wilson (Ed.), Objective measurement: Theory into practice (vol.
2). Norwood, NJ: Ablex. Luria, A. R. (1966). Higher cortical functions in man. New York: Basic Books. Luria, A. R. (1970). The functional organization of the brain. Scienti�ic American, 222, 66–78. Luria, A. R. (1973). The working brain. New York: Basic Books. Lynn, R. (1987). Japan: Land of the rising IQ. A reply to Flynn. Bulletin of the British Psychological Society, 40, 464–468. Lynn, R. (2009). What has caused the Flynn effect? Secular increases in the Development Quotients of infants. Intelligence, 37, 16–24. Lyon, G. R. (1996b). Special education for students with disabilities. The Future of Children, 6, 1–19. Lyon, G. R., (1996a). Learning disabilities. Special Education for Students With Disabilities, 6, 1–18. MacAndrew, C. (1965). The differentiation of male alcoholic out-patients from nonalcoholic psychiatric patients by means of the MMPI. Quarterly Journal of
Studies on Alcohol, 26, 238–246. Machover, K. (1949). Personality projection in the drawing of the human �igure. Spring�ield, IL: Charles C. Thomas. Machover, K. (1951). Drawing of the human �igure: A method of personality investigation. In H. Anderson & G. Anderson (Eds.), An introduction to projective
techniques. New York: Prentice Hall. Mack, J., & Patterson, M. (1995). Executive dysfunction and Alzheimer’s disease: Performance on a test of planning ability, the Porteus Maze Test.
Neuropsychology, 9, 556–564. Mackenzie Ross, S. J., Brewin, C., Curran, H. V., & others. (2010). Neuropsychological and psychiatric functioning in sheep farmers exposed to low levels of organo-
phosphate pesticides. Neurotoxicology and Teratology, 32, 452–459. MacPhillamy, D. J., & Lewinsohn, P. M. (1982). The Pleasant Events Schedule: Studies on reliability, validity, and scale intercorrelation. Journal of Consulting and
Clinical Psychology, 50, 363–380. Maddi, S. R. (2000). Personality theories: A comparative analysis (6th ed.). Prospective Heights, IL: Waveland Press. Mahoney, M., & Arnkoff, D. (1978). Cognitive and self-control therapies. In S. Gar�ield & A. Bergin (Eds.), Handbook of psychotherapy and behavior change: An
empirical analysis. New York: Wiley. Main, M., & Hesse, E. (1990). Parents’ unresolved traumatic experiences are related to infant disorganized attachment status: Is frightened and/or frightening
parental behavior the linking mechanism? In M. Greenberg, D. Cicchetti, & E. Cummings (Eds.), Attachment in the preschool years (pp. 161–182). Chicago: University of Chicago Press.
Main, M., & Solomon, J. (1986). Discovery of a new, insecure-disorganized/disoriented attachment pattern. In T. B. Brazelton & M. W. Yogman (Eds.), Affective development in infancy (pp. 95–124). Norwood, NJ: Ablex Publishing.
Majnemer, A., & Mazer, B. (1998). Neurologic evaluation of the newborn infant: De�inition and psychometric properties. Developmental Medicine and Child Neurology, 40, 708–715.
Malgady, R. G., Constantino, G., & Rogler, L. H. (1984). Development of a Thematic Apperception Test (TEMAS) for urban Hispanic children. Journal of Consulting and Clinical Psychology, 52, 986–996.
Maloney, M. P., & Ward, M. P. (1979). Mental retardation and modern society. New York: Oxford University Press. Man, D., Chung, J., & Mak, M. (2009). Development and validation of the Online Rivermead Behavioral Memory Test (OL-RBMT) for people with stroke.
Neurorehabilitation, 24, 231–236. Manly, T., Nimmo-Smith, I., Watson, P., & others. (2001). The differential assessment of children’s attention: The Test of Everyday Attention for Children (TEA-Ch),
normative sample and ADHD performance. Journal of Child Psychology and Psychiatry, 42, 1065–1081. Manning, W. H., & Jackson, R. (1984). College entrance examinations: Objective selection or gatekeeping for the economically privileged. In C. R. Reynolds & R. T.
Brown (Eds.), Perspectives on bias in mental testing. New York: Plenum Press. Manto, M., & Pandolfo, M. (Eds.). (2002). The cerebellum and its disorders. New York: Cambridge University Press.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 54/70
Marcus, D. K., Fulton, J. J., & Clarke, E. J. (2010). Lead and conduct problems: A meta-analysis. Journal of Clinical Child and Adolescent Psychology, 39, 234–241. Mardell, C., & Goldenberg, D. (2011). Developmental indicators for the assessment of learning—Fourth edition (DIAL-4). San Antonio, TX: Pearson. Marks, P. A., & Seeman, W. (1963). The actuarial description of abnormal personality. Baltimore: Williams & Wilkins. Markwardt, F. C. (1997). Peabody Individual Achievement Test-Revised/Normative Update. Circle Pines, MN: American Guidance Service. Marnic, L. R. (2011). Evaluating the Bender Visual Motor Gestalt Test II as a diagnostic screening instrument among clinically referred children and adolescents.
Dissertation Abstracts International: Section B: The Sciences and Engineering, 72(5-B), 3118. Martell, D. A. (1992). Forensic neuropsychology and the criminal law. Law and Human Behavior, 16, 313–336. Martin, J. C. (1994). Birth defects. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Martin, R. (2003). Sense of Humor. In S. Lopez & C. R. Snyder (Eds.), Positive psychological assessment: A handbook of models and measures. Washington, DC:
American Psychological Association. Martin, R. A. (1996). The Situational Humor Response Questionnaire (SHRQ) and Coping Humor Scale (CHS): A decade of research �indings. Humor: International
Journal of Humor Research, 9, 251–272. Martin, R. A., & Lefcourt, H. M. (1983). Sense of humor as a moderator of the relation between stressors and moods. Journal of Social and Personality Psychology,
45, 1313–1324. Martin, R. A., & Lefcourt, H. M. (1984). Situational Humor Response Questionnaire: Quantitative measure of sense of humor. Journal of Social and Personality
Psychology, 47, 145–155. Martin, R. A., Puhlik-Doris, P., Larsen, G., Gray, J., & Weir, K. (2003). Individual differences in uses of humor and their relation to psychological well-being:
Development of the Humor Styles Questionnaire. Journal of Research in Personality, 37, 48–75. Martin, S. (2010). The internet’s ethical challenges. APA Monitor, 41(7), 32. Martindale, C. (1981). Cognition and consciousness. Homewood, IL: Dorsey. Martuza, V. R. (1977). Applying norm-referenced and criterion-referenced measurement in education. Boston: Allyn and Bacon. Masten, A. S., Best, K. M., & Garmezy, N. (1990). Resilience and development: Contributions from the study of children who overcame adversity. Development and
Psychopathology, 2, 425–444. Masters, K. S., & Hooker, S. A. (2012, November 12). Religiousness/spirituality, cardiovascular disease, and cancer: Cultural integration for health research and
intervention. Journal of Consulting and Clinical Psychology, online publication. Matarazzo, J. D. (1972). Wechsler’s measurement and appraisal of adult intelligence (5th ed.). Baltimore: Williams & Wilkins. Matarazzo, J. D. (1990). Psychological assessment versus psychological testing: Validation from Binet to the school, clinic, and courtroom. American Psychologist,
45, 999–1017. Matarazzo, J. D. (1992). Psychological testing and assessment in the 21st century. American Psychologist, 47, 1007–1018. Matson, J. (Ed.). (2007). Handbook of assessment in persons with intellectual disability. London: Academic Press. Matson, J. L., & Tureck, K. (2012). Early diagnosis of autism: Current status of the Baby and Infant Screen for Children with Autism Traits (BISCUIT-Parts 1, 2, and
3). Research in Autism Spectrum Disorders, 6, 1135–1141. Matson, J. L., Boisjoli, J. A., & Wilkins, J. (2007). Baby and Infant Screen for Children with Autism Traits (BISCUIT). Baton Rouge, La: Disability Consultants, LLC. Matson, J. L., Boisjoli, J. A., Hess, J. A., & Wilkins, J. (2010). Factor structure and diagnostic �idelity of the Baby and Infant Screen for Children with Autism Traits-
Part 1 (BISCUIT-Part 1). Developmental Neurorehabilitation, 13, 72–79. Matson, J. L., Wilkins, J., & Fodstad, J. C. (2011). The validity of the Baby and Infant Screen for Children with Autism Traits: Part 1 (BISCUIT: Part 1). Journal of
Autism and Developmental Disorders, 41, 1139–1146. Matthews, G., Zeidner, M., & Roberts, R. (2002). Emotional intelligence: Science and myth. Cambridge, MA: MIT Press. Mattis, S. (2001). Dementia Rating Scale-2. Lutz, FL: Psychological Assessment Resources. Maxwell, J. K., & Wise, F. (1984). PPVT IQ validity in adults: A measure of vocabulary, not of intelligence. Journal of Clinical Psychology, 40, 1048–1053. May, P. A., Gossage, J. P., Kalberg, W. O., & others. (2009). Prevalence and epidemiologic characteristics of FASD from various research methods with an emphasis
on recent in-school studies. Developmental Disabilities Research Reviews, 15, 176–192. Mayer, J. D. (2007–2008). The big questions of personality psychology: De�ining common pursuits of the discipline. Imagination, Cognition and Personality, 27, 3–
26. Mayer, J., & Salovey, P. (1993). The intelligence of emotional intelligence. Intelligence, 17, 433–442. Mayer, J., Salovey, P., & Caruso, D. (2002). Mayer-Salovey-Caruso Emotional Intelligence Test (MSCEIT) user’s manual. Toronto, ON: Multi-Health Systems. Mayer, J., Salovey, P., & Caruso, D. (2004). Emotional intelligence: Theory, �indings, and implications. Psychological Inquiry, 15, 197–215. Mayer, J., Salovey, P., & Caruso, D. (2008). Emotional intelligence: New ability or eclectic traits? American Psychologist, 63, 503–517. Mayer, J., Salovey, P., Caruso, D., & Sitarenios, G. (2003). Measuring emotional intelligence with the MSCEIT V2.0. Emotion, 3, 97–105. Mayers, L., & Redick, T. S. (2012). Clinical utility of ImPACT assessment for postconcussion return-to-play counseling: Psychometric issues. Journal of Clinical and
Experimental Neuropsychology, 34, 235–242. Mayeux, R., & Kandel, E. R. (1991). Disorders of language: The aphasias. In E. R. Kandel, J. H. Schwartz, & T. M. Jessel (Eds.), Principles of neural science (3rd ed.).
New York: Elsevier. McAllister, L. W. (1986). A practical guide to CPI interpretation. Palo Alto, CA: Consulting Psychologists Press. McCall, R. B. (1976). Toward an epigenetic conception of mental development in the �irst three years of life. In M. Lewis (Ed.), Origins of intelligence: Infancy and
early childhood. New York: Plenum Press. McCall, R. B. (1979). The development of intellectual functioning in infancy and the prediction of later IQ. In J. D. Osofsky (Ed.), Handbook of infant development.
New York: Wiley. McCall, W. A. (1939). Measurement. New York: Macmillan. McCallum, R. S. (1990). Determining the factor structure of the Stanford-Binet: Fourth Edition—the right choice. Journal of Psychoeducational Assessment, 8, 436–
442. McCoy, B. (2000). Quack! Tales of medical fraud from the museum of questionable medical devices. Santa Monica, CA: Santa Monica Press. McCrae, R. R. (1985). Review of the De�ining Issues Test. Ninth mental measurements yearbook. Lincoln: University of Nebraska Press. McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the �ive-factor model of personality across instruments and observers. Journal of Personality and Social
Psychology, 2, 81–90. McCrae, R., Costa, P., & Martin, T. (2005). The NEO-PI-3: A more readable revised NEO Personality Inventory. Journal of Personality Assessment, 84, 261–270.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 55/70
McCullough, M. E., Emmons, R. A., & Tsang, J. (2002). The Grateful disposition: A conceptual and empirical topography. Journal of Personality and Social Psychology, 82, 112–127.
McDermott, B. E., & Sokolov, G. (2009). Malingering in a correctional setting: The use of the structured interview of reported symptoms in a jail sample. Behavioral Sciences & the Law, 27, 753–765.
McDonald, A., Nussbaum, D., & Bagby, R. (1992). Reliability, validity, and utility of the Fitness Interview Test. Canadian Journal of Psychiatry, 36, 480–484. McDonald, R. P. (1999). Test theory: A uni�ied approach. Mahwah, NJ: Erlbaum. McGee, R., Clark, S., & Symons, D. (2000). Does the Conners’ continuous performance test aid in ADHD diagnosis? Journal of Abnormal Child Psychology, 28, 415–
424. McGlynn, F. D., & Rose, M. P. (1998). Assessment of anxiety and fear. In A. S. Bellack & M. Hersen (Eds.), Behavioral assessment: A practical handbook (4th ed.).
Boston: Allyn and Bacon. McGrath, E., Wypij, D., Rappaport, L., Newburger, J., & Bellinger, C. (2004). Prediction of IQ and achievement at age 8 from neurodevelopmental status at age 1 in
children with D-transposition of the great arteries. Pediatrics, 114, 572–576. McGrath, R., Pogge, D., Stokes, J., & others. (2005). Field reliability of Comprehensive System scoring in an adolescent inpatient sample. Assessment, 12, 199–209. McGrew, K. S. (1997). Analysis of the major intelligence batteries according to a proposed comprehensive Gf-Gc framework. In D. P. Flanagan, J. L. Genshaft, & P. L.
Harrison (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (pp. 151–179). New York: Guilford. McGrew, K. S., & Flanagan, D. P. (1998). The intelligence test desk reference (ITDR): Gf-Gc cross-battery assessment. Boston: Allyn & Bacon. McGue, M., Bouchard, T., Iacono, W., & Lykken, D. (1993). Behavior genetics of cognitive ability: A lifespan perspective. In R. Plomin & G. McClearn (Eds.), Nature,
nurture, and psychology. Washington, DC: American Psychological Association. McGurk, F. C. J. (1953a). On white and Negro test performance and socio-economic factors. Journal of Abnormal and Social Psychology, 48, 448–450. McGurk, F. C. J. (1953b). Socioeconomic status and culturally-weighted test scores of Negro subjects. Journal of Applied Psychology, 37, 276–277. McGurk, F. C. J. (1975). Race differences—twenty years later. Homo, 26, 219–239. McKee, A. C., Cantu, R. C., Nowinski, C. J., & others. (2009). Chronic traumatic encephalopathy in athletes: Progressive tauopathy following repetitive head injury.
Journal of Neuropathology and Experimental Neurology, 68, 709–735. McKee, A. C., Stein, T. D., & Nowinski, C. J. (2012, October 1). The spectrum of disease in chronic traumatic encephalopathy. Brain, online publication McKee-Ryan, F. M., Song, Z., Wanberg, C., & Kinicki, A. (2005). Psychological and physical well-being during unemployment: A meta-analytic study. Journal of
Abnormal Psychology, 90, 53–76. McKey, R. H., & others. (1985). The impact of Head Start on children, families and communities. Washington, DC: U.S. Government Printing Of�ice. McKinley, J. C., & Hathaway, S. R. (1940). A Multiphasic Personality Schedule (Minnesota): II. A differential study of hypochondriasis. Journal of Psychology, 10,
255–268. McKinley, J. C., & Hathaway, S. R. (1944). The MMPI: V. Hysteria, hypomania and psychopathic deviate. Journal of Applied Psychology, 28, 153–174. McKinley, J. C., Hathaway, S. R., & Meehl, P. E. (1948). The MMPI: VI. The K scale. Journal of Consulting Psychology, 12, 20–31. McLean, C. P., Asnaani, A., Litz, B. T., & Hofmann, S. G. (2011). Gender differences in anxiety disorders: Prevalence, course of illness, comorbidity and burden of
illness. Journal of Psychiatric Research, 45, 1027–1035. McMillan, D., Hastings, R., & Coldwell, J. (2004). Clinical and actuarial prediction of physical violence in a forensic intellectual disability hospital: A longitudinal
study. Journal of Applied Research in Intellectual Disabilities, 17, 255–265. McNulty, J., Graham, J., Ben-Porath, Y., & Stein, L. (1997). Comparative validity of MMPI-2 scores of African American and caucasian mental health center clients.
Psychological Assessment, 9, 464–470. McReynolds, P., & Ludwig, K. (1984). Christian Thomasius and the origin of psychological rating scales. Isis, 75, 546–553. McReynolds, P., & Ludwig, K. (1987). On the history of rating scales. Personality and Individual Differences, 8, 281–283. Mednick, S. (1962). The associative basis of the creative process. Psychological Review, 3, 220–232. Mednick, S., & Mednick, M. (1966). Manual: Remote Associates Test. Boston: Houghton Mif�lin. Medoff-Cooper, B., & Ratcliffe, S. (2005). Development of preterm infants: Feeding behaviors and Brazelton Neonatal Behavioral Assessment Scale at 40 and 44
weeks’ post-conceptual age. Advances in Nursing Science, 28, 356–363. Meehl, P. E. (1954). Clinical versus statistical prediction. Minneapolis: University of Minnesota Press. Meehl, P. E. (1965). Seer over sign: The �irst good example. Journal of Experimental Research in Personality, 1, 29–32. Meehl, P. E. (1986). Causes and effects of my disturbing little book. Journal of Personality Assessment, 50, 370–375. Megargee, E. (1972). The California Psychological Inventory handbook. San Francisco: Jossey-Bass. Meichenbaum, D. (1977). Cognitive-behavior modi�ication: An integrative approach. New York: Plenum Press. Meier, S. T. (1984). The construct validity of burnout. Journal of Occupational Psychology, 57, 211–219. Meier, V. J., & Hope, D. A. (1998). Assessment of social skills. In A. S. Bellack & M. Hersen (Eds.), Behavioral assessment: A practical handbook (4th ed.). Boston:
Allyn and Bacon. Meijer, E., Verschuere, B., Merckelbach, H., & Crombez, G. (2008). Sex offender management using the polygraph: A critical review. International Journal of Law
and Psychiatry, 31, 423–429. Meisels, S., & Atkins-Burnett, S. (2005). Developmental screening in early childhood: A guide (5th ed.). Washington, DC: National Association for the Education of
Young Children. Meisels, S., Marsden, D., Wiske, M., & Henderson, L. (1997). Early Screening Inventory-Revised. San Antonio, TX: The Psychological Corporation. Meisels, S., Wiske, M., & Henderson, L. (2008). Early Screening Inventory—Revised. San Antonio, TX: The Psychological Corporation. Melton, G. B. (1995). Review of the Ackerman-Schoendorf Scales for Parent Evaluation of Custody. The Twelfth mental measurements yearbook. Lincoln:
University of Nebraska Press. Melton, G. B., Petrila, J., Poythress, N., & Slobogin, C. (1998). Psychological evaluation for the courts (2nd ed.). New York: Guilford. Mendez, M., Licht, E., & Saul, R. E. (2008). The Frontal Systems Behavior Scale in the evaluation of dementia. International Journal of Geriatric Psychiatry, 23,
1203–1204. Menzies, G. (2003). 1421: The year China discovered America. New York: William Morrow. Mercer, J. R., & Lewis, J. F. (1978). System of Multicultural Pluralistic Assessment. San Antonio, TX: The Psychological Corporation. Merenda, P. F. (1985). Comrey Personality Scales. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques (vol. 4). Kansas City, MO: Test Corporation of America.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 56/70
Messiah, A., Encrenaz, G., Sapinho, D., & others. (2007). Paradoxical increase of positive answers to the Cut-down, Annoyed, Guilt, Eye-opener (CAGE) questionnaire during a period of decreasing alcohol consumption: Results from two population-based surveys in Ile-de-France, 1991 and 2005. Addiction, 103, 598–603.
Messick, S. (1980). Test validity and the ethics of assessment. American Psychologist, 35, 1012–1027. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scienti�ic inquiry into score
meaning. American Psychologist, 50, 741–749. Mevarech, Z. (1995). Metacognition, general ability, and mathematical understanding. Early Education and Development, 6, 155–168. Meyer, G. J. (1997). Assessing reliability: Critical corrections for a critical examination of the Rorschach Comprehensive System. Psychological Assessment, 9, 480–
489. Meyer, G. J., & Eblin, J. J. (2012). An overview of the Rorschach Performance Assessment System (R-PAS). Psychological Injury and Law, 5, 107–121. Meyer, G. J., & Handler, L. (1997). The ability of the Rorschach to predict subsequent outcome: A meta-analysis of the Rorschach Prognostic Rating Scale. Journal
of Personality Assessment, 69, 1–38. Meyer, G. J., Viglione, D. J., Mihura, J. L., Erard, R. E., & Erdberg, P. (2011). Rorschach Performance Assessment System: Administration, coding, interpretation, and
technical manual. Toledo, OH: Rorschach Performance Assessment System. Mickley, J. (1990). Spiritual well-being, religiousness, and hope: Some relationships in a sample of women with breast cancer. Unpublished master’s thesis,
University of Maryland, School of Nursing, College Park, MD. Middleton, H., Keene, R., & Brown, G. (1990). Convergent and discriminant validities of the Scales of Independent Behavior and the revised Vineland Adaptive
Behavior Scales. American Journal of Mental Retardation, 94, 669–673. Miele, F. (1979). Cultural bias in the WISC. Intelligence, 3, 149–164. Milkman, K. L., Chugh, D., & Bazerman, M. H. (2009). How can decision making be improved? Perspectives on Psychological Science, 4, 379–383. Miller, F. G., & Lazowski, L. (1999). The adult SASSI-3 manual. Springville, IN: The SASSI Institute. Miller, F. G., Roberts, J., Brooks, M., & Lazowski, L. (1997). SASSI-3 user’s guide: A quick reference for administration and scoring. Bloomington, IN: Baugh
Enterprises. Miller, G. (2012). The smartphone psychology manifesto. Perspectives on Psychological Science, 7, 221–237. Miller, L. K. (1989). Musical savants: Exceptional skill in the mentally retarded. Hillsdale, NJ: Erlbaum. Miller, S. D., & Duncan, B. L. (2000). Outcome and Session Rating Scales: Administration and scoring manual. Chicago: Institute for the Study of Therapeutic Change. Miller, S. D., Duncan, B. L., Brown, J., Sparks, J., & Claud, D. (2003). The Outcome Rating Scale: A preliminary study of the reliability, validity, and feasibility of a
brief visual analog measure. Journal of Brief Therapy, 2, 91–100. Miller, T. R. (1991). Personality: A clinician’s experience. Journal of Personality Assessment, 57, 415–433. Millman, J., & Greene, J. (1989). The speci�ication and development of tests of achievement and ability. In R. L. Linn (Ed.), Educational measurement (3rd ed.). New
York: ACE/Macmillan. Millon, T. (1969). Modern psychopathology: A biosocial approach to maladaptive learning and functioning. Philadelphia: Saunders. Millon, T. (1981). Disorders of personality: DSM-III, Axis II. New York: Wiley. Millon, T. (1983). Millon Clinical Multiaxial Inventory manual (2nd ed.). Minneapolis, MN: National Computer Systems. Millon, T. (1986). A theoretical derivation of pathological personalities. In T. Millon & G. Klerman (Eds.), Contemporary directions in psychopathology: Toward the
DSM-IV. New York: Guilford. Millon, T. (1987). Manual for the Millon Clinical Multi-axial Inventory-II (MCMI-II) (2nd ed.). Minneapolis, MN: National Computer Systems. Millon, T. (1994). Manual for the Millon Clinical Multi-axial Inventory-III (MCMI-III) (3rd ed.). Minneapolis, MN: National Computer Systems. Millon, T., & Davis, R. (1996). The Millon Clinical Multiaxial Inventory-III (MCMI-III). In C. S. Newmark (Ed.), Major psychological assessment instruments (2nd
ed.). Boston: Allyn and Bacon. Mills, C., & Tissot, S. (1995). Identifying academic potential in students from underrepresented populations: Is using the Raven’s Progressive Matrices a good
idea? Gifted Child Quarterly, 39, 209–217. Mills, C., & Tissot, S. (1995). Identifying academic potential in students from underrepresented populations: Is using the Raven’s Progressive Matrices a good
idea? Gifted Child Quarterly, 39, 209–217. Mills, C., Potenza, M., Fremer, J., & Ward, W. (2002). Computer-based testing: Building the foundation for future assessments. Mahwah, NJ: Erlbaum. Milner, B. (1968). Disorders of memory after brain lesions in man. Neuropsychologia, 6, 175–179. Mischel, W. (1968). Personality and assessment. New York: Wiley. Mischel, W., Shoda, Y., & Mendoza-Denton, R. (2002). Situation-behavior pro�iles as a locus of consistency in personality. Current Directions in Psychological
Science, 11, 50–54. Mitchell, T. W., & Klimoski, R. J. (1986). Estimating the validity of cross-validity estimation. Journal of Applied Psychology, 71, 311–317. Mitchell, V. (2007). Earning a secure attachment style: A narrative of personality change in adulthood. In R. Josselson, A. Lieblich, & D. P. McAdams (Eds.), The
meaning of others: Narrative studies of relationships (pp. 93–116). Washington, DC: American Psychological Association. Moberg, D. O. (1971). Spiritual well-being: Background and issues. Washington, DC: White House Conference on Aging. Montague, M., & Bos, C. S. (1990). Cognitive and meta-cognitive characteristics of eighth grade students’ mathematical problem solving. Learning and Individual
Differences, 2, 371–388. Moore, E. G. J. (1986). Family socialization and the IQ-test performance of traditionally and transracially adopted children. Developmental Psychology, 22, 317–
326. Moore, R. C., Viglione, D. J., Rosenfarb, I. S., Patterson, T. L., & Mausbach, B. T. (2012, November 12). Rorschach measures of cognition relate to everyday and social
functioning in schizophrenia. Psychological Assessment, online publication. Moore, W. P. (1994). The devaluation of standardized testing: One district’s response. Applied Measurement in Education, 7, 343–368. Moreland, K. L. (1992). Computer-assisted psychological assessment. In M. Zeidner & R. Most (Eds.), Psychological testing: An inside view. Palo Alto, CA:
Consulting Psychologists Press. Moreno, K. E., & Segall, D. O. (1997). Reliability and construct validity of the CAT-ASVAB In W. A. Sands, B. K. Waters, & J. R. McBride (Eds.), Computerized adaptive
testing: From inquiry to operation. Washington, DC: American Psychological Association. Morgan, C. D., & Murray, H. A. (1935). A method for investigating phantasies: The Thematic Apperception Test. Archives of Neurology and Psychiatry, 34, 289–306.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 57/70
Morgan, C. D., Shoenberg, M., Dorr, D., & Burke, M. (2002). Overreport on the MCMI-III: Concurrent validation with the MMPI-2 using a psychiatric inpatient sample. Journal of Personality Assessment, 78, 288–300.
Mori, L., & Armendariz, G. (2001). Analogue assessment of child behavior problems. Psychological Assessment, 13, 36–45. Morrison, M. W., Gregory, R. J., & Paul, J. J. (1979). Reliability of the Finger Tapping Test and a note on sex differences. Perceptual and Motor Skills, 48, 139–142. Morrison, T., & Morrison, M. (1995). A meta-analytic assessment of the predictive validity of the quantitative and verbal components of the Graduate Record
Examination with graduate grade point average representing the criterion of graduate success. Educational and Psychological Measurement, 55, 309–316. Morrow, C., Bandstra, E., Anthony, J., & others. (2001). In�luence of prenatal cocaine exposure on full-term infant neurobehavioral functioning. Neurotoxicology
and Teratology, 23, 533–544. Moruzzi, G., & Magoun, H. W. (1949). Brain stem and reticular formation and activation of the EEG. Electroencephalography and Clinical Neurophysiology, 1, 455–
473. Motowidlo, S. J., Carter, G., Dunnette, M., Tippins, N., Werner, S., Burnett, J., & Vaughan, M. (1992). Studies of the structured behavioral interview. Journal of
Applied Psychology, 77, 571–587. Motta, R. W., Little, S., & Tobin, M. (1993). The use and abuse of human �igure drawings. School Psychology Quarterly, 8, 162–169. Mount, M., Witt, L., & Barrick, M. (2000). Incremental validity of empirically keyed biodata scales over GMA and the �ive factor personality constructs. Personnel
Psychology, 53, 299–323. Mountain, M., & Snow, W. (1993). Wisconsin Card Sorting Test as a measure of frontal pathology: A review. Clinical Neuropsychologist, 7, 108–118. Muchinsky, P. (2003). Psychology applied to work: An introduction to industrial and organizational psychology (7th ed.). Belmont, CA: Wadsworth. Murphy, K. R. (1984). Review of Armed Services Vocational Aptitude Battery. In D. Keyser & R. Sweetland (Eds.), Test critiques (vol. 1). Kansas City, MO: Test
Corporation of America. Murphy, K. R. (1992). Review of TONI-2. The eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Murphy, K. R., & Davidshofer, C. O. (1988). Psychological testing. Englewood Cliffs, NJ: Prentice Hall. Murphy, K. R., & Davidshofer, C. O. (2004). Psychological testing (6th ed.). Englewood Cliffs, NJ: Prentice Hall. Murphy, K. R., & Pardaffy, V. A. (1989). Bias in behaviorally anchored rating scales: Global or scale-speci�ic? Journal of Applied Psychology, 74, 343–346. Murphy, K. R., Jako, R., & Anhalt, R. (1993). Nature and consequences of halo error: A critical analysis. Journal of Applied Psychology, 78, 218–225. Murray, H. A. (1938). Explorations in personality. New York: Oxford University Press. Murray, H. A. (1943). Thematic Apperception Test—Manual. Cambridge, MA: Harvard University Press. Museum of Modern Art. (1955). The family of man. New York: Maco Magazine Corporation. Myers, D. (2002). Social psychology (7th ed.). New York: McGraw-Hill. Myers, I. B., & McCaulley, M. H. (1985). Manual: A guide to the development and use of the Myers-Briggs Type Indicator. Palo Alto, CA: Consulting Psychologists
Press. Myers, I., & McCaulley, M. (1985). A guide to the development and use of the Myers-Briggs Type Indicator. Palo Alto, CA: Consulting Psychologists Press. Myrtek, M. (2007). Type a behavior and hostility as independent risk factors for coronary heart disease. In J. Jordan, B. Bardé, & A. M. Zeiher (Eds.), Contributions
toward evidence-based psychocardiology: A systematic review of the literature (pp. 159–183). Washington, DC: American Psychological Association. Naglieri, J. A. (1981). Concurrent validity of the Revised Peabody Picture Vocabulary Test. Psychology in the Schools, 18, 286–289. Naglieri, J. A. (1988). Draw A Person: A quantitative scoring system. San Antonio, TX: The Psychological Corporation. Naglieri, J. A., & Das, J. P. (2005a). Planning, Attention, Simultaneous, and Successive (PASS) cognitive processes as a model for intelligence. Journal of
Psychoeducational Assessment, 8, 303–337. Naglieri, J. A., & Das, J. P. (2005b). Planning, Attention, Simultaneous, Successive (PASS) theory: A revision of the concept of intelligence. In D. P. Flanagan & P. L.
Harrison (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (pp. 120–135). New York: Guilford Press. Naglieri, J. A., & Paolitto, A. (n.d.). Attention de�icit diagnosis and treatment: Current status/future directions. Unpublished paper available at:
www.riverpub.com/products/cas/cas_add.html (http://www.riverpub.com/products/cas/cas_add.html) . Naglieri, J. A., & Pfeiffer, S. (1983). Stability, concurrent and predictive validity of the PPVT-R. Journal of Clinical Psychology, 39, 965–967. Naglieri, J. A., & Pfeiffer, S. (1992). Performance of disruptive behavior disordered and normal samples on the Draw A Person: Screening Procedure for Emotional
Disturbance. Psychological Assessment, 4, 156–159. Naglieri, J. A., & Rojahn, J. (2001). Intellectual classi�ication of Black and White children in special education programs using the WISC—III and the cognitive
assessment system. American Journal on Mental Retardation, 106, 359–367. Naglieri, J. A., & Yazzie, C. (1983). Comparison of the WISC-R and PPVT-R with Navajo children. Journal of Clinical Psychology, 39, 598–600. Naglieri, J. A., Das, J. P., & Goldstein, S. (2012). Cognitive Assessment System—Second edition. Austin, TX: PRO-ED. Naglieri, J. A., Rojahn, J., Matto, H. C., & Aquilino, S. A. (2005). Black-White differences in cognitive processing: A study of the planning, attention, simultaneous,
and successive theory of intelligence. Journal of Psychoeducational Assessment, 23, 146–160. Naglieri, J. A., Taddei, S., & Williams, K. M. (2012, September 17). Multigroup con�irmatory factor analysis of U.S. and Italian children’s performance on the PASS
theory of intelligence as measured by the Cognitive Assessment System. Psychological Assessment, online publication. Naglieri, J., & Das, J. (1990). Planning, attention, successive, and simultaneous cognitive processes as a model for intelligence. Journal of Psychoeducational
Assessment, 8, 165–170. Naglieri, J., McNeish, T., & Bardos, A. (1991). Draw-A-Person: Screening Procedure for Emotional Disturbance. Austin, TX: ProEd. National Association of School Psychologists. (1992). Principles for professional ethics. Silver Spring, MD: Author. National Association of School Psychologists. (2010). Principles for professional ethics. Silver Springs, MD: Author. National Joint Committee on Learning Disabilities. (1988). A position paper of the National Trust Committee on Learning Disabilities. Journal of Learning
Disabilities, 21, 53–55. Naugle, R. I., Chelune, G., & Tucker, G. (1993). Validity of the Kaufman Brief Intelligence Test. Psychological Assessment, 5, 182–186. Nauta, W. J. H. (1971). The problem of the frontal lobe. Journal of Psychiatric Research, 8, 167–187. Naveh-Benjamin, M., McKeachie, W. J., & Lin, Y. (1987). Two types of test-anxious students: Support for an information processing model. Journal of Educational
Psychology, 79, 131–136. Needleman, H. L., Gunnoe, C., Leviton, A., Reed, R., Peresie, H., Maher, C., & Barrett, P. (1979). De�icits in psychologic and classroom performance of children with
elevated dentine lead levels. The New England Journal of Medicine, 300, 689–695.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 58/70
Needleman, H. L., Schell, A., Bellinger, D., Leviton, A., & Allred, E. (1990). The long-term effects of exposure to low doses of lead in childhood. New England Journal of Medicine, 322, 83–88.
Neisser, U. (Ed.). (1998). The rising curve: Long-term gains in IQ and related measures. Washington, DC: American Psychological Association. Neisser, U., Boodoo, G., & Bouchard, T., & others. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51, 77–101. Nelson, R., & Piedmont, R. L. (2008, August). Psychometric utility of the ASPIRES Scales in non-Christian samples. Paper presented at the American Psychological
Association Conference, Boston. Nester, M. A. (1994). Psychometric testing and reasonable accommodation for persons with disabilities. In S. M. Bruyere & J. O’Keeffe (Eds.), Implications of the
Americans with Disabilities Act for psychology. New York: Springer. Nestor, P. G., & Schutt, R. K. (2012). Research methods in psychology: Investigating human behavior. Thousand Oaks, CA: SAGE. Nettelbeck, T., & Wilson, C. (2004). The Flynn effect: Smarter, not faster. Intelligence, 32, 85–93. Netter, B., & Viglione, D., Jr. (1994). An empirical study of malingering schizophrenia on the Rorschach. Journal of Personality Assessment, 62, 45–57. Nevo, B. (1985). Face validity revisited. Journal of Educational Measurement, 22, 287–293. Nevo, B. (1992). Examinee feedback: Practical guidelines. In M. Zeidner & R. Most (Eds.), Psychological testing: An inside view. Palo Alto, CA: Consulting
Psychologists Press. Newland, T. E. (1971). Blind Learning Aptitude Test. Champaign: University of Illinois Press. Newsome, S., Day, A., & Catano, V. (2000). Assessing the predictive validity of emotional intelligence. Personality and Individual Differences, 29, 1005–1016. Nichols, S. L., Glass, G. V., & Berliner, D. C. (2006). High-stakes testing and student achievement: Problems for the No Child Left Behind Act. Tempe, AZ: Education
Policies Study Laboratory. Nieuwenhuis-Mark, R. E. (2010). The death knoll for the MMSE: Has it outlived its purpose? Journal of Geriatric Psychiatry and Neurology, 23, 151–157. Nihira, K., Leland, H., & Lambert, N. (1993). Adaptive Behavior Scale-Residential and Community (2nd ed.). Washington, DC: American Association on Mental
Retardation. Nijenhuis, J., & van der Flier, H. (1997). Comparability of GATB scores for immigrants and majority group members: Some Dutch �indings. Journal of Applied
Psychology, 82, 675–687. Nijenhuis, J., Evers, A., & Mur, J. (2000). Validity of the Differential Aptitude Test for the assessment of immigrant children. Educational Psychology, 20, 99–115. Nisan, M., & Kohlberg, L. (1982). Universality and cross-cultural variation in moral development: A longitudinal and cross-sectional study in Turkey. Child
Development, 53, 865–876. Nisbett, R. E., Aronson, J., Blair, C., & others. (2012). Intelligence: New �indings and theoretical developments. American Psychologist, 67, 130–139. Norris, G., & Tate, R. (2000). The Behavioural Assessment of the Dysexecutive Syndrome (BADS): Ecological, concurrent and construct validity.
Neuropsychological Rehabilitation, 10, 33–45. Nottingham, E. J., & Mattson, R. E. (1981). A validation study of the Competency Screening Test. Law and Human Behavior, 5, 329–335. Nunnally, J. (1967). Psychometric theory. New York: McGraw-Hill. Nunnally, J. C. (1978). Psychometric theory (2nd ed.). New York: McGraw-Hill. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). New York: McGraw-Hill. O’Neill, J., Jacobson, S., & Jacobson, J. (1994). Evidence of observer reliability for the Fagan Test of Infant Intelligence (FTII). Infant Behavior and Development, 17,
465–469. Oakes, L. M. (2009). The “Humpty Dumpty Problem” in the study of early cognitive development: Putting the infant back together again. Perspectives on
Psychological Science, 4, 352–358. Ochse, R. (1990). Before the gates of excellence. Cambridge, England: Cambridge University Press. Oei, T., Evans, L., & Crook, G. M. (1990). Utility and validity of the STAI with anxiety disorder patients. British Journal of Clinical Psychology, 29, 429–432. Offer, D., & Sabshin, M. (1966). Normality: Theoretical and clinical concepts of mental health. New York: Basic Books. Ogg, Brinkman, T. M., Dedrick, R. F., & Carlson, J. S. (2010). Factor structure and invariance across gender of the Devereux Early Childhood Assessment Protective
Factor Scale. School Psychology Quarterly, 25, 107–118. Ogloff, J. R., Wong, S., & Greenwood, A. (1990). Treating criminal psychopaths in a therapeutic community program. Behavioral Science and the Law, 8, 181–190. Oles, H. J., & Davis, G. D. (1977). Publishers violate APA standards on test distribution. Psychological Reports, 41, 713–714. Ollendick, T. H. (1983). Reliability and validity of the Revised Fear Survey Schedule for Children (FSSC-R). Behavior Research and Therapy, 21, 685–692. Olson, H. C. (1994). Fetal alcohol syndrome. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Olson-Buchanan, J., Drasgow, F., Moberg, P., Mead, A., Keenan, P., & Donovan, M. (1998). Interactive video assessment of con�lict resolution skills. Personnel
Psychology, 51, 1–24. Ornberg, B., & Zalewski, C. (1994). Assessment of adolescents with the Rorschach: A critical review. Assessment, 1, 209–217. Ortner, T. (2008). Effects of changed item order: A cautionary note to practitioners on jumping to computerized adaptive testing for personality assessment.
International Journal of Selection and Assessment, 16, 249–257. Ortner, T. M., & Caspers, J. (2011). Consequences of test anxiety on adaptive versus �ixed item testing. European Journal of Psychological Assessment, 27, 157–163. OSS Assessment Staff. (1948). Assessment of men: Selection of personnel for the Of�ice of Strategic Services. New York: Rinehart. Otis, A. S. (1918). An absolute point scale for the group measure of intelligence. Journal of Educational Psychology, 9, 238–261, 333–348. Ottinger, R., & Kurzon, C. (2007, May 21). Biodata: The measure of an applicant?. New York Law Journal, online publication (3 pp.). Owens, W. A. (1976). Background data. In M. D. Dunnette (Ed.), Handbook of industrial and organizational psychology. Chicago: Rand McNally. Ownby, R. L. (1991). Psychological reports: A guide to report writing in professional psychology (2nd ed.). Brandon, VT: Clinical Psychology Publishing Co. Paloutzian, R. F., & Ellison, C. W. (1982). Loneliness, spiritual well-being and the quality of life. In L. A. Peplau & D. Perlman (Eds.), Loneliness: A sourcebook of
current theory, research and therapy. New York: Wiley. Panigua, F. (1994). Assessing and treating culturally diverse clients: A practical guide. Thousand Oaks, CA: Sage. Park, N., & Peterson, C. (2009). Achieving and sustaining a good life. Perspectives on Psychological Science, 4, 422–428. Parsons, F. (1909). Choosing a vocation. Boston: Houghton Mif�lin. Parsons, T. D., Rizzo, A. A., Brennan, J., Bittman, M., & Zelinski, E. (2008). Assessment of executive functioning using virtual reality: Virtual Environment Grocery
Store. Gerontechnology, 7, 189–189.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 59/70
Patterson, C. (1980). An alternative perspective—lead pollution in the human environment. In Lead in the human environment. Washington, DC: National Academy of Sciences.
Patton, J. R., Payne, J. S., & Beirne-Smith, M. (1986). Mental retardation (2nd ed.). Columbus, OH: Merrill. Paul, A. M. (2004). The cult of personality. New York: Free Press. Paul, L. K., Brown, W. S., Adolphs, R., & others. (2007). Agenesis of the corpus callosum: Genetic, developmental and functional aspects of connectivity. Nature, 8,
287–299. Paulhus, D., Fridhandler, B., & Hayes, S. (1997). Psychological defense: Contemporary theory and research. In R. Hogan, J. Johnson, & S. Briggs (Eds.), Handbook of
personality psychology. San Diego: Academic Press. Paulman, R. G., & Kennelly, K. J. (1984). Test anxiety and ineffective test taking: Different names, same construct? Journal of Educational Psychology, 76, 279–288. Payne, A. F. (1928). Sentence completions. New York: New York Guidance Clinic. Pearson, K. (1914, 1924, 1930ab). The life, letters, and labours of Francis Galton (Volumes I, II, III, IIIb). Cambridge: Cambridge University Press. Pedersen, N. L., Plomin, R., Nesselroade, J., & McClearn, G. (1992). A quantitative genetic analysis of cognitive abilities during the second half of the life span.
Psychological Science, 3, 346–353. Pen�ield, W. (1958). Functional localization in temporal and deep sylvian areas. Research Publication, Association of Nervous and Mental Disease, 36, 210–217. Pen�ield, W., & Evans, J. (1935). The frontal lobe in man: A clinical study of maximum removals. Brain, 58, 115–133. Pen�ield, W., & Jasper, H. (1959). Epilepsy and the functional anatomy of the human brain. Boston: Little, Brown. Peretz, H., & Fried, Y. (2012). National cultures, performance appraisal practices, and organizational absenteeism and turnover: A study across 21 countries.
Journal of Applied Psychology, 97, 448–459. Perry, J. C. (1990). The Defense Mechanism Rating Scales (5th ed.). Cambridge, MA: J. C. Perry. Perry, J. C., & Henry, M. (2004). Studying defense mechanisms in psychotherapy using the defense mechanism rating scales. In U. Hentschel, G. Smith, J. Draguns,
& W. Ehlers (Eds.), Defense mechanisms: Theoretical, research and clinical perspectives (pp. 165–192). Oxford, England: Elsevier. Perry, J. C., Beck, S. M., Constantinides, P., & Foley, J. (2009). Studying change in defensive functioning in psychotherapy using the defense mechanism rating
scales: Four hypotheses, four cases. In R. A. Levy & J. S. Ablon (Eds.), Handbook of evidence-based psychodynamic psychotherapy: Bridging the gap between science and practice (pp. 121–153). Totowa, NJ, US: Humana Press.
Pervin, L. A. (1993). Personality: Theory and research (6th ed.). New York: Wiley. Petersen, N. S., Kolen, M. J., & Hoover, H. D. (1989). Scaling, norming, and equating. In R. L. Linn (Ed.), Educational measurement (3rd ed.). New York: American
Council on Education/Macmillan. Peterson, C. (2000). Optimistic explanatory style and health. In J. Gillham (Ed.), The science and optimism of hope (pp. 145–162). Philadelphia: Templeton
Foundation Press. Pfeiffer, E. (1975). A short portable mental status questionnaire for the assessment of organic brain de�icit in elderly patients. Journal of the American Geriatrics
Society, 23, 433–441. Phelps, L., & Ensor, A. (1986). Concurrent validity of the WISC-R using deaf norms and the Hiskey-Nebraska. Psychology in the Schools, 23, 138–141. Phillips, S. E. (1994). High-stakes testing accommodations: Validity versus disabled rights. Applied Measurement in Education, 7, 93–120. Piaget, J. (1932). The moral judgment of the child. London: Kegan Paul. Piaget, J. (1972). The psychology of intelligence. Totowa, NJ: Little�ield Adams. Piedmont, R. L. (1999). Does spirituality represent the sixth factor of personality? Spiritual transcendence and the Five-Factor Model. Journal of Personality, 67,
985–1013. Piedmont, R. L. (2001). Spiritual transcendence and the scienti�ic study of spirituality. Journal of Rehabilitation, 67, 4–14. Piedmont, R. L. (2004). Spiritual transcendence as a predictor of psychosocial outcome from an outpatient substance abuse program. Psychology of Addictive
Behaviors, 18, 213–222. Piedmont, R. L. (2010). Assessment of Spirituality and Religious Sentiments (ASPIRES): Technical manual (2nd ed.). Timonium, MD: Author. Piedmont, R. L., & Weinstein, H. P. (1993). A psychometric evaluation of the new NEO-PIR Facet Scales for Agreeableness and Conscientiousness. Journal of
Personality Assessment, 60, 302–318. Piedmont, R. L., Werdel, M., & Fernando, M. (2009). The utility of the Assessment of Spirituality and Religious Sentiments (ASPIRES) scale with Christians and
Buddhists in Sri Lanka. Research in the Social Scienti�ic Study of Religion, 20, 131–143. Piersma, H., & Boes, J. (1997). MCMI-III as a treatment outcome measure for psychiatric inpatients. Journal of Clinical Psychology, 53, 825–832. Piirto, J. (1998). Understanding those who create. Scottsdale, AZ: Gifted Psychology Press. Pinals, D., Tillbrook, C., & Mumley, D. (2006). Practical application of the MacArthur Competence Assessment Tool—Criminal Adjudication (MacCAT-CA) in a
public sector forensic setting. Journal of the American Academy of Psychiatry and Law, 34, 179–188. Pintner, R. (1917). The mentality of the dependent child. Journal of Educational Psychology, 8, 220–238. Pintner, R. (1921). Intelligence. In E. L. Thorndike (Ed.), Intelligence and its measurement: A symposium. Journal of Educational Psychology, 12, 123–147, 195–
216. Piotrowski, C. (1996). The status of Exner’s Comprehensive System in contemporary research. Perceptual and Motor Skills, 82, 1341–1342. Piotrowski, Z. A. (1964). A digital computer administration of inkblot test data. Psychiatric Quarterly, 38, 1–26. Pirozzolo, F. J., Hansch, E., Mortimer, J., Webster, D., & Kuskowski, A. (1982). Dementia in Parkinson disease: A neuropsychological analysis. Brain and Cognition, 1,
71–83. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57, 210–221. Plaisted, J. R., & Golden, C. J. (1982). Test-retest reliability of the clinical, factor and localization scales of the Luria-Nebraska Neuropsychological Battery.
International Journal of Neuroscience, 17, 163–167. Plaud, J. J., & Eifert, G. (Eds.). (1998). From behavior theory to behavior therapy. Boston: Allyn and Bacon. Polivy, J., & Herman, C. P. (1993). Etiology of binge eating: Psychological mechanisms. In C. G. Fairburn & G. T. Wilson (Eds.), Binge eating: Nature, assessment, and
treatment (pp. 173–205). New York: Guilford Press. Pollack, R. H. (1971). Binet on perceptual-cognitive development or Piaget-come-lately. Journal of the History of the Behavioral Sciences, 7, 370–374. Pollens, R., McBratnie, B., & Burton, P. (1988). Beyond cognition: Executive functions. Cognitive Rehabilitation, 6, 26–33. Poortinga, Y. H., & Van de Vijver, F. J. R. (2004). Cultures and cognition: Performance differences and invariant structures. In R. J. Sternberg & E. L. Grigorenko
(Eds.), Culture and competence: Contexts of life success (pp. 139–162). Washington, DC: American Psychological Association.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 60/70
Pope, K. S. (1992). Responsibilities in providing psychological test feedback to clients. Psychological Assessment, 4, 268–271. Popham, W. J. (1978). Criterion-referenced measurement. Englewood Cliffs, NJ: Prentice Hall. Porch, B. (2001). Porch Index of Communicative Ability—2001 Revision. Austin, TX: Pro-Ed. Porteus, S. D. (1931). The psychology of a primitive people: A study of the Australian aborigine. London: Edward Arnold & Co. Porteus, S. D. (1965). Porteus Maze Test. Fifty years’ application. Palo Alto, CA: Paci�ic Books. Powers, D. (2004). Validity of Graduate Record Examinations (GRE) General Test scores for admissions to colleges of Veterinary Medicine. Journal of Applied
Psychology, 89, 208–219. Powers, K., & Hagans-Murillo, K. (2004). Twenty-�ive years after Larry P.: The California response to over-representation of African-Americans in special
education. California School Psychologist, 9, 145–158. Poythress, N., Monahan, J., Bonnie, R., Otto, R., & Hoge, S. (2002). Adjudicative competence: The MacArthur studies. New York: Kluwer/Plenum. Prentky, R. (2001). Mental illness and roots of genius. Creativity Research Journal, 13, 95–104. Prewett, P. N. (1995). A comparison of two screening tests (the Matrix Analogies Test-Short Form and the Kaufman Brief Intelligence Test) with the WISC-III.
Psychological Assessment, 7, 69–72. Prout, H., & Schwartz, J. (1984). Validity of the PPVT-R with mentally retarded adults. Journal of Clinical Psychology, 40, 584–587. Psychological Corporation. (1994). WISC-III Writer manual. San Antonio, TX: Author. Purish, A. (2001). Misconceptions about the Luria-Nebraska Neuropsychological Battery. Neurorehabilitation, 16, 275–280. Pyle, W. H. (1913). The examination of school children. New York: Macmillan. Qu, C. (1997). Reliability and validity of the Hiskey-Nebraska Test of Learning Aptitude (H-NTLA) in testing China’s deaf children. Chinese Mental Health Journal,
11, 70–72. Quek, K. F., Low, W. Y., Razack, A. H., Loh, C. S., & Chuak, C. B. (2004). Reliability and validity of the Spielberger State-Trait Anxiety Inventory (STAI) among
urological patients: A Malaysian study. Medical Journal of Malaysia, 59, 258–267. Ramey, C. T., & Ramey, S. (1998). Early intervention and early experience. American Psychologist, 53, 109–10. Ramos, E., Alfonso, V. C., & Schermerhorn, S. M. (2009). Graduate students’ administration and scoring errors on the Woodcock-Johnson III Tests of Cognitive
Abilities. Psychology in the Schools, 46, 650–657. Ranseen, J., Campbell, D., & Baer, R. (1998). NEO PI-R pro�iles of adults with attention de�icit disorder. Assessment, 5, 19–24. Rasch, G. (1960). Probabilistic models for some intelligence and attainment tests. Copenhagen: Denmarks Paedagogiske Institut. Raven, J. (2000). The Raven’s Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41, 1–48. Raven, J. C. (1938). Progressive Matrices. London: Lewis. Raven, J. C. (1965). The Coloured Progressive Matrices Test. London: Lewis. Raven, J. C., & Summers, B. (1986). Manual for Raven’s Progressive Matrices and Vocabulary Scales—research supplement no. 3. London: Lewis. Raven, J. C., Court, J. H., & Raven, J. (1983). Manual for Raven’s Progressive Matrices and Vocabulary Scales (Section 3)—Standard Progressive Matrices (1983
edition). London: Lewis. Raven, J. C., Court, J. H., & Raven, J. (1986). Manual for Raven’s Progressive Matrices and Vocabulary Scales (Section 2)—Coloured Progressive Matrices (1986
edition, with U.S. norms). London: Lewis. Raven, J. C., Court, J. H., & Raven, J. (1992). Standard Progressive Matrices. 1992 Edition. Oxford: Oxford Psychologists Press. Reddon, J. R., & Jackson, D. N. (1989). Readability of three adult personality tests: Basic Personality Inventory, Jackson Personality Inventory, and Personality
Research Form-E. Journal of Personality Assessment, 53, 180–183. Reeves, D., & Wedding, D. (1994). The clinical assessment of memory: A practical guide. New York: Springer. Regenwetter, M. (2009). Perspectives on preference aggregation. Perspectives on Psychological Science, 4, 403–407. Rehm, L. P. (1984). Self-management therapy for depression. Advances in Behavior Research and Therapy, 6, 83–98. Rehm, L. P., Kornblith, S. J., O’Hara, M. W., & others. (1981). An evaluation of major components in a self-control therapy program for depression. Behavior
Modi�ication, 5, 459–490. Reid-Arndt, S. A., Nehl, C., & Hinkebein, J. (2007). The Frontal Systems Behavior Scale (frSBe) as a predictor of community integration following a traumatic brain
injury. Brain Injury, 21, 1361–1369. Reilly, R. R., & Chao, G. T. (1982). Validity and fairness of some alternative employee selection procedures. Personnel Psychology, 35, 1–63. Reise, S., Ainsworth, A., & Haviland, M. (2005). Item response theory: Fundamentals, applications, and promise in psychological research. Current Directions in
Psychological Science, 14, 95–101. Reitan, R. M. (1984). Aphasia and sensory perceptual de�icits in adults. Tucson, AZ: Neuropsychology Press. Reitan, R. M., & Wolfson, D. (1993). The Halstead-Reitan Neuropsychological Test Battery: Theory and clinical interpretation (2nd ed.). Tucson, AZ:
Neuropsychology Press. Reppermund, S., Brodaty, H., Crawford, J. D., & others. (2011). The relationship of current depressive symptoms and past depression with cognitive impairment
and instrumental activities of daily living in an elderly population: The Sydney Memory and Ageing Study. Journal of Psychiatric Research, 45, 1600–1607. Reschly, D., Myers, T., & Hartel, C. (2002). Mental retardation: Determining eligibility for Social Security bene�its. Washington, DC: National Academies Press. Rest, J. R. (1979). The De�ining Issues Test: Manual. Minneapolis: University of Minnesota Press. Rest, J. R. (1986). Moral research methodology. In S. Modgil & C. Modgil (Eds.), Lawrence Kohlberg: Consensus and controversy. Philadelphia: Taylor & Francis. Rest, J. R., & Thoma, S. J. (1985). Relation of moral judgment to formal education. Developmental Psychology, 21, 709–714. Rest, J. R., Thoma, S., Narvaez, D., & Bebeau, M. (1997). Alchemy and beyond: Indexing the De�ining Issues Test. Journal of Educational Psychology, 89, 498–507. Rey, A. (1964). L’examen clinique en psychologie. Paris: Presses Universitaires de France. Reynolds, C. R. (1994). Bias in testing. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Reynolds, C. R. (1998). Cultural bias in testing of intelligence and personality. In A. Bellack & M. Hersen (Series Eds.) & C. Belar (Vol. Ed.), Comprehensive clinical
psychology: Sociocultural and individual differences. New York: Elsevier Science. Reynolds, C. R., & Brown, R. T. (1984a). Bias in mental testing: An introduction to the issues. In Reynolds, C. R., & Brown, R. T. (Eds.), Perspectives on bias in mental
testing. New York: Plenum Press. Reynolds, C. R., & Brown, R. T. (Eds.). (1984b). Perspectives on bias in mental testing. New York: Plenum Press.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 61/70
Reynolds, C. R., Chastain, R. L., Kaufman, A. S., & McLean, J. E. (1987). Demographic characteristics and IQ among adults: Analysis of the WAIS-R standardization sample as a function of the strati�ication variables. Journal of School Psychology, 25, 323–342.
Reynolds, C. R., Lowe, P. A., & Saenz, A. L. (1999). The problem of bias in psychological assessment. In C. R. Reynolds & T. B. Gutkin (Eds.), The handbook of school psychology (3rd ed.). New York: Wiley.
Riccio, C., Reynolds, C., & Lowe, P. (2001). Clinical applications of continuous performance tests: Measuring attention and impulsive responding in children and adults. New York: Wiley.
Richards, P. S. (1991). The relation between conservative religious ideology and principled moral reasoning: A review. Review of Religious Research, 32, 359–368. Richards, P. S., & Bergin, A. E. (2005). Religious and spiritual assessment. In P. S. Richards & A. E. Bergin (Eds.), A spiritual strategy for counseling and
psychotherapy (2nd ed., pp. 219–249). Washington, DC: American Psychological Association. Richards, P. S., & Davison, M. L. (1992). Religious bias in moral development research: A psychometric investigation. Journal for the Scienti�ic Study of Religion, 31,
467–485. Rieber, R. W. (Ed.). (1980). Wilhelm Wundt and the making of a scienti�ic psychology. New York: Plenum Press. Rinas, J., & Clyne-Jackson, S. (1988). Professional conduct and legal concerns in mental health practice. Norwalk, CT: Appleton & Lang. Ritter, N., Kilinc, E., Navruz, B., & Bae, Y. (2011). Test review: Test of Nonverbal Intelligence-4 (TONI-4). Journal of Psychoeducational Assessment, 29, 384–388. Ritzler, B. A., Sharkey, K. J., & Chudy, J. (1980). A comprehensive projective alternative to the TAT. Journal of Personality Assessment, 44, 358–362. Ritzler, B., Erard, R., & Pettigrew, G. (2002). Protecting the integrity of Rorschach expert witnesses: A reply to Grove and Barden (1999) Re: The admissibility of
testimony under Daubert/Kumho analyses. Psychology, Public Policy, and Law, 8, 201–215. Roberts, B. W., Walton, K. E., & Viechtbauer, W. (2006). Patterns of mean-level change in personality traits across the life course: A meta-analysis of longitudinal
studies. Psychological Bulletin, 131, 1–25. Robertson, I. H., Ward, T., Ridgeway, V., & NimmoSmith, I. (1994). Test of Everyday Attention (TEA). Gaylord, MI: National Rehabilitation Services. Robertson, I. H., Ward, T., Ridgeway, V., & NimmoSmith, I. (1996). The structure of normal human attention: The Test of Everyday Attention. Journal of the
International Neuropsychological Society, 2, 525–534. Robertson, I., & Smith, M. (2001). Personnel selection. Journal of Occupational and Organizational Psychology, 74, 441–472. Robins, D. L. (2008). Screening for autism in primary care settings. Autism, 12, 537–556. Robins, D. L., & Dumont-Mathieu, T. (2006). The Modi�ied Checklist for Autism in Toddlers (M-CHAT): A review of current �indings and future directions. Journal
of Developmental and Behavioral Pediatrics, 27, S111–S119. Robins, D. L., Fein, D., & Barton, M. (1999). The Modi�ied Checklist for Autism in Toddlers (M-CHAT). Storrs, CT: University of Connecticut. Roese, N. J., & Amir, E. (2009). Human-android interaction in the near and distant future. Perspectives on Psychological Science, 4, 429–434. Rogers, B. (1989). Review of Metropolitan Achievement Test, Sixth Edition. The tenth mental measurements yearbook. Lincoln: University of Nebraska Press. Rogers, B. G. (1992). Review of GED. The eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Rogers, C. R. (1951). Client-centered therapy: Its current practice, implications, and theory. Boston: Houghton Mif�lin. Rogers, C. R. (1961). On becoming a person: A therapist’s view of psychotherapy. Boston: Houghton Mif�lin. Rogers, C. R. (1980). A way of being. Boston: Houghton Mif�lin. Rogers, C. R., & Dymond, R. F. (Eds.). (1954). Psychotherapy and personality change: Co-ordinated research studies in the client-centered approach. Chicago:
University of Chicago Press. Rogers, R. (1984). Rogers Criminal Responsibility Assessment Scales. Odessa, FL: Psychological Assessment Resources. Rogers, R. (1986). Conducting insanity evaluations. New York: Van Nostrand Reinhold. Rogers, R. (1986). Conducting insanity evaluations. Odessa, FL: Psychological Assessment Resources. Rogers, R. (2001). Schedule of Affective Disorders and Schizophrenia (SADS). In R. Rogers (Ed.), Handbook of diagnostic and structured interviewing. New York:
Guilford. Rogers, R. (Ed.). (2008). Clinical assessment of malingering and deception (3rd ed.). New York: Guilford. Rogers, R., & Johansson-Love, J. (2009). Evaluating competency to stand trial with evidence-based practice. Journal of the American Academy of Psychiatry & Law,
37, 450–460. Rogers, R., & Sewell, K. (1999). The R-CRAS and insanity evaluations: A re-examination of construct validity. Behavioral Sciences and the Law, 17, 181–194. Rogers, R., Bagby, M., & Dickens, S. (1992). Structured Interview of Reported Symptoms (SIRS) manual. Odessa, FL: Psychological Assessment Resources. Rogers, R., Jackson, R., & Cashel, M. (2004). The Schedule for Affective Disorders and Schizophrenia (SADS). In M. J. Hilsenroth & D. L. Segal (Eds.), Comprehensive
handbook of psychological assessment (vol. 2). New York: John Wiley. Rogers, R., Sewell, K., & Goldstein, A. (1994). Explanation models of malingering: A prototypical analysis. Law and Human Behavior, 18, 543–552. Rogoff, B. (1984). What are the interrelations among the three subtheories of Sternberg’s triarchic theory of intelligence? Behavioral and Brain Sciences, 7, 300–
301. Roid, G. (2002, August). New Stanford-Binet Intelligence Scales, Fifth Edition: Author’s Overview. Paper presented at the Annual Convention of the American
Psychological Association, Chicago. Roid, G. (2003). Stanford-Binet Intelligence Scales (5th ed.). Itasca, IL: Riverside Publishing. Roid, G. (2005). Stanford-Binet Intelligence Scales for Early Childhood (5th ed.). Itasca, IL: Riverside Publishing. Roid, G. H., & Johnson, W. B. (1998). Computer assisted psychological assessment. In A. S. Bellack & M. Hersen (Eds.), Comprehensive clinical psychology (vol. 4).
Amsterdam: Elsevier. Roid, G., & Miller, L. (1997). Leiter-R Manual. Wood Dale, IL: Stoelting Co. Roldán-Tapia, L., Parrón, T., & Sánchez-Santed, F. (2005). Neuropsychological effects of long-term exposure to organophosphate pesticides. Neurotoxicology and
Teratology, 27, 259–266. Rorschach, H. (1921). Psychodiagnostik. Berne: Birchen. Rosenberg, S., Ryan, J., & Pri�itera, A. (1984). Rey Auditory-Verbal Learning Test performance of patients with and without memory impairment. Journal of
Clinical Psychology, 40, 785–787. Ross, S. M., Gottfredson, D. K., Christensen, P., & Weaver, R. (1986). Cognitive self-statements in depression: Findings across clinical populations. Cognitive
Therapy and Research, 10, 159–166. Rossier, J., de Stadelhofen, F., & Berthoud, S. (2004). The hierarchical structures of the NEO-PI-R and the 16 PF 5. European Journal of Psychological Assessment,
20, 27–38.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 62/70
Rosvold, H. E., Mirsky, A. E., Sarason, I., & others. (1956). A continuous performance test of brain damage. Journal of Consulting Psychology, 20, 343–350. Rotter, J. B. (1966). Generalized expectancies for internal versus external control of reinforcement. Psychological Monographs, 80 (Whole No. 609). Rotter, J. B. (1972). Beliefs, social attitudes, and behavior: A social learning analysis. In J. B. Rotter, J. Chances, & E. J. Phares (Eds.), Applications of a social learning
theory of personality. New York: Holt, Rinehart and Winston. Rotter, J. B., & Rafferty, J. E. (1950). Manual for the Rotter Incomplete Sentences Blank: College Form. New York: The Psychological Corporation. Rotter, J. B., Lah, M., & Rafferty, J. (1992). Manual—Rotter Incomplete Sentences Blank (2nd ed.). Orlando, FL: The Psychological Corporation. Rotter, J. B., Rafferty, J. E., & Schachtitz, E. (1965). Validation of the Rotter Incomplete Sentences Test. In B. I. Murstein (Ed.), Handbook of projective techniques.
New York: Basic Books. Rozin, P. (2009). What kind of empirical research should we publish, fund, and reward? A different perspective. Perspectives on Psychological Science, 4, 435. Rubenzer, S., Faschingbauer, T., & Ones, D. (2000). Assessing the U.S. presidents using the Revised NEO Personality Inventory. Assessment, 7, 403–420. Rubin, M. (1999). Emotional intelligence and its role in mitigating aggression. Unpublished doctoral dissertation, Immaculata College, Immaculata, Pennsylvania. Rule, W. R., & Traver, M. D. (1983). Test-retest reliabilities of State-Trait Anxiety Inventory in a stressful social analogue situation. Journal of Personality
Assessment, 47, 276–277. Rushton, J. P., & Jensen, A. R. (2005). Thirty years of research on race differences in cognitive ability. Psychology, Public Policy, and Law, 11, 235–294. Russell, M., Martier, S., Sokol, R., & others. (1994). Screening for pregnancy risk-drinking. Alcoholism: Clinical and Experimental Research, 18, 1156–1161. Russo, J. (1994). Thurstone’s scaling model applied to the assessment of self-reported depressive severity. Psychological Assessment, 6, 159–171. Rust, J., & Lindstrom, A. (1996). Concurrent validity of the WISC-III and Stanford-Binet-IV. Psychological Reports, 79, 618–620. Ryan, A. M., & Sackett, P. R. (1987). Pre-employment honesty testing: Fakability, reactions of test takers, and company image. Journal of Business and Psychology,
1, 248–256. Ryan, J. J., Sattler, J. M., & Tree, H. A. (2009, August). Exploratory factor analysis of the WAIS-IV. Paper presented at the Annual Convention of the American
Psychological Association, Toronto, Canada. Ryan, M. (1985). Review of the Minnesota Clerical Test. The ninth mental measurements yearbook (vol. I). Lincoln: University of Nebraska Press. Ryan, R. M. (1987). Thematic Apperception Test. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques compendium. Kansas City, MO: Test Corporation of America. Saccuzzo, D. P., & Johnson, N. E. (1995). Traditional psychometric tests and proportionate representation: An intervention and program evaluation study.
Psychological Assessment, 7, 183–194. Sackett, P. R., Borneman, M. J., & Connelly, B. S. (2008). High stakes testing in higher education and employment: Appraising the evidence for validity and fairness.
American Psychologist, 63, 215–227. Sadock, B., & Sadock, V. (2004). Kaplan and Sadock’s comprehensive textbook of psychiatry (8th ed.). Philadelphia: Lippincott, Williams and Wilkins. Sala, F. (2002). Emotional Competence Inventory: Technical manual. Philadelphia: McClelland Center for Research, HayGroup. Salovey, P., & Mayer, J. (1989–1990). Emotional intelligence. Imagination, Cognition, and Personality, 9, 185–211. Salter, D., Forney, D., & Evans, N. (2005). Two approaches to examining the stability of Myers-Briggs Type Indicator scores. Measurement and Evaluation in
Counseling and Development, 37, 208–219. Salvia, J., & Ysseldyke, J. (2001). Assessment (8th ed). Boston: Houghton Mif�lin. Samelson, F. (1977). World War I intelligence testing and the development of psychology. Journal of the History of the Behavioral Sciences, 13, 274–282. Sandford, J. A., & Turner, A. (1997). Intermediate Visual and Auditory Continuous Performance Test (IVA). Los Angeles: Western Psychological Services. Sarason, I. G. (1961). Test anxiety, experimental instructions, and verbal learning. American Psychologist, 16, 374. Sashidharan, T., Pawlow, L. A., & Pettibone, J. C. (2012). An examination of racial bias in the Beck Depression Inventory-II. Cultural Diversity and Ethnic Minority
Psychology, 18, 203–209. Sattler, J. M. (1988). Assessment of children (3rd ed.). San Diego, CA: Jerome M. Sattler, Publisher. Sattler, J. M. (2001). Assessment of children: Cognitive applications. San Diego, CA: Jerome M. Sattler, Publisher. Sattler, J. M. (2008). Assessment of children: Cognitive foundations (5th ed.). La Mesa, CA: Jerome M. Sattler, Publisher. Saulle, M., & Greenwald, B. D. (2012). Chronic Traumatic Encephalopathy: A review. Rehabilitation Research and Practice, online journal, Article ID 816069, 9
pages. Savickas, M., Taber, B., & Spokane, A. (2002). Convergent and discriminant validity of �ive interest inventories. Journal of Vocational Behavior, 61, 139–184. Scarr, S. (1981). Testing for children: Assessment and the many determinants of intellectual competence. American Psychologist, 36, 1159–1168. Scarr, S. (1987). Foreward. In R. Elliott (Ed.), Litigating intelligence: IQ tests, special education and social science in the courtroom. Dover, MA: Auburn House. Scarr, S. (1994). Culture-Fair and Culture-Free tests. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Scarr, S., & Weinberg, R. A. (1976). IQ test performance of black children adopted by white families. American Psychologist, 31, 726–739. Scarr, S., & Weinberg, R. A. (1983). The Minnesota Adoption Studies: Genetic differences and malleability. Child Development, 54, 260–267. Scarr-Salapatek, S. (1971). Unknowns in the IQ equation. Science, 174, 1223–1228. Schaie, K. W. (1958). Rigidity-�lexibility and intelligence: A cross-sectional study of the adult life span from 20–70. Psychological Monographs, 72, no. 9 (Whole
No. 462). Schaie, K. W. (1977). Quasi-experimental designs in the psychology of aging. In J. E. Birren & K. W. Schaie (Eds.), Handbook of the psychology of aging. New York:
Van Nostrand Reinhold. Schaie, K. W. (1978). Review of Senior Apperception Techniques. The eighth mental measurements yearbook. Lincoln: University of Nebraska Press. Schaie, K. W. (1980). Cognitive development in aging. In L. K. Obler & M. Alpert (Eds.), Language and communication in the elderly. Lexington, MA: Heath. Schaie, K. W. (1985). Manual for the Schaie-Thurstone Adult Mental Abilities Test (STAMAT). Palo Alto, CA: Consulting Psychologists Press. Schaie, K. W. (1996). Intellectual development in adulthood: The Seattle Longitudinal Study. New York: Cambridge University Press. Schaie, K. W. (2005). Developmental in�luences on adult intelligence: The Seattle longitudinal study. New York: Oxford University Press. Schaie, K. W. (2011). Historical in�luences on aging and behavior. In K. W. Schaie & S. L. Willis (Eds.), Handbook of the psychology of aging (7th ed., pp. 41–55). San
Diego, CA: Elsevier. Schaie, K. W., & Willis, S. L. (1986). Adult development and aging. Boston: Little, Brown. Schaie, K. W., Caskie, G., Revell, A., & others. (2005). Extending neuropsychological assessments in the Primary Mental Ability space. Aging, Neuropsychology, and
Cognition, 12, 245–277.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 63/70
Schalock, R. L., Borthwick-Duffy, S. A., Buntinx, W., & others. (2010). Intellectual disability: De�inition, classi�ication, and systems of supports (11th ed.). Washington, DC: American Association on Intellectual and Developmental Disability.
Schalock, R., Luckasson, R., Shogren, K., & others. (2007). The renaming of Mental Retardation: Understanding the change to the term Intellectual Disability. Intellectual and Developmental Disabilities, 45, 116–124.
Schatz, P., Pardini, J., Lovell, M. R., Collins, M. W., & Podell, K. (2006). Sensitivity and speci�icity of the ImPACT test battery for concussion in athletes. Archives of Clinical Neuropsychology, 21, 91–99.
Schear, J. M., & Craft, R. B. (1989). Examination of the concurrent validity of the California Verbal Learning Test. Clinical Neuropsychologist, 3, 162–168. Scheier, M., Carver, C., & Bridges, M. (1994). Distinguishing optimism from neuroticism (and trait anxiety, self-mastery, and self-esteem): A reevaluation of the
Life Orientation Test. Journal of Personality and Social Psychology, 67, 1063–1078. Scheuneman, J. D. (1987). An argument opposing Jensen on test bias: The psychological aspects. In S. Modgil & C. Modgil (Eds.), Arthur Jensen: Consensus and
controversy. New York: Falmer Press. Schmidt, F. (2002). The role of general cognitive ability and job performance: Why there cannot be a debate. Human Performance, 15, 187–211. Schmidt, F. L., Hunter, J. E., McKenzie, R. C., & Muldrow, T. W. (1979). Impact of valid selection procedures on work-force productivity. Journal of Applied
Psychology, 64, 609–626. Schmidt, F., & Zimmerman, R. (2004). A counterintuitive hypothesis about employment interview validity and some supporting evidence. Journal of Applied
Psychology, 89, 553–561. Schmidt, K. S., & Gallo, J. L. (2007). Behavioral and Psychological Assessment of Dementia (BPAD). Lutz, FL: Psychological Assessment Resources. Schmitt, N. (1995). Review of the Differential Aptitude Tests, Fifth Edition. The twelfth mental measurements yearbook. Lincoln: University of Nebraska Press. Schmitt, N. (1996). Uses and abuses of coef�icient alpha. Psychological Assessment, 8, 350–353. Schmitt, N., & Kunce, C. (2002). The effects of required elaboration of answers to biodata questions. Personnel Psychology, 55, 569–587. Schmitt, N., & Robertson, I. (1990). Personnel selection. Annual Review of Psychology, 41, 289–320. Schoenberg, M., Dawson, K., Duff, K., & others. (2006). Test performance and classi�ication statistics for the Rey Auditory Verbal Learning Test in selected clinical
samples. Archives of Clinical Neuropsychology, 21, 693–703. Schroeder, M. L., Schroeder, K. G., & Hare, R. D. (1983). Generalizability of a checklist for assessment of psychopathy. Journal of Consulting and Clinical Psychology,
51, 511–516. Schroffel, A. (2012). The use of in-basket exercises for the recruitment of advanced social service workers. Public Personnel Management, 41, 151–160. Schuldberg, D. (1988). The MMPI is less sensitive to the automated testing format than it is to repeated testing: Item and scale effects. Computers in Human
Behavior, 4, 285–298. Schuler, M. (1999). Brief report: Frequency of maternal cocaine use during pregnancy and infant neurobehavioral outcome. Journal of Pediatric Psychology, 24,
511–514. Schwab, L. O. (1979). The Nebraska assessment for independent living (Project 93–013). Lincoln: Department of Human Development and the Family, University
of Nebraska. Seashore, C. E. (1938). The psychology of musical talent. Boston: Silver, Burdett. Segal, N. (2012). Born together—Reared apart: The landmark Minnesota Twin Study. Cambridge, MA: Harvard University Press. Seligman, M. E. P., & Csikszentmihalyi, M. (2000). Positive psychology: An introduction. American Psychologist, 55, 5–14. Seligman, M. E. P., & Kahana, M. (2009). Unpacking intuition: A conjecture. Perspectives on Psychological Science, 4, 399–402. Seligman, M. E. P., Abramson, L. Y., Semmel, A., & Von Baeyer, C. (1979). Depressive attributional style. Journal of Abnormal Psychology, 88, 242–247. Sellbom, M., Fishler, G., & Ben-Porath, Y. (2007). Identifying MMPI-2 predictors of police of�icer integrity and misconduct. Criminal Justice and Behavior, 34, 985–
1004. Shapiro, E. S. (1996). Academic skills problems workbook. New York: Guilford. Sharkey, K. J., & Ritzler, B. A. (1985). Comparing diagnostic validity of the TAT and a new Picture Projective Test. Journal of Personality Assessment, 49, 406–412. Shaughnessy, M., & Moore, J. (1994). The KAIT with developmental students, honor students, and freshmen. Psychology in the Schools, 31, 286–287. Shaw, S., Cullen, J., McGuire, J., & Brinckerhoff, L. (1995). Operationalizing a de�inition of learning disabilities. Journal of Learning Disabilities, 28, 586–597. Shayer, M., Ginsburg, D., & Coe, R. (2007. Thirty years on—a large anti-Flynn effect? The Piagetian test Volume & Heaviness norms 1975–2003. British Journal of
Educational Psychology, 77, 25–41. Sheldon, W., & Stevens, S. (1942). The varieties of temperament: A psychology of constitutional differences. New York: Harper & Brothers. Shen, H., & Comrey, A. (1997). Predicting medical students’ academic performances by their cognitive abilities and personality characteristics. Academic
Medicine, 72, 781–786. Sheshlow, D., & Adams, W. (2006). Wide Range Assessment of Memory and Learning (2nd Ed.). Lutz, FL: Psychological Assessment Resources. Shiffman, S., & Hufford, M. (2001). Ecological momentary assessment. Applied Clinical Trials, 10, 42–48. Shiffman, S., Hufford, M., & Paty, J. (2001). The patient experience movement. Applied Clinical Trials, 10, 48–56. Shiffman, S., Hufford, M., Hickcox, M., & others. (1997). Remember that? A comparison of real-time versus retrospective recall of smoking lapses. Journal of
Consulting and Clinical Psychology, 65, 292–300. Shurrager, H. C. (1961). A haptic intelligence scale for adult blind. Chicago: Illinois Institute of Technology. Shurrager, H. C., & Shurrager, P. S. (1964). Manual for the Haptic Intelligence Scale for the Blind. Chicago: Psychology Research Technology Center, Illinois Institute
of Technology. Siegman, A. W. (1956). The effect of manifest anxiety on a concept formation task, a nondirected learning task, and on timed and untimed intelligence tests.
Journal of Consulting Psychology, 20, 176–178. Silver, J. M., McAllister, T. W., & Yudofsky, S. C. (Eds.). (2011). Textbook of traumatic brain injury (2nd ed.). Washington, DC: American Psychiatric Association. Silverstein, A. B. (1986). Organization and Structure of the Detroit Tests of Learning Aptitude (DTLA-2). Educational and Psychological Measurement, 46, 1061–
1066. Silvia, P. J., Wigert, B., Reiter-Palmon, R., & Kaufman, J. C. (2012). Assessing creativity with self-report scales: A review and empirical evaluation. Psychology of
Aesthetics, Creativity, and the Arts, 6, 19–34. Simpson, J. A., Rholes, W. S., & Nelligan, J. S. (1992). Support seeking and support giving within couples in an anxiety-provoking situation: The role of attachment
styles. Journal of Personality and Social Psychology, 62, 434–446.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 64/70
Sipps, G. J., Berry, G. W., & Lynch, E. M. (1987). WAIS-R and social intelligence: A test of established assumptions that uses the CPI. Journal of Clinical Psychology, 43, 499–504.
Sisson, E. D. (1948). Forced-choice: The new Army rating. Personnel Psychology, 1, 365–381. Sivan, A. B. (1991). Revised Visual Retention Test: Clinical and experimental applications (5th ed.). San Antonio, TX: The Psychological Corporation. Skinner, B. F. (1953). Science and human behavior. New York: Macmillan. Skinner, B. F. (1974). About behaviorism. New York: Knopf. Smith, A. (1960). Changes in Porteus Maze scores of brain-operated schizophrenics after an eight year interval. Journal of Mental Science, 106, 967–978. Smith, A. (1973). Symbol Digit Modalities Test. Manual. Los Angeles: Western Psychological Services. Smith, A., & Kinder, E. (1959). Changes in psychological test performances of brain-operated subjects after eight years. Science, 129, 149–150. Smith, G. T. (2009). Why do different individuals progress along different life trajectories? Perspectives on Psychological Science, 4, 415–421. Smith, J. (2001). Detroit Tests of Learning Aptitude, Fourth Edition. Fourteenth mental measurements yearbook. Lincoln: University of Nebraska Press. Smith, M., Delves, T., Lansdown, R., Clayton, B., & Graham, P. (1983). The effects of lead exposure on urban children: The Institute of Child Health/Southampton
Study. Developmental Medicine and Child Neurology, 25, 1–54. Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied
Psychology, 47, 149–155. Smith, T. W., Follick, M. J., Ahern, D. K., & Adams, A. (1986). Cognitive distortion and disability in chronic low back pain. Cognitive Therapy and Research, 10, 201–
210. Smyth, J., Wonderlich, S., Crosby, R., & others. (2001). The use of ecological momentary assessment approaches in eating disorder research. International Journal
of Eating Disorders, 30, 83–95. Snow, J. H. (1992). Review of Luria-Nebraska Neuropsychological Battery: Forms I and II. The eleventh mental measurements yearbook. Lincoln: University of
Nebraska Press. Snyder, C. R., & Lopez, S. (2007). Positive psychology: The scienti�ic and practical explorations of human strengths. Thousand Oaks, CA: Sage. Snyder, D. K., Lachar, D., & Wills, R. M. (1988). Computer-based interpretation of the Marital Satisfaction Inventory: Use in treatment planning. Journal of Marital
and Family Therapy, 14, 397–409. Society for Industrial and Organizational Psychology, Inc. (1987). Principles for the validation and use of personnel selection procedures (3rd ed.). College Park,
MD: Author. Society for Research in Child Development. (2010). Social policy report brief: Protecting children from lead exposure. Sharing Youth and Child Development
Knowledge, 24(1). Sokol, R. J., & Clarren, S. K. (1989). Guidelines for use of terminology describing the impact of prenatal alcohol on the offspring. Alcoholism: Clinical and
Experimental Research, 13, 597–598. Sonne, J. L. (2012). Mental status examination. In J. L. Sonne (Ed.), PsycEssentials: A pocket resource for mental health practitioners (pp. 47–56). Washington, DC:
American Psychological Association. Sontag, L. W., Baker, C., & Nelson, V. (1958). Mental growth and personality development: A longitudinal study. Monographs of the Society for Research in Child
Development, 23 (Whole No. 68). Sotile, W. M., Julian, A., Henry, S. E., & Sotile, M. O. (1988). Family Apperception Test manual. Los Angeles: Western Psychological Services. Soto, C. J., John, O. P., Gosling, S. D., & Potter, J. (2011). Age differences in personality traits from 10 to 65: Big Five domains and facets in a large cross-sectional
sample. Journal of Personality and Social Psychology, 100, 330–348. Spearman, C. (1904). “General intelligence,” objectively determined and measured. American Journal of Psychology, 15, 201–293. Spearman, C. (1923). The nature of ‘intelligence’ and the principles of cognition. London: Macmillan. Spearman, C. (1927). The abilities of man. New York: Macmillan. Specht, J., Egloff, B., & Schmuckle, S. C. (2011). Stability and change of personality across the life course: The impact of age and major life events on mean-level
and rank-order stability of the big �ive. Journal of Personality and Social Psychology, 101, 862–882. Special Education Today. (1985). ACALD de�inition of learning disabilities. 2, 1–20. Sperry, R. W. (1964). The great cerebral commissure. Scienti�ic American, 210, 42–52. Spielberger, C. D. (1973). Manual for the State-Trait Anxiety Inventory for children. Palo Alto: Consulting Psychologists Press. Spielberger, C. D. (1983). Manual for the State-Trait Anxiety Inventory (form y). Menlo Park, CA: Mind Garden. Spielberger, C. D. (1989). State-Trait Anxiety Inventory (STAI): A comprehensive bibliography (Revised). Menlo Park, CA: Mind Garden. Spielberger, C. D., & Vagg, P. R. (Eds.). (1995). Test anxiety: Theory, assessment, and treatment. Philadelphia: Taylor & Francis. Spielberger, C. D., Gonzalez, H. P., Taylor, C. J., & others. (1980). Test Anxiety Inventory. Palo Alto, CA: Consulting Psychologists Press. Spielberger, C. D., Gorsuch, R. L., & Lushene, R. E. (1970). The State-Trait Anxiety Inventory: Test manual. Palo Alto, CA: Consulting Psychologist Press. Spitzer, R., & Endicott, J. (1978). Research diagnostic criteria: Rationale and reliability. Archives of General Psychiatry, 35, 773–782. Spohr, H., & Steinhausen, H. (Eds.). (1996). Alcohol, pregnancy, and the developing child. Cambridge: Cambridge University Press. Spreen, O. (2001). Learning disabilities and their neurological foundations, theories, and subtypes. In A. Kaufman & N. Kaufman (Eds.), Speci�ic learning
disabilities and dif�iculties in children and adolescents. Cambridge, England: Cambridge University Press. Spreen, O., & Strauss, E. (1998). A compendium of neuro-psychological tests: Administration, norms, and commentary (2nd ed.). New York: Oxford University Press. Springer, S., & Deutsch, G. (1997). Left brain, right brain (5th ed.). San Francisco: W. H. Freeman. Sreenivasan, S., Walker, S., Weinberger, L., Kirkish, P., & Garrick, T. (2008). Four-facet PCL-R structure and cognitive functioning among high violent criminal
offenders. Journal of Personality Assessment, 90, 197–200. Stafford-Clark, D. (1971). What Freud really said. New York: Schocken Books. Stanley, J. C. (1971). Reliability. In R. L. Thorndike (Ed.), Educational measurement. Washington, DC: American Council on Education. Steele, C. M. (1997). A threat in the air: How stereotypes shape intellectual identity and performance. American Psychologist, 6, 613–629. Steele, C. M., & Aronson, J. (1995). Stereotype threat and the intellectual test performance of African Americans. Journal of Personality and Social Psychology, 69,
797–811. Steer, R. A., Beck, A. T., & Brown, G. (1989). Sex differences on the Revised Beck Depression Inventory for outpatients with affective disorders. Journal of
Personality Assessment, 53, 693–702. Steers, R. M., & Rhodes, S. R. (1978). Major in�luences on employee attendance: A process model. Journal of Applied Psychology, 63, 391–407.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 65/70
Stefan, S. (2001). Unequal rights: Discrimination against people with mental disabilities and the Americans with Disabilities Act. Washington, DC: American Psychological Association.
Stehouwer, R. S. (1987). Beck Depression Inventory. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques compendium. Kansas City, MO: Test Corporation of America.
Steinweg, D. L., & Worth, H. (1993). Alcoholism: The keys to the CAGE. American Journal of Medicine, 94, 520–523. Stenner, A. J. (2001). The Lexile Framework: A common metric for matching readers and text. California School Library Association Journal, 25, 41–42. Stephenson, W. (1953). The study of behavior: Q-technique and its methodology. Chicago: University of Chicago Press. Steptoe, A., Wright, C., Kunz-Ebrecht, S., & Iliffe, S. (2006). Dispositional optimism and health behaviour in community-dwelling older people: Associations with
healthy ageing. British Journal of Health Psychology, 11, 71–84. Stern, R., & White, T. (2003a). Neuropsychological Assessment Battery: Administration, scoring, and interpretive manual. Lutz, FL: Psychological Assessment
Resources. Stern, R., & White, T. (2003a). Neuropsychological Assessment Battery: Psychometric and technical manual. Lutz, FL: Psychological Assessment Resources. Stern, W. L. (1912). Uber die psychologischen Methoden der Intelligenzprufung. American translation by G. M. Whipple (1914). The psychological methods of
testing intelligence. Educational Psychology Monographs, no. 13, Baltimore: Warwick & York. Sternberg, R. J. (1981). Intelligence and nonentrenchment. Journal of Educational Psychology, 73, 1–16. Sternberg, R. J. (1985a). Componential analysis: A recipe. In D. K. Detterman (Ed.), Current topics in human intelligence (vol. 1). Norwood, NJ: Ablex. Sternberg, R. J. (1985b). Beyond IQ: A triarchic theory of human intelligence. Cambridge: Cambridge University Press. Sternberg, R. J. (1986). Intelligence applied: Understanding and increasing your intellectual skills. San Diego, CA: Harcourt Brace Jovanovich. Sternberg, R. J. (1993). Sternberg Triarchic Abilities Test (Level H). Unpublished test. Sternberg, R. J. (1994). The triarchic theory of intelligence. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Sternberg, R. J. (1996). Successful intelligence. New York: Simon & Schuster. Sternberg, R. J. (2002). Creativity as a decision. American Psychologist, 57, 376. Sternberg, R. J. (Ed.). (1994). Encyclopedia of human intelligence (vols. 1, 2). New York: Macmillan. Sternberg, R. J., & Detterman, D. K. (Eds.). (1986). What is intelligence? Contemporary viewpoints on its nature and de�inition. Norwood, NJ: Ablex. Sternberg, R. J., & Kaufman, J. C. (1998). Human abilities. Annual Review of Psychology, 49, 479–502. Sternberg, R. J., & Williams, W. (1997). Does the Graduate Record Examination predict meaningful success in the graduate training of psychologists? A case study.
American Psychologist, 52, 630–641. Sternberg, R. J., & Zhang, L. (1995). What do we mean by giftedness? A pentagonal implicit theory. Gifted Child Quarterly, 39, 88–94. Sternberg, R. J., Castejon, J., Prieto, M., Hautamaki, J., & Grigorenko, E. (2001). Con�irmatory factor analysis of the Sternberg Triarchic Abilities Test in three
international samples. European Journal of Psychological Assessment, 17, 1–16. Sternberg, R. J., Conway, B. E., Ketron, J. L., & Bernstein, M. (1981). People’s conceptions of intelligence. Journal of Personality and Social Psychology, 41, 37–55. Sternberg, R., & Lubart, T. (1992). Buy low and sell high: An investment approach to creativity. Current Directions in Psychological Research, 1, 1–5. Stevens, S. S. (1946). On the theory of scales and measurement. Science, 103, 677–680. Stewart, G., Dustin, S., Barrick, M., & Darnold, T. (2008). Exploring the handshake in employment interviews. Journal of Applied Psychology, 93, 1139–1146. Stewart, P., Reihman, J., Lonky, E., Darvill, T., & Pagano, J. (1999). Prenatal PCB exposure and neonatal behavioral assessment scale (NBAS) performance.
Neurotoxicology and Teratology, 22, 21–29. Stockwell, S., Schaeffer, B., & Lowenstein, J. (1991). The SAT coaching coverup. Cambridge, MA: Fairtest. Stokes, G., & Cooper, L. (2001). Content/construct approaches in life history form development for selection. International Journal of Selection and Assessment, 9,
138–151. Stokes, G., & Cooper, L. (2004). Biodata. In J. Thomas (Ed.), Comprehensive handbook of psychological assessment, Vol. 4: Industrial and organizational assessment
(pp. 243–268). Hoboken, NJ: John Wiley. Stokes, G., Mumford, M., & Owens (Eds.). (1994). Biodata handbook: Theory, research, and use of biographical information in selection and performance prediction.
Palo Alto, CA: Consulting Psychologists Press. Stone, B. J. (1994). Group ability test versus teachers’ ratings for predicting achievement. Psychological Reports, 75, 1487–1490. Storandt, M., & Hill, R. D. (1989). Very mild senile dementia of the Alzheimer type: 2. Psychometric test performance. Archives of Neurology, 46, 383–386. Stout, J. C., Ready, R. E., Grace, J., Malloy, P. F., & Paulsen, J. S. (2006). Factor Analysis of Frontal Systems Behavior Scale (frSBe). Assessment, 10, 79–85. Strauss, E., Sherman, E. M. S., & Spreen, O. (2006). A compendium of neuropsychological tests: Administration, norms, and commentary (3rd ed.). New York: Oxford
University Press. Strauss, E., Sherman, E., & Spreen, O. (2006). A compendium of neuropsychological tests: Administration, norms, and commentary (3rd ed.). New York: Oxford
University Press. Strayhorn, J. C., & Strayhorn, J. M. (2012). Lead exposure and the 2010 achievement test scores of children in New York counties. Child and Adolescent Psychiatry
and Mental Health, 6, 4. Streiner, D. L., Goldberg, J. O., & Miller, H. R. (1993). MCMI-II item weights: Their lack of effectiveness. Journal of Personality Assessment, 60, 471–476. Streissguth, A., Bookstein, F., & Barr, H. (1996). A dose-response study of the enduring effects of prenatal alcohol exposure: birth to 14 years. In H. Spohr & H.
Steinhausen (Eds.), Alcohol, pregnancy, and the developing child. Cambridge: Cambridge University Press. Streissguth, A., Bookstein, F., Barr, H., & others. (2004). Risk factors for adverse life outcomes in fetal alcohol syndrome and fetal alcohol effects. Developmental
and Behavioral Pediatrics, 25, 226–238. Streissguth, A., Martin, D., Barr, H., & Sandman, B. (1984). Intrauterine alcohol and nicotine exposure: Attention and reaction time in 4-year-old children.
Developmental Psychology, 20, 533–541. Strong, E. K. (1927). Vocational Interest Blank. Stanford, CA: Stanford University Press. Strong, E. K. (1955). Vocational interests 18 years after college. Minneapolis: University of Minnesota Press. Strong, E. K., Hansen, J., & Campbell, D. (1994). Strong Interest Inventory. Palo Alto, CA: Consulting Psychologists Press. Stroop, J. R. (1935). Studies of interference in serial verbal reaction. Journal of Experimental Psychology, 18, 643–662. Strub, R. L., & Black, F. W. (2000). The mental status examination in neurology (5th ed.). Philadelphia: F. A. Davis. Strutt, A. M., Scott, B. M., Lozano, V. J., Tieu, P. G., & Peery, S. (2012). Assessing sub-optimal performance with the Test of Memory Malingering in Spanish speaking
patients with TBI. Brain Injury, 26, 853–863.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 66/70
Sumi, K. (2006). Correlations between optimism and social relationships. Psychological Reports, 99, 938–940. Sundet, J., Barlaug, D., & Torjussen, T. (2004). The end of the Flynn effect? A study of secular trends in mean intelligence test scores of Norwegian conscripts
during half a century. Intelligence, 32, 349–362. Sundet, J., Borren, I., & Tambs, K. (2008). The Flynn effect is partly caused by changing fertility patterns. Intelligence, 36, 183–191. Super, D. E. (1953). A theory of vocational development. American Psychologist, 8(5), 185–190. Super, D. E. (1990). Career choice and development: Applying contemporary theories to practice. San Francisco: Jossey-Bass. Super, D. E. (1994). A life-span, life-space perspective on convergence. In M. L. Savika & R. W. Lent (Eds.), Convergence in career development theories: Implications
for science and practice (pp. 63–74). Palo Alto, CA: Consulting Psychologists Press. Super, D. E., Savickas, M. L., & Super, C. M. (1996). The life-span, life-space approach to careers. In D. Brown, L. Brooks, & Associates (Eds.), Career choice and
development (3rd ed., pp. 121–177). San Francisco: Jossey-Bass. Sweeney, J., Slade, H., Ivins, R., & others. (2007). Scienti�ic investigation of brain-behavior relationships using the Halstead-Reitan Battery. Applied
Neuropsychology, 14, 65–72. Swenson, W. M., Rome, H., Pearson, J., & Brannick, T. (1965). A totally automated psychological test: Experience in a medical center. Journal of the American
Medical Association, 191, 925–927. Tabachnick, B. G., & Fidell, L. S. (1989). Using multivariate statistics (2nd ed.). New York: Harper & Row. Tallent, N. (1993). Psychological report writing (4th ed.). Englewood Cliffs, NJ: Prentice Hall. Tamkin, A. S., & Scherer, I. W. (1957). What is measured by the “Cannot Say” scale of the group MMPI? Journal of Consulting Psychology, 21, 413–417. Tan, J. E., Hultsch, D. F., Hunter, M. A., & Strauss, E. (2010). Psychometric investigation of the modi�ied Scales of Independent Behavior-Revised in an elderly
sample. Clinical Gerontologist: The Journal of Aging and Mental Health, 33, 69–83. Tanner, B. A. (1992). Computer-aided reporting of the results of neuropsychological evaluations of traumatic brain injury. Computers in Human Behavior, 9, 51–
56. Tasbihsazan, R., Nettelbeck, T., & Kirby, N. (2003). Predictive validity of the Fagan Test of Infant Intelligence. British Journal of Developmental Psychology, 21, 585–
597. Tasto, D. L., Hickson, R., & Rubin, S. E. (1971). Scaled pro�ile analysis of fear survey schedule factors. Behavior Therapy, 2, 543–549. Tate, R. L. (2010). A compendium of tests, scales, and questionnaires: The practitioner’s guide to measuring outcomes after acquired brain impairment. Hove,
UK: Psychology Press. Tauszcik, Y. R., & Pennebaker, J. W. (2010). The psychological meaning of words: LIWC and computerized text analysis methods. Journal of Language and Social
Psychology, 29, 24–54. Taylor, F. S. (1942). The origin of the thermometer. Annals of Science, 5, 129–156. te Nijenhuis, J., Cho, S. H., Murphy, R., & Lee, K. H. (2012). The Flynn effect in Korea: Large gains. Personality and Individual Differences, 53, 147–151. Teacher, Administrator, and Counselor Manual: Iowa Tests of Educational Development. Forms X-8 and Y-8. 1988. Teare, J. F., & Thompson, R. W. (1982). Concurrent validity of the Perkins-Binet tests of intelligence for the blind. Journal of Visual Impairment and Blindness, 76,
279–280. Teasdale, G., & Jennett, B. (1974). The Glasgow Coma Scale. Lancet, 2, 81. Teasdale, T., & Owen, D. (2005). A long-term rise and recent decline in intelligence test performance: The Flynn effect in reverse. Personality and Individual
Differences, 39, 837–843. Teichner, G., Golden, C., Bradley, J., & Crum, T. (1999). Internal consistency and discriminant validity of the Luria Nebraska Neuropsychological Battery-III.
International Journal of Neuroscience, 98, 141–152. Tellegen, A., & Ben-Porath, Y. (1992). The new uniform T scores for the MMPI-2: Rationale, derivation, and appraisal. Psychological Assessment, 4, 145–155. Tellegen, A., & Ben-Porath, Y. S. (2008). MMPI-2-RF (Minnesota Multiphasic Personality Inventory-2 Restructured Form): Technical manual. Minneapolis: University
of Minnesota Press. Temple, R., & Zgaljardic, D. (2009). Ecological validity of the Neuropsychological Assessment Battery Screening Module in post-acute brain injury rehabilitation.
Brain Injury, 23, 45–50. Templeton, A. R. (2002). The genetic and evolutionary signi�icance of human races. In J. Fish (Ed.), Race and intelligence: Separating science from myth. Mahwah,
NJ: Erlbaum. Tendler, A. D. (1930). A preliminary report on a test for emotional insight. Journal of Applied Psychology, 14, 123–126. Teng, S. (1942–43). Chinese in�luence on the western examination system. Harvard Journal of Asiatic Studies, 7, 267–312. Terman, L. M. (1916). The measurement of intelligence. Boston: Houghton Mif�lin. Terman, L. M., & Oden, M. H. (1959). Genetic studies of genius: The gifted group at mid-life. Stanford, CA: Stanford University Press. Terrell, F., Terrell, S., & Taylor, J. (1981). Effect of race of examiner and cultural mistrust on the WAIS performance of Black students. Journal of Consulting and
Clinical Psychology, 49, 750–751. Thoma, S. (2006). Research on the de�ining issues test. In M. Killen & J. Smetana (Eds.), Handbook of moral development (pp. 67–91). Mahwah, NJ: Erlbaum. Thomas, M., & Watkins, P. (2003, May). Measuring the grateful trait: Development of revised GRAT. Paper presented at the Annual Convention of the Western
Psychological Association, Vancouver, BC. Thompson, C. (1949). The Thompson modi�ication of the Thematic Apperception Test. Journal of Projective Techniques, 13, 469–478. Thorndike, E. L. (1912). The permanence of interests and their relation to abilities. Popular Science Monthly, 81, 449–456. Thorndike, E. L. (1918). The seventeenth yearbook of the National Society for the Study of Education. Pt. II. Bloomington, IL: Public School Publishing Co. Thorndike, E. L. (1920). Intelligence and its uses. Harper’s Magazine, 140, 227–235. Thorndike, E. L. (1920a). A constant error in psychological ratings. Journal of Applied Psychology, 4, 25–29. Thorndike, E. L. (1920b). Intelligence and its use. Harper’s Magazine, 140, 227–235. Thorndike, E. L. (Ed.). (1921). Intelligence and Its Measurement: A Symposium. Journal of Educational Psychology, 12, 123–147, 195–216. Thorndike, R. L., & Stein, S. (1937). An evaluation of the attempts to measure social intelligence. Psychological Bulletin, 34, 275–285. Thorndike, R. L., Hagen, E. P., & Sattler, J. M. (1986). The Stanford-Binet Intelligence Scale: Fourth Edition, Guide for administering and scoring. Chicago: Riverside. Thurstone, L. L. (1921). Intelligence. In E. L. Thorndike (Ed.), Intelligence and Its Measurement: A Symposium. Journal of Educational Psychology, 12, 123–147,
195–216. Thurstone, L. L. (1925). A method of scaling psychological and educational tests. Journal of Educational Psychology, 16, 433–451.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 67/70
Thurstone, L. L. (1929). Theory of attitude measurement. Psychological Review, 36, 222–241. Thurstone, L. L. (1931). Multiple factor analysis. Psychological Review, 38, 406–427. Thurstone, L. L. (1938). Primary mental abilities. Psychometric Monographs, no. 1. Chicago: University of Chicago Press. Thurstone, L. L. (1947). Multiple factor analysis. Chicago: University of Chicago Press. Thurstone, L. L., & Thurstone, T. (1930). A neurotic inventory. Journal of Social Psychology, 1, 3–30. Thurstone, L. L., & Thurstone, T. (1941). Factorial studies in intelligence. Psychometric Monographs, No. 2. Chicago: University of Chicago Press. Tif�in, J. (1968). Purdue Pegboard Examiner’s Manual. Chicago: Science Research Associates. Tinius, T. (2003). The Intermediate Visual and Auditory Continuous Performance Test as a neuropsychological measure. Archives of Clinical Neuropsychology, 18,
199–214. Tombaugh, T. (1997). The test of memory malingering (TOMM): Normative data from cognitively intact and cognitively impaired individuals. Psychological
Assessment, 9, 260–268. Tombaugh, T., McDowell, I., Kristjansson, B., & Hubley, A. (1996). Mini-Mental State Examination (MMSE) and the Modi�ied MMSE (3MS): A psychometric
comparison and normative data. Psychological Assessment, 8, 48–59. Tomkins, S. S. (1947). The Thematic Apperception Test. New York: Grune & Stratton. Tong, E., Bishop, G., Enkelmann, H., & others. (2005). The use of ecological momentary assessment to test appraisal theories of emotion. Emotion, 5, 508–512. Torgerson, J. (2009). The response to intervention instructional model: Some outcomes from a large-scale implementation in Reading First schools. Child
Development Perspectives, 3, 38–40. Torrance, E. P. (1966). The Torrance Tests of Creative Thinking: Norms—Technical Manual (Research Edition). Princeton, NJ: Personnel Press. Torrance, E. P. (1974). The Torrance Tests of Creative Thinking Norms—Technical Manual Research Edition—Verbal Tests, Forms A & B. Princeton, NJ: Personnel
Press. Torrance, E. P. (1998). The Torrance Tests of Creative Thinking: Norms—Technical Manual Figural (Streamlined) Forms A & B. Bensenville, IL: Scholastic Testing
Service. Totsika, V., & Sylva, K. (2004). The Home Observation for Measurement of the Environment revisited. Child and Adolescent Mental Health, 9, 25–35. Traxler, A. E. (1951). Administering and scoring the objective test. In E. F. Lindquist (Ed.), Educational measurement. Washington, DC: American Council on
Education. Treffert, D. A. (1989). Extraordinary people. London: Bantam Press. Tref�linger, D. (1985). Review of the Torrance Tests of Creative Thinking. In J. V. Mitchell, Jr., (Ed.), The ninth mental measurements yearbook (pp. 1632–1634). Trevisan, M. S. (1992). Review of GED. The eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Trinidad, D., & Johnson, C. (2002). The association between emotional intelligence and early adolescent tobacco and alcohol use. Personality and Individual
Differences, 32, 95–105. Tröster, A. (2012). Understanding Parkinson’s: Cognition and Parkinson’s. New York: Parkinson’s Disease Foundation. Trull, T. J., Useda, J., Costa, Jr., P., & McCrae, R. (1995). Comparison of the MMPI-2 Personality Psychopathology Five (PSY-5), the NEO-PI, and the NEO-PI-R.
Psychological Assessment, 7, 508–516. Trull, T. J., Widiger, T., Useda, J., & others. (1998). A structured interview for the assessment of the �ive-factor model of personality. Psychological Assessment, 10,
229–240. Tsai, L., & Tsuang, M. (1979). The Mini-Mental State Test and computerized tomography. American Journal of Psychiatry, 136, 436–439. Turk, A. A., Brown, W. S., Symington, M., & Paul, L. K. (2010). Social narratives in agenesis of the corpus callosum: Linguistic analysis of the Thematic
Apperception Test. Neuropsychologia, 48, 43–50. Turkheimer, E., Haley, A., Waldron, M., D’Onofrio, B., & Gottesman, I. I. (2003). Socioeconomic status modi�ies heritability of IQ in young children. Psychological
Science, 4, 623–628. Tzeng, O. C. S. (1987). Strong-Campbell Interest Inventory. In D. J. Keyser & R. C. Sweetland (Eds.), Test critiques compendium. Kansas City, MO: Test Corporation of
America. Tzeng, O., Ware, R., & Chen, J. (1989). Measurement and utility of continuous unipolar ratings for the Myers-Briggs Type Indicator. Journal of Personality
Assessment, 53, 727–738. U.S. Department of Education. (1977). De�inition and criteria for de�ining students as learning disabled. Federal Register, 42(250), 65083. U.S. Department of Education. (1992). Fourteenth Annual Report to Congress on the Implementation of the Individuals with Disabilities Education Act. Washington,
DC: Author. Uematsu, S., Lesser, R., Fisher, R. S., & others. (1992). Motor and sensory cortex in humans: Topography studied with chronic subdural stimulation. Neurosurgery,
31(1), 59–71. Ulrich, L., & Trumbo, D. (1965). The selection interview since 1949. Psychological Bulletin, 63, 100–116. United States Employment Service. (1970). Manual for the USES General Aptitude Test Battery. Washington, DC: United States Department of Labor. Urquhart Hagie, M., Gallipo, P., & Svien, L. (2003). Traditional culture versus traditional assessment for American Indian Students: An investigation of potential
test item bias. Assessment for Effective Intervention, 29, 15–25. Vaillant, G. (1971). Theoretical hierarchy of adaptive ego mechanisms. Archives of General Psychiatry, 24, 107–118. Vaillant, G. (1977). Adaptation to life: How the best and the brightest came of age. Boston: Little, Brown. Vaillant, G. (1992). Ego mechanisms of defense: A guide for clinicians and researchers. Washington, DC: American Psychiatric Press. Vaillant, G., & Vaillant, C. (1990). Natural history of male psychosocial health, XII: A 45-year study of predictors of successful aging at age 65. American Journal of
Psychiatry, 147, 31–37. Van de Vijver, F., & Harsveld, M. (1994). The incomplete equivalence of the paper-and-pencil and computerized versions of the General Aptitude Test Battery.
Journal of Applied Psychology, 79, 852–859. Van Gorp, W. (1992). Review of Luria-Nebraska Neuropsychological Battery: Forms I and II. The eleventh mental measurements yearbook. Lincoln: University of
Nebraska Press. Van Iddekinge, C. H., Roth, P. L., Raymark, P. H., & Odle-Dusseau, H. N. (2012). The criterion-related validity of integrity tests: An updated meta-analysis. Journal of
Applied Psychology, 97, 499–530. Vance, B., Kitson, D., & Singer, M. (1985). Relationship between the standard scores of PPVT-R and Wide Range Achievement Test. Journal of Clinical Psychology,
41, 691–693.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 68/70
VanderVeer, B., & Schweid, E. (1974). Infant assessment: Stability of mental functioning in young retarded children. American Journal of Mental De�iciency, 79, 1– 4.
Varma, A., DeNisi, A., & Peters, L. (1996). Interpersonal affect and performance appraisal: A �ield study. Personnel Psychology, 49, 341–360. Vaughn, S., & Haager, D. (1994). The measurement and assessment of social skills. In G. R. Lyon (Ed.), Frames of reference for the assessment of learning
disabilities: New views on measurement issues. Baltimore: Brookes Publishing. Vautier, S., & Pohl, S. (2009). Do balanced scales assess bipolar construct? The case of the STAI scales. Psychological Assessment, 21, 187–193. Vernon, M. C., & Alles, B. F. (1986). Psychoeducational assessment of deaf and hard-of-hearing children and adolescents. In P. J. Lazarus & S. S. Strichart (Eds.),
Psychoeducational evaluation of children and adolescents with low-incidence handicaps. New York: Grune & Stratton. Vernon, M. C., & Brown, D. W. (1964). A guide to psychological tests and testing procedures in the evaluation of deaf and hard-of-hearing children. Journal of
Speech and Hearing Disorders, 29, 414–423. Vernon, P. A. (2000). Recent studies of intelligence and personality using Jackson’s Multidimensional Aptitude Battery and Personality Research Form. In R.
Gof�in & E. Helmes (Eds.), Problems and solutions in human assessment: Honoring Douglas N. Jackson at seventy. New York: Kluwer Academic/Plenum Publishers.
Vernon, P. A., Martin, R., Schermer, J., & Mackie, A. (2008). A behavioral genetic investigation of humor styles and their correlations with the Big-5 personality dimensions. Personality and Individual Differences, 44, 1116–1125.
Vernon, P. E. (1950). The structure of human abilities. London: Methuen. Vernon, P. E. (1979). Intelligence: Heredity and environment. San Francisco: Freeman. Viglione, D. J., Blume-Marcovici, A. C., Miller, H. L., Giromini, L., & Meyer, G. (2012). An inter-rater reliability study for the Rorschach Performance Assessment
System. Journal of Personality Assessment, 94, 607–612. Vince, J. (2004). Introduction to virtual reality. New York: Springer Publishing. Vincent, A., Roebuck-Spencer, T., Gilleland, K., & Schlegel, R. (2012). Automated Neuropsychological Assessment Metrics (v4) Traumatic Brain Injury Battery:
Military normative data. Military Medicine, 177, 256–269. Viswesvaran, C., Ones, D., & Schmidt, F. (1996). Comparative analysis of the reliability of job performance ratings. Journal of Applied Psychology, 81, 557–574. Wagner, R. (1949). The employment interview: A critical review. Personnel Psychology, 2, 17–46. Wainer, H. (Ed.). (2000). Computerized adaptive testing: A primer (2nd ed.). Mahwah, NJ: Erlbaum. Walker, C. (2006). Cognitive improvement and alcoholism recovery [fact sheet]. Center City, MN: Hazelden Publishing. Wallas, G. (1926). The art of thought. New York: Harcourt, Brace. Wallbrown, F. H., Carmin, C. N., & Barnett, R. W. (1988). Investigating the construct validity of the Multidimensional Aptitude Battery. Psychological Reports, 62,
871–878. Walls, R. T., Zane, T., & Thvedt, J. E. (1979). The Independent Living Behavior Checklist. Dunbar: West Virginia Research and Training Center. Walsh, B. D. (1996, March). The psychometric characteristics of the Career Beliefs Inventory. Dissertation Abstracts International, Section A: Humanities and Social
Sciences, 56(9-A), 3516. Walsh, B. D., Thompson, T., & Kapes, J. (1997). The construct validity of scores on the Career Beliefs Inventory. Journal of Career Assessment, 5, 31–46. Walsh, W. B., & Holland, J. L. (1992). A theory of personality types and work environments. In W. Walsh, R. Price, & K. Craik (Eds.), Person-environment
psychology: Models and perspectives. Hillsdale, NJ: Erlbaum. Wanek, J. (1999). Integrity and honesty testing: What do we know? How do we use it? International Journal of Selection and Assessment, 7, 183–195. Wang, J., & Kaufman, A. (1993). Changes in �luid and crystallized intelligence across the 20- to 90-year age range on the K-BIT. Journal of Psychoeducational
Assessment, 11, 29–37. Wang, L. (1995). Differential Aptitude Tests. Measurement and Evaluation in Counseling and Development, 28, 168–171. Wang, L., Beckett, G. H., & Brown, L. (2006). Controversies of standardized assessment in school accountability reform: A critical synthesis of multidisciplinary
research evidence. Applied Measurement in Education, 19, 306–328. Wang, M. C., Haertel, G. D., & Walberg, H. J. (1990). What in�luences learning? A content analysis of review literature. Journal of Educational Research, 84, 30–43. Washington, J., & Craig, H. (1999). Performance of at-risk, African American preschoolers on the Peabody Picture Vocabulary Test-III. Language, Speech, &
Hearing Services in Schools, 30, 75–82. Wasylkiw, L., & Fekken, G. (2002). Personality and self-reported health: Matching predictors and criteria. Personality and Individual Differences, 33, 607–620. Watkins, C., Campbell, V., Nieberding, R., & Hallmark, R. (1995). Contemporary practice of psychological assessment by clinical psychologists. Professional
Psychology: Research and Practice, 26, 54–60. Watkins, P., Woodward, K., Stone, T., & Kolts, R. (2003). Gratitude and happiness: Development of a measure of gratitude and relationships with subjective well-
being. Social Behavior and Personality, 31, 431–452. Watson, B. (1983). Test-retest stability of the Hiskey-Nebraska Test of Learning Aptitude in a sample of hearing-impaired children and adolescents. Journal of
Speech and Hearing Disorders, 48, 145–149. Watson, B. U., & Goldgar, D. E. (1985). A note on the use of the Hiskey-Nebraska Test of Learning Aptitude with deaf children. Language, Speech, and Hearing
Services in the Schools, 16, 53–57. Watson, C. G., Thomas, D., & Anderson, P. (1992). Do computer-administered Minnesota Multiphasic Personality Inventories underestimate booklet-based
scores? Journal of Clinical Psychology, 48, 744–748. Wechsler, D. (1932). Analytic use of the Army Alpha examination. Journal of Applied Psychology, 16, 254–256. Wechsler, D. (1939). The measurement of adult intelligence. Baltimore: Williams & Wilkins. Wechsler, D. (1941). The measurement of adult intelligence (2nd ed.). Baltimore: Williams & Wilkins. Wechsler, D. (1944). Measurement of adult intelligence (3rd ed.). Baltimore: Williams & Wilkins. Wechsler, D. (1949). Manual for the Wechsler Intelligence Scale for Children. New York: The Psychological Corporation. Wechsler, D. (1952). The range of human capacities (2nd ed.). Baltimore: Williams & Wilkins. Wechsler, D. (1955). Manual for the Wechsler Adult Intelligence Scale. New York: The Psychological Corporation. Wechsler, D. (1974). Manual for the Wechsler Intelligence Scale for Children-Revised. San Antonio, TX: The Psychological Corporation. Wechsler, D. (1981). Manual for the Wechsler Adult Intelligence Scale-Revised. San Antonio, TX: The Psychological Corporation. Wechsler, D. (1989). Manual for the Wechsler Preschool and Primary Scale of Intelligence-Revised. San Antonio, TX: The Psychological Corporation. Wechsler, D. (1991). Manual for the Wechsler Intelligence Scale for Children-III. San Antonio, TX: The Psychological Corporation.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 69/70
Wechsler, D. (1997). Manual for the Wechsler Adult Intelligence Scale-III. San Antonio, TX: The Psychological Corporation. Wechsler, D. (2003). WISC-IV: Technical and interpretive manual. San Antonio, TX: Psychological Corporation. Wechsler, D. (2008). Manual for the Wechsler Adult Intelligence Scale—Fourth Edition. San Antonio, TX: Pearson. Wechsler, D., Coalson, D., & Raiford, S. (2008). WAIS-IV technical and interpretive manual. San Antonio, TX: Pearson. Weekes, N. Y. (1994). Sex differences in the brain. In D. W. Zaidel (Ed.), Neuropsychology (2nd ed.). San Diego, CA: Academic Press. Weiner, I. B. (1994). The Rorschach Inkblot Method (RIM) is not a test: Implications for theory and practice. Journal of Personality Assessment, 62, 498–504. Weiner, I. B. (1996). Some observations on the validity of the Rorschach inkblot method. Psychological Assessment, 8, 206–213. Weiner, I. B., & Kuehnle, K. (1998). Projective assessment of children and adolescents. In A. S. Bellack, & M. Hersen (Eds.), Comprehensive clinical psychology, (vol.
4). Amsterdam: Elsevier. Weiss, D. J. (1985). Adaptive testing by computer. Journal of Consulting and Clinical Psychology, 53, 774–789. Weiss, D. J. (Ed.). (1983). New horizons in testing: Latent trait theory and computerized adaptive testing. New York: Academic Press. Weiss, D. J., & Vale, C. D. (1987). Computerized adaptive testing for measuring abilities and other psychological variables. In J. N. Butcher (Ed.), Computerized
psychological assessment: A practitioner’s guide. New York: Basic Books. Weiss, D. S., Zilberg, N. J., & Genevro, J. L. (1989). Psychometric properties of Loevinger’s Sentence Completion Test in an adult psychiatric outpatient sample.
Journal of Personality Assessment, 53, 478–486. Weiss, R. A., Rosenfeld, B., & Farkas, M. R. (2011). The utility of the Structured Interview of Reported Symptoms in a sample of individuals with intellectual
disabilities. Assessment, 18, 284–290. Weller, C. E., & Fields, J. (2011). The Black and White labor gap in America: Why African Americans struggle to �ind jobs and remain employed compared to Whites.
Washington, DC: Center for American Progress. Wertheimer, M. (1945). Productive thinking. New York: Harper & Row. Wesman, A. G. (1971). Writing the test item. In R. L. Thorndike (Ed.), Educational measurement (2nd ed.). Washington, DC: American Council on Education. Westbrook, B. W., & Bane, K. D. (1992). Review of De�ining Issues Test. Eleventh mental measurements yearbook. Lincoln: University of Nebraska Press. Whipple, G. M. (1910). Manual of mental and physical tests. Baltimore: Warwick and York. Whitney, D. R., Malizio, A. G., & Patience, W. M. (1985). The reliability and validity of the GED Tests. American Council on Education GED Research Brief, May, No. 6. Whyte, J., Polansky, M., Cavallucci, C., Fleming, M., Lhulier, J., & Coslett, H. (1996). Innattentive behaviour arfter traumatic brain injury. Journal of the International
Neuropsychological Society, 2, 274–281. Wiesner, W. H., & Cronshaw, S. F. (1988). A meta-analytic investigation of the impact of interview format and degree of structure on the validity of the
employment interview. Journal of Occupational Psychology, 61, 275–290. Wiggins, J. (1997). In defense of traits. In R. Hogan, J. Johnson, & S. Briggs (Eds.), Handbook of personality psychology. San Diego, CA: Academic Press. Wilkinson, G. S. (1993). Wide Range Achievement Test-III: Administration manual. Wilmington, DE: Wide Range. Wilkinson, G., & Robertson, G. (2006). Wide Range Achievement Test—Fourth Edition. Lutz, FL: Psychological Assessment Resources. Williams, M. (1979). Brain damage, behaviour, and the mind. New York: Wiley. Williams, R. L. (1970). Danger: Testing and dehumanizing Black children. Clinical Child Psychology Newsletter, 9, 5–6. Williamson, L., Campion, J., Malo, S., & others. (1997). Employment interview on trial: Linking interview structure with litigation outcomes. Journal of Applied
Psychology, 82, 900–912. Willingham, W. W., Ragosta, M., Bennett, R., & others. (1988). Testing handicapped people. Boston: Allyn and Bacon. Wilson, B. A., Cockburn, J., & Baddeley, A. (1991). The Rivermead Behavioral Memory Test (2nd ed.). Suffolk, UK: Thames Valley Test Company. Wilson, B., Alderman, N., Burgess, P., Emslie, H., & Evans, J. (1996). Behavioral Assessment of the Dysexecutive Syndrome. Bury St. Edmunds, England: Thames
Valley Test Company. Wilson, M. N. (1994). African Americans. In R. J. Sternberg (Ed.), Encyclopedia of human intelligence. New York: Macmillan. Wilson, M., & Reschly, D. (1996). Assessment in school psychology training and practice. School Psychology Review, 25, 9–23. Wilson, R. S. (1983). The Louisville Twin Study: Developmental synchronies in behavior. Child Development, 54, 298–316. Wilson, T. D. (2009). Know thyself. Perspectives on Psychological Science, 4, 384–389. Wing, H. (1992). Review of the Bennett Mechanical Comprehension Test. The eleventh mental measurements Yearbook. Lincoln: University of Nebraska Press. Winter, D. G., & Stewart, A. J. (1977). Power motive reliability as a function of retest instructions. Journal of Consulting and Clinical Psychology, 45, 436–440. Wirt, R. D., & Broen, W. E., Jr. (1958). Booklet for the Personality Inventory for Children. Minneapolis, MN: Authors. Wirt, R. D., Lachar, D., Klinedinst, J. K., & Seat, P. D. (1984). Multidimensional description of child personality: A manual for the Personality Inventory for Children,
Revised 1984. Los Angeles: Western Psychological Services. Wisniewski, J. J., & Naglieri, J. A. (1989). Validity of the Draw A Person: A Quantitative Scoring System with the WISC-R. Journal of Psychoeducational Assessment,
7, 346–351. Wissler, C. (1901). The correlation of mental and physical tests. The Psychological Review, Monograph Supplement 3(6). Witchalls, C. (2012, September 27). James R. Flynn: Are we really getting smarter every year? The Independent. Witelson, S. (2007). Sex and the single hemisphere: Specialization of the right hemisphere for spatial processing. In G. Einstein (Ed.), Sex and the brain (pp. 541–
544). Cambridge, MA: MIT Press. Wolf, A. W., Schubert, D., Patterson, M., Grande, T., & Pendleton, L. (1990). The use of the MacAndrew Alcoholism Scale in detecting substance abuse and
antisocial personality. Journal of Personality Assessment, 54, 747–755. Wolf, T. H. (1973). Alfred Binet. Chicago: The University of Illinois Press. Wolff, K. C., & Gregory, R. J. (1992). The effects of a temporary dysphoric mood upon selected WAIS-R subtests. Journal of Psychoeducational Assessment, 9, 340–
344. Wolpe, J. (1958). Psychotherapy by reciprocal inhibition. Stanford, CA: Stanford University Press. Wolpe, J. (1973). The practice of behavior therapy (2nd ed.). New York: Pergamon. Wolpe, J., & Lang, P. J. (1977). Manual for the Fear Survey Schedule (revised). San Diego, CA: Educational and Industrial Testing Service. Wonderlic, E. F. (1983). Wonderlic Personnel Test manual. North�ield, IL: E. F. Wonderlic & Associates. Wood, J. M., Nezworski, M., & Stejskal, W. (1996). The Comprehensive System for the Rorschach: A critical examination. Psychological Science, 7, 3–10.
7/30/2019 Print
https://content.ashford.edu/print/Gregory.8055.17.1?sections=ch04,ch04lev1sec1,ch04lev1sec2,ch04lev1sec3,ch04lev1sec4,ch04lev1sec5,ch04lev… 70/70
Wood, J., Garb, H., & Nezworski, M. T. (2007). Psychometrics: Better measurement makes better clinicians. In S. O. Lilienfeld & W. T. O’Donohue (Eds.), The great ideas of clinical science: 17 principles that every mental health professional should understand. New York: Routledge.
Woodcock, R. W., McGrew, K. S., & Werder, J. K. (1994). Mini-Battery of Achievement: Examiner’s manual. Chicago: Riverside. Woodcock, R., McGrew, K., & Mather, N. (2001). Woodcock-Johnson III Tests of Achievement. Itasca, IL: Riverside. Woodworth, R. S. (1919). Examination of emotional �itness for warfare. Psychological Bulletin, 16, 59–60. Wortman, J., Lucas, R. E., & Donnellan, M. B. (2012, July 9). Stability and change in the Big Five personality domains: Evidence from a longitudinal study of
Australians. Psychology and Aging, online publication. Wrightsman, L., Nietzel, M., Fortune, W., & Greene, E. (2002). Psychology and the legal system (5th ed.). Paci�ic Grove, CA: Brooks/Cole. Wulff, D. M. (1996). The psychology of religion: An overview. In E. P. Shafranske (Ed.), Religion and the clinical practice of psychology. Washington, DC: American
Psychological Association. Wundt, W. (1862). Die Geschwindigkeit des Gedankens. Gartenlaube, 263–265. Yalisove, D. (2004). Introduction to alcohol research: Implications for treatment, prevention, and policy. Boston: Allyn & Bacon. Yama, M. (1990). The usefulness of human �igure drawings as an index of overall adjustment. Journal of Personality Assessment, 54, 78–86. Yerkes, R. M. (1919). Report of the psychology Committee of the National Research Council. Psychological Review, 26, 83–149. Yerkes, R. M. (Ed.). (1921). Psychological examining in the United States Army. Memoirs of the National Academy of Sciences, vol. 15. Yuan, Y. (2002). Development of the norm for the Fagan Test of Infant Intelligence in a town near Changsha. Chinese Mental health Journal, 16, 320–322. Zapf, P., & Roesch, R. (1997). Assessing �itness to stand trial: A comparison of institution-based evaluations and a brief screening interview. Canadian Journal of
Community Mental Health, 16, 53–66. Zapf, P., & Roesch, R. (2009). Evaluation of competence to stand trial. New York: Oxford University Press. Zapf, P., Skeem, J., & Golding, S. (2005). Factor structure and validity of the MacArthur Competence Assessment Tool—Criminal Adjudication. Psychological
Assessment, 17, 433–445. Zavala, A. (1965). Development of the forced-choice rating scale technique. Psychological Bulletin, 63, 117–124. Zeidner, M., Roberts, R., & Matthews, G. (2008). The science of emotional intelligence: Current consensus and controversies. European Psychologist, 13, 64–78. Zhai, F., Brooks-Gunn, J., & Waldfogel, J. (2011). Head Start and urban children’s school readiness: A birth cohort study in 18 cities. Developmental Psychology, 47,
134–152.