1-1
Psychological Testing and Assessment An Introduction to Tests and Measurement
TENTH EDITION
Ronald Jay Cohen RJ COHEN CONSULTING
W. Joel Schneider TEMPLE UNIVERSITY
Renée M. Tobin TEMPLE UNIVERSITY
coh37025_fm_i-xxviii.indd 1 12/01/21 4:02 PM
PSYCHOLOGICAL TESTING AND ASSESSMENT
Published by McGraw Hill LLC, 1325 Avenue of the Americas, New York, NY 10121. Copyright ©2022 by McGraw Hill LLC. All rights reserved. Printed in the United States of America. No part of this publication may be reproduced or distributed in any form or by any means, or stored in a database or retrieval system, without the prior written consent of McGraw Hill LLC, including, but not limited to, in any network or other electronic storage or transmission, or broadcast for distance learning.
Some ancillaries, including electronic and print components, may not be available to customers outside the United States.
This book is printed on acid-free paper.
1 2 3 4 5 6 7 8 9 LWI 26 25 24 23 22 21
ISBN 978-1-265-79973-1 MHID 1-265-79973-3
Cover Image: rimom/Shutterstock
All credits appearing on page or at the end of the book are considered to be an extension of the copyright page.
The Internet addresses listed in the text were accurate at the time of publication. The inclusion of a website does not indicate an endorsement by the authors or McGraw Hill LLC, and McGraw Hill LLC does not guarantee the accuracy of the information presented at these sites.
mheducation.com/highered
coh99733_ISE_ii.indd 2 11/01/21 7:59 PM
This book is dedicated with love to the memory of Edith and Harold Cohen.
© 2017 Ronald Jay Cohen. All rights reserved.
coh37025_fm_i-xxviii.indd 3 12/01/21 4:02 PM
iv
Preface xiii
P A R T I� An Overv iew
��Psychological Testing and Assessment��
TESTING AND ASSESSMENT 1 Psychological Testing and Assessment Defined 2
THE TOOLS OF PSYCHOLOGICAL ASSESSMENT 8 The Test 8 The Interview 10 The Portfolio 12 Case History Data 13 Behavioral Observation 13 Role-Play Tests 14 Computers as Tools 15 Other Tools 18
WHO, WHAT, WHY, HOW, AND WHERE? 18 Who Are the Parties? 19 In What Types of Settings Are Assessments Conducted, and Why? 21 How Are Assessments Conducted? 27 Where to Go for Authoritative Information: Reference Sources 33
CLOSE�UP�Behavioral Assessment Using Smartphones 5 MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Alan Ogle 25 EVERYDAY PSYCHOMETRICS�Everyday Accommodations 32 SELF�ASSESSMENT�36 REFERENCES�36
��Historical, Cultural, and Legal/Ethical Considerations���
A HISTORICAL PERSPECTIVE 41 Antiquity to the Nineteenth Century 41 The Twentieth Century 44
CULTURE AND ASSESSMENT 47 Evolving Interest in Culture-Related Issues 47 Some Issues Regarding Culture and Assessment 52 Tests and Group Membership 58
LEGAL AND ETHICAL CONSIDERATIONS 60 The Concerns of the Public 60 The Concerns of the Profession 68 The Rights of Testtakers 74
Contents
coh37025_fm_i-xxviii.indd 4 12/01/21 4:02 PM
Contents v
CLOSE�UP�The Controversial Career of Henry Herbert Goddard 49 MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Neil Krishan Aggarwal 56 EVERYDAY PSYCHOMETRICS�Life-or-Death Psychological Assessment 71 SELF�ASSESSMENT�79 REFERENCES�80
P A R T II� The Sc ience o f Psycho log ica l Measuremen t
��A Statistics Refresher���
SCALES OF MEASUREMENT 86 Nominal Scales 88 Ordinal Scales 89 Interval Scales 90 Ratio Scales 91 Measurement Scales in Psychology 91
DESCRIBING DATA 93 Frequency Distributions 93 Measures of Central Tendency 98 Measures of Variability 101 Skewness 105 Kurtosis 105
THE NORMAL CURVE 106 The Area Under the Normal Curve 107
STANDARD SCORES 110 z Scores 110 T Scores 111 Other Standard Scores 111
CORRELATION AND INFERENCE 113 The Concept of Correlation 114 The Pearson r 116 The Spearman Rho 118 Graphic Representations of Correlation 119 Meta-Analysis 123
EVERYDAY PSYCHOMETRICS�Consumer (of Graphed Data), Beware! 97 CLOSE�UP�The Normal Curve and Psychological Tests 108 MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Joni L. Mihura 124 SELF�ASSESSMENT�126 REFERENCES�127
��Of Tests and Testing����
SOME ASSUMPTIONS ABOUT PSYCHOLOGICAL TESTING AND ASSESSMENT� 130
Assumption 1: Psychological Traits and States Exist� 130
coh37025_fm_i-xxviii.indd 5 12/01/21 4:02 PM
vi���Contents
Assumption 2: Psychological Traits and States Can Be Quantified and Measured� 132 Assumption 3: Test-Related Behavior Predicts Non-Test-Related Behavior� 133 Assumption 4: All Tests Have Limits and Imperfections� 133 Assumption 5: Various Sources of Error Are Part of the Assessment Process� 134 Assumption 6: Unfair and Biased Assessment Procedures Can Be Identified and
Reformed 134 Assumption 7: Testing and Assessment Offer Powerful Benefits to Society� 135
WHAT’S A “GOOD TEST”?� 136 Reliability� 136 Validity� 137 Other Considerations� 137
NORMS 140 Sampling to Develop Norms 140 Types of Norms 146 Fixed Reference Group Scoring Systems 149 Norm-Referenced versus Criterion-Referenced Evaluation 150 Culture and Inference 153
EVERYDAY PSYCHOMETRICS� Putting Tests to the Test 138 CLOSE�UP� How “Standard” Is Standard in Measurement? 141 MEET AN ASSESSMENT PROFESSIONAL� Meet Dr. Steve Julius and Dr. Howard W. Atlas 152 SELF�ASSESSMENT� 154 REFERENCES� 155
��Reliability����
MEASUREMENT ERROR 157 TRUE SCORES VERSUS CONSTRUCT SCORES 158 THE CONCEPT OF RELIABILITY 159
Sources of Error Variance 160 RELIABILITY ESTIMATES 163
Test-Retest Reliability Estimates 163 Parallel-Forms and Alternate-Forms Reliability Estimates 164 Split-Half Reliability Estimates 167 Other Methods of Estimating Internal Consistency 170 Measures of Inter-Scorer Reliability 172
USING AND INTERPRETING A COEFFICIENT OF RELIABILITY 174 The Purpose of the Reliability Coefficient 175 The Nature of the Test 176 The True Score Model of Measurement and Alternatives to It 179
RELIABILITY AND INDIVIDUAL SCORES 183 The Standard Error of Measurement 183 The Standard Error of the Difference Between Two Scores 187
CLOSE�UP�Psychology’s Replicability Crisis 165 EVERYDAY PSYCHOMETRICS�The Importance of the Method Used for Estimating Reliability 173 MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Bryce B. Reeve 184
coh37025_fm_i-xxviii.indd 6 12/01/21 4:02 PM
Contents vii
SELF�ASSESSMENT�189 REFERENCES�190
��Validity����
THE CONCEPT OF VALIDITY 193 Face Validity 195 Content Validity 196
CRITERION-RELATED VALIDITY 200 What Is a Criterion? 200 Concurrent Validity 202 Predictive Validity 202
CONSTRUCT VALIDITY 205 Evidence of Construct Validity 206
VALIDITY, BIAS, AND FAIRNESS 211 Test Bias 211 Test Fairness 214
MEET AN ASSESSMENT PROFESSIONAL� Meet Dr. Adam Shoemaker 197 CLOSE�UP� The Preliminary Validation of a Measure of� Individual Differences in Constructive
versus Unconstructive Worry 212 EVERYDAY PSYCHOMETRICS� Adjustment of Test Scores by Group Membership: Fairness in Testing
or Foul Play? 216 SELF�ASSESSMENT� 218 REFERENCES� 218
��Utility����
WHAT IS TEST UTILITY?� 222 Factors That Affect a Test’s Utility 222
UTILITY ANALYSIS � 227 What Is a Utility Analysis?� 227 How Is a Utility Analysis Conducted?� 228 Some Practical Considerations 242
METHODS FOR SETTING CUT SCORES 245 The Angoff Method 246 The Known Groups Method 246 IRT-Based Methods 247 Other Methods 248
MEET AN ASSESSMENT PROFESSIONAL� Meet Dr. Delphine Courvoisier 225 CLOSE�UP� Utility Analysis: An Illustration 229 EVERYDAY PSYCHOMETRICS� The Utility of Police Use of Body Cameras 239 SELF�ASSESSMENT�248 REFERENCES� 249
��Test Development����
TEST CONCEPTUALIZATION 252 Some Preliminary Questions 254
coh37025_fm_i-xxviii.indd 7 12/01/21 4:02 PM
viii���Contents
Pilot Work 256 TEST CONSTRUCTION 256
Scaling 256 Writing Items 261 Scoring Items 268
TEST TRYOUT 268 What Is a Good Item? 269
ITEM ANALYSIS 270 The Item-Difficulty Index 270 The Item-Reliability Index 271 The Item-Validity Index 272 The Item-Discrimination Index 272 Item-Characteristic Curves 275 Other Considerations in Item Analysis 278 Qualitative Item Analysis 280
TEST REVISION 282 Test Revision as a Stage in New Test Development 282 Test Revision in the Life Cycle of an Existing Test 284 The Use of IRT in Building and Revising Tests 288
INSTRUCTOR-MADE TESTS FOR IN-CLASS USE 291 Addressing Concerns About Classroom Tests 291
CLOSE�UP�Creating and Validating a Test of Asexuality 253 MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Scott Birkeland 276 EVERYDAY PSYCHOMETRICS�Adapting Tools of Assessment for Use with Specific Cultural Groups 283 SELF�ASSESSMENT�293 REFERENCES�294
P A R T III�The Assessmen t o f I n te l l i gence
��Intelligence and Its Measurement����
WHAT IS INTELLIGENCE? 297 Perspectives on Intelligence 299
MEASURING INTELLIGENCE� 312 Some Tasks Used to Measure Intelligence 312 Some Tests Used to Measure Intelligence� 314
ISSUES IN THE ASSESSMENT OF INTELLIGENCE 334 Culture and Measured Intelligence 335 The Flynn Effect 340 The Construct Validity of Tests of Intelligence 341
A PERSPECTIVE� 341 CLOSE�UP� Factor Analysis 302 MEET AN ASSESSMENT PROFESSIONAL� Meet Dr. Rebecca Anderson 315 EVERYDAY PSYCHOMETRICS� The Armed Services Vocational Aptitude Battery (ASVAB): A Test
You Can Take 330
coh37025_fm_i-xxviii.indd 8 12/01/21 4:02 PM
Contents ix
SELF�ASSESSMENT� 342 REFERENCES� 343
���Assessment for Education����
THE ROLE OF TESTING AND ASSESSMENT IN EDUCATION� 349 THE CASE FOR AND AGAINST EDUCATIONAL TESTING IN THE SCHOOLS� 350 THE COMMON CORE STATE STANDARDS 351
Response to Intervention (RtI) 352 Dynamic Assessment� 358
ACHIEVEMENT TESTS� 360 Measures of General Achievement 360 Measures of Achievement in Specific Subject Areas 361
APTITUDE TESTS 363 The Preschool Level 365 The Elementary-School Level 370 The Secondary-School Level 372 The College Level and Beyond 373
DIAGNOSTIC TESTS 376 Reading Tests 377 Math Tests 378
PSYCHOEDUCATIONAL TEST BATTERIES 378 The Kaufman Assessment Battery for Children, Second Edition Normative Update
(KABC-II NU) 378 The Woodcock-Johnson IV (WJ IV) 380
OTHER TOOLS OF ASSESSMENT IN EDUCATIONAL SETTINGS 381 Performance, Portfolio, and Authentic Assessment 381 Peer Appraisal Techniques 383 Measuring Study Habits, Interests, and Attitudes 384
EVERYDAY PSYCHOMETRICS� The Common Core Controversy 353 MEET AN ASSESSMENT PROFESSIONAL� Meet Eliane Keyes, M.A. 357 CLOSE�UP� Educational Assessment: An Eastern Perspective 371 SELF�ASSESSMENT� 385 REFERENCES� 385
P A R T IV�The Assessmen t o f Pe rsona l i t y
���Personality Assessment: An Overview����
PERSONALITY AND PERSONALITY ASSESSMENT� 390 Personality� 390 Personality Assessment� 391 Traits, Types, and States� 391
PERSONALITY ASSESSMENT: SOME BASIC QUESTIONS� 395 Who? 396
coh37025_fm_i-xxviii.indd 9 12/01/21 4:02 PM
x���Contents
What?� 402 Where?� 404 How?� 404
DEVELOPING INSTRUMENTS TO ASSESS PERSONALITY 413 Logic and Reason 413 Theory� 416 Data Reduction Methods 416 Criterion Groups 419
PERSONALITY ASSESSMENT AND CULTURE� 431 Acculturation and Related Considerations� 431
CLOSE�UP� The Personality of Gorillas 397 EVERYDAY PSYCHOMETRICS� Some Common Item Formats 408 MEET AN ASSESSMENT PROFESSIONAL� Meet Dr. Rick Malone 414 SELF�ASSESSMENT� 435 REFERENCES� 435
���Personality Assessment Methods����
OBJECTIVE METHODS� 444 How Objective Are Objective Methods of Personality Assessment?� 445
PROJECTIVE METHODS� 445 Inkblots as Projective Stimuli� 447 Pictures as Projective Stimuli 453 Words as Projective Stimuli 461 Sounds as Projective Stimuli 464 The Production of Figure Drawings 465 Projective Methods in Perspective 468
BEHAVIORAL ASSESSMENT METHODS� 472 The Who, What, When, Where, Why, and How of It 474 Varieties of Behavioral Assessment 478 Issues in Behavioral Assessment 485
A PERSPECTIVE 487 MEET AN ASSESSMENT PROFESSIONAL� Meet Dr. Monica Webb Hooper 476 EVERYDAY PSYCHOMETRICS� Confessions of a Behavior Rater 479 CLOSE�UP� General (g) and Specific (s) Factors in the Diagnosis of Personality
Disorders� 488 SELF�ASSESSMENT�490 REFERENCES�490
P A R T V�Tes t ing and Assessmen t i n Ac t i on
���Clinical and Counseling Assessment����
AN OVERVIEW� 499 The Diagnosis of Mental Disorders 501
coh37025_fm_i-xxviii.indd 10 12/01/21 4:02 PM
Contents xi
The Interview in Clinical Assessment 504 Case History Data 511 Psychological Tests 511
CULTURALLY INFORMED PSYCHOLOGICAL ASSESSMENT 513 Cultural Aspects of the Interview 515
SPECIAL APPLICATIONS OF CLINICAL MEASURES 518 The Assessment of Addiction and Substance Abuse 518 Forensic Psychological Assessment 520 Diagnosis and evaluation of emotional injury 526 Profiling 526 Custody Evaluations 527
CHILD ABUSE AND NEGLECT 530 Elder Abuse and Neglect 532 Suicide Assessment 534
THE PSYCHOLOGICAL REPORT 535 The Barnum Effect 535 Clinical Versus Mechanical Prediction 537
MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Stephen Finn 507 CLOSE�UP�PTSD in Returning Veterans and Military Culture 516 EVERYDAY PSYCHOMETRICS�Measuring Financial Competency 524 SELF�ASSESSMENT�539 REFERENCES�540
�� Neuropsychological Assessment ���
THE NERVOUS SYSTEM AND BEHAVIOR 550 Neurological Damage and the Concept of Organicity 551
THE NEUROPSYCHOLOGICAL EVALUATION 554 When a Neuropsychological Evaluation Is Indicated 554 General Elements of a Neuropsychological Evaluation 556 The Physical Examination 559
NEUROPSYCHOLOGICAL TESTS 565 Tests of General Intellectual Ability 565 Tests to Measure the Ability to Abstract 567 Tests of Executive Function 568 Tests of Perceptual, Motor, and Perceptual-Motor Function 572 Tests of Verbal Functioning 573 Tests of Memory 573 Neuropsychological Test Batteries 576
OTHER TOOLS OF NEUROPSYCHOLOGICAL ASSESSMENT 580 MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Jeanne P. Ryan 566 EVERYDAY PSYCHOMETRICS�Medical Diagnostic Aids and Neuropsychological Assessment 581 CLOSE�UP�A Typical In-Office Dementia Evaluation 583 SELF�ASSESSMENT�584 REFERENCES�584
coh37025_fm_i-xxviii.indd 11 12/01/21 4:02 PM
xii���Contents
�� Assessment, Careers, and Business ���
CAREER CHOICE AND CAREER TRANSITION 590 The Structure of Vocational Interests 590 Measures of Interest 592 Measures of Ability and Aptitude 594 Measures of Personality 596 Other Measures 599
SCREENING, SELECTION, CLASSIFICATION, AND PLACEMENT 601 The Résumé and the Letter of Application 602 The Application Form 602 Letters of Recommendation 602 Interviews 603 Portfolio Assessment 604 Performance Tests 604 Physical Tests 609
COGNITIVE ABILITY, PRODUCTIVITY, AND MOTIVATION MEASURES 611 Measures of Cognitive Ability 611 Productivity 612 Motivation 613
JOB SATISFACTION, ORGANIZATIONAL COMMITMENT, AND ORGANIZATIONAL CULTURE 617
Job Satisfaction 617 Organizational Commitment 618 Organizational Culture 619
OTHER TOOLS OF ASSESSMENT FOR BUSINESS APPLICATIONS 619 Consumer Psychology 620 The Measurement of Attitudes 621 Surveys 623 Motivation Research Methods 625
EVERYDAY PSYCHOMETRICS�The Selection of Personnel for the Office of Strategic Services (OSS): Assessment and Psychometrics in Action 606
MEET AN ASSESSMENT PROFESSIONAL�Meet Dr. Jed Yalof 620 SELF�ASSESSMENT�629 REFERENCES�629
Name Index I-1 Glossary/Index I-22 Timeline T-1
coh37025_fm_i-xxviii.indd 12 12/01/21 4:02 PM
xiii
Preface e are proud to welcome instructors of a measurement course in psychology to this tenth edition of Psychological Testing and Assessment. Thank you for the privilege of assisting in the exciting task of introducing the world of tests and measurement to your students. In this preface, we impart our vision for a measurement textbook, as well as the philosophy that has driven, and that continues to drive, the organization, content, writing style, and pedagogy of this book. We’ll briefly look back at this book’s heritage and discuss what is new and distinctive about this tenth edition. Of particular interest to instructors, this preface will overview the authors’ general approach to the course content and distinguish how that approach differs from other measurement textbooks. For students who happen to be curious enough to read this preface (or ambitious enough to read it despite the fact that it was not assigned), we hope that your takeaway from it has to do with the authors’ genuine dedication to making this book the far-and-away best available textbook for your measurement course.
Our Vision for a Textbook on Psychological Testing and Assessment
First and foremost, let’s get out there that the subject matter of this course is psychological testing and assessment—a fact that is contrary to the message conveyed by an array of would-be competitor books, all distinguished by their anachronistic “psychological testing” title. Of course we cover tests and testing, and no available textbook does it better or more comprehensively. But it behooves us to observe that we are now well into the twenty-first century and it has long been recognized that tests are only one tool of assessment. Psychological testing is a process that can be—perhaps reminiscent of those books with the same title—impersonal, noncreative, uninspired, routine, and even robotic in nature. By contrast, psychological assessment is a human, dynamic, custom, creative, and collaborative enterprise. These aspects of the distinction between psychological testing and psychological assessment are not trivial.
Paralleling important differences between our book’s title and that of other books in this area are key differences in the way that the subject matter of the course is approached. In routine writing and through a variety of pedagogical tools, we attempt to draw students into the world of testing and assessment by humanizing the material. Our approach to the course material stands in stark contrast to the “by-the-numbers” approach of some of our competitors; the latter approach can easily alienate readers, prompting them to “tune out.” Let’s briefly elaborate on this critical point.
Although most of our competitors begin by organizing their books with an outline that for the most part mimics our own—right down to the inclusion of the Statistics Refresher that we innovated some 30 years ago—the way that they cover that subject matter, and the pedagogical tools they rely on to assist student learning, bear only cosmetic resemblance to our approach. We take every opportunity to illustrate the course material by putting a human face to it, and by providing practical, “every day” examples of the principles and procedures at work. This approach differs in key ways from the approach of other books in the area, in which a “practical approach” may instead be equated with the intermingling of statistical or other exercises within every chapter of the book. Presumably, according to the latter vision, a textbook is a simultaneous delivery system for both course-related information and course-related exercises. Students are expected to read their textbooks until such time that their reading is interrupted by an exercise. After the completion of the exercise, students are expected to go back to the reading, but only until they happen upon another exercise. It is thus the norm to interrupt absorption in assigned reading on a relatively random (variable ratio) schedule in order to have students complete
W
coh37025_fm_i-xxviii.indd 13 12/01/21 4:02 PM
xiv���Preface
general, one-size-fits-all exercises. Students using such a book are not encouraged to concentrate on assigned reading; they may even be tacitly encouraged to do the opposite. The emphasis given to students having to complete exercises scattered within readings seems especially misplaced when, as is often the case with such one-size-fits-all tasks, some of the exercises will be way too easy for students in some classes and way too difficult for students in others. This situation brings to mind our own experience with testing-related exercises being assigned to varied groups of introductory students.
For several years and through several editions, our textbook was published with a supplementary exercises workbook. After extensive feedback from many instructors, some of whom used our book in their classes and some of whom did not, we determined that matters related to the choice, content, and level of supplementary exercises were better left to individual instructors as opposed to textbook authors. In general, instructors preferred to assign their own supplementary exercises, which could be custom-designed for the needs of their particular students and the goals of their particular course. A workbook of exercises, complete with detailed, step-by-step, illustrated solutions of statistical and psychometric problems, was determined by us to add little value to our textbook and it is therefore no longer offered. What we learned, and what we now believe, is that there is great value to supplementary, ancillary exercises for students taking an introductory course in measurement. However, these exercises are of optimal use to the student when they are custom-designed (or selected) by the instructor based on factors such as the level and interest of the students in the class, and the students’ in-class and out-of-class study schedule. To be clear, supplemental exercises randomly embedded in a textbook work, in our view, not to facilitate students’ immersion and concentration in assigned reading, but to obliterate it.1
Given that decisions regarding supplementary exercises are best left to individual instructors, the difference between our own approach to the subject matter of the course and that of other approaches are even more profound. In this tenth edition, we have concentrated our attention and effort to crafting a textbook that will immerse and involve students in assigned readings and motivate them to engage in critical and generative thinking about what they have read. Contrast that vision with one in which author effort is divided between writing text and writing nonsupplementary exercises. Could the net result of the latter approach be a textbook that divides student attention between assigned readings and assigned (or unassigned) exercises? Seasoned instructors may concur with our view that most students will skip the intrusive and distracting exercises when they are not specifically assigned for completion by the instructor. In the case where the exercises are assigned, students may well skim the reading to complete the exercises.
No available textbook is more focused on being practical, timely, and “real-life” oriented than our book is. Further, no other textbook provides students in an introductory course with a more readable or more comprehensive account of how psychological tests and assessment- related procedures are used in practice. That has been the case for some 30 years and it most certainly is the case today. With that as background, let’s briefly sum up some of our concerns with regard to certain members of the current community of “psychological testing” books.
Especially with regard to a textbook at the introductory level, what is critical is the breadth and depth of coverage of how tests and other tools of assessment are actually used in practice. Practice-level proficiency and hands-on experience are always nice, but may in some cases be too ambitious. For example, a practical approach to factor analysis in a textbook for an introductory measurement course need not equip the student to conduct a factor analysis.
1. We urge any instructors curious about this assertion to informally evaluate it by asking a student or two how they feel about the prospect of scattering statistical exercises in their assigned reading. If the assigned reading is at all immersive, the modal response may be something like “maddening.”
coh37025_fm_i-xxviii.indd 14 12/01/21 4:02 PM
Preface xv
Rather, the coverage ideally provides the student with a sound grounding in what this widely used set of techniques are, as well as how and why they are used. Similarly a practical approach to test utility, as exemplified in Chapter 7, provides students with a sound grounding in what that construct is, as well as how and why it is applied in practice.
Of course when it comes to breadth and depth of coverage of how tests and other tools of assessment are actually used in practice, we have long been the standard by which other books are measured. Consider in this context a small sampling of what is new, timely, and relevant in this tenth edition. The subject of our Chapter 1 Close-Up is behavioral assessment using smart phones. The subject of our Everyday Psychometrics in Chapter 7 on utility is the utility of police use of body cameras.2 Terrorism is a matter of worldwide concern and in Chapter 11, the professional profiled in our Meet an Assessment Professional feature is Colonel Rick Malone of the United States Army’s Criminal Investigation Command. Dr. Malone shares some intriguing insights regarding his area of expertise: threat assessment. Much more about our vision for this textbook and its supplements, as well as more previews of what is new and exciting in this tenth edition, is presented in what follows.
Organization
From the first edition of our book forward, we have organized the information to be presented into five major sections. Part I, An Overview, contains two chapters that do just that. Chapter�1 provides a comprehensive overview of the field, including some important definitional issues, a general description of tools of assessment, and related important information couched as answers to questions regarding the who, what, why, how, and where of the enterprise.
The foundation for the material to come continues to be laid in the second chapter of the overview, which deals with historical, cultural, and legal/ethical issues. The material presented in Chapter 2 clearly sets a context for everything that will follow. To relegate such material to the back of the book (as a kind of elective topic, much like the way that legal/ethical issues are treated in some books), or to ignore presentation of such material altogether (as most other books have done with regard to cultural issues in assessment), is, in our estimation, a grave error. “Back page infrequency” (to borrow an MMPI-2 term) is too often the norm, and relegation of this critically important information to the back pages of a textbook too often translates to a potential shortchanging of students with regard to key cultural, historical, and legal/ethical information. The importance of exposure early on to relevant historical, cultural, and legal/ethical issues cannot be overemphasized. This exposure sets a context for succeeding coverage of psychometrics and creates an essential lens through which to view and process such material.
Part II, The Science of Psychological Measurement, contains Chapters 3 through 8. These six chapters were designed to build—logically and sequentially—on the student’s knowledge of psychometric principles. Part II begins with a chapter reviewing basic statistical principles and ends with a chapter on test construction. In between, there is extensive discussion of assumptions inherent in the enterprise, the elements of good test construction, as well as the concepts of norms, correlation, inference, reliability, and validity. All of the measurement
2. This essay is an informative and timely discussion of the utility of police-worn body cameras in reducing use-of-force complaints. Parenthetically, let’s share our view that the concept of utility seems lost in, or at least given inadequate coverage in other measurement books. It seems that we may have caught many of those “psychological testing” books off-guard by devoting a chapter to this construct beginning with our seventh edition—this at a time when utility was not even an indexed term in most of them. Attempts to compensate have ranged from doing nothing at all to doing near nothing at all by equating “utility” with “validity.” For the record, although utility is related to validity, much as reliability is related to validity, we believe it is misleading to even intimate that “utility” and “validity” are synonymous.
coh37025_fm_i-xxviii.indd 15 12/01/21 4:02 PM
xvi���Preface
textbooks that came before us were written based on the assumption that every student taking the course was up to speed on all of the statistical concepts that would be necessary to build on learning about psychometrics. In theory, at least, there was no reason not to assume this previous knowledge; statistics was a prerequisite to taking the course. In practice, a different picture emerged. It was simply not the case that all students were adequately and equally prepared to begin learning statistics-based measurement concepts. Our remedy for this problem, some 30 years ago, was to include a “Statistics Refresher” chapter early on, just prior to building on students’ statistics-based knowledge. The rest, as they say, is history...
Our book forever changed for the better the way the measurement course was taught and the way all subsequent textbooks for the course would be written. Our unique coverage of the assessment of intelligence and personality, as well as our coverage of assessment for various applications (ranging from neuropsychological to business and organizational applications), made relics of the typical “psychological testing” course outline as it existed prior to the publication of our first edition in 1988.
In our seventh edition, in response to increasing general interest in test utility, we added a chapter on this important construct right after our chapters on the constructs of reliability and validity. Let’s note here that topics such as utility and utility analysis can get extremely complicated. However, we have never shied away from the presentation of complicated subject matter. For example, we were the first introductory textbook to present detailed information related to factor analysis. As more commercial publishers and other test users have adopted the use of item response theory (IRT) in test construction, our coverage of IRT has kept pace. As more test reviews have begun to evaluate tests not only in terms of variables such as reliability and validity but in terms of utility, we saw a need for the inclusion of a chapter on that topic.
Of course, no matter how “difficult” the concepts we present are, we never for a moment lose sight of the appropriate level of presentation. This book is designed for students taking a first course in psychological testing and assessment. Our objective in presenting material on methods such as IRT and utility analysis is simply to acquaint the introductory student with these techniques. The depth of the presentation in these and other areas has always been guided and informed by extensive reviews from a geographically diverse sampling of instructors who teach measurement courses. For users of this textbook, what currently tends to be required is a conceptual understanding of commonly used IRT methods. We believe our presentation of this material effectively conveys such an understanding. Moreover, it does so without unnecessarily burdening students with level-inappropriate formulas and calculations.
Part III of this book, The Assessment of Abilities and Aptitudes, contains two chapters, one on intelligence and its assessment, and the other on assessment in schools and other educational settings. In past editions of this book, two chapters were devoted to the assessment of intelligence. To understand why, it is instructive to consider what the coverage of intelligence testing looked like in the then available introductory measurement textbooks three decades ago. While the books all covered tests of intelligence, they devoted little or no attention to defining and discussing the construct of intelligence. We called attention to this problem and attempted to remedy it by differentiating our book with a chapter devoted to imparting a conceptual understanding of intelligence. Although revolutionary at the time, the logic of our approach had widespread appeal. Before long, the typical “psychological testing” course of the 1980s was being restructured to include conceptual discussions of concepts such as “intelligence” and “personality” before proceeding to discuss their measurement. The “psychological testing” textbooks of the day also followed our lead. And so, to the present day, two-chapter-coverage of the assessment of intelligence (with the first chapter providing a discussion of the construct of intelligence) has become the norm.
In retrospect, it seems reasonable to conclude that our addition of a chapter on the nature of intelligence, much like our addition of a statistics refresher, did more than remedy a serious drawback in existing measurement textbooks; it forever revolutionized the way that the
coh37025_fm_i-xxviii.indd 16 12/01/21 4:02 PM
Preface xvii
measurement course was taught in classrooms around the world. It did this first of all by making the teaching of the course more logical. This is so because the logic of our guiding principle—fully define and discuss the psychological construct being measured before discussing its measurement—had wide appeal. In our first edition, we also extended that logic to the discussion of the measurement of other psychological constructs such as personality. Another benefit we saw in adding the conceptual coverage was that such coverage would serve to “humanize” the content. After all, “Binet” was more than just the name of a psychological test; it was the name of a living, breathing person.
Also, since our first edition, we have revolutionized textbook coverage of psychological tests— this by a philosophy of “less is more” when it comes to such coverage. Back in the 1980s, the “psychological testing” books of the day had elements reminiscent of Tests in Print. They provided reliability, validity, and related psychometric data on dozens of psychological tests. But we raised the question, “Why duplicate in a textbook information about dozens of tests that is readily available from reference sources?” We further resolved to limit detailed coverage of psychological tests to a handful of representative tests. Once again, the simple logic of our approach had widespread appeal, and other textbooks in the area—both then, and to the present day—all followed suit.
There is another trend in textbook coverage of the measurement course that also figured prominently in our decision to cover the assessment of intelligence in a single chapter. This trend has to do with the widespread availability of online resources to supplement coverage of a specific topic. We have long taken advantage of this fact by making available various supplementary materials online to our readers, or by supplying links to such materials.
Some three decades after we revolutionized the organization of textbook coverage of the measurement course in so many significant ways, it was time to re-evaluate whether two chapters to cover the subject of intelligence assessment was still necessary. We gave thoughtful consideration to this question and sought-out the opinion of trusted colleagues. In the end, we determined that coverage of the construct and assessment of intelligence could be accomplished in a single chapter. And so, in the interest of streamlining this book in length, Chapter 9 in the ninth edition incorporated text formerly in Chapters 9 and 10 of the eighth edition. This combined chapter was maintained in the tenth edition of the textbook.
Part IV, The Assessment of Personality, contains two chapters, which respectively overview how personality assessments are conducted, and the various methods used.
Part V, Testing and Assessment in Action, is designed to convey to students a sense of how a sampling of tests and other tools of assessment are actually used in clinical, counseling, business, and other settings.
Content
In addition to a logical organization that sequentially builds on student learning, we view content selection as another key element of our appeal. The multifaceted nature and complexity of the discipline affords textbook authors wide latitude in terms of what material to elaborate on, what material to ignore, and what material to highlight, exemplify, or illustrate. In selecting content to be covered for chapters, the primary question for us was most typically “What do students need to know?” So, for example, since the publication of previous editions of this book, the field of educational evaluation has been greatly influenced by the widespread implementation of the Common Core Standards. Accordingly, we take cognizance of these changes in the K-through-12 education landscape and their implications for evaluation in education. Students of educational assessment need to know about the Common Core Standards and relevant coverage of these standards can be found in this tenth edition in our chapter on educational assessment.
While due consideration is given to creating content that students need to know, consideration is also given to relevant topics that will engage interest and serve as stimuli for
coh37025_fm_i-xxviii.indd 17 12/01/21 4:02 PM
xviii���Preface
critical or generative thinking. In the area of neuropsychological assessment, for example, the topic of Alzheimer’s disease is one that generates a great deal of interest. Most students have seen articles or feature stories in the popular media that review the signs and symptoms of this disease. However, while students are aware that such patients are typically referred to a neurologist for formal diagnosis, many questions remain about how a diagnosis of Alzheimer’s disease is clinically made. The Close-Up in our chapter on neuropsychological assessment addresses those frequently asked questions. It was guest-authored by an experienced neurologist and written especially for students of psychological assessment reading this textbook.
Let’s note here that in this tenth edition, more than in any previous edition of this textbook, we have drawn on the firsthand knowledge of psychological assessment experts from around the world. Specifically, we have asked these experts to guest-author brief essays in the form of Close-Up, Everyday Psychometrics, or Meet an Assessment Professional features. For example, in one of our chapters that deal with personality assessment, two experts on primate behavior (including one who is currently working at Dian Fossey’s research center in Karisoke, in Rwanda) prepared an essay on evaluating the personality of gorillas. Written especially for us, this Close-Up makes an informative contribution to the literature on cross-species personality assessment. In our chapter on test construction, an Australian team of behavioral scientists guest-authored a Close-Up entitled “Adapting Tools of Assessment for Use with Specific Cultural Groups.” This essay recounts some of the intriguing culture-related challenges inherent in the psychological assessment of clients from the Aboriginal community.
Sensitivity to cultural issues in psychological testing and assessment is essential, and this textbook has long set the standard for coverage of such issues. Coverage of cultural issues begins in earnest in Chapter 2, where we define culture and overview the importance of cultural considerations in everything from test development to standards of evaluation. Then, much like an identifiable musical theme that recurs throughout a symphony, echoes of the importance of culture repeat in various chapters throughout this book. For example, the echo is heard in Chapter 4 where, among other things, we continue a long tradition of acquainting students with the “do’s and don’ts” of culturally informed assessment. In Chapter 13, our chapter on assessment in clinical and counseling settings, there is a discussion of acculturation and culture as these issues pertain to clinical assessment. Also in that chapter, students will find a thought- provoking Close-Up entitled, “PTSD in Veterans and the Idealized Culture of Warrior Masculinity.” Guest-authored especially for us by Duncan M. Shields, this timely contribution to the clinical literature sheds light on the diagnosis and treatment of post-traumatic stress disorder (PTSD) from a new and novel, cultural perspective.
In addition to standard-setting content related to cultural issues, mention must also be made of our leadership role with respect to coverage of historical and legal/ethical aspects of measurement in psychology. Our own appreciation for the importance of history is emphasized by the listing of noteworthy historical events that is set within the front and back covers of this textbook. As such, readers may be greeted with some aspect of the history of the enterprise on every occasion that they open the book. Although historical vignettes are distributed throughout the book to help set a context or advance understanding, formal coverage begins in Chapter 2. Important historical aspects of testing and assessment may also be found in Close-Ups. See, for example, the fascinating account of the controversial career of Henry Goddard found in Chapter 2. In a Close-Up in Chapter 15, students will discover what contemporary assessment professionals can learn from World War II-vintage assessment data collected by the Office of Strategic Services (OSS). In this engrossing essay, iconic data meets contemporary data analytic methods with brilliant new insights as a result. This Close-Up was guest-authored by Mark F. Lenzenweger, who is a State University of New York (SUNY) Distinguished Professor in the Department of Psychology at the State University of New York at Binghamton.
Much like content pertaining to relevant historical and culture-related material, our discussion of legal–ethical issues, from our first edition through to the present day, has been standard-setting.
coh37025_fm_i-xxviii.indd 18 12/01/21 4:02 PM
Preface xix
Discussion of legal and ethical issues as they apply to psychological testing and assessment provides students not only with context essential for understanding psychometric principles and practice, but another lens through which to filter understanding of tests and measurement. In the first edition, while we got the addition of this pioneering content right, we could have done a better job in terms of placement. In retrospect, the first edition would have benefitted from the discussion of such issues much earlier than the last chapter. But in response to the many compelling arguments reviewers and users of that book, discussion of legal/ethical issues was prioritized in Chapter 2 by the time that our second edition was published. The move helped ensure that students were properly equipped to appreciate the role of legal and ethical issues in the many varied settings in which psychological testing and assessment takes place.
Another element of our vision for the content of this book has to do with the art program; that is, the photos, drawings, and other types of illustrations used in a textbook. Before the publication of our ground-breaking first edition, what passed for an art program in the available “psychological testing” textbooks were some number-intensive graphs and tables, as well as photos of test kits or test materials. In general, photos and other illustrations seemed to be inserted more to break up text than to complement it. For us, the art program is an important element of a textbook, not a device for pacing. Illustrations can help draw students into the narrative, and then reinforce learning by solidifying meaningful visual associations to the written words. Our figures and graphics bring concepts to life. Photos can be powerful tools to stir the imagination. See, for example, the photo of Army recruits being tested in Chapter 1, or the photo of Ellis Island immigrants being tested in Chapter 2. Photos can bring to life and “humanize” the findings of measurement-related research. See, for example, the photo in Chapter 3 regarding the study that examined the relationship between grades and cell phone use in class. Photos of many past and present luminaries in the field (such as John Exner, Jr. and Ralph Reitan), and photos accompanying the persons featured in our Meet an Assessment Professional boxes all serve to breathe life into their respective accounts and descriptions.
In the world of textbooks, photos such as the sampling of the ones described here may not seem very revolutionary. However, in the world of measurement textbooks, our innovative art program has been and remains quite revolutionary. One factor that has always distinguished us from other books in this area is the extent to which we have tried to “humanize” the course subject matter; the art program is just another element of this textbook pressed into the service of that objective.
“Humanization” of Content� This tenth edition was conceived with a commitment to continuing our three-decade tradition of exemplary organization, exceptional writing, timely content, and solid pedagogy. Equally important was our desire to spare no effort in making this book as readable and as involving for students as it could possibly be. Our “secret sauce” in accomplishing this is, at this point, not much of a secret. We have the highest respect for the students for whom this book is written. We try to show that respect by never underestimating their capacity to become immersed in course-relevant narratives that are presented clearly and straightfor- wardly. With the goal of further drawing the student into the subject matter, we make every effort possible to “humanize” the presentation of topics covered. So, what does “humanization” in this context actually mean?
While other authors in this discipline impress us as blindly intent on viewing the field as Greek letters to be understood and formulas to be memorized, we view an introduction to the field to be about people as much as anything else. Students are more motivated to learn this material when they can place it in a human context. Many psychology students simply do not respond well to endless presentations of psychometric concepts and formulas. In our opinion, to not bring a human face to the field of psychological testing and assessment, is to risk perpetuating all of those unpleasant (and now unfair) rumors about the course that first began circulating long before the time that the senior author himself was an undergraduate.
coh37025_fm_i-xxviii.indd 19 12/01/21 4:02 PM
xx���Preface
Our effort to humanize the material is evident in the various ways we have tried to bring a face (if not a helping voice) to the material. The inclusion of Meet an Assessment Professional is a means toward that end, as it quite literally “brings a face” to the enterprise. Our inclusion of interesting biographical facts on historical figures in assessment is also representative of efforts to humanize the material. Consider in this context the photo and brief biographical statement of MMPI-2 senior author James Butcher in Chapter 11 (p. 426). Whether through such images of historical personages or by other means, our objective has been made to truly involve students via intriguing, real-life illustrations of the material being discussed. See, for example, the discussion of life-or-death psychological assessment and the ethical issues involved in the Close-Up feature of Chapter 2. Or check out the candid “confessions” of a behavior rater in the Everyday Psychometrics feature in Chapter 12.
So how has our “humanization” of the material in this discipline been received by some of its more “hard core” and “old school” practitioners? Very well, thank you—at least from all that we have heard, and the dozens of reviews that we have read over the years. What stands out prominently in the mind of the senior author (RJC) was the reaction of one particular psychometrician whom I happened to meet at an APA convention not long after the first edition of this text was published. Lee J. Cronbach was quite animated as he shared with me his delight with the book, and how refreshingly different he thought that it was from anything comparable that had been published. I was so grateful to Lee for his encouragement, and felt so uplifted by that meeting, that I subsequently requested a photo from Lee for use in the second edition. The photo he sent was indeed published in the second edition of this book—this despite the fact that at that time, Lee had a measurement book that could be viewed as a direct competitor to ours. Regardless, I felt it was important not only to acknowledge Lee’s esteemed place in measurement history, but to express my sincere gratitude in this way for his kind, inspiring, and motivating words, as well as for what I perceived as his most valued “seal of approval.”
Pedagogical Tools
The objective of incorporating timely, relevant, and intriguing illustrations of assessment-related material is furthered by several pedagogical tools built into the text. One pedagogical tool we created several editions ago is Everyday Psychometrics. In each chapter of the book, relevant, practical, and “everyday” examples of the material being discussed are highlighted in an Everyday Psychometrics box. For example, in the Everyday Psychometrics presented in Chapter 1 (“Everyday Accommodations”), students will be introduced to accommodations made in the testing of persons with handicapping conditions. In Chapter 4, the Everyday Psychometrics feature (“Putting Tests to the Test”) equips students with a working overview of the variables they need to be thinking about when reading about a test and evaluating how satisfactory the test really is for a particular purpose. In Chapter 5, the subject of the Everyday Psychometrics is how the method used to estimate diagnostic reliability may affect the obtained estimate of reliability.
A pedagogical tool called Meet an Assessment Professional was first introduced in the seventh edition. This feature provides a forum through which everyday users of psychological tests from various fields can share insights, experiences, and advice with students. The result is that in each chapter of this book, students are introduced to a different test user and provided with an intriguing glimpse of their professional life—this in the form of a Meet an Assessment Professional (MAP) essay. For example, in Chapter 4, students will meet a team of test users, Drs. Steve Julius and Howard Atlas, who have pressed psychometric knowledge into the service of professional sports. They provide a unique and fascinating account of how application of their knowledge of was used to improve the on-court of achievement of the Chicago Bulls. A MAP essay from Stephen Finn, the well-known proponent of therapeutic assessment is presented in Chapter 13. Among the many MAP essays in this edition are essays from two mental-health professionals serving in the military.
coh37025_fm_i-xxviii.indd 20 12/01/21 4:02 PM
Preface xxi
Dr. Alan Ogle introduces readers to aspects of the work of an Air Force psychologist in Chapter 1. In Chapter 11, army psychiatrist Dr.� Rick Malone shares his expertise in the area of threat assessment. The senior author of an oft-cited meta-analysis that was published in Psychological Bulletin shares her insights on meta-analytic methods in Chapter 3, while a psychiatrist who specializes in cultural issues introduces himself to students in Chapter 2.
Our use of the pedagogical tool referred to as a “Close-Up” is reserved for more in-depth and detailed consideration of specific topics related to those under discussion. The Close-Up in our chapter on test construction, for example, acquaints readers with the trials and tribulations of test developers working to create a test to measure asexuality. The Close-Up in one of our chapters on personality assessment raises the intriguing question of whether it is meaningful to speak of general (g) and specific (s) factors in the diagnosis of personality disorders.
There are other pedagogical tools that readers (as well as other textbook authors) may take for granted—but we do not. Consider, in this context, the various tables and figures found in every chapter. In addition to their more traditional use, we view tables as space-saving devices in which a lot of information may be presented. For example, in the first chapter alone, tables are used to provide succinct but meaningful comparisons between the terms testing and assessment, the pros and cons of computer-assisted psychological assessment, and the pros and cons of using various sources of information about tests.
Critical thinking may be defined as “the active employment of judgment capabilities and evaluative skills in the thought process” (Cohen, 1994, p. 12). Generative thinking may be defined as “the goal-oriented intellectual production of new or creative ideas” (Cohen, 1994, p. 13). The exercise of both of these processes, we believe, helps optimize one’s chances for success in the academic world as well as in more applied pursuits. In the early editions of this textbook, questions designed to stimulate critical and generative thinking were raised “the old-fashioned way.” That is, they were right in the text, and usually part of a paragraph. Acting on the advice of reviewers, we made this special feature of our writing even more special beginning with the sixth edition of this book; we raised these critical thinking questions in the margins with a Just Think heading. Perhaps with some encouragement from their instructors, motivated students will, in fact, give thoughtful consideration to these (critical and generative thought-provoking) Just Think questions.
In addition to critical thinking and generative thinking questions called out in the text, other pedagogical aids in this book include original cartoons created by the authors, original illustrations created by the authors (including the model of memory in Chapter 14), and original acronyms created by the authors.3 Each chapter ends with a Self-Assessment feature that students may use to test themselves with respect to key terms and concepts presented in the text.
The tenth edition of Psychological Testing and Assessment is now available online with Connect, McGraw-Hill Education’s integrated assignment and
assessment platform. Connect also offers SmartBook for the new edition, which is the first adaptive reading experience proven to improve grades and help students study more effectively. All of the title’s website and ancillary content is also available through Connect, including:
� An Instructor’s Manual for each chapter. � A full Test Bank of multiple choice questions that test students on central concepts and
ideas in each chapter. � Lecture Slides for instructor use in class.
3. By the way, our use of the French word for black (noir) as an acronym for levels of measurement (nominal, ordinal, interval, and ratio) now appears in other textbooks.
Cohen, R. J. (1994). Psychology & adjustment: Values, culture, and change. Allyn & Bacon.
coh37025_fm_i-xxviii.indd 21 12/01/21 4:02 PM
A�ordable solutions, added value Make technology work for you with LMS integration for single sign-on access, mobile access to the digital textbook, and reports to quickly show you how each of your students is doing. And with our Inclusive Access program you can provide all these tools at a discount to your students. Ask your McGraw Hill representative for more information.
��% Less Time Grading
Checkmark: Jobalou/Getty ImagesPadlock: Jobalou/Getty Images
Instructors: Student Success Starts with You
Laptop: McGraw Hill; Woman/dog: George Doyle/Getty Images
Tools to enhance your unique voice Want to build your own course? No problem. Prefer to use our turnkey, prebuilt course? Easy. Want to make changes throughout the semester? Sure. And you’ll save time with Connect’s auto- grading too.
Solutions for your challenges A product isn’t a solution. Real solutions are a�ordable, reliable, and come with training and ongoing support when you need it and how you want it. Visit www. supportateverystep.com for videos and resources both you and your students can use throughout the semester.
Study made personal Incorporate adaptive study resources like SmartBook® �.� into your course and help your students be better prepared in less time. Learn more about the powerful personalized learning experience available in SmartBook �.� at www.mheducation.com/highered/connect/smartbook
coh37025_fm_i-xxviii.indd 22 12/01/21 4:02 PM
E�ective tools for e�cient studying Connect is designed to make you more productive with simple, �exible, intuitive tools that maximize your study time and meet your individual learning needs. Get learning that works for you with Connect.
Everything you need in one place Your Connect course has everything you need—whether reading on your digital eBook or completing assignments for class, Connect makes it easy to get your work done.
“I really liked this app—it made it easy to study when you don't have your text - book in front of you.”
- Jordan Cunningham, Eastern Washington University
Study anytime, anywhere Download the free ReadAnywhere app and access your online eBook or SmartBook �.� assignments when it’s convenient, even if you’re o�ine. And since the app automatically syncs with your eBook and SmartBook �.� assignments in Connect, all of your work is available every time you open it. Find out more at www.mheducation.com/readanywhere
Top: Jenner Images/Getty Images, Left: Hero Images/Getty Images, Right: Hero Images/Getty Images
Calendar: owattaphotos/Getty Images
Students: Get Learning that Fits You
Learning for everyone McGraw Hill works directly with Accessibility Services Departments and faculty to meet the learning needs of all students. Please contact your Accessibility Services O�ce and ask them to email [email protected], or visit www.mheducation.com/about/accessibility for more information.
coh37025_fm_i-xxviii.indd 23 12/01/21 4:03 PM
xxiv���Preface
Remote Proctoring & Browser-Locking Capabilities
New remote proctoring and browser-locking capabilities, hosted by Proctorio within Connect, provide control of the assessment environment by enabling security options and verifying the identity of the student.
Seamlessly integrated within Connect, these services allow instructors to control students’ assessment experience by restricting browser activity, recording students’ activity, and verifying students are doing their own work.
Instant and detailed reporting gives instructors an at-a-glance view of potential academic integrity concerns, thereby avoiding personal bias and supporting evidence-based claims.
Writing Assignment
Available within McGraw-Hill Connect® and McGraw-Hill Connect® Master, the Writing Assignment tool delivers a learning experience to help students improve their written communication skills and conceptual understanding. As an instructor you can assign, monitor, grade, and provide feedback on writing more efficiently and effectively.
Writing Style
What type of writing style or author voice works best with students being introduced to the field of psychological testing and assessment? Instructors familiar with the many measurement books that have come (and gone) may agree with us that the “voice” of too many authors in this area might best be characterized as humorless and academic to the point of arrogance or pomposity. Students do not tend to respond well to textbooks written in such styles, and their eagerness and willingness to spend study time with these authors (and even their satisfaction with the course as a whole) may easily suffer as a consequence.
In a writing style that could be characterized as somewhat informal and—to the extent possible, given the medium and particular subject being covered—“conversational,” we have made every effort to convey the material to be presented as clearly as humanly possible. In practice, this means:
� keeping the vocabulary of the presentation appropriate (without ever “dumbing-down” or trivializing the material);
� presenting so-called difficult material in step-by-step fashion where appropriate, and always preparing students for its presentation by placing it in an understandable context;
� italicizing the first use of a key word or phrase and then bolding it when a formal definition is given;
� providing a relatively large glossary of terms to which students can refer; � supplementing material where appropriate with visual aids, tables, or other illustrations. � supplementing material where appropriate with intriguing historical facts (as in the
Chapter 12 material on projectives and the projective test created by B. F. Skinner); � incorporating timely, relevant, and intriguing illustrations of assessment-related material
in the text as well as in the online materials.
coh37025_fm_i-xxviii.indd 24 12/01/21 4:03 PM
Preface xxv
In addition, we have interspersed some elements of humor in various forms (original cartoons, illustrations, and vignettes) throughout the text. The judicious use of humor to engage and maintain student interest is something of a novelty among measurement textbooks. Where else would one turn for pedagogy that employs an example involving a bimodal distribution of test scores from a new trade school called The Home Study School of Elvis Presley Impersonators? As readers learn about face validity, they discover why it “gets no respect” and how it has been characterized as “the Rodney Dangerfield of psychometric variables.” Numerous other illustrations could be cited here. But let’s reserve those smiles as a pleasant surprise when readers happen to come upon them.
Acknowledgments
Thanks to the members of the academic community who have wholeheartedly placed their confidence in this book through all or part of its tenth-edition life-cycle to date. Your trust in our ability to help your students navigate the complex world of measurement in psychology is a source of inspiration to us. We appreciate the privilege of assisting you in the education and professional growth of your students, and we will never take that privilege for granted.
Every edition of this book has begun with blueprinting designed with the singular objective of making this book far-and-away best in the field of available textbooks in terms of organization, content, pedagogy, and writing. Helping the authors to meet that objective were developmental editor Erin Guendelsberger and project supervisor Jamie Laferrera along with a number of guest contributors who graciously gave of their time, talent, and expertise. To be the all-around best textbook in a particular subject area takes, as they say, “a village.” On behalf of the authors, a hearty “thank you” is due to many “villagers” in the academic and professional community who wrote or reviewed something for this book, or otherwise contributed to it. First and foremost, thank you to all of the following people who wrote essays designed to enhance and enrich the student experience of the course work. In order of appearance of the tenth edition chapter that their essay appeared in, we say thanks to the following contributors of guest-authored Meet an Assessment Professional, Everyday Psychometrics, or Close-Up:
Alan, D. Ogle of the 559th Medical Group, Military Training Consult Service of the United States Air Force;
Dror Ben-Zeev of the Department of Psychiatry of the Geisel School of Medicine at Dartmouth;
Neil Krishan Aggarwal of the New York State Psychiatric Institute;
Joni L. Mihura of the Department of Psychology of the University of Toledo;
Michael Chmielewski of the Department of Psychology of Southern Methodist University;
Jason M. Chin of the University of Toronto Faculty of Law;
Ilona M. McNeill of the University of Melbourne (Australia);
Patrick D. Dunlop of the University of Western Australia;
Delphine Courvoisier of Beau-Séjour Hospital, Geneva, Switzerland;
Alex Sutherland of RAND Europe, Cambridge, United Kingdom;
Barak Ariel of the Institute of Criminology of the University of Cambridge (United Kingdom);
Lori A. Brotto of the Department of Gynaecology of the University of British Columbia;
Morag Yule of the Department of Gynaecology of the University of British Columbia;
Sivasankaran Balaratnasingam of the School of Psychiatry and Clinical Neurosciences of the University of Western Australia;
Zaza Lyons of the School of Psychiatry and Clinical Neurosciences of the University of Western Australia;
coh37025_fm_i-xxviii.indd 25 12/01/21 4:03 PM
xxvi���Preface
Aleksander Janca of the School of Psychiatry and Clinical Neurosciences of the University of Western Australia;
Yuanbo Gu of the School of Psychology of Shaanxi Normal University (China);
Ning He of the School of Psychology of Shaanxi Normal University (China);
Xuqun You of the School of Psychology of Shaanxi Normal University (China);
Chengting Ju of the School of Psychology of Shaanxi Normal University (China);
Rick Malone of the U.S. Army Criminal Investigation Command, Quantico, VA;
Winnie Eckardt of The Dian Fossey Gorilla Fund International, Atlanta, GA;
Alexander Weiss of the Department of Psychology of the University of Edinburgh (UK);
Monica Webb Hooper of the Case Comprehensive Cancer Center at Case Western Reserve University;
Carla Sharp of the Department of Psychology at the University of Houston (TX);
Liliana B. Sousa of the Faculty of Psychology and Educational Sciences of the University of Coimbra (Portugal);
Duncan M. Shields of the Faculty of Medicine of the University of British Columbia;
Eric Kramer of Medical Specialists of the Palm Beaches, (Neurology), Atlantis, Florida;
Jed Yalof of the Department of Graduate Psychology of Immaculata University;
Mark F. Lenzenweger of the Department of Psychology of the State University of New York at Binghamton;
Jessica Klein of the Department of Psychology of the University of Florida (Gainesville);
Anna Taylor of the Department of Psychology of Illinois State University;
Suzanne Swagerman of the Department of Biological Psychology of Vrije Universiteit (VU), Amsterdam, The Netherlands;
Eco J.C. de Geus of the Department of Biological Psychology of Vrije Universiteit (VU), Amsterdam, The Netherlands;
Kees-Jan Kan of the Department of Biological Psychology of Vrije Universiteit (VU), Amsterdam, The Netherlands;
Dorret I. Boomsma of the Department of Biological Psychology of Vrije Universiteit (VU), Amsterdam, The Netherlands;
Faith Miller of the Department of Educational Psychology of the University of Minnesota; and, Daniel Teichman formerly of the Department of Computer and Information Science and Engineering of the University of Florida (Gainesville).
For their enduring contribution to this and previous editions of this book, we thank Dr. Jennifer Kisamore for her work on the original version of our chapter on test utility, and Dr. Bryce Reeve who wrote a Meet an Assessment Professional essay. Thanks to the many assessment professionals who, whether in a past or the current edition, took the time to introduce students to what they do. For being a potential source of inspiration to the students who they “met” in these pages, we thank the following assessment professionals: Dr. Rebecca Anderson, Dr. Howard W. Atlas, Dr. Scott Birkeland, Dr. Anthony Bram, Dr. Stephen Finn, Dr. Chris Gee, Dr. Joel Goldberg, Ms. Eliane Hack, Dr. Steve Julius, Dr. Nathaniel V. Mohatt, Dr. Barbara C. Pavlo, Dr. Jeanne P. Ryan, Dr. Adam Shoemaker, Dr. Benoit Verdon, Dr. Erik Viirre, and Dr. Eric A. Zillmer. Thanks also to Dr. John Garruto for his informative contribution to Chapter 10.
coh37025_fm_i-xxviii.indd 26 12/01/21 4:03 PM
Preface xxvii
While thanking all who contributed in many varied ways, we remind readers that the present authorship team takes sole responsibility for any possible errors that may have somehow found their way into this tenth edition.
Meet the Authors
Ronald Jay Cohen, Ph.D., ABPP, ABAP, is a Diplomate of the American Board of Professional Psychology in Clinical Psychology, and a Diplomate of the American Board of Assessment Psychology. He is licensed to practice psychology in New York and Florida, and a “scientist-practitioner” and “scholar-professional” in the finest traditions of each of those terms. During a long and gratifying professional career in which he has published numerous journal articles and books, Dr. Cohen has had the privilege of personally working alongside some of the luminaries in the field of psychological assessment, including David Wechsler (while Cohen was a clinical psychology intern at Bellevue Psychiatric Hospital in New York City) and Doug Bray (while working as an assessor for AT&T in its Management Progress Study). After serving his clinical psychology internship at Bellevue, Dr. Cohen was appointed Senior Psychologist there, and his clinical duties entailed not only psychological assessment but the supervision and training of others in this enterprise. Subsequently, as an independent practitioner in the New York City area, Dr. Cohen taught various courses at local universities on an adjunct basis, including undergraduate and graduate courses in psychological assessment. Asked by a colleague to conduct a qualitative research study for an advertising agency, Dr.�Cohen would quickly become a sought-after qualitative research consultant with a client list of major companies and organizations—among them Paramount Pictures, Columbia Pictures, NBC Television, the Campbell Soup Company, Educational Testing Service, and the College Board. Dr. Cohen’s approach to qualitative research, referred to by him as dimensional qualitative research, has been emulated and written about by qualitative researchers around the world. Dr.�Cohen is a sought-after speaker and has delivered invited addresses at the Sorbonne in Paris, Peking University in Beijing, and numerous other universities throughout the world. It was Dr. Cohen’s work in the area of qualitative assessment that led him to found the scholarly journal Psychology & Marketing. Since the publication of the journal’s first issue in 1984, Dr.�Cohen has served as its Editor-in-Chief.
W. Joel Schneider, Ph.D., is Associate Professor of Counseling and School Psychology in the Department of Psychological Studies in Education at Temple University in Philadelphia. He completed his doctorate in clinical psychology at Texas A&M University. Dr. Schneider spent 15 years on the faculty at Illinois State University before joining the faculty at Temple. He is lead author of the second edition of the 2018 book, Essentials of Assessment Report Writing. He regularly teaches graduate-level courses in assessment and oversees clinic-based practicum students. Dr. Schneider has participated as an examiner in the standardization of several psychological tests, including the RIAS, CASE, and CASE-R. His primary research interests are psychological assessment of cognitive abilities and personality, psychometrics, statistics, and research methods, and psychotherapy with individuals, groups, couples, and families. Dr. Schneider has been involved in a number of research projects funded by the Pennsylvania Department of Education and the U.S. Department of Education. His work may be found in the Archives of Clinical Neuropsychology, Psychological Methods, Journal of Psychoeducational Assessment, Applied Neuropsychology, and Best Practices in School Psychology. He served as a test reviewer for the Mental Measurements Yearbook and currently serves as an Associate Editor for Journal of Psychoeducational Assessment. He is also an editorial board member for Journal of Intelligence and Journal of School Psychology.
coh37025_fm_i-xxviii.indd 27 12/01/21 4:03 PM
xxviii���Preface
Renée M. Tobin, Ph.D., is Professor and Chair of the Department of Psychological Studies in Education at Temple University in Philadelphia. She completed her master’s degree in social psychology and her doctorate in school psychology at Texas A&M University. Dr. Tobin spent 15 years on the faculty at Illinois State University before joining the faculty at Temple. She is co-author of the 2015 book, DSM-5 Diagnosis in the Schools. She regularly teaches graduate- level courses in assessment, counseling, and consultation. She served as an examiner for the standardization of several psychological tests, including the RIAS, CASE, and CASE-R. Her primary research interests center broadly on personality and social development. Dr. Tobin has extensive experience conducting mixed-methods research with children, adolescents, young adults, and their families, particularly among diverse populations in various school contexts. She has been involved in a number of program evaluation projects since 2010, which include serving as co-leader of the evaluation team for the Livingston County Children’s Network (funded by the Illinois Children’s Healthcare Foundation) and as coordinator of the continuous quality improvement team for the Champaign Area Relationship Education for Youth (CARE4U) grant program (funded by the U.S. Department of Health and Human Services). Her work may be found in the Journal of Personality and Social Psychology, Journal of Personality, Psychological Science, School Psychology Quarterly, and Best Practices in School Psychology. She served as a test reviewer for the Mental Measurements Yearbook and served as an Associate Editor for Journal of Psychoeducational Assessment for over 10 years. She is currently an editorial board member for Journal of School Psychology.
And on a Personal Note . . .
I think back to the time when we were just wrapping up work on the sixth edition of this book. At that time, I received the unexpected and most painful news that my mother had suffered a massive and fatal stroke. It is impossible to express the sense of sadness and loss experienced by myself, my brother, and my sister, as well as the countless other people who knew this gentle, loving, and much-loved person. To this day, we continue to miss her counsel, her sense of humor, and just knowing that she’s there for us. We continue to miss her genuine exhilaration, which in turn exhilarated us, and the image of her welcoming, outstretched arms whenever we came to visit. Her children were her life, and the memory of her smiling face, making each of us feel so special, survives as a private source of peace and comfort for us all. She always kept a copy of this book proudly displayed on her coffee table, and I am very sorry that a copy of more recent editions did not make it to that most special place. My dedication of this book is one small way I can meaningfully acknowledge her contribution, as well as that of my beloved, deceased father, to my personal growth. As in the sixth edition, I am using my parents’ wedding photo in the dedication. They were so good together in life. And so there Mom is, reunited with Dad. Now, that is something that would make her very happy.
As the reader might imagine, given the depth and breadth of the material covered in this textbook, it requires great diligence and effort to create and periodically re-create an instructional tool such as this that is timely, informative, and readable. Thank you, again, to all of the people who have helped through the years. Of course, I could not do it myself were it not for the fact that even through ten editions, this truly Herculean undertaking remains a labor of love.
Ronald Jay Cohen, Ph.D., ABPP, ABAP Diplomate, American Board of Professional Psychology (Clinical) Diplomate, American Board of Assessment Psychology
coh37025_fm_i-xxviii.indd 28 12/01/21 4:03 PM
�
C H A P T E R �
Psychological Testing and Assessment
ll fields of human endeavor use measurement in some form, and each field has its own set of measuring tools and measuring units. For example, you become aware of unique measurement units when making major purchases. When buying a new smartphone or computer, measurements of speed (e.g., gigahertz), screen resolution (e.g., 12 megapixels), and storage (e.g., 512 gigabytes) are salient, whereas the 4 Cs (i.e., cut, color, clarity, and carat) become relevant measurement terms when considering a marriage proposal. You also witnessed the worldwide importance of developing faster measurement tools to identify asymptomatic virus carriers during the COVID-19 pandemic. As a student of psychological measurement, you need a working familiarity with some of the commonly used units of measure in psychology as well as knowledge of some of the many measuring tools employed. In the pages that follow, you will gain that knowledge as well as an acquaintance with the history of measurement in psychology and an understanding of its theoretical basis.
Good helpers take time to understand the situation before helping a person. Great helpers make time to understand the person who needs help. Psychological assessment applies scientific rigor to the gentle art of understanding people before helping them. Psychological assessment encompasses a wide variety of methods, including direct observation, interviews, questionnaires, tests, and case file reviews.
Tests have been used by educators since ancient times, but psychological tests were developed only after psychology emerged as a formal scientific discipline in the late 1800s. Whereas educational testing tells us how much a person has learned, psychological assessment tells us what can be learned about a person. The experience of being closely listened to and deeply understood is itself a great comfort to many individuals who have sought the help of psychological assessment providers.
Testing and Assessment
The roots of contemporary psychological testing and assessment can be found in early twentieth- century France. In 1905, Alfred Binet and a colleague published a test designed to help place Paris schoolchildren in appropriate classes. The first society-wide application of psychological testing resulted from an attempt by Parisian educators and lawmakers to live up to the ideals inscribed on public buildings all over France: liberté, égalité, fraternité (liberty, equality, fraternity). In a series of sweeping educational reforms in the 1870s–1890s, France became one of the first countries to mandate free public education for all its children. Of course, mandating high-quality education for everyone is not the same as educating everyone equally well. Not long after the laws went into effect, French educational institutions were confronted with the full
A
coh37025_ch01_001-040.indd 1 12/01/21 4:03 PM
����Part 1: An Overview
magnitude of human diversity. Children with intellectual disabilities need higher levels of support. In previous generations, children with intellectual disabilities were given intensive education only if their families could pay for such services. No longer.
How does one meet the complex educational needs of students with the severest of disabilities while also treating students equally? French educational administrators wanted an efficient, accurate, and fair method of deciding which children were best served by learning in separate, special classes with slower, more intensive instruction. The Minister of Public Instruction commissioned a study of the matter, and the committee asked Alfred Binet and his colleague Theodore Simon to create a test that would help school personnel make placement decisions. Binet and Simon warned that without objective scientific rigor, decisions are made haphazardly, “which are subjective, and consequently uncontrolled. […] Some errors are excusable in the beginning, but if they become too frequent, they may ruin the reputations of these new [public school] institutions” (Binet & Simon, 1905, pp.�11–12).
Binet and Simon created a series of tests designed to forecast which students would likely fall ever further behind their peers without additional support. Although the Binet– Simon test became known as an “intelligence test,” its designers specifically warned that the test did not measure intelligence in its totality. Rather, the test was designed for the narrow purpose of identifying intellectually disabled children who needed additional help. Subsequent research found that the tests achieved their stated design goals reasonably well. Binet’s test would have consequences well beyond the Paris school district. Within a decade an English- language version of Binet’s test was prepared for use in schools in the United States. When the United States declared war on Germany and entered World War I in 1917, the military needed a way to screen large numbers of recruits quickly for intellectual and emotional problems. Psychological testing provided this methodology. During World War II, the military would depend even more on psychological tests to screen recruits for service. Following the war, more and more tests purporting to measure an ever-widening array of psychological variables were developed and used. There were tests to measure not only intelligence but also personality, brain functioning, performance at work, and many other aspects of psychological and social functioning.
William Stern, who developed a refined method of scoring Binet’s test—the Intelligence Quotient (IQ)—was horrified when Binet’s tests were later used by many institutions as tools of oppression rather than for their original purpose of liberation. He wrote movingly about how IQ tests should not be used to degrade individuals (Stern, 1933, as translated by Lamiell, 2003):
Under all conditions, human beings are and remain the centers of their own psychological life and their own worth. In other words, they remain persons, even when they are studied and treated from an external perspective with respect to others’ goals. ... Working “on” a human being must always entail working “for” a human being. (pp.�54–55)
We adopt Stern’s ideals and share his vision that with proper ethical safeguards, psychological tests can fulfill their original purpose—helping individuals and creating a more just society for everyone.
Psychological Testing and Assessment Defined
The world’s receptivity to Binet’s test in the early twentieth century spawned not only more tests but more test developers, more test publishers, more test users, and the emergence of what, logically enough, has become known as a testing enterprise. “Testing” was the term used to refer to everything from the administration of a test (as in “Testing in progress”) to the interpretation of a test score (“The testing indicated that� .� .� .”). During World War I, the term “testing” aptly described the group screening of thousands of military recruits. We suspect that it was then that the term gained a powerful foothold in the vocabulary of
coh37025_ch01_001-040.indd 2 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment �
professionals and laypeople. The use of “testing” to denote everything from test administration to test interpretation can be found in postwar textbooks (such as Chapman, 1921; Hull, 1922; Spearman, 1927) as well as in various test-related writings for decades thereafter. However, by World War II a semantic distinction between testing and a more inclusive term, “assessment,” began to emerge.
Military, clinical, educational, and business settings are but a few of the many contexts that entail behavioral observation and active integration by assessors of test scores and other data. In such situations, the term assessment may be preferable to testing. In contrast to testing, assessment acknowledges that tests are only one type of tool used by professional assessors (along with other tools, such as the interview), and that the value of a test, or of any other tool of assessment, is intimately linked to the knowledge, skill, and experience of the assessor.
The semantic distinction between psychological testing and psychological assessment is blurred in everyday conversation. Somewhat surprisingly, the distinction between the two terms still remains blurred in some published “psychological testing” textbooks. Yet the distinction is important. Society at large is best served by a clear definition of and differentiation between these two terms as well as related terms such as psychological test user and psychological assessor. Clear distinctions between such terms may also help avoid the turf wars now brewing between psychology professionals and members of other professions seeking to use various psychological tests. In many psychological evaluation contexts, conducting an assessment requires greater education, training, and skill than simply administering a test.
We define psychological assessment as the gathering and integration of psychology-related data for the purpose of making a psychological evaluation that is accomplished through the use of tools such as tests, interviews, case studies, behavioral observation, and specially designed apparatuses and measurement procedures. We define psychological testing as the process of measuring psychology-related variables by means of devices or procedures designed to obtain a sample of behavior. Some of the differences between these two processes are presented in Table 1–1.1
Varieties of assessment�The term assessment may be modified in a seemingly endless number of ways, each such modification referring to a particular variety or area of assessment. Sometimes the meaning of the specialty area can be readily discerned just from the word or term that modifies “assessment.” For example, the term “therapeutic psychological assessment” refers to assessment that helps individuals understand and solve their problems. Also intuitively obvious, the term educational assessment refers to, broadly speaking, the use of tests and other tools to evaluate abilities and skills relevant to success or failure in a school or pre-school context. Intelligence tests, achievement tests, and reading comprehension tests are some of the evaluative tools that may spring to mind with the mention of the term “educational assessment.” But what springs to mind with the mention of other, less common assessment terminology? Consider, for example, terms like retrospective assessment, remote assessment, and ecological momentary assessment.
J U S T T H I N K � . � . � .
Describe a situation in which testing is more appropriate than assessment. By contrast, describe a situation in which assessment is more appropriate than testing.
1. Especially when discussing general principles related to the creation of measurement procedures, as well as the creation, manipulation, or interpretation of data generated from such procedures, the word test (as well as related terms, such as test score) may be used in the broadest and most generic sense; that is, “test” may be used in shorthand fashion to apply to almost any procedure that entails measurement (including, e.g., situational performance measures). Accordingly, when we speak of “test development” in Chapter 8, many of the principles set forth will apply to the development of other measurements that are not, strictly speaking, “tests” (such as situational performance measures, as well as other tools of assessment). Having said that, let’s reemphasize that a real and meaningful distinction exists between the terms psychological testing and psychological assessment, and that effort should continually be made not to confuse the meaning of these two terms.
coh37025_ch01_001-040.indd 3 12/01/21 4:03 PM
����Part 1: An Overview
For the record, the term retrospective assessment is defined as the use of evaluative tools to draw conclusions about psychological aspects of a person as they existed at some point in time prior to the assessment. There are unique challenges and hurdles to be overcome when conducting retrospective assessments regardless if the subject of the evaluation is alive (Teel et al., 2016) or is deceased (Reyman & Shankar, 2015). Remote assessment refers to the use of tools of psychological evaluation to gather data and draw conclusions about a subject who is not in physical proximity to the person or people conducting the evaluation. One example of how psychological assessments may be conducted remotely was provided in this chapter’s Close-Up feature. In each chapter of this book, we will spotlight one topic for “a closer look.”
Table �–� Testing in Contrast to Assessment
In contrast to the process of administering, scoring, and interpreting psychological tests (psychological test- ing), psychological assessment is a problem-solving process that can take many different forms. How psychological assessment proceeds depends on many factors, not the least of which is the reason for assessing. Different tools of evaluation—psychological tests among them—might be marshaled in the process of assessment, depending on the particular objectives, people, and circumstances involved as well as on other variables unique to the particular situation. Admittedly, the line between what constitutes testing and what constitutes assessment is not always as clear as we might like it to be. However, by acknowledging that such ambiguity exists, we can work to sharpen our definition and use of these terms. It seems useful to distinguish the differences between testing and assessment in terms of the typical objective, process, and outcome of an evaluation and also in terms of the role and skill of the evaluator. Keep in mind that, although these are useful distinctions to con- sider, exceptions can always be found.
Testing Assessment
Objective
To obtain some gauge, usually numerical in nature, with regard to an ability or attribute.
To answer a referral question, solve a problem, or arrive at a decision through the use of tools of evaluation.
Process
Testing may be conducted individually or in groups. After test administration, the tester adds up “the number of correct answers or the number of certain types of responses�.�.� . with little if any regard for the�how or mechanics of such content” (Maloney & Ward, ����, p.���).
Assessment is individualized. In contrast to testing, assessment focuses on how an individual processes rather than simply the results of that processing.
Role of Evaluator
The tester is not key to the process; one tester may be substituted for another tester without appreciably a�ecting the evaluation.
The assessor is key to the process of selecting tests and/or other tools of evaluation as well as in drawing conclusions from the entire evaluation.
Skill of Evaluator
Testing requires technician-like skills in administering and scoring a test as well as in interpreting a test result.
Assessment requires an educated selection of tools of evaluation, skill in evaluation, and thoughtful organization and integration of data.
Outcome
Testing yields a test score or series of test scores. Assessment entails a logical problem-solving approach that brings to bear many sources of data designed to shed light on a referral question.
coh37025_ch01_001-040.indd 4 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment �
positioning system (GPS). When the user is outdoors, the GPS generates geospatial coordinates helpful in determining the daily distance covered, as well as the amount of time spent at speci�c locations. When the research team conducts studies with individuals who do not move from one location to another, such as hospitalized patients in closed psychiatric units, they place microbluetooth beacons in di�erent rooms throughout the venue. As the subject moves from one room to another, the smartphone’s bluetooth sensor receives signals sent by the beacons, and records the subject’s precise position in the unit.
A typical smartphone also comes equipped with accelerometers; these devices are designed to detect motion. Ben-Zeev’s monitoring system collects the accelerometer data to�determine whether the individual is or is not active.
The smartphone system collects and stores all of the sensor data and transmits it periodically to a secure study server. There,�the information is processed and displayed on a digital dashboard. By means of this system, multidimensional data from�faraway places can be viewed online to help clinicians and�researchers better understand experiences that cause changes in stress level and general mental health. One smartphone-sensing study conducted with college undergraduate and graduate student subjects over a ��-week period included pre- and post-measures of depression. The data suggested that social engagement (as measured by the speech detection software) and daily geospatial activity (as measured by GPS) were significantly related to changes in level of depression (Ben-Zeev et al., ����a).
C L O S E � U P
Behavioral Assessment Using Smartphones*
uch like the state of one’s physical health, the state of one’s mental health and functioning is changing and �uid. Varied internal factors (such as neurochemistry and hormonal shifts), external factors (such as marital discord and job pressures), or�combinations thereof may a�ect mental health and functioning. This �uctuation is as true for people with no diagnosis of mental disorder as it is for patients su�ering from chronic psychiatric illnesses.
Changes in people’s mental health status rarely come “out of the blue” (or, without warning). Behavioral signs that someone is experiencing increased stress and mental health di�culties may include changes in sleep and eating patterns, social engagement, and physical activity. Because these changes may emerge gradually over time, they can go unnoticed by family members, close friends, or even the a�ected individuals themselves. By the time most people seek support or professional care, their mental health and functioning may have deteriorated substantially. Identifying behavioral patterns that are associated with increased risk for underlying mental health di�culties is a �rst step toward more e�cient treatment, perhaps even prevention.
Dr. Dror Ben-Zeev and his colleagues have begun to identify problematic behavioral patterns using a device that is�already in the hands of billions: the smartphone. The smartphone (or, a mobile phone that features computational capacity) comes equipped with multiple embedded sensors that measure variables such as acoustics, location, and movement. Ben- Zeev’s team uses sophisticated smartphone software that enables them to repurpose these sensors and capture an abundance of information about the smartphone user’s environment and behavior. Their program activates the smartphone’s microphone every few minutes to capture ambient sound. If the software detects human conversation, it remains active for the duration of the conversation. To protect user’s privacy, the speech detection system does not record raw audio. It processes the data in real-time to extract and store conversation-related data while actual conversations cannot be reconstructed. The software calculates both the number of conversations and the average length of a conversation engaged in during a ��-hour period.
In addition to re-purposing the microphone in a cell phone, Ben-Zeev’s system repurposes the smartphone’s global
M
(continued)
*This Close-Up was written by Dror Ben-Zeev of the Department of Psychiatry of the Geisel School of Medicine at Dartmouth.
GaudiLab/Shutterstock
coh37025_ch01_001-040.indd 5 12/01/21 4:03 PM
����Part 1: An Overview
(Ben-Zeev et al., ����b). Patients and mental health professionals alike appreciate the promise of this potentially useful method for detecting emerging high-risk patterns that require preventative or immediate treatment.
As technology evolves, one can imagine a future in which at-risk individuals derive bene�t from smartphones repurposed to serve as objectively scalable measures of behavior (Ben-Zeev, ����). Used in a clinically skilled fashion and with appropriate protections of patient privacy, these ubiquitous devices, now repurposed to yield behavioral data, may be instrumental in creating meaningful diagnostic insights and pro�les. In turn, such minute-to-minute assessment data may yield highly personalized and e�ective treatment protocols.
Used with permission of Dror Ben-Zeev.
Of course, tracking someone via their smartphone without their awareness and consent would be unethical. However, for people who may be at risk for mental health problems, or for those who already struggle with psychiatric conditions and need support, this unobtrusive approach may have value. Explaining to patients (or their representatives) what the technology is, how it works, and how data from it may be used for patient bene�t, may well allay any privacy concerns. Preliminary research has suggested that even patients with severe mental illness can understand and appreciate the potential bene�ts of remote assessment by means of the smartphone tracking system (Ben-Zeev et al., ����). Most of�the subjects studied stated that they would have no objection to�using a system that could not only passively detect when they were not doing well, but o�er them helpful and timely suggestions for improving their mental state
C L O S E � U P
Behavioral Assessment Using Smartphones (continued)
In this chapter, the Close-Up box explored how the smartphone revolution in communication may also signal a revolution in the way that psychological assessments are conducted.
Psychological assessment by means of smartphones also serves as an example of an approach to assessment called ecological momentary assessment (EMA). EMA refers to the “in the moment” evaluation of specific problems and related cognitive and behavioral variables at the exact time and place that they occur. Using various tools of assessment, EMA has been used to help tackle diverse clinical problems including post-traumatic stress disorder (Black et�al., 2016), problematic smoking (Ruscio et al., 2016), chronic abdominal pain in children (Schurman & Friesen, 2015), and attention-deficit/hyperactivity symptoms (Li & Lansford,�2018).
The process of assessment� In general, the process of assessment begins with a referral for assessment from a source such as a teacher, parent, school psychologist, counselor, judge, clinician, or corporate human resources specialist. Typically one or more referral questions are put to the assessor about the assessee. Some examples of referral questions are: “Can this child function in a general education environment?,” “Is this defendant competent to stand trial?,” and “How well can this employee be expected to perform if promoted to an executive position?”
The assessor may meet with the assessee or others before the formal assessment in order to clarify aspects of the reason for referral. The assessor prepares for the assessment by selecting the tools of assessment to be used. For example, if the assessment occurs in a corporate or
military setting and the referral question concerns the assessee’s leadership ability, the assessor may wish to employ a measure (or two) of leadership. Typically, the assessor’s own past experience, education, and training play a key role in the specific tests or other tools to be employed in the assessment. Sometimes an institution
J U S T T H I N K � . � . � .
What qualities makes a good leader? How might these qualities be measured?
coh37025_ch01_001-040.indd 6 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment �
in which the assessment is taking place has prescribed guidelines for which instruments can and cannot be used. In almost every assessment situation, particularly situations that are relatively novel to the assessor, the tool selection process is informed by some research in preparation for the assessment. For example, in the assessment of leadership, the tool selection procedure might be informed by reviewing publications dealing with behavioral studies of leadership (Derue et al., 2011), psychological studies of leaders (Kouzes�& Posner, 2007), cultural issues in leadership (Byrne & Bradley, 2007), or whatever aspect of leadership on which the assessment will be focused (Carnevale et al., 2011; Elliott, 2011; Rosenman et al., 2015).
Subsequent to the selection of the instruments or procedures to be employed, the formal assessment will begin. After the assessment, the assessor writes a report of the findings that is designed to answer the referral question. More feedback sessions with the assessee and/or interested third parties (such as the assessee’s parents and the referring professional) may also be scheduled.
Different assessors may approach the assessment task in different ways. Some assessors approach the assessment with minimal input from assessees themselves. Other assessors view the process of assessment as more of a collaboration between the assessor and the assessee. For example, in one approach to assessment, referred to (logically enough) as collaborative psychological assessment, the assessor and assessee may work as “partners” from initial contact through final feedback (Finello, 2011; Fischer, 1978, 2004, 2006). The assessment provider encourages collaboration by asking questions like, “After this assessment is finished, what would you like to know that you do not know already?” One variety of collaborative assessment includes an element of therapy as part of the process. Stephen Finn and his colleagues (Finn, 2003, 2009, 2011; Finn & Martin, 1997; Finn & Tonsager, 2002; Fischer & Finn, 2014) have described a collaborative approach to assessment called therapeutic psychological assessment. In traditional psychological evaluations, the assessment is designed to have its intended benefits at the end of the process: The examiner explains the results, summarizes the case conceptualization, and shares a list of recommendations designed to help the examinee.
In contrast, therapeutic psychological assessment aims to be helpful throughout the assessment process. The results are not revealed at the end, but shared immediately so that both the assessor and the assessee can co-develop an interpretation of the results and decide what questions require further assessment. In this way, therapeutic self-discovery and new understandings are encouraged throughout the assessment process.
Another approach to assessment that seems to have picked up momentum in recent years, most notably in educational settings, is referred to as dynamic assessment (Poehner & van Compernolle, 2011). The term dynamic may suggest that a psychodynamic or psychoanalytic approach to assessment is being applied, but that is not the case. As used in the present context, dynamic is used to describe the interactive, changing, or varying nature of the assessment. In general, dynamic assessment refers to an interactive approach to psychological assessment that usually follows a model of (1) evaluation, (2) intervention of some sort, and (3) evaluation. Dynamic assessment is most typically employed in educational settings, although it may be employed in correctional, corporate, neuropsychological, clinical, and most any other setting as well.
Intervention between evaluations, sometimes even between individual questions posed or tasks given, might take many different forms, depending upon the purpose of the dynamic assessment (Haywood & Lidz, 2007). For example, an assessor may intervene in the course of an evaluation of an assessee’s abilities with increasingly more explicit feedback or hints. The purpose of the intervention may be to provide assistance with mastering the task at hand. Progress in mastering the same or similar tasks is then measured. In essence, dynamic assessment provides a means for evaluating how the assessee processes or benefits from some type of intervention (feedback, hints, instruction, therapy, and so forth) during the course of evaluation. In some educational contexts, dynamic assessment may be viewed as a way of measuring not just learning but “learning potential,” or “learning how to learn” skills. Computers are one tool used to help meet the objectives of dynamic assessment (Wang, 2011). There are others�.� .� .
coh37025_ch01_001-040.indd 7 12/01/21 4:03 PM
����Part 1: An Overview
The Tools of Psychological Assessment
The Test
A test is defined simply as a measuring device or procedure. When the word test is prefaced with a modifier, it refers to a device or procedure designed to measure a variable related to that modifier. Consider, for example, the term medical test, which refers to a device or procedure designed to measure some variable related to the practice of medicine (including a wide range of tools and procedures, such as X-rays, blood tests, and testing of reflexes). In a like manner, the term psychological test refers to a device or procedure designed to measure variables related to psychology (such as intelligence, personality, aptitude, interests, attitudes, or values). Whereas a medical test might involve analysis of a sample of blood, tissue, or the like, a psychological test almost always involves analysis of a sample of behavior. The behavior sample could range from responses to a pencil-and-paper questionnaire, to verbal responses to questions related to the performance of some task. The behavior sample could be elicited by the stimulus of the test itself, or it could be naturally occurring behavior (observed by the assessor in real time as it occurs, or it can be recorded and observed at a later time).
Psychological tests and other tools of assessment may differ with respect to a number of variables, such as content, format, administration procedures, scoring and interpretation procedures, and technical quality. The content (subject matter) of the test will, of course, vary with the focus of the particular test. But even two psychological tests purporting to measure the same thing—for example, personality—may differ widely in item content. This difference is, in part, because two test developers might have entirely different views regarding what is important in measuring “personality”; different test developers employ different definitions of “personality.” Additionally, different test developers come to the test development process with different theoretical orientations. For example, items on a psychoanalytically oriented personality test may have little resemblance to those on a behaviorally oriented personality test, yet both are personality tests. A psychoanalytically oriented personality test might be chosen for use by a psychoanalytically oriented assessor, and an existentially oriented personality test might be chosen for use by an existentially oriented assessor.
The term format pertains to the form, plan, structure, arrangement, and layout of test items as well as to related considerations such as time limits. Format is also used to refer to the form in which a test is administered: computerized, pencil-and- paper, or some other form. When making specific reference to a computerized test, the format may also involve the form of the software: local or online/cloud-based software and storage. The term format is not confined to tests. Format is also used to denote the form or structure of other evaluative tools and processes, such as the guidelines for creating a portfolio work sample.
Tests differ in their administration procedures. Some tests, particularly those designed for administration on a one-to-one basis, may require an active and knowledgeable test administrator. The test administration may involve demonstration of various kinds of tasks demanded of the assessee, as well as trained observation of an assessee’s performance. Alternatively, some tests, particularly those designed for administration to groups, may not even require the test administrator to be present while the testtakers independently complete the required tasks.
Tests differ in their scoring and interpretation procedures. To better understand how and why, let’s define score and scoring. Sports enthusiasts are no strangers to these terms. For them, these terms refer to the number of points accumulated by competitors and the process of accumulating those points. In testing and assessment, we formally define score as a code
J U S T T H I N K � . � . � .
Imagine you wanted to develop a test for a personality trait you termed “goth.” How would you de�ne this trait? What kinds of items would you include in the test? Why would you include those kinds of items? How would you distinguish this personality trait from others?
coh37025_ch01_001-040.indd 8 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment �
or summary statement, usually but not necessarily numerical in nature, that reflects an evaluation of performance on a test, task, interview, or some other sample of behavior. Scoring is the process of assigning such evaluative codes or statements to performance on tests, tasks, interviews, or other behavior samples. In the world of psychological assessment, many different types of scores exist. Some scores result from the simple summing of responses (such as the summing of correct/incorrect or agree/disagree responses), and some scores are derived from more elaborate procedures.
Scores themselves can be described and categorized in many different ways. For example, one type of score is the cut score. A cut score (also referred to as a cutoff score or simply a cutoff) is a reference point, usually numerical, derived by judgment and used to divide a set of data into two or more classifications. Some action will be taken or some inference will be made on the basis of these classifications. Cut scores on tests, usually in combination with other data, are used in schools in many contexts. For example, they may be used in grading, and in making decisions about the class or program to which children will be assigned. Cut scores are used by employers as aids to decision making about personnel hiring, placement, and advancement. State agencies use cut scores as aids in licensing decisions. There are probably more than a dozen different methods that can be used to formally derive cut scores (Dwyer, 1996). If you’re curious about what some of those different methods are, stay tuned; we cover that in an upcoming chapter.
Sometimes no formal method is used to arrive at a cut score. Some teachers use an informal “eyeball” method to proclaim, for example, that a score of 65 or more on a test means “pass” and a score of 64 or below means “fail.” Whether formally or informally derived, cut scores typically take into account, at least to some degree, the values of those who set them. Consider, for example, two professors who teach the same course at the same college. One professor might set a cut score for passing the course that is significantly higher (and more difficult for students to attain) than the other professor. There is also another side to the human equation as it relates to cut scores, one that is seldom written about in measurement texts. This phenomenon concerns the emotional consequences of “not making the cut” and “just making the cut” (see Figure 1–1).
Tests differ widely in terms of their guidelines for scoring and interpretation. Some tests are self-scored by the testtakers themselves, others are scored by computer, and others require scoring by trained examiners. Some tests, such as most tests of intelligence, come with test manuals that are explicit not only about scoring criteria but also about the nature of the interpretations that can be made from the scores. Other tests, such as the Rorschach Inkblot Test, are sold with no manual at all. The (presumably qualified) purchaser buys the stimulus materials and then selects and uses one of many available guides for administration, scoring, and interpretation.
Tests differ with respect to their psychometric soundness or technical quality. Synonymous with the antiquated term psychometry, psychometrics is defined as the science of psychological measurement. Variants of these words include the adjective psychometric (which refers to measurement that is psychological in nature) and the nouns psychometrist and psychometrician (both terms referring to a professional who uses, analyzes, and interprets psychological test data). One speaks of the psychometric soundness of a test when referring to how consistently and how accurately a psychological test measures what it purports to measure. Assessment professionals also speak of the psychometric utility of a particular test or assessment method. In this context, utility� refers to the usefulness or practical value that a test or other tool of assessment has for a particular purpose. These concepts are elaborated on in subsequent chapters. Now, returning to our discussion of tools of assessment, meet one well-known tool that, as they say, “needs no introduction.”
J U S T T H I N K � . � . � .
How might one test of intelligence have more utility than another test of intelligence in the same school setting?
coh37025_ch01_001-040.indd 9 12/01/21 4:03 PM
�����Part 1: An Overview
The Interview
In everyday conversation, the word interview conjures images of face-to-face talk. But the interview as a tool of psychological assessment typically involves more than talk. If the interview is conducted face-to-face, then the interviewer is probably taking note of not only the content of what is said but also the way it is being said. More specifically, the interviewer is taking note of both verbal and nonverbal behavior. Nonverbal behavior may include the interviewee’s “body language,” movements, and facial expressions in response to the interviewer, the extent of eye contact, apparent willingness to cooperate, and general reaction to the demands of the interview. The interviewer may also take note of the way the interviewee is dressed. Here, variables such as neat versus sloppy, and appropriate versus inappropriate, may be noted.
Because of a potential wealth of nonverbal information to be gained, interviews are ideally conducted face-to-face. However, face-to-face contact is not always possible and interviews may be conducted in other formats. In an interview conducted by telephone, for example, the interviewer may still be able to gain information beyond the responses to questions by being sensitive to variables such as changes in the interviewee’s voice pitch or the extent to which
Figure �–� Emotion engendered by categorical cuto�s.
People who just make some categorical cutoff may feel better about their accomplishment than those who make the cutoff by a substantial margin. But those who just miss the cutoff may feel worse than those who miss it by a substantial margin. Evidence consistent with this view was presented in research with Olympic athletes (Medvec et al., 1995; Medvec & Savitsky, 1997). Bronze medalists were—somewhat paradoxically—happier with the outcome than silver medalists. Bronze medalists might say to themselves “at least I won a medal” and be happy about it. By contrast, silver medalists might feel frustrated that they tried for the gold and missed winning it. Jean Catuffe/Getty Images
coh37025_ch01_001-040.indd 10 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
particular questions precipitate long pauses or signs of emotion in response. Of course, interviews need not involve verbalized speech, as when they are conducted in sign language. Interviews may also be conducted by various electronic means, as would be the case with online interviews, e-mail interviews, and interviews conducted by means of text messaging. In its broadest sense, then, we can define an interview as a method of gathering information through direct communication involving reciprocal exchange.
Interviews differ with regard to many variables, such as their purpose, length, and nature. Interviews may be used by psychologists in various specialty areas to help make diagnostic, treatment, selection, or other decisions. So, for example, school psychologists may use an interview to help make a decision about the appropriateness of various educational interventions or class placements. A court-appointed psychologist may use an interview to help guide the court in determining whether a defendant was insane at the time of a commission of a crime. A specialist in head injury may use an interview to help shed light on questions related to the extent of damage to the brain that was caused by the injury. A psychologist studying consumer behavior may use an interview to learn about the market for various products and services, as well as how best to advertise and promote them. A police psychologist may instruct eyewitnesses to serious crimes to close their eyes when they are interviewed about details related to the crime. They do so because there is suggestive evidence that the responses will have greater relevance to the questions posed if the witness’s eyes are closed (Vredeveldt et al., 2015).
An interview may be used to help professionals in human resources to make more informed recommendations about the hiring, firing, and advancement of personnel. In some instances, what is called a panel interview (also referred to as a board interview) is employed. Here, more than one interviewer participates in the assessment. A presumed advantage of this personnel assessment technique is that any idiosyncratic biases of a lone interviewer will be minimized (Dipboye, 1992). A disadvantage of the panel interview relates to its utility; the cost of using multiple interviewers may not be justified (Dixon et al., 2002).
Some interviewing, especially in the context of clinical and counseling settings, has as its objective not only the gathering of information from the interviewee, but a targeted change in� the interviewee’s thinking and behavior. A therapeutic technique called motivational interviewing, for example, is used by counselors and clinicians to gather information about some problematic behavior, while simultaneously attempting to address it therapeutically (Bundy, 2004; Miller & Rollnick, 2002, 2012). Motivational interviewing may be defined as a therapeutic dialogue that combines person-centered listening skills such as openness and empathy, with the use of cognition-altering techniques designed to positively affect motivation and effect therapeutic change. Motivational interviewing has been employed to address a relatively wide range of problems (Hoy et al., 2016; Kistenmacher & Weiss, 2008; Miller & Rollnick, 2009; Pollak et� al., 2016; Rothman & Wang, 2016; Shepard et al., 2016) and has been successfully employed in intervention by means of telephone (Lin et al., 2016), Internet chat (Skov-Ettrup et al., 2016), and text messaging (Shingleton et al., 2016).
The popularity of the interview as a method of gathering information extends far beyond psychology. Just try to think of one day when you were not exposed to an interview on television, radio, or the Internet! Regardless of the medium through which it is conducted, an interview is a reciprocal affair in that the interviewee reacts to the interviewer and the interviewer reacts to the interviewee. The quality, if not the quantity, of useful information produced by an interview depends in no small part on the skills of the interviewer. Interviewers differ in many ways: their pacing of
J U S T T H I N K � . � . � .
What type of interview situation would you envision as ideal for being carried out entirely through the medium of text-messaging?
J U S T T H I N K � . � . � .
What types of interviewing skills must the host of a talk show possess to be considered an e�ective interviewer? Do these skills di�er from those needed by a professional in the �eld of psychological assessment? If so, how?
coh37025_ch01_001-040.indd 11 12/01/21 4:03 PM
�����Part 1: An Overview
interviews, their rapport with interviewees, and their ability to convey genuineness, empathy, and humor. Keeping these differences firmly in mind, consider Figure 1–2. How might the distinctive personality attributes of these two celebrities affect responses of interviewees? Which of these two interviewers do you think is better at interviewing? Why?
The Portfolio
Students and professionals in many different fields of endeavor ranging from art to architecture keep files of their work products. These work products—whether retained on paper, canvas, film, video, audio, or some other medium—constitute what is called a portfolio. As samples of one’s ability and accomplishment, a portfolio may be used as a tool of evaluation. Employers of commercial artists, for example, will make hiring decisions based, in part, on the impressiveness of an applicant’s portfolio of sample drawings. As another example, consider the employers of on-air radio talent. They, too, will make hiring decisions that are based partly upon their judgments
of (audio) samples of the candidate’s previous work. The appeal of portfolio assessment as a tool of evaluation
extends to many other fields, including education. Some have argued, for example, that the best evaluation of a student’s writing skills can be accomplished not by the administration of a test, but by asking the student to compile a selection of writing samples. Also in the field of education, portfolio assessment has been
Figure �–� On interviewing and being interviewed.
Different interviewers have different styles of interviewing. How would you characterize the interview style of Jimmy Fallon as compared to that of Howard Stern? Theo Wargo/Getty Images
J U S T T H I N K � . � . � .
If you were to prepare a portfolio representing “who you are” in terms of your educational career, your hobbies, and your values, what would you include in your portfolio?
coh37025_ch01_001-040.indd 12 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
employed as a tool in the hiring of instructors. An instructor’s portfolio may consist of various documents such as lesson plans, published writings, and visual aids developed expressly for teaching certain subjects. All of these materials can be extremely useful to those who must make hiring decisions.
Case History Data
Case history data refers to records, transcripts, and other accounts in written, pictorial, or other form that preserve archival information, official and informal accounts, and other data and items relevant to an assessee. Case history data may include files or excerpts from files maintained at institutions and agencies such as schools, hospitals, employers, religious institutions, and criminal justice agencies. Other examples of case history data are letters and written correspondence including email, photos and family albums, newspaper and magazine clippings, home videos, movies, audiotapes, work samples, artwork, doodlings, and accounts and pictures pertaining to interests and hobbies. Postings on social media such as Facebook, Instagram, or Twitter may also serve as case history data. Employers, university admissions departments, healthcare providers, forensic investigators, and others may collect data from postings on social media to help inform inference and decision making (Lis et al., 2015; Pirelli et al., 2016).
Case history data is a useful tool in a wide variety of assessment contexts. In a clinical evaluation, for example, case history data can shed light on an individual’s past and current adjustment as well as on the events and circumstances that may have contributed to any changes in adjustment. Case history data can be of critical value in neuropsychological evaluations, where it often provides information about neuropsychological functioning prior to the occurrence of a trauma or other event that results in a deficit. School psychologists rely on case history data for insight into a student’s current academic or behavioral standing. Case history data is also useful in making judgments concerning future class placements.
The assembly of case history data, as well as related data, into an illustrative account is referred to by terms such as case study or case history. We may formally define a case study (or case history) as a report or illustrative account concerning a person or an event that was compiled on the basis of case history data. A case study might, for example, shed light on how one individual’s personality and a particular set of environmental conditions combined to produce a successful world leader. A case study of an individual who attempted to assassinate a high-ranking political figure could shed light on what types of individuals and conditions might lead to similar attempts in the future. Work on a social psychological phenomenon referred to as groupthink contains rich case history material on collective decision making that did not always result in the best decisions (Janis, 1972). Groupthink arises as a result of the varied forces that drive decision-makers to reach a consensus (such as the motivation to reach a compromise in positions).
Case history data, usually in combination with other intelligence (informative data), also play an important role in military or political threat assessment (Bolante & Dykeman, 2015; Borum, 2015; Dietz et� al., 1991; Gardeazabal & Sandler, 2015; Malone, 2015; Mrad et al., 2015). The United States Secret Service has long relied on such information to help protect the President as well its other protectees (Coggins et al., 1998; Institute of Medicine, 1984; Takeuchi et al., 1981; Vossekuil & Fein, 1997).
Behavioral Observation
If you want to know how someone behaves in a particular situation, observe the individual’s behavior in that situation. Such “down-home” wisdom underlies at least one approach to evaluation. Behavioral observation, as it is employed by assessment professionals, is defined as
J U S T T H I N K � . � . � .
What are the pros and cons of using case history data as a tool of assessment?
coh37025_ch01_001-040.indd 13 12/01/21 4:03 PM
�����Part 1: An Overview
monitoring the actions of others or oneself by visual or electronic means while recording quantitative and/or qualitative information regarding those actions. Behavioral observation is often used as a diagnostic aid in various settings such as inpatient facilities, behavioral research laboratories, and classrooms. Behavioral observation may be used for purposes of selection or placement in corporate or organizational settings. In such instances, behavioral observation may be used as an aid in identifying personnel who best demonstrate the abilities required to perform a particular task or job. Sometimes researchers venture outside of the confines of clinics, classrooms, workplaces, and research laboratories in order to observe behavior of humans in a natural setting—that is, the setting in which the behavior would typically be expected to occur. This variety of behavioral observation is referred to as naturalistic observation. So, for example, to study the socializing behavior of children with autism spectrum disorders with same-age peers, one research team opted for natural settings rather than a controlled, laboratory environment (Bellini et al., 2007; Dekker et al., 2016; Handen et al., 2018).
Behavioral observation as an aid to designing therapeutic intervention is extremely useful in institutional settings such as schools, hospitals, prisons, and group homes. Using published or self-constructed lists of targeted behaviors, staff can observe firsthand the behavior of individuals and design interventions accordingly. In a school situation, for example, naturalistic observation on the playground of a culturally different child
suspected of having linguistic problems might reveal that the child has the necessary English language skills but is unwilling—for reasons of shyness, cultural upbringing, or whatever—to demonstrate those abilities to adults.
In practice, behavioral observation, and especially naturalistic observation, tends to be used most frequently by researchers in settings such as classrooms, clinics, prisons, and other types of facilities where observers have ready access to assessees. For private practitioners, it is typically not practical or economically feasible to spend hours out of the consulting room observing clients as they go about their daily lives. Still, there are some mental health professionals, such as those in the field of assisted living, who find great value in behavioral observation of patients outside of their institutional environment. For them, it may be necessary to accompany a patient outside of the institution’s walls to learn if that patient is capable of independently performing activities of daily living. In this context, a tool of assessment that relies heavily on behavioral observation, such as the Test of Grocery Shopping Skills (see�Figure�1–3), may be extremely useful.
Role-Play Tests
Role play may be defined as acting an improvised or partially improvised part in a simulated situation. A role-play test is a tool of assessment wherein assessees are directed to act as if they were in a particular situation. Assessees may then be evaluated with regard to their expressed thoughts, behaviors, abilities, and other variables. (Note that role play is hyphenated when used as an adjective or a verb but not as a noun.)
Role play is useful in evaluating various skills. For example, grocery shopping skills (Figure 1–3) could conceivably be evaluated through role play. Depending upon how the task is set up, an actual trip to the supermarket could or could not be required. Of course, role play
may not be as useful as “the real thing” in all situations. Still, role play is used quite extensively, especially in situations where it is too time-consuming, too expensive, or simply too inconvenient to assess in a real situation. For example, astronauts in training may be required to role-play many situations “as if” in outer space. Such “as if” scenarios for training purposes result in truly “astronomical” savings.
J U S T T H I N K . � . � .
What are the pros and cons of role play as a tool of assessment? In your opinion, what type of presenting problem would be ideal for assessment by role play?
J U S T T H I N K . � . � .
What are the advantages and disadvantages of naturalistic observation as tools of assessment?
coh37025_ch01_001-040.indd 14 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
Individuals being evaluated in a corporate, industrial, organizational, or military context for managerial or leadership ability may routinely be placed in role-play situations. They may be asked, for example, to mediate a hypothetical dispute between personnel at a work site. The format of the role play could range from “live scenarios” with live actors, or computer- generated simulations. Outcome measures for such an assessment might include ratings related to various aspects of the individual’s ability to resolve the conflict, such as effectiveness of approach, quality of resolution, and number of minutes to resolution.
Role play as a tool of assessment may also be used in various clinical contexts. For example, it is routinely employed in many interventions with substance abusers. Clinicians may attempt to obtain a baseline measure of substance abuse, cravings, or coping skills by administering a role-play test prior to therapeutic intervention. The same test is then administered again subsequent to completion of treatment. Role play can thus be used as both a tool of assessment and a measure of outcome.
Computers as Tools
We have already made reference to the role computers play in contemporary assessment in the context of generating simulations. They may also help in the measurement of variables that in the past were quite difficult to quantify. But perhaps the more obvious role as a tool of assessment is their role in test administration, scoring, and interpretation.
As test administrators, computers do much more than replace the “equipment” that was so widely used in the past (e.g., a number 2 pencil). Computers can serve as test administrators (online or off) and as highly efficient test scorers. Within seconds they can derive not only test scores but
Figure �–� Price (and judgment) check in aisle �.
Designed primarily for use with persons with psychiatric disorders, the context-based Test of Grocery Shopping Skills (Brown et al., 2009; Hamera & Brown, 2000) may be very useful in evaluating a skill necessary for independent living. Dave and Les Jacobs LLC/Blend Images
coh37025_ch01_001-040.indd 15 12/01/21 4:03 PM
�����Part 1: An Overview
patterns of test scores. Scoring may be done on-site (local processing) or conducted at some central location (central processing). If processing occurs at a central location, test-related data may be sent to and returned from this central facility by means of the Internet, phone lines (teleprocessing), mail, or courier. Whether processed locally or centrally, an account of a testtaker’s performance can range from a mere listing of a score or scores (a simple scoring report) to the more detailed extended scoring report, which includes statistical analyses of the testtaker’s performance. A step up from scoring reports is the interpretive report, which is distinguished by its inclusion of numerical or narrative interpretive statements in the report. Some interpretive reports contain relatively little interpretation and simply call attention to certain high, low, or unusual scores. At the high end of interpretive reports is what is sometimes referred to as a consultative report. This type of report, usually written in language appropriate for communication between assessment professionals, may provide expert opinion concerning analysis of the data. Yet another type of computerized scoring report is designed to integrate data from sources other than the test itself into the interpretive report. Such an integrative report will employ previously collected data (such as medication records or behavioral observation data) into the test report.
An acronym you may come across is CAT, which stands for computer adaptive testing. The adaptive in this term is a reference to the computer’s ability to tailor the test to the testtaker’s ability or test-taking pattern. For example, on a computerized test of academic abilities, the computer might be programmed to switch from testing math skills to English skills after three consecutive failures on math items. Another way a computerized test could be programmed to adapt is by providing the testtaker with score feedback as the test proceeds. Score feedback in the context of CAT may, depending on factors such as intrinsic motivation and external incentives, positively affect testtaker engagement as well as performance (Arieli-Attali & Budescu, 2015).
Another acronym, CAPA, refers to the term computer-assisted psychological assessment. In this case, the word assisted typically refers to the assistance computers provide to the test user, not the testtaker. One specific brand of CAPA, for example, is Q-Interactive. Available from Pearson Assessments, this technology allows test users to administer tests by means of two iPads connected by bluetooth (one for the test administrator and one for the testtaker). Test administrators may record testtakers’ verbal responses and may make written notes using a stylus with the iPad. Scoring is immediate. Sweeney (2014) reviewed Q-Interactive and was favorably impressed. He liked the fact that it obviated the need for many essentials of paper-and-pencil test administration (including test kits and a stopwatch). However, he did point out that only a limited number of tests are available to administer, and that no Android or Windows edition of the software has been made available.
Also, despite the publisher’s promise of freedom from test kits, the reviewer often found himself “going back to the manual” (Sweeney, 2014, p. 19). Since the time of the Sweeney (2014) review, a total of 20 assessment tools have been added to the Q-Interactive testing system, which continues to be available exclusively on iPads. Vrana and Vrana (2017) carefully examined the elements of the Wechsler individual intelligence tests, arguing for viability of completely computer-administered assessment in the near future.
CAPA opened a world of possibilities for test developers, enabling them to create psychometrically sound tests using mathematical procedures and calculations so complicated that they may have taken weeks or months to use in a bygone era. It opened a new world to test users, enabling the construction of tailor-made tests with built-in scoring and interpretive capabilities previously unheard of. For many test users, CAPA was a great advance over the past, when they had to personally administer tests and possibly even place the responses in some other form prior to analysis (such as by manually using a scoring template or other device). And even after doing all of that, they would then begin the often laborious tasks of scoring and interpreting the resulting data. Still, every rose has its thorns; some of the pros and cons of CAPA are summarized in Table 1–2. The number of tests in this format is burgeoning, and test
J U S T T H I N K . � . � .
Describe a test that would be ideal for computer administration. Then describe a test that would not be ideal for computer administration.
coh37025_ch01_001-040.indd 16 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
users must take extra care in selecting the right test given factors such as the objective of the testing and the unique characteristics of the test user (Zygouris & Tsolaki, 2015).
The APA Committee on Psychological Tests and Assessment was convened to consider the pros and cons of computer-assisted assessment, and assessment using the Internet (Naglieri et�al., 2004). Among the advantages over paper-and-pencil tests cited were (1) test administrators have greater access to potential test users because of the global reach of the Internet, (2) scoring and interpretation of test data tend to be quicker than for paper-and-pencil tests, (3) costs associated with Internet testing tend to be lower than costs associated with paper-and-pencil tests, and (4)� the Internet facilitates the testing of otherwise isolated populations, as well as people with disabilities for whom getting to a test center might prove a hardship. We might add that Internet testing tends to be “greener,” as it may conserve paper, shipping materials, and so forth. Further, there is probably less chance for scoring errors with Internet-based tests as compared to paper-and-pencil tests.
Although Internet testing appears to have many advantages, it is not without potential pitfalls, problems, and issues. One basic issue has to do with what Naglieri et al. (2004) termed “test-client integrity.” In part this term refers to the verification of the identity of the testtaker when a test is administered online. It also refers, in more general terms, to the sometimes varying interests of the testtaker versus that of the test administrator. Depending upon the conditions of the administration, testtakers may have unrestricted access to notes, other Internet resources, and
Table �–� CAPA: Some Pros and Cons
Pros Cons
CAPA saves professional time in test administration, scoring, and interpretation.
Professionals must still spend signi�cant time reading software and hardware documentation and even ancillary books on the test and its interpretation.
CAPA results in minimal scoring errors resulting from human error or lapses of attention or judgment.
With CAPA, the possibility of software or hardware error is ever present, from di�cult-to-pinpoint sources such as software glitches or hardware malfunction.
CAPA ensures standardized test administration to all testtakers with little, if any, variation in test administration procedures.
CAPA leaves those testtakers who are unable to employ familiar test-taking strategies (previewing test, skipping questions, going back to previous question, etc.) at a disadvantage.
CAPA yields standardized interpretation of �ndings due to elimination of unreliability traceable to di�ering points of view in professional judgment.
CAPA’s standardized interpretation of �ndings based on a set, unitary perspective may not be optimal; interpretation could pro�t from alternative viewpoints.
Computers’ capacity to combine data according to rules is more accurate than that of humans.
Computers lack the �exibility of humans to recognize the exception to a rule in the context of the “big picture.”
Nonprofessional assistants can be used in the test administration process, and the test can typically be administered to groups of testtakers in one sitting.
Use of nonprofessionals leaves diminished, if any, opportunity for the professional to observe the assessee’s test-taking behavior and note any unusual extra-test conditions that may have a�ected responses.
Professional groups such as APA develop guidelines and standards for use of CAPA products.
Pro�t-driven nonprofessionals may also create and distribute tests with little regard for professional guidelines and standards.
Paper-and-pencil tests may be converted to CAPA products with consequential advantages, such as a shorter time between the administration of the test and its scoring and interpretation.
The use of paper-and-pencil tests that have been converted for computer administration raises questions about the equivalence of the original test and its converted form.
Security of CAPA products can be maintained not only by traditional means (such as locked �ling cabinets) but by high-tech electronic products (such as �rewalls).
Security of CAPA products can be breached by computer hackers, and integrity of data can be altered or destroyed by untoward events such as introduction of computer viruses.
Computers can automatically tailor test content and length based on responses of testtakers.
Not all testtakers take the same test or have the same test-taking experience.
coh37025_ch01_001-040.indd 17 12/01/21 4:03 PM
�����Part 1: An Overview
other aids in test-taking—despite the guidelines for the test administration. At least with regard to achievement tests, there is some evidence that unproctored Internet testing leads to “score inflation” as compared to more traditionally administered tests (Carstairs & Myors, 2009).
A related aspect of test-client integrity has to do with the procedure in place to ensure that the security of the Internet-administered test is not compromised. What will prevent other
testtakers from previewing past—or even advance—copies of the test? Naglieri et al. (2004) reminded their readers of the distinction between testing and assessment, and the importance of recognizing that Internet testing is just that—testing, not assessment. As such, Internet test users should be aware of all of the possible limitations of the source of the test scores.
Other Tools
The next time you have occasion to stream a video, fire-up that Blu-ray player, or even break- out an old DVD, take a moment to consider the role that video can play in assessment. In fact, specially created videos are widely used in training and evaluation contexts. For example, corporate personnel may be asked to respond to a variety of video-presented incidents of sexual harassment in the workplace. Police personnel may be asked how they would respond to various types of emergencies, which are presented either as reenactments or as video recordings of actual occurrences. Psychotherapists may be asked to respond with a diagnosis and a treatment plan for each of several clients presented to them on video. Graduate students in psychology programs may use interactive online programs like Theravue to develop their basic counseling skills. The list of video’s potential applications to assessment is endless. The next generation of video assessment is the assessment that employs virtual reality (VR) technology. Assessment using VR technology is fast finding its way into a number of psychological specialty areas (Anbro et al., 2020; Morina et al., 2015; Sharkey & Merrick, 2016).
Many items that you may not readily associate with psychological assessment may be pressed into service for just that purpose. For example, psychologists may use many of the tools traditionally associated with medical health, such as thermometers to measure body temperature and gauges to measure blood pressure. Biofeedback equipment is sometimes used to obtain measures of bodily reactions (such as muscular tension) to various sorts of stimuli. And then there are some less common instruments, such as the penile plethysmograph. This instrument, designed to measure male sexual arousal, may be helpful in the diagnosis and treatment of sexual predators. Impaired ability to identify odors is common in many disorders in which there is central nervous system involvement, and simple tests of smell may be
administered to help determine if such impairment is present. In general, there has been no shortage of innovation on the part of psychologists in devising measurement tools, or adapting existing tools, for use in psychological assessment.
To this point, our introduction has focused on some basic definitions, as well as a look at some of the “tools of the (assessment) trade.” We now raise some fundamental questions regarding the who, what, why, how, and where of testing and assessment.
Who, What, Why, How, and Where?
Who are the parties in the assessment enterprise? In what types of settings are assessments conducted? Why is assessment conducted? How are assessments conducted? Where does one go for authoritative information about tests? Think about the answer to each of these important questions before reading on. Then check your own ideas against those that follow.
J U S T T H I N K . � . � .
When is assessment using video a better approach than using a paper-and-pencil test? What are the pitfalls, if any, to using video in assessment?
J U S T T H I N K . � . � .
What cautions should Internet test users keep in mind regarding the source of their test data?
coh37025_ch01_001-040.indd 18 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
Who Are the Parties?
Parties in the assessment enterprise include developers and publishers of tests, users of tests, and people who are evaluated by means of tests. Additionally, we may consider society at large as a party to the assessment enterprise.
The test developer�Test developers and publishers create tests or other methods of assessment. The American Psychological Association (APA) has estimated that more than 20,000 new psychological tests are developed each year. Among these new tests are some that were created for a specific research study, some that were created in the hope that they would be published, and some that represent refinements or modifications of existing tests. Test creators bring a wide array of backgrounds and interests to the test development process.
Test developers and publishers appreciate the significant influence that test results can have on people’s lives. Accordingly, a number of professional organizations have published standards of ethical behavior that specifically address aspects of responsible test development and use. Perhaps the most detailed document addressing such issues is one jointly written by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education (NCME). Referred to by many psychologists simply as “the Standards,” Standards for Educational and Psychological Testing covers issues related to test construction and evaluation, test administration and use, and special applications of tests, such as special considerations when testing linguistic minorities. Initially published in 1954, revisions of the Standards were published in 1966, 1974, 1985, 1999, and 2014. The Standards is an indispensable reference work not only for test developers but for test users as well.
The test user�Psychological tests and assessment methodologies are used by a wide range of professionals, including clinicians, counselors, school psychologists, human resources personnel, consumer psychologists, industrial-organizational psychologists, experimental psychologists, and social psychologists. In fact, with respect to the job market, the demand for psychologists with measurement expertise far outweighs the supply (Dahlman & Geisinger, 2015). Still, questions remain as to who exactly is qualified to use psychological tests.
The Standards and other published guidelines from specialty professional organizations have had much to say in terms of identifying just who is a qualified test user and who should have access to (and be permitted to purchase) psychological tests and related tools of psychological assessment (American Psychological Association, 2017). Still, controversy exists about which professionals with what type of training should have access to which tests. Members of various professions, with little or no psychological training, have sought the right to obtain and use psychological tests. In many countries, no ethical or legal regulation of psychological test use exists (Leach & Oakland, 2007).
So who are (or should be) test users? Should occupational therapists, for example, be allowed to administer psychological tests? What about employers and human resources executives with no formal training in psychology?
So far, we’ve listed a number of controversial Who? questions that knowledgeable assessment professionals still debate. Fortunately, there is at least one Who? question about which there is very little debate: the one regarding who the testtaker or assessee is.
The testtaker�We have all been testtakers. However, we have not all approached tests in the same way. On the day a test is to be administered, testtakers may vary with respect to numerous variables, including these:
� The amount of test anxiety they are experiencing and the degree to which that test anxiety might significantly affect their test results
J U S T T H I N K � . � . � .
In addition to psychologists, who should be permitted access to, as well as the privilege of using, psychological tests?
coh37025_ch01_001-040.indd 19 12/01/21 4:03 PM
�����Part 1: An Overview
� The extent to which they understand and agree with the rationale for the assessment � Their capacity and willingness to cooperate with the examiner or to comprehend written
test instructions � The amount of physical pain or emotional distress they are experiencing � The amount of physical discomfort brought on by not having had enough to eat, having
had too much to eat, or other physical conditions � The extent to which they are alert and wide awake as opposed to nodding off � The extent to which they are predisposed to agree or disagree when presented with
stimulus statements � The extent to which they have received prior coaching � The importance they may attribute to portraying themselves in a good (or bad) light � The extent to which they are, for lack of a better term, “lucky” and can “beat the odds”
on a multiple-choice achievement test (even though they may not have learned the subject matter).
In the broad sense in which we are using the term “testtaker,” anyone who is the subject of an assessment or an evaluation can be a testtaker or an assessee. As amazing as it sounds, this means that even a deceased individual can be considered an assessee. True, a deceased person is the exception to the rule, but there is such a thing as a psychological autopsy. A psychological autopsy is defined as a reconstruction of a deceased individual’s psychological profile on
the basis of archival records, artifacts, and interviews previously conducted with the deceased assessee or people who knew the person well. For example, using psychological autopsies, Townsend (2007) explored the question of whether suicide terrorists were indeed suicidal from a classical psychological perspective. She concluded that they were not. Other researchers have provided fascinating postmortem psychological evaluations of people from various walks of life in many different cultures (Bhatia et al., 2006; Chan et�al., 2007; Dattilio, 2006; Fortune et al., 2007; Foster, 2011; Giner et al., 2007; Goldstein et al., 2008; Goodfellow et al., 2020; Heller et al., 2007; Knoll & Hatters Friedman, 2015; McGirr et al., 2007; Nock et al., 2017; Owens et al., 2008; Palacio et al., 2007; Phillips et al., 2007; Pouliot & De Leo, 2006; Ross et al., 2017; Rouse et al., 2015; Sanchez, 2006; Thoresen et al., 2006; Vento et al., 2011; Zonda, 2006).
Society at large
The uniqueness of individuals is one of the most fundamental characteristic facts of life.�.�.�. At all periods of human history men have observed and described differences between individuals.�.�.�. But educators, politicians, and administrators have felt a need for some way of organizing or systematizing the many-faceted complexity of individual differences. (Tyler, 1965, p. 3)
The societal need for “organizing” and “systematizing” has historically manifested itself in such varied questions as “Who is a witch?,” “Who is schizophrenic?,” and “Who is qualified?” The specific questions asked have shifted with societal concerns. The methods used to determine the answers have varied throughout history as a function of factors such as intellectual sophistication and religious preoccupation. Proponents of palmistry, podoscopy, astrology, and phrenology, among other pursuits, have argued that the best means of understanding and predicting human behavior was through the study of the palms of the hands, the feet, the stars, bumps on the head, tea leaves, and so on. Unlike such pursuits, the assessment enterprise has roots in science. Through systematic and replicable means that can produce compelling evidence, the assessment enterprise responds to what Tyler (1965, p. 3) described as society’s demand for “some way of organizing or systematizing the many-faceted complexity of individual differences.”
J U S T T H I N K . � . � .
What recently deceased public �gure would you like to see a psychological autopsy done on? Why? What results might you expect?
coh37025_ch01_001-040.indd 20 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
Society at large exerts its influence as a party to the assessment enterprise in many ways. As society evolves and as the need to measure different psychological variables emerges, test developers respond by devising new tests. Through elected representatives to the legislature, laws are enacted that govern aspects of test development, test administration, and test interpretation. Similarly, by means of court decisions, as well as less formal means (see Figure 1–4), society at large exerts its influence on various aspects of the testing and assessment enterprise.
Other parties�Beyond the four primary parties we have focused on here, let’s briefly make note of others who may participate in varied ways in the testing and assessment enterprise. Organizations, companies, and governmental agencies sponsor the development of tests for various reasons, such as to certify personnel. Companies and services offer test-scoring or interpretation services. In some cases these companies and services are simply extensions of test publishers, and in other cases they are independent. There are people whose sole responsibility is the marketing and sales of tests. Sometimes these people are employed by the test publisher; sometimes they are not. There are academicians who review tests and evaluate their psychometric soundness. All of these people, as well as many others, are parties to a greater or lesser extent in the assessment enterprise.
Having introduced you to some of the parties involved in the Who? of psychological testing and assessment, let’s move on to tackle some of the What? and Why? questions.
In What Types of Settings Are Assessments Conducted, and Why?
Educational settings�You are probably no stranger to the many types of tests administered in the classroom. As mandated by law, tests are administered early in school life to help identify children who may have special needs. In addition to school ability tests, another type of test commonly given in schools is an achievement test, which evaluates accomplishment or the degree of learning that has taken place. Some of the achievement tests you have taken in school were constructed by your teacher. Other achievement tests were constructed for more widespread use by educators working with measurement professionals. In the latter category, initialisms such as SAT and GRE may ring a bell.
Figure �–� Public feedback regarding an educational testing program.
In recent years there have been many public demonstrations against various educational testing programs. Strident voices have called for banishing such programs, or for parents to “opt out” of having their children tested. As you learn more about the art and science of testing, assessment, and measurement, you will no doubt develop an informed opinion about whether tests do more harm than good, or vice�versa. Eric Crama/Shutterstock
coh37025_ch01_001-040.indd 21 12/01/21 4:03 PM
�����Part 1: An Overview
You know from your own experience that a diagnosis may be defined as a description or conclusion reached on the basis of evidence and opinion. Typically this conclusion is reached through a process of distinguishing the nature of something and ruling out alternative conclusions. Similarly, the term diagnostic test refers to a tool of assessment used to help narrow down and identify areas of deficit to be targeted for intervention. In educational settings, diagnostic tests of reading, mathematics, and other academic subjects may be administered to assess the need for educational intervention as well as to establish or rule out eligibility for special education programs.
Schoolchildren receive grades on their report cards that are not based on any formal assessment. For example, the grade next to “Works and plays well with others” is probably
based more on the teacher’s informal evaluation in the classroom than on scores on any published measure of social interaction. We may define informal evaluation as a typically non- systematic assessment that leads to the formation of an opinion or attitude.
Informal evaluation is, of course, not limited to educational settings; it is a part of everyday life. In fact, many of the tools of evaluation we have discussed in the context of educational settings (such as achievement tests, diagnostic tests, and informal evaluations) are also administered in various other settings. And some of the types of tests we discuss in the� context of the settings described next are also administered in educational settings. So please keep in mind that the tools of evaluation and measurement techniques that we discuss in one context may well be used in other contexts. Our objective at this early stage in our survey of the field is simply to introduce a sampling (not a comprehensive list) of the types of tests used in different settings.
Clinical settings�Tests and many other tools of assessment are widely used in clinical settings such as public, private, and military hospitals, inpatient and outpatient clinics, private-practice consulting rooms, schools, and other institutions. These tools are used to help screen for or diagnose behavior problems. What types of situations might prompt the employment of such tools? Here’s a small sample:
� A private psychotherapy client wishes to be evaluated to see if the assessment can provide any nonobvious clues regarding his maladjustment.
� A school psychologist clinically evaluates a child experiencing learning difficulties to determine what factors are primarily responsible for it.
� A psychotherapy researcher uses assessment procedures to determine if a particular method of psychotherapy is effective in treating a particular problem.
� A psychologist-consultant retained by an insurance company is called on to give an opinion as to the reality of a client’s psychological problems; is the client really experiencing such problems or just malingering?
� A court-appointed psychologist is asked to give an opinion as to a defendant’s competency to stand trial.
� A prison psychologist is called on to give an opinion regarding the extent of a convicted violent prisoner’s rehabilitation.
The tests employed in clinical settings may be intelligence tests, personality tests, neuropsychological tests, or other specialized instruments, depending on the presenting or suspected problem area. The hallmark of testing in clinical settings is that the test or measurement technique is employed with only one individual at a time. Group testing is used primarily for screening—that is, identifying those individuals who require further diagnostic evaluation.
J U S T T H I N K . � . � .
What kinds of issues do psychologists have to consider when assessing prisoners in contrast to assessing workplace managers?
J U S T T H I N K . � . � .
What tools of assessment could be used to evaluate a student’s social skills?
coh37025_ch01_001-040.indd 22 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
Counseling settings� Assessment in a counseling context may occur in environments as diverse as schools, prisons, and governmental or privately owned institutions. Regardless of the particular tools used, the ultimate objective of many such assessments is the improvement of the assessee in terms of adjustment, productivity, or some related variable. Measures of social and academic skills and measures of personality, interest, attitudes, and values are among the many types of tests that a counselor might administer to a client. Referral questions to be answered range from “How can this child better focus on tasks?” to “For what career is the client best suited?” to “What activities are recommended for retirement?” Having mentioned retirement, let’s hasten to introduce another type of setting in which psychological tests are used extensively.
Geriatric settings�In the United States, more than 14.2 million adults are currently in the age range of 75 to 84; this is about 18 times more people in this age range than there were in 1900. More than six million adults in the United States are currently 85 years old or older, which is a 52-fold increase in the number of people of that age since 1900. People in the United States are living longer, and the population as a whole is getting older.
Older Americans may live at home, in special housing designed for independent living, in housing designed for assisted living, or in long-term care facilities such as hospitals and hospices. Wherever older individuals reside, they may at some point require psychological assessment to evaluate cognitive, psychological, adaptive, or other functioning. At issue in many such assessments is the extent to which assessees are enjoying as good a quality of life as possible. The definition of quality of life has varied as a function of perspective in different studies. In some research, for example, quality of life is defined from the perspective of an observer; in other research it is defined from the perspective of assessees themselves and refers to an individual’s own self-report regarding lifestyle-related variables. However defined, what is typically assessed in quality of life evaluations are variables related to perceived stress, loneliness, sources of satisfaction, personal values, quality of living conditions, and quality of friendships and other social support.
Generally speaking, from a clinical perspective, the assessment of older adults is more likely to include screening for cognitive decline and dementia than the assessment of younger adults (Gallo & Bogner, 2006; Gallo & Wittink, 2006). Dementia is a loss of cognitive functioning (which may affect memory, thinking, reasoning, psychomotor speed, attention, and related abilities, as well as personality) that occurs as the result of damage to or loss of brain cells. Perhaps the best known of the many forms of dementia that exist is Alzheimer’s disease. The road to diagnosis by the clinician is complicated by the fact that severe depression in the elderly can contribute to cognitive functioning that mimics dementia, a condition referred to as pseudodementia (Madden et al., 1952). It is also true that the� majority of individuals suffering from dementia exhibit depressive symptoms (Strober & Arnett, 2009). Clinicians rely on a variety of different tools of assessment to make a diagnosis of dementia or pseudodementia.
Business and military settings�In business, as in the military, various tools of assessment are used in sundry ways, perhaps most notably in decision making about the careers of personnel. A wide range of achievement, aptitude, interest, motivational, and other tests may be employed in the decision to hire as well as in related decisions regarding promotions, transfer, job satisfaction, and eligibility for further training. For a prospective air traffic controller, successful performance on a test of sustained attention to detail may be one requirement of employment. For promotion to the rank of officer in the military, successful performance on a series of leadership tasks may be essential.
Another application of psychological tests involves the engineering and design of products and environments. Engineering psychologists employ a variety of existing and specially devised
J U S T T H I N K . � . � .
Tests are used in geriatric, counseling, and other settings to help improve quality of life. But are there some aspects of quality of life that a psychological test just can’t measure?
coh37025_ch01_001-040.indd 23 12/01/21 4:03 PM
�����Part 1: An Overview
tests in research designed to help people at home, in the workplace, and in the military. Products ranging from home computers to office furniture to jet cockpit control panels benefit from the work of such research efforts.
Using tests, interviews, and other tools of assessment, psychologists who specialize in the marketing and sale of products are involved in taking the pulse of consumers. They help corporations predict the public’s receptivity to a new product, a new brand, or a new advertising or marketing campaign. Psychologists working in the area of marketing help “diagnose” what is wrong (and right) about brands, products, and campaigns. On the basis of such assessments, these psychologists might make recommendations regarding how new brands and products can be made appealing to
consumers, and when it is time for older brands and products to be retired or revitalized. Have you ever wondered about the variety of assessments conducted by a psychologist in
the military? In this chapter’s Meet an Assessment Professional (MAP) feature, we meet U.S. Air Force psychologist, Lt. Col. Alan Ogle, Ph.D., and learn about his wide range of professional duties. Note that each chapter of this book contains a “MAP” feature allowing readers unprecedented access to the “real world life” of a mental health professional who uses psychological tests and other tools of psychological assessment. Each of the featured assessment professionals were asked to write a brief essay in which they shared a thoughtful and educational perspective on their assessment-related activities.
Governmental and organizational credentialing�One of the many applications of measurement is in governmental licensing, certification, or general credentialing of professionals. Before they are legally entitled to practice medicine, physicians must pass an examination. Law school graduates cannot present themselves to the public as attorneys until they pass their state’s bar examination. Psychologists, too, must pass an examination before adopting the official title “psychologist.”
Members of some professions have formed organizations with requirements for membership that go beyond those of licensing or certification. For example, physicians can take further specialized training and a specialty examination to earn the distinction of being “board certified” in a particular area of medicine. Psychologists specializing in certain areas may be evaluated for a diploma from the American Board of Professional Psychology (ABPP) to recognize excellence in the practice of psychology. Another organization, the American Board of Assessment Psychology (ABAP), awards its diploma on the basis of an examination to test users, test developers, and others who have distinguished themselves in the field of testing and assessment.
Academic research settings�Conducting any sort of research typically entails measurement of some kind, and any academician who ever hopes to publish research should ideally have a sound knowledge of measurement principles and tools of assessment. To emphasize this simple fact of research life, imagine the limitless number of questions that psychological researchers could conceivably raise, and the tools and methodologies that might be used to find answers to those questions. For example, Thrash et al. (2010) wondered about the role of inspiration in the writing process. Herbranson and Schroeder (2010) raised the question “Are pigeons smarter than mathematicians?” Milling et al. (2010) asked whether one’s level of hypnotizability predicts
responses to pain-lessening hypnotic suggestions. Angie et al. (2011) explored whether the potential for violence of an ideological group can be assessed by studying the group’s website.
Other settings� Many different kinds of measurement procedures find application in a wide variety of settings. For example, the
J U S T T H I N K . � . � .
What research question would you like to see studied? What tools of assessment might be used in that research?
J U S T T H I N K . � . � .
Assume the role of a consumer psychologist. What ad campaign do you �nd particularly e�ective in terms of pushing consumer “buy” buttons? What ad campaign do you �nd particularly ine�ective in this regard? Why?
coh37025_ch01_001-040.indd 24 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
M E E T A N A S S E S S M E N T P R O F E S S I O N A L
Here, the “best �t” would be those candidates who not only are free of vulnerabilities in psychological health and psychosocial circumstances that might impair performance and possess the requisite quali�cations but also excel in job-relevant skills and characteristics for success in a speci�c unit and mission set.
One example of psychological assessment for a special duty is the program developed and utilized for selection of Military Training Instructors (MTIs) for USAF Basic Military Training (BMT). Called drill instructors or�drill sergeants in other services, these are noncommissioned o�cers (NCOs) with seven or more years of service in their primary career �eld (e.g., aircraft maintenance, security forces, intelligence) selected for this special duty assignment. This is a position of challenge and tremendous trust, tasked with engaging and transforming young civilian volunteers from diverse backgrounds and motivations through a highly intensive training regimen into capable military members. Training can devolve dangerously when not well managed by the instructor—intense training coupled with the power di�erential between MTI and recruits may lead to errors in decision making, overly a�ective responses, maltreatment, or maltraining. Assigning the right instructors, those best skilled and suited for this special duty, is paramount to the success and safety of the training.
Meet Dr. Alan Ogle
arrived at my �rst duty station on �th September, ����, having completed doctoral training at a civilian university followed by an internship at Wright-Patterson Air Force Medical Center. An amazing, challenging, and rewarding career has ensued, with assignments at various bases in the United States, the United Kingdom, and Afghanistan.
As a clinical psychologist for the Air Force, I provide assessment and treatment to military personnel and their families, as well as consultation to military commanders regarding psychological health, substance abuse prevention, and combat and operational stress control. A postdoctoral fellowship and additional military coursework has quali�ed me to also support various other military activities such as high-risk survival, evasion, resistance, and escape (SERE) training, reintegration support services for military and civilians returning from isolation or captivity, human performance optimization, and the evaluation and selection of personnel for special assignments.
The use of clinical assessment measures in the military is comparable to civilian practice. Commonly used measures include brief symptom screeners (such as the Patient Health Questionnaire-� and the Generalized Anxiety Disorder scale-�). We also administer, as indicated, measures of personality and cognitive functioning (such as the current versions of the MMPI and Wechsler tests) to identify treatment needs, monitor progress, and/or assess �tness for military service.
Unlike many other military selection assignments, assessment of military personnel for special missions may entail both “select-in” as well as “select-out” options. Here, the tools of assessment are used to identify psychological or psychosocial concerns that would indicate risk to job candidates (or their families) if selected for a challenging assignment as well as to identify areas that might make a challenging assignment as well as to identify areas that might make a candidate a liability to a mission. Beyond helping to “select out” candidates deemed to be at risk, psychologists assist in helping to “select in” candidates deemed to be the best for a particular unit and mission.
I
Alan Ogle, Ph.D., Lieutenant Colonel, U.S. Air Force
Alan Ogle
(continued)
coh37025_ch01_001-040.indd 25 12/01/21 4:03 PM
�����Part 1: An Overview
Responses are con�dential and not released to the candidate or other coworkers. There is also a component of the MD��� completed by the candidate that includes self- assessment of relevant skills, personality and attitude scales, and a situational judgment test developed speci�c to types of challenges faced in�MTI duty. A concurrent validation study of the self-assessment measures found signi�cant relationships of several attitudes to performance in leadership, mentorship, and risk for maltreatment by MTIs. Based on results of the interview and MD���, a�recommendation is made regarding strengths and�any concerns regarding suitability for MTI duty, including nonrecommend (select out) as well as recommend with su�cient characterization of skills for prioritization of candidates.
At least equally important to “getting the right people” are e�orts to su�ciently train, supervise, and support MTIs through their challenging duties. A�team titled the USAF BMT Military Training Consult Service was established, providing ongoing assessment and support to serving MTIs, as well as training in appropriate use of stress inoculation training of recruits. Additionally, training and command consultation is provided to mitigate risks of behavioral drift inherent to the positional power dynamics of the instructor–recruit relationship. The goal is to support safe, e�ective training of new military members as well as excellence in instructor�sta�.
Students considering service in the military are encouraged to research opportunities, either in uniform or civilian positions. The U.S. Air Force, Army, and Navy each o�er APA-approved internships at multiple sites, for those meeting medical and other requirements, then requiring completion of one assignment. I have been honored to remain in service beyond the initial obligation, thoroughly enjoying the opportunities for training, broad responsibilities from early on in my psychology career, and service with national purpose.
Used with permission of Dr. Alan Ogle.
I had the opportunity to serve on a working group of psychologists to develop an empirically derived, standardized psychological screening protocol of candidates for entry into MTI duty. Job analytic studies were conducted to identify knowledge, skills, abilities, and other characteristics (KSAOs) important to serving successfully in MTI duty, with emphasis on both identi�cation of factors important to safe, e�ective performance, as well as potential “red �ag” warning signs for this position of trust and power over a vulnerable population of trainees. An assessment protocol was developed including an interview by a mental health provider meeting with the MTI candidate and their signi�cant other (if partnered). With awareness that a large body of research indicates clinicians are at risk to overestimate clinical judgment’s accuracy for predicting behavior and job success, the interview is structured by behaviorally anchored rating scales for each of the job-critical areas. Ratings for the domain of judgment/self-control, for instance, include consideration of history of childhood delinquency behaviors (such as skipping school, or �ghting), adult discipline and legal issues, and interview questions such as “What are some choices or mistakes that you particularly regret?” Assessment of Family Stability/ Support includes interview of the candidate and partner regarding questions such as “What would be the most challenging changes for your family in this assignment?” Cognitive screening is required and a brief screening tool is used for time e�ciency.
An additional component of the assessment protocol we developed is the Multidimensional ��� Assessment (MD���), which collects input from a candidate’s coworkers regarding MTI-relevant work performance behaviors and potential “red �ags.” As�examples, subordinates, peers, and supervisors provide ratings about the candidate on items such as, “Remains focused, on task, and decisive in stressful situations,” “Leads others in a fair and consistent manner,” and, “Avoids inappropriate personal relationships (such as �irting or fraternization).”
M E E T A N A S S E S S M E N T P R O F E S S I O N A L
Meet Dr. Alan Ogle (continued)
coh37025_ch01_001-040.indd 26 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
courts rely on psychological test data and related expert testimony as one source of information to help answer important questions such as “Is this defendant competent to stand trial?” and “Did this defendant know right from wrong at the time the criminal act was committed?”
Measurement may play an important part in program evaluation, whether it is a large-scale government program or a small-scale, privately funded one. Is the program working? How can the program be improved? Are funds being spent in the areas where they ought to be spent? How sound is the theory on which the program is based? These are the types of general questions that tests and measurement procedures used in program evaluation are designed to answer.
Tools of assessment can be found in use in research and practice in every specialty area within psychology. For example, consider health psychology, a discipline that focuses on understanding the role of psychological variables in the onset, course, treatment, and prevention of illness, disease, and disability (Cohen, 1994). Health psychologists are involved in teaching, research, or direct-service activities designed to promote good health. Individual interviews, surveys, and paper-and-pencil tests are some of the tools that may be employed to help assess current status with regard to some disease or condition, gauge treatment progress, and evaluate outcome of intervention. One general line of research in health psychology focuses on aspects of personality, behavior, or lifestyle as they relate to physical health. The methodology employed may entail reporting on measurable respondent variables as they change in response to some intervention, such as education, therapy, counseling, change in diet, or change in habits. Measurement tools may be used to compare one naturally occurring group of research subjects to another such group (such as smokers compared to nonsmokers) with regard to some other health-related variable (such as longevity). Many of the questions raised in health-related research have real, life-and-death consequences. All of these important questions, like the questions raised in other areas of psychology, require that sound techniques of evaluation be employed.
How Are Assessments Conducted?
If a need exists to measure a particular variable, a way to measure that variable will be devised. As Figure 1–5 just begins to illustrate, the ways in which measurements can be taken are limited only by imagination. Keep in mind that this figure illustrates only a small sample of the many methods used in psychological testing and assessment. The photos are not designed to illustrate the most typical kinds of assessment procedures. Rather, their purpose is to call attention to the wide range of measurement tools that have been created for varied uses.
Responsible test users have obligations before, during, and after a test or any measurement procedure is administered. For purposes of illustration, consider the administration of a paper-and-pencil test. Before the test, ethical guidelines dictate that when test users have discretion with regard to the tests administered, they should select and use only the test or tests that are most appropriate for the individual being tested. Before a test is administered, the test should be stored in a way that reasonably ensures that its specific contents will not be made known to the testtaker in advance. Another obligation of the test user before the test’s administration is to ensure that a prepared and suitably trained person administers the test properly.
The test administrator (or examiner) must be familiar with the test materials and procedures and must have at the test site all the materials needed to properly administer the test. Materials needed might include a stopwatch, a supply of pencils, and a sufficient number of test protocols. By the way, in everyday, non-test-related conversation, protocol refers to diplomatic etiquette. A less common use of the word is a synonym for the first copy or rough draft of a treaty or other official document before its ratification. With reference to testing and assessment, protocol typically refers to the form, sheet, or booklet on which a testtaker’s responses are entered. The term may also be used to refer to a description of a set of test- or assessment-related procedures, as in the sentence, “The examiner dutifully followed the complete protocol for the stress interview.”
Test users have the responsibility of ensuring that the room in which the test will be conducted is suitable and conducive to the testing. To the extent possible, distracting conditions
coh37025_ch01_001-040.indd 27 12/01/21 4:03 PM
�����Part 1: An Overview
At least since the beginning of the nineteenth century, military units throughout the world have relied on psychological and other tests for personnel selection, program validation, and related reasons (Hartmann et al., 2003). In some cultures where military service is highly valued, students take preparatory courses with hopes of being accepted into elite military units. This is the case in Israel, where rigorous training such as that pictured here prepares high-school students for physical and related tests that only 1 in 60 military recruits will pass. Gil Cohen-Magen/AFP/Getty Images
Evidence suggests that some people with eating disorders may actually have a self-perception disorder; that is, they see themselves as heavier than they really are (Thompson & Smolak, 2001). Thompson and his associates devised the adjustable light-beam apparatus to measure body image distortion. Assessees adjust four beams of light to indicate what they believe is the width of their cheeks, waist, hips, and thighs. A measure of accuracy of these estimates is then obtained. Joel Thompson
Figure �–� The wide world of measurement.
Herman Witkin and his associates (Witkin & Goodenough, 1977) studied personality-related variables in some innovative ways. For example, they identified field (or context)-dependent and field-independent people by means of this specially constructed tilting room–tilting chair device. Assessees were asked questions designed to evaluate their dependence on or independence of visual cues. Source: Witkin, H. A., & Goodenough, D. R. (1977). Field dependence and interpersonal behavior. Psychological Bulletin, 84, 661–689.
coh37025_ch01_001-040.indd 28 12/01/21 4:03 PM
Chapter 1: Psychological Testing and Assessment ��
Pictures such as these sample items from the Meier Art Judgment Test might be used to evaluate people’s aesthetic perception. Which of these two renderings do you find more aesthetically pleasing? The difference between the two pictures involves the positioning of the objects on the shelf. Norman C. Meier Papers, University of Iowa Libraries, Iowa City, Iowa.
Impairment of certain sensory functions can indicate neurological deficit. For purposes of diagnosis, as well as measuring progress in remediation, the neurodevelopment training ball can be useful in evaluating one’s sense of balance. Fotosearch/Getty Images
Some college admissions officers are evaluating the notebook doodles of applicants in their search for “authentic and imperfect” (as opposed to “ideal”) candidates for admission (Gray, 2016). As a result, profiles created on social media platforms such as ZeeMee may increasingly be used by applicants to convey “a side of themselves that might not come through in the typical mix of transcripts, essays and teacher recommendations” (Gray, 2016, p. 48).
coh37025_ch01_001-040.indd 29 12/01/21 4:04 PM
�����Part 1: An Overview
such as excessive noise, heat, cold, interruptions, glaring sunlight, crowding, inadequate ventilation, and so forth should be avoided. Of course, creating an ideal testing environment is not always something every examiner can do (see Figure 1–6).
During test administration, and especially in one-on-one or small-group testing, rapport between the examiner and the examinee is critically important. In this context, rapport may be defined as a working relationship between the examiner and the examinee. Such a working relationship can sometimes be achieved with a few words of small talk when the examiner and examinee are introduced. If appropriate, some words about the nature of the test and why it is important for examinees to do their best may also be helpful. In other instances—for example, with a frightened child—the achievement of rapport might involve more elaborate techniques such as engaging the child in play or some other activity until the child has acclimated to the examiner and the surroundings. It is important that attempts to establish rapport with the testtaker not compromise any rules of the test administration instructions.
After a test administration, test users have many obligations as well. These obligations range from safeguarding the test protocols to conveying the test results in a clearly understandable fashion. If third parties were present during testing or if anything else that might be considered out of the ordinary happened during testing, it is the test user’s responsibility to make a note of such events on the report of the testing. Test scorers have obligations as well. For example, if a test is to be scored by people, scoring needs to conform to pre- established scoring
Figure �–� Less-than-optimal testing conditions.
In 1917, new Army recruits sat on the floor as they were administered the first group tests of intelligence—not ideal testing conditions by current standards. Time Life Pictures/US Signal Corps/The LIFE Picture Collection/Getty Images
J U S T T H I N K . � . � .
What unforeseen incidents could conceivably occur during a test session? Should such incidents be noted on the report of that session?
coh37025_ch01_001-040.indd 30 12/01/21 4:04 PM
Chapter 1: Psychological Testing and Assessment ��
criteria. Test users who have responsibility for interpreting scores or other test results have an obligation to do so in accordance with established procedures and ethical guidelines.
Assessment of people with disabilities�People with disabilities are assessed for exactly the same reasons people with no disabilities are assessed: to obtain employment, to earn a professional credential, to be screened for psychopathology, and so forth. A number of laws have been enacted that affect the conditions under which tests are administered to people with disabling conditions. For example, one law mandates the development and implementation of “alternate assessment” programs for children who, as a result of a disability, could not otherwise participate in state- and district-wide assessments. Defining exactly what “alternate assessment” meant was left to the individual states or their local school districts. These authorities define who requires alternate assessment, how such assessments are to be conducted, and how meaningful inferences are to be drawn from the assessment data.
In general, alternate assessment is typically accomplished by means of some accommodation made to the assessee. The verb to accommodate is defined as “to adapt, adjust, or make suitable.” In the context of psychological testing and assessment, accommodation is defined as the adaptation of a test, procedure, or situation, or the substitution of one test for another, to make the assessment more suitable for an assessee with exceptional needs.
At first blush, the process of accommodating students, employees, or other testtakers with special needs might seem straightforward. For example, the individual who has difficulty reading the small print of a particular test may be accommodated with a large-print version of the same test or with a specially lit test environment. A student with a hearing impairment may be administered the test in sign language. An individual with ADHD might have an extended evaluation time, with frequent breaks during periods of evaluation. Although this may all seem simple at first, it can actually become quite complicated.
Consider, for example, the case of a student with a visual impairment who is scheduled to be given a written, multiple-choice test. There are several possible alternate procedures for test administration. For example, the test could be translated into Braille and administered in that form, or the test could be administered by means of audiotape. However, some students may do better with a Braille administration and others with audiotape. Students with superior short-term attention and memory skills for auditory stimuli would seem to have an advantage with the audiotaped administration. Students with superior haptic (sense of touch) and perceptual-motor skills might have an advantage with the Braille administration. And so, even in this relatively simple example, it can be readily appreciated that a testtaker’s performance (and score) on a test may be affected by the manner of the alternate administration of the test. This reality of alternate assessment raises important questions about how equivalent such methods really are. Indeed, because the alternate procedures have been individually tailored, there is seldom compelling research to support equivalence. Governmental guidelines for alternate assessment will evolve to include ways of translating measurement procedures from one format to another. Other guidelines may suggest substituting one assessment tool for another. Currently there are many ways to accommodate people with disabilities in an assessment situation (see this chapter’s Everyday Psychometrics), and many different definitions of alternate assessment. For the record, we offer our own, general definition of that elusive term. Alternate assessment is an evaluative or diagnostic procedure or process that varies from the usual, customary, or standardized way a measurement is derived, either by virtue of some special accommodation made to the assessee or by means of alternative methods designed to measure the same variable(s).
Having considered some of the who, what, how, and why of assessment, let’s now consider sources for more information with regard to all aspects of the assessment enterprise.
J U S T T H I N K . � . � .
Are there some types of assessments for which no alternate assessment procedure should be developed?
coh37025_ch01_001-040.indd 31 12/01/21 4:04 PM
�����Part 1: An Overview
E V E R Y D A Y P S Y C H O M E T R I C S
Everyday Accommodations
s many as one in seven Americans has a disability that interferes with activities of daily living. In recent years society has acknowledged more than ever before the special needs of citizens challenged by physical and/or mental disabilities. The e�ects of this ever-increasing acknowledgment are visibly evident: special access ramps alongside �ights of stairs, captioned television programming for the hearing-impaired, and large-print newspapers, books, magazines, and size-adjustable online media for the visually impaired. In general, there has been a trend toward altering environments to make individuals with handicapping conditions feel less challenged.
Depending on the nature of a testtaker’s disability and other factors, modi�cations—referred to as accommodations—may need to be made in a psychological test (or measurement procedure) in order for an evaluation to proceed. Accommodation may take many di�erent forms. One general type of accommodation involves the form of the test as presented to the testtaker, as when a written test is set in larger type for presentation to a visually impaired testtaker. Another general type of accommodation concerns the way responses to the test are obtained. For example, a speech-impaired individual might be allowed to write out responses in an examination rather than saying aloud their responses during administration. Students with learning disabilities may be accommodated by being permitted to read test questions aloud (Fuchs et al., ����).
Modi�cation of the physical environment in which a test is conducted is yet another general type of accommodation. For example, a test that is usually group-administered at a central location may on occasion be administered individually to a disabled person at home. Modi�cations of the interpersonal environment in which a test is conducted is another possibility (see Figure �).
Which of many di�erent types of accommodation should be employed? An answer to this question is typically approached by consideration of at least four variables:
�. the capabilities of the assessee;
�. the purpose of the assessment;
�. the meaning attached to test scores; and
�. the capabilities of the assessor.
The Capabilities of the Assessee
Which of several alternate means of assessment is best tailored to the needs and capabilities of the assessee? Case history data, records of prior assessments, and interviews with friends, family, teachers, and others who know the assessee all can provide a
A
wealth of useful information concerning which of several alternate means of assessment is most suitable.
The Purpose of the Assessment
Accommodation is appropriate under some circumstances and inappropriate under others. In general one looks to the purpose of the assessment and the consequences of the accommodation in order to judge the appropriateness of modifying a test to accommodate a person with a disability. For example, modifying a written driving test—or a road test—so a blind person could be
Figure � Modi�cation of the interpersonal environment.
An individual testtaker who requires the aid of a helper or service dog may require the presence of a third party (or animal) if a particular test is to be administered. In some cases, because of the nature of the testtaker’s disability and the demands of a particular test, a more suitable test might have to be substituted for the test usually given if a meaningful evaluation is to be conducted. Huntstock/Getty Images
coh37025_ch01_001-040.indd 32 12/01/21 4:04 PM
Chapter 1: Psychological Testing and Assessment ��
tested for a driver’s license is clearly inappropriate. For their own as well as the public’s safety, the blind are prohibited from driving automobiles. In contrast, changing the form of most other written tests so that a blind person could take them is another matter entirely. In general, accommodation is simply a way of being true to a social policy that promotes and guarantees equal opportunity and treatment for all citizens.
The Meaning Attached to Test Scores
What happens to the meaning of a score on a test when that test has not been administered in the manner that it was designed to be? More often than not, when test administration instructions are modi�ed (some would say “compromised”), the meaning of scores on that test becomes questionable at best. Test users are left to their own devices in interpreting such data. Professional judgment, expertise, and, quite frankly, guesswork can all enter into the process of drawing inferences from scores on modi�ed tests. Of course, a precise record of just how a test was modi�ed for accommodation purposes should be made on the test report.
The Capabilities of the Assessor
Although most persons charged with the responsibility of assessment would like to think that they can administer an
assessment professionally to almost anyone, it is not always the case. It is important to acknowledge that some assessors may experience a level of discomfort in the presence of people with particular disabilities, and this discomfort may a�ect their evaluation. It is also important to acknowledge that some assessors may require additional training prior to conducting certain assessments, including supervised experience with members of certain populations. Alternatively, the assessor may refer such assessment assignments to another assessor who has had more training and experience with members of a particular population.
A burgeoning scholarly literature has focused on various aspects of accommodation, including issues related to general policies (Burns, ����; Nehring, ����; Shriner, ����; Simpson et�al., ����), method of test administration (Calhoon et al., ����; Danford & Steinfeld, ����), score comparability (Elliott et al., ����; Johnson, ����; Pomplun & Omar, ����, ����), documentation (Schulte et al., ����), and the motivation of testtakers to request accommodation (Baldridge & Veiga, ����). Before a decision about accommodation is made for any individual testtaker, due consideration must be given to issues regarding the meaning of scores derived from modi�ed instruments and the validity of the inferences that can be made from the data derived (Guthmann et al., ����; Reesman et al., ����; Toner et al., ����).
Where to Go for Authoritative Information: Reference Sources
Many reference sources exist for learning more about published tests and assessment-related issues. These sources vary with respect to detail. Some merely provide descriptions of tests, others provide detailed information on technical aspects, and still others provide critical reviews complete with discussion of the advantages and disadvantages of usage.
Test catalogues� Perhaps one of the most readily accessible sources of information is a catalogue distributed by the publisher of the test. Because most test publishers make available catalogues of their offerings, this source of test information can be tapped by a simple Internet search, telephone call, email, or note. As you might expect, however, publishers’ catalogues usually contain only a brief description of the test and seldom contain the kind of detailed technical information that a prospective user might require, although publishers are increasingly providing more information in online catalogues, presumably because they are not limited by the space or the cost of printing. It is important to remember, however, that the catalogue’s objective is to sell the test. For this reason, highly critical reviews of a test are seldom, if ever, found in a publisher’s test catalogue.
Test manuals�Detailed information concerning the development of a particular test and technical information relating to it should be found in the test manual, which usually can be purchased from the test publisher. However, for security purposes the test publisher will typically require documentation of professional training before filling an order for a test manual. The chances are good that your university maintains a collection of popular test manuals, perhaps in the library or counseling center. If the test manual you seek is not available there, ask your instructor how best to obtain a reference copy. In surveying the various test manuals, you are likely to see that they vary not only in the details of how the tests were developed and deemed psychometrically sound but also in the candor with which they describe their own test’s limitations.
coh37025_ch01_001-040.indd 33 12/01/21 4:04 PM
�����Part 1: An Overview
Professional books�Many books written for an audience of assessment professionals are available to supplement, reorganize, or enhance the information typically found in the manual of a very widely used psychological test. For example, a book that focuses on a particular test may contain useful information about the content and structure of the test, and how and why that content and structure is superior to a previous version or edition of the test. The book might shed new light on how or why the test may be used for a particular assessment purpose, or administered to members of some special population. The book might provide helpful guidelines for planning a pre-test interview with a particular assessee, or for drawing conclusions from, and making inferences about, the data derived from the test. The book may alert potential users of the test to common errors in test administration, scoring, or interpretation, or to well-documented cautions regarding the use of the test with members of specific cultural groups. In sum, books devoted to an in-depth discussion of a particular test can systematically provide students of assessment, as well as assessment professionals, with the thoughtful insights and actionable knowledge of more experienced practitioners and test users.
Reference volumes�The Buros Center for Testing provides “one-stop shopping” for a great deal of test-related information. The initial version of what would evolve into the Mental Measurements Yearbook series was compiled by Oscar Buros in 1938. This authoritative compilation of test reviews is currently updated about every three years. The Buros Center also publishes Tests in Print, which lists all commercially available English-language tests in print. This volume, which is also updated periodically, provides detailed information for each test listed, including test publisher, test author, test purpose, intended test population, and test administration time.
Journal articles�Articles in current journals may contain reviews of the test, updated or independent studies of its psychometric soundness, or examples of how the instrument was used in either research or an applied context. Such articles may appear in a wide array of behavioral science journals, such as Psychological Bulletin, Psychological Review, Professional Psychology: Research and Practice, Journal of Personality and Social Psychology, Psychology & Marketing, Psychology in the Schools, School Psychology, and School Psychology Review. There are also journals that focus more specifically on matters related to testing and assessment. For example, take a look at journals such as the Journal of Psychoeducational Assessment, Psychological Assessment, Educational and Psychological Measurement, Applied Measurement in Education, and the Journal of Personality Assessment. Journals such as Psychology, Public Policy, and Law and Law and Human Behavior frequently contain highly informative articles on legal and ethical issues and controversies as they relate to psychological testing and assessment. Journals such as Computers & Education, Computers in Human Behavior, and Cyberpsychology, Behavior, and Social Networking frequently contain insightful articles on computer and Internet-related measurement.
Online databases� One of the most widely used bibliographic databases for test-related publications is that maintained by the Educational Resources Information Center (ERIC). Funded by the U.S. Department of Education and operated out of the University of Maryland, the ERIC website at www.eric.ed.gov contains a wealth of resources and news about tests, testing, and assessment. There are abstracts of articles, original articles, and links to other useful websites. ERIC strives to provide balanced information concerning educational assessment and to provide resources that encourage responsible test use.
The American Psychological Association (APA) maintains a number of databases useful in locating psychology-related information in journal articles, book chapters, and doctoral dissertations. Of most relevance to testing and assessment is the APA database, PsycTESTS®. This database of over 58,000 items provides a detailed description as well as development and administration information for each test or assessment. PsycINFO is a database of abstracts dating back to 1887. ClinPSYC is a database derived from PsycINFO that focuses on abstracts of a clinical nature. PsycSCAN: Psychopharmacology contains abstracts of articles concerning psychopharmacology. PsycARTICLES is a database of full-length articles dating back to 1894. Health and Psychosocial Instruments (HAPI) contains a listing of measures created or modified for specific research studies
coh37025_ch01_001-040.indd 34 12/01/21 4:04 PM
Chapter 1: Psychological Testing and Assessment ��
but not commercially available; it is available at many college libraries through BRS Information Technologies. For more information on any of these databases, visit APA’s website at www.apa.org.
The world’s largest private measurement institution is Educational Testing Service (ETS). This company, based in Princeton, New Jersey, maintains a staff of over 3,200 people, including about 1,000 measurement professionals and education specialists. These are the folks who bring you the Scholastic Aptitude Test (SAT) and the Graduate Record Exam (GRE), among many other tests. Descriptions of these and the numerous other tests developed by this company can be found at their website, www.ets.org.
Other sources�A source for exploring the world of unpublished tests and measures is the Directory of Unpublished Experimental Mental Measures (Goldman & Mitchell, 2008). Also, as a service to psychologists and other test users, ETS maintains a list of unpublished tests. This list can be accessed at http://www.ets.org/testcoll/. Some pros and cons of the various sources of information we have listed are summarized in Table 1–3.
Table �–� Sources of Information About Tests: Some Pros and Cons
Information Source Pros Cons
Test catalogue available from the publisher of the test as well as a�liated distributors of the test
Contains general description of test, including what it is designed to do and with whom it is designed to be used. Readily available online or in hard copy to anyone who requests one.
Primarily designed to sell the test to test users and seldom contains any critical reviews. Information not detailed enough for basing a decision to use the test.
Test manual Usually the most detailed source available for information regarding the standardization sample and test administration instructions. May also contain useful information regarding the theory on which the�test is based if that is the case. Typically contains at least some information regarding psychometric soundness of the test.
Details regarding the test’s psychometric soundness are usually self-serving and written on the basis of studies conducted by the test author and/or test publisher. A test manual itself may be di�cult for students to obtain, as its distribution may be restricted to quali�ed professionals.
Professional books May contain one-of-a-kind, authoritative insights of a highly experienced assessment professional regarding the structure and content of the test, as well as more practical insights regarding the administration, scoring, and interpretation of the test.
Be on the lookout for a professional book author who is strongly allied with a unique theoretical perspective with regard to the test. Although useful to know, this theoretical perspective may not be widely accepted. Also, caution is advised when an author expresses strong but idiosyncratic views about the value of a test (or its lack thereof) with assessees who are members of a particular cultural group.
Reference volumes such as the Mental Measurements Yearbook, available in bound book form or online
Much like Consumer Reports for tests, contain descriptions and critical reviews of a test written by third parties who presumably have nothing to gain or lose by praising or criticizing the instrument, its standardization sample, and its psychometric soundness.
Few disadvantages if reviewer is genuinely trying to be objective and is knowledgeable, but as with any review, can provide a misleading picture if this is not the case. Also, for very detailed accounts of the standardization sample and related matters, it is best to consult the test manual itself.
Journal articles Up-to-date source of reviews and studies of psychometric soundness. Can provide practical examples of how an instrument is used in research or applied contexts.
As with reference volumes, reviews are valuable to the�extent that they are informed and, as far as possible, unbiased. Reader should research as many articles as possible when attempting to learn how the instrument is actually used; any one article alone may provide an atypical picture.
Online databases Widely known and respected online databases such as the ERIC database are virtual “gold mines” of useful information containing varying amounts of detail. Although some legitimate psychological tests may be available for self-administration and scoring online, the vast majority are not.
Consumer beware! Some sites masquerading as databases for psychological tests are designed more to entertain or to sell something than to inform. These sites frequently o�er tests you can take online. As you learn more about tests, you will probably become more critical of the value of these self-administered and self-scored “psychological tests.”
coh37025_ch01_001-040.indd 35 12/01/21 4:04 PM
�����Part 1: An Overview
Many university libraries also provide access to online databases, such as PsycINFO, and electronic journals. Most scientific papers can be downloaded straight to one’s computer using such an online service. This service is an extremely valuable resource to students, as non- subscribers to such databases may be charged hefty access fees for such access.
Armed with a wealth of background information about tests and other tools of assessment, we’ll explore historical, cultural, and legal/ethical aspects of the assessment enterprise in the following chapter.
Self-Assessment
Test your understanding of elements of this chapter by seeing if you can explain each of the following terms, expressions, and abbreviations:
accommodation achievement test alternate assessment behavioral observation CAPA case history case history data case study central processing collaborative psychological
assessment consultative report cut score dementia diagnosis diagnostic test dynamic assessment ecological momentary assessment educational assessment extended scoring report format
groupthink health psychology informal evaluation integrative report interpretive report interview local processing motivational interviewing naturalistic observation panel interview portfolio protocol pseudodementia psychological assessment psychological autopsy psychological test psychological testing psychometrician psychometrics psychometric soundness psychometrist
Q-Interactive quality of life rapport remote assessment retrospective assessment role play role-play test score scoring scoring report simple scoring report teleprocessing test test catalogue test developer test manual testtaker test user therapeutic psychological
assessment utility
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (1999). Standards for educational and psychological testing. Author.
American Psychological Association. (2017). 2003 ethical principles of psychologists and code of conduct, as amended 2010 and 2016. Retrieved July 6, 2020, from https://www.apa.org/ethics/code/ethics-code -2017.pdf
Anbro, S. J., Szarko, A. J., Houmanfar, R. A., Maraccini, A. M., Crosswell, L. H., Harris, F. C., Rebaleati, M., & Starmer, L. (2020). Using virtual simulations to assess situational awareness and communication in medical and nursing education: A technical feasibility study. Journal of Organizational Behavior Management. https://doi.org/10.1080/01608061.2020. 1746474
Angie, A. D., Davis, J. L., Allen, M. T., et al. (2011). Studying ideological groups online: Identification and assessment of risk factors for violence. Journal of Applied Social Psychology, 41, 627–657.
Arieli-Attali, M., & Budescu, D. V. (2015). Effects of score feedback on test-taker behavior in self-adapted testing. Multivariate Behavioral Research, 50(6), 724–725.
Baldridge, D. C., & Veiga, J. F. (2006). The impact of anticipated social consequences on recurring disability. Journal of Management, 32(1), 158–179.
Bellini, S., Akullian, J., & Hopf, A. (2007). Increasing social engagement in young children with autism spectrum disorders using video self-monitoring. School Psychology Review, 36(1), 80–90.
Ben-Zeev, D. (2017). Technology in mental health: Creating new knowledge and inventing the future of
coh37025_ch01_001-040.indd 36 12/01/21 4:04 PM
Chapter 1: Psychological Testing and Assessment ��
services. Psychiatric Services, 68, 107–108. https://doi.org/10.1176/appi.ps.201600520
Ben-Zeev, D., McHugo, G. J., Xie, H., Dobbins, K., & Young, M. A. (2012). Comparing retrospective reports to real-time/real-place mobile assessments in individuals with schizophrenia and a nonclinical comparison group. Schizophrenia Bulletin, 38, 396–404. https://doi.org/10.1093/schbul/sbr171
Ben-Zeev, D., Scherer, E. A., Wang, R., Xie, H., Campbell, A. T. (2015a). Next-generation psychiatric assessment: Using smartphone sensors to monitor behavior and mental health. Psychiatric Rehabilitation Journal, 38(3), 218–226.
Ben-Zeev, D., Wang, R., Abdullah, S., et al. (2015b). Mobile behavioral sensing in outpatients and inpatients with schizophrenia. Psychiatric Services, 67(5), 558–561.
Bhatia, M. S., Verma, S. K., & Murty, O. P. (2006). Suicide notes: Psychological and clinical profile. International Journal of Psychiatry in Medicine, 36(2), 163–170.
Binet, A., & Simon, T. (1905). Méthodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L’Année Psychologique, 11, 191–244.
Black, A. C., Cooney, N. L., Justice, A. C., et al. (2016). Momentary assessment of PTSD symptoms and sexual risk behavior in male OEF/OIF/OND Veterans. Journal of Affective Disorders, 190, 424–428.
Bolante, R., & Dykeman, C. (2015). Threat assessment in community colleges. Journal of Threat Assessment and Management, 2(1), 23–32.
Borum, R. (2015). Assessing risk for terrorism involvement. Journal of Threat Assessment and Management, 2(2), 63–87.
Brown, C., Rempfer, M., & Hamera, E. (2009). The Test of Grocery Shopping Skills. AOTA.
Bundy, C. (2004). Changing behaviour: Using motivational interviewing techniques. Journal of the Royal Society of Medicine, 97, 43–47.
Burns, E. (1998). Test accommodations for students with disabilities. Springfield, IL: Charles C Thomas.
Buros, O. K. (1938). The 1938 mental measurements yearbook. Rutgers University Press.
Byrne, G. J., & Bradley, F. (2007). Culture’s influence on leadership efficiency: How personal and national cultures affect leadership style. Journal of Business Research, 60(2), 168–175.
Calhoon, M. B., Fuchs, L. S., & Hamlett, C. L. (2000). Effects of computer-based test accommodations on mathematics performance assessments for secondary students with learning disabilities. Learning Disability Quarterly, 23, 271–282.
Carnevale, J. J., Inbar, Y., & Lerner, J. S. (2011). Individual differences in need for cognition and decision-making competence among leaders. Personality and Individual Differences, 51, 274–278.
Carstairs, J., & Myors, B. (2009). Internet testing: A natural experiment reveals test score inflation on a high-stakes, unproctored cognitive test. Computers in Human Behavior, 25, 738–742.
Chan, S. S., Lyness, J. M., & Conwell, Y. (2007). Do cerebrovascular risk factors confer risk for suicide in later life? A case control study. American Journal of Geriatric Psychiatry, 15(6), 541–544.
Chapman, J. C. (1921). Trade tests. Holt. Coggins, M. H., Pynchon, M. R., & Dvoskin, J. A.
(1998). Integrating research and practice in federal law
enforcement: Secret Service applications of behavioral science expertise to protect the president. Behavioral Sciences and the Law, 16, 51–70.
Cohen, R. J. (1994). Psychology & adjustment: Values, culture, and change. Allyn & Bacon.
Dahlman, K. A., & Geisinger, K. F. (2015). The prevalence of measurement in undergraduate psychology curricula across the United States. Scholarship of Teaching and Learning in Psychology, 1(3), 189–199.
Danford, G. S., & Steinfeld, E. (1999). Measuring the influences of physical environments on the behaviors of people with impairments. In E. Steinfeld & G. S. Danford (Eds.), Enabling environments: Measuring the impact of environment on disability and rehabilitation (pp. 111–137). Kluwer Academic/Plenum.
Dattilio, F. M. (2006). Equivocal death psychological autopsies in cases of criminal homicide. American Journal of Forensic Psychology, 24(1), 5–22.
Dekker, V., Nauta, M. H., Mulder, E. J., Sytema, S., & de Bildt, A. (2016). A fresh pair of eyes: A blind observation method for evaluating social skills of children with ASD in a naturalistic peer situation in school. Journal of Autism and Developmental Disorders, 46, 2890–2904. https://doi.org/10.1007/s10803-016-2829-y
Derue, D. S., Nahrgang, J. D., Wellman, N., & Humphrey, S. E. (2011). Trait and behavioral theories of leadership: An integration and meta-analytic test of their relative validity. Personnel Psychology, 64, 7–52.
Dietz, P. E., Matthews, D. B., Van Duyne, C., et al. (1991). Threatening and otherwise inappropriate letters to Hollywood celebrities. Journal of Forensic Sciences, 36, 185–209.
Dipboye, R. L. (1992). Selection interviews: Process perspectives. South-Western Publishing.
Dixon, M., Wang, S., Calvin, J., et al. (2002). The panel interview: A review of empirical research and guidelines for practice. Public Personnel Management, 31, 397–428.
Dwyer, C. A. (1996). Cut scores and testing: Statistics, judgment, truth, and error. Psychological Assessment, 8, 360–362.
Elliott, R. (2011). Utilising evidence-based leadership theories in coaching for leadership development: Towards a comprehensive integrating conceptual framework. International Coaching Psychology Review, 6, 46–70.
Elliott, S. N., Katochwill, T. R., & McKevitt, B. C. (2001). Experimental analysis of the effects of testing accommodations on the scores of students with and without disabilities. Journal of School Psychology, 39, 3–24.
Finello, K. M. (2011). Collaboration in the assessment and diagnosis of preschoolers: Challenges and opportunities. Psychology in the Schools, 48, 442–453.
Finn, S. E. (2003). Therapeutic assessment of a man with “ADD.” Journal of Personality Assessment, 80, 115–129.
Finn, S. E. (2009). The many faces of empathy in experiential, person-centered, collaborative assessment. Journal of Personality Assessment, 91, 20-23. https://doi.org/10.1080/00223890802483391
Finn, S. E. (2011). Therapeutic assessment on the front lines: Comment on articles from Westcoast Children’s Clinic. Journal of Personality Assessment, 93(1), 23–25.
Finn, S. E., & Martin, H. (1997). Therapeutic assessment with the MMPI-2 in managed health care. In J. N. Butcher (Ed.), Objective psychological assessment in managed health care: A practitioner’s guide (pp. 131–152). Oxford University Press.
coh37025_ch01_001-040.indd 37 12/01/21 4:04 PM
�����Part 1: An Overview
Finn, S. E., & Tonsager, M. E. (2002). How therapeutic assessment became humanistic. Humanistic Psychologist, 30(1–2), 10–22.
Fischer, C. T. (1978). Collaborative psychological assessment. In C. T. Fischer & S. L. Brodsky (Eds.), Client participation in human services: The Prometheus principle (pp. 41–61). Transaction.
Fischer, C. T. (2004). In what sense is collaborative psychological assessment collaborative? Some distinctions. SPA Exchange, 16(1), 14–15.
Fischer, C. T. (2006). Qualitative psychological research and individualized/collaborative psychological assessment: Implications of their similarities for promoting a life-world orientation. Humanistic Psychologist, 34(4), 347–356.
Fischer, C. T., & Finn, S. E. (2014). Developing the life meanings of psychological test data: Collaborative and therapeutic approaches. In R. P. Archer & S. R. Smith (Eds.), Personality assessment (2nd ed., pp. 401–431). Routledge/Taylor & Francis.
Fortune, S., Stewart, A., Yadav, V., & Hawton, K. (2007). Suicide in adolescents: Using life charts to understand the suicidal process. Journal of Affective Disorders, 100(1–3), 199–210.
Foster, T. (2011). Adverse life events proximal to adult suicide: A synthesis of findings from psychological studies. Archives of Suicide Research, 15, 1–15.
Fuchs, L. S., Fuchs, D., Eaton, S. B., et al. (2000). Using objective data sources to enhance teacher judgments about test accommodations. Exceptional Children, 67, 67–81.
Gallo, J. J., & Bogner, H. R. (2006). The context of geriatric care. In J. J. Gallo, H. R. Bogner, T. Fulmer, & G. J. Paveza (Eds.), The handbook of geriatric assessment (4th ed., pp. 3–13). Jones & Bartlett.
Gallo, J. J., & Wittink, M. N. (2006). Cognitive assessment. In J. J. Gallo, H. R. Bogner, T. Fulmer, & G. J. Paveza (Eds.), The handbook of geriatric assessment (4th ed., pp. 105–151). Jones & Bartlett.
Gardeazabal, J., & Sandler, T. (2015). INTERPOL’s surveillance network in curbing transnational terrorism. Journal of Policy Analysis and Management, 34(4), 761–780.
Giner, L., Carballo, J. J., Guija, J. A., et al. (2007). Psychological autopsy studies: The role of alcohol use in adolescent and young adult suicides. International Journal of Adolescent Medicine and Health, 19(1), 99–113.
Goldman, B. A., & Mitchell, D. F. (2008). Directory of unpublished experimental mental measures (Vol. 9). APA Books.
Goldstein, T. R., Bridge, J. A., & Brent, D. A. (2008). Sleep disturbance preceding completed suicide in adolescents. Journal of Consulting and Clinical Psychology, 76(1), 84–91.
Goodfellow, B., Kõlvesa, K., Selefenc, A., Massainb, T., Amadéod, S., & De Leoa, D. (2020). The WHO/ START study in New Caledonia: A psychological autopsy case series. Journal of Affective Disorders 262, 366–372. https://doi.org/10.1016/j. jad.2019.11.020
Gray, E. (2016). The new college application. Time, 187(14), 47–49, 51.
Guthmann, D., Lazowski, L. E., Moore, D., et al. (2012). Validation of the Substance Abuse Screener in American Sign Language (SAS-ASL). Rehabilitation Psychology, 57(2), 140–148.
Hamera, E., & Brown, C. E. (2000). Developing a context-based performance measure for persons with schizophrenia: The test of grocery shopping skills. American Journal of Occupational Therapy, 54, 20–25.
Handen, B. L., Mazefsky, C. A., Gabriels, R. L., Pedersen, K. A., Wallace, M., & Siegel, M. (2018). Risk factors for self-injurious behavior in an inpatient psychiatric sample of children with autism spectrum disorder: A naturalistic observation study. Journal of Autism and Developmental Disorders, 48, 3678–3688. https://doi .org/10.1007/s10803-017-3460-2
Hartmann, E., Sunde, T., Kristensen, W., & Martinussen, M. (2003). Psychological measures as predictors of military training performance. Journal of Personality Assessment, 80, 87–98.
Haywood, H., & Lidz, C. S. (2007). Dynamic assessment in practice: Clinical and educational applications. Cambridge University Press.
Heller, T. S., Hawgood, J. L., & De Leo, D. (2007). Correlates of suicide in building industry workers. Archives of Suicide Research, 11(1), 105–117.
Herbranson, W. T., & Schroeder, J. (2010). Are birds smarter than mathematicians? Pigeons (Columba livia) perform optimally on a version of the Monty Hall Dilemma. Journal of Comparative Psychology, 124(1), 1–13.
Hoy, J., Natarajan, A., & Petra, M. M. (2016). Motivational interviewing and the transtheoretical model of change: Under-explored resources for suicide intervention. Community Mental Health Journal, 52(5), 559–567.
Hull, C. L. (1922). Aptitude testing. World Book. Institute of Medicine. (1984). Research and training for
the Secret Service: Behavioral science and mental health perspectives: A report of the Institute of Medicine (IOM Publication No. IOM-84-01). National Academy Press.
Janis, I. L. (1972). Victims of groupthink. Houghton Mifflin.
Johnson, E. S. (2000). The effects of accommodation on performance assessments. Remedial and Special Education, 21, 261–267.
Kistenmacher, B. R, & Weiss, R. L. (2008). Motivational interviewing as a mechanism for change in men who batter: A randomized controlled trial. Violence and Victims, 23(5), 558–570.
Knoll, J. L., & Hatters Friedman, S. (2015). The homicide- suicide phenomenon: Findings of psychological autopsies. Journal of Forensic Sciences, 60(5), 1253–1257.
Kouzes, J. M., & Posner, B. Z. (2007). The leadership challenge (4th ed.). Jossey-Bass.
Lamiell, J. T. (2003). Beyond individual and group differences: Human individuality, scientific psychology, and William Stern’s critical personalism. Sage.
Leach, M. M., & Oakland, T. (2007). Ethics standards impacting test development and use: A review of 31 ethics codes impacting practices in 35 countries. International Journal of Testing, 7(1), 71–88.
Li, J. J., & Lansford, J. E. (2018). A smartphone-based ecological momentary assessment of parental behavioral consistency: Associations with parental stress and child ADHD symptoms. Developmental Psychology, 54, 1086–1098. https://doi.org/10.1037/ dev0000516
coh37025_ch01_001-040.indd 38 12/01/21 4:04 PM
Chapter 1: Psychological Testing and Assessment ��
Lin, C-H., Chiang, S-L, Heitkemper, M. M., et al. (2016). Effects of telephone-based motivational interviewing in lifestyle modification program on reducing metabolic risks in middle-aged and older women with metabolic syndrome: A randomized controlled trial. International Journal of Nursing Studies, 60, 12–23.
Lis, E., Wood, M. A., Chiniara, C., et al. (2015). Psychiatrists’ perceptions of Facebook and other social media. Psychiatric Quarterly, 86(4), 597–602.
Madden, J. J., Luhan, J. A., Kaplan, L. A., et al. (1952). Non dementing psychoses in older persons. Journal of the American Medical Association, 150, 1567–1570.
Malone, R. (2015). Protective intelligence: Applying the intelligence cycle model to threat assessment. Journal of Threat Assessment and Management, 2(1), 53–62.
Maloney, M. P., & Ward, M. P. (1976). Psychological assessment. Oxford University Press.
McGirr, A., Renaud, J., Seguin, M., et al. (2007). An examination of DSM-IV depressive symptoms and risk for suicide completion in major depressive disorder: A psychological autopsy study. Journal of Affective Disorders, 97(1–3), 203–209.
Medvec, V. H., Madey, S. F., & Gilovich, T. (1995). When less is more: Counterfactual thinking and satisfaction among Olympic medalists. Journal of Personality and Social Psychology, 69, 603–610.
Medvec, V. H., & Savitsky, K. (1997). When doing better means feeling worse: The efforts of categorical cutoff points on counterfactual thinking and satisfaction. Journal of Personality and Social Psychology, 72, 1284–1296.
Miller, W. R., & Rollnick, S. (2002). Motivational interviewing: Preparing people for change (2nd ed.). Guilford Press.
Miller, W. R., & Rollnick, S. (2009). Ten things that motivational interviewing is not. Behavioural and Cognitive Psychotherapy, 37, 129–140.
Miller, W. R., & Rollnick, S. (2012). Motivational interviewing: Helping people change (3rd ed.). Guilford.
Milling, L. S., Coursen, E. L., Shores, J. S., & Waszkiewicz, J. A. (2010). The predictive utility of hypnotizability: The change in suggestibility produced by hypnosis. Journal of Consulting and Clinical Psychology, 78(1), 126–130.
Morina, N., Ijntema, H., Meyerbröker, K., & Emmelkamp, P. M. G. (2015). Can virtual reality exposure therapy gains be generalized to real-life? A meta-analysis of studies applying behavioral assessments. Behaviour Research and Therapy, 74, 18–24.
Mrad, D. F., Hanigan, A. J. S., & Bateman, J. R. (2015). A model of service and training: Threat assessment on a community college campus. Psychological Services, 12(1), 16–19.
Naglieri, J. A., Drasgow, F., Schmit, M., et al. (2004). Psychological testing on the Internet: New problems, old issues. American Psychologist, 59, 150–162.
Nehring, W. M. (2007). Accommodations for school and work. In C. L. Betz & W. M. Nehring (Eds.), Promoting health care transitions for adolescents with special health care needs and disabilities (pp. 97–115). Brookes.
Nock, M. K., Dempsey, C. L., Aliaga, P. A., Brent, D. A., Heeringa, S. G., Kessler, R. C., Stein, M. B., Ursano, R. J., Benedek, D., & On behalf of the Army STARRS Collaborators. (2017). Psychological autopsy study comparing suicide decedents, suicide ideators, and propensity score matched controls: results from the
study to assess risk and resilience in service members (Army STARRS). Psychological Medicine, 47, 2663–2674. https://doi.org/10.1017/S0033291717001179
Owens, C., Lambert, H., Lloyd, K., & Donovan, J. (2008). Tales of biographical disintegration: How parents make sense of their son’s suicides. Sociology of Health & Illness, 30(2), 237–254.
Palacio, C., García, J., Diago, J., et al. (2007). Identification of suicide risk factors in Medellin, Columbia: A case-control study of psychological autopsy in a developing country. Archives of Suicide Research, 11(3), 297–308.
Phillips, M. R., Shen, Q., Liu, X., et al. (2007). Assessing depressive symptoms in persons who die of suicide in mainland China. Journal of Affective Disorders, 98(1–2), 73–82.
Pirelli, G., Otto, R. K., & Estoup, A. (2016). Using Internet and social media data as collateral sources of information in forensic evaluations. Professional Psychology: Research and Practice, 47(1), 12–17.
Poehner, M. E., & van Compernolle, R. A. (2011). Frames of interaction in Dynamic Assessment: Developmental diagnoses of second language learning. Assessment in Education: Principles, Policy & Practice, 18, 183–198.
Pollak, K. I., Coffman, C. J., Tulsky, J. A., et al. (2016). Teaching physicians motivational interviewing for discussing weight with overweight adolescents. Journal of Adolescent Health, 59(1), 96–103.
Pomplun, M., & Omar, M. H. (2000). Score comparability of a state mathematics assessment across students with and without reading accommodations. Journal of Applied Psychology, 85, 21–29.
Pomplun, M., & Omar, M. H. (2001). Score comparability of a state reading assessment across selected groups of students with disabilities. Structural Equation Modeling, 8, 257–274.
Pouliot, L., & De Leo, D. (2006). Critical issues in psychological autopsy studies. Suicide and Life- Threatening Behavior, 36(5), 491–51.
Reesman, J. H., Day, L. A., Szymanski, C. A., et al. (2014). Review of intellectual assessment measures for children who are deaf or hard of hearing. Rehabilitation Psychology, 59(1), 99–106.
Reyman, F., & Shankar, C. (2015). Retrospective assessment of testamentary capacity. Journal of the American Academy of Psychiatry and the Law, 43(1), 116–118.
Rosenman, E. D., Ilgen, J. S., Shandro, J. R., et al. (2015). A systematic review of tools used to assess team leadership in health care action teams. Academic Medicine, 90(10), 1408–1422.
Ross, V., Kõlves, K., & De Leo, D. (2017). Beyond psychopathology: A case–control psychological autopsy study of young adult males. International Journal of Social Psychiatry, 63, 151–160. https://doi.org/10.1177/0020764016688041
Rothman, E. F., & Wang, N. (2016). A feasibility test of a brief motivational interview intervention to reduce dating abuse perpetration in a hospital setting. Psychology of Violence, 6(3), 433–441.
Rouse, L. M., Frey, R. A., López, M., et al. (2015). Law enforcement suicide: Discerning etiology through psychological autopsy. Police Quarterly, 18(1), 79–108.
Ruscio, A. C., Muench, C., Brede, E., & Waters, A. J. (2016). Effect of brief mindfulness practice on self-reported affect, craving, and smoking: A pilot
coh37025_ch01_001-040.indd 39 12/01/21 4:04 PM
�����Part 1: An Overview
randomized controlled trial using ecological momentary assessment. Nicotine & Tobacco Research, 18(1), 64–73.
Sanchez, H. G. (2006). Inmate suicide and the psychological autopsy process. US Department of Justice Jail Suicide/Mental Health Update, 15(2), 5–11.
Schulte, A. A., Gilbertson, E., Kratochwil, T. R. (2000). Educators’ perceptions and documentation of testing accommodations for students with disabilities. Special Services in the Schools, 16, 35–56.
Schurman, J. V., & Friesen, C. A. (2015). Identifying potential pediatric chronic abdominal pain triggers�using ecological momentary assessment. Clinical Practice in Pediatric Psychology, 3(2), 131–141.
Sharkey, P. M., & Merrick, J. (Eds.), (2016). Recent advances in using virtual reality technologies for rehabilitation. Nova Science Publishers.
Shepard, D. S., Lwin, A. K., Barnett, N. P., et al. (2016). Cost-effectiveness of motivational intervention with significant others for patients with alcohol misuse. Addiction, 111(5), 832–839.
Shingleton, R. M., Pratt, E. M., Gorman, B., et al. (2016). Motivational text message intervention for eating disorders: A single-case alternating treatment design using ecological momentary assessment. Behavior Therapy, 47(3), 325–338.
Shriner, J. G. (2000). Legal perspectives on school outcomes assessment for students with disabilities. Journal of Special Education, 33, 232–239.
Simpson, R. L., Griswold, D. E., & Myles, B. S. (1999). Educators’ assessment accommodation preferences for students with autism. Focus on Autism and Other Developmental Disabilities, 14(4), 212–219, 230.
Skov-Ettrup, L. S., Dalum, P., Bech, M., & Tolstrup, J. S. (2016). The effectiveness of telephone counselling and internet- and text-message-based support for smoking cessation: Results from a randomized controlled trial. Addiction, 111(7), 1257–1266.
Spearman, C. (1927). The abilities of man: Their nature and measurement. Macmillan.
Stern, W. (1933). Der personale Faktor in Psychotechnik und praktischer Psychologie [The personal factor in psychotechnics in practical psychology]. Zeitschrift für angewandte Psychologie, 44, 52–63.
Strober, L. B., & Arnett, P. A. (2009). Assessment of depression in three medically ill, elderly populations: Alzheimer’s disease, Parkinson’s disease, and stroke. Clinical Neuropsychologist, 23, 205–230.
Sweeney, C. (2014). Assess: A review of Pearson’s Q-Interactive program. The Ohio School Psychologist, 59(2), 17–20.
Takeuchi, J., Solomon, F., & Menninger, W. W. (Eds.). (1981). Behavioral science and the Secret Service: Toward the prevention of assassination. National Academy.
Teel, E., Gay, M., Johnson, B., & Slobounov, S. (2016). Determining sensitivity/specificity of virtual reality- based neuropsychological tool for detecting residual abnormalities following sport-related concussion. Neuropsychology, 30(4), 474–483.
Thompson, J. K., & Smolak, L. (Eds.). (2001). Body image, eating disorders, and obesity in youth: Assessment, prevention, and treatment. APA Books.
Thoresen, S., Mehlum, L., Roysamb, E., & Tonnessen, A. (2006). Risk factors for completed suicide in veterans of peacekeeping: Repatriation, negative life events, and marital status. Archives of Suicide Research, 10(4), 353–363.
Thrash, T. M., Maruskin, L. A., Cassidy, S. E., et al. (2010). Mediating between the muse and the masses: Inspiration and the actualization of creative ideas. Journal of Personality and Social Psychology, 98(3), 469–487.
Toner, C. K., Reese, B. E., Neargarder, S., et al. (2012). Vision-fair neuropsychological assessment in normal aging, Parkinson’s disease and Alzheimer’s disease. Psychology and Aging, 27(3), 785–790.
Townsend, E. (2007). Suicide terrorists: Are they suicidal? Suicide and Life-Threatening Behavior, 37(1), 35–49.
Tyler, L. E. (1965). The psychology of human differences (3rd ed.). Appleton-Century-Crofts.
Vento, A. E., Schifano, F., Corkery, J. M., et al. (2011). Suicide verdicts as opposed to accidental deaths in substance-related fatalities (UK, 2001–2007). Progress in Neuro-Psychopharmacology & Biological Psychiatry, 35, 1279–1283.
Vossekuil, B., & Fein, R. A. (1997). Final report: Secret Service Exceptional Case Study Project. Washington, DC: U.S. Secret Service, Intelligence Division.
Vrana, S. R., & Vrana, D. T. (2017). Can a computer administer a Wechsler intelligence test? Professional Psychology: Research and Practice, 48, 191–198. https://psycnet.apa.org/doi/10.1037/pro0000128
Vredeveldt, A., Tredoux, C. G., Nortje, A., et al. (2015). A field evaluation of the Eye-Closure Interview with witnesses of serious crimes. Law and Human Behavior, 39(2), 189–197.
Wang, T.-H. (2011). Implementation of Web-based dynamic assessment in facilitating junior high school students to learn mathematics. Computers & Education, 56, 1062–1071.
Witkin, H. A., & Goodenough, D. R. (1977). Field dependence and interpersonal behavior. Psychological Bulletin, 84, 661–689.
Zonda, T. (2006). One-hundred cases of suicide in Budapest: A case-controlled psychological autopsy study. Crisis: The Journal of Crisis Intervention and Suicide Prevention, 27(3), 125–129.
Zygouris, S., & Tsolaki, M. (2015). Computerized cognitive testing for older adults: A review. American Journal of Alzheimer’s Disease and Other Dementias, 30(1), 13–28.
coh37025_ch01_001-040.indd 40 12/01/21 4:04 PM
��
C H A P T E R �
Historical, Cultural, and Legal/Ethical Considerations
e continue our broad overview of the field of psychological testing and assessment with a look backward, the better to appreciate the historical context of the enterprise. We also present “food for thought” regarding cultural and legal/ethical matters. Consider this presentation only as an appetizer; material on historical, cultural, and legal/ethical considerations is interwoven where appropriate throughout this book.
A Historical Perspective
Antiquity to the Nineteenth Century
It is believed that tests and testing programs first came into being in China as early as 2200 B.C.E., though the selection of government officials was still mostly based on political and familial ties (DuBois, 1966, 1970). Beginning in 196 B.C.E., the former system of selecting government officials mostly by heredity was replaced by a system of recommendation and investigation. Local aristocrats recommended qualified candidates to be sent to the capital where they underwent a series of interviews in which they were questioned about how they would solve various problems of politics and governance. Hoping to make the selection of officials more efficient, formal, and meritocratic, emperors of the Sui dynasty created the imperial examination system in the seventh century. Every three years, examinees who had passed local and provincial exams from all over the empire arrived at the capital to undergo rigorous testing about a wide variety of subjects. Generally, only a small percentage passed the exams and was given positions of authority in the government. This system became one of the most durable institutions in world history, operating with few interruptions over the next 13 centuries until it was replaced by political reform efforts in the Qing dynasty in 1906 (Wang, 2012).
On what were applicants for jobs in ancient China tested? As might be expected, the content of the examination changed over time and with the cultural expectations of the day—as well as with the values of the ruling dynasty. Some tests were directly related to the knowledge a civil servant would need. For example, examinees needed to demonstrate they could read, write, keep records, and perform the kinds of arithmetic calculations needed to collect taxes. They needed deep knowledge of civil law and had to demonstrate proficiency in geography, agriculture, and military strategy—all of which were vital to serving in a large agricultural
W
coh37025_ch02_041-084.indd 41 12/01/21 4:04 PM
�����Part 1: An Overview
society that was frequently at war. Some test subjects may seem surprising to modern sensibility: archery, horsemanship, religious rites, classical literature, and poetry writing. According to cultural ideals, a government official should be a soldier-scholar ready to serve the ruling dynasty with physical prowess, moral rectitude, and a deep knowledge of accumulated cultural wisdom from the past (Wang, 2012).
Over its long history, the examination system was at times more rigorous and fair and at other times more lax and corrupt. Societal elites typically bent the system so that less privileged members of the society were either less likely to be able to pass the exams or were prevented outright from taking them. Poor families generally lacked the resources needed to give their sons an extended education. Even so, the historical record has many instances of talented young men from poorer families who were able to vastly improve their lot in life by passing the state-sponsored examinations. Aside from a brief period in the nineteenth century, women—even women from aristocratic families—were not allowed to take the examinations.
In dynasties with state-sponsored examinations for official positions (referred to as imperial examination), the privileges of making the grade varied. During some periods, those who
passed the examination were entitled not only to a government job but also to wear special garb; this entitled them to be accorded special courtesies by anyone they happened to meet. In some dynasties, passing the examinations could result in exemption from taxes. Passing the examination might even exempt one from government-sponsored interrogation by torture if the individual was suspected of committing a crime. Clearly, it paid to do well on these difficult examinations.
Also intriguing from a historical perspective are ancient Greco-Roman writings indicative of attempts to categorize people in terms of personality types. Such categorizations typically included reference to an overabundance or deficiency in some bodily fluid (such as blood or phlegm) as a factor believed to influence personality. During the Middle Ages, a question of critical importance was “Who is in league with the Devil?” and various measurement procedures were devised to address this question. It would not be until the Renaissance that psychological assessment in the modern sense began to emerge. By the eighteenth century, Christian von
Wolff (1732, 1734) had anticipated psychology as a science and psychological measurement as a specialty within that science.
In 1859, the book On the Origin of Species by Means of Natural Selection by Charles Darwin (1809–1882) was published. In this important, far-reaching work, Darwin argued that chance variation in species would be selected or rejected by nature according to adaptivity and survival value. He further argued that humans had descended from the ape as a result of
such chance genetic variations. This revolutionary notion aroused interest, admiration, and a good deal of enmity. The enmity came primarily from religious individuals who interpreted Darwin’s ideas as an affront to the biblical account of creation in Genesis. Still, the notion of an evolutionary link between human beings and animals conferred a new scientific respectability on experimentation with animals. It also raised questions about how animals and humans compare with respect to states of consciousness—questions that would beg for answers in laboratories of future behavioral scientists.1
J U S T T H I N K � . � . � .
What parallels in terms of privileges and bene�ts can you draw between doing well on examinations in ancient China and doing well on modern-day civil service examinations?
J U S T T H I N K � . � . � .
Among the most critical “diagnostic” questions during the Middle Ages was “Who is in league with the Devil?” What is one of the most critical diagnostic questions today?
1. The influence of Darwin’s thinking is also apparent in the theory of personality formulated by Sigmund Freud. In this context, Freud’s notion of the primary importance of instinctual sexual and aggressive urges can be better understood.
coh37025_ch02_041-084.indd 42 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
History records that it was Darwin who spurred scientific interest in individual differences. Darwin (1859) wrote:
The many slight differences which appear in the offspring from the same parents . . . may be called individual differences. . . . These individual differences are of the highest importance�.�.�. [for they] afford materials for natural selection to act on. (p. 125)
Indeed, Darwin’s writing on individual differences kindled interest in research on heredity by his half cousin, Francis Galton. In the course of his efforts to explore and quantify individual differences between people, Galton became an extremely influential contributor to the field of measurement (Forrest, 1974). Galton (1869) aspired to classify people “according to their natural gifts” (p. 1) and to ascertain their “deviation from an average” (p. 11). Along the way, Galton would be credited with devising or contributing to the development of many contemporary tools of psychological assessment, including questionnaires, rating scales, and self-report inventories.
Galton’s initial work on heredity was done with sweet peas, in part because there tended to be fewer variations among the peas in a single pod. In this work Galton pioneered the use of a statistical concept central to psychological experimentation and testing: the coefficient of correlation. Although Karl Pearson (1857–1936) developed the product-moment correlation technique, its roots can be traced directly to the work of Galton (Magnello & Spies, 1984). From heredity in peas, Galton’s interest turned to heredity in humans and various ways of measuring aspects of people and their abilities.
At an exhibition in London in 1884, Galton displayed his Anthropometric Laboratory, where for a few pence you could be measured on variables such as height (standing), height (sitting), arm span, weight, breathing capacity, strength of pull, strength of squeeze, swiftness of blow, keenness of sight, memory of form, discrimination of color, and steadiness of hand. Through his own efforts and his urging of educational institutions to keep anthropometric records on their students, Galton excited widespread interest in the measurement of psychology-related variables.
Assessment was also an important activity at the first experimental psychology laboratory, founded at the University of Leipzig in Germany by Wilhelm Max Wundt (1832–1920), a medical doctor whose title at the university was professor of philosophy. Wundt and his students tried to formulate a general description of human abilities with respect to variables such as reaction time, perception, and attention span. In contrast to Galton, Wundt focused on how people were similar, not different. In fact, Wundt viewed individual differences as a frustrating source of error in experimentation, and he attempted to control all extraneous variables in an effort to reduce error to a minimum. As we will see, such attempts are fairly routine in contemporary assessment. The objective is to ensure that any observed differences in performance are indeed due to differences between the people being measured and not to any extraneous variables. Manuals for the administration of many tests provide explicit instructions designed to hold constant or “standardize” the conditions under which the test is administered. This is so that any differences in scores on the test are due to differences in the testtakers rather than to differences in the conditions under which the test is administered. In Chapter 4, we will elaborate on the meaning of terms such as standardized and standardization as applied to tests.
In spite of the prevailing research focus on people’s similarities, one of Wundt’s students at Leipzig, an American named James McKeen Cattell (Figure 2–1), completed a doctoral dissertation that dealt with individual differences—specifically, individual differences in reaction time. After receiving his doctoral degree from Leipzig, Cattell returned to the United States, teaching at Bryn Mawr and then at the University of Pennsylvania, before leaving for Europe to teach at Cambridge. At Cambridge, Cattell came in contact with Galton, whom he later described as “the greatest man I have known” (Roback, 1961, p. 96).
J U S T T H I N K � . � . � .
Which orientation in assessment research appeals to you more, the Galtonian orientation (researching how individuals di�er) or the Wundtian (researching how individuals are the same)? Why? Do you think researchers arrive at similar conclusions despite these two contrasting orientations?
coh37025_ch02_041-084.indd 43 12/01/21 4:04 PM
�����Part 1: An Overview
Inspired by his interaction with Galton, Cattell returned to the University of Pennsylvania in 1888 and coined the term mental test in an 1890 publication. Boring (1950, p. 283) noted that “Cattell more than any other person was in this fashion responsible for getting mental testing underway in America, and it is plain that his motivation was similar to Galton’s and that he was influenced, or at least reinforced, by Galton.” Cattell went on to become professor and chair of the psychology department at Columbia University. Over the next 26 years, he not only trained many psychologists but also founded a number of publications (such as the Psychological Review, Science, and American Men of Science). In 1921, Cattell was instrumental in founding the Psychological Corporation, which named 20 of the country’s leading psychologists as its directors. The goal of the corporation was the “advancement of psychology and the promotion of the useful applications of psychology.”2
Other students of Wundt at Leipzig included Charles Spearman, Victor Henri, Emil Kraepelin, E. B. Titchener, G. Stanley Hall, and Lightner Witmer. Spearman is credited with originating the concept of test reliability as well as building the mathematical framework for the statistical technique of factor analysis. Victor Henri was the Frenchman who collaborated with Alfred Binet on papers suggesting how mental tests could be used to measure higher mental processes (e.g., Binet & Henri, 1895a, 1895b, 1895c). Psychiatrist Emil Kraepelin was an early experimenter with the word association technique as a formal test (Kraepelin, 1892, 1895). Lightner Witmer received his Ph.D. from Leipzig and went on to succeed Cattell as director of the psychology laboratory at the University of Pennsylvania. Witmer is cited as the “little-known founder of clinical psychology” (McReynolds, 1987), owing at least in part to his being challenged to treat a “chronic bad speller” in March of 1896 (Brotemarkle, 1947). Later that year Witmer founded the first psychological clinic in the United States at the University of Pennsylvania. In 1907, Witmer founded the journal Psychological Clinic. The first article in that journal was entitled “Clinical Psychology” (Witmer, 1907).
The Twentieth Century
Much of the nineteenth-century testing that could be described as psychological in nature involved the measurement of sensory abilities, reaction time, and the like. Generally the public
2. Today, many of the products and services of what was once known as the Psychological Corporation have been absorbed under the “PsychCorp” brand of a corporate parent, Pearson Assessment, Inc.
Figure �–� James McKeen Cattell (����–����).
The psychologist who is credited with coining the term “mental test” is James McKeen Cattell. Among his many accomplishments, Cattell was a founding member of the American Psychological Association and that organization’s fourth president. JHU Sheridan Libraries/Gado/Archive Photos/Getty Images
coh37025_ch02_041-084.indd 44 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
was fascinated by such testing. However, there was no widespread belief that testing for variables such as reaction time had any applied value. But all of that changed in the early 1900s with the birth of the first formal tests of intelligence. These were tests that were useful for reasons readily understandable to anyone who had school-age children. Public receptivity to psychological tests would shift from mild curiosity to outright enthusiasm as more and more instruments that purportedly quantified mental ability were introduced. Soon there were tests to measure sundry mental characteristics such as personality, interests, attitudes, values, and widely varied mental abilities. It all began with a single test designed for use with young Paris pupils.
The measurement of intelligence� As early as 1895, Alfred Binet (1857–1911) and his colleague Victor Henri published several articles in which they argued for the measurement of abilities such as memory and social comprehension. Ten years later, Binet and collaborator Theodore Simon published a 30-item “measuring scale of intelligence” designed to help identify Paris schoolchildren with intellectual disability (Binet & Simon, 1905). The Binet test would subsequently go through many revisions and translations—and, in the process, launch both the intelligence testing movement and the clinical testing movement. Before long, psychological tests were being used with regularity in such diverse settings as schools, hospitals, clinics, courts, reformatories, and prisons (Pintner, 1931).
In 1939 David Wechsler, a clinical psychologist at Bellevue Hospital in New York City, introduced a test designed to measure adult intelligence. For Wechsler, intelligence was “the aggregate or global capacity of the individual to act purposefully, to think rationally, and to deal effectively with his environment” (Wechsler, 1939, p. 3). Originally christened the Wechsler-Bellevue Intelligence Scale, the test was subsequently revised and renamed the Wechsler Adult Intelligence Scale (WAIS). The WAIS has been revised several times since then, and versions of Wechsler’s test have been published that extend the age range of testtakers from early childhood through senior adulthood.
A natural outgrowth of the individually administered intelligence test devised by Binet was the group intelligence test. Group intelligence tests came into being in the United States in response to the military’s need for an efficient method of screening the intellectual ability of World War I recruits. This same need again became urgent as the United States prepared for entry into World War II. Psychologists would again be called upon by the government service to develop group tests, administer them to recruits, and interpret the test data.
After the war, psychologists returning from military service brought back a wealth of applied testing skills that would be useful in civilian as well as governmental applications. Psychological tests were increasingly used in diverse settings, including large corporations and private organizations. New tests were being developed at a brisk pace to measure various abilities and interests as well as personality.
The measurement of personality�Public receptivity to tests of intellectual ability spurred the development of many other types of tests (Garrett & Schneck, 1933; Pintner, 1931). Only eight years after the publication of Binet’s scale, the field of psychology was being criticized for being too test oriented (Sylvester, 1913). By the late 1930s, approximately 4,000 different psychological tests were in print (Buros, 1938), and “clinical psychology” was synonymous with “mental testing” (Institute for Juvenile Research, 1937; Tulchin, 1939).
World War I had brought with it not only the need to screen the intellectual functioning of recruits but also the need to screen for recruits’ general adjustment. A governmental Committee on Emotional Fitness chaired by psychologist Robert S. Woodworth was assigned
J U S T T H I N K � . � . � .
In the early ����s, the Binet test was being used worldwide for various purposes far beyond identifying exceptional Paris schoolchildren. What were some of the other uses of the test? How appropriate do you think it was to use this test for these other purposes?
J U S T T H I N K � . � . � .
Should the de�nition of intelligence change as one moves from infancy through childhood, adolescence, adulthood, and late adulthood?
coh37025_ch02_041-084.indd 45 12/01/21 4:04 PM
�����Part 1: An Overview
the task of developing a measure of adjustment and emotional stability that could be administered quickly and efficiently to groups of recruits. The committee developed several experimental versions of what were, in essence, paper-and-pencil psychiatric interviews. To disguise the true purpose of one such test, the questionnaire was labeled as a “Personal Data Sheet.” Draftees and volunteers were asked to indicate yes or no to a series of questions that probed for the existence of various kinds of psychopathology. For example, one of the test questions was, “Are you troubled with the idea that people are watching you on the street?”
The Personal Data Sheet developed by Woodworth and his colleagues never went beyond the experimental stages, for the treaty of peace rendered the development of this and other tests less urgent. After the war, Woodworth developed a personality test for civilian use that was based on the Personal Data Sheet. He called it the Woodworth Psychoneurotic Inventory. This instrument was the first widely used self-report measure of personality. In general, self-report refers to a
process whereby assessees themselves supply assessment-related information by responding to questions, keeping a diary, or self- monitoring thoughts or behaviors.
Personality tests that employ self-report methodologies have both advantages and disadvantages. On the face of it, respondents are arguably the best-qualified people to provide answers about themselves. However, there are also compelling arguments against respondents supplying such information. For example, respondents may have poor insight into themselves. People might honestly
believe some things about themselves that in reality are not true. And regardless of the quality of their insight, some respondents are unwilling to reveal anything about themselves that is personal or that could show them in a negative light. Given these shortcomings of the self-report method of personality assessment, there was a need for alternative types of personality tests.
Various methods were developed to provide measures of personality that did not rely on self-report. One such method or approach to personality assessment came to be described as projective in nature. A projective test is one in which an individual is assumed to “project” onto some ambiguous stimulus his or her own unique needs, fears, hopes, and motivation. The ambiguous stimulus might be an inkblot, a drawing, a photograph, or something else. Perhaps the best known of all projective tests is the Rorschach, a series of inkblots developed by the Swiss psychiatrist Hermann Rorschach. The use of pictures as projective stimuli was popularized in the late 1930s by Henry A. Murray, Christiana D. Morgan, and their colleagues
at the Harvard Psychological Clinic. When pictures or photos are used as projective stimuli, respondents are typically asked to tell a story about the picture they are shown. The stories told are then analyzed in terms of what needs and motivations the respondents may be projecting onto the ambiguous pictures. Projective and many other types of instruments used in personality assessment will be discussed in Chapter 12.
The academic and applied traditions�Like the development of its parent field, psychology, the development of psychological measurement can be traced along two distinct threads: the academic and the applied. In the tradition of Galton, Wundt, and other scholars, researchers at universities throughout the world use the tools of assessment to help advance knowledge and understanding of human and animal behavior. Yet there is also an applied tradition, one that dates at least back to ancient China and the examinations developed there to help select applicants for various positions on the basis of merit. Today, society relies on the tools of psychological assessment to help answer important questions. Who is best for this job? In which class should this child be placed? Who is competent to stand trial? Tests and other tools of assessment, when used in a competent manner, can help provide answers.
J U S T T H I N K � . � . � .
Describe an ideal situation for obtaining personality-related information by means of self-report. In what type of situation might it be inadvisable to rely solely on an assessee’s self-report?
J U S T T H I N K � . � . � .
What potential problems do you think might attend the use of picture story-telling tests to assess personality?
coh37025_ch02_041-084.indd 46 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
Contemporary test users hold a keen appreciation for the role of culture in the human experience. Whether in academic or applied settings, assessment professionals recognize the need for cultural sensitivity in the development and use of the tools of psychological assessment. In what follows, we briefly overview some of the issues that such cultural sensitivity entails.
Culture and Assessment
Culture is defined as “the socially transmitted behavior patterns, beliefs, and products of work of a particular population, community, or group of people” (Cohen, 1994, p. 5). As taught to us by parents, peers, and societal institutions such as schools, culture prescribes many behaviors and ways of thinking. Spoken language, attitudes toward elders, and techniques of child rearing are but a few critical manifestations of culture. Culture teaches specific rituals to be performed at birth, marriage, death, and other momentous occasions. Culture imparts much about what is to be valued or prized as well as what is to be rejected or despised. Culture teaches a point of view about what it means to be born of one or another gender, race, or ethnic background. Culture teaches us something about what we can expect from other people and what we can expect from ourselves. Indeed, the influence of culture on an individual’s thoughts and behavior may be a great deal stronger than most of us would acknowledge at first blush.
Professionals involved in the assessment enterprise have shown increasing sensitivity to the role of culture in many different aspects of measurement. This sensitivity is manifested in greater consideration of cultural issues with respect to every aspect of test development and use, including decision making on the basis of test data. Unfortunately, it was not always that way.
Evolving Interest in Culture-Related Issues
Soon after Alfred Binet introduced intelligence testing in France, the U.S. Public Health Service began using such tests to measure the intelligence of people seeking to immigrate to the United States (Figure 2–2). Henry H. Goddard, who had been highly instrumental in getting Binet’s test adopted for use in various settings in the United States, was the chief researcher assigned to the project. Early on, Goddard raised questions about how meaningful such tests are when used with people from various cultural and language backgrounds. Goddard (1913) used interpreters in test administration, employed a bilingual psychologist, and administered mental tests to selected immigrants who appeared to have intellectual disability to trained observers. Although seemingly sensitive to cultural issues in assessment, Goddard’s legacy with regard to such sensitivity is, at best, controversial. Goddard found most immigrants from various nationalities to be mentally deficient when tested. In one widely quoted report, 35 Jews, 22 Hungarians, 50 Italians, and 45 Russians were selected for testing among the masses of immigrants being processed for entry into the United States at Ellis Island. Reporting on his findings in a paper entitled “Mental Tests and the Immigrant,” Goddard (1917) concluded that, in this sample, 83% of the Jews, 80% of the Hungarians, 79% of the Italians, and 87% of the Russians were feebleminded. Although Goddard had written extensively on the genetic nature of mental deficiency, it is to his credit that he did not summarily conclude that these test findings were the result of hereditary. Rather, Goddard (1917) wondered aloud whether the findings were due to “hereditary defect” or “apparent defect due to deprivation” (p. 243). In reality, the findings were largely the result of using a translated Binet test that overestimated mental deficiency in native English-speaking populations, let alone immigrant populations (Terman, 1916).
J U S T T H I N K � . � . � .
Can you think of one way in which you are a product of your culture? How about one way this fact might come through on a psychological test?
coh37025_ch02_041-084.indd 47 12/01/21 4:04 PM
�����Part 1: An Overview
Goddard’s research, although leaving much to be desired methodologically, fueled the fires of an ongoing nature–nurture debate about what intelligence tests actually measure. On one side were those who viewed intelligence test results as indicative of some underlying native ability. On the other side were those who viewed such data as indicative of the extent to which
knowledge and skills had been acquired. More details about the highly influential Henry Goddard and his most controversial career are presented in this chapter’s Close-Up.
If language and culture did indeed have an effect on mental ability test scores, then how could a less confounded or “pure” measure of intelligence be obtained? One way that early test
J U S T T H I N K . � . � .
What safeguards must be �rmly in place before meaningful psychological testing with immigrants can take place?
Figure �–� Psychological testing at Ellis Island.
Immigrants coming to America via Ellis Island were greeted not only by the Statue of Liberty, but also by immigration officials ready to evaluate them with respect to physical, mental, and other variables. Here, a block design test, one measure of intelligence, is administered to a would-be American. Immigrants who failed physical, mental, or other tests were returned to their country of origin at the expense of the shipping company that had brought them. Critics would later charge that at least some of the immigrants who had fared poorly on mental tests were sent away from our shores not because they were actually mentally deficient but simply because they did not understand English well enough to follow instructions. Critics also questioned the criteria on which these immigrants from many lands were being evaluated. Everett Collection Inc./Alamy Stock Photo
coh37025_ch02_041-084.indd 48 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
psychology laboratories in Europe. It is a matter of historical interest that on this tour he did not visit Binet at the Sorbonne in Paris. Rather, it happened that a Belgian psychologist (Ovide Decroly) informed Goddard of Binet’s work and gave him a copy of the Binet-Simon Scale. Few people at the time could appreciate just how momentous the Decroly–Goddard meeting would be nor how in�uential Goddard would become in terms of launching the testing movement. Returning to New Jersey, Goddard oversaw the translation of Binet’s test and distributed thousands of copies of it to professionals working in various settings. Before long, Binet’s test would be used in schools, hospitals, and clinics to help make diagnostic and treatment decisions. The military would use the test, as well as other newly created intelligence tests, to screen recruits. Courts would even begin to mandate the use of intelligence tests to aid in making determinations as to the intelligence of criminal defendants. Such uses of psychological tests were very “cutting edge” at the time.
C L O S E � U P
The Controversial Career of Henry Herbert Goddard
orn to a devout Quaker family in Maine, Henry Herbert Goddard (����–����) was the �fth and youngest child born to farmer Henry Clay Goddard and Sarah Winslow Goddard. The elder Goddard was gored by a bull and succumbed to the injuries he sustained when young Henry was �. Sarah would subsequently marry a missionary, and she and her new husband would travel the United States and abroad preaching. Young Henry attended boarding school at Oak Grove Seminary in Maine and the Friends School in Providence, Rhode Island. After earning his bachelor’s degree from Haverford College, a Quaker-founded school just outside of Philadelphia, he set o� to California to visit an older sister. While there, he accepted a temporary teaching post at the University of Southern California (USC) that included coaching the school’s football team. And so it came to pass that, among Herbert H. Goddard’s many lifelong achievements, he could list the distinction of being USC’s �rst football coach (along with a co-coach; see Pierson, ����).
Goddard returned to Haverford in ���� to earn a master’s degree in mathematics and then took a position as a teacher, principal, and prayer service conductor at a small Quaker school in Ohio. In August of that year, he married Emma Florence Robbins; the couple never had children. Goddard enrolled to study psychology at Clark University and by ���� had earned a doctorate under G. Stanley Hall. Goddard’s doctoral dissertation, a blending of his interests in faith and science, was entitled, “The Effects of Mind on Body as Evidenced in Faith Cures.”
Goddard became a professor at the State Normal School in West Chester, Pennsylvania, a teacher’s college, where he cultivated an interest in the growing child-welfare movement. As a result of his interest in studying children, Goddard had occasion to meet Edward Johnstone, the superintendent of the New Jersey Home for Feeble-Minded Children in Vineland, New Jersey. In ����, Goddard and Johnstone, along with educator Earl Barnes, founded a “Feebleminded Club,” which—despite its misleading name by current standards—served as an interdisciplinary forum for the exchange of ideas regarding special education. By ����, Goddard felt frustrated in his teaching position. His friend Johnstone created the position of Director of Psychological Research at the Vineland facility and so Goddard moved to New Jersey.
In ����, with a newfound interest in the study of “feeblemindedness” (mental de�ciency), Goddard toured
B
(continued)
Fine Art Images/Heritage Images/Hulton Archive/Getty Images
coh37025_ch02_041-084.indd 49 12/01/21 4:04 PM
�����Part 1: An Overview
At the Vineland facility, Goddard found that Binet’s test appeared to work well in terms of quantifying degrees of mental de�ciency. Goddard devised a system of classifying assessees by their performance on the test, coining the term moron and using other such terms that today are out of favor and not in use. Goddard fervently believed that one’s placement on the test was revealing in terms of many facets of one’s life. He believed intelligence tests held the key to answers to questions about everything from what job one should be working at to what activities could make one happy. Further, Goddard came to associate low intelligence with many of the day’s most urgent social problems, ranging from crime to unemployment to poverty. According to him, addressing the problem of low intelligence was a prerequisite to addressing prevailing social problems.
Although previously disposed to believing that mental deficiency was primarily the result of environmental factors, Goddard’s perspective was radically modified by exposure to the views of biologist Charles Davenport. Davenport was a strong believer that heredity played a role in mental deficiency and was a staunch advocate of eugenics, the science of improving the qualities of a breed (in this case, humans) through intervention with factors related to heredity. Davenport collaborated with Goddard in collecting hereditary information on children at the Vineland school. At Davenport’s urgings, the research included a component whereby a “eugenic field worker,” trained to identify mentally deficient individuals, would be sent out to research the mental capabilities of relatives of the residents of the Vineland facility.
The data Goddard and Davenport collected were used to argue the case that mental deficiency was caused by a recessive gene and could be inherited, much like eye color is inherited. Consequently, Goddard believed that—in the interest of the greater good of society at large—mentally deficient individuals should be segregated or institutionalized (at places such as Vineland) and not be permitted to reproduce. By publicly advocating this view, Goddard, along with Edward Johnstone, “transformed their obscure little institution in rural New Jersey into a center of international influence—a model school famous for its advocacy of special education, scientific research, and social reform” (Zenderland, ����, p. ���).
Goddard traced the lineage of one of his students at the Vineland school back �ve generations in his �rst (and most
famous) book, The Kallikak Family: A Study in the Heredity of Feeble-Mindedness (����). In this book Goddard sought to prove how the hereditary “menace of feeble-mindedness” manifested itself in one New Jersey family. “Kallikak” was the �ctional surname given to the Vineland student, Deborah, whose previous generations of relatives were from distinctly “good” (from the Greek kalos) or “bad” (from the Greek kakos)�genetic inheritance. The book traced the family lineages�resulting from the legitimate and illegitimate unions of�a Revolutionary War soldier given the pseudonym “Martin Kallikak.” Martin had fathered children both with a mentally defective waitress and with the woman he married—the latter being a socially prominent and reportedly normal (intellectually) Quaker. Goddard determined that feeblemindedness ran in the�line of descendants from the illegitimate tryst with the waitress. Deborah Kallikak was simply the latest descendant in�that line of descendants to manifest that trait. By contrast, the line of descendants from Martin and his wife contained primarily �ne citizens. But how did Goddard come to this conclusion?
One thing Goddard did not do was administer the Binet to all of the descendants on both the “good” and the “bad” sides of Martin Kallikak’s lineage over the course of some ��� years. Instead Goddard employed a crude case study approach ranging from analysis of o�cial records and documents (which�tended to be scarce) to reports of neighbors (later characterized by critics as unreliable gossip). Conclusions regarding the feeblemindedness of descendants were likely to be linked to any evidence of alcoholism, delinquency, truancy, criminality, prostitution, illegitimacy, or economic dependence. Some of Martin Kallikak’s descendants, alive at the time the�research was being conducted, were classi�ed as feebleminded solely on the basis of their physical appearance. Goddard (����) wrote, for example:
The girl of twelve should have been at school, according to the law, but when one saw her face, one realized that it made no di�erence. She was pretty, with olive complexion and dark, lan- guid eyes, but there was no mind there. (pp. ��–��)
Although well received by the public, the lack of sophistication in the book’s research methodology was a cause for concern for many professionals. In particular, psychiatrist Abraham Myerson (����) attacked the Kallikak study, and the eugenics movement in general, as pseudoscience (see also
C L O S E � U P
The Controversial Career of Henry Herbert Goddard (continued)
coh37025_ch02_041-084.indd 50 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
resulted in the misclassi�cation and consequential repatriation of countless would-be citizens.
Despite an impressive list of career accomplishments, the light of history has not shone favorably on Henry Goddard. Goddard’s (����) recommendation for segregation of the mentally de�cient and his calls for their sterilization tend to be viewed as, at best, misguided. The low esteem in which Goddard is generally held today is perhaps compounded by the fact that Goddard’s work has traditionally been held in high esteem by some groups with radically o�ensive views, such as the Nazi party. During the late ����s and early ����s, more than ��,��� people were euthanized by Nazi physicians simply because they were deemed mentally de�cient. This action preceded the horri�c and systematic mass murder of more than � million innocent civilians by the Nazi military. The alleged “genetic defect” of most of these victims was that they were Jewish. Clearly, eugenicist propaganda fed to the German public was being used by the Nazi party for political gains. The purported goal was to “purify German blood” by limiting or totally eliminating the ability of people from various groups to reproduce.
It is not a matter of controversy that Goddard used ill- advised research methods to derive many of his conclusions; he himself acknowledged this sad fact in later life. At the least Goddard could be criticized for being too easily influenced by the (bad) ideas of others, for being somewhat naive in terms of how his writings were being used, and for not being up to the task of executing methodologically sound�research. Goddard focused on the nature side of the nature–nurture controversy not because he was an ardent eugenicist�at heart but rather because the nature side of the coin was where researchers at the time all tended to focus. Responding to a critic some years later, Goddard (letter to Nicolas Pastore dated April �, ����, quoted in J. D. Smith, ����) wrote, in part, that he had “no inclination to deemphasize environment . . . [but] in those days environment was not being considered.”
The conclusion of Leila Zenderland’s relatively sympathetic biography of Goddard leaves one with the impression that he was basically a decent and likable man who was a product of his times. He harbored neither evil intentions nor right-wing prejudices. For her, a review of the life of Henry Herbert Goddard should serve as a warning not to re�exively jump to the conclusion that “bad science is usually the product of bad motives or, more broadly, bad character” (����, p. ���).
Trent, ����). Myerson reanalyzed data from studies purporting to support the idea that various physical and mental conditions could be inherited, and he criticized those studies on statistical grounds. He especially criticized Goddard for making sweeping and unfounded generalizations from questionable data. Goddard’s book became an increasing cause for concern because it was used (along with related writings on the menace of feeblemindedness) to support radical arguments in favor of eugenics, forced sterilization, restricted immigration, and other social causes. Goddard classi�ed many people as feebleminded based on undesirable social status, illegitimacy, or “sinful” activity. This fact has left some scholars wondering how much Goddard’s own religious upbringing—along with biblical teachings linking children’s problems with parents’ sins—may have been inappropriately emphasized in what was supposed to be strictly scienti�c writing.
After �� years at Vineland, Goddard left under conditions that have been the subject of some speculation (Wehmeyer & Smith, ����). From ���� through ����, Goddard was director of the Ohio Bureau of Juvenile Research. From ���� until his retirement in ����, Goddard was a psychology professor at the Ohio State University. In ���� Goddard moved to Santa Barbara, California, where he lived until his death at the age of ��. His remains were cremated and interred at the Vineland school, along with those of his wife, who had predeceased him in ����.
Goddard’s accomplishments were many. It was largely through his e�orts that state mandates requiring special education services �rst became law. These laws worked to the bene�t of many mentally de�cient as well as many gifted students. Goddard’s introduction of Binet’s test to American society attracted other researchers, such as Lewis Terman, to see what they could do in terms of improving the test for various applications. Goddard’s writings certainly had a momentous heuristic impact on the nature–nurture question. His books and papers stimulated many others to research and write, if only to disprove Goddard’s conclusions. Goddard advocated for court acceptance of intelligence test data into evidence and for the limitation of criminal responsibility in the case of mentally defective defendants, especially with respect to capital crimes. He personally contributed his time to military screening e�orts during World War I. Of more dubious distinction, of course, was the Ellis Island intelligence testing program he set up to screen immigrants. Although ostensibly well intentioned, this e�ort
coh37025_ch02_041-084.indd 51 12/01/21 4:04 PM
�����Part 1: An Overview
developers attempted to deal with the impact of language and culture on tests of mental ability was, in essence, to “isolate” the cultural variable. So-called culture-specific tests, or tests designed for use with people from one culture but not from another, soon began to appear on the scene. Representative of the culture-specific approach to test development were early versions of some of the best-known tests of intelligence. For example, the 1937 revision of the Stanford- Binet Intelligence Scale, which enjoyed widespread use until it was revised in 1960, included no racially, ethnically, socioeconomically, or culturally diverse children in the research that went into its formulation. Similarly, the Wechsler-Bellevue Intelligence Scale, forerunner of a widely used measure of adult intelligence, contained no racially, ethnically, socioeconomically, or culturally diverse members in the samples of testtakers used in its development. Although “a large number” of Blacks had, in fact, been tested (Wechsler, 1944), those data had been omitted from the final test manual because the test developers “did not feel that norms derived by mixing the populations could be interpreted without special provisos and reservations.” Hence, Wechsler (1944) stated at the outset that the Wechsler-Bellevue norms could not be used for “the colored
populations of the United States.” In like fashion, the inaugural edition of the Wechsler Intelligence Scale for Children (WISC), first published in 1949 and not revised until 1974, contained no racially, ethnically, socioeconomically, or culturally diverse children in its development.
Even though many published tests were purposely designed to be culture-specific, it soon became apparent that the tests were being administered—improperly—to people from different cultures. Perhaps not surprisingly, racially, ethnically,
socioeconomically, or culturally diverse testtakers tended to score lower as a group than people from the group for whom the test was developed. Illustrative of the type of problems encountered by test users was this item from the 1949 WISC: “If your mother sends you to the store for a loaf of bread and there is none, what do you do?” Many Hispanic children were routinely sent to the store for tortillas and so were not familiar with the phrase “loaf of bread.”
Today test developers typically take many steps to ensure that a major test developed for national use is indeed suitable for such use. Those steps might involve administering a preliminary version of the test to a tryout sample of testtakers from various cultural backgrounds, particularly from those whose members are likely to be administered the final version of the test. Examiners who administer the test may be asked to describe their impressions with regard to various aspects of testtakers’ responses. For example, subjective impressions regarding testtakers’ reactions to the test materials or opinions regarding the clarity of instructions will be noted. All of the accumulated test scores from the tryout sample will be analyzed to determine if any individual item seems to be biased with regard to race, gender, or culture. In addition, a panel of independent reviewers may be asked to go through the test items and screen them for possible bias. A revised version of the test may then be administered to a large sample of testtakers that is representative of key variables of the latest U.S. Census data (such as age, gender, ethnic background, and socioeconomic status). Information from this large-scale test administration will also be used to root out any identifiable sources of bias, often using sophisticated statistical techniques designed for this purpose. More details regarding the contemporary process of test development will be presented in Chapter 8.
Some Issues Regarding Culture and Assessment
Communication between assessor and assessee is a most basic part of assessment. Assessors must be sensitive to any differences between the language or dialect familiar to assessees and the language in which the assessment is conducted. Assessors must also be sensitive to the degree to which assessees have been exposed to the dominant culture and the extent to which they have
J U S T T H I N K . � . � .
Try your hand at creating one culture-speci�c test item on any subject. Testtakers from what culture would probably succeed in responding correctly to the item? Testtakers from what culture would not?
coh37025_ch02_041-084.indd 52 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
made a conscious choice to become assimilated. Next, we briefly consider assessment-related issues of communication, both verbal and nonverbal, in a cultural context.
Verbal communication� Language, the means by which information is communicated, is a key yet sometimes overlooked variable in the assessment process. Most obviously, the examiner and the examinee must speak the same language. This common language is necessary not only for the assessment to proceed but also for the assessor’s conclusions regarding the assessment to be reasonably accurate. If a test is in written form and includes written instructions, then the testtaker must be able to read and comprehend what is written. When the language in which the assessment is conducted is not the assessee’s primary language, the assessee may not fully comprehend the instructions or the test items. The danger of such misunderstanding may increase as infrequently used vocabulary or unusual idioms are employed in the assessment. All of the foregoing presumes that the assessee is making a sincere and well-intentioned effort to respond to the demands of the assessment. Although this is frequently presumed, it is not always the case. In some instances, assessees may purposely attempt to use a language deficit to frustrate evaluation efforts (Stephens, 1992).
When an assessment is conducted with the aid of a translator, different challenges may emerge. Depending upon the translator’s skill and professionalism, subtle nuances of meaning may be lost in translation, or unintentional hints to the correct or more desirable response may be conveyed. Whether translated “live” by a translator or in writing, translated items may be either easier or more difficult than the original. Some vocabulary words may change meaning or have dual meanings when translated.
Interpreters may have limited understanding of mental health issues. In turn, an assessor may have little experience in working with a translator. For these reasons, when possible, it is desirable to have some pretraining for interpreters on the relevant issues, and some pretraining for assessors on working with translators (Searight & Searight, 2009).
In interviews or other situations in which an evaluation is made on the basis of a spoken exchange between two parties, a trained examiner may detect through verbal or nonverbal means that the examinee’s grasp of a language or a dialect is too deficient to proceed. A trained examiner might not be able to detect this when the test is in written form. In the case of written tests, it is clearly essential that the examinee be able to read and comprehend what is written. Otherwise the evaluation may be more about language or dialect competency than whatever the test purports to measure. Even when examiner and examinee speak the same language, miscommunication and consequential effects on test results may result owing to differences in dialect (Wolfram, 1971).
In the assessment of an individual whose proficiency in the English language is limited or nonexistent, some basic questions may need to be raised: What level of proficiency in English must the testtaker have, and does the testtaker have that proficiency? Can a meaningful assessment take place through a trained interpreter? Can an alternative and more appropriate assessment procedure be devised to meet the objectives of the assessment? In addition to linguistic barriers, the contents of tests from a particular culture are typically laden with items and material—some obvious, some subtle—that draw heavily from that culture. Test performance may, at least in part, reflect not only whatever variables the test purports to measure but also one additional variable: the degree to which the testtaker has assimilated the culture.
Nonverbal communication and behavior�Humans communicate not only through verbal means but also through nonverbal means. Facial expressions, finger and hand signs, and shifts in one’s position in space may all convey messages. Of course, the messages conveyed by such body language may be different from culture to culture. In American culture, for example, one who fails to look
J U S T T H I N K � . � . � .
What might an assessor do to make sure that a prospective assessee’s language competence is su�cient to administer the test in that language to that assessee?
coh37025_ch02_041-084.indd 53 12/01/21 4:04 PM
�����Part 1: An Overview
another person in the eye when speaking may be viewed as deceitful or having something to hide. However, in other cultures, failure to make eye contact when speaking may be a sign of respect.
If you have ever gone on or conducted a job interview, you may have developed a firsthand appreciation of the value of nonverbal communication in an evaluative setting. Interviewees who show enthusiasm and interest have the edge over interviewees who appear to be drowsy or bored. In clinical settings, an experienced evaluator may develop hypotheses to be tested from the nonverbal behavior of the interviewee. For example, a person who is slouching, moving slowly, and exhibiting a sad facial expression may be depressed. Then again, such an individual may be experiencing physical discomfort from any number of sources, such as a muscle spasm or an arthritis attack. It remains for the assessor to determine which hypothesis best accounts for the observed behavior.
Certain theories and systems in the mental health field go beyond more traditional interpretations of body language. For example, in psychoanalysis, a theory of personality and psychological treatment developed by Sigmund Freud, symbolic significance is assigned to many nonverbal acts. From a psychoanalytic perspective, an interviewee’s fidgeting with a wedding band during an interview may be interpreted as a message regarding an unstable marriage. As evidenced by his thoughts on “the first chance actions” of a patient during a therapy session, Sigmund Freud believed he could tell much about motivation from nonverbal behavior:
The first . . . chance actions of the patient . . . will betray one of the governing complexes of the neurosis. . . . A young girl . . . hurriedly pulls the hem of her skirt over her exposed ankle; she has betrayed the kernel of what analysis will discover later; her narcissistic pride in her bodily beauty and her tendencies to exhibitionism. (Freud, 1913/1959, p. 359)
This quote from Freud is also useful in illustrating the influence of culture on diagnostic and therapeutic views. Freud lived in Victorian Vienna. In that time and in that place, sex was
not a subject for public discussion. In many ways Freud’s views regarding a sexual basis for various thoughts and behaviors were a product of the sexually repressed culture in which he lived.
An example of a nonverbal behavior in which people differ is the speed at which they characteristically move to complete tasks. The overall pace of life in one geographic area, for example, may tend to be faster than in another. In a similar vein, differences in pace of life across cultures may enhance or detract from test scores on tests involving timed items (Gopaul-McNicol, 1993). In a more general sense, Hoffman (1962) questioned the value of timed tests of ability, particularly those tests that employed multiple-choice items. He believed such tests relied too heavily on testtakers’ quickness of response and as such discriminated against the individual who is characteristically a “deep, brooding thinker.”
Culture exerts effects over many aspects of nonverbal behavior. For example, a child may present as noncommunicative
and having only minimal language skills when verbally examined. This finding may be due to the fact that the child is from a culture where elders are revered and where children speak to adults only when they are spoken to—and then only in as short a phrase as possible. Clearly, it is incumbent upon test users to be knowledgeable about aspects of an assessee’s culture that are relevant to the assessment.
Standards of evaluation�Suppose an international contest was held to crown “the best chicken soup in the world.” Who do you think would win? The answer to that question hinges on the evaluative standard to be employed. If the sole judge of the contest was the owner of a kosher delicatessen on the Lower East Side of Manhattan, it is conceivable that the entry that came
J U S T T H I N K . � . � .
Play the role of a therapist in the Freudian tradition and cite one example of a student’s or an instructor’s public behavior that you believe may be telling about that individual’s private motivation. No naming names!
J U S T T H I N K . � . � .
What type of test is best suited for administration to people who are “deep, brooding thinkers”? How practical for group administration would such tests be?
coh37025_ch02_041-084.indd 54 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
closest to the “Jewish mother homemade” variety might well be declared the winner. However, other judges might have other standards and preferences. For example, soup connoisseurs from Arabic cultures might prefer chicken soup with fresh lemon juice in the recipe. Judges from India might be inclined to give their vote to a chicken soup flavored with coriander and cumin. For Japanese and Chinese judges, soy sauce might be viewed as an indispensable ingredient. Ultimately, the judgment of which soup is best will probably be very much a matter of personal preference and the standard of evaluation employed.
Somewhat akin to judgments concerning the best chicken soup recipe, judgments related to certain psychological traits can also be culturally relative. For example, whether specific patterns of behavior are considered to be male- or female-appropriate will depend on the prevailing societal standards regarding masculinity and femininity. In some societies, for example, it is role- appropriate for women to fight wars and put food on the table while the men are occupied in more domestic activities. Whether specific patterns of behavior are considered to be psychopathological also depends on the prevailing societal standards. In Sudan, for example, there are tribes that live among cattle because they regard the animals as sacred. Judgments as to who might be the best employee, manager, or leader may differ as a function of culture, as might judgments regarding intelligence, wisdom, courage, and other psychological variables.
Cultures differ from one another in the extent to which they are individualist or collectivist (Markus & Kitayama, 1991). Generally speaking, an individualist culture (typically associated with the dominant culture in countries such as the United States and Great Britain) is characterized by value being placed on traits such as self-reliance, autonomy, independence, uniqueness, and competitiveness. In a collectivist culture (typically associated with the dominant culture in many countries throughout Asia, Latin America, and Africa), value is placed on traits such as conformity, cooperation, interdependence, and striving toward group goals. As a consequence of being raised in one or another of these types of cultures, people may develop certain characteristic aspects of their sense of self. Markus and Kitayama (1991) believe that people raised in Western culture tend to see themselves as having a unique constellation of traits that are stable over time and through situations. The person raised in an individualist culture exhibits behavior that is “organized and made meaningful primarily by reference to one’s own internal repertoire of thoughts, feelings, and action, rather than by reference to the thoughts, feelings, and actions of others” (Markus & Kitayama, 1991, p. 226). By contrast, people raised in a collectivist culture see themselves as part of a larger whole, with much greater connectedness to others. And rather than seeing their own traits as stable over time and through situations, the person raised in a collectivist culture believes that “one’s behavior is determined, contingent on, and, to a large extent organized by what the actor perceives to be the thoughts, feelings, and actions of others in the relationship” (Markus & Kitayama, 1991, p. 227, emphasis in the original).
Consider in a clinical context, for example, a psychiatric diagnosis of dependent personality disorder. To some extent the description of this disorder reflects the values of an individualist culture in deeming overdependence on others to be pathological. Yet the clinician making such a diagnosis would, ideally, be aware that such a belief foundation is contradictory to a guiding philosophy for many people from a collectivist culture wherein dependence and submission may be integral to fulfilling role obligations (Chen et al., 2009). In the workplace, individuals from collectivist cultures may be penalized in some performance ratings because they are less likely to attribute success in their jobs to themselves. Rather, they are more likely to be self-effacing and self-critical (Newman et al., 2004). The point is clear: Cultural differences carry with them important implications for assessment.
A challenge inherent in the assessment enterprise concerns tempering test- and assessment-related outcomes with good judgment regarding the cultural relativity of those outcomes. In practice, this means raising questions about the applicability of assessment-related findings to specific individuals.
J U S T T H I N K � . � . � .
When considering tools of evaluation that purport to measure the trait of assertiveness, what are some culture-related considerations that should be kept in mind?
coh37025_ch02_041-084.indd 55 12/01/21 4:04 PM
�����Part 1: An Overview
M E E T A N A S S E S S M E N T P R O F E S S I O N A L
stereotyping patients based on group identities (such as race or ethnicity). The tool of assessment that I use in clinical practice is the DSM-� core Cultural Formulation Interview (CFI). The CFI consists of �� questions, and is based on a comprehensive literature review of ��� publications in seven languages. Field tested with ��� patients by �� clinicians in six countries, the CFI has been revised through patient and clinician feedback (Lewis- Fernández et al., ����). The �� questions cover topics of enduring interest in mental health such as patients’ explanations of illness (de�nitions for their presenting problem, preferred idiomatic terms, level of severity, causes), perceived social stressors and supports, the role of cultural identity in their lives and in relation to the presenting problem, individual coping mechanisms, past help-seeking behaviors, personal barriers to care, current expectations of treatment, and potential di�erences between patients and clinicians that can impact rapport. In recognition of this instrument’s scienti�c value, the American Psychiatric Association has made the CFI available to all users. It may be accessed, free-of- charge, at https://www.psychiatry.org/File%�� Library/Psychiatrists/Practice/DSM/APA_DSM�_ Cultural-Formulation-Interview.pdf.
Meet Dr. Neil Krishan Aggarwal
ultural assessment informs every aspect of my work, from the medical students and psychiatry resident trainees whom I teach at C.U., the mental health clinicians whom I train to conduct culturally competent interviews with patients for my research at N.Y.S.P.I., and the patients I treat in private practice. The fact that an understanding of culture is essential to understand all aspects of mental health has been recognized increasingly over the years by the American Psychiatric Association in its Diagnostic and Statistical Manual (DSM).
In my subspecialty of cultural psychiatry, it has long been recognized that culture in�uences when, where, how, and to whom patients narrate their experiences of distress, the patterning of symptoms recognized as illnesses, and the models clinicians use to interpret symptoms through diagnoses (Kirmayer, ����; Kleinman, ����). Culture also shapes perceptions of care such as expectations around appropriate healers (medical or non-medical), the duration and types of acceptable treatments, and anticipated improvements in quality of life (Aggarwal, Pieh, et al., ����). The American Psychiatric Association and the American Psychological Association now have professional guidelines that encourage cultural competence training for all clinicians with the recognition that all patients—not just those from racial or ethnic minority groups—have cultural concerns that impact diagnosis and treatment. Despite the growing appreciation that cultural competence training for clinicians can reduce disparities in treatment (O�ce of the Surgeon General, ����), many well- intentioned clinicians are too often trained only in making a diagnosis, developing a treatment plan, or administering therapies without systematically re�ecting on a patient’s cultural needs.
Mental health clinicians need an assessment tool that comprehensively accounts for all relevant cultural factors in su�cient depth, and can be used in a standardized way in diverse clinical settings with di�erent populations. Ideally, such an instrument would be focused on the cultural identity of the individual patient, the better to avoid the risk of
C
Neil Krishan Aggarwal, M.D., M.A., Assistant Professor of Clinical Psychiatry at Columbia University (C.U.), Research Psychiatrist at the New York State Psychiatric Institute (N.Y.S.P.I.), and psychiatrist in private practice.
Neil Krishan Aggarwal, M.D., M.A.
coh37025_ch02_041-084.indd 56 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
are to be raised, in order, during the initial intake interview (prior to taking the medical or psychiatric history). Sometimes this feels too rigid, especially when a patient’s responses to CFI questions seem to naturally lead to questions about the medical or psychiatric history. Second, some patients in acute illness cannot answer the questions. For example, people with acute substance intoxication, psychosis, or cognition-impairing conditions such as Alzheimer’s or Parkinson’s disease cannot always answer questions directly. Finally, the CFI builds from the meaning-centered approach to culture in medical anthropology that mostly relies on patient interviews (Lewis-Fernández et al., ����). The CFI thus has all of the drawbacks one would expect from a self-report instrument that lacks a behavioral component. Accordingly, the CFI is perhaps best viewed as a beginning, and not an end, to a conversation about culture and mental health with new patients.
The CFI builds from and contributes to an ongoing movement across the health disciplines that patient care should be culturally competent and individually tailored. Today, all clinical stakeholders— patients, clinicians, administrators, families, and health advocates—recognize that cultural assessment is one of the few ways to emphasize the patient’s own narrative of su�ering within a health care environment that has too often prioritized diagnostic assessment and billing considerations. Budding psychiatrists and psychologists can help advance the science and practice of cultural assessments in mental health by using, critiquing, and re�ning standardized instruments such as the CFI. In the continued absence of con�rmatory laboratory or radiological tests that we can order, diagnosis and treatment planning are acts of interpretation in mental health: Patients must �rst interpret their symptoms through the use of language and we must interpret their colloquial language in scienti�c terms (Kleinman, ����). Cultural assessments such as the CFI can remind psychiatrists and psychologists that our own professional cultures—systems of knowledge, concepts, rules, and practices that are learned and transmitted across generations—mold our scienti�c interpretations that may not re�ect the realities of health and illness in our patients’ lives.
Used with permission of Dr. Neil Krishan Aggarwal.
The data derived from a CFI administration can introduce clinicians to fundamental ways that culture and mental health interrelate for the individual patient. Responses can yield important clinical insights as to when, where, how, and to whom patients narrate experiences of illness, and the healers whom they approach for care. It can provide useful information regarding the duration and types of treatments that the individual patient would �nd acceptable. On average, the complete interview takes about �� to �� minutes—well within the time typically allotted for an initial intake session (Aggarwal, Jiménez-Solomon, et al., ����). The use of the CFI can improve health communication as it provides patients with an open-ended opportunity to narrate what is most at stake for them during illness in an open-ended way (Aggarwal et al., ����).
Several versions of the CFI, all based on the core format, are available. These alternative versions are variously designed for use with informants and caregivers, and for use with children, adolescents, older adults, and immigrants and refugees (Lewis- Fernández et al., ����). I particularly �nd useful the CFI supplementary interviews on level of functioning, cultural identity, and spirituality, religion, and moral traditions because they help me better situate patients in their environment.
Consistent with recommendations from the latest version of the DSM (the DSM-�), I use the CFI with all patients whenever I do an initial intake interview. A patient’s responses to the questions can be particularly helpful in formulating a diagnosis when the presenting symptoms seem to di�er from formal DSM criteria. The data may also be instructive with regard to judging impairments in academic, occupational, and social functioning, and in negotiating a treatment plan around the length and types of treatments deemed necessary. Additionally, the data may have value in formulating a treatment plan that is devoid of approaches to therapy, including certain medications, that an individual patient is not predisposed to respond to favorably. In cases where patients develop resistance to therapy protocols, it may be useful to revisit CFI data as a way of reminding patients of what was previously agreed upon, or open a door to renegotiation of the therapeutic contract.
No tool of assessment is perfect, and the CFI certainly has its shortcomings. First, the DSM-� encourages the use of all �� questions. The questions
coh37025_ch02_041-084.indd 57 12/01/21 4:04 PM
�����Part 1: An Overview
It therefore seems prudent to supplement questions such as “How intelligent is this person?” or “How assertive is this individual?” with other questions, such as: “How appropriate are the norms or other standards that will be used to make this evaluation?” “To what extent has the assessee been assimilated by the culture from which the test is drawn, and what influence might such assimilation (or lack of it) have on the test results?” “What research has been done on the test to support the applicability of findings with it for use in evaluating this particular asssessee?” These are the types of questions that are being raised by responsible test users such as this chapter’s guest assessment professional, Dr. Neil Krishan Aggarwal (see Meet an Assessment Professional). They are also the types of questions being increasingly raised in courts of law.
Tests and Group Membership
Tests and other evaluative measures administered in vocational, educational, counseling, and other settings leave little doubt that people differ from one another on an individual basis and also from group to group on a collective basis. What happens when groups systematically differ in terms of scores on a particular test? The answer, in a word, is conflict.
On one hand, questions such as “Which student is best qualified to be admitted to this school?” or “Which job candidate should get the job?” are rather straightforward. On the other hand, societal concerns about fairness both to individuals and to groups of individuals have made the answers to such questions matters of heated debate, if not lawsuits and civil disobedience. Consider the case of a person who happens to be a member of a particular group— cultural or otherwise—who fails to obtain a desired outcome (such as attainment of employment or admission to a university). Suppose it is further observed that most other people from that same group have also failed to obtain that same prized outcome. What may well happen is that the criteria being used to judge attainment of the prized outcome becomes the subject of intense scrutiny, sometimes by a court or a legislature.
In vocational assessment, test users are sensitive to legal and ethical mandates concerning the use of tests with regard to hiring, firing, and related decision making. If a test is used to evaluate a candidate’s ability to do a job, one point of view is that the test should do just that—regardless of the group membership of the testtaker. According to this view, scores on a test of job ability should be influenced only by job-related variables. That is, scores should not be affected by variables such as group membership, hair length, eye color, or any other variable extraneous to the ability to perform the job. Although this rather straightforward view of the role of tests in personnel selection may seem consistent with principles of equal opportunity, it has attracted charges of unfairness and claims of discrimination. Why?
Claims of test-related discrimination made against major test publishers may be best understood as evidence of the great complexity of the assessment enterprise rather than as a conspiracy to use tests to discriminate against individuals from certain groups. In vocational assessment, for example, conflicts may arise from disagreements about the criteria for performing a particular job. The potential for controversy looms over almost all selection criteria that an employer sets, regardless of whether the criteria are physical, educational, psychological, or experiential.
The critical question with regard to hiring, promotion, and other selection decisions in almost any work setting is: “What criteria must be met to do this job?” A state police department may require all applicants for the position of police officer to meet certain physical requirements, including a minimum height of 5 feet 4 inches. A person who is 5 feet 2 inches tall is therefore barred from applying. Because such police force evaluation policies have the effect of systematically excluding members of cultural groups where the average height of adults is less than 5 feet 4 inches, the result may be a class-action lawsuit charging discrimination. Whether the police department’s height requirement is reasonable and job related, and whether discrimination actually occurred, are complex questions that are usually left for the courts to resolve. Compelling arguments may be presented on both sides, as benevolent, fair-minded,
coh37025_ch02_041-084.indd 58 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
knowledgeable, and well-intentioned people may have honest differences about the necessity of the prevailing height requirement for the job of police officer.
Beyond the variable of height, it seems that variables such as appearance and religion should have little to do with what job one is qualified to perform. However, it is precisely such factors that keep some group members from entry into many jobs and careers. Consider in this context observant Jews. Their appearance and dress is not mainstream. The food they eat must be kosher. They are unable to work or travel on weekends. Given the established selection criteria for many positions in corporate America, candidates who are members of the group known as observant Jews are effectively excluded, regardless of their ability to perform the work (Korman, 1988; Mael, 1991; Zweigenhaft, 1984).
General differences among groups of people also extend to psychological attributes such as measured intelligence. Unfortunately, the mere suggestion that such differences in psychological variables exist arouses skepticism if not charges of discrimination, bias, or worse. These reactions are especially true when the observed group differences are deemed responsible for blocking one or another group from employment or educational opportunities.
If systematic differences related to group membership were found to exist on job ability test scores, then what, if anything, should be done? One view is that nothing needs to be done. According to this view, the test was designed to measure job ability, and it does what it was designed to do. In support of this view is evidence suggesting that group differences in scores on professionally developed tests do reflect differences in real- world performance (Gottfredson, 2000; Halpern, 2000; Hartigan & Wigdor, 1989; Kubiszyn et al., 2000; Neisser et al., 1996; Schmidt, 1988; Schmidt & Hunter, 1992).
A contrasting view is that efforts should be made to “level the playing field” between groups of people. The term affirmative action refers to voluntary and mandatory efforts undertaken by federal, state, and local governments, private employers, and schools to combat discrimination and to promote equal opportunity for all in education and employment (American Psychological Association, 1996, p. 2). Affirmative action seeks to create equal opportunity actively, not passively. One impetus to affirmative action is the view that “policies that appear to be neutral with regard to ethnicity or gender can operate in ways that advantage individuals from one group over individuals from another group” (Crosby et al., 2003, p. 95).
In assessment, one way of implementing affirmative action is by altering test-scoring procedures according to set guidelines. For example, an individual’s score on a test could be revised according to the individual’s group membership (McNemar, 1975). While proponents of this approach view such remedies as necessary to address past inequities, others condemn manipulation of test scores as introducing “inequity in equity” (Benbow & Stanley, 1996).
As sincerely committed as they may be to principles of egalitarianism and fair play, test developers and test users must ultimately look to society at large—and, more specifically, to laws, administrative regulations, and other rules and professional codes of conduct—for guidance in the use of tests and test scores.
Psychology, tests, and public policy�Few people would object to using psychological tests in academic and applied contexts that obviously benefit human welfare. Then again, few people are aware of the everyday use of psychological tests in such ways. More typically, members of the general public become acquainted with the use of psychological tests in high-profile contexts, such
J U S T T H I N K . � . � .
What might be a fair and equitable way to determine the minimum required height, if any, for police o�cers in your community?
J U S T T H I N K � . � . � .
What should be done if a test adequately assesses a skill required for a job but is discriminatory?
J U S T T H I N K � . � . � .
What are your thoughts on the manipulation of test scores as a function of group membership to advance certain social goals? Should membership in a particular cultural group trigger an automatic increase (or decrease) in test scores?
coh37025_ch02_041-084.indd 59 12/01/21 4:04 PM
�����Part 1: An Overview
as when an individual or a group has a great deal to gain or to lose as a result of a test score. In such situations, tests and other tools of assessment are portrayed as instruments that can have a momentous and immediate impact on one’s life. In such situations, tests may be perceived by the everyday person as tools used to deny people things they want or need. Denial of educational advancement, dismissal from a job, denial of parole, and denial of custody are some of the more threatening consequences that the public may associate with psychological tests and assessment procedures.
Members of the public call upon government policy-makers to protect them from perceived threats. Legislators pass laws, administrative agencies make regulations, judges hand down rulings, and citizens call for referenda regarding prevailing public policies. In the section that follows, we broaden our view of the assessment enterprise beyond the concerns of the profession. Legal and ethical considerations with regard to assessment are a matter of concern to the public at large.
Legal and Ethical Considerations
Laws are rules that individuals must obey for the good of the society as a whole—or rules thought to be for the good of society as a whole. Some laws are and have been relatively uncontroversial. For example, the law that mandates driving on the right side of the road has not been a subject of debate, a source of emotional soul-searching, or a stimulus to civil disobedience. For safety and the common good, most people are willing to relinquish their freedom to drive all over the road. Even visitors from countries where it is common to drive on the other side of the road will readily comply with this law when driving in the United States.
Although rules of the road may be relatively uncontroversial, there are some laws that are controversial. Consider in this context laws pertaining to abortion, capital punishment, euthanasia, affirmative action, busing . . . the list goes on. Exactly how laws regulating matters like these should be written and interpreted are issues of heated controversy. So too is the role of testing and assessment in such matters.
Whereas a body of laws is a body of rules, a body of ethics is a body of principles of right, proper, or good conduct. Thus, for example, an ethic of the Old West was “Never shoot ‘em in the back.” Two well-known principles subscribed to by seafarers are “Women and children leave first in an emergency” and “A captain goes down with his ship.” The ethics of journalism dictate that reporters present all sides of a controversial issue. A principle of ethical research is that the researcher should never fudge data; all data must be reported accurately.
To the extent that a code of professional ethics is recognized and accepted by members of a profession, it defines the standard of care expected of members of that profession. In this context,
we may define standard of care as the level at which the average, reasonable, and prudent professional would provide diagnostic or therapeutic services under the same or similar conditions.
Members of the public and members of the profession have not always been on “the same side of the fence” with respect to issues of ethics and law. Let’s review how and why this disagreement has been the case.
The Concerns of the Public
The assessment enterprise has never been well understood by the public, and even today we might hear criticisms based on a misunderstanding of testing (e.g., “The only thing tests measure is the ability to take tests”). Possible consequences of public misunderstanding include fear, anger, legislation, litigation, and administrative regulations. In recent years, the testing-related provisions of the No Child Left Behind Act of 2001 (re-authorized in 2015 as the Every Student
J U S T T H I N K . � . � .
List �ve ethical guidelines that you think should govern the professional behavior of psychologists involved in psychological testing and assessment.
coh37025_ch02_041-084.indd 60 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
Succeeds Act or ESSA) and the 2010 Common Core State Standards (jointly drafted and released by the National Governor’s Association Center for Best Practices and the Council of Chief State School Officers) have generated a great deal of controversy. The Common Core State Standards was the product of a state-led effort to bring greater interstate uniformity to what constituted proficiency in various academic subjects. To date, however, Common Core has probably been more at the core of public controversy than anything else. Efforts to dismantle these standards have taken the form of everything from verbal attacks by politicians, to local demonstrations by consortiums of teachers, parents, and students. In Chapter 10, Educational Assessment, we will take a closer look at the pros and cons of Common Core.
Concern about the use of psychological tests first became widespread in the aftermath of World War I, when various professionals (as well as nonprofessionals) sought to adapt group tests developed by the military for civilian use in schools and industry. Reflecting growing public discomfort with the burgeoning assessment industry were popular magazine articles featuring stories with titles such as “The Abuse of Tests” (see Haney, 1981). Less well known were voices of reason that offered constructive ways to correct what was wrong with assessment practices.
The nationwide military testing during World War II in the 1940s did not attract as much popular attention as the testing undertaken during World War I. Rather, an event that took place on the other side of the globe had a far more momentous effect on testing in the United States: the launching of a satellite into space by the country then known as the Union of Soviet Socialist Republics (USSR or Soviet Union). This unanticipated action on the part of a cold-war enemy immediately compounded homeland security concerns in the United States. The prospect of a Russian satellite orbiting Earth 24 hours a day was most unsettling, as it magnified feelings of vulnerability. Perhaps on a positive note, the Soviet launch of Sputnik (the name given to the satellite) had the effect of galvanizing public and legislative opinion around the value of education in areas such as math, science, engineering, and physics. More resources would have to be allocated toward identifying the gifted children who would one day equip the United States to successfully compete with the Soviets.
About a year after the launch of Sputnik, Congress passed the National Defense Education Act, which provided federal money to local schools for the purpose of testing ability and aptitude to identify gifted and academically talented students. This event triggered a proliferation of large-scale testing programs in the schools. At the same time, the use of ability tests and personality tests for personnel selection increased in government, the military, and business. The wide and growing use of tests led to renewed public concern, reflected in magazine articles such as “Testing: Can Everyone Be Pigeonholed?” (Newsweek, July 20, 1959) and “What the Tests Do Not Test” (New York Times Magazine, October 2, 1960). The upshot of such concern was congressional hearings on the subject of testing (Amrine, 1965).
The fires of public concern about testing were again fanned in 1969 when widespread media attention was given to the publication of an article, in the prestigious Harvard Educational Review, entitled “How Much Can We Boost IQ and Scholastic Achievement?” Its author, Arthur Jensen, argued that “genetic factors are strongly implicated in the average Negro–white intelligence difference” (1969, p. 82). What followed was an outpouring of public and professional attention to nature-versus-nurture issues in addition to widespread skepticism about what� intelligence tests were really measuring. By 1972 the U.S. Select Committee on Equal Education Opportunity was preparing for hearings on the matter. However, according to Haney (1981), the hearings “were canceled because they promised to be too controversial” (p. 1026).
The extent of public concern about psychological assessment is reflected in the extensive involvement of the government in many aspects of the assessment process in recent decades. Assessment has been affected in numerous and important ways by activities of the legislative, executive, and judicial branches of federal and state governments. A sampling of some landmark legislation and litigation is presented in Table 2–1.
coh37025_ch02_041-084.indd 61 12/01/21 4:04 PM
�����Part 1: An Overview
Table �–� Some Signi�cant Legislation and Litigation
Legislation Signi�cance
Americans with Disabilities Act of ���� Employment testing materials and procedures must be essential to the job and not discrimi- nate against persons with handicaps.
Civil Rights Act of ���� (amended in ����), also known as the Equal Opportunity Employment Act
It is an unlawful employment practice to adjust the scores of, use di�erent cuto� scores for, or otherwise alter the results of employment-related tests on the basis of race, religion, sex, or national origin.
Family Education Rights and Privacy Act (����)Parents and eligible students must be given access to school records, and have a right to challenge �ndings in records by a hearing.
Health Insurance Portability and Accountability Act of ���� (HIPAA)
New federal privacy standards limit the ways in which health care providers and others can use patients’ personal information.
Education for All Handicapped Children (PL ��-���) (���� and then amended several times thereafter, including IDEA of ���� and ����)
Screening is mandated for children suspected to have mental or physical handicaps. Once iden- ti�ed, an individual child must be evaluated by a professional team quali�ed to determine that child’s special educational needs. The child must be reevaluated periodically. Amended in ���� to extend disability-related protections downward to infants and toddlers.
Individuals with Disabilities Education Act (IDEA) Amendments of ���� (PL ���-��)
Children should not be inappropriately placed in special education programs due to cultural di�erences. Schools should accommodate existing test instruments and other alternate means of assessment for the purpose of gauging the progress of special education students as measured by state- and district-wide assessments.
Every Student Succeeds Act (ESSA) (����) This reauthorization of the Elementary and Secondary Education Act of ����, commonly known as No Child Left Behind (NCLB), was designed to “close the achievement gaps between minority and nonminority students and between disadvantaged children and their more advantaged peers” by, among other things, setting strict standards for school accountability and establishing periodic assessments to gauge the progress of school districts in improving academic achievement. The “battle cry” driving this legislation was “Demographics are not destiny!” However, by ����, it was clear that many, perhaps the majority of states, sought or will seek waivers to opt out of NCLB and what has been viewed as its demanding bureau- cratic structure, and overly ambitious goals.
Hobson v. Hansen (����) U.S. Supreme Court ruled that ability tests developed on whites could not lawfully be used to track Black students in the school system. To do so could result in resegregation of deseg- regated schools.
Taraso� v. Regents of the University of California (����)
Therapists (and presumably psychological assessors) must reveal privileged information if a third party is endangered. In the words of the Court, “Protective privilege ends where the public peril begins.”
Larry P. v. Riles (���� and rea�rmed by the same judge in ����)
California judge ruled that the use of intelligence tests to place Black children in special classes had a discriminatory in�uence because the tests were “racially and culturally biased.”
Debra P. v. Turlington (����) Federal court ruled that minimum competency testing in Florida was unconstitutional because it perpetuated the e�ects of past discrimination.
Griggs v. Duke Power Company (����) Black employees brought suit against a private company for discriminatory hiring practices. The U.S. Supreme Court found problems with “broad and general testing devices” and ruled that tests must “fairly measure the knowledge or skills required by a particular job.”
Albemarle Paper Company v. Moody (����) An industrial psychologist at a paper mill found that scores on a general ability test predicted measures of job performance. However, as a group, whites scored better than Blacks on the test. The U.S. District Court found the use of the test to be su�ciently job related. An appeals court did not. It ruled that discrimination had occurred, however unintended.
Regents of the University of California v. Bakke (����)
When Alan Bakke, who had been denied admission, learned that his test scores were higher than those of students from a “minority group” (in this case, Blacks, Chicanos, Asians, and American Indians) who had gained admission to the University of California at Davis medi- cal school, he sued. A highly divided U.S. Supreme Court agreed that Bakke should be admitted, but it did not preclude the use of diversity considerations in admission decisions.
Allen v. District of Columbia (����) Blacks scored lower than whites on a city �re department promotion test based on speci�c aspects of �re�ghting. The court found in favor of the �re department, ruling that “the pro- motional examination . . . was a valid measure of the abilities and probable future success of those individuals taking the test.”
coh37025_ch02_041-084.indd 62 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
Legislation�Although the legislation summarized in Table 2–1 was enacted at the federal level, states also have passed legislation that affects the assessment enterprise. In the 1970s numerous states enacted minimum competency testing programs: formal testing programs designed to be used in decisions regarding various aspects of students’ education. The data from such programs was used in decision making about grade promotions, awarding of diplomas, and identification of areas for remedial instruction. These laws grew out of grassroots support for the idea that high-school graduates should have, at the very least, “minimal competencies” in areas such as reading, writing, and arithmetic.
Truth-in-testing legislation was also passed at the state level beginning in the 1980s. The primary objective of these laws was to give testtakers a way to learn the criteria by which they are being judged. To meet that objective, some laws mandate the disclosure of answers to postsecondary and professional school admissions tests within 30 days of the publication of test scores. Some laws require that information relevant to a test’s development and technical soundness be kept on file. Some truth-in-testing laws require providing descriptions of (1) the test’s purpose and its subject matter, (2) the knowledge and skills the test purports to measure, (3) procedures for ensuring accuracy in scoring, (4) procedures for notifying testtakers of errors in scoring, and (5) procedures for ensuring the testtaker’s confidentiality. Truth-in-testing laws create special difficulties for test developers and publishers, who argue that it is essential for them to keep the test items secret. They note that there may be a limited item pool for some tests and that the cost of developing an entirely new set of items for each succeeding administration of a test is prohibitive.
Some laws mandate the involvement of the executive branch of government in their application. For example, Title VII of the Civil Rights Act of 1964 created the Equal Employment Opportunity Commission (EEOC) to enforce the act. The EEOC has published sets of guidelines concerning standards
Legislation Signi�cance
Adarand Constructors, Inc. v. Pena et al. (����)A construction �rm competing for a federal contract brought suit against the federal government after it lost a bid to a competitor from a diverse background, which the government had retained instead in the interest of a�rmative action. The U.S. Supreme Court, in a close (�–�) decision, found in favor of the plainti�, ruling that the government’s a�rmative action policy violated the equal protection clause of the ��th Amendment. The Court ruled, “Government may treat people di�erently because of their race only for the most compelling reasons.”
Ja�ee v. Redmond (����) Communication between a psychotherapist and a patient (and presumably a psychological assessor and a client) is privileged in federal courts.
Grutter v. Bollinger (����) In a highly divided decision, the U.S. Supreme Court approved the use of race in admissions decisions on a time-limited basis to further the educational bene�ts that �ow from a diverse student body.
Mitchell v. State, ��� P.�d ��� (Nev. ����) Does a court order for a compulsory psychiatric examination of the defendant in a criminal trial violate that defendant’s Fifth Amendment right to avoid self-incrimination? Given the particular circumstances of the case (see Leahy et al., ����), the Nevada Supreme Court ruled that the defendant’s right to avoid self-incrimination was not violated by the trial court’s order to have him undergo a psychiatric evaluation.
Ricci v. DeStefano (����) The ruling of the U.S. Supreme Court in this case had implications for the ways in which gov- ernment agencies can and cannot institute race-conscious remedies in hiring and promo- tional practices. Employers in the public sector were forbidden from e-hiring or promoting personnel using certain practices (such as altering a cuto� score to avoid adverse in�uence) unless the practice has been demonstrated to have a “strong basis in evidence.”
Table �–� Some Signi�cant Legislation and Litigation (continued)
J U S T T H I N K � . � . � .
How might truth-in-testing laws be modi�ed to better protect both the interest of testtakers and that of test developers?
coh37025_ch02_041-084.indd 63 12/01/21 4:04 PM
�����Part 1: An Overview
to be met in constructing and using employment tests. In 1978 the EEOC, the Civil Service Commission, the Department of Labor, and the Justice Department jointly published the Uniform Guidelines on Employee Selection Procedures. Here is a sample guideline:
The use of any test which adversely affects hiring, promotion, transfer or any other employment or membership opportunity of classes protected by Title VII constitutes discrimination unless (a) the test has been validated and evidences a high degree of utility as hereinafter described, and (b) the person giving or acting upon the results of the particular test can demonstrate that alternative suitable hiring, transfer or promotion procedures are unavailable for . . . use.
Note that here the definition of discrimination as exclusionary coexists with the proviso that a valid test evidencing “a high degree of utility” (among other criteria) will not be considered discriminatory. Generally, however, the public has been quick to label a test as unfair and discriminatory regardless of its utility. As a consequence, a great public demand for proportionality by group membership in hiring and college admissions now coexists with a great lack of proportionality in skills across groups. Gottfredson (2000) noted that although selection standards can often be improved, the manipulation of such standards “will produce only lasting frustration, not enduring solutions.” She recommended that enduring solutions be
sought by addressing the problems related to gaps in skills between groups. She argued against addressing the problem by lowering hiring and admission standards or by legislation designed to make hiring and admissions decisions a matter of group quotas.
In Texas, state law was enacted mandating that the top 10% of graduating seniors at each Texas high school be admitted to a state university regardless of SAT scores. This means that, regardless of the
quality of education in any particular Texas high school, a senior in the top 10% of the graduating class is guaranteed college admission regardless of how the student might score on a nationally administered measure. In California, the use of skills tests in the public sector decreased following the passage of Proposition 209, which banned racial preferences (Rosen, 1998). One consequence has been the deemphasis on the Law School Admissions Test (LSAT) as a criterion for being accepted by the University of California at Berkeley law school. Additionally, the law school stopped weighing grade point averages from undergraduate schools in their admission criteria, so that a 4.0 from any California state school “is now worth as much as a 4.0 from Harvard” (Rosen, 1998, p. 62).
Gottfredson (2000) makes the point that those who advocate reversal of achievement standards obtain “nothing of lasting value by eliminating valid tests.” For her, lowering standards amounts to hindering progress “while providing only the illusion of progress.” Rather than reversing achievement standards, society is best served by action to reverse other trends with deleterious effects (such as trends in family structure). In the face of consistent gaps between members of various groups, Gottfredson emphasized the need for skills training, not a lowering of achievement standards or an unfounded attack on tests.
State and federal legislatures, executive bodies, and courts have been involved in many aspects of testing and assessment. There has been little consensus about whether validated tests on which there are racial differences can be used to assist with employment-related decisions. Courts have also been grappling with the role of diversity in criteria for admission to colleges, universities, and professional schools. For example, in 2003 the question before the U.S. Supreme Court in the case of Grutter v. Bollinger was “whether diversity is a compelling interest that can justify the narrowly tailored use of race in selecting applicants for admission to public universities.” One of the questions to be decided in that case was whether the University of Michigan Law School was using a quota system, a selection procedure whereby a fixed number or percentage of applicants from certain backgrounds were selected.
Many of the cases brought before federal courts under Title VII of the Civil Rights Act are employment discrimination cases. In this context, discrimination may be defined as the
J U S T T H I N K . � . � .
How can government and the private sector address problems related to gaps in skills between groups?
coh37025_ch02_041-084.indd 64 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
practice of making distinctions in hiring, promotion, or other selection decisions that tend to systematically favor members of a majority group regardless of actual qualifications for positions. Discrimination may occur as the result of intentional or unintentional action on the part of an employer. As an example of unintentional discrimination, consider the hiring practice of a municipal fire department that required applicants to weigh not less than 135 pounds, and not more than 225 pounds. This job requirement might unintentionally discriminate against, and systematically screen-out, applicants from members of cultural groups whose average weight fell below the required minimum. In all likelihood, the fire department would be challenged in a court of law by a member of the excluded cultural group. Accordingly, the municipality would be required to document why weighing a minimum of 135 pounds should be a requirement for joining that particular fire department.
Typically, when a Title VII charge of discrimination in the workplace is leveled at an employer, a claim is made that hiring, promotion, or some related employment decisions are systematically being made not on the basis of job-related variables, but rather on the basis of some non-job- related variable (such as race, gender, sexual orientation, religion, or national origin). Presumably, the selection criteria favors members of the majority group. In some instances, however, it is members of the majority group who are compelled to make a claim of reverse discrimination. In this context, reverse discrimination may be defined as the practice of making distinctions in hiring, promotion, or other selection decisions that systematically tend to favor racially, ethnically, socioeconomically, or culturally diverse persons regardless of actual qualifications for positions.
In both discrimination and reverse discrimination cases, the alleged discrimination may occur as the result of intentional or unintentional employer practices. The legal term disparate treatment refers to the consequence of an employer’s hiring or promotion practice that was intentionally devised to yield some discriminatory result or outcome. Possible motivations for disparate treatment include racial prejudice and a desire to maintain the status quo. By contrast, the legal term disparate impact refers to the consequence of an employer’s hiring or promotion practice that unintentionally yielded a discriminatory result or outcome. Because disparate impact is presumed to occur unintentionally, it is not viewed as the product of motivation or�planning.
As you will discover as you learn more about test construction and the art and science of testing, a job applicant’s score on a test or other assessment procedure is, at least ideally, a reflection of that applicant’s underlying ability to succeed at the job. Exactly how well that score actually reflects the job applicant’s underlying ability depends on a number of factors. One factor it surely depends on is the quality of the test or selection procedure. When a claim of discrimination (or reverse discrimination) is made, an evaluation of the quality of a test or selection procedure will typically entail scrutiny of a number of variables including, for example: (a) the competencies actually assessed by the test and how related those competencies are to the job; (b) the differential weighting, if any, of items on the test or the selection procedures; (c) the psychometric basis for the cutoff score in effect (is a score of 65 to pass, e.g., really justified?); (d) the rationale in place for rank-ordering candidates; (e) a consideration of potential alternative evaluation procedures that could have been used; and (f) an evaluation of the statistical evidence that suggests discrimination or reverse discrimination occurred.
Many large companies and organizations, as well as government agencies, hire experts in assessment to help make certain that their hiring and promotion practices result in neither disparate treatment nor disparate impact. They do so because the mere allegation of discrimination can be a source of great expense for any private or public employer. An employer accused of discrimination under Title VII will typically have to budget for a number of expenses including the costs of attorneys, consultants, and experts, and the retrieval, scanning, and storage of records. The consequences of losing such a lawsuit can add additional, sometimes staggering, costs. Included here, for example, are the costs of the plaintiff’s attorney fees, the costs attendant to improving and restructuring hiring and promotion protocols, and the costs
coh37025_ch02_041-084.indd 65 12/01/21 4:04 PM
�����Part 1: An Overview
of monetary damages to all present and past injured parties. Additionally, new hiring may be halted and pending promotions may be delayed until the court is satisfied that the new practices put into place by the offending employer do not and will not result in disparate treatment or impact. In some cases, a lawsuit will be momentous not merely for the number of dollars spent, but for the number of changes in the law that are a direct result of the litigation.
Litigation� Rules governing citizens’ behavior stem not only from legislatures but also from interpretations of existing law in the form of decisions handed down by courts. In this way, law resulting from litigation (the court-mediated resolution of legal matters of a civil, criminal, or administrative nature) can influence our daily lives. Examples of some court cases that have affected the assessment enterprise were presented in Table 2–1 under the “Litigation” heading. It is also true that litigation can result in bringing an important and timely matter to the attention of legislators, thus serving as a stimulus to the creation of new legislation. This is exactly what happened in the cases of PARC v. Commonwealth of Pennsylvania (1971) and Mills v. Board of Education of District of Columbia (1972). In the PARC case, the Pennsylvania Association for Retarded Children brought suit because children with intellectual disability in that state had been denied access to public education. In Mills, a similar lawsuit was filed on behalf of children with behavioral, emotional, and learning impairments. Taken together, these two cases had the effect of jump-starting similar litigation in several other jurisdictions and alerting Congress to the need for federal law to ensure appropriate educational opportunities for children with disabilities.
Litigation has sometimes been referred to as “judge-made law” because it typically comes in the form of a ruling by a court. And although judges do, in essence, create law by their rulings, these rulings are seldom made in a vacuum. Rather, judges typically rely on prior rulings and on other people—most notably, expert witnesses—to assist in their judgments. A psychologist acting as an expert witness in criminal litigation may testify on matters such as the competence of a defendant to stand trial, the competence of a witness to give testimony, or the sanity of a defendant entering a plea of “not guilty by reason of insanity.” A psychologist acting as an expert witness in a civil matter could conceivably offer opinions on many different types of issues ranging from the parenting skills of a parent in a divorce case to the capabilities of a factory worker prior to sustaining a head injury on the job. In a malpractice case, an expert witness might testify about how reasonable and professional the actions taken by a fellow psychologist were and whether any reasonable and prudent practitioner would have engaged in the same or similar actions (Cohen, 1979).
The issues on which expert witnesses can be called upon to give testimony are as varied as the issues that reach courtrooms for resolution. And so, some important questions arise with respect to expert witnesses. For example: Who is qualified to be an expert witness? How much weight should be given to the testimony of an expert witness? Questions such as these have themselves been the subject of litigation.
A landmark case heard by the U.S. Supreme Court in June 1993 has implications for the admissibility of expert testimony in court. The case was Daubert v. Merrell Dow Pharmaceuticals. The origins of this case can be traced to Mrs. Daubert’s use of the prescription drug Bendectin to relieve nausea during pregnancy. The plaintiffs sued the manufacturer of this drug, Merrell Dow Pharmaceuticals, when their children were born with birth defects. They claimed that Mrs. Daubert’s use of Bendectin had caused their children’s birth defects.
Attorneys for the Dauberts were armed with research that they claimed would prove that Bendectin causes birth defects. However, the trial judge ruled that the research failed to meet the criteria for admissibility. In part because the evidence the Dauberts wished to present was not deemed admissible, the trial judge ruled against the Dauberts.
The Dauberts appealed to the next higher court. That court, too, ruled against them and in favor of Merrell Dow. Once again, the plaintiffs appealed, this time to the U.S. Supreme Court. A question before the Court was whether the judge in the original trial had acted
coh37025_ch02_041-084.indd 66 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
properly by not allowing the plaintiffs’ research to be admitted into evidence. To understand whether the trial judge acted properly, it is important to understand (1) a ruling that was made in the 1923 case of Frye v. the United States and (2) a law subsequently passed by Congress, Rule 702 in the Federal Rules of Evidence (1975).
In Frye, the Court held that scientific research is admissible as evidence when the research study or method enjoys general acceptance. General acceptance could typically be established by the testimony of experts and by reference to publications in peer-reviewed journals. In short, if an expert witness claimed something that most other experts in the same field would agree with then, under Frye, the testimony could be admitted into evidence. Rule 702 changed that by allowing more experts to testify regarding the admissibility of the original expert testimony. Beyond expert testimony indicating that some research method or technique enjoyed general acceptance in the field, other experts were now allowed to testify and present their opinions with regard to the admissibility of the evidence. So, an expert might offer an opinion to a jury concerning the acceptability of a research study or method regardless of whether that opinion represented the opinions of other experts. Rule 702 was enacted to assist juries in their fact-finding by helping them to understand the issues involved.
Presenting their case before the Supreme Court, the attorneys for the Dauberts argued that Rule 702 had wrongly been ignored by the trial judge. The attorneys for the defendant, Merrell Dow Pharmaceuticals, countered that the trial judge had ruled appropriately. The defendant argued that high standards of evidence admissibility were necessary to protect juries from “scientific shamans who, in the guise of their purported expertise, are willing to testify to virtually any conclusion to suit the needs of the litigant with resources sufficient to pay their retainer.”
The Supreme Court ruled that the Daubert case be retried and that the trial judge should be given wide discretion in deciding what does and does not qualify as scientific evidence. In effect, federal judges were charged with a gatekeeping function with respect to what expert testimony would or would not be admitted into evidence. The Daubert ruling superseded the long-standing policy, set forth in Frye, of admitting into evidence only scientific testimony that had won general acceptance in the scientific community. Opposing expert testimony, whether such testimony had won general acceptance in the scientific community, would be admissible.
Copyright 2016 Ronald Jay Cohen. All rights reserved.
coh37025_ch02_041-084.indd 67 12/01/21 4:04 PM
�����Part 1: An Overview
In Daubert, the Supreme Court viewed factors such as general acceptance in the scientific community or publication in a peer-reviewed journal as only some of many possible factors for judges to consider. Other factors judges might consider included the extent to which a theory or technique had been tested and the extent to which the theory or technique might be subject to error. In essence, the Supreme Court’s ruling in Daubert gave trial judges a great deal of leeway in deciding what juries would be allowed to hear.
Subsequent to Daubert, the Supreme Court has ruled on several other cases that in one way or another clarify or slightly modify its position in Daubert. For example, in the case of General Electric Co. v. Joiner (1997), the Court emphasized that the trial court had a duty to exclude unreliable expert testimony as evidence. In the case of Kumho Tire Company Ltd. v. Carmichael (1999), the Supreme Court expanded the principles expounded in Daubert to include the testimony of all experts, regardless of whether the experts claimed scientific research as a basis for their testimony. Thus, for example, a psychologist’s testimony based on personal experience in independent practice (rather than findings from a formal research study) could be admitted into evidence at the discretion of the trial judge (Mark, 1999).
Whether Frye or Daubert will be relied on by the court depends on the individual jurisdiction in which a legal proceeding occurs. Some jurisdictions still rely on the Frye standard when it comes to admitting expert testimony, and some subscribe to Daubert. As an example, consider the Missouri case of Zink v. State (2009). After David Zink rear-ended a woman’s car in traffic, Zink kidnapped the woman, and then raped, mutilated, and murdered her. Zink was subsequently caught, tried, convicted, and sentenced to death. In an appeal proceeding, Zink argued that the death penalty should be set aside because of his mental disease. Zink’s position was that he was not adequately represented by his attorney, because during the trial, his defense attorney had failed to present “hard” evidence of a mental disorder as indicated by a PET scan (a type of neuroimaging tool that will be discussed in Chapter 14). The appeals court denied Zink’s claim, noting that the PET scan failed to meet the Frye standard for proving mental disorder (Haque & Guyer, 2010).
The implications of Daubert for psychologists and others who might have occasion to provide expert testimony in a trial are wide ranging (Ewing & McCann, 2006). More specifically, discussions of the implications of Daubert for psychological experts can be found in cases involving mental capacity (Bumann, 2010; Frolik, 1999; Poythress, 2004), claims of emotional distress (McLearen et al., 2004), personnel decisions (Landy, 2007), child custody and termination of parental rights (Bogacki & Weiss, 2007; Gould, 2006; Krauss & Sales, 1999), and numerous other matters (Grove & Barden, 1999; Lipton, 1999; Mossman, 2003; Posthuma et al., 2002; Saldanha, 2005; Saxe & Ben-Shakhar, 1999; Slobogin, 1999; Stern, 2001; Tenopyr, 1999). One concern is that Daubert has not been applied consistently across jurisdictions and within jurisdictions (Sanders, 2010).
The Concerns of the Profession
As early as 1895 the American Psychological Association (APA), in its infancy, formed its first committee on mental measurement. The committee was charged with investigating various aspects of the relatively new practice of testing. Another APA committee on measurement was formed in 1906 to further study various testing-related issues and problems. In 1916 and again in 1921, symposia dealing with various issues surrounding the expanding uses of tests were sponsored (Mentality Tests, 1916; Intelligence and Its Measurement, 1921). In 1954, APA published its Technical Recommendations for Psychological Tests and Diagnostic Tests, a document that set forth testing standards and technical recommendations. The following year, another professional organization, the National Educational Association (working in collaboration with the National Council on Measurements Used in Education—now known as the National Council on Measurement) published its Technical Recommendations for Achievement Tests.
coh37025_ch02_041-084.indd 68 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
Collaboration between these professional organizations led to the development of rather detailed testing standards and guidelines that would be periodically updated in future years.
Expressions of concern about the quality of tests being administered could also be found in the work of several professionals, acting independently. Anticipating the present-day Standards, Ruch (1925), a measurement specialist, proposed a number of standards for tests and guidelines for test development. He also wrote of “the urgent need for a fact-finding organization which will undertake impartial, experimental, and statistical evaluations of tests” (Ruch, 1933). History records that one team of measurement experts even took on the (overly) ambitious task of attempting to rank all published tests designed for use in educational settings. The result was a pioneering book (Kelley, 1927) that provided test users with information needed to compare the merits of published tests. However, given the pace at which test instruments were being published, this resource required regular updating. And so, Oscar Buros was not the first measurement professional to undertake a comprehensive testing of the tests. He was, however, the most tenacious in updating and revising the information.
The APA and related professional organizations in the United States have made available numerous reference works and publications designed to delineate ethical, sound practice in the field of psychological testing and assessment.3 Along the way, these professional organizations have tackled a variety of thorny questions, such as the questions cited in the next Just Think.
Test-user quali�cations� Should anyone be allowed to purchase and use psychological test materials? If not, then who should be permitted to use psychological tests? As early as 1950 an APA Committee on Ethical Standards for Psychology published a report called Ethical Standards for the Distribution of Psychological Tests and Diagnostic Aids. This report defined three levels of tests in terms of the degree to which the test’s use required knowledge of testing and psychology.
Level A: Tests or aids that can adequately be administered, scored, and interpreted with the aid of the manual and a general orientation to the kind of institution or organization in which one is working (for instance, achievement or proficiency tests).
Level B: Tests or aids that require some technical knowledge of test construction and use and of supporting psychological and educational fields such as statistics, individual differences, psychology of adjustment, personnel psychology, and guidance (e.g., aptitude tests and adjustment inventories applicable to normal populations).
Level C: Tests and aids that require substantial understanding of testing and supporting psychological fields together with supervised experience in the use of these devices (for instance, projective tests, individual mental tests).
The report included descriptions of the general levels of training corresponding to each of the three levels of tests. Although many test publishers continue to use this three-level classification, some do not. In general, professional standards promulgated by professional organizations state that psychological tests should be used only by qualified persons. Furthermore, there is an ethical mandate to take reasonable steps to prevent the misuse of the tests and the information they provide. The obligations of professionals to testtakers are set forth in a document called the Code of Fair Testing Practices in Education. Jointly authored and/or sponsored by the Joint Committee of Testing Practices (a coalition of APA, AERA, NCME, the American
J U S T T H I N K � . � . � .
Who should be privy to test data? Who should be able to purchase psychological test materials? Who is quali�ed to administer, score, and interpret psychological tests? What level of expertise in psychometrics quali�es someone to administer which types of test?
3. Unfortunately, although organizations in many other countries have verbalized concern about ethics and standards in testing and assessment, relatively few organizations have taken meaningful and effective action in this regard (Leach & Oakland, 2007).
coh37025_ch02_041-084.indd 69 12/01/21 4:04 PM
�����Part 1: An Overview
Association for Measurement and Evaluation in Counseling and Development, and the American Speech-Language Hearing Association), this document presents standards for educational test developers in four areas: (1) developing/selecting tests, (2) interpreting scores, (3) striving for fairness, and (4) informing testtakers.
Beyond promoting high standards in testing and assessment among professionals, APA has initiated or assisted in litigation to limit the use of psychological tests to qualified personnel. Skeptics label such measurement-related legal action as a kind of jockeying for turf, done solely for financial gain. A more charitable and perhaps more realistic view is that such actions benefit society at large. It is essential to the survival of the assessment enterprise that certain assessments be conducted by people qualified to conduct them by virtue of their education, training, and experience.
A psychologist licensing law designed to serve as a model for state legislatures has been available from APA since 1987. The law contains no definition of psychological testing. In the interest of the public, the profession of psychology, and other professions that employ psychological tests, it may now be time for that model legislation to be rewritten—with terms such as psychological testing and psychological assessment clearly defined and differentiated. Terms such as test-user qualifications and psychological assessor qualifications must also be clearly defined and differentiated. It seems that legal conflicts regarding psychological test usage partly stem from confusion of the terms psychological testing and psychological assessment. People who are not considered professionals by society may be qualified to use
psychological tests (psychological testers). However, these same people may not be qualified to engage in psychological assessment. As we argued in Chapter 1, psychological assessment requires certain skills, talents, expertise, and training in psychology and measurement over and above that required to engage in psychological testing. In the past, psychologists have been lax in differentiating psychological
testing from psychological assessment. However, continued laxity may prove to be a costly indulgence, given current legislative and judicial trends.
Testing people with disabilities�Challenges analogous to those concerning culturally and linguistically diverse testtakers are present when testing people with disabling conditions. Specifically, these challenges may include (1) transforming the test into a form that can be taken by the testtaker, (2) transforming the responses of the testtaker so that they are scorable, and (3) meaningfully interpreting the test data.
The nature of the transformation of the test into a form ready for administration to the individual with a disabling condition will, of course, depend on the nature of the disability. Then, too, some test stimuli do not translate easily. For example, if a critical aspect of a test item contains artwork to be analyzed, there may be no meaningful way to translate this item for use with testtakers who are blind. With respect to any test converted for use with a population for which the test was not originally intended, choices must inevitably be made regarding exactly how the test materials will be modified, what standards of evaluation will be applied, and how the results will be interpreted. Professional assessors do not always agree on the answers to such questions.
Another complex issue—this one, ethically charged—has to do with a request by a terminally ill individual for assistance in quickening the process of dying. In Oregon, the first state to enact “Death with Dignity” legislation, a request for assistance in dying may be granted only contingent on the findings of a psychological evaluation; life or death literally hangs in the balance of such assessments. Some ethical and related issues surrounding this phenomenon are discussed in greater detail in this chapter’s Everyday Psychometrics.
J U S T T H I N K . � . � .
Why is it essential for the terms psychological testing and psychological assessment to be de�ned and di�erentiated in state licensing laws?
J U S T T H I N K . � . � .
If the form of a test is changed or adapted for a speci�c type of administration to a particular individual or group, can the scores obtained by that individual or group be interpreted in a “business as usual” manner?
coh37025_ch02_041-084.indd 70 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
E V E R Y D A Y P S Y C H O M E T R I C S
Life-or-Death Psychological Assessment
he state of Oregon has the distinction—dubious to some people, depending on one’s values—of having enacted the nation’s �rst aid-in-dying law. Oregon’s Death with Dignity Act (ODDA) provides that a patient with a medical condition thought to give that patient � months or less to live may end his or her own life by voluntarily requesting a lethal dose of medication. The law requires that two physicians corroborate the terminal diagnosis and stipulates that either may request a psychological evaluation of the patient by a state-licensed psychologist or psychiatrist in order to ensure that the patient is competent to make the life-ending decision and to rule out impaired judgment due to psychiatric disorder. Assistance in dying will be denied to persons “su�ering from a psychiatric or psychological disorder, or depression causing impaired judgement” (ODDA, ����). Since ����, similar legislation has been enacted in other states (California, Montana, New Mexico, Vermont, and Washington), and a number of other states are actively considering such “death with dignity” (otherwise known as “physician-aid-in-dying”) legislation. Although our focus here is on the ODDA as it a�ects�psychological assessors who are called upon to make life-and-death evaluations, many of the complex issues surrounding such legislation are the same or similar in other jurisdictions. More detailed coverage of the complex legal and�values-related issues can be found in sources such as Johnson et al. (����, ����), Reynolds (����), Smith et al. (����), and White (����).
The ODDA was hotly debated prior to its passage by referendum, and it remains controversial today. Critics of the law question whether suicide is ever a rational choice under any circumstances, and they fear that state-condoned aid in dying will serve to destigmatize suicide in general (Callahan, ����; see also Richman, ����). It is argued that the �rst duty of health and mental health professionals is to do no harm (Jennings, ����). Some fear that professionals willing to testify to almost anything (so-called hired guns) will corrupt the process by providing whatever professional opinion is desired by those who will pay their fees. Critics also point with concern to the experience of the Dutch death-with-dignity legislation. In the Netherlands, relatively few individuals requesting physician-assisted suicide are referred for psychological assessment. Further, the highest court of that land ruled that “in rare cases, physician-assisted suicide is possible even for individuals su�ering only from mental problems rather than from physical illnesses” (Abeles & Barlev, ����, p. ���). On moral and religious grounds, it has
T
been argued that death should be viewed as the province solely of Divine, not human, intervention.
Supporters of death-with-dignity legislation argue that life-sustaining equipment and methods can extend life beyond a time when it is meaningful and that the �rst obligation of health
(continued)
Sigmund Freud (����–����)
It has been said that Sigmund Freud made a “rational decision” to end his life. Suffering from terminal throat cancer, having great difficulty in speaking, and experiencing increasing difficulty in breathing, the founder of psychoanalysis asked his physician for a lethal dose of morphine. For years it has been debated whether a decision to die, even made by a terminally ill patient, can ever truly be “rational.” Today, in accordance with death-with-dignity legislation, the responsibility for evaluating just how rational such a choice is falls on mental health professionals. Time Life Pictures/Mansell/The LIFE Picture Collection/Getty Images
coh37025_ch02_041-084.indd 71 12/01/21 4:04 PM
�����Part 1: An Overview
control over the dying process. Couched in these terms, the sober duty of the clinician drawn into the process may be made more palatable or even ennobled.
The ODDA provides for various records to be kept regarding patients who die under its provisions. Each year since the Act �rst took e�ect, the collected data is published in an annual report. So, for example, in the ���� report we learn that the reasons most frequently cited for seeking to end one’s life were loss of autonomy, decreasing ability to participate in activities that made life enjoyable, loss of dignity, and loss of control of bodily functions. In ����, �� prescriptions for lethal medications were prescribed and �� people had opted to end their life by ingesting the medications.
Psychologists and psychiatrists called upon to make death-with-dignity competency evaluations may accept or decline the responsibility (Haley & Lee, ����). Judging from one survey of ��� psychologists in clinical practice in Oregon (Fenn & Ganzini, ����), many of the psychologists who could be asked to make such a life-or-death assessment might decline to do so. About one-third of the sample responded that an ODDA assessment would be outside the scope of their practice. Another ��% of the sample said they would either refuse to perform the assessment and take no further action or refuse to perform the assessment themselves and refer the patient to a colleague.
Guidelines for the ODDA assessment process were o�ered by Farrenkopf and Bryan (����), and they are as follows.
and mental health professionals is to relieve su�ering (Latimer, ����; Quill et al., ����; Weir, ����). Additionally, they may point to the dogged determination of people intent on dying and to stories of how many terminally ill people have struggled to end their lives using all kinds of less-than-sure methods, enduring even greater su�ering in the process. In marked contrast to such horror stories, the �rst patient to die under the ODDA is said to have described how the family “could relax and say what a wonderful life we had. We could look back at all the lovely things because we knew we �nally had an answer” (cited in Farrenkopf & Bryan, ����, p. ���).
Professional associations such as the American Psychological Association and the American Psychiatric Association have long promulgated codes of ethics requiring the prevention of suicide. The enactment of the law in Oregon has placed clinicians in that state in a uniquely awkward position. Clinicians who for years have devoted their e�orts to suicide prevention have been thrust into the position of being a potential party to, if not a facilitator of, physician-assisted suicide—regardless of how the aid-in-dying process is referred to in the legislation. Note that the Oregon law scrupulously denies that its objective is the legalization of physician-assisted suicide. In fact, the language of the act mandates that action taken under it “shall not, for any purpose, constitute suicide, assisted suicide, mercy killing or homicide, under the law.” The framers of the legislation perceived it as a means by which a terminally ill individual could exercise some
E V E R Y D A Y P S Y C H O M E T R I C S
Life-or-Death Psychological Assessment (continued)
�. Review of Records and Case History With the patient’s consent, the assessor will gather records from all relevant sources, including medical and mental health records. A goal is to understand the patient’s current functioning in the context of many factors, ranging from the current medical condition and prognosis to the e�ects of medication and substance use.
�. Consultation with Treating Professionals With the patient’s consent, the assessor may consult with the patient’s physician and other professionals involved in the case to better understand the patient’s current functioning and current situation.
�. Patient Interviews Sensitive but thorough interviews with the patient will explore the reasons for the aid-in-dying request, including the pressures and values motivating the request. Other areas to explore
include: (a) the patient’s understanding of his or her medical condition, the prognosis, and the treatment alternatives; (b) the patient’s experience of physical pain, limitations of functioning, and changes over time in cognitive, emotional, and perceptual functioning; (c) the patient’s characterization of his or her quality of life, including exploration of related factors including personal identity, role functioning, and self-esteem; and (d) external pressures on the patient, such as personal or familial �nancial inability to pay for continued treatment.
�. Interviews with Family Members and Signi�cant Others With the permission of the patient, separate interviews should be conducted with the patient’s family and signi�cant others. One objective is to explore from their perspective how the patient has adjusted in the past to adversity and how the patient has changed and adjusted to his or her current situation.
The ODDA Assessment Process
coh37025_ch02_041-084.indd 72 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
�. Assessment of Competence Like the other elements of this overview, this aspect of the assessment is complicated, and only the barest of guidelines can be presented here. In general, the assessor seeks to understand the patient’s reasoning and decision-making process, including all information relevant to the decision and its consequences. Some formal tests of competency are available (Appelbaum & Grisso, ����a, ����b; Lavin, ����), but the clinical and legal applicability of such tests to an ODDA assessment has yet to be established.
�. Assessment of Psychopathology To what extent is the decision to end one’s life a function of pathological depression, anxiety, dementia, delirium, psychosis, or some other pathological condition? The assessor addresses this question using not only interviews but formal tests. Examples of the many possible instruments the assessor might
employ include intelligence tests, personality tests, neuropsychological tests, symptom checklists, and depression and anxiety scales; refer to the appendix in Farrenkopf and Bryan (����) for a complete list of these tests.
�. Reporting Findings and Recommendations Findings, including those related to the patient’s mental status and competence, family support and pressures, and anything else relevant to the patient’s aid-in-dying request, should be reported. If treatable conditions were found, treatment recommendations relevant to those conditions may be made. Nontreatment types of recommendations may include recommendations for legal advice, estate planning, or other resources. In Oregon, a Psychiatric/Psychological Consultant’s Compliance Form with the consultant’s recommendations should be completed and sent to the Oregon Health Division.
Computerized test administration, scoring, and interpretation� Computer-assisted psychological assessment (CAPA) has become more the norm than the exception. An ever-growing number of psychological tests can be purchased on disc or administered and scored online. In many respects, the relative simplicity, convenience, and range of potential testing activities that computer technology brings to the testing industry have been a great boon. Of course, every rose has its thorns.
For assessment professionals, some major issues with regard to CAPA are as follows.
� Comparability of pencil-and-paper and computerized versions of tests. Many tests once available only in a paper-and-pencil format are now available in computerized form as well. In many instances the comparability of the traditional and the computerized forms of the test has not been researched or has only insufficiently been researched. With questionnaire assessments such as the MMPI-2, the results are generally comparable across formats (Forbey & Ben-Porath, 2007; Nyquist & Forbey, 2018). Preliminary evidence suggests that it is possible to create computerized ability tests that are largely comparable across formats (Wahlstrom et al., 2019) though equivalence cannot be assumed in all cases (Krach et al., 2020). Some studies have found that participants find tablet-based testing to be more engaging than the comparable pencil-and-paper version of the test (Marble-Flint et al., 2019; Noland, 2017).
� The value of computerized test interpretations. Many tests available for computerized administration also come with computerized scoring and interpretation procedures. Although computerized scoring is generally more accurate than hand scoring (e.g., Allard et al., 1995), the comparative accuracy of computerized interpretation versus clinician interpretation is often not known (e.g., Stolberg, 2018).
� Unprofessional, unregulated “psychological testing” online. A growing number of Internet sites purport to provide, usually for a fee, online psychological tests. Yet the vast majority of the tests offered would not meet a psychologist’s standards. Assessment professionals wonder about the long-term effect of these largely unprofessional
J U S T T H I N K � . � . � .
What di�erences in the test results may exist as a result of the same test being administered orally, online, or by means of a paper-and-pencil examination? What di�erences in the testtaker’s experience may exist as a function of test administration method?
coh37025_ch02_041-084.indd 73 12/01/21 4:04 PM
�����Part 1: An Overview
and unregulated “psychological testing” sites. Might they, for example, contribute to more public skepticism about psychological tests?
Imagine being administered what has been represented to you as a “psychological test,” only to find that the test is not bona fide. The online availability of myriad tests of uncertain quality that purport to measure psychological variables increases the possibility of this happening. To help remedy such potential problems, a Florida-based organization called the International Test Commission developed the “International Guidelines on Computer-Based and Internet-Delivered Testing” (Coyne & Bartram, 2006). These guidelines address technical, quality, security, and related issues. Although not without limitations (Sale, 2006), these guidelines clearly are a step forward in nongovernmental regulation. Other guidelines are written to inform the rendering of professional services to members of certain populations.
Guidelines with respect to certain populations�From time to time, the American Psychological Association (APA) has published special guidelines for professionals who have occasion to assess, treat, conduct research with, or otherwise consult with members of certain populations. In general, the guidelines are designed to assist professionals in providing informed and developmentally appropriate services. Note that there exists a distinction between APA guidelines and standards. Although standards must be followed by all psychologists, guidelines are more aspirational in nature (Reed et al., 2002). In late 2015, for example, APA published its Guidelines for Psychological Practice with Transgender and Gender Nonconforming (TGNC) People. The document lists and discusses 16 guidelines. To get a sense of what these guidelines say, the first guideline is: “Psychologists understand that gender is a non-binary construct that allows for a range of gender identities and that a person’s gender identity may not align with sex assigned at birth.” The last guideline, Guideline 16, is: “Psychologists seek to prepare trainees in psychology to work competently with TGNC people.” In 2012, APA also published guidelines for working with lesbian, gay, and bisexual clients. Further, APA (2017) offered a broader, ecological approach in its guidelines for multicultural practice in addressing context, identity, and intersectionality.
Various other groups and professional organizations also publish documents that may be helpful to mental health professionals vis-à-vis the provision of services to members of specific populations. For example, the Intercollegiate Committee of the Royal College of Psychiatrists publishes a list of “good practices” for the assessment and treatment of people with gender dysphoria (Wylie et al., 2014). Other groups have their own “best practices” (Goodrich et al., 2013) or simply “practices” (Beek et al., 2015; Bouman et al., 2014; de Vries et al., 2014; Dhejne et al., 2016; Sherman et al., 2014) that may inform professional practice. Additional practice-related resources that may be of particular interest to assessment professionals include special issues of journals devoted to the topic of interest (such as Borden, 2015), and publications that specifically focus on the topic from an assessment perspective (da Silva et al., 2016; Dèttore et al., 2015; Johnson et al., 2004; Luyt, 2015; Rönspies et al., 2015).
The Rights of Testtakers
As prescribed by the Standards and in some cases by law, some of the rights that test users accord to testtakers are the right of informed consent, the right to be informed of test findings, the right to privacy and confidentiality, and the right to the least stigmatizing label.
The right of informed consent�Testtakers have a right to know why they are being evaluated, how the test data will be used, and what (if any) information will be released to whom. With full knowledge of such information, testtakers give their informed consent to be tested. The disclosure of the information needed for consent must, of course, be in language the testtaker
coh37025_ch02_041-084.indd 74 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
can understand. Thus, for a testtaker as young as 2 or 3 years of age or an individual who has an intellectual disability with limited language skills, a disclosure before testing might be worded as follows: “I’m going to ask you to try to do some things so that I can see what you know how to do and what things you could use some more help with” (APA, 1985, p. 85).
Competency in providing informed consent has been broken down into several components: (1) Being able to evidence a choice as to whether one wants to participate; (2) demonstrating a factual understanding of the issues; (3) being able to reason about the facts of a study, treatment, or whatever it is to which consent is sought, and (4) appreciating the nature of the situation (Appelbaum & Roth, 1982; Roth et al., 1977).
Competency to provide consent may be assessed informally, and in fact many physicians engage in such informal assessment. Marson et al. (1997) cautioned that informal assessment of competency may be idiosyncratic and unreliable. As an alternative, many standardized instruments are available (Sturman, 2005). One such instrument is the MacArthur Competence Assessment Tool-Treatment (Grisso & Appelbaum, 1998). Also known as the MacCAT-T, it consists of structured interviews based on the four components of competency listed above (Grisso et al., 1997). Other instruments have been developed that are performance based and yield information on decision-making competence (Finucane & Gullion, 2010).
Another consideration related to competency is the extent to which persons diagnosed with psychopathology may be incompetent to provide informed consent (Sturman, 2005). For example, individuals diagnosed with dementia, bipolar disorder, and schizophrenia are likely to have competency impairments that may affect their ability to provide informed consent. By contrast, individuals with major depression may retain the competency to give truly informed consent (Grisso & Appelbaum, 1995; Palmer et al., 2007; Vollmann et al., 2003). Competence to provide informed consent may be improved by training (Carpenter et al., 2000; Dunn et al., 2002; Palmer et al., 2007). Therefore, clinicians should not necessarily assume that patients are not capable of consent based solely on their diagnosis.
If a testtaker is incapable of providing an informed consent to testing, such consent may be obtained from a parent or a legal representative. Consent must be in written rather than spoken form. The written form should specify (1) the general purpose of the testing, (2) the specific reason it is being undertaken in the present case, and (3) the general type of instruments to be administered. Many school districts now routinely send home such forms before testing children. Such forms typically include the option to have the child assessed privately if a parent so desires. In instances where testing is legally mandated (as in a court-ordered situation), obtaining informed consent to test may be considered more of a courtesy (undertaken in part for reasons of establishing good rapport) than a necessity.
One gray area with respect to the testtaker’s right of fully informed consent before testing involves research and experimental situations wherein the examiner’s complete disclosure of all facts pertinent to the testing (including the experimenter’s hypothesis and so forth) might irrevocably contaminate the test data. In some instances, deception is used to create situations that occur relatively rarely. For example, a deception might be created to evaluate how an emergency worker might react under emergency conditions. Sometimes deception involves the use of confederates to simulate social conditions that can occur during an event of some sort.
For situations in which it is deemed advisable not to obtain fully informed consent to evaluation, professional discretion is in order. Testtakers might be given a minimum amount of information before the testing. For example, “This testing is being undertaken as part of an experiment on obedience to authority.” A full disclosure and debriefing would be made after the testing. Various professional organizations have created policies and guidelines regarding deception in research. For example, the APA Ethical Principles of Psychologists and Code of Conduct (2017) provides that
J U S T T H I N K . � . � .
Describe a scenario in which knowledge of the experimenter’s hypotheses would probably invalidate the data gathered.
coh37025_ch02_041-084.indd 75 12/01/21 4:04 PM
�����Part 1: An Overview
psychologists (a) do not use deception unless it is absolutely necessary, (b) do not use deception at all if it will cause participants emotional distress, and (c) fully debrief participants.
The right to be informed of test �ndings� In a bygone era, the inclination of many psychological assessors, particularly many clinicians, was to tell testtakers as little as possible about the nature of their performance on a particular test or test battery. In no case would they disclose diagnostic conclusions that could arouse anxiety or precipitate a crisis. This orientation was reflected in at least one authoritative text that advised testers to keep information about test results superficial and focus only on “positive” findings. This was done so that the examinee would leave the test session feeling “pleased and satisfied” (Klopfer et al., 1954, p. 15). But all that has changed, and giving realistic information about test performance to examinees is not only ethically and legally mandated but may be useful from a therapeutic perspective as well. Testtakers have a right to be informed, in language they can understand, of the nature of the findings with respect to a test they have taken. They are also entitled to know what recommendations are being made as a consequence of the test data. If the test results, findings, or recommendations made on the basis of test data are voided for any reason (such as irregularities in the test administration), testtakers have a right to know that as well.
Because of the possibility of untoward consequences of providing individuals with information about themselves—ability, lack of ability, personality, values—the communication of results of a psychological test is a most important part of the evaluation process. With sensitivity to the situation, the test user will inform the testtaker (and the parent or the legal representative or both) of the purpose of the test, the meaning of the score relative to those of other testtakers, and the possible limitations and margins of error of the test. And regardless of whether such reporting is done in person or in writing, a qualified professional should be available to answer any further questions that testtakers (or their parents or legal representatives) have about the test scores. Ideally, counseling resources will be available for those who react adversely to the information presented.
The right to privacy and con�dentiality�The concept of the privacy right “recognizes the freedom of the individual to pick and choose for himself the time, circumstances, and particularly the extent to which he wishes to share or withhold from others his attitudes, beliefs, behavior, and opinions” (Shah, 1969, p. 57). When people in court proceedings “take the Fifth” and refuse to answer a question put to them on the grounds that the answer might be self-incriminating, they are asserting a right to privacy provided by the Fifth Amendment to the Constitution. The information withheld in such a manner is termed privileged; it is information that is protected by law from disclosure in a legal proceeding. State statutes have extended the concept of privileged information to parties who communicate with each other in the context of certain relationships, including the lawyer–client relationship, the doctor–patient relationship, the priest–penitent relationship, and the husband–wife relationship. In most states, privilege is also accorded to the psychologist–client relationship.
Privilege is extended to parties in various relationships because it has been deemed that the parties’ right to privacy serves a greater public interest than would be served if their communications were vulnerable to revelation during legal proceedings. Stated another way, it is for the social good if people feel confident that they can talk freely to their attorneys, clergy, physicians, psychologists, and spouses. Professionals such as psychologists who are parties to
such special relationships have a legal and ethical duty to keep their clients’ communications confidential.
Confidentiality may be distinguished from privilege in that, whereas “confidentiality concerns matters of communication outside the courtroom, privilege protects clients from disclosure in judicial proceedings” (Jagim et al., 1978, p. 459). Privilege is not absolute. There are occasions when a court can deem the
J U S T T H I N K . � . � .
Psychologists may be compelled by court order to reveal privileged communications. What types of situations might result in such a court order?
coh37025_ch02_041-084.indd 76 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
disclosure of certain information necessary and can order the disclosure of that information. Should the psychologist or other professional so ordered refuse, the professional does so under the threat of going to jail, being fined, and other legal consequences.
Privilege in the psychologist–client relationship belongs to the client, not the psychologist. The competent client can direct the psychologist to disclose information to some third party (such as an attorney or an insurance carrier), and the psychologist is obligated to make the disclosure. In some rare instances the psychologist may be ethically (if not legally) compelled to disclose information if that information will prevent harm either to the client or to some endangered third party. An illustrative case is the situation in which a client details a plan to die by suicide or commit homicide. In such an instance the psychologist would be legally and ethically compelled to take reasonable action to prevent the client’s intended outcome from occurring. Here, the preservation of life is deemed an objective more important than the nonrevelation of privileged information. Matters of ethics are seldom straightforward; questions will inevitably arise, and reasonable people may differ as to the answers to those questions. One such assessment-related ethics question has to do with the extent to which third-party observers should be allowed to be part of an assessment (see Figure 2–3). Some have argued that third parties are necessary and should be allowed, whereas others have argued that the presence of
Figure �–� Ethical issues when third-parties observe or participate in assessments.
Two necessary parties to any assessment are an assessor and an assessee. A third party might be an observer/supervisor of the assessor, a friend or relative of the assessee, a legal representative of the assesse or the institution in which the assessment is being conducted, a translator, or someone else. Ethical questions have been raised regarding the extent to which assessment data gathered in the presence of third parties is compromised due to a process of social influence (Duff & Fisher, 2005). Thomas Barwick/Stone/Getty Images
coh37025_ch02_041-084.indd 77 12/01/21 4:04 PM
�����Part 1: An Overview
the third party changes the dynamics of the assessment by a social influence process that may result in spurious increases or decreases in the assessee’s observed performance (Aiello & Douthitt, 2001; Gavett et al., 2005; McCaffrey, 2007; McCaffrey et al., 2005; Vanderhoff et al., 2011; Yantz & McCaffrey, 2005, 2009). Advocates of the strict enforcement of a policy that prohibits third-party observers during psychological assessment argue that alternatives to such observation either exist (e.g., unobtrusive electronic observation) or must be developed.
Another important confidentiality-related issue has to do with what a psychologist must keep confidential versus what must be disclosed. A wrong judgment on the part of the clinician regarding the revelation of confidential communication may lead to a lawsuit or worse. A landmark U.S. Supreme Court case in this area was the 1974 case of Tarasoff v. Regents of the University of California. In that case, a therapy patient had made known to his psychologist his intention to kill an unnamed but readily identifiable girl two months before the murder. The Court held that “protective privilege ends where the public peril begins,” and so the therapist had a duty to warn the endangered girl of her peril. Clinicians may have a duty to warn endangered third parties not only of potential violence but of potential infection from an HIV-positive client (Buckner & Firestone, 2000; Melchert & Patterson, 1999) as well as other threats to physical well-being.
Another ethical mandate with regard to confidentiality involves the safekeeping of test data. Test users must take reasonable precautions to safeguard test records. If these data are stored in a filing cabinet, then the cabinet should be locked and preferably made of steel. If these data are stored in a computer, electronic safeguards must be taken to ensure only authorized access. The individual or institution should have a reasonable policy covering the length of time that records are stored and when, if ever, the records will be deemed to be outdated, invalid, or useful only from an academic perspective. In general, it is not a good policy to maintain all records in perpetuity. Policies in conformance with privacy laws should
also be in place governing the conditions under which requests for release of records to a third party will be honored. Some states have enacted law that describes, in detail, procedures for storing and disposing of patient records.
Relevant to the release of assessment-related information is the Health Insurance Portability and Accountability Act of 1996 (HIPAA), which took effect in April 2003. These federal
privacy standards limit the ways that health care providers, health plans, pharmacies, and hospitals can use patients’ personal medical information. For example, personal health information may not be used for purposes unrelated to health care.
In part due to the decision of the U.S. Supreme Court in the case of Jaffee v. Redmond (1996), HIPAA singled out “psychotherapy notes” as requiring even more stringent protection than other records. The ruling in Jaffee affirmed that communications between a psychotherapist and a patient were privileged in federal courts. The HIPAA privacy rule cited Jaffee and defined privacy notes as “notes recorded (in any medium) by a health care provider who is a mental health professional documenting or analyzing the contents of conversation during a private counseling session or a group, joint, or family counseling session and that are separated from the rest of the individual’s medical record.” Although “results of clinical tests” were specifically excluded in this definition, we would caution assessment professionals to obtain specific consent from assessees before releasing assessment-related information. This is particularly essential with respect to data gathered using assessment tools such as the interview, behavioral observation, and role play.
The right to the least stigmatizing label�The Standards advise that the least stigmatizing labels should always be assigned when reporting test results. To better appreciate the need for this standard, consider the case of Jo Ann Iverson.4 Jo Ann was 9 years old and experiencing
J U S T T H I N K . � . � .
Describe key features of a model law designed to guide psychologists in the storage and disposal of patient records.
4. See Iverson v. Frandsen, 237 F. 2d 898 (Idaho, 1956) or Cohen (1979), pp. 149–150.
coh37025_ch02_041-084.indd 78 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
claustrophobia when her mother brought her to a state hospital in Blackfoot, Idaho, for a psychological evaluation. Arden Frandsen, a psychologist employed part-time at the hospital, conducted an evaluation of Jo Ann, during the course of which he administered a Stanford-Binet Intelligence Test. In his report, Frandsen classified Jo Ann as “feeble-minded, at the high-grade moron level of general mental ability.” Following a request from Jo Ann’s school guidance counselor, a copy of the psychological report was forwarded to the school—and embarrassing rumors concerning Jo Ann’s mental condition began to circulate.
Jo Ann’s mother, Carmel Iverson, brought a libel (defamation) suit against Frandsen on behalf of her daughter.5 Mrs. Iverson lost the lawsuit. The court ruled in part that the psychological evaluation “was a professional report made by a public servant in good faith, representing his best judgment.” But although Mrs. Iverson did not prevail in her lawsuit, we can certainly sympathize with her anguish at the thought of her daughter going through life with a label such as “high-grade moron”—this despite the fact that the psychologist had probably merely copied that designation from the test manual. We would also add that the Iversons may have prevailed in their lawsuit had the cause of action been breach of confidentiality and had the defendant been the school counselor; there was uncontested testimony that it was from the school counselor’s office, and not that of the psychologist, that the rumors concerning Jo Ann first emanated.
While on the subject of the rights of testtakers, let’s not forget about the rights—of sorts— of students of testing and assessment. Having been introduced to various aspects of the assessment enterprise, you have the right to learn more about technical aspects of measurement. Exercise that right in the succeeding chapters.
Self-Assessment
Test your understanding of elements of this chapter by seeing if you can explain each of the following terms, expressions, abbreviations, events, or names in terms of their significance in the context of psychological testing and assessment:
affirmative action Albemarle Paper Company v. Moody Alfred Binet James McKeen Cattell Charles Darwin Code of Fair Testing Practices in
Education code of professional ethics collectivist culture confidentiality culture culture-specific test Debra P. v. Turlington discrimination disparate impact disparate treatment ethics eugenics
Francis Galton Henry H. Goddard Griggs v. Duke Power Company HIPAA hired gun Hobson v. Hansen individualist culture informed consent Jaffee v. Redmond Larry P. v. Riles laws litigation minimum competency testing
programs Christiana D. Morgan Henry A. Murray ODDA Karl Pearson
privacy right privileged information projective test psychoanalysis Public Law 105-17 quota system reverse discrimination Hermann Rorschach self-report Sputnik standard of care Tarasoff v. Regents of the University
of California truth-in-testing legislation David Wechsler Lightner Witmer Robert S. Woodworth Wilhelm Max Wundt
5. An interesting though tangential aspect of this case was that Iverson had brought her child in with a presenting problem of claustrophobia. The plaintiff questioned whether the administration of an intelligence test under these circumstances was unauthorized and beyond the scope of the consultation. However, the defendant psychologist proved to the satisfaction of the Court that the administration of the Stanford-Binet was necessary to determine whether Jo Ann had the mental capacity to respond to psychotherapy.
coh37025_ch02_041-084.indd 79 12/01/21 4:04 PM
�����Part 1: An Overview
References
Abeles, N., & Barlev, A. (1999). End of life decisions and assisted suicide. Professional Psychology: Research and Practice, 30, 229–234.
Aggarwal, N. K., DeSilva, R., Nicasio, A. V., Boiler, M., & Lewis-Fernández, R. (2015). Does the Cultural Formulation Interview for the fifth revision of the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) affect medical communication? A qualitative exploratory study from the New�York site. Ethnicity & Health, 20, 1–28.
Aggarwal, N. K., Jiménez-Solomon, O., Lam, P. C., Hinton, L., & Lewis-Fernández, R. (2016). The core and informant Cultural Formulation Interviews in DSM-5. In R. Lewis-Fernández, N. K. Aggarwal, L. Hinton, D. E. Hinton, & L. J. Kirmayer (Eds.), DSM-5 Handbook on the Cultural Formulation Interview (pp.�27–44). American Psychiatric Association.
Aggarwal, N. K., Pieh, M. C., Dixon, L., Guarnaccia, P., Alegría, M., & Lewis-Fernández, R. (2016). Clinician descriptions of�communication strategies to improve treatment engagement by racial/ethnic minorities in mental health services: A systematic review. Patient Education and Counseling, 99, 198–209.
Aiello, J. R., & Douthitt, E. A. (2001). Social facilitation from Triplett to electronic performance monitoring. Group Dynamics, 5, 163–180.
Allard, G., Butler, J., Faust, D., & Shea, M. T. (1995). Errors in hand scoring objective personality tests: The case of the Personality Diagnostic Questionnaire— Revised (PDQ—R). Professional Psychology: Research and Practice, 26(3), 304–308.
American Educational Research Association, Committee on Test Standards, and National Council on Measurements Used in Education. (1955). Technical recommendations for achievement tests. American Educational Research Association.
American Psychological Association. (1950). Ethical standards for the distribution of psychological tests and diagnostic aids. American Psychologist, 5, 620–626.
American Psychological Association. (1985). Standards for educational and psychological testing. Author.
American Psychological Association. (1996). Affirmative action: Who benefits? Author.
American Psychological Association. (2015). Guidelines for psychological practice with transgender and gender nonconforming people. American Psychologist, 70(9), 832–864.
American Psychological Association (2017). 2003 ethical principles of psychologists and code of conduct, as amended 2010 and 2016. Retrieved July 6, 2019 from https://www.apa.org/ethics/code/ethics-code-2017.pdf
Amrine, M. (Ed.). (1965). Special issue. American Psychologist, 20, 857–991.
Appelbaum, P., & Grisso, T. (1995a). The MacArthur Treatment Competence Study: I. Mental illness and competence to consent to treatment. Law and Human Behavior, 19, 105–126.
Appelbaum, P., & Grisso, T. (1995b). The MacArthur Treatment Competence Study: III. Abilities of patients to consent to psychiatric and medical treatments. Law and Human Behavior, 19, 149–174.
Appelbaum, P. S., & Roth, L. H. (1982). Competency to consent to research: A psychiatric overview. Archives of General Psychiatry, 39, 951–958.
Beek, T. F., Kreukels, B. P. C., Cohen–Kettenis, P. T., & Steensma, T. D. (2015). Partial treatment requests and underlying motives of applicants for gender affirming interventions. Journal of Sexual Medicine, 12(11), 2201–2205.
Benbow, C. P., & Stanley, J. C. (1996). Inequity in equity: How “equity” can lead to inequity for high-potential students. Psychology, Public Policy, and Law, 2, 249–292.
Binet, A., & Henri, V. (1895a). La mémoire des mots. L’Année Psychologique, 1, 1–23.
Binet, A., & Henri, V. (1895b). La mémoire des phrases. L’Année Psychologique, 1, 24–59.
Binet, A., & Henri, V. (1895c). La psychologie individuelle. L’Année Psychologique, 2, 411–465.
Binet, A., & Simon, T. (1905). Méthodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L’Année Psychologique, 11, 191–244.
Bogacki, D. F., & Weiss, K. J. (2007). Termination of parental rights: Focus on defendants. Journal of Psychiatry & Law, 35(1), 25–45.
Borden, K. A. (2015). Introduction to the special section transgender and gender noncomforming individuals: Issues for professional psychologists. Professional Psychology: Research and Practice, 46(1), 1–2.
Boring, E. G. (1950). A history of experimental psychology (rev. ed.). Appleton-Century-Crofts.
Bouman, W. P., Richards, C., Addinall, R. M., et al. (2014). Yes and yes again: Are standards of care which require two referrals for genital reconstructive surgery ethical? Sexual and Relationship Therapy, 29(4), 377–389.
Brotemarkle, R. A. (1947). Clinical psychology, 1896– 1946. Journal of Consulting and Clinical Psychology, 11, 1–4.
Buckner, F., & Firestone, M. (2000). “Where the public peril begins”: 25 years after Tarasoff. Journal of Legal Medicine, 21, 187–222.
Bumann, B. (2010). The Future of Neuroimaging in Witness Testimony. Virtual Mentor: American Medical Association Journal of Ethics, 12, 873–878.
Buros, O. K. (1938). The 1938 mental measurements yearbook. Rutgers University Press.
Callahan, J. (1994). The ethics of assisted suicide. Health and Social Work, 19, 237–244.
Carpenter, W. T., Gold, M. J., Lahti, A. C., et al. (2000). Decisional capacity for informed consent in schizophrenia research. Archives of General Psychiatry, 57, 533–538.
Chen, Y., Nettles, M. E., & Chen, S.-W. (2009). Rethinking dependent personality disorder: Comparing different human relatedness in cultural contexts. Journal of Nervous and Mental Disease, 197, 793–800.
Code of Fair Testing Practices in Education. (2004). Joint Committee on Testing Practices.
Cohen, R. J. (1979). Malpractice: A guide for mental health professionals. Free Press.
Cohen, R. J. (1994). Psychology & adjustment: Values, culture, and change. Allyn & Bacon.
Coyne, I., & Bartram, D. (2006). Design and development of the ITC guidelines on computer-based and internet- delivered testing. International Journal of Testing, 6(2), 133–142.
coh37025_ch02_041-084.indd 80 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
Crosby, F. J., Iyer, A., Clayton, S., & Downing, R. A. (2003). Affirmative action: Psychological data and the policy debates. American Psychologist, 58, 93–115.
da Silva, D. C., Schwarz, K., Fontanari, A. M. V., et al. (2016). Whoqol-100 before and after sex reassignment surgery in Brazilian male-to-female transsexual individuals. Journal of Sexual Medicine, 13(6), 988–993.
Darwin, C. (1859). On the origin of species by means of natural selection. Murray.
Daubert v. Merrell Dow Pharmaceuticals, 113 S. Ct. 2786 (1993).
de Vries, A. L. C., McGuire, J. K.. Steensma, T. D., et al. (2014). Young adult psychological outcome after puberty suppression and gender reassignment. Pediatrics, 134(4), 696–704.
Dèttore, D., Ristori, J., Antonelli, P., et al. (2015). Gender dysphoria in adolescents: The need for a shared assessment protocol and proposal of the AGIR protocol. Giornale di Psicopatologia [Journal of Psychopathology], 21(2), 152–158.
Dhejne, C., Van Vlerken, R., Heylens, G., & Arcelus, J. (2016). Mental health and gender dysphoria: A review of the literature. International Review of Psychiatry, 28(1), 44–57.
DuBois, P. H. (1966). A test-dominated society: China 1115 B.C.E–1905 A.D. In A. Anastasi (Ed.), Testing problems in perspective (pp. 29–36). American Council on Education.
DuBois, P. H. (1970). A history of psychological testing. Allyn & Bacon.
Duff, K., & Fisher, J. M. (2005). Ethical dilemmas with third party observers. Journal of Forensic Neuropsychology, 4(2), 65–82.
Dunn, L. B., Lindamer, L. A., Palmer, B. W., et al. (2002). Improving understanding of research consent in middle-aged and elderly patients with psychotic disorders. American Journal of Geriatric Psychology, 10, 142–150.
Ewing, C. P., and McCann, J. T. (2006). Minds on trial: Great cases in law and psychology. Oxford University Press.
Farrenkopf, T., & Bryan, J. (1999). Psychological consultation under Oregon’s 1994 Death With Dignity Act: Ethics and procedures. Professional Psychology: Research and Practice, 30, 245–249.
Federal Rules of Evidence. (1975). West Group. Fenn, D. S., & Ganzini, L. (1999). Attitudes of Oregon
psychologists toward physician-assisted suicide and the Oregon Death With Dignity Act. Professional Psychology: Research and Practice, 30, 235–244.
Finucane, M. L., & Gullion, C. M. (2010). Developing a tool for measuring the decision-making competence of older adults. Psychology and Aging, 25, 271–288.
Forbey, J. D., & Ben-Porath, Y. S. (2007). Computerized adaptive personality testing: A review and illustration with the MMPI-2 computerized adaptive version. Psychological Assessment, 19(1), 14–24. https://doi. org/10.1037/1040-3590.19.1.14
Forrest, D. W. (1974). Francis Galton: The life and works of a Victorian genius. Taplinger.
Freud, S. (1913/1959). Further recommendations in the technique of psychoanalysis. In E. Jones (Ed.) and J. Riviere (Trans.), Collected papers (Vol. 2). Basic Books.
Frolik, L. A. (1999). Science, common sense, and the determination of mental capacity. Psychology, Public Policy, and Law, 5, 41–58.
Frye v. United States, 293 Fed. 1013 (D.C. Cir. 1923). Garrett, H. E., & Schneck, M. R. (1933). Psychological
tests, methods and results. Harper. Gavett, B. E., Lynch, J. K., & McCaffrey, R. J. (2005).
Third party observers: The effect size is greater than you might think. Journal of Forensic Neuropsychology, 4(2), 49–64.
General Electric Co. v. Joiner, 118 S. Ct. 512 (1997). Goddard, H. H. (1912). The Kallikak family: A study in
the heredity of feeble-mindedness. The MacMillan Company.
Goddard, H. H. (1913). The Binet tests in relation to immigration. Journal of Psycho-Asthenics, 18, 105–107.
Goddard, H. H. (1917). Mental tests and the immigrant. Journal of Delinquency, 2, 243–277.
Goodrich, K. M., Harper, A. J., Luke, M., & Singh, A. A. (2013). Best practices for professional school counselors working with LGBTQ youth. Journal of LGBT Issues in Counseling, 7(4), 307–322.
Gopaul-McNicol, S. (1993). Working with West Indian families. Guilford.
Gottfredson, L. S. (2000). Skills gaps, not tests, make racial proportionality impossible. Psychology, Public Policy, and Law, 6, 129–143.
Gould, J. W. (2006). Conducting scientifically crafted child custody evaluations (2nd ed.). Professional Resource Press/Professional Resource Exchange.
Grisso T., & Appelbaum, P. S. (1995). MacArthur Treatment Competence Study. Journal of American Psychiatric Nurses Association, 1, 125–127.
Grisso, T., & Appelbaum, P. S. (1998). Assessing competence to consent to treatment: A guide for physicians and other health professionals. Oxford University Press.
Grisso, T., Appelbaum, P. S., & Hill-Fotouhi, C. (1997). The MacCAT-T: A clinical tool to assess patients’ capacities to make treatment decisions. Psychiatric Services, 48, 1415–1419.
Grove, W. M., & Barden, R. C. (1999). Protecting the integrity of the legal system: The admissibility of testimony from mental health experts under Daubert/ Kumho analyses. Psychology, Public Policy, and Law, 5, 224–242.
Haley, K., & Lee, M. (Eds.). (1998). The Oregon Death With Dignity Act: A guidebook for health care providers. Oregon Health Sciences University, Center for Ethics in Health Care.
Halpern, D. F. (2000). Validity, fairness, and group differences: Tough questions for selection testing. Psychology, Public Policy, & Law, 6, 56–62.
Haney, W. (1981). Validity, vaudeville, and values: A short history of social concerns over standardized testing. American Psychologist, 36, 1021–1034.
Haque, S., & Guyer, M. (2010). Neuroimaging studies in diminished-capacity defense. Journal of the American Academy of Psychiatry and the Law, 38(4), 605–607.
Hartigan, J. A., & Wigdor, A. K. (1989). Fairness in employment testing: Validity generalization, minority issues, and the General Aptitude Test Battery. The National Academies Press. https://doi.org/10 .17226/1338.
Hoffman, B. (1962). The tyranny of testing. Crowell- Collier.
Institute for Juvenile Research. (1937). Child guidance procedures, methods and techniques employed at the Institute for Juvenile Research. Appleton-Century.
coh37025_ch02_041-084.indd 81 12/01/21 4:04 PM
�����Part 1: An Overview
Iverson v. Frandsen, (Idaho, 1956), 237 F. 2d 898. Jaffee v. Redmond (1996), 518 U.S. 1; 116 S. Ct. (1923). Jagim, R. D., Wittman, W. D., & Noll, J. O. (1978).
Mental health professionals’ attitudes towards confidentiality, privilege, and third-party disclosure. Professional Psychology, 9, 458–466.
Jennings, B. (1991). Active euthanasia and forgoing life- sustaining treatment: Can we hold the line? Journal of Pain, 6, 312–316.
Jensen, A. R. (1969). How much can we boost IQ and scholastic achievement? Harvard Educational Review, 39, 1–123.
Johnson, L. L., Bradley, S. J., Birkenfeld-Adams, A. S., et al. (2004). A parent-report gender identity questionnaire for children. Archives of Sexual Behavior, 33(2), 105–116.
Johnson, S. M., Cramer, R. J., Conroy, M. A., & Gardner, B. O. (2014). The role of and challenges for psychologists in physician assisted suicide. Death Studies, 38(9), 582–588.
Johnson, S. M., Cramer, R. J., Gardner, B. O., & Nobles, M. R. (2015). What patient and psychologist characteristics are important in competency for physician-assisted suicide evaluations? Psychology, Public Policy, and Law, 21(4), 420–431.
Kelley, T. L. (1927). Interpretation of educational measurements. World Book.
Kirmayer, L. K. (2006). Beyond the “new cross-cultural psychiatry”: Cultural biology, discursive psychology and the ironies of globalization. Transcultural Psychiatry, 43, 126–144.
Kleinman, A. (1988). Rethinking psychiatry: From cultural category to personal experience. The Free Press.
Klopfer, B., Ainsworth, M., Klopfer, W., & Holt, R. R. (1954). Developments in the Rorschach technique: Vol. 1. Technique and theory. World.
Korman, A. K. (1988). The outsiders: Jews and corporate America. Lexington.
Krach, S. K., McCreery, M. P., Dennis, L., Guerard, J., & Harris, E. L. (2020). Independent evaluation of Q-Interactive: A paper equivalency comparison using the PPVT-4 with preschoolers. Psychology in the Schools, 57(1), 17–30.
Kraepelin, E. (1892). Uber die Beeinflussing einfacher psychischer Vorgange durch einige Arzneimittel. Fischer.
Kraepelin, E. (1895). Der psychologische versuch in der psychiatrie. Psychologische Arbeiten, 1, 1–91.
Krauss, D. A., & Sales, B. D. (1999). The problem of “helpfulness” in applying Daubert to expert testimony: Child custody determinations in family law as an exemplar. Psychology, Public Policy, and Law, 5, 78–99.
Kubiszyn, T. W., Meyer, G. J., Finn, S. E., Eyde, L. D., Kay, G. G., Moreland, K. L., Dies, R. R., & Eisman, E. J. (2000). Empirical support for psychological assessment in clinical health care settings. Professional Psychology: Research and Practice, 31, 119–130. https://doi.org/10.1037/0735-7028.31.2.119
Kumho Tire Co. Ltd. v. Carmichael, 119 S. Ct. 1167 (1999). Landy, F. J. (2007). The validation of personnel decisions
in the twenty-first century: Back to the future. In S. M. McPhail (Ed.), Alternative validation strategies: Developing new and leveraging existing evidence (pp. 409–426). Wiley.
Latimer, E. J. (1991). Ethical decision-making in the care of the dying and its applications to clinical practice. Journal of Pain and Symptom Management, 6, 329–336.
Lavin, M. (1992). The Hopkins Competency Assessment Test: A brief method for evaluating patients’ capacity to give informed consent. Hospital and Community Psychiatry, 646, 132–136.
Leach, M. M., & Oakland, T. (2007). Ethics standards impacting test development and use: A review of 31 ethics codes impacting practices in 35 countries. International Journal of Testing, 7(1), 71–88.
Leahy, M. M., Easton, C. J., & Edwards, L. M. (2010). Compulsory psychiatric testing. Journal of the American Academy of Psychiatry and the Law, 38(1), 126–128.
Lewis-Fernández, R., Aggarwal, N. K., Hinton, L., Hinton, D. E., & Kirmayer, L. J. (2016). DSM-5 handbook on the Cultural Formulation Interview. American Psychiatric Association.
Lipton, J. P. (1999). The use and acceptance of social science evidence in business litigation after Daubert. Psychology, Public Policy, and Law, 5, 59–77.
Luyt, R. (2015). Beyond traditional understanding of gender measurement: The gender (re)presentation approach. Journal of Gender Studies, 24(2), 207–226.
Mael, F. A. (1991). Career constraints of observant Jews. Career Development Quarterly, 39, 341–349.
Magnello, M. E., & Spies, C. J. (1984). Francis Galton: Historical antecedents of the correlation calculus. In B. Laver (Chair), History of mental measurement: Correlation, quantification, and institutionalization. Paper session presented at the 92nd annual convention of the American Psychological Association, Toronto.
Marble-Flint, K. J., Strattman, K. H., & Schommer- Aikins, M. A. (2019). Comparing iPad® and paper assessments for children with ASD: An initial study. Communication Disorders Quarterly, 40(3), 152–155.
Mark, M. M. (1999). Social science evidence in the courtroom: Daubert and beyond? Psychology, Public Policy, and Law, 5, 175–193.
Markus, H., & Kitayama, S. (1991). Culture and the self: Implications for cognition, emotion, and motivation. Psychological Review, 98, 224–253.
Marson, D. C., McInturff, B., Hawkins, L., Bartolucci, A., & Harrell, L. E. (1997). Consistency of physician judgments of capacity to consent in mild Alzheimer’s disease. American Geriatrics Society, 45, 453–457.
McCaffrey, R. J. (2007). Participant. In A. E. Puente (Chair), Third party observers in psychological and neuropsychological forensic psychological assessment. Symposium presented at the 115th Annual Convention of the American Psychological Association, San Francisco, CA.
McCaffrey, R. J., Lynch, J. K., & Yantz, C. L. (2005). Third party observers: Why all the fuss? Journal of Forensic Neuropsychology, 4(2), 1–15.
McLearen, A. M., Pietz, C. A., & Denney, R. L. (2004). Evaluation of psychological damages. In W. T. O’Donohue & E. R. Levensky (Eds.), Handbook of forensic psychology (pp. 267–299). Elsevier.
McNemar, Q. (1975). On so-called test bias. American Psychologist, 30, 848–851.
McReynolds, P. (1987). Lightner Witmer: Little-known founder of clinical psychology. American Psychologist, 42, 849–858.
Melchert, T. P., & Patterson, M. M. (1999). Duty to warn and interventions with HIV-positive clients. Professional Psychology: Research and Practice, 30, 180–186.
Mills v. Board of Education of the District of Columbia, 348 F. Supp 866 (D. DC 1972).
coh37025_ch02_041-084.indd 82 12/01/21 4:04 PM
Chapter 2: Historical, Cultural, and Legal/Ethical Considerations ��
Mossman, D. (2003). Daubert, cognitive malingering, and test accuracy. Law and Human Behavior, 27(3), 229–249.
Myerson, A. (1925). The inheritance of mental disease. Williams & Wilkins.
Neisser, U., Boodoo, G., Bouchard, T. J., Boykin, A. W., Brody, N., Ceci, S. J., Halpern, D. F., Loehlin, J. C., Perloff, R., Sternberg, R. J., & Urbina, S. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51, 77–101.
Newman, D. A., Kinney, T., & Farr, J. L. (2004). Job performance ratings. In J. C Thomas (Ed.), Comprehensive handbook of psychological assessment, Volume 4: Industrial and organizational assessment (pp. 373–389). Wiley.
Noland, R. M. (2017). Intelligence testing using a tablet computer: Experiences with using Q-interactive. Training and Education in Professional Psychology, 11(3), 156–163. https://doi.org/10.1037/tep0000149
Nyquist, A. C., & Forbey, J. D. (2018). An investigation of a computerized sequential depression module of the MMPI-2. Assessment, 25(8), 1084–1097.
Oregon Death With Dignity Act, 2 Ore. Rev. Stat. §§127.800–127.897 (1997).
Palmer, B. W., Dunn, L. B., Depp, C. A., Eyler, L. T., & Jeste, D. V. (2007). Decisional capacity to consent to research among patients with bipolar disorder: comparison with schizophrenia patients and healthy subjects. Journal of Clinical Psychiatry, 68, 689–696.
PARC v. Pennsylvania, 334 F. Supp. 1257 (E.D. PA 1972). Pierson, D. (1974). The Trojans: Southern California
football. H. Regnery Co. Pintner, R. (1931). Intelligence testing. Holt. Posthuma, A., Podrouzek, W., & Crisp, D. (2002). The
implications of Daubert on neuropsychological evidence in the assessment of remote mild traumatic brain injury. American Journal of Forensic Psychology, 20(4), 21–38.
Poythress, N. G. (2004). Editorial. “Reasonable medical certainty:” Can we meet Daubert standards in insanity cases? Journal of the American Academy of Psychiatry and the Law, 32, 228–230.
Quill, T. E., Cassel, C. K., & Meier, D. E. (1992). Care of the hopelessly ill: Proposed clinical criteria for physician-assisted suicide. New England Journal of Medicine, 327, 1380–1384.
Reed, G. M., McLaughlin, C. J., & Newman, R. (2002). American Psychological Association policy in context: The development and evaluation of guidelines for professional practice. American Psychologist, 57, 1041–1047.
Reynolds, L. (2014). Losing the quality of life: The move towards society’s understanding and acceptance of physician aid-in-dying and the Death with Dignity Act. New England Law Review, 48(2), 343–370.
Richman, J. (1988). The case against rational suicide. Suicide & Life-Threatening Behavior, 18, 285–289.
Roback, A. A. (1961). History of psychology and psychiatry. Philosophical Library.
Rönspies, J., Schmidt, A. F., Melnikova, A., et al. (2015). Indirect measurement of sexual orientation: Comparison of the implicit relational assessment procedure, viewing time, and choice reaction time tasks. Archives of Sexual Behavior, 44(5), 1483–1492.
Rosen, J. (1998, February 23/March 2). Damage control. New Yorker, 74, 64–68.
Roth, L. H., Meisel, A., & Lidz, C. W. (1977). Tests of competence to consent to treatment. American Journal of Psychiatry, 134, 279–284.
Ruch, G. M. (1925). Minimum essentials in reporting data on standard tests. Journal of Educational Research, 12, 349–358.
Ruch, G. M. (1933). Recent developments in statistical procedures. Review of Educational Research, 3, 33–40.
Saldanha, C. (2005). Daubert and suicide risk of antidepressants in children. Academy of Psychiatry and the Law, 33(1), 123–125.
Sale, R. (2006). International guidelines on computer- based and internet-delivered testing: A practitioner’s perspective. International Journal of Testing, 6(2), 181–188.
Sanders, J. (2010). Applying Daubert Inconsistently?: Proof of individual causation in toxic tort and forensic cases. 75 Brooklyn Law Review, 1367, 1370–1374.
Saxe, L., & Ben-Shakhar, G. (1999). Admissibility of polygraph tests: The application of scientific standards post-Daubert. Psychology, Public Policy, and Law, 5, 203–223.
Schmidt, F. L. (1988). The problem of group differences in ability scores in employment selection. Journal of Vocational Behavior, 33, 272–292.
Schmidt, F. L., & Hunter, J. E. (1992). Development of a causal model of processes determining job performance. Current Directions in Psychological Science, 1, 89–92.
Searight, H. R., & Searight, B. K. (2009). Working with foreign language interpreters: Recommendations for psychological practice. Professional Psychology: Research and Practice, 40, 454–451.
Shah, S. A. (1969). Privileged communications, confidentiality, and privacy: Privileged communications. Professional Psychology, 1, 56–59.
Sherman, M. D., Kauth, M. R., Shipherd, J. C., et al. (2014). Provider beliefs and practices about assessing sexual orientation in two Veterans Health Affairs hospitals. LGBT Health, 1(3), 185–191.
Slobogin, C. (1999). The admissibility of behavioral science information in criminal trials: From primitivism to Daubert to Voice. Psychology, Public Policy, and Law, 5, 100–119.
Smith, J. D. (1985). Minds made feeble: The myth and legacy of the Kallikaks. Pro–Ed.
Smith, K. A., Harvath, T. A., Goy, E. R., & Ganzini, L. (2015). Predictors of pursuit of physician-assisted death. Journal of Pain and Symptom Management, 49(3), 555–561.
Stephens, J. J. (1992). Assessing ethnic minorities. SPA Exchange, 2(1), 4–6.
Stern, B. H. (2001). Admissability of neuropsychological testimony after Daubert and Kumho. NeuroRehabilitation, 16(2), 93–101.
Stolberg, R. A. (2018). Influence of the Caldwell report, and other computer generated interpretive reports, in child custody evaluations: A brief report. Journal of Child Custody, 15(4), 369–378.
Sturman, E. D. (2005). The capacity to consent to treatment and research: A review of standardized assessment tools and potentially impaired populations. Clinical Psychology Review, 25, 954–974.
Sylvester, R. H. (1913). Clinical psychology adversely criticized. Psychological Clinic, 7, 182–188.
Tarasoff v. Regents of the University of California, 17 Cal. 3d 425, 551 P.2d 334, 131 Cal. Rptr. 14 (Cal. 1976).
coh37025_ch02_041-084.indd 83 12/01/21 4:04 PM
�����Part 1: An Overview
Tenopyr, M. L. (1999). A scientist-practitioner’s viewpoint on the admissibility of behavioral and social scientific information. Psychology, Public Policy, and Law, 5, 194–202.
Terman, L. M. (1916). The measurement of intelligence: An explanation of and a complete guide for the use of the Stanford revision and extension of the Binet-Simon Intelligence Scale. Houghton Mifflin.
Trent, J. W. (2001). “Who shall say who is a useful person?” Abraham Myerson’s opposition to the eugenics movement. History of Psychiatry, 12(45, Pt.1), 33–57.
Tulchin, S. H. (1939). The clinical training of psychologists and allied specialists. Journal of Consulting Psychology, 3, 105–112.
United States. Public Health Service. Office of the Surgeon General, Center for Mental Health Services (U.S.), National Institute of Mental Health (U.S.), United States. Substance Abuse, & Mental Health Services Administration. (2001). Mental health: Culture, race, and ethnicity: A supplement to mental health: A report of the Surgeon General (Vol. 2). Department of Health and Human Services, U.S. Public Health Service.
Vanderhoff, H., Jeglic, E. L., & Donovick, P. J. (2011). Neuropsychological assessment in prisons: Ethical and practical challenges. Journal of Correctional Health Care, 17, 51–60.
Vollmann, J., Bauer, A., Danker-Hopfe, H., & Helmchen, H. (2003). Competence of mentally ill patients: a comparative empirical study. Psychological Medicine, 33, 1463–1471.
von Wolff, C. (1732). Psychologia empirica. von Wolff, C. (1734). Psychologia rationalis. Wahlstrom, D. A., Daniel, P. M., & Weiss, L. G. (2019).
Digital assessment with Q-interactive. In A. Prifitera, D. H. Saklofske, L. G. Weiss, & J. A. Holdnack (Eds.), WISC-V: Clinical Use and Interpretation (2nd ed.; pp. 417–446). Academic Press.
Wang, R. (2012). The Chinese imperial examination system: An annotated bibliography. Scarecrow Press.
Wechsler, D. (1939). The measurement of adult intelligence. Williams & Wilkins.
Wechsler, D. (1944). The measurement of adult intelligence (3rd ed.). Williams & Wilkins.
Wehmeyer, M. L., & Smith, J. D. (2006). Leaving the garden: Henry Herbert Goddard’s exodus from the Vineland Training School. Mental Retardation, 44, 150–155.
Weir, R. F. (1992). The morality of physician-assisted suicide. Law, Medicine and Health Care, 20, 116–126.
White, C. (2015). Physician aid-in-dying. Houston Law Review, 53(2), 595–629.
Witmer, L. (1907). Clinical psychology. Psychological Clinic, 1, 1–9.
Wolfram, W. A. (1971). Social dialects from a linguistic perspective: Assumptions, current research, and future directions. In R. Shuy (Ed.), Social dialects and interdisciplinary perspectives. Center for Applied Linguistics.
Wylie, K., Barrett, J., Besser, M., et al. (2014). Good practice guidelines for the assessment and treatment of adults with gender dysphoria. Sexual and Relationship Therapy, 29(2), 154–214.
Yantz, C. L., & McCaffrey, R. J. (2005). Effects of a supervisor’s observation on memory test performance of the examinee: Third party observer effect confirmed. Journal of Forensic Neuropsychology, 4(2), 27–38.
Yantz, C. L., & McCaffrey, R. J. (2009). Effects of parental presence and child characteristics on children’s neuropsychological test performance: Third party observer effect confirmed. Clinical Neuropsychologist, 23, 118–132.
Zenderland, L. (1998). Measuring minds: Henry Herbert Goddard and the origins of American intelligence testing. Cambridge University Press.
Zink v. State, 278 SW 3d 170–Mo: Supreme Court 2009.
Zweigenhaft, R. L. (1984). Who gets to the top? Executive suite discrimination in the eighties. American Jewish Committee Institute of Human Relations.
coh37025_ch02_041-084.indd 84 12/01/21 4:04 PM
��
rom the red-pencil number circled at the top of your first spelling test to the computer printout of your college entrance examination scores, tests and test scores touch your life. They seem to reach out from the paper and shake your hand when you do well and punch you in the face when you do poorly. They can point you toward or away from a particular school or curriculum. They can help you to identify strengths and weaknesses in your physical and mental abilities. They can accompany you on job interviews and influence your career choices.
In your role as a student, you have probably found that your relationship to tests has been primarily that of a testtaker. But as a psychologist, teacher, researcher, or employer, you may find that your relationship with tests is primarily that of a test user—the person who breathes life and meaning into test scores by applying the knowledge and skill to interpret them appropriately. You may one day create a test, whether in an academic or a business setting, and then have the responsibility for scoring and interpreting it. In that situation, or even from the perspective of someone who would take that test, it is essential to understand the theory underlying test use and the principles of test-score interpretation.
Test scores are frequently expressed as numbers, and statistical tools are used to describe, make inferences from, and draw conclusions about numbers.1 In this statistics refresher, we cover scales of measurement, tabular and graphic presentations of data, measures of central tendency, measures of variability, aspects of the normal curve, and standard scores. If these statistics-related terms look painfully familiar to you, we ask your indulgence and ask you to remember that overlearning is the key to retention. Of course, if any of these terms appear unfamiliar, we urge you to learn more about them. Feel free to supplement the discussion here with a review of these and related terms in any good elementary statistics text. The brief review of statistical concepts that follows can in no way replace a sound grounding in basic statistics gained through an introductory course in that subject.
F
C H A P T E R �
A Statistics Refresher
J U S T T H I N K � . � . � .
For most people, test scores are an important fact of life. But what makes those numbers so�meaningful? In general terms, what information, ideally, should be conveyed by a test score?
1. Of course, a test score may be expressed in other forms, such as a letter grade or a pass–fail designation. Unless stated otherwise, terms such as test score, test data, test results, and test scores are used throughout this book to refer to numeric descriptions of test performance.
coh37025_ch03_085-128.indd 85 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
Scales of Measurement
We may formally define measurement as the act of assigning numbers or symbols to characteristics of things (people, events, whatever) according to rules (Stevens, 1946). The rules used in assigning numbers are guidelines for representing the magnitude (or some other
characteristic) of the object being measured. Here is an example of a measurement rule: Assign the number 12 to all lengths that are exactly the same length as a 12-inch ruler. A scale is a set of numbers (or other symbols) whose properties model empirical properties of the objects to which the numbers are assigned.2
The sample space of a variable refers to the values that a variable can take on. For example, if you collect data on study participants’ gender, the sample space might be {male, female, nonbinary}. The sample space for participants’ age in years might be natural integers. In theory, the natural integers extend to positive infinity {0, 1, 2, .� .� .}, but in practice few participants will be older than 100. The sample space for participants’ height in centimeters might be any positive real number [0,+�], even though no one has a height near 0 or much higher than 200 cm.
There are various ways in which a scale can be categorized. One important distinction between scales is whether the variable is discrete or continuous (see Figure 3–1). A discrete scale has a sample space that can be counted. A categorical variable like year in high school has four members in its sample space: {freshman, sophomore, junior, senior}. Quantitative variables like a patient’s number of previous hospitalizations are discrete because the sample space is countable: {0, 1, 2, 3, .�.�.}. In discrete variables, numbers between the sample space members are not allowed. For example, a patient cannot have 2.5 previous hospitalizations.
In a continuous scale, the values can be any real number in the scale’s sample space. Continuous scales therefore can have fractions or numbers with as many decimals as needed. In theory, a continuous scale could have irrational numbers like the square root of 2 or a transcendental number like �. In practice, measurements have to be rounded.
In general, it is best to round continuous scales so that the numbers do not convey unwarranted precision. For example, a scale should not round to the nearest hundredth of a gram if it is not accurate enough to detect a change of 0.01 g. When such precision is not possible, rounding to the nearest tenth or nearest integer would be better. If rounding to the
J U S T T H I N K � . � . � .
What is another example of a measurement�rule?
2. David L. Streiner reflected, “Many terms have been used to describe a collection of items or questions—scale, test, questionnaire, index, inventory, and a host of others—with no consistency from one author to another” (2003, p. 217, emphasis in the original). Streiner proposed to refer to questionnaires of theoretically like or related items as scales and those of theoretically unrelated items as indexes. He acknowledged that counterexamples of each term could readily be found.
Figure �–� Discrete vs. continuous variables.
In this example, the discrete variable can only take on the natural numbers from 0 to 10. By contrast, continuous variables can be any real number within a specified range.
Continuous
Discrete
0 1 2 3 4 5 6 7 8 9 10
Possible Values a Variable Can Assume
coh37025_ch03_085-128.indd 86 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
nearest integer still implies more precision than the scale has, then rescaling the variable to a larger metric might be needed (e.g., converting grams to kilograms).
In everyday usage, the word “error” implies that someone made a mistake. In the context of scientific measurement, error has a broader meaning. In the language of assessment, error refers to the collective influence of all of the factors on a test score or measurement beyond those specifically measured by the test or measurement. As we will see, there are many different sources of error in measurement, most of which have more to do with uncertainty of the measurement than they do with mistakes. Consider, for example, the score someone received on a test in American history. We might conceive of part of the score as reflecting the testtaker’s knowledge of American history and part of the score as reflecting measurement error. The error part of the test score may be due to many different factors. One source of error might have been a distracting thunderstorm going on outside at the time the test was administered. Another source of error was the particular selection of test items the instructor chose to use for the test. Had a different item or two been used in the test, the testtaker’s score on the test might have been higher or lower. Error is an element of all measurement, and it is an element for which any theory of measurement must surely account.
Measurement using continuous scales always involves error. To illustrate why, imagine that you are ordering blinds for a window. If the ruler measures to the nearest tenth inch, then a width read as 35.5 inches might really be 35.484 inches. The measuring scale is conveniently marked off in grosser gradations of measurement. Most scales used in psychological and educational assessment are continuous and therefore can be expected to contain this sort of error. The number or score used to characterize the trait being measured on a continuous scale should be thought of as an approximation of the “real” number. Thus, for example, a score of 25 on a test of anxiety should not be thought of as a precise measure of anxiety. Rather, it should be thought of as an approximation of the real anxiety score had the measuring instrument been calibrated to yield such a score. In such a case, perhaps the score of 25 is an approximation of a real score of, say, 24.7 or 25.44.
Beyond the continuous versus discrete distinction, it is convenient to distinguish between levels of measurement as first proposed by Stevens (1946). Within these levels or scales of measurement, assigned numbers convey different kinds of information. Accordingly, certain statistical manipulations may or may not be appropriate, depending upon the level or scale of measurement.3
The French word for black is noir (pronounced “‘nwa� re”). We bring this up here only to call attention to the fact that this word is a useful acronym for remembering the four levels or scales of measurement shown in Figure 3–2. Each letter in noir is the first letter of the succeedingly more rigorous levels: N stands for nominal, o for ordinal, i for interval, and r for ratio scales.
J U S T T H I N K . � . � .
Assume the role of a test creator. Now write some instructions to users of your test that are designed to reduce to the absolute minimum any error associated with test scores. Be sure to include instructions regarding the preparation of the site where the test will be administered.
J U S T T H I N K � . � . � .
The scale with which we are all perhaps most familiar is the common bathroom scale. How are a psychological test and a bathroom scale alike? How are they di�erent? Your answer may change as you read on.
J U S T T H I N K . � . � .
Acronyms like noir are useful memory aids. As you continue in your study of psychological testing and assessment, create your own acronyms to help remember related groups of information. Hey, you may even learn some French in the process.
3. For the purposes of our statistics refresher, we present what Nunnally (1978) called the “fundamentalist” view of measurement scales, which “holds that 1. there are distinct types of measurement scales into which all possible measures of attributes can be classified, 2. each measure has some ‘real’ characteristics that permit its proper classification, and 3. once a measure is classified, the classification specifies the types of mathematical analyses that can be employed with the measure” (p. 24). Nunnally and others have acknowledged that alternatives to the “fundamentalist” view may also be viable.
coh37025_ch03_085-128.indd 87 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
Nominal Scales
Nominal scales are the simplest form of measurement. These scales involve classification or categorization based on one or more distinguishing characteristics, where all things measured must be placed into mutually exclusive and exhaustive categories. For example, researchers studying college students might ask what their current major is. For the sake of convenience, college majors might be listed alphabetically (e.g., Accounting, Biology, Chemistry, .� .� .), but there is no inherent order to college majors.
Many demographic variables like gender, race, or place of birth are nominal because they are categories with no defined order. Although numbers usually indicate quantities, it is possible for numbers to serve as unique identifiers such as telephone numbers, zip codes, and social security numbers. These are nominal variables, not quantities. For example, although we might sort telephone numbers for convenience, there is no sense in which one telephone number is “higher” than another.
An example of a “numeric” nominal variable in assessment can be found in the Diagnostic and Statistical Manual of Mental Disorders. Each disorder listed in that manual is assigned its own number. In a past version of that manual, the version really does not matter for the purposes of this example, the number 303.00 identified alcohol intoxication, and the number 307.00 identified stuttering. But these numbers were used exclusively for classification purposes and could not be meaningfully added, subtracted, ranked, or averaged. Hence, the middle number between these two diagnostic codes, 305.00, did not identify an intoxicated stutterer.
Individual test items may also employ nominal scaling, including yes/no responses. For example, consider the following test items:
Instructions: Answer either yes or no.
Are you actively contemplating suicide? __________
Are you currently under professional care for a psychiatric disorder? _______
Have you ever been convicted of a felony? _______
In each case, a yes or no response results in the placement into one of a set of mutually exclusive groups: suicidal or not, under care for psychiatric disorder or not, and felon or not. In Figure 3–2, measurement with nominal variables consists of assigning individuals to one and only one category. The possible operations for nominal
variables are verifying that two objects are alike (i.e., the equality operation, =) or different (i.e., the inequality operation, �). Nominal data can also be counted for the purpose of determining how many cases fall into each category and a resulting determination of proportion or percentages.4
Figure �–� Levels of measurement.
J U S T T H I N K � . � . � .
What are some other examples of nominal scales?
4. Other ways to analyze nominal data exist (Gokhale & Kullback, 1978; Kranzler & Moursund, 1999). However, let’s leave the discussion of these advanced methods for another time (and another book).
N������ O������ I������� R����
=� <> +� ×÷
Distinct Categories
Ordered Categories
Meaningful Distances
Absolute Zero
Level
Defining Feature
Operations
Categorical Quantitative
Variable
coh37025_ch03_085-128.indd 88 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
Ordinal Scales
Like nominal scales, ordinal scales assign people to categories. Unlike nominal scales, ordinal scales have categories with a clear and uncontroversial order. For example, questionnaire items often ask how often you engage in a behavior by giving you options like {Never, Sometimes, Often}. A personality item like “I am a thrill seeker.” might offer answer choices like {strongly disagree, disagree, agree, strongly agree}.
Measurements in which people are ranked are ordinal scales. In business and organizational settings, job applicants may be rank-ordered according to their desirability for a position. In clinical settings, people on a waiting list for psychotherapy may be rank-ordered according to their need for treatment. In these examples, individuals are compared with others and assigned a rank (perhaps 1 to the best applicant or the most needy wait-listed client, 2 to the next, and so forth).
Although he never used the term ordinal scale, Alfred Binet, a developer of the intelligence test that today bears his name, believed strongly that the data derived from an intelligence test are ordinal in nature. He emphasized that what he tried to do with his test was not to measure people (as one might measure a person’s height), but merely to classify (and rank) people on the basis of their performance on the tasks. He wrote:
I have not sought .� .� .� to sketch a method of measuring, in the physical sense of the word, but only a method of classification of individuals. The procedures which I have indicated will, if perfected, come to classify a person before or after such another person, or such another series of persons; but I do not believe that one may measure one of the intellectual aptitudes in the sense that one measures a length or a capacity. Thus, when a person studied can retain seven figures after a single audition, one can class him, from the point of his memory for figures, after the individual who retains eight figures under the same conditions, and before those who retain six. It is a classification, not a measurement�.�.�.�we do not measure, we classify. (Binet, cited in Varon, 1936, p. 41)
Assessment instruments applied to the individual subject may also use an ordinal form of measurement. The Rokeach Value Survey uses such an approach. In that test, a list of personal values—such as freedom, happiness, and wisdom—are put in order according to their perceived importance to the testtaker (Rokeach, 1973). If a set of 10 values is rank ordered, then the testtaker would assign a value of “1” to the most important and “10” to the least important.
In Figure 3–2, ordinal scales permit relational operators (i.e., <, �, >, �), which allow us to compare positions or ranks. For example, on the ordinal scale of high school year, Senior > Junior > Sophomore > Freshman. Ordinal scales imply nothing about how much greater one ranking is than another. Even though ordinal scales may employ numbers or “scores” to represent the rank ordering, the numbers do not indicate units of measurement. So, for example, the performance difference between the first-ranked job applicant and the second-ranked applicant may be small while the difference between the second- and third-ranked applicants may be large. On the Rokeach Value Survey, the value ranked “1” may be handily the most important in the mind of the testtaker. However, ordering the values that follow may be difficult to the point of being almost arbitrary.
Ordinal scales have no absolute zero point. In the case of a test of job performance ability, every testtaker, regardless of standing on the test, is presumed to have some ability. No testtaker is presumed to have zero ability. Zero is without meaning in such a test because the number of units that separate one testtaker’s score from another’s is simply not known. The scores are ranked, but the actual number of units separating one score from the next may be many, just a few, or practically none. Because there is no zero point on an ordinal scale, the ways in
J U S T T H I N K � . � . � .
What are some other examples of ordinal scales?
coh37025_ch03_085-128.indd 89 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
which data from such scales can be analyzed statistically are limited. One cannot average the qualifications of the first- and third-ranked job applicants, for example, and expect to come out with the qualifications of the second-ranked applicant.
Interval Scales
In addition to the features of nominal and ordinal scales, interval scales have meaningful distances between numbers. Each unit on the scale is exactly equal to any other unit on the scale. Because distance has a consistent meaning on interval scales, it is possible to add and subtract scores, which allows for calculating means and standard deviations. But like ordinal scales, interval scales contain no absolute zero point. An absolute zero indicates the absence of a quantity. Temperature is usually measured as an interval scale. When the temperature is 0°C, the zero does not mean that there is no heat. Thus, temperature in degrees Celsius does not represent magnitudes of heat. Rather, they are simply distances from the temperature at which water freezes at sea level on Earth. Even so, subtracting any two temperatures gives a consistent meaning in terms of how much energy is required to change from one temperature to another.
By contrast, temperature on the Kelvin scale has an absolute zero because 0 K indicates the complete absence of heat. Thus, temperatures in degrees Kelvin represent temperature as true magnitudes. We can say that 200 K is twice as hot as 100 K, but we cannot say that 200°C is twice as hot as 100°C. At best we can say that 200°C is twice as far from water’s freezing point as 100°C.
Clear examples of true interval scales are few in number, but there are a few that we encounter in daily life: calendar year, piano notes, and color hues.
� The distance between calendar years has a consistent meaning (e.g., the time from 500 C.E. to 600 C.E. is the same as the time from 1900 C.E. to 2000 C.E. However, the Gregorian calendar year does not measure time as a magnitude, but as a distance from the time Jesus of Nazareth is believed to have been born.
� Notes on a piano have a consistent distance as measured by half steps, but there is no absolute zero note on a piano.
� Color can be separated into three components: hue, saturation, and brightness. Whereas saturation and brightness have absolute zeros corresponding to no color (white) and no light (total darkness), hue (a smooth gradient from red to orange to yellow and so on to violet) has no true zero. For convenience, we locate hue’s zero at red which corresponds to light with the longest wavelength we can perceive. Perceptually, however, red is adjacent to and blends seamlessly with violet, which corresponds to light with the shortest wavelength we can perceive. When choosing colors that blend well or colors that can be easily distinguished, designers often consider the numerical distances between each color’s hue.
Well-designed tests of ability, personality, and psychopathology generally consist of ordinal test items. However, the test items are combined to produce a total score that behaves like an interval scale. Though such psychological tests are not true interval scales, for most purposes, they can be treated as if they were. For example, the difference in intellectual ability represented by IQs of 80 and 100, for example, is thought to be similar to that existing between IQs of 100 and 120. However, if an individual were to achieve an IQ of 0 (something that is not even possible, given
the way most intelligence tests are structured), that would not be an indication of zero (the total absence of) intelligence. Because interval scales contain no absolute zero point, a presumption inherent in their use is that no testtaker possesses none of the ability or trait (or whatever) being measured. Because interval scales are
J U S T T H I N K � . � . � .
What are some other examples of interval scales?
coh37025_ch03_085-128.indd 90 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
not magnitudes, they cannot be compared as ratios, proportions, or percentages. For example, one cannot meaningfully say that an IQ of 100 is “twice as high” as an IQ of 50. Although it may not be obvious, this statement makes no more sense than to assume that all second-place finishers in a foot race take twice as long as first-place winners because 2 is “twice as large” as 1. The number 100 is certainly twice as large as 50, but the quantity being measured is not double in size. Likewise, an IQ of 110 is not “10% higher” than an IQ of 100. This statement makes no more sense than to say that a zip code of 00110 is 10 percent larger than a zip code of 00100. Interval scales cannot be compared in this way because to do so involves division, which has no meaning for interval scales. To compare the relative size of people’s intelligence, we would need a consensus definition of what it would mean to have zero intelligence. Although at first glance you might imagine such a definition would be easy to generate, it has proved elusive whenever scholars attempt to give it a rigorous definition that we can all agree on.
Ratio Scales
In addition to all the properties of nominal, ordinal, and interval measurement, a ratio scale has a true zero point, which indicates the absence of the thing being measured. For example, 0 siblings means the absence of siblings. For countable quantities, negative numbers are meaningless (e.g., to say that one has �3 siblings is meaningless nonsense). However, for some quantities, negative numbers are possible. For example, a savings account balance is a ratio variable because having a balance of $0 means there is no money in the account. A negative balance means that the account is overdrawn and the bank is owed money. All mathematical operations can meaningfully be performed because there exist equal intervals between the numbers on the scale as well as a true or absolute zero point. The ratio scale values represent the magnitude of the quantity being measured. These magnitudes can be compared as ratios and proportions. It is possible that one person weighs twice as much as another person weighs or that a person’s income is 10% larger than it was in the previous year.
In psychology, ratio-level measurement is employed in some types of tests and test items, perhaps most notably those involving assessment of neurological functioning. One example is a test of hand grip, where the variable measured is the amount of pressure a person can exert with one hand (see Figure 3–3). Another example is a timed test of perceptual-motor ability that requires the testtaker to assemble a jigsaw-like puzzle. In such an instance, the time taken to successfully complete the puzzle is the measure that is recorded. Because there is a true zero point on this scale (or, 0 seconds), it is meaningful to say that a testtaker who completes the assembly in 30 seconds has taken half the time of a testtaker who completed it in 60�seconds. In this example, it is meaningful to speak of a true zero point on the scale—but in theory only. Why? Just think� .� .� .
No testtaker could ever obtain a score of zero on this assembly task. Stated another way, no testtaker, not even The Flash (a comic- book superhero whose power is the ability to move at superhuman speed), could assemble the puzzle in zero seconds.
Measurement Scales in Psychology
The ordinal level of measurement is most frequently used in psychology. As Kerlinger (1973, p. 439) put it: “Intelligence, aptitude, and personality test scores are, basically and strictly speaking, ordinal. These tests indicate with more or less accuracy not the amount of intelligence, aptitude, and personality traits of individuals, but rather the rank-order positions of the individuals.” Kerlinger allowed that “most psychological and educational scales approximate interval equality fairly well,” though he cautioned that if ordinal measurements are treated as if they were interval measurements, then the test user must “be constantly alert to the possibility of gross inequality of intervals” (pp. 440–441).
J U S T T H I N K � . � . � .
What are some other examples of ratio scales?
coh37025_ch03_085-128.indd 91 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
Figure �–� Ratio-level measurement in the palm of one’s hand.
Pictured above is a dynamometer, an instrument used to measure strength of hand grip. The examinee is instructed to squeeze the grips as hard as possible. The squeezing of the grips causes the gauge needle to move and reflect the number of pounds of pressure exerted. The highest point reached by the needle is the score. This is an example of ratio-level measurement. Someone who can exert 10 pounds of pressure (and earns a score of 10) exerts twice as much pressure as a person who exerts 5 pounds of pressure (and earns a score of 5). On this test it is possible to achieve a score of 0, indicating a complete lack of exerted pressure. Although it is meaningful to speak of a score of 0 on this test, we have to wonder about its significance. How might a score of 0 result? One way would be if the testtaker genuinely had paralysis of the hand. Another way would be if the testtaker was uncooperative and unwilling to comply with the demands of the task. Yet another way would be if the testtaker was attempting to malinger or “fake bad” on the test. Ratio scales may provide us “solid” numbers to work with, but some interpretation of the test data yielded may still be required before drawing any “solid” conclusions. BanksPhotos/Getty Images
Why would psychologists want to treat their assessment data as interval when those data would be better described as ordinal? Why not just say that they are ordinal? The attraction of interval measurement for users of psychological tests is the flexibility with which such data can be manipulated statistically. “What kinds of statistical manipulation?” you may ask.
In this chapter we discuss the various ways in which test data can be described or converted to make those data more manageable and understandable. Some of the techniques we will describe, such as the computation of an average, can be used if the data are assumed to be interval- or ratio-level data, but not if they are ordinal- or nominal-level data. Other techniques, such as those involving the creation of graphs or tables, may be used with ordinal- or even nominal-level data.
coh37025_ch03_085-128.indd 92 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
Describing Data
Suppose you have magically changed places with the professor teaching this course and that you have just administered an examination that consists of 100 multiple-choice items (where 1 point is awarded for each correct answer). The distribution of scores for the 25 students enrolled in your class could theoretically range from 0 (none correct) to 100 (all correct). A distribution may be defined as a set of test scores arrayed for recording or study. The 25�scores in this distribution are referred to as raw scores. As its name implies, a raw score is a straightforward, unmodified accounting of performance that is usually numerical. A raw score may reflect a simple tally, as in number of items responded to correctly on an achievement test. As we will see later in this chapter, raw scores can be converted into other types of scores. For now, let’s assume it’s the day after the examination and that you are sitting in your office looking at the raw scores listed in Table 3–1. What do you do next?
One task at hand is to communicate the test results to your class. You want to do that in a way that will help students understand how their performance on the test compared to the performance of other students. Perhaps the first step is to organize the data by transforming it from a random listing of raw scores into something that immediately conveys a bit more information. Later, as we will see, you may wish to transform the data in other ways.
Frequency Distributions
The data from the test could be organized into a distribution of the raw scores. One way the scores could be distributed is by the frequency with which they occur. In a frequency distribution, all scores are listed alongside the number of times each score occurred. The scores might be listed in tabular or graphic form. Table 3–2 lists the frequency of occurrence of each score in one column and the score itself in the other column.
Often, a frequency distribution is referred to as a simple frequency distribution to indicate that individual scores have been used and the data have not been grouped. Another kind of
J U S T T H I N K � . � . � .
In what way do most of your instructors convey test-related feedback to students? Is there a better way they could do this?
Student Score (number correct)
Judy �� Joe �� Lee-Wu �� Miriam �� Valerie �� Diane �� Henry �� Esperanza �� Paula �� Martha �� Bill �� Homer �� Robert �� Michael �� Jorge �� Mary �� “Mousey” �� Barbara �� John �� Donna �� Uriah �� Leroy �� Ronald �� Vinnie �� Bianca ��
Table �–� Data from Your Measurement Course Test
coh37025_ch03_085-128.indd 93 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
frequency distribution used to summarize data is a grouped frequency distribution. In a grouped frequency distribution, test-score intervals, also called class intervals, replace the actual test scores. The number of class intervals used and the size or width of each class interval (or, the range of test scores contained in each class interval) are for the test user to decide. But how?
In most instances, a decision about the size of a class interval in a grouped frequency distribution is made on the basis of convenience. Of course, virtually any decision will represent a trade-off of sorts. A convenient, easy-to-read summary of the data is the trade-off for the loss of detail. To what extent must the data be summarized? How important is detail? These types of questions must be considered. In the grouped frequency distribution in Table�3–3, the test scores have been grouped into 12 class intervals, where each class interval is equal to 5�points.5 The highest class interval (95–99) and the lowest class interval (40–44) are referred to, respectively, as the upper and lower limits of the distribution. Here, the need for convenience in reading the data outweighs the need for great detail, so such groupings of data seem logical.
Frequency distributions of test scores can also be illustrated graphically. A graph is a diagram or chart composed of lines, points, bars, or other symbols that describe and illustrate
Score f (frequency)
�� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� � �� �
Table �–� Frequency Distribution of Scores from Your Test
Table �–� A Grouped Frequency Distribution
Class Interval f (frequency)
��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� � ��–�� �
5. Technically, each number on such a scale would be viewed as ranging from as much as 0.5 below it to as much as 0.5 above it. For example, the “real” but hypothetical width of the class interval ranging from 95 to 99 would be the difference between 99.5 and 94.5, or 5. The true upper and lower limits of the class intervals presented in the table would be 99.5 and 39.5, respectively.
coh37025_ch03_085-128.indd 94 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
data. With a good graph, the place of a single score in relation to a distribution of test scores can be understood easily. Three kinds of graphs used to illustrate frequency distributions are the histogram, the bar graph, and the frequency polygon (Figure 3–4). A histogram is a graph
Figure �–� Graphic illustrations of data from Table �–�.
A histogram (a), a bar graph (b), and a frequency polygon (c) all may be used to graphically convey information about test performance. Of course, the labeling of the bar graph and the specific nature of the data conveyed by it depend on the variables of interest. In (b), the variable of interest is the number of students who passed the test (assuming, for the purpose of this illustration, that a raw score of 65 or higher had been arbitrarily designated in advance as a passing grade).
Returning to the question posed earlier—the one in which you play the role of instructor and�must communicate the test results to your students—which type of graph would best serve your purpose? Why?
As we continue our review of descriptive statistics, you may wish to return to your role of professor and formulate your response to challenging related questions, such as “Which measure(s) of central tendency shall I use to convey this information?” and “Which measure(s) of variability would convey the information best?”
Scores
41–45 46–50 51–55 56–60 61–65 66–70 71–75 76–80 81–85 86–90 91–95 96–100
N um
be r
of c
as es
0
1
2
3
4
5
(a)
Pass Fail
N um
be r
of c
as es
0
4
8
12
16
20
(b)
0
(c)
Scores
41-45 46-50 51-55 56-60 61-65 66-70 71-75 76-80 81-85 86-90 91-95 96-100
N um
be r
of c
as es
1
2
3
4
5
coh37025_ch03_085-128.indd 95 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
with vertical lines drawn at the true limits of each test score (or class interval), forming a series of contiguous rectangles. It is customary for the test scores (either the single scores or the midpoints of the class intervals) to be placed along the graph’s horizontal axis (also referred to as the abscissa or X-axis) and for numbers indicative of the frequency of occurrence to be placed along the graph’s vertical axis (also referred to as the ordinate or Y-axis). In a bar graph, numbers indicative of frequency also appear on the Y-axis, and reference to some categorization (e.g., yes/no/maybe, male/female) appears on the X-axis. Here the rectangular bars typically are not contiguous. Data illustrated in a frequency polygon are expressed by a continuous line connecting the points where test scores or class intervals (as indicated on the X-axis) meet frequencies (as indicated on the Y-axis).
Graphic representations of frequency distributions may assume any of a number of different shapes (Figure 3–5). Regardless of the shape of graphed data, it is a good idea for the consumer of the information contained in the graph to examine it carefully—and, if need be, critically. Consider, in this context, this chapter’s Everyday Psychometrics.
As we discuss in detail later in this chapter, one graphic representation of data of particular interest to measurement professionals is the normal or bell-shaped curve. Before getting to that, however, let’s return to the subject of distributions and how we can describe and
Figure �–� Shapes that frequency distributions can take.
E: Uniform F: Exponential growth
C: Negative skewness D: Positive skewness
A: Normal B: Bimodal
0 5 10 15 20 0 10 20 30 40
0.0 0.5 1.0 1.5 2.0 0.0 0.5 1.0 1.5 2.0
�4 �2 0 2 4 �4 �2 0 2 4 0.0
0.1
0.2
0.3
0.4
0
1
2
3
0
500
1000
1500
0.0
0.1
0.2
0.3
0.4
0
1
2
3
0.0
0.5
1.0
1.5
2.0
Sample space
F re
qu en
cy
coh37025_ch03_085-128.indd 96 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
E V E R Y D A Y P S Y C H O M E T R I C S
Consumer (of Graphed Data), Beware!
ne picture is worth a thousand words, and one purpose of representing data in graphic form is to convey information at a glance. However, although two graphs may be accurate with respect to the data they represent, their pictures—and the impression drawn from a glance at them—may be vastly di�erent. As an example, consider the following hypothetical scenario involving a hamburger restaurant chain we’ll call “The Charred House.”
The Charred House chain serves charbroiled, microscopically thin hamburgers formed in the shape of little triangular houses. In the ��-year period since its founding in ����, the company has sold, on average, ��� million burgers per year. On the chain’s tenth anniversary, The Charred House distributes a press release proudly announcing “Over a Billion Served.”
Reporters from two business publications set out to research and write a feature article on this hamburger restaurant chain. Working solely from sales �gures as compiled from annual reports to the shareholders, Reporter � focuses her story on the di�erences in yearly sales. Her article is entitled “A Billion Served—But Charred House Sales Fluctuate from Year to Year,” and its graphic illustration is reprinted here.
Quite a di�erent picture of the company emerges from Reporter �’s story, entitled “A Billion Served—And Charred House Sales Are as Steady as Ever,” and its accompanying graph. The latter story is based on a diligent analysis of comparable data for the same number of hamburger chains in the same areas of the country over the same time period. While researching the story, Reporter � learned that yearly �uctuations in sales are common to the entire industry and that the annual �uctuations observed in the Charred House �gures were— relative to other chains—insigni�cant.
Compare the graphs that accompanied each story. Although both are accurate insofar as they are based on the correct numbers, the impressions they are likely to leave are quite di�erent.
Incidentally, custom dictates that the intersection of the two axes of a graph be at � and that all the points on the Y-axis be in equal and proportional intervals from �. This custom is followed in Reporter �’s story, where the �rst point on the ordinate is ���units more than �, and each succeeding point is also �� more units away from �. However, the custom is violated in Reporter
O
�’s story, where the �rst point on the ordinate is �� units more than �, and each succeeding point increases only by �. The fact that the custom is violated in Reporter �’s story should serve as a warning to evaluate pictorial representations of data all the more critically.
(a)
Year
N um
be r
of h
am bu
rg er
s so
ld (
in m
ill io
ns )
93 94 95 96 97 98 99 00 01 02 0
10
20
30
40
50
60
70
80
90
100
110
The Charred House Sales over a 10-Year Period
Reporter 2
(b)
Year
N um
be r
of h
am bu
rg er
s so
ld (
in m
ill io
ns )
93 94 95 96 97 98 99 00 01 02 0
95
96
97
98
99
100
101
102
103
104
105
The Charred House Sales over a 10-Year Period
Reporter 1
coh37025_ch03_085-128.indd 97 12/01/21 4:06 PM
�����Part 2: The Science of Psychological Measurement
characterize them.�One way to describe a distribution of test scores is by a measure of central tendency.
Measures of Central Tendency
A measure of central tendency is a statistic that indicates the average or midmost score between the extreme scores in a distribution. The center of a distribution can be defined in different ways. Perhaps the most commonly used measure of central tendency is the arithmetic mean (or, more simply, mean), which is referred to in everyday language as the “average.” The mean takes into account the actual numerical value of every score. In special instances, such as when there are only a few scores and one or two of the scores are extreme in relation to the remaining ones, a measure of central tendency other than the mean may be desirable. Other measures of central tendency we review include the median and the mode. Note that, in the formulas to follow, the standard statistical shorthand called “summation notation” (summation meaning “the sum of”) is used. The Greek uppercase letter sigma, �, is the symbol used to signify “sum”; if X represents a test score, then the expression � X means “add all the test scores.”
The arithmetic mean� The arithmetic mean, denoted by the symbol ̄ X (and pronounced “X bar”), is equal to the sum of the observations (or test scores, in this case) divided by the number of observations. Symbolically written, the formula for the arithmetic mean is ̄ X = �( X/n), where n equals the number of observations or test scores. The arithmetic mean is typically the most appropriate measure of central tendency for interval or ratio data when the distributions are believed to be approximately normal. An arithmetic mean can also be computed from a frequency distribution. The formula for doing this is
̄ X = �( f X)
_______ n
where �( f X) means “multiply the frequency of each score by its corresponding score and then sum.” An estimate of the arithmetic mean may also be obtained from a grouped frequency distribution using the same formula, where X is equal to the midpoint of the class interval. Table 3–4 illustrates a calculation of the mean from a grouped frequency distribution. After doing the math you will find that, using the grouped data, a mean of 71.8 (which may be rounded to 72) is calculated. Using the raw scores, a mean of 72.12 (which also may be rounded to 72) is calculated. Frequently, the choice of statistic will depend on the required degree of precision in measurement.
The median� The median, defined as the middle score in a distribution, is another commonly used measure of central tendency. We determine the median of a distribution of scores by ordering the scores in a list by magnitude, in either ascending or descending order. If the total number of scores ordered is an odd number, then the median will be the score that is exactly in the middle, with one-half of the remaining scores lying above it and the other half of the remaining scores lying below it. When the total number of scores ordered is an even number, then the median can be calculated by determining the arithmetic mean of the two middle scores. For example, suppose that 10 people took a preemployment word-processing test at The
J U S T T H I N K � . � . � .
Imagine that a thousand or so engineers took an extremely di�cult pre-employment test. A�handful of the engineers earned very high scores but the vast majority did poorly, earning extremely low scores. Given this scenario, what are the pros and cons of using the mean as a measure of central tendency for this test?
coh37025_ch03_085-128.indd 98 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ��
Table �–� Calculating the Arithmetic Mean from a Grouped Frequency Distribution
Class Interval f X (midpoint of class interval) fX
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
��–�� � �� ���
� f = �� � (fX) = �,���
To estimate the arithmetic mean of this grouped frequency distribution,
̄ X = �( f X)
_____ n = 1795 _____ 25
= 71.80
To calculate the mean of this distribution using raw scores,
̄ X = �X
____ n = 1803 _____ 25
= 72.12
Rochester Wrenchworks (TRW) Corporation. They obtained the following scores, presented here in descending order:
66
65
61
59
53
52
41
36
35
32
The median of these data would be calculated by obtaining the average (or, the arithmetic mean) of the two middle scores, 53 and 52 (which would be equal to 52.5). The median is an appropriate measure of central tendency for ordinal, interval, and ratio data. The median may be a particularly useful measure of central tendency in cases where relatively few scores fall at the high end of the distribution or relatively few scores fall at the low end of the distribution.
Suppose not 10 but rather tens of thousands of people had applied for jobs at The Rochester Wrenchworks. It would be impractical to find the median by simply ordering the data and finding the midmost scores, so how would the median score be identified? For our purposes, the answer
coh37025_ch03_085-128.indd 99 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
is simply that there are advanced methods for doing so. There are also techniques for identifying the median in other sorts of distributions, such as a grouped frequency distribution and a distribution wherein various scores are identical. However, instead of delving into such new and complex territory, let’s resume our discussion of central tendency and consider another such measure.
The mode� The most frequently occurring score in a distribution of scores is the mode.6 As an example, determine the mode for the following scores obtained by another TRW job applicant, Bruce. The scores reflect the number of words Bruce word-processed in seven 1-minute trials:
43 34 45 51 42 31 51
It is TRW policy that new hires must be able to word-process at least 50 words per minute. Now, place yourself in the role of the corporate personnel officer. Would you hire Bruce? The most frequently occurring score in this distribution of scores is 51. If hiring guidelines gave you the freedom to use any measure of central tendency in your personnel decision making, then it would be your choice as to whether or not Bruce is hired. You could hire him and justify this decision on the basis of his modal score (51). You also could not hire him and justify this decision on the basis of his mean score (below the required 50 words per minute). Ultimately, whether Rochester Wrenchworks will be Bruce’s new home away from home will depend on other job-related factors, such as the nature of the job market in Rochester and the qualifications of�competing applicants. Of course, if company guidelines dictate that only the mean score be used in hiring decisions, then a career at TRW is not in Bruce’s immediate future.
Distributions that contain a tie for the designation “most frequently occurring score” can have more than one mode. Consider the following scores—arranged in no particular order— obtained by 20 students on the final exam of a new trade school called the Home Study School of Elvis Presley Impersonators:
51 49 51 50 66 52 53 38 17 66 33 44 73 13 21 91 87 92 47 3
These scores are said to have a bimodal distribution because there are two scores (51�and 66) that occur with the highest frequency (of two). Except with nominal data, the mode tends not to be a very commonly used measure of central tendency. Unlike the arithmetic mean, which has to be calculated, the value of the modal score is not calculated; one simply counts and determines which score occurs most frequently. Because the mode is arrived at in this manner, the modal score may be totally atypical—for instance, one at an extreme end of the distribution—which nonetheless occurs with the greatest frequency. In fact, it is theoretically possible for a bimodal distribution to have two modes, each of which falls at the high or the low end of the distribution—thus violating the expectation that a measure of central tendency should be� .� .� .�well, central (or indicative of a point at the middle of the distribution).
Even though the mode is not calculated in the sense that the mean is calculated, and even though the mode is not necessarily a unique point in a distribution (a distribution can have two, three, or even more modes), the mode can still be useful in conveying certain types of information. The mode is useful in analyses of a qualitative or verbal nature. For example, when assessing consumers’ recall of a commercial by means of interviews, a researcher might be interested in which word or words were mentioned most by interviewees.
The mode can convey a wealth of information in addition to the mean. As an example, suppose you wanted an estimate of the number of journal articles published by clinical psychologists in the United States in the past year. To arrive at this figure, you might total the number of journal articles accepted for publication written by each clinical psychologist in the
6. If adjacent scores occur equally often and more often than other scores, custom dictates that the mode be referred to as the average.
coh37025_ch03_085-128.indd 100 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
United States, divide by the number of psychologists, and arrive at the arithmetic mean. This calculation would yield an indication of the average number of journal articles published. Whatever that number would be, we can say with certainty that it would be more than the mode. It is well known that most clinical psychologists do not write journal articles. The mode for publications by clinical psychologists in any given year is zero. In this example, the arithmetic mean would provide us with a precise measure of the average number of articles published by clinicians. However, what might be lost in that measure of central tendency is that, proportionately, very few of all clinicians do most of the publishing. The mode (in this case, a mode of zero) would provide us with a great deal of information at a glance. It would tell us that, regardless of the mean, most clinicians do not publish.
Because the mode is not calculated in a true sense, it is a nominal statistic and cannot legitimately be used in further calculations. The median is a statistic that takes into account the order of scores and is itself ordinal in nature. The mean, an interval-level statistic, is generally the most stable and useful measure of central tendency.
Measures of Variability
Variability is an indication of how scores in a distribution are scattered or dispersed. As Figure 3–6 illustrates, two or more distributions of test scores can have the same mean even though differences in the dispersion of scores around the mean can be wide. In both distributions A and B, test scores could range from 0 to 100. In distribution A, we see that the mean score was 50 and the remaining scores were widely distributed around the mean. In distribution B, the mean was also 50 but few people scored higher than 60 or lower than 40.
Statistics that describe the amount of variation in a distribution are referred to as measures of variability. Some measures of variability include the range, the interquartile range, the semi-interquartile range, the average deviation, the standard deviation, and the variance.
The range� The range of a distribution is equal to the difference between the highest and the lowest scores. We could describe distribution B of Figure 3–5, for example, as having a range of 8 if we knew that the highest score in this distribution was 4 and the lowest score was �4 (4 � (�4) = 8). With respect to distribution D, if we knew that the lowest score was 0 and the highest score was 2, the range would be equal to 2 � 0, or 2. The range is the simplest
J U S T T H I N K � . � . � .
Devise your own example to illustrate how the mode, and not the mean, can be the most useful measure of central tendency.
Figure �–� Two distributions with di�erences in variability.
50
Test score
Distribution A Distribution B
0 100
F re
qu en
cy
50
Test score
0 40 60 100
F re
qu en
cy
X X
J U S T T H I N K � . � . � .
Devise two distributions of test scores to illustrate how the range can overstate or understate the degree of variability in the scores.
coh37025_ch03_085-128.indd 101 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
measure of variability to calculate, but its potential use is limited. Because the range is based entirely on the values of the lowest and highest scores, one extreme score (if it happens to be the lowest or the highest) can radically alter the value of the range. For example, suppose distribution B included a score of 90. The range of this distribution would now be equal to 90 � (�4), or 94. Yet, in looking at the data in the graph for distribution B, it is clear that the vast majority of scores tend to be between �4 and 4.
As a descriptive statistic of variation, the range provides a quick but gross description of the spread of scores. When its value is based on extreme scores in a distribution, the resulting description of variation may be understated or overstated. Better measures of variation include the interquartile range and the semi-interquartile range.
The interquartile and semi-interquartile ranges�A distribution of test scores (or any other data, for that matter) can be divided into four parts such that 25% of the test scores occur in each quarter. As illustrated in Figure 3–7, the dividing points between the four quarters in the distribution are the quartiles. There are three of them, respectively labeled Q1, Q2, and Q3. Note that quartile refers to a specific point whereas quarter refers to an interval. An individual score may, for example, fall at the third quartile or in the third quarter (but not “in” the third quartile or “at” the third quarter). It should come as no surprise to you that Q2 and the median are exactly the same. And just as the median is the midpoint in a distribution of scores, so are quartiles Q1 and Q3 the quarter-points in a distribution of scores. Formulas may be employed to determine the exact value of these points.
The interquartile range is a measure of variability equal to the difference between Q3 and Q1. Like the median, it is an ordinal statistic. A related measure of variability is the semi-interquartile range, which is equal to the interquartile range divided by 2. Knowledge of the relative distances of Q1 and Q3 from Q2 (the median) provides the seasoned test interpreter with immediate information as to the shape of the distribution of scores. In a perfectly symmetrical distribution, Q1 and Q3 will be exactly the same distance from the median. If
Figure �–� A quartered distribution.
Q1
First quartile
Q2
Second quartile (Median)
Q3
Third quartile
Test scores
F re
qu en
cy
coh37025_ch03_085-128.indd 102 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
these distances are unequal then there is a lack of symmetry. This lack of symmetry is referred to as skewness, and we will have more to say about that shortly.
The mean absolute deviation (MAD)� Another tool that could be used to describe the amount of variability in a distribution is the mean absolute deviation, or MAD for short. Its formula is
MAD = ��X � ̄ X �
_________ n
The bars on each side of X — ̄ X indicate that it is the absolute value of the deviation
score (ignoring the positive or negative sign and treating all deviation scores as positive). All the deviation scores are then summed and divided by the total number of scores (n) to arrive at the average deviation. As an exercise, calculate the average deviation for the following distribution of test scores:
85 100 90 95 80
Begin by calculating the arithmetic mean. Next, obtain the absolute value of each of the five deviation scores and sum them. As you sum them, note what would happen if you did not ignore the plus or minus signs: All the deviation scores would then sum to 0. Divide the sum of the deviation scores by the number of measurements (5). Did you obtain a MAD of 6? The MAD tells us that the five scores in this distribution varied, on average, 6 points from the mean.
The average deviation is rarely used. Perhaps this is so because the deletion of algebraic signs renders it a useless measure for purposes of any further operations. Why, then, discuss it here? The reason is that a clear understanding of what an average deviation measures provides a solid foundation for understanding the conceptual basis of another, more widely used measure: the standard deviation. Keeping in mind what an average deviation is, what it tells us, and how it is derived, let’s consider its more frequently used “cousin,” the standard deviation.
The standard deviation�Recall that, when we calculated the average deviation, the problem of the sum of all deviation scores around the mean equaling zero was solved by employing only the absolute value of the deviation scores. In calculating the standard deviation, the same problem must be dealt with, but we do so in a different way. Instead of using the absolute value of each deviation score, we use the square of each score. With each score squared, the sign of any negative deviation becomes positive. Because all the deviation scores are squared, we know that our calculations will not be complete until we go back and obtain the square root of whatever value we reach.
We may define the standard deviation as a measure of variability equal to the square root of the average squared deviations about the mean. More succinctly, it is equal to the square root of the variance. The variance is equal to the arithmetic mean of the squares of the differences between the scores in a distribution and their mean. The formula used to calculate the variance (s2) using deviation scores is
s2 = � (X � ̄ X )2
__________ n
Simply stated, the variance is calculated by squaring and summing all the deviation scores and then dividing by the total number of scores.
J U S T T H I N K � . � . � .
After reading about the standard deviation, explain in your own words how an understanding of the average deviation can provide a “stepping-stone” to better understanding the concept of a standard deviation.
coh37025_ch03_085-128.indd 103 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
The variance is a widely used measure in psychological research. To make meaningful interpretations, the test-score distribution should be approximately normal. We’ll have more to say about “normal” distributions later in the chapter. At this point, think of a normal distribution as a distribution with the greatest frequency of scores occurring near the arithmetic mean. Correspondingly fewer and fewer scores relative to the mean occur on both sides of it.
For some hands-on experience with—and to develop a sense of mastery of—the concepts of variance and standard deviation, why not allot the next 10 or 15 minutes to calculating the standard deviation for the test scores shown in Table 3–1? Using deviation scores, your calculations should look similar to these:
s2 = �(X � ̄ X )2
__________ n
s2 = [(78 � 72.12)2 + (67 � 72.12)2 + … + (79 � 72.12)2]
___________________________________________ 25
s2 = 4972.64 _______ 25
s2 = 198.91
The standard deviation is the square root of the variance (s2). According to our calculations, the standard deviation of the test scores is 14.10. If s = 14.10, then 1 standard deviation unit is approximately equal to 14 units of measurement or (with reference to our example and rounded to a whole number) to 14 test-score points. The test data did not provide a good normal curve approximation. Test professionals would describe these data as “positively skewed.” Skewness, as well as related terms such as negatively skewed and positively skewed, are covered in the next section. Once you are “positively familiar” with terms like positively skewed, you’ll appreciate all the more the section later in this chapter entitled “The Area Under the Normal Curve.” There you will find a wealth of information about test-score interpretation in the case when the scores are not skewed—that is, when the test scores are approximately normal in distribution.
The symbol for standard deviation has variously been represented as s, S, SD, and the lowercase Greek letter sigma (�). One custom (the one we adhere to) has it that s refers to the sample standard deviation and � refers to the population standard deviation. The number of observations in the sample is n, and the denominator n � 1 is sometimes used to calculate what is referred to as an “unbiased estimate” of the population value (though it’s actually only less biased; see Hopkins & Glass, 1978). Unless n is 10 or less, the use of n or n � 1 tends not to make a meaningful difference.
Whether the denominator is more properly n or n � 1 has been a matter of debate. Lindgren (1983) has argued for the use of n � 1, in part because this denominator tends to make correlation formulas simpler. By contrast, most texts recommend the use of n � 1 only when the data constitute a sample; when the data constitute a population, n is preferable. For Lindgren (1983), it doesn’t matter whether the data are from a sample or a population. Perhaps the most reasonable convention is to use n either when the entire population has been assessed or when no inferences to the population are intended. So, when considering the examination scores of one class of students—including all the people about whom we’re going to make inferences—it seems appropriate to use n.
Having stated our position on the n versus n � 1 controversy, our formula for the population standard deviation follows. In this formula, X
_ represents a sample mean and the Greek
letter � (mu) represents a population mean:
�
_________
�(X � �) 2
_________ n
coh37025_ch03_085-128.indd 104 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
The standard deviation is a useful measure of variation because each individual score’s distance from the mean of the distribution is factored into its computation. You will come across this measure of variation frequently in the study and practice of measurement in psychology.
Skewness
Distributions can be characterized by their skewness, or the nature and extent to which symmetry is absent. Skewness is an indication of how the measurements in a distribution are distributed. A distribution has a positive skew when relatively few of the scores fall at the high end of the distribution. Positively skewed examination results may indicate that the test was too difficult. More items that were easier would have been desirable in order to better discriminate at the lower end of the distribution of test scores. A distribution has a negative skew when relatively few of the scores fall at the low end of the distribution. Negatively skewed examination results may indicate that the test was too easy. In this case, more items of a higher level of difficulty would make it possible to better discriminate between scores at the upper end of the distribution. (Refer to Figure 3–5 for graphic examples of skewed distributions.)
The term skewed carries with it negative implications for many students. We suspect that skewed is associated with abnormal, perhaps because the skewed distribution deviates from the symmetrical or so-called normal distribution. However, the presence or absence of symmetry in a distribution (skewness) is simply one characteristic by which a distribution can be described. Consider in this context a hypothetical Marine Corps Ability and Endurance Screening Test administered to all civilians seeking to enlist in the U.S. Marines. Now look again at the graphs in Figure 3–5. Which graph do you think would best describe the resulting distribution of test scores? (No peeking at the next paragraph before you respond.)
No one can say with certainty, but if we had to guess, then we would say that the Marine Corps Ability and Endurance Screening Test data would look like graph C, the positively skewed distribution in Figure 3–5. We say this assuming that a level of difficulty would have been built into the test to ensure that relatively few assessees would score at the high end of the distribution. Most of the applicants would probably score at the low end of the distribution. All of this is quite consistent with recruiters advertised objective of selecting “The Few. The Proud, The Marines.” An older recruiting slogan was “If Everybody Could Get In The Marines, It Wouldn’t Be The Marines.” Now, a question regarding this positively skewed distribution: Is the skewness a good thing? A bad thing? An abnormal thing? In truth, it is probably none of these things—it just is.
Various formulas exist for measuring skewness. One way of gauging the skewness of a distribution is through examination of the relative distances of quartiles from the median. In a positively skewed distribution, Q3 � Q2 will be greater than the distance of Q2 � Q1. In a negatively skewed distribution, Q3 � Q 2 will be less than the distance of Q2 � Q 1. In a�distribution that is symmetrical, the distances from Q1 and Q3 to the median are the same.
Kurtosis
The term testing professionals use to refer to the steepness of a distribution in its center is kurtosis. To the root kurtic is added to one of the prefixes platy-, lepto-, or meso- to describe the peakedness/flatness of three general types of curves (Figure 3–8). Distributions are generally described as platykurtic (relatively flat), leptokurtic (relatively peaked), or—somewhere in the middle— mesokurtic. Distributions that have high kurtosis are characterized by a high peak and “fatter” tails compared to a normal distribution. In contrast, lower kurtosis values indicate a distribution with a
J U S T T H I N K � . � . � .
Like skewness, reference to the kurtosis of a distribution can provide a kind of “shorthand” description of a distribution of test scores. Imagine and describe the kind of test that might yield a distribution of scores that form a platykurtic curve.
coh37025_ch03_085-128.indd 105 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
rounded peak and thinner tails. Many methods exist for measuring kurtosis. According to the original definition, the normal bell-shaped curve (see graph A from Figure 3–5) would have a kurtosis value of 3. In other methods of computing kurtosis, a normal distribution would have kurtosis of 0, with positive values indicating higher kurtosis and negative values indicating lower kurtosis. It is important to keep the different methods of calculating kurtosis in mind when examining the values reported by researchers or computer programs. So, given that this can quickly become an advanced-level topic and that this book is of a more introductory nature, let’s move on. It’s time to focus on a type of distribution that happens to be the standard against which all other distributions (including all of the kurtic ones) are compared: the normal distribution.
The Normal Curve
Before delving into the statistical, a little bit of the historical is in order. Development of the concept of a normal curve began in the middle of the eighteenth century with the work of Abraham DeMoivre and, later, the Marquis de Laplace. At the beginning of the nineteenth century, Karl Friedrich Gauss made some substantial contributions. Through the early nineteenth century, scientists referred to it as the “Laplace-Gaussian curve.” Karl Pearson is credited with being the first to refer to the curve as the normal curve, perhaps in an effort to be diplomatic to all of the people who helped develop it. Somehow the term normal curve stuck—but don’t be surprised if you’re sitting at some scientific meeting one day and you hear this distribution or curve referred to as Gaussian.
Theoretically, the normal curve is a bell-shaped, smooth, mathematically defined curve that is highest at its center. From the center it tapers on both sides approaching the X-axis asymptotically (meaning that it approaches, but never touches, the axis). In theory, the distribution of the normal curve ranges from negative infinity to positive infinity. The curve is perfectly symmetrical, with no skewness. If you folded it in half at the mean, one side would lie exactly on top of the other. Because it is symmetrical, the mean, the median, and the mode all have the same exact value.
Figure �–� The kurtosis of curves.
–3 –2
Mesokurtic
Leptokurtic
Platykurtic
–1 0 +1 +2 +3
z scores
coh37025_ch03_085-128.indd 106 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
Why is the normal curve important in understanding the characteristics of psychological tests? Our Close-Up provides some answers.
The Area Under the Normal Curve
The normal curve can be conveniently divided into areas defined in units of standard deviation. A hypothetical distribution of National Spelling Test scores with a mean of � = 50 and a standard deviation of � = 15 is illustrated in Figure 3–9. In this example, a score equal to 1 standard deviation above the mean would be equal to 65 (� + 1� = 50 + 15 = 65).
Before reading on, take a minute or two to calculate what a score exactly at 3 standard deviations below the mean would be equal to. How about a score exactly at 3 standard deviations above the mean? Were your answers 5 and 95, respectively? The graph tells us that 99.74% of all scores in these normally distributed spelling-test data lie between ±3 standard deviations. Stated another way, 99.74% of all spelling test scores lie between 5 and 95. This graph also illustrates the following characteristics of all normal distributions.
� 50% of the scores occur above the mean and 50% of the scores occur below the mean. � Approximately 34% of all scores occur between the mean and 1 standard deviation
above the mean. � Approximately 34% of all scores occur between the mean and 1 standard deviation
below the mean. � Approximately 68% of all scores occur between the mean and ±1 standard deviation. � Approximately 95% of all scores occur between the mean and ±2 standard deviations.
Figure �–� The area under the normal curve.
±1� �68%
±2� �95%
±3� �99.7%
2.3% 13.6% 34.1% 34.1% 13.6% 2.3%
5 � 3�
20 � 2�
35 � 1�
50 �
65 +1�
80 +2�
95 +3�
Raw scores z-scores
coh37025_ch03_085-128.indd 107 12/01/21 4:06 PM
������ Part 2: The Science of Psychological Measurement
C L O S E � U P
The Normal Curve and Psychological Tests
cores on many psychological tests are often approximately normally distributed, particularly when the tests are administered to large numbers of subjects. Few, if any, psychological tests yield precisely normal distributions of test scores (Micceri, ����). As a general rule (with ample exceptions), the larger the sample size and the wider the range of abilities measured by a particular test, the more the graph of the test scores will approximate the normal curve. A classic illustration of this was provided by E. L. Thorndike and his colleagues (����). They compiled intelligence test scores from several large samples of students. As you can see in Figure �, the distribution of scores closely approximated the normal curve.
Following is a sample of more varied examples of the wide range of characteristics that psychologists have found to be approximately normal in distribution.
� The strength of handedness in right-handed individuals, as measured by the Waterloo Handedness Questionnaire (Tan,�����).
S � Scores on the Women’s Health Questionnaire, a scale measuring a variety of health problems in women across a wide age range (Hunter, ����).
� Responses of both college students and working adults to a measure of intrinsic and extrinsic work motivation (Amabile et�al., ����).
� The intelligence-scale scores of girls and women with eating disorders, as measured by the Wechsler Adult Intelligence Scale–Revised and the Wechsler Intelligence Scale for Children–Revised (Ranseen & Humphries, ����).
� The intellectual functioning of children and adolescents with cystic �brosis (Thompson et al., ����).
� Decline in cognitive abilities over a one-year period in people with Alzheimer’s disease (Burns et al., ����).
� The rate of motor-skill development in developmentally delayed preschoolers, as measured by the Vineland Adaptive Behavior Scale (Davies & Gavin, ����).
Figure � Graphic representation of Thorndike et al. data.
The solid line outlines the distribution of intelligence test scores of sixth-grade students (N = 15,138). The dotted line is the theoretical normal curve (Thorndike et al., 1927).
–3 –2 –1 0
z scores
+1 +2 +3
coh37025_ch03_085-128.indd 108 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
normal distribution of scores. Why? One bene�t of a normal distribution of scores is that it simpli�es the interpretation of individual scores on the test. In a normal distribution, the mean, the median, and the mode take on the same value. For example, if we know that the average score for intellectual ability of children with cystic �brosis is a particular value and that the scores are normally distributed, then we know quite a bit more. We know that the average is the most common score and the score below and above which half of all the scores fall. Knowing the mean and the standard deviation of a scale and that it is approximately normally distributed tells us that (�) approximately two-thirds of all testtakers’ scores are within a standard deviation of the mean and (�) approximately ��% of the scores fall within � standard deviations of the mean.
The characteristics of the normal curve provide a ready model for score interpretation that can be applied to a wide range of test results.
� Scores on the Swedish translation of the Positive and Negative Syndrome Scale, which assesses the presence of positive and negative symptoms in people with schizophrenia (von Knorring & Lindstrom, ����).
� Scores of psychiatrists on the Scale for Treatment Integration of the Dually Diagnosed (people with both a drug problem and another mental disorder); the scale examines opinions about drug treatment for this group of patients (Adelman et al., ����).
� Responses to the Tridimensional Personality Questionnaire, a measure of three distinct personality features (Cloninger et al., ����).
� Scores on a self-esteem measure among undergraduates (Addeo et al., ����).
In each case, the researchers made a special point of stating that the scale under investigation yielded something close to a
A normal curve has two tails. The area on the normal curve between 2 and 3 standard deviations above the mean is referred to as a tail. The area between �2 and �3 standard deviations below the mean is also referred to as a tail. Let’s digress here momentarily for a “real-life” tale of the tails to consider along with our rather abstract discussion of statistical concepts.
As observed in a thought-provoking article entitled “Two Tails of the Normal Curve,” an intelligence test score that falls within the limits of either tail can have momentous consequences in terms of the tale of one’s life:
Individuals who are mentally retarded or gifted share the burden of deviance from the norm, in both a developmental and a statistical sense. In terms of mental ability as operationalized by tests of intelligence, performance that is approximately two standard deviations from the mean (or, IQ of 70–75 or lower or IQ of 125–130 or higher) is one key element in identification. Success at life’s tasks, or its absence, also plays a defining role, but the primary classifying feature of both gifted and retarded groups is intellectual deviance. These individuals are out of sync with more average people, simply by their difference from what is expected for their age and circumstance. This asynchrony results in highly significant consequences for them and for those who share their lives. None of the familiar norms apply, and substantial adjustments are needed in parental expectations, educational settings, and social and leisure activities. (Robinson et al., 2000, p. 1413)
Robinson et al. (2000) convincingly demonstrated that knowledge of the areas under the normal curve can be quite useful to the interpreter of test data. This knowledge can tell us not only something about where the score falls among a distribution of scores but also something about a person and perhaps even something about the people who share that person’s life. This knowledge might also convey something about how impressive, average, or lackluster the individual is with respect to a particular discipline or ability. For example, consider a high-school student whose score on a national, well-respected spelling test is close to 3 standard deviations above the mean. It’s a good bet that this student would know how to spell words like asymptotic and leptokurtic.
coh37025_ch03_085-128.indd 109 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
Just as knowledge of the areas under the normal curve can instantly convey useful information about a test score in relation to other test scores, so can knowledge of standard scores.
Standard Scores
Simply stated, a standard score is a raw score that has been converted from one scale to another scale, where the latter scale has some arbitrarily set mean and standard deviation. Why convert raw scores to standard scores?
Raw scores may be converted to standard scores because standard scores are more easily interpretable than raw scores. With a standard score, the position of a testtaker’s performance relative to other testtakers is readily apparent.
Different systems for standard scores exist, each unique in terms of its respective mean and standard deviations. We will briefly describe z scores, T scores, stanines, and some other standard scores. First for consideration is the type of standard score scale that may be thought of as the zero plus or minus one scale. That is, it has a mean set at 0 and a standard deviation set at 1. Raw scores converted into standard scores on this scale are more popularly referred to as z scores.
z Scores
A z score results from the conversion of a raw score into a number indicating how many standard deviation units the raw score is below or above the mean of the distribution. Let’s use an example from the normally distributed “National Spelling Test” data in Figure 3–9 to demonstrate how a raw score is converted to a z score. We’ll convert a raw score of 65 to a z score by using the formula
z = X � ̄ X ______ s = 65 � 50 _______ 15
= 15 ___ 15
= 1
In essence, a z score is equal to the difference between a particular raw score and the mean divided by the standard deviation. In the preceding example, a raw score of 65 was found to be equal to a z score of +1. Knowing that someone obtained a z score of 1 on a spelling test provides context and meaning for the score. Drawing on our knowledge of areas under the normal curve, for example, we would know that only about 16% of the other testtakers obtained higher scores. By contrast, knowing simply that someone obtained a raw score of 65 on a spelling test conveys virtually no usable information because information about the context of this score is lacking.
In addition to providing a convenient context for comparing scores on the same test, standard scores provide a convenient context for comparing scores on different tests. As an example, consider that Crystal’s raw score on the hypothetical Main Street Reading Test was 24 and that her raw score on the (equally hypothetical) Main Street Arithmetic Test was 42. Without knowing anything other than these raw scores, one might conclude that Crystal did better on the arithmetic test than on the reading test. Yet more informative than the two raw scores would be the two z scores.
Converting Crystal’s raw scores to z scores based on the performance of other students in her class, suppose we find that her z score on the reading test was 1.32 and that her z score on the arithmetic test was �0.75. Thus, although her raw score in arithmetic was higher than in reading, the z scores paint a different picture. The z scores tell us that, relative to the other students in her class (and assuming that the distribution of scores is relatively
coh37025_ch03_085-128.indd 110 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
normal), Crystal performed above average on the reading test and below average on the arithmetic test. An interpretation of exactly how much better she performed could be obtained by reference to tables detailing distances under the normal curve as well as the resulting percentage of cases that could be expected to fall above or below a particular standard deviation point (or z score).
T Scores
If the scale used in the computation of z scores is called a zero plus or minus one scale, then the scale used in the computation of T scores can be called a fifty plus or minus ten scale; that is, a scale with a mean set at 50 and a standard deviation set at 10. Devised by W. A. McCall (1922, 1939) and named a T score in honor of his professor E. L. Thorndike, this standard score system is composed of a scale that ranges from 5 standard deviations below the mean to 5 standard deviations above the mean. Thus, for example, a raw score that fell exactly at 5 standard deviations below the mean would be equal to a T score of 0, a raw score that fell at the mean would be equal to a T of 50, and a raw score 5 standard deviations above the mean would be equal to a T of 100. One advantage in using T scores is that none of the scores is negative. By contrast, in a z score distribution, scores can be positive and negative; this characteristic can make further computation cumbersome in some instances.
Other Standard Scores
Numerous other standard scoring systems exist. Researchers during World War II developed a standard score with a mean of 5 and a standard deviation of approximately 2. Divided into nine units, the scale was christened a stanine, a term that was a contraction of the words standard and nine.
Stanine scoring may be familiar to many students from achievement tests administered in elementary and secondary school, where test scores are often represented as stanines. Stanines are different from other standard scores in that they take on whole values from 1 to 9, which represent a range of performance that is half of a standard deviation in width (Figure 3–10). The 5th stanine indicates performance in the average range, from 1/4
Figure �–�� Stanines and the normal curve.
4% 7% 12% 17% 20% 17% 12% 7% 4% Low
Below average
Average
Above average
High
Stanine 1 2 3 4 5 6 7 8 9
coh37025_ch03_085-128.indd 111 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
standard deviation below the mean to 1/4 standard deviation above the mean, and captures the middle 20% of the scores in a normal distribution. The 4th and 6th stanines are also 1/2 standard deviation wide and capture the 17% of cases below and above (respectively) the 5th stanine.
Have you ever heard the term IQ used as a synonym for one’s score on an intelligence test? Of course you have. What you may not know is that what is referred to variously as IQ, deviation IQ, or deviation intelligence quotient is yet another kind of standard score. For most IQ tests, the distribution of raw scores is converted to IQ scores, whose distribution typically has a mean set at 100 and a standard deviation set at 15. Let’s emphasize typically because there is some variation in standard scoring systems, depending on the test used. The typical mean and standard deviation for IQ tests results in approximately 95% of deviation IQs ranging from 70 to 130, which is 2 standard deviations below and above the mean. In the context of a normal distribution, the relationship of deviation IQ scores to the other standard scores we have discussed so far (z, T, and A scores) is illustrated in Figure 3–11.
Figure �–�� Some standard score equivalents.
Note that the values presented here for the IQ scores assume that the intelligence test scores have a mean of 100 and a standard deviation of 15. This is true for many, but not all, intelligence tests. If a particular test of intelligence yielded scores with a mean other than 100 and/or a standard deviation other than 15, then the values shown for IQ scores would have to be adjusted accordingly.
Very low Low Below
average Average Above
average High Very high
�4 �3 �2 �1 0 1 2 3 4
40 55 70 85 100 115 130 145 160
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
10 20 30 40 50 60 70 80 90
0.003 0.1 2 16 50 84 98 99.9 99.997
z-Scores (µ = 0, � = 1)
IQ Scores (µ = 100, � = 15)
T Scores (µ = 50, � = 10)
Scaled Scores (µ = 10, � = 3)
Percentile Rank
coh37025_ch03_085-128.indd 112 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
Standard scores converted from raw scores may involve either linear or nonlinear transformations. A standard score obtained by a linear transformation is one that retains a direct numerical relationship to the original raw score. The magnitude of differences between such standard scores exactly parallels the differences between corresponding raw scores. Sometimes scores may undergo more than one transformation. For example, the creators of the SAT did a second linear transformation on their data to convert z scores into a new scale that has a mean of 500 and a standard deviation of 100.
A nonlinear transformation may be required when the data under consideration are not normally distributed yet comparisons with normal distributions need to be made. In a nonlinear transformation, the resulting standard score does not necessarily have a direct numerical relationship to the original, raw score. As the result of a nonlinear transformation, the original distribution is said to have been normalized.
Normalized standard scores� Many test developers hope that the test they are working on will yield a normal distribution of scores. Yet even after very large samples have been tested with the instrument under development, skewed distributions result. What should be done?
One alternative available to the test developer is to normalize the distribution. Conceptually, normalizing a distribution involves “stretching” the skewed curve into the shape of a normal curve and creating a corresponding scale of standard scores, a scale that is technically referred to as a normalized standard score scale.
Normalization of a skewed distribution of scores may also be desirable for purposes of comparability. One of the primary advantages of a standard score on one test is that it can readily be compared with a standard score on another test. However, such comparisons are appropriate only when the distributions from which they derived are the same. In most instances, they are the same because the two distributions are approximately normal. But if, for example, distribution A were normal and distribution B were highly skewed, then z scores in these respective distributions would represent different amounts of area subsumed under the curve. A z score of �1 with respect to normally distributed data tells us, among other things, that about 84% of the scores in this distribution were higher than this score. A z score of �1 with respect to data that were very positively skewed might mean, for example, that only 62% of the scores were higher.
For test developers intent on creating tests that yield normally distributed measurements, it is generally preferable to fine-tune the test according to difficulty or other relevant variables so that the resulting distribution will approximate the normal curve. This approach usually is a better bet than attempting to normalize skewed distributions. This is so because there are technical cautions to be observed before attempting normalization. For example, transformations should be made only when there is good reason to believe that the test sample was large enough and representative enough and that the failure to obtain normally distributed scores was due to the measuring instrument.
Correlation and Inference
Central to psychological testing and assessment are inferences (deduced conclusions) about how some things (such as traits, abilities, or interests) are related to other things (such as
J U S T T H I N K � . � . � .
Apply what you have learned about frequency distributions, graphing frequency distributions, measures of central tendency, measures of variability, and the normal curve and standard scores to the question of the data listed in Table �–�. How would you communicate the data from Table �–� to the class? Which type of frequency distribution might you use? Which type of graph? Which measure of central tendency? Which measure of variability? Might reference to a normal curve or to standard scores be helpful? Why or why not?
coh37025_ch03_085-128.indd 113 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
behavior). A coefficient of correlation (or correlation coefficient) is a number that provides us with an index of the strength of the relationship between two things. An understanding of the concept of correlation and an ability to compute a coefficient of correlation is therefore central to the study of tests and measurement.
The Concept of Correlation
Simply stated, correlation is an expression of the degree and direction of correspondence between two things. A coefficient of correlation (r) expresses a linear relationship between two (and only two) variables, usually continuous in nature. It reflects the degree of concomitant variation between variable X and variable Y. The coefficient of correlation is the numerical index that expresses this relationship: It tells us the extent to which X and Y are “co-related.”
The meaning of a correlation coefficient is interpreted by its sign and magnitude. If a correlation coefficient were a person asked “What’s your sign?,” it would not answer anything like “Leo” or “Pisces.” It would answer “plus” (for a positive correlation), “minus” (for a negative correlation), or “none” (in the rare instance that the correlation coefficient was exactly equal to zero). If asked to supply information about its magnitude, it would respond with a number between �1 and +1. And here is a rather intriguing fact about the magnitude of a correlation coefficient: It is judged by its absolute value. This fact means that to the extent that we are impressed by correlation coefficients, a correlation of �.99 is every bit as impressive as a correlation of +.99. To understand why, you need to know a bit more about correlation.
“Ahh�.�.�.�a perfect correlation! Let me count the ways.” Well, actually there are only two ways. The two ways to describe a perfect correlation between two variables are as either +1 or �1. If a correlation coefficient has a value of +1 or �1, then the relationship between the two variables being correlated is perfect—without error in the statistical sense. And just as
perfection in almost anything is difficult to find, so too are perfect correlations. It is challenging to try to think of any two variables in psychological work that are perfectly correlated. Perhaps that is why, if you look in the margin, you are asked to “just think” about it.
If two variables simultaneously increase or simultaneously decrease, then those two variables are said to be positively (or
directly) correlated. The height and weight of normal, healthy children ranging in age from birth to 10 years tend to be positively or directly correlated. As children get older, their height and their weight generally increase simultaneously. A positive correlation also exists when two variables simultaneously decrease. For example, the less a student prepares for an examination, the lower that student’s score on the examination. A negative (or inverse) correlation occurs when one variable increases while the other variable decreases. For example, there tends to be an inverse relationship between the number of miles on your car’s odometer (mileage indicator) and the number of dollars a car dealer is willing to give you on a trade-in allowance; all other things being equal, as the mileage increases, the number of dollars offered on trade-in decreases. And by the way, we all know students who use cell phones during class to text, tweet, check e-mail, or otherwise be engaged with their phone at a questionably appropriate time and place. What would you estimate the correlation to be between such daily, in-class cell phone use and test grades? See Figure 3–12 for one such estimate (and kindly refrain from sharing the findings on Instagram during class).
If a correlation is zero, then absolutely no relationship exists between the two variables. And some might consider “perfectly no correlation” to be a third variety of perfect correlation; that is, a perfect noncorrelation. After all, just as it is nearly impossible in psychological
J U S T T H I N K � . � . � .
Can you name two variables that are perfectly correlated? How about two psychological variables that are perfectly correlated?
coh37025_ch03_085-128.indd 114 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
work to identify two variables that have a perfect correlation, so it is nearly impossible to identify two variables that have a zero correlation. Most of the time, two variables will be fractionally correlated. The fractional correlation may be extremely small but seldom “perfectly” zero.
As we stated in our introduction to this topic, correlation is often confused with causation. It must be emphasized that a correlation coefficient is merely an index of the relationship between two variables, not an index of the causal relationship between two variables. If you were told, for example, that from birth to age 9 there is a high positive correlation between hat size and spelling ability, would it be appropriate to conclude that hat size causes spelling ability? Of course not. The period from birth to age 9 is a time of maturation in all areas, including physical size and cognitive abilities such as spelling. Intellectual development parallels physical development during these years, and a relationship clearly exists between physical and mental growth. Still, this doesn’t mean that the relationship between hat size and spelling ability is causal.
Figure �–�� Cell phone use in class and class grade.
Current students may be the “wired” generation, but some college students are clearly more wired than others. They seem to be on their cell phones constantly, even during class. Their gaze may be fixed on Mech Commander when it should more appropriately be on Class Instructor. Over the course of two semesters, Chris Bjornsen and Kellie Archer (2015) studied 218 college students, each of whom completed a questionnaire on their cell phone usage right after class. Correlating the questionnaire data with grades, the researchers reported that cell phone usage during class was significantly, negatively correlated with grades. Gorodenkoff/Shutterstock
J U S T T H I N K � . � . � .
Could a correlation of zero between two variables also be considered a “perfect” correlation? Can you name two variables that have a correlation that is exactly zero?
J U S T T H I N K � . � . � .
Bjornsen and Archer (����) discussed the implications of their cell phone study in terms of the e�ect of cell phone usage on student learning, student achievement, and post- college success. What would you anticipate those implications to be?
coh37025_ch03_085-128.indd 115 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
Although correlation does not imply causation, there is an implication of prediction. Stated another way, if we know that there is a high correlation between X and Y, then we should be able to predict—with various degrees of accuracy, depending on other factors—the value of one of these variables if we know the value of the other.
The Pearson r
Many techniques have been devised to measure correlation. The most widely used of all is the Pearson r, also known as the Pearson correlation coefficient and the Pearson product-moment coefficient of correlation. Devised by Karl Pearson (Figure 3–13), r can be the statistical tool of choice when the relationship between the variables is linear and when the two variables being correlated are continuous (or, they can theoretically take any value). Other correlational techniques can be employed with data that are discontinuous and where the relationship is nonlinear. The formula for the Pearson r takes into account the relative position of each test score or measurement with respect to the mean of the distribution.
A number of formulas can be used to calculate a Pearson r. One formula requires that we convert each raw score to a standard score and then multiply each pair of standard scores. A mean for the sum of the products is calculated, and that mean is the value of the Pearson r. Even from this simple verbal conceptualization of the Pearson r, it can be seen that the sign of the resulting r would be a function of the sign and the magnitude of the standard scores used. If, for example, negative standard score values for measurements of X always corresponded with negative standard score values for Y scores, the resulting r would be positive (because the product of two negative values is positive). Similarly, if positive standard score values on X always corresponded with positive standard score values on Y, the resulting correlation would also be positive. However, if positive standard score values for X corresponded with negative standard score values for Y and vice versa, then an inverse relationship would exist and so a negative correlation would result. A zero or near-zero correlation could result when some products are positive and some are negative.
Figure �–�� Karl Pearson (����–����).
Karl Pearson’s name has become synonymous with correlation. History records, however, that it was actually Sir Francis Galton who should be credited with developing the concept of correlation (Magnello & Spies, 1984). Galton experimented with many formulas to measure correlation, including one he labeled r. Pearson, a contemporary of Galton’s, modified Galton’s r, and the rest, as they say, is history. The Pearson r eventually became the most widely used measure of correlation. The History Collection/Alamy Stock Photo
coh37025_ch03_085-128.indd 116 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
The formula used to calculate a Pearson r from raw scores is
r = �(X � ̄ X )(Y � ̄ Y )
______________________ �
____________________ [�(X � ̄ X )2][�(Y � ̄ Y )2]
This formula has been simplified for shortcut purposes. One such shortcut is a deviation formula employing “little x,” or x in place of X — X
_ , and “little y,” or y in place of Y — Y
_ :
r = �xy ___________
� _________
(�x 2)(�y 2)
Another formula for calculating a Pearson r is
r = N�XY � (�X)(�Y)
___________________________ �
___________ N�X 2� (�X) 2 �
____________ N �Y 2 � (�Y) 2
Although this formula looks more complicated than the previous deviation formula, it is easier to use. Here N represents the number of paired scores; � XY is the sum of the product of the paired X and Y scores; � X is the sum of the X scores; � Y is the sum of the Y scores; � X2 is the sum of the squared X scores; and � Y2 is the sum of the squared Y scores. Similar results are obtained with the use of each formula.
The next logical question concerns what to do with the number obtained for the value of� r. The answer is that you ask even more questions, such as “Is this number statistically significant, given the size and nature of the sample?” or “Could this result have occurred by chance?” At this point, you will need to consult tables of significance for Pearson r—tables that are probably in the back of your old statistics textbook. In those tables you will find, for�example, that a Pearson r of .899 with an N = 10 is significant at the .01 level (using a two-tailed test). You will recall from your statistics course that significance at the .01 level tells you, with reference to these data, that a correlation such as this could have been expected to occur merely by chance only one time or less in a hundred if X and Y are not correlated in the population. You will also recall that significance at either the .01 level or the (somewhat less rigorous) .05 level provides a basis for concluding that a correlation does indeed exist. Significance at the .05 level means that the result could have been expected to occur by chance alone five times or less in a hundred.
The value obtained for the coefficient of correlation can be further interpreted by deriving from it what is called a coefficient of determination, or r2. The coefficient of determination is an indication of how much variance is shared by the X- and the Y-variables. The calculation of r2 is quite straightforward. Simply square the correlation coefficient and multiply by 100; the result is equal to the percentage of the variance accounted for. If, for example, you calculated r to be .9, then r2 would be equal to .81. The number .81 tells us that 81% of the�variance is accounted for by the X- and Y-variables. The remaining variance, equal to 100(1 � r2), or 19%, could presumably be accounted for by chance, error, or otherwise unmeasured or unexplainable factors.7
Before moving on to consider another index of correlation, let’s address a logical question sometimes raised by students when they hear the Pearson r referred to as the product-moment coefficient of correlation. Why is it called that? The answer is a little complicated, but here goes.
7. On a technical note, Ozer (1985) cautioned that the actual estimation of a coefficient of determination must be made with scrupulous regard to the assumptions operative in the particular case. Evaluating a coefficient of determination solely in terms of the variance accounted for may lead to interpretations that underestimate the magnitude of a relation.
coh37025_ch03_085-128.indd 117 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
In the language of psychometrics, a moment describes a deviation about a mean of a distribution. Individual deviations about the mean of a distribution are referred to as deviates. Deviates are referred to as the first moments of the distribution. The second moments of the distribution are the moments squared. The third moments of the distribution are the moments cubed, and so forth. The computation of the Pearson r in one of its many formulas entails multiplying corresponding standard scores on two measures. One way of conceptualizing standard scores is as the first moments of a distribution because standard scores are deviates about a mean of zero. A formula that entails the multiplication of two corresponding standard scores can therefore be conceptualized as one that requires the computation of the product of corresponding moments. And there you have the reason r is called product-moment correlation. It is probably all more a matter of psychometric trivia than anything else, but we think it is cool to know. Further, you can now understand the rather “high-end” humor contained in the cartoon (below).
The Spearman Rho
The Pearson r enjoys such widespread use and acceptance as an index of correlation that if for some reason it is not used to compute a correlation coefficient, mention is made of the statistic that was used. There are many alternative ways to derive a coefficient of correlation. One commonly used alternative statistic is variously called a rank-order correlation coefficient, a rank-difference correlation coefficient, or simply Spearman’s rho. Developed by Charles Spearman, a British psychologist (Figure 3–14), this coefficient of correlation is frequently used when the sample size is small (fewer than 30 pairs of measurements) and especially when both sets of measurements are in ordinal (or rank-order) form. Special tables are used to determine whether an obtained rho coefficient is or is not significant.
Copyright 2016 Ronald Jay Cohen. All rights reserved.
coh37025_ch03_085-128.indd 118 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
Graphic Representations of Correlation
One type of graphic representation of correlation is referred to by many names, including a bivariate distribution, a scatter diagram, a scattergram, or—our favorite—a scatterplot. A scatterplot is a simple graphing of the coordinate points for values of the X-variable (placed along the graph’s horizontal axis) and the Y-variable (placed along the graph’s vertical axis). Scatterplots are useful because they provide a quick indication of the direction and magnitude of the relationship, if any, between the two variables. Figures 3–15 and 3–16 offer a quick course in eyeballing the nature and degree of correlation by means of scatterplots. To distinguish positive from negative correlations, note the direction of the curve. And to estimate the strength of magnitude of the correlation, note the degree to which the points form a straight line.
Scatterplots are useful in revealing the presence of curvilinearity in a relationship. As you may have guessed, curvilinearity in this context refers to an “eyeball gauge” of how curved a graph is. Remember that a Pearson r should be used only if the relationship between the variables is linear. If the graph does not appear to take the form of a straight line, the chances are good that the relationship is not linear (Figure 3–17, left panel). When the relationship is nonlinear, other statistical tools and techniques may be employed.8
Figure �–�� Charles Spearman (����–����).
Charles Spearman is best known as the developer of the Spearman rho statistic and the Spearman-Brown prophecy formula, which is used to “prophesize” the accuracy of tests of different sizes. Spearman is also credited with being the father of a statistical method called factor analysis, discussed later in this text. Keystone Press/Alamy Stock Photo
8. The specific statistic to be employed will depend at least in part on the suspected reason for the nonlinearity. For example, if it is believed that the nonlinearity is due to one distribution being highly skewed because of a poor measuring instrument, then the skewed distribution may be statistically normalized and the result may be a correction of the curvilinearity. If—even after graphing the data—a question remains concerning the linearity of the correlation, a statistic called “eta squared” (�2) can be used to calculate the exact degree of curvilinearity.
coh37025_ch03_085-128.indd 119 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
Figure �–�� Scatterplots and correlations for positive values of r.
Correlation coe�cient = 0 Correlation coe�cient = .40
Correlation coe�cient = .60 Correlation coe�cient = .80
0
(a) (b)
(c) (d)
(e) (f)
0
1
2
3
4
5
6
1 2 3 4 5 6 0 0
1
2
3
4
5
6
1 2 3 4 5 6
0 0
1
2
3
4
5
6
1 2 3 4 5 6 0 0
1
2
3
4
5
6
1 2 3 4 5 6
0 0
1
2
3
4
5
6
1 2 3 4 5 6
Correlation coe�cient = .90 Correlation coe�cient = .95
0 0
1
2
3
4
5
6
1 2 3 4 5 6
coh37025_ch03_085-128.indd 120 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
Figure �–�� Scatterplots and correlations for negative values of r.
Correlation coe�cient = –.30 Correlation coe�cient = –.50
Correlation coe�cient = – .70 Correlation coe�cient = – .90
0
(a) (b)
(c) (d)
(e) (f)
0
1
2
3
4
5
6
1 2 3 4 5 6 0 0
1
2
3
4
5
6
1 2 3 4 5 6
0 0
1
2
3
4
5
6
1 2 3 4 5 6 0 0
1
2
3
4
5
6
1 2 3 4 5 6
0 0
1
2
3
4
5
6
1 2 3 4 5 6
Correlation coe�cient = –.95 Correlation coe�cient = – .99
0 0
1
2
3
4
5
6
1 2 3 4 5 6
coh37025_ch03_085-128.indd 121 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
A graph also makes the spotting of outliers relatively easy. An outlier is an extremely atypical point located at a relatively long distance—an outlying distance—from the rest of the coordinate points in a scatterplot (Figure 3–17, right panel). Outliers stimulate interpreters of test data to speculate about the reason for the atypical score. For example, consider an outlier on a scatterplot that reflects a correlation between hours each member of a fifth-grade class spent studying and their grades on a 20-item spelling test. And let’s say that one student studied for 10 hours and received a failing grade. This outlier on the scatterplot might raise a red flag and compel the test user to raise some important questions, such as “How effective are this student’s study skills and habits?” or “What was this student’s state of mind during the test?”
In some cases, outliers are simply the result of administering a test to a small sample of testtakers. In the example just cited, if the test were given statewide to fifth-graders and the sample size were much larger, perhaps many more low scorers who put in large amounts of study time would be identified.
As is the case with low raw scores or raw scores of zero, outliers can sometimes help identify a testtaker who did not understand the instructions, was not able to follow the instructions, or was simply oppositional and did not follow the instructions. In other cases, an outlier can provide a hint of some deficiency in the testing or scoring procedures.
People who have occasion to use or make interpretations from graphed data need to know if the range of scores has been restricted in any way. To understand why this is so necessary to know, consider Figure 3–18. Let’s say that graph A describes the relationship between Public University entrance test scores for 600 applicants (all of whom were later admitted) and their grade point averages at the end of the first semester. The scatterplot indicates that the relationship between entrance test scores and grade point average is both linear and positive. But what if the admissions officer had accepted only the applications of the students who scored within the top half or so on the entrance exam? To a trained eye, this scatterplot (graph B) appears to indicate a weaker correlation than that indicated in graph A—an effect attributable exclusively to the restriction of range. Graph B is less a straight line than graph A, and its direction is not as obvious.
Figure �–�� Scatterplot showing nonlinear and linear relationships.
In the lower left side of the right panel, the isolated point is an outlier.
Outlier
Nonlinear relationship Linear relationship
0 2 4 6 �4 �2 0 2 4
�3
�2
�1
0
1
2
3
X
Y
coh37025_ch03_085-128.indd 122 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
Meta-Analysis
Generally, the best estimate of the correlation between two variables is most likely to come not from a single study alone but from analysis of the data from several studies. One option to facilitate understanding of the research across a number of studies is to present the range of statistical values calculated from a number of different studies of the same phenomenon. Viewing all of the data from a number of studies that attempted to determine the correlation between variable X and variable Y, for example, might lead the researcher to conclude that “The correlation between variable X and variable Y ranges from .73 to .91.” Another option might be to combine statistically the information across the various studies; that is what is done using a statistical technique called meta-analysis. Using this technique, researchers raise (and strive to answer) the question: “Combined, what do all of these studies tell us about the matter under study?” For example, Imtiaz et al. (2016) used meta-analysis to draw some conclusions regarding the relationship between cannabis use and physical health. Bolger (2015) used meta-analysis to study the correlations of use-of-force decisions among American police officers. Yang et al. (2020) used meta-analysis to examine pre-existing medical conditions (i.e., hypertension, respiratory system disease, and cardiovascular disease) as predictors of severe reactions to the 2019 coronavirus (COVID-19).
Meta-analysis may be defined as a family of techniques used to statistically combine information across studies to produce single estimates of the data under study. The estimates derived, referred to as effect size, may take several different forms. In most meta-analytic studies, effect size is typically expressed as a correlation coefficient.9 Meta-analysis facilitates the drawing
Figure �–�� Two scatterplots illustrating unrestricted and restricted ranges.
9. More generally, effect size refers to an estimate of the strength of the relationship (or the size of the differences) between groups. In a typical study using two groups (an experimental group and a control group) effect size, ideally reported with confidence intervals, is helpful in determining the effectiveness of some sort of intervention (such as a new form of therapy, a drug, a new management approach, and so forth). In practice, many different procedures may be used to determine effect size, and the procedure selected will be based on the particular research situation.
Unrestricted range Restricted range
0 25 50 75 100 0 25 50 75 100
0
1
2
3
4
College entrance test scores
C ol
le ge
g ra
de p
oi nt
a ve
ra ge
coh37025_ch03_085-128.indd 123 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
M E E T A N A S S E S S M E N T P R O F E S S I O N A L
students take their �rst psychology course, they are often surprised how much of the �eld is based on research �ndings rather than just “common sense.” Even so, because undergraduate textbooks have numerous topics about which they cannot cite all of the research, it can appear that the textbook is relying on just one or two studies as the “proof.” Therefore, you might be surprised just how many psychological research studies actually exist! Conducting a quick search in the PsycINFO database shows that over a million psychology journal articles are classi�ed as empirical studies— and that excludes chapters, theses, dissertations, and many other studies not listed in PsycINFO.
But, good news or bad news, a signi�cant challenge with many research studies is how to summarize results. The classic example of such a dilemma and the eventual solution is a fascinating one that comes from the psychotherapy literature. In ����, Hans Eysenck published a classic article
Meet Dr. Joni L. Mihura
i, my name is Joni Mihura, and my research expertise is in psychological assessment, with a special focus on the Rorschach. To tell you a little about me, I was the only woman* to serve on the Research Council for John E. Exner’s Rorschach Comprehensive System (CS) until he passed away in ����. Due to the controversy around the Rorschach’s validity, I began reviewing the research literature to ensure I was teaching my doctoral students valid measures to assess their clients. That is, the controversy about the Rorschach has not been that it is a completely invalid test—the critics have endorsed several Rorschach scales as valid for their intended purpose—the main problem that they have highlighted is that only a small proportion of its scales had been subjected to “meta-analysis,” a systematic technique for summarizing the research literature. To make a long story short, I eventually published my review of the Rorschach literature in the top scienti�c review journal in psychology (Psychological Bulletin) in the form of systematic reviews and meta-analyses of the �� main Rorschach CS variables (Mihura et al., ����), therefore making the Rorschach the psychological test with the most construct validity meta-analyses for its scales!
My meta-analyses also resulted in two other pivotal events. They formed the backbone for a new scientifically based Rorschach system of which I am a codeveloper—the Rorschach Performance Assessment System (R-PAS; Meyer et�al., ����), and they resulted in the Rorschach critics removing the “moratorium” they had recommended for the Rorschach (or, Garb, ����) for the scales they deemed had solid support in our meta-analyses (Wood et al., ����; also see our reply, Mihura et al., ����).
I’m very excited to talk with you about meta-analysis. First, to set the stage, let’s take a step back and look at what you might have experienced so far when reading about psychology. When
H
Joni L. Mihura, Ph.D., ABAP, is Professor of Psychology at the University of Toledo in Toledo, Ohio
© Joni L. Mihura, Ph.D.
*I have also edited the Handbook of Gender and Sexuality in Psychological Assessment (Brabender & Mihura, 2016).
coh37025_ch03_085-128.indd 124 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
summarize the research �ndings for a particular topic. In ����, Mary Lee Smith and Gene V. Glass published the �rst meta-analysis of psychotherapy outcomes. They found strong support for the e�cacy of psychotherapy. Subsequently, others tried to challenge Smith and Glass’ �ndings. However, the systematic rigor of their meta-analytic technique produced �ndings that were consistently replicated by others. Today there are thousands of psychotherapy studies, and many meta-analysts ready to research speci�c, therapy-related questions (like “What type of psychotherapy is best�for what type of problem?”).
What does all of this mean for psychological testing and assessment? Meta-analytic methodology can be used to glean insights about speci�c tools of assessment, and testing and assessment procedures. However, meta-analyses of information related to psychological tests brings new challenges owing, for example, to the sheer number of articles to be analyzed, the many variables on which tests�di�er, and the speci�c methodology of the meta-analysis. Consider, for example, that multiscale personality tests may contain over ��, and sometimes over ���, scales that each need to be evaluated separately. Furthermore, some popular multiscale personality tests, like the MMPI-� and Rorschach, have had over a thousand research studies published on them. The studies typically report �ndings that focus on varied aspects of the test (such as the utility of speci�c test scales, or other indices of test reliability or validity). In order to make the meta-analytic task manageable, meta- analyses for multiscale tests will typically focus on one or another of these characteristics or�indices.
In sum, a thoughtful meta-analysis of research on a speci�c topic can yield important insights of both theoretical and applied value. A meta-analytic review of the literature on a particular psychological test can even be instrumental in the formulation of revised ways to score the test and interpret the �ndings (just ask Meyer et al., ����). So, the next time a question about psychological research arises, students are advised to respond to that question with their own question, namely “Is there a meta-analysis on that?”
Used with permission of Dr. Joni L. Mihura.
entitled “The E�ects of Psychotherapy: An Evaluation,” in which he summarized the results of a few studies and concluded that psychotherapy doesn’t work! Wow! This �nding had the potential to shake the foundation of psychotherapy and even ban its existence. After all, Eysenck had cited research that suggested that the longer a person was in therapy, the worse-o� they became. Notwithstanding the psychotherapists and the psychotherapy enterprise, Eysenck’s publication had sobering implications for people who had sought help through psychotherapy. Had they done so in vain? Was there really no hope for the future? Were psychotherapists truly ill-equipped to do things like reduce emotional su�ering and improve peoples’ lives through psychotherapy?
In the wake of this potentially damning article, several psychologists—and in particular Hans H. Strupp—responded by pointing out problems with Eysenck’s methodology. Other psychologists conducted their own reviews of the psychotherapy literature. Somewhat surprisingly, after reviewing the same body of research literature on psychotherapy, various psychologists drew widely di�erent conclusions. Some researchers found strong support for the e�cacy of psychotherapy. Other researchers found only modest support for the e�cacy of psychotherapy. Yet other researchers found no support for it at all.
How can such di�erent conclusions be drawn when the researchers are reviewing the same body of literature? A comprehensive answer to this important question could �ll the pages of this book. Certainly, one key element of the answer to this question had to do with a lack of systematic rules for making decisions about including studies, as well as lack of a widely acceptable protocol for statistically summarizing the �ndings of the various studies. With such rules and protocols absent, it would be all too easy for researchers to let their preexisting biases run amok. The result was that many researchers “found” in their analyses of the literature what they believed to be true in the �rst place.
A fortuitous bi-product of such turmoil in the research community was the emergence of a research technique called “meta-analysis.” Literally, “an analysis of analyses,” meta-analysis is a tool used to systematically review and statistically
coh37025_ch03_085-128.indd 125 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
of conclusions and the making of statements like, “the typical therapy client is better off than 75% of untreated individuals” (Smith & Glass, 1977, p. 752), there is “about 10% increased risk for antisocial behavior among children with incarcerated parents, compared to peers” (Murray et al., 2012), and “GRE and UGPA [undergraduate grade point average] are generalizably valid predictors of graduate grade point average, 1st-year graduate grade point average, comprehensive examination scores, publication citation counts, and faculty ratings” (Kuncel et al., 2001, p. 162).
A key advantage of meta-analysis over simply reporting a range of findings is that, in meta-analysis, more weight can be given to studies that have larger numbers of subjects. This weighting process results in more accurate estimates (Hunter & Schmidt, 1990). Some advantages to meta-analyses are: (1) meta-analyses can be replicated; (2) the conclusions of meta-analyses tend to be more reliable and precise than the conclusions from single studies; (3) there is more focus on effect size rather than statistical significance alone; and (4) meta- analysis promotes evidence-based practice, which may be defined as professional practice that is based on clinical and research findings (Sánchez-Meca & Marín-Martínez, 2010). Despite these and other advantages, meta-analysis is, at least to some degree, art as well as science (Hall & Rosenthal, 1995). The value of any meta-analytic investigation is a matter of the skill and ability of the meta-analyst (Kavale, 1995), and use of an inappropriate meta-analytic method can lead to misleading conclusions (Kisamore & Brannick, 2008).
It may be helpful at this time to review this statistics refresher to make certain that you indeed feel “refreshed” and ready to continue. We will build on your knowledge of basic statistical principles in the chapters to come, and it is important to build on a rock-solid foundation.
Self-Assessment
Test your understanding of elements of this chapter by seeing if you can explain each of the following terms, expressions, and abbreviations:
arithmetic mean average deviation bar graph bimodal distribution bivariate distribution coefficient of correlation coefficient of determination correlation curvilinearity distribution dynamometer effect size error evidence-based practice frequency distribution frequency polygon graph grouped frequency distribution histogram interquartile range interval scale kurtosis
leptokurtic linear transformation mean measurement measure of central tendency measure of variability median mesokurtic meta-analysis mode negative skew nominal scale nonlinear transformation normal curve normalized standard score scale normalizing a distribution ordinal scale outlier Pearson r platykurtic positive skew quartile
range rank-order/rank-difference
correlation coefficient ratio scale raw score scale scatter diagram scattergram scatterplot semi-interquartile range skewness Spearman’s rho standard deviation standard score stanine T score tail variability variance z score
coh37025_ch03_085-128.indd 126 12/01/21 4:06 PM
Chapter 3: A Statistics Refresher ���
References
Addeo, R. R., Greene, A. F., & Geisser, M. E. (1994). Construct validity of the Robson Self-Esteem Questionnaire in a sample of college students. Educational and Psychological Measurement, 54, 439–446.
Adelman, S. A., Fletcher, K. E., Bahnassi, A., & Munetz, M. R. (1991). The Scale for Treatment Integration of the Dually Diagnosed (STIDD): An instrument for assessing intervention strategies in the pharmacotherapy of mentally ill substance abusers. Drug and Alcohol Dependence, 27, 35–42.
Amabile, T. M., Hill, K. G., Hennessey, B. A., & Tighe, E. M. (1994). The Work Preference Inventory: Assessing intrinsic and extrinsic motivational orientations. Journal of Personality and Social Psychology, 66, 950–967.
Bjornsen, C. A., & Archer, K. J. (2015). Relations between college students’ cell phone use during class and grades. Scholarship of Teaching and Learning in Psychology, 1(4), 326–336.
Bolger, P. C. (2015). Just following orders: A meta- analysis of the correlates of American police officer use of force decisions. American Journal of Criminal Justice, 40(3), 466–492.
Brabender, V. M., & Mihura, J. L. (Eds.) (2016). Handbook of gender and sexuality in psychological assessment. Routledge.
Burns, A., Jacoby, R., & Levy, R. (1991). Progression of cognitive impairment in Alzheimer’s disease. Journal of the American Geriatrics Society, 39, 39–45.
Cloninger, C. R., Przybeck, T. R., & Svrakis, D. M. (1991). The Tridimensional Personality Questionnaire: U.S. normative data. Psychological Reports, 69, 1047–1057.
Davies, P. L., & Gavin, W. J. (1994). Comparison of individual and group/consultation treatment methods for preschool children with developmental delays. American Journal of Occupational Therapy, 48, 155–161.
Eysenck, H. J. (1952). The effects of psychotherapy: An evaluation. Journal of Consulting Psychology, 16, 319–324.
Garb, H. N. (1999). Call for a moratorium on the use of the Rorschach Inkblot Test in clinical and forensic settings. Assessment, 6(4), 313–317. https://doi. org/10.1177/107319119900600402
Gokhale, D. V., & Kullback, S. (1978). The information in contingency tables. Marcel Dekker.
Hall, J. A., & Rosenthal, R. (1995). Interpreting and evaluating meta-analysis. Evaluation and the Health Professions, 18, 393–407.
Hopkins, K. D., & Glass, G. V. (1978). Basic statistics for the behavioral sciences. Prentice-Hall.
Hunter, J. E., & Schmidt, F. L. (1990). Methods of meta- analysis. Sage.
Hunter, M. S. (1992). The Women’s Health Questionnaire: A measure of mid-aged women’s perceptions of their emotional and physical health. Psychology and Health, 7, 45–54.
Imtiaz, S., Shield, K. D., Roerecke, M., et al. (2016). The burden of disease attributable to cannabis use in Canada in 2012. Addiction, 111(4), 653–662.
Kavale, K. A. (1995). Meta-analysis at 20: Retrospect and prospect. Evaluation and the Health Professionals, 18, 349–369.
Kerlinger, F. N. (1973). Foundations of behavioral research (2nd ed.). Holt.
Kisamore, J. L., & Brannick, M. T. (2008). An illustration of the consequences of meta-analysis model choice. Organizational Research Methods, 11, 35–53.
Kranzler, G., & Moursund, J. (1999). Statistics for the terrified (2nd ed.). Prentice-Hall.
Kuncel, N. R., Hezlett, S. A., & Ones, D. S. (2001). A comprehensive meta-analysis of the predictive validity of the Graduate Record Examinations: Implications for graduate student selection and performance. Psychological Bulletin, 127, 162–181.
Lindgren, B. (1983, August). N or N–1? [Letter to the editor]. American Statistician, p. 52.
Magnello, M. E., & Spies, C. J. (1984). Francis Galton: Historical antecedents of the correlation calculus. In B. Laver (Chair), History of mental measurement: Correlation, quantification, and institutionalization. Paper session presented at the 92nd annual convention of the American Psychological Association, Toronto.
McCall, W. A. (1922). How to measure in education. Macmillan.
McCall, W. A. (1939). Measurement. Macmillan. Meyer, G. J., Viglione, D. J., Mihura, J. L., Erard, R. E.,
& Erdberg, P. (2011). Rorschach Performance Assessment System: Administration, coding, interpretation, and technical manual. Rorschach Performance Assessment System.
Micceri, T. (1989). The unicorn, the normal curve and other improbable creatures. Psychological Bulletin, 105, 156–166.
Mihura, J. L., Meyer, G. J., Bombel, G., & Dumitrascu, N. (2015). Standards, accuracy, and questions of bias in Rorschach meta-analyses: Reply to Wood, Garb, Nezworski, Lilienfeld, and Duke (2015). Psychological Bulletin, 141, 250–260.
Mihura, J. L., Meyer, G. J., Dumitrascu, N., & Bombel, G. (2013). The validity of individual Rorschach variables: Systematic reviews and meta-analyses of the comprehensive system. Psychological Bulletin, 139, 548–605.
Murray J., Farrington D. P., & Sekol I. (2012). Children’s antisocial behavior, mental health, drug use, and educational performance after parental incarceration: a systematic review and meta-analysis. Psychological Bulletin, 138(2), 175–210. https://doi.org/10.1037 /a0026407
Nunnally, J. C. (1978). Psychometric theory (2nd ed.). McGraw-Hill.
Ozer, D. J. (1985). Correlation and the coefficient of determination. Psychological Bulletin, 97, 307–315.
Ranseen, J. D., & Humphries, L. L. (1992). The intellectual functioning of eating disorder patients. Journal of the American Academy of Child and Adolescent Psychiatry, 31, 844–846.
Robinson, N. M., Zigler, E., & Gallagher, J. J. (2000). Two tails of the normal curve: Similarities and differences in the study of mental retardation and giftedness. American Psychologist, 55, 1413–1424.
coh37025_ch03_085-128.indd 127 12/01/21 4:06 PM
������Part 2: The Science of Psychological Measurement
Rokeach, M. (1973). The nature of human values. Free Press.
Sánchez-Meca, J. & Marín-Martínez, F. (2010). Meta- analysis in psychological research. International Journal of Psychological Research, 3, 151–163.
Smith, M. L., & Glass, G. V. (1977). Meta-analysis of psychotherapy outcome studies. American Psychologist, 32, 752–760.
Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103, 677–680. https://doi. org/10.1126/science.103.2684.677
Streiner, D. L. (2003). Being inconsistent about consistency: When coefficient alpha does and doesn’t matter. Journal of Personality Assessment, 80, 217–222.
Tan, U. (1993). Normal distribution of hand preference and its bimodality. International Journal of Neuroscience, 68, 61–65.
Thompson, R. J., Gustafson, K. E., Meghdadpour, S., & Harrell, E. S. (1992). The role of biomedical and psychosocial processes in the intellectual and academic functioning of children and adolescents with cystic fibrosis. Journal of Clinical Psychology, 48, 3–10.
Thorndike, E. L., Bregman, E. O., Cobb, M. V., Woodward, E., & the staff of the Division of Psychology of the Institute of Educational Research of Teachers College, Columbia University. (1927). The measurement of intelligence. Bureau of Publications, Teachers College, Columbia University.
Varon, E. J. (1936). Alfred Binet’s concept of intelligence. Psychological Review, 43, 32–49.
von Knorring, L., & Lindstrom, E. (1992). The Swedish version of the Positive and Negative Syndrome Scale (PANSS) for schizophrenia: Construct validity and interrater reliability. Acta Psychiatrica Scandinavica, 86, 463–468.
Wood, J. M., Garb, H. N., Nezworski, M. T., Lilienfeld, S. O., & Duke, M. C. (2015). A second look at the validity of widely used Rorschach indices: Comment on Mihura, Meyer, Dumitrascu, and Bombel (2013). Psychological Bulletin, 141, 236–249.
Yang, J., Zheng, Y., Gou, X., Pu, K., Chen, Z., Guo, Q., . . . Zhou, Y. (2020). Prevalence of comorbidities and its effects in patients infected with SARS-CoV-2: a systematic review and meta-analysis. International Journal of Infectious Diseases, 94, 91–95. https://doi .org/10.1016/j.ijid.2020.03.017
coh37025_ch03_085-128.indd 128 12/01/21 4:06 PM
���
C H A P T E R �
Of Tests and Testing
A patient bursts into the emergency room announcing he is the archangel Dustin, with the power to heal the sick. When approached by staff members, he bursts into tears and sobs inconsolably. What is this patient’s diagnosis?
A young man with a long history of academic and conduct problems was arrested for armed robbery. His court-appointed lawyer is frustrated because his client is barely able to understand the proceedings against him, and he appears unable to assist in his own defense. Should the court�order a psychological evaluation to determine whether the young man is competent to stand�trial?
A large corporation hires thousands of entry-level employees. Who should be hired, transferred, promoted, or fired?
A college admissions office is considering the credentials of hundreds of applicants. Which individual should gain entry to this special program or be awarded a scholarship?
Each parent in the process of a bitter divorce has accused the other of negligent and abusive behavior toward their three children. The children’s mother has been arrested for shoplifting twice in the last year. Their father has been heard yelling loudly by the family’s neighbors. Who shall be granted custody of the children?
very day, throughout the world, critically important questions like these are addressed through the use of tests. The answers to these kinds of questions are likely to have a significant impact on many lives.
If they are to sleep comfortably at night, assessment professionals must have confidence in the tests and other tools of assessment they employ. They need to know, for example, what does and does not constitute a “good test.”
Our objective in this chapter is to overview the elements of a good test. As background, we begin by listing some basic assumptions about assessment. Aspects of these fundamental assumptions will be elaborated later on in this chapter as well as in subsequent chapters.
E
J U S T T H I N K . � . � .
What’s a “good test”? Outline some elements or features that you believe are essential to a good test before reading on.
coh37025_ch04_129-156.indd 129 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
Some Assumptions About Psychological Testing and Assessment
Assumption 1: Psychological Traits and States Exist
We humans are neither wholly predictable nor completely erratic, and this interplay of order and chaos makes the study of human behavior endlessly fascinating. To communicate efficiently, scholars have developed a technical vocabulary to describe components of stability and change in our behavior. A trait has been defined as “any distinguishable, relatively enduring way in which one individual varies from another” (Guilford, 1959, p. 6). States also distinguish one person from another but are relatively less enduring (Chaplin et al., 1988). The trait term that an observer applies, as well as the strength or magnitude of the trait presumed to be present, is based on observing a sample of behavior. Samples of behavior may be obtained in a number of ways, ranging from direct observation to the analysis of self-report statements given via pencil-and-paper test answers or via electronic responses.
The term psychological trait, much like the term trait alone, covers a wide range of possible characteristics. Thousands of psychological trait terms can be found in the English language (Allport & Odbert, 1936). Among them are psychological traits that relate to intelligence, specific intellectual abilities, cognitive style, adjustment, interests, attitudes, sexual orientation and preferences, psychopathology, personality in general, and specific personality traits. New concepts or discoveries in research may bring new trait terms to the fore. For example, a trait term seen in the professional literature on human sexuality is androgynous (referring to an absence of primacy of male or female characteristics). Cultural evolution may bring new trait terms into common usage, such as the term gender non-binary to refer to individuals who do not classify themselves on the masculine-feminine or male-female continuum.
Few people deny that psychological traits exist. Yet there has been a fair amount of controversy regarding just how they exist (McCabe & Fleeson, 2016; Sherman et al., 2015). For example, do traits have a physical existence, perhaps as a circuit in the brain? Although some have argued in favor of such a conception of psychological traits (Allport, 1937; Holt, 1971), compelling evidence to support such a view has been difficult to obtain. For our purposes, a psychological trait exists only as a construct—an informed, scientific concept developed or constructed to describe or explain behavior. We can’t see, hear, or touch constructs, but we can infer their existence from overt behavior. In this context, overt behavior refers to an observable action or the product of an observable action, including test- or assessment-related responses. A challenge facing test developers is to construct tests that are at least as telling as observable behavior such as that illustrated in Figure 4–1.
The phrase relatively enduring in our definition of trait is a reminder that a trait is not expected to be manifested in behavior 100% of the time So, for example, we may become more agreeable and conscientious as we age, and perhaps become less prone to “sweat the small stuff” (Lüdtke et al., 2009; Roberts et al., 2003, 2006). Yet even as personality evolves, it is partially stable over the lifespan. For example, energetic children tend to become active adults, even though most adults move about less than they did when they were younger. This stability of traits over time is evidenced by relatively high correlations between trait scores at different time points (Damian et al., 2019; Lüdtke et al., 2009; Roberts & DelVecchio, 2000).
Whether a trait manifests itself in observable behavior, and to what degree it manifests, is presumed to depend not only on the strength of the trait in the individual but also on the nature of the situation. Stated another way, exactly how a particular trait manifests itself is, at least to some extent, situation-dependent. For example, a violent parolee may be prone to behave in a rather subdued way with her parole officer and much more violently in the presence of her family and friends. John may be viewed as dull and cheap by his wife but as charming and extravagant by his business associates, whom he keenly wants to impress.
coh37025_ch04_129-156.indd 130 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
The context within which behavior occurs also plays a role in helping us select appropriate trait terms for observed behavior. Consider how we might label the behavior of someone who is kneeling and praying aloud. Such behavior might be viewed as either religious or deviant, depending on the context in which it occurs. A person who is doing this inside a church or upon a prayer rug may be described as religious, whereas another person engaged in the exact same behavior at a venue such as a sporting event or a movie theater might be viewed as deviant or paranoid.
The definitions of trait and state we are using also refer to a way in which one individual varies from another. Attributions of a trait or state term are relative. For example, in describing one person as shy, or even in using terms such as very shy or not shy, most people are making an unstated comparison with the degree of shyness they could reasonably expect the average person to exhibit under the same or similar circumstances. In psychological assessment, assessors may also make such comparisons with respect to the hypothetical average person. Alternatively, assessors may make comparisons among people who, because of their membership in some group or for any number of other reasons, are decidedly not average.
Figure �–� Measuring sensation seeking.
The psychological trait of sensation seeking has been defined as “the need for varied, novel, and complex sensations and experiences and the willingness to take physical and social risks for the sake of such experiences” (Zuckerman, 1979, p. 10). A 22-item Sensation-Seeking Scale (SSS) seeks to identify people who are high or low on this trait. Assuming the SSS actually measures what it purports to measure, how would you expect a random sample of people lining up to bungee jump to score on the test as compared with another age-matched sample of people shopping at the local mall? What are the comparative advantages of using paper-and-pencil measures, such as the SSS, and using more performance-based measures, such as the one pictured here? Vitalii Nesterchuk/Shutterstock
J U S T T H I N K . � . � .
Give another example of how the same behavior in two di�erent contexts may be viewed in terms of two di�erent traits.
coh37025_ch04_129-156.indd 131 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
As you might expect, the reference group with which comparisons are made can greatly influence one’s conclusions or judgments. For example, suppose a psychologist administers a test of shyness to a 22-year-old male who earns his living as an exotic dancer. The interpretation of the test data will almost surely differ as a function of the reference group with which the testtaker is compared—that is, other males in his age group or other male exotic dancers in his age group.
Assumption 2: Psychological Traits and States Can Be Quantified and Measured
E.L. Thorndike (1918, p.16) famously declared, “Whatever exists at all exists in some amount. To know it thoroughly involves knowing its quantity as well as its quality.” Professor Thorndike may have overstated his thesis, but we agree with him that most psychological traits and states vary by degree, and thus in theory can be quantified. Once it’s acknowledged that psychological traits and states do exist, the specific traits and states to be measured and quantified need to be carefully defined. Test developers and researchers, much like people in general, have many different ways of looking at and defining the same phenomenon. Just think, for example, of the different ways the term aggressive is used. We speak of an aggressive salesperson, an aggressive killer, and an aggressive waiter, to name but a few contexts. In each of these different contexts, aggressive carries with it a different meaning. If a personality test yields a score purporting to provide information about how aggressive a testtaker is, a first step in understanding the meaning of that score is understanding how aggressive was defined by the test developer. More specifically, what types of behaviors are presumed to be indicative of someone who is aggressive as defined by the test? One test developer may define aggressive behavior as “the number of self-reported acts of physically harming others.” Another test developer might define it as the number of observed acts of aggression, such as pushing, hitting, or kicking, that occur in a playground setting. Other test developers may define “aggressive behavior” in vastly different ways, such as socially aggressive acts like gossiping, ostracism, and slander. Ideally, the test developer has provided test users with a clear operational definition of the construct under study.
Once having defined the trait, state, or other construct to be measured, a test developer considers the types of item content that would provide insight into it. From a universe of behaviors presumed to be indicative of the targeted trait, a test developer has a world of possible items that can be written to gauge the strength of that trait in testtakers.1 For example, if the test developer deems knowledge of American history to be one component of intelligence in U.S. adults, then the item Who was the second president of the United States? may appear on the test. Similarly, if social judgment is deemed to be indicative of adult intelligence, then it might be reasonable to include the item Why should guns in the home always be inaccessible to children?
Suppose we agree that an item tapping knowledge of American history and an item tapping social judgment are both appropriate for an adult intelligence test. One question that arises is: Should both items be given equal weight? That is, should we place more importance on—and award more points for—an answer keyed “correct” to one or the other of these two items? Perhaps a correct response to the social judgment question should earn more credit than a correct response to the American history question. Weighting the
J U S T T H I N K . � . � .
Is the strength of a particular psychological trait the same across all situations or environments? What are the implications of one’s answer to this question for assessment?
1. In the language of psychological testing and assessment, the word domain is substituted for world in this context. Assessment professionals speak, for example, of domain sampling, which may refer to either (1) a sample of behaviors from all possible behaviors that could conceivably be indicative of a particular construct or (2) a sample of test items from all possible items that could conceivably be used to measure a particular construct.
coh37025_ch04_129-156.indd 132 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
comparative value of a test’s items comes about as the result of a complex interplay among many factors, including technical considerations, the way a construct has been defined for the purposes of the test, and the value society (and the test developer) attaches to the behaviors evaluated.
Measuring traits and states by means of a test entails developing not only appropriate test items but also appropriate ways to score the test and interpret the results. For many varieties of psychological tests, some number representing the score on the test is derived from the examinee’s responses. The test score is presumed to represent the strength of the targeted ability or trait or state and is frequently based on cumulative scoring.2 In cumulative scoring, a trait is measured by a series of test items. Each response to a test item is converted to a number according to a test “key” (e.g., correct = 1 and incorrect = 0). The magnitude of the trait is assumed to correspond in some way to the sum of the keyed responses. You were probably first introduced to cumulative scoring early in elementary school when you observed that your score on a weekly spelling test had everything to do with how many words you spelled correctly or incorrectly. The score reflected the extent to which you had successfully mastered the spelling assignment for the week. On the basis of that score, we might predict that you would spell those words correctly if called upon to do so. And in the context of such prediction, consider the next assumption.
Assumption 3: Test-Related Behavior Predicts Non-Test-Related Behavior
Many tests involve tasks such as blackening little grids with a number 2 pencil, pressing keys on a computer keyboard, or tapping the screen of your cell phone. The objective of such tests typically has little to do with predicting future grid-blackening key-pressing, or screen tapping behavior. Rather, the objective of the test is to provide some indication of other aspects of the examinee’s behavior. For example, patterns of answers to a test of personality can be used in decision making regarding mental disorders.
The tasks in some tests mimic the actual behaviors that the test user is attempting to understand. By their nature, however, such tests yield only a sample of the behavior that can be expected to be emitted under nontest conditions. The obtained sample of behavior is typically used to make predictions about future behavior, such as work performance of a job applicant. In some forensic (legal) matters, psychological tests may be used not to predict behavior but to postdict it—that is, to aid in the understanding of behavior that has already taken place. For example, there may be a need to understand a criminal defendant’s state of mind at the time of the commission of a crime. It is beyond the capability of any known testing or assessment procedure to reconstruct someone’s state of mind. Still, behavior samples may shed light, under certain circumstances, on someone’s state of mind in the past. Additionally, other tools of assessment—such as case history data or the defendant’s personal diary during the period in question—might be of great value in such an evaluation.
Assumption 4: All Tests Have Limits and Imperfections
Competent test users understand a great deal about the tests they use. They understand, among other things, how a test was developed, the circumstances under which it is appropriate to
2. Other models of scoring are discussed in Chapter 8.
J U S T T H I N K . � . � .
On an adult intelligence test, what type of item should be given the most weight? What type of item should be given the least weight?
J U S T T H I N K . � . � .
In practice, tests have proven to be good predictors of some types of behaviors and not-so-good predictors of other types of behaviors. For example, tests have not proven to be as good at predicting violence as had been hoped. Why do you think it is so di�cult to predict violence by means of a test?
coh37025_ch04_129-156.indd 133 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
administer the test, how the test should be administered and to whom, and how the test results should be interpreted. Competent test users understand and appreciate the limitations of the tests they use as well as how those limitations might be compensated for by data from other sources. All of this may sound quite commonsensical, and it probably is. Yet this deceptively simple assumption—that test users know the tests they use and are aware of the tests’ limitations—is emphasized repeatedly in the codes of ethics of associations of assessment professionals.
Assumption 5: Various Sources of Error Are Part of the Assessment Process
In everyday conversation, we use the word error to refer to mistakes, miscalculations, and the like. In the context of assessment, error need not refer to a deviation, an oversight, or something that otherwise violates expectations. To the contrary, error traditionally refers to something that is more than expected; it is actually a component of the measurement process. More specifically, error refers to a long-standing assumption that factors other than what a test attempts to measure will influence performance on the test. Test scores are always subject to questions about the degree to which the measurement process includes error. For example, an intelligence test score could be subject to debate concerning the degree to which the obtained score truly reflects the examinee’s intelligence and the degree to which it was due to factors other than intelligence. Because error is a variable that must be taken account of in any assessment, we often speak of error variance, that is, the component of a test score attributable to sources other than the trait or ability measured.
There are many potential sources of error variance. Whether an assessee has the flu when taking a test is a source of error variance. In a more general sense, then, assessees themselves are sources of error variance. Assessors, too, are sources of error variance. For example, some assessors are more professional than others in the extent to which they follow the instructions governing how and under what conditions a test should be administered. In addition to assessors and assessees, measuring instruments themselves are another source of error variance. Some tests are simply better than others in measuring what they purport to measure. Some error is random, or, for lack of a better term, just a matter of chance. To illustrate, consider the weather outside, right now, as you are reading this chapter. If it is daytime, would you characterize the weather as unambiguously sunny, unambiguously rainy, or mixed? Now, consider the weather at another random time—the day that happens to be the one that a personality test is being administered. Might the weather on the day that one takes a personality test affect that person’s test scores? According to Beatrice Rammstedt and her colleagues (2015), the answer is “blowing in the wind” (see Figure 4–2).
Instructors who teach the undergraduate measurement course will occasionally hear a student refer to error as “creeping into” or “contaminating” the measurement process. Yet measurement professionals tend to view error as simply an element in the process of measurement, one for which any theory of measurement must surely account. In Chapter 5, we will explore various ways in which measurement error is measured and how it can be minimized. We estimate measurement error in part because it puts limits on how confident we can be in our test score interpretation.
Assumption 6: Unfair and Biased Assessment Procedures Can Be Identified and Reformed
If we had to pick the one of these seven assumptions that is more controversial than the remaining six, this one is it. Decades of court challenges to various tests and testing programs have sensitized test developers and users to the societal demand for fair tests used in a fair manner. Today all major test publishers strive to develop instruments that are fair when used in strict accordance with guidelines in the test manual. Assessment experts have developed a set of
coh37025_ch04_129-156.indd 134 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
sophisticated procedures to identify and correct test bias and a thoughtful list of ethical guidelines to ensure test fairness. However, despite the best efforts of many professionals, fairness-related questions and problems do occasionally arise. One source of fairness-related problems is the test user who attempts to use a particular test with people whose background and experience are different from the background and experience of people for whom the test was intended. Some potential problems related to test fairness are more political than psychometric. For example, heated debate on selection, hiring, and access or denial of access to various opportunities often surrounds affirmative action programs. In many cases the real question for debate is not “Is this test or assessment procedure fair?” but rather “What do we as a society wish to accomplish by the use of this test or assessment procedure?” In all questions about tests with regard to fairness, it is important to keep in mind that tests are tools. And just like other, more familiar tools (hammers, ice picks, wrenches, and so on), they can be used properly or improperly.
Assumption 7: Testing and Assessment Offer Powerful Benefits to Society
At first glance, the prospect of a world devoid of testing and assessment might seem appealing, especially from the perspective of a harried student preparing for a week of
FIGURE �–� Weather and self-concept.
There is research to suggest that self-reported personality ratings may differ depending upon the weather on the day that the self-report was made (Rammstedt et al., 2015). For example, people rate themselves as less disciplined and dutiful on sunny days compared to rainy days, perhaps because sunny days offer more opportunities to relax and enjoy oneself. This research is instructive regarding the extent to which random situational conditions (such as the weather on the day of an assessment) may affect the expression of traits. Andrei Mayatnik/Shutterstock
J U S T T H I N K . � . � .
Do you believe that testing can be conducted in a fair and unbiased manner?
coh37025_ch04_129-156.indd 135 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
midterm examinations. Yet a world without tests would most likely be more a nightmare than a dream. In such a world, people could present themselves as surgeons, bridge builders, or airline pilots regardless of their background, ability, or professional credentials. In a world without tests or other assessment procedures, personnel might be hired on the basis of nepotism rather than documented merit. In a world without tests, teachers and school administrators could arbitrarily place children in different types of special classes simply because that is where they believed the children belonged. In a world without tests, there
would be a great need for instruments to diagnose educational difficulties in reading and math and point the way to remediation. In a world without tests, there would be no instruments to diagnose neuropsychological impairments. In a world without tests, there would be no practical way for the military to screen thousands of recruits with regard to many key variables.
Considering the many critical decisions that are based on testing and assessment procedures, we can readily appreciate the need for tests, especially good tests. And that, of course, raises one critically important question . . .
What’s a “Good Test”?
Logically, the criteria for a good test would include clear instructions for administration, scoring, and interpretation. It would also seem to be a plus if a test offered economy in the time and money it took to administer, score, and interpret it. Most of all, a good test would seem to be one that measures what it purports to measure.
Beyond simple logic, there are technical criteria that assessment professionals use to evaluate the quality of tests and other measurement procedures. Test users often speak of the psychometric soundness of tests, two key aspects of which are reliability and validity.
Reliability
A good test or, more generally, a good measuring tool or procedure is reliable. As we will explain in Chapter 5, the criterion of reliability involves the consistency of the measuring tool: the precision with which the test measures and the extent to which error is present in measurements. In theory, the perfectly reliable measuring tool consistently measures in the same way.
To exemplify reliability, visualize three digital scales labeled A, B, and C. To determine if they are reliable measuring tools, we will use a standard 1-pound gold bar that has been certified by experts to indeed weigh 1 pound and not a fraction of an ounce more or less. Now, let the testing begin.
Repeated weighings of the 1-pound bar on Scale A register a reading of 1 pound every time. No doubt about it, Scale A is a reliable tool of measurement. On to Scale B. Repeated weighings of the bar on Scale B yield a reading of 1.3 pounds. Is this scale reliable? It sure is! It may be consistently inaccurate by three-tenths of a pound, but there’s no taking away the fact that it is reliable. Finally, Scale C. Repeated weighings of the bar on Scale C register a different weight every time. On one weighing, the gold bar weighs in at 1.7 pounds. On the next weighing, the weight registered is 0.9 pound. In short, the weights registered are all over the map. Is this scale reliable? Hardly. This scale is neither reliable nor accurate. Contrast it to Scale B, which also did not record the weight of the gold standard correctly. Although inaccurate, Scale B was consistent in terms of how much the registered weight deviated from the true weight. By contrast, the weight registered by Scale C deviated from the true weight of the bar in seemingly random fashion.
Whether we are measuring gold bars, behavior, or anything else, unreliable measurement is to be avoided. We want to be reasonably certain that the measuring tool or test that we are
J U S T T H I N K . � . � .
How else might a world without tests or other assessment procedures be di�erent from the world today?
coh37025_ch04_129-156.indd 136 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
using is consistent. That is, we want to know that it yields the same numerical measurement every time it measures the same thing under the same conditions. Psychological tests, like other tests and instruments, are reliable to varying degrees. As you might expect, however, reliability is a necessary but not sufficient element of a good test. In addition to being reliable, tests must be reasonably accurate. In the language of psychometrics, tests must be valid.
Validity
A test is considered valid for a particular purpose if it does, in fact, measure what it purports to measure. In the gold bar example cited earlier, the scale that consistently indicated that the 1-pound gold bar weighed 1 pound is a valid scale. Likewise, a test of reaction time is a valid test if it accurately measures reaction time. A test of intelligence is a valid test if it truly measures intelligence. Well, yes, but . . .
Although there is relatively little controversy about the definition of a term such as reaction time, a great deal of controversy exists about the definition of intelligence. Because there is controversy surrounding the definition of intelligence, the validity of any test purporting to measure this variable is sure to be closely scrutinized by critics. If the definition of intelligence on which the test is based is sufficiently different from the definition of intelligence on other accepted tests, then the test may be condemned as not measuring what it purports to measure.
Questions regarding a test’s validity may focus on the items that collectively make up the test. Do the items adequately sample the range of areas that must be sampled to adequately measure the construct? Individual items will also come under scrutiny in an investigation of a test’s validity. How do individual items contribute to or detract from the test’s validity? The validity of a test may also be questioned on grounds related to the interpretation of resulting test scores. What do these scores really tell us about the targeted construct? How are high scores on the test related to testtakers’ behavior? How are low scores on the test related to testtakers’ behavior? How do scores on this test relate to scores on other tests purporting to measure the same construct? How do scores on this test relate to scores on other tests purporting to measure opposite types of constructs?
We might expect one person’s score on a valid test of introversion to be inversely related to that same person’s score on a valid test of extraversion; that is, the higher the introversion test score, the lower the extraversion test score, and vice versa. As we will see when we discuss validity in greater detail in Chapter 6, questions concerning the validity of a particular test may be raised at every stage in the life of a test. From its initial development through the life of its use with members of different populations, assessment professionals may raise questions regarding the extent to which a test is measuring what it purports to measure.
Other Considerations
A good test is one that trained examiners can administer, score, and interpret with a minimum of difficulty. A good test is a useful test, one that yields actionable results that will ultimately benefit individual testtakers or society at large. In “putting a test to the test,” there are a number of ways to evaluate just how good a test really is (see this chapter’s Everyday Psychometrics).
If the purpose of a test is to compare the performance of the testtaker with the performance of other testtakers, then a “good test” is one that contains adequate norms. Also referred to as normative data, norms provide a standard with which the results of measurement can be compared. Let’s explore the important subject of norms in a bit more detail.
J U S T T H I N K . � . � .
Why might a test shown to be valid for use for a particular purpose with members of one population not be valid for use for that same purpose with members of another population?
coh37025_ch04_129-156.indd 137 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
E V E R Y D A Y P S Y C H O M E T R I C S
Putting Tests to the Test
or experts in the �eld of testing and assessment, certain questions occur almost re�exively in evaluating a test or measurement technique. As a student of assessment, you may not be expert yet, but consider the questions that follow when you come across mention of any psychological test or other measurement technique.
Why Use This Particular Instrument or Method?
Typically there will be a choice of measuring instruments when it comes to measuring a particular psychological or educational variable, and the test user must therefore choose from many available tools. Why use one over another? Answering this question typically entails raising other questions, such as: What is the objective of using a test and how well does the test under consideration meet that objective? Who is this test designed for use with (age of testtakers? reading level? etc.) and how appropriate is it for the targeted testtakers? How is what the test measures de�ned? For example, if a test user seeks a test of “leadership,” how is “leadership” de�ned by the test developer (and how close does this de�nition match the test user’s de�nition of leadership for the purposes of the assessment)? What type of data will be generated from using this test, and what other types of data will it be necessary to generate if this test is used? Do alternate forms of this test exist? Answers to questions about speci�c instruments may be found in published sources of information (such as test catalogues, test manuals, and published test reviews) as well as unpublished sources (correspondence with test developers and publishers and with colleagues who have used the same or similar tests). Answers to related questions about the use of a particular instrument may be found elsewhere—for example, in published guidelines. This brings us to another question to “put to the test.”
Are There Any Published Guidelines for the Use of This Test?
Measurement professionals make it their business to be aware of published guidelines from professional associations and related organizations for the use of tests and measurement techniques. Sometimes a published guideline for the use of a particular test will list other measurement tools that should also be used along with it. For example, consider the case of psychologists called upon to provide input to a court in the matter of a child custody decision. More specifically, the court has asked the psychologist for
F expert opinion regarding an individual’s parenting capacity. Many psychologists who perform such evaluations use a psychological test as part of the evaluation process. However, the psychologist performing such an evaluation is— or should be—aware of the guidelines promulgated by the American Psychological Association’s Committee on Professional Practice and Standards. These guidelines describe three types of assessments relevant to a child custody decision: (�) the assessment of parenting capacity, (�) the assessment of psychological and developmental needs of the child, and (�) the assessment of the goodness of fit between the parent’s capacity and the child’s needs. According to these guidelines, an evaluation of a parent—or even of two parents—is not sufficient to arrive at an opinion regarding custody. Rather, an educated opinion about who should be awarded custody can be arrived at only after evaluating (�) the parents (or others seeking custody), (�) the child, and (�) the goodness of fit between the needs and capacity of each of the parties.
In this example, published guidelines inform us that any instrument the assessor selects to obtain information about parenting capacity must be supplemented with other instruments or procedures designed to support any expressed opinion, conclusion, or recommendation. In everyday practice, these other sources of data will be derived using other tools of psychological assessment such as interviews, behavioral observation, and case history or document analysis. Published guidelines and research may also provide useful information regarding how likely the use of a particular test or measurement technique is to meet standards set by courts (see, e.g., Yañez & Fremouw, ����).
Is This Instrument Reliable?
Earlier we introduced you to the psychometric concept of reliability and noted that it concerned the consistency of measurement. An assessor’s due diligence to determine whether a particular instrument is reliable starts with a careful reading of the test’s manual and of published research on the test, test reviews, and related sources. However, it does not necessarily end with such research.
Measuring reliability is not always a straightforward matter. For example, we might want to measure a person’s current a�ective state—what in everyday language we would call mood. We want to be sure that the measurement of emotional states is reliable in the sense that we measure states accurately and with
coh37025_ch04_129-156.indd 138 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
precision. We would not want the measurement to indicate that a person experiencing bliss is feeling distressed or that a currently distressed person is feeling bliss. However, emotional states can change quickly from one moment to the next. Thus, if emotional state scores have a low retest reliability coe�cient, the scores are not necessarily inaccurate. To estimate the reliability of emotional states—or any other construct we expect is not stable—we can measure it twice over short intervals (e.g.,��� seconds) or we can measure it with a short series of test items. There are statistical procedures that can estimate the reliability of the measurement from the consistency of item responses. We will discuss these measures of “internal consistency” in Chapter �.
Is This Instrument Valid?
Validity, as you have learned, refers to the extent to which a test measures what it purports to measure. And as was the case with questions concerning a particular instrument’s reliability, research to determine whether a particular instrument is valid starts with a careful reading of the test’s manual as well as published research on the test, test reviews, and related sources. Once again, as you might have anticipated, there will not necessarily be any simple answers at the end of this preliminary research.
As with reliability, questions related to the validity of a test can be complex and colored more in shades of gray than black or white. For example, interrater reliability is the degree to which different respondents give similar evaluations of a behavior or trait. In the assessment of childhood behavior problems, parents and teachers often give discrepant ratings. That is, parents might report that their child has high levels of anxiety, whereas the teacher might report the child has typical levels of anxiety. Who is right? Many children are anxious in one setting but not in another. Thus, it is possible that “low interrater reliability” is not a problem because it reflects reality (De Los Reyes et al., ����). The need for multiple sources of data on which to base an opinion stems not only from the ethical mandates published in the form of guidelines from professional associations but also from the practical demands of meeting a burden of proof in court. In sum, what starts as research to determine the validity of an individual instrument for a particular objective may end with research as to which combination of instruments will best achieve that objective.
Is This Instrument Cost-E�ective?
During World Wars I and II, the military needed to quickly screen hundreds of thousands of recruits for intelligence. It
may have been desirable to individually administer a Binet intelligence test to each recruit, but it would have taken a great deal of time—too much time, given the demands of war— and it would not have been very cost-e�ective. Instead, the armed services developed group measures of intelligence that could be administered quickly and that addressed its needs more e�ciently than an individually administered test. In this instance, it could be said that group tests had greater utility than individual tests.
What Inferences May Reasonably Be Made from This Test Score, and How Generalizable Are the Findings?
In evaluating a test, it is critical to consider the inferences that may reasonably be made as a result of administering that test. Will we learn something about a child’s readiness to begin �rst grade? about whether one is harmful to oneself or others? about whether an employee has executive potential? These queries represent but a small sampling of critical questions for which answers must be inferred on the basis of test scores and other data derived from various tools of assessment.
Intimately related to considerations regarding the inferences that can be made are those regarding the generalizability of the findings. As you learn more and more about test norms, for example, you will discover that the population of people used to help develop a test has a great effect on the generalizability of findings from an administration of the test. Many other factors may affect the generalizability of test findings. For example, if the items on a test are worded in such a way as to be less comprehensible by members of a specific group, then the use of that test with members of that group could be questionable. Another issue regarding the generalizability of findings concerns how a test was administered. Most published tests include explicit directions for testing conditions and test administration procedures that must be followed to the letter. If a test administration deviates in any way from these directions, the generalizability of the findings may be compromised. Culture is a variable that must be taken account of in the development of new tests as well as the administration, scoring, and interpretation of any test. The role of culture, too often overlooked in testing and assessment, will be emphasized and elaborated on at various points throughout this book.
Although you may not yet be an expert in measurement, you are now aware of the types of questions experts ask when evaluating tests. It is hoped that you can now appreciate that simple questions such as “What’s a good test?” don’t necessarily have simple answers.
coh37025_ch04_129-156.indd 139 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
Norms
We may define norm-referenced testing and assessment as a method of evaluation and a way of deriving meaning from test scores by evaluating an individual testtaker’s score and comparing it to scores of a group of testtakers. In this approach, the meaning of an individual test score is understood relative to other scores on the same test. A common goal of norm-referenced tests is to yield information on a testtaker’s standing or ranking relative to some comparison group of testtakers.
Norm in the singular is used in the scholarly literature to refer to behavior that is usual, average, normal, standard, expected, or typical. Reference to a particular variety of norm may be specified by means of modifiers such as age, as in the term age norm. Norms is the plural form of norm, as in the term gender norms. In a psychometric context, norms are the test performance data of a particular group of testtakers that are designed for use as a reference when evaluating or interpreting individual test scores. As used in this definition, the “particular group of testtakers” may be defined broadly (e.g., “a sample representative of the adult population of the United States”) or narrowly (e.g., “female inpatients at the Bronx Community Hospital with a primary diagnosis of depression”). A normative sample is that group of people whose performance on a particular test is analyzed for reference in evaluating the performance of individual testtakers.
Whether broad or narrow in scope, members of the normative sample will all be typical with respect to some characteristic(s) of the people for whom the particular test was designed. A test administration to this representative sample of testtakers yields a distribution (or distributions) of scores. These data constitute the norms for the test and typically are used as a reference source for evaluating and placing into context test scores obtained by individual testtakers. The data may be in the form of raw scores or converted scores.
The verb to norm, as well as related terms such as norming, refer to the process of deriving norms. Norming may be modified to describe a particular type of norm derivation. For example, race norming is the controversial practice of norming on the basis of race or ethnic background. Race norming was once engaged in by some government agencies and private organizations, and the practice resulted in the establishment of different cutoff scores for hiring by cultural group. Members of one cultural group would have to attain one score to be hired, whereas members of another cultural group would have to attain a different score. Although initially instituted in the service of affirmative action objectives (Greenlaw & Jensen, 1996), the practice was outlawed by the Civil Rights Act of 1991. If decision makers cannot use different norms for different groups, is it legal to adjust the test items so that different groups are, on average, more likely to obtain similar scores? A number of scholars have developed procedures designed to make equitable hiring practices more likely (e.g., Song et al., 2017).
Norming a test, especially with the participation of a nationally representative normative sample, can be an expensive proposition. For this reason, some test manuals provide what are variously known as user norms or program norms, which “consist of descriptive statistics based on a group of testtakers in a given period of time rather than norms obtained by formal sampling methods” (Nelson, 1994, p. 283). Understanding how norms are derived through “formal sampling methods” requires some discussion of the process of sampling.
Sampling to Develop Norms
The process of administering a test to a representative sample of testtakers for the purpose of establishing norms is referred to as standardization or test standardization. As will be clear from this chapter’s Close-Up, a test is said to be standardized when it has clearly specified procedures for administration and scoring, typically including normative data. To understand how norms are derived, an understanding of sampling is necessary.
coh37025_ch04_129-156.indd 140 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
C L O S E � U P
How “Standard” Is Standard in Measurement?
he foot, a unit of distance measurement in the United States, probably had its origins in the length of a British king’s foot used as a standard—one that measured about �� inches, give or take. It wasn’t so very long ago that di�erent localities throughout the world all had di�erent “feet” to measure by. We have come a long way since then, especially with regard to standards and standardization in measurement . . . haven’t we?
Perhaps. However, in the �eld of psychological testing and assessment, there’s still more than a little confusion when it comes to the meaning of terms like standard and standardization. Questions also exist concerning what is and is not standardized. To address these and related questions, a close-up look at the word standard and its derivatives seems in�order.
The word standard can be a noun or an adjective, and in either case it may have multiple (and quite di�erent) de�nitions. As a noun, standard may be de�ned as that which others are compared to or evaluated against. One may speak, for example, of a test with exceptional psychometric properties as being “the standard against which all similar tests are judged.” An exceptional textbook on the subject of psychological testing and assessment—take the one you are reading, for example— may be judged “the standard against which all similar textbooks are judged.” Perhaps the most common use of standard as a noun in the context of testing and assessment is in the title of that well-known manual that sets forth ideals of professional behavior against which any practitioner’s behavior can be judged: The Standards for Educational and Psychological Testing, usually referred to simply as the Standards.
As an adjective, standard often refers to what is usual, generally accepted, or commonly employed. One may speak, for example, of the standard way of conducting a particular measurement procedure, especially as a means of contrasting it to some newer or experimental measurement procedure. For example, a researcher experimenting with a new, multimedia approach to conducting a mental status examination might conduct a study to compare the value of this approach to the standard mental status examination interview.
In some areas of psychology, there has been a need to create a new standard unit of measurement in the interest of better understanding or quantifying particular phenomena. For example, in studying alcoholism and associated problems, many researchers have adopted the concept of a standard drink. The notion of a “standard drink” is designed to facilitate
T
communication and to enhance understanding regarding alcohol consumption patterns (Aros et al., ����; Gill et al., ����), intervention strategies (Hwang, ����; Podymow et al., ����), and costs associated with alcohol consumption (Farrell, ����). Regardless of whether it is beer, wine, liquor, or any other alcoholic beverage, reference to a “standard drink” immediately conveys information to the knowledgeable researcher about the amount of alcohol in the beverage.
The verb “to standardize” refers to making or transforming something into something that can serve as a basis of comparison or judgment. One may speak, for example, of the e�orts of researchers to standardize an alcoholic beverage that contains �� milliliters of alcohol as a “standard drink.” For many of the variables commonly used in assessment studies, there is
Figure � Ben’s Cold Cut Preference Test (CCPT).
Ben owns a small “deli boutique” that sells 10 varieties of private-label cold cuts. Ben had read somewhere that if a test has clearly specified methods for test administration and scoring, then it must be considered “standardized.” He then went on to create his own “standardized test”—the Cold Cut Preference Test (CCPT). The CCPT consists of only two questions: “What would you like today?” and a follow-up question, “How much of that would you like?” Ben scrupulously trains his only colleague (his wife—it’s literally a “mom and pop” business) on “test administration” and “test scoring” of the CCPT. So, just think: Does the CCPT really qualify as a “standardized test”? DreamPictures/Pam Ostrow/Blend Images LLC
(continued)
coh37025_ch04_129-156.indd 141 12/01/21 4:07 PM
������ Part 2: The Science of Psychological Measurement
an attempt to standardize a de�nition. As an example, Anderson (����) sought to standardize exactly what is meant by “creative thinking.” Well known to any student who has ever taken a nationally administered achievement test or college admission examination is the standardizing of tests. But what does it mean to say that a test is “standardized”? Some “food for thought” regarding an answer to this deceptively simple question can be found in Figure �.
Test developers standardize tests by developing replicable procedures for administering the test and for scoring and interpreting the test. Also part of standardizing a test is developing norms for the test. Well, not necessarily . . . whether norms for the test must be developed in order for the test to be deemed “standardized” is debatable. It is true that almost any “test” that has clearly speci�ed procedures for administration, scoring, and interpretation can be considered “standardized.” So even Ben the deli guy’s CCPT (described in Figure �) might be deemed a “standardized test” according to some because the test is “standardized” to the extent that the “test items” are clearly speci�ed (presumably along with “rules” for “administering” them and rules for “scoring and interpretation”). Still, many assessment professionals would hesitate to refer to Ben’s CCPT as a “standardized test.” Why?
Traditionally, assessment professionals have reserved the term standardized test for those tests that have clearly speci�ed procedures for administration, scoring, and interpretation in addition to norms. Such tests also come with manuals that are as much a part of the test package as the test’s items. Ideally, the test manual, which may be published in one or more booklets, will provide potential test users with all of the information they need to use the test in a responsible fashion. The test manual enables the test user to administer the test in the “standardized” manner in which it was designed to be administered; all test users should be able to replicate the test administration as prescribed by the test developer. Ideally, there will be little deviation from examiner to examiner in the way that a standardized test is administered, owing to the rigorous preparation and training that all potential users of the test have undergone prior to administering the test to testtakers.
If a standardized test is designed for scoring by the test user (in contrast to computer scoring), the test manual will ideally contain detailed scoring guidelines. If the test is one of ability that has correct and incorrect answers, the manual will ideally
contain an ample number of examples of correct, incorrect, or partially correct responses, complete with scoring guidelines. In like fashion, if it is a test that measures personality, interest, or any other variable that is not scored as correct or incorrect, then ample examples of potential responses will be provided along with complete scoring guidelines. We would also expect the test manual to contain detailed guidelines for interpreting the test results, including samples of both appropriate and inappropriate generalizations from the �ndings.
Also from a traditional perspective, we think of standardized tests as having undergone a standardization process. Conceivably, the term standardization could be applied to “standardizing” all the elements of a standardized test that need to be standardized. Thus, for a standardized test of leadership, we might speak of standardizing the de�nition of leadership, standardizing test administration instructions, standardizing test scoring, standardizing test interpretation, and so forth. Indeed, one de�nition of standardization as applied to tests is “the process employed to introduce objectivity and uniformity into test administration, scoring and interpretation” (Robertson, ����, p. ��). Another and perhaps more typical use of standardization, however, is reserved for that part of the test development process during which norms are developed. It is for this very reason that the terms test standardization and test norming have been used interchangeably by many test professionals.
Assessment professionals develop and use standardized tests to bene�t testtakers, test users, and/or society at large. Although there is conceivably some bene�t to Ben in gathering data on the frequency of orders for a pound or two of bratwurst, this type of data gathering does not require a “standardized test.” So, getting back to Ben’s CCPT . . . although some writers would staunchly defend the CCPT as a “standardized test” (simply because any two questions with clearly speci�ed guidelines for administration and scoring would make the “cut”), practically speaking this acceptance of the CCPT as a standardized test is simply not the case from the perspective of most assessment professionals.
There are a number of other ambiguities in psychological testing and assessment when it comes to the use of the word standard and its derivatives. Consider, for example, the term standard score. Some test manuals and books reserve the term standard score for use with reference to z scores. Raw scores (as well as z scores) linearly transformed to any
C L O S E � U P
How “Standard” Is Standard in Measurement? (continued)
coh37025_ch04_129-156.indd 142 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
other type of standard scoring systems—that is, transformed to a scale with an arbitrarily set mean and standard deviation—are differentiated from z scores by the term standardized. For these authors, a z score would still be referred to as a “standard score” whereas a T score, for example, would be referred to as a “standardized score.”
For the purpose of tackling another “nonstandard” use of the word standard, let’s digress for just a moment to images of the great American pastime of baseball. Imagine, for a moment, all of the di�erent ways that players can be charged with an error. There really isn’t one type of error that could be characterized as
standard in the game of baseball. Now, back to psychological testing and assessment—where there also isn’t just one variety of error that could be characterized as “standard.” No, there isn’t one . . . there are lots of them! One speaks, for example, of the standard error of measurement (also known as the standard error of a score) the standard error of estimate (also known as the standard error of prediction), the standard error of the mean, and the standard error of the di�erence. A table brie�y summarizing the main di�erences between these terms is presented here, although they are discussed in greater detail elsewhere in this book.
Type of “Standard Error” What Is It?
Standard error of measurement A statistic used to estimate the extent to which an observed score deviates from a true score
Standard error of estimate In regression, an estimate of the degree of error involved in predicting the value of one variable from another
Standard error of the mean A measure of sampling error
Standard error of the di�erence A statistic used to estimate how large a di�erence between two scores should be before the di�erence is considered statistically signi�cant
We conclude by encouraging the exercise of critical thinking upon encountering the word standard. The next time you encounter the word standard in any context, give some thought to how standard that “standard” really is.
Certainly with regard to this word’s use in the context of psychological testing and assessment, what is presented as “standard” usually turns out to be not as standard as we might expect.
Sampling�In the process of developing a test, a test developer has targeted some defined group as the population for which the test is designed. This population is the complete universe or set of individuals with at least one common, observable characteristic. The common observable characteristic(s) could be just about anything. For example, it might be high-school seniors who aspire to go to college, or the 16 boys and girls in Ms. Perez’s day-care center, or all athletes who have run a marathon.
To obtain a distribution of scores, the test developer could have the test administered to every person in the targeted population. If the total targeted population consists of something like the 16 boys and girls in Ms. Perez’s day-care center, it may well be feasible to administer the test to each member of the targeted population. However, for tests developed to be used with large or wide-ranging populations, it is usually impossible, impractical, or simply too expensive to administer the test to everyone, nor is it necessary.
The test developer can obtain a distribution of test responses by administering the test to a sample of the population—a portion of the universe of people deemed to be representative of the whole population. The size of the sample could be as small as one person, though samples that approach the size of the population reduce the possible sources of error due to insufficient sample size. The process of selecting the portion of the universe deemed to be representative of the whole population is referred to as sampling.
Subgroups within a defined population may differ with respect to some characteristics, and it is sometimes essential to have these differences proportionately represented in the sample. Thus, for example, if you devised a public opinion test and wanted to sample the
coh37025_ch04_129-156.indd 143 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
opinions of Manhattan residents with this instrument, it would be desirable to include in your sample people representing different subgroups (or strata) of the population, such as Blacks,
whites, Asians, other non-whites, males, females, non-binary persons, the poor, the middle class, the rich, professional people, business people, office workers, skilled and unskilled laborers, the unemployed, homemakers, Catholics, Jews, members of other religions, and so forth—all in proportion to the current occurrence of these strata in the population of people who reside on the
island of Manhattan. Such sampling, termed stratified sampling, would help prevent sampling bias and ultimately aid in the interpretation of the findings. If such sampling were random (or, if every member of the population had the same chance of being included in the sample), then the procedure would be termed stratified-random sampling.
Two other types of sampling procedures are purposive sampling and incidental sampling. If we arbitrarily select some sample because we believe it to be representative of the population, then we have selected what is referred to as a purposive sample. Manufacturers of products frequently use purposive sampling when they test the appeal of a new product in one city or market and then make assumptions about how that product would sell nationally. For example, the manufacturer might test a product in a market such as Cleveland because, on the basis of experience with this particular product, “how goes Cleveland, so goes the nation.” The danger in using such a purposive sample is that the sample, in this case Cleveland residents, may no longer be representative of the nation. Alternatively, this sample may simply not be�representative of national preferences with regard to the particular product being test-marketed.
Often a test user’s decisions regarding sampling wind up pitting what is ideal against what is practical. It may be ideal, for example, to use 50 chief executive officers from any of the Fortune 500 companies (or, the top 500 companies in terms of income) as a sample in an experiment. However, conditions may dictate that it is practical for the experimenter only to use 50 volunteers recruited from the local Chamber of Commerce. This important distinction between what is ideal and what is practical in sampling brings us to a discussion of what has been referred to variously as an incidental sample or a convenience sample.
Ever hear the old joke about a drunk searching for money he lost under the lamppost? He may not have lost his money there, but that is where the light is. Like the drunk searching for money under the lamppost, a researcher may sometimes employ a sample that is not necessarily the most appropriate but is simply the most convenient. Unlike the drunk, the researcher employing this type of sample is doing so not as a result of poor judgment but because of budgetary limitations or other constraints. An incidental sample or convenience sample is one that is convenient or available for use. You may have been a party to incidental sampling if you have ever been placed in a subject pool for experimentation with introductory psychology students. It’s not that the students in such subject pools are necessarily the most appropriate subjects for the experiments, it’s just that they are the most available. Generalization of findings from incidental samples must be made with caution.
If incidental or convenience samples were clubs, they would not be considered exclusive clubs. By contrast, there are many samples that are exclusive, in a sense, because they contain many exclusionary criteria. Consider, for example, the group of children and adolescents who served as the normative sample for one well-known children’s intelligence test. The sample was selected to reflect key demographic variables representative of the U.S. population according to the latest available census data. Still, some groups were deliberately excluded from participation. Who?
� Persons tested on any intelligence measure in the six months prior to the testing � Persons not fluent in English or who are primarily nonverbal � Persons with uncorrected visual impairment or hearing loss
J U S T T H I N K . � . � .
Truly random sampling is relatively rare. Why do you think this is so?
coh37025_ch04_129-156.indd 144 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
� Persons with upper-extremity disability that affects motor performance � Persons currently admitted to a hospital or mental or psychiatric facility � Persons currently taking medication that might depress test performance � Persons previously diagnosed with any physical condition or illness that might depress
test performance (such as stroke, epilepsy, or meningitis)
Our general description of the norming process for a standardized test continues in what follows and, to varying degrees, in subsequent chapters. A highly recommended way to supplement this study and gain a great deal of firsthand knowledge about norms for intelligence tests, personality tests, and other tests is to peruse the technical manuals of major standardized instruments. By going to the library and consulting a few of these manuals, you will discover not only the “real life” way that normative samples are described but also the many varied ways that normative data can be presented.
Developing norms for a standardized test� Having obtained a sample, the test developer administers the test according to the standard set of instructions that will be used with the test. The test developer also describes the recommended setting for giving the test. This instruction may be as simple as making sure that the room is quiet and well lit or as complex as providing a specific set of toys to test an infant’s cognitive skills. Establishing a standard set of instructions and conditions under which the test is given makes the test scores of the normative sample more comparable with the scores of future testtakers. For example, if a test of concentration ability is given to a normative sample in the summer with the windows open near people mowing the grass and arguing about whether the hedges need trimming, then the normative sample probably won’t concentrate well. If a testtaker then completes the concentration test under quiet, comfortable conditions, that person may well do much better than the normative group, resulting in a high standard score. That high score would not be helpful in understanding the testtaker’s concentration abilities because it would reflect the differing conditions under which the tests were taken. This example illustrates how important it is that the normative sample take the test under a standard set of conditions, which are then replicated (to the extent possible) on each occasion the test is administered.
After all the test data have been collected and analyzed, the test developer will summarize the data using descriptive statistics, including measures of central tendency and variability. In addition, it is incumbent on the test developer to provide a precise description of the standardization sample itself. Good practice dictates that the norms be developed with data derived from a group of people who are presumed to be representative of the people who will take the test in the future. After all, if the normative group is different from future testtakers, the basis for comparison becomes questionable at best. In order to best assist future users of the test, test developers are encouraged to “provide information to support recommended interpretations of the results, including the nature of the content, norms or comparison groups, and other technical evidence” (Code of Fair Testing Practices in Education, 2004, p. 4).
In practice, descriptions of normative samples vary widely in detail. Test authors wish to present their tests in the most favorable light possible. Shortcomings in the standardization procedure or elsewhere in the process of the test’s development therefore may be given short shrift or totally overlooked in a test’s manual. Sometimes, although the sample is scrupulously defined, the generalizability of the norms to a particular group or individual is questionable. For example, a test carefully normed on school-age children who reside within the Los Angeles school district may be relevant only to a lesser degree to school-age children who reside within the Dubuque, Iowa, school district. How many children in the standardization sample were English speaking? How many were of Hispanic origin? How does the elementary school
J U S T T H I N K . � . � .
Why do you think each of these groups of people were excluded from the standardization sample of a nationally standardized intelligence test?
coh37025_ch04_129-156.indd 145 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
curriculum in Los Angeles differ from the curriculum in Dubuque? These are the types of questions that must be raised before the Los Angeles norms are judged to be generalizable to the children of Dubuque. Test manuals sometimes supply prospective test users with guidelines for establishing local norms (discussed shortly), one of many different ways norms can be categorized.
One note on terminology is in order before moving on. When the people in the normative sample are the same people on whom the test was standardized, the phrases normative sample and standardization sample are often used interchangeably. Increasingly, however, new norms for standardized tests for specific groups of testtakers are developed some time after the original standardization. That is, the test remains standardized based on data from the original standardization sample; it’s just that new normative data are developed based on an administration of the test to a new normative sample. Included in this new normative sample may be groups of people who were underrepresented in the original standardization sample data. For example, with the changing demographics of a state such as California, and the increasing numbers of people identified as “Hispanic” in that state, an updated normative sample for a California-statewide test might well include a higher proportion of individuals of Hispanic origin. In such a scenario, the normative sample for the new norms clearly would not be identical to the standardization sample, so it would be inaccurate to use the terms standardization sample and normative sample interchangeably.
Types of Norms
Some of the many different ways we can classify norms are as follows: age norms, grade norms, national norms, national anchor norms, local norms, norms from a fixed reference group, subgroup norms, and percentile norms. Percentile norms are the raw data from a test’s standardization sample converted to percentile form. To better understand them, let’s backtrack for a moment and review what is meant by percentiles.
Percentiles� In our discussion of the median, we saw that a distribution could be divided into quartiles where the median was the second quartile (Q2), the point at or below which 50% of the scores fell and above which the remaining 50% fell. Instead of dividing a distribution of scores into quartiles, we might wish to divide the distribution into deciles, or 10 equal parts. Alternatively, we could divide a distribution into 100 equal parts—100 percentiles. In such a distribution, the xth percentile is equal to the score at or below which x% of scores fall. Thus, the 15th percentile is the score at or below which 15% of the scores in the distribution fall. The 99th percentile is the score at or below which 99% of the scores in the distribution fall. If 99% of a particular standardization sample answered fewer than 47 questions on a test correctly, then we could say that a raw score of 47 corresponds to the 99th percentile on this test. It can be seen that a percentile is a ranking that conveys information about the relative position of a score within a distribution of scores. More formally defined, a percentile is an expression of the percentage of people whose score on a test or measure falls below a particular raw score.
Intimately related to the concept of a percentile as a description of performance on a test is the concept of percentage correct. Note that percentile and percentage correct are not synonymous. A percentile is a converted score that refers to a percentage of testtakers. Percentage correct refers to the distribution of raw scores—more specifically, to the number of items that were answered correctly multiplied by 100 and divided by the total number of items.
Because percentiles are easily calculated, they are a popular way of organizing all test-related data, including standardization sample data. Additionally, they lend themselves to use with a wide range of tests. Of course, every rose has its thorns. A problem with using percentiles with normally distributed scores is that real differences between raw scores may be minimized near the ends of the distribution and exaggerated in the middle of the distribution. This distortion may be even worse with highly skewed data. In the normal distribution, the highest frequency of raw scores occurs in the middle. That being the case, the differences between all those scores that cluster in the middle might be quite small, yet even the smallest differences
coh37025_ch04_129-156.indd 146 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
will appear as differences in percentiles. The reverse is true at the extremes of the distributions, where differences between raw scores may be great, though we would have no way of knowing that from the relatively small differences in percentiles.
Age norms�Also known as age-equivalent scores, age norms indicate the average performance of different samples of testtakers who were at various ages at the time the test was administered. If the measurement under consideration is height in inches, for example, then we know that scores (heights) for children will gradually increase at various rates as a function of age up to the middle to late teens. With the graying of America, there has been increased interest in performance on various types of psychological tests, particularly neuropsychological tests, as a function of advancing age.
Carefully constructed age norm tables for physical characteristics such as height enjoy widespread acceptance and are virtually noncontroversial. This is not the case, however, with respect to age norm tables for psychological characteristics such as intelligence. Ever since the introduction of the Stanford-Binet to this country in the early twentieth century, the idea of identifying the “mental age” of a testtaker has had great intuitive appeal. The child of any chronological age whose performance on a valid test of intellectual ability indicated that the child had intellectual ability similar to that of the average child of some other age was said to have the mental age of the norm group in which the child’s test score fell. The reasoning here was that, irrespective of chronological age, children with the same mental age could be expected to read the same level of material, solve the same kinds of math problems, reason with a similar level of judgment, and so forth.
Increasing sophistication about the limitations of the mental age concept has prompted assessment professionals to be hesitant about describing results in terms of mental age. The problem is that “mental age” as a way to report test results is too broad and too inappropriately generalized. To understand why, consider the case of a 6-year-old who, according to the tasks sampled on an intelligence test, performs intellectually like a 12-year-old. Regardless, the 6-year-old is likely not to be similar at all to the average 12-year-old socially, psychologically, and in many other key respects. Beyond such obvious faults in mental age analogies, the mental age concept has also been criticized on technical grounds.3
Grade norms� Designed to indicate the average test performance of testtakers in a given school grade, grade norms are developed by administering the test to representative samples of children over a range of consecutive grade levels (such as first through sixth grades). Next, the mean or median score for children at each grade level is calculated. Because the school year typically runs from September to June—10 months—fractions in the mean or median are easily expressed as decimals. Thus, for example, a sixth-grader performing exactly at the average on a grade-normed test administered during the fourth month of the school year (December) would achieve a grade-equivalent score of 6.4. Like age norms, grade norms have great intuitive appeal. Children learn and develop at varying rates but in ways that are in some aspects predictable. Perhaps because of this fact, grade norms have widespread application, especially to children of elementary school age.
Now consider the case of a student in 12th grade who scores “6” on a grade-normed spelling test. Does this mean that the student has the same spelling abilities as the average sixth-grader? The answer is no. What this finding means is that the student and a hypothetical, average sixth-grader answered the same fraction of items correctly on that test. Grade norms do not provide information
3. For many years, IQ (intelligence quotient) scores on tests such as the Stanford-Binet were calculated by dividing mental age (as indicated by the test) by chronological age. The quotient would then be multiplied by 100 to eliminate the fraction. The distribution of IQ scores had a mean set at 100 and a standard deviation of approximately 16. A child of 12 with a mental age of 12 had an IQ of 100 (12/12 × 100 = 100). The technical problem here is that IQ standard deviations were not constant with age. At one age, an IQ of 116 might be indicative of performance at 1 standard deviation above the mean, whereas at another age an IQ of 121 might be indicative of performance at 1 standard deviation above the mean.
coh37025_ch04_129-156.indd 147 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
as to the content or type of items that a student could or could not answer correctly. Perhaps the primary use of grade norms is as a convenient, readily understandable gauge of how one student’s performance compares with that of fellow students in the same grade.
One drawback of grade norms is that they are useful only with respect to years and months of schooling completed. They have little or no applicability to children who are not yet in school or to children who are out of school. Further, they are not typically designed for use with adults who have returned to school. Both
grade norms and age norms are referred to more generally as developmental norms, a term applied broadly to norms developed on the basis of any trait, ability, skill, or other characteristic that is presumed to develop, deteriorate, or otherwise be affected by chronological age, school grade, or stage of life.
National norms�As the name implies, national norms are derived from a normative sample that was nationally representative of the population at the time the norming study was conducted. In the fields of psychology and education, for example, national norms may be obtained by testing large numbers of people representative of different variables of interest such as age, gender, racial/ethnic background, socioeconomic strata, geographical location (such as North, East, South, West, Midwest), and different types of communities within the various parts of the country (such as rural, urban, suburban).
If the test were designed for use in the schools, norms might be obtained for students in every grade to which the test aimed to be applicable. Factors related to the representativeness of the school from which members of the norming sample were drawn might also be criteria for inclusion in or exclusion from the sample. For example, is the school the student attends publicly funded, privately funded, religiously oriented, military, or something else? How representative are the pupil/teacher ratios in the school under consideration? Does the school have a library, and if so, how many books are in it? These are only a sample of the types of questions that could be raised in assembling a normative sample to be used in the establishment of national norms. The precise nature of the questions raised when developing national norms will depend on whom the test is designed for and what the test is designed to do.
Norms from many different tests may all claim to have nationally representative samples. Still, close scrutiny of the description of the sample employed may reveal that the sample differs in many important respects from similar tests also claiming to be based on a nationally representative sample. For this reason, it is always a good idea to check the manual of the tests under consideration to see exactly how comparable the tests are. Two important questions that test users must raise as consumers of test-related information are “What are the differences between the tests I am considering for use in terms of their normative samples?” and “How comparable are these normative samples to the sample of testtakers with whom I will be using the test?”
National anchor norms�Even the most casual survey of catalogues from various test publishers will reveal that, with respect to almost any human characteristic or ability, there exist many different tests purporting to measure the characteristic or ability. Dozens of tests, for example, purport to measure reading. Suppose we select a reading test designed for use in grades 3 to 6, which, for the purposes of this hypothetical example, we call the Best Reading Test (BRT). Suppose further that we want to compare findings obtained on another national reading test designed for use with grades 3 to 6, the hypothetical XYZ Reading Test, with the BRT. An equivalency table for scores on the two tests, or national anchor norms, could provide the tool for such a comparison. Just as an anchor provides some stability to a vessel, so national anchor norms provide some stability to test scores by anchoring them to other test scores.
The method by which such equivalency tables or national anchor norms are established typically begins with the computation of percentile norms for each of the tests to be compared.
J U S T T H I N K . � . � .
Some experts in testing have called for a moratorium on the use of grade-equivalent as well as age-equivalent scores because such scores may so easily be misinterpreted. What is your opinion on this issue?
coh37025_ch04_129-156.indd 148 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
Using the equipercentile method, the equivalency of scores on different tests is calculated with reference to corresponding percentile scores. Thus, if the 96th percentile corresponds to a score of 69 on the BRT and if the 96th percentile corresponds to a score of 14 on the XYZ, then we can say that a BRT score of 69 is equivalent to an XYZ score of 14. We should note that the national anchor norms for our hypothetical BRT and XYZ tests must have been obtained on the same sample—each member of the sample took both tests, and the equivalency tables were then calculated on the basis of these data.4 Although national anchor norms provide an indication of the equivalency of scores on various tests, technical considerations entail that it would be a mistake to treat these equivalencies as precise equalities (Angoff, 1964, 1966, 1971).
Subgroup norms�A normative sample can be segmented by any of the criteria initially used in selecting subjects for the sample. What results from such segmentation are more narrowly defined subgroup norms. Thus, for example, suppose criteria used in selecting children for inclusion in the XYZ Reading Test normative sample were age, educational level, socioeconomic level, geographic region, community type, and handedness (whether the child was right-handed or left-handed). The test manual or a supplement to it might report normative information by each of these subgroups. A community school board member might find the regional norms to be most useful, whereas a psychologist doing exploratory research in the area of brain lateralization and reading scores might find the handedness norms most useful.
Local norms� Typically developed by test users themselves, local norms provide normative information with respect to the local population’s performance on some test. A local company personnel director might find some nationally standardized test useful in making selection decisions but might deem the norms published in the test manual to be far afield of local job applicants’ score distributions. Individual high schools may wish to develop their own school norms (local norms) for student scores on an examination that is administered statewide. A school guidance center may find that locally derived norms for a particular test—say, a survey of personal values— are more useful in counseling students than the national norms printed in the manual. Some test users use abbreviated forms of existing tests, which requires new norms. Some test users substitute one subtest for another within a larger test, thus creating the need for new norms. There are many different scenarios that would lead the prudent test user to develop local norms.
Fixed Reference Group Scoring Systems
Norms provide a context for interpreting the meaning of a test score. Another type of aid in providing a context for interpretation is termed a fixed reference group scoring system. Here, the distribution of scores obtained on the test from one group of testtakers—referred to as the fixed reference group— is used as the basis for the calculation of test scores for future administrations of the test. Perhaps the test most familiar to college students that has historically exemplified the use of a fixed reference group scoring system is the SAT. This test was first administered in 1926. Its norms were then based on the mean and standard deviation of the people who took the test at the time. With passing years, more colleges became members of the College Board, the sponsoring organization for the test. It soon became evident that SAT scores tended to vary somewhat as a function of the time of year the test was administered. In an effort to ensure perpetual comparability and continuity of scores, a fixed reference group scoring system was put into place in 1941. The distribution of scores from the 11,000 people who took the SAT in 1941 was immortalized as a standard to be used in the conversion of raw scores on future administrations of the test.5 A new fixed reference group, which consisted of
4. When two tests are normed from the same sample, the norming process is referred to as co-norming.
5. Conceptually, the idea of a fixed reference group is analogous to the idea of a fixed reference foot, the foot of the English king that also became immortalized as a measurement standard (Angoff, 1962).
coh37025_ch04_129-156.indd 149 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
the more than 2�million testtakers who completed the SAT in 1990, began to be used in 1995. A score of 500� on the SAT corresponds to the mean obtained by the 1990 sample, a score of 400 corresponds to a score that is 1 standard deviation below the 1990 mean, and so forth. As an example, suppose John took the SAT in 1995 and answered 50 items correctly on a particular scale. And let’s say Mary took the test in 2021 and, just like John, answered 50 items correctly. Although John and Mary may have achieved the same raw score, they would not necessarily achieve the same scaled score. If, for example, the 2021 version of the test was judged to be somewhat easier than the 1995 version, then scaled scores for the 2021 testtakers would be calibrated downward. This would be done so as to make scores earned in 2021 comparable to scores earned in 1995.
Test items common to each new version of the SAT and each previous version of it are employed in a procedure (termed anchoring) that permits the conversion of raw scores on the new version of the test into fixed reference group scores. Like other fixed reference group scores, including Graduate Record Examination scores, SAT scores are most typically interpreted by local decision-making bodies with respect to local norms. Thus, for example, college admissions officers usually rely on their own independently collected norms to make selection decisions. They will typically compare applicants’ SAT scores to the SAT scores of students in their school who completed or failed to complete their program. Of course, admissions decisions are seldom made on the basis of the SAT (or any other single test) alone. Various criteria are typically evaluated in admissions decisions.
Norm-Referenced versus Criterion-Referenced Evaluation
One way to derive meaning from a test score is to evaluate the test score in relation to other scores on the same test. As we have pointed out, this approach to evaluation is referred to as norm-referenced. Another way to derive meaning from a test score is to evaluate it on the basis of whether some criterion has been met. We may define a criterion as a standard on which a judgment or decision may be based.
Criterion-referenced testing and assessment may be defined as a method of evaluation and a way of deriving meaning from test scores by evaluating an individual’s score with reference to a set standard. Some examples: � To be eligible for a high-school diploma, students must demonstrate at least a sixth-grade
reading level. � To earn the privilege of driving an automobile, would-be drivers must take a road test
and demonstrate their driving skill to the satisfaction of a state-appointed examiner. � To be licensed as a psychologist, the applicant must achieve a score that meets or
exceeds the score mandated by the state on the licensing test. � To conduct research using human subjects, many universities and other organizations require
researchers to successfully complete an online course that presents testtakers with ethics- oriented information in a series of modules, followed by a set of forced-choice questions.
The criterion in criterion-referenced assessments typically derives from the values or standards of an individual or organization. For example, in order to earn a black belt in karate, students must
demonstrate a black-belt level of proficiency in karate and meet related criteria such as demonstrating self-discipline and focus. Each student is evaluated individually to see if all of these criteria are met. Regardless of the level of performance of all testtakers, only students who meet all criteria will leave the dojo (training room) with a brand-new black belt.
Criterion-referenced testing and assessment goes by other names. Because the focus in the criterion-referenced approach is on how scores relate to a particular content area or domain, the approach has also been referred to as domain- or content-referenced
J U S T T H I N K . � . � .
List other examples of a criterion that must be met in order to gain privileges or access of some sort.
coh37025_ch04_129-156.indd 150 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
testing and assessment.6 One way of conceptualizing the difference between norm-referenced and criterion-referenced approaches to assessment has to do with the area of focus regarding test results. In norm-referenced interpretations of test data, a usual area of focus is how an individual performed relative to other people who took the test. In criterion-referenced interpretations of test data, a usual area of focus is the testtaker’s performance: what the testtaker can or cannot do; what the testtaker has or has not learned; whether the testtaker does or does not meet specified criteria for inclusion in some group, access to certain privileges, and so forth. Because criterion-referenced tests are frequently used to gauge achievement or mastery, they are sometimes referred to as mastery tests. The criterion-referenced approach has enjoyed widespread acceptance in the field of computer- assisted education programs. In such programs, mastery of segments of materials is assessed before the program user can proceed to the next level.
“Has this flight trainee mastered the material she needs to be an airline pilot?” This question is an example of what an airline personnel office might seek to address with a mastery test on a flight simulator. If a standard, or criterion, for passing a hypothetical “Airline Pilot Test” (APT) has been set at 85% correct, then trainees who score 84% correct or less will not pass. It matters not whether they scored 84% or 42%. Conversely, trainees who score 85% or better on the test will pass whether they scored 85% or 100%. All who score 85% or better are said to have mastered the skills and knowledge necessary to be an airline pilot. Taking this example one step further, another airline might find it useful to set up three categories of findings based on criterion-referenced interpretation of test scores:
85% or better correct = pass
75% to 84% correct = retest after a two-month refresher course
74% or less = fail
How should cut scores in mastery testing be determined? How many and what kinds of test items are needed to demonstrate mastery in a given field? The answers to these and related questions have been tackled in diverse ways (Cizek & Bunch, 2007; Ferguson & Novick, 1973; Geisenger & McCormick, 2010; Glaser & Nitko, 1971; Panell & Laabs, 1979).
Critics of the criterion-referenced approach argue that if it is strictly followed, potentially important information about an individual’s performance relative to other testtakers is lost. Another criticism is that although this approach may have value with respect to the assessment of mastery of basic knowledge, skills, or both, it has little or no meaningful application at the upper end of the knowledge/skill continuum. Thus, the approach is clearly meaningful in evaluating whether pupils have mastered basic reading, writing, and arithmetic. But how useful is it in evaluating doctoral-level writing or math? Identifying stand-alone originality or brilliant analytic ability is not the stuff of which criterion- oriented tests are made. By contrast, brilliance and superior abilities are recognizable in tests that employ norm-referenced interpretations. They are the scores that trail off all the way to the right on the normal curve, past the third standard deviation.
Norm-referenced and criterion-referenced are two of many ways that test data may be viewed and interpreted. However, these terms are not mutually exclusive, and the use of one approach with a set of test data does not necessarily preclude the use of the other approach for another application.
6. Although acknowledging that content-referenced interpretations can be referred to as criterion-referenced interpretations, the 1974 edition of the Standards for Educational and Psychological Testing also noted a technical distinction between interpretations so designated: “Content-referenced interpretations are those where the score is directly interpreted in terms of performance at each point on the achievement continuum being measured. Criterion-referenced interpretations are those where the score is directly interpreted in terms of performance at any given point on the continuum of an external variable. An external criterion variable might be grade averages or levels of job performance” (p. 19; footnote in original omitted).
J U S T T H I N K . � . � .
For licensing of physicians, psychologists, engineers, and other professionals, would you advocate that your state use criterion- or norm-referenced assessment? Why?
coh37025_ch04_129-156.indd 151 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
In a sense, all testing is ultimately normative, even if the scores are as seemingly criterion-referenced as pass–fail, because even in a pass–fail score there is an inherent acknowledgment of a continuum of abilities. At some point in that continuum, a dichotomizing cutoff point has been applied. We should also make the point that some so-called norm-referenced assessments are made with subject samples wherein “the norm is hardly the norm.” In a similar vein, when dealing with special or extraordinary populations, the criterion level that is set by a test may also be “far from the norm” in the sense of being average with regard to the general population. To get a sense of what we mean by such statements, think of the norm for everyday skills related to playing basketball, and then imagine how those norms might be with a subject sample limited exclusively to players on NBA teams. Now, meet two sports psychologists who have worked in a professional assessment capacity with the Chicago Bulls in this chapter’s Meet an Assessment Professional.
M E E T A N A S S E S S M E N T P R O F E S S I O N A L
Meet Dr. Steve Julius and Dr. Howard W. Atlas
he Chicago Bulls of the ����s is considered one of the great dynasties in sports, as witnessed by their six world championships in that decade. . . .
The team bene�ted from great individual contributors, but like all successful organizations, the�Bulls were always on the lookout for ways to maintain a competitive edge. The Bulls . . . were one of the �rst NBA franchises to apply personality testing and behavioral interviewing to aid in the selection of college players during the annual draft, as well as in the evaluation of goodness-of-�t when considering the addition of free agents. The purpose of this e�ort was not to rule out psychopathology, but rather to evaluate a range of competencies (e.g., resilience, relationship to authority, team orientation) that were deemed necessary for success in the league, in general, and the Chicago Bulls, in particular.
The team utilized commonly used and well- validated personality assessment tools and techniques from the world of business (e.g., ��PF–�fth edition). . . . Eventually, su�cient data were collected to allow for the validation of a regression formula, useful as a prediction tool in its own right. In addition to selection, the information collected on the athletes often is used to assist the coaching sta� in their e�orts to motivate and instruct players, as well as to�create an atmosphere of collaboration.
Used with permission of Steve Julius and Howard W. Atlas.
T
Steve Julius, Ph.D., Sports Psychologist, Chicago Bulls Steve Julius
Howard W. Atlas, Ed.D., Sports Psychologist, Chicago Bulls Howard W. Atlas
coh37025_ch04_129-156.indd 152 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
Culture and Inference
Along with statistical tools designed to help ensure that prediction and inferences from measurement are reasonable, there are other considerations. It is incumbent upon responsible test users not to lose sight of culture as a factor in test administration, scoring, and interpretation. In selecting a test for use, the responsible test user does some advance research on the test’s available norms to determine how appropriate they are for use with the targeted testtaker population. In interpreting data from psychological tests, it is frequently helpful to know about the culture of the testtaker, including something about the era or “times” that the testtaker experienced. In this regard, think of the words of the famous anthropologist Margaret Mead (1978, p. 71), who, in recalling her youth, wrote: “We grew up under skies which no satellite had flashed.” In interpreting assessment data from assessees of different generations, it would seem useful to keep in mind whether “satellites had or had not flashed in the sky.” In other words, historical context should be taken into consideration in evaluation (Rogler, 2002).
It seems appropriate to conclude a chapter entitled “Of Tests and Testing” with the introduction of the term culturally informed assessment and with some guidelines for accomplishing it (Table 4–1). Think of these guidelines as a list of themes that may be repeated in different ways as you continue to learn about the assessment enterprise. To supplement this list, see the updated guidelines published by the American Psychological Association (2017). For now, let’s continue to build a sound foundation in testing and assessment with a discussion of the psychometric concept of reliability in Chapter 5.
J U S T T H I N K . � . � .
What event in recent history may have relevance when interpreting data from a psychological assessment?
Table �–� Culturally Informed Assessment: Some “Do’s” and “Don’ts”
Do Do Not
Be aware of the cultural assumptions on which a test is basedTake for granted that a test is based on assumptions that impact all groups in much the same way
Consider consulting with members of particular cultural communities regarding the appropriateness of particular assessment techniques, tests, or test items
Take for granted that members of all cultural communities will automatically deem particular techniques, tests, or test items appropriate for use
Strive to incorporate assessment methods that complement the worldview and lifestyle of assessees who come from a speci�c cultural and linguistic population
Take a “one-size-�ts-all” view of assessment when it comes to evaluation of persons from various cultural and linguistic populations
Be knowledgeable about the many alternative tests or measurement procedures that may be used to ful�ll the assessment objectives
Select tests or other tools of assessment with little or no regard for the extent to which such tools are appropriate for use with a particular assessee.
Be aware of equivalence issues across cultures, including equivalence of language used and the constructs measured
Simply assume that a test that has been translated into another language is automatically equivalent in every way to the original
Score, interpret, and analyze assessment data in its cultural context with due consideration of cultural hypotheses as possible explanations for �ndings
Score, interpret, and analyze assessment in a cultural vacuum
coh37025_ch04_129-156.indd 153 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
Self-Assessment
Test your understanding of elements of this chapter by seeing if you can explain each of the following terms, expressions, and abbreviations:
age-equivalent scores age norms construct content-referenced testing and
assessment convenience sample criterion criterion-referenced testing and
assessment cumulative scoring developmental norms domain-referenced testing and
assessment domain sampling equipercentile method
error variance fixed reference group scoring system grade norms incidental sample local norms national anchor norms national norms norm normative sample norming norm-referenced testing and
assessment overt behavior percentage correct percentile
program norms purposive sampling race norming sample sampling standardization standardized test state stratified-random sampling stratified sampling subgroup norms test standardization trait true score theory user norms
coh37025_ch04_129-156.indd 154 12/01/21 4:07 PM
Chapter 4: Of Tests and Testing ���
References
Allport, G. W. (1937). Personality: A psychological interpretation. Holt.
Allport, G. W., & Odbert, H. S. (1936). Trait-names: A psycho-lexical study. Psychological Monographs, 47 (Whole No. 211).
American Psychological Association. (2017). Multicultural guidelines: An ecological approach to context, identity, and intersectionality. http://www.apa .org/about/policy/multicultural-guidelines.pdf
American Psychological Association, American Educational Research Association, & National Council on Measurement in Education. (1974). Standards for educational and psychological tests. Author.
Anderson, D. (2007). A reciprocal determinism analysis of the relationship between naturalistic media usage and the development of creative-thinking skills among college students. Dissertation Abstracts International Section A: Humanities and Social Sciences, 67(7-A), 2007, 2459.
Angoff, W. H. (1962). Scales with nonmeaningful origins and units of measurement. Educational and Psychological Measurement, 22, 27–34.
Angoff, W. H. (1964). Technical problems of obtaining equivalent scores on tests. Educational and Psychological Measurement, 1,�11–13.
Angoff, W. H. (1966). Can useful general-purpose equivalency tables be prepared for different college admissions tests? In A. Anastasi (Ed.), Testing problems in perspective (pp. 251–264). American Council on Education.
Angoff, W. H. (1971). Scales, norms, and equivalent scores. In R. L. Thorndike (Ed.), Educational measurement (2nd ed., pp. 508–560). American Council on Education.
Aros, S., Mills, J. L., Torres, C., et al. (2006). Prospective identification of pregnant women drinking four or more standard drinks (= 48 g) of alcohol per day. Substance Use & Misuse, 41(2), 183–197.
Chaplin, W. F., John, O. P., & Goldberg, L. R. (1988). Conceptions of state and traits: Dimensional attributes with ideals as prototypes. Journal of Personality and Social Psychology, 54, 541–557.
Cizek, G. J., & Bunch, M. B. (2007). Standard setting: A guide to establishing and evaluating performance standards on tests. Sage.
Code of Fair Testing Practices in Education. (2004). Joint Committee on Testing Practices.
Damian, R. I., Spengler, M., Sutu, A., & Roberts, B.W. (2019). Sixteen going on sixty-six: A longitudinal study of personality stability and change across 50 years. Journal of Personality and Social Psychology, 117(3), 674–695. https://doi .org/10.1037/pspp0000210
De Los Reyes, A., Augenstein, T. M., Wang, M., Thomas, S. A., Drabick, D. A. G., Burgers, D. E., & Rabinowitz, J. (2015). The validity of the multi- informant approach to assessing child and adolescent mental health. Psychological Bulletin, 141(4), 858–900. https://doi.org/10.1037/a0038498
Farrell, S. F. (1998). Alcohol dependence and the price of alcoholic beverages. Dissertation Abstracts International: Section B: The Sciences and Engineering, 59(4-B), 1606.
Ferguson, R. L., & Novick, M. R. (1973). Implementation of a Bayesian system for decision analysis in a program of individually prescribed instruction. ACT Research Report, No.�60.
Geisenger, K. F., & McCormick, C. M. (2010). Adopting cut scores: Post-standard-setting panel considerations for decision makers. Educational Measurement: Issues and Practice, 29, 38–44.
Gill, J. S., Donaghy, M., Guise, J., & Warner, P. (2007). Descriptors and accounts of alcohol consumption: Methodological issues piloted with female undergraduate drinkers in Scotland. Health Education Research, 22(1), 27–36.
Glaser, R., & Nitko, A. J. (1971). Measurement in learning and instruction. In R. L. Thorndike (Ed.), Educational measurement (2nd ed.). American Council on Education.
Greenlaw, P. S., & Jensen, S. S. (1996). Race-norming and the Civil Rights Act of 1991. Public Personnel Management, 25, 13–24.
Guilford, J. P. (1959). Personality. McGraw-Hill. Holt, R. R. (1971). Assessing personality. Harcourt Brace
Jovanovich. Hwang, S. W. (2006). Homelessness and harm reduction.
Canadian Medical Association Journal, 174(1), 50–51. Lüdtke, O., Trautwein, U. & Husemann, N. (2009).
Goal and personality trait development in a transitional period: Testing principles of personality development. Personality and Social Psychology Bulletin, 35, 428–441.
McCabe, K. O., & Fleeson, W. (2016). Are traits useful? Explaining trait manifestations as tools in the pursuit of goals. Journal of Personality and Social Psychology, 110(2), 287–301.
Mead, M. (1978). Culture and commitment: The new relationship between the generations in the 1970s (rev. ed.). Columbia University Press.
Nelson, L. D. (1994). Introduction to the special section on normative assessment. Psychological Assessment, 4, 283.
Panell, R. C., & Laabs, G. J. (1979). Construction of a criterion-referenced, diagnostic test for an individualized instruction program. Journal of Applied Psychology, 64, 255–261.
Podymow, T., Turnbull, J., Coyle, D., et al. (2006). Shelter-based managed alcohol administration to chronically homeless people addicted to alcohol. Canadian Medical Association Journal, 174(1), 45–49.
Rammstedt, B., Mutz, M., & Farmer, R. F. (2015). The answer is blowing in the wind: Weather effects on personality ratings. European Journal of Psychological Assessment, 31(4), 287–293.
Roberts, B. W., & DelVecchio, W. F. (2000). The rank- order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126, 3–25.
Roberts, B. W., Robins, R. W., Caspi, A., & Trzesniewski., K. (2003). Personality trait development in adulthood. In J. Mortimer & M. Shanahan (Eds.), Handbook of the life course (pp. 579–598). Kluwer Academic.
Roberts, B. W., Walton, K., & Viechtbauer, W. (2006). Personality changes in adulthood: Reply to Costa & McCrae (2006). Psychological Bulletin, 132, 29–32.
coh37025_ch04_129-156.indd 155 12/01/21 4:07 PM
������Part 2: The Science of Psychological Measurement
Robertson, G. J. (1990). A practical model for test development. In C. R. Reynolds & R. W. Kamphaus (Eds.), Handbook of psychological and educational assessment of children: Intelligence & achievement (pp. 62–85). Guilford.
Rogler, L. H. (2002). Historical generations and psychology: The case of the Great Depression and World War II. American Psychologist, 57, 1013–1023.
Sherman, R. A., Rauthmann, J. F., Brown, N. A., et al. (2015). The independent effects of personality and situations on real-time expressions of behavior and emotion. Journal of Personality and Social Psychology, 109(5), 872–888.
Song, Q. C., Wee, S., & Newman, D. A. (2017). Diversity shrinkage: Cross-validating pareto-optimal weights to
enhance diversity via hiring practices. Journal of Applied Psychology, 102(12), 1636–1657. https://doi .org/10.1037/apl0000240
Thorndike, E. L. (1918). The nature, purposes, and general methods of measurements of educational products. In G. M. Whipple (Ed.), Yearbook of the national society for the study of education (pp. 16–24). Public School Publishing.
Yañez, Y. T., & Fremouw, W. (2004). The application of the Daubert standard to parental capacity measures. American Journal of Forensic Psychology, 22(3), 5–28.
Zuckerman, M. (1979). Traits, states, situations, and uncertainty. Journal of Behavioral Assessment, 1, 43–54.
coh37025_ch04_129-156.indd 156 12/01/21 4:07 PM
I-��
A AAS. See Addiction Acknowledgment Scale
(AAS) ABAP. See American Board of Assessment
Psychology (ABAP) ability/aptitude, measures of, 594–596 ABLE. See Adult Basic Learning
Examination (ABLE) ABPP. See American Board of Professional
Psychology (ABPP) absolute cut scores, 244 abstraction ability tests, 567–568 abuse: (1) Infliction of or allowing the
infliction of physical injury or emotional impairment that is nonaccidental; (2) the creation of or allowing the creation of substantial risk of physical injury or emotional impairment that is nonaccidental; (3) the committing of or allowing the committing of a sexual offense against a child; contrast with neglect, 530–535
academic achievement, 206 academic research settings, 24 accommodation: (1) In Piagetian theory,
one of two basic mental operations through which humans learn, this one involving change from what is already known, perceived, or thought to fit with new information (contrast with assimilation); (2) in assessment, the adaptation of a test, procedure, or situation, or the substitution of one test for another in order to make the assessment more suitable for an assessee with exceptional needs; (3) in the workplace, modification of or adjustments to job functions or circumstances, 31, 32–33, 319
Accounting Program Admission Test (APAT), 376
acculturation: The process by which an individual’s thoughts, behaviors, values, identity, and worldview develop in relation to the general thinking, behavior, customs, and values of a particular cultural group, 431–435
achievement batteries, 360, 361 achievement test: Evaluation of
accomplishment or the degree of learning that has taken place, usually with regard to an academic area, 21, 360–363
acquiescent response style, 425 ACT Assessment, 373 actuarial assessment: An approach to
evaluation characterized by the application of empirically demonstrated statistical rules as a determining factor in the assessor’s judgment and actions; contrast with clinical assessment, 538
actuarial prediction: An approach to predicting behavior based on the application of empirically demonstrated statistical rules and probabilities; contrast with clinical prediction and mechanical prediction, 538
adaptive testing: An examination method or procedure characterized by individually tailoring presentation of items to the testtaker; also referred to as tailored testing, sequential testing, branched testing, and response- contingent testing, 321. See also computerized adaptive testing (CAT)
adaptive treatment, 241 Adarand Constructors, Inc. v. Pena et al., 63 addiction, 518–520 Addiction Acknowledgment Scale (AAS),
519 Addiction Potential Scale (APS), 519 Addiction Severity Index (ASI), 519 additional materials stage, 244 ADHD. See attention deficit hyperactivity
disorder (ADHD) Adjective Check List, 409 adjective checklist format, 409 adjustable light-beam apparatus, 28 administration error, 293 administration procedures, 8 ADRESSING: A purposely misspelled
word but easy-to-remember acronym to remind assessors of the following sources of cultural influence: age, disability, religion, ethnicity, social status, sexual orientation, indigenous heritage, national origin, and gender, 515
Adult Basic Learning Examination (ABLE), 362
aesthetic perception, 29 affirmative action: Voluntary and
mandatory efforts undertaken by federal, state, and local governments, private employers, and schools to combat discrimination and to promote equal opportunity in education and employment for all, 59
AFQT. See Armed Forces Qualification Test (AFQT)
AGCT. See Army General Classification Test (AGCT)
age-based scale, 257 age-equivalent scores. See age norms age norms: Norms specifically designed to
compare a testtaker’s score with those of same-age peers; contrast with grade norms, 147
age scale: A test with items organized by the age at which most testtakers are believed capable of responding in the way keyed correct; contrast with point scale, 318
AHPAT. See Allied Health Professions Admission Test (AHPAT)
Airman Qualifying Exam, 329 Albermarle Paper Company v. Moody, 62 alcohol abuse, 518–520 ALI standard: American Law Institute
standard of legal insanity, which provides that a person is not responsible for criminal conduct if, at the time of such conduct, the person lacked substantial capacity either to appreciate the criminality of the conduct or to conform the conduct to the requirements of the law; contrast with the Durham standard and the M’Naghten standard, 523
Allen v. District of Columbia, 62 Allied Health Professions Admission Test
(AHPAT), 376 alternate assessment: An evaluative or
diagnostic procedure or process that varies from the usual, customary, or standardized way a measurement is derived, either by some special accommodation made to the assessee or by alternative methods designed to measure the same variable(s), 31
alternate forms: Different versions of the same test or measure; contrast with parallel forms, 164
alternate-forms reliability: An estimate of the extent to which item sampling and other errors have affected scores on two versions of the same test; contrast with parallel-forms reliability, 164, 176
alternate item: A test item to be administered only under certain conditions to replace the administration of an existing item on the test, 316
Alzheimer’s disease, 23, 583–584
Glossary/Index
coh37025_sndx_i22-i50.indd 22 12/01/21 4:15 PM
Note: This is a brief, decidedly noncomprehensive overview of historical events perceived important by the authors. Consult other authoritative historical sources for more detailed and comprehensive descriptions of these and other events.
2200 �.�.�. Proficiency testing is known to have been conducted in China. The Emperor has public officials evaluated periodically.
1115 �.�.�. Open and competitive civil service examinations in China are common during the Chang Dynasty. Proficiency is tested in areas such as arithmetic, writing, geography, music, agriculture, horsemanship, and cultural rites and ceremonies.
400 �.�.�. Plato suggests that people should work at jobs consistent with their abilities and endowments—a sentiment that will be echoed many times through the ages by psychologists, human resource professionals, and parents.
175 �.�.�. Claudius Galenus (otherwise known as Galen) designs experiments to show that the brain, not the heart, is the seat of the intellect.
200 The so-called Dark Ages begin and society forces science to take a (temporary) backseat to faith and superstition.
1484 Interest in individual differences centers primarily on questions such as “Who is in league with Satan?” and “Are they in voluntary or involuntary league?” The Hammer of Witches is a primitive, diagnostic manual of sorts with tips on interviewing and identifying persons suspected of having strayed from the righteous path.
1550 The Renaissance witnesses a rebirth in philosophy, and German physician Johann Weyer writes that those accused of being witches may have been suffering from mental or physical disorders. For the faithful, Weyer is seen as advancing Satan’s cause.
1600 The pendulum begins to swing away from a religion-dominated view of the world to one that is more philosophical and scientific in nature.
1700 The cause of philosophy and science is advanced with the writings of the French philosopher René Descartes, the German philosopher Gottfried Leibniz, and a group of English philosophers (John Locke, George Berkeley, Dave Hume, and David Hartley) referred to collectively as “the British empiricists.” Descartes, for example, raised intriguing questions regarding the relationship between the mind and the body. These issues would be explored in a less philosophical and more physical way by Pierre Cabanis, a physiologist. For humanitarian purposes, Cabanis personally observed the state of consciousness of guillotine victims of the French Revolution. He concluded that the mind and body were so intimately linked that the guillotine was probably a painless mode of execution.
1734 Christian von Wolff authors two books, Psychologia Empirica [Empirical Psychology] (1732) and Psychologia Rationalis [Rational Psychology] (1734), which anticipate psychology as a science. A student of Gottfried Leibniz, von Wolff also elaborated on Leibniz’s idea that there exist perceptions below the threshold of awareness, thus anticipating Freud’s notion of the unconscious.
1780 Franz Mesmer “mesmerizes” not only Parisian patients but some members of the European medical community with his use of what he once referred to as “animal magnetism” to effect cures. Mesmerism (or hypnosis as we know it today) would go on to become a tool of psychological assessment; the hypnotic interview is one of many alternative techniques for information gathering.
1823 The Journal of Phrenology is founded to further the study of Franz Joseph Gall’s notion that ability and special talents are localized in concentrations of brain fiber that press outward. Extensive experimentation eventually discredits phrenology, and the journal folds by the early twentieth century. By the mid- twentieth century, evaluation of “bumps” in paper profiles would be preferable to examination of bumps on the head for obtaining information about ability and talents.
1829 In Analysis of the Phenomena of the Human Mind, English philosopher James Mill argued that the structure of mental life consists of sensation and ideas. Mill anticipates an approach to experimental psychology called structuralism, the goal of which would be to explore the components of the structure of the mind.
Psychological Testing and Assessment: A Timeline Spanning ���� B.C.E to the Present
T-�
coh37025_timeline_T1-T5.indd 1 11/01/21 7:58 PM
- Cover
- Title
- Copyright
- Contents
- Preface