HELP WITH PSYCHOLOGY CLASS

profileadityabokadia
EssentialsofTestingandAssessmentbyEdwardSNeukrugz-lib.org.pdf

Essentials of Testing and Assessment A Practical Guide to Counselors, Social Workers, and Psychologists

THIRD EDITION

Edward S. Neukrug Old Dominion University

R. Charles Fawcett Director, Region Ten Fluvanna Counseling Center

Australia • Brazil • Mexico • Singapore • United Kingdom • United States

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

This is an electronic version of the print textbook. Due to electronic rights restrictions, some third party content may be suppressed. Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. The publisher reserves the right to remove content from this title at any time if subsequent rights restrictions require it. For valuable information on pricing, previous editions, changes to current editions, and alternate formats, please visit www.cengage.com/highered to search by ISBN#, author, title, or keyword for materials in your areas of interest.

Essentials of Testing and Assessment: A Practical Guide to Counselors, Social Workers, and Psychologists, Third Edition Edward S. Neukrug and R. Charles Fawcett

Product Director: Jon-David Hague

Product Manager: Julie Martinez

Content Coordinator: Sean Cronin

Product Assistant: Kyra Kane

Media Developer: Audrey Espy

Outsource Development Manager: Jeremy Judson

Outsource Development Coordinator: Joshua Taylor

Associate Marketing Manager: Shannon Shelton

Content Project Manager: Matthew Ballantyne

Art and Cover Direction, Production Management, and Composition: PreMediaGlobal

Manufacturing Planner: Judy Inouye

Rights Acquisitions Specialist: Dean Dauphinais

Cover Image: © John Grant/Getty Images

© 2015, 2010, 2006 Cengage Learning

ALL RIGHTS RESERVED. No part of this work covered by the copyright herein may be reproduced, transmitted, stored, or used in any form or by any means graphic, electronic, or mechanical, including but not limited to photocopying, recording, scanning, digitizing, taping, Web distribution, information networks, or information storage and retrieval systems, except as permitted under Section 107 or 108 of the 1976 United States Copyright Act, without the prior written permission of the publisher.

For product information and technology assistance, contact us at Cengage Learning Customer & Sales Support, 1-800-354-9706.

For permission to use material from this text or product, submit all requests online at www.cengage.com/permissions.

Further permissions questions can be e-mailed to [email protected].

Library of Congress Control Number: On File

ISBN-13: 978-1-285-45424-5

ISBN-10: 1-285-45424-3

Cengage Learning 200 First Stamford Place, 4th Floor Stamford, CT 06902 USA

Cengage Learning is a leading provider of customized learning solutions with office locations around the globe, including Singapore, the United Kingdom, Australia, Mexico, Brazil, and Japan. Locate your local office at www.cengage.com/global.

Cengage Learning products are represented in Canada by Nelson Education, Ltd.

To learn more about Cengage Learning Solutions, visit www.cengage.com.

Purchase any of our products at your local college store or at our preferred online store www.cengagebrain.com.

Printed in the United States of America 1 2 3 4 5 6 7 17 16 15 14 13

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

WCN: 02-200-203

To my father and my brother, the real math experts in the family. —Ed Neukrug

To my loving wife Laura, who makes my life richer. —Charlie Fawcett

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Brief Contents

SECTION I UNDERSTANDING THE ASSESSMENT PROCESS: HISTORY, ETHICAL AND PROFESSIONAL ISSUES, DIAGNOSIS, AND THE ASSESSMENT REPORT 1

1 History of Testing and Assessment 3

2 Ethical, Legal, and Professional Issues in Assessment 21

3 Diagnosis in the Assessment Process 43

4 The Assessment Report Process: Interviewing the Client and Writing the Report 59

SECTION II TEST WORTHINESS AND TEST STATISTICS 81

5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 83

6 Statistical Concepts: Making Meaning Out of Raw Scores 110

7 Statistical Concepts: Creating New Scores to Interpret Test Data 127

iv

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SECTION III COMMONLY USED ASSESSMENT TECHNIQUES 151

8 Assessment of Educational Ability: Survey Battery, Diagnostic, Readiness, and Cognitive Ability Tests 157

9 Intellectual and Cognitive Functioning: Intelligence Testing and Neuropsychological Assessment 190

10 Career and Occupational Assessment: Interest Inventories, Multiple Aptitude, and Special Aptitude Tests 221

11 Clinical Assessment: Objective and Projective Personality Tests 247

12 Informal Assessment: Observation, Rating Scales, Classification Methods, Environmental Assessment, Records and Personal Documents, and Performance-Based Assessment 281

BRIEF CONTENTS v

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Contents

PREFACE XV

SECTION I Understanding the Assessment Process: History, Ethical and Professional Issues, Diagnosis, and the Assessment Report 1

1 History of Testing and Assessment 3 Distinguishing Between Testing and Assessment 4

The History of Assessment 5 Ancient History 5 Precursors to Modern-Day Test Development 5 The Emergence of Ability Tests (Testing in the Cognitive Domain) 7

Intellectual and Cognitive Functioning: Intelligence Testing and Neuropsychological Assessment 7

Group Tests of Ability (Group Testing in the Cognitive Domain) 8 The Emergence of Personality Tests (Testing in the Affective Realm) 12

Interest Inventories and Vocational Assessment 12 Objective Personality Assessment 12 Projective Testing 12

The Emergence of Informal Assessment Procedures 13 Modern-Day Use of Assessment Procedures 14

Questions to Consider When Assessing Individuals 17

Summary 17

Chapter Review 18

References 19

vi

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

2 Ethical, Legal, and Professional Issues in Assessment 21 Ethical Issues in Assessment 22

Overview of the Assessment Sections of the ACA and APA Ethical Codes 22 Choosing Appropriate Assessment Instruments 22 Competence in the Use of Tests 23 Confidentiality 23 Cross-Cultural Sensitivity 24 Informed Consent 24 Invasion of Privacy 24 Proper Diagnosis 24 Release of Test Data 24 Test Administration 25 Test Security 25 Test Scoring and Interpretation 25

Standards for Responsible Testing Practices 26 Making Ethical Decisions 27

Legal Issues in Assessment 28 The Family Education Rights and Privacy Act (FERPA) of 1974 29 The Health Insurance Portability and Accountability Act (HIPAA) 29 Privileged Communication 30 The Freedom of Information Act 30 Civil Rights Acts (1964 and Amendments) 30 Americans with Disabilities Act (ADA) (PL 101-336) 32 Individuals with Disabilities Education Act (IDEA) 32 Section 504 of the Rehabilitation Act 33 Carl Perkins Career and Technical Education Improvement Act of 2006 33

Professional Issues 33 Professional Associations 34 Accreditation Standards of Professional Associations 34 Forensic Evaluations 34 Assessment as a Holistic Process 35 Cross-Cultural Issues in Assessment 36 Embracing Testing and Assessment Procedures 38

Summary 38

Chapter Review 39

References 40

3 Diagnosis in the Assessment Process 43 The Importance of Diagnosis 44

The Diagnostic and Statistical Manual (DSM): A Brief History 45

The DSM-5 46 Single-Axis vs. Multiaxial Diagnosis 47 Making and Reporting Diagnosis 47

Ordering Diagnoses 47 Subtypes, Specifiers, and Severity 47 Provisional Diagnosis 48 Other Specified Disorders and Unspecified Disorders 48

Specific Diagnostic Categories 49

CONTENTS vii

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Other Medical Considerations 52 Psychosocial and Environmental Considerations 53 Cultural Considerations 54 Final Thoughts on DSM-5 in the Assessment Process 54

Summary 56

Chapter Review 56

Answers to Exercise 57

References 57

4 The Assessment Report Process: Interviewing the Client and Writing the Report 59 Purpose of the Assessment Report 60

Gathering Information for the Report: Garbage In, Garbage Out 60

Structured, Unstructured, and Semi-Structured Interviews 61 Computer-Driven Assessment 63

Choosing an Appropriate Assessment Instrument 64

Writing the Report 65 Demographic Information 65 Presenting Problem or Reason for Referral 66 Family Background 66 Significant Medical/Counseling History 67 Substance Use and Abuse 67 Educational and Vocational History 68 Other Pertinent Information 68 Mental Status 68

Appearance and Behavior 68 Emotional State 68 Thought Components 70 Cognition 71

Assessment Results 72 Diagnosis 73 Summary and Conclusions 74

Recommendations 75

Summarizing the Writing of an Assessment Report 75

Summary 77

Chapter Review 78

References 78

SECTION II Test Worthiness and Test Statistics 81

5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 83 Correlation Coefficient 84

Coefficient of Determination (Shared Variance) 86

viii CONTENTS

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Validity 87 Content Validity 87 Criterion-Related Validity 89

Concurrent Validity 89 Predictive Validity 89

Construct Validity 91 Experimental Design Validity 91 Factor Analysis 92 Convergence with Other Instruments (Convergent Validity) 92 Discrimination with Other Measures (Discriminant Validity) 92

Visual Representation of Types of Validity 93

Reliability 94 Test-Retest Reliability 95 Alternate, Parallel, or Equivalent Forms Reliability 96 Internal Consistency 96

Split-Half or Odd-Even Reliability 96 Cronbach’s Coefficient Alpha and Kuder–Richardson 97

Visual Representation of Types of Reliability 97 Item Response Theory: Another Way of Looking at Reliability 98

Cross-Cultural Fairness 99

Practicality 102 Time 102 Cost 102 Format 102 Readability 103 Ease of Administration, Scoring, and Interpretation 103

Selecting and Administering a Good Test 103 Step 1: Determine the Goals of Your Client 103 Step 2: Choose Instrument Types to Reach Client Goals 104 Step 3: Access Information About Possible Instruments 104 Step 4: Examine Validity, Reliability, Cross-Cultural Fairness, and Practicality of the Possible Instruments 105 Step 5: Choose an Instrument Wisely 106

Summary 106

Chapter Review 107

References 107

6 Statistical Concepts: Making Meaning Out of Raw Scores 110 Raw Scores 111

Frequency Distributions 112

Histograms and Frequency Polygons 113

Cumulative Distributions 115

Normal Curves and Skewed Curves 116 The Normal Curve 116 Skewed Curves 117

CONTENTS ix

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Measures of Central Tendency 118 Mean 118 Median 118 Mode 119

Measures of Variability 119 Range 120 Interquartile Range 120 Standard Deviation 122

Remembering the Person 124

Summary 125

Chapter Review 126

Answers to Items 5 through 8 126

Reference 126

7 Statistical Concepts: Creating New Scores to Interpret Test Data 127 Norm Referencing versus Criterion Referencing 128

Normative Comparisons and Derived Scores 129 Percentiles 130 Standard Scores 130

z-Scores 131 T-scores 133 Deviation IQ 133 Stanines 135 Sten Scores 136 Normal Curve Equivalents (NCE) Scores 136 College and Graduate School Entrance Exam Scores 136 Publisher-Type Scores 138

Developmental Norms 139 Age Comparisons 139 Grade Equivalents 139

Putting It All Together 140

Standard Error of Measurement 141

Standard Error of Estimate 143

Scales of Measurement 145 Nominal Scale 146 Ordinal Scale 146 Interval Scale 146 Ratio Scale 146

Summary 147

Chapter Review 148

Answers to Items 4 through 10 149

References 149

x CONTENTS

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SECTION III Commonly Used Assessment Techniques 151

8 Assessment of Educational Ability: Survey Battery, Diagnostic, Readiness, and Cognitive Ability Tests 157 Defining Assessment of Educational Ability 158

Survey Battery Achievement Testing 159 National Assessment of Educational Progress (NAEP) 160 Stanford Achievement Test 162 Iowa Test of Basic Skills (ITBS) 163 Metropolitan Achievement Test 165

Diagnostic Testing 165 The Wide Range Achievement Test 4 (WRAT4) 166 Wechsler Individual Achievement Test—Third Edition (WIAT-III) 167 Peabody Individual Achievement Test (PIAT-R/NU) 167 Woodcock-Johnson® III 170 KeyMath3™ Diagnostic Assessment 170

Readiness Testing 171 Kindergarten Readiness Test (KRT) (Anderhalter & Perney) 172 Kindergarten Readiness Test (KRT) (Larson and Vitali) 172 Metropolitan Readiness Test (MRT6) 173 Gesell Developmental Observation—Revised 173

Cognitive Ability Tests 174 Otis-Lennon School Ability Test, Eighth Edition (OLSAT 8) 174 The Cognitive Ability Test (CogAT) 177 College and Graduate School Admission Exams 178

ACT 178 SAT 179 GRE General Test 180 GRE Subject Tests 181 Miller Analogy Test (MAT) 181 Law School Admission Test (LSAT) 181 Medical College Admission Test (MCAT) 182

The Role of Helpers in the Assessment of Educational Ability 182

Final Thoughts About the Assessment of Educational Ability 183

Summary 184

Chapter Review 185

References 186

9 Intellectual and Cognitive Functioning: Intelligence Testing and Neuropsychological Assessment 190 A Brief History of Intelligence Testing 191

Defining Intelligence Testing 191

Models of Intelligence 192 Spearman’s Two-Factor Approach 192 Thurstone’s Multifactor Approach 193 Vernon’s Hierarchical Model of Intelligence 193

CONTENTS xi

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Guilford’s Multifactor/MultiDimensional Model 193 Cattell’s Fluid and Crystal Intelligence 193 Piaget’s Cognitive Development Theory 195 Gardner’s Theory of Multiple Intelligences 196 Sternberg’s Triarchic Theory of Successful Intelligence 197 Cattell-Horn-Carroll (CHC) Integrated Model of Intelligence 199 Theories of Intelligence Summarized 199

Intelligence Testing 199 Stanford-Binet, Fifth Edition 202 Wechsler Scales 203 Kaufman Assessment Battery for Children 208 Nonverbal Intelligence Tests 208

Comprehensive Test of Nonverbal Intelligence, Second Edition (CTONI-2) 209 Universal Nonverbal Intelligence Test (UNIT) 209 Wechsler Nonverbal Scale of Ability (WNV) 209

Neuropsychological Assessment 210 A Brief History of Neuropsychological Assessment 210 Defining Neuropsychological Assessment 210 Methods of Neuropsychological Assessment 211

Fixed Battery Approach and the Halstead-Reitan Battery 211 Flexible Battery Approach and the Boston Process Approach (BPA) 213

The Role of Helpers in the Assessment of Intellectual and Cognitive Functioning 214

Final Thoughts on the Assessment of Intellectual and Cognitive Functioning 215

Summary 215

Chapter Review 217

References 218

10 Career and Occupational Assessment: Interest Inventories, Multiple Aptitude, and Special Aptitude Tests 221 Defining Career and Occupational Assessment 222

Interest Inventories 222 Strong Interest Inventory® 223

General Occupational Themes 223 Basic Interest Scales 225 Occupational Scales 225 Personal Style Scales 225 Response Summary 225

Self-Directed Search 228 COPSystem 228

Career Occupational Preference System Interest Inventory (COPS) 229 Career Ability Placement Survey (CAPS) 230 Career Orientation Placement and Evaluation Survey (COPES) 230

O*NET and Career Exploration Tools 230 Other Common Interest Inventories 233

Multiple Aptitude Testing 233 Factor Analysis and Multiple Aptitude Testing 233 Armed Services Vocational Aptitude Battery and Career Exploration Program 234 Differential Aptitude Tests 237

xii CONTENTS

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Special Aptitude Testing 238 Clerical Aptitude Tests 238 Mechanical Aptitude Tests 239 Artistic Aptitude Tests 240 Musical Aptitude Tests 241

The Role of Helpers in Occupational and Career Assessment 241

Final Thoughts Concerning Occupational and Career Assessment 242

Summary 242

Chapter Review 244

References 244

11 Clinical Assessment: Objective and Projective Personality Tests 247 Defining Clinical Assessment 248

Objective Personality Testing 248 Common Objective Personality Tests 249

Minnesota Multiphasic Personality Inventory-2 249 Millon Clinical Multiaxial Inventory, Third Edition 252 Personality Assessment Inventory® 253 Beck Depression Inventory-II 256 Beck Anxiety Inventory 256 Myers-Briggs Type Indicator® 257 Sixteen Personality Factors Questionnaire (16PF)® 260 The Big Five Personality Traits and the NEO PI-3™ and NEO-FFI-3™ 262 Conners 3rd Edition 264 Substance Abuse Subtle Screening Inventory (SASSI®) 265 Other Common Objective Personality Tests 266

Projective Testing 266 Common Projective Tests 267

The Thematic Apperception Test and Related Instruments 267 Rorschach Inkblot Test 269 Bender Visual-Motor Gestalt Test, Second Edition 270 House-Tree-Person and Other Drawing Tests 272 Sentence Completion Tests 273

The Role of Helpers in Clinical Assessment 274

Final Thoughts on Clinical Assessment 274

Summary 275

Chapter Review 277

References 277

12 Informal Assessment: Observation, Rating Scales, Classification Methods, Environmental Assessment, Records and Personal Documents, and Performance-Based Assessment 281 Defining Informal Assessment 282

Types of Informal Assessment 282 Observation 283

CONTENTS xiii

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Rating Scales 284 Numerical Scales 285 Likert-Type Scales (Graphic Scales) 286 Semantic Differential Scales 286 Rank-Order Scales 287

Classification Methods 287 Behavior Checklists 287 Feeling Word Checklists 288 Other Classification Methods 288

Environmental Assessment 290 Direct Observation 290 Situational Assessment 291 Sociometric Assessment 291 Environmental Assessment Instruments 291

Records and Personal Documents 293 Biographical Inventories 293 Cumulative Records 295 Anecdotal Information 295 Autobiography 296 Journals and Diaries 296 Genograms 296

Performance-Based Assessment 297

Test Worthiness of Informal Assessment 299 Validity 299 Reliability 300 Cross-Cultural Fairness 301 Practicality 302

The Role of Helpers in the Use of Informal Assessment 302

Final Thoughts on Informal Assessment 302

Summary 303

Chapter Review 304

References 305

Appendix A Websites of Codes of Ethics of Select Mental Health Professional Associations 306

Appendix B Assessment Sections of ACA’s and APA’s Codes of Ethics 308

Appendix C Code of Fair Testing Practices in Education 314

Appendix D Sample Assessment Report 321

Appendix E Supplemental Statistical Equations 327

Appendix F Converting Percentiles from z-Scores 329

GLOSSARY 331

INDEX 340

xiv CONTENTS

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Preface

We have been delighted to update Essentials of Testing and Assessment to its third edition. We believe this new edition brings some important and critical changes, yet we have kept the basic content of the book as is. We hope that by obtaining good input from reviewers and by adding important and new issues that have arisen, the third edition adds value to the previous two editions. Despite the changes and additions, we were careful to keep the core of the book the same and to only cover what can reasonably be learned in a one-semester course on testing and assessment—not more. But don’t let that fool you, as Essentials of Testing and Assessment includes quite a bit of information! We believe this book is written in a down-to-earth fashion that you will hopefully find interesting. In addition, we offer stories and vignettes that highlight learning. Some of the overarching changes include:

• Rearranging the sections of the book. Although the book has kept the basic content from the original 12 chapters, the chapters have been redistributed in the book to better suit classroom teaching and student learning (see descrip- tion of chapters that follows).

• Inclusion of information on two national studies, one on counselors use of assessment instruments and one that focuses on the kinds of assessment instruments that are taught in counseling programs.

• Update of citations and research. • Additional information on cross-cultural assessment. • The addition of new and updated assessment instruments. • Updating the chapter on diagnosis to reflect DSM-5.

The following gives a brief description of the content and the changes to the 3 sections and 12 chapters of Essentials of Testing and Assessment.

xv

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SECTION I: UNDERSTANDING THE ASSESSMENT PROCESS: HISTORY, ETHICAL AND PROFESSIONAL ISSUES, DIAGNOSIS, AND THE ASSESSMENT REPORT

Section I introduces the reader to a broad range of issues related to understanding testing and assessment. Chapter 1 provides definitions of assessment and the his- tory of assessment, Chapter 2 provides the reader with important professional, eth- ical, and legal issues in assessment. Chapter 3 offers information on diagnosis, one critical aspect of the assessment process, whereas Chapter 4 presents information on how to write a test report.

Chapter 1: History of Testing and Assessment Chapter 1 begins with a discussion of the differences between testing and assess- ment and then goes on to highlight the historical development of testing and assessment from ancient times to modern-day assessment instruments. Along the way, we discuss some of the people who were critical to the development of assess- ment measures and examine some of the many controversial issues that arose. The chapter nears its conclusion by reviewing the current categories of assessment instruments, including ability testing (achievement and aptitude testing), personal- ity assessment, and informal assessment. We finish by raising a number of con- cerns that continue to face us as we administer assessment instruments. Some of the changes in this chapter include adding the following: additional examples of early types of assessment instruments, quick definitions of types of assessment pro- cedures, and a figure that helps to understand the various kinds of ability, person- ality, and informal assessment procedures.

Chapter 2: Ethical, Legal, and Professional Issues in Assessment Chapter 2 focuses on the complex ethical, legal, and professional issues that are faced by individuals who are assessing others. We begin by discussing the com- plexity of ethical decision making and then identify ethical codes and professional standards critical to testing and assessment. We go on to discuss the importance of wise ethical decision making, and we identify a number of laws that have been passed and lawsuits resolved that impinge on the use of tests in the assessment pro- cess. The chapter concludes with a discussion of professional associations that address assessment, a brief discussion of accrediting bodies that address assess- ment, an introduction to the growing field of forensic evaluation, the importance of viewing assessment as a holistic process, a discussion cross-cultural assessment, and the importance of embracing the testing and assessment process. Some changes to this chapter include updating information from revisions of ethical codes, updat- ing and streamlining important assessment standards, updating relevant laws related to the administration and interpretation of tests, and expanding the discus- sion on cross-cultural issues related to assessment.

xvi PREFACE

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Chapter 3: Diagnosis in the Assessment Process Chapter 3 (formerly Chapter 11) has been dramatically changed to reflect the transition from DSM-IV to DSM-5. This chapter starts with a discussion on the importance of making a diagnosis then goes on to offer a brief history of the Diagnostic and Statistical Manual of Mental Disorders. An overview of DSM-5 is then offered, which includes the change to a single axis system; the use of a dimensional assessment (mild, moderate, severe, and very severe); the use of wide spectrum disorders, brief explanations of the specific diagnostic categories; a description of how DSM-5 addresses cross-cultural issues; an explanation of how medical, psychosocial, and environmental conditions can impact diagnosis; and some examples of how to make a diagnosis.

Chapter 4: The Assessment Report Process: Interviewing the Client and Writing the Report Chapter 4 (formerly Chapter 12) was moved to the first section of the book because it seemed to better fit the content of Chapters 1 and 2. Also, many faculty were having students write test reports; thus, we thought it would be best to high- light the test report-writing process near the beginning of the book so students would have an early understanding of how to begin this important project. How- ever, since this chapter and Chapter 3 can stand on their own, you may want to continue to cover them later in the course. Chapter 4 begins with a definition of the purpose of the assessment report. Then, we discuss the importance of accurately iden- tifying the breadth and depth of a client’s issues so that one can make smart decisions in choosing which assessment procedures to use. In this chapter, we also distinguish between conducting a structured, unstructured, and semi-structured interview, and we point out how computers have become increasingly important in the writing of test reports. Finally, we give a detailed description of the categories of a test report and offer suggestions on how to write a report. Included in this description is an expanded explanation of the mental status exam and a new section on assessing lethality. Using a fictitious client, we offer an example of how to right a report. The resulting five-page assessment report can be found in its entirety in Appendix D of the book.

SECTION II: TEST WORTHINESS AND TEST STATISTICS Section II of the book addresses test worthiness and test statistics. The three chap- ters that make up this section examine how tests are created, scored, and inter- preted. These chapters all use test statistics in some manner to explain the concepts being presented. In this section, we demonstrate how collecting and inter- preting test data is a deliberate and planned process that involves a scientific approach to the understanding of differences among people.

Chapter 5: Test Worthiness: Validity, Reliability, Practicality, and Cross-Cultural Fairness Chapter 5 examines four critical areas of test worthiness: (1) validity: whether a test measures what it is supposed to measure; (2) reliability: whether the score an

PREFACE xvii

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

individual receives on a test is an accurate measure of his or her true score; (3) cross-cultural fairness: whether the score the individual has obtained is a true reflection of the individual and not a function of cultural bias inherent in the test, and (4) practicality: whether it makes sense to use a test in a particular situation. After examining these four factors, we conclude with a discussion of five steps to use to assure test worthiness when selecting a test to administer. The chapter also provides an explanation of two statistics: correlation coefficient and coefficient of determination. We present these statistics as they are foundational to understand- ing much of what is presented in this and future chapters. In this chapter, we enhanced the section on negative correlation to improve readability, and we added figures to visually depict representations of different types of validity.

Chapter 6: Statistical Concepts: Making Meaning Out of Raw Scores Chapter 6 starts by noting that raw scores generally provide little meaningful information about a set of scores. We then provide ways that we can manipulate raw scores to make sense out of a set of data, and we examine how the following concepts are used to help us understand raw scores: frequency distributions; histo- grams and frequency polygons; cumulative distributions; the normal curve; skewed curves; measures of central tendency such as the mean, median and mode; and measures of variability, such as the range, semi-interquartile range, and standard deviation. A couple of the updates to this chapter include an explanation why cal- culators sometimes give a slightly different standard deviation than the standard deviation you will get when you use the formula in the book and a second formula has been added for the calculation of standard deviation.

Chapter 7: Statistics Concepts: Creating New Scores to Interpret Data In Chapter 7 we examine how raw scores are converted to what are called “derived” scores so that individuals can more easily understand what their raw scores mean. We first distinguish between norm-referenced and criterion-referenced testing and then go on to discuss the following derived scores: (1) percentiles; (2) standard scores, including z-scores, T-scores, deviation IQs, stanines, normal curve equivalents (NCE’s), sten scores, college entrance exam scores (e.g., SATs and ACTs), and publisher-type scores; and (3) developmental norms, such as age comparisons and grade equivalents. The chapter then examines standard error of measurement and standard error of estimate, both of which offer ways of understanding the range in which one’s “true score” actually falls. The chapter concludes with a discussion of nominal, ordinal, interval, and ratio scales in which we explore each scale’s unique attributes and note that different kinds of assessment instruments use different kinds of scales. In this chapter, we have updated critical information on a number of standard scores, such as the SATs, GREs, and ACTs.

xviii PREFACE

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SECTION III: COMMONLY USED ASSESSMENT TECHNIQUES Section III, which includes Chapters 8 through 12, examines some commonly used assessment procedures. Throughout these chapters, we have carefully updated information on many of the tests discussed. At the beginning of this section, we offer an overview of some studies conducted with counselors and psychologists regarding the kinds of testing and assessment instruments they use and also offer a table that identifies the instruments we will cover in this section of the book. We note that most of the instruments we cover are used by a large portion of counselors and psychologists. Each of the chapters in this section has a specific focus that helps to delineate the kinds of assessment procedures used.

Chapter 8: Assessment of Educational Ability: Survey Battery, Diagnostic, Readiness, and Cognitive Ability Tests Chapter 8 examines tests typically given in schools to measure what students have learned and what they are capable of learning. After defining the various categories of assessment of educational ability, we give examples of survey battery achievement tests, diagnostic tests, readiness tests, and cognitive ability tests. Throughout the chapter, we offer insight into the world of high-stakes testing and some of the important issues raised as the result of such laws as No Child Left Behind. In this chapter, we have added information about the Wechsler Individual Achievement Test (WIAT) and the Woodcock-Johnson® III, expanded the discussion of readiness testing, and provided the latest information about college and graduate school tests.

Chapter 9: Intellectual and Cognitive Functioning: Intelligence Testing and Neuropsychological Assessment Chapter 9 explores intellectual and cognitive functioning. Here, we offer a brief his- tory of intelligence testing and of neuropsychological assessment, define these two areas, and describe differences and similarities between them. The chapter also offers an overview of some of the more popular models of intelligences, and we present some common verbal and nonverbal tests of intelligence as well as types of neuropsy- chological assessments. In this edition, we have added two new theories of intelli- gence, the Cattell-Horn-Carroll (CHC) integrated model of intelligence and the triarchic theory of successful intelligence; added information about the Wechsler Nonverbal Scale of Ability (WNV), and provided an updated discussion of the use of neuropsychological assessment as applied to individuals with traumatic brain injury, such as returning service members from Iraq and Afghanistan.

Chapter 10: Career and Occupational Assessment: Interest Inventories, Multiple Aptitude, and Special Aptitude Tests Chapter 10 examines the kinds of tests that can help individuals make decisions about their occupational or career path. We begin with an examination of interest

PREFACE xix

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

inventories, which are a type of personality assessment that generally look at indi- viduals’ likes and dislikes as well their personality orientation toward the world of work. We next explore multiple aptitude testing, which can help individuals iden- tify the range of skills and abilities that may be important in choosing an occupa- tion. We conclude with a look at some of the more popular special aptitude tests that look at focused areas of ability, such as clerical skills, mechanical skills, artis- tic skills, and musical ability. In this edition, we added information on O*NET and its Career Exploration Tools and discuss the new computerized version of the Armed Services Vocational Aptitude Battery (CAT-ASVAB).

Chapter 11: Clinical Assessment: Objective and Projective Personality Tests Chapter 11 examines the process of using tests for clinical assessment. We start by underscoring the fact that such assessment has a wide variety of applications and can be an important tool for the clinician or researcher. The chapter then defines clinical assessment and speaks of some of the critical issues that such assessment can address. The meat of the chapter is a review of a number of the more well known and used objective and projective clinical assessment procedures. This edi- tion has added information on the Conners 3 instrument for assessing ADHD and the Beck Anxiety Inventory (BAI).

Chapter 12: Informal Assessment: Observation, Rating Scales, Classification Methods, Environmental Assessment, Records and Personal Documents, and Performance-Based Assessment This final chapter of the book covers informal assessment procedures. We start by defining informal assessment and then identify a number of different kinds of informal assessment techniques, including observation, rating scales, classification methods, environmental assessments, records and personal documents, and perfor- mance-based assessment. The second part of the chapter offers a discussion con- cerning test worthiness of informal assessment. In this edition, we provide an expanded genogram as a model to use to inquire about patterns in family and family assessment.

GLOSSARY AND APPENDICES At the end of the book, we offer a glossary of all of the major words and terms used in the text. This should be helpful in remembering some of the basic concepts you read about. Also, within each chapter, you will find important concepts highlighted in the margins. Essentials of Testing and Assessment offers six appen- dices to enhance some of the issues identified in the text. Appendix A offers the Web sites of major professional associations and their accompanying ethical codes, Appendix B presents the assessment sections of ACA’s and APA’s ethical codes, Appendix C provides the Code of Fair Testing Practices in Education,

xx PREFACE

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Appendix D offers an example of a test report that should be helpful as a guide when you write your own report. Appendix E offers supplemental statistical equa- tions to help guide you in some of aspects of tests statistics, and Appendix F pro- vides a conversion table of percentiles from z-scores

FINAL THOUGHTS We believe this book is comprehensive without being overbearing. The text pro- vides an overview of testing and assessment in a readable, and we think, enjoyable fashion. We hope that after reading it, you come away with a new appreciation of testing and assessment.

ACKNOWLEDGEMENTS Every book has a number of people behind the scenes supporting its development. In this edition of Essentials of Testing and Assessment, we would like to acknowl- edge some of these individuals. From Cengage Publishing we would like to thank Jon-David Hague, product director, who helped us get this revision off-the-ground and set direction for the new edition. Also from Cengage, a very special thanks goes to Julie Martinez, product manager, who has given us the dad-to-day support and encouragement needed to complete this book. In addition, Kyra Kane, product assistant from Cengage, assisted with numerous behind the scenes tasks that are needed to be completed to finish any book. Other people were also critical to the book’s completion. Greg Johnson, senior project manager from PreMediaGlobal (PMG), was essential in helping us go through each chapter and ensuring that everything was on time, correct, and ready to go out the door. A special thanks goes to Greg. We would also like to thank the copyeditor from PMG who spent countless hours copyediting our text. A shout-out also goes to all those people who were behind the scenes ensuring completion of the book and whose names we never get to see.

In addition to the individuals just noted, for every revision we rely on a number of faculty to review the text and offer us important feedback about potential changes that need to be made. For this edition, the following individuals were criti- cal to this process: Eric Bruns, Campbellsville University; Laurie Carlson, Colorado State University; David Carter, University of Nebraska Omaha; Mary Ann Coup- land, Sinte Gleska University; Aaron Hughey, Western Kentucky University; Sheri Pickover, University of Detroit Mercy; Clarrice Rapisarda, The University of North Carolina at Charlotte; and Todd Whitman, Shippensburg University. In addition, a special thanks to Peg Jensen for her careful read of Chapter 9: Intellectual and Cognitive Functioning, which led us to make substantial changes in that chapter. Finally, thanks to Martin Baggarly, my nephew, who drew the amazing turtle and shoe for the “Test of Artistic Potential” in Chapter 10.

PREFACE xxi

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

S E C T I O N IUnderstanding theAssessment Process: History, Ethical and Professional Issues, Diagnosis, and the Assessment Report

Section I of this book includes four chapters: Chapter 1: History of Testing and Assessment; Chapter 2: Ethical, Legal, and Professional Issues in Assessment; Chapter 3: Diagnosis in the Assessment Process; and Chapter 4: The Assessment Report Process: Interviewing the Client and Writing the Assessment Report.

In the first chapter, we define testing and assessment and examine the history of testing and assessment starting with ancient times and working our way to the development of modern-day assessment instruments. Near the end of the chapter, we examine current categories of assessment instruments, including ability testing (testing in the cognitive realm), personality assessment (testing in the affective realm), and informal assessment techniques. The chapter concludes by raising a number of concerns that continue to face us today as we administer assessment instruments.

Chapter 2 examines the many complex ethical, legal, and professional issues that confront individuals who are assessing others. We identify aspects of major ethical codes that focus on assessment; summarize standards that have been devel- oped to help guide the practitioner when administering, scoring, and interpreting assessment procedures; and discuss the process of ethical decision making. We then examine legal issues that have impinged on the use of tests, professional asso- ciations that focus on assessment, and organizations that accredit programs that teach assessment. A discussion on the professional field of forensic evaluation then ensues and the chapter concludes with a discussion of assessment as a holistic pro- cess and issues of bias in testing.

1

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

In Chapter 3, we begin by discussing the importance of making a diagnosis. We then go on to introduce the Diagnostic and Statistical Manual of Mental Disorders, 5th edition (DSM-5), which is the common diagnostic tool used by almost all clinicians. We offer a brief history of the DSM, discuss the recent changes that took place in the development of DSM-5 from DSM-IV-TR, offer a brief overview of the various diagnostic classifications, and demonstrate how DSM-5 is used in assessing a client. We conclude the chapter with a discussion of the importance of diagnosis in the total assessment process.

Chapter 4 highlights the importance of the assessment report and suggests the report is the “deliverable” or “end product” of the assessment process. Here, we offer guidelines for conducting an effective interview, and we distinguish among structured, unstructured, and semi-structured interviews. In Chapter 4, we discuss the importance of choosing assessment measures that match the client’s presenting issues, and we stress the importance of considering the breadth and depth of cli- ents’ issues when choosing assessment procedures. The rest of Chapter 4 is dedi- cated to ways of writing effective test reports, and we delineate topics that should be covered to create a thorough assessment of a client. An example of a fictitious client is offered as we show how each section of the assessment report should be written. The complete assessment report can be found in Appendix D of the book.

2 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 1History of Testing andAssessment

You walk into the room, prepared to take the test and know that the results will impact your future. You can feel your heart begin to pound and your stomach being to churn. “OMG, I hope I can do well, you say to yourself.”

With millions of children and adults frightened by the thought of taking a test, this is not a pretty picture. But is there value in this sometimes-terrifying experience? We’ll let you answer that question after you have finished reading this book. But how did test-taking start? That question will be answered in this chapter.

(Ed Neukrug)

In this chapter we will examine the history of testing and assessment. First, we will explore the differences between testing and assessment and point out how their cur- rent definitions are directly related to their history. We will then take a ride through the history of assessment, starting with ancient history and working our way to the development of modern-day assessment instruments. Along the way, we will highlight some of the people who were pioneers in the development of assessment measures and discuss some of the controversial issues that arose. As the chapter nears its conclusion, we will examine the current categories of assess- ment instruments, and we will finish by raising a number of ongoing concerns sur- rounding the use of assessment instruments.

3

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

DISTINGUISHING BETWEEN TESTING AND ASSESSMENT Today, the term assessment includes a broad array of evaluative procedures that yield information about a person (Hunsley, 2002). Assessment procedures include the clinical interview; informal assessment techniques such as observation, rating scales, classification methods, environmental assessment, records and personal documents, and performance-based assessment; personality tests such as objective tests, projective tests, and interest inventories; and ability tests such as achievement tests and aptitude tests (see Figure 1.1).

Tests are a subset of assessment techniques that yield scores based on the gath- ering of collective data (e.g., finding the sum of correct items on a multiple-choice exam). Assessment procedures can be formal, which means they have been well- researched and shown to be scientifically sound, valid, and reliable, or informal, which implies that such rigor has not been demonstrated, although the procedure might still yield some valuable information.

Generally, the greater the number of procedures used in assessing an individ- ual, the greater the likelihood that they will yield a clearer snapshot of the client. Thus, using multiple assessment procedures, or a holistic approach to assessment,

FIGURE 1.1 | Assessment Procedures

Assessment A broad array of eval- uative procedures

Tests Instruments that yield scores based on col- lected data—a subset of assessment

© Ce ng ag e Le ar ni ng

4 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

should always be considered when making important decisions about a client’s life (Association for Assessment and Research in Counseling, 2012; Joint Committee on Testing Practices, 2004). In this text we will examine a broad array of formal and informal assessment procedures, all of which can be used in the decision- making process. But let’s start at the beginning and see how events of the past have moved us toward our current use of assessment instruments.

THE HISTORY OF ASSESSMENT Although the modern era of assessment began near the beginning of the twentieth century, assessment procedures can be found in ancient times. Let’s examine some of the changes in assessment that have taken place over the centuries.

Ancient History He said, “Take your son, your only son Isaac, whom you love, and go to the land of Moriah, and offer him there as a burnt offering on one of the mountains that I shall show you” (Genesis 22:5, New Revised Standard Version).

Assessment has been around for as long as humans have walked the earth. In fact, one might say that Abraham’s loyalty was assessed when God asked him to kill his son Isaac. From a more down-to-earth perspective, the Chinese government is given credit for developing one of the first widely used tests when it began to assess indi- viduals for fitness to work in government positions in approximately 2200 B.C.E. (DuBois, 1970; Higgins & Sun, 2002). With testing done under grueling conditions in hundreds of small cubicles or huts, it was not unusual that examinees would die from exhaustion (Cohen, Swerdlik, & Sturman, 2012). This kind of testing was not abolished until 1905. In the Western world, passages from Plato’s (428–327 B.C.E.) writings indicate the Greeks assessed both the intellectual and physical abil- ity of men when screening for state service (Doyle, 1974).

Precursors to Modern-Day Test Development As experimental and controlled research spread throughout the scientific commu- nity during the 1800s, physicians and philosophers began to apply these research principles to the understanding of people, particularly in the area of cognitive func- tioning. For instance, working in mental asylums, the French physician Jean Esquirol (1772–1840) examined how language ability of individuals with intellec- tual disabilities was related to intelligence (Zusne, 1984; Drummond, 2009). Seen as having a condition called “idiocy,” (Esquirol, 1838, p. 38), these individuals were viewed as having intellectual deficits as compared to a “normal” person reared in a similar environment. Esquirol’s focus on language ability is often seen as the beginning of what later became known as the assessment of “verbal IQ.” At around the same time, Edouard Seguin (1812–1880), also from France, sug- gested that the prognosis regarding intellectual deficits in children was worse if such deficits were associated with physiological problems. He suggested that physi- cians should “watch for a swinging walk, ‘automatically busy’ hands, saliva drip- ping from a ‘meaningless mouth,’ a ‘lustrous and empty’ look, and ‘limited’ or

Multiple assessment procedures should always be considered

Jean Esquirol Used language to identify intelligence— forerunner of “verbal IQ”

Edouard Seguin Developed the form board to increase motor control— forerunner of “performance IQ”

CHAPTER 1 History of Testing and Assessment 5

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

‘repetitive’ speech” (Zenderland, 1987, p. 54). Eventually, Seguin developed the form board to increase his patients’ motor control and sensory discrimination and to compare children and individuals with severe intellectual deficits to average chil- dren at different age groups (see Figure 1.2). Considered by some to be the forerun- ner of “performance IQ” measures (DuBois, 1970), versions of the form board, which is similar to the toy in which children place shapes in their respective grooves, are still used today in some performance-oriented IQ test.

Meanwhile, intrigued by Charles Darwin’s (1809–1882) theory of evolution, scientists during the mid-1800s became engrossed in trying to understand the devel- opment of the human species (Juve, 2008; Kerr, 2008). For instance, Sir Francis Galton (1822–1911), Darwin’s cousin, became fascinated by differences among people and eventually came to believe that people inherited physical and mental characteristics (Gillham, 2001; Murdoch, 2007). He hypothesized that some inher- ited physical traits, such as reaction time and stronger grip strength, might be related to superior intellectual ability. His curiosity led him to examine the relation- ship among such characteristics, and his research spurred others to develop the sta- tistical concept of the correlation coefficient, which describes the strength of the relationship among variables (DuBois, 1970; Kerr, 2008). Calculating the correla- tion coefficient has become an important tool in the development and refinement of tests.

Wilhelm Wundt (1832–1920), another scientist intrigued with human nature, set out to create “a new domain of science” that he called physiological psychol- ogy. Around 1875, at the University of Leipzig in Germany, Wundt developed one of the first psychological laboratories that used experimental research (Nicolas, Gyselinck, Murray, & Bandomir, 2002). Many of the experiments in Wundt’s lab- oratory studied reaction time of hearing, sight, and other senses in response to sti- muli (Watson, 1968). A number of students who worked with Wundt helped to

FIGURE 1.2 | Reproduction of Sequin’s Form Board Task: Children, or individuals with intellectual disabilities, would be given ten blocks, in three piles, and asked to place them in the slots as fast as they can. They would then determine intel- lectual age by finding which age group the individual was most similar to.

Sir Francis Galton Examined relationship of sensory motor responses to intelligence

Wilhelm Wundt Developed one of the first psychological laboratories

© Ce ng ag e Le ar ni ng

20 15

6 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

foster in the new age of psychological science. For instance, James McKeen Cattell (1860–1944), a doctoral student under Wundt who was later greatly inspired by Galton, became one of the earliest American psychologists to use statistical concepts in understanding the person (Goodwin, 2008; Roback, 1961). Cattell’s main emphasis became the assessment of what he termed mental tests and included exam- ining individual differences of such things as memory span and reaction time. Another important figure, G. S. Hall (1844–1924), also worked with Wundt and eventually set up his own experimental lab at Johns Hopkins University. Hall became a mentor to other great American psychologists and was the founder and first president of the American Psychological Association in 1892 (Benjamin, 2008).

The Emergence of Ability Tests (Testing in the Cognitive Domain) Influenced by the new scientific approach to understanding human nature, researchers at the beginning of the twentieth century began to develop instruments that could scientifically measure an individual’s abilities. This era saw the emer- gence of ability tests, including individual intelligence tests, neuropsychological assessments, and group tests of ability.

Intellectual and Cognitive Functioning: Intelligence Testing and Neuro- psychological Assessment Although commonplace today, the first intelligence tests were developed by Alfred Binet (1857–1911) who, in 1904, was commissioned by the Ministry of Public Education in Paris, to construct a test that could be of assistance in integrating what they called “subnormal” children into the schools (Binet & Simon, 1916). Highly critical of the manner in which “mental deficiency” was diagnosed in children, Binet and his colleague Theophile Simon developed a scale that could be administered one-on-one and which would measure higher men- tal processes by assessing responses to a variety of different kinds of tasks (e.g., tracking a light, asking the individual to distinguish between different types of words) (Ryan, 2008). The information gained from their observations was then used to develop the first modern-day intelligence test (Watson, 1968). A relatively short time later, Lewis Terman (1877–1956), from Stanford University, began ana- lyzing and methodically gathering extensive normative data on Binet and Simon’s scale from hundreds of children in the Stanford, CA area (Jolly, 2008; Kerr, 2008). Based on the data, Terman made a number of revisions to the Binet and Simon scale. Originally called the Stanford Revision of the Binet and Simon scale, the test later became known as the Stanford-Binet, the name by which the revised version continues to be known as today. Terman was the first to use the term intelligence quotient, or “IQ,” which used a ratio of mental age to chronological age (see Box 1.1).

Intelligence tests are sometimes used in neuropsychological assessment, which examines changes in brain function as the result of injury or disease process. Interest in how the brain impacts cognitive and behavioral functions, however, can be traced back to early Egypt where observations of behavioral changes following head inju- ries are recorded in 5,000-year-old Egyptian medical documents (Hebben & Milberg, 2009). In modern times, research conducted during World War I examined

James McKeen Cattell Brought statistics to mental testing— coined term mental test

G. S. Hall Early experimental psychologist. First president of APA

Alfred Binet Created first modern intelligence test

Lewis Terman Enhanced Binet’s work to create Stanford-Binet intelligence test

Intelligence quotient Mental age divided by chronological age

CHAPTER 1 History of Testing and Assessment 7

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

behaviors of soldiers who suffered from brain injuries. At that time, with the test- ing movement in full swing, it is not surprising that new screening and diagnostic measures to assess behavioral changes due to brain trauma were created (Lezak, Howieson, Bigler, & Tranel, 2012). During the twentieth century, as individuals became more interested in the nature of the brain and spurred on by the use of X-rays and other forms of brain imaging, the field of neuropsychology, or the study of brain function as it relates to behavior, was established. Today, when suspected changes occur in brain function due to disease, accidents, or violence, a “neuropsych” assessment, which sometimes includes an intelligence test, is often used.

Group Tests of Ability (Group Testing in the Cognitive Domain) Realizing the importance of obtaining accurate information from examinees, early test developers, such as Terman and others, devised standardized directions to use in testing and stressed the importance of having trained examiners administer tests individually (Geisinger, 1994; Jolly, 2008). However, it was soon evident that individual testing, such as that conducted when doing intelligence testing, often took a particularly long time and was costly. During World War I, these practical concerns came to a head as it became critical to quickly administer tests of cognitive ability in order to place large groups of recruits in the military. At that time, Robert Yerkes, the president of the American Psychological Association, chaired a special committee to create a screening test for these new recruits. The committee, composed of many well- known psychologists, including Terman, prepared a draft of the test in just four months (Geisinger, 2000; Jones, 2007). The original test the committee developed was known as the Army Alpha (see Box 1.2 and Illustration 1.1).

Although the Army Alpha clearly had its problems, it was a large step toward the mass use of tests in decision-making and was administered to more than 1.7 million recruits in less than two years (Haney, 1981). Since there were many foreign-born recruits and large numbers of people who could not read, a second

BOX 1.1 Developing the Notion of “IQ”

Lewis Terman wanted to develop a logical and rela- tively easy way of expressing an individual’s intelli- gence. Using the data from his research, he quickly realized that he could compute a ratio score for each child by dividing a child’s mental score (the age score at which the child performed) by the child’s actual age. Thus, if a child was performing at the level of the average 12-year-old but was actually 9 years old, the ratio would be 12/9 or 1.33. Multiply- ing this number by 100 to eliminate the decimal point would yield an intelligence quotient (“IQ”) of 133*.

Use this method to determine the IQs of the children below, based on their mental age scores and their actual ages.

Child 1: mental age of 6 and chronological age of 8.

Child 2: mental age of 16 and chronological age of 16.

Child 3: mental age of 10 and chronological age of 9.

Answers: Child 1: 75 (6/8 3 100), Child 2: 100 (16/16 3 100), Child 3: 111 (10/9 3 100).

*Note: IQ is no longer determined in this manner, and the current method of calculation will be discussed later in the text.

© Ce ng ag e Le ar ni ng

Robert Yerkes Chairman of the com- mittee that developed the Army Alpha

Army Alpha First modern group test—used during WWI

8 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

BOX 1.2 The Army Alpha Test

The Army Alpha test was created to place recruits in the military (Jones, 2007; Mckean, 1985). Based on this test, it was found that average mental age of the

recruit was 13. Take the test below in the 3-minute time allotment given, then consider potential issues of bias and cultural fairness of the questions on the test.

Source: McKean, K. (1985). Intelligence: New ways to measure the wisdom of man. Discover Magazine, 6(10), 28. Reprinted by permission of Disney Publications.

The Army Alpha was used to determine placement in the armed forces during WWI. Below is an adaptation of the test, as printed in Discover magazine. Take the test and discuss your thoughts about it.

The average mental age of the recruits who took the Army Alpha test during WWI was approxi- mately 13. Could you do better? You have three minutes to complete these sample questions, drawn verbatim from the original exam. (McKean, 1985)

The following sentences have been disarranged but can be unscrambled to make sense. Rearrange them and then answer whether each is true or false.

1. Bible earth the says inherit the the shall meek. true false 2. a battle in racket very tennis useful is true false

Answer the following questions: 3. If a train goes 200 yards in a sixth of a minute, how many feet does it go in a fifth of a second? 4. A U-boat makes 8 miles an hour under water and 15 miles on the surface. How long will it take to

cross a 100-mile channel if it has to go two-fifths of the way under water? 5. The spark plug in a gas engine is found in the: crank case manifold cylinder carburetor 6. The Brooklyn Nationals are called the: Giants Orioles Superbas Indians 7. The product advertised as 99.44 per cent pure is:

Arm & Hammer Baking Soda Crisco Ivory Soap Toledo 8. The Pierce-Arrow is made in: Flint Buffalo Detroit Toledo 9. The number of Zulu legs is: two four six eight

Are the following words the same or opposite in meaning? 10. vesper–matin same opposite 11. aphorism–maxim same opposite

Find the next number in the series: 12. 74, 71, 65, 56, 44, Answer: 13. 3, 6, 8, 16, 18, Answer:

14. Select the image that belongs in the mirror: 15. & 16. What’s missing in these pictures?

Answers: 1. true, 2, false, 3. twelve feet, 4, nine hours, 5. cylinder, 6. superbas, 7. Ivory Soap, 8. Buffalo, 9. two, 10. opposite, 11. same, 12. 29, 13. 36, 14. A, 15, spoon, 16. gramophone horn Scoring: All items except 3, 4, 10, and 11 = 1.25 points. Items 3 and 4 = 1.875 points, Items 10 & 11 = .625 points. Add them all up, they equal your mental age. What is wrong with this test? Examine it for problems with content, history, cross-cultural contamination, and so forth.

15. 16.

B

A

C

D

CHAPTER 1 History of Testing and Assessment 9

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

language-free version of the test, known as the Army Beta, was created. The Army Beta applied the use of form boards and mazes, and directions were given by pan- tomime so recruits could take the entire test without reading (Goodwin, 2008). Although crude by today’s standards, the Army Alpha and Army Beta ushered in the era of group tests of ability.

In contrast with neuropsychological assessments and individual intelligence tests, which are given one-to-one, group tests of cognitive ability tend to be multi- ple choice and true-false measures given to groups of individuals simultaneously in an effort to assess the academic promise of each individual in the group. Probably the most well known of these has been the Scholastic Aptitude Test (now the SAT Reasoning Test, or SAT). Developed by the Educational Testing Service after World War II, the test in many ways was the brain child of James Bryant Conant, president of Harvard. Believing in a democratic, classless society, Conant thought that such tests could identify the ability of individuals and ultimately help to equal- ize educational opportunities (Frontline, 1999). Unfortunately, many have argued that instead of fostering equality, the SATs have been used to separate the social classes, and many in the testing movement were not as magnanimous as James Bryant Conant (see Box 1.3).

ILLUSTRATION 1.1 | Recruits Taking an Examination at Camp Lee, 1917 Source: U.S. Signal Corps photo number 11-SC-386 in the National Archives.

James Bryant Conant Developed SAT to equalize educational opportunities

10 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Paralleling the rise of group tests of cognitive ability was the administration of group tests of achievement in the schools. Traditionally, such tests had been given orally, and later in essay fashion, but the practicality of administering objective tests of academic performance to large groups of students was obvious. With the new scientific approach to testing on the rise, Edward Thorndike (1874–1949), one of the pioneers of modern-day educational and psychological assessment, and others, thought that such tests could be given in a format that was more reliable than previous tests. This move toward group testing culminated with the develop- ment of the Stanford Achievement Test in 1923 (Armstrong, 2006). Today, these tests are commonplace and are given to students en masse in school systems throughout America.

One last area where group testing became popular was in vocational counsel- ing. With Frank Parsons (1909/1989) at its helm, vocational counseling became increasingly important at the turn of the twentieth century as large numbers of peo- ple moved to the bigger cities in search of employment. With vocational counseling seen as a process of (1) acquiring self-knowledge, (2) acquiring knowledge of the world of work, and (3) finding a suitable match through a process called “true rea- soning,” thousands of individuals were anxious to discover what jobs might be a suitable match for them. And tests that could quickly measure large groups of peo- ples’ likes and dislikes, as well as their abilities, could help do that. Thus, we began to see the rise of “multiple aptitude” tests. For example, the General Aptitude Test

BOX 1.3 Eugenics and the Testing Movement

As scientists tried to make sense of the relatively new theory of evolution and as they began to examine differences among people, a number of them began attributing these differences to genetics. This information was used to bolster support for the emerging Eugenics Movement, whose members espoused a belief in selective breeding in order to improve the human race.

Individuals such as Galton, Terman, and Yerkes believed that the data retrieved from tests could help distinguish those who were naturally bright from those who, they argued, were less fortunate. Test results were used to advocate for providing incen- tives for the upper class to procreate and in finding methods to prevent the lower classes from having children (Gillham, 2001). Based on flimsy evidence and misguided thinking, this movement is also seen today as having racist undertones.

Believing that the Army Alpha and Army Beta measured innate ability, Terman, Yerkes, and others

used the results of these tests to support the Eugenics Movement. However, the tests were a far cry from being a measurement of intelligence, as they were saturated with cultural bias and were largely based on measuring achievement, or what had been learned, as opposed to some kind of raw, inherent intelligence. Unfortunately, their beliefs about the test and what should be done as a result of the test data may have been one of many influ- ences that led the U.S. government to manipulate whom it would allow to immigrate to the United States (Sokal, 1987). As a result, thousands— perhaps millions—of individuals were unable to emi- grate from tyrannical governments in Europe and other parts of the world (Gould, 1996).

Question to Ponder: Do you prefer partnering with someone who is bright? And, if so, are you practicing your own, individualized eugenics?

Edward Thorndike Developer of the Stanford Achievement Test

Frank Parsons Leader in vocational counseling

© Ce ng ag e Le ar ni ng

CHAPTER 1 History of Testing and Assessment 11

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Battery (GATB) was developed by the United States Employment Service to mea- sure abilities in a number of specific areas. These areas of ability could be directly matched with job characteristics to identify appropriate occupational choices.

The Emergence of Personality Tests (Testing in the Affective Realm) Paralleling the rise of tests in the cognitive domain, personality tests (or tests in the affective realm) began to be devised. Thus, around the turn of the twentieth cen- tury three types of personality assessment instruments were developed: interest inventories, objective personality tests, and projective personality tests. Let’s take a brief look at each of these areas.

Interest Inventories and Vocational Assessment In 1912, Edward Thorndike, one of the first to conduct research in the field of vocational assessment, published the results of a study that examined the interests of 100 students as they progressed from elementary school through college (DuBois, 1970). As the relationship of inter- ests to vocational choice became more obvious, in 1922, J. B. Miner developed one of the first formal interest blanks (inventories) that was used to assist large groups of high school students in selecting occupations. Miner (1922) understood that his test was only part of the total assessment process, and he explained that his inven- tory was “the basis for individual interviews with vocational counselors” (p. 312).

On the heels of Miner’s interest blank, in the mid-1920s Edward Strong (1884–1963) teamed up with a number of other researchers to develop what was to become the most well-known interest inventory (Cowdery, 1926; DuBois, 1970; Strong, 1926). Known as the Strong Vocational Interest Blank, the original inven- tory consisted of 420 items. Strong spent the rest of his life perfecting his voca- tional interest inventory. Having undergone numerous revisions over the years, this inventory continues to be one of the most widely used instruments in career counseling. Today, interest inventories like the Strong are often used in conjunction with multiple aptitude tests as part of the career counseling process.

Objective Personality Assessment Although Emil Kraeplin developed a crude word association test to study schizophrenia in the 1880s, Woodworth’s Personal Data Sheet is considered to be the ancestor of most modern-day personality inven- tories. Woodworth’s instrument, which was developed to screen WWI recruits for their susceptibility to mental health problems (DuBois, 1970), had 116 items to which individuals were asked to respond by underlining “yes” or “no” to indicate whether or not the statement represented them (see Box 1.4). These rather obvious questions were then related to certain types of neuroses and pathologies. Although the test had questionable validity compared to today’s instruments, it became an early model for a number of other, better-refined instruments including the Minnesota Multiphasic Personality Inventory (MMPI).

Projective Testing Experiments such as these allow an unexpected amount of illumination to enter into the deepest recess of the character, which are opened and bared by them like the

GATB Developed by U.S. Employment Service to measure multiple aptitudes

J. B. Miner Developed one of the first group interest inventories

Edward Strong Founder of the Strong Vocational Interest Blank—derivative still used today

Emil Kraeplin Developed early word association test

Woodworth’s Personal Data Sheet First modern person- ality inventory—used during WWI

12 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

anatomy of an animal under the scalpel of a dissector in broad daylight (Galton, 1879, p. 431).

These words of Galton speak to the premise of projective testing: present a stim- ulus to an individual in an attempt to tap into the unconscious mind and discover the inner world of that person. Recognizing the importance of Galton’s work, Cat- tell examined the kinds of associations that mentally healthy individuals made to a standard list of words (DuBois, 1970).

By 1904, Carl Jung (1875–1961) had come up with 156 stimulus words that he used in one of the earliest word association tests. Examining the responses of his cli- ents to his list of words, Jung came up with the word complex to describe sets of unusual and delayed responses that individuals had to these stimulus words that seemed to point to a problematic or neurotic area in their lives (Jung, 1918/1969; Jung & Riklin, 1904; Storr, 1973). However, it was Herman Rorschach (1884–1922), a student of Jung, who developed the most well-known projective test—the Rorschach Inkblot test. Rorschach created this test by selecting ten inkblots “thrown on a piece of paper, the paper folded, and the ink spread between the two halves of the sheet” (Rorschach, 1942, p. 1). He believed the interpretation of an individual’s reactions to these forms could tell volumes about the individual’s unconscious life. This test was the precursor to many other kinds of projective tests, such as Henry Murray’s Thematic Apperception Test, or TAT, which asks a subject to view a number of standard pictures and create a story that explains the situation.

The Emergence of Informal Assessment Procedures The twentieth century saw the increased use of informal assessment procedures, which are assessment instruments that are often developed by the user and designed to meet a particular testing situation. For instance, as business and industry expanded during the 1930s, the situational test became more prevalent. In these tests, businesses

BOX 1.4 Items from Woodworth’s Personal Data Sheet

Although crude by today’s standards, Woodworth’s Personal Data Sheet was one of the first instruments that attempted to assess one’s personality. Below are some of the original 116 items.

Answer the questions by underlining “Yes” when you mean yes, and by underlining “No” when you mean no. Try to answer every question.

1. Do you usually feel well and strong?

YES NO

3. Are you often frightened in the middle of the night?

YES NO

27. Have you ever been blind, half-blind, deaf or dumb for a time?

YES NO

51. Have you hurt yourself by masturbation (self abuse)?

YES NO

80. At night are you troubled by the idea that some- body is following you?

YES NO

112. Has any of your family been a drunkard?

YES NO

Source: Adapted from A History of Psychological Testing (pp. 160–163), by P. DuBois, 1970, Boston: Allyn and Bacon.

Carl Jung Used word associa- tions to identify mental illness

Herman Rorschach Developed famous Rorschach Inkblot test

Henry Murray Developed Thematic Apperception Test (TAT)

Informal assessment procedures User-created and situational

CHAPTER 1 History of Testing and Assessment 13

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

took individuals who were potential hires or candidates for promotion and placed them into “contrived naturalistic situations” to assess their ability to respond to real- life situations. Meanwhile, as treatment of mental health clients improved, another informal procedure, the clinical interview, became prominent. The clinical interview became especially important as clients were increasingly being assessed for a diagno- sis through the use of the Diagnostic and Statistical Manual of Mental Disorders (DSM), first developed by the American Psychiatric Association in 1952 (Neukrug & Schwitzer, 2006).

During the 1960s and 1970s the use of tests in schools greatly increased and laws were passed that called for the assessment of students with disabilities. It was then that a number of informal techniques became popular, including observation, rating scales, classification techniques, and the review of records and personal docu- ments to assess learning problems of children. Conducting an environmental assess- ment where many of these informal tools are used has also become popular in recent years. For instance, knowing whether a particular home (the “environment”) is a healthy place to raise a child has become an important aspect of the Child Protective Service worker’s role. Finally, in recent years, performance-based assessment as an alternative to the more traditional cognitive-based assessments (e.g., multiple choice tests) has become increasingly popular. Today, informal assessment techniques, such as those already mentioned, are used in a variety of settings in numerous ways.

Modern-Day Use of Assessment Procedures As complex statistical analyses became possible through the use of computers, the quality of assessment instruments advanced rapidly. Today, assessment instruments can be found in every aspect of society and their uses have been vastly expanded. Although one can categorize such instruments in many ways, we have found it helpful to classify them into the following groups: (1) testing in the cognitive domain, often called “ability testing,” (2) testing in the affective domain, usually called “personality assessment,” and (3) informal assessment procedures. Figures 1.3–1.5 are graphic displays of these domains and are followed in Box 1.5 by short definitions of the various categories.

In Section III of the text, we will demonstrate how all of the assessment catego- ries noted in Figures 1.3–1.5 and Box 1.5 are used today in a variety of ways. And the following shows how the chapters in Section III of the text correspond to the categories:

Chapter 8: Assessment of Educational Ability: Survey Battery, Diagnostic, Readiness, and Cognitive Ability Tests

Chapter 9: Intellectual and Cognitive Functioning: Intelligence Testing and Neuropsychological Assessment

Chapter 10: Career and Occupational Assessment: Interest Inventories, Multiple Aptitude, and Special Aptitude Tests

Chapter 11: Clinical Assessment: Objective and Projective Personality Tests

Chapter 12: Informal Assessment: Observation, Rating Scales, Classification Methods, Environmental Assessment, Records and Personal Documents, and Performance-Based Assessment

14 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

ASSESSMENT OF ABILITY (What one can do)

ACHIEVEMENT TESTING (What one has learned)

APTITUDE TESTING (What one is capable of learning)

Survey Battery

Diagnostic Readiness

Intelligence

Testing

Neuropsychological

Assessment

Intellectual and Cognitive

Functioning

Cognitive Ability

Multiple Aptitude

Special Aptitude

FIGURE 1.3 | Assessment in the Cognitive Domain

PERSONALITY ASSESSMENT (Temperament, Habits, Likes, Disposition, Nature)

Objective ("Paper and Pencil")

Projective (Unstructured Responses)

Interests (Likes and Dislikes)

FIGURE 1.4 | Assessment in the Affective Domain

INFORMAL ASSESSMENT INSTRUMENTS

Environmental Assessment

Performance- Based

Assessment

Records and Personal

Documents

Classification methods

Observation Rating scales

FIGURE 1.5 | Informal Assessment Procedures

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 1 History of Testing and Assessment 15

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

BOX 1.5 Brief Definitions of Assessment Categories

Assessment of Ability: Tests that measure what a person can do in the cognitive realm.

Achievement Testing: Tests that measure what one has learned.

Survey Battery Tests: Tests, usually given in school settings, which measure broad content areas. Often used to assess progress in school.

Diagnostic Tests: Tests that assess problem areas of learning. Often used to assess learning disabilities.

Readiness Tests: Tests that measure one’s readiness for moving ahead in school. Often used to assess readiness to enter first grade.

Aptitude Testing: Tests that measure what one is capable of learning.

Intellectual and Cognitive Functioning: Tests that measure a broad range of cognitive functioning in the following domains: general intelligence, intel- lectual disabilities, giftedness, and changes in overall cognitive functioning. Includes intelligence testing that leads to an “IQ” score and neuro- psychological assessment that assesses changes in cognitive functioning over time.

Cognitive Ability Tests: Tests that measure a broad range of cognitive ability. These tests are usually based on what a student has learned in school and are useful in making predictions about the future (e.g., whether an individual might succeed in college).

Special Aptitude Tests: Tests that measure one aspect of ability. Often useful in determining the likelihood of success in a vocation (e.g., a mechanical aptitude test to determine potential success as a mechanic).

Multiple Aptitude Tests: Tests that measure many aspects of ability. Often useful in deter- mining the likelihood of success in a number of vocations.

Personality Assessment: Tests in the affective realm used to assess habits, temperament, likes and dislikes, character, and sim- ilar behaviors.

Objective Personality Testing: Multiple choice and true-false tests that assess various aspects of personality. Often used to increase client insight, to identify psychopathology, and to assist in treatment planning.

Projective Personality Tests: Tests that present a stimulus to which individuals can respond. Personality factors are interpreted based on the individual’s response. Often used to identify psy- chopathology and to assist in treatment planning.

Interest Inventories: Tests that measure likes and dislikes as well as one’s personality orien- tation toward the world of work. Generally used in career counseling.

Informal Assessment Instruments: Often developed by the user, these tests tend to assess broad areas of ability or personality and tend to be specific to the testing situation.

Observation: Observing an individual in order to develop a deeper understanding of one or more specific behaviors (e.g., observing a stu- dent’s acting-out behavior in class or assessing a client’s ability to perform eye-hand coordina- tion tasks as a means of determining potential vocational placement).

Rating Scales: Scales developed to assess any of a number of attributes of the examinee. Can be rated by the examinee or someone who knows the examinee well (e.g., rating a faculty member’s teaching ability or a student’s ability to make empathic responses).

Classification Methods: A tool whereby an individual identifies whether he or she has, or does not have, specific attributes or character- istics (e.g., from a list, checking adjectives that seem to be most like you).

16 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

QUESTIONS TO CONSIDER WHEN ASSESSING INDIVIDUALS It is clear that today’s assessment instruments have widespread applications. With the knowledge that many individuals have used assessment instruments for less than honorable reasons (e.g., the Eugenics Movement), it is critical that we remain vigilant about the use of such instruments. Keeping this in mind, we should contin- ually be asking ourselves some important questions regarding the use of assessment instruments, including the following:

1. How valid is the information gained from assessment instruments, and how should that information be applied?

2. How do assessment instruments invade an individual’s privacy, and does the government have, at times, the right to insist that an individual be assessed?

3. Can the use of some assessment instruments lead to labeling, and what are the implications for individuals who are “labeled”?

4. Are assessment procedures used to foster equality for all people, or do they tend to reinforce existing societal divisions based on class?

SUMMARY We began this chapter by defining assessment and pointing out that assessment encompasses a broad range of techniques, including testing. We noted that modern-day assessment has been greatly influenced by the long history of assessment. Going back to 2200 B.C.E., we pointed out that the Chinese developed one of the first widely used tests and that hundreds of years ago the Greeks assessed intellectual and physical ability of men for state service. As we neared the modern era of testing, we pointed out that individuals such as Esquirol examined the relationship between language ability and intelligence, while others, such as Seguin, looked at the relationship between

motor control and intelligence. We noted that Darwin’s theory of evolution spurred on others such as Galton, Wundt, Cattell, and Hall to examine individual differences, a focus that would be critical to the nature of assessment.

We next pointed out that some of the first ability tests were neuropsychological assessments and intelligence testing. We noted that interest in brain functioning could be traced by to early Egypt, around 2500 B.C. and moving forward in time, we noted that the late 1800s saw Alfred Binet develop the first intelligence test. Later revised by Terman, the Stanford-Binet compared an individual’s chronological age to his or her

Environmental Assessment: A naturalistic and systemic approach to assessment in which information about clients is collected from their home, work, school, or other places through observation, self-reports, and checklists.

Records and Personal Documents: Items such as diaries, personal journals, genograms, and school records that are examined

to gain a broader understanding of an individual.

Performance-Based Assessment: The evalu- ation of an individual using informal assessment procedures based on real-world activities that are not highly loaded for cognitive skills. These procedures are seen as an alternative to stan- dardized testing (e.g., a portfolio).

© Ce ng ag e Le ar ni ng

CHAPTER 1 History of Testing and Assessment 17

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

mental age. We pointed out that intelligence test- ing was also sometimes used as one aspect of neuropsychological assessments, which became more important with the advent of brain- imaging techniques and the use of additional methods to assess the relationship of brain func- tion to behavior.

As the chapter continued, we noted that the early 1900s saw the development of group tests of ability, including the Army Alpha and Army Beta, achievement tests in the schools, and multiple aptitude tests. We then noted that during the 1900s, individuals such as Galton, Yerkes, and Terman were influential in the Eugenics Movement. This misguided venture attempted to use test data to show intellectual differences among cultural groups and ended up influencing government policy, including laws regarding who would be allowed to immigrate to the United States.

We pointed out that many of the first person- ality tests paralleled the development of ability tests. For instance, the early 1900s saw Thorndike, Miner, and Strong research the area of vocational assessment and develop some of the first interest inventories. Kraeplin developed a crude word association test, and Woodworth developed his

“personal data sheet,” which many consider the precursor of modern-day personality inventories. These early assessment instruments were soon followed by the development of one of the first projective tests, a word association test by Carl Jung who sought to uncover what kinds of reac- tions caused complexes. Other early projective tests included Rorschach’s Inkblot Test and Murray’s Thematic Apperception Test. As the twentieth century continued, a number of informal assessment instruments were developed, including observational techniques, rating scales, classifica- tion schemes, records and personal documents, environmental assessments, and performance- based assessments.

Near the end of the chapter, we reviewed the definitions of the various assessment categories, including those assessment techniques that are found in the following domains: ability testing (achievement and aptitude), personality testing, and informal assessment instruments. The chapter concluded by highlighting a number of important issues in assessment, including test validity, inva- sion of privacy, caution regarding labeling, and the importance of assuring that assessment proce- dures foster equality.

CHAPTER REVIEW 1. Identify some of the ancient precursors to

assessment. 2. Identify some of the precursors to modern-

day assessment during the 1800s. 3. Discuss how the work of Darwin, Galton,

Wundt, and Cattell has influenced the devel- opment of modern-day testing.

4. Identify some of the individuals involved and describe the precursors to modern-day intel- ligence testing.

5. What was the Eugenics Movement, and how did it influence government policy in the United States?

6. Identify some of the early group tests of ability, the main players in their development, and their uses.

7. Identify some of the early personality tests and the main players involved in their development.

8. Describe the contributions of some of the early developers of projective testing.

9. Draw three diagrams (see Figures 1.3–1.5) that list the various kinds of achievement testing, aptitude testing, personality assess- ment, and informal assessment. Define each type of assessment category on your diagrams.

10. Make a list of every historical figure discussed in this chapter, and define each person’s contribution to testing and assessment.

18 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

REFERENCES Armstrong, T. (2006). The best schools: How human

development research should inform educational practice. Alexandria, VA: Association for Supervision and Curriculum Development.

Association for Assessment and Research in Counseling. (2012). Standards for multicultural assessment (4th ed.). Retrieved from http://www.theaaceonline .com/AACE-AMCD.pdf

Benjamin, L. T. (2008). Psychology before 1900. In S. F. Davis, & W. Buskit (Eds.), 21st century psy- chology: A reference handbook (Vol. 1, pp. 2–11). Thousand Oaks, CA: Sage Publications.

Binet, A., & Simon, T. (1916). The development of intelligence in children: The Binet-Simon Scale. (E. S. Kite, Trans.). Baltimore: Williams & Wilkins Company.

Cohen, J., Swerdlik, M., & Sturman, E. (2012). Psychological testing and assessment: An introduc- tion to tests and measurement (8th ed.). Columbus, OH: McGraw Hill.

Cowdery, K. (1926). Measurement of professional atti- tudes: Differences between lawyers, physicians and engineers. Journal of Personnel Research, 5, 131–141.

Doyle, K. (1974). Theory and practice of ability testing in ancient Greece. Journal of the History of Behav- ioral Sciences, 10, 202–212. doi:10.1002/1520-6696 (197404)10:2<202::AID-JHBS2300100208>3.0. CO;2-Q

Drummond, R. J. (2009). Assessment procedures for counselors and helping professionals (7th ed.). Upper Saddle River, NJ: Pearson.

DuBois, P. (1970). A history of psychological testing. Boston: Allyn and Bacon.

Esquirol, J. (1838). Des maladies mentales considerees sous les rapports medical, hygienique, et medicole- gal. Paris: Bailliere.

Frontline. (1999). The secrets of the SATs. Retrieved November 23, 2008, from http://www.pbs.org/wgbh /pages/frontline/shows/sats/etc/script.html

Galton, F. (1879). Psychometric facts. Nineteenth Century, 5, 425–433.

Geisinger, K. (1994). Psychometric issues in testing stu- dents with disabilities. Applied Measurement in Education, 7, 121–140.

Geisinger, K. (2000). Psychological testing at the end of the millennium: A brief historical review. Profes- sional Psychology, Research, and Practice, 31, 117–119. doi:10.1037/0735-7028.31.2.117

Gillham, N. (2001). A life of Sir Francis Galton: From African exploration to the birth of eugenics. New York: Oxford University Press.

Goodwin, C. J. (2008). Psychology in the 20th century. In S. F. Davis & W. Buskit (Eds.), 21st century psy- chology: A reference handbook (Vol. 1, pp. 12–20). Thousand Oaks, CA: Sage Publications.

Gould, S. J. (1996). The mismeasure of man (revised and expanded). New York: Norton.

Haney, W. (1981). Validity, vaudeville, and values: A short history of social concerns over standardized testing. American Psychologist, 36, 1021–1033 .doi:10.1037/0003-066X.36.10.1021

Hebben, N., & Milberg, W. (2009). Essentials of neuropsychological assessment (2nd ed.) John Wiley & Sons: New York.

Higgins, L. T., & Sun, C. H. (2002). The development of psychological testing in China. International Journal of Psychology, 37, 246–254.

Hunsley, J. (2002). Psychological testing and psycho- logical assessment: A closer examination. American Psychologist, 57, 139–140. doi:10.1037/0003-066X .57.2.139

Joint Committee on Testing Practices. (2004). Code of fair testing practices in education. Washington, DC: American Psychological Association.

Jolly, J. L. (2008). Lewis Terman: Genetic study of genius—elementary school students. Gifted Child Today, 31(1), 27–33.

Jones, L. V. (2007). Some lasting consequences of U.S. psychology programs in World Wars I and II. Multivariate Behavioral Research, 42, 593–608. doi:10.1080/00273170701382542

Jung, C. G. (1969) Studies in word-association: Experi- ments in the diagnosis of psychopathological condi- tions carried out at the psychiatric clinic of the University of Zurich under the direction of C. G. Jung. New York: Russell & Russell. (Original work published in 1918)

Jung, C. G., & Riklin, F. (1904). Untersuchungen iiber assoziationen gesunder. Journal fur Psychologie uiind Neurologie, 3, 55–83.

Juve, J. (2008). Testing and assessment. In S. F. Davis, & W. Buskit (Eds.), 21st century psychology: A ref- erence handbook (Vol. 1, pp. 383–391). Thousand Oaks, CA: Sage Publications.

Kerr, M. S. (2008). Psychometrics. In S. F. Davis, & W. Buskit (Eds.), 21st century psychology: A reference

CHAPTER 1 History of Testing and Assessment 19

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

handbook (Vol. 1, pp. 374–382). Thousand Oaks, CA: Sage Publications.

Lezak, M., Howieson, D., Bigler, E. D. & Tranel, D. (2012). Neuropsychological assessment (5th ed.). New York: Oxford University Press

McKean, K. (1985). Intelligence: New ways to measure the wisdom of man. Discover Magazine, 6(10), 28. Reprinted by permission of Disney Publications.

Miner, J. (1922). An aid to the analysis of vocational interest. Journal of Educational Research, 5, 311–323.

Murdoch, S. (2007). IQ: A smart history of a failed idea. Hoboken, NJ: John Wiley & Sons.

Neukrug, E., & Schwitzer, A. (2006). Skills and tools for today’s counselors and psychotherapists: From natural helping to professional counseling. Pacific Grove, CA: Brooks/Cole.

Nicolas, S., Gyselinck, V., Murray, D., & Bandomir, C. A. (2002). French descriptions of Wundt’s labora- tory in Leipzig in 1886. Psychological Research, 66, 208–214.

Parsons, F. (1989). Choosing a vocation. Garrett Park, MD: Garrett Park. (Original work published in 1909)

Roback, A. (1961). History of psychology and psychi- atry. New York: Philosophical Library.

Rorschach, H. (1942). Psycho diagnostics. (P. V. Lamkau, Trans.). Bern, Switzerland: Verlag Hans Huber.

Ryan, J. J. (2008). Intelligence. In S. F. Davis, & W. Buskit (Eds.), 21st century psychology: A refer- ence handbook (Vol. 1, pp. 413–421). Thousand Oaks, CA: Sage Publications.

Sokal, M. M. (1987). Introduction: Psychological testing and historical scholarship-questions, contrasts, and context. In M. M. Sokal (Ed.), Psychological testing and American society: 1890-1930 (pp. 1–20). New Brunswick, NJ: Rutgers University Press.

Storr, A. (1973). C. G. Jung. New York: The Viking Press.

Strong, E. (1926). An interest test for personnel man- agers. Journal of Personnel Research, 5, 194–203.

Watson, R. (1968). The great psychologists from Aristotle to Freud. Philadelphia: Lippincott.

Zenderland, L. (1987). The debate over diagnosis: Henry Herbert Goddard and the medical acceptance of intelligence testing. In M. M. Sokal (Ed.), Psycho- logical testing and American society: 1890-1930 (pp. 46–74). New Brunswick, NJ: Rutgers University Press.

Zusne, L. (1984). Biographical dictionary of psychol- ogy. Westport, CT: Greenwood.

20 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 2Ethical, Legal, andProfessional Issues in Assessment

A friend of mine, a psychologist, complained to the ethics committee of her state licensing board that a colleague of hers was continually writing incompetent test reports. She reported her colleague as incompetent, and referred to her ethical code which stated the following:

“Psychologists provide services, teach and conduct research with populations and in areas only within the boundaries of their competence, based on their education, train- ing, supervised experience, consultation, study or professional experience. (American Psychological Association, [APA], 2010, Section 2.01a), and that Psychologists under- take ongoing efforts to develop and maintain their competence” (Section 2.03).

My friend felt a sense of self-righteousness for having reported this psychologist. Much to her surprise, the colleague decided to report my friend on the grounds that she had not first approached her and tried to work out the situation informally. The colleague of my friend was not sanctioned, but my friend was!

“When psychologists believe that there may have been an ethical violation by another psychologist, they attempt to resolve the issue by bringing it to the attention of that individual …” (Section 1.04).

Whether it’s psychology, counseling, or social work, the professional associations’ ethical codes are in agreement in suggesting that if you have a concern about a col- league’s ethical behavior, if reasonable, you should first speak to that person prior to reporting him or her (American Counseling Association [ACA], 2005; APA, 2010; National Association of Social Workers, [NASW], 2008).

21

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Whether a concern about testing or another area, ethical dilemmas are by their very nature, complex. Thus, in this chapter we examine the sometimes per- plexing ethical, legal, and professional issues that confront individuals who are assessing others. We start by examining some of the ethical codes and standards that drive ethical decision-making and then we review the ethical decision- making process as it relates to assessment. We next examine some of the many laws and legal issues that have affected testing and assessment over the years. The chapter concludes with an overview of some professional issues, including a quick look at associations and accreditation bodies that focus on assessment, a review of forensic evaluation, the importance of viewing testing as a holistic process, a look at a number of important cross-cultural issues related to assess- ment, and a discussion about the fear that many clinicians have concerning their use of assessment techniques.

ETHICAL ISSUES IN ASSESSMENT This section of the chapter offers (1) a brief overview of some aspects of the American Counseling Association’s (ACA) and American Psychological Association’s (APA) ethical codes that highlight assessment, (2) a quick look at the many standards that have been developed to help guide the practitioner in the ethical use of assess- ment procedures, and (3) a discussion on the process of ethical decision-making.

Overview of the Assessment Sections of the ACA and APA Ethical Codes The ethical codes of our professional associations provide us with guidelines about how to respond under certain situations. For instance, both the ACA (2005) and APA (2010) codes include guidelines that specifically address issues of testing and assessment. The following discussion summarizes some of the more salient aspects of the assessment portions of these codes, including choos- ing appropriate assessment instruments, competence in the use of assessment instruments, confidentiality, cross-cultural sensitivity, informed consent, invasion of privacy, proper diagnosis, release of test data, test administration, test secu- rity, and test scoring and interpretation. Please review the ethical guidelines of the professional associations, the Web sites of which can be found in Appendix A, as well as the actual portions of the assessment guidelines of ACA and APA that can be found in Appendix B.

Choosing Appropriate Assessment Instruments Ethical codes stress the importance of professionals choosing assessment instruments that show test worthi- ness, which has to do with the reliability (consistency), validity (the test measuring what it’s supposed to measure), cross-cultural fairness, and practicality of a test. Professionals must take appropriate actions when issues of test worthiness arise during an assessment so that the results of the assessment are not misconstrued. Test worthiness and how to choose an appropriate assessment instrument will be examined in detail in Chapter 5.

Ethical codes Professional guidelines for appropriate behavior

Attend to test worthiness in choosing assessment instruments

22 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Competence in the Use of Tests Competence to use tests accurately is another aspect that is stressed in the codes. The codes declare that professionals should have adequate knowledge about testing and familiarity with any test they may use. To establish who is qualified to give specific tests, the APA, in 1954, adopted a three-tier system for establishing test user qualifications. Although APA reevaluated this system and currently provides rather extensive guidelines of test user qualifications (see Turner, DeMers, Fox, & Reed, 2001), many test publishers continue to use the original three-tier system, although they take into account the more involved guidelines. The original system labeled some tests according to the following three levels:

• Level A tests are those that can be administered, scored, and interpreted by responsible nonpsychologists who have carefully read the test manual and are familiar with the overall purpose of testing. Educational achievement tests fall into this category.

• Level B tests require technical knowledge of test construction and use and appropriate advanced coursework in psychology and related courses (e.g., statistics, individual differences, and counseling).

• Level C tests require an advanced degree in psychology or licensure as a psychologist and advanced training/supervised experience in the particular test. (American Psychological Association, 1954, pp. 146–148)

More specifically, an individual with a bachelor’s degree who has some knowl- edge of assessment and is thoroughly versed with the test manual can give a Level A test and, under some limited circumstances, a Level B test. For example, teachers can administer most survey battery achievement tests. The master’s-level helping professional who has taken a basic course in tests and measurement can administer Level B tests. For example, counselors can administer a wide range of personality tests, including most interest inventories and many objective person- ality tests. However, they cannot administer tests that require additional training, such as most individual tests of intelligence, most projective tests, and many diagnostic tests. These Level C tests are reserved for those who have a minimum of a master’s degree, a basic testing course, and advanced training in the special- ized test (e.g., school psychologists, learning disabilities specialists, clinical and counseling psychologists, and master’s-level counselors who have gained addi- tional training) (ACA, 2003).

Confidentiality Whether giving one test or conducting a broad assessment of a client, keeping information confidential is a critical part of the assessment process and follows similar guidelines to how one would keep information con- fidential in a therapeutic relationship. The importance of keeping information confidential is a professional responsibility, but sometimes it is contradicted by the law. Thus, professionals must sometimes grapple with difficult dilemmas pitting professional responsibility against the law (Remley & Herlihy, 2014). When can one reveal confidential information? Although one should always check the laws of the state in which one practices, practitioners are generally

Competence in using tests Requires adequate knowledge and train- ing in administering an instrument

Confidentiality Ethical guideline to protect client information

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 23

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

on fairly solid professional and legal ground in revealing information under the following conditions:

1. If a client is in danger of harming himself or herself or someone else; 2. If a client is a minor or is legally incompetent and the law states that parents or

guardians have a right to information about that person; 3. If a client asks you to break confidentiality (e.g., your testimony is needed in court); 4. To defend oneself against charges filed by a client; 5. If you are asked by the court to break confidentiality and privileged communi-

cation does not exist (privileged communication is when there is a statute ensuring clients that information shared is confidential);

6. To reveal information about your client to clerical help, colleagues, or supervi- sor to benefit the client;

7. When you have a written agreement from your client to reveal information to specified sources (e.g., the court has asked you to send a test report to them).

Cross-Cultural Sensitivity In reference to cross-cultural sensitivity, the codes tend to focus on the potential biases of assessment procedures when selecting, administering, and interpreting assessment instruments. In addition, they stress the importance of professionals being aware of and attending to the effects of age, color, cultural identity, disability, ethnicity, gender, religion, sexual orienta- tion, and socioeconomic status on test administration and test interpretation. Later in this chapter, and in Chapter 5, we discuss this important topic in greater detail.

Informed Consent Informed consent involves ensuring that clients obtain infor- mation about the nature and purpose of all aspects of the assessment process and for clients to give their permission to be assessed. Although there are times when informed consent does not need to be obtained, such as when it is implied (e.g., giving an achievement test in schools), or when testing is mandated by the courts (e.g., custody battles), generally test administrators should give their clients infor- mation about the assessment process and have obtained the client’s permission to be assessed.

Invasion of Privacy The codes generally acknowledge that, to some degree, all tests invade one’s privacy and they highlight the importance of clients under- standing how their privacy might be infringed upon. Concerns about invasion of privacy are lessened if clients give informed consent, have real choice in accepting or refusing testing, and know the limits of confidentiality, as noted earlier.

Proper Diagnosis Due to the delicate nature of diagnoses, the codes emphasize the important role that professionals play when deciding which assessment techni- ques to use in forming a diagnosis for a mental disorder and the ramifications of making such a diagnosis.

Release of Test Data Test data can be, and have been, misused. Thus, the codes assert that data should only be released to others if clients have given

Cross-cultural sensitivity Ethical guideline to protect clients from discrimination and bias in testing

Informed consent Permission given by client after assessment process is explained

Invasion of privacy Testing is an invasion of a person’s privacy

Proper diagnosis Choose appropriate assessment techni- ques for accurate diagnosis

Release of test data Test data are protected—client release required

24 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

their consent. The release of such data is generally only given to individuals who can adequately interpret the test data and to those who will not misuse the information.

Test Administration As you might guess, the codes reinforce the notion that tests should be administered in a manner that is in accord with the way that they were established and standardized. Alterations to this process should be noted and interpretations of test data adjusted if testing conditions were not ideal.

Test Security The codes remind professionals that it is their responsibility to make reasonable efforts to ensure the integrity of test content and the security of the test itself. Professionals should not duplicate tests or change test material with- out the permission of the publisher.

Test Scoring and Interpretation Ethical codes highlight the fact that when scoring tests and interpreting their results, professionals should reflect on how test worthiness (reliability, validity, cross-cultural fairness, and practicality) might affect the results. Results should always be couched in terms that reflect any potential problems with test interpretation. (These issues will be discussed in detail in Chapter 5).

Test administration Use established and standardized methods

Test security Ensure integrity of test content and test itself

Test scoring and interpretation Take into consider- ation problems with tests

Cl ay

Be nn et t/ ©

19 98

Th e Ch ris tia n Sc ie nc e M on ito r (w w w .C SM

on ito r.c om

). Re pr in te d w ith

pe rm is si on .

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 25

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Standards for Responsible Testing Practices Partly a response to criticism of the manner in which tests have been used and abused over the years, a number of standards in assessment to help guide practices have been developed in recent years. These standards delineate the proper use of assessment techniques in educational, agency, and private practice settings. As an example of one standard, we have included the Code of Fair Testing Practices (JCTP, 2004) in Appendix C. Table 2.1 defines the purpose of a number of well- known standards and offers ways of accessing the standards should you want to examine them in more detail:

• Standards for Qualifications of Test Users • Responsibilities of Users of Standardized Tests • Standards for Multicultural Assessment • Code of Fair Testing Practices • Rights and Responsibilities of Test Takers • Standards for Educational and Psychological Testing • Competencies for testing in School Counseling; Mental Health Counseling;

Marriage, Couples, and Family Counseling; Career Counseling, and Substance Abuse Counseling

TABLE 2.1 Standards of Responsible Test Usage

Name of Standard Purpose of Standard

Developed/ Endorsed by* and Citation for

Standards for Qualifications of Test Users (5 pages)

Addresses qualifications of counselors to use assessment instruments in following areas: test theory, test construction, test worthiness, test statistics, administration and interpre- tation of tests, cross-cultural issues, knowledge of standards and ethical codes.

ACA (2003)

Responsibilities of Users of Standardized Tests (RUST) (5 pages)

“The intent of ‘RUST’ is to help counselors and other edu- cators implement responsible testing practices … in the following areas: qualifications of test users, technical knowl- edge, test selection, test administration, test scoring, inter- preting test results, communicating test results.” (p. 1)

AARC (2003)

Standards for Multicultural Assessment (8 pages)

Focuses on advocating for proper use of instruments and taking steps to address systemic issues, ensuring those who conduct assessments understand test biases and the bar- riers facing diverse populations; choosing appropriate instruments for diverse clients, considering clients’ cultural backgrounds when administering and scoring instruments, providing culturally accurate interpretation of instruments, and ensuring cultural competence in the training and super- vision of testing.

AARC (2012a)

(Continued)

Standards in Assessment Used to further ethical application of assessment techniques

26 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Making Ethical Decisions Because ethical codes can be limiting in their ability to guide a practitioner who is faced with a thorny ethical dilemma, it is important that other avenues are avail- able to aid in ethical decision-making. For instance, some practitioners might use moral models in guiding their ethical decision-making process. One moral model, described by Kitchener (1984, 1986; Urofsky, Engels, & Engebretson, 2008), suggests that there are six critical moral principles one should consider when making difficult ethical decisions. They include autonomy, which has to do with protecting the independence, self-determination, and freedom of choice of clients; nonmaleficense is the concept of “do no harm” when working with clients; benefi- cence is related to promoting the good of society, which can be at least partially

TABLE 2.1 Standards of Responsible Test Usage (continued)

Name of Standard Purpose of Standard

Developed/ Endorsed by* and Citation for

Code of Fair Testing Practices (6 pages)

Intended to ensure that professionals “provide and use tests that are fair to all test takers regardless of age, gender, dis- ability, race, ethnicity, national origin, religion, sexual orientation, linguistic background, or other personal charac- teristics” (para. 1). Provides guidance for test developers and test users in four areas: developing and selecting appropriate tests, administering and scoring tests, reporting and inter- preting test results, and informing test takers.

JCTP (2004)

Rights and Responsibilities of Test Takers (10 pages)

Delineates and clarifies “the expectations that test takers may reasonably have about the testing process, and the expectations that those who develop, administer, and use tests may have of test takers.” ( p. l)

JCTP (1998)

Standards for Educational and Psychological Testing (194 pages)

“… to provide criteria for the evaluation of tests, testing practices, and the effects of test use” (p. 2). Extensive explanation of testing practices in three broad areas: test construction, evaluation, and documentation; fairness in testing; and testing applications.

AERA/APA/ NCME (1999)

Competencies in the following counseling areas: School; Mental Health; Marriage, Couple, and Family; Career; and Substance Abuse

Each competency tends to be a few-pages long and covers such things as how to adequately choose assess- ment instruments, proper administration and scoring of instruments, appropriate interpretation of results, diversity issues and assessment, and the use of instruments in diagnosis.

ASCA AMHCA IAMFC NCDA IAAOC AARC (2012b)

*ACA ¼ American Counseling Association; AARC ¼ Association of Assessment in Research in Counseling; JCTP ¼ Joint Commission on Testing Practices; AERA ¼ American Educational Research Association; APA ¼ American Psychological Association; ASCA ¼ American School Counseling Association; AMHCA ¼ American Mental Health Counselors Association; IAMFC ¼ International Association of Marriage and Family Counseling; NCDA ¼ National Career Development Association; NCME ¼ National Council on Measurement in Education; IAAOC ¼ International Association of Addiction and Offender Counselors.

Moral model Consider moral princi- ples involved in ethical decision-making

© Ce ng ag e Le ar ni ng

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 27

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

accomplished by promoting the client’s well-being; justice refers to providing equal and fair treatment to all clients; fidelity is related to maintaining trust (e.g., keeping conversations confidential) in the counseling relationship and being committed to the client within that relationship; and veracity has to do with being truthful and genuine with the client, within the context of the counseling relationship. Consider these principles if you had just assessed a client and had determined that she poten- tially might cause harm to her children. How might each of these moral principles play into the decisions you make regarding your client. For instance, after consider- ing each of the principles, how and to whom would you communicate your results? To make things a bit more complicated, Remley and Herlihy (2014) go on to note that the culture of the client might impact your understanding of your results and how you apply the principles. Autonomy, for individuals from some cultures, may have to do with individual behaviors whereas individuals from other cultures might view autonomy within the context of their extended family or community. As you can see, ethical decision-making can be a complex and difficult process.

In addition to the moral model just noted, a number of other ethical decision- making models exist (Neukrug, 2012). One hands-on, practical, problem-solving model espoused by Corey, Corey, Corey, and Callanan (2015) suggests that the practi- tioner go through the following eight steps when making complex ethical decisions:

1. Identify the problem or dilemma 2. Identify the potential issues involved 3. Review the relevant ethical guidelines 4. Know the applicable laws and regulations 5. Obtain consultation 6. Consider possible and probable courses of action 7. Enumerate the consequences of various decisions 8. Decide on what appears to be the best course of action

Finally, in addition to the moral and practical models mentioned earlier, some suggest that regardless of the approach one takes in ethical decision-making, the ability to make wise ethical decisions may well be influenced by the clinician’s level of ethical, moral, and cognitive development (Linstrum, 2005; Neukrug, Lovell, & Parker, 1996) (see Exercise 2.1). Those who are at higher levels of cogni- tive development, they state, view ethical decision-making in more complex ways than others. Certainly, this has broad implications for the training that takes place in clinical programs, as it would be hoped that students are challenged to make decisions that are comprehensive and thoughtful (McAuliffe & Eriksen, 2010).

LEGAL ISSUES IN ASSESSMENT A number of laws about testing have been passed and lawsuits resolved that impinge on the use of tests. Most of these legal decisions speak to issues of confidentiality, fairness, and test worthiness (reliability, validity, practicality, and cross-cultural fairness) (Swen- son, 1997; Greene & Heilbrun, 2011). In this section of the chapter, we summarize some of the more important laws that have been passed and legal cases resolved over the years, including the Family Educational Rights and Privacy Act (FERPA); the Health Insurance Portability and Accountability Act (HIPAA); privileged communication laws;

Corey, Corey, Corey, and Callanan Recommend an eight- step decision-making model

Wise ethical decisions reflect higher cognitive development

Laws about testing Created to protect the client or examinee

28 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

the Freedom of Information Act; various Civil Rights Acts (1964 and amendments); the Americans with Disabilities Act (ADA) (PL 101-336); the Individuals with Disabilities Education Improvement Act (IDEA), which was an expansion of PL 94-142; Section 504 of the Rehabilitation Act; and the Carl Perkins Act (PL 98-24).

The Family Education Rights and Privacy Act (FERPA) of 1974 Passed in 1974 and sometimes called the Buckley Amendment, FERPA assures the right of privacy of student records, including test records, and applies to almost all schools that receive federal funds (K-12 and institutions of higher education) (U.S. Department of Education, n.d.a). The law also gives the right to parents of students, or to the stu- dents themselves once they reach the age of 18 and are beyond high school, to review their records. If parents or the eligible student believe that school records are incorrect or misleading, they have the right to challenge them. If the school decides not to change the records, the parent or eligible student can ask for a formal hearing.

The law also states that if records are to be released, schools must receive written permission from parents or the eligible student, except in some specific circumstances (e.g., school officials acting in the educational interest of the student, access by organizations, to comply with a subpoena, for purposes of school evaluation, and other reasons).

The Health Insurance Portability and Accountability Act (HIPAA) In 1996 Congress passed the HIPAA, and three main rules that underscore HIPAA’s purpose subsequently went into effect (privacy, transaction rule, and secu- rity rules) (U.S. Department of Health and Human Services, n.d.; Zuckerman, 2008). In general, HIPAA restricts the amount of information that can be shared

Exercise 2.1 Making Ethical Decisions

After reading the section on ethical decision-making, in small groups in class or as a homework assignment, use the moral principles identified and Corey’s model of ethical decision-making to decide on your course of action in the following situation. Share your answers in class. You will have the opportunity to respond to additional vignettes in Exercise 2.2 near the end of the chapter.

Situation: You have been asked to provide a broad per- sonality assessment of a 17-year-old high school student who has been truant from school and has a history of acting-out behaviors. After meeting with her, conducting a clinical interview, and administering a number of

projective tests and an MMPI-2 (an objective personality test), you find evidence that an uncle, who is five years older than the client, had sexually molested her when she was 12 years old. In addition, you believe that the young woman has been involved in some petty crimes, such as shoplifting and stealing audio equipment from the school. In writing your report, what should you include regarding her being molested and the crimes she has allegedly committed? Are you obligated to report this case to Child Protective Services? Do you have any obligation to contact the police? What are your obligations to this young person’s parents, to the school, and to society?

FERPA Affirms right to access test records in the school

HIPAA Ensures privacy of medical and counsel- ing records

© Ce ng ag e Le ar ni ng

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 29

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

without clients’ consent and allows clients to have access to their records, except for process notes used in counseling. In fact, HIPAA requires agencies to show how they have complied with this act. Although HIPAA has far-reaching implica- tions for the privacy and security of records and how confidential transactions are made, relative to testing, some of the more salient features of the law includes ensuring that (1) clients are given information regarding their right to privacy, (2) agencies have clear procedures established to protect client privacy, (3) employees are trained to protect the privacy of clients, and (4) client records are secure.

Privileged Communication Privileged communication is a conversation conducted with someone who state or federal law identifies as a person with whom conversations may legally be kept confidential (i.e., therapist-patient, attorney-client, doctor-patient, clergy-penitent, husband-wife, etc.). In the case of clinicians, the goal of the law is to encourage the client to engage in conversations without fear that the clinician will reveal the contents of the conversation (e.g., in a court of law), thus ensuring the privacy and efficacy of the counseling relationship. The privilege belongs to the client, and only the client can waive that privilege (Remley & Herlihy, 2014). Privileged communi- cation should not be confused with confidentiality, which is the ethical, not legal, obligation of the counselor to keep conversations confidential (Glosoff, Herlihy, & Spence, 2000).

A 1996 ruling upheld the right to privileged communication (Jaffee v. Redmond, 1996). In this case, the Supreme Court upheld the right of a licensed social worker to keep her case records confidential. Describing the social worker as a “therapist” and “psychotherapist,” the ruling has bolstered the right of all licensed therapists to have privileged communication (Remley, Herlihy, & Herlihy, 1997; Remley & Herlihy, 2014) (see Box 2.1).

The Freedom of Information Act Originally passed in 1967 and amended in 2002, this law assures the right of indi- viduals to access their federal records, including test records, if an individual makes a request in writing (U.S. Department of Justice, 2011). All states have enacted sim- ilar laws that assure the right to access state records.

Civil Rights Acts (1964 and Amendments) In 1964 the federal government passed the first of a number of laws to ban a broad spectrum of discrimination in the United States. This first act was focused on banning racial segregation in schools, public places, and employment. Since that time, a number of far-reaching amendments of the original law have been passed. Relative to testing, the laws assert that any test used for employment or promotion must be shown to be suitable and valid for the job in question. If not, alternative means of assessment must be provided. Differential test cutoffs are not allowed unless such scores can be shown to be based on valid educational principles (see Box 2.2).

Privileged communication Legal right to maintain privacy of conversation

Jaffee v. Redmond Affirms privileged communication laws

Freedom of Information Act Affirms right to access federal and state records

Civil Rights Act Test must be valid for job in question

High-Stakes Testing The use of cutoff scores in important client decisions

30 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

BOX 2.1 Jaffee v. Redmond

“Mary Lu Redmond, a police officer in a village near Chicago, responded to a ‘fight in progress’ call at an apartment complex on June 27, 1991. At the scene, she shot and killed a man she believed was about to stab another man he was chasing. The family of the man she had killed sued Redmond, the police department, and the village, alleging that Officer Redmond had used excessive force in violation of the deceased’s civil rights.

When the plaintiff’s lawyers learned that Redmond had sought and received counseling from a licensed social worker employed by the village, they sought to compel the social worker to turn over her case notes and records and testify at the trial. Redmond and the social worker claimed that their communications were privileged under an Illinois statute.

They both refused to reveal the substance of their counseling sessions even though the trial judge rejected their argument that the communications were privileged. The judge then instructed jurors that they could assume that the information withheld would have been unfavorable to the policewoman, and the jury awarded the plaintiffs $545,000.” (Remley et al., 1997, p. 214)

After a series of appeals, the Supreme Court heard the case on February 26, 1996. The court decided that the licensed therapist did indeed hold privilege and that the judge’s instruction to the jury was therefore unwarranted.

BOX 2.2 The Civil Rights of High Stakes Testing

Out of concern that college athletes were being recruited to play sports without regard for how they might do academically, in 1986, the National Collegiate Athletic Association (NCAA) implemented a rule stating that high school athletes would have to exceed specific scores on SAT or ACT college entrance exams to be eligible for athletic-based scholarships and interscho- lastic competition (Waller, 2003). This affected the abil- ity of many high school athletes, particularly African- American athletes, to obtain sports scholarships and be accepted into college. On the other hand, a higher percentage of African-American athletes were found to graduate college after this rule went into effect.

With federal civil rights law asserting that the use of tests could not disproportionately affect one group unless such use could be shown to be educationally sound, the NCAA was sued by African-American stu- dents for discrimination. Although the court has never definitively settled the cases, the uproar over the cutoff scores seems to have affected how the NCAA applied the rule. Now the NCAA uses a floating score: The higher your high school GPA, the lower the SAT or ACT score that the athlete is allowed (Paskus, 2012).

With continuing concerns about the use of cutoff scores for athletic scholarships, advancing in public school, and admission to college, the battle in the courts will probably continue. What do you think about the use of such high-stakes testing?

1. Should SAT and ACT scores be used to prevent student athletes from playing sports so that they can spend more time on their studies?

2. Should standardized test scores (e.g., SATs, ACTs, GREs, state achievement tests) be used to make educational decisions about placement into grade level, college, or graduate schools? Why or why not?

3. As the result of laws like No Child Left Behind that mandate that school systems assure that all stu- dents achieve, there has been a slight increase in the average scores of minorities on some tests. What are the drawbacks and benefits of the continued use of such high-stakes testing?

4. How might the use of high-stakes testing affect those who are involved in the test-preparing process, such as teachers and administrators?

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 31

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Americans with Disabilities Act (ADA) (PL 101-336) In the United States, close to 19% of noninstitutionalized Americans have a disability (U.S. Census Bureau, 2012), including 15.6 million (6.4%) with a sight or hearing disability, 38.3 million (12.6%) with a severe disability, and 12.3 million (4.4%) with a disability that affects their ability to dress, bathe, or get around their home. Despite such large numbers, prejudice and discrimination has been commonplace for individuals with disabilities. As a result, in 1990, Congress passed the ADA.

Implemented in 1992, the ADA bans discrimination in employment, public ser- vices and public transportation, public accommodations, and telecommunications for individuals with disabilities. One of its biggest impacts was in the area of employment, where the law prohibits employers “from discriminating against qual- ified individuals with disabilities in job application procedures, hiring, firing, advancement, compensation, job training, and other terms, conditions and privi- leges of employment” (U.S. Equal Employment Opportunity Commission, 2008, para. 1). Because testing is often used in the hiring and promotion of individuals in jobs, the law asserts that accommodations must be made for individuals with disabilities who are taking tests for employment and that testing must be shown to be relevant to the job in question.

Individuals with Disabilities Education Act (IDEA) In 1975 Public Law 94-142 was passed and guaranteed the right to a free and pub- lic education, with appropriate accommodations, for children with disabilities (U.S. Department of Education, 2007a). Over the years, the law was expanded and even- tually became known as the Individuals with Disabilities Education Act (IDEA). IDEA is far reaching and impacts individuals with disabilities from birth through 21 in many ways. Part B of the law focuses on testing in the schools and assures the right of a student between the age of 3 and 21 to be tested, at the school sys- tem’s expense, if he or she is suspected of having a disability that interferes with learning (U.S. Department of Education, n.d.b). The law asserts that if a student is assessed and found to have a disability, schools must ensure that the student is given accommodations for his or her disability and taught within the “least restric- tive environment,” which is often a regular classroom.

Students who are suspected of having a disability are generally referred for medi- cal, psychological, communication, and/or vision and hearing evaluations. Any assess- ment that is conducted should be cross-culturally appropriate, and parents should be informed of the kinds of tests being used and give permission for the child to be tested. If the assessments indicate that special education eligibility requirements are met and if parental consent given, an individualized education plan (IEP) must be developed within 30 calendar days of the evaluation (Assistance to the States for the Education of Children with Disabilities, 2011). The plan should address what services are needed and how they can be provided within the least restrictive environment. The team that develops the plan often includes the parent(s), the child’s teacher(s), a district represen- tative who is able to provide or oversee the delivery of special education services, the child (when appropriate), representatives from the evaluation team, possible service providers, and other relevant people chosen by the parent or school.

ADA Accommodations for testing must be made

IDEA Assures right to be tested for learning disabilities in schools

32 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Section 504 of the Rehabilitation Act This act applies to organizations and employers that receive financial assistance from the federal government and was established to create a “level playing field” and prevent discrimination based on disability. Based on this law, employers and organizations cannot:

• Deny qualified individuals the opportunity to participate in or benefit from federally funded programs, services, or other benefits.

• Deny access to programs, services, benefits, or opportunities to participate as a result of physical barriers.

• Deny employment opportunities, including hiring, promotion, training, and fringe benefits, for which they are otherwise entitled or qualified. (U.S. Department of Health and Human Services, 2006, p. 2)

Relative to assessment, any instrument used to measure appropriateness for a pro- gram or service must measure the individual’s ability, not be a reflection of his or her disability.

Carl Perkins Career and Technical Education Improvement Act of 2006 This federally funded program, which gives financial grants to the states, assures that individuals in six identified “special populations” can have access to voca- tional assessment, counseling, and placement so that these individuals will be more likely to succeed in technical education programs and in their careers (U.S. Depart- ment of Education, 2007b). Originally passed in 1984, this fourth reiteration of the Carl Perkins Act applies to the following populations: (a) individuals with disabil- ities; (b) individuals from economically disadvantaged families, including foster children; (c) individuals preparing for nontraditional fields; (d) single parents, including single pregnant women; (e) displaced homemakers; and (f) individuals with limited English proficiency.

PROFESSIONAL ISSUES This section of the chapter begins by highlighting two professional associations in the field of assessment: The Association for Assessment and Research in Counseling (AARC) and Division 5 of the American Psychological Association: Evaluation, Measurement, and Statistics. We then briefly mention two accrediting bodies that focus on assessment. This is followed by a discussion about the growing profes- sional field of forensic evaluation. Then, we underscore the importance of recogniz- ing that assessment works best when it is a holistic process; that is, when consideration is given for the use of multiple procedures while assessing an individ- ual. The use of assessment techniques with minorities and women is then highlighted, and the chapter concludes with a discussion about importance of embracing the use of testing and assessment.

Section 504 Assessment for pro- grams must measure ability—not disability

Carl Perkins Act Ensures access to vocational assess- ment, counseling, and placement

Major professional associations AARC of ACA and Division 5 of APA

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 33

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Professional Associations Although there are literally dozens of professional associations in the field that might pique your interest (see Appendix A), few specifically focus on assessment. Although we encourage you to join the professional association with which you feel the closest affinity, if you have a strong interest in assessment you might want to consider joining the following two associations.

The Association for Assessment and Research in Counseling (AARC), a divi- sion of ACA, “is an organization of counselors, educators, and other professionals that advances the counseling profession by promoting best practices in assessment, research, and evaluation in counseling” (AARC, 2012c, para. 1). Consider joining AARC if you are interested in testing, diagnosis, the training, and supervision of those who do assessment, and if you are fascinated by the process of developing and validating assessment products and procedures. AARC publishes two journals, Counseling Outcome Research and Evaluation (CORE) and Measurement and Evaluation in Counseling and Development (MECD). AARC also publishes a newsletter called NewsNotes.

Division 5 of the American Psychological Association: Evaluation, Measurement, and Statistics is devoted to “promoting high standards in both research and practical application of psychological assessment, evaluation, measurement, and statistics” (APA, 2013a, para. 1). Division 5 publishes two journals, Psychological Assessment, which is geared toward assessment, and Psychological Methods, which is research- oriented. Division 5 also publishes a quarterly newsletter, The Score, which focuses on business issues of the association as well as new issues in the field.

Accreditation Standards of Professional Associations A number of the professional associations have accreditation standards that specifi- cally speak to curriculum issues in the area of assessment. Such standards help to establish a common core of experience for students who enter their programs, regardless of the institution in which the program is housed. Thus, we find organiza- tions such as the APA (2009), the National Association of School Psychologists (2010), and the Council for the Accreditation of Counseling and Related Educational Programs (CACREP, 2013), setting standards that drive the curriculum for their graduate programs. In fact, much of what is covered in this book is a result of the authors examining these curriculum standards and trying to ensure that we cover them as fully as possible. If you get a chance, you may want to visit the Web sites listed in the references and examine the standards for each of these organizations.

Forensic Evaluations Increasingly, counselors, psychologists, and other mental health professionals are becoming involved in a wide range of evaluation activities related to law enforce- ment and legal issues. To do this competently, they must have accurate knowledge of assessment, legal issues, and ethical issues when dealing with such cases (Packer, 2008; Patterson, 2006; Roesch & Zapf, 2013). What areas do such specialists deal with? Packer identifies a wide range of common areas that such experts will often testify about (see Box 2.3).

Accreditation bodies Assist in setting cur- riculum standards (APA, NASP, and CACREP)

Forensic evaluations Completed by forensic health evaluators and forensic psychologists

34 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Forensic evaluators and forensic psychologists need to know how to conduct forensic evaluations, which include the use of specific tests for the situation at hand, interviewing techniques that are focused on the goals of the court case, knowledge of ethical and legal issues relevant to expert testimony and the specific case, and how to write forensic reports that will be used in court. Although many of the tests and interviewing techniques used in this book can be applied to such evaluations, know- ing how and when to use them is crucial. In addition, other assessment procedures are often involved in forensic evaluations. Thus, to assure the accuracy of such eva- luations, the National Board of Forensic Evaluators (NBFE, 2009) certifies mental health counselors, marriage and family counselors, and social workers as forensic health evaluators. In addition, psychologists interested in forensics can now do resi- dencies in forensic psychology and also become board certified in forensic psychol- ogy (American Board for Forensic Psychology, 2013). Finally, Specialty Guidelines for Forensic Psychologists have been established by the APA “to improve the quality of forensic psychological services; enhance the practice and facilitate the systematic development of forensic psychology; encourage a high level of quality in professional practice; and encourage forensic practitioners to acknowledge and respect the rights of those they serve” (APA, 2013b, para. 3).

Assessment as a Holistic Process As we stress throughout this book, assessment of clients is much broader than sim- ply giving an individual a test. In fact, one should generally “avoid using a single test score as the sole determinant of decisions …” and one should “Interpret test scores in conjunction with other information about individuals” (JCTP, 2004, Section C-5). A good assessment will often involve a number of different kinds of instruments, including formal tests, informal assessment instruments, and a clinical interview. In addition, tests tend not to measure certain aspects of a person, such as an individual’s motivation, intention, and focus of attention and one should always consider how cultural factors can impact test results (AARC, 2012a; Preston,

BOX 2.3 Common Areas for Forensic Health Evaluators and Forensic Psychologists

Civil Criminal

• Child custody • Juvenile waiver (from juvenile to adult court)

• Civil commitment • Juvenile sentencing

• Deprivation of parental rights • Competence to stand trial

• Divorce mediation • Competence to be sentenced

• Employment litigation • Competence to waive Miranda rights

• Guardianship • Insanity defense

• Personal injury • Diminished capacity

• Testamentary capacity • Sentencing (including special issues related to capital sentencing)

• Harassment and discrimination • Evaluation of violence risk

Holistic process The importance of using multiple measures

© Ce ng ag e Le ar ni ng

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 35

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

2005). Ultimately, the examiner should take into account test scores, multiple assessment procedures, individual traits, and cultural factors that can impact results when making suggestions and decisions about the client. This holistic process can help us obtain a broader and more accurate view of the client than if we were to rely on just one test (Mislevy, 2004; Moss, 2004).

In addition, assessment is not a static process; instead, it should be seen as continu- ous and ongoing. The assessment of a client occurs at a specific point in time; however, if we believe that a person can change, that point represents only a small sample of an individual’s total functioning. Viewing assessment in this way allows us to understand that an individual’s cognitive functioning and personality may, and probably will, change significantly throughout the life span. Seeing a client in this manner helps us develop treatment plans for the client in the here and now, while reminding us not to be held hostage to labels and diagnoses—for the individual may, and probably will, change.

Cross-Cultural Issues in Assessment [Mental health professionals] need to examine the quality and usefulness of available assessment activities relative to how such assessments may negatively affect clients, no matter the population group to which they may belong. Needed is a reorientation of assessment practices to promote the development of human talent and resources for all who are assessed. (Loesch, 2007, p. 204)

A review of contemporary articles on testing and assessment will quickly reveal that there has been increased attention to bias in testing and how it disproportionately affects minorities and women. Lawsuits have questioned the accuracy of some tests (see Box 2.2), laws have been passed preventing the use of other tests, and research has been conducted that demonstrates the negative impact that some tests can have on minorities and women. In response to these problems, as you saw earlier in this chap- ter, ethical codes and standards that address assessment always include statements con- cerning how to choose, administer, and interpret tests and assessment instruments for minorities and women. Such codes and standards have stressed that individuals who give tests and use assessment procedures should, at the minimum, keep the following in mind (International Test Commission, 2008; AARC, 2012a):

1. Assume that all tests hold some bias. 2. Be in touch with your own biases and prejudices. 3. Only use tests that have been shown to be constructed using sound research

procedures. 4. Only use tests that have good validity and reliability. 5. Know there are times when it is appropriate to test and times when it is not. 6. Know how to choose good tests that are relevant to the situation at hand. 7. Know how to administer, score, and interpret tests within the cultural context

of the client. 8. View assessment as a holistic process, and whenever possible and reasonable,

include interviews, formal testing, and informal testing. 9. Know and consider the implications that testing may have for the client.

10. Advocate for clients when tests are shown to be biased. 11. Treat people humanely during the assessment process.

Assessment is a snapshot; clients continually change

Cross-cultural issues Bias in tests, bias in examiner, understanding affects on client

36 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

In Chapter 5, we will discuss a number of matters related to cross-cultural issues and test worthiness. We will highlight the fact that a test should accurately measure a construct across such attributes as class, race, religion, or gender. We will discuss how the U.S. court system, public law, federal acts, and constitu- tional amendments have all supported the notion that testing must be fair for all groups of people and free from bias. We will note how tests cannot be used to track students and that as a result of the Supreme Court case of Griggs v. Duke Power Company (1971), tests used for hiring and advancement at work must show that they can predict job performance for all groups. Also, as we have highlighted in this chapter and will discuss in Chapter 5, a number of laws have been passed that impinge on the use of tests and assert the rights of all individuals to be tested fairly. In the future, we are likely to see an increased emphasis on understanding the inherent bias in tests, the creation of new tests that are less biased, and new efforts to properly administer, score, and interpret tests with the understanding that they will, to some degree, have bias.

Exercise 2.2 Making Ethical Decisions

Review the situations below, and then using the moral principles identified in the chapter, Corey’s models of ethical decision-making, and your knowledge of legal and professional issues decide on your probable course of action. Share your answers with the rest of the class.

Situation 1: A graduate-level mental health professional with no training in career development is giving interest inventories as she counsels individuals for career issues. Can she do this? Is this ethical? Professional? Legal? If this professional happened to be a colleague of yours, what, if anything, would you do?

Situation 2: During the taking of some routine tests for promotion, a company learns that there is a high prob- ability that one of the employees is abusing drugs and is a pathological liar. The firm decides not to promote him and instead fires him. He comes to see you for counsel- ing because he is depressed. Has the company acted ethically? Legally? What responsibility do you have toward this client?

Situation 3: An African-American mother is concerned that her child may have an attention deficit problem. She goes to the teacher, who supports her concerns, and they go to the assistant principal requesting testing for a possible learning disorder. The mother asks if the child could be given an individual intelligence test that can screen for such problems, and the assistant principal

states, “Those tests have been banned for minority stu- dents because of concerns about cross-cultural bias.” The mother states that she will give her permission for such testing, but the assistant principal says, “I’m sorry, we’ll have to make do with some other tests and with observation.” Is this ethical? Professional? Legal? If you were a school counselor or school psychologist and this mother came to see you, what would you tell her?

Situation 4: A test that has not been researched to show that it is predictive of success for all potential graduate students in social work is used as part of the program’s admission process. When challenged on this by a poten- tial student, the head of the program states that the test has not been shown to be biased and the program uses other, additional criteria for admission. You are a member of the faculty at this program. Is this ethical? Professional? Legal? What is your responsibility in this situation?

Situation 5: An individual who is physically challenged and wheelchair bound applies for a job at a national fast-food chain. When he goes in to take the test for a mid-level job at this company, he is told that he cannot be given this test because it has not been assessed for its predictive ability for individuals with his disability. You are hired by the company to do the testing. What is your responsibility, if any, to this individual and to the company?

© Ce ng ag e Le ar ni ng

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 37

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Embracing Testing and Assessment Procedures It’s not a secret that many students are fearful of, and maybe even dreading, a course in assessment (Davis, Chang, & McGlothlin, 2005; Wood & D’Agostino, 2010). And, if many students have this fear, one has to wonder if they also leave a training program with an aversion to using tests and assessment procedures in their professional practice. If you are one of those students, then this speaks directly to you. If you’re not, our hats go off to you!

The bottom line: If you have an aversion to using assessment procedures, and you act on it, then you are not being fair to your clients. This is because testing and assessment procedures can help clients know who they are, why they behave the way they do, what they are good at, and what choices to make in their future. It can help clinicians decide on the most appropriate treatment plans, and it can help friends and families support our clients in loving and effective ways. Probably, multiple assessment procedures should always be considered for our clients if we are to give them the most effective services possible. To not do so is bringing your fears and biases into the helping relationship.

SUMMARY This chapter examined ethical, legal, and profes- sional issues involved in assessment. We began by exploring ethical concerns, and first summarized ACA’s and APA’s assessment sections of their eth- ical codes, including (1) choosing assessment instruments, (2) competence in the use of tests, (3) confidentiality, (4) cross-cultural sensitivity, (5) informed consent, (6) invasion of privacy, (7) proper diagnosis, (8) release of test data, (9) test administration, (10) test security, and (11) test scoring and interpretation. We next briefly highlighted the purposes of a number of important standards in assessment, including Standards for Qualifications of Test Users, Responsibilities of Users of Standardized Test, Standards for Multicultural Assessment, Code of Fair Testing Practices, Rights and Responsibilities of Test Takers, Standards for Educational and Psychological Testing, and competencies in school counseling; mental health counseling; marriage, couple and family counseling; career counseling; and substance abuse counseling.

As the section on ethical issues continued, we discussed the fact that good ethical decision- making is more than just relying on a code of ethics. We presented a moral model of ethical decision-making, which has to do with the

importance of focusing on the autonomy, or self-determination, of the client; nonmaleficence, or ensuring that you “do no harm”; beneficence, or promoting the well-being of society; justice, or providing equal and fair treatment; fidelity, or being loyal and faithful to your clients; and veracity, which means dealing honestly with your clients. We noted that one’s use of moral models should take into account the cultural dif- ferences of clients. We then presented Corey’s eight-step model, which includes identifying the problem or dilemma, identifying the potential issues involved, reviewing the relevant ethical guidelines, knowing the applicable laws and reg- ulations, obtaining consultation, considering possible and probable courses of action, enumer- ating the consequences of various decisions, and deciding on what appears to be the best course of action. We suggested that the ability to make wise ethical decisions might well be influenced by the counselor’s level of ethical, moral, and cognitive development.

The next part of the chapter examined impor- tant legal issues involving the use of tests, including the FERPA; the HIPAA; privileged communication laws; the Freedom of Information Act; various Civil Rights Acts (1964 and amendments); the

Use assessment instruments! They allow clients better understanding of self

38 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

ADA (PL 101-336); the IDEA, which was an expansion of PL 94-142; Section 504 of the Rehabilitation Act; and the Carl Perkins Act (PL 98-24). Most of these laws have to do with issues of confidentiality, fairness, and test worthiness. This section of the chapter also offered a discussion of some of the implications of high-stakes testing.

In this chapter, we examined a number of professional issues. First, we highlighted two professional associations in the field of assess- ment: the AARC, which is a division of ACA; and Division 5 of the APA. We next highlighted three accrediting bodies in the helping profes- sions that address curriculum standards in assessment: APA, NASP, and CACREP. We then talked about the increasingly important field of forensics and the significance of becom- ing properly trained as a forensic evaluator or psychologist if one is to conduct such evaluations accurately. We next spoke of the importance of viewing assessment as a holistic process that should involve assessing the client in multiple ways, including the use of formal tests, informal

assessment instruments, and the clinical inter- view. We pointed out that assessment should be viewed as an ongoing process because people change as they live and learn.

Relative to cross-cultural issues, we stressed that individuals who give tests and use assessment procedures should, at the minimum, only use tests that have been shown to be constructed using sound research procedures; only use tests that have good validity and reliability; know there are times when it is appropriate to test and times it is not; know how to choose good tests that are relevant to the situa- tion at hand; know how to administer, score, and interpret tests within the cultural context of the cli- ent; view assessment as a holistic process; know and consider the implications that testing may have for the client; advocate for clients when tests are shown to be biased; and treat people humanely during the assessment process. We noted that Chapter 5 would expand on the discussion of cross-cultural issues in assessment. Finally, we suggested that clinicians should remain open to using assessment instruments as part of their counseling process.

CHAPTER REVIEW 1. Relative to testing and assessment, identify

and discuss some of the major themes addressed in ethical codes.

2. In addition to the ethical codes, other standards have been developed to guide individuals in test selection, administration, and interpretation. Describe some of these standards.

3. Describe APA’s levels of test user competence.

4. Describe the moral model and Corey’s problem-solving model of ethical decision- making.

5. Compare and contrast how individuals at different cognitive developmental levels would go about making ethical decisions.

6. Identify some of the major legal issues that have affected the selection, administration,

and interpretation of assessment instruments.

7. List and briefly describe two professional associations that specifically address assess- ment issues.

8. What is the role of accreditation in the deliv- ery of curriculum content in the area of assessment?

9. What are some of the unique issues involved in forensic evaluation?

10. What is needed for a good assessment of an individual?

11. Why should assessment procedures often be considered when working with clients?

12. Describe the importance of having an under- standing of cross-cultural issues when using assessment procedures.

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 39

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

REFERENCES American Board for Forensic Psychology. (2013).

Apply for APFP/ABPP forensic board certification. Retrieved from http://abfp.com/certification.asp

American Counseling Association. (2003). Standards for qualifications of test users. Alexandria, VA: Author. Retrieved from http://aarc-counseling.org /resources

American Counseling Association. (2005). Code of ethics. Retrieved from http://www.counseling.org /knowledge-center/ethics

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (1999). Standards for educational and psychological testing. Washington, DC: AERA.

American Psychological Association. (1954). Technical recommendations for psychological tests and diag- nostic techniques. Washington, DC: Author.

American Psychological Association. (2009). Guide- lines for principles and accreditation of programs in professional psychology. Washington, DC: Author.

American Psychological Association. (2010). Ethical principles of psychologists and code of conduct. Retrieved from http://www.apa.org/ethics/code /index.aspx

American Psychological Association. (2013a). Division 5: Evaluation, Measurements and Statistics. Retrieved from http://www.apa.org/about/division /div5.html

American Psychological Association. (2013b). Specialty guidelines for forensic psychology. Retrieved from http://www.apa.org/practice/guidelines/forensic- psychology.aspx?item=1

Assistance to the States for the Education of Children with Disabilities. (2011). 34, C.F.R. pt III.

Association for Assessment and Research in Counseling. (2003). Responsibilities of users of standardized tests (RUST) (3rd ed.). Retrieved from http:// aarc-counseling.org/resources

Association for Assessment and Research in Counseling. (2012a). Standards for multicultural assessment (4th ed.). Retrieved from http://aarc-counseling.org /resources

Association for Assessment and Research in Counseling. (2012b). AARC sponsored Assessment standards and statements. Retrieved from http://aarc-counseling .org/resources

Association for Assessment and Research in Counseling. (2012c). AARC statements of purpose. Retrieved from http://aarc-counseling.org/about-us

Corey, G., Corey, M. S., Corey, C. & Callanan, P. (2015). Issues and ethics in the helping professions (9th ed.). Belmont, CA: Brooks/Cole, Cengage Learning.

Council for Accreditation of Counseling and Related Education Programs (CACREP). (2013). Resources: Download a copy of the 2009 CACREP Standards. Retrieved from http://www.cacrep.org/template /index.cfm

Davis, K. M., Chang, C. Y., & McGlothlin, J. M. (2005). Teaching assessment and appraisal: Human- istic strategies and activities for counselor educators. Journal of Humanistic Counseling, Education and Development, 44, 94–101. doi:10.1002/j.2164-490X .2005.tb00059.x

Glosoff, H. L., Herlihy, B., & Spense, B. E. (2000). Privileged communication in the counselor-client relationship. Journal of Counseling and Development, 78, 454–462. doi:10.1002/j.1556-6676.2000.tb01 929.x

Greene, E., & Heilbrun, K. (2011). Wrightsman’s psychology and the legal system (7th ed.). Belmont, CA: Cengage.

Griggs V. Duke Power Company, 401 U. S. 424 (1971). International Test Commission. (2008). Guidelines.

Retrieved from http://www.intestcom.org/guide- lines/index.php

Jaffee v. Redmond, 518 U.S. 1 (U.S. Supreme Ct., 1996). Joint Committee on Testing Practices. (1998). Rights

and responsibilities of test takers: Guidelines and expectations. Retrieved from http://aarc-counseling .org/resources

Joint Committee on Testing Practices. (JCTP). (2004). Code of fair testing practices in education. Washing- ton, DC: American Psychological Association.

Kitchener, K. S. (1984). Intuition, critical evaluation and ethical principles: The foundation for ethical decisions in counseling psychology. The Counseling Psychologist, 12(3), 43–45. doi:10.1177/0011000084123005

Kitchener, K. S. (1986). Teaching applied ethics in counselor education: An integration of psychologi- cal processes and philosophical analysis. Journal of Counseling and Development, 64, 306–311. doi:10.1177/0011000084123005

40 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Linstrum, K. S. (2005). The effects of training on ethi- cal decision making skills as a function of moral development and context in master-level counseling students. Dissertation Abstracts International Section A: Humanities & Social Sciences, 65(9-A), 3289.

Loesch, L. C. (2007). Fair access to and the use of assess- ment in counseling. In. C. C. Lee (Ed.), Counseling for social justice (2nd ed., pp. 201–222). Alexandria, VA: American Counseling Association.

McAuliffe, G., & Eriksen, K. (Eds.). (2010). Handbook of counselor preparation. Thousand Oaks, CA: Sage Publications.

Mislevy, R. J. (2004). Can there be reliability without “reliability?” Journal of Educational and Behavioral Statistics, 29, 241–244. doi:10.3102/107699860290 02241

Moss, P. (2004). The meaning and consequences of reliability. Journal of Educational and Behavioral Statistics, 29, 245–249. doi:10.3102/107699860290 02245

National Association of School Psychologists (NASP). (2010). NASP Professional standards/training. Retrieved from http://www.nasponline.org/standards /2010standards.aspx

National Association of Social Workers. (NASW). (2008). Code of ethics. Retrieved from http://www .naswdc.org/pubs/code/default.asp

National Board of Forensic Evaluators. (2009). About NBFE. Retrieved from http://www.nbfe.net/overview .php

Neukrug, E. (2012). The world of the counselor. Belmont, CA: Brooks/Cole.

Neukrug, E., Lovell, C, & Parker, R. (1996). Employ- ing ethical codes and decision-making models: A developmental process. Counseling and Values, 40, 98–106. doi:10.1002/j.2161-007X.1996.tb00843.x

Packer, I. K. (2008). Specialized practice in forensic psychology: Opportunities and obstacles. Professional Psychology: Research and Practice, 39, 245–249. doi:10.1037/0735-7028.39.2.245

Paskus, T. S. (2012). A summary and commentary on the quantitative results of current NCAA academic reforms. Journal of Intercollegiate Sports, 5, 41–52.

Patterson, J. (2006). Attaining specialty credentials can enhance any counsellor’s career and knowledge base. Counseling Today, 49(5), 20.

Preston, P. (2005). Testing children: A practitioner’s guide to the assessment of mental development in infants and young children. Kirkland, WA: Hogrefe & Huber.

Remley, T. P., & Herlihy, B. (2014). Ethical and pro- fessional issues in counseling (4th ed.). Boston: Pearson.

Remley, T. P., Herlihy, B., & Herlihy, S. B. (1997). The U.S. Supreme Court decision in Jaffe v. Redmond: Implications for counselors. Journal of Counseling and Development, 75, 213–218. doi:10.1002 /j.1556-6676.1997.tb02335.x

Roesch, R., & Zapf, P. A. (2013). Forensic assessments in criminal and civil law. New York: Oxford University Press.

Swenson, L. (1997). Psychology and law for the helping professions (2nd ed.). Pacific Grove, CA: Brooks/Cole.

Turner, S. M., DeMers, S. T., Fox, H. R., & Reed. G. M. (2001). APA’s guidelines for test user qualifications: An executive summary. American Psychologist, 56, 1099–1113. doi:10.1037/0003-066X.56.12.1099

Urofsky, R., Engels, D., & Engebretson, K. (2008). Kitchener’s principle ethics: Implications for counseling practice and research. Counseling and Values, 53, 67–78. doi:10.1002/j.2161-007X.2009 .tb00114.x

U.S. Census Bureau. (2012). Americans with disabil- ities: 2010. Retrieved from http://www.census.gov /prod/2012pubs/p70-131.pdf

U.S. Department of Education. (n.d.a). Family Educational Rights and Privacy Act (FERPA). Retrieved from http://www2.ed.gov/policy/gen/guid /fpco/ferpa/index.html

U.S. Department of Education (n.d.b). Building the leg- acy: IDEA 2004. Retrieved from http://idea.ed.gov/

U.S. Department of Education. (2007a). History: Twenty-five years of progress in educating children with disabilities through IDEA. Retrieved from http://www2.ed.gov/policy/speced/leg/idea/history .html

U.S. Department of Education (2007b). Carl D. Perkins Career and Technical Education Act of 2006. Retrieved from http://www.ed.gov/policy/sectech /leg/perkins/index.html#intro

U.S. Department of Health and Human Services. (n.d.). Understanding health information privacy. Retrieved from http://www.hhs.gov/ocr/privacy/hipaa /understanding/index.html

U.S. Department of Health and Human Services, (2006). Fact sheet: Your rights under section 504 of the Rehabilitation Act. Retrieved from http:// www.hhs.gov/ocr/504.html

U.S. Department of Justice. (2011). What is FOIA? Retrieved from http://www.foia.gov/about.html

CHAPTER 2 Ethical, Legal, and Professional Issues in Assessment 41

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

U.S. Equal Employment Opportunity Commission. (2008). Facts about the Americans with Disabilities Act. Retrieved from http://www.eeoc.gov/facts/fs-ada .html

Waller, J. M. (2003). A necessary evil: Proposition 16 and its impact on academics and athletics in the NCAA. DePaul journal of Sports Law and Contemporary Problems, 1, 189–206.

Wood, C. & D’Agostino, J. V. (2010). Assessment in counseling: A tool for social justice work. In

M. J. Ratts, R. L. Toporek, & J. A. Lewis (Eds.), ACA advocacy competencies: A social justice frame- work for counselors (pp. 151–159). Alexandria, VA: American Counseling Association.

Zuckerman, E. (2008). The paper office: Forms, guide- lines, and resources to make your practice work eth- ically, legally, and profitably (4th ed.). New York: Guilford Press.

42 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 3Diagnosis in theAssessment Process*

It was 1975, and part of my job as an outpatient therapist at a mental health center entailed answering the crisis counseling phones every ninth night. I would sleep at the center and answer a very loud phone that would ring periodically throughout the night, usually with a person in crisis on the other end. Every once in a while, a former client of the center would call in and start to read aloud from his case notes, which he had stolen from the center. Parts of these notes were a description of his diagnosis from what was then the second edi- tion of the Diagnostic and Statistical Manual (DSM-II). In a sometimes angry, sometimes funny tone, he would read these clinical terms that were supposed to be describing him. I could understand his frustration when reading these notes over the phone, as in some ways, the diagnosis seemed removed from the person—a label. “Was this really describing the person, and how was it helpful to him?” I would often wonder.

(Ed Neukrug)

An important aspect of the clinical assessment and appraisal process is skillful diagnosis. Today, the use of diagnosis permeates the mental health professions, and although there continues to be some question as to its helpfulness, it is clear that making diagnoses and using them in treatment planning has become an integral part of what all mental health professionals do. Thus, in this chapter we examine the use of diagnosis.

We begin this chapter by discussing the importance of diagnosis in the assess- ment process and then provide a brief overview of the history of the Diagnostic and Statistical Manual of Mental Disorders (DSM) and its evolution over the past

*Updated and revised by Katherine A. Heimsch & Gina B. Polychronopoulos

Diagnosis adds clarity to the assessment process

43

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

several decades. We then introduce the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5) and note some of the differences from previous versions, such as the use of a single axis and factors that now come into play when making and reporting diagnosis. Next, we highlight the DSM-5 diagnostic categories and follow up with other important considerations when making a diag- nosis, such as medical concerns, psychosocial and environmental concerns, and cross-cultural issues. There are several case studies and exercises that will help to hone some of your diagnostic skills. At the end of the chapter, we relate the impor- tance of formulating a diagnosis within the overall assessment process.

THE IMPORTANCE OF DIAGNOSIS • John is in fifth grade and has been assessed as having a conduct disorder and

attention-deficit/hyperactivity disorder (ADHD). John’s mother has panic disorder and is taking antianxiety medication. His father has bipolar and is taking lithium. Jill is John’s school counselor. John’s individualized education plan (IEP) states that he will work with Jill individually and in small groups to address behavior, attention, and social skills deficits. Jill must also periodically consult with John’s mother, father, and teachers.

• Tamara has just started college. After breaking up with her boyfriend, she became severely depressed and unable to concentrate on her schoolwork; her grades have dropped from As to Cs. She comes to the college counseling center and sobs during most of her first session with her counselor. She admits having always struggled with depression but states that “This is worse than ever; I need to get better if I am going to stay in school. Can you give me any medication to help me so I won’t have to drop out?”

• Benjamin goes daily to the day treatment center at the local mental health center. He seems fairly coherent and generally in good spirits. He has been hospitalized for schizophrenia on numerous occasions and now takes risperdone to relieve his symptoms. He admits to Jordana, one of his counse- lors, that he doesn’t take his medication because he believes that computers have consciousness and are conspiring through the World Wide Web to take over the world. His insurance company pays for his treatment. He will not receive treatment unless Jordana specifies a diagnosis on the insurance form.

As you can see from these examples, diagnosis is an essential tool for profes- sionals in a wide range of settings. In fact, current research suggests that up to 20% of all children and adults struggle with a diagnosable mental disorder each year (Centers for Disease Control and Prevention [CDC], 2013; Substance Abuse and Mental Health Services Administration, 2012), and approximately 50% of adults in the United States will experience mental illness in their lifetime (CDC, 2011). There- fore, all persons serving in helping roles will encounter persons dealing with a mental disorder and will need to be familiar with a common diagnostic language to best serve these individual and to effectively communicate with other professionals. Today, the importance of an accurate diagnosis is related to a number of changes that have occurred over the past years. Some of these include the following:

1. Interventions and accommodations for children with emotional, behavioral, and learning disorders are now required by federal and state laws (e.g., PL94-142, Individuals with Disabilities Education Act [IDEA]) and a diagnosis is generally

There are many reasons to diagnose

44 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

necessary if professionals are to identify students with such disorders. Today, teachers, school counselors, school psychologists, child study team members, and other school professionals are often the first to recognize and diagnose young people with these disorders.

2. Today, a diagnosis is viewed as one aspect of holistically understanding the client. Along with testing, interviews, and other measures, it can be used to help conceptualize client problems and assist in the accurate development of treatment plans.

3. Due to laws like the Americans with Disabilities Act (e.g., U.S. Department of Justice, n.d.), employers are now required to make reasonable accommodations for individuals with disabilities, including those with mental disorders. Mental health professionals must know about diagnosis if they are to help individuals maintain themselves at work and assist employers in understanding the condi- tions of individuals with mental disorders.

4. In the past 50 years, a mental disorder diagnosis has generally become mandatory if medical insurance is to reimburse for treatment. Accurate diagnosing is important because the insurance carrier often allows only a certain number of treatments per particular diagnosis.

5. The diagnostic nomenclature of the DSM has increasingly become an essential and effective way of communicating with community partners who may be part of the client’s same treatment team (e.g., other mental health professionals, doctors, representatives of the legal system).

6. It has become increasingly evident that accurately and appropriately commu- nicating a mental health diagnosis to a client can help the individual under- stand his or her prognosis and aid in forming reasonable expectations for treatment.

These items show why it is important for a wide range of professionals to understand diagnosis. Although the DSM-IV-TR (4th ed. with text revisions; American Psychiatric Association [APA], 2000) had been the most well-known diagnostic classification sys- tem, with the recent release of DSM-5 (APA, 2013), a revised nomenclature was devel- oped. But what is the DSM and how does it work?

THE DIAGNOSTIC AND STATISTICAL MANUAL (DSM): A BRIEF HISTORY

Derived from the Greek words dia (apart) and gnosis (to perceive or to know), the term diagnosis refers to making an assessment of an individual from an outside, or objective, viewpoint (Segal & Coolidge, 2001). One of the first attempts to classify mental illness occurred during the mid-1800s when the U.S. Census Bureau started counting the incidence of “idiocy” and “insanity” (Smith, 2012). However, it was not until 1943 that a formal classification system called the Medical 203 was devel- oped by the U.S. War Department (Houts, 2000). Revised over the next few years, in 1952 this publication became the basis for APA’s first DSM (DSM-I), which included 106 diagnoses in three broad categories (APA, 1952; Houts, 2000). In 1968 DSM-II was released (APA, 1968), which created 11 diagnostic categories with 185 discrete diagnoses and included a large increase in childhood diagnoses. In an effort to improve the science behind diagnosis as well as increase the

DSM-I First edition published in 1952 with three broad categories

DSM-III Introduced multiaxial diagnosis in 1980

CHAPTER 3 Diagnosis in the Assessment Process 45

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

compatibility with the American Medical Association’s International Classification of Disease (ICD) manual, the third edition of the DSM was released in 1980 (APA, 1980), which included 265 diagnoses and a multiaxial approach to diagnosis. In 1994 DSM-IV was released, and in 2000 an additional text revision of DSM-IV became available (DSM-IV-TR) and contained 365 diagnoses (APA, 1994, 2000). Although there were many critics of the DSM-IV-TR (Beutler & Malik, 2002; Thyer, 2006; Zalaquett, Fuerth, Stein, Ivey, & Ivey, 2008), it became the most widely utilized diag- nostic classification system for mental health disorders (Seligman, 1999, 2004). A DSM-IV diagnosis consisted of five axes that included clinical disorders, personality disorders and mental retardation, medical conditions, psychosocial and environmental factors, and a global assessment of functioning (GAF) scale (see Table 3.1).

The practice of utilizing the multiaxial diagnostic system allowed mental health professionals to present a thorough description of clients and communicate their concerns and symptoms to other professionals (Neukrug & Schwitzer, 2006). However, there were drawbacks to a multiaxial approach and the DSM-5 moved toward a one-axis approach.

THE DSM-5 The newest diagnostic manual, DSM-5 (APA, 2013), was under development from 1999 to 2013 (Smith, 2012) and was first published in May of 2013. The DSM-5 includes a sleeker, more computer-friendly name, which replaces the Roman numeral tradition of the DSM. Subsequent editions, like computer software, will follow with editions 5.1, 5.2, 5.3, and so on. In addition to the print version of DSM-5, an online component (www.psychiatry.org/dsm5) is now available for supplemental materials such as assessment measures, but it also includes related news articles, fact sheets, and audiovisual materials. Another important change that has been made to the DSM-5 is an effort to align it with the ICD-9, and later, the ICD-10 (release date: October 1, 2014). This serves to unify the diagnostic and billing process between psychological and medical professions. Thus the DSM-5 gives both the ICD-9 and ICD-10 codes, and when making a diagnosis, one may want to list the ICD-9 code first and place the ICD-10 code in parenthesis. Clearly, it is important to know which version of the ICD is being used when making your diagnosis.

TABLE 3.1 Former Five-Axis Diagnostic System

Axis Category Examples

Axis I Clinical disorders Depression, anxiety, bipolar, schizophrenia, etc.

Axis II Personality disorders and mental retardation

Borderline personality disorder, antisocial personality disorder, etc.

Axis III General medical conditions High blood pressure, diabetes, sprained ankle, etc.

Axis IV Psychosocial and environmental factors

Recent loss of job, recent divorce, homelessness, etc.

Axis V Global assessment of functioning A single score from 1 to 100 summarizing one’s functioning and symptoms

DSM-IV-TR Used a five-axis diagnosis

DSM-5 Accepted diagnostic classification system for mental disorders

© Ce ng ag e Le ar ni ng

20 15

46 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Single-Axis vs. Multiaxial Diagnosis Perhaps the most significant change in the DSM-5 was the return to a single-axis diagnosis (APA, 2013; Wakefield, 2013). This was done for a number of reasons. First, the separation of personality disorders to Axis II under DSM-IV gave these disorders undeserved status and the misguided belief that they were untreatable (Good, 2012; Krueger & Eaton, 2010). Clients who met the criteria for an Axis II diagnosis may now find it easier to navigate mental health treatment as they will no longer be seen as having a diagnosis that is more difficult to treat than a host of other disorders. In DSM-5, medical conditions are no longer listed on a separate axis (Axis III in DSM-IV). Thus, they will likely take a more significant role in mental health diagnosis as they can be listed side-by-side with the mental disorder (Wakefield, 2013). Also, psychosocial and environmental stressors, previously listed on Axis IV of DSM-IV, will be listed alongside mental disorders and physical health issues. In fact, DSM-5 has increased the number of “V codes” (Z codes in ICD-10), which are considered nondisordered conditions that sometimes are the focus of treatment and often are reflective of a host of psychosocial and environ- mental issues (e.g., homelessness, divorce, etc.). As for the GAF score, previously on Axis V of DSM-IV, the APA intended to replace this historically unreliable tool with a different scaling assessment altogether. One assessment instrument, now being researched, is the World Health Organization Disability Assessment Schedule 2.0 (WHODAS 2.0). This 36-item, self-administered questionnaire assesses a cli- ent’s functioning in six domains: understanding and communicating, getting around, self-care, getting along with people, life activities, and participation in soci- ety (APA, 2013). Disorders and other assessments that are under review for further research can be found in Section III of the DSM-5.

Making and Reporting Diagnosis In the next section of the chapter, we discuss specific diagnostic categories, but first let’s look at other factors involved in making and reporting diagnoses, including how to order the diagnoses; the use of subtypes, specifiers, and severity; making a provisional diagnosis; and use of “other specified” or “unspecified” disorders.

Ordering Diagnoses Individuals will often have more than one diagnosis, so it is important to consider their ordering. The first diagnosis is called the principal diagnosis. In an inpatient setting, this would be the most salient factor that resulted in the admission (APA, 2013). In an outpatient environment, this would be the reason for the visit or the main focus of treatment. The secondary and tertiary diagnosis should be listed in order of need for clinical attention. If a mental health diagnosis is due to a general medical condition, the ICD coding rules require listing the medical condition first, followed by the psychiatric diagnosis, due to the general medical condition.

Subtypes, Specifiers, and Severity Subtypes for a diagnosis can be used to help communicate greater clarity. They can be identified in the DSM-5 by the instruction “Specify whether” and represent mutually exclusive groupings of symptoms (i.e., the clinician can only pick one). For example, ADHD has three different subtypes to choose from: predominantly inattentive, predominantly hyperactive/impulsive, or a combined presentation. Specifiers, on the other hand, are not mutually exclusive, so

Single-axis diagnosis Attempt to make clini- cal diagnoses and personality disorders on par with medical codes

Medical conditions Can list, using ICD codes, along with DSM diagnosis

Principal diagnosis The reason the person came to treatment is listed first

Subtype “Specify whether”— only choose one

Specifier “Specify if”—pick as many as apply

CHAPTER 3 Diagnosis in the Assessment Process 47

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

more than one can be used. The clinician chooses which specifiers apply, if any, and they are listed in the manual as “Specify if.” The ADHD diagnosis offers only one specifier that is “in partial remission” (APA, 2013, p. 60). Some diagnoses will offer an opportunity to rate the severity of the symptoms. These are identified in the DSM as “Specify current severity.” Referencing the ADHD diagnosis, there are three options of severity: mild, moderate, or severe. The DSM-5 authors have attempted to offer greater flexibility in rating severity through dimensional diagnosis. For exam- ple, some diagnoses offer greater options when rating severity. The Autism Spectrum Disorder has “Table 2 Severity levels of autism spectrum disorder” (APA, 2013, p. 52), which classifies autism on three levels of severity “requiring support,” “requiring substantial support,” and “requiring very substantial support.” Similarly, schizophrenia has the user go to a “Clinician-Rated Dimensions of Psychosis Symp- tom Severity” chart (pp. 743–744) to rate symptoms on a five-point Likert scale. It is easy to see how insurance companies might use severity classification as one method of determining which clients they will fund for treatment. In summary, the three types of specifiers are identified by:

• Subtype: “Specify whether”—only choose one, • Specifier: “Specify if”—pick as many as apply, and • Severity: “Specify current severity”—choose the most accurate level of

symptomology.

Provisional Diagnosis Sometimes, the clinician has a strong inclination that a client will meet the criteria for a diagnosis, but does not yet have enough informa- tion to make the diagnosis. This is when the clinician can make a provisional diagnosis. Once the criteria are later confirmed, the provisional label can be removed. These situations often occur when a client is not able to give an adequate history or further collateral information is required. In addition, there are informal diagnostic labels not listed in the DSM-5 that are helpful in communicating addi- tional information. They are generally found in a diagnostic summary or when communicating informally with other clinicians. They include the following:

• Rule-out—the client meets many of the symptoms but not enough to make a diagnosis at this time; it should be considered further (e.g., rule-out major depressive disorder).

• Traits—this person does not meet criteria; however, he or she presents with many of the features of the diagnosis (e.g., borderline traits or cluster B traits).

• By history—previous records (another provider or hospital) indicate this diagnosis; records can be inaccurate or outdated (e.g., alcohol dependence by history).

• By self-report—the client claims this as a diagnosis; it is currently unsubstantiated; these can be inaccurate (e.g., bipolar by self-report).

For example, you may receive a fax from a hospital or other provider that might say, “Provisional Borderline Personality Disorder. Bipolar Diagnosis by self-report—no manic symptoms identified.”

Other Specified Disorders and Unspecified Disorders The DSM-IV had a diagnosis of not otherwise specified (NOS) to capture symptomology that did not fit well into a structured category. In lieu of the NOS diagnosis, the DSM-5 offers two

Dimensional diagnosis Offers ability to note symptom severity

Severity “Specify current severity”—choose the most accurate level of symptomology

Provisional diagnosis Used when strong inclination but can’t yet confirm

48 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

options when these situations arise. The other specified and unspecified disorders should be used when a provider believes an individual’s impairment to functioning or distress is clinically significant; however, it does not meet the specific diagnostic criteria in that category. The “other specified” should be used when the clinician wants to communi- cate specifically why the criteria do not fit. The “unspecified disorder” should be used when he or she does not wish, or is unable to, communicate specifics. For example, if someone appeared to have significant panic attacks but only had three of the four required criteria, the diagnosis could be “Other Specified Panic Disorder—due to insuf- ficient symptoms.” Otherwise, the clinician would report “Unspecified Panic Disorder.”

Specific Diagnostic Categories Section II of DSM-5 offers an in-depth discussion of 22 broad diagnostic categories and their subtypes as well as descriptions of medication-induced disorders and what is called “other conditions that may be a focus of clinical attention.” The fol- lowing offers a brief description of these disorders and is summarized from DSM-5 (APA, 2013). Please refer to the DSM-5 for an in-depth review of each disorder. When you finish reviewing these diagnoses, the class may want to do Exercise 3.1.

• Neurodevelopmental Disorders. This group of disorders typically refers to those that manifest during early development, although diagnoses are sometimes not assigned until adulthood. Examples of neurodevelopmental disorders include intel- lectual disabilities, communication disorders, autism spectrum disorders (incorpo- rating the former categories of autistic disorder, Asperger’s disorder, childhood disintegrative disorder, and pervasive developmental disorder), ADHD, specific learning disorders, motor disorders, and other neurodevelopmental disorders.

• Schizophrenia Spectrum and Other Psychotic Disorders. The disorders that belong to this section all have one feature in common: psychotic symptoms, that is, delusions, hallucinations, grossly disorganized or abnormal motor behavior, and/or negative symptoms. The disorders include schizotypal person- ality disorder (which is listed again, and explained more comprehensively, in the category of personality disorders in the DSM-5), delusional disorder, brief psychotic disorder, schizophreniform disorder, schizophrenia, schizoaffective disorder, substance/medication-induced psychotic disorders, psychotic disorders due to another medical condition, and catatonic disorders.

• Bipolar and Related Disorders. The disorders in this category refer to distur- bances in mood in which the client cycles through stages of mania or mania and depression. Both children and adults can be diagnosed with bipolar disor- der, and the clinician can work to identify the pattern of mood presentation, such as rapid-cycling, which is more often observed in children. These disor- ders include bipolar I, bipolar II, cyclothymic disorder, substance/medication- induced, bipolar and related disorder due to another medical condition, and other specified or unspecified bipolar and related disorders.

• Depressive Disorders. Previously grouped into the broader category of “mood disorders” in the DSM-IV-TR, these disorders describe conditions where depressed mood is the overarching concern. They include disruptive mood dys- regulation disorder, major depressive disorder, persistent depressive disorder (also known as dysthymia), and premenstrual dysphoric disorder.

Other specified disorder Doesn’t fit a standard diagnosis with an explanation why not

Unspecified disorder Doesn’t fit a standard diagnosis without explanation

Diagnostic categories Twenty-two categories included in one axis

CHAPTER 3 Diagnosis in the Assessment Process 49

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

• Anxiety Disorders. There are a wide range of anxiety disorders, which can be diagnosed by identifying a general or specific cause of unease or fear. This anxiety or fear is considered clinically significant when it is excessive and per- sistent over time. Examples of anxiety disorders that typically manifest earlier in development include separation anxiety and selective mutism. Other exam- ples of anxiety disorders are specific phobia, social anxiety disorder (also known as social phobia), panic disorder, and generalized anxiety disorder.

• Obsessive-Compulsive and Related Disorders. Disorders in this category all involve obsessive thoughts and compulsive behaviors that are uncontrollable and the client feels compelled to perform them. Diagnoses in this category include obsessive- compulsive disorder, body dysmorphic disorder, hoarding disorder, trichotilloma- nia (or hair-pulling disorder), and excoriation (or skin-picking) disorder.

• Trauma- and Stressor-Related Disorders. A new category for DSM-5, trauma and stress disorders emphasize the pervasive impact that life events can have on an individual’s emotional and physical well-being. Diagnoses include reac- tive attachment disorder, disinhibited social engagement disorder, posttrau- matic stress disorder, acute stress disorder, and adjustment disorders.

• Dissociative Disorders. These disorders indicate a temporary or prolonged dis- ruption to consciousness that can cause an individual to misinterpret identity, surroundings, and memories. Diagnoses include dissociative identity disorder (formerly known as multiple personality disorder), dissociative amnesia, deper- sonalization/derealization disorder, and other specified and unspecified disso- ciative disorders.

• Somatic Symptom and Related Disorders. Somatic symptom disorders were previously referred to as “somatoform disorders” and are characterized by the experiencing of a physical symptom without evidence of a physical cause, thus suggesting a psychological cause. Somatic symptom disorders include somatic symptom disorder, illness anxiety disorder (formerly hypochondriasis), conver- sion (or functional neurological symptom) disorder, psychological factors affecting other medical conditions, and factitious disorder.

• Feeding and Eating Disorders. This group of disorders describes clients who have severe concerns about the amount or type of food they eat to the point that serious health problems, or even death, can result from their eating beha- viors. Examples include avoidant/restrictive food intake disorder, anorexia ner- vosa, bulimia nervosa, binge eating disorder, pica, and rumination disorder.

• Elimination Disorders. These disorders can manifest at any point in a person’s life, although they are typically diagnosed in early childhood or adolescence. They include enuresis, which is the inappropriate elimination of urine, and encopresis, which is the inappropriate elimination of feces. These behaviors may or may not be intentional.

• Sleep-Wake Disorders. This category refers to disorders where one’s sleep pat- terns are severely impacted, and they often co-occur with other disorders (e.g., depression or anxiety). Some examples include insomnia disorder, hypersom- nolence disorder, restless legs syndrome, narcolepsy, and nightmare disorder. A number of sleep-wake disorders involve variations in breathing, such as sleep- related hypoventilation, obstructive sleep apnea hypopnea, or central sleep apnea. See the DSM-5 for the full listing and descriptions of these disorders.

50 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

• Sexual Dysfunctions. These disorders are related to problems that disrupt sex- ual functioning or one’s ability to experience sexual pleasure. They occur across sexes and include delayed ejaculation, erectile disorder, female orgasmic disorder, and premature (or early) ejaculation disorder, among others.

• Gender Dysphoria. Formerly termed, “gender identity disorder,” this category includes those individuals who experience significant distress with the sex they were born and with associated gender roles. This diagnosis has been separated from the category of sexual disorders, as it is now accepted that gender dys- phoria does not relate to a person’s sexual attractions.

• Disruptive, Impulse Control, and Conduct Disorders. These disorders are characterized by socially unacceptable or otherwise disruptive and harmful behaviors that are outside of the individual’s control. Generally, more common in males than in females, and often first seen in childhood, they include oppo- sitional defiant disorder, conduct disorder, intermittent explosive disorder, antisocial personality disorder (which is also coded in the category of person- ality disorders), kleptomania, and pyromania.

• Substance-Related and Addictive Disorders. Substance use disorders include disruptions in functioning as the result of a craving or strong urge. Often caused by prescribed and illicit drugs or the exposure to toxins, with these dis- orders the brain’s reward system pathways are activated when the substance is taken (or in the case of gambling disorder, when the behavior is being per- formed). Some common substances include alcohol, caffeine, nicotine, canna- bis, opioids, inhalants, amphetamine, phencyclidine (PCP), sedatives, hypnotics or anxiolytics. Substance use disorders are further designated with the follow- ing terms: intoxication, withdrawal, induced, or unspecified.

• Neurocognitive Disorders. These disorders are diagnosed when one’s decline in cognitive functioning is significantly different from the past and is usually the result of a medical condition (e.g., Parkinson’s or Alzheimer’s disease), the use of a substance/medication, or traumatic brain injury, among other phenomena. Examples of neurocognitive disorders (NCD) include delirium, and several types of major and mild NCDs such as frontotemporal NCD, NCD due to Parkin- son’s disease, NCD due to HIV infection, NCD due to Alzheimer’s disease, substance- or medication-induced NCD, and vascular NCD, among others.

• Personality Disorders. The 10 personality disorders in DSM-5 all involve a pattern of experiences and behaviors that are persistent, inflexible, and deviate from one’s cultural expectations. Usually, this pattern emerges in adolescence or early adulthood and causes severe distress in one’s interpersonal relationships. The personality disorders are grouped into the three following clusters based on similar behaviors: • Cluster A: Paranoid, schizoid, and schizotypal. These individuals seem

bizarre or unusual in their behaviors and interpersonal relations. • Cluster B: Antisocial, borderline, histrionic, and narcissistic. These indivi-

duals seem overly emotional, are melodramatic, or unpredictable in their behaviors and interpersonal relations.

• Cluster C: Avoidant, dependent, and obsessive-compulsive (not to be con- fused with obsessive-compulsive disorder). These individuals tend to appear anxious, worried, or fretful in their behaviors.

CHAPTER 3 Diagnosis in the Assessment Process 51

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

In addition to these clusters, one can be diagnosed with other specified or unspecified personality disorder, as well as a personality change due to another medical condition, such as a head injury.

• Paraphilic Disorders. These disorders are diagnosed when the client is sexually aroused to circumstances that deviate from traditional sexual stimuli and when such behaviors result in harm or significant emotional distress. The disorders include exhibitionistic disorder, voyeuristic disorder, frotteurisitc disorder, sexual sadism and sexual masochism disorders, fetishistic disorder, transvestic disorder, pedophilic disorder, and other specified and unspecified paraphilic disorders.

• Other Mental Disorders. This diagnostic category includes mental disorders that did not fall within one of the previously mentioned groups and do not have uni- fying characteristics. Examples include other specified mental disorder due to another medical condition, unspecified mental disorder due to another medical condition, other specified mental disorder, and unspecified mental disorder.

• Medication-Induced Movement Disorders and Other Adverse Effects of Medications. These disorders are the result of adverse and severe side effects to medications, although a causal link cannot always be shown. Some of these disorders include neuroleptic-induced parkinsonism, neuroleptic malignant syndrome, medication-induced dystonia, medication-induced acute akathisia, tardive dyskinesia, tardive akathisia, medication-induced postural tremor, other medication-induced movement disorder, antidepressant discontiunation syndrome, and other adverse effect of medication.

• Other Conditions That May Be a Focus of Clinical Assessment. Reminiscent of Axis IV of the previous edition of the DSM, this last part of Section II ends with a description of concerns that could be clinically significant, such as abuse/neglect, relational problems, psychosocial, personal, and environmental concerns, educational/occupational problems, housing and economic problems, and problems related to the legal system. These conditions, which are not con- sidered mental disorders, are generally listed as V codes, which correspond to ICD-9, or Z codes, which correspond to ICD-10.

Sometimes, mental health conditions can co-occur or be “comorbid.” For example, suppose a client presents with an anxiety disorder but also abuses alcohol. In this situa- tion, it would be appropriate to denote both disorders when making a diagnosis (e.g., generalized anxiety disorder and alcohol abuse). Sometimes disorders can even exacer- bate each other. An example of this could be someone who meets the criteria for depression, but his or her symptoms only present while withdrawing from cocaine use. Rather than diagnosing this as a major depressive episode, it is more appropriate that he or she be diagnosed with a substance-induced mood disorder (see Exercise 3.1).

Other Medical Considerations Sometimes, physical symptoms caused by a medical condition may look a lot like one or more of the mental disorders. For example, some of the symptoms for depres- sion include appetite disturbance (increase or decrease), irritability or restlessness, hyper or insomnia (i.e., sleeping too much or too little), difficulty concentrating, and fatigue or decreased energy. Interestingly, all of these symptoms can also be attrib- uted to hypothyroidism or underactive thyroid. Thus, in addition to clients being assessed for mental health problems, it is also important for them to be assessed for

Co-occurring disorders Disorders may coexist and can sometimes exacerbate one another

Be cognizant of medi- cal factors that may influence mental health

52 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

potential medical problems. One way to address this is to obtain specific details about when the client began experiencing his or her symptoms. Such information will help determine whether symptoms began while a medical condition was present and whether it is likely that the medical condition was the cause of the mental disor- der. For instance, suppose you have a client who is presenting with all of the criteria for an anxiety disorder (such as restlessness, irritability, and insomnia), but you know that these symptoms began when the client’s thyroid began declining and he or she found out it was underactive. If the client’s anxiety disorder only came about because of the hypothyroidism, then it would be appropriate to designate it as such, that is, anxiety disorder due to a general medical condition, hypothyroidism. Of course, it is always prudent to refer a client to his or her primary care physician if there is any suspicion that a medical problem may be the source of a psychological issue (see Exercise 3.2). If reporting a medical problem, the ICD code for the particu- lar problem can be used along with the DSM-5 mental health disorder diagnosis.

Psychosocial and Environmental Considerations As a part of a complete diagnosis, it is imperative for the clinician to assess the client’s psychosocial and environmental stressors. Such a focus promotes a holistic view of the client, provides important diagnostic clues, and can help to identify important issues in treatment planning. Not considered mental disorders, some of the many psychosocial and environmental concerns may include problems with the client’s primary support group, social environment, education, occupation, housing, economic situation, access to health care, crime or the legal system, or other significant psychosocial and environ- mental considerations (APA, 2013). Whereas these concerns were previously listed on Axis IV of DSM-IV, they are now denoted in the single-axis system, are now mostly listed under “Other conditions that may be a focus of clinical attention” discussed ear- lier, and are correlated to V codes in DSM-5 (which matches ICD-9) or Z codes (which matches ICD-10 ) (e.g., Z59.0 Homelessness; Z65.1 Imprisonment; Z55.3 Underachievement in school).

To illustrate the importance of psychosocial and environmental considerations, consider a 48-year-old male experiencing severe anxiety and depression. He explains

Exercise 3.1 Diagnosing a Disorder

Once the class has become familiar with the various disorders, the instructor may ask the students to prac- tice identifying and diagnosing disorders by performing

role-plays in dyads or small groups. You may want to use DSM-5 as a guide for acting out the criteria.

Exercise 3.2 Diagnosing Medical Conditions

After practicing formulating a diagnosis in Exercise 3.1, the instructor may ask the students to role-play again, incorporating a medical condition this time. Identify the

medical condition as well as the mental health condition, and be sure to note whether it is separate or if it is the cause of the mental health diagnosis.

Psychosocial/ environmental factors Stressors that are crucial for under- standing the whole person

V & Z codes Allows psychosocial stressors to be listed in diagnosis

© Ce ng ag e Le ar ni ng

20 15

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 3 Diagnosis in the Assessment Process 53

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

that his symptoms started immediately after a tornado caused severe damage to his home and neighboring farm (natural disaster). The man and his family have been staying with relatives about 70 miles away from home (homelessness), and have had no source of income for the past three months (economic issues) since their crop of soy was also destroyed in the tornado (occupational problem). By understanding the client’s psychosocial and environmental considerations, his anxiety and depression can be viewed in the context of his life circumstances.

Cultural Considerations Because people from diverse cultures may express themselves in different ways, symptomatology may vary as a function of culture (Mezzich & Caracci, 2008). Thus, some have argued that although diagnosis can be helpful in treatment plan- ning, it can lead to the misdiagnosis of culturally oppressed groups when clinicians do not fully take into account cultural, gender, and ethnic differences (Eriksen & Kress, 2005, 2006, 2008; Kress, Eriksen, Rayle, & Ford, 2005; Madesen & Leech, 2007; Rose & Cheung, 2012).

The APA (2013) has attempted to combat some of these problems by asking clin- icians to understand and acknowledge “culturally patterned differences in symptoms” (p. 758). For example, Latin American culture acknowledges that ataque de nervios (“attack of nerves”) is a common disorder related to difficult and burdensome life experiences and may exhibit itself through “headaches and ‘brain aches’ (occipital neck tension), irritability, stomach disturbances, sleep difficulties, nervousness, easy tearfulness, inability to concentrate, trembling, tingling sensations, and mareos (dizzi- ness with occasional vertigo-like exacerbations)” (p. 835). A clinician who ignores the client’s culture could easily misdiagnose a client who presents with symptoms like this and begin to treat the client with inappropriate strategies. Best practice for multicul- tural counseling suggests that the clinician have some understanding of differences in cross-cultural expression of symptoms and that the clinician explore the client’s culture with him or her when deciding on appropriate treatment strategies.

Finally, DSM-5 offers a section entitled Cultural Formulation Interview (CFI) that helps clinicians understand the kinds of values, experiences, and influences that have come to shape the client’s worldview and provides an outline for how to appropriately interview clients from diverse backgrounds. In addition, DSM-5 offers definitions of some cross-cultural symptoms and identifies how cross-cultural issues impact a wide- range of diagnoses.

Final Thoughts on DSM-5 in the Assessment Process DSM-5 is one additional piece of the total assessment process. Along with the clini- cal interview, the use of tests, and informal assessment procedures, it can provide a broad understanding of the client and can be a critical piece in the treatment plan- ning process. Consider what it might be like to establish a treatment plan if only one test were used. Then, consider what it would be like if two tests were used, then two tests and an informal assessment procedure; then two tests, an informal assessment procedure, and a clinical interview; and finally, two tests, an informal assessment procedure, a clinical interview, and a diagnosis. Clearly, the more “pieces of evidence” we can gather, the clearer our snapshot of our client becomes and this, in turn, yields better treatment planning (see Exercise 3.3).

Cultural considerations Use CFI and info from DSM to understand differences in symptomatology

Some “abnormal” behaviors may be considered “normal” in other cultures

54 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Exercise 3.3 Practice Making a Diagnosis

On your own, in pairs, or as a class, read the following case studies and formulate a diagnosis for each person using the DSM-5 as a guide. Discuss how you came to this diagnosis, what other diagnoses you considered but ruled out, and what additional information would have been helpful in assessing the scenario. Answers can be found at the end of the chapter.

I. Mikayla Mikayla is an 8-year-old girl in the second grade. She lives with her parents and younger brother, and her mother describes her as “a handful, but very sweet.” Mikayla had to repeat second grade due to behavioral issues in class, which also resulted in lower test scores and poor grades. She is a popular child among her peers, but she continually struggles with her teacher to follow directions, stay on task, and remain seated. Mikayla’s teacher has consulted with her kindergarten and first-grade teachers and does not think that Mikayla has communication issues or a specific learning disorder because she performs above grade-level expectations in small group or with one-on-one attention. In a large classroom situation, she is in constant motion, shouts when she should be talking quietly, and is easily distracted, which prevents her from meeting expectations. Most recently, Mikayla was referred to the school counselor after she broke the classroom fish tank during a silent read- ing activity. “I was just trying to feed Flipper,” she explained.

II. Tracey Tracey is a 25-year-old single working mother. Her daughter, Alicia, is 3 years old and in day care during the work week. Tracey was recently divorced from Alicia’s father and has sole custody of their child because her ex-husband was physically abusive. In the past few years, when the marital problems began, Tracey became overwhelmed with anxiety but was so busy that she stated she just didn’t have time to deal with it. She starts her day at 5:30 a.m. to get Alicia dressed, packed, and ready for day care so that she can get to work by 7:00 a.m. Tracey usually has break- fast on the road, and she frequents the drive-through on her way to work for convenience. At work it’s “go, go, go,” and Tracey doesn’t usually have time to break

for lunch. By the time she picks up Alicia from day care and gets home, it’s about 6:00 p.m. Tracey cooks din- ner by 7:00 p.m., which usually consists of a healthy, balanced meal. Once she gives Alicia a bath and puts her to bed, Tracey finally gets a breather to relax on the couch and watch TV. Now that she is alone, she feels an uncontrollable urge to snack and often goes through a large bag of potato chips followed by a quart of ice cream before she realizes it. Sometimes, she finishes eating that amount before her favorite half-hour sitcom is over. “I just can’t stop. It’s like I zone out, and I don’t even realize how much I’ve eaten. I feel like I can’t con- trol myself. Usually, I feel physically sick by the end of it and just pass out, like a food coma.” Tracey doesn’t like to eat junk food in front of others because she’s ashamed that she has gained so much weight since the divorce and feels self-conscious. She’s been eating in secret like this for the past year since the divorce, and it happens almost every night. It’s gotten to the point where she has begun isolating herself, preferring to go home and snack all night in front of the TV instead of spending time with family and friends.

III. Alan Alan is a 37-year-old banker, who was divorced from his wife two years ago. Alan reported that his wife left him after he became disengaged from the marriage. He recalled that he and his wife were college sweethearts and were previ- ously very active in their community. Then, approximately five years ago, Alan said, “I just ran out of steam.” He has since been constantly irritable, started sleeping excessively, gained about 45 pounds, and lost interest in being social and engaging in pleasurable activities. Alan smokes mari- juana approximately two to four times per day and drinks vodka nightly to “relax and take my mind off of things.” He recently was arrested for possession of marijuana and driv- ing under the influence, put on probation for one year, and was referred to counseling by the court. Alan admits to being mildly depressed, but insists, “It’s nothing I can’t handle.” He does not wish to discontinue his marijuana or alcohol use but has thoughts about stopping due to monthly drug screens, which will soon be required by his probation officer.

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 3 Diagnosis in the Assessment Process 55

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SUMMARY We began this chapter by discussing the impor- tance that diagnosis plays in the professional life of mental health professionals. We noted that a large percentage of Americans are diagnosed, yearly, with a mental disorder, and highlighted some reasons that diagnosis has become impor- tant for professionals, such as their significance in identifying children in schools with emotional, behavioral, and learning disorders; the fact that a diagnosis can help in case conceptualization and treatment planning; the fact that professionals can help employers make accommodations for and understand individuals with mental disorders; because they are critical to insurance reimburse- ment; because they assist professionals in accurately communicating to one another; and because they can help clients understand their prognosis and expectations for treatment planning.

Next, we offered a brief history of the DSM starting with the U.S. Census Bureau counting those who were “idiots” and “insane” during the mid-1800s. However, we noted that it wasn’t until 1943, with the military’s Medical 203, that a formal classification system was developed. We then noted that DSM-I was developed in 1952 and underwent a number of revisions up through the most recent edition, the DSM-5 in 2013. We then introduced the DSM-5 and began noting some of the differences from its predecessors, particularly the move from a five-axes system in DSM-IV to a one-axis system in DSM-5.

A large portion of the chapter described DSM-5. We began by explaining why DSM-5 moved to a one-axis system and discussed making and reporting a diagnosis. In this process, we dis- cussed how to order diagnoses; the use of sub- types, specifiers, and severity; how to make a provisional diagnosis; and the use of other speci- fied or unspecified disorders. Next, we offered very brief descriptions of the 22 diagnostic cate- gories. We also offered a brief discussion about co-occurring or comorbid disorders. This was fol- lowed by a discussion about the importance of understanding how medical conditions can cause or exacerbate a diagnosis. We then noted that whereas psychosocial and environmental consid- erations were placed on Axis IV of DSM-IV, they are now correlated to an ICD code, are often given a V or Z code, and included in the single- axis system. Finally, we noted that individuals may present symptoms in varying ways as a func- tion of their culture and talked about the impor- tance of taking into consideration one’s cultural background when making a diagnosis. We pointed out that DSM-5 offers a Cultural Formu- lation Interview (CFI) that can help in the process of understanding diverse clients, provides some examples of cross-cultural symptoms, and identi- fies how cross-cultural issues impact a wide range of diagnoses. We ended the chapter by noting that DSM-5 is one piece in the total assessment process and offered an exercise where students could try to diagnose three hypothetical clients.

CHAPTER REVIEW 1. Consider how a mental health diagnosis

could be beneficial to a client. What might be potential harm from a diagnosis?

2. Why is it important for clinicians, medical doctors, legal professionals, and so on to use a common diagnostic language?

3. Explain why it is an ethical responsibility for clinicians to be knowledgeable about diagnosis.

4. Give examples of how you might utilize a diagnosis when formulating a treatment plan.

5. Describe how medical conditions can be rel- evant to a mental health diagnosis.

6. Describe how psychosocial and environ- mental considerations are now included in DSM-5.

7. Explore how you can include multicultural considerations into a diagnosis.

56 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

8. Discuss the difference between use of subtypes, specifiers, and severity.

9. Explain how provisional diagnoses can be made.

10. Discuss the use of other specified or unspecified disorders.

11. Describe how a one-axis system can be used to encapsulate all of the five axes from DSM-IV.

12. Identify a diagnosis from any category that makes you personally feel uncomfortable. Explore where these feelings come from and how you might go about working with a cli- ent who has this diagnosis.

ANSWERS TO EXERCISE I. 314.01 (F90.2) Attention-deficit hyperactivity

disorder, combined presentation; V62.3 (Z55.9) Academic underachievement.

II. 307.51 (F50.8) Binge eating disorder, moderate; V61.03 (Z63.5) Disruption of family by divorce (recent); V60.2 (Z59.7) Low income; V62.9 (Z60.9) Unspecified problem related to social environment: Social isolation.

III. 300.4 (F34.1) Persistent depressive disorder (dysthymia); 303.90 (F10.20) Alcohol use disorder, moderate; 304.30 (F12.20) Cannabis use disorder, moderate; V61.03 (Z63.5) Disruption of Family by Divorce (two years ago); V62.5 (Z65.0) Conviction in civil or criminal proceedings without imprisonment: Probation.

REFERENCES American Psychiatric Association. (1952). Diagnostic

and statistical manual of mental disorders. Washington, DC: Author.

American Psychiatric Association. (1968). Diagnostic and statistical manual of mental disorders (2nd ed.). Washington, DC: Author.

American Psychiatric Association. (1980). Diagnostic and statistical manual of mental disorders (3rd ed.). Washington, DC: Author.

American Psychiatric Association. (1994). Diagnostic and statistical manual of mental disorders (4th ed.). Washington, DC: Author.

American Psychiatric Association. (2000). Diagnostic and statistical manual of mental disorders (4th ed., text revision). Washington, DC: Author.

American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). Washington, DC: Author.

Beutler, L. E., & Malik, M. L. (Eds.). (2002). Rethink- ing the DSM: A psychological perspective. Washing- ton, DC: American Psychological Association.

Centers for Disease Control and Prevention. (2011). U.S. adult mental illness surveillance report. Retrieved from http://www.cdc.gov/Features/ MentalHealthSurveillance/

Centers for Disease Control and Prevention. (2013). Mental health surveillance among children—United States, 2005–2011. Retrieved from http://www.cdc. gov/mmwr/preview/mmwrhtml/su6202a1.htm?s_cid= su6202a1_w

Eriksen, K., & Kress, V. (2005). Beyond the DSM story: Ethical quandaries, challenges, and best prac- tices. Thousand Oaks, CA: Sage.

Eriksen, K., & Kress, V. (2006). The DSM and the professional counseling identity: Bridging the gap. Journal of Mental Health Counseling, 28, 202–217.

Eriksen, K., & Kress, V. (2008). Gender and diagnosis: Struggles and suggestions for counselors. Journal of Counseling and Development, 86, 152–162.

Good, E. M. (2012). Personality disorders in the DSM-5: Proposed revisions and critiques. Journal of Mental Health Counseling, 34, 1–13.

Houts, A. C. (2000). Fifty years of psychiatric nomen- clature: Reflections on the 1943 War Department Technical Bulletin, Medical 203. Journal of Clinical Psychology, 56, 935–967. doi.org/10.1002/1097- 4679(200007)56:7<935::AID-JCLP11>3.0.CO;2-8

Kress, V., Eriksen, K., Rayle, A., & Ford, S. (2005). The DSM-IV-TR and culture: Considerations for

CHAPTER 3 Diagnosis in the Assessment Process 57

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

counselors. Journal of Counseling and Develop- ment, 83, 97–104.

Krueger, R. F., & Eaton, N. R. (2010). Personality traits and the classification of mental disorders: Toward a more complete integration in DSM-5 and an empirical model of psychopathology. Per- sonality Disorders: Theory, Research, and Treat- ment, 1(2), 97–118. doi:10.1037/a0018990

Madesen, K., & Leech, P. (2007). The ethics of labeling in mental health. Jefferson, NC: MacFarland & Company.

Mezzich, J. E., & Caracci, G. (Eds.) (2008). Cultural formation: A reader for psychiatric diagnosis. Lan- ham, MD: Jason Aronson.

Neukrug, E., & Schwitzer, A. (2006). Skills and tools for today’s counselors and psychotherapists: From natural helping to professional counseling. Pacific Grove, CA: Brooks/Cole.

Rose, A. L., & Cheung, M. (2012). DSM-5 research: Assessing the mental health needs of older adults from diverse ethnic backgrounds. Journal of Ethnic & Cultural Diversity in Social Work: Innovation in Theory, Research & Practice, 21, 144–167. doi:10.1080/15313204.2012.673437

Segal, D. L., & Coolidge, F. L. (2001). Diagnosis and classification. In M. Hersen & V. B. Van Hasselt, (Eds.), Advanced abnormal psychology (2nd ed., pp. 5–22). New York: Kluwer Academic/Plenum Publishers.

Seligman, L. (1999). Twenty years of diagnosis and the DSM. Journal of Mental Health Counseling, 21, 229–239.

Seligman, L. (2004). Diagnosis and treatment planning in counseling (3rd ed.). New York: Plenum.

Smith, T. A. (2012, October 15). Revolutionizing diagnosis & treatment using DSM-5. Workshop presented at CMI Education Institute, Newport News Marriott, Newport News, VA.

Substance Abuse and Mental Health Services Adminis- tration. (2012). Results from the 2011 National Survey on Drug Use and Health: Mental Health Findings. Retrieved from http://www.samhsa.gov/ data/NSDUH/2k11MH_FindingsandDetTables/2K11 MHFR/NSDUHmhfr2011.htm#2.1

Thyer, B. A. (2006). It is time to rename the DSM. Ethical Human Psychology and Psychiatry, 8, 61–67. doi:10.1891/ehpp.8.1.61

U.S. Department of Justice. (n.d.). Information and technical assistance on the Americans with Disabil- ities Act. Retrieved from http://www.ada.gov/

Wakefield, J. C. (2013). DSM-5: An overview of changes and controversies. Clinical Social Work Journal, 41, 139–154. doi:10.1007/s10615-014-0445-2

Zalaquett, C. P., Fuerth, K. M., Stein, C., Ivey, A. E., & Ivey, M. B. (2008). Reframing the DSM-IV-TR from a multicultural/social justice perspective. Jour- nal of Counseling & Development, 86, 364–371. doi:10.1002/j.1556-6678.2008.tb00521.x

58 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 4The Assessment ReportProcess: Interviewing the Client and Writing the Report

When I began working with clients in the early 1970s, I was dutiful about writing my case notes on lined paper and typing more involved client reports. Soon it was the mid-70s, and I got a job as an outpatient therapist at a mental health center. I was compulsive about conducting thorough clinical interviews, and religious about dictating my case notes, my intake summaries, my quarterly summaries, and my test reports. My dictations were care- fully typed out by secretaries and placed in client folders. Ten years later, now in private practice, I was still compulsive about conducting a thorough interview, about writing my case notes, and about typing my client reports. But soon an innovation arrived I hadn’t expected—the computer. I started doing my notes and my reports on the computer—that was a relief for me, although not everyone felt the same way. Soon laws were passed that protected the client’s right to confidentiality of his or her case notes and reports and the rights of clients to view them. Over the years, the methods we use to write reports and the ways we secure them have certainly changed dramatically, but one thing has remained the same—the examiner decides what the report will say!

(Ed Neukrug)

As you can see from the vignette, over the years the process of writing notes and reports has changed and the laws protecting clients’ right to view notes and reports have evolved. Despite these changes, there is little question that “paperwork,” whether it be on lined paper, typed, or written on the computer, continues to be a major part of what all practitioners do. In this chapter, we examine one important aspect of the paperwork process—how to write an assessment report. First, we

59

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

discuss the purpose of the assessment report and then discuss the process of gather- ing information for the report, including conducting the interview and choosing appropriate assessment techniques. Then, we suggest ways to write effective test reports and supply an example of such a report.

PURPOSE OF THE ASSESSMENT REPORT The assessment report is the “deliverable” or “end product” of the assessment pro- cess whose purpose is to synthesize an assortment of assessment techniques so that a deeper understanding of the examinee is made and recommended courses of action are offered (Goldfinger & Pomerantz, 2010; Lichtenberger, Mather, Kaufman, & Kaufman, 2004; Spores, 2013). Such courses of action can vary dramatically, based on the reason the individual is being assessed. For instance, some of the many purposes of reports include the following:

1. To respond to the referral questions being asked; 2. To provide insight to clients for therapy; 3. To assist in the case-conceptualization process; 4. To develop treatment options in counseling (e.g., type of counseling, use of

medications, and so on); 5. To suggest educational services for students with special needs (e.g., for stu-

dents who are mentally retarded, learning disabled, or gifted); 6. To offer direction when providing vocational rehabilitation services; 7. To offer insight about and treatment options for individuals who have incurred

a cognitive impairment (e.g., brain injury, senility); 8. To assist the courts in making difficult decisions (e.g., custody decisions, sanity

defenses, and determination of guilt or innocence); 9. To provide evidence for placement in schools and at jobs; and

10. To challenge decisions made by institutions and agencies (social security disability, Individual Educational Plans (IEPs) in schools).

Because complex decisions regarding clients’ lives are often based on the assess- ment report, synthesizing the information gathered and placing it in the report is only accomplished after the examiner can conduct interviews successfully, adminis- ter assessment procedures proficiently, and write reports skillfully. One of the first steps in this process is to ensure that any information gathered is directly related to the purpose of the assessment and is of high quality.

GATHERING INFORMATION FOR THE REPORT: GARBAGE IN, GARBAGE OUT

Gathering information for the report is as important as writing the report, because your report will reflect the methods you used to obtain your information. If you choose inappropriate instruments or conduct a poor interview (“garbage in”), your report will be filled with error and bias (“garbage out”). To help ensure that the information you are gathering is of high quality, you should always take into account the breadth and depth of your assessment procedures.

Assessment report Written summary, synthesis, and recom- mendations from the assessment

60 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The breadth of the assessment has to do with casting a wide enough net to ensure that the examiner has done all that is necessary to adequately assess what he or she is looking for. Breadth should be based on the purpose of the assessment. For instance, if a middle school student came to see a school counselor to examine career possibilities, the counselor would likely conduct an interview focused on vocational interests and offer a broad-based career interest inventory to gather information from the student about his or her general interests. However, if an adult came to a counseling center as a result of depression, anxiety, and general discontent in life, a very broad assessment might be called for to help establish a diagnosis and deter- mine treatment goals. In this case, it would not be unusual to conduct a clinical interview, administer a number of objective and projective personality tests, and per- haps interview others to assess the client’s relationships at home and at work.

The depth of the assessment has to do with ensuring that one is using techni- ques that reflect the intensity of the issue(s) being examined. As with breadth, depth is also dependent on the purpose for which the client is being assessed. For instance, conducting an in-depth clinical interview and offering a rather complex interest inventory that helps the middle school student determine a career would be too involved—too much depth. On the other hand, offering a personality inven- tory like the Myers–Briggs for the individual suffering from depression, anxiety, and discontentment in life would not entail enough depth. You could simply miss too much because you have not delved more deeply into this client’s issues.

In establishing the breadth and depth of the interview, it is important that you are able to establish trust and rapport and assure confidentiality within the limits of the purpose the individual is being assessed. The better the interviewer is able to build trust, the more likely the information obtained will be reliable. With these points in mind, examiners need to determine whether a structured, unstructured, or semi-structured interview would be best when assessing the client.

STRUCTURED, UNSTRUCTURED, AND SEMI-STRUCTURED INTERVIEWS

Determining the kind of interview to conduct is critical to gathering information successfully because the clinical interview accomplishes a number of tasks not pos- sible through the use of other assessment techniques. For instance, the interview

1. sets the tone for the types of information that will be covered during the assessment process,

2. allows the client to become desensitized to information that can be very inti- mate and personal,

3. allows the examiner to assess the nonverbal signals of the client while he or she is talking about sensitive information, thus giving the examiner a sense of what might be important to focus on,

4. allows the examiner to learn firsthand the problem areas of the client and place them in perspective, and

5. gives the client and examiner the opportunity to study each other’s personality style to assure that they can work together.

Breadth Covering all important or relevant issues; wide net

Depth Extent and seriousness of a concern

Clinical interview Offers ability to obtain reliable information from clients

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 61

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Interviewers generally have a choice between three kinds of interviews: struc- tured, unstructured, or semi-structured. With advantages and disadvantages to both, which one to choose is not always easy (Bruchmuller, Margraf, Suppiger, & Schneider, 2011; Goldfinger & Pomerantz, 2010; Lichtenberger et al., 2004). For instance, completed verbally or in response to written items, the structured interview has the examiner ask the examinee to respond to preestablished items. This kind of interview can provide the following benefits:

• It offers broad enough areas of content to cover topics a practitioner may oth- erwise have missed or forgotten to ask (assures breadth of coverage).

• It increases the reliability of results by ensuring that all prescribed items will be covered.

• It ensures that the examiner will cover all of the items because they are listed in detail and there is an expectation that they all will be covered.

• It ensures that items will not be missed due to interviewer or interviewee embarrassment.

On the other hand, the structured interview can have the following drawbacks:

• The examiner may miss information due to the fact that items are predeter- mined and the examiner does not feel free to go off on a tangent or a “hunch.”

• Clients may experience the interview as dehumanizing. • Clients, particularly minorities, may misinterpret or be unfamiliar with certain

items. • Follow-up by the examiner to alleviate any confusion on the part of the exam-

inee is less likely as compared to other kinds of interviewing. • It does not always allow for depth of information to be covered because the

interviewer is more concerned with gathering all the information than going into detail about one potentially sensitive area.

Contrast the structured interview with the unstructured interview, where the examiner does not have a preestablished list of items or questions to which the client can respond. In this case, examinee responses to inquiries will set the direc- tion for follow-up questioning. The unstructured interview offers the following advantages:

• It creates an atmosphere that is more conducive to building rapport. • It allows the client to feel as if he or she is directing the interview, thus allow-

ing the client to discuss items that he or she deems important. • It offers the potential for greater depth of information because the clinician can

focus on a potentially sensitive area and possibly uncover underlying issues that the client might otherwise avoid revealing.

On the other hand, the unstructured interview may have the following disadvantages:

• Because it does not allow for breadth of coverage, the interviewer might miss information because he or she is “caught up” in the client’s story instead of following a prescribed set of questions.

• The interviewer may end up spending more time on some items than he or she might like.

Structured interview Uses preestablished questions to assess broad range of behaviors

Unstructured interview Examiner asks questions based on client responses

62 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Drawing from the advantages of both the structured and unstructured interview, examiners will often conduct a semi-structured interview. This kind of interview uses prescribed items, which allows the examiner to obtain the necessary information within a relatively short amount of time. However, it also gives leeway to the exam- iner should the client need to “drift” during the interview process. Allowing the client to discuss potentially emotion-filled topics can be cathartic, open up new issues of importance, and be an important tool in the rapport-building process. The skilled examiner can easily flow back and forth between structured and unstructured approaches. If time is not an issue, a semi-structured interview can provide the breadth and depth needed when interviewing clients while still allowing the examiner to focus in on building the relationship (see Box 4.1).

Computer-Driven Assessment Today, computers are frequently used when conducting a structured or semi-structured interview and in the report generation. For instance, many agencies now use electronic health records (EHR) where information about a client is stored on a computer. Inter- viewers and interviewees can jointly sit down and complete specific items included in the EHR, and this information can be stored with other, related, medical and psycho- logical information (Cimino, 2013). There are also programs for purchase that can assist in the interview process. One such program has the interviewer or the client com- plete 120 items that requests information about a wide range of personal issues, and the user receives a computer-generated report that describes the client’s presenting problems, legal issues, current living situation, tentative diagnosis, emotional state, treatment recommendations, mental status, health and habits, disposition, and behavioral/physical descriptions (see Schinka, 2012). Computer-assisted questioning is as reliable, or some- times more reliable, than structured interviews and can provide an accurate diagnosis at a minimal cost (Farmer, McGuffin, & Williams, 2002).

In addition to assisting in the interviewing process, computers can generate test reports, and oftentimes pieces of these reports can be moved directly into the exam- iner’s written assessment report (Berger, 2006; Michaels, 2006). Final assessment reports generated by computer-driven programs have become so sophisticated that most well-trained clinicians cannot tell them apart from reports written by sea- soned professionals. However, whether it is a computer-generated report that

BOX 4.1 Missing Substance Abuse

When I first started doing counseling I tended to use unstructured interview style. I believed it was critical to let the client take the interview where he or she wanted it to go. However, after a number of years of missing alcohol abuse as well as other “hidden” issues, I slowly began to make the switch to a more semi-structured interview style where I would go through a list of predetermined items and also have clients complete a genogram (see

Chapter 12). I had learned that clients were often embarrassed about revealing some very important information that sometimes ended up being the focus of treatment. The switch to a semi- structured interview style allowed me to quickly pick up on these issues and immediately address what in the past had been hidden.

—Ed Neukrug

Semi-structured interview Allows flow between structured and unstructured approaches

Computer-driven assessment Can result in sophisticated assessment and well-written report

© Ce ng ag e Le ar ni ng

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 63

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

resulted from a client interview or a report that includes aspects of computer- generated tests reports, it is still up to the examiner to make sure that the correct questions are being asked of the client, the correct assessment procedures are being used, and that the material used in the development of the examiner’s assessment report is chosen wisely by the examiner.

While the computer may administer these [tests] and even prepare a report, what the test looks like, how it responds, what it achieves and any reports generated, are prede- termined by the author, and it is a person who puts it all together to make the interpre- tation. (Berger, 2006, p. 70)

CHOOSING AN APPROPRIATE ASSESSMENT INSTRUMENT

After the interview, clinicians can consider the purpose of their assessment as well as the breadth and depth of other needed information. Then clinicians can choose from a broad array of assessment instruments, such as those we will examine later in this book, including assessment of educational ability (Chapter 8), assessment of intellectual and cognitive functioning through intelligence testing and neuropsychological assess- ment (Chapter 9), career and occupational assessment (Chapter 10), clinical assessment (Chapter 11), and informal assessment (Chapter 12). During this process, it is important that clinicians carefully reflect on which are the most appropriate instruments to use, for it is unethical to assess an individual using instruments that are not related to the pur- pose of the assessment being undertaken (see American Counseling Association [ACA], 2005; American Psychological Association [APA], 2010) (see Box 4.2).

Counselors are responsible for the appropriate application, scoring, interpretation, and use of assessment instruments relevant to the needs of the client, whether they score and inter- pret such assessments themselves or use technology or other services. (ACA, Section E.2.b)

and

Psychologists administer, adapt, score, interpret, or use assessment techniques, interviews, tests, or instruments in a manner and for purposes that are appropriate in light of the research on or evidence of the usefulness and proper application of the techniques. (APA, Section 9.02.a)

BOX 4.2 Proper Assessment of a Client

I was once hired to write an assessment report for the expressed purpose of challenging the denial of a client’s social security disability payments despite the fact that she had been diagnosed as having a multiple personality disorder (also discussed in Chapter 11). Aware of her disorder, and having worked hard to integrate her various personalities in therapy, she was, however, quite depressed and at

times would dissociate. Because the task at hand was to assess the client’s ability to work rather than to affirm her diagnosis, it was important to choose instruments that would only address the assessment question: Could this client effectively hold down employment?

—Ed Neukrug

Only choose assessment technique to match purpose of testing

© Ce ng ag e Le ar ni ng

64 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

WRITING THE REPORT After you have conducted a thorough assessment of your client, you will be ready to write your report. Reports are scrutinized today more than ever before because they are the mechanism used by the interviewer to communicate his or her assessment to stake- holders and are often used by funding agencies and supervisors when evaluating a clini- cian’s work. In addition, as a result of laws passed over the years, such as the Family Educational Rights and Privacy Act (FERPA), the Freedom of Information Act, and the Health Insurance Portability and Accountability Act (HIPAA) (see Chapter 2), clients will generally have access to their records if they choose to review them. Keeping all of this in mind, Box 4.3 offers a summary of a number suggestions concerning how to write a report (Lichtenberger et al, 2004; Wiener & Costaris, 2012).

Although many clinicians these days are asked to write involved reports, the actual format of the report tends to vary from setting to setting. For example, a large mental health clinic may specify a preferred or required format for its thera- pists. Similarly, a social worker in private practice may be driven to a particular format by insurance provider requirements, while a school counselor may have to use an established format required by the system-wide school counseling director.

Although report formats vary, they will often include some or all of the follow- ing sections: (1) demographic information, (2) presenting problem or reason for refer- ral, (3) family background, (4) significant medical/counseling history, (5) substance use and abuse, (6) educational and vocational history, (7) other pertinent information, (8) mental status, (9) assessment results, (10) diagnosis, (11) summary and conclusions, and (12) recommendations. Let’s take a look at each of these areas in more detail.

Demographic Information In this section, we find basic information about the client, including such items as the client’s name, address, phone number, e-mail address, date of birth, age, sex, ethnicity, and date of interview. Also, it is in this section that the name of the

BOX 4.3 Fifteen Suggestions for Writing Reports

1. Omit passive verbs. 2. Be nonjudgmental. 3. Reduce the use of jargon. 4. Do not use a patronizing tone. 5. Increase the use of subheadings. 6. Reduce the use of and define acronyms. 7. Minimize the number of difficult words. 8. Try to use shorter rather than longer words. 9. Make sure paragraphs are concise and flow

well. 10. Point out strengths and weaknesses of your

client.

11. Don’t try to dazzle the reader of your report with your brilliance.

12. When possible, describe behaviors that are representative of client issues.

13. Only label when it is necessary and valuable to do so for the client’s well-being.

14. Write the report so a non–mental-health pro- fessional can understand it (e.g., a teacher).

15. Don’t be afraid to take a stand if you feel strongly that the information warrants it (e.g., the information leads you to believe a client is in danger of harming self).

Writing assessment reports Know “tips” to good report writing

FERPA, HIPAA, Freedom of Information Act Increase clients’ ability to access records

© Ce ng ag e Le ar ni ng

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 65

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

interviewer is placed. Often, this information is included at the top of the report. The following is an example of the demographic information gathered from a ficti- tious client, Mr. Unclear.

Name: Eduardo (Ed) Unclear Address: 223 Confused Lane

Coconut Creek, Florida Phone: 954-969-5555 E-mail: [email protected] Name of Interviewer: Sigmund Freud, MD

DOB: 1/8/1966 Age: 48 Sex: Male Ethnicity: Hispanic

(Cuban-American) Date of Interview: 10/22/2014

Presenting Problem or Reason for Referral In this section, the person who referred the client is generally noted (e.g., self- referred, physician, counselor, etc.) and an explanation is given as to why the indi- vidual has come for counseling and/or why the examiner has been asked to do the assessment. For instance, here it might be explained that a social worker has been asked to do a court assessment of a child for a custody hearing; a school psycholo- gist has been asked to assess a child who has been exhibiting severe behavioral pro- blems at school, for a possible diagnosis of emotional disturbance; or a licensed clinician in private practice might suggest to a client an assessment to help sort out a diagnosis and to set treatment goals. Continuing with our example of Mr. Unclear, we might include the following information:

Eduardo Unclear is a 48-year-old Hispanic male of average stature and build. He was self- referred to counseling due to stress and inability to sleep. The client reported feeling anxious for approximately two years and intermittently depressed for approximately seven or eight years. He states that he feels discontent with his marriage and confused about his future. Mr. Unclear appeared appropriately dressed and was attentive during the session. An assessment was conducted to determine differential diagnosis and the course of treatment.

Family Background The family background section of the report is an opportunity to give the reader an understanding of possible factors concerning the client’s upbringing that may be related to his or her presenting problem. Trivial bits of information should be left out of this section, and opinions regarding this information should be saved for the summary and conclusions section of the report.

In this section, it is often useful to mention where the individual grew up, sexes and ages of siblings, whether the client came from an intact family, who were the major caretakers, and significant others who may have had an impact on the client’s life. The examiner may also want to relay important stories from childhood that have affected how the client defines himself or herself. For adults, one should also include such items as marital status, marital or relationship issues, ages and sexes of any children, and sig- nificant others. Using our example, we might include the following information:

Mr. Unclear was raised in Miami, Florida. When he was five years old, his parents fled from Cuba on a fishing boat with him and his two brothers, José, who is two years older, and Juan, who is two years younger. Mr. Unclear comes from an intact family. He reports that his father was a bookkeeper and his mother was a stay-at-home mom.

66 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

He states that his parents were “loving but strict” and notes that his father was “in charge” of the family and would often “take a belt to me.” He reports that he and his brothers were always close and that both brothers currently live within 1 mile of his home. He states that his younger brother is married and has two children. He describes his other brother as single and “gay but not out.” He and his brothers went to Catholic school, and he states that he was a good student and had the “normal” number of friends. His father died approximately four years ago of a “heart disorder.” His mother currently resides in a retirement community in North Miami Beach.

Mr. Unclear notes that he met his wife Carla in college when he was 20. They married when he was 21 and quickly had two children, Carlita and Carmen, who are now 27 and 26. Both daughters are college-educated, have professional jobs, and are married. Carlita has two children aged 3 and 4, while Carmen has one child aged 5. He notes that both daughters and their families live close to him, and he maintains positive relationships with them. He states that although his marriage was “good” for the first 20 years, in recent years he has found himself feeling unloved and depressed. He wonders if he should remain in the marriage.

Significant Medical/Counseling History This section of the report delineates any significant medical history, especially any physical conditions that may be affecting the client’s psychological state. Any pre- scribed medication, with dosage, should be noted. In addition, any history of counseling should be noted in this section. Mr. Unclear’s medical and counseling history is summarized below:

Mr. Unclear reports that approximately four years ago he was in a serious car accident that subsequently left him with chronic back pain. Although he is prescribed medication for the pain (Flexeril, 5mg), he prefers not to take it, stating that he mostly tries to “live without drugs.” He notes that he often feels fatigued and has trouble sleeping, usually sleeping around four hours a night. He reports that a recent medical exam revealed no apparent medical reason for his fatigue and sleep difficulties. He notes that in the past two years he has had obsessive worry related to fears of dying of a heart attack. He describes his eating habits as “normal” and reports no other significant medical history.

Mr. Unclear explained that after the birth of his second child, his wife required surgery to repair vaginal tears. He states that since that time she has experienced pain during intercourse and their level of intimacy has significantly decreased. He notes that he and his wife attended couples counseling for about two months approximately 15 years ago. He feels that counseling did not help, and he reports that it “particularly did nothing to help our sex life.”

Substance Use and Abuse This section reports the use and abuse of any legal or illegal substances that may be addictive or potentially harmful to the client. Thus, the interviewer should note the use or abuse of food, cigarettes, alcohol, prescription medication, and illegal drugs. In reference to Mr. Unclear, we include the following information:

Mr. Unclear states that he does not smoke cigarettes but does occasionally smoke cigars, adding that he “will never smoke a Cuban cigar.” He describes himself as a moderate alcohol user, stating that he has a “couple of beers a day” but rarely drinks “hard liquor.” He reports taking prescription medication intermittently for chronic back pain, and he denies the use of illegal substances.

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 67

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Educational and Vocational History This section describes the client’s educational background and delineates his or her job path and career focus. For Mr. Unclear, we include the following information:

Mr. Unclear attended Catholic school in Miami, Florida. He reports that he excelled in math but had difficulty with reading and spelling. After high school, he attended college at the University of Miami, where he majored in business administration. After gradu- ating with his bachelor’s degree, he obtained a job as an accountant at a major tobacco import company, where he worked for 17 years. During that time, he began to work on his master’s in business administration but stated he never finished his degree because it was “boring.” Approximately eight years ago he changed jobs to “make more money.” He obtained employment as an accountant at a local new car company.

Mr. Unclear states that as an accountant, his “books were always perfect,” although he went on to note that he was embarrassed by his inability to prepare a well-written report. He expresses dissatisfaction with his career path and wants to “do something more meaningful with his life.” He adds, however, that “I am probably too old to change careers now.”

Other Pertinent Information This “catch-all” category addresses any significant information that has not been noted elsewhere. Issues that might be addressed in this section could be related to sexual orientation, changes in sexual desires, sexual dysfunction; current or past legal problems that may be affecting functioning; and financial problems the client may be having. For Mr. Unclear we include the following information:

Mr. Unclear states that he is unhappy with his sex life and reports limited intimacy with his wife. He denies an extramarital affair but states “I would have one if I met the right person.” He notes that he is “just making it” financially and that it was difficult to support his two children through college. He denies any problems with the law.

Mental Status A mental status exam is an assessment of the client’s appearance and behavior, emotional state, thought components, and cognitive functioning. This assessment is used to assist the interviewer in making a diagnosis and in treatment planning (Akiskal, 2008; Polanski & Hinkle, 2000; Sommers-Flanagan & Sommers- Flanagan, 2012). A short synopsis of each of the four areas of the mental status exam follows and definitions of common words used in mental status exam can be found in Table 4.1.

Appearance and Behavior This part of the mental status exam reports the client’s observable appearance and behaviors during the clinical interview. Thus, such items as manner of dress, hygiene, body posture, tics, significant nonverbal behaviors (eye contact or the lack thereof, wringing of hands, swaying), and manner of speech (e.g., stuttering, tone) are often reported.

Emotional State When assessing emotional state, the examiner describes the cli- ent’s affect and mood. The affect is the client’s current, prevailing feeling state (e.g.,

Mental status exam Assesses appearance and behavior, emo- tional state, thought, and cognitive functioning

68 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TABLE 4.1 Common Terms and Definitions or Descriptions Used in the Mental Status Exam

Category Term Definition or Description

Appearance and Behavior Appearance Appropriate or baseline, eccentric or odd, abnormal movement or gait, good or poor grooming, or hygiene

Eye contact Good or poor

Speech Within normal limits, loud, soft, pressured, hesitant

Emotional State—Affect Appropriate or inappropriate Appropriate or inappropriate to mood (e.g., laughing while talking of recent death)

Full and reactive Full range of emotions correctly associated with the conversation

Labile Uncontrollable crying or laughing

Blunted Reduced expression of emotional intensity

Flat No or very little expression of emotional intensity

Emotional State—Mood Euthymic Normal mood Depressed Sad, dysphoric, discontent Euphoric Extreme happiness or joy Anxious Worried Anhedonic Unable to derive pleasure from previously

enjoyable activities

Angry/hostile Annoyed, irritated, irate, etc.

Alexithymic Unable to describe mood

Thought Components—Content

Hallucinations False perception of reality: may be auditory,

visual, tactile (touch), olfactory (smell), or taste

Ideas of reference Misinterpreting casual and external events as being related to self (e.g., newspaper headlines, TV stories, or song lyrics are about the client)

Delusions False belief (e.g., “satellites are tracking me”); may be grandiose, persecutory (to be harmed), somatic (physical symptom with no medical condition), erotic

Derealization External world seems unreal (e.g., watching it like a movie)

Depersonlization Feeling detached from self often with no control (e.g., “I feel like I’m living a dream”)

Suicidality and homicidality Ranges from none, ideation, plan, means, preparation, rehearsal, and intent

(Continued)

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 69

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

happy, sad, joyful, angry, depressed, etc.) and may also be reported as constricted or full, appropriate or inappropriate to content, labile, flat, blunted, exaggerated, and so forth. The client’s mood, on the other hand, represents the long-term, underlying emotional well-being of the client and is usually assessed through client self-report. Thus, a client may seem anxious and sad during the session (affect) and report that his or her mood has been depressed.

Thought Components The manner in which a client thinks can reveal much about how he or she comes to understand and make meaning of the world. Thought components are generally broken down into the content and the process of thinking. Clinicians will often make statements about thought content by addressing whether the client has delusions, distortions of body image, hallucina- tions, obsessions, suicidal or homicidal ideation (see Box 4.4), and so forth. The kinds of thought processes often identified include circumstantiality, coherence, flight of ideas, logical thinking, intact as opposed to loose associations, organiza- tion, and tangentiality.

TABLE 4.1 Common Terms and Definitions or Descriptions Used in the Mental Status Exam (continued)

Category Term Definition or Description

Thought Components—Process

Logical and organized Normal state where one’s thoughts are rational and structured

Poverty Lack of verbal content or brief responses

Blocking Difficulty or unable to complete statements Clang Emphasis on using words that rhyme together

rather than on meaning Echolalia “Echoing” clients own speech or your speech;

repeating Flight of ideas Rapid thoughts almost incoherent Perseveration Thoughts keep returning to the same idea Circumstantial Explanations are long and often irrelevant but

eventually get to the point Tangential Responses never get to the point of the question Loose Thoughts have little or no association to the con-

versation or to each other Redirectable Responses may get off track, but you can direct

them back to the topic

Cognition Orientation Knows who they are, where they are, and date Memory Ability to remember events from recent, immediate,

and long-term Insight Ability to recognize his or her mental illness; good,

limited, or none Judgment Ability to make sound decisions; good, fair,

or poor

© Ce ng ag e Le ar ni ng

20 15

70 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Cognition Cognition includes a statement as to whether the client is oriented to time, place, and person (knows what time it is, where he or she is, and who he or she is); an assessment of the client’s short- and long-term memory; an evaluation of the client’s knowledge base and intellectual functioning; and a statement about the client’s level of insight and ability to make judgments.

Although much more can be said about each of these four areas, generally, when incorporating a mental status into a report, all four areas are collapsed into a one- or two-paragraph statement about the client’s presentation. Usually, a state- ment about the client’s demeanor, orientation, affect, intellectual functioning, judg- ment, insight, and suicidal or homicidal ideation are included. Other areas are generally reported only if they are deemed significant (see Exercise 4.1). A descrip- tion of Mr. Unclear’s mental status follows:

Eduardo Unclear appeared for his appointment casually but neatly dressed and groomed. He was able to maintain appropriate eye contact and was oriented to time, place, and person. Visual acuity appeared to be within normal limits; audition and speech were unremarkable. During the interview he appeared anxious, often rubbing his hands together. Mr. Unclear was cooperative with the examiner and demonstrated satisfactory levels of motivation, interest, and energy. He is currently prescribed pain medication, which he only takes occasionally for chronic back pain.

He stated that he often feels fatigued because he usually sleeps approximately four hours a night. He described himself as feeling intermittently depressed over the past seven or eight years. He appeared to be of above average intelligence and his memory

BOX 4.4 Assessment of Lethality

It is important to assess for the risk of suicide or homicide. Both of these can be thought of occurring along a continuum ranging from ideation (thinking about it), developing a plan, means to carry out the plan, prep- aration stage, rehearsing the plan, and finally acting it out (Commonwealth of Virginia Knowledge Center [COVKC], 2010). Hence, determining where the client might be on the continuum is helpful in determining how much risk he or she may have for harming self or others.

It can also be important to evaluate both risk and protective factors (COVKC). Risk factors can include having a history of psychiatric treatment, non-medication compliance, substance abuse, prior attempts, recent losses, or other significant problems. Protective factors to consider are strong family or social supports, adherence to med- ications, stable employment, children under the age of 18, religious beliefs, or fear of killing one’s. Finally, deter- mine if this individual will contract for safety in verbal or written format.

A c t

N o t h o u g

h ts

T h o u g

h ts

P la

n

M e a n s

P re

p a ra

ti o n

R e h e a rs

a l

© Ce ng ag e Le ar ni ng

20 15

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 71

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

was intact. His judgment seemed fair and his insight fair to good. He stated that he has some suicidal ideation but denies he has a plan or would kill himself, noting that it is “against my religion.” He denied homicidal ideation.

Assessment Results It is often helpful to begin this section with a simple list of the assessment proce- dures that were used. Next, the results of the assessment procedures are generally presented. When presenting test results, it is important to not just give out raw scores. Instead, offering converted or standardized test scores that the reader will understand is usually more helpful (discussed in Chapter 7). It should be remem- bered that the client, parents, or some other nonprofessional may read these results, so it is important to state the results in language that is unbiased and understand- able to the reader. An example might be, “Johnny scored a 300 on his SAT, which is 2 standard deviations below the mean and places him at about the 2nd percentile as compared to his peers in the 12th grade.”

The results of the assessment should be concise, yet cover all items that are clearly relevant to the presenting concerns or that stand out as a result of the assessment. Results should be presented objectively, and interpretations should be kept to a mini- mum, if used at all. In the summary and conclusions section, the examiner will have the opportunity to hypothesize about what is happening with the client. The following is an example of the assessment section of the report for Mr. Unclear. (We will discuss these specific tests in Section III of this book. For now pay attention to the formatting as these examples will be more meaningful near the end of the semester.)

Mr. Unclear was administered a battery of objective and projective personality tests, including the Beck Depression Inventory-II (the BDI-II), the Minnesota Multiphasic Personality Inventory-II (the MMPI-II), the Rorschach Inkblot Test (the Rorschach), the Thematic Apperception Test (TAT), the Kinetic Family Drawing (KFD), the Sentence Completion Test, the Strong Interest Inventory (the Strong), and the Wide Range Achievement Test-4 (the WRAT).

Through self-report of the past two weeks, Mr. Unclear’s score on the BDI-II indicates that he has moderate depression (raw score ¼ 24). His responses showed some evidence of possible suicidal ideation. The BDI-II is not only able to diagnose depression due to its consistency with the DSM diagnostic criteria, but it can also determine the severity of depressive symptoms. The MMPI-II supports this finding of moderate to severe depression and also indicates some mild anxiety. The MMPI-II, which reveals dissatisfaction with one’s life, demonstrates that Mr. Unclear is generally “discontent with the world” and feels a lack of intimacy in his life. It suggests assessing for possible suicidal ideation.

Exercise 4.1 Writing the Mental Status Report

Your instructor may ask a student to role-play a client being interviewed. (It might also help if the student chose to reflect a diagnosis in the DSM-5.) After the role-play is complete, all other students in class should write a mental-status report. Share your reports with the

instructor, and come up with one mental status report for the class. Compare your own report to the final ver- sion produced in class. This type of role-play can be repeated in small groups if you would like to gain further practice writing mental-status exams.

Assessment results Provide results that will be understandable to reader

© Ce ng ag e Le ar ni ng

72 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The Rorschach and the TAT are projective assessment tools used to evaluate psychological functioning. Both demonstrated that Mr. Unclear is grounded in reality and open to testing, as evidenced by his willingness to readily respond to initial inkblots and TAT cards, his ability to complete stories in the TAT, and the fact that many of his responses were “common” responses. Feelings of depression and hopelessness are evident in a number of responses, such as not readily seeing color in many of the responses to the “color cards” of the Rorschach and making a number of pessimistic stories that generally had depressive endings to them on the TAT.

When he was administered the KFD, a projective test that asks the client to draw his family all doing something together, Mr. Unclear placed his father as an angel in the sky and included his wife, mother, children, and grandchildren. His mother was standing next to him while his wife was off to the side with the grandchildren. He also placed himself in a chair and when describing the picture stated, “I’m sitting because my back hurts.” The picture showed the client and his family at his mother’s house having a Sunday dinner while it rained outside. Rain could be indicative of depressive feelings. A cross was prominent in the background and was larger than most of the people in the picture, which is likely an indica- tion of strong religious beliefs and could also indicate a need to be taken care of.

On the Sentence Completion Test Mr. Unclear made a number of references to missing his father, such as “The thing I think most about is missing my father.” He also referenced continual back pain. Finally, he noted discontent with his marriage, including the statement, “Sex is nonexistent.”

On the Strong, a self-report assessment tool used to evaluate both personality and career interest, Mr. Unclear’s two highest personality codes were Conventional and Enterprising, respectively. All the other codes were significantly lower. Individuals of the Conventional type are stable, controlled, conservative, sociable, and like to follow instructions. Enterprising individuals are self-confident, adventurous, and sociable. They have good persuasive skills, and prefer positions of leadership. Careers in business and industry where persuasive skills are important are good choices for these individuals.

On the WRAT-4, Mr. Unclear scored at the 86th percentile in math, 75th percentile in reading, 64th percentile in sentence comprehension, and 42nd percentile in spelling. His reading composite score was at the 69th percentile. These results could indicate a possible learning disorder in spelling, although cross-cultural considerations should be taken into account due to the fact that Mr. Unclear was an immigrant to this country at a young age.

Diagnosis This is the section where a clinical diagnosis is generally made using the criteria from the Diagnostic and Statistical Manual for Mental Disorder, fifth edition (DSM-5; APA, 2013) (see Chapter 3). The diagnosis is an outgrowth of the whole assessment process and is based on the integration of all of the knowledge gained (Seligman, 2004). As mentioned in Chapter 3, the older DSM-IV-TR (APA, 2000) offered sepa- rate axes that included medical conditions, psychosocial and environmental condi- tions, and global assessment of functioning. Although the current DSM uses a single mental disorder axis, you may want to include V (or Z) codes, which reflect psycho- social and environmental conditions. In addition, you may want to provide medical diagnoses from the International Classification of Disease, ninth revision (ICD-9) or ICD-10, dependent on the audience of the report. For example, if this report is for a court or some portion of the medical community, it might be useful to include the ICD diagnosis and codes. However, if the report is going to a mental health profes- sional or directly to a client, reporting medical conditions in layman’s terms may be

Older DSM-IV diagnosis Be familiar with for reading charts prior to DSM-5

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 73

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

more helpful. There is a fine balance between making the report accurate and profes- sional using correct terminology while also keeping it readable and understandable for the end user. Below we list Mr. Unclear’s diagnosis. Note that the first diagnostic number is the ICD-9 code and ICD-10 is followed in parenthesis.

296.22 (F32.1) Major depression, single episode, moderate

309.28 (F43.23) Rule out: Adjustment disorder with mixed Anxiety and depressed mood

V62.29 (Z56.9) Problems related to employment

V61.10 (Z63.0) Relationship distress with spouse

722.0 Displacement cervical intervertebral disc (Chronic Back Pain)

In this example, the diagnosis describes a client who is experiencing symptoms of a moderate episode of major depression, including ongoing feelings of depres- sion, fatigue, and sleep problems; may be having difficulty adjusting to a new life circumstance; has problems related to work and his relationship with his wife, and has difficulties with chronic back pain.

Summary and Conclusions This section is the examiner’s chance to pull together all of the information that has been gathered. Often, this is the only section of the report that is read by others, so it is important that it is accurate and does not leave out any main points. However, it should not be excessively long. Making it accurate, succinct, and rele- vant is the key to writing a good summary. One major error in writing summaries, we have found, is adding information that has not been included elsewhere. The summary should have no new information. Although inferences can be made in this section, they must be logical, sound, defendable, and based on facts that are mentioned in your report. We also generally recommend writing a paragraph or two about the strengths of the individual. All too often, we have found that this is left out of reports. The following might be a summary and conclusions section based on the information we have gathered from Mr. Unclear:

Mr. Unclear is a 48-year-old married male who was self-referred due to feelings of depression, anxiety, and discontent with his job and his marriage. Mr. Unclear fled from Cuba to Miami, Florida with his parents and two siblings when he was 5 years old. He describes his family as close, and he continues to live near his children, siblings, and mother. His father died approximately four years ago. He married while in college. He and his wife subsequently raised two girls who are now in their mid-20s, married, and have their own children.

Mr. Unclear finished college with a degree in business and has been working as an accountant for the past 25 years. He reports feeling dissatisfied in his career and states that he wants to “do something more meaningful with his life.” He also reports marital discord, which he attributes partly to medical problems his wife had after the birth of their second child. These problems, he states, resulted in diminishing sexual relations with his wife.

Mr. Unclear was oriented during the session but appeared anxious and talked about feelings of depression. He noted that he often feels fatigued, has difficulty sleep- ing, and has fleeting thoughts of suicide, which he states he would not act upon. Recently, he has had obsessive worries about having a heart attack, although there is

Report summary should have no new information

74 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

no medical reason to support his concerns. Chronic back pain due to a car accident a few years ago seems to exacerbate his current feelings of depression.

Throughout testing the consistent themes of depression, isolation, and hopelessness emerged. High scores on the BDI-II and the MMPI-II depression scale evidenced this. This was additionally indicated by specific responses to the Rorschach, TAT cards, the KFD, and the sentence completion.

Dissatisfaction with his marriage, sadness about the loss of his father, and chronic pain were also major themes that arose during testing. Testing also revealed a person whose career is a good match for his personality. However, he might be more chal- lenged if he entered a position requiring additional responsibilities and leadership skills. Such a change may be disadvantageous if Mr. Unclear does not receive treatment for his depression. Finally, testing also shows a possible learning disability in spelling, although cross-cultural issues may have affected his score.

On a positive note, testing and the clinical interview showed a man who was neatly dressed and open to collaborating with the examiner. He has worked hard in his life and is proud of the family he has raised. He was grounded in reality, willing to engage interpersonally, and showed fair to good judgment and insight. He seems to be aware of many of his most pressing concerns and showed some willingness to address them.

Recommendations This last section of the report should be based on all of the information gathered. It should make logical sense to the reader. Although some prefer writing this section in paragraph form, we prefer listing each recommendation, as we believe this for- mat is clearer to the reader. The signature of the examiner generally follows this last section. The following might be some recommendations for Mr. Unclear:

1. Counseling, 1 hour a week for depression, possible anxiety, marital discord, and career dissatisfaction.

2. Possible marital counseling with particular focus on sexual relations of the couple. 3. Referral to a physician/psychiatrist for medication, possibly antidepressants. 4. Possible further assessment for learning problems. 5. Long-term consideration of a career move following alleviation of depressive feel-

ings and addressing possible learning problems. 6. Possible orthopedic reevaluation of back problems.

Signature of the Examiner

SUMMARIZING THE WRITING OF AN ASSESSMENT REPORT As you can see, a great deal of information is gathered from the client, and much of it is included in the report. Although one could probably write a short novel about a client after gathering information from an in-depth interview, generally the skilled examiner will keep the report between two and five pages, single- spaced. Box 4.5 summarizes the major points that should be gathered in an assess- ment report, and in Appendix D you can see Mr. Unclear’s report in its entirety.

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 75

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

BOX 4.5 Summary of Assessment Report

The following categories are generally assessed in a report:

Demographic Information Name: DOB.: Address: Age: Phone: Sex: Ethnicity: E-mail address: Date of interview: Name of interviewer:

Presenting Problem or Reason for Referral 1. Who referred the client to the agency? 2. What is the main reason the client contacted the

agency? 3. Reason for assessment

Family Background 1. Significant factors from family of origin 2. Significant factors from current family 3. Some specific issues that may be mentioned:

where the individual grew up, sexes and ages of siblings, whether the client came from an intact family, who were the major caretakers, important stories from childhood, sexes and ages of current children, significant others, and marital concerns

Significant Medical/Counseling History 1. Significant medical history, particularly anything

related to the client’s assessment (e.g., psychi- atric hospitalization, heart disease leading to depression)

2. Types and dates of previous counseling

Substance Use and Abuse 1. Use or abuse of food, cigarettes, alcohol, pre-

scription medication, or illegal drugs 2. Counseling related to use and abuse

Educational and Vocational History 1. Educational history (e.g., level of education and

possibly names of institutions) 2. Vocational history and career path (names and

types of jobs) 3. Satisfaction with educational level and career

path

4. Significant leisure activities

Other Pertinent Information 1. Legal concerns and history of problems with the

law 2. Issues related to sexuality (e.g., sexual orienta-

tion, sexual dysfunction) 3. Financial problems 4. Other concerns

The Mental Status Exam 1. Appearance and behavior (e.g., dress, hygiene,

posture, tics, nonverbals, and manner of speech) 2. Emotional state (e.g., affect and mood) 3. Thought components (e.g., content and pro-

cess: delusions, distortions of body image, hal- lucinations, obsessions, suicidal or homicidal ideation, circumstantiality, coherence, flight of ideas, logical thinking, intact as opposed to loose associations, organization, and tangentiality)

4. Cognitive functioning (e.g., orientation to time, place, and person; short- and long-term mem- ory; knowledge base and intellectual function- ing; insight and judgment)

Assessment Results 1. List assessment and test instruments used 2. Summarize results 3. Avoid raw scores and state results in unbiased

manner 4. Consider using standardized test scores and

percentiles

Diagnosis 1. DSM-5 diagnoses 2. Include V and/or Z codes if appropriate 3. Include other diagnoses such as medical,

rehabilitation, or other salient factors

Summary and Conclusions 1. Integration of all previous information 2. Accurate, succinct, and relevant 3. No new information

76 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SUMMARY This chapter began by describing some of the pur- poses of the assessment report, including respond- ing to the referral question, providing insight to clients, developing case conceptualization, devel- oping treatment options, suggesting educational services, providing vocational rehabilitation ser- vices, offering insight and treatment options for those with cognitive impairment, assisting courts in decision-making, helping with placement in schools and jobs, and challenging decisions made by institutions and agencies.

We pointed out that putting “garbage” into the assessment report process leads to a flawed report (garbage out) and that it is important to assure that the breadth and depth of the report is being attended to. Thus, we noted the impor- tance of providing a wide enough net to appropri- ately address the breadth of the intended assessment and to provide testing instruments that reflect the intensity or depth of the issue(s) being examined. We also stated that it is critical to establish trust and rapport and assure confi- dentiality if one is to gather reliable information.

We noted that examiners can generally choose from three kinds of interviews: structured, unstruc- tured, and semi-structured. We pointed out that interviews set the tone for the gathering of informa- tion process and allow clients to become desensitized to sometimes very personal information, allow the examiner the ability to assess clients nonverbally, allow the examiner to understand the client “first- hand” and place the client’s issues in perspective, and afford the examiner and client the opportunity to see if they can work together.

Distinguishing among the different kinds of interviews, we noted that a structured interview

asks the examinee to respond to preestablished items or questions, whereas an unstructured inter- view is more open-ended. The semi-structured interview uses prescribed items but also gives lee- way to the examiner should the client need to “drift” during the interview process. We pointed out some strengths and weaknesses of each of these approaches, especially as they relate to gath- ering information that has breadth and depth. We also noted that computer-assisted assessment, such as electronic health records (EHR) and computer-driven assessment reports, can help in the information-gathering process and can be integrated into the examiner’s assessment report.

We next examined the process of selecting an appropriate assessment technique, which often will be based on the clinical interview you had with your client. We stressed the importance of considering the breadth and depth of information provided by the chosen instruments and noted that clinicians will often choose from a broad array of instruments, such as the ones examined in this book. We pointed out that clinicians should carefully reflect on which are the most appropriate instruments and that it is unethical to assess an individual using instruments that are not related to the purpose of the assessment being undertaken.

As the chapter continued, we discussed how to write the actual assessment report. We pointed out that laws such as FERPA, HIPAA, and the Freedom of Information Act mean that clients often will have access to their records. We also noted that reports today are scrutinized more than ever before and stressed the importance of developing a writing style that is clear, concise, and easy to understand. We suggested an array of points to consider when

4. Inferences that are logical, sound, defendable, and based on facts in the report

5. At least one paragraph that speaks to the cli- ent’s strengths

Recommendations 1. Based on all the information gathered 2. Should make logical sense to reader 3. In paragraph form or as a listing 4. Usually followed by signature of examiner

© Ce ng ag e Le ar ni ng

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 77

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

writing a report, including omitting passive verbs, being nonjudgmental, reducing jargon, not being patronizing, using subheadings, reducing acronyms, minimizing difficult words, using shorter sentences, being concise and having good flow, pointing out client strengths and weaknesses, not dazzling the reader with your brilliance, describing behaviors, avoiding labeling if possible, writing for a nonmen- tal health professional, and being able to take a stand when necessary.

In the last section of the chapter, we exam- ined the format of the report. We noted that the

areas addressed when gathering information from a client often parallel the areas included in the actual report. They include (1) demographic information, (2) presenting problem or reason for the report, (3) family background, (4) significant medical/ counseling history, (5) substance use and abuse, (6) vocational and educational history, (7) other pertinent background information, (8) mental sta- tus, (9) assessment or test results, (10) diagnosis, (11) summary and conclusions,and (12) recommen- dations. We discussed each area and offered a case example.

CHAPTER REVIEW 1. Describe some purposes of the assessment

report. 2. Relative to the assessment report process,

describe what is meant by “garbage in, gar- bage out.”

3. Describe what it means to choose an assess- ment instrument with “breadth and depth.”

4. Discuss some of the tasks associated with conducting clinical interviews that are not associated with other kinds of assessment techniques.

5. Compare and contrast structured, semi-structured, and unstructured interview techniques.

6. What place can computer-generated reports take in the assessment report process?

7. Explain the impact of laws such as the Family Educational Rights and Privacy Act (FERPA), the Freedom of Information Act, and the Health Insurance Portability and Accountability Act (HIPAA) on the preparation of assessment reports.

8. Write down as many of the suggestions for writing a report that you can remember.

9. Make a list of the kinds of information gen- erally obtained in an assessment report.

10. Describe the four components of a mental-status exam. Interview a student in class and then write a one- or two-paragraph mental-status exam.

11. Conduct an assessment of a client and use the categories listed in this chapter to write an assessment report.

REFERENCES Akiskal, H. S. (2008). The mental status examination.

In S. H. Fatemi, & P. J. Clayton (Eds.), The medical basis of psychiatry (pp. 3–16). Totowa, NJ: Humana. doi.org/10.1007/978-1-59745-252-6_1

American Counseling Association (ACA). (2005). Code of ethics. (Rev. ed.). Alexandria, VA: Author.

American Psychiatric Association. (2000). Diagnostic and statistical manual of mental disorders (4th ed., text revision). Washington, DC: Author.

American Psychiatric Association (2013). Diagnostic and statistical manual of mental disorders (5th ed). Washington, DC, Author.

American Psychological Association. (2010). Ethical principles of psychologists and code of conduct including 2010 amendments. Retrieved from http:// www.apa.org/ethics/code/index.aspx

Berger, M. (2006). Computer assisted clinical assess- ment. Child and Adolescent Mental Health, 11(2), 64–75. doi.org/10.1111/j.1475-3588.2006.00394.x

Bruchmuller, K., Margraf, J., Suppiger, A., & Schneider, S. (2011). Popular or unpopular? Therapists’ use of structured interviews and their estimation of patient acceptance. Behavior Therapy, 42, 634–643. doi: 10.1016/j.beth.2011.02.003

78 SECTION I Understanding the Assessment Process

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Cimino, J. J. (2013). Improving the electronic health record—Are clinicians getting what they wished for? The Journal of the American Medication Association, 309, 991–992. doi:10.1001/jama.2013.890.

Commonwealth of Virginia Knowledge Center (2010). Department of Behavioral Health and Developmental Services. Assessing the risk of serious harm to self module 10. Retrieved from https://covkc.virginia .gov/dbhds/external/Kview/CustomCodeBehind/base /courseware/scorm/scorm12courseframe.aspx

Farmer, A., McGuffin, P., & Williams, J. (2002). Measuring psychopathology. New York: Oxford University Press.

Goldfinger, K., & Pomerantz, A. M. (2010). Psychological assessment and report writing. Los Angeles, CA: Sage.

Lichtenberger, E. O., Mather, N., Kaufman, N. L., & Kaufman, A. L. (2004). Essentials of assessment report writing. Hoboken, NJ: John Wiley & Sons.

Michaels, M. H. (2006). Ethical considerations in writing psychological assessment reports. Journal of Clinical Psychology, 62(1), 47–58. doi.org/10.1002/jclp.20199

Polanski, P. J., & Hinkle, J. S. (2000). The mental status examination: Its use by professional counse- lors. Journal of Counseling and Development, 78, 357–364. doi.org/10.1002/j.1556-6676.2000. tb01918.x

Schinka, J. A. (2012). Mental status checklist-adult. Lutz, FL: Psychological Assessment Resources.

Seligman, L. (2004). Diagnosis and treatment planning in counseling (3rd ed.). New York: Plenum. doi.org/10.1007/978-1-4419-8927-7

Sommers-Flanagan, J., & Sommers-Flanagan, R. (2012). Clinical interviewing. Hoboken, NJ: John Wiley & Sons.

Spores, J. M. (2013). Clinician’s guide to psychological testing and assessment: With forms and templates for effective practice. New York: Springer.

Wiener, J., & Costaris, L. (2012). Teaching psycholog- ical report writing: Content and process. Canadian Journal of School Psychology, 27(2), 119–135. doi: 10.1177/0829573511418484

CHAPTER 4 The Assessment Report Process: Interviewing the Client and Writing the Report 79

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

S E C T I O N IITest Worthiness and TestStatistics

In Section II, we examine how tests are created, scored, and interpreted. Chapters 5, 6, and 7 explain a number of important concepts related to test statistics. Being able to understand these statistics and how they are applied to the assessment process is crucial if one is to interpret test data appropriately. As you read these chapters, we will show you that the development of tests and the interpretation of test data is a deliberate and planned process that involves a scientific approach to the understanding of differences among people. We try to present this information in a down-to-earth and comprehensible fashion.

Chapter 5 defines test worthiness as an involved, objective analysis of a test. To do this objective analysis, however, we first highlight the importance of correla- tion coefficient and show how this important statistic is used to conduct an analy- sis of many aspects of the four factors of test worthiness that are examined in this chapter: (1) validity: whether the test measures what it’s supposed to measure; (2) reliability: whether the score an individual has received on a test is an accurate measure of his or her true score; (3) cross-cultural fairness: whether the score the individual has obtained is a true reflection of the individual, and not a function of cultural bias inherent in the test or in the examiner; and (4) practicality: whether it makes sense to use a test in a particular situation. The chapter concludes with a list of five steps to use in test selection to assure test worthiness.

Chapter 6 starts with the basics: an examination of raw scores. We first show that raw scores generally provide little meaningful information about a set of scores, and then we look at various ways that we can manipulate raw scores to

81

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

make sense out of a set of data. Thus, in this chapter we examine how a variety of basic statistics and graphs can help us understand raw scores. Some of these include frequency distributions; histograms and frequency polygons; cumulative distributions; the normal curve; skewed distributions; measures of central tendency such as the mean, median, and mode; and measures of variability such as the range, interquartile range, and standard deviation.

Chapter 7 is a natural extension of Chapter 6, and in it we examine derived scores and explore how they are used to help us understand raw scores. We start by distinguishing norm-referenced testing and criterion-referenced testing because these two ways of understanding test scores are quite different, and derived scores are mostly associated with norm-referenced testing. Next, we discuss specific types of derived scores: percentiles; standard scores, including z-scores, T-scores, devia- tion IQs, stanines, sten scores, college and graduate school entrance exam scores (e.g., SATs, GREs, and ACTs), NCE scores, and publisher-type scores; and devel- opmental norms such as age comparisons and grade equivalents. The chapter nears its conclusion with a discussion of standard error of measurement and stan- dard error of estimate. The chapter ends with a brief discussion of nominal, ordi- nal, interval, and ratio scales of measurement as we describe how each type of scale has unique attributes that may limit the statistical calculations one can per- form, and we mention that different kinds of assessment instruments use different kinds of scales.

82 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 5Test Worthiness: Validity,Reliability, Cross-Cultural Fairness, and Practicality

I’m walking down a Cincinnati street, and a man walks up to me and asks, “Want to take a test?” “Sure,” I reply. He takes me to a storefront, gives me a test, and then goes to a back room to score it. A few minutes later he reappears and tells me, “Well, you’re pretty bright and have a fairly good personality, but if you take Ron Hubbard’s course in Scientology, you will be brighter and have a better personality.” I tell him, “No thanks.” A few years later, I’m walking down a street in Minneapolis and a man comes up to me and inquires, “Want to take a test?” This time I say, “If you can show me that this test is a good test— that is, it has good reliability and validity—I’ll take it.” He says, “I’m sure it has good reliability and validity.” I say, “Show me.” He says, “Well, I know our New York office must have that information.” I say, “Well, I tell you what, I’ll buy Ron Hubbard’s book, Dianetics, and if and when you can get me the information from the New York office, I’ll read the book.” I gave him my name and address. I never heard from him again.

(Ed Neukrug)

As you might expect from this example, this chapter is about test worthiness, or how good a test actually is. Demonstrating test worthiness requires an involved, objective analysis of a test in four critical areas: (1) validity: whether it measures what it’s supposed to measure; (2) reliability: whether the score an individual receives on a test is an accurate measure of his or her true score; (3) cross-cultural fairness: whether a person’s score is a true reflection of the individual and not a function of cultural bias inherent in the test; and (4) practicality: whether it makes sense to use a test in a particular situation. After examining these four factors, we

Test worthiness Based on validity, reliability, cross- cultural fairness, and practicality

83

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

conclude the chapter with a discussion of five steps to use in test selection to assure test worthiness. However, prior to examining the four critical factors of validity, reliability, cross-cultural fairness, and practicality, and before discussing test selec- tion, we examine the concept of correlation coefficient because it is central to understanding much of what is presented in this chapter.

CORRELATION COEFFICIENT Correlation coefficient, which shows the relationship between two sets of scores, is a statistical concept frequently used in discussions of the critical factors just listed. Correlation coefficients range from �1.00 to þ1.00 and generally are reported in decimals of one-hundredth. A positive correlation shows a tendency for scores to be related in the same direction. For instance, if a group of individuals took two tests, a positive correlation would show a tendency for those who obtained high scores on the first test to obtain high scores on the second test, for those who obtained low scores on the first test to obtain low scores on the second test, and so forth. On the other hand, a negative correlation shows an inverse relationship between sets of scores; for instance, individuals who obtain high scores on the first test would be likely to obtain low scores on the second test.

Let’s take a look at real world examples of positive and negative correlations. For instance, researchers have generally found positive correlations between social connect- edness and other variables such as subjective well-being (Yoon, Lee, & Goh, 2008), longevity (Zunzunegui, Beland, Sanchez, & Otero, 2009), and life satisfaction (Park, 2009). This makes sense as people who are more connected, tend to be happier, live lon- ger, and are more satisfied. These are positive correlations that move in the same direc- tion. (See Figure 5.1.) In contrast, we can expect a negative correlation between social connectedness and depression (Glass, Mendes de Leon, Bassuk, & Berkman, 2006); that is, as people connect more with others, they tend to be less depressed. Since social connectedness and depression move in opposite directions, this makes them inversely related or negatively correlated. In the school setting, students with attention deficit hyperactivity disorder (ADHD) are likely to have a negative correlation with their school performance. The more severe their ADHD symptoms are, the more likely aca- demic achievement will decrease. These variables move in opposite directions that make them correlate negatively or inversely as seen in Figure 5.1.

A correlation that approaches �1.00 to þ1.00 demonstrates a strong relation- ship, while a correlation that approaches 0 shows little or no relationship between two measures or variables. (See Figure 5.2.) For instance, if I wanted to show that SAT scores predict college performance, I would have to show that the correlation coefficient does not approach zero and is significant enough to warrant its use. (See Kobrin & Patterson, 2011; Maruyama, 2012.) Similarly, if I wanted to show that my newly made test of depression was worthwhile, I could correlate scores on my new test with an established test of depression. In this case, I would expect to find a relatively high correlation coefficient as evidenced by the fact that individuals who scored high on my test would have a tendency to score high on the established test, individuals who scored low on my test would have a tendency to score low on the established test, and so forth. As we discuss the four critical factors of validity, reliability, cross-cultural fairness, and practicality, you will see that correlation

Correlation coefficient Relationship between two sets of test scores

Positive correlations Move in the same direction

Negative correlations Move in opposite direction

84 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

coefficient often plays an important role in many of them. If you would like to find out how to calculate a correlation coefficient, see Appendix E.

A correlation between two sets of variables, or test scores, can also be plotted on a graph. By placing an individual’s first score on the horizontal (x) axis and second score on the vertical (y) axis, you can plot this person’s scores on the graph. If you continue doing this for the remaining members of a group of people, each of whom has two sets of scores, you will end up with what’s called a scatterplot. (See Figure 5.3.)

As illustrated in Figure 5.3, as the dots come close to forming a diagonal line, the scores on the x-axis are more closely related to scores on the y-axis, or the

0 to 10.3 5 weak 0 to 20.3 5 weak

10.4 to 10.6 5 medium 20.4 to 20.6 5 medium

10.7 to 11.0 5 strong 20.7 to 21.0 5 strong

StrongStrong

21.0 0 11.0

Inverse Direct

Weak

FIGURE 5.2 | Correlation Coefficient

Positive Correlations Move in the Same Direction

Negative Correlations Move in Opposite Directions

Social Connectedness

Social Connectedness

Social Connectedness

Subjective Well-Being

Life Satisfaction

Subjective Well-Being

Life Satisfaction

OR

OR Social

Connectedness

Social

Connectedness

Social

Connectedness

Attention Deficit

Hyperactivity

Disorder

Depression

Academic

Achievement

Depression

Academic

Achievement

OR

OR

Attention Deficit

Hyperactivity

Disorder

FIGURE 5.1 | Examples of Positive and Negative Correlations

Scatterplot Graph showing two or more sets of test scores

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 85

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

correlation becomes closer to plus or minus 1.0. As the dots become more random (little relationship between scores on the two tests), the correlation approaches zero. Additionally, it is called a positive correlation if the general slope of the dots rises from left to right and a negative correlation if it rises from the right to the left.

COEFFICIENT OF DETERMINATION (SHARED VARIANCE)

Squaring the correlation coefficient gives you the coefficient of determination, or shared variance between two variables. This variance is a statement about underly- ing factors that account for the relationship between two variables. Thus, if the correlation is 0.70, the square is 0.49, which represents a percentage of shared variance—in this case, 49 percent. For instance, in one study, a correlation of 0.85 was found between scores on a test of depression and scores on a test that mea- sured anxiety (Cole, Truglio, & Peeke, 1997). Therefore, in this case it can be said that 72% of the variance is shared variance (0.85 squared), or, in other words, a large percentage of similar factors underlie feelings of depression and anxiety as measured by these tests. (See Figure 5.4.) What might some of these factors be? An educated guess might lead us to think that they could be environmental stres- sors (e.g., job loss, relationship problems, etc.), cognitive schemata (ways of under- standing the world), chemical imbalance, and so forth. All of these factors could trigger feelings of depression and of anxiety. Of course, we need to keep in mind that 28% of the variance is not shared, which means that other factors differen- tially affect whether one feels depressed or anxious. Keep the concept of coefficient of determination in mind when you read the next section on validity.

r 5 1.0 r 5 0.70

r 5 0.0 r 5 20.40

r 5 0.30

r 5 21.0

FIGURE 5.3 | Scatterplot Charts and Correlational Estimates

Coefficient of determination Common factors that account for a relationship; square of the correlation coefficient

© Ce ng ag e Le ar ni ng

86 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

VALIDITY Validity is a unitary concept. It is the degree to which all of the accumulated evidence supports the intended interpretation of test scores for the intended purpose. (American Educational Research Association, [AERA], 1999, p. 11)

How well does a test measure what it’s supposed to measure? That is the primary question that validity attempts to answer. Over the years, a number of methods have been developed to determine the validity of a test. Logically, it makes sense that the more methods one can use to provide evidence that a test is valid, the stronger the case one can make that the interpretation of test scores is accurate for the manner in which the test is being used (AERA, 1999). Various types of validity that help to provide the evidence needed to demonstrate the worthiness of a test include content validity; criterion-related validity, which includes con- current validity and predictive validity; and construct validity, which includes methods of experimental design, factor analysis, convergent validity, and dis- criminant validity.

Content Validity Probably the most basic form of validity is content validity. As with all types of validity, its name is reflective of what it attempts to show: Is the content of the test valid for the kind of test it is? In assuring content validity, a process is used to develop items for a test, such as examining established books in the field, gathering information from experts, and examining curriculum guides. In demonstrating con- tent validity, test publishers need to do the following:

Step 1: Show that the test developer adequately surveyed the domain (e.g., examining books and curriculum guides, consulting with experts, etc.)

Step 2: Show that the content of the test matches what was found in the survey of the domain.

Step 3: Show that test items accurately reflect the content.

Depression Anxiety

Shared trait variance

FIGURE 5.4 | Shared Variance Between Depression and Anxiety(r ¼ 0.85, r 2 ¼ 0.72)

© Ce ng ag e Le ar ni ng

Validity Evidence supporting the use of test scores

Content validity Evidence that test items represent the proper domain

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 87

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Step 4: Show that the number of items for each content area matches the relative importance of these items as reflected in the survey of the domain (Wolfe & Smith, 2007).

To illustrate this process, let’s look at the creation of a fourth-grade math achievement test to be used nationally. In demonstrating content validity, the test developer should do the following:

Step 1: Show that information for the test was gathered from places such as fourth-grade math books, teachers of fourth-grade math, curriculum specialists in the schools, professors at colleges who are training fourth-grade math tea- chers, and so forth.

Step 2: Show that the content was chosen based on the information that was gath- ered (e.g., addition, subtraction, multiplication, division, decimals, and fractions).

Step 3: Show that the items reflect the content chosen.

Step 4: Show that the number of items for each content area reflects the rela- tive importance of that area (e.g., multiplication and division might be empha- sized more than the other items). (See Figure 5.5.)

Despite one’s painstaking attempts to assure content validity, not all fourth-grade teachers would be teaching the same math content, so the test might hold more valid- ity for some fourth-grade classes than for others. Thus, content validity is somewhat contextual and depends on who is taking the test (Goodwin, 2002a, 2002b).

Face validity, which is not considered an actual type of validity, is sometimes confused with content validity. Face validity has to do with how the test superfi- cially looks. If you were examining the test, would it appear to be measuring what it is supposed to measure? Although most tests should have face validity, some tests could be valid yet not have it. For instance, there might be some items on a

Survey of domain

Item 1 Item 2 Item 3 Item 4 Item 5

3 3

3 2

3 1

3 2

3 2.5

Step 3: Test items reflect content

Step 4: Adjusted for relative importance

Content matches domain

Step 1: Survey the domain

Step 2: Content matches domain

FIGURE 5.5 | Establishing Content Validity

Face validity Superficial appearance of a test— not true validity

© Ce ng ag e Le ar ni ng

20 15

88 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

personality test that, on the surface, do not seem to be measuring a quality that the test is attempting to assess. Look, for example, at the following hypothetical item below that is assessing whether or not an individual may have a panic disorder:

In the past week, check the following symptoms that you have experienced: _____a. sweaty _____c. racing heart _____e. distracted myself _____b. left event early _____d. trouble breathing _____f. avoided events

Even though the symptoms listed in the item might not seem obvious to a per- son unfamiliar with panic disorder, they are often associated with the disorder, as noted in the Diagnostic and Statistical Manual of Mental Disorders, fifth edition (DSM-5; American Psychiatric Association, 2013). On the surface, this item might not seem to be measuring the construct, but individuals with panic disorder often have some or all of these symptoms, and thus such a question is conceivable and might be important to ask. Therefore, the item might be important to the content validity of the test. (See Exercise 5.1.)

Criterion-Related Validity What is the relationship between a test and a criterion (external source) that the test should be related to? That is the essence of criterion-related validity. Two types of criterion-related validity generally described are concurrent validity and predictive validity.

Concurrent Validity Concurrent, or “here and now,” validity occurs when a test is shown to be related to an external source that can be measured at around the same time the test is being given. For instance, suppose I develop a test to measure the ten- dency toward alcohol abuse. I might give this test to 500 individuals and then have the examinee’s friends and family rank order the examinee’s use of alcohol to see if there is a correlation between his or her alcohol abuse score and use of alcohol (e.g., num- ber of drinks per day). Clearly, I would suspect a high correlation (the higher the score, the more one drank), and if I did not find one, my test would be suspect.

Predictive Validity Whereas concurrent validity relates a test to a “here and now” measure, predictive validity relates a test to a criterion in the future. This kind of validity is clearly important if a test should be predicting something about an individ- ual. For instance, the Graduate Record Examine (GREs) should have predictive validity for grade point average (GPA) in graduate school; otherwise the scores would not be valuable and should not be used. In actuality, the correlation between the GREs and GPA in graduate school is about 0.34, which is not very high. (See Table 5.1.) However, when placed in the context of other possible predictors of

Exercise 5.1 Demonstrating Content Validity

In small groups, discuss how you might show content validity for a test that measures depression. Share your ideas in class.

Criterion-related validity Relationship between test scores and another standard

Concurrent validity Relationship between test scores and another currently obtainable benchmark

Predictive validity Relationship between test scores and a future standard

© Ce ng ag e Le ar ni ng

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 89

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

graduate school success (e.g., undergraduate GPA, extracurricular activities, and inter- views), it’s not so bad. And, if we combine the use of predictors of graduate school GPA, we increase our ability to make a prediction about students’ success. For instance, the correlation between undergraduate GPA and grades in graduate school is 0.37, and when we combine the GRE scores and GPA in undergraduate school, we find a correlation of 0.46 with grades in graduate school. (See Table 5.1.)

A practical application of predictive validity is the standard error of the esti- mate (SEest), which, based on one variable, allows us to predict a range of scores on a second variable. For example, if we know students’ GRE scores and the corre- lation between GRE scores and first-year grades in graduate school, we can predict the range of GPAs students are likely to obtain in graduate school. This range of scores is sometimes referred to as a confidence interval (e.g., where students GPAs will fall 68% of the time). Actually “seeing” how well one set of scores predicts a second set of scores can give us a more realistic sense of the whether or not such scores should be used in making important decisions about individuals’ lives. In Chapter 7, we will show how a SEest is actually determined.

Another concept in the application of predictive validity is false positives and false negatives. A false positive is when an instrument incorrectly predicts a test taker will have an attribute or be successful when he or she will not. A false negative is when a test forecasts an individual will not have an attribute or will be unsuccessful when in fact he or she will. For example, let’s say you are working in a high school and use the adolescent version of the Substance Abuse Subtle Screening Inventory (SASSI, 2008–2009) to flag students who may have an addiction. After giving the instrument to 100 students, 10 are identified as having a high likelihood of having a substance dependency. If two of those students do not, in fact, have an addiction, that would be considered a false positive. If one of the 90 students whose scores indicated

TABLE 5.1 Average Estimated Correlations of GRE General Test (Verbal, Quantitative, and Analytical) Scores and Undergraduate Grade Point Average with Graduate First-Year Grade Point Average by Department Type

Type of Department

Number of Departments

Number of Examinees

Predictors

V Q A U VQA* VQAU*

All Departments 1,038 12,013 0.30 0.29 0.28 0.37 0.34 0.46

Natural Sciences 384 4,420 0.28 0.27 0.26 0.36 0.31 0.44

Engineering 87 1,066 0.27 0.22 0.24 0.38 0.30 0.44

Social Sciences 352 4,211 0.33 0.32 0.30 0.38 0.37 0.48

Humanities and Arts 115 1,219 0.30 0.33 0.27 0.37 0.34 0.46

Education 86 901 0.31 0.30 0.29 0.35 0.36 0.47

Business 14 196 0.28 0.28 0.25 0.39 0.31 0.47

V ¼ GRE verbal, Q ¼ GRE quantitative, A ¼ GRE analytical, U ¼ undergraduate grade point average *Combination of individual predictors.

Source: Graduate Record Examinations, 2004–2005. GRE materials selected from 2004–2005. Guide to the Use of Scores, p. 22. Reprinted by permission of Educational Testing Service, the copyright owner. Copyright © 2013 Educational Testing Service. www.ets.org

90 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

a low probability of a dependency, but in actuality had a diagnosable substance dependence, that would be considered a false negative. Although it would be better to have neither false positives or negatives, in the case of the SASSI, the designers chose to err toward having a greater chance of false positives than false negatives. Why do you think that is? (See Exercise 5.2.)

Construct Validity What is intelligence? When teaching an assessment class, we have found that this often becomes a “hot” topic for discussion, with students arguing over what they believe actually makes up this construct. Construct validity is the scientific basis for showing that a construct (idea, concept, or model), such as intelligence, is being measured by a test. Showing evidence of construct validity is particularly important when developing tests that measure abstract domains, such as intelli- gence, self-concept, depression, anxiety, empathy, and many other personality char- acteristics. On the other hand, construct validity becomes much less of an issue when one is measuring well-defined areas, such as achievement in geometry.

To demonstrate construct validity, it is often useful to provide multiple sources of evidence that the construct we are measuring is indeed being exposed. Although some have argued that almost any kind of validity is evidence that the construct being measured exists (Goodwin, 2002a), a more restrictive definition of construct validity includes an analysis of a test through one or more of the following methods: (1) experimental design, (2) factor analysis, (3) convergence with other instruments, and/or (4) discrimination with other measures.

Experimental Design Validity You’ve just developed your new test of depres- sion, and you’re very proud of it. Of course, you want to show that it is indeed valid. You already have shown that it has content validity by developing items through an examination of DSM-5, scholarly journal articles, and consultation with experts. However, now you want to show that the test indeed works—that it measures your construct. So you approach a number of expert clinicians who work with depressed clients and ask them to identify new clients on their caseloads who are depressed. You then request that they administer your depression test prior to and at the end of six months of treatment. What should you expect to find? Clearly, if the test is good and does measure the construct depression, it should be

Exercise 5.2 Demonstrating Criterion-Related Validity

In class, form small groups, some of which will examine concurrent validity while others will examine predictive validity. In your groups, choose a test from the following list or come up with another test. Next, discuss criteria with which you might compare your test.

Possible Tests:

High school achievement test

Test of anxiety

SATs, GREs, MATs Test of hypochondriasis

Intelligence test Clerical aptitude test

Depression test First-grade reading test

Construct validity Evidence that an idea or concept is being measured by a test

Experimental design validity Using experimentation to show that a test measures a concept

© Ce ng ag e Le ar ni ng

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 91

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

able to accurately reflect the change in these clients—in this case, a significant decrease in depression. If not, you need to go back to the drawing board. Experi- mentally based construct validity thus confirms the hypothesis you developed to scientifically show that your construct exists. Sometimes, when a number of studies are completed, you will find authors conducting a “meta-analysis,” which statisti- cally analyzes all of these studies in an effort to offer broader evidence of the exis- tence of the construct being examined.

Factor Analysis Factor analysis, another method to show construct validity, demonstrates the statistical relationship among subscales or items of a test. Suppose, for instance, that your depression test has subscales that measure hopelessness, suicidal ideation, and self-esteem. Theoretically, you would expect them all to be related to your larger test score that measures depression. In addition you would expect them to be somewhat related, but not largely related, to one another. After all, they are mea- suring something different from one another (although they all make up the larger test “depression”). So after giving your new depression test to a sample of 100 people, you run a statistical program that shows that each of the items measuring the domain of hopelessness correlates fairly highly with the domain; that is, people who score high on hopelessness tend to score high on each of the items in that domain. You find the same for suicidal ideation. However, you find one of the items that is supposed to be measuring self-esteem does not correlate highly with that domain (i.e., doesn’t load for that factor). So you decide to delete that item or rewrite it and try again. In addition, you also find that the three subscales do not correlate particularly highly with one another, thus showing that they are discrete, or separate scales (although, they all cor- relate highly with the total test score, which is a combination of the three scales).

Convergence with Other Instruments (Convergent Validity) Believing that your test does measure depression, you would, of course, expect your test to be related to other existing well-known valid measures of depression. Thus, you decide to correlate your test with the Beck Depression Inventory II (BDI-II) (Beck, Steer, & Brown, 2004), a well-known test of depression. You give the BDI-II and your test to 500 subjects and you find a relatively high correlation between the two scores, perhaps 0.75. You’re satisfied, perhaps even glad you didn’t get a higher correlation. After all, isn’t your test different (and a little better!) than the well-known existing test? Thus, you expect your correlation to be less than perfect. Thus, convergent validity occurs when you find a significant pos- itive correlation between your test and other existing measures of a similar nature. Sometimes this relationship involves very similar kinds of instruments, but other times you would be looking for correlations between your test and variables that may seem only somewhat related. For instance, if you correlate your test with a test to measure despair, you would expect a positive correlation. However, because despair is theoreti- cally different from depression, you would expect a lower correlation—maybe 0.4.

Discrimination with Other Measures (Discriminant Validity) With convergent validity you’re expecting to find a relationship between your test and other variables of a similar nature. Discriminant validity, in a sense, is the opposite in that you’re looking to find little or no relationship between your test and measures of constructs that are not theoretically related to your test. For instance, with your test of depres- sion, you might want to compare your test scores with an existing test that measures

Factor analysis Statistically examining the relationship between subscales and the larger construct

Convergent validity Relationship between a test and other similar tests

Discriminant validity Showing a lack of relationship between a test and other dissimilar tests

92 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

anxiety. In this case, you give 500 subjects your test of depression as well as a valid test to measure anxiety, looking to find little or no relationship. Consider how important discriminant validity might be if you’re working with clients. Clients who are dealing with high amounts of anxiety will often present with depressive features. If you were able to give your client a test that could discriminate depression from anxiety, you would get a better picture of your client. Making a more accurate diag- nosis could dramatically affect your treatment plan, as approaches to working with depressed clients will vary from approaches to working with anxious clients.

Visual Representation of Types of Validity Figure 5.6 is a visual representation of the seven types of validity across the three broader categories. For content validity, you will notice a previous figure that

Content Validity

Criterion-Related Validity

Construct Validity

New Test

• Item 1 • Item 2 • Item 3 • Item 4

Content Matches Domain

Concurrent Validity Predictive Validity

Experimental Design Validity Factor Analysis

Convergent Validity Discriminant Validity

New Test

• Item 1 • Item 2 • Item 3 • Item 4

Total

Items Factors

New Test

• Item 1 • Item 2 • Item 3 • Item 4

New Test

• Item 1 • Item 2 • Item 3 • Item 4

Future Event

New Test

• Item 1 • Item 2 • Item 3 • Item 4

New Test

• Item 1 • Item 2 • Item 3 • Item 4

r

Existing Similar Test • Item 1 • Item 2 • Item 3

New Test

• Item 1 • Item 2 • Item 3 • Item 4

r

Existing Dissimilar Test • Item 1 • Item 2 • Item 3

Survey of domain

Item 1 Item 2 Item 3 Item 4 Item 5 3 3

3 2

3 1

3 2

3 2.5

Content matches domain

FIGURE 5.6 | Visual Representation of Types of Validity

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 93

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

exemplifies the test developer demonstrating he or she properly surveyed the field of information on the topic and rigorously developed the items from that survey. Both concurrent and predictive validity are represented under criterion-related validity. Concurrent validity is symbolized by a new test being compared to a known quantity (a ruler), and predictive validity hopefully allows a test to estimate a future event. Four types of construct validity are mentioned in this book. In experimental design validity, we give our new test to a group of people receiving an intervention hoping our instrument captures the desired change. The factor analysis figure represents the statistical process where similar test items should correlate with other like items within the instrument creating subgroups called factors or dimensions. In convergent validity, our new test will hopefully have a high correlation (r) with an existing instrument meant to measure the same construct. Divergent validity is the opposite, and we want our new test has to have a low correlation with another test that measures a different construct. (See Exercise 5.3.)

RELIABILITY Test reliability can be compared to eating at the same restaurant over and over again. If it’s a highly reliable restaurant that serves great food, you can order the same meal multiple times and it will always taste the same, and it will always be great. On the other hand, a restaurant that has poor reliability can’t be trusted. Perhaps each cook makes it differently or uses slightly different or bad ingredients, and maybe the restau- rant has an atmosphere that turns you off and makes you not enjoy your meal.

Similarly, a test with high reliability is put together well—it has the right ingredients—and when you take the test, it’s in an environment that is conducive to you producing your best results. The test is made well and the environment is optimal for the test-taking situation. Hypothetically, if your knowledge base stayed the same and you took this test over and over again, you would score about the same each time (e.g., if it were an achievement test, you didn’t look up the answers or study the content prior to taking the test again). On the other hand, if you were to take a test with poor reliability, over and over again, your scores would fluctuate—be higher or lower each time you took it. In this case, problems with the test and with the testing environment would cause you to answer items in a dif- ferent manner every time you took the test.

Reliability can be defined as “the degree to which test scores are free from errors of measurement” (AERA, 1999, p. 180). Hypothetically, if we had a perfect

Exercise 5.3 Demonstrating Construct Validity

In small groups, show how you would develop the con- struct validity of a test that measures self-actualizing values (e.g., being in touch with feelings, being spontane- ous, being accepting and nonjudgmental, and showing

empathy). Try to touch on many of the methods noted in this section of the chapter, including experimental design, factor analysis, convergent validity, and discrimi- nant validity. Present your group’s proposal to the class.

Reliability Amount of freedom from measurement error—consistency of test scores

© Ce ng ag e Le ar ni ng

94 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

test and the perfect environment, and a person who remained exactly the same each time he or she took the test, this individual would always score the same on the test, even if he or she took it 1000 times. As we all know, there are no perfect tests or environments, so we always end up with measurement error—or factors that affect one’s score on a test. Clearly, the more we can reduce error on tests, the more accurate will be the examinee’s score.

Have you ever taken a test and received a score that you felt didn’t reflect how you actually performed on the test? Some of that difference may be attributed to measurement error. Many factors can cause such error, including poorly worded questions, poor test-taking instructions, test-taker anxiety or fatigue, and distrac- tions in the testing room, to name just a few. On the other hand, part of the dis- crepancy may be due to lack of knowledge of the subject or lack of awareness of self, and sometimes it’s easier to blame a test than ourselves for scores that we did not expect. If we know the amount of error on a test, as measured by reliability estimates, then we should be able to determine, to some degree, whether our unex- pected score was due to measurement error or to performance issues.

Test creators will evaluate their new instrument to determine if the scores it pro- duces are reliable (consistent and dependable) and publish this information in the test manual, usually reported in the form of a reliability (correlation) coefficient. The closer the reliability estimate is to 1.0, the less error there is on the test. A good reli- ability estimate is sometimes a function of the kind of measurement being assessed, although there is no magic cutoff point for what makes a good reliability estimate (Heppner, Wampold, & Kivlighan, 2008). For instance, teacher-made tests generally will have lower reliability estimates, perhaps around 0.7, because they do not go through the same type of scrutiny as some other kinds of tests, such as national achievement tests that often have reliability estimates in the 0.90s. Also, personality tests generally have lower reliability estimates than ability tests because the construct being measured is more abstract and because personality tends to fluctuate, thus making these constructs more difficult to define and measure. Some ways of measur- ing reliability include test-retest, alternative forms, and internal consistency. After looking at these types of reliability, this section concludes with a discussion about item response theory, a more recent method of examining reliability.

Test-Retest Reliability A relatively simple way to determine whether an instrument is reliable is to give the test twice to the same group of people. For example, imagine you have 500 people who take a test. A day later they take the same test at the same place. You would then correlate the scores from the first test with those from the second test. The closer the two sets of scores, the more reliable the test. Although reliability coefficients pro- vide individual test-takers with estimates about the stability of their individual scores, the actual process of gathering reliability information has to do with average fluctua- tion of many people’s scores. Thus, although one person’s score might fluctuate a lot, an instrument can still have high reliability if most scores have little fluctuation.

Accuracy of the test-retest method can be affected by several factors. For instance, depending on the time between test administrations, people may forget information. On the other hand, sometimes people might learn more about a specific

Test-retest reliability Relationship between scores from one test given at two different administrations

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 95

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

subject by reading a book, listening to the radio, searching on the Internet, and so on. Also, some people might score higher on a second test because they are more familiar with the test format. As you might expect, test-retest reliability tends to be more effective in areas that are less likely to change over time, such as intelligence.

Alternate, Parallel, or Equivalent Forms Reliability Another method for determining reliability is to make two or more alternate, parallel, or equivalent forms of the same test. These alternate forms are created to mimic one another, yet are different enough to eliminate some of the problems found in test- retest reliability (e.g., looking up an answer). In this case, rather than giving the same test twice, the examiner gives the alternate form the second time. Clearly, a challenge involved in this kind of reliability is to assure that both forms of the test use the same or very similar directions, format, and number of questions, and are equal in diffi- culty and content. Although creating alternate forms eliminates some of the problems found in test-retest reliability, difficulty can arise in ensuring that the two forms are truly parallel. Also, because creating a parallel form is labor-intensive and costly, it is often not a practical kind of reliability to implement. If this kind of reliability is used, the burden is on the test creator to prove that both tests are truly equal.

Internal Consistency A third type of reliability is called internal consistency. In this case, a determination is made as to how scores on individual items relate to each other or to the test as a whole. For instance, it would make sense that individuals who score high on a test of depression should, on average, respond to all the items on the test in a manner that indicates depressive ideation. If not, one might wonder why a particular item is not assessing depression. This kind of reliability is called internal consistency because you are looking within the test itself, or not going “outside of the test” to determine reliabil- ity estimates as you would be with test-retest or parallel forms reliability. Some types of internal consistency reliability include split-half (or odd-even), Cronbach’s coefficient alpha, and Kuder–Richardson, each of which is discussed in the following sections.

Split-Half or Odd-Even Reliability The most basic form of internal consistency reliability is called split-half or odd-even reliability. This method, which requires only one form and one administration of the test, splits the test in half and correlates the scores of one half of the test with the other half. Again, imagine 500 individuals who come in to take a test. After they have finished, we gather their responses, split each of their tests in half, and score each half, almost as if each individual had taken two different tests. The scores on the two halves are then correlated with one another. Obviously, advantages of this kind of reliability include only having to give the test once and not having to create a separate alternate form. One potential pitfall of this form of reliability would arise if the two halves of the test were not parallel or equivalent, such as when a test gets progressively more difficult. In that case, you might end up correlating the first half of the test with the second nonequivalent half, which can give you an inaccurately low estimate of your reliability. One com- mon method to alleviate this potential error is to split the test into odd-numbered

Alternate forms reliability Relationship between scores from two similar versions of the same test

Internal consistency Reliability measured statistically by going “within” the test

Split-half reliability Correlating one half of a test against the other half

96 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

items and even-numbered items, although even this method runs the risk of not producing parallel halves.

Another disadvantage of this kind of reliability is that by turning one test into two, you have made it half as long. Generally speaking, one obtains a more accurate reading of the reliability of a test the longer it is; consequently, shortening the test may decrease its reliability. One common method to mathematically compensate for shortening the number of correlations is to use the Spearman–Brown equation (Brown, 1910; Spearman, 1910). The generally used Spearman–Brown formula is

Spearman ¼ Brown reliability ¼ 2rhh 1 þ rhh

where rhh is the split-half or odd-even reliability estimate. For example, if the reli- ability is 0.70, you would multiply 0.70 by 2, and then divide by the sum of 1 þ 0.70. This would be 1.40 divided by 1.70, which gives you a reliability esti- mate of 0.82 for the whole test. So if a test manual states that split-half reliability was used, check to see if the Spearman–Brown formula was applied. If it was not, the test might be somewhat more reliable than actually noted.

Cronbach’s Coefficient Alpha and Kuder–Richardson Although used for dif- ferent kinds of tests, coefficient alpha and Kuder–Richardson forms of reliability are also considered internal consistency methods that attempt to estimate the reli- ability of all the possible split-half combinations (Cronbach, 1951). In brief, they do this by correlating the scores for each item on the test with the total score on the test and finding the average correlation for all of the items (Salkind, 2011). Clearly, individuals who are producing depressed scores should be responding to test items in a manner that indicates depression, and individuals who are producing scores that indicate they are not depressed should be responding to test items in a manner that do not indicate depression. If particular items are not able to discrimi- nate between depressed and nondepressed individuals, there are likely some pro- blems with those items and this will be reflected in an overall lower-internal consistency estimate. Obtaining the correlation between each item and the whole test score for each person, and then finding the average of all those correlations, used to be a tedious process. However, with computers, such reliability estimates can now be configured in a millisecond and therefore have become very popular. The difference between these two formulas is that Kuder–Richardson can be used only with tests that have right and wrong answers, such as achievement tests, whereas coefficient alpha can be applied to assessment instruments that result in varied types of responses, such as rating scales. If you are interested in the actual formulas for Kuder–Richardson and coefficient alpha, you can find them in Appendix E.

Visual Representation of Types of Reliability Figure 5.7 is a visual representation of the three types of reliability. Notice that test-retest is represented by the same test being given again with a time difference, alternative forms is represented by two similar forms that are not affected by time, and two types of internal consistency are shown: split-half, which is represented by a test being split, and coefficient alpha and Kuder–Richardson being shown as a

Coefficient alpha or Kuder–Richardson Reliability based on a mathematical comparison of individual items with one another and total score

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 97

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

grid that represents each item being related to the whole test. These visual represen- tations can help you remember each form.

Item Response Theory: Another Way of Looking at Reliability Most of the concepts in this book fall under the description of classical test theory. Item response theory (IRT) is an extension of this (Bechger, Maris, Verstralen, & Beguin, 2003; Reid, Kolakowsky-Hayner, Lewis, & Armstrong, 2007). Classical test theory, which originated with Spearman (1904), assumes that measurement error exists in every instrument. Hence, a respondent’s true score is equal to his or her observed score plus or minus measurement error.

True score ¼ observed score � measurement error Reliability is the opposite of measurement error, so we can see why it is important in classical test theory. As the reliability increases, the measurement error decreases. If we have a reliability � of 0.60, then we have a lot of measurement error, which makes it more difficult to estimate a true score. An instrument with 0.93 reliability has little error, allowing us to feel more confident. In classical test theory, test items are viewed as a whole, with the desire to reduce overall measurement error (increase reliability) so that we can estimate the true score.

IRT, on the other hand, examines items individually for their ability to mea- sure the trait being examined (Ostini & Nering, 2006). This provides IRT test developers with a slight advantage over classical test theorists for improving reli- ability, because IRT provides more sophisticated information regarding individual items (MacDonald & Paunonen, 2002). However, there appears to be some hesi- tancy to switch from classical test theory to IRT since it is slightly more complex (Progar, Socan, & Pec, 2008). The item characteristic curve is a tool used in IRT that assumes that as people’s abilities increase, their probability of answering an

Internal Consistency

Split-Half Cronbach , KR-20, KR-21

A A

Test-Retest Alternate Forms

A BTime A A

FIGURE 5.7 | Visual Representation of Types of Reliability

IRT Examining each item for its ability to discriminate as a function of the con- struct being measured

© Ce ng ag e Le ar ni ng

98 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

item correctly increases. We can plot examinees’ ability along the x-axis and the probability of getting the item correct along the y-axis. If we were developing an IQ test using IRT, we could graph our sample data for each question using the item characteristic curve as shown in Figure 5.8. We can see that if someone has average ability (i.e., IQ of 100) along the bottom of the graph, he or she has a 50% chance of getting the item correct on the left side of the chart. If an individual has less abil- ity, say an IQ of 85, his or her chance of getting the item correct has been reduced to 25%. Similarly, you can see as one’s ability increases along the bottom toward the right, the chance of getting items correct begins to approach 100%.

The item characteristic curve provides a lot of information. We can see that if the shape of the “S” flattens out, the item has less ability to discriminate or provide a range of probabilities of getting a correct or incorrect response. If the S is tall, it means the item is creating strong differentiation across ability. IRT allows test developers to write items targeting specific ranges of ability. For example, we could write items that would have 50% probability of answering correctly at the 115 IQ level. If we were administering our test via computer, we could provide dif- ferent questions for different abilities dependent on how the previous question was answered. If Jo Anne got our first question correct at the 100 IQ level, we could then give her a question at the 115 level. If she missed this item, we could give her one at the 107 level. We can continually fine-tune Jo Anne’s score this way.

CROSS-CULTURAL FAIRNESS Although inextricably related to the validity and reliability of a test, cross-cultural fairness deserves a separate section due to the importance we as a nation place on issues of fairness, especially with regard to diversity. Awareness of cross-cultural factors and how they impact the development, administration, scoring, and interpretation of assessment procedures are critically important and should always be considered when assessing an individual. When using tests, mental

0.50

1.0

0.0 100 115 130 145857055

0.25

IQ Ability

0.75

P ro

b a

b il it y o

f C

o rr

e c

t A

n s w

e r

FIGURE 5.8 | Item Characteristic Curve

Cross-cultural fairness Degree to which cultural background, class, disability, and gender do not affect test results

© Ce ng ag e Le ar ni ng

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 99

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

health professionals should take into account how such issues may impact a broad array of minorities and individuals with disabilities. For instance, the Stan- dards for Qualifications of Test Users (American Counseling Association, 2003) states that those involved with testing should

be committed to fairness in every aspect of testing. Information gained and decisions made about the client or student are valid only to the degree that the test accurately and fairly assesses the client’s or student’s characteristics. Test selection and inter- pretation are done with an awareness of the degree to which items may be culturally biased or the norming sample not reflective or inclusive of the client’s or student’s diversity. Test users understand that age and physical disability differences may impact the client’s ability to perceive and respond to test items. Test scores are interpreted in light of the cultural, ethnic, disability, or linguistic factors that may impact an individual’s score. These include visual, auditory, and mobility disabilities that may require appropriate accommodation in test administration and scoring. Test users understand that certain types of norms and test score interpretation may be inappropriate, depending on the nature and purpose of the testing. (Standard 6)

Although much emphasis is placed on cross-cultural fairness today, this has not always been the case (recall the cultural bias of the Army Alpha described in Chapter 1). In fact, issues of bias in testing did not garner much attention until the civil rights movement of the 1960s when there was concern that African Americans and Hispanics were being compared unfairly against the White majority. In a series of court decisions, it was decided that some tests, such as IQ and achievement tests, could not be used to track students because minority students were being disproportionately placed in lower-achieving classrooms as the result of tests that may not have been accurately assessing their ability. (See Hobson v. Hansen, 1967; Moses v. Washington Parish School Board, 1969.) And in 1971, the U.S. Supreme Court case of Griggs v. Duke Power Company asserted that tests used for hiring and advancement at work must show that they can predict job performance for all groups. Over the years, a number of laws have been passed that impinge on the use of tests and assert the rights of all individuals to be tested fairly. (See Chapter 2 for expanded definitions of some of these laws.) These laws, and their relationship to assessment, are briefly described here.

Americans with Disabilities Act. This law states that accommodations must be made for individuals who are taking tests for employment and that testing must be shown to be relevant to the job in question.

The Family Education Rights and Privacy Act (FERPA) of 1974. Also known as the Buckley Amendment, this law affirms the right of all individuals to review their school records, including test records.

Carl Perkins Act (PL 98-524). This law assures that individuals who have disabilities or who are disadvantaged have access to vocational assessment, counseling, and placement.

Civil Rights Acts (1964 and Amendments). This series of laws asserts that any test used for employment or promotion must be shown to be suitable and

Many laws have affirmed the rights of minorities and others to fairness in testing

100 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

valid for the job in question. If not, alternative means of assessment must be provided. Differential test cutoffs are not allowed.

The Freedom of Information Act. This law assures the right of individuals to access their federal records, including test records. Most states have expanded this law so that it also applies to state records.

The Individuals with Disabilities Education Act (IDEA) (fka PL 94-142). These legislative acts assure the right of students to be tested, at a school system’s expense, if they are suspected of having a disability that interferes with learning. These laws assert that schools must make accommodations, within the least restrictive environment, for students with learning disabilities.

Section 504 of the Rehabilitation Act. Relative to assessment, any instrument used to measure appropriateness for a program or service must be measuring the individual’s ability, not be a reflection of his or her disability.

Today, it is critical that cognitive differences among groups of individuals rep- resent differences in ability, not differences that result from cultural identification, gender, age, or so forth, and that tests used for predictive purposes show that they predict accurately for all groups of people (e.g., the SATs, MATs, GREs) (Berry, Clark, & McClure, 2011; Bobko & Roth, 2012; Hartman, McDaniel, & Whetzel, 2004). However, it should also be stressed that such differences do exist, and one of the greatest challenges today is to understand why they exist to develop ways to eliminate such differences (Hartman et al., 2004). These days, cognitive differ- ences are generally traced to environmental factors, as is evidenced by the No Child Left Behind Act, which assumes that all children can succeed in school and demands that school systems show that all children pass minimal competencies, regardless of their gender, culture, or disability (U.S. Department of Education, 2005, 2008a, 2008b). (See Exercise 5.4.)

BOX 5.1 The Use of Intelligence Tests with Minorities: Confusion and Bedlam

The use of intelligence tests with culturally diverse populations has long been an area of controversy. Over the years, states have found intelligence tests biased and banned their use in certain circum- stances with some groups (Gold, 1987; Swenson, 1997). One case in California in 1987 highlighted this controversy. Ms. Mary Amaya was concerned that her son was being recommended for remedial courses he did not need. Having had an older son who was found to not need such assistance only after he was tested, Ms. Amaya requested testing

with an intelligence test for her other son. However, since the incident with her first son, California decided that intelligence tests were culturally biased and thus banned their use for members of certain groups. Despite the fact that Ms. Amaya was requesting the use of an intelligence test, it was found that she had no legislative right to have the test given to her son. Although California subse- quently reversed its ban, concerns about racial bias in testing continue today (Ortiz, Ochoa, & Dynda, 2012).

© Ce ng ag e Le ar ni ng

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 101

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

PRACTICALITY Although tests should have good validity and reliability, as well as cross-cultural fair- ness, it is also important for the test to be practical. For instance, would it make sense to give a 2-hour exam to first- or second-graders, or might a shorter test that may be somewhat less valid do nearly as well? Or, is it feasible to use a Wechsler Scale of Intelligence as prescreening for learning disabilities to hundreds of possible learning- disabled students in a school system? Such a test is given individually, takes 1 to 2 hours to administer, and requires another 1 to 2 hours to write up. Decisions about test selection should be made carefully because they have an impact on the person being tested, the examiner, and at times, the institutions that are requiring testing. Examples of a few of the major practical concerns examiners face include time, cost, format, readability, and ease of administration, scoring, and interpretation.

Time As just noted, the amount of time it takes to administer a test can clearly affect whether or not it is used. Time factors tend to be related to the attention span of the client you are testing, to the amount of time allotted for testing in a particular setting, and to the final cost of testing (time is money!).

Cost With increasingly limited insurance reimbursements as well as funding cutbacks, counselors and therapists in private practice, public and private agencies, and school systems are forced to hold down expenses. Thus, the cost of testing is an important factor in making a decision about which test to use. For instance, it would be nice to have all high school seniors take an interest inventory to help them in their career decision-making process; however, given that this might cost about $10 per student, it may not be a wise economic decision for many school systems. And what decision would you make if you were working with a person who had just lost his or her job, was short on finances, and needed to take a battery of tests that might cost $500 when other tests, perhaps somewhat less valid and reliable, might cost $100?

Format The format of a test should also be considered when deciding which test to use. Some format issues include clarity of the print, print size, sequencing of questions, and the type of questions being used. Although there is a debate regarding the optimal number of distracters in multiple-choice questions, some have found that for many individuals, multiple-choice questions tend to lessen test anxiety while tests requiring constructive

Exercise 5.4 Differences in Ability Scores as a Function of Culture

In small groups, discuss why there might be differences among cultural groups on their ability scores. Then

discuss ways in which such differences could be ame- liorated. Discuss your reasons and solutions in class.

Practicality Feasibility considerations in test selection and administration

© Ce ng ag e Le ar ni ng

102 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

responses (i.e., open-ended, essay questions) generally increased test-taker anxiety, anx- iety that could reduce cognitive clarity (Hudson & Treagust, 2013; Rodriguez, 2005). Also, it appears that online formats produce less test anxiety than paper-and-pencil meth- ods (Stowell & Bennett, 2010). In addition, the format of a test could affect scores differ- entially as a function of gender. In fact, in one study men perceived testing as more fair when only given a multiple choice format while women perceived testing as more fair when given a choice between multiple-choice and essay-type responses (Mauldin, 2009). The format of a test should always be considered when choosing which test to use.

Readability Readability, or the ability of the examinee to comprehend what he or she is read- ing, is of course critical for all tests other than vocabulary or reading-level tests (Hewitt & Homan, 2004). No one should be surprised that the reliability of test scores is reduced when readability is not controlled for. Thus, it has been suggested that each item on a test be scrutinized for readability.

Ease of Administration, Scoring, and Interpretation In deciding how practical a test is to use, the ease of test administration, scoring, and interpretation should always be considered. This dimension involves a number of factors, including the following:

1. the ease of understanding and using test manuals and related information; 2. the number of individuals taking the test and whether or not their numbers

affect the ease of administering the instrument; 3. the kind of training and education needed to administer the test, score the test,

and interpret test results; 4. the “turnaround time” in scoring the test and obtaining the results; 5. the amount of time needed to explain test results to examinees; and 6. associated materials that may be helpful in explaining test scores to examinees

(e.g., printed sheets generated by the publisher).

SELECTING AND ADMINISTERING A GOOD TEST Now you know that a test should be valid, reliable, cross-culturally fair, and prac- tical. But, with thousands of tests to choose from, how exactly does one find a test that meets the specific needs of the testing situation? A number of steps will assist you in the process.

Step 1: Determine the Goals of Your Client The information-gathering process, which was discussed in detail in Chapter 4, is critical to one’s ability to determine what is happening with a client. As informa- tion is gathered from a client, client goals will become clearer and these goals will help you determine what assessment instruments might be valuable in helping to reach client goals.

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 103

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Step 2: Choose Instrument types to Reach Client Goals Choosing the right instrument is not always an easy process. However, as soon as you have identified client goals, you have gained a sense of the kinds of instru- ments that may be helpful. For instance, if a client’s goals include developing a career path and making a decision about a job, you might want to consider admin- istering an interest inventory and a multiple aptitude test.

Step 3: Access Information About Possible Instruments A number of sources exist today to help you choose possible assessment instru- ments. Some of these are described here.

Source books on testing. Source books on testing provide important informa- tion about tests, such as the names of the author(s); the publication date; the purpose of the test; bibliographical information; information on test con- struction; the name and address of the publisher; the costs of the test; infor- mation on validity, reliability, cross-cultural issues, and practicality; reviews of the test. (See Box 5.2.) Information on scoring; and other basic informa- tion about the test.

BOX 5.2 Mental Measurement Review of the WAIS-IV

The following is a section from the second reviewer of the Wechsler Adult Intelligence Scale-IV (WAIS-IV), as found in the Mental Measurement Yearbook, 18th edition. The actual two reviews run about 11 pages, single-spaced.

DESCRIPTION. The Wechsler Adult Intelligence Scale- Fourth Edition (WAIS-IV) provides an individually administered test of intelligence for individuals between 16 and 90 years of age. The WAIS-IV provides a com- posite of intellectual functioning using 15 separate subtests that are combined into four cognitive skill cat- egories, including Verbal Comprehension (4 four subt- ests), Perceptual Reasoning (5 subtests), Working Memory (3 subtests), and Processing Speed (3 subt- ests). The total composite scale based on all subtests is interpreted as a measure of general intellectual ability.

The 15 subtests are administered in a prescribed order, which require 2 hours or less to complete for most examinees. Administration guidelines and dis- continue rules are stated clearly in the 258-page administration and scoring manual. Detailed scoring rules are provided as well.

DEVELOPMENT. The WAIS-IV is a revision of its immediate predecessors, most notably the WAIS-III. The test has been in continuous use via updated versions since 1939 and retains the same general theoretical and administrative structure. All of the WAIS tests are based on hierarchical models of intelligence such as Spearman’s g (Spearman, 1923), and the two-factor theory of Cattell (1963), which distinguishes between fluid and crystallized intelligence. Fluid ability is thought to be biologically driven and represents general ability to reason on novel tasks and unfamiliar contexts, whereas crystallized ability repre- sents reasoning and problem solving related to task- specific knowledge and schooling (Ackerman & Lohman, 2006; Carroll, 1993). Fluid and crystallized ability may be combined to create a single composite that measures a general intelligence factor, g.

Theoretically, the 15 subtests load on four cognitive factors (i.e., Verbal Comprehension, Perceptual Reasoning, Working Memory, and Processing Speed), which load on fluid and crystallized forms of intelligence, which load on the composite intellectual ability factor known as g…. (Schraw, 2010, p. 151)

© Ce ng ag e Le ar ni ng

20 15

104 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Two of the most important source books are the Buros Mental Measurements Yearbook (MMY) and its companion volume Tests in Print (Spies, Carlson, & Geisinger, 2010; Murphy, Geisinger, Carlson, & Spies, 2010). The MMY, which is in its 18th edition, offers reviews of more than 2700 tests, instruments, and screening devices. Because there are thousands of possible tests, not all of them can be included in any one edition; consequently, it is often necessary to go back to previous editions to find a specific test that you might be looking for. Most large universities carry hardcover and online editions of this source book. The MMY test reviews are classified into 18 major categories: achievement, behavior assess- ment, developmental, education, English and language, fine arts, foreign language, intelligence and general aptitude, mathematics, miscellaneous, neuropsychological, personality, reading, science, sensory-motor, social studies, speech and hearing, and vocations. Online searches of the MMY offer a quick mechanism for obtain- ing information about assessment instruments. For example, a quick search from the first to the 18th editions shows that the MMY holds 775 tests in the area of intelligence and general aptitude. Plug in the word love, and we find that 16 tests address this construct in some manner.

Publisher resource catalogs. Test publishing companies freely distribute cata- logs that describe the tests they sell. Additional information about the tests usually can be purchased from the publisher (e.g., sample kits, technical man- uals, and so on).

Journals in the field. Professional journals, especially those associated with measurement, will often describe tests used in research or give reviews of new tests in the field.

Books on testing. Textbooks that present an overview of testing are usually fairly good at highlighting a number of the better-known tests that are in use today.

Experts. School psychologists, learning disability specialists, experts at a school system’s central testing office, psychologists at agencies, and professors are some of the many experts who can be called upon as providers of testing information.

The Internet. Today, publishing companies have home pages that offer infor- mation about tests they sell. In addition, the Internet has become an increas- ingly important place in which to search for information about testing. Of course, one needs to ensure that any information from the Internet is accurate.

Step 4: Examine Validity, Reliability, Cross-Cultural Fairness, and Practicality of the Possible Instruments As you gather information about tests you might use, you will hopefully obtain information about their validity, reliability, cross-cultural fairness, and practicality. However, if the sources do not offer this information, you can do a search of jour- nal articles that may have examined these tests and you can contact the publisher to purchase a copy of the technical manual for the test. This manual should offer you all of the necessary information to make an informed judgment concerning whether or not to use the instrument in question.

Mental Measurements Yearbook A tremendous resource in finding and selecting tests

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 105

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Step 5: Choose an Instrument Wisely You’ve gone through your steps, and ideally you have some good options available to you. Most likely you narrowed your choices down to a few possible instruments. Now, it’s time for you to be wise. Examine the technical information, consider the purpose of the testing, reflect on the ease and cost of the instrument, and make a wise choice—a choice that you will feel comfortable with and a choice that is right for your client.

SUMMARY In this chapter, we examined four components of test worthiness: validity, reliability, cross-cultural fairness, and practicality. We began by explaining the concept of correlation coefficient, and we noted that it ranges between �1.0 and þ1.0 and described the strength and direction of the rela- tionship between two sets of variables such as test scores. We then noted that the coefficient of determination, or shared variance, is the square of the correlation (r2), and is a statement about the commonality between variables.

Next, we examined validity, a broad evidence- based concept that attempts to verify that a test is measuring what it is supposed to measure. Content validity is based on evidence that test items accu- rately reflect the examinee’s knowledge of the domain being measured. Criterion validity has two forms: concurrent, which demonstrates that the instrument is related to some criterion in the “now” (e.g., depression scores with clinical ratings of cli- ents), and predictive, which examines how well the instrument forecasts some future event. We discussed how the SEest and false positives and false negatives are considerations in predictive validity. Construct validity shows whether the test is properly measuring the correct concept, model, or schematic idea. Evi- dence for construct validity is often available in the form of research studies, factor analysis, and conver- gence and discrimination with other existing tests.

In the next part of the chapter, we examined reliability, or the degree to which test scores are free from measurement error (i.e., consistent and dependable). We first examined three of the most common forms of reliability: test-retest; alter- nate, parallel, or equivalent forms; and internal consistency, which includes split-half or odd- even reliability, Cronbach’s coefficient alpha,

and Kuder-Richardson reliability. Test-retest reliability is calculated by giving the same instru- ment twice to the same group of people and cor- relating the test scores. Alternate form reliability involves administering two equivalent forms of the test to the same group and correlating the scores. One type of internal consistency reliabil- ity, called split-half or odd-even reliability, is used when one-half of the test items are corre- lated with the other half. Other forms of internal consistency, such as Cronbach � and Kuder– Richardson reliability, use more complex statisti- cal calculations to find the average correlation of all test items. Finally, we concluded the section with a brief look at Item Response Theory (IRT), which is a relatively new method of examining reliability. With IRT, each item is examined indi- vidually for its ability to discriminate from other items based on the construct being measured (e.g., intelligence).

The next area of test worthiness we discussed is cross-cultural fairness. We noted that a test should accurately measure a construct regardless of one’s membership in a class, race, disability, religion, or gender. We saw how the legal system has supported the notion that testing must be fair for all groups of people and free from bias. We specifically highlighted the fact that tests that may be biased could not be used to track students and that as a result of the Supreme Court case of Griggs v. Duke Power Company (1971), tests used for hiring and advancement at work must be predictive of job performance for all groups. We also noted that a number of laws have been passed that assert the rights of all individuals to be tested fairly, including the Americans with Disabilities Act, the FERPA, the Carl Perkins Act (PL 98-524), Civil

106 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Rights Acts (1964 and Amendments), the Freedom of Information Act, the IDEA (and PL94-142), and Section 504 of the Rehabilitation Act.

Our last area of test worthiness was practical- ity, which includes the amount of time it takes to

give the test, the cost of the test, the readability of the instrument, test format, and the ease of administration, scoring, and interpretation. We concluded the chapter by offering five steps for selecting a good test.

CHAPTER REVIEW 1. What are the four cornerstones of test wor-

thiness? Briefly define each of the four types of test worthiness.

2. Explain what is meant by a correlation coef- ficient and describe how it is applied to the understanding of test validity and test reliability.

3. Describe the three main types and the various subtypes of validity.

4. Why is validity known as a “unitary concept”?

5. Why is “face validity” not considered a type of validity?

6. Describe the three main types and various subtypes of reliability.

7. Explain how item response theory (IRT) is different from the more commonly used types of reliability from classical test theory.

8. Why have cross-cultural issues taken on such importance in the realm of tests and assessment?

9. Relative to cross-cultural issues and assess- ment, provide a brief explanation for each of the following:

a. Griggs v. Duke Power Company b. Americans with Disabilities Act c. FERPA (The Buckley Amendment) d. Carl Perkins Act (PL 98-524) e. Civil Rights Acts (1964) and Amendments f. The Freedom of Information Act g. PL 94-142 and the IDEA h. Section 504 of the Rehabilitation Act

10. Describe the main issues involved in assessing the practicality of an assessment instrument.

11. Describe the five steps critical to the selection of a good assessment instrument.

REFERENCES American Counseling Association. (2003). Standards

for qualifications of test users. Alexandria, VA: Author.

American Educational Research Association. (1999). Standards for educational and psychological testing. Washington, DC: AERA.

American Psychiatric Association (2013). Diagnostic and statistical manual of mental disorders (5th ed). Washington, DC, Author.

Beck, A. T., Steer, R. A., & Brown, G. K. (2004). Beck depression inventory, II. San Antonio, TX: Harcourt Assessment.

Bechger, T. M., Maris, G., Verstralen, H. H., & Beguin, A. A. (2003). Using classical test theory in combination with item response theory. Applied Psychological Measurement, 27, 319–334. doi:10.1177/0146621603257518

Berry, C. M., Clark, M. A., & McClure, T. K. (2011). Racial/ethnic differences in the criterion-related validity of cognitive ability tests: A qualitative and quantitative review. Journal of Applied Psychology, 96, 881–906. doi:10.1037/a0023222

Bobko, P., & Roth, P. L. (2012). Reviewing, categoriz- ing, and analyzing the literature on black-white dif- ferences for predictors of job performance: Verifying some perceptions and updating/correcting others. Personnel Psychology. doi:10.111peps. [12007 online version before inclusion into an issue.]

Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3, 296–322.

Cole, D., Truglio, R., & Peeke, L. (1997). Relation between symptoms of anxiety and depression in children: A multitrait-multimethod-multigroup

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 107

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

assessment. Journal of Consulting and Clinical Psychology, 65, 110–119. doi:10.1037/0022-006X. 65.1.110

Cronbach, L. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16, 297–334. doi:10.1007/BF02310555

Glass, T., Mendes de Leon, C., Bassuk, S., & Berkman, L. (2006). Social engagement and depressive symp- toms in late life: Longitudinal findings. Journal of Aging and Health, 18, 604–628. doi:10.1177/ 0898264306291017

Gold, D. L. (1987). Civil-rights panel investigating I.Q. test ban in California. Education Week. Retrieved from http://www.edweek.org/ew/articles/1987/10/ 28/07330001.h07. html?print=l

Goodwin, L. (2002a). Changing conceptions of mea- surement validity: An update on the new standards. Journal of Nursing Education, 41, 100–106.

Goodwin, L. (2002b). The meaning of validity. Journal of Pediatric and Gastroenterology and Nutrition, 35, 6–7. doi:10.1097/00005176-200207000-00003

Griggs v. Duke Power Company, 401 U.S. 424 (1971).

Hartman, N., McDaniel, M., & Whetzel, D. (2004). Racial and ethnic difference in performance. In J. Wall, & G. Watz (Eds.), Measuring up: Assess- ment issues for teachers, counselors, and administra- tors (pp. 99–115). Greensboro, NC: CAPS press.

Heppner, P. P., Wampold, B. E., & Kivlighan, D. M. (2008). Research design in counseling (3rd ed.). Belmont, CA: Thompson Brooks/Cole.

Hewitt, M. A., & Homan, S. P. (2004). Readability level of standardized test items and student perfor- mance: The forgotten validity variable. Reading Research and Instruction, 43, 1–16. doi:10.1080/ 19388070409558403

Kobrin, J. L., & Patterson, B. F. (2011). Contextual factors associated with the validity of SAT scores and high school GPA for predicting first-year college grades. Educational Assessment, 16, 207–226. doi:10.1080/10627197.2011.635956

Hobson v. Hansen, 269 F. Supp. 401 (D.D.C., 1967). Hudson, R. D., & Treagust, D. F. (2013). Which form

of assessment provides the best information about student performance in chemistry examinations? Research in Science & Technology Education, (ahead-of-print), 1–17. DOI:10.1080/02635143. 2013.764516

MacDonald, P., & Paunonen, S. V. (2002). A Monte Carlo comparison of item and person statistics based on item response theory versus classical

test theory. Educational and Psychological Measurement, 62, 921–943. doi:10.1177/ 0013164402 238082

Maruyama, G. (2012). Assessing college readiness: Should we be satisfied with ACT or other threshold scores? Educational Researcher, 41, 252–261. doi:10.3102/0013189X12455095

Mauldin, R. K. (2009). Gendered perceptions of learn- ing and fairness when choice between exam types is offered. Active Learning in Higher Education, 10, 253–264. doi:10.1177/1469787409343191

Moses v. Washington Parish School Board, 302 F.Supp. 362, 367 (ED La.1969).

Murphy, L. L., Geisinger, K. F., Carlson, J. F., & Spies, R. A. (2010). Tests in print VIII: An index to tests, test reviews, and the literature on specific tests. Lincoln, NB: Buros Institute of Mental Measurements.

Ortiz, S. O., Ochoa, S. H., & Dynda, A. M. (2012). Testing with culturally and linguistically diverse populations. In D. P. Flanagan, & P. L. Harrison (Eds.), Contemporary Intellectual Assessment: Theories, Tests, and Issues (3rd ed., pp. 526–552). New York: Guilford Press.

Ostini, R., & Nering, M. L. (2006). Polytomous item response theory models. Thousand Oaks, CA: Sage Publishing.

Park, N. S. (2009). The relationship of social engagement to psychological well-being of older adults in assisted living facilities. Journal of Applied Gerontology, 28, 461–481. DOI:10.1177/0733464808328606

Progar, S., Socan, G., & Pec, M. (2008). An empirical comparison of item response theory and classical test theory. Horizons of Psychology, 17(3), 5–24.

Reid, C. A., Kolakowsky-Hayner, S. A., Lewis, A. N., & Armstrong, A. J. (2007). Modern psychometric methodology: Applications of item response theory. Rehabilitation Counseling Bulletin, 50(3), 177–188.

Rodriguez, M. C. (2005). Three options are optimal for multiple-choice items: A meta-analysis for 80 years of research. Educational Measurement: Issues and Practice, 24(2), 3–13. doi:10.1111/j.1745-3992. 2005.00006.x

Salkind, N. J. (2011). Statistics for people who (think they) hate statistics (4th ed.). Thousand Oaks, CA: Sage Publications.

SASSI [Substance Abuse Subtle Screening Inventory]. (2008-2009).Welcome to SASSI. Retrieved from http://www.sassi.com/

Schraw, G. (2010). Review of the Wechsler adult intel- ligence scale, fourth edition. In R. A. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The Eighteenth

108 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Mental Measurement Yearbook. Retrieved from the Burros Institute’s Mental Measurements Yearbook online database.

Spearman, C. (1904). The proof and measurement of association between two things. American Journal of Psychology, 15, 72–101. doi:10.2307/1412159

Spearman, C. (1910). Correlation calculated from faulty data. British Journal of Psychology, 3, 271–295.

Spies, R. A., Carlson, J. F., & Geisinger, K. F. (2010). (Eds.). The eighteenth mental measurements year- book. Lincoln, NE: Buros Institute of Mental Measurements.

Stowell, J. R., & Bennett, D. (2010). Effects of online testing on student exam performance and test anxi- ety. Journal of Educational Computing Research, 42, 161–171. doi:10.2190/EC.42.2.b

Swenson, L. (1997). Psychology and law for the helping professions (2nd ed.). Pacific Grove, CA: Brooks/Cole.

U.S. Department of Education. (2005). Stronger accountability: The facts about making progress.

Retrieved from http://www.ed.gov/nclb/accountability /ayp/testing.html

U.S. Department of Education. (2008a). No child left behind: State and local implementation of the No Child Left Behind Act. Jessup, MD: Education Publications Center.

U.S. Department of Education. (2008b). Fact sheets. Retrieved from http://www.ed.gov/news/opeds /factsheets/index.html

Wolfe, E. W., & Smith, E. V., Jr. (2007). Instrument development tools and activities for measure valida- tion using Rasch models: Part II—validation activi- ties. Journal of Applied Measurement, 8, 204–234.

Yoon, E., Lee, R., & Goh, M. (2008). Acculturation, social connectedness, and subjective well-being. Cultural Diversity and Ethnic Minority Psychology, 14, 246–255. doi.10.1037/1099-9809.14.3.246

Zunzunegui, M. V., Beland, F., Sanchez, M. T., & Otero, A. (2009). Longevity and relationships with children: The importance of the Parental role. BMC Public Health, 9(351), 1–10.

CHAPTER 5 Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality 109

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R6 Statistical Concepts: MakingMeaning Out of Raw Scores

When my two daughters were nine and four, we had some pretty interesting conversations around the dinner table. One night we were talking about testing, and I was sharing the diffi- culty I sometimes have explaining the value of testing. Suddenly, my older daughter Hannah said, “Testing is good because teachers know what level you are at and then know what to teach you.” I looked at her and said, “That’s right.” Then Emma chimed in, “Test scores are important because when somebody is sad, you should help them.” They made it so simple!

(Ed Neukrug)

This chapter helps us understand how numbers are used to determine, as Hannah noted, what “level” individuals are at. Understanding this concept will enable us to do our job: help others. Whether it be an achievement test score that shows deficien- cies in reading, or a personality test that indicates clinical depression (or as Emma noted, sadness), test scores can help us identify problem areas in a person’s life.

To understand the importance of test scores, we start with the basics: an exam- ination of raw scores. After concluding that raw scores generally provide little meaningful information, we look at various ways that we can manipulate those scores to make sense out of a set of data. Thus, in this chapter we examine how the following are used to help us understand raw scores: frequency distributions; histograms and frequency polygons; cumulative distributions; the normal curve; skewed curves; measures of central tendency such as the mean, median, and mode; and measures of variability, such as the range, interquartile range, and

110

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

standard deviation. In Chapter 7, which is a natural extension of this chapter, we will examine derived scores and see how they too are used to help us understand raw scores. But let’s start with the basics: raw scores.

RAW SCORES Let’s say Jeremiah receives a score of 47 on a test, and Elise receives a score of 95 on a different test. How have they done? If Jeremiah’s score was 47 out of 52, one might assume he had done fairly well. But if 1,000 people had taken this test and all others had received a higher score than Jeremiah, we might view his score somewhat differ- ently. And to make things even more complicated, what if a high score represents an undesirable trait (e.g., cynicism, depression, schizophrenia)? Then clearly, the lower he scores compared to his norm (peer) group, the better he has done, and vice versa.

What about Elise’s score? Is a score of 95 good? What if it is out of a possible score of 200, or 550, or 992? If 1,000 people take the test and Elise’s score is the highest and a higher score is desirable, we might say that she did well, at least com- pared with her norm group. But if her score is on the lower end of the group of scores, then comparatively she did not do well.

Because raw scores provide little information, we need to do something to make them meaningful. One simple procedure is to add up the various types of responses an individual makes (e.g., number of right or wrong items on an achievement test; number of different kinds of personality traits chosen). Although this can provide us with a general idea of the types of responses the person has made, comparing an indi- vidual’s responses to those of his or her norm group can usually give us more in-depth information. Norm group comparisons are helpful for the following reasons:

• They tell us the relative position, within the norm group, of a person’s score. For instance, Jeremiah and Elise, or others interested in Jeremiah’s or Elise’s scores (e.g., teachers, counselors), can compare their scores with those of peo- ple who took the same test and are like them in some important way (e.g., similar age, same grade in school).

• They allow us to compare the results among test-takers who took the same test but are in different norm groups. For instance, a parent who has two children, two grades apart, could determine which one is doing better in reading relative to his or her norm group (e.g., one child might score at the 50th percentile, the other at the 76th, relative to their respective norm groups). Or a school coun- selor might be interested in the self-esteem scores of all the fifth-graders as compared to all the third-graders.

RULE NUMBER 1 Raw Scores Are Meaningless

Raw scores alone tell us little, if anything, about how a person has done on a test. We must take an

individual’s raw score and do something to it to give it meaning.

Raw score Untreated score be- fore manipulation or processing

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 111

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

• They allow us to compare test results on two or more different tests taken by the same individual. For instance, it is sometimes valuable for a teacher to know an individual’s score on an achievement test and on an aptitude test because a dis- crepancy between the two tests might indicate the presence of a learning disability. Or, when conducting a personality assessment, it is not unusual to give a number of different tests with an effort made to find similar themes running through the various tests (e.g., high indications of anxiety on a number of different tests).

To help us make some sense out of raw scores, a number of procedures have been developed to allow normative comparisons. In the rest of this chapter, we look at some of these procedures.

FREQUENCY DISTRIBUTIONS One method of understanding test scores is to develop a frequency distribution. Such distributions order a set of scores from highest to lowest and list the corre- sponding frequency of each score (see Table 6.1). A frequency distribution allows

TABLE 6.1 A Frequency Distribution

Score Frequency (f)

66 1

63 1

60 1

58 2

57 3

55 2

54 5

53 3

51 5

50 5

49 7

48 6

47 4

45 5

43 4

42 1

40 2

38 2

37 1

35 1

32 1

Frequency distribution List of scores and number of times a score occurred

© Ce ng ag e Le ar ni ng

112 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

one to easily identify the most frequent scores and is helpful in identifying where an individual’s score falls relative to the rest of the group. For instance, by examining Table 6.1, we can easily see that most of the scores occurred around the 50 range, with fewer occurring as we move to the higher or lower ends. If an individual scored a 60, he or she can quickly see that such a score is higher than most.

HISTOGRAMS AND FREQUENCY POLYGONS Once a frequency distribution has been developed, it is relatively easy to convert it to either a histogram (similar to a bar graph) or to a frequency polygon (comparable to a line graph). These visual representations of the frequency distribution are often easier for individuals to understand. To develop a histogram or frequency polygon, one has to first determine what class intervals are to be used.

Class intervals are derived from your frequency distribution and are groupings of scores that have a predetermined range. For example, we could take the frequency dis- tribution from Table 6.1 and rearrange or group the scores by fives (the predetermined range), starting with the lowest score. The first class of scores would then be 32 to 36 (i.e., 32, 33, 34, 35, or 36). The next interval would be 37 to 41, and so forth. Table 6.2 illustrates the frequency distribution of the class intervals using a range of 5.

Of course, the class interval does not have to be 5. You could arrange the class interval to have any predetermined range, such as 3, 4, 6, 10, or whatever might be useful (see Box 6.1).

After you have constructed your class interval, the scores can be placed onto a graph in the form of a histogram or frequency polygon. To create a histogram, one places the class intervals along the x-axis and the frequency of scores along the y-axis (see Figure 6.1). In this case, a vertical line is placed at the beginning and end of each class interval and a horizontal line connects these two lines at the height (as measured along the vertical axis). The horizontal line represents the respective frequency of the specific interval being drawn.

To create a frequency polygon of the class interval of scores from Table 6.2, we would simply place a dot at the center of each class interval across from its respective frequency and then connect the dots (see Figure 6.2).

TABLE 6.2 Scores Arranged with a Class Interval of 5

Class Interval Frequency (f)

62–66 2

57–61 6

52–56 10

47–51 27

42–46 10

37–41 5

32–36 2

Histogram Bar graph of class intervals and frequency of a set of scores

Frequency polygon Line graph of class intervals and frequency of a set of scores

Class interval Grouping scores by a predetermined range

© Ce ng ag e Le ar ni ng

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 113

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

32–36 37–41 42–46 47–51

Class Intervals

F re

q u e n c y

52–56 57–61 62–66 0

5

10

15

20

25

30

FIGURE 6.1 | Histogram with Intervals That Have a Range of 5 (Based onTable 6.2)

32–36 37–41 42–46 47–51 52–56 57–61 62–66 0

5

10

15

20

25

30

Class Intervals

F re

q u e n c y

FIGURE 6.2 | Frequency Polygon with Intervals That Have a Range of 5(Based on Table 6.2)

BOX 6.1 Configuring the Number of Intervals You Want for Your Graph

To determine how many numbers should be in your class interval, follow these instructions (note: calcula- tions are based on the data in Table 6.1)*:

1. Subtract the lowest number in the series of scores from the highest number: 66 � 32 ¼ 34

2. Divide this number by the number of class intervals you want to end up with: 34/7 ¼ 4.86

3. Round off the number you obtained in step 2: 4.86 is rounded off to 5

4. Starting with the lowest number in your series of scores, use the number obtained in step 3 (the number 5 in this case) as the number of scores you should place in each interval: 32–36, 37–41, 42–46, 47–51, 52–56, 57–61, 62–66

*Please note that this is a crude method of configuring the number of intervals you want and sometimes can be off by one interval.

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

114 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

CUMULATIVE DISTRIBUTIONS Another method of visually displaying data from a frequency distribution is the cumulative distribution, sometimes called an ogive curve. Although the histogram and frequency polygon are the preferred methods for observing the shape of a distri- bution, the cumulative distribution is better for conveying information about the per- centile rank. To create a cumulative distribution, you convert the frequency within each class interval to a percentage and then add it to the previous cumulative percent- age of the distribution. In Table 6.3, we expand the scores and class intervals from Table 6.2 to include a column for the percentage and cumulative percentage.

We can now graph the cumulative percentage along the y-axis and our class inter- vals along the x-axis. As can be seen in Figure 6.3, we can quickly approximate the

TABLE 6.3 Percentages Calculated for a Cumulative Distribution

Class Interval f % Cumulative Percentage

62–66 2 3 100

57–61 6 10 97

52–56 10 16 87

47–51 27 44 71

42–46 10 16 27

37–41 5 8 11

32–36 2 3 3

Total 62

0

10

20

30

40

50

60

70

80

90

100

32–36 37–41 42–46 47–51 52–56 57–61 62–66

Class Intervals

C u m

u la

ti v e P

e rc

e n ta

g e

FIGURE 6.3 | Cumulative Distribution

Cumulative distribution Line graph to examine percentile rank of a set of scores

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 115

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

percentage of any point of the distribution. If these were classroom test scores for a statewide proficiency achievement exam and the cutoff score for passing was 42, we could see from this distribution that about 19% of our students did not pass.

NORMAL CURVES AND SKEWED CURVES As we collect data from test scores and create histograms or frequency polygons using class intervals, we will see that our distributions are sometimes skewed (asymmetrical) and other times represent what is called a normal or bell-shaped curve. Each of these types of curves has different implications for how we compare scores. Let’s take a look at each of these types of curves and then see how measures of central tendency and measures of variability, two important measurement con- cepts, are applied to them.

The Normal Curve A number of years ago, while I was visiting the Boston Museum of Science, I (Ed Neukrug) saw a device called a “quincunx” (also known as Galton’s board) through which hundreds of balls would be dropped onto a series of protruding points (see Figure 6.4). Each ball had a 50/50 chance of falling left or right every time it would hit one of the protruding objects. After all the balls were dropped, they would be collected and automatically dropped again. All day long this machine would drop the balls over and over again. Now, this process in and of itself was not so amazing. However, what did seem extraordinary was the fact that each time those balls were dropped, they would distribute themselves, more

FIGURE 6.4 | A Quincunx

Quincunx Board developed by Sir Francis Galton to demonstrate bell- shaped curve

© Ce ng ag e Le ar ni ng

116 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

or less, in the shape of a normal curve (also called the bell-shaped curve). Now, mind you, they were not distributing themselves in that manner because they were being sent in that direction (see http://ww2.odu.edu/~eneukrug/galton.htm). In point of fact, the resulting bell-shaped curve is a product of the natural laws of the universe and is explained through the laws of probability. So perfect are these natural laws that some people have given religious connotations to them (Bryson, 2003). However, no matter how you explain these results, it is amazing that such a predictable pattern occurs over and over again. What does this have to do with testing? Like the balls dropping in this device, when we measure most traits and abilities of people, the scores tend to approximate a bell-shaped distribution. This is very convenient, for the symmetry of this curve allows us to understand measures of variability, particularly standard deviation, an important concept we examine later in this chapter.

Skewed Curves Sometimes a distribution of scores does not fall in a symmetrical shape or a normal curve. When this happens, the curve is called skewed or asymmetrical. If the major- ity of the scores fall toward the upper end, the curve is called negatively skewed. A positively skewed curve occurs when the majority of scores fall near the lower end (see Figure 6.5). If you split a normal curve in half, you will find the same number of scores in the first half of the curve as in the second half. This is not true of skewed curves. For instance, in a negatively skewed curve, there are more scores toward the high end of the curve as compared to the low end, and in a positively skewed curve, there are more scores at the low end.

RULE NUMBER 2 “God Does Not Play Dice with the Universe.” (Einstein Paraphrased*)

So perfect are the laws of nature that some see this as the creation of a perfect God who has set in motion the wheels of the universe. Whether or not you believe this is a God-inspired phenomenon, it

has great implications for testing. Concepts such as the bell-shaped curve are crucial to our under- standing of norm-referenced testing and allow us to compare individuals.

*As reported by Bryson (2003).

Negatively skewed Normal curve Positively skewed

FIGURE 6.5 | Skewed and Normal Curves

Normal curve Bell-shaped distribu- tion that human traits tend to fall along

Skewed curve Test scores not falling along a normal curve

Negatively skewed curve Majority of scores at upper end

Positively skewed curve Majority of scores at lower end

© Ce ng ag e Le ar ni ng

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 117

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

MEASURES OF CENTRAL TENDENCY Measures of central tendency tell us what is occurring in the midrange or “center” of a group of scores. Thus, if you know a person’s score, you can compare that score to one of three scores that represent the middle of the group. Three measures of central tendency are the mean, median, and mode. Although measures of central tendency tell you nothing about the range of scores or how much scores vary, they do give you a sense of how close a score is to the middle of the distribution.

Mean The most commonly used measure of central tendency is the mean, which is the arith- metic average of a set of scores. The mean is calculated by adding together all of the scores and dividing by the number of scores. (See the following formula, where M is the mean, �X is the sum of all the scores, and N is the number of scores.)

M ¼ ∑X N

The first column of Table 6.4 shows how one determines the mean. By sum- ming all of the scores and dividing by 11, we find the mean to be 85.55 (941/11).

Median The median is the middle score, or the score at which 50% of scores fall above and 50% fall below. In Table 6.4, we can see that there are 11 total scores, and the middle score is the one that has 5 scores above it and 5 scores below it; hence, we can see that in this case the middle score, or median, is 87. If there were an even number of scores, we would find not one middle score but two, and in this case we would simply take

TABLE 6.4 Determining Measures of Central Tendency

Mean Median Mode

97 97 97

94 94 94

92 92 92

89 89 89

89 89 5 x

89

87 87 Middle # 87

84 84 5 y

84

82 82 82

79 79 79

75 75 75

73 73 73

Sum ¼ 941 N ¼ 11 M ¼ 85.55

Median ¼ 87 Mode ¼ 89

Mean Arithmetic average of a set of scores

Median Score where 50% fall above and 50% below

© Ce ng ag e Le ar ni ng

118 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

the average of the two middle scores. In situations involving a skewed set of scores, the median is generally considered a more accurate measure of central tendency because any unusually low or high scores do not distort the median as it would with the mean (because all scores are averaged to get the mean). For example, one can use the mean or median when trying to understand “average” salaries. If our community happens to have a relatively small number of individuals who make extremely large incomes, these incomes will be included when computing the mean. On the other hand, extremely large incomes would not affect the resulting median. Which measure of central ten- dency do you think would be more important if you were considering moving to a com- munity where there were a few houses that cost very little or a few houses that were very expensive and you were trying to get a sense for what houses generally cost?

Mode The mode, the final measure of central tendency we examine in this section, is the score that occurs most often. In Table 6.4, you can quickly see that the mode is 89 because it is the only score that occurs twice. In smaller groups of scores, the mode may be erratic and not very accurate. However, with larger sets of numbers, the mode captures the peak or top of the curve. In addition, sometimes you can have multiple modes, such as when two numbers occur most often and the same number of times.

In examining the mode, consider our earlier example of incomes in a community. One could conceivably have a mean that is higher than the median (due to a few extremely high incomes), and a mode that is lower than the median (because the most commonly found incomes could fall below the median) (see Figure 6.6, curve C). As you review the skewed and normal curves in Figure 6.6, consider the reason why the mean, median, and mode are placed where they are, and then read Box 6.2.

MEASURES OF VARIABILITY Measures of variability tell us how much scores vary in a distribution. Three mea- sures of variability are the range, or the number of scores between the highest and lowest scores on a distribution; the interquartile range, which measures the range of the middle 50% of a group of scores around the median; and the standard deviation, or the manner in which scores deviate around the mean in a standard fashion.

Negatively skewed curve A

Normal curve curve B

Positively skewed curve C

Mean Median Mode

Median

Mean MedianMode

Mean

Mode

FIGURE 6.6 | Respective Positions of Measures of Central Tendency with Skewed and NormalCurves

In a skewed distribu- tion, the median is a better measure of central tendency

Mode Most frequently occurring score

© Ce ng ag e Le ar ni ng

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 119

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Range The simplest measure of variability, called the range, is calculated by subtracting the lowest score from the highest score and adding 1. Although the range tells you the distance from the highest to lowest score, it does little to identify where most of the scores fall. For example, if the highest score on an exam is 98 and the lowest score is 62, the range is 37 (98 � 62 þ 1 ¼ 37). However, this range gives no indication where the majority of scores may fall. Thus, the range is limited in function. A more complex and informative measure of variability is the inter- quartile range.

Interquartile Range The interquartile range provides the range of the middle 50% of scores around the median. Because it eliminates the top and bottom quartiles, the interquartile range is most useful with skewed curves because it offers a more representative picture of where a large percentage of the scores fall (see Figure 6.7).

To calculate the interquartile range, after developing our distribution of scores from high to low such as in Table 6.5, we subtract the score that is 1/4 of the way

BOX 6.2 Which to Use: Mean or Median?

A school system has asked you to determine the average reading score for its 50,000 fifth-grade stu- dents. The assistant superintendent wants to make sure the school system looks “as good as possible” because the press will get hold of the results. On the other hand, you realize that a sizable population of

students read poorly and you want to make sure that their scores are included in your data so that their learning needs are focused on. Which measure of central tendency do you use? What concerns do you have?

Median

25% 25%25% 25%

First quarter

Second quarter

Third quarter

Fourth quarter

Interquartile range

FIGURE 6.7 | Interquartile Range with Skewed Curve

Range Difference between highest and lowest score plus 1

Interquartile range Middle 50% of scores around the median

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

120 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

from the bottom from the score that is 3/4 of the way from the bottom and divide by 2. Next, we add and subtract this number to the median. For example, Table 6.5 has a set of 12 test scores. Let’s find the interquartile range.

Since we have 12 scores, we first find the score that is 1/4 from the bottom (or 1/4 of 12). Since 1/4 of 12 ¼ 3, we find the third score, or 81 (see Table 6.5) (round off if 1/4 of the N is not a whole number). Next, we find the score that is 3/4 from the bottom (or 3/4 of 12). Since 3/4 of 12 ¼ 9, we find the ninth score, or 92 (see Table 6.5) (round off if 3/4 of the N is not a whole number). Next, we subtract our 1/4 score form our 3/4 score, or 92 � 81 ¼ 11. Finally, we divide 11 by 2 and add and subtract this number to the median (87.5 � 5.5). So, our interquartile range is 82.0 through 93. A formula to find the interquartile range is

Median � 3 4 Nðfind that scoreÞ � 1

4 Nðfind that scoreÞ

2

Using this example with the formula we find

87:5 � 3 4 ð12 scoresÞ � 1

4 ð12 scoresÞ

2

or

87:5 � 92 � 81 2

(Where “92” is the 9th score and “81” is the 3rd score) or 87.5 � 5.5 (82 through 93).

TABLE 6.5 Example of Calculating an Interquartile Range

Test Scores

98

97

95 3 4 ðNÞ ¼ 3

4 ð12Þ ¼ 9th score 92

90

88

87

85

83 1 4 ðNÞ ¼ 1

4 ð12Þ ¼ 3rd score 81

80

79

�������!

�������!

© Ce ng ag e Le ar ni ng

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 121

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Standard Deviation In contrast to the interquartile range, which examines the spread of scores around the median, the standard deviation is a measure of variability that describes how scores vary around the mean. With normally distributed curves, the standard devi- ation is a powerful tool in helping us understand test scores. The standard devia- tion (“SD”) is important because in all normal curves the percentage of scores between standard deviation units is the same. For instance, between the mean and þ1 standard deviation, we find approximately 34% of the scores (see Figure 6.8). Similarly, between the mean and �1 standard deviation, we find 34% of the scores, as each side of the curve is the mirror image of the opposite side. In addi- tion, as you can see in Figure 6.8, approximately 13.5% of the group falls between þ1 and þ2 standard deviations and another 13.5% between �1 and �2 standard deviations. Most of the rest of the group is between þ2 and þ3 standard devia- tions (2.25%), or �2 and �3 standard deviations (another 2.25%). As you can also see, approximately 68% of the group is between �1 and þ1 standard devia- tions, and about 95% of scores will fall between �2 SD and þ2 SD (13.5% þ 34% þ 34% þ 13.5% ¼ 95%), and so forth. Although standard deviations con- tinue (e.g., 4 SDs, 5 SDs, and so forth), since 99.5% of people will fall within the first three standard deviations, we tend to focus only on these.

To understand standard deviation, let’s suppose for a moment that a test has a mean of 52 and a standard deviation of 20. For this test, most of the scores (about 68%) would range between 32 and 72 (plus and minus one standard deviation; see Figure 6.8). On the other hand, a second test that had a mean of 52 and a standard deviation of 5 would have 68% of its scores between 47 and 57 (again, plus and minus one standard deviation). Clearly, even though the two tests have the same mean, the range of scores around the mean varies considerably. An individual’s

68%

95%

99.5%

23 SD 22 SD

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

21 SD 0 Mean

11 SD 12 SD 13 SD

FIGURE 6.8 | Standard Deviation and the Normal Curve

Standard deviation How scores vary from the mean

© Ce ng ag e Le ar ni ng

122 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

score of 40 would be in the average range on the first test but well below average on the second test. This example shows why measures of central tendency and measures of variability are both so important in understanding test scores. To actu- ally determine the standard deviation of a group of scores, one uses the following formula (an alternative formula can be found in Appendix D):

SD ¼ ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ∑ðX � MÞ2

N

s

In this formula, X is each individual test score, M is the mean, and N is the number of scores. For example, let’s say we want to find how far one standard deviation is from the mean for a set of test scores in Table 6.6. First, we generate a column for the X � M for each score. Next, we generate another column to calculate the (X � M)2 for each score. Once this step is complete, we can sum all of the (X � M)2, divide by the number of scores (N), and take the square root.

Using the example in Table 6.6, we calculate the mean (M) of the scores by sum- ming them (42) and dividing by N, which is 6, so our M ¼ 7. In the next column, we subtract the mean from each score to get X � M. Now in the third column, we square our results from column 2 to get (X � M)2. Next, we sum our (X � M)2 column and get 20. Now, we place the numbers in our equation as follows:

SD ¼ ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ∑ðX � MÞ2

N

s ¼

ffiffiffiffiffiffi 20 6

r ¼

ffiffiffiffiffiffiffiffiffiffi 3:33

p ¼ 1:83

This result tells us that each standard deviation is 1.83 from the mean. Applying this to a normal curve would give us Figure 6.9.

Consequently, we can see that someone who scored a 5 falls slightly below �1 standard deviation (or at approximately the 15th percentile). Similarly, a person with a score of 11 is slightly above þ2 standard deviations (or at approximately the 98th percentile). Standard deviation is a particularly important statistical concept when examining test scores relative to normal curves. In Chapter 7 we will examine how standard deviation is used to help us interpret a wide variety of test score data.

TABLE 6.6 Calculating the Standard Deviation

X X � M (X � M)2

10 10 � 7 ¼ 3 32 ¼ 9 8 8 � 7 ¼ 1 12 ¼ 1 7 7 � 7 ¼ 0 02 ¼ 0 7 7 � 7 ¼ 0 02 ¼ 0 6 6 � 7 ¼ �1 �l2 ¼ 1 4 4 � 7 ¼ �3 �32 ¼ 9

� ¼ 42 M ¼ 42/6 ¼ 7 � ¼ 20

© Ce ng ag e Le ar ni ng

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 123

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

REMEMBERING THE PERSON It is important to remember that individuals will have their own opinions about their performance. Thus, what is perceived as a high score by one person could be seen as a low score by another. The individual who consistently scores at the high- est percentiles on a math test may feel upset about a score that is in the average range, while an individual who consistently scores low may feel good about that same score. Similarly, an individual who has struggled with lifelong depression

BOX 6.3 Why Does My Calculator Give Me a Slightly Different Standard Deviation?

Students often check their standard deviation math using a calculator and find they get a slightly different result. Why is this? The equation we provide in the book is the “population” standard deviation that divides by N. Most calculators use the “sample” standard deviation formula that divides by N � 1

SDsample ¼ ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ðX � MÞ2 N � 1

s0 @

1 A:The latter formula is often

used with smaller samples to adjust for missing extreme scores that will result in underestimating the standard deviation. With larger samples, the differ- ence between the two formulas becomes negligible. Although most other textbooks teach the population formula, some professors may prefer for you to use the sample standard deviation equation.

23 SD 22 SD 21 SD Mean 11 SD 12 SD 13 SD 1.51 12.493.34 10.665.17 8.837.00

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

SD 5 1.83

FIGURE 6.9 | A Normal Distribution (M ¼ 7, SD ¼ 1.83)

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

124 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

may feel good about a moderate depression score on the Beck Depression Inventory, while another person might be concerned about such a score. As a helper, it is criti- cal to always ask for feedback from the person who was assessed so that you can determine how he or she perceives the score that was received. Similarly, we should be careful to avoid letting our biases about what is a “good” or “bad” score inter- fere with hearing how an examinee feels about his or her score.

SUMMARY We began this chapter by examining the useful- ness of raw scores. First, we noted that raw scores are not particularly meaningful unless we manip- ulate them in some fashion, such as by comparing an individual’s scores to his or her norm group. Such comparisons (1) tell the relative position, within the norm group, of a person’s score, (2) allow for a comparison of the results among test-takers who took the same test but are in dif- ferent norm groups, and (3) allow for a compari- son of test results on two or more different tests taken by the same individual.

Next, we reviewed a series of mechanisms that one could use to make normative compari- sons. We first looked at how to create a frequency distribution, because such a distribution can help us understand where most individuals fall. We then showed that data can be placed in class inter- vals and graphed as a frequency polygon or a his- togram, which are visual representations of the norm group. We found that the cumulative distri- bution is helpful for gathering information about percentile rank.

Next, we examined the normal curve and dis- tinguished it from skewed curves. We highlighted the amazing fact that due to the natural laws of the universe, many qualities, when measured, approximate the normal curve. We pointed out that the symmetry of the normal curve allows us to apply certain statistical concepts to it, such as measures of central tendency and measures of var- iability. We noted that in contrast, negatively and positively skewed curves are not symmetrical and have more scores at the higher (negatively skewed) or lower (positively skewed) ends of the distribution.

We defined three measures of central tendency: mean, or arithmetic average; median, or middle score; and mode, or most frequent score. We noted that in normal curves the mean, median, and mode are at the midpoint, cutting the curve into two halves. We contrasted this with a skewed curve where the mode is the highest point, the mean is drawn out toward the endpoint, and the median is between the mean and mode. Next, we discussed three types of variability: the range, or the differ- ence between the highest and lowest score þ1; the interquartile range, or the middle 50% of scores around the median; and the standard deviation, or how scores deviate around the mean.

Relative to the normal curve, we explained how standard deviation can be used to under- stand where test scores fall. We noted that the percentage of scores between standard deviation units on a normally distributed curve is constant, with about 34% of scores falling between 0 and þ1 standard deviation and another 34% between 0 and �1 standard deviation (68% total), that approximately 13.5% of the scores fall between þ1 and þ2 and another 13.5% between �1 and �2 standard deviations (27% total), and that approximately 2.5% of the scores fall between þ2 and þ3 and another 2.5% between �2 and �3 (5% total). Thus, if we know an individual’s raw score as well as the mean and standard devi- ation of a test, we can approximate where on the curve the person’s score falls.

As the chapter concluded, we highlighted the importance of always asking the individual who was assessed his or her perception of the score that was received. A high score for one person is a low score for someone else.

CHAPTER 6 Statistical Concepts: Making Meaning Out of Raw Scores 125

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

CHAPTER REVIEW* 1. Describe how norm group comparisons can

make raw scores meaningful. 2. Using the scores below, create a frequency

distribution.

1 2 4 6 12 16 14 17 7 21 4 3 11 4 10

12 7 9 3 2 1 3 6 1 3 6 5 10 3

3. From the numbers in item 2, develop a histogram that uses class intervals which has 3 numbers in each interval, a frequency polygon that uses class intervals with 4 numbers in each interval, and a cumulative distribution that has class intervals that has 4 numbers in each interval.

4. Describe the relationship between the mean, median, and mode on a negatively skewed curve, a positively skewed curve, and a normal curve.

5. Using the numbers in item 2, determine the mean, median, and mode of this group of scores (measures of central tendency).

6. Using the numbers in item 2, determine the range, interquartile range, and standard deviation of this group of scores (measures of variability).

7. If the scores in item 2 were the scores of a class of graduate students in counseling who took a national test of depression, and if the mean and standard deviation of the test nationally was 14 and 5, respectively, for a group of moderately depressed clients, what statement might you be able to make about this group of students?

8. Using the national mean and standard devia- tion for moderately depressed clients (from item 7), at what approximate percentile has a person scored if he or she obtained a raw score of 9? What if he or she has a raw score of 19?

9. How might you feel if you scored a 19 on this test?

10. Describe the relationship between how a person does on a test, compared to a national mean and standard deviation, and how a person might feel about his or her score. For example, if I had a history of major depres- sion and scored a 15 on the test just dis- cussed, how might my interpretation of my score differ from an individual who has never dealt with depression and receives a score of 15?

ANSWERS TO ITEMS 5 THROUGH 8 5. Mean: 7; Mode: 3; Median: 6 6. Range: 21; Interquartile range: 6 � 3.5, or

2.5 through 9.5; SD: 5.21 7. The mean was significantly lower than those

individuals who were moderately depressed but the SD was about the same. Thus, many

students scored lower on depression, but there was a bit of overlap between the two groups.

8. (9 – 14)/5 ¼ –1 (p ¼ 16); (19 – 14)/5 ¼ 1 (p ¼ 84)

REFERENCE Bryson, B. (2003). A short history of nearly everything. New York: Broadway Books.

*Your instructor can log-in to the Instructor Resource Center (login.cengage.com) to find answers to the questions found in this text.

126 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 7Statistical Concepts:Creating New Scores to Interpret Test Data

I remember taking an organic chemistry test in college and thinking afterwards that there was no doubt I had failed. This would have been the only test I had ever failed, and I was devastated. As I waited for my score, I pondered my future. Suddenly, a friend of mine burst into my dorm room and announced, “You passed. You got a D.” I knew that even a score of D would have been impossible, so I asked, “What did I get on the test?” His response was, “You got a 17.” By far, this was the lowest grade I had ever received. I responded, “Well, how could I have passed if I received a 17?” My friend explained, “He took the square root of each person’s score and multiplied it by 10. Then he applied his usual ‘curve,’ and your score of 41 was a D.” In essence, what had happened was that the professor had “converted” my score. I then understood that things are not always what they appear to be. As it was with my converted organic chemistry test score, many raw scores in the world of testing are converted so that they have new meaning. That is the topic of this chapter— varying ways that test publishers and others convert test scores for easier interpretation.

(Ed Neukrug)

In this chapter, we examine how raw scores are converted to what are called “derived” scores to make them easier to understand. We start by distinguishing between norm-referenced testing and criterion-referenced testing, which are two different ways of understanding test scores. As we will see, derived scores are mostly associated with norm-referenced testing. Next, we discuss specific types of derived scores, such as (1) percentiles; (2) standard scores, including z-scores, T-scores, deviation IQs (DIQ), stanines, sten scores, NCE scores, college and

127

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

graduate school entrance exam scores (e.g., SATs, GREs, and ACTs), and publisher-type scores; and (3) developmental norms, such as age comparisons and grade equivalents. As we near the end of the chapter, we discuss standard error of measurement (SEM), which is a mechanism for estimating how close an indivi- dual’s obtained score is to his or her true score. We then discuss a related concept, standard error of estimate, which helps us predict a range of scores based on a criterion score. Finally, the chapter concludes with a short discussion of scales of measurement, which allow us to assign numerical values to nonnumeric character- istics for purposes of developing tests and when using tests in research.

NORM REFERENCING VERSUS CRITERION REFERENCING Two testing experts pass each other in the hallway, and one turns to the other and says, “How are you doing?” The second expert replies, “Compared to whom?”

The terms norm referencing and criterion referencing are used to describe two dif- ferent ways of understanding an individual’s score. In norm referencing, each indi- vidual’s test score is compared to the average score of a group of individuals, called the norm group or peer group. Thus, one could compare his or her score to the aggregate scores of students who took an organic chemistry test (as in the example at the beginning of the chapter), to the aggregate scores of thousands of individuals who took a national achievement test, or to a representative group of individuals who took a personality test. A large percentage of nationally made standardized tests are norm-referenced.

Criterion referencing, on the other hand, compares test scores to a predeter- mined value or a set criterion. For example, an instructor may decide that a test score of 90% to 100% correct is an A, 80% correct to 89.9% correct is a B, 70% correct to 79.9% correct is a C, and so forth.

Criterion-referenced testing is generally used by state departments of motor vehicles (DMVs), which often require a minimum score of 70% correct to pass their written exam. If the DMV used norm-referencing testing and decided that people would pass the test if they received a score above −1 standard deviation, approximately 16% of the individuals taking the test would consistently fail, even if they answered more than 70% of the questions correctly. In fact, if suddenly everyone who took this test was studying more diligently and the overall scores were higher, using norm-referenced testing would likely mean that 16% of the group would still fail. So those who fail might have higher scores than those who had previously passed. This could create an administrative mess for many DMVs and also result in a lot of angry people.

Many states have begun using criterion-referenced testing of students as a result of the No Child Left Behind (NCLB) legislation. In this case, the federal government mandated that all students must achieve minimum preset scores on statewide exams (e.g., 75% correct in reading) (see Box 7.1).

Table 7.1 shows some common tests that are norm-referenced or criterion- referenced. Obviously, the choice of whether to use norm referencing or criterion referencing when considering test results is quite important.

Norm referencing Comparison of indi- vidual test scores to average score of a group

Criterion referencing Comparison of test scores to a predeter- mined standard

128 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

NORMATIVE COMPARISONS AND DERIVED SCORES As discussed in Chapter 6, the relative position of an individual in his or her norm group is a reflection of how a person has performed. To review this point, consider the following question: If John scores a 52 on a test and Marietta scores a 98, who has done better? Reflecting back to Chapter 6, our answer must be based on a number of factors, including the following:

1. The number of items on the test and the highest possible score. 2. The relative positions of a score of 52 and a score of 98 compared with the

rest of the group. If out of 1,000 examinees, the vast majority scored above 98, then the real difference between a 98 and 52 may be minimal.

BOX 7.1 High-Stakes Testing: Using Criterion-Referenced Testing to Assure that All Students Will Achieve

As you read the following paragraph, substitute the word criterion for the italicized word or phrase and you will see why No Child Left Behind (NCLB) is based on criterion-referenced testing.

As a result of the NCLB federal initiative, testing has become a critical method of assuring that all students will achieve in schools. In general, NCLB encourages each state to set a passing score, known as the “starting point,” which is based on the performance of its lowest-achieving demo- graphic group or of the lowest-achieving schools in the state, whichever is higher. The state then sets the level of student achievement that a school must attain after two years to continue to show “adequate

yearly progress.” Subsequent thresholds must be raised at least once every 3 years until at the end of 12 years all students in the state are achieving at the proficiency level on state assess- ments in reading/language arts and math (U.S. Department of Education, 2005, 2011). Although this is a noble effort, a major criticism of NCLB is the fact that the federal government has threat- ened to remove aid to schools if these thresholds are not met, even though little funding has been provided to assist school systems to meet higher standards of achievement (National Education Association, 2002–2013). The future of NCLB is precarious.

TABLE 7.1 Examples of Norm-Referenced and Criterion-Referenced Tests

Norm-Referenced Criterion-Referenced

GREs, SATs, ACTs, MCATS, etc. BDI-2 (Beck Depression Inventory)

IQ tests (Wechsler, Stanford-Binet, etc.) College writing entrance or exit exam

Personality inventories (MBTI, CPI, MMPI-2, etc.) Driver’s licensing exam

Career inventories (Strong Interest Inventory, Self-Directed Search, etc.)

MAST (Michigan Alcohol Screening Test)

College exam scored on a “curve” College exams scored against a standard (A ¼ 90% right, B ¼ 80% right, C ¼ 70% right, etc.)

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 7 Statistical Concepts: Creating New Scores 129

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

3. Whether higher scores are better than lower scores. For instance, on a depres- sion test, lower scores might be better.

4. How an individual feels about his or her score. A low score for some indivi- duals may be a high score for others (e.g., some might be ecstatic with a score at the 80th percentile on a math test, while others might be disappointed).

As you can see from items 1 and 2, understanding an individual’s relative posi- tion in a group is critical if we are to make judgments about that individual’s score. Of course, relative position does not tell us whether a lower or higher score is bet- ter and does not indicate how a person feels about his or her score (items 3 and 4).

To determine an individual’s relative position in his or her norm group, raw scores are often converted to frequency distributions, histograms, and polygons to provide a visual representation of what is happening with all the scores. We dis- cussed this process in Chapter 6 and also introduced the concept of measures of central tendency and measures of variability to begin to examine how scores are related to the normal curve. In this chapter, we introduce the concept of derived scores, which are conversions of raw scores that allow us to make further compar- isons of an individual’s score with those of his or her norm group. Derived scores include percentiles; standard scores such as z-scores, T-scores, DIQs, stanines, sten scores, normal curve equivalents (NCEs), college and graduate school entrance exam scores (e.g., SATs, GREs, and ACTs), and publisher-type scores; and devel- opmental norms, such as age comparisons and grade equivalents.

Percentiles Perhaps the simplest and most common method of comparing raw scores to a norm group is to use percentile rank, often just called percentile. A percentile repre- sents the percentage of people falling below an obtained score and ranges from 1 to 99, with 50 being the mean. For example, if the median score on a sociology exam was a 45, and an individual scored a 45, his or her percentile score would be 50 (p ¼ 50), meaning that 50% of the individuals who took the test scored below a 45. If the top score for that exam was 75, then that person would be at the 99th percentile. Be careful not to confuse percentile scores with the term percentage cor- rect, which refers to the number of correct items. In the example just given, the per- son with a percentile score of 99 could have answered 75 out of 125 questions correctly, which would have given that person a “percentage correct” score of 60 (75/125). Percentile ranking is considered norm referencing because it compares an individual’s score with the scores of a larger group (the normative group).

By examining Figure 7.1, you can see that on the normal curve percentiles break down in the following approximate manner: �3 SD is a percentile that is less than 1, �2 SD is a percentile that is close to 2, �1 SD is a percentile of about 16, 0 SD is a percentile of 50, þ 1 SD is a percentile of about 84, þ 2 SD is a per- centile close to 98, and þ 3 SD is a percentile that is over 99.

Standard Scores Standard scores represent a number of different kinds of scores that are derived by converting an individual’s raw score to a new score that has a new mean and new

Derived score Converted score based against a norm group

Percentiles Percentage of people falling at or below a score

Standard scores Derived score based on mean and standard deviation

130 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

standard deviation. Standard scores are generally used to make interpretation of test material easier for the examinee. Some of the more common types of standard scores include z-scores, T-scores, DIQ, stanines, sten scores, NCE scores, college and graduate school entrance exam scores (e.g., SATs, GREs, MATs, ACTs), and publisher-type scores.

z-Scores The most fundamental standard score is called a z-score, which is a simple conversion of an individual’s raw score to a new score that has a mean of 0 and a standard deviation of 1. Thus, if an individual scored at the mean, the z-score would be 0; if an individual scored 1 standard deviation above the mean, the z-score would be 1; if an individual scored 1 standard deviation below the mean, the z-score would be �1; and so forth. Figure 7.2 demonstrates where z-scores lie on the normal curve.

Converting a raw score to a z-score is almost always the first step to take to understand the meaning of the raw score an individual obtained. Once the raw score has been converted to a z-score, almost any other type of derived score can be found, including percentiles, T-scores, DIQ, stanines, and so forth.

23 SD 22 SD

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

21 SD 0 Mean

11 SD

1 405 6020 30 95807010 999050

12 SD 13 SD

Percentiles

FIGURE 7.1 | The Normal Curve with Percentile Scores

RULE NUMBER 3 z-Scores are Golden*

z-Scores are great for helping us see where an indi- vidual’s raw score falls on a normal curve and are helpful for converting a raw score to other kinds of

derived scores. That is why we like to keep in mind that z-scores are golden and can often be used to help us understand the meaning of scores.

z-score Standard score with mean of 0 and SD of 1

*Rules 1 and 2 were introduced in Chapter 6: Rule 1: Raw scores are meaningless. Rule 2: God does not play dice with the universe.

© Ce ng ag e Le ar ni ng

CHAPTER 7 Statistical Concepts: Creating New Scores 131

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

That is why z-scores are so critical to our understanding of an individual’s raw score (see Rule Number 3).

The formula for converting a raw score to a z-score is

z ¼ X � M SD

where X is the raw score, M is the mean score, and SD is the standard deviation. For example, let’s say an individual takes a psychology exam and the mean turns out to be 45 while the standard deviation is 10. To convert an individual’s raw score of 65 on this test to a z-score, we would use the conversion formula in the following manner:

X ¼ 65 ðraw scoreÞ M ¼ 45 ðmeanÞ SD ¼ 10 ðstandard deviationÞ

We plug these values into our formula and get

z ¼ X � M 10

¼ 65 � 45 10

¼ 20 10

¼ 2:0

From this example, we can see that the individual who scored a 65 on the exam has a z-score of þ 2.0. By examining Figure 7.2, you can see that this person has a percentile score of approximately 98.

Using the same formula, see if you can determine what the z-score would be for an individual who had a raw score of 30 on this same test. If you followed the formula correctly, you should have had a z-score of �1.5. Can you determine the approximate percentile that this person obtained? (All one need to do is trace a line from the z-score to the percentile rank line on Figure 7.2.) In this case, you

23.0 22.0 21.0 0 11.0 12.0 13.0

23 SD 22 SD

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

21 SD 0 Mean

11 SD

1 405 6020 30 95807010 999050

12 SD 13 SD

Percentiles

z - score

FIGURE 7.2 | z-Scores on the Normal Curve

© Ce ng ag e Le ar ni ng

132 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

can see that the approximate percentile is a 7. If you would like an exact method for determining a percentile from a z-score, refer to Appendix F, which offers such a conversion formula. Additionally, Appendix F also includes a look-up table for quicker and easier conversions of z-scores to percentiles.

The main value of z-scores is to assist the test administrator in understanding where on the curve an individual falls compared with his or her peers, and as noted earlier, is the first step toward configuring other kinds of standard scores that are more readily understandable by clients. Explaining to a client that he or she had a z-score of �1.5 might not only be useless, but even counterproductive. However, we can convert the z-score to other kinds of derived scores (e.g., percen- tiles, stanines, DIQs, T-scores, and others) that are often more palatable for clients.

T-scores One type of standard score that can be easily converted from a z-score is the T-score. T-scores have a mean of 50 and a standard deviation of 10 and are generally used with personality tests. Figure 7.3 shows T-scores on the normal curve.

To convert a z-score to a T-score, we use the following formula:

Conversion score ¼ zðSDof new desired scoreÞ þ Mof new desired score where the conversion score is the new standard score to which you are converting (e.g., T-score). In converting, one uses the standard deviation and mean of the score to which one is converting (in this case, T-scores; the SD ¼ 10 and the M ¼ 50). Continuing with our earlier example, the individual who has received a z-score of �1.5 has a T-score of 35 (�1.5 � 10 þ 50 ¼ 35). What about the person who had a raw score of 65 that converted to a z-score of 2.0? What would his or her T-score be? If you came up with a T-score of 70, you would be correct (2 � 10 þ 50 ¼ 70).

Deviation IQ Another commonly used standard score is DIQ. The DIQ has a mean of 100 and a standard deviation of 15. To apply this to the normal curve, see Figure 7.4.

23.0 22.0 21.0 0 11.0 12.0 13.0

20 30 40 50 60 70 80

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

T- score

FIGURE 7.3 | T-Scores on the Normal Curve

T-score Standard score with mean of 50 and SD of 10

Deviation IQ Standard score with mean of 100 and SD of 15

© Ce ng ag e Le ar ni ng

CHAPTER 7 Statistical Concepts: Creating New Scores 133

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Although most intelligence tests employ this standard of scoring, some caution should be exercised. For example, older versions of the Stanford-Binet Intelligence Scales (fourth edition) used a mean of 100 and a standard deviation of 16; however, the most recent version (fifth edition; Houghton Mifflin Harcourt, n.d) uses a mean of 100 and a standard deviation of 15. To convert a z-score to a DIQ, we utilize the same conversion formula noted earlier [Conversion score ¼ zðSDof new desired scoreÞ þ Mof new desired score]. However, in this case we use the stan- dard deviation of 15 and the mean of 100. Using the example earlier in which an individual obtains a z-score of �1.5, we would find a DIQ of 77.5 (DIQ ¼ �1.5 � 15 þ 100). What is an individual’s DIQ if his or her z-score is 2.0? If you came up with a DIQ of 130, you would be correct (2 � 15 þ 100 ¼ 130). And, finally, what would a person’s DIQ be if he or she had a z-score of 6 (see Box 7.2)?

23.0 22.0 21.0 0 11.0 12.0 13.0

55 70 85 100 115 130 145

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

DIQ

FIGURE 7.4 | Comparing the Deviation IQ to the Normal Curve

BOX 7.2 IQ, Population, and the Normal Curve

Have you ever heard anyone say he or she knows someone with an IQ of 200? I have. My grandmother told me that one of my cousins had been tested and found to be brilliant, with an IQ over 200. Using the look-up table for percentiles (see Appendix F), you can see how rare people are at the outer edges of the normal or bell-shaped curve. Looking at the table, you will note that the fourth standard deviation (z ¼ 4.0), which corresponds to an IQ of about 160, is extremely rare; only 3 out of 100,000 people are at this level. An IQ of 175 is found in only 3 out of 10 million people. That would mean there are only 94 people living in the United States with this

IQ assuming a U.S. population of 314 million (U.S. Census Bureau, 2013). An IQ of 190 (z ¼ þ 6.0) is found in only one out of a billion people, which means there are approximately seven people in the world with this IQ. As a matter of fact, I was speaking with Senior Project Director of the Stanford-Binet Intelligence Scale (personal commu- nication, Dr. Andrew Carson, October 7, 2004), and he said they have failed to ever find a person with an IQ above 160 because it is so rare. So the next time you hear of someone with a 200 IQ, you might want to question them further . . . apparently they don’t know a lot about the normal curve! —Charlie Fawcett

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

134 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Stanines Another standard scoring system frequently used in the schools is stanines, which stands for “standard nines.” Often used with achievement tests, stanines have a mean of 5 and a standard deviation of 2, and range from 1 to 9. For example, an individual who scores one standard deviation above the mean would have a stanine of 7 (M ¼ 5 þ 1 SD of 2). Similarly, if a student has a z-score of �1, his or her stanine would be 5 minus 2, which is 3. Figure 7.5 shows where stanines fall in comparison to z-scores and percentiles.

Unlike other forms of standard scoring, we have examined thus far where a score could be identified with a specific z-score and percentile (e.g., T-score of 60 equals a z-score of 1 and percentile of 84), stanines represent a range of z-scores and percentiles. For instance, a stanine of 5 runs from a z-score of −0.25 to a z-score of þ 0.25 and from the 40th to 60th percentiles (see Figure 7.5). Similarly, a stanine of 6 runs from a z-score of þ 0.25 to a z-score of þ 0.75 and from the 60th to 77th percentiles, and a stanine of 7 runs from a z-score of þ 0.75 to z-score of þ 1.25 and from the 77th to 89th percentiles, and so forth. As you can see from Figure 7.5, any z-score that is below �1.75 or above þ 1.75 is a stanine of 1 or 9, respectively.

Changing z-scores to stanines is still done with our conversion formula z (SD) þ M. Applying the formula for a z-score of −1.5, we would get

z ¼ �1:5ðgivenÞ SD ¼ 2 M ¼ 5; so

Stanine ¼ zðSDÞ þ M ¼ �1:5ð2Þ þ 5 ¼ �3 þ 5 ¼ 2 Consequently, a z-score of �1.5 converts to a stanine of 2. Similarly, a z-score of 2 converts to a stanine of 9 (stanine ¼ 2 � 2 þ 5). Because stanines are reported only as whole numbers, in those cases where the use of the z-score in the conversion for- mula results in a fraction, the fraction should be rounded off to the nearest whole number. For example, a z-score of 0.87 converts to a stanine of 6.74, which is rounded off to 7 (0.87 � 2 þ 5 ¼ 6.74, or in this case, 7).

As with most forms of standard scoring, stanines are an attempt to explain scoring so that the test-taker, or his or her parents, can easily understand test results. Unless you have happened to take a course like this one, it is usually easier

23.0 22.0 21.0 0 11.0 12.0 13.0

4 96897760402311

1 87 96

Stanines

5432

Percentiles

z - score

FIGURE 7.5 | Stanines Compared to z-Scores and Percentiles

Stanine Standard score with mean of 5 and SD of 2

© Ce ng ag e Le ar ni ng

CHAPTER 7 Statistical Concepts: Creating New Scores 135

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

to tell a parent that his or her daughter scored a 7 out of a range from 1 to 9 than to explain that the daughter has a z-score of 0.87.

Sten Scores Sten scores are derived from the name “standard ten” and are com- monly used on personality inventories and questionnaires. Stens have a mean of 5.5 and a standard deviation of 2. Stens divide the scale into 10 units, each of which is one-half of a z-score, except for the 1st sten, which represents all scores below �2 z-scores, and the 10th sten, which represents all scores above þ 2 z-scores. Stens are similar to stanines in that they represent a range of scores rather than an absolute point. Figure 7.6 shows the relationship between stens, z-scores, and percentiles.

As you might expect, the conversion formula of z(SD) þ M also applies to sten scores. To use our example of a z-score of �1.5, we would determine our sten score to be 2.5:

z ¼ �1:5ðgivenÞ SD ¼ 2 M ¼ 5:5; so

Sten ¼ zðSDÞ þ M ¼ �1:5ð2Þ þ 5:5 ¼ �3 þ 5:5 ¼ 2:5 However, because stens, like stanines, are reported in whole numbers, we would round up a score of 2.5 to a sten of 3. What would be the sten score of an individ- ual who obtained a z-score of 2? If you got 10, you would be correct (2 � 2 þ 5.5 ¼ 9.5, rounded off to 10).

Normal Curve Equivalents (NCE) Scores A form of standard scoring frequently used in the educational community is NCE. The NCE has a mean of 50 and a standard deviation of 21.06. These units range from 1 to 99 in equal units along the bell-shaped curve. As can be seen in Figure 7.7, percentile ranks range from 1 to 99 but are not arranged in equal units along the curve. Consequently, the NCE and per- centile ranks are only the same at 1, 50, and 99 (see Figure 7.7 to compare NCEs with the normal curve). The standard formula for converting scores still applies to NCEs except that the limits are 1 and 99. The formula is NCE ¼ z(21.06)þ 50.

College and Graduate School Entrance Exam Scores* Probably, almost all of you have taken the SATs or ACT prior to entering college, and maybe you wondered how the score was derived. In this case, the SAT has three sections

23.0 22.0 21.0 0 11.0 12.0 13.0

2 93 9884695031167

1 87 96

Sten

5432

Percentiles

z - scores 10

FIGURE 7.6 | Sten Scores Compared to z-Scores and Percentiles

*Please note that the SATs and ACTs are approximate as their means and standard deviations tend to float depending on the year taken.

Sten score Standard score with mean of 5.5 and SD of 2

Normal curve equivalent Standard score with mean of 50 and SD of 21.06

SAT type score Standard score with mean around 500 and SD of 100

© Ce ng ag e Le ar ni ng

136 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

(critical reading, mathematics, and writing), which use a standard score that has a standard deviation of 100 and a mean of 500, although the actual mean and stan- dard deviation from year to year will vary somewhat. This is because they compare each year’s group to a 1990 group of students. If the current year students’ have done better than the 1990 group of students, the mean will be higher than 500. If they have done worse, then it will be lower than 500. Similarly, the standard devia- tion will fluctuate some for the group. However, rather than giving you a percentile as compared with that norm group, they determine an individual’s percentile based on students who have taken the test over the past three years. Thus, an individual’s

23.0 22.0 21.0 0 11.0 12.0 13.0

20 30 40 50 60 70 80

200 300 400 500 600 700 800

23 SD 22 SD

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

21 SD 0 Mean

11 SD

1 405 6020 30 95807010 999050

12 SD 13 SD

55 70 85 100 115 130 145

Percentiles

T - score

Deviation IQ

Stanine

z - score

1 87 965432

1 87 965432 10 Sten score

SAT

6 11

4010 50 80201 6030 90 9970

16 21 26 31 36 ACT

NCE

FIGURE 7.7 | The Normal Curve and Types of Standard Scores*

© Ce ng ag e Le ar ni ng

*Please note that the SATs and ACTs are approximate as their means and standard deviations tend to float depending on the year taken.

CHAPTER 7 Statistical Concepts: Creating New Scores 137

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

standard score is compared to the 1990 norm group, but an individual’s percentile is based on a more current norm group. Using the more current norm group for an individual’s percentile allows colleges and universities to see where an individual falls compared to other current students who are applying to college. Using our con- version formula, see if you can convert z-scores of �1.5 and 2 to an SAT-type score. If you got 350 and 700, you would be correct [SAT ¼ �1.5 � 100 þ 500 ¼ 350, and 2 � 100 þ 500 ¼ 700]

The ACT offers scores in four main subcategories, which are English, mathe- matics, reading, and science plus an optional writing exam (ACT, Inc., 2007). Like the SATs, the ACTs has a composite score that is created by converting an indivi- dual’s raw score to a standard score, which in this case uses a mean of 21 and a standard deviation of 5 for college-bound students. Like the SATs, an individual’s raw score is compared to an earlier norm group, in this case 1995, and the percen- tile is determined from students who have taken the test over the past three years. Using the conversion formula, see if you can convert z-scores of �1.5 and 2 to an ACT-type score. If you got 14 and 31, you would be correct [ACT ¼ �1.5 � 5 þ 21 ¼ 13.5 or 14 rounded off, and 2 � 5 þ 21 ¼ 31].

Although not depicted on Figure 7.7, the GRE General Test currently has a mean of about 151 and a standard deviation of about 5.6, although this will vary somewhat year to year. The GRE subject tests scores run from 200 through 900, and their means and standard deviations vary considerably depending on the test. The Miller Analogy Test has a mean of 400 although the standard deviation will vary and percentiles are given along with the individual’s score.

Publisher-Type Scores As you probably have come to realize, the conversion formula can be used to create a standard score with any prechosen mean and stan- dard deviation. Consequently, sometimes test developers generate their own stan- dard score. That is why it is common to see standardized achievement tests using unique test publisher scores that employ means and standard deviations of the pub- lisher’s choice. For instance, if Marietta took a standardized reading achievement test and received a raw score of 48, the scores on her actual profile sheet might look something like what you see in Box 7.3.

In the example in Box 7.3, Marietta’s raw score is equivalent to a percentile score of 77, which is equivalent to a standard score of 588, which is equivalent to a stanine of 7. How did the publisher get these scores? First, they took her raw

BOX 7.3 Marietta’s Reading Score

Marietta Smartgirl Grade 5.2 Age: 10.5

Raw Score* Percentile Stanine Standard Score**

48 77 7 588

© Ce ng ag e Le ar ni ng

ACT score Standard score with mean around 21 and SD of 5

Publisher-type score Test developer creates own standard score

*Mean of raw score ¼ 45; standard deviation of raw score ¼ 4 **Mean of publisher’s standard score ¼ 550; standard deviation ¼ 50

138 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

score and compared it to her norm group’s mean and standard deviation to obtain a z-score (remember, z-scores are golden!). So, hypothetically, we may have seen something like the following: z ¼ (48 � 45)/4, which equals a z-score of þ 0.75. By looking at our graph (see Figure 7.2), we can see that a z-score of 0.75 is approxi- mately a percentile of 77. Also, using our conversion formula, we can determine that a z-score of 0.75 is equal to a stanine of 7 (0.75 � 2 þ 5 ¼ 6.5; rounded to 7). Finally, we can use our conversion formula to determine the publisher’s standard score if we know the mean and standard deviation of the publisher’s standard score (generally found in the publisher’s manual and in this case given in Box 7.3). For instance, with a z-score of 0.75 we could determine that the publisher simply plugged the mean and standard deviation into the conversion formula: 0.75 � 50 þ 550 ¼ 587.5 (rounded to 588).

Unfortunately, test printouts usually only list the publisher’s standard score without telling you the mean and standard deviation. This makes it quite difficult to determine what the score actually is based on. We sometimes call these publisher standard scores “magical scores” because they seem to magically appear with little explanation on the part of the publishing company. Thus, we usually suggest using other derived scores, such as percentiles or stanines, when interpreting test data for clients.

Developmental Norms As opposed to standard scores, which in some manner convert an individual’s raw score to a score that has a new mean and standard deviation, developmental norms, such as age comparisons and grade equivalents, directly compare an indivi- dual’s score to the average scores of others at the same age or grade level.

Age Comparisons Remember when you were a kid and you were always being compared to others of your age on those weight and height charts? The doctor was looking at where your height and weight fell compared to your norm group. Basically, the doctor would compare your height and weight to a bunch of kids who were the same age as you. So, if you were a 9-year, 4-month-old girl, and you were 55 inches tall, and the mean height for girls of your age was 52.5 inches with a standard deviation of 2.4 inches, your z-score would be þ 1.04 [(55 � 52.5)/2.4)]. This z-score converts to a percentile of approximately 85. One other way of using age norms is to see how your performance compares to the average performance of individuals at other age levels. Thus, for the 9-year, 4-month-old girl, we could say that she is at the average height for a 10-year, 11-month-old girl. If we were looking at an intelligence test, we might find that a 12-year, 5-month old has the mental age of the average 15-year, 4-month old (see Box 7.4).

Grade Equivalents Similar to age norms, grade equivalents compare an indivi- dual’s score to the average score of children at the same grade level. Thus, if a stu- dent who was in the second month of the third grade (grade 3.2) took a reading test and scored at the mean, that student’s grade equivalent would be 3.2. Unfortu- nately, this is where the comparison to age norms stops. Rather than actually

Age comparison score Comparison of indi- vidual score to aver- age score of others at the same age

Grade equivalent Comparison of indi- vidual score to aver- age score of others at the same grade level

CHAPTER 7 Statistical Concepts: Creating New Scores 139

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

comparing the raw score of a student to the raw scores of students at differing grade levels (e.g., comparing the raw score of the 3.2 grade student to the raw scores of students in the second grade, fourth grade, fifth grade, and so on), pub- lishing companies usually extrapolate an individual’s score. Thus, the student in the 3.2 grade who obtained a score at the 75th percentile when compared to her grade might end up with a grade equivalent of 4.5 despite the fact that she was never actually compared to students in the fourth grade, fifth month.

So, what can be said about a student’s grade equivalent when it is either higher or lower than the mean score of students at his or her grade level? The best inter- pretation would be to state that a student has scored either lower or higher than students in his or her grade. For instance, a student in the 3.2 grade who obtains a grade equivalent of 5.4 has not mastered many of the concepts of the average stu- dent in grade 5.4; however, such a student is clearly doing much better than a large percentage of students in the third grade. Thus, it would make sense to say that this student is performing above the average for students at the 3.2 grade, but it may not be accurate to state that this student is performing at the 5.4 grade level. Similarly, another student in the 3.2 grade who obtains a grade equivalent of 2.2 has mastered many concepts that the average student at the 2.2 grade level has not even examined. Thus, it would make sense to say that this student is perform- ing below the average for students at the 3.2 grade level, but probably not accurate to state that this student is performing at the 2.2 grade level.

You can see why it is important for test interpreters to read and understand how the test developer normed and calculated the grade equivalent score. It should be easy to see why having a strong background in testing and assessment is impor- tant in communicating test results.

PUTTING IT ALL TOGETHER Now that we have learned that norm-referenced scores are based on the normal curve, we can use Figure 7.7 to examine the relationships between most of the var- ious norm-referenced scores that we introduced in the chapter. You are encouraged to become familiar with these scores, as you will come across them often through- out your professional life.

BOX 7.4 Is Your Head Too Big?

When my older daughter, Hannah, was two months old, the pediatrician measured her head size. He immediately said, “It’s at the 98th percentile. Let’s measure it again in a month.” I asked him if he was concerned, knowing full well that her head size was “outside the norm” or not within the average range for her age group, which could suggest a number of

medical problems. Thankfully there were no pro- blems, and it turned out that some members of the family just have “big heads.” However, such age norm comparisons can be critical to the identification of possible early developmental problems.

—Ed Neukrug

© Ce ng ag e Le ar ni ng

140 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

STANDARD ERROR OF MEASUREMENT As we saw in Chapter 5, all tests have a certain amount of measurement error, which results in an individual’s score being an approximation of his or her true score. By using some simple statistics, we can calculate the range or band in which a person’s true score might actually lie. Known as the standard error of measurement (SEM), this range of scores is where we would expect a person’s score to fall if he or she took the instrument over and over again (and was the same person each time he or she took it—e.g., no learning had taken place). As you will recall from Chapter 5, reliability is the degree to which a test is free of measurement error. Consequently, we can use the reliability to determine the SEM if we know the standard deviation of the raw score or of the standard score. The formula for calculating SEM is:

SEM ¼ SD ffiffiffiffiffiffiffiffiffiffiffi 1 � r

p

where SD is the standard deviation of the raw score or of the standard score and r is the reliability coefficient of the test. As an example, let’s say Latisha’s “true” DIQ score is 120 (i.e., if she were to take this test over and over again, the mean of all her scores would be 120). From reading the published data about the instru- ment, we know that it has a reliability coefficient of 0.95; we also know that DIQs have a standard deviation of 15. Applying this information to our formula, we get the following result:

SEM ¼ SD ffiffiffiffiffiffiffiffiffiffiffi 1 � r

p

¼ 15 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � :95

p

¼ 15 ffiffiffiffiffiffiffi :05

p

¼ 15 � :22 ¼ 3:35ðplus & minusÞ

Now we know that if Latisha were to take the intelligence test over and over again, her score would fall plus or minus 3.35 of her score of 120, or between 116.65 and 123.35. As you may recall, the area under the normal curve of plus and minus one standard deviation equals 68% of the scores. Hypothetically, then, if Latisha were to take the test 1,000 times, 68% of the time she would score between 116.65 and 123.35 (see Figure 7.8).

If we wanted greater accuracy, we could use plus or minus two 2 SEMs (two standard deviations around the mean), and that would tell us where her score would fall approximately 95% of the time. If we went out to 3 SEMs, we would know where her score would fall 99.5% of the time:

SEM ¼ 2ðSDÞ ffiffiffiffiffiffiffiffiffiffiffi 1 � r

p

SEM ¼ 3ðSDÞ ffiffiffiffiffiffiffiffiffiffiffi 1 � r

p

So Latisha’s 2 SEM score would be 3.35 times 2, or �6.7; that is, 95% of the time her score would fall between 113.30 and 126.70, and 99.5% of the time her score would fall in the range of 109.95 to 130.05 (3 � 3.35 ¼ � 10.05).

Standard error of measurement Range where a “true” score might lie

CHAPTER 7 Statistical Concepts: Creating New Scores 141

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

At this point, you are probably beginning to see that there is an inverse rela- tionship between the reliability coefficient and SEM; that is, as the reliability decreases, the SEM (the range of true scores) increases. Following the example of Latisha, let’s say the intelligence test was not very reliable, with r ¼ 0.70. How would this affect the SEM? Her SEM would increase. Calculating the SEM for the new reliability coefficient of 0.70, we would find:

SEM ¼ SD ffiffiffiffiffiffiffiffiffiffiffi 1 � r

p

¼ 15 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � :70

p

¼ 15 ffiffiffiffiffiffiffi :30

p

¼ 15 � :55 ¼ �8:22

So, if Latisha were to take this new intelligence test with a reliability coefficient of 0.70, 68% of the time her true score would fall in the range of 111.78 to 128.22. That range is quite a bit larger than the previous example when we used r ¼ 0.95 and her SEM was determined to be between 116.35 and 123.35.

If you have observed test reports from the larger test publishers, you may have seen individual test scores with an “X” and with a line on either side of the “X” (see Figure 7.9). This line represents the SEM, or the range in which the true score might fall. Figure 7.9 shows how these bands may look on test reports showing Latisha’s score on tests with reliability coefficients of 0.95 versus 0.70.

You can see that SEM is particularly important for the interpretation of test scores, because the larger the SEM, the more error and the larger the range where an individual’s true score might fall (see Box 7.5)

1 SEM (68.0%)

2 SEM (95.0%)

3 SEM (99.5%)

23 SD 22 SD

2.25%

13.5% 13.5%

34.0% 34.0%

2.25%

21 SD 0 Mean

11 SD 12 SD 13 SD

FIGURE 7.8 | Standard Error of Measurement and the Normal Curve

Reliability coefficient and SEM Have an inverse relationship

© Ce ng ag e Le ar ni ng

142 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

STANDARD ERROR OF ESTIMATE A similar concept to SEM is the standard error of estimate (SEest), which was discussed briefly in Chapter 5. However, rather than giving a confidence interval around an obtained score as is done with SEM, SEest gives us a confidence inter- val around a predicted score. Based on the scores received on one variable (e.g., SATs), SEest allows us to predict a range of scores an individual might obtain on a second variable (e.g., first-year college GPA). The formula for calculating the standard error of estimate (SEest) for two variables in which there is a correlation is

SEest ¼ SDY ffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � r2

p

BOX 7.5 Determining Standard Error of Measurement

Let’s say Ed obtained a score on a test of depres- sion that was a T-score of 60, and the test-retest reliability of the instrument was 0.75. Assuming that the higher the score, the more depressed a person is, answer the following questions:

1. What is Ed’s percentile score on this test? 2. Determine what the SEM is 68% and 95% of

the time. 3. What would be the range of Ed’s T-score and

percentile scores 68% of the time? 4. What would be the range of Ed’s T-score and

percentile scores 95% of the time? 5. What implications does the standard error of

measurement have for interpreting Ed’s score?

Answers:

1. About a percentile of 84. 2. SEM 68% of the time is � 5: (10

ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � :75

p ); SEM

95% of the time is � 10: (10 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � :75

p ) � 2.

3. Sixty-eight percent of the time Ed’s score would fall between T-scores of 55 and 65 (percentiles of 69 to 93) (60 � 5).

4. Ninety-five percent of the time he would fall between T-scores of 50 to 70 (percentiles of 50 to 98) (60 � 10).

5. Error has great implications for the interpretation of Ed’s scores. The greater the error, the less we can rely on his score to be an indication of high levels of depression. At first, it might seem that he has a fairly high level of depression (T-score of 60, percentile of 84), but as we consider the error and the range of his “true score,” we lower our confidence that the score truly indicates somewhat high levels of depression.

If r � 0.95

If r � 0.70

100 110 120 130 145

X

X

FIGURE 7.9 | Example of SEM on Two Test Reports with Different Reliabilities

Standard error of estimate Range where a pre- dicted score might lie

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 7 Statistical Concepts: Creating New Scores 143

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

where SDY is the standard deviation of the variable to which you are predicting, (in this case, GPA) and r is the correlation between the two variables.

As an example, we know the correlation between the SATs and GPA of first- year college students is about a 0.5. If we knew that the mean and standard deviation of a fictional department was 3.1 and 0.2 respectively, we can deter- mine the SEest:

SEest ¼ SDY ffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � r2

p

¼ :2 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � :502

p ¼ :2

ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � :25

p

¼ :2 ffiffiffiffiffiffiffi :75

p

¼ :2 � :87 ¼ :17

Next, if we know the z-score of a student’s SAT score, we can convert the student’s SAT score to a GPA score using our conversion formula [z(SDof new desired score) þ Mof new desired score]. For instance, if the mean and standard deviation of the SAT is 500 and 100, and if an individual has obtained a 600, his or her z-score is 1 [z ¼ (X � M)/SD ¼ (600 � 500)/100 ¼ 1). And, if we know the mean and standard deviation of the GPA, we can use our conversion formula to convert the SAT score to this person’s predicted GPA:

GPAconversion ¼ zðSDÞ þ M GPA ¼ 1ð0:2Þ þ 3:1 ¼ 3:3

Now, if we add and subtract the obtained SEest (0.17), we can see where this per- son’s predicted GPA is likely to fall 68% of the time: 3.3 � .17 ¼ 3.13 through 3.47. And, similar to the standard error of measurement, if we add and subtract it one more time, we can determine where this individual’s GPA is likely to fall 95% of the time [3.3 � .17(2) ¼ 2.96 through 3.64].

Now, let’s do the same thing with a person who scored a 300 on the SATs. His or her z-score is

�2 z ¼ X � M SDx

¼ 300 � 500 100

¼ �200 100

¼ �2 � �

:

Now, converting this z-score to a GPA score we get 2.7 [z(SD) þ M ¼ �2 � 0.2 þ 3.1 ¼ 2.7]. Now, if we add and subtract 0.17 to this score, we can see that 68% of the time this individual is likely to have a GPA between 2.53 and 2.87 and 95% of the time a GPA between 2.36 and 3.04. You can see why a program is more likely to take a risk with students who have higher SAT scores when you look at the differences in predicted ranges of likely GPA (Box 7.6 breaks this down into simple steps for you). (Read Rule Number 4.)

144 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SCALES OF MEASUREMENT Now that we have finished discussing statistical concepts related to interpreting test scores and we are about to discuss commonly used assessment techniques (Section III), it is important to consider that not all scores are the same. For instance, we can assign people to one of four categories of depression, such as none, low, medium, or high, and give each category a number, such as 0, 1, 2, or 3, respectively. Or we can take the same group of people and have them take a test for depression and examine the group scores, which would fall on a continuum from low to high, based on the number of items on the test. In the first example, we can only say that one group is higher or lower than another group. In the second example, however, we can make some judgment about the amount of depression an individual has, rela- tive to another individual, and we can add and subtract the various scores.

To distinguish between different kinds of test scores and subsequently know what kinds of statistics can be applied to them, four kinds of scales of measurement have been identified: nominal scales, ordinal scales, interval scales, and ratio scales. Know- ing which scale of measurement is being used for a specific instrument has profound

BOX 7.6 Steps for Calculating Standard Error of Estimate

Step 1: Obtain the correlation between your obtained scores and the scores you want to predict to (e.g., SAT and GPA).

Step 2: Obtain the standard deviation (SD) of the scores you want to predict to (e.g., GPA)

Step 3: Obtain your standard error of estimate by plugging the numbers you obtained in Steps 1 and 2 into the following formula:

SEest ¼ SDY ffiffiffiffiffiffiffiffiffiffiffiffiffi 1 � r2

p

Step 4: Calculate the z-score for your obtained score (e.g., SAT)

Step 5: Convert your obtained score to your predictor score (e.g., score on GPA) using your conversion formula:

zðSD of new desired scoreÞ þ Mof new desired score Step 6: Add and subtract your SEest in Step 3 to your obtained score in Step 5 to calculate your range of scores for 68% of the time. Add and subtract the SEest a second time to obtain your range of scores for 95% of the time.

RULE NUMBER 4 Don’t Mix Apples and Oranges

As you practice various formulas in class, it is easy to use the wrong score, mean, or standard deviation. For instance, in determining the SEest you can see how easy it would be to use the wrong standard deviation (SAT) instead of GPA—your predictor

variable. This, obviously, would give you an errone- ous answer. Thus, whenever you are asked to figure out a problem, remember to use the correct set of numbers (don’t mix apples and oranges); otherwise your answer will be incorrect.

© Ce ng ag e Le ar ni ng

CHAPTER 7 Statistical Concepts: Creating New Scores 145

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

implications for how the resulting scores can be manipulated statistically, and these implications are particularly important when conducting research. Although the kinds of tests we will examine in Section III result in scores that are mostly of the interval type, some of the assessment procedures we will examine result in scores that fall into the nominal, ordinal, and ratio scales range. As you read through the text of Section III and review the various assessment instruments, you might want to consider which scale of measurement is being used for the instrument being discussed.

Nominal Scale The most basic or simple measurement scale is the nominal scale. In this scale, numbers are arbitrarily assigned to represent different categories or variables. For example, race might be recorded as 1 ¼ Asian, 2 ¼ Latino, 3 ¼ African American, 4 ¼ Caucasian, and so on. The assignment of numbers to these categories does not represent magnitude, so normal statistical calculations cannot be performed. All you can do is count the occurrences or calculate the mode.

Ordinal Scale In the ordinal scale, magnitude or rank order is implied; however, the distance between measurements is unknown. An example of this would be asking someone how much they agree with the statement “The counseling I received was helpful in obtaining the goals I came in for” and then asking them to choose from: 1 ¼ strongly disagree, 2 ¼ somewhat disagree, 3 ¼ neither agree or disagree, 4 ¼ somewhat agree, 5 ¼ strongly agree. These numbers represent rank or magni- tude, but it is impossible to know the true distance between “somewhat agree” and “strongly agree.”

Interval Scale The interval scale establishes equal distances between measurements but has no absolute zero reference point. The SAT test you may have taken for admission into college is an interval scale. A score of 530 is 20 equal units above a score of 510; however, there is no true zero point since the minimum possible score is 200. Some basic statistical analysis is appropriate, such as determining how many stan- dard deviations an individual is from the mean, but one cannot say that a student who scores a 700 is twice as likely to succeed in college as a student who scores a 350.

Ratio Scale The ratio scale has a meaningful zero point and equal intervals; therefore, it can be manipulated by all mathematical principles. Very few behavioral measures fall into this category. Units such as height, weight, and temperature on the Celsius scale are all ratio scales since they have a true zero reference point. An example of a ratio scale would be the measurement of reaction times. If a researcher is attempting to measure brake response time by individuals with varying blood alcohol content

Nominal scale Numbers arbitrarily assigned to represent categories

Ordinal scale Numbers with rank order but unequal distances between

Interval scale Numbers with equal distances between but no zero

Ratio scale Numbers with equal intervals and meaningful zero

146 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

(BAC) levels, both the BAC and the response time would be ratio scales. A BAC level of 0.00 (sober) might provide an average reaction time of 0.6 second, while a BAC level of 0.10 might correspond to an average reaction time of 1.2 seconds. Both the BAC and reaction time have true zero points, and therefore they are con- sidered ratio scales. In other words, a reaction time of 0.6 second is twice as fast as 1.2 seconds.

SUMMARY In this chapter, we examined how raw scores are converted to what are called “derived” scores so that individuals can more easily understand the meaning of the raw scores. We started the chapter by pointing out that in norm-referenced testing, test scores are compared to a group of individuals called the norm group or peer group, but in criterion-referenced testing, test scores are com- pared to a predetermined value or a set criterion. We noted that although some tests are criterion- referenced, a large percentage are norm- referenced and employ a range of different kinds of derived scores in an attempt to make sense out of raw scores.

Relative to norm-referenced testing, we noted some factors that can affect our understanding of a raw score, including the number of items on a test and knowledge of what is the highest score possible, the relative position of an individual’s score as compared to the scores of others who took the test, whether higher scores are better than lower scores, and how a person feels about his or her score on a test. Next, we identified a number of derived scores that are used to help us understand an individual’s relative position in the norm group, such as (1) percentiles; (2) standard scores, including z-scores, T-scores, DIQs, sta- nines, sten scores, NCE scores, college and gradu- ate school entrance exam scores (e.g., SATs, GREs, MATs, and ACTs), and publisher-type scores; and (3) developmental norms, such as age comparisons and grade equivalents.

In examining derived scores, we first defined percentiles and distinguished percentiles from the concept of “percentage correct.” We noted that

percentiles represent the percentage of people fall- ing below an individual’s obtained score and range from 1 to 99.

Next, we examined standard scores, which are obtained by converting a raw score mean and standard deviation to a scale that has a new mean and standard deviation. First, we highlighted the z-score, which is a standard score that has a mean of 0 and a standard devia- tion of 1, and offered a formula for obtaining it. We noted that an individual’s z-score reflects where he or she falls on the normal curve. We pointed out that z-scores are so important to our understanding of derived scores that we might say “z-scores are golden,” meaning that configuring a z-score is often the first critical step to finding all other derived scores.

We showed how to convert a z-score to a number of standard scores, including T-scores (SD ¼ 10, M ¼ 50), DIQs (SD ¼ 15, M ¼ 100), stanines (SD ¼ 2, M ¼ 5), sten scores (SD ¼ 2, M ¼ 5.5), SATs (SD ¼ 100, M ¼ 500), ACTs (SD ¼ 5, M ¼ 21), NCEs (SD ¼ 21.06, M ¼ 50), and publisher-type scores, where the standard deviation and mean vary depending on the pub- lisher and the test.

As the chapter continued, we noted that developmental norms are also used to help us understand the relative position of an individual’s raw score as compared with his or her norm group. We noted that age norms compare an indi- vidual’s performance to the average performance of others in his or her age group (e.g., height, weight, and mental ability), or to the average per- formance of individuals in other age groups (e.g.,

CHAPTER 7 Statistical Concepts: Creating New Scores 147

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

a 12-year-old might have the mental ability of the average 15-year-old). Grade equivalents, we noted, compare an individual’s performance to the average performance of other students in his or her grade. We cautioned that sometimes grade equivalents are misinterpreted because it is falsely assumed that a higher grade equivalent means an individual can perform at that higher grade level or that a lower grade equivalent means that the individual has not learned concepts at his or her own grade level.

Near the end of the chapter, we discussed standard error of measurement (SEM) which is a mechanism for estimating how close an indivi- dual’s obtained score is to his or her true score. We noted that by using the formula to find SEM, we can determine the range of scores in which an individual’s true score is likely to fall 68% of the time (�1 SEM), 95% of the time (�2 SEM), or 99.5% of the time (�3 SEM). We noted that SEM

has great implications for the interpretation of test data, because the higher the error, the less confi- dence we have about where the individual’s true score actually lies. We then discussed standard error of estimate SEest, which like SEM can give a range of scores within which an individual might fall. However, in this case, it looks at the range of predicted scores as opposed to actual obtained scores. In discussing SEest, we highlighted our fourth rule: “Don’t Mix Apples and Oranges”; that is, when computing any problem, make sure that you are using the correct sets of scores (e.g., correct mean and standard deviation).

Finally, we had a brief discussion of the four scales of measurement: nominal, ordinal, interval, and ratio. We noted that each type of scale has unique attributes that may limit the statistical and/or mathematical calculations we can perform and that different kinds of assessment instruments use different kinds of scales.

CHAPTER REVIEW* 1. When might a criterion-referenced test be

more appropriate to use than a norm- referenced test?

2. Discuss the strengths and weaknesses of using a criterion-referenced test instead of a norm- referenced test to measure progress as defined by No Child Left Behind or other high-stakes testing standards.

3. Distinguish between the following kinds of derived scores: percentiles, standard scores, and developmental norms.

4. An individual receives a raw score of 62 on a national standardized test. Given that the mean and standard deviation of the test were 58 and 8, respectively, find the individual’s z-score.

5. Using the z-score found in Item 4, find the following derived scores: a. Percentile (approximate)

b. T-score c. Deviation IQ d. Stanine e. Sten score f. Normal curve equivalent (NCE) g. SAT-type score h. ACT score i. A publisher-type score that has a mean of 75 and standard deviation of 15

6. Define the term developmental norms. 7. Explain what a grade equivalent is. What is a

potential major downfall of using a grade equivalent type of score?

8. Find the z-score and approximate percentile of a 5.5-year-old child who is 46 inches tall when the mean and standard deviation of height for a 5.5-year-old child are 44 inches and 3 inches, respectively.

*Your instructor can log-in to the Instructor Resource Center (login.cengage.com) to find answers to the questions found in this text.

148 SECTION II Test Worthiness and Test Statistics

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

9. Referring to Item 5, b through i, find what the standard error of measurement would be 68% of the time for each of the scores if the reliability of the test is 0.84. Also determine what the individual’s score would be for each of the items (b through i) 95% of the time.

10. If a student received a score of 24 on the ACT, what is the range of his or her likely

GPA at the end of the first year in college given the following: (GPA Mean ¼ 3.1, SD ¼ 0.3; correlation between ACT and GPA is 0.45)?

11. Identify the four different types of scales of measurement and give examples of situations when some might be more appropriate to use than others.

ANSWERS TO ITEMS 4 THROUGH 10 4. (62–58)/8 ¼ .5 5. p: 69; T-score: 55; DIQ: 107.5; Stanine: 6;

Sten: 6.5 becomes 7 rounded off; NCE: 60.53; SAT: 550; ACT: 23.5; Publisher: 82.5

8. z ¼ .67; p ¼ 74 9. T-score: 55 � 4; DIQ: 107.5 � 6; Stanine:

6 � .8 then round off (5–7); Sten: 6.5 � .8 then round off (6–7); NCE: 60.53 � 8.42;

SAT: 550 � 40; ACT: 23.5 � 2; Publisher: 82.5 � 6. Second part of question: Multiply the answer times 2 and add to the derived score.

10. Step 3 of Box 5.6: .3(square root of 1 – .452) ¼ .3(.89) ¼ .27; Step 4 of Box 5.6: (24–21)/5 ¼ .6; Step 5: .6(.3) þ 3.1 ¼ 3.28; Step 6: 3.28 � .27.

REFERENCES ACT, Inc. (2007). The ACT technical manual.

Retrieved from http://www.act.org/aap/pdf/ACT_ Technical_Manual.pdf

Houghton Mifflin Harcourt. (n.d.). The Stanford-Binet intelligence scales (SB5) (5th ed.). Retrieved from http://www.riverpub.com/products/sb5/index.html

National Education Association. (2002–2013). No Child Left Behind Act (NCLB). Retrieved from http://www.nea.org/esea/policy.html

U.S. Census Bureau. (2013). U.S. and world population clocks. Retrieved from http://www.census.gov/main/ www/popclock.html

U.S. Department of Education. (2005). Stronger accountability: The facts about making progress. Retrieved from http://www.ed.gov/nclb/accountabil- ity/ayp/testing.html

U.S. Department of Education. (2011). No Child Left Behind legislation and policies. Retrieved from http://www2.ed.gov/policy/elsec/guid/states/index .html#nclb

CHAPTER 7 Statistical Concepts: Creating New Scores 149

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

S E C T I O N IIICommonly Used AssessmentTechniques

Section III of the book examines some commonly used techniques for assessing educational ability, intellectual and cognitive functioning, career and occupational interests and aptitudes, clinical issues, and conducting informal assessment. Within this section we wanted to identify and describe some of the more popular tests often used. We were partly driven by an examination of a number of studies that looked at some of the tests most frequently used by practicing counselors and psy- chologists and taught by counselor educators (see Tables 1 and 2). Thus, many of the tests listed in Tables 1 and 2 will be discussed in the chapters in this section of the book.

In Chapter 8, we examine those tests typically given to assess what we have learned in school and are used when making a variety of educational decisions (see Table 3). In the chapter, we first identify the purpose of the assessment of edu- cational ability. We then define the different kinds of achievement and aptitude tests generally given to assess our learning and to make decisions about our futures, including survey battery achievement tests, diagnostic tests, readiness tests, and cognitive ability tests. Next, we examine some of the more popular tests in these categories.

Chapter 9 explores intellectual and cognitive functioning and the tests associ- ated with it. The chapter starts with a brief history of intelligence testing and the models of intelligence that have impacted the development of intelligence tests. We then provide some examples of the major verbal and nonverbal tests of intelligence (see Table 3). The chapter then provides a short history of neuropsychological assessment and defines how it is different from merely giving an intellectual assess- ment. We then highlight two broad categories of neuropsychological assessment: the Fixed Battery Approach and Flexible Battery Approach (see Table 3).

Chapter 10 looks at how interest inventories, and special and multiple aptitude tests are important in the career counseling process. We note how interest invento- ries look at an individual’s likes and dislikes, whereas special and multiple aptitude

151

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

T A B L E 1

F re q u e n c y o f T e st s U se

d b y A ll C o u n se

lo rs , S c h o o l C o u n se

lo rs , M e n ta l H e a lt h C o u n se

lo rs , a n d

T a u g h t b y C o u n se

lo r E d u c a to rs

R a n k in g

A ll C o u n se

lo rs

S c h o o l C o u n se

lo rs

M e n ta l H e a lt h

C o u n se

lo rs

C o u n se

lo r E d u c a to rs

1 B ec

k D ep

re ss io n

In ve n to ry

(B D I)

S A T an

d /o r P S A T

B ec

k D ep

re ss io n

In ve n to ry

(B D I)

B ec

k D ep

re ss io n In ve n to ry

(B D I)

2 M ye rs

B rig

g s T yp

e In d ic at o r (M

B T I)

A C T

B ec

k A n xi et y In ve n to ry

(B A I)

M ye rs

B rig

g s T yp

e In d ic at o r

(M B T I)

3 S tr o n g In te re st

In ve n to ry

C o n n er ’s

R at in g S ca

le s

S u b st an

ce A b u se

S u b tle

S cr ee

n in g In ve n to ry

(S A S S I)

S tr o n g In te re st

In ve n to ry

4 S el f- D ire ct ed

S ea

rc h

W ec

h sl er

In te lli g en

ce S ca

le fo r C h ild re n

M in iM

en ta lS

ta te

E xa m

(M M S E )

S el f- D ire ct ed

S ea

rc h

5 A C T

W o o d co

ck -J o h n so

n T es t

o f C o g n iti ve

A b ili tie s

M ye rs

B rig

g s T yp

e In d ic at o r (M

B T I)

M in n es o ta

M u lti p h as ic

P er so

n al ity

In ve n to ry

(M M P I)

6 S A T an

d /o r P S A T

W o o d co

ck -J o h n so

n T es t

o f A ch

ie ve m en

t M in n es o ta

M u lti p h as ic

P er so

n al ity

In ve n to ry

(M M P I)

W ec

h sl er

A d u lt In te lli g en

ce S ca

le (W

A IS )

7 W ec

h sl er

In te lli g en

ce S ca

le fo r C h ild re n (W

IS C )

S tr o n g In te re st

In ve n to ry

H o u se

T re e P er so

n (H T P )

W ec

h sl er

In te lli g en

ce S ca

le fo r C h ild re n (W

IS C )

8 C o n n er ’s

R at in g S ca

le s

Io w a T es ts

o fB

as ic S ki lls /

E d u ca

tio n al

D ev el o p m en

t S ym

p to m

C h ec

kl is t

(S C L -9 0 -R

) M in iM

en ta lS

ta te

E xa m

(M M S E )

9 B ec

k A n xi et y In ve n to ry

(B A I)

M ye rs

B rig

g s T yp

e In d ic at o r (M

B T I)

S tr o n g In te re st

In ve n to ry

B ec

k A n xi et y In ve n to ry

(B A I)

1 0

O *N

E T S ys te m

an d

C ar ee

r E xp

lo ra tio

n T o o ls

A rm

ed S er vi ce

s V o ca

tio n al

A p tit u d e

B at te ry

(A S V A B )

C o n n er ’s

R at in g S ca

le s

S ix te en

p er so

n al ity

F ac

to rs

(1 6 P F )

1 1

W ec

h sl er

A d u lt In te lli -

g en

ce S ca

le (W

A IS )

B ec

k D ep

re ss io n

In ve n to ry

(B D I)

W ec

h sl er

A d u lt In te lli -

g en

ce S ca

le (W

A IS )

S u b st an

ce A b u se

S u b tle

S cr ee

n in g In ve n to ry

(S A S S I)

1 2

S u b st an

ce A b u se

S u b tle

S cr ee

n in g In ve n to ry

(S A S S I)

A C T

T ra u m a S ym

p to m

C h ec

kl is t

T h em

at ic

A p p er ce

p tio

n T es t

(T A T )

So u rc e:

P et er so n , C . H ., L o m as , G . I. , N eu k ru g E . S. , &

B o n n er , M . W . (2 0 1 4 ). A ss es sm

en t u se

b y co u n se lo rs

in th e U n it ed

St at es : Im

p li ca ti o n s fo r p o li cy

an d p ra ct ic e.

Jo u rn al

o f

C o u n se li n g an

d D ev el o p m en t. N eu k ru g,

E ., P et er so n , C ., B o n n er , M ., &

L o m as , G . (2 0 1 3 ). A

n at io n al

su rv ey

o f as se ss m en t in st ru m en ts

ta u gh

t b y co u n se lo r ed u ca to rs . C o u n se lo r

E d u ca ti o n an

d Su

p er vi si o n , 5 2 , 2 0 7 – 2 2 1 .

152 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

T A B L E 2

F re q u e n tl y U se

d T e st s b y D iff e re n t T yp

e s o f P sy

c h o lo g is ts

R a n k in g

C lin

ic a l

C o u n se

lin g

N e u ro

S c h o o l

1 W ec

h sl er

A d u lt In te lli -

g en

ce S ca

le (W

A IS )

M M P I

M M P I

W IS C

2 M in n es o ta

M u lti p h as ic

P er so

n al ity

In ve n to ry

(M M P I)

S tr o n g In te re st

In ve n to ry *

W ec

h sl er

A d u lt In te lli -

g en

ce S ca

le (W

A IS )

P ea

b o d y In d iv id u al

3 W ec

h sl er

In te lli g en

ce S ca

le fo r C h ild re n (W

IS C )

W ec

h sl er

A d u lt In te lli -

g en

ce S ca

le (W

A IS )

W ec

h sl er

M em

o ry

S ca

le H o u se -T re e- P er so

n *

4 R o rs ch

ac h

S en

te n ce

C o m p le tio

n T ra il M ak

in g T es t

W ec

h sl er

A d u lt In te lli g en

ce S ca

le (W

A IS )

5 B en

d er

V is u al -M

o to r

G es ta lt

W ec

h sl er

In te lli g en

ce S ca

le fo r C h ild re n (W

IS C )*

F A S W o rd

F lu en

cy T es t

H u m an

F ig u re

D ra w in g s*

6 T h em

at ic

A p p er ce

p tio

n T es t

T h em

at ic

A p p er ce

p tio

n T es t

F in g er

T ap

p in g T es t

D ev el o p m en

ta lT

es t o f

V is u al -M

o to r In te g ra tio

n

7 W R A T

S ix te en

P F

G ro o ve d P eg

b o ar d T es t

W R A T

8 H o u se -T re e- P er so

n B en

d er

V is u al -M

o to r

G es ta lt

B o st o n N am

in g T es t

K in et ic

F am

ily D ra w in g

9 W ec

h sl er

M em

o ry

S ca

le W id e R an

g e A ch

ie ve m en

t T es t (W

R A T )*

C at eg

o ry

T es t

W o o d co

ck -J o h n so

n

1 0

M ill o n

M ill o n

W id e R an

g e A ch

ie ve m en

t T es t (W

R A T )*

C h ild

B eh

av io r C h ec

kl is t

1 1

B ec

k D ep

re ss io n

In ve n to ry *

H o u se -T re e- P er so

n *

B ec

k D ep

re ss io n

In ve n to ry

B en

d er

V is u al -M

o to r G es ta lt*

1 2

T ra il M ak

in g T es t

R o rs ch

ac h

R ey

C o m p le x F ig u re

P ea

b o d y P ic tu re

V o ca

b u la ry

* T ie d w it h p re vi o u s te st

So u rc e:

H o ga n , T . P . (2 0 0 5 ). W id el y u se d p sy ch o lo gi ca l te st s. In

G . P . K o o ch er , J. C . N o rc ro ss , &

S. S.

H il l (E d s. ), P sy ch o lo gi st s’ d es k re fe re n ce

(2 n d ed ., p p . 1 0 1 – 1 0 4 ). N ew

Y o rk :

O x fo rd

U n iv er si ty

P re ss .

SECTION III Commonly Used Assessment Techniques 153

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

T A B L E 3

K e y A ss

e ss

m e n t In st ru m e n ts

b y C h a p te r a n d C a te g o ry

C H A P T E R

8 : A s s e s s m e n t

o f E d u c a ti o n a l A b ili ty

C H A P T E R

9 :

In te lle

c tu a l a n d

C o g n it iv e F u n c ti o n in g

C H A P T E R

1 0 : C a re e r

a n d O c c u p a ti o n a l

A s s e s s m e n t

C H A P T E R

1 1 : C lin

ic a l

A s s e s s m e n t

C H A P T E R

1 2 : In fo rm

a l

A s s e s s m e n t

S u rv e y A c h ie ve

m e n t T e st s

• N at io n al

A ss es sm

en t o f

E d u ca

tio n al

P ro g re ss

• S ta n fo rd

A ch

ie ve m en

t T es t

• Io w a T es t o f B as ic

S ki lls

• M et ro p o lit an

A ch

ie ve m en

t T es t

D ia g n o st ic

T e st s

• W id e R an

g e A ch

ie ve m en

t T es t4

• W ec

h sl er

In d iv id u al

A ch

ie ve m en

t T es t- III

• P ea

b o d y In d iv id u al

A ch

ie ve m en

t T es t

• W o o d co

ck -J o h n so

n (r ) III

• K ey M at h 3 D ia g n o st ic

A ss es sm

en t

In d iv id u a l In te lli g e n c e

T e st s

• S ta n fo rd -B

in et , 5 th

ed .

• W ec

h sl er

S ca

le s o f

In te lli g en

ce •

K au

fm an

A ss es s-

m en

t B at te ry

fo r

C h ild re n

N o n ve

rb a l In te lli g e n c e

T e st s

• C o m p re h en

si ve

T es t

o f N o n ve rb al

In te lli -

g en

ce (C T O N I)

• U n iv er sa lI n te lli g en

ce T es t (U N IT ).

• W ec

h sl er

N o n ve rb al

T es t o f In te lli g en

ce (W

N V )

In te re st

In ve

n to ri e s

• S tr o n g V o ca

tio n al In te re st

In ve n to ry

• S el f- D ire ct ed

S ea

rc h

• C ar ee

r O cc

u p at io n al

P re fe re n ce

S ys te m

S p e c ia l A p ti tu d e T e st s

• C le ric al

T es t B at te ry

• M in n es o ta

C le ric al

A ss es sm

en t B at te ry

• U .S . P o st al

S er vi ce

’s 4 7 0 B at te ry

E xa m

• F ed

er al

C le ric al

E xa m

• S ki lls

P ro fil er

S er ie s

M e c h a n ic a l A p ti tu d e T e st

• T ec

h n ic al

T es t B at te ry

• W ie se n T es t o f

M ec

h an

ic al

A p tit u d e

O b je c ti ve

P e rs o n a lit y

T e st s

• M in n es o ta

M u lti p h as ic

P er so

n al ity

In ve n to ry

(M M P I- 2 )

• M ill o n C lin ic al

M u lti ax ia l

In ve n to ry

(M C M I- III )

• B ec

k D ep

re ss io n

In ve n to ry

(B D I- II)

• M ye rs -B

rig g s

T yp

e In d ic at o r

(M B T I)

• T h e 1 6 P F

• N E O

P I- R

• N E O -F F I

• T h e P A I

• T h e S A S S I

In fo rm

a l A ss

e ss

m e n t

In st ru m e n ts

• O b se rv at io n

• R at in g sc al es

• C la ss ifi ca

tio n

m et h o d s

• E n vi ro n m en

ta l

as se ss m en

t •

R ec

o rd s an

d p er -

so n al

d o cu

m en

ts •

P er fo rm

an ce

-b as ed

m et h o d s

(C o n tin u ed

)

154

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

155

R e a d in e ss

T e st s

• K in d er g ar te n R ea

d in es s

T es t (1 )

• K in d er g ar te n R ea

d in es s

T es t (2 )

• T h e M et ro p o lit an

R ea

d in es s T es t

• G es el lD

ev el o p m en

ta l

O b se rv at io n

C o g n it iv e A b ili ty

T e st s

• T h e O tis

L en

n o n S ch

o o l

A b ili ty

T es t- 8 (O L S A T 8 )

• T h e C o g n iti ve

A b ili ty

T es t

• C o lle g e an

d G ra d u at e

S ch

o o lA

d m is si o n s E xa m s

(A C T , S A T , G R E , M A T ,

L S A T , M C A T )

N e u ro p sy

c h o lo g ic a l

A ss

e ss

m e n t

• H al st ed

-R ei ta n

B at te ry

• L u ria -N

eb ra sk a

N eu

ro p sy ch

o lo g ic al

B at te ry

• B o st o n P ro ce

ss A p p ro ac

h

• A R C O

M ec

h an

ic al

A p ti-

tu d e & S p at ia lR

el at io n s

T es ts

• B en

n et t T es t o f M ec

h an

i- ca

lC o m p re h en

si o n

• M u si c A p tit u d e P ro fil e

• Io w a T es t o f M u si c

L ite ra cy

• G ro u p T es t o f M u si ca

l A b ili ty

• A d va n ce

d M ea

su re s o f

M u si c A u d ia tio

n

M u lt ip le

A p ti tu d e T e st s

• A rm

ed S er vi ce

s V o ca

tio n al

A p tit u d e

B at te ry

• D iff er en

tia lA

p tit u d e T es t

P ro je c ti ve

P e rs o n a lit y

T e st s

• T h em

at ic A p p er ce

p tio

n T es t (T A T )

• R o rs ch

ac h In kb

lo t T es t

• B en

d er

V is u al -M

o to r

G es ta lt T es t

• H o u se -T re e- P er so

n •

K in et ic -

H o u se -T re e- P er so

n T es t

• S en

te n ce

C o m p le tio

n S er ie s

• E P S S en

te n ce

C o m p le tio

n •

K in et ic

F am

ily D ra w n

• D ra w -A -M

an /

D ra w -A -W

o m an

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

tests examine an individual’s abilities. Combined, these instruments can be very powerful in the career counseling process (see Table 3).

Our focus in Chapter 11 is on a broad arrange of tests that examine personal- ity traits and temperament (see Table 3). We start by defining clinical assessment and identifying its uses. Then we present a number of objective personality tests that range from tests that can be used to help build client insight in counseling to tests that are helpful in diagnosing and identify psychopathology. The second part of the chapter examines projective tests, which are tests in which individuals proj- ect their inner world onto an unstructured stimuli and an interpretation is made from that projection by the test examiner. These tests can be quite powerful, although as you might imagine, their reliability and validity is lower than the objec- tive tests.

The last chapter in this section is Chapter 12, and this chapter is about infor- mal assessment techniques (see Table 2). By their very nature, these techniques have lower reliability and validity than many other techniques we examine in this book because they are “homegrown” and do not go through the rigor of many of the standardized assessment instruments covered in other chapters. However, they can be particularly helpful by focusing in on a particular aspect of client behavior and thus can be useful because they add one more instrument when making a broad, holistic assessment of a person.

In Chapter 12, we introduce various techniques of informal assessment and identify some positive and negative aspects of their use. We then offer an overview of a number of such instruments, but stress that the ones we demonstrate are by no means an exhaustive list (see Table 3). Near the end of the chapter, we discuss how to assess reliability and validity of informal assessment procedures.

Near the end of Chapters 8–12, we discuss the role that helpers play in using the assessment procedures being discussed in the particular chapter. Also, each chapter in this section concludes with a section that places the assessment proce- dures within the context of the whole person and generally stresses the importance of using wisdom and sensitivity when assessing any person.

156 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 8Assessment of EducationalAbility: Survey Battery, Diagnostic, Readiness, and Cognitive Ability Tests

I woke up early, ready to take my test. I had studied hard, very hard, as hard as I ever had. I had taken an intensive course—10 hours a day for six days to make sure I was covering what I was supposed to cover. And now it was time to take “the test.” I was confident— well, not really. Let’s say I was as ready as I thought I could be. The big day was here. I walked slowly to the testing center in the heart of Boston. Having stayed at a friend’s house the night before, I noticed my hip hurting as I marched my way to the test center. I walk in, “ID please,” says a burly looking character in a rough voice. I reach for my ID, thinking she does not have to be mean. After all, I am “choosing” to take this test. It starts late, and we’re all getting more anxious. Finally, it comes. I take the test, and a few hours later I’m done. I soon notice that my hip pain is gone—caused from my stress. A few weeks later I get my results—above the mean—slightly. I’m in. I’m now am a licensed psychologist. I’m excited, but was it worth it? Probably, but does it really prove that I’m “better” than the 50 percent of doctoral level examinees who did not pass?

(Ed Neukrug)

Mass-produced tests of educational ability—we have all taken them, we have all sweated them, and they have affected all of us in some way. That’s what this chap- ter is about—the types of widely used tests that are given to assess what we have learned and to make decisions about our futures. First, we identify and define the different kinds of achievement and aptitude tests generally given, including survey battery achievement tests, diagnostic tests, readiness tests, and cognitive ability tests. Next, we look at some of the more popular kinds of tests in these categories.

157

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

We then provide an overview of some of the roles played by helpers in the assess- ment of educational ability and conclude with some thoughts about assessment in this important domain.

DEFINING ASSESSMENT OF EDUCATIONAL ABILITY Whether it is the “Iowa,” the “Stanford,” the “SAT,” the “ACT,” or some other widely used test of educational ability, most students have taken many of these kinds of tests that are used to determine how we have been doing in school and to make decisions about our future. Tests of educational ability have broad applica- tions. For instance, it is not unusual to find these tests being used for the following purposes:

1. to determine how well a student is learning; 2. to assess how well a class, grade, school, school system, or state is learning

content knowledge; 3. as one method of detecting learning problems; 4. as one method of identifying giftedness; 5. to help determine if a child is ready to move to the next grade level; 6. as one measure to assess teacher effectiveness; 7. to help determine readiness or placement in college, graduate school, or pro-

fessional schools; and 8. to determine if an individual has mastered content knowledge for professional

advancement (e.g., credentialing exams).

This chapter provides an overview of three kinds of achievement tests (survey battery, diagnostic, and readiness tests) and one kind of aptitude test (cognitive ability test). Together, these make up the four domains of educational assessment (see the shaded domains in Figure 8.1). Individual intelligence tests, and special and multiple aptitude tests, which largely do not focus on educational ability, are also sometimes used in the schools and have other broad applications. These tests will be covered in Chapter 9 and Chapter 10, respectively.

As you might recall from Chapter 1, we defined survey battery, diagnostic, readiness, and cognitive ability tests in the following ways:

Survey Battery Tests: Tests, usually given in school settings, which measure broad content areas and often used to assess progress in school.

Diagnostic Tests: Tests that assess problem areas of learning and often used to assess learning disabilities.

Readiness Tests: Tests that measure one’s readiness for moving ahead in school and often used to assess readiness to enter first grade.

Cognitive Ability Tests: Tests that measure a broad range of cognitive ability. These tests are usually based on what one has learned in school and are useful in making predictions about the future (e.g., whether an individual might suc- ceed in school or in college).

You might also remember from Chapter 1 that sometimes the difference between Achievement Testing and Aptitude Testing has more to do with how the

158 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

test is being used than what the test is actually measuring. For instance, although the SAT seems to be measuring content knowledge or what you have learned in school, it is used to predict success in college and is therefore often listed in the Aptitude Section of Ability Testing. That is why there is a double-headed arrow in Figure 8.1—to remind us that many of these tests share much in common with one another—at least in terms of what they are measuring. In the rest of this chapter, we examine a number of survey battery, diagnostic, readiness, and cognitive ability tests and conclude with some final thoughts about the use and misuse of such tests.

SURVEY BATTERY ACHIEVEMENT TESTING With millions of children taking achievement tests every year, achievement testing has become a huge industry. Not surprisingly, a number of publishing companies spend a great deal of money on the development and refinement of survey battery achievement tests. Also, with “No Child Left Behind” (NCLB) mandating that states must show that “adequate yearly progress” is being made toward all students achieving at state specified academic standards (U.S. Department of Education, 2005), survey battery achievement testing has taken on a bigger role as they’re used to document progress toward this goal (see Box 8.1).

Survey Battery Achievement Tests can be helpful on a number of levels. For instance, by providing individual profile reports, they can help a student, his or her parents, and his or her teachers identify strengths and weaknesses and develop strategies for working on weak academic areas. Similarly, profile reports at the classroom, school, or school system level can make it possible for teachers, princi- pals, administrators, and the public see how students are doing and help deter- mine which students, staff, and schools might obtain needed resources to improve test scores. This section of the chapter first reviews the testing process of the National Assessment of Educational Progress (NAEP) that uses achievement testing to assess how each state is doing compared to other states. Then, we examine three

ASSESSMENT OF ABILITY ACHIEVEMENT AND APTITUDE TESTING

ACHIEVEMENT TESTING (What one has learned)

APTITUDE TESTING (What one is capable of learning)

Survey Battery

Diagnostic Readiness

Intelligence

Testing

Neuropsychological

Assessment

Intellectual and Cognitive

Functioning

Cognitive Ability

Multiple Aptitude

Special Aptitude

FIGURE 8.1 | Tests in the Cognitive Domain

Survey battery tests Paper-and-pencil tests measuring broad knowledge content

NCLB Federal law ensuring all children succeed in school

© Ce ng ag e Le ar ni ng

CHAPTER 8 Assessment of Educational Ability 159

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

of the most frequently used survey battery achievement tests often given by various states, including the Stanford Achievement Test, the Iowa Test of Basic Skills (ITBS), and the Metropolitan Achievement Test.

National Assessment of Educational Progress (NAEP) Partly because each state has a unique way of assessing its students, there was a push to develop one assessment procedure that could make national comparisons across all states. Sometimes called the National Report Card, the NAEP, which is sponsored by the U.S. Department of Education, samples students from all the states and compares them on a variety of subjects. Results are not provided for spe- cific students, classes, or schools and states cannot use the NAEP to show that ade- quate progress has been made toward NCLB.

All states are required to participate in the NAEP assessment in math and read- ing that occurs every two years, mostly at grades 4 and 8, and most states also par- ticipate in periodic testing in writing and science. In addition, other subjects such as the arts, civics, economics, geography, U.S. history, and starting in 2014, technol- ogy and engineering literacy (TEL), are sometimes assessed on a voluntary basis. NAEP carefully selects a representative sample of 3,000 students from 100 public schools from each state (National Center for Education Statistics, 2012) for their testing data. Results are typically given as the percentage of students who scored above basic, at proficient, or at the advanced achievement level (See Figure 8.2). In addition, scaled scores that can be compared year to year within the state, to other

BOX 8.1 High-Stakes Testing: Achievement Testing and No Child Left Behind

The “No Child Left Behind Act” (NCLB) is a federal law requiring each state to have a plan to show how, by the year 2014, all students will have obtained profi- ciency in reading/language arts and math (also see Box 5.1) (U.S. Department of Education, 2011). Not surprisingly, scores on achievement tests are gener- ally used to measure success toward NCLB. Thus, the development of new achievement tests, or the use of existing achievement tests, has become more important than ever in showing that all children are succeeding. For instance, Virginia has developed its own test, called the Standards of Learning, which they use to show that progress has been made toward reaching proficiency as identified in NCLB.

NCLB ties federal funding to the success of school districts and defines a number of actions

that states must perform to show that attempts are being made to increase student scores. For good or bad, the NCLB puts a lot of pressure on the lower- performing school districts, which tend to be located in poorer neighborhoods (Baker, Robinson, Danner, & Neukrug, 2001; Lloyd, 2008).

Test scores not only allow schools to identify students who are not meeting specified levels of achievement but also demonstrate how well each class, grade, school, districts, or state is performing. As you might guess, there is a lot of pressure on teachers, principals, and superintendents, for if their class, school, or district does not show adequate progress, their jobs can be at stake. The upside: It is hoped that as a result of NCLB learning will improve, particularly learning among those who have traditionally been most disenfranchised.

NAEP The “nation’s report card”

© Ce ng ag e Le ar ni ng

160 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

FIGURE 8.2 | The National Report Card: Virginia Mathematics Source: U.S. Department of Education. (2012). The nation’s report card: Mathematics. Virginia, Grade 8. Retrieved from http://nces.ed.gov/nationsreportcard/pdf/stt2011/2012451VA8.pdf

161

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

states, and to a national sample are provided. Scaled scores cannot be compared between subjects. Results can also be sorted by subject area, gender, ethnicity, and eligibility for national school lunch programs. In Figure 8.2 what trends do you see occurring in Virginia on the mathematics NAEP?

In addition to the state assessment, a long-term national assessment is also completed every four years. This assessment looks at student performance in math- ematics and reading for ages 9, 13, and 17. For the national data, a sample of approximately 9,000 students at ages 9, 13, and 17 in public and nonpublic schools are assessed in mathematics and reading.

Both the state and national long-term assessments appear to undergo a rigor- ous process to assure that their tests are assessing national curriculum and that they are reliable. The NAEP has become an important test that helps states exam- ine whether or not they are achieving at adequate levels.

Stanford Achievement Test As mentioned in Chapter 1, the Stanford Achievement Test is one of the oldest sur- vey battery achievement tests, having been introduced in 1923 (Carney, 2005; Morse, 2005). Since that time, it has gone through a number of revisions prior to publication of its 10th edition in 2003. The Stanford Achievement Test, 10th Edi- tion (Stanford 10), is given to students in grades K–12 and has been normed against hundreds of thousands of students.

This latest edition of the Stanford Achievement Test offers many options such as full-length or abbreviated versions as well as content modules, which are tests for specific subjects such as reading, language, spelling, mathematics, science, social studies, and writing (science and social science are merged into a heading called “environment” at the lower levels) (Carney, 2005). The Stanford Achievement Test also has sections that can be completed in open-ended format, requiring stu- dents to fill in the blank, respond with short answers, or write an essay that will be scored by the classroom teacher using criterion grading.

Reliability of the Stanford Achievement Test appears sound, with most subtests showing KR-20 internal consistency estimates in the mid 0.80s to low 0.90s. How- ever, reliability estimates for the open-ended sections generally fell to the 0.60s to 0.80s and in some cases to the mid 0.50s (Harcourt Assessment, 2004). The Stanford Achievement Test appears to have sound validity, with content validity being established by working with content experts, teachers, editors, and measure- ment specialists. Evidence for construct validity was established by comparing scores on subtests with the Otis-Lennon School Ability Test (OLSAT), which produced a fairly high correlation. Criterion-related validity was addressed by numerous studies and is thorough and reasonable.

The Stanford Achievement Test, like other nationally made survey battery achievement tests, offers a number of interpretive reports, including Individual Profile Reports, Class Grouping Reports, Grade Grouping Reports, and School System Grouping Reports. Figure 8.3 shows an example of an Individual Student Profile, and Figure 8.4 shows a Class Group Report from 15 students in a fictitious fourth-grade class. Information can also be disaggregated by ethnicity, socioeco- nomic status, limited English proficiency (LEP), and whether students have an

Stanford 10 Assesses subject areas in school

162 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Individualized Education Plan (IEP). Also, the test is sometimes given with the Otis- Lennon Ability test so that scores on how one has done in specific content areas (the achievement test) can be compared to an individual’s potential (an aptitude test) (see Box 8.2). This is discussed later in this chapter, in more detail, when we look at cognitive ability tests.

Iowa Test of Basic Skills (ITBS) One of the oldest and best-known achievement tests is the Iowa Test of Basic Skills (ITBS). Initially developed in 1935, it has changed greatly from its early inception. Today, the test emphasizes the basic skills necessary to make satisfac- tory progress through school. The purposes of the instrument are to “(a) to obtain information that can support instructional decisions made by teachers in the classroom, (b) to provide information to students and their parents for moni- toring the student’s growth from grade to grade, and (c) to examine yearly prog- ress of grade groups as they pass through the school’s curriculum” (Engelhard, 2007, para. 4).

FIGURE 8.3 | Individual Student Profile Report on the Stanford 10 Source: Pearson. (2012). Sample student report. Retrieved from http://www.pearsonassessments.com/HAIWEB /Cultures/en-us/Productdetail.htm?Pid=SAT10C

IEP Mandatory plan to ac- commodate students with special needs

ITBS Measures skills to “satisfactorily” prog- ress through school

CHAPTER 8 Assessment of Educational Ability 163

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The versions currently used are Form A (2001), Form B (2003), and Form C (2007). These versions of the test are geared for K through 8 and have numerous subtests, depending on the grade level, which may include language, reading com- prehension, vocabulary, listening, word analysis, math, social studies, science, and sources of information (the ability of students to use maps, dictionaries, reference

FIGURE 8.4 | Class/Grade Grouping Report on the Stanford 10 Source: Pearson. (2012). Sample group report. Retrieved from http://www.pearsonassessments.com/HAIWEB/Cultures /en-us/Productdetail.htm?Pid=SAT10C

BOX 8.2 Survey Battery Achievement Tests and Cognitive Ability Tests

Sometimes, students are given a cognitive ability test at around the same time they take a survey battery achievement test. Then, teachers, school counse- lors, school psychologists, and others can look at differences in the scores between the two. Since survey battery achievement tests measure what one has learned at a specific grade level and cogni- tive ability tests assess overall potential, differences between the two types of test can be an important signal to problems in learning. What do you think might cause a student to do significantly lower on a

survey battery achievement test as compared to a cognitive ability test? One of the major causes is a learning disability. Other causes can be problems at home, problems at school, poor peer relationships, poor teaching, and more. What do you think might be the result of a person scoring significantly higher on a survey battery achievement test as compared to a cognitive ability test? When we examine cogni- tive ability tests later in the chapter, you may want to revisit the differences between these two types of tests.

© Ce ng ag e Le ar ni ng

20 15

164 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

materials, charts, and so forth). Testing can run from 30 minutes for a single test to over 6 hours when a total battery is administered.

The authors show strong evidence of content validity that was guided by what is typically covered across the country, current teaching methods and text- books used, and a review of national curriculum (Engelhard, 2007; Lane, 2007). Reliability of most subtests are in the mid 0.80s to low 0.90s, and the authors offer a Guide to Research and Development Manual (GRD), which stresses the strengths and weaknesses of the instrument. Finally, the publisher demonstrates correlations from the low 0.70s to the low 0.80s with the Cognitive Ability Test (CogAT), offering additional validity evidence for the instrument and providing a mechanism to possibly identify learning problems if scores on the ITBS are considerably lower than those on the CogAT (see Box 8.2 and CogAT later in the chapter).

Metropolitan Achievement Test First published during the 1930s, The Metropolitan Achievement Test is another popular survey battery achievement test, the newest version of which is the eighth edition (Harwell, 2005; Lukin, 2005). It is designed to test students in grades K–12 for knowledge in a broad range of subjects such as reading, language arts, mathematics, science, and social studies. It has 13 test levels (K–12) and can be given in a short form that takes 90 minutes or the complete form that can take up to 5 hours. Test items consist of multiple-choice questions, which are graded correct or incorrect, and open-ended items, which are scored as 0 to 3. Although research based on extensive sampling data has been fairly exhaustive, some have suggested these samples might be too heavily weighted for rural classrooms and underrepresent urban classrooms. As with the Stanford and ITBS, reliability esti- mates are quite high, usually between 0.8 and 0.9 or higher for most subtests, and content, criterion, and construct validity are sound.

DIAGNOSTIC TESTING With the passage of Public Law 94-142 (PL94-142) in 1975, as well as the more recent Individuals with Disabilities Education Act (IDEA), millions of children and young adults between the ages of 3 and 21 who were found to have a learning dis- ability were assured the right to an education within the least restrictive environment (Federal Register, 1977; U.S. Department of Education, 2005). These laws also assert that any individual who is suspected of having one of many disabilities that interfere with learning has the right to be tested, at the school system’s expense, for the dis- ability. Thus, diagnostic testing, usually administered by the school psychologist or learning disability specialist, has become one of the main ways to determine who might be learning disabled. These laws also state that a school team should review the test results and other assessment information obtained, and that any student identified as learning-disabled would be given an IEP describing services that should be offered to assist the student with his or her learning problem.

Although diagnostic testing was certainly in existence prior to PL94-142 and IDEA, its use, as well as the development of new diagnostic tests, greatly expanded

Metropolitan Achievement Test Assesses subject areas in school—has option for open-ended questions

PL94-142 Asserts right to be tested for learning disabilities

IDEA Extension of PL94- 142

Diagnostic tests Used to assess learn- ing disabilities or difficulties

CHAPTER 8 Assessment of Educational Ability 165

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

as the result of these laws. In this section, we review five of the more common diagnostic tests of achievement often used to assess learning problems: the Wide Range Achievement Test 4 (WRAT4), the Wechsler Individual Achievement Test (WIAT), the Peabody Individual Achievement Test (PIAT), the Woodcock- Johnson, and the Key Math Diagnostic Arithmetic Test. Throughout our discus- sion, keep in mind that dozens of diagnostic tests are in use today.

The Wide Range Achievement Test 4 (WRAT4) The Wide Range Achievement Test 4 (WRAT4) was developed to assess basic read- ing, spelling, math, and sentence comprehension skills. It is “intended for use by those professionals who need a quick, simple, psychometrically sound assessment of important fundamental academic skills” (Wilkinson & Robertson, 2006, p. 3). It is called “wide-range” because it can be used with individuals from ages 5 to 94. The test takes between 15 and 45 minutes, depending on the age of the individ- ual, and is administered individually because some sections are read aloud by the examinee. There are two equivalent forms of the exam called “blue” and “green” (PAR, 2012).

The WRAT4 attempts to assure that the test is assessing the fundamentals of reading, spelling, and arithmetic, as opposed to comprehension, which is often the case when examinees are asked to read multiple-choice questions or paragraphs. Thus, the test is fairly simple to administer as the individual is asked by the exam- iner to “read” (pronounce) words, to spell words, to figure out a number of math problems, and with its latest revision, to provide a missing word or words to simple sentences to show you understand the meaning of the sentence. The test includes the original three subtests: word reading, spelling, and math computation as well as the new, fourth subset called sentence comprehension. Combining the word reading and sentence comprehension subtests provides a reading composite score. The spelling comprehension and math computation subtests can be given in group format but the word reading and sentence comprehension must be adminis- tered individually.

Scores are presented in a number of ways including grade equivalents, percen- tiles, NCEs, stanines, and a standard score that compares the individual by grade level or by age and uses a deviation IQ (DIQ; mean of 100 and standard deviation of 15) (see Table 8.1). Confidence intervals, which are based on the WRAT4’s standard error of measurement (SEM), are also provided for the standard score and suggest where the individual’s “true score” falls 95% of the time. The DIQ is used so that the WRAT4 can be compared to scores on an intelligence test, such as the Wechsler Intelligence Scale for Children—Fourth Edition (WISC-IV). This is important as a significantly lower score on any of the scales of the WRAT4, as compared to the overall DIQ, could indicate the presence of a learning disability. Generally, more intensive testing for a learning disability would follow such findings.

Internal consistency reliability estimates are impressive and generally run in the 0.90s, and alternate form reliability averages in the mid 0.80s (Wilkinson & Robertson, 2006). The authors provide a rationale for the content of the test and demonstrate evidence of construct and criterion validity, such as moderate

WRAT4 Assesses basic learn- ing problems in read- ing, spelling, math, and sentence comprehension

Significant differences between WRAT4 and IQ may indicate learn- ing disability

166 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

correlations with the WISC-III, the Wechsler Adult Intelligence Test—Revised (WAIS-R), and the Stanford Achievement Test.

Wechsler Individual Achievement Test— Third Edition (WIAT-III) The Wechsler Individual Achievement Test—Third Edition (WIAT-III) is an indi- vidually administered achievement test most used for Prekindergarten through 12th grade, although norms also exist for ages 4 through 50 (Willse, 2010). The intended purpose of the WIAT-III is to “Identify the academic strengths and weak- nesses of a student; inform decisions regarding eligibility for educational services, educational placement, or diagnosis of a specific learning disability; design instruc- tional objectives and plan intervention” (Pearson, 2012a, “overview”). Thus, this test is designed to provide assessment and intervention strategies for those with spe- cific learning disabilities. The test consists of 16 subtests that comprise 7 composite scores (see Figure 8.5).

Internal consistency and test-retest reliability tends to be high, ranging mostly in the 0.80s and 0.90s for the subtests and in the 0.90s for the composite scores. The test manual shows evidence of content validity and of convergent validity with other similar tests. It also shows some evidence of construct validity by com- paring special populations (e.g., learning disabled, gifted individuals) with match control groups. Testing time is from 1 minute to 17 minutes, depending on the age and grade level of the person being and the number of areas being assessed. Figure 8.6 gives a sample clinical report of a seventh-grade student.

Peabody Individual Achievement Test (PIAT-R/NU) The 1998 edition of Peabody Individual Achievement Test—Revised–Normative Update (PIAT-R/NU) provides academic screening for children in grades K–12 and covers six content areas: general information, reading recognition, reading

TABLE 8.1 Score Summary Table (Green Test Form)

Subtest/Composite Raw Score

Standard Score Norms: Age

Norms Confidence Interval 95%

Percentile Rank

Grade Equivalent NCE Stanine

Word Reading 34 84 76–93 14 3.5 28 3

Sentence Completion 7 68 61–77 2 1.3 5 1

Spelling 27 85 76–95 16 3.7 29 3

Math Computation 37 105 94–115 63 7.3 57 6

Reading Composite* 152 74 69–80 4 N/A 13 2

*Reading Composite Raw Score ¼ Word Reading Standard Score þ Sentence Completion Standard Score Source: Reynolds, C. R., Kamphaus, R., & PAR Staff. (2007). RIASTM/WRAT4 discrepancy report. Retrieved from http://www4.parinc.com /WebUploads/samplerpts/RIASWRAT4_TBDiscrepancy_Pred.pdf

WIAT-III Diagnostic test to screen broad areas of ability

PIAT Six content areas for screening K–12 students

CHAPTER 8 Assessment of Educational Ability 167

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

comprehension, mathematics, spelling, and written expression (Cross 2001; Fager, 2001). Although re-normed, the test is the same as the 1989 version. Except for general information and written expression, the test is multiple-choice. Individually administered, the instrument takes about 1 hour to give. Multiple derived scores can be obtained, including a standard score using DIQ, age equivalents, grade equivalents, percentile ranks, and normal curve equivalents (NCEs). A developmen- tal scaled score is used for the written expression subtest (mean of 8, standard devi- ation of about 3).

A wide variety of reliability estimates for the revised version show a median of approximately 0.94, although the written portion subtest, which is hand-graded, has significantly lower interrater reliability (Cross, 2001; Fager, 2001). Content validity was established by using a number of school curriculum guides, and the manual shows evidence of criterion and construct validity. The new normative

FIGURE 8.5 | Subtests and Composite Scores on the WIAT-III Source: Pearson Education. (2012). Subtests and composite scores. Retrieved from http://www.pearsonassessments .com/hai/images/products/wiat-iii/wiat-iii-chart.gif

168 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

FIGURE 8.6 | WIAT-III Clinician Report Source: Pearson Education. (2009). WIAT-III Clinician report. Retrieved from http://www.pearsonassessments.com /pai/ca/training/webinars/WIAT-IIIWebinar.htm

CHAPTER 8 Assessment of Educational Ability 169

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

sample ranged from 2,809 students for the mathematics subtest to 1,285 students for the written expression subtest. The instrument sampled well for sex, ethnicity, parental educational level, special education, and gifted students, but the Northeast and West are somewhat underrepresented.

Woodcock-Johnson® III A comprehensive, individually administered diagnostic test, the Woodcock- Johnson® III is designed to assess “cognitive abilities, skills, and academic knowl- edge most recognized as comprising human intelligence and routinely encountered in school and other settings” (Cizek, 2003, para. 1). Although normed and applica- ble to individuals between the ages of 2 and 90, the instrument is generally used for students around the age of 10. The instrument, which shows strong evidence of test worthiness, actually consists of two batteries: (1) the Woodcock-Johnson Tests of Achievement, which examines academic strengths, and (2) the Woodcock-Johnson Tests of Cognitive Abilities, which looks at specific and general cognitive abilities. These tests, which can be given in standard and extended versions, were recently renormed with a diverse 8,800-subject sample.

KeyMath3™ Diagnostic Assessment The KeyMath3™ has been described as “an individually administered test that is well developed and provides scores that can inform the design of individual student intervention programs and monitor performance over time” (Graham, Lane, & Moore, 2010, “summary”). Often used as a follow-up when there is a suspected math learning disability, the test is generally administered and interpreted by a per- son who has been well versed in the test and in learning disabilities.

The KeyMath3 has 10 subtests grouped under three broad math content areas: basic concepts (conceptual knowledge), operations (computational knowledge), and applications (problem solving) (Pearson, 2012b) (see a sample applications item in Figure 8.7).

The test is appropriate for children in kindergarten through ninth grade or for individuals between the ages of 4.5 to 21 if they are believed to be functioning between the K–9 grade levels. The test is untimed, but generally takes between 30 and 90 minutes to finish. Standard scores, scaled scores, percentiles, and grade or age equivalents are reported. This newest version was updated to reflect current curriculum standards and normed with a sample of 4,000 individuals to reflect gender, ethnicity, socioeconomic status, region of the country, and wide range of learning disabilities.

Internal consistency reliability estimates range from the 0.60s to the mid 0.90s, although most are in the 0.80s. To develop content validity, the author’s reviewed state and national curriculum standards and information retrieved from experts in math education (Graham et al., 2010). To address issues of cross- cultural bias, a broad review of items based on sex, race, ethnicity, culture, and geographic region was completed (KeyMath—3 DA Publication Summary Form, 2007).

Woodcock- Johnson® III Broad assessment of ability, ages 2–90

KeyMath3™ Comprehensive test to assess for learning disabilities in math

170 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

READINESS TESTING Tonight, I propose working with states to make high-quality preschool available to every child in America. (Barack Obama, 2013 State of the Union Address)

Although President Obama’s goal of high-quality preschool for all children may be lofty, it’s grounded in the belief that if we are all to do well in life, we must start with a safe and intellectually stimulating environment in our early years (Marjanovic, Kranjc, Fekonja, & Bajc, 2008). Unfortunately, today, too many of our children are not ready to learn when they begin kindergarten. One of the ways that we can assess a student’s readiness to learn as he or she enters kindergarten or first grade is through a readiness test. Over the years, such tests have tended to be classified as either measurements of ability, often in reading or math achievement, or those that assess developmental level, such as psychomotor ability, language ability, and social maturity.

Because children change so rapidly at these young ages, and because the pre- dictive ability of these tests tends to be weak, readiness testing has always been a questionable practice (Harris, 2007). In addition, because these assessments often carry cultural and language biases, children from low-income families, minority groups, and homes where English is not the first language will often obtain scores lower than their true ability. Thus, these tests need to be administered with care, if given at all. Despite problems with these tests, when a child’s readiness to enter kindergarten or first grade is unclear, these instruments can sometimes be helpful and are used in many preschools and elementary schools. Generally, there are two

FIGURE 8.7 | Sample Item for the Applied Section of the KeyMath3 Source: Pearson. (2012). Applied problem solving: Sample items. Retrieved from http://psychcorp.pearsonassessments.com/hai/images/pa /products/keymath3_da/km_forma_aps_web.pdf

Readiness tests Assesses readiness for kindergarten or first grade

CHAPTER 8 Assessment of Educational Ability 171

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

categories of readiness testing—those that measure ability and those that measure developmental level, with ability readiness tests making up the bulk of readiness tests. Three ability-based readiness tests we examine include two tests, with the same name, called the Kindergarten Readiness Test (KRT), and a test called the Metropolitan Readiness Test. One developmental test we will look at is the Gesell Developmental Observation instrument.

Kindergarten Readiness Test (KRT) (Anderhalter & Perney) The KRT (Anderhalter & Perney, 2006) is designed to assess competencies in six areas for children who are finishing preschool or at the very beginning of kinder- garten. The authors hope that the information can be used to help determine readi- ness for kindergarten and to help in determining curriculum planning for individual students. The test, which takes 25 to 30 minutes, can be given individually or in groups of two or three students. In addition to obtaining six subtest scores (vocab- ulary, identifying letters, visual discrimination, phonemic awareness, comprehen- sion and interpretation, and mathematical knowledge), a score for the entire test is also obtained and converted to a total readiness score of “not ready without atten- tion,” “marginally ready,” average degree of readiness,” and “above average degree of readiness.”

Although internal consistency reliability of the subtests and of the overall score seems reasonable, there is little evidence that the test can adequately be used for modifying curriculum for students who have taken the test and little evidence of predictive validity; that is, the ability of the test to predict how well a student will do in kindergarten (Johnson, 2010; Swerdlike & Hoff, 2010). This makes the instrument somewhat questionable for its intended purpose.

Kindergarten Readiness Test (KRT) (Larson and Vitali) Yes, you are seeing the same name of a test used again. However, this readiness test is developed by Larson and Vitali (1988). Similar to the first KRT, this KRT assesses whether children, between the ages of 4 and 6, are “developmentally or maturationally ready to begin kindergarten” by assessing five skills areas: under- standing, awareness, and interactions with one’s environment; judgment and reasoning in problem solving; numerical awareness; visual and fine-motor coordina- tion; and auditory attention span and concentration (Slosson Educational Publications, n.d., para. 2).

The test takes between 15 and 20 minutes to give, and although administrative procedures and item quality are acceptable, items sometimes seem out of sequence and reliability and validity information is minimal (Beck, 1995; Sutton & Knight, 1995). Also, the sample population is drawn entirely from a four-state region in the Midwestern United States, which is not representative of the nation (Sutton & Knight, 1995). Despite these drawbacks, this instrument may be useful in determining whether a student is ready to begin kindergarten if the user believes the content of the test matches the curriculum of the school the child will be attending.

Kindergarten Readiness Test Assesses broad range of cognitive and sen- sory motor skills

Kindergarten Readiness Test Developmental or maturational readiness in five skill areas

172 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Metropolitan Readiness Test (MRT6) The Metropolitan Readiness Test, Sixth Edition (MRT6), is designed to assess beginning educational skills in preschoolers, kindergarteners, and first graders (Novak, 2001). Level 1 of the test is administered individually and assesses literacy development for preschoolers and beginning kindergarteners, while Level 2 assesses reading and mathematics development of kindergarteners through beginning first graders and is usually given in a group setting. The test generally takes between 80 and 100 minutes to administer, and the results are often used as an aid in deter- mining whether a student should be placed in first or second grade. Results of the Metropolitan Readiness Test are generally reported using raw scores, stanines, or percentiles.

Reliability estimates for the composite test are strong and mostly hover around the 0.90s. Individual subtest reliabilities tend to deteriorate, often ranging from 0.53 to 0.77. Some have questioned the validity evidence of the MRT6 and its earlier versions. In fact, Kamphaus (2001) states that the Level 1 test shows no evidence of validity and that Level 2 has “virtually none” and is “unacceptable.” However, others are more forgiving and find the MRT6 “useful in determining early academic or ‘readiness’ skills in reading and math” (Novak, 2001, “Summary”).

Gesell Developmental Observation—Revised The Gesell Developmental Observation—Revised is “a standardized assessment sensitive to the development of the whole child (social-emotional, physical, cogni- tive including language, and adaptive)” for children between the ages of 2.5 and 9 (Gesell Institute, 2011, p. 4). The test, which was most recently revised in 2010, is based on the work of Arnold Gesell, who spent years examining the normal devel- opment of children (Gesell Institute of Child Development, 2012). The test is administered in a nonthreatening and comfortable environment by a highly trained examiner who observes the child’s developmental maturity to assess the child’s readiness to excel in different settings. This contrasts greatly with most other tests of readiness that focuses mostly on cognitive ability that use chronological age or intelligence in making decisions on child readiness (Bradley, 1985). The humanistic approach that the examiner takes when administering the test is appealing to many; however, there is quite a bit of examiner discretion in how the test is scored and interpreted.

Recent reviews of the newly revised test have not been available; however, reviews of past editions have indicated that although the test provides some useful descriptive information about how children might perform in certain types of situa- tions, overall, the test has been weak in providing adequate information about its validity and reliability (Bradley, 1985; Waters, 1985). These reviews have also noted that although the test manual describes age-appropriate responses, it does not clearly specify how placement recommendations are made once an assessment has been completed. Therefore, school systems are not likely to use the instrument. Hopefully, some of these issues have been addressed in the more recent revision, and despite these problems, the test is sometimes used because it offers a view of readiness different from those that are based strictly on achievement in a content

MRT6 Assesses literacy development, reading, and mathematics

Gesell Development Observation Assesses develop- ment of the whole child

CHAPTER 8 Assessment of Educational Ability 173

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

area. As far as Gesell was concerned, “achievement” involved more than getting high scores on a reading or math test (see Box 8.3).

COGNITIVE ABILITY TESTS As noted earlier, cognitive ability tests are aptitude tests that measure what one is capable of doing and are often used to assess a student’s potential to succeed in grades K–12, college, or graduate school. First, we look at two K–12 tests: the Otis-Lennon School Ability Test (OLSAT) and the Cognitive Ability Test (CogAT). Then we look at tests used to assess potential ability in college and graduate school: the American College Testing Assessment (ACT), the SAT, the Graduate Record Exam (GRE), the Miller Analogy Test (MAT), the Law School Admission Test (LSAT), and the Medical College Admission Test (MCAT).

Otis-Lennon School Ability Test, Eighth Edition (OLSAT 8) The Otis-Lennon School Ability Test, Eighth Edition (OLSAT 8), is one of the more common cognitive ability tests. “The OLSAT 8 supplies educators with valu- able information to enhance the insights gained from traditional achievement tests” (Pearson, 2012c, “overview”). Usually given in large group format and for students

BOX 8.3 Measuring Assessment: Ability or Developmental Maturity?

When my daughter, Hannah, entered kindergarten, we asked her principal if she could be assessed for readiness for first grade, since her birthday missed the school system’s deadline to enter first grade by only two days. Our daughter had gone to a very humanistically oriented preschool where achieve- ment was not stressed but optimizing developmental level was. Between the preschool and some pretty good parenting (well, I think we did a good job!), our daughter clearly was mature for her age, socially adept, and seemed to be above average for most developmental milestones such as language and motor development. However, was she “smart enough” to enter first grade?

Hannah’s public school was happy to give her a reading readiness test to see if she would “fit” into first grade. After taking the test, which was given by a reading specialist, we were told she would be a star in kindergarten and have some “catching up” to do in first grade. We were also told that it was our decision whether or not to move her into first

grade. So what to do—go with our sense of her developmental level or go with the reading test?

First, we consulted with her preschool teachers and experts in child development (after all, I do work in a College of Education). Boy did we get mixed opinions. Knowing that such a decision should be based on many factors, not just one test, we consid- ered Hannah holistically and hesitatingly decided to place her in first grade. We moved her to first grade, and as the reading specialist predicted, she did have some catching up to do. However, our decision seems to have been a good one, since she has excelled in school and socially since that time. Maybe we made the best decision based on all the evidence, and maybe we just lucked out. Who knows? But Hannah’s example shows how making a decision about early readiness is complex and highlights the importance of considering a variety of factors when in the decision-making process.

—Ed Neukrug

Cognitive ability tests Measure what one is capable of doing

OLSAT 8 Assesses abstract thinking and reasoning skills via verbal and nonverbal sections

© Ce ng ag e Le ar ni ng

174 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

in K–12, the test assesses different clusters in the verbal and nonverbal realms. For instance, the two clusters for verbal ability include verbal comprehension and ver- bal reasoning. The three clusters for nonverbal ability include pictorial reasoning, figural reasoning, and quantitative reasoning. As you might expect, different grade levels are given different clusters, and each cluster has different subtests. Testing time is between 60 and 75 minutes depending on the age of the student.

A variety of scores can be used in describing OLSAT 8 results, including a School Ability Index (SAI), which uses a mean of 100 and SD of 16, percentile ranks based on age and grade, stanines, scaled scores, and NCEs based on grade. In addition, when given along with the Stanford 10, an Achievement/ Ability Comparison (AAC) score can be obtained to give teachers insights into how students are actually doing in school compared to their potential. Signifi- cantly higher scores on a cognitive ability test (e.g., the OLSAT 8) as compared to an achievement test (e.g., Stanford 10) could be an indication of a learning disability (see Box 8.2). Figure 8.8 shows an Individual Profile Report from the

FIGURE 8.8 | OLSAT 8 Individual Profile Report Source: OLSAT 8 Results Online. Retrieved from http://pearsonassess.com/haiweb/Cultures/en-US/Site/ProductsAndServices/Products/OLSAT8 /OLSAT+8+Results+Online.htm

CHAPTER 8 Assessment of Educational Ability 175

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

OLSAT 8 and Figure 8.9 shows 20 of 56 students from a Grade Profile Report of the OLSAT 8.

The norm groups of the OLSAT 8 consisted of 275,500 students in the spring of 2002 and an additional 135,000 in the fall of 2002 (Morse, 2010). Internal con- sistency measures of reliability based on the KR-20 for the composite score ranged from 0.89 to 0.94. Reliabilities for individual subtests using the KR-21 ranged from 0.52 to 0.82, with most falling in the 0.60s and 0.70s. Evidence of test con- tent validity is somewhat vague as the publisher notes that each user must deter- mine if the content fits the population they are testing. Correlation coefficients for the OLSAT 8 composite scores with OLSAT 7 scores demonstrated coefficients in the range of 0.74 to 0.85, depending on grade level. Similarly, correlations among

FIGURE 8.9 | OLSAT 8 Grade Profile Report Source: OLSAT 8 Results Online. Retrieved from http://pearsonassess.com/haiweb/Cultures/en-US/Site/ProductsAndServices/Products/OLSAT8 /OLSAT+8+Results+Online.htm

176 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

different levels of the OLSAT 8 were adequate. The test also showed reasonable correlations with the Stanford Achievement Test (Stanford 10). Finally, the test is often given around the same time as the Stanford Achievement Test and scores on each test can be listed on the profile reports thus making comparisons relatively easy (see Box 8.2 and Figures 8.3 and 8.4).

The Cognitive Ability Test (CogAT) Another common cognitive ability, the Cognitive Ability Test (CogAT) has a name befitting of its test category. Although a Form 7 has been developed, technical information about it is not readily available, so this review examines Form 6, which is very similar to the new form.

Form 6 of the CogAT is designed to assess cognitive skills of children from kin- dergarten through 12th grade (Riverside Publishing, 2010). The purpose of this test is threefold, and includes helping a teacher understand the ability of each child to optimize instruction, to provide a different means of measuring cognitive ability than traditional achievement tests, and to identify students who might have large discrepancies between their cognitive ability testing and their achievement testing (DiPerna, 2005). Such discrepancies can be indicative of learning problems, lack of motivation, problems at home, problems at school, or self-esteem issues. Teachers and support staff should be cognizant of such discrepancies and make appropriate referrals as necessary (see Box 8.2 and 8.4).

The CogAT measures three broad areas of ability: verbal ability, quantitative ability, and nonverbal ability and also offers a composite score. It is constructed with two models of intelligence in mind: Vernon’s hierarchical abilities and Cattell’s fluid and crystallized abilities (DiPerna, 2005; Rogers, 2005) (see Chapter 9). However, cognitive ability tests should never be viewed as substitutes for individ- ual intelligence tests, as the manner in which they are created and administered is

BOX 8.4 Identifying Needed Services

When I was in private practice, I worked with a third- grader who, a couple of years prior to my first meeting him, had been identified as having a math learning disability. This disability was first hypothe- sized after a large discrepancy was found between his cognitive ability score in math and his math achievement in school. Further testing verified the disability, and he was soon given an Individualized Education Plan that included three one-hour ses- sions of individualized assistance with math each week. After receiving this help, he soon began to do much better in math. However, after he had been getting higher math scores for about a year,

the school discontinued these services. I met him soon after this, following his parents’ divorce, when his math scores had once again dropped. The scores had likely dropped due to the chaos at home as well as the removal of services. I immediately realized that this young man was not being given the extra help he was legally entitled to receive. I contacted the school, which agreed that he should be given assistance in math based on his Individualized Education Plan (IEP). Within a few weeks after obtaining this assistance, his math scores once again improved, and he was a noticeably happier child.

—Ed Neukrug

CogAT Assesses verbal, quantitative, and non- verbal reasoning; uses Vernon’s and Cattell’s models of intelligence

© Ce ng ag e Le ar ni ng

CHAPTER 8 Assessment of Educational Ability 177

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

vastly different from intelligence tests and they tend to focus primarily on tradi- tional knowledge as obtained in school, particularly verbal and mathematical ability.

CogAT scores can be converted to standard scores that use a mean of 100 and standard deviation of 16, percentile ranks, and stanines. The entire test takes between 2 and 3 hours and is given in multiple administrations, depending on the age range of the student. In 2005, the CogAT 6 updated its norm group that included a more robust and varied sample. Internal consistency reliability of the CogAT 6 ranges from 0.86 to 0.96 for the verbal, quantitative, and nonverbal sections and from 0.94 to 0 .98 from the composite score (Lohman & Hagen, 2002).

Because it is difficult to define a knowledge base (content) that predicts future ability well, all cognitive ability tests have a difficult time establishing content valid- ity. However, the CogAT does offer a rationale by stating that the content domain was defined logically and through the sampling of student textbooks (Lohman & Hagen, 2002). Concurrent validity with the ITBS was 0.83, and CogAT fourth grade scores correlated 0.79 with ninth-grade scores on the ITBS (Lohman & Hagen, 2002). Studies are in progress to correlate the CogAT Form 6 with the Woodcock-Johnson III and the Wechsler Intelligence Scale for Children III. Although some have questioned whether the instrument should be used to deter- mine classroom learning, the test does show promise as an instrument in identify- ing student strengths and weaknesses and possible learning problems (DiPerna, 2005; Rogers, 2005).

College and Graduate School Admission Exams A number of cognitive ability tests are used to predict achievement in college and graduate school. Some of the ones with which you might be familiar include the ACT, the SATs, the GRE General Test and Subject Tests, the Miller Analogies Test (MAT), the LSAT, and the MCAT. Despite much consternation over the use of these tests, research indicates that such tests tend to predict performance in undergraduate and graduate school about as well or better than other indicators and are especially useful when combined with grade predictors (e.g., grades in high school or college grades) (Kobrin, Patterson, Shaw, Mattern, & Barbuti, 2008; Kuncel & Hezlett, 2007).

ACT The ACT, along with the SAT (see next section), are the two most widely used admission exams at the undergraduate level (Straus, 2012). The ACT assesses four skill areas that are based on what one has learned in high school: English, math, reading, and science. In addition, there is a composite score. The test con- tains 215 multiple-choice questions and takes 3.5 hours to complete, and the mean ACT composite score for graduating seniors tends to be about a 21, with a stan- dard deviation around 5 (ACT, 2012). The SEM for the composite score is about 1 (ACT, 2007). Publishers of the ACT performed a major norm sampling in 1988 with more than 100,000 high school students and another resampling update in 1995 with 24,000 students stratified against the U.S. Department of Education fig- ures for region, school size, affiliation, and ethnicity.

ACT Assesses educational development and abil- ity to complete college work

178 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Reliability of the ACT has ranged between 0.85 and 0.91 for the four skill areas and is 0.96 for the composite score. Evidence of content validity is shown through the test development process by consistently showing that test items are related to how students have “developed the academic skills and knowledge that are important for success in college” (ACT, 2007, p. 62). The ACT publishers also performed studies that showed a sound correlation between students’ ACT scores and their high school GPAs. Predictive validity studies correlating ACT scores and first-year college GPA had a median of 0.42. Combining ACT scores with high school GPA increased the predictive validity to 0.53. Numerous other studies of validity are available in the ACT technical manual (see Box 8.5).

SAT The other major undergraduate college exam is the SAT. The SAT measures critical thinking and problem solving skills in three areas: reading, mathematics, and a writing section that has multiple-choice questions as well as a writing sam- ple (College Board, 2013a). On each of the three sections students earn a score that ranges between 200 and 800 as well as a percentile score that compares the examinee’s results to those of students who have taken the test recently. On the writing section, students also receive a writing score that is based on their essay and ranges between 1 and 6 (6 being a better score) and is evaluated by two or three readers.

On the mathematics and critical reading sections, longitudinal comparisons can be made as test scores are compared to a 1990 norm group that had its mean “set” at about 500 and its standard deviation “set” at about 100. Thus, if the mean mathematics score of students in a recent year is 514 and the standard deviation is

BOX 8.5 Use of College Admission Exams: Selective or Oppressive?

Today, tests like the SATs and ACTs, along with a student’s GPA and other materials, are used to determine college readiness. Supporters of these college admission exams believe the test “equals the playing field” in that it allows all students to com- pete on the same test, whereas high school grades and rank can vary dramatically as a function of the student’s high school. However, others have sug- gested that they actively prevent some students from gaining access to some colleges.

For instance, consider the student who has the ability to do exceptionally well but attended a disad- vantaged school that lacked resources to promote college readiness? Or what about the student who does not have the economic resources to hire a tutor or to pay for an SAT or ACT study course? Or, think

about the student who hopes to be the first in his or her family to go to college and compare the kind of intellectual stimulation he or she gets at home to the child of highly educated parents. Unfortunately, all education is not created equal, particularly if you live in poverty.

At best, college admissions exams are a tool that standardizes the admission process by allowing students from various academic backgrounds to compete equally. At worst, they are tools that have the potential to widen the educational gap by dis- couraging, excluding, and stigmatizing those that lack access to adequate educational resources. What do you think?

—Michelle Reaves, Graduate Student, Old Dominion University

SAT Assesses reading, math, and writing— predicts mildly well for college grades

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 8 Assessment of Educational Ability 179

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

108, it can be said that this group is doing better than an earlier norm group with a lower mean and similar standard deviation. Such comparisons make it possible to determine whether today’s students are doing better or worse in mathematics and reading than students in past years.

Internal consistency reliability estimates for a national sample of the SAT ran- ged from 0.90 to 0.93 for critical reading, 0.91 to 0.93 for mathematics, and 0.88 to 0.91 for the composite writing section (including the essay) while the SEM was between 30 and 35 points for each of the three sections (College Board, 2012). On the different sections of the test, predictive validity with first year GPA and with fourth year GPA ranges between 0.47 and 0.54, respectively, and is between 0.53 and 0.56 for a combined score of the three tests (College Board, 2013b). As you might expect, combining high school GPA with the SATs offers a better pre- diction and is between 0.62 and 0.64 for first year college GPA and 0.64 for fourth year college GPA. The test predicts a little better for White students than “underrepresented groups” for first year GPA groups in college with correlations ranging from 0.46 to 0.51 for White students and from 0.40 to 0.46 for underrepre- sented groups (Mattern, Patterson, Shaw, Kobrin, & Barbuti, 2008). Correlations were slightly higher for females as compared to males. Interestingly, socioeconomic status seems to play little role in predicting college admissions when using the SATs (Sacket et al., 2012)

GRE General Test The GRE General Test is a cognitive ability test frequently required by U.S. graduate schools. The General Test contains three sections: verbal reasoning, quantitative reasoning, and analytical writing (Educational Testing Service, [ETS], 2007, 2013a). The scaled scores for the verbal and quantitative rea- soning sections used to be similar to that of the SAT, but recently have changed. Now, the verbal and quantitative sections range from scores of 130 to 170, while the analytical writing section is scored on a scale of 0 to 6 with one-half point increments. The analytical section is scored by two trained readers, and scores range from 0 to 6 (a third reader is brought in if the two scores are more than 1 point apart). Recent scaled score means and standard deviations for the verbal and quantitative sections of the test hovered around 151 and 8.6, respectively, while the mean and standard deviation of the writing section was 3.7 and 0.9, respectively. The Educational Testing Service does not “set” the mean or standard deviation and instead uses a scaled score mean and standard deviation that floats over time. Thus, it is probably prudent to look at percentile ranks as opposed to scaled scores as they allow examinees to compare themselves to those who are currently taking the same exam.

Reliability estimates for the GREs are 0.92 for the verbal reasoning and 0.95 for quantitative reasoning while it is a lower, 0.82 for the analytical writing section (ETS, 2013—2014). Correlations for the GRE General Test (verbal and quantita- tive) with a number of different predictors are as follows: with first-year graduate GPA (0.34 and 0.46), with comprehensive exam scores is (0.44 and 0.51), with fac- ulty ratings (0.42 and 0.50) (ETS, 2007). Correlations between combined scores on the verbal and quantitative sections with graduate school GPA for specific graduate school majors are often much higher (e.g., 0.51 for psychology; 0.66 for education (Burton & Wang, 2005).

GRE General Test Assesses verbal and quantitative reasoning and analytical writing; predicts graduate school success

180 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

GRE Subject Tests In addition to the GRE General Test, a number of subject tests are provided for those graduate programs that wish to assess more specific ability. The available subject tests include biochemistry, cell and molecular biology; biology; chemistry; computer science; literature in English; mathematics; physics; and psychology. Like the GRE General Test, the subject tests use a floating mean and standard deviation; however, the Subject Test’s scaled score ranges between 200 and 990 (ETS, 2013b). Means and standard deviations can vary dramatically among subject tests, so subject scores should not be compared to each other, although they can be compared to themselves over time. For instance, the mean and standard deviation for those examinees who took the biochemistry, cell and molecular biology test between July 2009 and June 2012 were 526 and 93, respec- tively, while the mean and standard deviation for psychology during that time were 616 and 102 (ETS, 2013–2014). Reliability of the subject tests tend to be in the low to mid 0.90s while the average correlation between subject tests and first-year graduate GPA is 0.45 (ETS, 2007).

Miller Analogy Test (MAT) Another test used for admission to graduate school is the Miller Analogies Test (MAT), which “measures your ability to recognize rela- tionships between ideas, your fluency in the English language, and your general knowledge of the humanities, natural sciences, mathematics, and social sciences” (MAT, 2013, p. 5). The test, which has 120 analogies, can be taken on computer or by hand and takes 1 hour to complete. A scaled score based on all who took the test between January 2008 and December 2011 are given and range from 200 to 600 with a mean set at approximately 400. Percentile scores are also compared to the examinee’s intended major and to the total group who took the test. Internal reliability coefficients range from 0.91 to 0.94. Predictive validity studies for stu- dents entering graduate school in 2005 showed correlations with first year GPA to be 0.27. Other correlations included 0.21 with the GRE verbal score, 0.27 with the GRE quantitative score, 0.11 with the GRE writing score, and 0.24 with under- graduate GPA. Although all but the writing score were statistically significant, these correlations reflect weak to moderate convergent and construct validity (Meagher, 2008).

Law School Admission Test (LSAT) The Law School Admissions Test (LSAT) is a half-day test that consists of five 35-minute sections used to determine admission to law school (Law School Admission Council, [LSAC], 2013). The test includes three multiple-choice sections measuring reading comprehension, analytical reason- ing, logical reasoning, and a fourth section that asks for a writing sample that is not scored but sent directly to law schools to which the applicant is applying. A fifth section is unscored and used to pretest new questions. The LSAT has a unique scoring method that ranges from 120 to 180 based on item difficulty and how the examinee has scored compared to others. Percentiles are also given. In 2010, the median of a number of correlations between LSAT scores and grades in the first year of law school were 0.36, and the correlation increases to 0.48 when the test is combined with undergraduate grade-point average (LSAC, n.d.). Reliability estimates are quite high, ranging from 0.90 to 0.95.

GRE Subject Test Predicts graduate school success in specific majors

MAT Uses analogies to as- sess analytical abilities; predicts graduate school success

LSAT Assesses acquired reading and verbal reasoning skills; pre- dicts grades in law school

CHAPTER 8 Assessment of Educational Ability 181

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Medical College Admission Test (MCAT) Most medical schools use the Medical College Admission Test (MCAT) as one factor in determining admissions. The test consists of four sections: physical sciences, verbal reasoning, biological sciences, and a volunteer trial section for future test questions (Association of Medical Colleges, [AAMC], 2012). Raw scores are obtained on all but the trial section of the test and converted to a scaled score that ranges from 1 to 15. The mean and standard deviation of the scaled scores can vary, and examinees are given a separate sheet that states the mean and standard deviation for the test they took as well as their percentile rankings. Reliability estimates for the MCAT are strong, but lower than many other similar kinds of tests. They are 0.85 for the biological section and for the physical sciences section, 0.80 for the verbal reasoning section, and 0.96 for the total multiple-choice score (AAMC, 2005). Although some research appears to support the predictive validity of the MCAT (see Callahan, Hojat, Veloski, Erdmann, & Gonnella, 2010; AAMC, 1995–2013), additional research that examines both its validity and reliability would be helpful.

THE ROLE OF HELPERS IN THE ASSESSMENT OF EDUCATIONAL ABILITY

In addition to teachers, a variety of helping professionals can play critical roles in assessing educational ability. For instance, school counselors, school psychologists, learning disabilities specialists, and school social workers often work together as members of the school’s special education team to assess eligibility for assessment of learning disabilities and to help determine a child’s Individualized Education Plan (IEP). School psychologists and learning disability specialists are generally the testing experts who are called upon to do the testing to identify learning problems. Sometimes outside experts in assessment, such as clinical and counseling psycholo- gists, are called upon to do additional assessments or to contribute a second opin- ion to the school’s assessment.

Because school counselors are some of the few experts in assessment who are housed at the school (school psychologists generally float from school to school), they will sometimes assist teachers in understanding and interpreting educational ability tests. In addition, by disaggregating the data from achievement tests, school

BOX 8.6 Assisting Teachers to Interpret Standardized Testing

A school system once hired me to run a series of workshops for their teachers on how to interpret the results of national standardized testing, such as survey battery achievement test scores. I was quite surprised to learn that few of them had ever taken a course on how to interpret such data and thus had no idea of the important information that could be gained from analyzing these data. I was, however,

impressed with the content knowledge that these teachers possessed. It was almost as though I was teaching them Greek (basic test statistics used in test interpretation) and they were teaching me Latin (the meaning of the content areas that were being reported on the test reports). It was certainly a learn- ing experience for all of us!

—Ed Neukrug

MCAT Assesses knowledge of physical sciences, verbal reasoning, bio- logical sciences; pre- dicts grades in medical school

© Ce ng ag e Le ar ni ng

182 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

counselors and others can play an important role in helping to identify students and classrooms that might need additional assistance in learning specific subject areas. Finally, licensed professionals in private practice need to know about the assessment of educational ability when working with children who are having pro- blems at school. In fact, it is often critical that these clinicians consult with profes- sionals in the schools to assure that the child is making adequate progress (see Box 8.6).

FINAL THOUGHTS ABOUT THE ASSESSMENT OF EDUCATIONAL ABILITY

As you can see from this chapter, the assessment of educational ability has become an important aspect of testing in the United States. Despite the widespread use of these tests, many criticisms have arisen, including the following:

1. Teachers are increasingly being forced to teach to the test. This prevents them from being creative and limits the kind of learning that takes place in the schools.

2. Testing leads to labeling and this can cause peers, teachers, parents, and others to treat the child as the label. For the child, the label sometimes becomes a self- fulfilling prophecy that prevents him or her from being able to transcend the label.

3. Some tests, particularly readiness tests and cognitive ability tests, are just a mechanism to allow majority children to move ahead and to keep minority children behind.

4. Testing fosters competitiveness and peer pressure, creating a failure identity for a large percentage of students.

On the other hand, many have spoken positively about tests of educational ability and have made the following points:

1. Tests allow us to identify children, classrooms, schools, and school systems that are performing poorly. This, in turn, allows us to address weaknesses in learning. In fact, evidence already exists that as a result of state standards of learning and the achievement testing associated with them, poor children, minority children, and others who traditionally have not done as well in schools have been doing better academically.

2. Without diagnostic testing, we could not identify a large portion of those chil- dren who have a learning disability, and we would not be able to offer them services to help them learn.

3. Testing allows a child to be accurately placed in his or her grade level. This ultimately provides a better learning environment for all children.

4. Testing helps children identify what they are good at and pinpoint weak areas that require added attention.

Probably both the criticism and praises of educational ability testing hold some truth. Perhaps as we realize the positive aspects of such testing we should also pay attention to the criticisms and find ways to address them.

CHAPTER 8 Assessment of Educational Ability 183

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SUMMARY In this chapter, we examined the assessment of educational ability. We started by noting that such tests have a variety of purposes, including determining if students are learning; assessing content knowledge of classes, schools, and school systems; identifying learning problems; determin- ing giftedness; deciding whether a child is ready to move to the next grade level; as a measure to assess teacher effectiveness; and to help determine readiness and placement in college and graduate school.

We identified four kinds of educational abil- ity testing, including three kinds of achievement tests (survey battery, diagnostic, and readiness) and one kind of aptitude test (cognitive ability testing). We defined each of these test categories and noted that often the difference between achievement testing and aptitude testing involves how the test is used as opposed to what it is measuring.

In this chapter, we looked at four types of survey battery achievement tests: the NAEP, which is sponsored by the U.S. Department of Education and offers a “National Report Card” on how schools are doing in educating our stu- dents. Next, we looked at three achievement tests often used by states in assessing how their students are doing: the Stanford Achievement Test, Iowa Test of Basic Skills (ITBS), and the Metropolitan Achievement Test. We noted that these tests tend to have very good validity and reliability and in recent years have become partic- ularly important in measuring progress in meeting states’ standards of learning and guidelines as set by No Child Left Behind (NCLB). In the age of high-stakes testing and NCLB, these tests have shown to be good at identifying the individual strengths and weaknesses of students and how well individual teachers, schools, and school sys- tems are doing at teaching content areas.

Next, we noted that diagnostic testing has become increasingly important in the identifica- tion and diagnosis of learning disabilities. This is partly due to the passage of PL94-142 and the Individual with Disabilities Education Act

(IDEA), which states that all students with a learning disability must be given accommodations for their disability within the least restrictive envi- ronment. We noted that Individualized Education Plans (IEPs) address any accommodations that need to be made to help the student learn.

In discussing the measurement of learning dis- abilities, we highlighted five diagnostic achieve- ment tests. First, we discussed the Wide Range Achievement Test 4 (WRAT4), which looks at whether an individual has learned the basic codes for reading, spelling, and arithmetic. Next, we discussed the Wechsler Individual Achieve- ment Test (WIAT-III) and the Peabody Individual Achievement Test (PIAT), both of which provide academic screening for pre-K–12 or K–12 stu- dents in broad areas. The Woodcock-Johnson®

III, we noted, is as a broad-based screening tool that can be used for individuals ages 2 to 90, although it is generally targeted for children around the age of 10, and the KeyMath3, offers scrutiny of a child’s math ability when he or she is suspected of having a math disability. As with the survey battery tests, these tests tend to have good validity and reliability.

The final area of achievement testing we looked at was readiness testing. Although the validity information on readiness testing is weak, these tests are sometimes helpful in determining whether a child is ready to begin kindergarten or first grade. Four readiness tests we examined included the Kindergarten Readiness Test (KRT), a second test also called the Kindergarten Readi- ness Test (KRT), the Metropolitan Readiness Test, and the Gesell Developmental Observation. The first three of these tests measure attributes somewhat related to traditional cognitive ability (e.g., language, numerical ability, attention) while the Gesell assesses personal and social skills, neurological and motor growth, language devel- opment, and overall adaptive behavior, or the ability of the child to adapt to new situations.

As the chapter continued, we reviewed a number of cognitive ability tests, starting with the Otis-Lennon School Ability Test 8 (OLSAT 8).

184 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

We noted that this is one of the more common cognitive ability tests and that it assesses students’ abstract thinking and reasoning skills via verbal and nonverbal sections. We also highlighted the Cognitive Ability Test (CogAT), which measures cognitive skills for children from kindergarten to 12th grade. Somewhat based on Vernon’s and Cattell’s theories of intelligence, this test provides scores in three areas: verbal, quantitative, and nonverbal reasoning abilities. We noted that cog- nitive ability tests such as these are particularly helpful in identifying students who have potential to do well in school but are not succeeding due to such things as learning disabilities, motivation, problems at home, problems at school, or self- esteem issues. We warned that tests such as these should not be confused with individualized intelligence tests.

Next, we examined cognitive ability tests that are used to predict achievement in college and graduate school. For instance, we identified the ACT and the SAT as two tests that do a fairly good job at predicting how well an individual will do in college. We pointed out that they have about the same level of predictive accuracy as high school GPA. On the graduate level, we iden- tified the Graduate Record Exam (GRE General Test and Subject Tests) and the Miller Analogies Test (MAT), all of which predict grades fair to

moderately well. We also noted that the Law School Admission Test (LSAT) and the Medical College Admission Test (MCAT) both are fair at predicting achievement in law school or medical school.

As the chapter neared its conclusion, we iden- tified some of the important roles played by school counselors, school social workers, school psychologists, and learning disabilities specialists in the schools relative to the assessment of educa- tional ability. We also highlighted the importance of private practice clinicians knowing about such types of assessment.

The chapter concluded with some final thoughts about the assessment of educational ability. We noted that such tests have been attacked for forcing teachers to teach to the test, labeling children, hindering progress of minority children, and fostering competitiveness and peer pressure. On the other hand, they have been praised because they can identify which children, classrooms, schools, and school systems are performing poorly; signal the pres- ence of learning disabilities; allow children to be accurately placed in their grade levels; and help students identify what areas they need to focus on. We stressed that probably both the criticism and praises of educational ability testing hold some truth.

CHAPTER REVIEW 1. Discuss applications of tests for assessing

educational ability. 2. Define the following types of tests of educa-

tional ability: survey battery tests, readiness tests, diagnostic tests, and cognitive ability tests.

3. Discuss some of the benefits and drawbacks of No Child Left Behind in terms of the “high- stakes testing” atmosphere it tends to promote.

4. From your reading in the chapter, identify two or three survey battery tests and discuss how they might be applied.

5. Relative to survey battery achievement tests, what are some of the uses of classroom, school, and school system profile reports?

6. Discuss the relevance of diagnostic testing to PL94-142 and to the IDEA.

7. Identify and describe two or three diagnostic tests discussed in this chapter.

8. Compare and contrast a readiness test such as the Metropolitan with one like the Gesell.

9. Identify one or more cognitive ability tests that are typically used in the schools and discuss their applications. In particular, how are they important for the identification of learning problems?

10. Discuss two or more types of cognitive ability tests that are typically used for admission into college or graduate school.

CHAPTER 8 Assessment of Educational Ability 185

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

11. Discuss the role of school counselors, school psychologists, learning disability specialists, school social workers, and private practice

clinicians in the assessment of educational ability.

REFERENCES ACT. (2007). ACT technical manual. Retrieved from

http://www.act.org/aap/pdf/ACT_Technical_Manual .pdf

ACT. (2012). ACT profile report—National: Graduate Class of 2012. Retrieved from http://www.act.org /newsroom/data/2012/pdf/profile/National2012.pdf

Anderhalter, O. F., & Perney, J. (2006). Kindergarten Readiness Test. Bensenville, IL: Scholastic Testing Service.

Association of Medical Colleges (AAMC). (2005). MCAT interpretive manual: A guide for understand- ing and using MCAT scores in admissions decisions. Retrieved from https://camcom.ngu.edu/Science /ScienceClub/Shared%20Documents/MCAT% 20Interpretive%20Manual.pdf

Association of Medical Colleges (AAMC). (2012). MCAT essentials. Retrieved from https://www.aamc .org/students/download/63060/data/mcatessentials .pdf

Association of Medical Colleges (AAMC). (1995–2013). Annotated bibliography of MCAT research. Retrieved from https://www.aamc.org/students/applying/mcat /admissionsadvisors/research/bibliography/85382 /mcat_bibliography.html

Baker, S. R., Robinson, J. E., Danner, M. J. E., & Neukrug, E. (2001). Community social disorganiza- tion theory applied to adolescent academic achieve- ment (Report No. UD034167). (ERIC Document Reproduction Service No. ED453301).

Beck, M. (1995). Review of the Kindergarten Readiness Test. In J. C. Conoley, & J. C. Impara (Eds.), The twelfth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Bradley, R. (1985). Review of the Gesell Readiness Test. In J. V. Mitchell, Jr. (Ed.), The ninth mental measurements yearbook (pp. 609–610). Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Burton, N. W., & Wang, M. (2005). Predicting long- term success in graduate school: A collaborative

validity study. Princeton, NJ: Educational Testing Service.

Callahan, C. A., Hojat, M., Veloski, J., Erdmann, J. B., & Gonnella, J. S. (2010). The predictive validity of three versions of the MCAT in relation to performance in medical school, residency, and licensing examinations: A longitudinal study of 36 classes of Jefferson Medical College. Academic Medicine, 85, 980–987. doi:10.1097/ACM.0b0 13e3181cece3d

Carney, R. N. (2005). Review of the Stanford Achieve- ment Test, Tenth edition. In R. A. Spies & B. S. Plake (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Cizek, G. J. (2003). Review of the Woodcock- Johnson® III. In B. S. Plake, J. C. Impara, & R. A. Spies (Eds.), The fifteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measure- ments Yearbook database.

College Board. (2012). Test characteristics of the SAT®: Reliability, difficult levels, completion rates: January 2011-December 2011. Retrieved from http://media.collegeboard.com/digitalServices/pdf /research/Test-Characteristics-of%20-SAT-2012.pdf

College Board. (2013a). Understanding your scores. Retrieved from http://sat.collegeboard.org/scores /understanding-sat-scores

College Board. (2013b). About the SAT. Retrieved from http://press.collegeboard.org/sat/about-the-sat

Cross, L. (2001). Review of the Peabody Individual Achievement Test—revised [1998 normative update]. In B. S. Plake, & J. C. Impara (Eds.), The fourteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

DiPerna, J. C. (2005). Review of the Cognitive Abilities Test, form 6. In R. A. Spies, & B. S. Plake (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements.

186 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Retrieved from Mental Measurements Yearbook database.

Educational Testing Service (ETS). (2007). GRE: Guide to the use of scores (2007–08). Retrieved from http:// www.ets.org/Media/Tests/GRE/pdf/994994.pdf

Educational Testing Service (ETS). (2013a). Verbal reasoning, quantitative reasoning, and analytical writing interpretive data used on score reports. Retrieved from http://www.ets.org/s/gre/pdf/gre_ guide_table1a.pdf

Educational Testing Service (ETS). (2013b). GRE: 2013-2014 interpreting your GRE scores. Retrieved from http://www.ets.org/s/gre/pdf/gre_interpreting_ scores.pdf

Educational Testing Service (ETS). (2013–2014). GRE: Guide to the use of scores. Retrieved from http:// www.ets.org/s/gre/pdf/gre_guide.pdf

Engelhard, G. (2007). Review of the Iowa Tests of basic skills: Forms A and B. In K. F. Geisinger, R. A. Spies, J. F. Carlson, & B. S. Plake (Eds.), The seventeenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Fager, J. J. (2001). Review of the Peabody Individual Achievement Test—revised [1998 normative update]. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Federal Register. (1977) Federal Register. (1977). Reg- ulation implementing Education for All Handi- capped Children Act of 1975 (PL94-142). Federal Register, 42(163), 42474–42518.

Gesell Institute. (2011). GDO-R review kit. Retrieved from http://www.gesellinstitute.org/pdf/2011Spring- GDO-R_ReviewKit.pdf

Gesell Institute of Child Development. (2012). Our his- tory. Retrieved from http://www.gesellinstitute.org /history.html

Graham, R., Lane, S., & Moore, D. (2010). Review of the KeyMath-3 Diagnostic Assessment. In R. S. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Harcourt Assessment. (2004). Stanford Achievement Test Series, tenth edition: Technical data report. San Antonio, TX: Author.

Harris, R. C. (2007). Motivation and school readiness: What is missing from current assessments of

preschooler’s readiness for kindergarten? NHSA Dia- log, 10, 151–163. doi:10.1080/15240750701741645

Harwell, M. (2005). Review of the Metropolitan Achievement Tests, eighth edition. In R. A. Spies & B. S. Plake (Eds.), The sixteenth mental measure- ments yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Johnson, K. M. (2010). Review of the Kindergarten Readiness Test. In R. S. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental mea- surements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Kamphaus. (2001). Review of the Metropolitan Readiness Test, sixth edition. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measure- ments yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

KeyMath—3 DA publication summary form (2007). Retrieved from http://www.pearsonassessments.com /hai/images/pa/products/keymath3_da/km3-da-pub- summary.pdf

Kobrin, J. L., Patterson, B. F., Shaw, E. J., Mattern, K. D., & Barbuti, S. (2008). Validity of the SAT for predicting first-year college grade point average (Research Report No. 2008-5). Retrieved from http://professionals.collegeboard.com/profdownload /Validity_of_the_SAT_for_Predicting_First_Year_ College_Grade_Point_Average.pdf

Kuncel, N. R., & Hezlett, S. A. (2007). Standardized tests predict graduate students’ success. Science, 315, 1080–1081. doi:10.1126/science.1136618

Lane, S. (2007). Review of the Iowa Tests of basic skills: Forms A and B. In K. F. Geisinger, R. A. Spies, J. F. Carlson, & B. S. Plake (Eds.), The seventeenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Larson, S. L., & Vitali, G. J. (1988). Kindergarten Readiness Test. East Aurora, NY: Slosson Educa- tional Publications.

Law School Admission Council (LSAC). (n.d.). LSAT scores as predictors of law school performance. Retrieved from http://www.lsac.org/jd/pdfs/LSAT- Score-Predictors-of-Performance.pdf

Law School Admission Council (LSAC). (2013). The LSAT. Retrieved from http://www.lsac.org/jd/lsat /about-the-lsat.asp

CHAPTER 8 Assessment of Educational Ability 187

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Lloyd, S. C. (2008). Stat of the week: Sanctions and low-performing schools. Retrieved from http://www .edweek.org/rc/articles/2008/03/04/sow0304.h27.html

Lohman, D., & Hagen, E. (2002). CogAT form 6 research handbook. Itasca, IL: Riverside Publishing.

Lukin, L. (2005). Review of the Metropolitan Achieve- ment Tests, eighth edition. In R. A. Spies & B. S. Plake (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Marjanovic, U. L., Kranjc, S., Fekonja, U., & Bajc, K. (2008). The effect of preschool on children’s school readiness. Early Child Development and Care, 178, 569–588. doi:10.1080/03004430600851280

Mattern, K. D., Patterson, B. F., Shaw, E. J., Kobrin, J. L., & Barbuti, S. (2008). Differential validity and predic- tion of the SAT (Research Report No. 2008-4). Retrieved from http://professionals.collegeboard.com /profdownload/Differential_Validity_and_Prediction_ of_the_SAT.pdf

Meagher, D. (2008). Miller Analogies Test: Predictive validity study. Retrieved from http://pearsonassess .com/NR/rdonlyres/423607BB-F273-467C-8232-0E26 13093D23/0/MAT_Whitepaper.pdf

Miller Analogies Test (MAT). (2013). Candidate infor- mation booklet. Retrieved from https://www.pearson assessments.com/hai/Images/dotCom/milleranalogies /pdfs/MAT2011CIB_FNL.pdf

Morse, D. T. (2005). Review of the Stanford Achieve- ment Test, tenth edition. In R. A. Spies & B. S. Plake (Eds.), The sixteenth mental measurements year- book. Lincoln, NE: Buros Institute of Mental Mea- surements. Retrieved from Mental Measurements Yearbook database.

Morse, D. T. (2010). Review of the Otis-Lennon School Ability Test. In R. S. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental mea- surements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

National Center for Education Statistics. (2012). National Assessment of Educational Progess (NAEP). Retrieved from http://nces.ed.gov/nations reportcard/faq.asp#ques26

Novak, C. (2001). Review of the Metropolitan Readi- ness Test, sixth edition. In S. Plake, & J. C. Impara (Eds.), The fourteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measure- ments Yearbook database.

PAR. (2012). Wide Range Achievement Test 4 (WRAT4). Retrieved from http://www4.parinc .com/Products/Product.aspx?ProductID=WRAT4

Pearson. (2012a). WIAT-III: Overview. Retrieved from https://www.pearsonassessments.com/HAIWEB /Cultures/en-us/Productdetail.htm?Pid=015-8984-609

Pearson. (2012b). KeyMath3 Diagnostic Assessment. Retrieved from http://psychcorp.pearsonassessments .com/HAIWEB/Cultures/en-us/Productdetail.htm?Pid =PAaKeymath3

Pearson. (2012c). Otis-Lennon School Ability Test, eighth edition. Retrieved from http://education. pearson assessments.com/HAIWEB/Cultures/en-us /Productdetail.htm?Pid=OLSAT

Riverside Publishing. (2010). Cognitive Abilities Test, Form 6. Retrieved from http://www.riverpub.com /products/cogAt/index.html

Rogers, B. G. (2005). Review of the Cognitive Abilities Test, Form 6. In R. A. Spies & B. S. Plake (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measure- ments. Retrieved from Mental Measurements Year- book database.

Sacket, P. R., Kuncel, N. R., Beatty, A. S., Rigdon, J. L., Shen, W., & Kiger, T. B. (2012). The role of socioeconomic status in SAT-grade relationships and in college admissions decisions. Psychological Science, 23, 1000–1007. doi:10.1177/09567976124 38732

Slosson Educational Publications. (n.d.). The Kinder- garten Readiness Test. Retrieved from http://www .slosson.com/onlinecatalogstore_i1003393.html

Straus, V. (2012, September 24). “How ACT overtook SAT as the top college entrance exam.” The Washington Post: Post Local. Retrieved from http://www.washingtonpost.com/blogs/answer-sheet /post/how-act-overtook-sat-as-the-top-college-entrance- exam/2012/09/24/d56df11c-0674-11e2-afff-d6c7f20a 83bf_blog.html

Sutton, R., & Knight, C. (1995). Review of the Kinder- garten Readiness Test. In J. C. Conoley, & J. C. Impara (Eds.), The twelfth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measure- ments Yearbook database.

Swerdlike, M. E., & Hoff, K. E. (2010). Review of the Kindergarten Readiness Test. In R. S. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

188 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

U.S. Department of Education. (2005). Stronger account- ability: The facts about … measuring progress. Retrieved from http://www.ed.gov/nclb/accountability /ayp/testing.html

U.S. Department of Education. (2011). No Child Left Behind legislation and policies. Retrieved from http://www2.ed.gov/policy/elsec/guid/states/index .html#nclb

Waters, E. (1985). Review of the Gesell Readiness Test. In J. V. Mitchell, Jr. (Ed.), The ninth mental mea- surements yearbook. Lincoln, NE: Buros Institute

of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Willse, J. T. (2010). Review of the Wechsler Individ- ual Achievement Test—third edition. In R. S. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Wilkinson, G. S., & Robertson, G. J. (2006). Wide Range Achievement Test professional manual. Lutz, Fl: PAR, Inc.

CHAPTER 8 Assessment of Educational Ability 189

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R9 Intellectual and CognitiveFunctioning: Intelligence Testing and Neuropsychological Assessment

John is a 9-year-old, fifth-grade student. His teachers reported that he was “a wiz at math and was the first student in the class to memorize all the state capitols.” On his fourth-grade report card, John received all As. This past year, John was riding his bike after school to a friend’s house. He was not wearing his helmet and was struck by an oncoming vehicle. John was rushed to the emergency room and sustained what was then considered to be a “minor head injury.” He was released a few days later from the hospital. When John returned to school, his teachers began noticing that John took longer to complete tasks, had difficulty staying on task, and had problems with his short-term memory. His grades dropped from As to mostly Bs and Cs. John’s parents were concerned about these changes and scheduled an appointment for John to be tested by a local clinical neuropsychologist.

This chapter examines one of the most interesting and perplexing aspects of what it’s like to be human—how we think and how our cognitive processes function— in other words, how our brains work! Within the chapter, we look at models of intelligence and how intelligence is measured, and then we examine the interesting world of neuropsychology, which looks at changes in brain functioning over time. For example, with John, we might want to administer an intelligence test and a neuropsychological battery. The intelligence test, which is sometimes included as part of a complete psychoeducational and/or neuropsychological workup, helps us see how John is doing compared to his peers and might uncover some cognitive impairments. A neuropsychological battery looks at a broad range of cognitive functioning and helps us make a determination about possible changes in

190

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

John’s brain functioning over time. If we did test John in this manner, we might find that the intelligence scores reflect average intellectual ability while the neuro- psychological battery shows a number of cognitive impairments that are reflective of changes since the bicycle accident. This case exemplifies how the neuropsycho- logical battery is more concerned with comparing one’s premorbid levels (prior to the accident) to one’s postmorbid levels (after the accident) rather than comparing one’s level of functioning to a normative group as is done in intelligence testing.

The chapter begins with a brief history and definition of intelligence testing and then provides an overview of some models of intelligence. Next, we examine some of the better-known individual intelligence tests, including measures of non- verbal intelligence. We then examine the field of neuropsychology and explain neuropsychological assessment. In this section, we provide a brief history and defi- nition of neuropsychology, discuss the domains assessed in neuropsychological assessment, and introduce two common methods used in neuropsychological assessment: a fixed battery and flexible battery approach. As the chapter nears its conclusion, we look at some of the roles that helpers have in the assessment of intellectual and cognitive functioning and conclude with some final thoughts about the assessment of these important domains.

A BRIEF HISTORY OF INTELLIGENCE TESTING As noted in Chapter 1, the emergence of ability testing can be traced back thousands of years and can be found in biblical passages, Chinese history, and ancient Greek writings. Centuries later, the first individual intelligence test was pioneered in France by Alfred Binet (1857–1911) and his colleague, Theophile Simon (Jolly, 2008; Kerr, 2008). Shortly thereafter, at Stanford University, Lewis Terman (1877–1956) gath- ered and analyzed normative data and made revisions to Binet’s measure. Terman’s alterations were extensive and the test later was renamed, the Stanford-Binet. This revised version (and many subsequent revisions since that time) has stood the test of time and continues to be used today in clinical and educational settings.

Over the years, new models of intelligence have emerged, and concurrently, new intelligence tests have been developed. The following sections explore current defini- tions of intelligence testing, introduce a variety of models that attempt to explain intelligence, and describe some of the more commonly used intelligence tests.

DEFINING INTELLIGENCE TESTING As mentioned in Chapter 1, intelligence testing is a subset of intellectual and cogni- tive functioning and assesses a broad range of cognitive capabilities that generally results in an “IQ” score (see Figure 1.3 and Box 1.5). Like special aptitude testing, multiple aptitude testing, and cognitive ability testing, intelligence testing measures aptitude, or what one is capable of doing. Special and multiple aptitude testing will be examined in Chapter 10 as types of tests sometimes used in occupational and career counseling. Cognitive ability tests were examined in Chapter 8 as a type of educational ability testing, and although intelligence tests are sometimes

CHAPTER 9 Intellectual and Cognitive Functioning 191

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

used to assess aspects of educational ability, they tend to have much broader appli- cations. For instance, intelligence tests are used:

1. to assist in determining giftedness; 2. to assess intellectual disabilities; 3. to identify certain types of learning disabilities; 4. to assess intellectual ability following an accident, the onset of dementia, sub-

stance abuse, disease processes, and trauma to the brain; 5. as part of the admissions process to certain private schools; and 6. as part of a personality assessment battery to aid in understanding the whole

person.

But before we examine intelligence tests, let’s take a look at some models of intelligence that have been used as templates for the development of some of the more frequently used intelligence tests.

MODELS OF INTELLIGENCE Some models of intelligence have been around for 100 years, and others are new. Many of them are complicated and all of them have contributed to the develop- ment of current day intelligence tests. What follows is a brief overview of some of the more prevalent models, including Spearman’s two-factor approach, Thurstone’s multifactor approach, Vernon’s hierarchical model of intelligence, Guilford’s multi- factor/multidimensional model, Cattell’s fluid and crystal intelligence, Piaget’s cog- nitive development theory, Gardner’s theory of multiple intelligences, Sternberg’s triarchic theory of successful intelligence, and the Cattell-Horn-Carroll (CHC) inte- grated model of intelligence.

Spearman’s Two-Factor Approach When Alfred Binet created the first widely used intelligence test, he developed a number of subtests that assessed a range of what he considered to be intellectual tasks. Then, he determined the average scores that individuals, at different age groups, obtained on these tasks. Consequently, when an individual was assessed on the Binet scale, the individual’s score could be compared to the average score of individuals at different age levels. Asserting that such a test was a hodgepodge or “promiscuous pooling” of factors, Charles Edward Spearman (1863–1945) was critical of Binet and others (Spearman, 1970, p. 71). Spearman, in other words, felt that Binet had lumped a number of different factors together in a spurious fashion.

Spearman (1970) believed in a two-factor approach to intelligence that included a general factor (g) and a specific factor (s), with the “weight” of g vary- ing as a function of what was being measured. For example, he stated that the “tal- ent for classics” (e.g., understanding the ancient worlds of Rome and Greece) had a ratio of g to s of 15 to 1; that is, general intelligence is much more significant than any specific ability in understanding the ancient worlds. Conversely, he purported that the ratio of general intelligence (g) to specific talent for music (s) was a ratio of 1 to 4, meaning that having music ability was much more significant in having

Spearman Known for his g and s factors of intelligence

192 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

talent for music than was general intelligence (Spearman, 1970). Although Spearman’s theory was one of the earlier models of intelligence, many still adhere to the concept that there is a g factor that mediates general intelligence and s factors that speaks to a variety of specific talents.

Thurstone’s Multifactor Approach Using multiple-factor analysis, Thurstone developed a model that included seven pri- mary factors or mental abilities. Although Thurstone’s research did not substantiate Spearman’s general factor (g), he did not rule out the possibility that it existed since there appeared to be some commonality among the seven factors (Thurstone, 1938). The seven primary mental abilities he recognized were verbal meaning, number ability, word fluency, perception speed, spatial ability, reasoning, and memory.

Vernon’s Hierarchical Model of Intelligence Perhaps one of the greatest and most widely adopted models of intelligence has been Vernon’s hierarchical model (Vernon, 1961). Philip Vernon believed that sub- components of intelligence could be added in a hierarchical manner to obtain a cumulative (g) factor score. His hierarchical model comprised four levels, with fac- tors from each lower level contributing to the next level on the hierarchy. Vernon’s top level was similar to Spearman’s general factor (g) and was considered to have the most variance of any of the factors, while the second level had two major fac- tors: v:ed, which stands for verbal and educational abilities, and k:m, which repre- sents mechanical-spatial-practical abilities. The third level is composed of what was called minor group factors, while the fourth level is made of what was identified as specific factors, as illustrated in Figure 9.1.

Guilford’s Multifactor/MultiDimensional Model Guilford (1967) originally developed a model of intelligence with 120 factors. As if that were not enough, he later expanded it to 180 factors (Guilford, 1988). His three-dimensional model can be represented as a cube and involves three kinds of cognitive ability: operations, or the general intellectual processes we use in under- standing; content, or how we apply our intellectual process; and the products, or how we apply our operations to our content (see Figure 9.2). Different mental abil- ities will require different combinations of processes, contents, and products. All of the possible combinations are combined to create the (6 � 6 � 5) ¼ 180 factors. Guilford’s multifactor model provides a broad view of intelligence (Guilford, 1967); however, his model is sometimes considered too unwieldy to implement and has not significantly influenced the testing community.

Cattell’s Fluid and Crystal Intelligence After attempting to remove cultural bias from intelligence tests, Raymond Cattell observed that as information based on learning was removed from such tests (the portion most affected by cultural influences), the raw or unlearned abilities provided

Thurstone Believed in seven dif- ferent mental abilities

Vernon Developed hierarchy approach that is still used by most tests today

Guilford Developed 180 factors in his model shaped as a cube

Cattel Differentiated fluid (innate) from crystal- lized (learned) intelligence

CHAPTER 9 Intellectual and Cognitive Functioning 193

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Major group factors

Minor group factors

Specific factors

Verbal-Educational (v:ed)

Verbal ability

Numerical ability

Vocabulary

Reading comprehension

Spelling

Logical reasoning

Arithmetic

Matrix reasoning

Mechanical ability

Spatial ability

Practical ability

Mechanical knowledge

Manipulation & dexterity

Object assembly

Picture arrangement

Block design

Symbol search

Perceptual

Memory

Clerical

Practical (k:m)

General (g)

FIGURE 9.1 | Diagram Illustrating Vernon’s Hierarchical Structure of Abilities

CONTENT

Visual

Auditory

Semantic

Behavioral

PRODUCTS

Units

Classes

Relations

Systems

Transformations

Implications OPERATIONS

Evaluation

Convergent production

Divergent production

Memory retention

Memory recording

Cognition

Symbolic

FIGURE 9.2 | Guilford’s Multifactor Model of Intelligence Source: Guilford, J. (1988). Some changes in the structure of the intellect model. Educational and Psychological Measurement, 48, 1–4, p. 3.

© Ce ng ag e Le ar ni ng

194 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

a different score (Cattell, 1971). He then considered the possibility that two “general factors” made up intelligence: “fluid” (gf) intelligence, or that culture-free portion of intelligence that is inborn and unaffected by new learning, and “crystallized” intelli- gence (gc), which is acquired as we learn and is affected by our experiences, school- ing, culture, and motivation (Cattell, 1979). He eventually estimated that heritability variance within families for fluid intelligence was about 0.92, which basically means if your parents have it, you are likely to have it (Cattell, 1980). Abilities such as memory and spatial capability are aspects of fluid intelligence.

As one might expect, crystallized intelligence will generally increase with age, while many research studies have found that fluid intelligence tends to decline slightly as we get older (see Box 9.1). Therefore, many theorists believe that overall intelligence (g) is maintained evenly across the lifespan (see Figure 9.3). As we look at specific intelligence tests later in the chapter, see if you can identify how Cattell’s ideas have influenced their development.

Piaget’s Cognitive Development Theory Piaget (1950) approached intelligence from a developmental perspective rather than a factors approach. Spending years observing how children’s cognitions were

BOX 9.1 Example of Fluid and Crystallized Intelligence

At our university, we have a good mix of traditional aged as well as older adult (35+-year-old) students. It is interesting to observe students when we give exams. The last five or six students left taking an exam will invariably be almost all of the older adult students. Sure, some of the reason is that they may be more careful, but the decrease in fluid intelli- gence could also be a major factor. Are the older

students’ scores lower? Absolutely not. As a matter fact they are at least equal to the younger students, if not higher. Why is this? It might be due to the fact that older students may be experiencing a decrease in fluid intelligence, but their crystallized intelligence is making up for the difference!

—Charlie Fawcett

Fluid intelligence (gf)

In te

ll ig

e n c e

Overall intelligence

Lifespan (time)

Crystallized intelligence (gc)

FIGURE 9.3 | Fluid and Crystallized Intelligence Across the Lifespan

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 9 Intellectual and Cognitive Functioning 195

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

shaped as they grew, he developed the now familiar four stages of cognitive devel- opment: sensorimotor, preoperational, concrete operational, and formal opera- tional. Piaget (1950) believed that cognitive development is adaptive; that is, as new information from our environment is presented, we are innately programmed to take it in and make sense of it in some manner to maintain a sense of order and equilibrium in our lives.

Piaget believed that we adapt our mental structures to maintain equilibrium through two methods: assimilation and accommodation. Assimilation is incorpo- rating new stimuli or information into existing cognitive structures. On the other hand, accommodation is creating new cognitive structures and/or behaviors based on new stimuli. For example, a parent might teach a young child that the word hot means that one should stay away from certain items (e.g., a stove, iron, etc.) because touching those items can result in something bad happening. The child then learns that stoves are “hot” and should be avoided. In addition, every time the child is near a hot object, he or she knows that this object should be avoided (e.g., match, coal, frying pan, etc.). The new object has been assimilated as some- thing to be avoided.

As the child grows older, he or she comes to realize that not all hot items are hot all the time, and the child accommodates to this new information. For instance, as the child comes to understand the stove is not hot all the time, he or she creates a new meaning for the concept “stove,” which now is seen as an object that can be hot or cold. Consequently, the child’s behavior around a stove will change as he or she has accommodated this new meaning into his or her mental framework.

Consider how assimilation and accommodation are also important in the learning of important concepts in school, such as the shift between addition and multiplication (multiplication being a type of advanced addition). Also, consider how you have assimilated or accommodated information from this text into your existing structures of understanding! Although Piaget’s understanding of cognitive development does not speak directly to the amount of learning taking place, it does highlight the process of learning—a critical concept for teachers and helpers to understand.

Gardner’s Theory of Multiple Intelligences Gardner (1999, 2011), who is vehemently opposed to most other models of intelli- gences and the manner in which intelligence tends to be measured, refers to the pre- dominant notion of intelligence as the “dipstick theory” of the mind; that is, it holds that there is a specific amount or level of intelligence in the brain, and if you could place a dipstick in the brain and pull it out, you should be able to accurately read how smart a person is (Gardner, 1996). In contrast to this approach, he believes that intelligence is much too vast and complex to be measured accurately by our current methods.

Based on his research of brain-damaged individuals, as well as literature in the areas of the brain, evolution, genetics, psychology, and anthropology, Gardner developed his theory of multiple intelligences, which asserts there are eight or nine intelligences, and with more research, others might even be found. Following are the nine identified intelligences, although research on the ninth type of intelligence,

Piaget Cognitive develop- mental model high- lights assimilation and accommodation in learning

Gardner Theory of multiple intelligences is novel but difficult to apply

196 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

existential intelligence, has not clearly established its validity at this point (Gardner, 2003).

1. Verbal-Linguistic Intelligence: well-developed verbal skills and sensitivity to the sounds, meanings, and rhythms of words

2. Mathematical-Logical Intelligence: ability to think conceptually and abstractly, and capacity to discern logical or numerical patterns

3. Musical Intelligence: ability to produce and appreciate rhythm, pitch, and timbre

4. Visual-Spatial Intelligence: capacity to think in images and pictures, to visualize accurately and abstractly

5. Bodily-Kinesthetic Intelligence: ability to control one’s body movements and to handle objects skillfully

6. Interpersonal Intelligence: capacity to detect and respond appropriately to the moods, motivations, and desires of others

7. Intrapersonal Intelligence: capacity to be self-aware and in tune with inner feelings, values, beliefs, and thinking processes

8. Naturalist Intelligence: ability to recognize and categorize plants, animals, and other objects in nature

9. Existential Intelligence: sensitivity and capacity to tackle deep questions about human existence, such as the meaning of life, why do we die, and how did we get here. (Educational Broadcasting Corporation, 2004, section: “What is …”)

Gardner’s understanding of intelligence is revolutionary and not yet main- stream. However, agreement with this theory appears to be growing in academic and nonacademic settings. Not only are his identified categories novel, but his understanding of how intelligence manifests itself is different, as noted in the fol- lowing list:

1. All human beings possess a certain amount of all of the intelligences. 2. All humans have different profiles or amounts of the multiple intelligences

(even identical twins!). 3. Intelligences are manifested by the way a person carries out a task in relation-

ship to his or her goals. 4. Intelligences can work independently or together, and each is located in a dis-

tinct part of the brain (Gardner, 2003, 2011).

Sternberg’s Triarchic Theory of Successful Intelligence Like Gardner, Sternberg (2009, 2012) has a novel view of intelligence that is based on the individual’s capacity at using one’s abilities and talents, navigate one’s envi- ronment, and adapt to new situations. He states that

successful intelligence is: (1) the use of an integrated set of abilities needed to attain success in life, however an individual defines it, within his or her sociocultural context. People are successfully intelligent by virtue of (2) recognizing their strengths and mak- ing the most of them, at the same time that they recognize their weaknesses and find ways to correct or compensate for them. Successfully intelligent people (3) adapt to, shape, and select environments through (4) finding a balance in their use of analytical, creative, and practical abilities. (Sternberg, 2012, pp. 156–157)

Triarchic theory Componential, experi- ential, and contextual subtheories

CHAPTER 9 Intellectual and Cognitive Functioning 197

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Sternberg’s model is composed of what he calls three subtheories (Sternberg, 2012), which he calls the componential, experiential, and contextual (see Figure 9.4).

In Sternberg’s theory, the componential subtheory, sometimes called analytical facet, is related to the more traditional types of intelligence and has to do with higher-order thinking (metacomponents), how one acts on our higher-order think- ing (performance), and the strategies one uses to store and use knowledge (knowl- edge acquisition).

The experiential subtheory, sometimes called the creative facet, is focused on the ability to deal with novel situations and one’s adeptness at automatically attending to tasks so that one can focus on other tasks (e.g., multitasking). Creative thinkers can focus on novel situations, attend to them, and eventually deal with similar situations in the future in an automatic manner as they no longer are novel.

Finally, Sternberg’s contexutal subtheory, sometimes called the practical facet, has to do with the ability to adapt to the ever-changing envirnoment, being able to shape one’s environment to meet one’s goals, and selecting a new envionrment (“renunciating the old environment”) if adaptation and shaping is not successful.

Although Sternberg believes that his theory is universal, how it is applied can vary in different cultures. For instance, what is seen as important higher-order think- ing in one culture, may be different in another culture, and the kinds of novel situa- tions faced in one’s environment will vary from culture to culture (see Exercise 9.1).

1. Metacomponents

2. Performance

3. Knowledge Acquisition

1. Novelty Ability

2. Automation Ability

1. Adaptation

2. Shaping

3. Selection

Triarchic Theory

Componential Subtheory

Contextual Subtheory

Experiential Subtheory

FIGURE 9.4 | Sternberg’s Triarchic Theory of Successful Intelligence

Exercise 9.1 Using Your Intelligence to Successfully Navigate Through School

In class, discuss the kinds of componential intelligence, experiential intelligence, and contextual intelligence that is needed to successfully go through your major or

graduate program? How does each type of intelligence impact one another?

© Ce ng ag e Le ar ni ng

20 15

© Ce ng ag e Le ar ni ng

20 15

198 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Cattell-Horn-Carroll (CHC) Integrated Model of Intelligence Like those who came before them, Horn and Cattell (1966) examined the idea of multiple abilities. Using factor analysis, they provided evidence that intelligence could be understood in reference to six factors that included fluid intelligence, crys- tallized intelligence, general visualization, general speediness, facility in the use of concept labels, and carefulness. This work was later expanded by (Carroll, 1993). More recently, Horn and Blankson (2012) theorized eight or nine factors and Carroll (2005) and Schneider & McGrew (2012) reported support for three addi- tional abilities (i.e., kinesthetic ability, olfactory ability, and tactile ability), which, they argued, helps to explain Gardner’s (2003, 2011) multiple intelligences and Sternberg’s (2009, 2012) triarchic theory of successful intelligence.

Using Cattell, Horn, and Carroll’s work, today, an integrated model that con- solidates their research has been developed (McGrew, 2009; Schneider & McGrew, 2012). The Cattell-Horn-Carroll (CHC) integrated model includes 16 broad ability factors, 6 of which are tentative (see Table 9.1). In addition, this approach suggests over 70 narrow abilities that tie into the different factors. Finally, although Carroll (1993) suggested that a g factor mediates the various abilities, Cattell and Horn suggested it did not (Horn & Blankson, 2012). Despite this difference, their theo- ries tie together nicely.

Theories of Intelligence Summarized As you can see, the manner in which intelligence is conceptualized varies depending on the theory. Thus, how an intelligence test is constructed will reflect the model one believes is most true to the nature of this construct. Table 9.2 summarizes some of the major points of each theory discussed in this chapter. Consider what you might measure if you were developing an intelligence test based on one or more of these theories.

INTELLIGENCE TESTING It would make sense that the models of intelligence just discussed are the basis for intelligence tests. So, it is not surprising that over the years a number of intelli- gence tests have been devised to measure such factors as general intelligence (g), specific intelligence (s), fluid and crystal intelligence, and other factors tradition- ally seen to be related to intellectual ability. However, one should keep in mind that intelligence tests only measure a portion of the competencies involved with human intelligence. In fact, such tests are best seen as estimates of performance in school, work, and in the broad range of life activities. Although there is likely some innate capacity being measured in intelligence tests, the variability of intelli- gence test scores is impacted by a wide range of family, cultural, and societal fac- tors (Rose, 2006), and to some degree, IQ is a reflection of how well individuals have mastered middle-class facts, concepts, and problem-solving strategies. This is highlighted by the fact that intelligence tests scores are not necessarily fixed, as some persons will exhibit significant increases or decreases in their measured intellectual abilities.

CHC theory Sixteen factors related to mental abilities

CHAPTER 9 Intellectual and Cognitive Functioning 199

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Although a number of intelligence tests are available today, the Stanford-Binet and the three Wechsler scales of intelligence are the most widely used. Thus, we examine these tests as well as another popular intelligence test known as the Kaufman Assessment Battery for Children. We also briefly introduce the concept of nonverbal intelligence tests and introduce the Comprehensive Test of Nonverbal Intelligence, Second Edition (CTONI-2), the Universal Intelligence Test (UNIT), and the Wechsler Nonverbal Scale of Ability (WNV).

TABLE 9.1 CHC Factors Listed Under Broad Domains*

Domain Free

Fluid Reasoning (Gf): “… the deliberate but flexible control of attention to solve novel, ‘on the spot’ problems that cannot be performed by relying exclusively on previously learned habits, schemas, and scripts.”

Memory

Short-Term Memory (Gsm): “… the ability to encode, maintain, and manipulate information in ones immediate awareness.”

Long-Term Storage & Retrieval (Glr): “… the ability to store, consolidate, and retrieve information over periods of time measured in minutes, hours, days, and years.”

General Speed

*Psychomotor Speed (Gps): “… the speed and fluidity with which physical body movements can be made.” Processing Speed (Gs): “… [t]he ability to perform simple repetitive cognitive tasks quickly and fluently.” Reaction and Decision Speed (Gt): “… [t]he speed of making very simple decisions or judgments when items

are presented one at a time.”

Motor

*Kinesthetic Abilities (Gk): “… the ability to detect and process meaningful information in proprioceptive sensations.”

*Psychomotor Abilities (Gp): “… the ability to perform physical body motor movements (e.g., movement of fingers, hands, legs) with precision, coordination, or strength.”

Sensory

*Olfactory Abilities (Go): “… the ability to detect and process meaningful information in odors.” *Tactile Abilities (Gh): “… the ability to detect and process meaningful information in haptic (touch)

sensations.” Visual Processing (Gv): “… the ability to make use of simulated mental imagery (often in conjunction with

currently perceived images) to solve problems.” Auditory Processing (Ga): “… the ability to detect and process meaningful nonverbal information in sound.”

Acquired Knowledge

Reading and Writing (Grw): “… [t]he depth and breadth of knowledge and skills related to written language.” Quantitative Knowledge (Gq): “… the depth and breadth of knowledge related to mathematics.” Comprehension-Knowledge (Gc):” … The depth and breadth of knowledge and skills that are valued by one’s

culture.” *Domain-Specific Knowledge (Gkn): “… the depth, breadth, and mastery of specialized knowledge (knowledge

not all members of a society are expected to have).” (McGrew & Schneider, 2012, pp. 111–137)

* ¼ Tentative Factors Source: McGrew, K. S., & Schneider, K. S. (2012). The Cattell-Horn-Carroll model of intelligence. In D. P. Flanagan & P. L. Harrison (Eds.), Contemporary Intellectual Assessment: Theories, Tests, and Issues (3rd ed., pp. 99–144). New York: Guilford Press.

200 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TABLE 9.2 Summary of Models of Intelligence

Theoretical Model

Number of Factors or Attributes Nature of Factors or Attributes

Spearman’s Two Factor Approach

Two: g and s The g factor mediates general intelligence and s factor mediates specific abilities. Ratio of g to s varies depending on ability.

Thurstone’s Multifactor Approach

Seven Research on multiple factor analysis led to the belief that there are seven primary factors that may be related to g. They include verbal meaning, number ability, word fluency, perception speed, spatial ability, reasoning, and memory.

Vernon’s Hierarchal Model of Intelligence

Four hierarchical levels The four levels include (1) g, the highest level, with the largest source of variance between individuals; (2) major group factors, including verbal-numerical-educational (v:ed) and practical- mechanical-spatial-physical (k:m) ability; (3) minor group factors; and (4) specific factors.

Guilford’s Multifactor/ Multi-dimensional Model

Three dimensional model with 180 factors

Three kinds of cognitive ability: operations, or the 6 processes we use in understanding; contents, or the 6 ways we perform our thinking process; and the product, or the 5 possible ways we apply our operations to our content (6 � 6 � 5 ¼ 180 possible paths).

Cattell’s Fluid and Crystal Intelligence

Two: fluid gf and crys- tallized intelligence gc

Two g factors: fluid intelligence (gf), culturally-free intelligence with which we are born, and crystallized intelligence (gc), acquired as we learn, and affected by experiences, schooling, culture, and motivation.

Piaget’s Cognitive Development Theory

Two: assimilation and accommodation

Process model, not a cognitive gain model. Assimilation is adapting new stimuli or information into existing cognitive structures, and accommodation is creating new cognitive structures and/or behaviors from new stimuli. We learn through these processes.

Gardner’s Theory of Multiple Intelligences

Eight, maybe nine Nontraditional model of eight or nine intelligences that all people have at different levels: verbal-linguistic, mathematical-logical, musical, visual-spatial, bodily-kinesthetic, interpersonal, intraper- sonal, naturalist, and maybe existential

Sternberg’s Triarchic Theory of Successful Intelligence

Three interrelated subtheories having two or three components

Three subtheories include componential (analytical), which has to do with higher-order thinking and how that gets processed; experiential (creative), which has to do with how one deals with novel situations as well as with the ability to do automated tasks; and contextual (practical), which has to do with the ability to adapt to, shape, or select new environments to successfully meet one’s goals.

Cattell-Horn-Carroll (CHC) Integrated Model

10, maybe 16 A total of 16 broad ability factors, 6 of which are tentative, and over 70 associated tasks that may or may not be related to a g factor, which include fluid reasoning (Gf), comprehension-knowledge (Gc), short-term memory (Gsm), visual processing (Gv), auditory proces- sing (Ga), long-term storage and retrieval (Glr), processing speed (Gs), reaction and decision speed (Gt), reading and writing (Grw), quantitative knowledge (Gq), domain-specific knowledge (Gkn), tactile abilities (Gh), kinesthetic abilities (Gk), olfactory abilities (Go), psycho- motor abilities (Gp), and psychomotor speed (Gps).

© Ce ng ag e Le ar ni ng

CHAPTER 9 Intellectual and Cognitive Functioning 201

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Stanford-Binet, Fifth Edition Probably the most well-known intelligence test is the Stanford-Binet, which dates back to the original work of Alfred Binet in 1904. As discussed in Chapter 1, Lewis Terman at Stanford University began improving on Binet’s scale in the early 1900s, and since that time, the test has been periodically revised until its most recent fifth edition in 2003.

The Stanford-Binet, Fifth Edition (SB5) takes between 45 and 75 minutes to administer individually and can be given to individuals from 2 to 85þ years old (Johnson & D’Amato, 2005; Kush, 2005). To administer the test in a reasonable amount of time to individuals of all ages, the Stanford-Binet uses a vocabulary routing test, almost a pretest, to determine where an individual should begin (Houghton Mifflin Harcourt, n.d.a; Roid, 2003). For example, a 35-year-old engi- neer does not need to answer all the questions in the mathematics section that would normally apply to an elementary school student. Instead, using the routing test, the examiner determines the age level at which to begin testing. Next, a basal level is determined, which is the highest point where the examinee is able to get all the questions right on two consecutive age levels. Testing continues until the ceiling level is reached, or the point where the individual misses 75% of the questions on two consecutive age levels. Basal and ceiling levels are important because they help the examinee avoid feeling bored, and they lessen the likelihood that the examinee will experience a sense of failure, which is more likely to happen if the examinee was asked to respond to all questions on the test. Basal and ceiling levels are also found in other kinds of assessment procedures and are an important concept to understand.

The SB5 measures verbal and nonverbal intelligence across five factors: fluid reasoning, knowledge, quantitative reasoning, visual-spatial processing, and work- ing memory (Houghton Mifflin Harcourt, n.d.a). These divisions create 10 subtests (2 domains � 5 factors) (see Figure 9.5). Discrepancies among scores on the subt- ests as well as between scores on the verbal and nonverbal factors can be an indica- tion of a learning disability.

Although previous versions of the Stanford-Binet used a mean of 100 and an SD of 16, the SB5 has now joined most other intelligence tests by using a mean of 100 and an SD of 15 (Johnson & D’Amato, 2005). Subtests use a mean of 10 and an SD of 3, and other standardized scoring methods are available, such as percen- tile ranks, grade equivalents, and a descriptive classification (average, low average, etc.). A nice feature of this latest edition is that the publishers also offer a compati- ble program called the SB5 Scoring Pro™. This Windows-based software program allows the examiner to enter raw scores that are converted to standard scores, and it provides various profiles related to areas that are assessed. Figure 9.6 provides example of the interpretive worksheet of the Stanford-Binet.

The SB5 norm group was based on a stratified random sample of 4,800 people gathered from the 2000 U.S. Census. Each item was examined for bias in gender, ethnicity, regional, and socioeconomic status (SES). Split-half and test-retest reli- abilities for individual subtests by age averaged between 0.66 and 0.93, while full- scale IQ reliability was very high and ranged between 0.97 and 0.98 for all age groups (Roid, 2003). Content validity included professional judgment of items

SB5 Uses routing test, basal and ceiling levels to determine start and stop points; measures verbal and nonverbal intelligence across five factors

202 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

over a seven-year development process and item analysis using advanced statistical methods. Criterion-related validity showed a “substantial” correlation of 0.90 with the previous version of the Stanford-Binet (SB IV). Sound predictive validity was also shown with a variety of special groups such as individuals who are gifted, intellectually disabled, and learning disabled. Correlations with other instruments ranged from 0.53 to 0.84 for the SB5 with the Wechsler scales, Woodcock- Johnson III Tests of Cognitive Abilities, Woodcock-Johnson III Tests of Achievement, and Wechsler Individual Achievement Test (WAIT) (Roid, 2003). Finally, the 217-page technical manual is filled with other advanced statistical analyses that show evidence of reliability and validity.

Wechsler Scales Although the Stanford-Binet may be the best-known intelligence test, the three Wechsler scales of intelligence are the most widely used intelligence tests today. In contrast to the Stanford-Binet, which assesses intelligence of individuals from a broad spectrum of ages, each Wechsler test measures a select age group.

Fluid Reasoning (Object Series/Matrices)

Knowledge (Procedural Knowledge,

Picture Absurdities)

Quantitative Reasoning (Nonverbal Quantitative

Reasoning)

Visual-Spatial Processing (Form Board, Form Patterns)

Working Memory (Delayed Response, Block

Span)

5 Nonverbal Subtests Fluid Reasoning

(Early Reasoning, Verbal Absurdities, Verbal Analogies)

Knowledge (Vocabulary)

Quantitative Reasoning (Verbal Quantitative

Reasoning)

Visual-Spatial Processing (Position and Direction)

Working Memory (Memory for Sentences, Last

Word)

5 Verbal Subtests

Fluid Reasoning

Knowledge

Quantitative Reasoning

Visual-Spatial Processing

Working Memory

5 Factor Indexes

Nonverbal IQ Verbal IQ

Full Scale IQ

FIGURE 9.5 | Organization of the Stanford-Binet, Fifth Edition

Wechsler scales Three different tests for three different age groups

© Ce ng ag e Le ar ni ng

CHAPTER 9 Intellectual and Cognitive Functioning 203

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

FIGURE 9.6 | The SB5 Interpretive Report On the above interpretive report you can see the following (going counter-clockwise): (1) basic demo- graphic information, (2) the routing score to help determine the basal age, (3) verbal and nonverbal subtest raw and scaled scores (M ¼ 10, SD ¼ 3), (4) comparison of verbal IQ, nonverbal IQ, full- scale IQ using DIQ, percentile rank, and confidence intervals (range of true scores), and (5) comparison of verbal IQ, nonverbal IQ, full-scale IQ, and subtest scores using DIQ.

Source: Houghton Mifflin Harcourt. (2003). The Stanford-Binet Intelligence Scales, Fifth Edition. Retrieved from http://www.riverpub.com/products/sb5/securesite/list.html

204

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

For instance, the WPPSI-III (Wechsler Preschool and Primary Scale of Intelligence— Third Edition) assess children between the ages of 2 years, 6 months to 7 years, 3 months, the WISC-IV (Wechsler Intelligence Scale for Children—Fourth Edition) is geared for children between the ages of 6 and 16, and the WAIS-IV (Wechsler Adult Intelligence Scale—Fourth Edition) was developed to assess adults, ages 16 through 90. Due to an overwhelming amount of research that provided evidence for the Cattell-Horn-Carroll (CHC) framework discussed earlier, the WISC-IV and WAIS-IV both underwent substantial revisions that aligned them more closely to this theory while also providing for more user-friendly assessments and improved psychometric statistics.

All three versions of the Wechsler tests are useful in assessing general cognitive functioning, in helping to determine intellectual disabilities and giftedness, and in assessing probable learning problems. All three tests share many similar features, such as similar subtests, except that they are geared for their unique age groups. In the development of each of the tests, the publishers used a random sample of approximately 2,000 children or adults who were stratified for age, sex, race, par- ent education level, and geographic region. In addition, they worked hard to assure content validity, based on their ongoing view of intellectual functioning, and showed a variety of different kinds of criterion-related and construct validity. Over- all reliability estimates for each of the tests tend to be high, in the mid 0.90s, although some of the subtests will dip down into the 0.70s (Hess, 2001; Madle, 2005; Pearson, 2008a).

Because the tests are similar in nature, to give you a sense of what these tests are like, the following offers an overview of the WISC-IV, the most widely used of the three Wechsler intelligence scales. In Table 9.3 you can see a short description of the 15 subtests of the WISC-IV (Wechsler, 2003), many of which are similar to those in the WPPSI-III and WAIS-IV. Each subtest measures a different aspect of cognitive functioning, although there is some overlap among certain subtests.

The WISC-IV provides a full-scale IQ as well as four additional composite score indexes in areas called verbal comprehension index (VCI), perceptual reason- ing index (PRI), working memory index (WMI), and processing speed index (PSI). The subtests that are associated with each of their respective composite score indexes are listed in Table 9.4. The 10 subtests not italicized in Table 9.4 are used to find the full-scale IQ, while each of the 5 italicized subtests can be used if a sub- test is invalidated for some reason (e.g., given incorrectly, not appropriate for a child due to his or her specific disability) or if additional information is warranted.

The four composite score indexes provide important information concerning the child being tested, including identifying strengths and weaknesses of a child as well as helping to identify a possible learning disability, when used in conjunction with an educational achievement measure. Subtest scores use scaled scores (stan- dard scores) that have a mean of 10 and a standard deviation of 3, while the full- scale IQ is reported with a mean of 100 and standard deviation of 15. The front page of the WISC-IV Record Form offers a summary of the child’s test scores. Figure 9.7 shows the form, which offers six parts, A–F.

Part A presents the child’s age at testing and the date of testing. Part B presents the raw score of each of the subtests given as well as the scaled (standard) scores. In this case, you can see that although two additional subtests were used to gather

WISC-IV Subtests measure broad range of cogni- tive ability; 10 subtests combine for a com- posite score (g)

Composite scores are useful in identifying learning disabilities

CHAPTER 9 Intellectual and Cognitive Functioning 205

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TABLE 9.3 Abbreviations and Descriptions of Subtests

Subtest Abbreviation Description

Block Design BD While viewing a constructed model or a picture in a Stimulus Book, the child uses red-and-white blocks to re-create the design within aspecified time limit. (Visual-Motor-Spatial Skills)

Similarities SI The child is presented two words that represent common objects or concepts and describes how they are similar. (Verbal Abstract Reasoning)

Digit Span DS For Digit Span Forward, the child repeats numbers in the same order as presented aloud by the examiner. For Digit Span Backward, the child repeats numbers in the reverse order of that presented aloud by the examiner. (Short-term Auditory Memory)

Picture Concepts PCn The child is presented with two or three rows of pictures and chooses one picture from each row to form a group with a common characteris- tic. (Visual Abstract Reasoning)

Coding CD The child copies symbols that are paired with simple geometric shapes or numbers. Using a key, the child draws each symbol in its corre- sponding shape or box within a specified amount of time. (Short-term visual memory and fine motor skills)

Vocabulary VC For Picture Items, the child names pictures that are displayed in the Stimulus Book. For Verbal Items, the child gives definitions for words that the examiner reads aloud. (Verbal Reasoning skills)

Letter-Number LN The child is read a sequence of numbers and letters and recalls the numbers in ascending order and the letters in alphabetical order. (Short- term memory & executive functioning)

Matrix Reasoning MR The child looks at an incomplete matrix and selects the missing portion from five response options. (Logical Sequencing skills-visual)

Comprehension CO The child answers questions based on his or her understanding of general principles of social situations. (Social Norms)

Symbol Search SS The child scans a search group and indicates whether the target symbol(s) matches any of the symbols in the search group within a specified time limit. (visual processing acuity)

(Picture Completion)

PCm The child views a picture and then points to or names the important part missing within a specified time limit.

(Cancellation) CA The child scans both a random and a structured arrangement of pictures and marks target pictures within a specified time limit.

(Information) IN The child answers questions that address a broad range of general knowledge topics.

(Arithmetic) AR The child mentally solves a series of orally presented arithmetic problems within a specified time limit.

(Word Reasoning)

WR The child identifies the common concept being described in a series of clues.

Note: Items in parentheses are not in the original.

Source: Wechsler, D. (2003). WISC-IV administration and scoring manual (pp. 2–3). San Antonio, TX: Harcourt Assessment, Inc.

206 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TABLE 9.4 Subtests and Composite Indexes*

VCI PRI PSI VMI

Similarities Block design Coding Digit span

Vocabulary Picture concepts Symbol search Letter-number sequencing

Comprehension Matrix reasoning Cancellation Arithmetic

Information Picture completion

Word reasoning

*Subtests not italicized are used to determine full scale IQ. Italicized subtests are used as “supplemental subtests” when there is a significant dif- ferences noted among scaled scores within an index composite. (e.g., three or more points difference) or to replace one of the nonitalicized tests due to examiner mistakes or inability of the individual to take the test due to a disability.

FIGURE 9.7 | WISC-IV Record Form Source: Wechsler, D. (2003). WISC-IV administration and scoring manual (p. 46). San Antonio, TX: Harcourt Assessment, Inc.

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 9 Intellectual and Cognitive Functioning 207

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

more information (cancellation and arithmetic), they are not included when deriv- ing the full-scale IQ. Sometimes, however, these subtests can be used in place of another subtest or in an effort to gain additional information about the individual. In these cases, they are considered when reporting the full-scale IQ. Part C offers the sum of the scaled scores. Part D offers a summary of the four composite indexes as well as the full-scale IQ. Here, the sums of all of the scaled scores are offered as well as converted scores using a mean of 100 and standard deviation of 15. Corresponding percentiles are also presented, and confidence intervals (stan- dard error of measurement) show the probability that 95% of the time the score is likely to fall within the range given. Part E shows the breakdown of the specific subtests grouped by composite index. This visual representation allows one to read- ily see an individual’s strengths and weaknesses as based on the composite indexes. Part F shows a visual representation of the composite indexes and the full-scale IQ using a mean of 100 and an SD of 15. Thus, comparisons can be made to each of the composite indexes with the full-scale IQ.

The Wechsler tests in general have good validity and excellent reliability. For example, the WISC-IV internal consistency for the full scale is 0.97. Individual sub- test reliabilities tend to average in the 0.80s (Wechsler, 2003). The WISC-IV techni- cal manual has 51 pages dedicated to providing evidence of validity, including evidence based on test content, internal structure, factor analysis, and relationships with numerous other variables (Wechsler, 2003). As you can see, the WISC-IV, as well as its first cousins the WPPSI-III and WAIS-IV, offers a comprehensive picture of the cognitive function of the individual. Such a test is useful in determining such things as intellectual disabilities, giftedness, learning problems, and other related cognitive deficiencies.

Kaufman Assessment Battery for Children The Kaufman Assessment Battery for Children, Second Edition (KABC-II) is an individually administered test of cognitive ability for children between the ages of 3 and 18. Depending on the age range, test times can vary from 25 to 70 minutes. Subtests and scoring allow for a choice between two theoretical models, one of which is CHC model discussed earlier. However, both methods examine visual processing, fluid reasoning, and short- and long-term memory. Scores are age- based and have a mean of 100 and an SD of 15 but also can be provided as a per- centile rank or age equivalent (Pearson, 2012a).

The norm group for the KABC-II was based on a sample of 3,025 that was strat- ified against the U.S. population for gender, race, SES, religion, and special education status. Reliability estimates are quite sound, ranging from 0.87 to 0.95 for composite score means. Subtest reliabilities are generally strong, falling within the 0.80s range. Additional psychometric data such as validity is available in the manual.

Nonverbal Intelligence Tests Nonverbal intelligence tests are different than traditional measures of intelligence in that they rely on little or no verbal expression and are often appropriate for chil- dren who may be disadvantaged by traditional verbal and language-based

KABC-II Measures cognitive ability for ages 3 to 18; provides choice of theoretical model of intelligence

Nonverbal intelligence tests Intelligence tests that rely on little or no ver- bal expression.

208 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

measures. Such tests assess intelligence for children with autism, intellectual disabilities, specific language-based learning disabilities, poor expressive abilities, hearing impair- ments, differences in cultural background, and certain psychiatric disorders (Aylward, 1996). Three commonly used nonverbal intelligence tests are the Comprehensive Test of Nonverbal Intelligence 2 (CTONI-2), the Universal Intelligence Test (UNIT), and the Wechsler Nonverbal Scale of Ability (WNV).

Comprehensive Test of Nonverbal Intelligence, Second Edition (CTONI-2) The Comprehensive Test of Nonverbal Intelligence (CTONI) is a nonverbal instrument designed to measure intellect from ages 6 years, 0 months to 89 years, 11 months (Hammill, Pearson, & Wiederholt, 2009). The test is composed of six subtests (pictorial analogies, geometric analogies, pictorial categories, geometric categories, pictorial sequences, and geometric sequences) that measure different nonverbal intellec- tual abilities. Individuals receive results in a number of ways including standard scores, percentiles, and age equivalents. The CTONI was renormed on a national sample of 2,827 individuals in 2007 and 2008, and the publisher currently reports reliabilities mostly in the 0.90s with some in the 0.80s. Convergent validity with popular intelli- gences tests tend to be in the high 0.70s, which is fairly good for a test of this kind.

Universal Nonverbal Intelligence Test (UNIT) The Universal Nonverbal Intelli- gence Test (UNIT) is unique because it is a completely nonverbal instrument designed to measure intelligence of children ages 5 to 17 years (Houghton Mifflin Harcourt, n.d.b). Composed of six subtests, the UNIT measures symbolic memory, analogic reasoning, object memory, spatial memory, cube design, and reasoning and offers scaled score with a mean of 10 and an SD of 3 for each subtest. In addi- tion, full-scale IQ, memory quotient, reasoning quotient, symbolic quotient, and nonsymbolic quotient scores are given. The test time is about 45 minutes.

The UNIT was standardized on a national sample of 2,100 children stratified against gender, race, Hispanic origin, region, parental educational attainment, com- munity setting, and classroom placement (regular or special education) (Bandalos, 2001). All versions of the UNIT show reliability estimates that range from the mid 0.80s to mid 0.90s, and the test showed a correlation of 0.80 with the WISC-III. Predictive validity coefficients were weaker, with correlations between the UNIT standard battery full-scale scores and the Woodcock Johnson-Revised broad knowledge scale to be 0.51, 0.56, and 0.66 for samples of learning disabled, intel- lectually disabled, and intellectually gifted students, respectively.

Wechsler Nonverbal Scale of Ability (WNV) The Wechsler Nonverbal Scale of Ability (WNV) is a nonverbal test that can be used for any individual but is particu- larly adaptable for those who are culturally diverse, non–English speaking, hard of hearing, special education individuals, and gifted individuals from linguistically diverse populations (Pearson, 2008b, 2012b). The test offers six subscales with matrices, coding, object assembly, and recognitions subtests being offered for 4- to 7-year-olds, and matrices, coding, spatial span, and picture arrangement being offered for 8- to 22-year-olds (notice the similarities to the earlier Wechsler intelli- gence scales). The test time is about 45 minutes and the test offers T-scores for the subtest report and a full-scale IQ score using deviation IQ (Pearson, 2008c).

CTONI, UNIT, WNV Three nonverbal tests of intelligence

CHAPTER 9 Intellectual and Cognitive Functioning 209

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Evidence of convergent validity with other tests of verbal and of nonverbal intelligence ranged from 0.57 to 0.73 that Maddux (2010) suggests is on the low- end. Reliability for the whole test was 0.91 while subscale reliability tended to be in the mid 0.70s or low 0.80s. The normative sample seems to be representative of the U.S. population and development of the test was carefully done to ensure that it was useful for those who were culturally and linguistically diverse. Despite some- what lackluster convergent validity, the test seems to be a fairly good test at mea- suring nonverbal intelligence.

NEUROPSYCHOLOGICAL ASSESSMENT Neuropsychological assessment is a new field compared to intelligence testing and offers a broad array of ways to examine the cognitive functioning of individuals. The following offers a brief history of neuropsychological assessment, defines the purposes and uses of this kind of assessment, and provides a quick look at a couple of the more common tests used when conducting a neuropsychological assessment.

A Brief History of Neuropsychological Assessment Although the professional identity of neuropsychology emerged in the early 1970s, interest in how the brain functions has been around for centuries. For instance, obser- vations of behavioral changes following head injuries are found in 5000-year-old Egyptian medical documents (Hebben-Milberg, 2009). In modern days, interest in brain injury was peaked during World War I as significant numbers of soldiers suf- fered brain trauma, and it was at this point that screening and diagnostic measures were created (Lezak, Howieson, Bigler, & Tranel, 2012). This early research on war- damaged veterans is said to be the catalyst for the birth of clinical neuropsychology.

During the 1950s, it was realized that what might be considered the same kind of brain damage in different individuals could, in fact, affect people differently. In other words, brain injury was found to be unique in the sense that it could lead to a wide variety of behavioral patterns (McGee, 2004). More recently, the advent of modern-day technology has significantly contributed to changes to the relatively new field of neuropsychology. For instance, the invention of diagnostic scanning devices such as the magnetic resonance imaging (MRI) and positron emission tomography (PET) makes many of the former neuropsychological assessments unnecessary (Kolb & Whishaw, 2009). However, the most sensitive measure of brain capacity is behavior, which is not measured by these scanning devices. Thus, neuropsychological assessment continues to play a significant role in explaining brain–behavior relationships. The following section defines the complex field of neuropsychology and neuropsychological assessment.

Defining Neuropsychological Assessment Neuropsychology is a domain of psychology that examines brain–behavior rela- tionships. A particular discipline of neuropsychology is clinical neuropsychology, which includes both the assessment of the central nervous system and interventions that may result from an assessment (Hebben & Milberg, 2009). Neurological

Neuropsychology A domain of psychol- ogy that examines brain–behavior relationships

Clinical neuropsychology The assessment and intervention principles related to the central nervous system

210 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

assessments are generally employed following a traumatic brain injury, an illness that affects brain function, or because of suspected changes in brain function from the aging process.

Neuropsychological assessment can measure a number of domains related to brain–behavior, including memory, intelligence, language, visual-perception, visual- spatial thinking, psychosensory and motor abilities, academic achievement, personal- ity, or psychological functioning (Lezak et al., 2012). Uses of such an assessment can vary dramatically. For example, results can be used in the following ways:

• as a diagnostic tool to identify the root of a condition and the extent of the brain damage;

• to measure change in one’s functioning (e.g., cognitive ability, movement, and reaction time);

• to compare changes in cognitive or functional status to others within the nor- mative sample;

• to provide specific rehabilitation treatment and planning guidelines for indivi- duals and families; and

• to provide specific guidelines for educational planning in the schools. (Kolb & Whishaw, 2009)

Neuropsychological assessment shares some common elements with intelligence testing. In fact, many neuropsychological assessments begin with a measure of gen- eral intelligence (Kolb & Whishaw, 2009). Further, some intelligence scales have been adapted for neuropsychological purposes; most notably the Wechsler Adult Intelligence Scale-Revised as a Neurological Instrument (WAIS-R-NI). In spite of these parallels, there are notable differences between an intellectual assessment and a neuropsychological evaluation. When conducting an intellectual assessment, the individual’s scores are compared to normative data for one of the purposes men- tioned earlier in the chapter (e.g., to determine learning disabilities, giftedness, intellectual disabilities, etc). In contrast, neuropsychological assessment compares an individual’s current levels of functioning to his/her estimated premorbid level of functioning (Lezak et al., 2012). In the following section, we take a closer examina- tion of what a neuropsychological assessment looks like.

Methods of Neuropsychological Assessment Early assessment procedures relied on single, independent instruments to assess brain damage. However, most modern approaches to neuropsychology have moved away from the idea that there is one single test for brain damage and rely on “battery approaches” to assessment (Kolb & Whishaw, 2009). With this being said, there is still no universal approach to neuropsychological assessment. That is, current assessment practices utilize a continuum of approaches, from a fixed bat- tery approach to a flexible battery approach. This section introduces each approach and also includes their respective strengths and weakness.

Fixed Battery Approach and the Halstead-Reitan Battery A fixed battery approach to neuropsychological assessment involves a standardized administration of a uniform group of instruments. That is, with fixed battery approaches all

Fixed battery The rigid and stan- dardized administra- tion of a uniform group of instruments

CHAPTER 9 Intellectual and Cognitive Functioning 211

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

individuals requiring assessment receive the same set of tests. The majority of fixed batteries have cut off scores that reflect the degree of severity of the impairment and differentiate between impaired and unimpaired individuals. Two common fixed neuropsychological batteries are the Halstead-Reitan Battery and the Luria- Nebraska Neuropsychological Battery. This section provides a description of the Halstead-Reitan Battery.

The Halstead-Reitan was developed by Ward Halstead in the 1950s and modi- fied by Halstead’s graduate student Ralph Reitan (Hebben & Millberg, 2009). Two children’s versions also exist: the Reitan Indiana Neuropsychological Test Bat- tery (ages 5 to 8) and the Halstead Neuropsychological Test Battery for Older Chil- dren (ages 9 to 14). The Halstead-Reitan provides individuals with a cut-off score or index of impairment, which discriminates brain damaged from normally func- tioning individuals. Information about specific areas of the brain that are damaged and information about the severity of the damage can also be garnered from this battery. The Halstead-Reitan battery takes approximately 5 to 6 hours to complete and consists of eight core tests as follows (Dean, 1985; Encyclopedia of Mental Disorders, n.d.; Hebben & Milberg, 2009):

1. Category Test: A total of 208 pictures consisting of geometric figures are pre- sented to the examinee. With each picture, individuals are asked whether the figure reminds them of the number 1, 2, 3, or 4. Test-takers then press a key corresponding to their response. This core test evaluates abstraction ability and the ability to draw specific conclusions from general information.

2. Tactual Performance Test: A form board consisting of 10 cut-out shapes and 10 wooden blocks matching those shapes are presented in front of the blind- folded test-taker. Individuals are instructed to place the blocks in their appro- priate square on the board, first with their dominant hand, then with their nondominant hand, and then with both hands. Then, the form board and blocks are removed. Without the blindfold removed, participants are then instructed to draw the form board and shapes in their proper locations. This core test evaluates sensory ability, memory for shapes and the special location of shapes, motor functions, and transferability between both hemispheres of the brain.

3. Trail Making Test: This is a two-part core test. Part A consists of a page with 25 circles numbered and randomly arranged. Individuals are instructed to draw lines, circle to circle, in chronological order until they reach the circle labeled “End.” Part B consists of a page with circles containing the letters A through L as well as 13 numbered circles intermixed and randomly arranged. Individuals are instructed to connect the circles by drawing lines alternating between num- bers and letters in sequential order, until they reach the circle labeled “End.” This core tests evaluates information processing speed, visual scanning ability, the integration of visual and motor functions, letter and number recognition and sequencing, and the ability to maintain two different trains of thought.

4. Finger Tapping Test: Individuals place the index finger on a lever that is attached to a counting device. Then, the test takers are instructed to tap their fingers as quickly as possible for 10 seconds. The trial is repeated 5 to 10 times for each hand to measure motor speed, manual dexterity, and hand

Halstead-Reitan A widely utilized fixed battery neuropsycho- logical assessment consisting of eight core tests

212 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

dominance. This test is helpful in determining the particular area(s) of the brain that may be damaged.

5. Rhythm Test: A total of 30 pairs of rhythmic beats are presented to the exam- inee. With each pair, the test-taker decides if the two sounds are the same or different. This core test evaluates auditory attention and concentration, and the ability to discriminate between nonverbal sounds.

6. Speech Sounds Perception Test: A total of 60 nonsense syllables containing the vowel sound “ee” are presented. After each syllable, the test-takers underline, from a set of four written alternatives, the spelling that represents the sound they just heard. This core test examines auditory attention and concentration and the ability to discriminate between verbal sounds.

7. Reitan-Indiana Aphasia Screening Test: Here, test-takers are presented with a variety of questions and tasks including naming pictures aloud, writing the name of the picture without saying the name aloud, reading material aloud, repeating words, drawing shapes without lifting the pencil, and so on. This core test detects possible signs of aphasia, which is the loss of ability to under- stand or use written or spoken language, usually as a result of brain damage, illness, or deterioration.

8. Reitan-Klove Sensory-Perceptual Examination: Test-takers are asked to partic- ipate in kinesthetic-related procedures. For example, (a) specify whether touch, sound, or visible movement is occurring on the right, left, or both sides of the body; (b) identify numbers “written” on their fingertips while eyes are closed; (c) recall numbers assigned to particular fingers (the examiner assigns numbers by touching each finger and stating the number with the individual’s eyes closed); and (d) identify the shape of a wooden block placed in one hand by pointing to its shape on a form board with the opposite hand. This core test detects whether individuals are unable to perceive stimulation on one side of the body when both sides are stimulated at the same time.

The Halstead-Reitan is perhaps the most widely used, most rigorous and most researched of all fixed neuropsychological batteries. However, this battery and other fixed battery approaches, have notable weaknesses. For example, the length of time required to complete the battery requires a great deal of patience and financial cost to the individual. Further, since fixed batteries give the same test to all individuals, some information may be missed or some unnecessary information may be gathered. In addition, the 102-page testing manual fails to provide psychometric data concerning the validity, reliability, and norming proce- dures of the Halstead-Reitan. Failure to include the information makes the inter- pretation of the instrument somewhat problematic. This lack of information should be a “red flag” and test administrators should utilize this battery with some amount of caution.

Flexible Battery Approach and the Boston Process Approach (BPA) In the flexible battery approach to neuropsychological assessment, the use of a combination of tests is dictated by the referral questions and the unique needs and behaviors of the client. In this case, a series of tests are chosen that may evaluate different areas of neuropsychological functioning. Generally speaking, clinicians will utilize an exhaustive

Flexible battery Combination of tests dictated by the referral questions and unique needs of the client

CHAPTER 9 Intellectual and Cognitive Functioning 213

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

list of tests within a flexible battery. The Boston Process Approach (BPA) is one example of the flexible battery approach. This approach requires careful observation of the test-taker during test administration. With the BPA and additional flexible bat- tery approaches, there is a strong emphasis on garnering qualitative data. For exam- ple, the clinician may be more concerned with how the test-taker solves the problem, may give additional time to the test-taker for answering the problem, and may adapt the exercise to the individual’s specific needs (Milberg, Hebben & Kaplan, 2009).

There are clear strengths and weaknesses of using the flexible battery approach in neuropsychological assessment. For instance, whereas this approach allows the test administrator to give a wide variety of tests, tailored to the presenting problem, some critical areas may be overlooked. Additionally, a flexible battery approach requires a great deal of training specific to neuropsychology (Vanderploeg, 2000). Flexible batteries have also been criticized because they gather limited psychometric data. Finally, the flexible battery approach has undergone a considerable amount of scrutiny from within the courts system. For example, certain neuropsychological evidence has been excluded from trials because they were deemed to not be as sci- entific as a fixed battery approach.

Both the fixed and flexible battery approaches to neuropsychological assess- ment requires specific training to adequately and ethically administer to others. As discussed earlier, this is also the case for intelligence tests (see Box 9.2).

THE ROLE OF HELPERS IN THE ASSESSMENT OF INTELLECTUAL AND COGNITIVE FUNCTIONING

Intellectual and neuropsychological assessments take advanced training, which is generally given in school psychology programs and many doctoral programs in counseling, clinical psychology, and clinical neuropsychology. Although other help- ers, such as learning disabilities specialists, licensed clinical social workers, and licensed professional counselors can also obtain such training, many graduate

BOX 9.2 Traumatic Brain Injury (TBI), Veterans, and Fixed or Flexible Approaches to Assessment

Since the beginning of the wars in Iraq and Afghanistan over 250,000 of our service members have been diagnosed with a traumatic brain injury (TBI) (Defense and Veterans Brain Injury Center, 2012). The brain is malleable and has great healing powers if adequate treatment is given. However, such treatment can only occur if our brave soldiers are assessed appropriately (Decker, Englund, & Roberts, 2012). When diffuse brain damage is sus- pected, a fixed battery approach that provides a

board standardized assessment may be called for. Other times, when focal damage is suspected, trained neuropsychological assessors may choose to use a flexible approach that uses combination of different tasks that focuses specifically on the sus- pected area of brain injury. You can see why training and experience in this field would be important in deciding whether to use a fixed or flexible approach and in choosing the appropriate assessment instru- ments to assess the individual.

Boston Process Approach A well-known flexible battery method of assessment

© Ce ng ag e Le ar ni ng

20 15

214 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

programs in special education, counseling, and social work do not automatically offer these courses.

When a helper does have this training, he or she can provide a wide range of services to individuals, including the assessment of intelligence to help determine learning problems, intellectual disabilities, giftedness, potential for learning, and possible neurological impairment. For those helpers who do not have the training in intellectual or neurological assessment, it is still imperative that they have suffi- cient knowledge of these kinds of tests so that they will know when to refer clients who might need such an assessment and be able to participate in the development of treatment plans for their clients who have received a cognitive assessment.

FINAL THOUGHTS ON THE ASSESSMENT OF INTELLECTUAL AND COGNITIVE FUNCTIONING

Intelligence testing and neuropsychological assessment can be used in a variety of ways, and if one is not careful, can be abused. Abuses over the years have been many, including the miscalculation of intelligence of minorities, overclassification of individuals who are learning disabled, the misguided belief that there was “proof” of racial differences of ability, and a means of differentiating social classes. However, the astute examiner knows that the assessment of cognitive functioning is complex and the examinee’s culture, environment, genetics, and biology should be considered in such assessments. Any conclusions about an individual should be made within the context of knowing the whole person as well as the complex soci- etal issues that are involved.

SUMMARY We started this chapter by looking at a young man who had neurological impairment as result of a bicycle accident. We noted that the use of both an intelligence test and a neurological assessment would be critical to understanding his cognitive functioning after the accident. Then we noted that this chapter would examine intelligence test- ing, models of intelligence, and the involved pro- cess of neuropsychological assessment.

The chapter then offered a brief review of the history of intelligence testing and a definition of intelligence testing. We noted that intelligence tests are types of aptitude tests that measure a range of intellectual ability and offer a broad assessment of one’s cognitive capabilities. We pointed out that intelligence tests are used in a

variety of ways, such as assisting in determining giftedness, intellectual disabilities, and learning disabilities; helping to assess intellectual ability following an accident, the onset of dementia, sub- stance abuse, disease processes, and trauma to the brain; as part of the admissions process for some private schools; and as part of a personality assessment to help understand the whole person.

Next, we offered a brief introduction of models of intelligence, including Spearman’s two-factor approach, Thurstone’s multifactor approach, Vernon’s hierarchical model of intelligence, Guilford’s multifactor/multidimensional model, Cattell’s fluid and crystal intelligence, Piaget’s cog- nitive development theory, Gardner’s theory of multiple intelligences, Sternberg’s triarchic model

CHAPTER 9 Intellectual and Cognitive Functioning 215

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

of successful intelligence, and the Cattell- Horn-Carroll (CHC) integrated model. A sum- mary of these models was presented in Table 9.2.

After reviewing the models of intelligence, we discussed three commonly used intelligence tests that are partly based on some of these models: the Stanford-Binet, Fifth Edition (SB5), the Wechsler tests (focusing on the WISC-IV), and the Kaufman Assessment Battery for Children (KABC). With the Stanford-Binet, we noted that the test mea- sures cognitive functioning of individuals from age 2 through 85+. We pointed out that it assesses verbal and nonverbal intelligence across five fac- tors: fluid reasoning, knowledge, quantitative rea- soning, visual-spatial processing, and working memory. We noted that it uses a vocabulary rout- ing test to help determine the basal age of the examinee, or the point where testing is started, and that ceiling ages are used to determine when the testing should be completed. We highlighted the fact that the test has been one of the most popular tests of cognitive functioning and con- tinues to be widely used today.

Next, we discussed the Wechsler tests, partic- ularly the WISC-IV. We identified 15 subtests that can be assessed when administering the WISC-IV, although only 10 are used in determining the examinee’s full-scale IQ. Subtests are specific to four composite score indexes, which include Verbal Comprehension Index (VCI), Perceptual Reasoning Index (PRI), Working Memory Index (WMI), and Processing Speed Index (PSI). These indexes are important in determining possible learning problems. The Wechsler tests are proba- bly the most widely used tests of mental ability and are used to assess intellectual disabilities, gift- edness, learning problems, and general cognitive functioning.

We then examined the Kaufman Assessment Battery for Children, Second Edition (KABC-II), which measures cognitive ability of children from ages 3 to 18 and specifically focuses on visual processing, fluid reasoning, and short- and long-term memory. We pointed out that the test can be scored based on a couple of models of

intelligences, including Cattell’s model of fluid and crystallized intelligence.

The last intelligence tests we examined were the Comprehensive Test of Nonverbal Intelligence (CTONI), the Universal Nonverbal Intelligence Test (UNIT), and the Wechsler Nonverbal Scale of Ability (WNV), three nonverbal intelligence measures. The CTONI was designed to measure nonverbal intelligence from ages 6 to 90 and is composed of 6 subtests that measure different non- verbal intellectual abilities. The UNIT was designed to measure nonverbal intelligence of children ages 5 year to 17 years, is composed of 6 subtests, and offers a full-scale IQ, memory quotient, reasoning quotient, symbolic quotient, and nonsymbolic quo- tient. The WNV is a nonverbal test of ability that can be used for any individual but is particularly adaptable for the culturally diverse, non–English speaking, hard of hearing, special education indivi- duals, and gifted individuals from linguistically diverse populations. The test is geared for ages 4 through 22 and four of six subscales are chosen based on the age of the examinee.

We then introduced the concept of neuropsy- chology. More specifically, this section focused on how neuropsychological assessment assesses a wide variety of problems related to brain–behavior rela- tionships. We noted that although the field of neu- ropsychology is relatively new, interest in brain functioning goes back thousands of years. We pointed out that neuropsychological assessments are generally called for following a traumatic brain injury, an illness that affects brain function, or because of suspected changes in brain function from the aging process. We highlighted the fact that the results of a neurological assessment can be used as a diagnostic tool in identifying root causes and the extent of brain damage, to measure changes in functioning, to compare cognitive functioning to a norm group, to help in treatment and planning for individuals and their families, and to offer guide- lines for educational planning.

Relative to neurological assessment, we highlighted two methods of assessment: a fixed and flexible battery approach. In a fixed battery

216 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

approach, all individuals requiring assessment receive the same set of standardized tests. Two well-known tests of this kind are the Halstead- Reitan Battery and the Luria-Nebraska Neuropsy- chological Battery. Some of the categories assessed in the Halstead-Reitan include Category Test, Tactual Performance Test, Trail Making Test, Finger Tapping Test, Rhythm Test, Speech Sounds Perception Test, Aphasia Screening Test, and the Sensory-Perceptual Examination. We also discussed the Boston Process Approach (BPA), which is one example of a flexible battery approach. This approach requires careful obser- vation of the test-taker during test administration and is much more qualitative in nature.

As the chapter neared its conclusion, we noted that only some helpers are given the

advanced training needed for the assessment of cognitive ability. For those helpers who do not have such training, we highlighted the reasons it is still imperative to have basic knowledge of these kinds of tests.

Finally, we noted that over the years, the examination of cognitive functioning has been abused by miscalculating the intelligence of minorities, overclassifying individuals with learn- ing disabilities, “proving” racial differences of ability, and differentiating social classes. We sug- gested that astute examiners know that the assess- ment of cognitive functioning is complex; is based on the environment, genetics, and biology; and that conclusions should be drawn tentatively and done within the context of knowing the whole person.

CHAPTER REVIEW 1. Discuss some of the many ways that intelli-

gence tests are used. 2. Highlight the major points of the following

theories of intelligence: a. Spearman’s two-factor approach b. Thurstone’s multifactor approach c. Vernon’s hierarchical model of intelligence d.Guilford’s multifactor/multidimensional model

e. Cattell’s fluid and crystal intelligence f. Piaget’s cognitive development theory g. Gardner’s theory of multiple intelligences h.Sternberg’s triarchic theory of successful intelligence

i. Cattell-Horn-Carroll’s (CHC) integrated model of intelligence

3. Why are Gardner’s theory of multiple intelli- gences and Sternberg’s triarchic theory of successful intelligence revolutionary as com- pared to the more traditional models of intelligence?

4. If you were to use Gardner’s theory or Sternberg’s theory to develop an intelligence test, what might it look like?

5. Describe the concepts of basal level and ceil- ing level. How are these levels used in the Stanford-Binet?

6. Explain the differences between the WAIS-IV, WISC-IV, and WPPSI-III.

7. Although there are many similarities between the Stanford-Binet and the Wechsler tests of intelligence (e.g., both measure global intelli- gence, both compare nonverbal and verbal intelligence), some major differences exist. Discuss these differences and the strengths and weaknesses of both tests.

8. Discuss the purpose of nonverbal intelligence tests. What populations are they suited for?

9. Define neuropsychological assessment. 10. What domains are assessed in neuropsycho-

logical assessment? 11. What are the two major assessment

approaches to neuropsychological assess- ment? Give description of each.

12. Discuss the role of helpers in the administra- tion and interpretation of intellectual assess- ments and of neuropsychological assessments.

CHAPTER 9 Intellectual and Cognitive Functioning 217

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

REFERENCES Aylward, G. P. (1996). Review of the Comprehensive

Test of Nonverbal Intelligence. In J. C. Impara, & B. S. Plake (Eds.), The thirteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Bandalos, D. L. (2001). Review of the Universal Non- verbal Intelligence Test. In B. S. Plake, & J. C. Impara (Eds.), The thirteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measure- ments Yearbook database.

Carroll, J. B. (1993). Human cognitive abilities: A survey of factor-analytic studies. New York: Cambridge University Press.

Carroll, J. B. (2005). The three-stratum theory of cog- nitive abilities. In D. P. Flanagan & P. L. Harrison (Eds.), Contemporary intellectual assessment: Theo- ries, tests, and issues (2nd ed., pp. 69–76). New York: Guilford Publications, Inc.

Cattell, R. (1971). Abilities: Their structure, growth, and action. Boston: Houghton Mifflin.

Cattell, R. (1979). Are culture fair intelligence tests pos- sible and necessary? Journal of Research and Devel- opment in Education, 12, 3–13.

Cattell, R. (1980). The heritability of fluid, gf, and crys- tallized, gc, intelligence, estimated by a least squares use of the MAVA method. British Journal of Educa- tional Psychology, 50, 253–265.

Dean, R. (1985). Review of the Halstead-Reitan Neuropsychological Battery. In J. V. Mitchell (Ed.), The ninth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Decker, S. L., Englund, J. A., & Roberts, A. M. (2012). Intellectual and neuropsychological assessment of individuals with sensory and physical disabilities and traumatic brain injury. In D. P. Flanagan, & P. L. Harrison (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (3rd ed., pp. 708–725). New York: Guilford Publications, Inc.

Defense and Veterans Brain Injury Center. (2012). DoD worldwide numbers for TBI. Retrieved from http://www.dvbic.org/dod-worldwide-numbers-tbi

Educational Broadcasting Corporation. (2004). Con- cept to classroom: What is the theory of multiple intelligences (M. I.)? Retrieved from http://www. thirteen.org/edonline/concept2class/mi/index.html

Encyclopedia of Mental DisordersEncyclopedia of Mental Disorders. (n.d.). Halstead-Reitan battery. Retrieved from http://www.minddisorders.com /Flu-Inv/Halstead-Reitan-Battery.html

Gardner, H. (1996). MI: Intelligence, understanding and the mind [motion picture]. (Available from Media into the classroom, 10573 W. Pico Blvd. #162, Los Angeles, CA 90064)

Gardner, H. (1999). Intelligence reframed: Multiple intelligence for the 21st century. New York, NY: Basic books.

Gardner, H. (2003, April 21). MI after twenty years. Paper presented at the American Educational Research Association, Chicago. Retrieved http:// howardgardner01.files.wordpress.com/2012/06/mi- after-twenty-years2.pdf

Gardner, H. (2011). Frames of mind. New York: Basic Books Inc. (Original work published 1983).

Guilford, J. (1967). The nature of human intelligence. New York: McGraw-Hill.

Guilford, J. (1988). Some changes in the structure of the intellect model. Educational and Psychological Measurement, 48, 1–4. doi:10.1177/001316448804800102

Hammill, D. D., Pearson, N. A., & Winderholt, J. L. (2009). CTONI 2: Comprehensive Test of Nonver- bal Intelligence: Examiners manual (2nd ed.). Aus- tin, TX: Pro-ed.

Hebben, N., & Milberg, W. (2009). Essentials of neuropsychological assessment (2nd ed.). Hoboken, NJ, New York: John Wiley & Sons.

Hess, A. K. (2001). Review of the Wechsler Adult Intelli- gence Scale. In B. S. Plake, & J. C. Impara (Eds.), The thirteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Horn, J. L., & Blankson, N. (2012). Foundations for better understanding of cognitive abilities. In D. P. Flanagan, & P. L. Harrison (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (2nd ed., pp. 73–98). New York: Guilford Publica- tions, Inc.

Horn, J. L., & Cattell, R. B. (1966). Refinement and test of the theory of fluid and crystallized intelli- gence. Journal of Educational Psychology, 57, 253–270. doi:10.1037/h0023816

Houghton Mifflin Harcourt (n.d.a). The Stanford-Binet Intelligence Scales (SB5), fifth edition. Retrieved

218 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

from http://www.riverpub.com/products/sb5/index. html

Houghton Mifflin Harcourt. (n.d.b). Universal Non- verbal Intelligence Test (UNIT). Retrieved from http://www.riverpub.com/products/unit/details.html# interpret

Johnson, J. A., & D’Amato, R. C. (2005). Review of the Stanford-Binet Intelligence Scales, fifth edition. In R. A. Spies, & B. S. Plake (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Jolly, J. L. (2008). Lewis Terman: Genetic study of genius—elementary school students. Gifted Child Today, 31(1), 27–33.

Kerr, M. S. (2008). Psychometrics. In S. F. Davis, & W. Buskit (Eds.), 21st century psychology: A refer- ence handbook (vol. 1, pp. 374–382). Thousand Oaks, CA: Sage Publications.

Kolb, B., & Whishaw, I. (2009). Fundamentals of human neuropsychology (6th ed.). New York: Worth Publishers.

Kush, J. C. (2005). Review of the Stanford-Binet Intel- ligence Scales, fifth edition. In R. A. Spies, & B. S. Plake (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measure- ments Yearbook database.

Lezak, M. D., Howieson, D. B., Bigler, E. D., & Tranel, D. (2012). Neuropsychological assessment (5th ed.). New York: Oxford University Press.

Madle, R. A. (2005). Review of the Wechsler Preschool and Primary Scale of Intelligence. In R. A. Spies, & B. S. Plake (Eds.), The sixteenth mental measure- ments yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Maddux, C. D. (2010). Review of Wecshler Nonverbal Scale of Ability. In R. S. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measure- ments yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

McGee, J. (2004). Neuroanatomy of behavior after brain injury or you don’t like my behavior? You’ll have to discuss that with my brain directly. Retrieved from http://bianj.org/Websites/bianj/images/ neuroanatomyofbehavior.pdf

McGrew, K. S. (2009). CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research.

Intelligence, 37, 1–10. doi:10.1016/j.intell.2008.08. 004

Milberg, W. P., Hebben, N., & Kaplan, E. (2009). The Boston Process Approach to neuropsychological assessment. In I. Grant, & K. M. Adams (Eds.), Neuropsychological assessment of neuropsychiatric disorders (3rd ed., pp. 42–55). New York: Oxford University Press.

Pearson. (2008a). Wechsler Adult Intelligence Scale, fourth edition: Now available from Pearson. Retrieved from http://pearsonassess.com/haiweb/ Cultures/en-US/Site/AboutUs/NewsReleases/NewsItem/ NewsRelease082808.htm

Pearson. (2008b). Reference materials: WNV Brochure. Retrieved from http://www.pearsonassessments.com/ HAIWEB/Cultures/en-us/Productdetail.htm?Pid= 015-8338-499

Pearson. (2008c). Sample report: Parent report. Retrieved from http://www.pearsonassessments. com/HAIWEB/Cultures/en-us/Productdetail.htm? Pid=015-8338-499

Pearson. (2012a). Kaufman Assessment Battery for Chil- dren, second edition (KABC-II). Retrieved from http://psychcorp.pearsonassessments.com/HAIWEB/ Cultures/en-us/Productdetail.htm?Pid=PAa21000

Pearson. (2012b). Wechsler Nonverbal Scale of Ability (WNV). Retrieved from http://www.pearson assessments.com/HAIWEB/Cultures/en-us/Product detail.htm?Pid=015-8338-499

Piaget, J. (1950). The psychology of intelligence (M. Piercy & D. Berlyne, Trans.). London: Routledge & Kegan.

Roid, G. (2003). Stanford-Binet Intelligence Scales, fifth edition, technical manual. Itasca, IL: Riverside Publishing.

Rose, S. P. R. (2006). Commentary: Heritability estimates—long past their sell-by date. International Journal of Epidemiology, 35, 525–527. doi:10.1093/ ije/dyl064

Schneider, W. J., & McGrew, K. S. (2012). The Cattell-Horn—Carroll model of intelligence. In D. P. Flanagan, & P. L. Harrison (Ed.), Contempo- rary intellectual assessment: Theories, tests, and issues (3rd ed., pp. 99–144). New York: Guilford Press.

Spearman, C. (1970). The abilities of man. New York: AMS Press. (Original work published in 1932).

Sternberg, R. J. (2009). Toward a triarchic theory of human intelligence. In J. C. Kaufman, & E. L. Grigorenko (Eds.), The essential Sternberg: Essays on intelligence, psychology, and education (pp. 33– 70). New York: Springer Publishing Company.

CHAPTER 9 Intellectual and Cognitive Functioning 219

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Sternberg, R. J. (2012). The triarchic theory of success- ful intelligence. In D. P. Flanagan, & P. L. Harrison (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (3rd ed., pp. 156–177). New York: Guilford Publications, Inc.

Thurstone, L. (1938). Primary mental abilities. Chicago: University of Chicago Press.

Vanderploeg, R. (Ed.). (2000). Clinician’s guide to neuropsychological assessment (2nd ed.). Mahwah, NJ: Lawrence Erlbaum Associates.

Vernon, P. (1961). The structure of human abilities (2nd ed.). London: Methuen & Co. LTD.

Wechsler, D. (2003). WISC-IV administration and scoring. San Antonio, TX: Harcourt Assessment.

220 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 10Career and OccupationalAssessment: Interest Inventories, Multiple Aptitude, and Special Aptitude Tests

A friend of mine was struggling with whether or not to continue working toward his doctoral degree in psychology. He already had a master’s degree, had worked a while in the field of mental health, and had begun his first semester toward his Ph.D. Every day he would come home from school and have panic attacks. He simply was not sure that this was the career path he wanted to take. He decided to take a Strong Interest Inven- tory, and his concerns were reinforced when the test revealed that he viewed learning more as a means to an end than as something to embrace. His personality orientation toward the world of work also did not quite fit those of most psychologists. He decided to quit his Ph.D. program and become a ski instructor. Later, he went into investment banking, and today he’s a very successful insurance broker—and much happier. What a switch!

When I was 35 and making a paltry salary as an assistant professor, I thought, “If I joined the Reserves, I could make some extra money pretty easily.” So I went down to the local recruiting center, where they told me, “First, you need to take this test.” It was the Armed Services Vocational Aptitude Test Battery (the ASVAB). I knew about it because I had taught a brief overview of it in my testing class. It measured general cognitive ability as well as specific vocational aptitudes. I said, “Sure.” A couple of weeks later, a recruiter called and said, “We want you.” I had scored high on most of the scales. I pondered joining the Reserves but then called him back and said, “Thanks, but I think I’ll pass.” Despite my ability, I decided I didn’t quite have the interest or inclination to be in the military.

(Ed Neukrug)

221

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

This chapter is about tests that can help an individual make a decision about his or her academic or vocational path. Whether it’s an interest inventory such as the one my friend took, or a multiple aptitude test such as the one I took, these instruments can be critical in helping an individual choose an occupation or a career. Thus, in this chapter we examine tests that are used in occupational and career counseling, including interest inventories, which are a type of personality assessment, and spe- cial aptitude and multiple aptitude tests, which help an individual determine what he or she is good at. We also examine some of the various roles helpers play in career and occupational assessment and conclude with some final thoughts about this important domain.

DEFINING CAREER AND OCCUPATIONAL ASSESSMENT Career and occupational assessment can take place at any point in an individual’s life, but it is often most critical at transition points such as when an adolescent moves into high school and begins to ponder an occupation or a career, when a young adult finds a job or goes on to college, when an adult decides to crystallize his or her career choices, when a middle-aged adult makes a career shift, or when an older adult con- siders shifting out of his or her career. Although talking with a counselor can make these transitions smoother, tests provide a vital adjunct to this counseling process. Three kinds of tests are sometimes used in the vocational counseling process: interest inventories, multiple aptitude tests, and special aptitude tests.

In Chapter 1 interest inventories were defined as a type of personality assess- ment, similar to objective personality tests that are discussed in Chapter 11, but generally classified separately because of their popularity and very specific focus (see Figure 1.4). Our definition from Chapter 1 suggested that interest inventories are used to determine a person’s likes and dislikes as well as an individual’s person- ality orientation toward the world of work, and are almost exclusively used in the career counseling process. In contrast, we noted in Chapter 1 that multiple aptitude tests measure a number of homogeneous abilities and are used to predict the likeli- hood of success in any of a number of vocations, whereas special aptitude tests usually measure one homogenous area of ability and are used to predict success in a specific vocational area. Interest inventories, which focus on likes and dislikes, and special and multiple aptitude tests, which focus on ability, complement each other and are sometimes given together in the career counseling process. In this chapter, we explore interest inventories, take a look at multiple aptitude tests, and conclude by discussing special aptitude tests.

INTEREST INVENTORIES With millions of interest inventories being administered each year (ACT, 2009; CPP, 2009a), there is little doubt that these instruments have become an important adjunct to the career counseling process and are big business in the world of assessment. This is partially due to the fact that interest inventories are fairly good at predicting job satisfaction based on occupational fit (Rottinghaus, Hees, & Conrath, 2009).

Interest inventories Determine likes and dislikes from a career perspective; good at predicting job satisfaction

222 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

For instance, if a person takes an interest inventory and chooses a job that seems to match his or her personality type, then that person is more likely to be satisfied in that occupation than a person whose personality type does not match his or her job. It should be noted, however, that interest in an area does not necessarily correlate with ability in that same area. For example, you might want to be a rock star, but if you lack certain musical abilities, your career may be short-lived.

Three of the most common interest inventories that we examine in this chapter include the Strong Vocational Interest Inventory (CPP, 2009a), the Self-Directed Search (SDS; PAR 2013, and the Career Occupational Preference System Interest Inventory (COPS) (EdITS, 2012a).

Strong Interest Inventory®* One of the most commonly used career inventories is the “Strong.” First developed in 1927 and published as the Strong Vocational Interest Blank (Strong, 1926), its latest version is called the Strong Interest Inventory (SII). The 291-item SII, which was reduced from 317 items in 2004, is given to people aged 16 or older, takes 35 to 40 minutes to administer, and can be administered individually or in a group setting (CPP, 2004, 2009a; Kelly, 2003). The latest version uses a five-point Likert scale that asks people to rate themselves from strongly like to strongly dislike in six broad areas (occupations, subject areas, activities, leisure activities, people, and characteristics). The SII test report offers five different types of interpretive scales or indexes in the following areas:

• General occupational themes • Basic interest scales • Occupational scales • Personal style scales • Response summary

General Occupational Themes The most commonly used score on the Strong is the one found on the General Occupational Themes, which offers a three-letter code based on Holland’s hexagon model (see Figure 10.1). Holland defines six per- sonality types: realistic, investigative, artistic, social, enterprising, and conventional. These classifications are further explained in Box 10.1. The SII identifies the test- taker’s top three Holland codes and places them in hierarchical order.

On the hexagon, codes adjacent to one another share more elements in com- mon than nonadjacent ones. For example, a typical counselor or social worker will have a Holland code of SAE, codes that are on adjacent corners of the hexa- gon and share many common elements. In contrast, people in the artistic realm tend to have little interest in jobs or tasks in the conventional category (located on the opposite side of the hexagon), such as accounting or bookkeeping.

*Strong Interest Inventory is a registered trademark of CPP, Inc.

Strong—General Occupational Themes Uses Holland’s theory of personality type (RIASEC)

CHAPTER 10 Career and Occupational Assessment 223

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Conventional Artistic

Enterprising Social

Realistic Investigative

FIGURE 10.1 | Holland’s Hexagon Model of Personality Types

BOX 10.1 Holland’s Personality and Work Type

Realistic: Realistic persons like to work with equip- ment, machines, computer hardware, or tools, often prefer to work outdoors, and are good at manipulat- ing concrete physical objects. These individuals pre- fer to avoid social situations, artistic endeavors, or abstract tasks. Occupational settings in which you might find realistic individuals include filling stations, farms, machine shops, construction sites, computer repair labs, and power plants.

Investigative: Investigative persons like to think abstractly, solve problems, and investigate. These individuals feel comfortable with the pursuit of knowl- edge and enjoy manipulating ideas and symbols. Investigative individuals prefer to avoid social situa- tions and see themselves as introverted. Some set- tings in which you might find investigative individuals include research laboratories, hospitals, universities, and government-sponsored research agencies.

Artistic: Artistic individuals like to express them- selves creatively, usually through artistic forms such as drama, art, music, and writing. They prefer unstructured activities in which they can use their imagination and express their creativity. Some set- tings in which you might find artistic individuals include the theater, concert halls, libraries, art or music studios, dance studios, orchestras, photogra- phy studios, newspapers, and restaurants.

Social: Social people are nurturers, helpers, and care- givers who have a high degree of concern for others. They are introspective and insightful and prefer work envi- ronments in which they can use their intuitive and caregiv- ing skills. Some settings in which you might find social people include government social service agencies, counseling offices, churches, schools, mental hospitals, recreational centers, personnel offices, and hospitals.

Enterprising: Enterprising individuals are self-confident, adventurous, bold, and enjoy working with other people. They have good persuasive skills and prefer positions of leadership. They tend to dominate conversations and enjoy work environments in which they can satisfy their need for recognition, power, and expression. Some set- tings in which you might find enterprising individuals include life insurance agencies, advertising agencies, political offices, real estate offices, new and used car lots, sales offices, and management positions.

Conventional: Individuals of the Conventional orienta- tion are stable, controlled, conservative, and coopera- tive. They prefer working on concrete tasks and like to follow instructions. They value the business world, cler- ical tasks, and tend to be good at computer program- ming or database operations. Some settings in which you might find conventional people include banks, business offices, accounting firms, computer software companies, and medical records departments.

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

224 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Basic Interest Scales The Basic Interest Scales identify the respondent’s interest in 30 broad areas, such as science, performing arts, marketing, teaching, law, and so on. The individual is shown his or her top five interest areas and can also view the 30 interest areas together with their corresponding Holland Codes (see Figure 10.2). Using T-scores, individuals can compare themselves to the average scores of people of their own gender. The higher the T-score (remember that 50 is the mean and 10 is the standard deviation), the more interest a person will have in a particular area as compared with others. Interest areas greater than one standard deviation (i.e., 60 or above) are considered very high.

Occupational Scales These scales provided the original basis for the 1927 Strong and allow an individual to compare his or her interests to the interests of individuals of the same sex who are satisfied in their jobs. Thus, respondents can compare how they scored on the Strong to the scores of individuals in 244 com- monly held jobs who also took the Strong. The Strong lists the 10 occupations to which the respondent is most similar (see Figure 10.3), and then separately lists the T-scores of the client as compared to all 244 occupations (not shown in the figure). The higher the T-score, the more similar are the respondent’s interests to those of satisfied people in the stated job.

Personal Style Scales These scales give an estimate as to how comfortable the test-taker is in certain activities, including work style (alone or with people), learn- ing environment (practical vs. academic), leadership style (taking charge vs. letting others take charge), risk taking/adventure (risk taker vs. nonrisk taker), and team orientation (working in a team vs. independently). T-scores are used to compare individuals to men, women, or a combined group.

Response Summary The Response Summary provides a percentage breakdown of the client’s responses across all of the six interest areas measured by the Strong (e.g., occupations, subject areas, etc.). The summary offers the total percentages of each of the Likert response types, varying from “strongly like” to “strongly dislike.” In addi- tion, there is a new typicality index that assists in identifying or flagging examinees who may be making random responses, which has been found to be 95% accurate. A score of 16 or lower indicates an inconsistent pattern of item selection (CPP, 2004).

The Strong can be scored in a number of ways. The test booklet can be mailed to the testing center, on-site scoring software can be purchased and used, or the assessment can be administered over the Internet to produce immediate scoring reports. Norm data were updated for the latest version of the Strong and contained an equal number of men and woman (n ¼ 2,250) and a similar distribution of racial and ethnic groups compared to U.S. Census data (CPP, 2004, 2009b). The sample included people representing 370 different occupations. Reliability coefficient alphas for the General Occupational Themes (Holland codes) were between 0.90 and 0.95 with test-retest reliabilities ranging from 0.88 to 0.95. The Basic Interest Scales had a median alpha of 0.87. The Occupational Scales had test-retest reliability median average of 0.86 between tests taken two to seven months apart. The five Personal Style Scales had coefficient alphas that ranged from 0.82 to 0.87. Evi- dence of convergent validity was demonstrated when scores from the most recent

Strong—Basic Interest Scales Identify broad areas of interest

Strong— Occupational Scales Compares interests to those of others with the same job

Strong—Personal Style Scales Assesses work style, learning environment, leadership style, risk taking, and team orientation

CHAPTER 10 Career and Occupational Assessment 225

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

FIGURE 10.2 | Strong Interest Inventory Profile Sheet for Basic Interest Scales Source: CPP. (2012). CPP sample reports. Strong Interest Inventory profile. Retrieved from https://www.cpp.com/pdfs/smp284108.pdf, p. 4.

226 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

version of the General Occupational Themes had a 0.95 correlation with the previ- ous version. Regression analysis of the Basic Interest Scales revealed that they, as a group, explained 68% to 78% of the variance in the broader occupational groups. Evidence of discriminant validity is suggested within the five Personal Style Scales through comparing the divergence of correlations between scales. Results indicate correlations from 0.03 to 0.55, suggesting that each Personal Style Scale is analyzing a unique aspect of preferred work environment.

FIGURE 10.3 | Strong Interest Inventory Profile Sheet for Occupational Scales Source: CPP. (2012). CPP sample reports. In Strong Interest Inventory profile. Retrieved from https://www.cpp.com/pdfs/smp284108.pdf, p. 5.

CHAPTER 10 Career and Occupational Assessment 227

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Self-Directed Search The Self-Directed Search, Fourth Edition (SDS; PAR, 2009) was created by John Holland and is based on the Holland hexagon shown in Figure 10.1. As the name of the instrument suggests, the SDS can be self-administered, self-scored, and self- interpreted, although it is suggested that a counselor guide a client in his or her exploration. Although the instrument is based primarily on interests, it also includes self-estimates of competencies and ability. Once the instrument is scored and the client obtains his or her three-letter Holland code, he or she can cross- reference the code with the Occupations Finder, which classifies more than 1,300 occupations by Holland type, with a book entitled the Dictionary of Holland Occupational Codes (Gottfredson & Holland, 1996), which lists more than 12,000 occupational codes, or by conducting an “Interest Search” on the U.S. gov- ernment sponsored O*NET Online system (see O*NET, n.d.a) (also see “O*NET and Career Exploration Tools” section).

The SDS is available in four forms. Form R (regular), is designed for high school students, college students, and adults; Form E (easy-to-read) is written at the fourth-grade level and can be used with students or adults who have limited read- ing ability; Form CE (career explorer) is for middle school and junior high stu- dents; and Form CP (Career Planning) is designed for professional-level employees (Brown, 2001; PAR, 2012). The SDS can be administered by hard copy booklet, computer software, and Form R is available on the World Wide Web at www.self- directed-search.com (PAR, 2009).

Scoring for the test is done by simply adding the raw scores of each of the six types of personality, with the three highest scores, from highest to lowest, provid- ing the individual’s “Holland code.” Norm data for Form R were based on 2,602 people between the ages of 17 and 65 from 25 different states. Internal reliability coefficients ranged from 0.90 to 0.94 for the combined scales (Brown, 2001). Form E was normed against 719 people from ages 15 to 72 and has slightly higher reliability coefficients than Form R. Form CP was administered to only 101 work- ing people aged 19 to 61, most of whom had been to college. This form had slightly lower reliability estimates than the other forms. Convergent validity of the SDS showed correlations greater than .94 between codes on an earlier version of the SDS with codes on the current version. In fact, half of individuals had the same three-letter Holland code for both, and two-thirds had the same first two Holland codes. Research on the Holland codes has consistently shown moderate to high correlations with job satisfaction (Rottinghaus et al., 2009).

COPSystem The COPSystem is a career measurement package that contains three instruments that can be used individually or together to aid in making career decisions. One instrument measures interests, another abilities, and the third, values (see Table 10.1; EdITS, 2012a). The tests can be hand scored, mailed to EdITS for machine scoring, or scored with on-site scoring software purchased from the pub- lisher. There is an additional web-based version that provides all three instruments in a package called the COPSystem 3C (EdITS, 2012b).

SDS This is a self- administered, scored, and interpreted test that was created by Holland; uses his personality types

COPSystem Three instruments that measure interests, abilities, and values

228 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Career Occupational Preference System Interest Inventory (COPS) The COPS is designed for individuals from seventh grade to adult and consists of 168 items that are based on high school and college curricula as well as sources of occupa- tional information. Scores on the instrument are related to a unique career cluster model that is used to guide the individual to a number of career areas (see Figure 10.4). Although the publisher suggests the instrument can be used for a wide age group, norms are based on high school or college samples (EdITS, 2012c). The instrument takes 20 to 30 minutes to administer and about 15 minutes to hand score. Coefficient alpha reliabilities range between 0.86 and 0.92 and convergent validity has been demonstrated between the COPS and other career inventories (EdITS, 2012c).

TABLE 10.1 COPSystem Assessments and Administration Times

Measurement Acronym Full Name Administration Time (in Minutes)

Interests COPS Career Occupational Preference System Interest Inventory

20 to 30

Abilities CAPS Career Ability Placement Survey 50

Work Values COPES Career Orientation Placement and Evaluation Survey

30 to 40

A R

T S

SE RV

ICE SCIENCE

O U

T D

O O

R

BUS INE

SS CLERICAL

C O

M M

U N

IC A T IO

N

T E C

H N

O L O

G Y

PROFE SS

IO NA

L

P R

O F E

S S

IO N

A L

P

RO FE

SS ION

AL PROFESSIONAL P R O

F E S

S IO

N A

L

SKIL LE

D

S K

IL L E D

SK ILL

ED SKILLED S

K IL

L E

D

C O

N S U

M E

R

E C

O N

O M

IC S

FIGURE 10.4 | COPS Career Clusters Source: EdITS. (2010). COPSystem: Home. Retrieved from http://www.edits.net/component/content/article/10 /352-copsystem-career-wheel-of-career-clusters.html

COPS Assesses interests along career clusters

© Ce ng ag e Le ar ni ng

CHAPTER 10 Career and Occupational Assessment 229

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Career Ability Placement Survey (CAPS) The CAPS test measures abilities across eight different dimensions that relate to the COPS career clusters (outer cir- cle of Figure 10.4). The test allows individuals to identify which career fields are best suited to their abilities or identify careers for which more training may be required. The instrument is designed for middle school students through adults. Norm data are available for eighth- through twelfth-graders and for college stu- dents. Each of the eight subtests takes five minutes to administer, and the total test time is about 50 minutes. Hand scoring requires an additional 15 to 20 minutes. Test-retest reliability has ranged from 0.70 to 0.95 (EdITS, 2012c).

Career Orientation Placement and Evaluation Survey (COPES) COPES is an assessment of values that are important in occupational selection and job satisfaction. Scales are based on eight dichotomous poles (Table 10.2), which then are keyed to the career clusters in Figure 10.4 to assist in using values in choosing a career.

As with the CAPS and COPS, this instrument is designed for middle school stu- dents through adults, however, norm data are available only from high school and college students. Alpha reliabilities are a bit lower for this instrument ranging from 0.70 to 0.83 (EdITS, 2012c).

O*NET and Career Exploration Tools The Occupational Information Network referred to as O*NET is a free online database that contains hundreds of occupational classifications and offers addi- tional self-directed career exploration tools for those contemplating career moves. It is produced and maintained by the U.S. Department of Labor/Employment and Training Administration (USDOL/ETA) and replaced the hard copy Dictionary of Occupational Titles (DOT), which was published from 1938 until the 1990s (Mariani, 1999; O*NET, n.d.b). The O*NET database is used by job seekers, human resource professionals, students, and researchers (O*NET, n.d.c).

The content model for O*NET is based on six major domains (O*NET, n.d.d). Three of the domains are focused on the worker, including worker characteristics, worker requirements, and experience requirements. The other three domains are job-oriented and include occupational requirements, workforce characteristics, and

TABLE 10.2 The COPES Scales

Investigative vs. Accepting

Practical vs. Carefree

Independence vs. Conformity

Leadership vs. Supportive

Orderliness vs. Flexibility

Recognition vs. Privacy

Aesthetic vs. Realistic

Social vs. Reserved

CAPS Measures abilities in the work environment that relate to career clusters

COPES Values in job selection related to career clusters

O*NET Government online database of job descriptions

© Ce ng ag e Le ar ni ng

230 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

occupation-specific information. The model also allows information to be applied across jobs and industries as well as within specific occupations. Figure 10.5 demonstrates how the top three domains revolve around the worker and the bot- tom three around the job. In Table 10.3 you can view the six O*NET domains and their corresponding subparts.

When using O*NET, online you can search for a specific job description, groups of similar careers, or use an advanced search to find all jobs that might use a particular skills, ability, or other domains from the model. If you perform a search for school counselor, you will find an occupation entitled “Educational, Guidance, School, and Vocational Counselors” with an associated standard occu- pational classification code of 21-1012.00. Once you select that title, you will find a summary report for this job providing details for all six of the model domains, which include projected growth and projected job openings. O*NET is not a resource for finding individual job openings but rather a large database about occupations.

The O*NET system offers a series of career exploration tools to assist people wanting to make a work-related transition as well as students preparing to enter the workforce (O*NET, n.d.e). The O*NET team describes these instruments as self- directed; however, most of them require a counselor or teacher to administer and score them. All are paper-n-pencil except for one, and they can be given individually or in a group. The instruments, answering sheets, user guides, and administration manuals are available for free in downloadable PDF formats. However, scannable

Worker-Oriented

Cross Occupation

Job-Oriented

Occupation Specific

Worker Characteristics

Abilities Occupational Interests

Work Values Work Styles

Worker Requirements

Skills-Knowledge Education

Experience Requirements

Experience and Training Skills-Entry Requirement

Licensing

Occupational Requirements

Generalized Work Activities Detailed Work Activities Organizational Context

Work Context

Workforce Characteristics

Labor Market Information Occupational Outlook

Occupation-Specific Information

Tasks Tools and Technology

FIGURE 10.5 | The O*NET Content Model Source: O*NET (n.d.). O*NET Resource Center. Retrieved from http://www.onetcenter.org/content.html

CHAPTER 10 Career and Occupational Assessment 231

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

answer sheets and test booklets can also be ordered for a fee for those wanting the convenience in group settings. The O*NET Career Exploration Tools are the

• Ability Profiler • Interest Profiler • Computerized Interest Profiler • Interest Profiler Short Form • Work Importance Locator • Work Importance Profiler

As you can see, these instruments attempt to capture abilities, interests, and values similar to the COPSystem (see Exercise 10.1).

TABLE 10.3 O*NET Domains and Subparts

Domain Subdomain

Worker Characteristics Abilities Occupational interests—Holland code Work values—theory of work adjustment Work styles

Worker Requirements Basic skills Cross-functional skills Knowledge Education

Experience Requirements Experience and training Basic skills Cross-functional skills Licensing

Occupation-Specific Information Tasks Tools and technology

Workforce Characteristics Labor market information Occupational outlook

Occupational Requirements Generalized work activities Detailed work activities Organizational context Work content

Exercise 10.1 Reviewing Occupations via O*NET Online

In class or individually view the profession you intend to enter (i.e., teacher, school counselor, mental health counselor, etc) using the O*NET online available at www.onetonline.org. Review the tasks, knowledge, skills, abilities, work activities, interest (Holland code),

work styles, work values, and wages and employment trend. Discuss in small groups or as a class. Are there any surprises? Any domains that cause anxiety? How does this fit with your personality? How could you use this in the future with a client?

© Ce ng ag e Le ar ni ng

20 15

© Ce ng ag e Le ar ni ng

20 15

232 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Other Common Interest Inventories Today, dozens of interest inventories are available for client use. However, a few stand out as being especially popular. For instance, the Campbell Interest and Skill Survey (CISS) is a self-report instrument measuring both interests and skills for a vari- ety of occupations. It is intended for people 15 years or older and is primarily for college-bound students or college-educated adults (Pearson, 2012a). The System of Integrated Guidance and Information, Third Edition (SIGI3) is a web-based career self-assessment program primarily for high school and college students. It provides up-to-date career information such as educational requirements, income, job satisfac- tion, and so on and offers assessments in values, interests, personality, and skills (Valpar, 2012). The Career Assessment Inventory—Enhanced Version is another instrument for career development and guidance for individuals 15 years and older. It also uses the familiar Holland codes (realistic, investigative, artistic, social, enter- prising, and conventional) (Pearson, 2012b). Clearly, experts in career and occupa- tional assessment have many good tests from which to choose.

MULTIPLE APTITUDE TESTING Multiple aptitude tests, as the name implies, measure several abilities at one time and are used to predict how an individual might perform in many different types of jobs. For example, the Armed Services Vocational Aptitude Battery (ASVAB), which will be discussed in greater detail, measures auto and shop knowledge, mechanical comprehension, general science, electronics knowledge, and more. This type of information helps military administrators make placement decisions and is useful for those who are considering entering the armed services or who want to get a better idea of the skills they possess. Obviously, someone who has high mechanical aptitude and enjoys that type of work might be better suited to repair- ing jet aircraft or tank engines than working as a desk clerk.

Factor Analysis and Multiple Aptitude Testing As you may recall from Chapter 5, test developers use factor analysis in demon- strating evidence of construct validity for their tests. This common statistical tech- nique allows researchers to analyze large data sets to determine patterns and calculate which variables have a high degree of commonality. Using this process allows multiple aptitude test developers to assess the differences and similarities in the many abilities they are attempting to measure. For example, if you wanted to create a multiple aptitude test that assessed a variety of sports abilities, you could develop a number of tests, such as a softball throw, a 100-meter dash, a long jump, a 5-km run, a baseball throwing accuracy test, and a 50-meter swim. After giving the tests to your norm group and applying factor analysis to your data, you may find that the softball throw and the baseball accuracy test have so much over- lap that you decide to remove one of them from your instrument or combine the two subtests into one. Two commonly used multiple aptitude tests that use factor analysis to show the purity of their component parts include the Armed Services Vocational Aptitude Battery (ASVAB) and the Differential Aptitude Test (DAT).

Multiple aptitude tests Measure abilities and predicts success in several fields

Factor analysis Helps developers determine differences and similarities between subtests

CHAPTER 10 Career and Occupational Assessment 233

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Armed Services Vocational Aptitude Battery and Career Exploration Program The Armed Services Vocational Aptitude Battery (ASVAB) is the most widely used multiple aptitude test in the world (ASVAB, n.d.a). Originally developed in 1968, it has gone through many improvements since that time. The latest revision is called the ASVAB Career Exploration Program (CEP) and consists of three primary com- ponents: the traditional ASVAB aptitude test available in both computerized and paper-n-pencil formats, an interest inventory called Find Your Interests (FYI), and a career exploration tool called OCCU-Find (ASVAB, n.d.b). The ASVAB CEP also includes other components such as a career exploration guide, online resources, and a short video program.

The ASVAB is now offered in two formats: the paper-n-pencil version and a new computerized format (CAT-ASVAB). Both versions produce similar results but are scored differently. The CAT-ASVAB uses an adaptive administration sys- tem that tailors the exam according to the examinees abilities (ASVAB, n.d.c). It pulls items from a pool of questions and provides more difficult or easier items depending on the previous answer the examinee made. For example, if you answer the first question on a section correct, the computer will give you a more difficult question next. If you miss the second question, then it provides an easier question, and so on. In Figure 10.6 you can see how the test can quickly narrow in on one’s ability. Hence, the CAT-ASVAB takes less time and can usually be done in 1.5 hours whereas the paper-n-pencil version takes about 3 hours (ASVAB).

The ASVAB (n.d.c) battery consists of 10 “power tests,” a term used to describe a test that has generous time limits (in contrast to a “speed test” such as a 1-minute typing test). The 10 subtests fall into one of four domains, which are ver- bal, math, science and technical, and spatial. In Table 10.4 you can see which tests contribute to which domains. On the paper-n-pencil version, the auto information (AI) and shop information (SI) are combined into one auto and shop (AS) score. The paper-n-pencil test contains 225 items whereas the CAT-ASVAB has 145 ques- tions (see typical sample questions, Box 10.2). The composite scores can be associ- ated with job classifications in the U.S. Department of Labor’s Occupational Information Network (O*NET). Although the ASVAB began as a test strictly for the military, over the years it has been altered and is now useful for both civilian and military occupations.

ASVAB scores were normed against two large samples that in combination contained approximately 10,700 men and women ranging from 10th grade up to age 23 from the Profile of American Youth project in 1997. The 18 and older group was oversampled for Hispanic and African-American youth (ASVAB, n.d.a). ASVAB developers employed item response theory (IRT) techniques to estimate reliabilities, which ranged between 0.69 and 0.88 for the eight subscales for 10th, 11th, and 12th grades. The three composite scores (math, verbal, and science and technical skills) ranged between 0.88 and 0.91. Content validity and content con- struction focused on knowledge required to perform specific jobs rather than the entire content domain. For example, instead of testing the complete domain of math, the math domain is geared toward knowledge and skills needed in the typi- cal work environment. Numerous validity studies show correlations that range

ASVAB Measures many abilities required for military and civilian jobs

234 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Initial Item

Harder Item

Yes

Yes

No

No

Yes

No

Correct?

Easier Item

Yes

No

Correct?

Yes

No

Correct?

Yes

No

Correct?

Yes

No

Correct?

Correct?

Easier Item

Harder Item

Correct?

Harder Item

Easier Item

FIGURE 10.6 | CAT-ASVAB Adaptive Item Process System Source: Armed Service Vocational Aptitude Battery. (n.d.c). ASVAB fact sheet (p. 3). Retrieved from http://official-asvab.com /docs/asvab_fact_sheet.pdf

TABLE 10.4 ASVAB Domains and Tests

Domain Test

Verbal Word knowledge (WK) Paragraph comprehension (PC)

Math Arithmetic reasoning (AR) Mathematics knowledge (MK)

Science and Technology General science (GS) Electronics information (EI) Auto information (AI) Shop information (SI) Mechanical comprehension (MC)

Spatial Assembling objects (AO)

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 10 Career and Occupational Assessment 235

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

BOX 10.2 Sample ASVAB Questions

General Science 1. An eclipse of the sun throws the shadow of the A. moon on the sun. B. moon on the earth. C. earth on the sun. D. earth on the moon.

2. Substances that hasten chemical reaction time without themselves undergoing change are called

A. buffers. B. colloids. C. reducers. D. catalysts.

Word Knowledge 3. The wind is variable today. A. mild B. steady C. shifting D. chilling

Math Knowledge 4. If 50% of X ¼ 66, then X ¼ A. 33 B. 66 C. 99 D. 132

5. What is the area of this square? A. 1 square foot B. 5 square feet C. 10 square feet D. 25 square feet

Electronics Information 6. Which of the following has the least resistance? A. wood B. iron C. rubber D. silver

7. In this circuit diagram, the resistance is 100 ohms, and the current is 0.1 amperes. The voltage is

A. 5 volts. B. 10 volts. C. 100 volts. D. 1,000 volts.

Automotive and Shop Information 8. A car uses too much oil when which of the

following parts are worn? A. pistons B. piston rings C. main bearings D. connecting rods

9. A chisel is used for A. prying. B. cutting. C. twisting. D. grinding.

Mechanical Comprehension

10. In this arrangement of pulleys, which pulley turns fastest?

A. A B. B C. C D. D

Solutions

1. B; 2. D, 3. C; 4. D; 5. D; 6. D; 7. B; 8. B; 9. B; 10. A

Source: Adapted from ASVAB Career Exploration Program (n.d.). The ASVAB test: Sample ASVAB test questions. Retrieved from http://www.asvabprogram.com/index.cfm?fuseaction= overview.testsample

5 FEET

R

AD

B

C

236 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

from a moderate 0.36 to an impressive 0.77 for predicting success in military occu- pations (ASVAB, 2005). Regarding construct validity, one study found a 0.79 cor- relation between ASVAB and ACT scores. Another study found correlations ranging from 0.70 to 0.86 with the California Achievement Test.

In addition to the traditional ASVAB, the “ASVAB CEP” offers the FYI, an interest inventory to assist high school students in determining what they would like to do rather than what they can do (ASVAB, n.d.b). It is based on Holland’s RIASEC personality types and provides individuals their three highest codes. FYI contains 90-items and can be taken online at the ASVAB Web site or with paper- n-pencil. The paper-n-pencil version takes approximately 15 minutes to complete and score (ASVAB, 2005; n.d.a). It uses a three-point Likert scale asking indivi- duals if they would like, be indifferent to, or dislike certain activities. Coefficient alpha reliabilities ranged from 0.92 to 0.94 whereas test-retest reliabilities ranged from 0.89 to 0.93 for each of the Holland codes (ASVAB, 2005; n.d.a). Evidence of validity has been demonstrated for the FYI through several methods. Adherence to proper development processes suggests adequate content validity. Factorial anal- ysis suggesting internal structure and correlation with the SII help demonstrate con- struct validity. Also, the FYI correlations range between 0.68 and 0.85 with the Strong General Occupation Themes, between 0.44 and 0.59 with the Strong Basic Interest Scales, and between 0.54 and 0.61 with the Strong Occupation Scales. Although useful for any student, the FYI can be particularly helpful for the approx- imately 60% of students who do not go on to college (National Center for Educa- tion Statistics, 2011), because they can be hopeful about their futures when they see occupations that are in range of their interests and skills.

Once a person has assessed his or her ability through the ASVAB and interests through the FYI, he or she can now explore career options through the OCCU- Find tool. OCCU-Find allows a user to compare his or her skills and interests with over 400 jobs from the O*NET database for both civilian and military equivalents (ASVAB, n.d.a). It uses the knowledge, skills, and abilities from O*NET to compare to respondents RIASEC code and ASVAB subtest scores with a relatively high degree of accuracy (ASVAB, 2011).

Differential Aptitude Tests The DAT, Fifth Edition, is a long-standing series of tests for students in grades 7 through 12 that measures adolescents’ abilities across vocational skills (Hattrup & Schmitt, 1995; Pearson, 2008). Often administered and interpreted by school coun- selors, the DAT takes approximately 1.5 to 2.5 hours to complete, depending on whether the full or abbreviated version is used. The DAT has eight separate tests of ability that measure verbal reasoning, numerical reasoning, abstract reasoning, perceptual speed and accuracy, mechanical reasoning, space relations, spelling, and language usage. The DAT also includes a Career Interest Inventory (CII) to deter- mine what careers a person might like. It is easy to see how both sets of these scores can be useful to a school counselor when assisting a student in determining a possible occupation or college major. Raw scores are converted to percentiles and stanines.

DAT Measures abilities and interests to assist with career decision making

CHAPTER 10 Career and Occupational Assessment 237

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Internal consistency reliability measures for the DAT are high, ranging between 0.80 and 0.95 for the different tests. Regarding construct validity, correlations with the DAT and several other major aptitude tests (ACT, ASVAB, SAT, and the California Achievement Test) range between 0.68 and 0.85 (Hattrup & Schmitt, 1995). Although correlations with DAT scores and high school grades were sound, no data have been provided regarding predictive validity in job perfor- mance. A healthy norm sample of approximately 170,000 students proportionately represented by geography, gender, socioeconomic status, ethnicities, and other fac- tors was used for the fifth edition.

An alternative version of the DAT for adults is called the Differential Aptitude Battery for Personnel and Career Assessment (DAT PCA). The DAT PCA measures ability and aptitude across eight different areas, similar to the DAT, and is often used for hiring purposes. It allows employers to determine one’s current ability and aptitude for learning new skills during training. Like the DAT, the DAT PCA uses percentiles and stanines. Reliability estimates are sound, ranging between 0.88 and 0.94 (Wilson & Wing, 1995). The norm sample consisted of 12th-grade stu- dents, which is a bit of a leap from the adult population for which the test is intended. One concern of this test is the fact that predictive validity correlations comparing DAT PCA scores with grades, job supervisors’ ratings, and job perfor- mance were low, ranging from 0.1 to 0.4.

SPECIAL APTITUDE TESTING As noted earlier, special aptitude tests measure a homogenous area of ability and are generally used to predict success in a specific vocational area of interest. Thus, they are frequently used as a screening process to assess one’s ability to perform a certain job or to master a new skill at work. Hiring and training employees is an expensive operation for any size company, so you can imagine how useful these tests might be. Similarly, specialized vocational training in areas such as art, music, plumbing, and mechanics all require aptitudes that not everyone possesses. Hence, educational institutions, vocational training institutes, and others frequently rely on special aptitude testing during the admission process.

The tests we briefly examine in this section include the Clerical Test Battery, the Minnesota Clerical Assessment Battery, the U.S. Postal Service’s (USPS) 473 Battery Examination, the Mechanical Aptitude Test, the Wiesen Test of Mechanical Aptitude (WTMA), the Arco Mechanical Aptitude and Spatial Relations Tests, the Technical Test Battery, the Bennett Test of Mechanical Comprehension, the Music Aptitude Profile, the Iowa Test of Music Literacy, the Keynotes Music Evaluation Software Kit, the Group Test of Musical Ability, and the Advanced Measures of Music Audiation.

Clerical Aptitude Tests Several tests are available to measure one’s aptitude at performing clerical tasks. The Clerical Test Battery (CTB2), published by Psytech (2010), measures clerical skill across a range of abilities, including verbal reasoning, numerical ability, cleri- cal checking, spelling, typing, and filing. The CTB2 takes only 27 minutes to

DAT PCA A form of the DATA used by employers to assess ability

Special aptitude tests Designed to predict success in a vocational area

Clerical aptitude tests Used for screening applicants for clerical jobs

238 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

administer, and it can be completed with paper and pencil or on a personal com- puter. Reliability estimates for the subtests are sound and fall between 0.81 and 0.90. Reasonable evidence of validity is provided in the test manual. A similar type of clerical test is the Minnesota Clerical Assessment Battery, which measures traits such as typing, proofreading, filing, business vocabulary, business math, and clerical knowledge. Reliability estimates for the subtests are acceptable; however, Fitzpatrick (2001) questioned whether the publisher had provided adequate evi- dence of validity.

One of the nation’s largest employers, the USPS, has over 700,000 employees (USPS, 2013). Ninety percent of newly hired postal workers are required to take their entrance exam, called Test 473, which has replaced earlier exams (Pathfinder, 2013). The entry-level jobs requiring the exam include positions such as clerk, mail handler, carrier, mark-up clerk, mail processor, flat sorter, and distribution clerk. The “473” is given in five sections that measures aptitudes of address checking, correctly completing forms, identifying correct zip codes, memorizing codes, and a personal inventory of job-related characteristics and experiences. The instrument has a total 398 items and examinees are given 129 minutes to complete the five sections. A score of 70% is considered passing, but those hired generally have scores 90% or greater (USPS).

Mechanical Aptitude Tests Have you ever known people who could look at almost any mechanical-related problem (stalled car, leak underneath the sink, malfunctioning household appliance) and fix it? Obviously there is some learned information (crystallized intelligence) involved, but some people just seem to have a knack for it (fluid intel- ligence). Mechanical aptitude is generally considered the ability to learn physical and mechanical principles and manipulate mechanical objects. Put another way, it is the ability to understand how mechanical things work. Some people seem to have it, and some don’t. (Ed and Charlie are in the second category!)

Many mechanical aptitude tests are available today, and small manufacturing companies, governmental agencies, and technical institutes frequently use them when they want to measure this ability before hiring or training someone. The Mechanical Aptitude Test (MAT-3C; Ramsay Corporation, 2013) was developed to reduce the impact of gender and race on the measurement of mechanical apti- tudes. The 36 items were written to relate to everyday things and focus on house- hold objects, work production, science and physics, and hand and power tools. This instrument is offered in paper-n-pencil format or online and can be taken in 20 minutes. Reliability alphas in one published sample were 0.72. The test has been shown to correlate with technical college GPAs, mechanical knowledge, man- ager ratings of employees (predictive validity), and other measures of mechanical aptitude (convergent validity).

The Wiesen Test of Mechanical Aptitude (WTMA) (Criteriacorp, 2012) is a similar test with a focus on predicting job performance in the operation and main- tenance of tools, machinery, and equipment. It has 60 items and takes 30 minutes to complete. It attempts to remove gender and race bias and has a Spanish version. The reliability with the norm sample was 0.97, and it had a high correlation with

Mechanical aptitude tests Measure ability to learn mechanical principles and manipulate me- chanical objects

CHAPTER 10 Career and Occupational Assessment 239

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

the Bennett Mechanical Comprehension Test. Other common mechanical ability tests include the Technical Test Battery (TTB2), the Arco Mechanical Aptitude and Spatial Relations Tests, Fifth Edition, and the Bennett Test of Mechanical Comprehension.

Artistic Aptitude Tests Rating artistic ability is not an easy task, and although some tests have been devel- oped (see Figure 10.7), they have not been widely used and have questionable reli- ability and validity (e.g., Meir Art Test, Graves Design Judgment Test). To demonstrate artistic ability, professional art schools require applicants to submit a portfolio. One of the drawbacks to this process is that, usually, a faculty member must subjectively score each portfolio, which can cause problems with reliability. One way to improve reliability is to have two or more individuals (raters) practice rating items similar to the ones they will eventually rate. Generally, as the parameters of what they are rating become more clearly defined (e.g., form, appeal, use of color, and so on), their ratings should increasingly be similar. Once they reach a high degree of agreement in their ratings, they are ready to rate the students’ pieces (called “interrater reliability,” this concept will be discussed in more detail in Chapter 12).

Movement, Depth, Height, Spacing

Personal Information

Age:

Formal Art Training:

How many years have you been drawing:

What would you like to achieve as an artist?

Address:

Email:

Phone:

Answer:

Answer:

Answer:

Answer:

FIGURE 10.7 | Commonly seen art test graphics from local art schools

© Ce ng ag e Le ar ni ng

20 15

Artistic aptitude tests Frequently used for art school admissions

240 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Musical Aptitude Tests One of the challenges of measuring musical aptitude is to be able to distinguish between one’s ability to learn a distinct knowledge base, such as musical theory, history, and rhythm and cadence, and one’s ability to actually play an instrument. Although several tests attempt to distinguish the ability to learn information from the ability to perform, they do so with some difficulty.

The Music Aptitude Profile is probably the most researched test of this type on the market (GIA Publications, 2013a). It is designed for students in grades 4 through 12, involves having the student listen to two musical excerpts and then answer a series of questions pertaining to the music. Traits that are measured include the student’s ability to assess harmony, tone, melody, tempo, balance, and style. Unfortunately, test administration is lengthy, running 2.5 hours. Split-half reliability coefficients were 0.90 for composite scores. Although evidence of validity appears fair, when comparing this test with what is defined as musical intelligence by Gardner (see Chapter 9), there is some question as to whether it actually con- forms to Gardner’s understanding of this trait (Cohen, 1995).

Another music test designed to measure students’ “audiation,” or the process of thinking about music in one’s mind without it actually being there, is the Iowa Test of Music Literacy (Revised). Although it may be classified as an achievement test, it can also be used to determine a student’s aptitude for music training. In this test, the exam- inee not only listens to music but is also required to read sheet music and compare what he or she is reading with what is being played (GIA Publications, 2013b; Radocy, 1998). Another music test, the Keynotes Music Evaluation Software Kit, is designed for classroom teachers to assess their students’ musical aptitude (Brunsman & Johnson, 2000). For students aged 9 to 14, this 65-item instrument is taken on a computer with speakers or headphones and provides scores for pitch discrimination, pattern recogni- tion, and the ability to read music. Unfortunately, no reliability or validity data is sup- plied in the manual (Brunsman & Johnson). Still other musical aptitude tests include the Group Test of Musical Ability and the Advanced Measures of Music Audiation.

THE ROLE OF HELPERS IN OCCUPATIONAL AND CAREER ASSESSMENT

A wide range of helpers provide occupational and career assessment. For instance, middle school counselors administer interest inventories to help students begin to examine their occupational likes and dislikes. High school counselors and college counselors provide such inventories to help students begin to think about occupa- tional choices and to help them make tentative choices about their college major. High school counselors can also be found orchestrating the administration of mul- tiple aptitude tests and helping students interpret those tests to identify possible vocational strengths. Similarly, private practice clinicians can be found giving inter- est inventories and aptitude tests to help their clients examine what they are good at and to identify which occupational areas might best fit their personalities. Today, the administration of career and occupational assessment has become big business, and we can even find businesses that specifically focus on such assess- ments. Although many of these tests are used by a wide range of helpers and do

Musical aptitude tests Assess knowledge of music

CHAPTER 10 Career and Occupational Assessment 241

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

not require advanced training to administer and interpret, any professional who administers such instruments should obtain the necessary basic training to be able to administer them accurately and interpret them wisely.

FINAL THOUGHTS CONCERNING OCCUPATIONAL AND CAREER ASSESSMENT

The tests discussed in this chapter can be fundamental in helping an individual make important decisions regarding occupational and career choices. However, as with all testing, occupational and career assessment should not be done in a vac- uum. Understanding the complexities of why one has specific interests and abilities and why one ultimately chooses a particular occupation or career can be related to a myriad of reasons including psychodynamic issues (e.g., parental influences), social pressures (e.g., racism, sexism, and peer pressure), environmental concerns (e.g., the economy), and family issues (e.g., sibling order). As colleges and universi- ties have increasingly moved toward “career management centers,” as opposed to career counseling centers, these issues have become paramount. Choosing a career or occupation should never be based solely on taking a test.

SUMMARY We began this chapter by highlighting the fact that career and occupational assessment can occur at any point in an individual’s life but is often most critical at transitional points. We noted that interest inventories, multiple aptitude tests, and special aptitude tests are often used to assist an individual during these transitions and that these tests are best used within a counseling context. We defined interest inventories as per- sonality tests that measure likes and dislikes as well as one’s personality orientation toward the world of work; multiple aptitude tests as tests that measure a number of homogeneous areas of ability and are used to predict likelihood of suc- cess in any of a number of vocational areas; and special aptitude tests as tests that usually measure one homogenous area of ability and are used to predict success in a specific vocational area. First, we looked at interest inventories, then examined multiple aptitude tests, and lastly took a quick look at some special aptitude tests.

The first interest inventory that we examined was the Strong, one of the oldest career interest inventories. Its newest version focuses on five areas. The General Occupational Theme uses a

Holland code to match personality type with occupations. The Basic Interest Scales also use the Holland code, this time to show how closely an individual’s personality aligns with 30 broad interest areas. The Occupational Scales allow an individual to compare his or her interests to the interests of same-sex individuals from 244 occu- pations. The Personal Style Scales give estimates as to the test-taker’s work style (alone or with people), learning environment (practical vs. aca- demic), leadership style (taking charge vs. letting others take charge), risk taking/adventure orienta- tion (risk taker vs. nonrisk taker), and team ori- entation (working in a team vs. independently). Finally, the Response Summary allows the exam- iner to see how the client responded to the test overall and can be important in flagging indivi- duals who are making random responses.

The SDS, which was created by John Holland, also uses the Holland Codes and provides a per- sonality orientation based on the client’s self- estimates of interests, competencies, and ability. This easy-to-use instrument comes in a number of forms depending on the age and professional level of the client and can be cross-referenced with

242 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

occupations having a Holland code similar to the client’s code.

Next, we examined the COPSystem, which is a career measurement package that contains three instruments: the Career Occupational Preference System Interest Inventory (COPS) measures inter- ests, the Career Ability Placement Survey (CAPS) measures ability, and the Career Orientation Placement and Evaluation Survey (COPES) mea- sures values. The scores on these tests can be cross-referenced with a number of career clusters that define occupations in key areas.

We also discussed the O*NET database, which is produced by the U.S. Department of Labor and contains detailed descriptions of hun- dreds of jobs. The database is built on six domains with three orienting around worker characteristics and three around job-related traits. O*NET also contains a series of career exploration tools, which can be downloaded for free or ordered for a small fee. The tools assess examinees abilities, interests, and work-related values. We ended this section by discussing a few other common interest inventories such as the Campbell Interest and Skill Survey (CISS), the System of Integrated Guidance and Information, Third Edition (SIGI3), and the Career Assessment Inventory—Enhanced Version.

In the next part of the chapter, we took a look at two Multiple Aptitude Tests: the Armed Services Vocational Aptitude Battery (ASVAB) Career Exploration Program and the Differential Aptitude Test (DAT). We pointed out that these assessments often use factor analysis to assure that the various tests that collectively make up the whole instrument are unique and share little in common with one another.

The ASVAB, which is the most widely used multiple aptitude test, is now available in a comput- erized version and the traditional paper-n-pencil for- mats. We learned the CAT-ASVAB uses an adaptive administration system, which allows for shorter administration periods. The exam consists of 10 power tests in the areas of general science, arithme- tic reasoning, word knowledge, paragraph compre- hension, mathematics knowledge, electronics information, auto information, shop information, mechanical comprehension, and assembling objects.

Scores for each of the 10 tests, as well as three career exploration (composite) scores, are provided to the examinee. The ASVAB CEP has an interest inventory called FYI (Find Your Interests), which uses the Holland RIASEC types, to assist students in determining their interests as a supplement to the ASVAB data regarding their abilities. The OCCU- Find tool allows users to compare their composite ASVAB scores to related jobs in the O*NET data- base to help them explore jobs they might enjoy.

The Differential Aptitude Test (DAT) is for students in grades 7 through 12 and consists of eight separate tests that measure verbal reasoning, numerical reasoning, abstract reasoning, percep- tual speed and accuracy, mechanical reasoning, space relations, spelling, and language usage. An alternative version of the DAT for adults called the Differential Aptitude Battery for Personnel and Career Assessment (DAT PCA) is also avail- able and allows employers to determine an indivi- dual’s current ability and aptitude for learning new skills during training.

In the last part of the chapter, we reviewed a number of different kinds of special aptitude tests that are frequently used as a screening process to assess one’s ability to perform a certain job or to master a new skill at work. We briefly examined special aptitude tests in the areas of clerical ability, mechanical ability, music ability, and artistic ability.

As the chapter neared its conclusion, we noted that a wide variety of helpers offer career and occupational assessment. Although advanced training is not needed to administer and interpret such instruments, these helpers must obtain the skills needed to be able to administer these tests accurately and interpret the results wisely.

The chapter concluded with a discussion of how occupational and career assessment should not be done in a vacuum. We pointed out that there are frequently a myriad of reasons why a person chooses an occupation or a career, includ- ing psychodynamic issues, social pressures, envi- ronmental concerns, and family issues. We concluded that choosing an occupation or a career should never be the result of only taking a test and that counseling, in addition to assess- ment, should often be used.

CHAPTER 10 Career and Occupational Assessment 243

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

CHAPTER REVIEW 1. Distinguish between interest inventories,

multiple aptitude tests, and special aptitude tests.

2. On the Strong Interest Inventory, there are five different types of interpretive scales or indexes. Describe each of them: a. General occupational themes b. Basic interest scales c. Occupational scales d. Personal style scales e. Response summary

3. The Self-Directed Search uses the Holland code to assist in identifying possible person- ality fit with occupations. Describe how it does this.

4. Describe some possible advantages of the COPSystem over the Strong and the Self- Directed Search.

5. What information does O*NET provide and why would someone want to use this database?

6. Discuss why factor analysis is important to the development of a good multiple aptitude test.

7. Compare and contrast the different subtests of the ASVAB and the DAT. Which test would you prefer to take?

8. Describe the advantage of a multiple aptitude test, like the DAT or ASVAB, having an interest inventory attached to it.

9. Identify three or four special aptitude tests and discuss why they might be used in lieu of a multiple aptitude test.

10. What might be some disadvantages of special aptitude tests that measure vague constructs like music and art?

11. Discuss the role of the helper in the selection, administration, and interpretation of career and occupational assessment instruments. What kind of training do you think is neces- sary for an examiner in this area of assessment?

REFERENCES ACT. (2009). ACT interest inventory technical man-

ual. Retrieved from http://www.act.org/research /researchers/pdf/ACTInterestInventoryTechnicalManual .pdf

Armed Service Vocational Aptitude Battery (ASVAB). (n.d.a). The ASVAB career exploration program: Counselor manual. Retrieved from http://asvabprogram .com/downloads/asvab_counselor_manual.pdf

Armed Service Vocational Aptitude Battery (ASVAB). (n.d.b). ASVAB career exploration program—program at-a-glance. Retrieved from http://asvabprogram.com /downloads/ASVAB_Fact%20Sheet.pdf

Armed Service Vocational Aptitude Battery (ASVAB). (n.d.c). ASVAB fact sheet. Retrieved from http:// official-asvab.com/docs/asvab_fact_sheet.pdf

Armed Service Vocational Aptitude Battery (ASVAB). (2005). ASVAB career exploration program: Coun- selor manual. Retrieved from http://www.dsusd.us /users/christopherg/asvab_counselor_manual.pdf

Armed Service Vocational Aptitude Battery (ASVAB). (2011). The ASVAB career exploration program: Theoretical and technical underpinnings of the

revised skill composites and OCCU-Find. Retrieved from http://asvabprogram.com/downloads/Technical _Chapter_2010.pdf

Brown, M. (2001). Review of the self-directed search, 4th edition [Forms R, E, and CP]. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measure- ments yearbook (pp. 1105–1107). Lincoln, NE: Buros Institute of Mental Measurements.

Brunsman, B. A., & Johnson, C. (2000). Review of the Keynotes Music Evaluation Software Kit. In B. S. Plake, J. C. Imapara, & R. A. Spies (Eds.), The fifteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements.

Cohen, A. (1995). Review of the Musical Aptitude Pro- file 1988 revision. In J. C. Conoley & J. C. Impara (Eds.), The twelfth mental measurements yearbook (pp. 663–666). Lincoln, NE: Buros Institute of Mental Measurements.

CPP. (2004). Technical brief for the newly revised Strong Interest Inventory Assessment: Content, reli- ability, and validity. Retrieved from https://www .cpp.com/Pdfs/StrongTechnicalBrief.pdf

244 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

CPP. (2009a). Strong Interest Inventory. Retrieved from https://www.cpp.com/products/strong/index.aspx

CPP. (2009b). Validity of the Strong Interest Inven- tory® instrument. Retrieved from https://www.cpp .com/Products/strong/strong_info.aspx

Criteriacorp. (2012). Wiesen Test of Mechanical Apti- tude (WTMA). Retrieved from http://www.criteria corp.com/solution/wtma.php

EdITS. (2012a). COPSystem Career Measurement Package. Retrieved from http://www.edits.net /products/copsystem.html

EdITS. (2012b). COPSystem 3C. Retrieved from http:// www.edits.net/products/copsystem/365-copsystem -3c-institutional.html

EdITS. (2012c). EdITS supplemental test information & resources. A brief summary of the reliability and validity of the COPSystem assessment. Retrieved from http://www.edits.net/resourcecenter/testing -supplementals/63-newsletter-1.html

Fitzpatrick R. (2001). Review of the Minnesota Clerical Assessment battery. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements year- book (pp. 769–771). Lincoln, NE: Buros Institute of Mental Measurements.

GIA Publications. (2013a). Musical Aptitude Profile. Retrieved from http://www.giamusic.com/products /P-musicaptitudeprofile.cfm

GIA Publications. (2013b). Iowa tests of music literacy—test manual. Retrieved from http://www .giamusic.com/search_details.cfm?title_id=5973

Gottfredsonn, G. D., & Holland, J. (1996). Dictionary of Holland occupational codes (3rd ed.). Lutz, FL: Psychological Assessment Resources.

Hattrup, K., & Schmitt, N. (1995). Review of the Dif- ferential Aptitude Test, Fifth edition. In J. C. Conoley & J. C. Impara (Eds.), The twelfth mental measure- ments yearbook (pp. 301–305). Lincoln, NE: Buros Institute of Mental Measurements.

Kelly, K. (2003). Review of the Strong Interest Inven- tory. In B. S. Plake, J. C. Imapara, & R. A. Spies (Eds.), The fifteenth mental measurements yearbook (pp. 893–897). Lincoln, NE: Buros Institute of Mental Measurements.

Mariani, M. (1999). Replace with a database: O*NET replaces the Dictionary of Occupational Titles. Occupational Outlook Quarterly, 43(1), 3–9. Retrieved from http://www.bls.gov/opub/ooq/1999 /Spring/art01.pdf

National Center for Education Statistics (2011). Fast facts. Retrieved from http://nces.ed.gov/fastFacts /display.asp?id=98

O*NET. (n.d.a). Browse by O*NET data: Interests. Retrieved from http://www.onetonline.org/find /descriptor/browse/Interests/

O*NET. (n.d.b). O*NET Resource Center: About O*NET. Retrieved from http://www.onetcenter .org/overview.html

O*NET. (n.d.c). O*NET Resource Center: Build your future with O*NET online. Retrieved from http:// www.onetonline.org/

O*NET. (n.d.d). O*NET Resource Center. The O*NET detailed content model. Retrieved from http://www .onetcenter.org/dl_files/ContentModel_Detailed.pdf

O*NET. (n.d.e). O*NET Resource Center: O*NET career exploration tools. Retrieved from http:// www.onetcenter.org/tools.html

PAR. (2009). Discover the careers that best match your interests and abilities. Retrieved from http://www .self-directed-search.com/default.aspx

PAR. (2012). Vocational (career interest/counseling). Retrieved from http://www4.parinc.com/products/Pro- ductListByCategory.aspx?Category=VOCATIONAL& SubCategory=CAREER_COUNSELING

PAR. (2013). John Holland’s Self Directed Search. Retrieved from http://www.self-directed-search .com/default.aspx

Pathfinder. (2013). Postalexam.com: Postal exam 473/473E. Retrieved from http://www.postalexam .com/postal-exams/details/postal-exam-473-473e/

Pearson. (2008). Differential Aptitude Tests, Fifth edi- tion (DAT®). Retrieved from http://pearsonassess .com/HAIWEB/Cultures/en-us/Productdetail.htm?Pid= 015406047X&Mode=summary

Pearson. (2012a). CISS: Campbell interest and skill survey. Retrieved from http://psychcorp.pearsonassessments .com/HAIWEB/Cultures/en-us/Productdetail.htm?Pid= PAg115

Pearson. (2012b). Career assessment inventory-the enhanced version. Retrieved from http://psychcorp .pearsonassessments.com/HAIWEB/Cultures/enus /Productdetail.htm?id=PAg112

Psytech. (2010). Clerical Test Battery 2 (CTB2). Psy- techInternational. Retrieved from http://psytech. com/assessments-CTB2.php

Radocy, R. (1998). Review of the Iowa test of music literacy revised. In J. C. Impara & B. S. Plake (Ed.), The thirteenth mental measurements year- book (pp. 552–555). Lincoln, NE: Buros Institute of Mental Measurements.

Ramsay Corporation. (2013). Mechanical Aptitude Test MAT-3C RR110-A. Retrieved from http:// www.ramsaycorp.com/catalog/view/?productid=25

CHAPTER 10 Career and Occupational Assessment 245

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Rottinghaus, P. J., Hees, C. K., & Conrath, J. A. (2009). Enhancing job satisfaction perspectives: Combining Holland themes and basic interests. Journal of Vocational Behavior, 75(2), 139–151. doi:10.1016/ j.jvb.2009.05.010

Strong, E. (1926). An interest test for personnel man- agers. Journal of Personnel Research, 5, 194–203.

U.S. Postal Service (USPS). (2013). United States Postal Service: About. Publication 60-A-Test 473 orienta- tion guide for major entry-level jobs. Retrieved from

http://about.usps.com/publications/pub60a/welcome .htm

Valpar. (2012). SIGI 3: Educational and career plan- ning software for the web. Retrieved from http:// www.sigi3.org/

Wilson, V., & Wing, H. (1995). Review of the Differential Aptitude Test for personnel and career assessment. In J. C. Conoley & J. C. Impara (Eds.), The twelfth men- tal measurements yearbook (pp. 305–309). Lincoln, NE: Buros Institute of Mental Measurements.

246 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 11Clinical Assessment:Objective and Projective Personality Tests

Several years ago I was asked to assess a client who had been diagnosed with a dissociative identity disorder. At our first meeting, she told me that she had 254 personalities. I was fas- cinated. However, I was not testing her for the disorder, as she was in treatment, accepted her disorder, and was working hard on integrating her personalities. However, she had been denied Social Security disability compensation, and I had been asked to assess her ability to work. Thus, I needed to carefully choose the tests that 1 would use to assess her. When I had finished administering the tests, I concluded that she was not able to work at that time, but in a year she might be able to do so. She was not happy with my conclusion.

One of my former students was doing her dissertation on the relationship between self- actualizing values and the number of years one had meditated. She went to an Ashram and obtained permission from dozens of meditators to participate in her study. Using a test to mea- sure self-actualizing values, she tested the meditators and correlated their test results with the number of years they had been meditating. She was quite surprised to find no relationship and subsequently came up to me and said, “There must be something wrong with this study, because I know there is a relationship!” I suggested that there was nothing wrong with her study and that not finding a relationship did not mean that meditation wasn’t worthwhile, and indeed, it had already been found to be related to many attributes such as reduced stress levels. She ended up a bit dejected, but with a broader understanding of the limitations of objective tests.

I once was asked to do a broad personality assessment for a high school student some teachers were concerned about. I conducted a clinical interview and gave him some objective and projective tests and was somewhat surprised when a number of the projective tests clearly had a theme of destruction. I duly noted this in my report and warned the school about what I had found. Shortly afterward, he was arrested for property destruction!

(Ed Neukrug)

247

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

This chapter is about clinical assessment, which often includes a clinical interview, objective testing, and/or projective testing. As you can see from the vignettes, such assessments have a wide variety of applications and can be an important tool for the clinician, educator, or researcher. In this chapter, we define clinical assessment and examine some of the major objective and projective tests used in the clinical assessment process.

DEFINING CLINICAL ASSESSMENT Clinical assessment involves assessing clients through one or more of the following methods: the clinical interview (see Chapter 4), the administration of informal assessment techniques (Chapter 12), and the administration of objective and pro- jective tests (this chapter). This process is used to gather information from clients for the following purposes:

1. to help clients gain greater insight; 2. to aid in case conceptualization and mental health diagnostic formulations; 3. to assist in making decisions concerning the use of psychotropic medications; 4. to assist in treatment planning; 5. to assist in court decisions (e.g., custody decisions, testing a defendant in a

child molestation case); 6. to assist in job placement decisions (e.g., candidates for high-security jobs); 7. to aid in diagnostic decisions for health-related problems (e.g., Alzheimer’s);

and 8. to identify individuals at risk (e.g., students at risk for suicide or students with

low self-esteem).

In this chapter, we look only at a small number of objective and projective tests used in the clinical assessment process, but remember that clinical assessment can involve a broad range of assessment instruments. Let’s start with a review of some of the more popular objective personality tests.

OBJECTIVE PERSONALITY TESTING As discussed in Chapter 1, objective personality testing is a type of personality assessment that mostly uses multiple-choice, true/false, and related types of for- mats to assess various aspects of personality. Each objective personality test mea- sures different aspects of an individual’s personality based on the specific constructs defined by the test developer. For example, the Minnesota Multiphasic Personality Inventory (MMPI) measures psychopathology and is used to assist in the diagnosis of emotional disorders. The Myers-Briggs Type Indicator (MBTI)®

measures personality based on a construct created by Carl Jung related to how people perceive and make judgments about the world, and the Substance Abuse Subtle Screening Inventory (SASSI)® is used to assess the probability of one having substance abuse issues. Although these three tests measure very different aspects of personality, they all can be useful in developing a picture of a client. Let’s take a look at some of the more common objective personality tests and explore how they are used.

Objective personality tests Paper-n-pencil tests to assess various as- pects of personality

248 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Common Objective Personality Tests Objective personality tests assess various aspects of personality and may increase client insight, identify pathology, and assist in treatment planning; however, how they do this can vary dramatically. In this section, we highlight nine objective personality tests, each of which has a slightly different emphasis (see Table 11.1). In addition, some other popular objective tests are briefly noted at the end of this section of the chapter.

Minnesota Multiphasic Personality Inventory-2 The most widely used diag- nostic personality test is the MMPI (Pearson, 2012a), which was originally devel- oped by Hathaway and McKinley in 1942. Today, there are currently three versions available: the full version, called the MMPI-2, which was introduced in 1989 (Butcher, Dahlstrom, Graham, Tellegen, & Kaemmer, 1989); an adolescent version, MMPI-A, released in 1992; and a shorter version called the Restructured Form (MMPI-2-RF), which became available in 2008. Both the MMPI-2 and the MMPI-A can be administered individually or in groups and require approximately 90 minutes for test-takers to complete the 567 items by hand. A computerized version is somewhat faster (Groth-Marnat, 2003; Pearson, 2012a). The MMPI-2-RF has only 338 items, so it can be taken by paper and pencil in about 45 minutes or completed on the com- puter in about 30 minutes (Pearson, 2012b). Although administration and scoring of the MMPIs is relatively straightforward, interpretation of the test is not. The MMPI-2 manual states, “Interpreting it demands a high level of psychometric, clin- ical, personological, and professional sophistication as well as a strong commitment to the ethical principles of test usage” (Butcher et al., 1989, p. 11). To be qualified to administer the test, examiners must have taken a minimum of a graduate-level

TABLE 11.1 Objective Personality Tests and Their General Use

Test Name General Use

Minnesota Multiphasic Personality (MMPI-2) Psychopathology, personal maladjustment, and a broad range of diagnoses

Millon Clinical Multiaxial Inventory (MCM-III) Psychopathology, pervasive personality disorders, and a broad range of diagnoses

Personality Assessment Inventory (PAI) Psychopathology, personality disorder features, and interpersonal traits

Beck Depression Inventory II (BDI-II) Presence and severity of depression

Beck Anxiety Inventory (BAI) Presence and severity of anxiety

Myers-Briggs Type Indicator (MBTI)® Personality types based on Jung’s theory of personality (nonclinical population)

16 Personality Factors (16PF)® General personality characteristics (nonclinical) based on Cattell’s work

NEO Personality Inventory-Revised (NEO PI-R)™ Assesses nonclinical personality dimensions using the Big Five personality traits

Conners 3 Assess attention-deficit/hyperactivity disorder and other comorbid behaviors

Substance Abuse Subtle Screening Inventory (SASSI) Detection of substance dependence

MMPI-2 Assists in identifying psychopathology; takes skill to interpret

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 249

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

course in psychological testing and a graduate-level course in psychopathology. There are three options for scoring the exam, which are local scoring on a com- puter, a mail-in scoring service, and hand scoring (Pearson, 2012a).

The MMPI-2 provides a large host of scoring scales. The core ones include 3 of the 9 validity scales, 10 basic (clinical) scales, and 15 content scales. There are numerous other scales available that include restructured clinical scales, clinical subscales, content component scales, negative treatment indicators, personality psychopathology scales, broad personality characteristics, generalized emotional distress, and so on. Figure 11.1 demonstrates a sample of the validity and content

120

110

100

90

80

70

60

50

40

30

Raw Score: 6 8 9 9 0 19 5 14 22 10 41 24 29 32 16 32 23 8 47

Response %: 100 100

Cannot Say (Raw): 1 F-K (Raw): Welsh Code: 27”406’811523/:9# F-L/K:

25

100 96 96 100 100 100 100 100 100 100 100 100

41 59

69.8

Percent True: Percent False:

Profile Elevation:

100 100 100 100 100

T Score (Plotted): 54 57F 64 79 41 69 56 47 47 62 95 57 79 62 72 92 68 33 75 Non-Gendered T Score:

54 57F 66 79 42 65 57 47 46 60 94 55 81 71 87 68 33 73

K Correction: 7 6 14 14 3

VR IN

TR IN F

FB S L K S Hs D Hy Pd M

F Pa P t

Sc M a SlF b F P

120 MMPI-2 VALIDITY AND CLINICAL SCALES PROFILE

110

100

90

80

70

60

50

40

30

The highest and lowest T scores possible on each scale are indicated by a "--".

FIGURE 11.1 | MMPI-2 Validity and Clinical Scales Extended Profile On the left side of this profile sheet are the validity scales and the right side the clinical scales. In Table 11.2, we discuss the L, F, and K scales, but as you can see there are additional scales of VRIN, TRIN, F(B), F(P), and S, which are also used to help assess if the individual is faking good or bad, over or under reporting symptoms, or answering items inconsistently. From the clinical scales on the right side, we can see this individual appears to have extremely high rates of depression (D; low Ma) and anxiety (Pt), as well as high levels of antisocial traits (Pd), social anxiety (Si), and paranoia (Pa).

Source: Excerpted from the MMPI®-2 (Minnesota Multiphasic Personality Inventory®-2) Manual for Administration, Scoring, and Interpretation, Revised Edition. Copyright © 2001 by the Regents of the University of Minnesota. Used by permission of the University of Minnesota Press. All rights reserved. “MMPI” and “Minnesota Multiphasic Person- ality Inventory” are trademarks owned by the Regents of the University of Minnesota.

A total of 3 of 9 validity scales are particularly important for interpre- tation, and 10 basic (clinical) scales are helpful in diagnosis

250 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

scales from an extended report, and Table 11.2 defines the most common validity scales and the 10 basic scales.

In interpreting the MMPI, it is important to understand the meaning of each scale. For example, a high L (Lie) score does not necessarily indicate compulsive lying but does indicate that the client is having trouble acknowledging his or her faults and that the entire test results are suspect and may have been “spoiled” (Butcher et al., 1989, p. 23). The basic or clinical scales are particularly useful in diag- nosis and treatment planning, and patterns of responses by clients are often used in making decisions about clients as opposed to examining an individual score on any

TABLE 11.2 Most Commonly Used Scales of the MMPI-2

Abbreviation/ Scale Number Name Brief Description

Validity Scales

L Lie Lacks ability to admit minor faults or character flaws; does not neces- sarily indicate “lying,” but that test scores may have been “spoiled.”

F Infrequency Reflects random scoring, which may indicate unwillingness to coop- erate, poor reading skills, or “faking bad” to gain special attention.

K Correction Tendency to “slant” or “spin” answers to minimize appearance of poor emotional control or personal ineffectiveness.

Basic (Clinical) Scales

Hs-1 Hypochondriasis Excessive concern regarding health with little or no organic basis, and rejecting reassurance of no physical problem.

D-2 Depression Depression and/or a depressive episode with feelings of discourage- ment, pessimism, and hopelessness.

Hy-3 Conversion hysteria

Conversion disorders where a sensory or motor problem has no organic basis; denial and lack of social anxiety often accompany symptoms.

Pd-4 Psychopathic deviant

Frequent hostility toward authority, law, or social convention, with no basis in cultural deprivation, subnormal intelligence, or other disorders.

Mf-5 Masculinity- femininity

Gender-role confusion and attempting to control homoerotic feelings; also emotions, interests, hobbies differing from one’s gender group.

Pa-6 Paranoia Paranoia marked by interpersonal sensitivities and tendency to mis- interpret intentions and motives of others.

Pt-7 Psychasthenia Obsessive-compulsive concerns (excessive worries and compulsive rituals) as well as generalized anxiety and distress.

Sc-8 Schizophrenia Wide range of strange beliefs, unusual experiences, and special sen- sitivities; often accompanied by social or emotional alienation.

Ma-9 Hypomania Displaying symptoms found in a manic episode such as hyperactivity, flight of ideas, euphoria, or emotional excitability.

Si-0 Social introversion

High scores indicate increasing levels of social shyness and desire for solitude; low scores indicate the opposite (social participation).

© Ce ng ag e Le ar ni ng

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 251

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

one scale. Because hundreds of patterns can arise from 10 scales, computerized scoring is helpful in quickly finding a diagnosis, interpreting client issues, and treatment planning. The MMPI-2 defines “clinical significance” as a T-score of 65 or greater, whereas the older MMPI established clinical significance as a T-score of 70 or higher.

The content scales identify 15 specific traits such as anxiety, fears, anger, cyni- cism, and low self-esteem, and are useful in creating a more detailed picture of the client as well as identifying other considerations for counseling. The extended report offers clinical subscales and content component scales, which can be useful in further delineating specific traits that may be of interest.

As compared to the original MMPI, the MMPI-2 has 82 rewritten items. How- ever, these items are so “psychometrically equivalent” to the original test that they have a 0.98 correlation; hence, the clinical scales are virtually unchanged (Graham, 2000, p. 189). To renorm the MMPI-2 against the nonclinical population, it was restandardized against a sample that included 2,600 people living in seven states and was fairly representative of the 1980 U.S. Census; however, both Hispanic and Asian-Americans were slightly underrepresented in the sample (Butcher et al., 1989). Test-retest reliabilities for the basic scales range between 0.67 and 0.92 for normal males and 0.58 and 0.91 for a similar sample of women. Internal consistency esti- mates using the Cronbach alpha for the basic scales ranged between 0.34 and 0.87. The only evidence of validity provided in the MMPI-2 manual is discriminant valid- ity. As you may recall, discriminant validity is the ability of the instrument to differ- entiate between different constructs. For example, MMPI-2 scores in depression should be minimally related to scores in hypochondriasis or conversion hysteria. However, the intercorrelations between scales were found to be quite high, primarily because many of the scales share test items (Groth-Marnat, 2003).

The 338 items in the MMPI-2-RF are a subset taken from the lengthier MMPI-2. Consequently, the developers were able to use the same normed data from the MMPI-2 for the newer Restructured Form. A major difference between the MMPI-2 and the MMPI-2-RF, besides the latter’s shorter length, is that the MMPI-2-RF has new validity scales, additional high-order scales, and restructured clinical scales, as well as somatic/cognitive and internalizing scales, externalizing, interpersonal and interest scales, and psychopathology scales (Pearson, 2012b). Despite the fact that the MMPI-2 does a decent job at assessing psychopathology, the test best known for measuring severe clinical syndromes including personality disorders is the Millon.

Millon Clinical Multiaxial Inventory, Third Edition The Millon Clinical Multi- axial Inventory (MCMI) has become the second most used objective personality test for measuring psychopathology after the MMPI-2 (Camara, Nathan, & Puente, 2000; Neukrug, Peterson, Bonner, & Lomas, 2013; Peterson, Lomas, Neukrug, & Bonner, 2014). The latest version of this test was designed to assess DSM-IV-TR per- sonality disorders and clinical symptomatology (Pearson, 2012c). The MCMI-III, which is generally called the “Millon” (Mill-on), focuses on DSM-IV-TR personality disorders (formerly called Axis II) in contrast to the MMPI’s focus on clinical disor- ders (formerly Axis I) (Groth-Marnat, 2003). The Millon has 175 true/false items, is written at an eighth-grade reading level, and is designed for individuals 18 years or older. Hence, there is also an adolescent version for ages 13 to 19 called the Millon Adolescent Clinical Inventory or MACI (Pearson, 2012d).

MCMI Used to assess per- sonality disorders (formerly Axis II) and clinical symptomatology

252 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The MCMI-III is much quicker to administer than the MMPI, requiring only 25 minutes (Pearson, 2012c). It can be taken via paper and pencil or on a com- puter. Scoring the Millon can be accomplished by using computer software, hand scoring, optical scan scoring, or a mail-in scoring service. The Millon has six differ- ent major scales: clinical personality pattern scales, severe personality pathology scales, clinical syndrome scales, severe clinical syndrome scales, modifying indices, and a validity index. The subscales to these major scales can be seen in Table 11.3.

The Millon uses a unique scoring method called the base rate (BR) that con- verts a raw score to a more meaningful standardized score based primarily on the psychiatric population. To do this, the publishers changed the actual median for a nonpsychiatric, or normal person, by setting it at 35, while the actual median for the psychiatric population is also changed and set at 60. A BR of 75 indicates that some of the features of that characteristic are present, while a BR of 85 indicates that the trait is clearly present (Groth-Marnat, 2003).

The norms for the interpretive report came from a clinical sampling of 752 men and women with a wide variety of diagnoses (Pearson, 2012c). There is also an additional norm group that can be used for a special corrections report that is based on 1,676 incarcerated inmates. The corrections report provides information that may be of value when doing a forensic evaluation, such as whether there is a need for mental health or anger management services, if the individual has suicidal tendencies, whether the person may be an escape risk, and so on.

Reliability coefficient alphas ranged from 0.67 to 0.90 for the scales and were predominantly in the 0.80s (Millon as cited in Groth-Marnat, 2003). Relative to convergent validity, the MCMI-2 scales have been correlated with several other scales and measures such as the MMPI and the Beck Depression Inventory (BDI). In general, most of the correlations were healthy and expected. One surprise was a low correlation (0.29) between the MCMI-III paranoid scale and the MMPI-2 paranoia scale. Other studies using the MCMI-2 have demonstrated moderate to high predictive validity for the instrument with DSM-IV-TR diagnoses.

Personality Assessment Inventory® The Personality Assessment Inventory (PAI®) is designed to aid in making clinical diagnoses, screen for psychopathology, and assist in treatment planning (Boyle & Kavan, 1995; PAR, 2012a). This instru- ment was designed in 1991 by Leslie Morey and continues to gain popularity with both clinicians and researchers. Some have argued that the PAI may be more effec- tive than the MMPI-2 (McDevitt-Murphy, Weathers, Flood, Eakin, & Benson, 2007; Weiss & Weiss, 2010; White, 1996).

The PAI is for use with adults 18 years and older, contains 344 items which are written at a fourth-grade reading level, and takes 50 to 60 minutes to complete (PAR, 2012a). Respondents rate items using a four-point ordinal scale with choices of false (not at all true), slightly true, mainly true, and very true. Scoring can be completed by hand in 15 to 20 minutes. It can also be scored on a computer or mailed to PAR, Inc., both of which will generate an interpretive report. Profile results provide 4 validity scales, 11 clinical scales, 5 treatment scales, and 2 inter- personal scales. Raw scores are converted to T-scores and percentiles (see Table 11.4), and two standard deviations or greater (T ≥ 70) are generally used as an indicator of scales requiring clinical attention.

PAI®

Aid in making diagno- sis, treatment plan- ning, and screening for psychopathology

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 253

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TABLE 11.3 The Millon Clinical Multiaxial Inventory-III Scales

Major Scales Subscale Name

Clinical Personality Pattern Scales

Scale 1 Schizoid

Scale 2A Avoidant

Scale 2B Depressive

Scale 3 Dependent

Scale 4 Histrionic

Scale 5 Narcissistic

Scale 6A Antisocial

Scale 6B Aggressive/sadistic

Scale 7 Compulsive

Scale 8A Passive-aggressive

Scale 8B Self-defeating

Severe Personality Pathology Scales

Scale S Schizotypal

Scale C Borderline

Scale P Paranoid

Clinical Syndrome Scales

Scale A Anxiety

Scale H Somatoform

Scale N Bipolar mania

Scale D Dysthymia

Scale B Alcohol dependence

Scale T Drug dependence

Scale R Posttraumatic stress disorder

Severe Clinical Syndrome Scales

Scale SS Thought disorder

Scale CC Major depression

Scale PP Delusional disorder

Modifying Indices

Scale X Disclosure

Scale Y Desirability

Scale Z Debasement

Validity Index

V Validity

© Ce ng ag e Le ar ni ng

254 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The PAI was normed using three samples: 1,000 adults stratified for gender, race, and age according to U.S. Census statistics; 1,296 patients from 69 clinical locations; and 1,051 college students (PAR, 2012a). Coefficient alphas for the median of the 22 scales were 0.81, 0.86, and 0.83 for these three samples, respectively (Boyle & Kavan, 1995). Test-retest reliabilities ranged between 0.68 and 0.85; however, inconsistency and infrequency validity scores were low, at 0.31 and 0.48, respectively. Numerous researchers have analyzed the PAI, in part, to assess its validity. As an example of concurrent validity, researchers examined the ability of the PAI Borderline features scale to assess 58 participants from an outpatient setting (Stein, Pinsker-Aspen, & Hilsenroth, 2007). The PAI was able to correctly classify 73% of the individuals as either having borderline personality disorder or not. In a study of 90 individuals who had been exposed to a trauma, the participants were classified into one of four categories: posttraumatic stress disorder (PTSD), depressive, social phobia, and well- adjusted (McDevitt-Murphy et al, 2007). Both the MMPI-II and PAI were adminis- tered to the groups. Results suggested that the PAI anxiety-related disorders scale was more accurate than the MMPI’s PTSD scale in identifying and differentiating the four groups.

TABLE 11.4 Personality Assessment Inventory Scales

Validity Scales Inconsistency

Infrequency

Negative impression

Positive Impression

Clinical Scales Somatic complaints

Anxiety

Anxiety-related disorders

Depression

Mania

Paranoia

Schizophrenia

Borderline features

Antisocial features

Alcohol problems

Drug problems

Treatment Scales Aggression

Suicidal ideation

Stress

Nonsupport

Treatment rejection

Interpersonal Scales Dominance

Warmth

© Ce ng ag e Le ar ni ng

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 255

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Beck Depression Inventory-II Originally introduced in 1961 by Aaron Beck and his colleagues, the BDI was designed to measure severity of depression (Beck, Steer, & Brown, 2003). The newest version, the BDI-II, was released in 1996, and today the inventory is ranked the number 1 assessment tool used by counselors and counselor educators (Neukrug et al., 2013; Peterson et al., 2014).

The BDI-II, which takes only 10 minutes to complete, asks clients to rate 21 questions on a scale from 0 to 3 based on depressive symptoms during the pre- ceding two weeks. Scores are obtained by adding up the total points from the series of answers and are interpreted based on the scales listed in Table 11.5.

When giving the BDI-II, special attention should be directed to questions 2 (hopelessness), and 9 (suicidal ideation), as scores of 2 or higher may indicate a higher risk of suicide (Beck et al., 2003). The BDI-II is quite useful in identifying and assessing the severity of symptoms of depression; however, as with all tests, it should not be used as a sole criterion for making a diagnosis. Due to its ease of administration, the instrument is also useful as a means of measuring client prog- ress by having the client take the instrument on an ongoing basis (Beck, 1995).

The norm group for the BDI-II included 500 outpatients who had been diag- nosed with depression using the DSM-III-R or the DSM-IV. Another smaller sam- ple of 120 Canadian college students was used as a normal comparative group (Beck et al., 2003). Using coefficient alpha, reliability was found to be 0.92 for out- patients and 0.93 for college students. As compared to the BDI, the content and criterion validity were increased in the BDI-II by having it more closely conform to the DSM-IV diagnosis criteria. In a study of convergent validity, a correlation of 0.93 was found for depressed outpatient clients who took both the BDI and BDI-II, and obtained means for the tests were 18.92 and 21.88, respectively (Beck et al., 2003). As was found in this study, scores on the BDI-II are generally about three points higher than those on the BDI. Finally, having clients with other disor- ders take the instrument, and finding scores that were not as high as those of clients with depression, showed discriminant validity.

Beck Anxiety Inventory The Beck Anxiety Inventory (BAI) was also developed by Aaron Beck and his colleagues. It is designed to be a simple and brief tool to assess anxiety for individuals from ages 17 to 80 (Dowd, 1998). Although it was originally published in 1993, it is still one of the top 10 assessment tools used by counselors and counselor educators (Neukrug et al., 2013; Peterson et al., in

TABLE 11.5 Interpreting Beck Depression Inventory Scores

Score Level of Depression

0 to 13 No or minimal depression

14 to 19 Mild depression

20 to 28 Moderate depression

29 to 63 (max) Severe depression

Below 4 Possible faking good

BDI-II Quick and easy method to assess depression

BAI Assessment of anxiety

© Ce ng ag e Le ar ni ng

256 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

press). The BAI is a self-report instrument that takes 5 to 10 minutes to administer and can be scored in moments (Pearson, 2012e). There are 21 questions using a four-point Likert-type scale ranging from 0 (not at all) to 3 (I could barely stand it). Scoring is done by adding up the total points, which then can be compared to scores, as listed in Table 11.6, to get a relative comparison of anxiety levels. The manual notes women tend to score higher than men and young people higher than older people, although how to adjust for this is vague.

The instrument was developed by combining three other anxiety measures that Beck had previously created (Waller, 1998). He and his team eliminated duplicate items and then used factor analysis to end up with the final 21 questions. The orig- inal norm group was made from 810 outpatients. Reliability and validity studies were done with three samples: one mixed diagnostic group, one group diagnosed with anxiety, and one that was nonclinical. Internal coefficient reliabilities ranged from 0.85 to 0.94 and a one-week test-retest reliability was found to be 0.75 (Dowd, 1998). Although factor analysis supported the anxiety construct, anxiety instruments tend to have a high correlation with depression instruments creating some uncertainty regarding the conceptual distinction between the two constructs.

Myers-Briggs Type Indicator® About two million people each year take the MBTI® test, making it the most widely used personality assessment for normal functioning (Quenk, 2009). Although the MBTI® instrument has enjoyed tremen- dous success and been used in a variety of settings, there has also been a “gap” between scientists’ and practitioners’ regard for the instrument, which the latest versions have attempted to mend (Mastrangelo, 2001). The theory is so widely used that there is a journal called Journal of Psychological Type with 72 volumes primarily dedicated to the MBTI as well as over 2,000 dissertations on the topic (ProQuest, 2013). The MBTI® instrument is based on the original work of Carl Jung (1921/1964) and his book Psychological Types. Through observation, Jung noted that people have basic characteristics along a diametrically opposed contin- uum regarding several factors, including extroversion or introversion, sensing or intuiting, and thinking or feeling.

After reading Jung’s book, Katharine Briggs and her daughter Isabel Briggs Myers became fascinated with Jung’s typology of people. They ultimately added a fourth dimension to Jung’s factors, which they called judging versus perceiving (Fleenor, 2001; Quenk, 2009). Believing this typology could help people with

TABLE 11.6 Interpreting Beck Anxiety Inventory Scores

Score Level of Depression

0 to 7 No or minimal anxiety

8 to 15 Mild anxiety

16 to 25 Moderate anxiety

26 to 63 (max) Severe anxiety

Women Tend to score 4 points higher than men

MBTI®

Popular method to assess normal per- sonality; based on Carl Jung’s psychological types

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 257

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

career selection and better understanding of self and others, they created an instru- ment that eventually became known as the MBTI® assessment. In 1975 CPP became the exclusive publisher of the MBTI® assessment and has managed and published the assessment since.

Today the MBTI® instrument is used in a wide variety of settings: in therapists’ offices to assist clients to develop a deeper understanding of self; in marriage and family counseling sessions and workshops to help couples and families examine dif- ferences and similarities of personality; in business and industry to help employees understand why individuals respond the way that they do; and in career counseling to help individuals find careers that match their personality types.

The four dimensions of the MBTI® instrument can be seen in Figure 11.2, and assessment results indicate preferences respondents favor across these four dichoto- mies. For example, social workers, counselors, and psychologists are often INFP, INFJ, and ENFJ types, while police officers are often ISTJ or ESTJ types. Although some use the MBTI® assessment as a career selection tool, the manual states that it should not be used as a screening tool for hiring employees (Mastrangelo, 2001).

The latest versions of the instrument are MBTI Step I® (form M) and MBTI Step II® (form Q). MBTI Step I, also known as form M, replaced form G in 1998 and contains 93 items with improved accuracy using Item Response Theory (IRT) (CPP, 2009a; Quenk, 2009). Form M only takes 15 to 25 minutes to administer, is written at a seventh-grade reading level, and is geared toward individuals 14 and older. It can be hand scored, computer scored, or mailed to the publisher for

Extraversion (E) Energy directed outward to people and objects

Introversion (I) Energy directed inward to ideas and concepts

Sensing (S) Perception comes mainly from the five senses

Intuition (N) Perception comes mainly from observing patterns and hunches

Thinking (T) Processing decisions based on logic, fact, and rationality

Feeling (F) Processing decision based on personal and social values

Judging (J) Makes decision quickly based on T or J; likes organization, planning, schedules

Perceiving (P) Processes decision based on S or N; likes spontaneity, flexibility, and diversions

FIGURE 11.2 | The four MBTI dichotomies

© Ce ng ag e Le ar ni ng

20 15

258 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

scoring. A computer-generated profile sheet for form M can be seen in Figure 11.3. There is an online version for form M that can be purchased called the MBTI®

Complete (R). It takes longer to complete, ranging from 45 to 60 minutes, but comes with a detailed interpretation of the results (CPP, 2009b).

MBIT Step II®, also known as form Q, replaced form K in 2001. It has 144 items, takes 25 to 35 minutes to administer, and is for adults 18 and older (CPP, 2009c). Form Q provides the basic scores along the four main dichotomies; how- ever, it also examines five facets to each of the four constructs (Quenk, 2009). For example, form Q breaks down the dichotomy of judging vs. perceiving into five opposite or polar subcomponents to further delineate these differences. The five subcomponents for judging-perceiving are systematic (J) vs. casual (P), planful (J) vs. open-ended (P), early starting (J) vs. pressure prompted (P), scheduled (J) vs. spontaneous (P), and methodological (J) vs. emergent (P). Form Q also uses IRT and the same norm group that form M uses (CPP, 2009c).

FIGURE 11.3 | Sample MBTI Profile for Form M Source: CPP. (2009). MBTI Profile – Form M. Retrieved from https://www.cpp.com/en/mbtiproducts.aspx?pc=11

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 259

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The norm group for forms M and Q was a national randomized sample of 3,200 adults. Coefficient alphas for form M are generally 0.90 or higher for each dichotomous scale. However, another study in the manual using test-retest reliability across all four dichotomies had only 65% of the participants receiving the same score after only four weeks (Mastrangelo, 2001). The MBTI® manual reports several studies providing some evidence of validity. Evidence is sound for the four separate scales, but it is lacking for the synergistic combination of the four scales. Another study showing evidence of validity found that 90% of the people taking the test agreed with the results on a self-report best fit. Further evidence of validity was found by comparing the MBTI® assessment with specific scales on the California Psychological InventoryTM assessment, a well-known personality instrument that measures basic personality traits somewhat similar to the 16PF discussed in the next section.

Sixteen Personality Factors Questionnaire (16PF)® The 16 Personality Factors Questionnaire (16PF)®, fifth edition, was developed based on Raymond Cattell’s research suggesting 16 primary personality components (Pearson, 2012f; Russell & Karol, 1994). As with the MBTI®, the 16PF is not a measure of pathology but rather a method to describe human behavior. The instrument has 185 items written at a fifth-grade level, and it takes about 45 minutes to administer by hand or 30 minutes by computer. Administration can be done individually or in a group, and scoring can be done by hand or computer.

The 16PF profile sheet provides the results in three sections: 16 primary fac- tors, 5 global factors, and 3 validity scales. The 16 primary factors are personality traits along a bipolar continuum. The profile sheet results use sten scores that range from 1 to 10. In Table 11.7, you can see the 16 personality factors and a descriptive summary of the ends of the bipolar continuum.

An “average” sten score will be between 4 and 7, capturing individuals between plus and minus one standard deviation, or the middle 68% of the popula- tion. Hence, scores from 1 to 3 or 8 to 10 indicate further deviation from the norm group toward either the left or right end of the bipolar scale and would merit extra attention. As you may remember from chapter seven, a sten score of 1 or 10 indi- cates a distance greater than two standard deviations from the norm.

The 16PF has five global factors: extraversion, anxiety, tough-mindedness, independence, and self-control. These are a result of factor analysis. As we briefly described in Chapters 5 and 10, factor analysis is a statistical method of determin- ing how individual items or factors correlate (i.e., hold together). For example, if we had written 10 items to capture the construct of sensitivity, we would hope those 10 items would combine (correlate together) to create a factor that is separate from the other items in the test. The 16PF global factors are secondary factors and are a result of how the 16 primary factors hold together. For example, the first global factor, extraversion, is created from combining the primary factors of warmth, liveliness, and social boldness. The global factor of anxiety is created from the primary factors of vigilance, apprehension, and tension. Table 11.8 shows the meaning of the bipolar continuums of the global factors.

The 16PF has three validity scales: impression management, infrequency, and acquiescence. The impression management scale is a measure of social desirability.

16PF Based on Cattell’s 16 bipolar personality traits

260 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

An unusually high score may indicate an individual’s inability to admit faults or perhaps a desire to “fake good.” A low score may indicate a desire to fake bad, or perhaps someone suffering from low self-esteem. The infrequency scale indicates that a respondent answered questions in an unusual way. A high infrequency score may represent reading comprehension problems, random responding, or try- ing to make the right impression. The acquiescence scale measures answering ten- dencies regardless of content. High acquiescence scores may indicate random responding, misunderstanding item content, or perhaps difficulty evaluating one’s self (Russell & Karol, 1994).

TABLE 11.7 The 16PF® Primary Personality Factors

Factor Meaning of Left End of Scale Meaning of Right End of Scale

A: Warmth Reserved, impersonal, distant Warm, outgoing, attentive to others

B: Reasoning Concrete Abstract

C: Emotional stability Reactive, emotionally changeable Emotionally stable, adaptive, mature

E: Dominance Deferential, cooperative, avoids conflict Dominant, forceful, assertive

F: Liveliness Serious, restrained, careful Lively, animated, spontaneous

G: Rule-consciousness Expedient, nonconforming Rule-conscious, dutiful

H: Social boldness Shy, threat-sensitive, timid Socially bold, venturesome, thick-skinned

I: Sensitivity Utilitarian, objective, unsentimental Sensitive, aesthetic, sentimental

L: Vigilance Trusting, unsuspecting, accepting Vigilant, suspicious, skeptical, wary

M: Abstractedness Grounded, practical, solution-oriented Abstracted, imaginative, idea-oriented

N: Privateness Forthright, genuine, artless Private, discreet, nondisclosing

0: Apprehension Self-assured, unworried, complacent Apprehensive, self-doubting, worried

Ql: Openness to change Traditional, attached to familiar Open to change, experimenting

Q2: Self-reliance Group-oriented, affiliative Self-reliant, solitary, individualistic

Q3: Perfectionism Tolerates disorder, unexacting, flexible Perfectionistic, organized, self-disciplined

Q4: Tension Relaxed, placid, patient Tense, high energy, impatient, driven

TABLE 11.8 The 16PF® Global Factors

Factor Meaning of Left End of Scale Meaning of Right End of Scale

Extraversion Introverted, socially inhibited Extraverted, socially participating

Anxiety Low anxiety, unperturbed High anxiety, perturbable

Tough-mindedness Receptive, open-minded, intuitive Tough-minded, resolute, unempathic

Independence Accommodating, agreeable, selfless Independent, persuasive, willful

Self-control Unrestrained, follows urges Self-controlled, inhibits urges

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 261

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The norm group for the 16PF, fifth edition, was taken from a sample of 4,449 people (Russell & Karol, 1994). To stratify the sample to match the U.S. population according to gender, race, age, and education, the norm group was reduced to 2,500. The reliability coefficient alphas for this sample ranged from 0.64 to 0.85 for the 16 factors. Two-week test-retest reliabilities from a separate sample of undergraduate students ranged from 0.69 to 0.86. Evidence of validity is presented in the adminis- trator’s manual in several forms. Factor analysis demonstrates individual items grouped (i.e., factored together) in accordance to the 16 personality factors the test developers had hoped for. Convergent validity was demonstrated through comparing 16PF with numerous other personality instruments such as the MBTI®, the California Psychological Instrument (CPI), the NEO Personality Inventory-Revised (NEO PI-R), and the Personality Research Form (PRF). Criterion-related validity was suggested through comparing 16PF scores with self-esteem (the Coopersmith Self-esteem Inventory), adjustment (Bell’s Adjustment Inventory), social skills (the Social Skills Inventory), empathy (CPI), creative potential, and leadership potential.

The Big Five Personality Traits and the NEO PI-3™ and NEO-FFI-3™ The Big Five personality traits, sometimes referred to as the five-factor model (NEO Five- Factor Inventory [NEO-FFI]), were first mentioned in the literature by Thurstone (1934). Although Raymond Cattell’s research suggested that personality can best be explained through 16 personality traits, numerous independent researchers found that personality variables can be further reduced to five constructs (Raad, 1998). Consequently, and unlike many other instruments, the five-factor model is based on research rather than theory. The five traits are openness, conscientious- ness, extraversion, agreeableness, and neuroticism (“OCEAN”). The Big Five can be summarized as:

• Openness—a willingness or desire to have new experiences, emotions, ideas, and be curious and imaginative; unconventional

• Conscientiousness—a sense of duty or self-discipline in one’s actions; prefers order and planning; goal driven

• Extraversion—having warmth, outgoingness, and a positive attitude; enjoys activity and excitement; high energy

• Agreeableness—being cooperative, kind, trusting, and altruistic; believes people are good

• Neuroticism—prone toward emotional distress such as anxiety, depression, anger, and impulsivity; emotionally reactive

The NEO Personality Inventory-3 (NEO-PI-3)™ is the latest instrument that measures personality across the five-factor model. The previous version, NEO Personality Inventory-Revised (NEO-PI-R), is, however, still available. The name, NEO, is a holdover from an earlier version when it was called the Neuroticism- Extroversion-Openness Inventory. The NEO-PI-3 updated 38 of the 240-items from the NEO-PI-R to lower the reading level for adolescents or adults who might not read as well (McCrae, Costa, & Martin, 2005; McCrae, Martin, & Costa, 2005; PAR 2012b). In one study this lowered the reading grade level from 8.3 to 4.4 and another to 5.3. The authors recommend the NEO-PI-3 for ages 12 to 99 and the NEO-PI-R for 17 to 89. Administration times range from 30 to 45 minutes,

NEO PI-3 Uses the Big Five model to assess per- sonality differences

262 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

and items are rated using a five-point Likert-type scale ranging from strongly dis- agree to strongly agree. Examples of items from the conscientiousness section are “I get chores done right away” and a reversed item, “I leave my belongings around.” There are two versions of each of these tests: One is for self-report (S) and the other is for a third-party observer or rater (R). Each of the Big Five factors consists of six facets or second-order factors that compose the larger constructs (see Table 11.9).

TABLE 11.9 The NEO PI-R™ Five Factors and 30 Facets

Big Five Factor Facets

Openness 1. Fantasy

2. Aesthetics

3. Feelings

4. Actions

5. Ideas

6. Values

Conscientiousness 1. Competence

2. Order

3. Dutifulness

4. Achievement striving

5. Self-discipline

6. Deliberation

Extraversion 1. Warmth

2. Gregariousness

3. Assertiveness

4. Activity

5. Excitement seeking

6. Positive emotion

Agreeableness 1. Trust

2. Aesthetics

3. Feelings

4. Actions

5. Ideas

6. Values

Neuroticism 1. Anxiety

2. Hostility

3. Depression

4. Self-consciousness

5. Impulsiveness

6. Vulnerability to stress

© Ce ng ag e Le ar ni ng

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 263

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Results for the factors and facets are reported using T-scores, which correspond to ranges as follows: below 35 is very low, 35 to 45 is low, 45 to 55 is average, 55 to 65 is high, and 65 and above being very high. In addition, a narrative describing the individual’s personality style and 10 personality-style graphs that describe 10 broad styles are offered to assist in interpretation of results.

The NE- PI-R and NEO-PI-3 have been used in numerous countries around the world and is translated into several languages. Reliability alphas have shown to remain consistent or slightly improved from the NEO-PI-3 (0.87 to 0.92) over the pre- vious NEO-PI-R (0.87 to 0.91) in samples (McCrae, Costa, & Martin, 2005). Exten- sive research has been conducted checking the convergent and discriminant validity of the NEO-PR-R and -3 with other instruments. For example, in a meta-analyses of 24 studies comparing the NEO Big Five factors and Holland’s six vocational domains (discussed in Chapter 10), researchers found five substantial correlations out of the potential 30 (5 factors � 6 domains) (Larson, Rottinghaus, & Borgen, 2002). Concur- rent validity is exemplified in a meta-analyses of 15 samples comparing the NEO-PI-R five-factor model and DSM-IV-TR personality disorders (Saulsman & Page, 2004). Results indicate neuroticism was associated with high emotional distress personalities such as paranoid, schizotypal, borderline, avoidant, and dependent. Extroversion cor- related with gregarious personality qualities such as histrionic and narcissistic, and lack of agreeableness related with personalities experiencing interpersonal difficulties like paranoid, schizotypal, antisocial, borderline, and narcissistic.

A shortened version, called the NEO Five Factor Inventory-3 (NEO-FFI-3)™ is available for individuals 12 or older. Containing only 60 items, it can be adminis- tered individually or in a group in only 10 to 15 minutes (PAR, 2012c). Reliabil- ities are decreased due to the shorter length than the 240-item NEO-PI-3. For example, coefficient alphas for the NEO-FFI-3 five factors using three samples (adults, adolescent, and middle school children) ranged between 0.71 and 0.87 for the self-report version and 0.66 to 0.88 for the observer-rater form (McCrae & Costa, 2007). Evidence of validity is suggested through correlations between the shortened NEO-FFI-3 and the full NEO PI-3 with the self-report versions correlat- ing between 0.83 and 0.97 for the five domains and the observer-rated versions demonstrated relationships from 0.81 to 0.97.

Conners 3rd Edition The Conners, 3rd Edition™ or Conners 3™, is an instru- ment geared to help in the diagnosis of attention-deficit/hyperactivity disorder (ADHD) and other comorbid problematic behaviors such as oppositional defiant disorder and conduct disorder (Multi-Health Systems, 2013). The Conners 3 is the replacement to the familiar Conner’s Rating Scale, Revised (CRS-R), which was discontinued at the end of 2012. The Conners 3 is for use with children between the ages of 6 and 18 and has forms to be completed by the teacher, parent, and self-report (if 8 years or older). There is a long version that takes approximately 20 minutes and a shorter one that takes about 10 minutes (Arffa, 2010; Dunn, 2010). The instrument can be completed online or using paper-n-pencil. Respon- dents answer a serious of questions on a four-point Likert-type scale ranging from 0 (never or seldom) to 3 (very true or very frequently). The long forms range between 99 to 115 items and the short forms between 41 and 45 (Arffa, 2010). Scoring and interpretation should be done by a professional with a graduate-level

Conners 3™

Assess ADHD and other problematic behaviors

264 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

course in testing and assessment. Results can be useful in making eligibility deter- mination decisions regarding special education as well as monitoring treatment interventions.

Content validity was demonstrated through the developer’s analysis of theory, research, previous instruments, current legislation, as well as focus groups using practitioners and scholars (Dunn, 2010). Construct validity was established through the use of exploratory factor analysis on original pilot data and confirma- tory factor analysis with a clinical sample of 731 youth. The use of comparing test results with previously diagnosed populations can also be considered a form of concurrent validity where test data is compared to a known quantity. The develo- pers created their standardization sample using 50 male and 50 female students for each of the age groups while generally matching U.S. census demographics. Internal consistency reliabilities for parents and teachers were above 0.90 and self- report for students were 0.85 or above.

Substance Abuse Subtle Screening Inventory (SASSI®) The SASSI is a helpful tool for identifying people who have a high probability of a substance-related disor- der. There are two versions: the SASSI-3 (Miller, 1999) for adults 18 and older and the SASSI-A2 for adolescents between the ages of 12 and 17 (Miller & Lazowski, 2001). These instruments suggest if an individual has a substance dependency with a 93% and 94% overall accuracy, respectively (Lazowski, Miller, Boye, & Miller, 1998; SASSI, 2001, 2009). The adolescent version varies in that it further classifies respondents between substance dependency and substance-abuse disorders. Because both instruments are reasonably similar, we will only discuss the adult version.

The SASSI-3 can be administered in about 30 minutes and scored in 5. The first section, which is the “subtle” component, contains 67 true/false statements, and the second section consists of 26 alcohol- and other drugs-related questions rated on a four-point scale ranging from “never” to “repeatedly.” Examples of questions from the first section are “Sometimes I have a hard time sitting still” or “I always feel sure of myself.” The second section is overt and asks the frequency of statements such as how often have you “argued with your family or friends because of your drinking?” Scoring generates nine subscales, which are compared against nine rules to determine if there is a low or high probability of a substance- dependence disorder. The instrument has a check for random answering. The sub- scales can be useful in diagnosis, treatment planning, and interpreting validity. For example, if an individual scores low in the face valid alcohol and other drug scales but scores high in defensiveness and subtle attributes, this indicates that he or she is guarded and may be minimizing his or her substance use. It is important to note that an individual may deny substance abuse, but due to the subtle attributes scale as well as other subscales, the instrument can continue to suggest dependency with high accuracy. If a respondent does not have a dependency, but the SASSI suggests he or she does, this is a false positive, as discussed in Chapter 5. A false negative is when a score suggests a low probability of dependency but, in fact, the individual does have an addiction. The SASSI-3 and SASSI-A2 are designed to err toward making false positives rather than false negatives (SASSI, 2001).

Original norming for the SASSI-3 began with a sample of 1,958 participants from clinical settings, prisons, and newspaper advertisements (Lazowski et al.,

SASSI A subtle instrument to screen for substance dependence

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 265

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

1998). The sample was reduced to 839 respondents that had full data, such as a DSM-IV diagnosis, and was split in two. One-half of the sample was used to develop the classification rules, and the other half used to test the accuracy of those rules. Reliability coefficient alphas from the larger sample (n ¼ 1,821) were 0.93 for the overall instrument. Test-retest reliabilities with 40 respondents after a two-week period ranged between 0.92 and 1.00. Evidence of validity is suggested through comparing the accuracy of the SASSI scores with DSM-IV criteria. As you may recall from Chapter 5, this is criterion-related validity because the developers compared their instrument (SASSI-3) with another criterion (DSM-IV diagnosis) to suggest validity. The comparison between SASSI scores and substance-dependency diagnosis in this sample was accurate 94% of the time (Lazowski et al.).

The SASSI-3 profile sheet provides nine subscales, which can be useful in pro- viding additional information for diagnosis and treatment planning. The nine scales are face valid alcohol, face valid other drugs, symptoms, obvious attributes, subtle attributes, defensiveness, supplemental addiction scale, family vs. controls, and cor- rectional. For example, if an individual has a relatively high face valid alcohol score, symptoms score, attribute score, and a low defensiveness score, this suggests that he or she is open about using, acknowledges that he or she may have a prob- lem, and is likely ready for treatment. Extremely low SASSI defensiveness scores are often associated with individuals who are experiencing depression (SASSI, 2001). If a profile has a low face valid scores, symptoms scores, and attribute scores, but a high defensiveness score, this suggests that person is likely guarded, resistant, faking good, and precontemplative regarding his or her readiness for change. You would likely approach this individual differently regarding his or her substance use than the previous example.

Other Common Objective Personality Tests Numerous other objective per- sonality tests are available, but we have room to describe just a couple of other personality instruments here. The Taylor-Johnson Temperament Analysis (TJTA, 2013) assesses 18 dimensions of personality that affects social, family, marital, work, and other environments. It can be used to assist in counseling emotionally normal adolescents and adults as individuals, couples, families, or in vocational situations (Axford & Boyle, 2005). Although the TJTA can be used for relationship assessment, other instruments have been created specifically to appraise marriages or partnerships. For instance, the Marital Satisfaction Inventory, Revised (MSI-R; WPS, n.d.) is an instrument to assess the severity and nature of the conflict in a relationship or partnership. The inventory takes only 25 minutes to administer, is inexpensive, and has respectable validity and reliability, making it a nice tool for therapists doing relationship counseling (Bernt & Frank, 2001; WPS, n.d.).

PROJECTIVE TESTING In Chapter 1, projective personality testing was defined as a type of personality assessment in which a client is presented a stimulus to which to respond, and sub- sequently, personality factors are interpreted based on the client’s responses. We noted that such testing is often used to identify psychopathology and to assist in treatment planning.

Projective tests Responses to stimuli are used to interpret personality factors

266 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

When interpretations about client responses are made, they are often based on normative data. However, with projective testing clients respond in an open-ended manner to vague stimuli that results in a wide-range of responses that often limits norm-group comparisons. Thus, the validity and reliability of these instruments is not as solid as most objective tests. Therefore, a battery of projective tests should never be used alone. Although they are a powerful tool that can elicit information often not found from objective tests, projective instruments must be paired with objective measures as well as the clinical interview and other collateral data to make a well-rounded assessment. Dozens of projective tests exist, and we explore some of the more prevalent ones in the following sections.

Common Projective Tests The most popular projective tests include the Thematic Apperception Test (TAT); the Rorschach Inkblot Test; the Bender Visual-Motor Gestalt Test, Second Edition; the House-Tree-Person (HTP); the Kinetic House-Tree-Person Test (KHTP); the Sentence Completion Series; and the Rotter Incomplete Sentence Blank. The following discussion offers a brief overview of these instruments.

The Thematic Apperception Test and Related Instruments The Thematic Apperception Test (TAT) was developed by Henry Murray and his colleagues in 1938 and consists of a series of 31 cards with vague pictures on them, although only 8 to 12 cards are generally used during an assessment depending on the age and gender of the client as well as the client’s presenting issues. Showing the cards one at a time, the examiner asks the client to create and describe a story that has a beginning, middle, and end. The storytelling process allows great access to the cli- ent’s inner world and shows how that world is affected by the client’s needs and by environmental forces, known as press.

The ambiguous pictures on the cards are more structured than inkblot tests such as the Rorschach; consequently, the TAT tends to draw out from the client issues related more to current life situations than deep-seated personality structures (Groth-Marnat, 2003) (see Figure 11.4). The TAT is based on Murray’s need-press personality theory, which states that people are driven by their internal desires, such as attitudes, values, goals, and so on (needs), or by external stimuli (press) from the environment. Therefore, individuals are constantly struggling to balance these two opposing forces.

The TAT has been extensively researched; however, it still lacks a level of stan- dardization that most objective personality tests have achieved (Groth-Marnat, 2003). There is no universally agreed upon scoring and interpretation method, although most clinicians use a qualitative process of interpreting responses. Hence there is considerable controversy over reliability and validity of the instru- ment. When scoring systems have been used in a controlled setting, interscorer reliability has been seen to range between 0.83 and 0.92 and test-retest reliabil- ities between 0.64 and 0.83 (Ronan, Gibbs, Dreer, & Lombardo, 2008); however, if responses are interpreted by clinicians outside of more controlled environments, these figures would likely drop (Groth-Marnat). Studies to show evidence of validity regarding the TAT are controversial with some arguing in favor and others against

TAT Clients create a story based on cards with vague pictures; TAT is based on Murray’s need-press personali- ty theory

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 267

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

(see Lilienfeld, Wood, & Garb, 2000 vs. Woike & McAdams, 2001). However, others contend that due to the nature of projective tests, the evidence of validity for these instruments is not as important as for objective measures (Karon, 2000). For instance, some suggest that the rich narrative detail developed through the TAT gives the therapist a unique window into the client’s psyche. Indeed, the value of the TAT seems to be supported by its widespread use. It is the sixth most frequently used test by both clinical and counseling psychologists (Camara et al., 2000) and heavily used by counselors and counselor educators (Neukrug, et al., 2013; Peterson, et al., 2014) (see Table 1 and Table 2 in the Section III Introduction).

Because the cards are dated, and the fact that the human figures in the cards are almost exclusively White, many of the cards may raise historical and cross- cultural bias. The development of Southern Mississippi’s TAT (SM-TAT) and the Apperceptive Personality Test (APT) are two attempts to counter some of these problems. The APT has only eight cards with multicultural pictures and an objective scoring method. Although the SM-TAT and APT are probably superior instruments (more modern, more rigorous methodology, and greater validity), the long tradition of the TAT will probably prevent its replacement (Groth-Marnat, 2003).

FIGURE 11.4 | TAT Figures Source: Reprinted by permission of the publisher from Henry A. Murray. Thematic Apperception Test, Plate 12F. Cambridge, MA: Harvard University Press. Copyright 1943 by the President and Fellows of Harvard College.

268 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

In addition to the TAT, the Children’s Apperception Test (CAT), designed for children ages 3 to 10, has been developed. Because children have shorter attention spans than adults, the instrument has only 10 cards, and due to the fact that chil- dren tend to relate more easily to animals than humans, the pictures depict ani- mals. In addition, a later version of the CAT, the CAT-H, was made that depicts humans. Despite the development of these instruments, the TAT is still frequently used with children, probably due to its familiarity among clinicians.

Rorschach Inkblot Test Herman Rorschach developed his famous inkblot test in 1921 by splattering ink onto pieces of paper and folding them in half to form a symmetrical image (see Figure 11.5). After much experimentation, he chose 10 cards to create the Rorschach Inkblot Test that is still used today. When giving the Rorschach, clinicians show clients the cards, one at a time, and ask them to talk about what they see on the card. A follow-up inquiry with clients addresses issues of what they actually saw, how they saw it, and where on the card it was seen. Ultimately, the clinician wants to see exactly what the client saw on the card.

Rorschach, a student of Carl Jung, believed the ambiguous shapes of the inkblots allowed the test-taker to project his or her unconscious mind onto these images. By 1959 the Rorschach had become the most frequently used instrument

FIGURE 11.5 | Inkblot Similar to Those on the Rorschach Source: Lambert/Archive Photos/Getty Images

Rorschach Inkblot Test Clients are asked to identify what they see in 10 inkblots; the un- conscious mind “pro- jects” itself onto the image

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 269

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

in clinical practice (Sundberg, 1961), and it continues to be one of the most frequently used projective personality test (Camara et al., 2000; Hogan, 2005; Neukrug et al., 2013; Peterson et al., 2014). Although it has had tremendous pop- ularity, it has also been closely scrutinized and criticized. By 1955, more than 3,000 journal articles had been written about it (Exner, 1974), and a recent ERIC and Psyclnfo database search shows more than 8,500 articles in which the Rorschach is cited. The greatest difficulty with the Rorschach has been providing adequate validity. Another challenge of the Rorschach test is that it requires exten- sive training and practice to use. However, we believe this instrument still has merit and can be a useful tool in the assessment process (see Box 11.1).

One of the most popular scoring systems for the Rorschach was developed by Exner (1974). This system uses three components: location, determinants, and con- tent. Location is the portion of the blot to which the response occurred, and the examinee’s responses are broken down into categories such as the whole blot (w), common details (D), unusual details (Dd), and white space details (S). Determinants are used to describe the manner in which the examinee understood what he or she saw, and these are broken down into (1) form (“that looks just like a bat”), (2) color (e.g., “it’s blood, because it’s red”) (3) shading (“it looks like smoke because it’s grayish-white”). Finally, content is scored based on 22 categories such as whole human, human detail, animal, art, blood, clouds, fire, household items, sex, and so on. Specific content can hold meaning; for instance, a goat can be an indication of a person being obstinate, or a number of animal responses by an adult could be an indication of immature psychosexual development (children tend to include lots of animals in their responses). Once all of the data have been recorded, a fairly complex series of calculations are used to create numerical ratios, percentages, and derivations with which to interpret the results. Scoring systems such as Exner’s are very complex and are important ways of managing the large amount of interpretive material the client is presenting.

Bender Visual-Motor Gestalt Test, Second Edition Lauretta Bender originally published the Bender Visual-Motor Gestalt Test in 1938, and after several revisions, it is now called the Bender Gestalt II. A brief test that takes only 5 to 10 minutes to

BOX 11.1 Rorschach Use in Clinical Practice

Card IV of the Rorschach is known by some as the “father card” because it shows what some see as a hefty, overbearing figure with a large penis. When I was giving the Rorschach to a 17-year-old female high school student, I obtained what is sometimes called a “shock response.” I showed her this card and although she had made fairly normal responses to the other cards, when seeing this card she said to me, rather emphatically, “I see nothing there.” Upon

inquiry later, I again asked her to give me a response to that card, at which point she firmly put the card face down and said, “I told you I didn’t see anything there.” When all testing was complete, I looked at the young woman and asked, “Were you molested?” at which point she broke down and started to sob.... This was the beginning of counseling for this young woman who had never shared this secret before.

—Ed Neukrug

Exner’s scoring sys- tem examines location, determinants, and content

Bender Gestalt II Assists in identifying developmental, psy- chological, or neuro- logical deficits

© Ce ng ag e Le ar ni ng

270 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

administer, it measures an individual’s developmental level and psychological func- tioning and is also used to assess neurological deficits after a traumatic brain injury (Pearson, 2012g). The test asks children aged 4 to 7 and individuals aged 8 to 85þ to draw the nine figures shown in Figure 11.6. Children 4 to 7 have four additional figures to replicate, and individuals 8 to 85 have three additional figures to copy.

The latest version of the instrument, the Bender Gestalt II, used a norm group of 4,000 individuals who were representative of the 2000 U.S. Census (Riverside Publishing, 2010). This version uses a new, global, five-point scoring system.

BENDER GESTALT FIGURES HANNAH'S DRAWINGS (age 31⁄2)

REBECCA's DRAWINGS (age 51⁄2)

Fig. A.

Fig. 1.

Fig. 2.

Fig. 3. Fig. 4. Fig. 5.

Fig. 6.

Fig. 7. Fig. 8.

Fig. A.

Fig. 4.

Fig. 5.

Fig. 6.

FIGURE 11.6 | The Bender Gestalt Figures Above are reproductions of four of nine figures of the Bender Visual Motor Gestalt from two children who are devel- opmentally on target. When Hannah couldn’t reproduce figure 6, she became very frustrated. Look at how much eas- ier it is for the older child, Rebecca, to reproduce the figures.

© Ce ng ag e Le ar ni ng

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 271

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

A score of 0 represents no resemblance or scribbling, and a score of 4 represents a nearly perfect drawing (Brannigan, Decker, & Madsen, 2004). This version pro- vides standard scores, T-scores, and percentile ranks. As you might suspect, the test can examine psychomotor development of children by comparing a child’s response to mean responses of children belonging to the child’s age group; person- ality factors such as compulsivity in completing the drawing accurately (graduate students!); and neurological impairment, such as might be evidenced when an indi- vidual cannot accurately place the diamond in drawing 8 (see Figure 11.6).

Accurately interpreting a number of these factors takes advanced training and should not be attempted without such preparation. The original version of the Bender Gestalt showed some evidence that it was measuring the factors it pur- ported to measure, and reliability data showed test-retest reliability at 0.84 while interrater reliability was shown to be in the low to mid 0.90s (Naglieri, Obrzut, & Boliek, 1992). Evidence of convergent validity is demonstrated in studies examining scores of the Bender-Gestalt II with visual-spatial subtests on the WISC-III (Decker, Allen, & Choca, 2006) and quantitative reasoning, fluid reasoning, and visual- spatial factors on the Stanford-Binet intelligence test (SB-5; Decker, Englund, Carboni, & Brooks, 2011). Predictive validity has been demonstrated by showing that kindergarten children’s Bender-Gestalt scores is predictive of academic achieve- ment, social adjustment, and emotional adjustment one-year later (Bart, Hajami, & Bar-Haim, 2007).

House-Tree-Person and Other Drawing Tests It’s not quite clear when the use of drawing tests began, but over the last 50 years they have become some of the most popular, simple, and effective projective devices. By asking clients to draw simple pictures, one can gain tremendous insight into a client’s life and perhaps unconscious undertow.

Buck (1948) introduced the original House-Tree-Person (HTP) drawing test when he simply asked clients to draw a house, a tree, and a person on three sepa- rate pieces of paper. The Kinetic House-Tree-Person Drawing Test (KHTP) is slightly different in that the client is asked to draw all the figures “with some kind of action” on one sheet of paper (8½ � 11 presented horizontally) (Burns, 1987, p. 5). Burns believes the tree is the universal metaphor for human development as seen in religion, myth, poetry, art, and sacred literature:

In drawing a tree, the drawer reflects his or her individual transformation process. In creating a person, the drawer reflects the self or ego functions interacting with the tree to create a larger metaphor. The house reflects the physical aspects of the drama. (p. 3)

Numerous books and materials describe how to specifically interpret the H-T-P and K-H-T-P drawings. Table 11.10 provides examples of a few interpretive sug- gestions (Burns, 1987).

Other drawing tests exist, such as the Draw-A-Man, Draw-A-Woman, or the Kinetic Family Drawing (KFD), which asks the individual to draw his or her family doing something together. These tests all try to tap into unconscious aspects of the individual’s self by focusing on slightly different content. Drawing tests do not require artistic prowess on the part of the client, are quickly administered, and can often produce important interpretive material for the clinician.

Drawing tests Quick, simple, and effective projective tests

272 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Sentence Completion Tests “This book is .” Completing a sen- tence has been used as a projective device since Galton and Jung (see Chapter 1), and although you are not likely to see a sentence stem asking about a specific book, you are likely to see sentence stems that ask you to describe your relation- ship with your mother, father, spouse, lover, friends, and so on. Two of the more common sentence completion tests are the Sentence Completion Series and the Rotter Incomplete Sentence Blank. In addition, clinicians will sometimes create their own sentence completion tests that can be used when giving a personality battery.

The Sentence Completion Series is a “semi-projective” series of tests for gather- ing personality and psychodiagnostic information from adolescents and adults. The series contains eight forms with 50 sentence stems per form (PAR, 2012d). Individ- ual forms address specific issues such as family, marriage, work, aging, and so on. An interpretive manual makes very broad suggestions about how to interpret the responses, such as assessing the examinees general tone and level of defensiveness; however, there is no objective scoring methodology. In addition, there is no reli- ability, validity, or norm data. It is certainly a useful instrument, but its greatest weakness is the lack of any psychometric data (Moreland & Werner, 2001).

TABLE 11.10 Sample of Suggested K-H-T-P Interpretations (Burns, 1987)

Characteristic Interpretation

General

Unusually large drawings Aggressive tendencies; grandiose tendencies; possible hyper or manic conditions

Unusually small drawings Feelings of inferiority, ineffectiveness or inadequacy; withdrawal tendencies; feelings of insecurity

Very short, circular, sketch strokes Anxiety, uncertainty, depression, and timidity

House

Large chimney Concerns about power, psychological warmth, or sexual masculinity

Very small door Reluctant accessibility; shyness

Absence of windows Suggests withdrawal and possible paranoid tendencies

Tree

Broken or cut-off branches Feelings of trauma

Upward-pointed branches Reaching for opportunities in the environment

Slender trunk Precarious adjustment or hold on life

Person

Unusually large head Overvaluation of intelligence; dissatisfaction with body; preoccupation with headaches

Hair emphasis on head, chest, or beard

Virility striving; sexual preoccupation

Wide stance Aggressive defiance and/or insecurity

Sentence completion tests Can reveal uncon- scious issues, but some question the validity and reliability

© Ce ng ag e Le ar ni ng

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 273

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

The Rotter Incomplete Sentence Blank®, Second Edition (RISB®-2; Pearson, 2012h), is designed for assessing the overall adjustment of high school students, college students, and adults. The instrument has 40 items, which start with one- or two-word sentences that need completing. It takes approximately 20 minutes to complete and has a semi-objective scoring method using a seven-point ordinal scale for rating (Boyle, 1995). The norm group in the manual was college students that brings into question its validity for use with high school students and adults. Researchers have shown internal consistency alphas of 0.78, split-half alphas of 0.76, and inter-rater agreement of 0.92 (Logan & Waehler, 2001; Weis, Toolis, & Cerankosky, 2008). Results also suggested convergent validity between the RISB-2 scores and self-report, parent report, and teacher report. Criterion validity was sug- gested when shown that students with scores of 140 or higher could be screened for maladaptive behaviors (Weis et al.).

Although questions about the validity and reliability of sentence completion tests and other projective measures remain, it is clear that these instruments can provide a quick method of obtaining a client’s feelings and unconscious thoughts about important issues in the client’s life (Youngstrom, 2013).

THE ROLE OF HELPERS IN CLINICAL ASSESSMENT Because clinical assessment can add one more piece of knowledge about a client, it should always be considered as an additional tool to use. Therefore, all helpers can, and perhaps should, be involved in some aspects of clinical assessment. For instance, an elementary school counselor might consider using a self-esteem inven- tory when working with young children to help identify students who are strug- gling with self-esteem issues, while a high school counselor might want to use any of a wide range of objective personality measures to help identify concerns and aid in setting goals for students. College counselors, agency clinicians, social workers, and private practice professionals all might use a wide range of clinical assessment tools as part of their repertoire to help identify issues and devise strategies for problem solving. Clinicians should reflect on what kinds of clinical assessment tools would benefit their clients at their particular setting and whether they have sufficient training to give such assessments (e.g., projective tests generally require advanced training).

FINAL THOUGHTS ON CLINICAL ASSESSMENT The clinical assessment process assists helpers in making decisions that often will affect clients’ lives in critical ways. Such decisions can result in persons being labeled, institutionalized, incarcerated, stigmatized, placed on medication, losing or gaining a job, being granted or denied access to their children, and more. Test givers should remember the impact that their decisions will have on clients and monitor the quality of the tests they use, their level of competence to administer tests, and their ability to make accurate interpretations of client material. Acting in any other fashion is practicing incompetence.

274 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SUMMARY We began this chapter by noting that clinical assessment examines the client through multiple methods, including the clinical interview, the administration of informal assessment techni- ques, and the administration of objective and projective tests. This process is used (1) as an adjunct to counseling, (2) for case conceptua- lization and diagnostic formulations, (3) to deter- mine psychotropic medications, (4) for treatment planning, (5) for court decisions, (6) for job placement, (7) to aid in the diagnosis of health- related problem (e.g., Alzheimer’s), and (8) to screen for people at risk (e.g., those who might commit suicide).

In the next part of the chapter, we examined a number of widely used objective personality tests. First, we looked at the MMPI-2, which provides a number of scales, of which the most commonly used are the three validity scales (lie, infrequency, and correction) and the 10 basic (clinical) scales: hypochondriasis, depression, conversion hysteria, psychopathic deviant, masculinity-femininity, paranoia, psychasthenia, schizophrenia, hypoma- nia, and social introversion.

In contrast to the MMPI-2, which assesses mostly clinical disorders (formerly Axis I), we noted that the Millon assesses personality disor- ders (formerly Axis II disorders) and clinical symptomatology. This test offers six major scales, including Clinical Personality Pattern Scales, Severe Personality Pathology Scales, Clinical Syndrome Scales, Severe Clinical Syndrome Scales, Modifying Indices, and a Validity Index. The Millon sets a Base Rate (BR) to determine responses that indicate a trait’s presence.

The Personality Assessment Inventory (PAI) is an instrument that offers information regarding both clinical disorders and personality disorders and is growing in popularity. The PAI contains 4 validity scales, 5 treatment scales, 2 inter- personal scales, and 11 clinical scales, which include somatic complaints, anxiety, anxiety- related disorders, depression, mania, paranoia,

schizophrenia, borderline features, antisocial fea- tures, alcohol problems, and drug problems.

The Beck Depression Inventory-II (BDI-II), another very popular objective test, asks 21 ques- tions related to depression, and can be quickly answered and scored. The test is valuable for assessing depression or suicidal ideation, and it can be used on an ongoing basis for the evalua- tion of positive changes in counseling. Similarly, the Beck Anxiety Inventory (BAI) has 21 ques- tions relating to anxiety symptomology. Exami- nees rate each one using a four-point Likert-type scale that can be speedily added up to assess the over-all level of anxiety severity.

The Myers-Briggs Type Indicator (MBTI®), which measures normal personality functioning, is based on Jung’s six psychological types as well as two additional types identified by the test’s creators, Katharine Briggs and Isabel Briggs Myers. The four opposing personality dichotomies are as follows: extroverted (E) vs. introverted (I); sensing (S) vs. intuiting (N); thinking (T) vs. feeling (F); judging (J) vs. perceiving (P). Some evidence shows that com- binations of the four dichotomies may indicate certain global personality traits.

The 16PF was developed based on the work of Raymond Cattell and is used to describe the 16 traits of personality that he identified. We learned that similar to the MBTI®, the 16 primary scales are bipolar in nature, and scores reflect an indivi- dual’s personality. The instrument additionally contains five global factors based on a factor analysis of the 16 primary scales and three valid- ity scales.

The NEO Personality Inventory-3 (NEO PI-3)™ and NEO Five-Factor Inventory-3 (NEO-FFI-3)™ use the research based on the Big Five personality model. The Big Five personality traits are openness, conscientiousness, extraver- sion, agreeableness, and neuroticism. Each of the five factors contains an additional six facets or subscales. We discovered that the NEO PI-3 is the full-length version of the instrument that takes

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 275

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

about 40 minutes to administer. The NEO-FFI-3 is the shortened version, taking 10 to 15 minutes to complete.

The Conners, 3rd Edition, which recently replaced the Conners Rating Scale, Revised, is used to assess children ages 6 to 18 for ADHD and other problematic behaviors. There are forms to be completed by the teacher, parents, and stu- dent. The long version takes about 20 minutes while the short version about 10.

The Substance Abuse Subtle Screening Inven- tory (SASSI) has separate versions for adolescents and adults to screen for potential substance dependence. We learned that a portion of the instrument is “subtle” and attempts to capture dependency regardless of whether the examinee is attempting to fake good (i.e., minimize or deny substance use). The subscales can be inter- preted to evaluate a potential client’s readiness for change.

In the last part of the chapter, we examined projective tests that are often used to identify psy- chopathology and to assist in treatment planning. The first projective test we examined was the The- matic Apperception Test (TAT). Developed by Henry Murray in 1938, the TAT consists of 31 cards with vague pictures on them. Showing select cards, the examiner asks the client to tell a story that has a beginning, middle, and end. Client responses are interpreted based on Murray’s need-press personality theory, which states that people are driven by their internal desires (needs) or by the external stimuli (press) from their environment. Because some view the TAT pictures as being dated and culturally biased, updated versions of the test have been developed. However, many clinicians continue to prefer the original TAT.

The Rorschach was the next projective test we examined. Developed in 1921 by Herman Rorschach, the examiner asks the client to reveal everything he or she sees on each of 10 inkblots. Interpretation of responses assumes that the cli- ent is projecting onto the inkblot his or her

unconscious thoughts. One of the most well- known scoring systems for the Rorschach, which was developed by Exner, looks at loca- tion, or portion of the blot where the response occurred; determinants, or the manner in which the examinee understood what he or she saw; and content, or what the examinee actually saw.

The Bender Visual-Motor Gestalt Test was originally published in 1938 by Lauretta Bender. Now called the Bender Gestalt II, it is a brief test that asks a client to copy a number of figures. Interpretation of client drawings reveals informa- tion about the client’s developmental level and psychological functioning, as well as neurological deficits after a traumatic brain injury.

As the chapter continued, we looked at a num- ber of drawing tests, including the House- Tree-Person (HTP), the Kinetic House-Tree-Person (KHTP), Draw-A-Man, Draw-A-Woman, and Kinetic Family Drawing (KFD). We pointed out that these drawings symbolize issues and develop- mental changes in a person’s life and that interpre- tation of such drawings takes a well-trained clinician who understands the multiple meanings of objects.

Sentence completion tests were the final type of projective test we examined in the chapter. We noted that such tests present stems to which the client can respond. We pointed out that sentence completion tests provide a quick method of acces- sing a client’s feelings and unconscious thoughts about important issues in his or her life.

As the chapter neared its conclusion, we highlighted the fact that clinicians in all settings should consider when it might be appropriate to use a clinical assessment tool. Such tools can add to the clinician’s understanding of the client and can aid in treatment planning. We concluded by noting that the clinical assessment process results in making decisions for clients that often will affect their lives in critical ways and that test givers should remember the impact these decisions will have on clients.

276 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

CHAPTER REVIEW 1. Describe some of the uses of clinical

assessment. 2. Distinguish between objective and projective

testing and compare and contrast how each can be used in clinical assessment.

3. For each of the following tests, describe its main purpose, the kinds of scales that are used in test interpretation, and the population for which it is geared. a. Minnesota Multiphasic Personality

Inventory (MMPI-2) b. Millon Clinical Multiaxial Inventory

(MCMI-III) c. Personality Assessment Inventory (PAI) d. Beck Depression Inventory II (BDI-II) e. Beck Anxiety Inventory (BAI) f. Myers-Briggs Type Indicator (MBTI®) g. 16 Personality Factors Questionnaire

(16PF) h. NEO Personality Inventory-3 (NEO PI-3) i. Conners 3

j. Substance Abuse Subtle Screening Inventory (SASSI-3)

4. For each of the following tests, briefly describe its main purpose, how it is given, the population for which it is geared, and how test results are interpreted. a. Thematic Apperception Test (TAT) b. Rorschach Inkblot Test c. Bender Visual-Motor Gestalt Test, Second

Edition d. House-Tree-Person Test e. Kinetic House-Tree-Person Test f. Sentence Completion Series g. Rotter Incomplete Sentences Blank

5. Because many projective tests do not have the level of test worthiness of objective personal- ity tests, some argue that their use should be curtailed. What are your thoughts on this subject?

6. Discuss the role of helpers in objective and projective testing.

REFERENCES Arffa, S. (2010). Review of Conners, third edition. In

R. A. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measure- ments. Retrieved from the Burros Institute’s Mental Measurements Yearbook online database.

Axford, S., & Boyle, G. (2005). Review of the Taylor- Johnson Temperament Analysis, 2002 edition. In J. C. Conoley & J. C. Impara (Eds.), The sixteenth mental measurements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from the Mental Measurements Yearbook database.

Bart, O., Hajami, D., & Bar-Haim, Y. (2007). School adjustment from motor abilities in kindergarten. Infant and Child Development, 16, 597–615. doi:10.1002/icd.514

Beck, A. T., Steer, R. A., & Brown, G. K. (2003). BDI-II manual. San Antonio, TX: Psychological Corporation.

Beck, J. (1995). Cognitive therapy: Basics and beyond. New York: Guilford Press.

Bernt, F., & Frank, M. L. (2001). Review of the marital satisfaction inventory—revised. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measure- ments yearbook (pp. 710–714). Lincoln, NE: Buros Institute of Mental Measurements.

Boyle, G. J. (1995). Review of the Rotter Incomplete Sentences Blank, second edition. In J. C. Conoley & J. C. Impara (Eds.), The twelfth mental measure- ments yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Mea- surements Yearbook database.

Boyle, G. J., & Kavan, M. G. (1995). Review of the Personality Assessment Inventory. In J. C. Conoley and J. C. Impara (Eds.), The Twelfth Mental Mea- surements Yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Mea- surements Yearbook database.

Brannigan, G., Decker, S., & Madsen, D. (2004). Inno- vative features of the Bender-Gestalt II and expanded guidelines for the use of the global scoring

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 277

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

system. Assessment Service Bulletin Number 1. Retrieved from http://www.riversidepublishing .com/products/bender/pdf/BenderII_ASB1.pdf

Buck, J. (1948). The H-T-P Test. Journal of Clinical Psy- chology, 4, 151–159. doi:10.1002/1097-4679(194804) 4:2<151::AID-JCLP2270040203>3.0.CO;2-O

Burns, R. (1987). Kinetic House-Tree-Person drawings (K-H-T-P): An interpretative manual. New York: Brunner/Mazel.

Butcher, J., Dahlstrom, W., Graham, J., Tellegen, A., & Kaemmer, B. (1989). Manual for administration and scoring: MMPI-2. Minneapolis, MN: University of Minnesota Press.

Camara, W., Nathan, J., & Puente, A. (2000). Psycholog- ical test usage: Implications in professional psychol- ogy. Professional Psychology: Research and Practice, 31, 141–154. doi:10.1037/0735-7028.31.2.141

CPP. (2009a). Myers-Briggs Type Indicator. Retrieved from https://www.cpp.com/products/mbti/index.aspx

CPP. (2009b). MBTI® complete (R). Retrieved from https://www.cpp.com/en/mbtiproducts.aspx?pc=157

CPP. (2009c). MBTI step II® profile—Form Q (R). Retrieved from https://www.cpp.com/en/mbtipro ducts.aspx?pc=49

Decker, S. L., Allen, R., & Choca, J. P. (2006). Con- struct validity of the Bender-Gestalt II: Comparison with the Wechsler intelligence scale for children-III. Perceptual and Motor Skills, 102, 133–141. doi:10.2466/pms.102.1.133-141

Decker, S. L., Englund, J. A., Carboni, J. A., & Brooks, J. H. (2011). Cognitive and developmental influ- ences in visual-motor integration skills in young children. Psychological Assessment, 23, 1010–1016. doi:10.1037/a0024079

Dowd, E. T. (1998). Review of the Beck Anxiety Inven- tory. In J. C. Impara & B. S. Plake (Eds.), The thir- teenth mental measurements yearbook. Retrieved from the Burros Institute’s Mental Measurements Yearbook online database.

Dunn, T. M. (2010). Review of Conners, third edition. In R. A. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements year- book. Retrieved from the Burros Institute’s Mental Measurements Yearbook online database.

Exner, J. (1974). The Rorschach: A comprehensive sys- tem. New York: Wiley.

Fleenor, J. (2001). Review of the Myers-Briggs Type Indicator Form M. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements

yearbook (pp. 816–818). Lincoln, NE: Buros Insti- tute of Mental Measurements.

Graham, J. (2000). MMPI-2: Assessing personality and psychopathology (3rd ed.). New York: Oxford Uni- versity Press.

Groth-Marnat, G. (2003). Handbook of psychological assessment (4th ed.). Hoboken, NJ: Wiley.

Hogan, T. P. (2005). Widely used psychological tests. In G. P. Koocher, J. C. Norcross, & S. S. Hill (Eds.), Psychologists’ desk reference (2nd ed., pp. 101–104). New York: Oxford University Press.

Jung, C. (1964). Psychological types (H. G. Baynes, Trans.). London: Pantheon. (Original work pub- lished in 1921.)

Karon, B. (2000). The clinical interpretation of the Thematic Apperception Test, Rorschach, and other clinical data: A reexamination of statistical versus clinical prediction. Professional Psychology: Research and Practice, 31, 230–233. doi:10.1037/0735-7028 .31.2.230

Larson, L. M., Rottinghaus, P. J., & Borgen, F. H. (2002). Meta-analyses of big six interests and big five personality factors. Journal of Vocational Behav- ior, 61, 217–239. doi:10.1006/jvbe.2001.1854

Lazowski, L. E., Miller, F. G., Boye, M. W., & Miller, G. A. (1998). Efficacy of the Substance Abuse Subtle Screening Inventory-3 (SASSI-3) in identifying sub- stance dependence disorders in clinical settings. Journal of Personality Assessment, 71, 114–128. doi:10.1207/s15327752jpa7101_8

Lilienfeld, S. O., Wood, J. M., & Garb, H. N. (2000). The scientific status of projective techniques. Psycho- logical Science in the Public Interest, 1(2), 27–66.

Logan, R. E., & Waehler, C. A. (2001). The Rotter incomplete sentence blank: Examining potential race differences. Journal of Personality Assessment, 76, 448–460. doi:10.1207/S15327752JPA7603_06

Mastrangelo, P. (2001). Review of the Myers-Briggs Type Indicator Form M. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements yearbook (pp. 818–819). Lincoln, NE: Buros Insti- tute of Mental Measurements.

McCrae, R. R., & Costa, P. T., Jr. (2007). Brief versions of the NEO-PI-3. Journal of Individual Differences, 28, 116–128. doi:10.1027/1614-0001.28.3.116

McCrae, R. R., Costa, P. T., Jr., & Martin, T. A. (2005). The NEO-PI-3: A more readably revised NEO per- sonality inventory. Journal of Personality Assessment, 84, 261–270. doi:10.1027/1614-0001.28.3.116

278 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

McCrae, R. R., Martin, T. A., & Costa, P. T., Jr. (2005). Age trends and age norms for the NEO Personality Inventory-3 in adolescents and adults. Assessment, 12, 363–373. doi:10.1177/1073191105279724

McDevitt-Murphy, M. E., Weathers, F. W., Flood, A. M., Eakin, D. E., & Benson, T. A. (2007). The util- ity of the PAI and the MMPI-2 for discriminating PTSD, depression, and social phobia in trauma- exposed college students. Assessment, 14, 181–195. doi:10.1177/1073191105279724

Miller, G. A. (1999). The SASSI manual: Substance abuse measures (2nd ed.). Springville, IN: SASSI Institute.

Miller, G. A., & Lazowski, L. E. (2001). Adolescent SASSI-A2 manual. Springville, IN: SASSI Institute.

Moreland, K., & Werner, P. (2001). Review of the Sen- tence Completion Series. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements yearbook (pp. 1109–1110). Lincoln, NE: Buros Institute of Mental Measurements.

Multi-Health Systems. (2013). Conners 3™. Retrieved from http://www.mhs.com/product.aspx?gr=cli&prod= conners3&id=overview

Naglieri, J., Obrzut, J. E., & Boliek, C. (1992). Review of the Bender Gestalt Test. In J. J. Kramer & J. C. Conoley (Eds.), The eleventh mental measurements yearbook (pp. 101–106). Lincoln, NE: Buros Insti- tute of Mental Measurements.

Neukrug, E., Peterson, C., Bonner, M., & Lomas, G. (2013). A national survey of assessment instruments taught by counselor educators. Counselor Educa- tion and Supervision, 52, 207–221.

PAR. (2012a). Personality Assessment Inventory™

(PAI®). Retrieved from http://www4.parinc.com /Products/Product.aspx?ProductID=PAI

PAR. (2012b). NEO Inventories: NEO Personality Inventory-3 (NEO-PI-3). Retrieved from http:// www4.parinc.com/Products/Product.aspx?ProductID= NEO-PI-3

PAR. (2012c). NEOTM Inventories: NEOTM Five-Factor Inventory-3 (NEOTM-FFI-3). Retrieved from http:// www4.parinc.com/Products/Product.aspx?ProductID =NEO-FFI-3

PAR. (2012d). Sentence Completion Series (SCS). Retrieved from http://www4.parinc.com/Products /Product.aspx?ProductID=SCS

Pearson. (2012a). Minnesota Multiphasic Personality Inventory®-2. Retrieved from http://psychcorp .pearsonassessments.com/HAIWEB/Cultures/en-us /Productdetail.htm?Pid=MMPI-2&Mode=summary

Pearson. (2012b). Minnesota Multiphasic Personality Inventory-2-RF®. Retrieved from http://psychcorp .pearsonassessments.com/haiweb/cultures/en-us /productdetail.htm?pid=PAg523

Pearson. (2012c). MCMI-III: Millon® Clinical Multiax- ial Inventory-Ill. Retrieved from http://psychcorp .pearsonassessments.com/HAIWEB/Cultures/en-us /Productdetail.htm?Pid=PAg505

Pearson. (2012d). MACI: Millon® Adolescent Clinical Inventory. Retrieved from http://psychcorp.pearson assessments.com/HAIWEB/Cultures/en-us/Product detail.htm?Pid=PAg501

Pearson. (2012e). Beck Anxiety Inventory. Retrieved from http://www.pearsonassessments.com/HAIWEB/ Cultures/en-us/Productdetail.htm?Pid=015-8018-400

Pearson. (2012f). 16PF®, fifth edition. Retrieved from http://www.pearsonassessments.com/HAIWEB /Cultures/en-us/Productdetail.htm?Pid=PAg101& Mode=summary

Pearson. (2012g). Bender Visual-Motor Gestalt Test (2nd ed.). Retrieved from http://psychcorp.pearso- nassessments.com/HAIWEB/Cultures/en-us/Product- detail.htm?Pid=015-8064-127&Mode=summary

Pearson. (2012h). Rotter Incomplete Sentence Blank®, second edition (RISB®-2). Retrieved from http://www. pearsonassessments.com/HAIWEB/Cultures/en-us /Productdetail.htm?Pid=RISB-2&Mode=summary

Peterson, C., Lomas, G., Neukrug, E., & Bonner, M. (2014). Assessment use by counselors in the United States: Implications for policy and practice. Journal of Counseling and Development, 92, 90–98.

ProQuest. (2013). ProQuest dissertations & theses full text. Search using the terms Myers Briggs. Retrieved from the ProQuest Dissertations and Theses Full Text database.

Quenk, N. (2009). Essentials of Myers-Briggs Type Indi- cator assessment. Retrieved from http://books.google .com/books?hl=en&lr=&id=ZqPFa8jy1s4C&oi=fnd& pg=PR9&dq=Myers-Briggs+Type+Indicator&ots=tJZD7 M1UcR&sig=CR0IwfUC6WYBTXXHIyjaNJmWMJA

Raad, B. D. (1998). Five big, big five issues: Rationale, content, structure, status, and crosscultural assess- ment. European Psychologist, 3, 113–124.

Riverside Publishing. (2010). Bender® Visual-Motor Gestalt Test (Bender-Gestalt II). Retrieved from http://www.riversidepublishing.com/products/bender /details.html#technical

Ronan, G. F., Gibbs, M. S., Dreer, L. E., & Lombardo, J. A. (2008). Personal problem-solving system— revised. In S. R. Jenkins (Ed.), A handbook of clinical

CHAPTER 11 Clinical Assessment: Objective and Projective Personality Tests 279

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

scoring systems for thematic apperception techniques (pp. 181–207). New York: Erlbaum.

Russell, M. T., & Karol, D. L. (1994). The 16PF fifth edition administrator’s manual. Champaign, IL: Institute for Personality and Ability Testing, Inc.

SASSI (Substance Abuse Subtle Screening Inventory). (2001). Clinical interpretation training outline. Author: Springville, IN.

SASSI (Substance Abuse Subtle Screening Inventory). (2009). The SASSI Institute—Adult Substance Abuse Subtle Screening Inventory—3. Retrieved from http:// www.sassi.com/products/SASSI3/shopS3-pp.shtml

Saulsman, L. M., & Page, A. C. (2004). The five-factor model and personality disorder empirical literature: A meta-analytic review. Clinical Psychology Review, 23, 1055–1085. doi:10.1016/j.cpr.2002.09.001

Stein, M. B., Pinsker-Aspen, J. H., & Hilsenroth, M. J. (2007). Borderline pathology and the Personality Assessment Inventory (PAI): An evaluation of crite- rion and concurrent validity. Journal of Personality Assessment, 88, 81–89.

Sundberg, N. (1961). The practice of psychological test- ing in clinical services in the United States. American Psychologist, 16, 79–83. doi:10.1037/h0040647

Thurstone, L. L. (1934). The vectors of the mind. Psy- chological Review, 41, 1–32. doi:10.1037/h0075959

TJTA. (2013). Taylor Johnson Temperament Analysis®. Retrieved from https://www.tjta.com/asp/index.asp

Waller, N. G. (1998). Review of the Beck Anxiety Inventory [1993 edition]. In J. C. Impara & B. S.

Plake (Eds.), The thirteenth mental measurements yearbook. Retrieved from the Buros Institute’s Men- tal Measurements Yearbook online database.

Weis, R., Toolis, E. E., & Cerankosky, B. C. (2008). Construct validity of the Rotter Incomplete Sen- tences Blank with clinic-related and nonreferred adolescents. Journal of Personality Assessment, 90, 564–573. doi:10.1080/00223890802388491

Weiss, W. U., & Weiss, P. A. (2010). Use of Personality Assessment Inventory in police and security personnel selection. In P. A. Weiss (Ed.), Personality assessment in police psychology: A 21st century perspective. Springfield, IL: Charles C. Thomas Publishers.

White, L. J. (1996). Review of the Personality Assessment Inventory (PAI-™): A new psychological test for clini- cal and forensic assessment. Australian Psychologist, 31, 38–40. doi:10.1080/00050069608260173

Woike, B. A., & McAdams, D. P. (2001). TAT-based personality measures have considerable validity. American Psychological Society Observer, 14, 10.

WPS. (n.d.). Marital Satisfaction Inventory, Revised (MSI-R). Retrieved from http://portal.wpspublish. com/portal/page?_pageid=53,103808&_dad=portal& _schema=PORTAL

Youngstrom, E. A. (2013). Future directions in psycho- logical assessment: Combining evidenced-based medicine innovations with psychology’s historical strengths to enhance utility. Journal of Clinical Child and Adolescent Psychology, 42, 139–159.

280 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C H A P T E R 12

Informal Assessment: Observation, Rating Scales, Classification Methods, Environmental Assessment, Records and Personal Documents, and Performance- Based Assessment

I was teaching a graduate course in testing, and because I knew that some students have trouble understanding basic test statistics, I went out of my way to make sure I was avail- able for them. I gave them my home number and cell phone number, and I met with them individually if they wanted extra help. In addition, I offered two extra classes on the week- end for those who might need additional assistance in understanding some of the more diffi- cult concepts. At the end of the semester, students had an opportunity to rate the class using a 6-point rating scale (1 is low, 6 is high). They rated seven aspects of the class, including such things as the amount they learned, how useful the information was, how helpful and sensitive the instructor was, and so forth. After this particular course was finished, I looked at my ratings, which had a mean of about 5.3. Not bad, I thought to myself. However, to my chagrin, I noticed that the rating for “helpfulness” was only 4.8, the lowest of all the items. I reflected on why this was so, and I couldn’t come up with an answer. I must admit it bothered me, especially since I had gone out of my way to offer additional help. The next semester I saw one of the students from the class and asked her about the low rating. She said, “Well, the class was at night (7:10–9:50 PM), and many of the students were upset that you went the whole time.” I said, “Well, I was supposed to go the whole time.” She said, “Yeah, I know, but a lot of them just wanted to get home.” I suddenly realized that some of the students may have given me a lower rating for doing what I was supposed to be doing—teaching in the allotted time period. “What a drag,” I said to myself. Well, I guess that’s what happens when you deal with the subjectivity of rating scales!

(Ed Neukrug)

281

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

This chapter is about informal assessment procedures. Whether it is an end-of-semester evaluation of faculty or a dream journal of a client, informal assessment techniques can have a huge impact on a person. In this chapter, we start by defining informal assess- ment and then identify a number of different kinds of informal assessment techniques, including observation, rating scales, classification methods, environmental assessment techniques, records and personal documents, and performance-based assessment. We then discuss the test worthiness of these kinds of techniques and talk about how to assess their validity, reliability, cross-cultural fairness, and practicality. Next, we discuss the helper’s role in the use of informal procedures. We conclude the chapter with some final thoughts regarding informal assessment.

DEFINING INFORMAL ASSESSMENT By their very nature, informal assessment techniques are subjective and thus have a unique role in the assessment process. Whereas the kinds of assessment we have examined up to this point in the text were, for the most part, created and produced in association with national publishing companies and used nationally, most of the assessments in this chapter are “homegrown”; that is, developed by individuals who have specific assessment needs. Being homegrown, the amount of time, money, and expertise put into their development is generally much less than it is for those nationally developed instruments. Thus, reliability, validity, and cross- cultural issues are generally not formally addressed and often are lacking. Despite this obvious drawback, such instruments can be a practical addition to the assess- ment process and supply valuable information when assessing an individual.

Although informal assessment techniques are generally not as test worthy as formal assessment procedures, they have some advantages over them:

1. They add to the total assessment process and thus increase our ability to better understand the whole person.

2. They can be designed to assess the exact attribute we want to measure, whereas formal assessment techniques provide a wide net of assessment— sometimes so wide that we do not gain enough information about the specific attribute being examined.

3. They can often be developed or gathered in a rather short amount of time, providing important information in a timely fashion.

4. They can be nonintrusive; that is, you are not directly gathering information from the client (e.g., a student’s cumulative record file at school). Thus, they provide a nonthreatening mechanism for gathering information about the individual.

5. They generally are free or low-cost procedures. 6. They tend to be easy to administer and relatively easy to interpret.

TYPES OF INFORMAL ASSESSMENT Dozens of informal assessment techniques can enhance our understanding of the whole person. In fact, because informal assessment techniques are, by their nature, informal, a creative examiner can often come up with unique procedures never

Informal assessment procedures “Homegrown” meth- ods developed to meet specific needs

Although informal assessment proce- dures lack test worthi- ness, they have distinct advantages

282 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

used before. However, several types of informal assessment techniques are frequently used, including observation, rating scales, classification methods, environmental assessment techniques, records and personal documents, and performance-based assessment. Let’s take a look at each of these.

Observation An important tool to help understand our clients, observation can be done by pro- fessionals who wish to observe the individual (e.g., observing students in the class- room, a family in the home); by significant others who have the opportunity to observe the individual in natural settings (e.g., parents observing a child at home); and even by oneself, such as when a client is asked to observe specific targeted behaviors he or she is working on changing (e.g., eating habits) (see Box 12.1). Observation comes in many forms, with two of the more important ones being event sampling and time sample.

Event sampling is the viewing and assessment of a targeted behavior without regard for time. Usually, in evaluating the targeted behavior, general comments are made about the behavior or a rating scale is used to evaluate the behavior. For instance, the school counselor who is interested in observing the “acting-out” behavior of a student could view the child for an entire school day and note when the acting-out behavior is exhibited. Of course, “acting out” would have to be clearly defined.

Whereas event sampling is conducted with little concern for time, time sampling is when an individual is observed for a predetermined and limited amount of time. For instance, because it would be inconvenient for a school counselor to spend a whole day observing one student, the counselor might choose three 15-minute periods during the day to observe a student. Ideally, this time sample would give the counselor a snapshot of the student’s typical kinds of behaviors. The upside is that the counselor does not have to spend his or her whole day

BOX 12.1 Disruptive Observation of a Third-Grade Class

I was once asked to “debrief” a third-grade class that had just finished a trial period in which a young boy, who was paraplegic and severely intellectually disabled, had been mainstreamed into their class- room. During this trial period, a stream of observers from a local university had visited the classroom. The observers would sit in the back and take notes about the interactions between the students. This informa- tion was to be used at a later date to decide whether it was beneficial to all involved to mainstream the student with the disability. When I met with the stu- dents, they clearly had adapted well to the presence

of this young boy in their classroom. Although the students seemed to have difficulty forming any close relationships with him, his presence seemed in no way to detract from their studies or from their other relationships in the classroom. However, almost without exception, the students noted that the constant stream of observers interfering with their daily schedule had been quite annoying. Perhaps, if a limited time sample had been used, the students would not have reacted so strongly.

—Ed Neukrug

Observation Conducted by profes- sionals, significant others, or clients themselves

Event sampling Observing a targeted behavior with no regard to time

Time sampling Observing behaviors during a set amount of time

© Ce ng ag e Le ar ni ng

CHAPTER 12 Informal Assessment 283

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

observing one student. The downside is that there is a chance the time sample will not be an accurate portrayal of the student’s typical ways of being.

For the sake of convenience, time and event sampling are often combined. For example, the school counselor mentioned earlier could decide to observe a student’s acting-out behaviors (event sample) for three 15-minute segments of time (time sample) during the day. Clearly, scheduling the observation times would free up the counselor to do other important tasks during the day.

Time and event sampling are often used when an instructor wants to view or listen to recordings of students practicing their clinical skills (e.g., a student work- ing with a client at an internship, a student teacher in a classroom). Instead of lis- tening to or viewing all of the students’ work, instructors will often listen to or view targeted responses (e.g., the ability to be empathic, the kinds of questions asked) for a specified number of times (e.g., three times) for a designated amount of minutes (e.g., 10 minutes). Listening to all of the recordings of all of the students in an internship would be nearly impossible as it would simply take up too much time (see Exercise 12.1).

Rating Scales A rating scale is an instrument used by an individual to evaluate or assess attributes or characteristics being presented to the rater. Whereas formal ratings scales go through an impressive validation process, informal rating scales, like the ones we discuss in this chapter, tend to be quite subjective as their results are based on the “inner judgment” of the rater (Sajatovic & Ramirez, 2012). However, because such scales are easily developed, quick to administer, usually focused to assess a specific area or attribute, and usually free or low cost, they are often used by men- tal health professionals. Of course, due to the fact that such scales are subjective, their results should be considered cautiously and understood within the context in which they were developed.

Although a number of sources of error exist for rating scales, two of the most frequently cited are the halo effect and generosity error (Gay, Mills, & Airasian, 2012). The halo effect, which was first alluded to almost 100 years ago (see Box 12.2), occurs when an overall impression of an individual clouds the rating of that person in one select area, while generosity error occurs when the individual doing the rating identifies with the person being rated and thus rates the individual inaccurately. An example of the halo effect is the supervisor who mistakenly rates an outstanding supervisee high on being punctual, despite the fact that the supervi- see consistently comes to work late. The supervisor’s overall favorable impression has caused an error in the rating of this one attribute. An example of a generosity error is the student who ranks a fellow student high on exhibiting effective

Exercise 12.1 Application of Observational Techniques

In small groups, devise ways that you might use obser- vational techniques. Try to incorporate the use of event

sampling and time sampling in your examples. Share your answers in class.

Time and event sampling Observing a targeted behavior for a set amount of time

Rating scales Subjective quantifica- tion of an attribute or characteristic

Halo effect Overall impression of client causes inaccu- rate rating

Generosity error Identification with cli- ent causes inaccurate rating

© Ce ng ag e Le ar ni ng

284 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

empathic responses because the first student identifies with the anxiety the second student feels about being under the microscope.

Despite some of the potential problems of error with rating scales, they are easily created and can be completed quite quickly, so they offer a convenient mechanism of assessing individuals. Commonly used rating scales include numerical scales, Likert- type scales (graphic scales), semantic differential scales, and rank-order scales.

Numerical Scales Numerical scales generally provide a written statement or question that can be rated from high to low on a number line. Such scales are com- monly used during counseling as a way to prioritize issues or assess progress. For instance, the technique of scaling, commonly used in solution-focused therapy, asks clients to subjectively rate themselves on a scale from 0 to 10 that assesses any of a number of experiences, feelings, or behaviors (De Jong & Berg, 2013). Although clients in solution-focused therapy generally do this verbally, it would be quite easy to have the scale written, which would allow for documentation of prog- ress made in therapy (see Box 12.3).

Similarly, the Subjective Unit of Discomfort (SUD) scale is often used in the behavioral therapy technique of systematic desensitization in determining whether to move on to the next level in a hierarchy when imagining anxiety-provoking situations. Using this scale, 0 equals the most relaxing experience imagined by a cli- ent and 100 equals the most anxiety-evoking experience imagined by the client.

BOX 12.2 A Human Tendency?

Edward L. Thorndike, one of the early pioneers in modern-day assessment, was one of the first to write of the halo effect way back in 1920, in an article about error in ratings. Our tendency to see people as all one way or another can affect the judgments we make:

In a study made in 1915 of employees of two large industrial corporations, it appeared that the estimates of the same man in a number of different traits such as intelligence, industry, technical skill, reliability, etc., etc, were very highly correlated and very evenly correlated.

It consequently appeared probable that those giving the ratings were unable to analyze out these different aspects of the person’s nature and achievement and rate each in independence of the others. Their ratings were apparently affected by a marked tendency to think of the person in general as rather good or rather inferior and to color the judgments of the qualities by this gen- eral feeling. This same constant error toward suffusing ratings of special features with a halo belonging to the individual as a whole appeared in the ratings of officers made by their superiors in the army. (Thorndike, 1920, p. 25; italics not in original)

BOX 12.3 Numerical Scale

With 0 being equal to your depression being the worst it ever was and 10 being equal to the best

you could possibly feel, can you tell me where on a scale of 0 – 10 you are today?

0 1 2 3 4 5 6 7 8 9 10

Worst Depression Best I Could Feel

Numerical scale Statement or question followed by a number line

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

20 15

CHAPTER 12 Informal Assessment 285

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Clients move up the hierarchy when they experience a 10 or less at their current level (Ferguson & Sgambati, 2008).

Likert-Type Scales (Graphic Scales) Likert-type scales, sometimes called graphic scales, contain a number of items that are being rated on the same theme and are anchored by both numbers and a statement that correspond to the numbers. Take a look at the items in Box 12.4, which represent a few behaviors drawn from some research conducted on mental health practitioners’ perceptions of ethical behaviors (Neukrug & Milliken, 2011).

Semantic Differential Scales Semantic differential scales provide a statement that is followed by one or more pairs of words that are across from one another along a line and that reflect opposing traits. The rater is asked to place a mark on the line that is reflective of how much of the quality that rater believes he or she has. A number line may or may not be associated with the dichotomous pairs of words. Box 12.5 shows a semantic differential scale with a number line.

BOX 12.4 Likert-Type Scale

Please indicate how strongly you agree or disagree with each of the following statements:

Strongly disagree

Somewhat disagree

Neither agree nor disagree

Somewhat agree

Strongly agree

* It is fine to view a client’s personal web page (e.g., Facebook, blog) without informing the client.

1 2 3 4 5

* It is okay to tell your client you are attracted to him or her.

1 2 3 4 5

* I have no problem counseling a terminally ill client about end-of-life decisions, including suicide.

1 2 3 4 5

* Self-disclosing to a client is ethical. 1 2 3 4 5

BOX 12.5 Semantic Differential Scale

Place an “X” on the line to represent how much of each quality you possess.

Sadness Happiness 1 2 3 4 5 6 7 8 9 10

Introverted Extroverted 1 2 3 4 5 6 7 8 9 10

Anxious Calm 1 2 3 4 5 6 7 8 9 10

Likert-type scale Items rated on same theme, anchored by numbers and a statement

Semantic differential scale Number line with opposite traits at each end

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

286 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

As you can see from the example in Box 12.5, a Semantic Differential Scale can easily be created if you were interested in quickly assessing the behaviors and/or affect of a client. Such a scale can be helpful in assessing clients’ needs and setting treatment goals. Of course, such scales can have many other broad educational and psychological applications.

Rank-Order Scales Rank-order scales provide a series of statements that the respondent can place in hierarchical order based on his or her preferences. For instance, the rank-order scale in Box 12.6 could be used as part of a larger instru- ment to determine preference for counseling style (see Exercise 12.2).

Classification Methods Whereas rating scales tend to assess a quantity of an attribute or characteristic or the preferences of attributes or characteristics (e.g., which is liked more than another), classification methods are all or nothing; they provide information about whether an individual has, or does not have, certain attributes or characteristics. Two com- mon classification methods include behavior checklists and feeling word checklists.

Behavior Checklists Behavior checklists allow an individual to identify those words that best describe typical or atypical behaviors he or she might exhibit. Behavior checklists are easy to develop, simple to give, and can quickly uncover patterns of behavior that are important to look at. Although their usefulness is clearly related to the kinds of behavior being surveyed and whether the client is

BOX 12.6 Rank-Order Scale

Please rank order your preferred method of doing counseling. Place a 1 next to the item that you most prefer, a 2 next to the item you second most prefer, and so on down to a 5 next to the item you prefer least.

I prefer listening to clients and then reflecting back what I hear from them to facilitate client self-growth.

I prefer advising clients and suggesting mechanisms for change.

I prefer interpreting client behaviors in the hope that they will gain insight into themselves.

I prefer helping clients identify which behaviors they would like to change.

I prefer helping clients identify which thoughts are causing problematic beha- viors and helping them to develop new ways of thinking about the world.

Exercise 12.2 Application of Rating Scales

In small groups, devise ways that you might be able to use rating scales. Try to incorporate the use of the

different types of rating scales discussed in the chapter. Share your answers in class.

Rank order A method for clients to order their preferences

Classification methods Information regarding presence or absence of attribute or characteristic

Behavior checklists Type of classification method that assesses behaviors

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

CHAPTER 12 Informal Assessment 287

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

honest in his or her responses, they should be considered as one more method of client assessment among the many that you might use. The number and types of behavior checklists are only limited by your imagination and what you might be able to find and freely use, for example, from the Internet. Box 12.7 is a behavior checklist that could be used to help identify abusive behaviors.

Feeling Word Checklists Like behavior checklists, feeling word checklists are another simple and practical classification method. By simply checking or circling the feeling on the list, a client can quickly identify a number of feelings that he or she has had, is currently experiencing, or hopes to feel. For instance, school counse- lors often use such feeling word checklists for younger children to help them begin to discriminate different kinds of feelings they may be having. Identification of feel- ings is the first step toward being able to effectively communicate one’s feelings to others. But identification of feelings is not only important for children. For instance, marital and relationship counselors might want each member of a couple to clearly understand what they are feeling so that each can effectively communi- cate with their partner. The ability to accurately identify what one is feeling is the first step toward the differentiation of the emotional from the intellectual process, often an important task in couples and family counseling (Guerin & Guerin, 2002). Table 12.1 is an example of a feeling word checklist that can be used in counseling to identify problem feelings. Like the behavior checklist, such lists are only limited by your ability to identify feeling words or find existing checklists.

Other Classification Methods The kinds of classification methods that one can create are innumerable, and as with behavior checklists, are only limited by your

BOX 12.7 Behavior Checklist of Abusive Behaviors

Check those behaviors you have exhibited toward your partner and your partner has exhibited toward you.

Exhibited by You to Your Partner

Exhibited by Your Partner to You

1. Hitting

2. Pulling hair

3. Throwing objects

4. Burning

5. Pinching

6. Choking

7. Slapping

8. Biting

9. Tying up

10. Hitting walls or other inanimate objects

11. Throwing objects with intent to break them

12. Restraining or preventing from leaving

Feeling word checklists Type of classification method that assesses feelings

© Ce ng ag e Le ar ni ng

288 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TABLE 12.1 Feeling Word Checklist to Identify Problematic Feelings

Abandoned Embarrassed Lying Stuck-up

Aggravated Empty Mean Stupid

Aggressive Envious Miserable Suppressed

Angry Exasperated Misunderstood Teased

Anxious Failure Nauseated Terrified

Appreciated Fearful Neglected Thoughtless

Apprehensive Forced Nervous Tormented

Argumentative Forlorn Obligated Traumatized

Arrogant Forsaken Oppressed Troubled

Ashamed Frigid Overwhelmed Troubling

Awful Frustrated Pained Unaccepted

Betrayed Futile Panicked Unconcerned

Bitter Grieving Paranoid Undesirable

Blind Guilty Pitiful Uneasy

Bored Hateful Powerless Unfriendly

Broken-hearted Haughty Pressured Unfulfilled

Burdened Helpless Provoked Unhappy

Claustrophobic Hopeless Punished Unhelpful

Concerned Humiliated Rejected Unneeded

Confused Hurried Repulsive Unloved

Criticized Hurt Resentful Unpleasant

Cut off Impatient Restless Unreliable

Deceived Imposed Restricted Unsociable

Defensive Impossible Sad Unsuccessful

Depressed Inadequate Scared Unwanted

Difficult Incapable Selfish Unworthy

Dirty Indecisive Shameful Upset

Disappointed Insecure Shattered Used

Discontented Intolerant Shocked Useless

Discouraged Irresponsible Shy Victimized

Disgusted Irritated Skeptical Vindictive

Disloyal Jealous Smug Wasted

Disrespected Left out Spiteful Wary

Distressed Let down Sorrowful Weary

Doubtful Lonely Sour Worthless

Drained Longing Stifled Wounded

Egotistical Lost Stubborn Wrong

© Ce ng ag e Le ar ni ng

CHAPTER 12 Informal Assessment 289

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

imagination. For instance, one can come up with a classification method that asks clients to examine and choose items that represent their “irrational thoughts”; an individual participating in a career assessment might be asked to check those jobs that look appealing; or an older person might be asked to pick, from a long list, any barriers to living fully that he or she faces (difficulty getting out of the bath, problems seeing, etc.).

Environmental Assessment An often neglected but valuable source of information is an individual’s environment. Frequently called environmental assessment, this kind of assessment includes collecting information from a client’s home, school, workplace, or other place of interest, usually through observation or self-reports. This form of appraisal is more naturalistic and contextual than in-office testing and can be eye-opening, because even when clients do not intentionally mislead their counselors, they will often present a distorted view based on their own inaccurate perceptions or because they are embarrassed about revealing some aspect of their lives (e.g., a person living in poverty might not want to reveal an unpleasant home situation). Environmental assessment can be conducted through direct observation, by conducting a situational assessment, by applying a sociometric assessment, or by using an environmental assessment instrument.

Direct Observation Visiting a client’s home, classroom, workplace, or other set- ting can yield a plethora of information that might otherwise evade even the most seasoned therapist. For instance, imagine if after administering your in-office assess- ment procedures and conducting your interview, you were to visit your client’s home and find that he had old magazines stacked to the ceiling. Or consider the 10- year-old little “angel” in your office who dramatically changes her demeanor when around other girls in her class. Or imagine your surprise when during an assessment of your client’s home situation, you find his wife to be kind, considerate, and attrac- tive, unlike the “raving, angry, and pathetic” spouse that he had portrayed her to be.

Home visits are almost always profitable, and it is not unusual to discover impor- tant information about your client that you would rarely ascertain in therapy (Yalom, 2002; see Box 12.8). Additionally, discussing the home visit with the client before- hand can generate productive conversations. For instance, you might notice that as you discuss an upcoming visit your client suddenly becomes anxious due to some

BOX 12.8 An Effective Home Visit

The master therapist Milton Erickson was known for practicing unusual interventions that were brief and often at the client’s home or workplace. For instance, there is the story of how he once spent less than an hour at the home of a woman who was depressed, lonely, and shy. By visiting her home, he realized she was an expert at cultivating African violets and that she was a committed member of her church. He quickly

came up with an intervention: Every time there was a birth, wedding, or death at the church she attended, she was to send one of her flowers. She did this reli- giously, which dramatically changed her life as other church members began to get to know and appreciate her. At her death, the newspaper noted that hundreds of people attended the funeral of the “African Violet Lady” (Gordon & Meyers-Anderson, 1981).

Environmental assessment Collecting information from a client’s home, school, or workplace via observation or self- report

Direct observation An environmental assessment through observation

© Ce ng ag e Le ar ni ng

290 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

reason not yet discussed. This anxiety can then be processed and explored during the session. Another benefit is the fact that the home visit can be construed by the client as an affirmation of your level of caring and commitment. This can do wonders in building rapport and trust. Visits to the workplace can also generate useful informa- tion; however, caution must be exercised to maintain the client’s confidentiality.

Finally, school and classroom visits are often useful in making assessments of children and adolescents. Such visits allow the helping professional to assess other environmental factors such as lighting, seat position, room layout, and distracting noises that may be related to a student’s underperformance. Understanding chil- dren’s social interactions can often be achieved only by observing the student in the classroom or on the playground.

Situational Assessment Another type of environmental tool sometimes used is a situational assessment. This kind of assessment uses simulated, but natural situa- tions to examine how a person is likely to respond in real-life situations. For instance, when applying to doctoral programs, I (Ed Neukrug) was asked to role- play a counselor with a faculty member. He role-played the same client with every potential student, and the faculty were then able to assess my clinical skills and compare them to those of other potential students in this contrived yet realistic role-play. Today, situational assessments are frequently used in business and indus- try to determine whether an employee has the skills to be promoted.

Sociometric Assessment Sociometric assessment is used to determine the rela- tive position and dynamics of individuals within a group, organization, or institu- tion. For instance, if I was interested in knowing how well a group of preschool students liked one another, I could ask each of them to privately tell me the name of the individual who was their best friend. Then I might ask them to tell me the name of their second-best friend. I could then map these rankings (see Figure 12.1).

Obviously, the kind of sociometric instrument that is used will vary based on the situation being assessed. For instance, you might want to ask a family: “Who holds the power in this family”; or you might want to ask employees: “Who is most willing to help out at work?”; or if issues of confidentiality and trust were assured, you could ask college students in a residence hall, “Who in your residence hall has most abused drugs and alcohol?”

Environmental Assessment Instruments To assist in environmental assess- ment, a number of instruments have been developed that are often used in conjunc- tion with simple observation. In contrast to the other instruments that we have looked at in this chapter, these tend to have more rigorous test construction pro- cesses. The following is a very brief list of some of them to give you a sense of how they assess the environment of the client (see Exercise 12.3).

• Comprehensive Assessment of School Environments Information Management System (CASE-IMS). This instrument is used to assess the entire school envi- ronment and climate. It uses self-report surveys of students, parents, teachers, and the principal. The data from the assessment indicate the school’s strengths and weaknesses as normed against other schools (Manduchi, 2001).

Situational assessments Role-play to determine how individuals might act

Sociometric assessment Used to assess the social dynamics of a group

Environmental assessment instruments Used with observation and more rigorously constructed than other informal instruments

CHAPTER 12 Informal Assessment 291

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

• Behavior Rating Inventory of Executive Function—Adult Version. This instru- ment is a 75-item questionnaire for adults between the ages of 18 and 90 that is used to assess higher-order cognitive functioning whose purpose is to assess one’s ability to regulate behavior, emotions, and thoughts. It is often used as part of an assessment for attention deficit disorders, learning disabilities, autism, brain injury, depression, and schizophrenia and related disorders that might impact executive functioning. It assesses nine areas labeled as inhibit, shift, emotional control, self-monitor, initiate, working memory, plan/organize, task monitor, and organization of materials. Raw scores are converted to T-scores based on the individual’s age group (Dean & Dean, 2007).

• Emotional or Behavior Disorder Scale-Revised. This instrument is completed by individuals who are “familiar with the student” and helps in the identifica- tion of students with emotional and behavioral disorders. There are two broad components that can be assessed through observation, usually by school per- sonnel: a behavioral component that has 64 items and a vocational component that has 54 items (Watson, 2005).

Solid line = Person liked most Social Isolate: Jermeiah Dotted line = Person liked second most Most-liked person: Hannah

Margo Martin

Emma Jeremiah

Hannah Brett

Izzi

FIGURE 12.1 | Sociometric Mapping at a Preschool

Exercise 12.3 Self-Reflection About Environmental Assessment

If you were in therapy and your therapist came to visit your home or work environment, would he or she learn something about you that you had not self-disclosed? In small groups, discuss why you might “neglect” to tell

your therapist certain things about your home or work environment and how useful an environmental assess- ment would be in revealing the “true” you.

© Ce ng ag e Le ar ni ng

© Ce ng ag e Le ar ni ng

292 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Records and Personal Documents Records and personal documents can help the examiner understand the beliefs, values, and behaviors of the person being assessed. Such materials can be obtained directly from the individual, from individuals close to the person (e.g., parents, loved ones), and from just about any institution with which the client has inter- acted, such as educational institutions, mental health agencies, and clients’ places of employment. Some of the more common records and personal documents include biographical inventories, cumulative records, anecdotal information, auto- biographies, journals and diaries, and genograms.

Biographical Inventories Biographical inventories provide a detailed picture of the individual from birth. They can be the result of an involved, structured inter- view that is conducted by the examiner, or they can be created by having the client complete a checklist or respond to a series of questions. Biographical inventories will often cover the same kinds of items that are found in a structured interview (see Chapter 4 for a description of the structured interview). The following kinds of information are often gathered in a biographical inventory:

• Demographic information • Age • Sex • Ethnicity • Address • Date of birth • E-mail address • Phone number(s) (home and cell)

• Presenting problem • Nature of problem • Duration of problem • Severity of problem • Previous treatment of problem

• Family of origin • Names and basic descriptions of primary caregivers in family of origin • Names and basic descriptions of siblings in family of origin • Names and basic descriptions of significant others involved with family of

origin • Placement in family of origin • Quality of relationships in family of origin • Traumatic events in family of origin

• Current family • Nature of relationship with current domestic partner • Length of relationship with current domestic partner • Names, ages, and basic descriptions of children, step-children, and foster

children • Names and basic descriptions of significant others involved with current family • Quality of relationships in current family • Traumatic events in current family

Records and personal documents Shed light on beliefs, values, and behaviors of the client

Biographical inventories Provide detailed picture of client’s life

CHAPTER 12 Informal Assessment 293

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

• Educational history • Highest degree attained by each parent and college major, if any • Highest degree attained by each spouse and college major, if any • Highest or current educational level of children

• Vocational background • Detailed work history from adolescence • Detailed work history of spouse/significant other • Work history of parents • Current salary of primary family members • Assessment of current job satisfaction of client and spouse or significant

other • Financial history

• Economic status/social class of family of origin • Current economic status/social class • History of any financial hardships

• History of counseling and mental illness • Detailed counseling/psychiatric history • Detailed counseling/psychiatric history of significant other • Detailed counseling/psychiatric history of children • History of mental illness in family of origin and extended family • Use of psychotropic medication in current family and family of origin

• Medical history • Significant medical problem • Significant medical problems of significant other • Significant medical problems of children, stepchildren, foster children • Significant medical problems in family of origin • Current status of medical problems in immediate family • Current use of medication: give names and dosages

• Substance use and abuse history • Cigarette smoking

None Number of years: Number of cigarettes per day:

• Alcohol use None Occasional Regular Heavy Binge

Number of years: • Illegal drugs

Type of drug(s) (Respond multiple times if more than one drug) Usage: None Occasional Regular Heavy Number of years:

• Prescription drug abuse Type of drug(s) (Respond multiple times if more than one drug) Usage: None Occasional Regular Heavy Number of years:

• History of legal issues • Nature of legal issue • Effect on self • Effect on others

294 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

• Check if changes have occurred in any of the following over the last six months:

Weight Appetite Sleep patterns Interest in sex Sexual activity General level of activity

Explain any items checked: • Sexual orientation

Heterosexual Bisexual Homosexual Transgender Unclear

• History of aggressive behaviors • Nature of past acting-out behaviors: • Current violent ideation: • Likelihood of acting out:

• History of self-injurious behaviors • Nature of past behaviors: • Current ideation (suicide or other): • Likelihood of injuring self:

• Affective and mental state (Check all that apply) Depressed Anxious Euphoric Tearful Sad Irritable Angry Passive Apathetic Delusions Hallucinations Emotional lability Panicky Compulsions Obsessive thoughts Phobic Passive Fearful Low self-esteem Guilt Memory problems

Explain any items checked (note intensity and duration):

Cumulative Records Almost assuredly, a cumulative record of significant beha- viors we have exhibited has been kept on most of us. For instance, cumulative records are commonplace in schools, where information about a child’s test scores, grades, behavioral problems, family issues, relationships with others, and other matters are stored. Also, most workplaces maintain some kind of cumulative record on each employee. These records can add vital information to our under- standing of the whole person and can generally be accessed with a written request form signed by the client.

Anecdotal Information Anecdotal information can sometimes be found in an individual’s cumulative record and generally includes behaviors of an individual that are consistent (e.g., Jonathan is always punctual) or inconsistent (Samantha, who generally gets along with her coworkers, had an altercation with one today). Anecdotal information can give us insight about the usual manner in which a per- son behaves or about inconsistent or rarely seen behaviors that may offer glimpses into the inner world of the client.

Cumulative records Collected documenta- tion from a school, employer, or mental health agency

Anecdotal information Subjective comments or notes in client’s re- cords regarding usual patterns or atypical behaviors

CHAPTER 12 Informal Assessment 295

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Autobiography In contrast to biographical inventories that help us collect in- depth and comprehensive information about a person, asking an individual to write an autobiography allows an examiner to gain subjective historical informa- tion that for some reason stands out in an individual’s life. In some ways, the infor- mation highlighted by an individual in his or her autobiography is a type of projective test in that the individual unconsciously chooses certain information to include—information that has affected the development of the individual’s sense of self. For creative and insightful individuals, writing an autobiography can be an enjoyable process that can reveal much about the self-awareness of the client.

Journals and Diaries Some individuals enjoy and benefit from keeping an ongo- ing journal or diary. For instance, dream journals can provide valuable insight into self by revealing unconscious drives and desires and by uncovering patterns that indicate issues in a client’s life. Clients can often learn how to “get in touch with their dreams” simply by keeping a dream journal next to their beds and writing down their memories of dreams as soon as they awake. In fact, it has been shown that individuals can be taught how to remember their dreams in this fashion (Gordon, 2007; Hill, 2004). Similarly, diaries can help uncover the inner world of the client and identify important patterns of behavior. By examining themes found in journals and diaries, helpers and clients may focus more attention on behaviors that may have seemed insignificant. Journals and diaries can add an important “inside” perspective to our understanding of the total person.

Genograms A popular and informative assessment tool is the genogram. It can also serve as a focal point or discussion aid during treatment. The genogram is a map of an individual’s family tree that may include the family history of illnesses, mental disorders, substance use, sexual orientation, relationships, cultural issues, and other items that might be of interest in counseling. Usually the therapist draws the genogram while asking the client questions; however, drawing or completing the genogram can be given as homework.

There are numerous symbols and items that can be included in the genogram and some typical ones are displayed at the bottom of Figure 12.2. Still others can be created at your discretion. Dates or ages are useful on the genogram. Some pre- fer adding the year born, and if applicable, the year deceased, near each indivi- dual’s name. Another method is to include the age of the individual inside his box or her circle. Similarly, the year married or the cumulative years married can be placed above the marriage relationship line, and a line separating two individuals is often used to identify a divorce. Sometimes, as in the genogram in Figure 12.2, individuals list geographic location of some or all of the family as location can be a sign of who gets along, who may have needed to distance self from others, and of cultural differences in styles of relating (individuals in the North tend to have different relating styles than those in the South). Identifying birth and death dates, physical problems, mental health problems, marriages and divorces, location of family members, and more can help highlight issues to work on in counseling and can help track family patterns that are passed down through the generations. Creating the genogram to at least the level of the client’s grandparents is important to fully capture these trends in the family (see Exercise 12.4).

Autobiography Asking a client to write his or her life story

Journals and diaries Having clients log their daily thoughts, ac- tions, or dreams

Genogram Map of client’s family relationships and rele- vant history

296 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Performance-Based Assessment Performance-based assessment, which has sometimes been offered as an alternative or addition to standardized testing, is the evaluation of an individual using a vari- ety of informal assessment procedures that are often based on real-world responsi- bilities. Occurring in many settings, including business, higher education, and for the purposes of credentialing (Feldman, Lazzara, Vanderbilt, & DiazGranado,

Walt

1920–2005

Virgina

1919

George1

1918–1992

Virginia5

1942

Eva

1908–1996

Benjamin

1905–1975

Williams

Ron

1941

Rebecca

1899–1989

Lovine

Melvin

1930

Artie

1925–2009

Herman

1891–1951

Mini

1898–1975

Neukrug

Sam4

1890–1985

Joseph4

1919–1976

Carole2

1946

Roy

1954

Joseph

1980

Edward 2, 4

1951

Hannah

1994

Emma

1999

Daniel

1991

Martin

2000

Stella

2009

Rich 1

1975

Kristina 1,5

1969

1, 2, 3, 4, 5 = Tags. These could represent any of a number of mental health problems, disorders, or physical disorders and should be clearly labeled (they are fictitious in this genogram). For instance, 1 = depression, 2 = schizophrenia, 3 = substance abuse, 4 = diabetes, etc.

Legend:

Kim 5

1965

Steve

1965 Ron

2

1967

Trang

1977

Amy

1956

Sarah

1996

Joseph

1994

Howard 2

1956

Eleanor 2

1921–2005

Heshey3

1927

Arkie

1925

Dietrich

= Male = Offspring= Stillborn or Abortion = Child in Utero= Female

Dietrich/Williams Cultural Considerations: Southern Christian (Catholic, Baptist, Methodist)

Levine/Neukrug Cultural Considerations: New York Jewish

= Marriage = Separation/Divorce = Death = Physical Custody

FIGURE 12.2 | Sample Genogram Source: Neukrug, E. (2007). The World of the Counselor. Belmont, CA: Brooks/Cole

Exercise 12.4 Application of Records and Personal Documents

In small groups, devise ways to use records and per- sonal documents. Try to incorporate the various types

of records and personal documents discussed in the chapter. Share your answers in class.

© Ce ng ag e Le ar ni ng

Performance-based assessment Assessment proce- dures based on real- world responsibilities

CHAPTER 12 Informal Assessment 297

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

2012), performance-based assessment provides a wide range of alternative methods of assessing an individual that is predictive of how well the individual will achieve at his or her current or eventual setting.

Standardized testing, which is highly loaded for the assessment of cognitive skills, has been shown to be predictive in a wide range of studies; however, the fact that these tests continue to show race-based differences raises many eyebrows (Sackett, Schmitt, Ellingson, & Kabin, 2001). Some suggest that such differences can be lessened if performance-based assessments that are predictive of desired out- comes are used in addition to or in lieu of such testing. This is particularly the case if the additional measures are based on interpersonal skills, personality measures, and other measures not highly loaded for cognitive skills (e.g., producing a successful counselor). The use of additional predictive methods would widen the number of individuals who could demonstrate competency in a specific area and increase the numbers of minorities who could potentially be chosen for a job, obtain a credential, or be admitted into an academic program. One kind of performance-based assess- ment that has become increasingly popular in recent years is the use of portfolios.

Portfolio assessments have become commonplace in teacher education pro- grams and are increasingly being used in training programs in the helping profes- sions (Cobia et al., 2005; Robles, 2012). Such portfolios ask students to pull together a number of items that demonstrate competencies in a wide range of areas that meet specific standards, such as accreditation standards. Oftentimes student portfolios include major projects that students have already finished and which have been previously evaluated. Sometimes, the projects include comments by teachers or supervisors who may have reviewed the student’s work. Students also typically have an opportunity to refine these projects before they are placed in the portfolio. Of course, additional items can be placed into a portfolio at any point.

As an example of what may be placed in a portfolio, a school counseling student might be asked to include any of the following in his or her portfolio: a resume, videos of the student’s work with clients, a supervisor’s assessment of the student’s work, a paper that highlights the student’s view of human nature, examples of how the student helped to build a multicultural school environment, ways that the student shows a commitment to the school counseling profession, a test report written by the student, or a project that shows how the student would build a comprehensive school counseling program. Although portfolios have, in the past, been a “paper” project, today’s portfolios are often placed on a CD or online (Robles, 2012).

The validity of portfolio assessment is based on a number of factors, includ- ing ensuring that the content in the portfolio reflects the competencies and/or standards that the portfolio is based on, a rubric or scoring system that is clear and based on the domains being assessed, the outcomes hoped for (e.g., counselor competency with clients) that should reflect the competencies the portfolio is based on, and that the portfolio should predict outcomes for what is being assessed as well as related skills (e.g., counselor success with clients) (vander Schaaf, & Stokking, 2008).

Performance-based assessment has become an alternative to traditional ways of assessing an individual and will likely to continue to gain in popularity. As you con- tinue on your career path and become responsible for the assessment of others, you might want to consider using some kinds of performance-based assessment procedures.

Portfolio assessment Performance-based assessment often found in higher education

298 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

TEST WORTHINESS OF INFORMAL ASSESSMENT As already noted earlier in this chapter, informal assessment techniques tend to be less valid, less reliable, and less cross-culturally sensitive than formal assessment techniques, although they do tend to be quite practical. But let’s take a look at what we mean when we talk about these attributes within the context of informal assessment techniques.

Validity Validity of informal assessment techniques has to do with how well the examiner is defining that which is being assessed. For instance, if I am concerned about the acting-out behavior of a middle school child, I need to clearly define the behavior identified as “acting out,” and I need to be specific on where acting-out behaviors are being exhibited. For example, does acting out include only inappropriate beha- viors in the classroom, or does it also include inappropriate behaviors in the hall- way, on the playground, on field trips, and at home? Also, exactly which “inappropriate” behaviors are we talking about? Does acting out include pushing, interrupting, making inappropriate nonverbal gestures, or withdrawing in class? The more clearly one is able to identify the kinds of behaviors and the places in which the behaviors are exhibited, the easier it will be to collect information about that domain, and consequently, the more valid will be the assessment. Thus, when referring to informal assessment techniques, we are generally not assessing validity in the traditional ways, but are more concerned about how clearly we are defining the domain being measured (see Exercise 12.5).

Exercise 12.5 Validity of Informal Assessment Techniques

In small groups, consider the various kinds of informal assessment techniques highlighted in the following list and demonstrate how you might show evidence that the information being assessed is valid.

Observation • Event sampling • Time sampling • Event and time sampling

Rating Scales • Numerical scales • Likert-type scales (graphic scales) • Semantic differential scales • Rank-order scales

Classification Systems • Behavior checklists • Feeling word checklists

• Other “homemade” classification systems

Environmental Assessment • Direct observation • Situational tests • Sociometric assessment • Environmental assessment instruments

Records and Personal Documents • Biographical inventories • Cumulative records • Anecdotal information • Autobiography • Journals and diaries • Genograms

Performance-based Assessment • Portfolio assessment

Validity of informal assessment techniques Based on clearly defining a set of behaviors

© Ce ng ag e Le ar ni ng

CHAPTER 12 Informal Assessment 299

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Reliability With informal assessment, the better we define the behavior being assessed, the more reliable our data collection will be. Thus, with informal assessment there is an intimate relationship between validity and reliability. Ideally, when conducting informal assessment, a minimum of two individuals who are highly trained in understanding the meaning of the data being collected should collect and categorize the data. With two raters, one can conduct interrater reliability that looks at the statistical concordance between the raters (Gwet, 2010). However, because it is often difficult to find a second person to do a rating or observation, and because training two people to record similar can take days if not weeks, this is rarely done. For instance, it would be nearly impossible to have two trained individuals write down the dreams of an individual unless that individual was in a controlled setting, like a sleep clinic. Similarly, it is quite unusual to have two trained raters review and rate the video of a counselor–client role-play by a counselor trainee. Or, consider the earlier example of the acting-out middle school student. In this case, it would be ideal if two individuals observed the child at the same time and knew exactly what acting-out behaviors they were looking for and in which con- text to look for them (the classroom, playground, home, etc.). Subsequent to observing the child, they could crosscheck their observations and see if they had collected similar information. If this were possible, when comparing such accuracy statistically, we would be examining the interrater reliability of the observers.

For instance, when I (Ed Neukrug) was in college, I and a friend worked with a psychology professor who was examining how quickly thirsty mice could learn that a drop of water was on a receptacle on the black side of a box. The box had removable sides, one of which was black and the other was white. By randomly moving the sides, but always keeping the water in the receptacle on the black side, the learning curve of thirsty mice could be assessed. My friend and I would place a mouse in the middle of the box and separately rate whether or not the mouse went to the correct side (the black side) and took a drink. However, to increase our validity, and conse- quently the reliability of our ratings, we had to be clear about what was considered a learned behavior by the mice. For instance, was walking over to the black side and venturing toward the water receptacle, but not touching it, considered a learned response? What if the mouse walked over and looked into the receptacle but did not drink? And what if the mouse walked over to the receptacle, looked in, stuck its ton- gue out and touched the water, but didn’t appear to swallow? Did the mouse exhibit the learned response in this case? A “correct” response had to be clearly defined for us, and the more clearly it was defined, the more similar would be our ratings. In fact, in this case, good interrater reliability was said to have been obtained when the correlation between our ratings reached 0.80 or higher.

In another example of interrater reliability, I (Ed Neukrug), had, two graduate students learn how to rate the empathic responses of dozens of individuals who had responded to a role-play client on an audiotape. These two graduate students were taught how to rate empathy on a five-point scale using increments of 0.5, from a low of 0.5 to a high of 5.0. Thus, I had to be clear on what was meant by empathy and how to determine whether a response deserved a 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, or 5 rating. In both cases, it took weeks to train the two raters in their

Reliability of informal techniques Based on how accu- rately a behavior is defined

Interrater reliability The agreement or consistency among two or more evaluators

300 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

understanding of empathy and the use of the scale to the point where they had an interrater reliability of 0.80. Based on this example, it is easy to see why the use of two or more raters is rarely used as the training that must take place to obtain high enough interrater reliability is an arduous process.

Cross-Cultural Fairness The nature of informal assessment procedures makes them particularly vulnerable to cultural bias. This can happen in two ways. First, as a result of unconscious or conscious bias, an examiner, observer, or rater may misinterpret the verbal or non- verbal behaviors of the individual being observed. Second, the examiner, observer, or rater may simply be ignorant about the verbal or nonverbal behaviors of a par- ticular group. For instance, an example of unconscious bias affecting one’s obser- vation might be an examiner who is asked to observe the behaviors of an Asian student and mistakenly assumes a particular student is bright because Asians as a group tend to do better academically than other ethnic groups. An example of ignorance is when the same observer falsely assumes the student has psychological issues because the Asian student does not readily express feelings. In fact, expres- sion of feelings is valued in many Western cultures but not regarded in the same manner in many Asian countries (Robinson-Wood, 2013).

Despite the fact that informal assessment techniques are easily affected by bias, they still can add much to the assessment of an individual. Because they are uniquely geared toward the particular behavior of an individual, one can pick and choose exactly which behaviors to observe. This can enhance our understanding of

ILLUSTRATION 12.1 | Interrater Reliability As Ed and Charlie observe the mouse, they hope that they attain high interrater reliability. But in reality, the mouse is the one observing them.

Cross-cultural fairness Possibility of bias must be recognized and addressed

© Ed w ar d N eu kr ug

CHAPTER 12 Informal Assessment 301

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

the individual. Thus, informal assessment can be a mixed bag when it comes to cross-cultural fairness, and results need to be examined carefully.

Practicality The practical nature of informal assessments makes them particularly useful. Such assessments are generally low-cost or cost-free, can be devised so they focus on the specific issue at hand, can be created or obtained in a short amount of time, are rela- tively easy to administer, and with the exception of possible cultural bias, are fairly easy to interpret. Thus, when completing a broad assessment of an individual, informal assessments should often be considered as one additional method of assessment.

THE ROLE OF HELPERS IN THE USE OF INFORMAL ASSESSMENT

The use of informal assessment techniques is only limited by our imaginations. Proba- bly, what is most important is considering whether or not the specific informal tech- nique is adding to our knowledge about what is being assessed. Using a technique that does not add to our body of knowledge about someone is not using it wisely. For instance, if I decided to assess a client’s cognitive ability by asking that client to develop a portfolio, and if the portfolio added little because we already had garnered good information about his or her ability from traditional standardized tests, then I would just be wasting my time and the client’s time. On the other hand, there are many times when informal techniques can be used wisely and add much.

How do helpers and others use such techniques? School counselors observe chil- dren in the classroom and parents can observe their children at home. Teachers and helpers can rate students’ and clients’ progress. Clinicians can have clients assess their progress by using rating scales, and can have them write autobiographies and use journals and diaries to encourage insight. Teachers and school counselors can use anecdotal records and cumulative records in helping to understand the learning style of a student. School social workers can visit the homes of students to better understand them in that important environment, and faculty can use portfolios to help students integrate what they have learned and forgo the traditional end- of-program assessment procedures (e.g., comps). There are an endless number of ways that we can use informal techniques. Can you think of some more?

FINAL THOUGHTS ON INFORMAL ASSESSMENT Although informal assessment techniques sometimes have questionable reliability, validity, and may not always be cross-culturally fair, they can provide an addi- tional means of understanding a person. Thus, when making decisions about a per- son, these techniques can add significantly to an assessment battery if they can be shown to accurately measure an important characteristic or quality. Therefore, always consider using an informal assessment technique in combination with other assessment techniques, while keeping in mind the potential impact that such a tech- nique will have on the client.

Practicality Informal procedures are inexpensive, easy to administer, and easy to interpret

302 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

SUMMARY This chapter offered an overview of informal assessment. We noted that by their very nature, informal assessment techniques are subjective and are developed by individuals who have specific assessment needs. We pointed out that because they are mostly homegrown, the amount of time, money, and expertise put into their development is generally much less than for nationally developed instruments. Thus, reliability, validity, and cross- cultural issues are often not formally addressed, and data often are lacking. These techniques do, however, have the following advantages: (1) They offer an additional way of assessing an individual; (2) they can provide a highly specific assessment measure; (3) they can be developed in a short amount of time; (4) they are relatively nonintru- sive; (5) they are free or low-cost; and (6) they are easy to administer and interpret.

The first informal assessment technique we discussed was observation. Two types of observa- tion highlighted were event sampling, which is the viewing and assessment of a targeted behavior without regard for time, and time sampling, which takes place when a specific amount of time is set aside for the observation. Sometimes time sampling and event sampling are combined.

Next, we examined rating scales, which are used to assess a quantity of an attribute being pre- sented to the rater. We noted that such scales are subjective and can be filled with bias and errors, including the halo effect and generosity error. Some common types of rating scales included numerical scales, Likert-type (graphic) scales, semantic differential scales, and rank-order scales. We then noted that in contrast to rating scales, classification methods provide information regarding the presence or absence of an attribute or characteristic. We specifically discussed two kinds of classification methods, including behavior checklists and feeling word checklists, and also noted that many other classification methods that address additional areas can easily be created.

The next kind of informal assessment instru- ment we discussed was environmental assessment, which includes the collection of information from a client’s home, school, or workplace via obser- vation or self-reports. The four kinds of envi- ronmental assessment we discussed were direct observation, situational tests, sociometric assess- ment, and environmental assessment instruments.

In this chapter, we also discussed the use of records and personal documents. These informal assessment tools can help the examiner understand the beliefs, values, and behaviors of the person being assessed and can be obtained from numerous sources including the client, individuals close to the client, institutions, and agencies. Some of the more common records and personal documents include biographical inventories, cumulative records, anec- dotal information, autobiographies, journals and diaries, and genograms.

The last kind of informal assessment we pre- sented was performance-based assessment. We noted that performance-based assessment involves the use of a variety of informal assessment proce- dures, often based on real-world responsibilities, to make decisions about individuals. As an alternative to the use of some tests, such as standardized tests that measure cognitive skills, it is hoped that the use of performance-based assessment procedures will broaden the pool of qualified individuals seek- ing a job, seeking a credential, or being admitted to school. One kind of performance-based assess- ment, the portfolio assessment, has become popu- lar in teacher education programs and has increasingly been used in training programs in the helping professions.

We noted that informal assessment techniques tend to be less valid, less reliable, and less cross- culturally sensitive than formal assessment techni- ques, although they are quite practical. We pointed out that the validity of informal assessment techni- ques has to do with how well the examiner is defining what is being assessed. We also stressed that the bet- ter we define the behavior being assessed (the more

CHAPTER 12 Informal Assessment 303

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

valid our definition is), the more reliable will be our data collection. Ideally, informal assessment techni- ques should be administered by a minimum of two individuals, and the interrater reliability of these two individuals should be at least 0.80. However, realis- tically, we noted that this is rarely done.

Relative to cross-cultural issues, we noted that the examiner may misinterpret the verbal or non- verbal behaviors of a member of a minority group as a result of unconscious or conscious bias. Second, the observer, examiner, or rater may sim- ply be ignorant about the verbal or nonverbal

behaviors of a particular minority group. Despite potential problems with bias, informal procedures are often useful because they are uniquely geared toward the particular behavior of an individual. Finally, we pointed out that informal assessment techniques are practical in that they are generally low-cost or cost-free, can be created or developed or obtained in a short amount of time, are rela- tively easy to administer, and are fairly easy to interpret. We stressed that when used wisely, infor- mal assessment techniques provide additional mechanisms for understanding the person.

CHAPTER REVIEW 1. Define informal assessment procedures and

describe some of the ways that they can be used.

2. Describe the two types of observation and give some examples of how you might use observation in a clinical setting.

3. Describe and give examples of the following types of rating scales: numerical scales, Likert-type scales, semantic differential scales, and rank-order scales.

4. Many sources of error exist in the application of rating scales. Discuss how the halo effect and generosity error can deleteriously affect the rating of individuals.

5. Give an example of how behavior and feeling word checklists, or other types of classi- fication systems, can be used in a clinical setting.

6. Discuss how the following environmental assessment instruments may be more benefi- cial than “in office” kinds of instruments: direct observation, situational tests, socio- metric assessment, and environmental assess- ment instruments.

7. Many types of records and personal docu- ments can be useful in the assessment process.

For each of the ones listed, describe how you might use it in a clinical setting: a. biographical inventory b. cumulative records c. anecdotal information d. autobiography e. journals and diaries f. genogram

8. In what ways can performance-based assess- ment broaden the number of highly qualified applicants for a job or admission to higher education?

9. Discuss how validity is determined with regard to informal assessment procedures.

10. Discuss interrater reliability and its impor- tance to the informal assessment process.

11. Discuss the strengths and weaknesses of informal assessment procedures relative to cross-cultural issues.

12. Because informal assessment procedures can be easily and quickly developed and administered, they are practical to use. However, this very fact means that they often are not developed in a manner that yields good reliability or valid- ity. Discuss the strengths and weaknesses of this aspect of informal assessment procedures.

304 SECTION III Commonly Used Assessment Techniques

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

REFERENCES Cobia, C. D., Carney, J. S., Buckhalt, J. A., Middleton,

R. A., Shannon, D. M., Trippany, R., et al. (2005). The doctoral portfolio: Centerpiece of a comprehen- sive system of evaluation. Counselor Education and Supervision, 44, 242–254.

Dean, G. J., & Dean, S. F. (2007). Review of the Behav- ior Rating Inventory of Executive Function—adult version. In K. F. Geisinger, R. A. Spies, J. F. Carlson, & B. S. Plake (Eds.), The seventeenth mental mea- surements yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

De Jong, P., & Berg, I. K. (2013). Interviewing for solu- tions (4th ed.). Pacific Grove, CA: Brooks/Cole.

Feldman, M., Lazzara, E. H., Vanderbilt, A. A., & DiazGranados, D. (2012). Rater training to support high-stakes simulation-based assessments. Journal of Continuing Education in the Health Professions, 32, 279–286.

Ferguson, K. E., & Sgambati, R. E. (2008). Relaxation. In W. O’Donohue & J. E. Fisher (Eds.), Cognitive behavior therapy: Applying empirically supported techniques in your practice (2nd ed., pp. 434–444). Hoboken, NJ: John Wiley & Sons.

Gay, L. R., Mills, G. E., & Airasian, P. W. (2012). Educational research: Competencies for analysis and applications (10th ed.). Upper Saddle River, NJ: Pearson.

Gordon, D. (2007). Mindful dreaming: A practical guide for emotional healing through transformative mythic journeys. Franklin Lakes, NJ: New Page Books.

Gordon, D., & Meyers-Anderson, M. (1981). Phoenix: Therapeutic patterns of Milton H. Erickson. Cupertino, CA: Meta Publications.

Guerin, P., & Guerin, K. (2002). Bowenian family ther- apy. In J. Carlson & D. Kjos (Eds.), Theories and strategies of family therapy (pp. 126–157). Boston: Allyn & Bacon.

Gwet, K. L. (2010). Handbook of inter-rater reliability: The definition guide to measuring the extent of agreement among raters (3rd ed.). Gaithersburg, MD: Advanced Analytics, LLC.

Hill, C. E. (Ed.). (2004). Dream work in therapy: Facil- itating exploration, insight, and action. Washington, DC: American Psychological Association.

Manduchi, J. R. (2001). Review of the Comprehensive Assessment of School Environments Information Management System. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements year- book. Lincoln, NE: Buros Institute of Mental Mea- surements. Retrieved from Mental Measurements Yearbook database.

Neukrug, E., & Milliken, T. (2011). Counselors’ per- ceptions of ethical behaviors. Journal of Counseling and Development, 89, 206–216.

Robinson-Wood, T. (2013). The convergence of race, ethnicity, and gender: Multiple identities in counsel- ing (4th ed.). Saddle River, NJ: Merrill.

Robles, A. C. M. O. (2012). Cyber portfolio: The inno- vative menu for 21st century technology. Psychol- ogy Research, 2(3), 143–150.

Sackett, P. R., Schmitt, N., Ellingson, J. E., & Kabin, M. B. (2001). High-stakes testing in employment, credentialing, and higher education. Prospects in a post-affirmative-action world. American Psycholo- gist, 56, 302–318.

Sajatovic, M., & Ramirez, L. F. (2012). Rating scales in mental health (3rd ed.). Baltimore, MD: The Johns Hopkins Press.

Thorndike, E. L. (1920). A constant error on psycho- logical rating. Journal of Applied Psychology, 4(1), 25–29.

vander Schaaf, M. F., & Stokking, K. M. (2008). Developing and validating a design for teacher port- folio assessment. Assessment and Evaluation in Higher Education, 33, 245–262.

Watson, R. S. (2005). Review of the Emotional or Behavior Disorder Scale—revised. In R. A. Spies & B. S. Plake (Eds.), The sixteenth mental measure- ments yearbook. Lincoln, NE: Buros Institute of Mental Measurements. Retrieved from Mental Measurements Yearbook database.

Yalom, I. (2002). The gift of therapy: An open letter to a new generation of therapists and their patients. New York: HarperCollins.

CHAPTER 12 Informal Assessment 305

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

AA P P E N D I X Websites of Codes of Ethicsof Select Mental Health Professional Associations

American Counseling Association (ACA)

Main website: http://www.counseling.org Code of ethics: http://www.counseling.org/Resources/aca-code-of-ethics.pdf

American Association of Marriage and Family Therapy (AAMFT)

Main website: http://www.aamft.org Code of ethics: http://www.aamft.org/imis15/Content/Legal_Ethics/Code_of_Ethics.aspx

American Association of Pastoral Counselors (AAPC)

Main website: http://www.aapc.org Code of ethics: http://www.aapc.org/policies/code-of-ethics.aspx

American Mental Health Counselors Association (AMHCA)

Main website: http://www.amhca.org Code of ethics: http://www.amhca.org/about/codetoc.aspx

American Psychiatric Association (APA)

Main website: http://www.psychiatry.org/ Code of ethics: http://www.psychiatry.org/practice/ethics/resources-standards

306

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

American Psychological Association (APA)

Main website: http://www.apa.org Code of ethics: http://www.apa.org/ethics/code/index.aspx

American Psychological Association: Division 5: Evaluation, Measurement, and Statistics

Main website: http://www.apa.org/about/division/div5.html Code of ethics: http://www.apa.org/ethics/code/index.aspx

American School Counselor Association (ASCA)

Main website: http://www.schoolcounselor.org Code of ethics: http://www.schoolcounselor.org/files/EthicalStandards2010.pdf

Association for Assessment and Research in Counseling (AARC)

Main website: http://aarc-counseling.org/ Code of ethics: Uses ACA’s (see previous page)

Certified Rehabilitation Counselors (CRC)

Main website: http://www.crccertification.com Code of ethics: http://www.crccertification.com/filebin/pdf/CRCCodeOfEthics.pdf

National Association of Social Workers (NASW)

Main website: http://www.naswdc.org Code of ethics: https://www.socialworkers.org/pubs/code/default.asp

National Organization of Human Services (NOHS)

Main website: http://www.nationalhumanservices.org/ Code of ethics: http://www.nationalhumanservices.org/ethical-standards-for-hs-professionals

APPENDIX A 307

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

BA P P E N D I X Assessment Sectionsof ACA’s and APA’s Codes of Ethics

American Counseling Association’s Ethical Code: Section E

Section E: Evaluation, Assessment, and Interpretation

Introduction Counselors use assessment instruments as one component of the counseling process, taking into account the client personal and cultural context. Counselors promote the well-being of individual cli- ents or groups of clients by developing and using appropriate educational, psychological, and career assessment instruments.

E.1. General E.1.a. Assessment The primary purpose of educational, psychological, and career assessment is to provide measurements that are valid and reliable in either comparative or absolute terms. These include, but are not limited to, measurements of ability, personality, interest,

intelligence, achievement, and performance. Coun- selors recognize the need to interpret the statements in this section as applying to both quantitative and qualitative assessments.

E.1.b. Client Welfare Counselors do not misuse assessment results and interpretations, and they take reasonable steps to pre- vent others from misusing the information these tech- niques provide. They respect the client’s right to know the results, the interpretations made, and the bases for counselors’ conclusions and recommendations.

E.2. Competence to Use and Interpret Assessment Instruments E.2.a. Limits of Competence Counselors utilize only those testing and assessment services for which they have been trained and are competent. Counselors using technology-assisted

308

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

test interpretations are trained in the construct being measured and the specific instrument being used prior to using its technology-based application. Coun- selors take reasonable measures to ensure the proper use of psychological and career assessment techni- ques by persons under their supervision. (See A.12.)

E.2.b. Appropriate Use Counselors are responsible for the appropriate appli- cation, scoring, interpretation, and use of assessment instruments relevant to the needs of the client, whether they score and interpret such assessments themselves or use technology or other services.

E.2.c. Decisions Based on Results Counselors responsible for decisions involving indivi- duals or policies that are based on assessment results have a thorough understanding of educational, psychological, and career measurement, including validation criteria, assessment research, and guide- lines for assessment development and use.

E.3. Informed Consent in Assessment E.3.a. Explanation to Clients Prior to assessment, counselors explain the nature and purposes of assessment and the specific use of results by potential recipients. The explanation will be given in the language of the client (or other legally authorized person on behalf of the client), unless an explicit exception has been agreed upon in advance. Counselors consider the client’s per- sonal or cultural context, the level of the client’s understanding of the results, and the impact of the results on the client. (See A.2., A.12.g.,F.l.c.)

E.3.b. Recipients of Results Counselors consider the examinee’s welfare, explicit understandings, and prior agreements in determining who receives the assessment results. Counselors include accurate and appropriate interpretations with any release of individual or group assessment results. (See B.2.C., B.5.)

E.4. Release of Data to Qualified Professionals Counselors release assessment data in which the client is identified only with the consent of the client

or the client’s legal representative. Such data are released only to persons recognized by counselors as qualified to interpret the data. (See B.I., B.3., B.6.b.)

E.5. Diagnosis of Mental Disorders E.5.a. Proper Diagnosis Counselors take special care to provide proper diag- nosis of mental disorders. Assessment techniques (including personal interview) used to determine cli- ent care (e.g., locus of treatment, type of treatment, or recommended follow-up) are carefully selected and appropriately used.

E.5.b. Cultural Sensitivity Counselors recognize that culture affects the man- ner in which clients’ problems are defined. Clients’ socioeconomic and cultural experiences are considered when diagnosing mental disorders. (See A.2.c.)

E.5.c. Historical and Social Prejudices in the Diagnosis of Pathology Counselors recognize historical and social prejudices in the misdiagnosis and pathologizing of certain indi- viduals and groups and the role of mental health pro- fessionals in perpetuating these prejudices through diagnosis and treatment.

E.5.d. Refraining From Diagnosis Counselors may refrain from making and/or report- ing a diagnosis if they believe it would cause harm to the client or others.

E.6. Instrument Selection E.6.a. Appropriateness of Instruments Counselors carefully consider the validity, reliability, psychometric limitations, and appropriateness of instruments when selecting assessments.

E.6.b. Referral Information If a client is referred to a third party for assessment, the counselor provides specific referral questions and sufficient objective data about the client to ensure that appropriate assessment instruments are utilized. (See A.9.b., B.3.)

APPENDIX B 309

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

E.6.c. Culturally Diverse Populations Counselors are cautious when selecting assess- ments for culturally diverse populations to avoid the use of instruments that lack appropriate psychomet- ric properties for the client population. (See A.2.c, E.5.b.)

E.7. Conditions of Assessment Administration (See A.12.b, A.12.d.)

E.7.a. Administration Conditions Counselors administer assessments under the same conditions that were established in their standardiza- tion. When assessments are not administered under standard conditions, as may be necessary to accommodate clients with disabilities, or when unusual behavior or irregularities occur during the administration, those conditions are noted in inter- pretation, and the results may be designated as invalid or of questionable validity.

E.7.b. Technological Administration Counselors ensure that administration programs function properly and provide clients with accurate results when technological or other electronic meth- ods are used for assessment administration.

E.7.c. Unsupervised Assessments Unless the assessment instrument is designed, intended, and validated for self-administration and/ or scoring, counselors do not permit inadequately supervised use.

E.7.d. Disclosure of Favorable Conditions Prior to administration of assessments, conditions that produce most favorable assessment results are made known to the examinee.

E.8. Multicultural Issues/Diversity in Assessment Counselors use with caution assessment techniques that were normed on populations other than that of the client. Counselors recognize the effects of age, color, culture, disability, ethnic group, gender, race, language preference, religion, spirituality, sexual orientation, and socioeconomic status on test

administration and interpretation, and place test results in proper perspective with other relevant factors. (See A.2.c., E.5.b.)

E.9. Scoring and Interpretation of Assessments E.9.a. Reporting In reporting assessment results, counselors indi- cate reservations that exist regarding validity or reliability due to circumstances of the assessment or the inappropriateness of the norms for the person tested.

E.9.b. Research Instruments Counselors exercise caution when interpreting the results of research instruments not having sufficient technical data to support respondent results. The specific purposes for the use of such instruments are stated explicitly to the examinee.

E.9.c. Assessment Services Counselors who provide assessment scoring and interpretation services to support the assessment process confirm the validity of such interpretations. They accurately describe the purpose, norms, valid- ity, reliability, and applications of the procedures and any special qualifications applicable to their use. The public offering of an automated test interpretations service is considered a professional-to-professional consultation. The formal responsibility of the con- sultant is to the consultee, but the ultimate and overriding responsibility is to the client. (See D.2.)

E.10. Assessment Security Counselors maintain the integrity and security of tests and other assessment techniques consistent with legal and contractual obligations. Counselors do not appropriate, reproduce, or modify published assessments or parts thereof without acknowledg- ment and permission from the publisher.

E.11. Obsolete Assessments and Outdated Results Counselors do not use data or results from assess- ments that are obsolete or outdated for the current purpose. Counselors make every effort to prevent

310 APPENDIX B

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

the misuse of obsolete measures and assessment data by others.

E.12. Assessment Construction Counselors use established scientific procedures, relevant standards, and current professional knowl- edge for assessment design in the development, publication, and utilization of educational and psy- chological assessment techniques.

E.13. Forensic Evaluation: Evaluation for Legal Proceedings E.13.a. Primary Obligations When providing forensic evaluations, the primary obli- gation of counselors is to produce objective findings that can be substantiated based on information and techniques appropriate to the evaluation, which may include examination of the individual and/or review of records. Counselors are entitled to form professional opinions based on their professional knowledge and expertise that can be supported by the data gathered in evaluations. Counselors will define the limits of their reports or testimony, especially when an examination of the individual has not been conducted.

E.13.b. Consent for Evaluation Individuals being evaluated are informed in writing that the relationship is for the purposes of an evaluation

and is not counseling in nature, and entities or indivi- duals who will receive the evaluation report are identi- fied. Written consent to be evaluated is obtained from those being evaluated unless a court orders evalua- tions to be conducted without the written consent of individuals being evaluated. When children or vulnera- ble adults are being evaluated, informed written con- sent is obtained from a parent or guardian.

E.13.c. Client Evaluation Prohibited Counselors do not evaluate individuals for forensic purposes they currently counsel or individuals they have counseled in the past. Counselors do not accept as counseling clients individuals they are evaluating or individuals they have evaluated in the past for forensic purposes.

E.13.d. Avoid Potentially Harmful Relationships Counselors who provide forensic evaluations avoid potentially harmful professional or personal relation- ships with family members, romantic partners, and close friends of individuals they are evaluating or have evaluated in the past.

Source: Copyright © ACA 2005. Reprinted by permission. No further reproduction authorized without written permis- sion of the American Counseling Association.

American Psychological Association Ethical Code: Section 9

Section 9: Assessment

9.01 Bases For Assessments

(a) Psychologists base the opinions contained in their recommendations, reports, and diagnostic or evaluative statements, including forensic testimony, on information and techniques sufficient to substan- tiate their findings. (See also Standard 2.04, Bases for Scientific and Professional Judgments.) (b) Except as noted in 9.01c, psychologists pro- vide opinions of the psychological characteristics of

individuals only after they have conducted an examination of the individuals adequate to support their statements or conclusions. When, despite rea- sonable efforts, such an examination is not practi- cal, psychologists document the efforts they made and the result of those efforts, clarify the probable impact of their limited information on the reliability and validity of their opinions, and appropriately limit the nature and extent of their conclusions or recommendations. (See also Standards 2.01,

APPENDIX B 311

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Boundaries of Competence, and 9.06, Interpreting Assessment Results.) (c) When psychologists conduct a record review or provide consultation or supervision and an individual examination is not warranted or necessary for the opinion, psychologists explain this and the sources of information on which they based their conclusions and recommendations.

9.02 Use of Assessments

(a) Psychologists administer, adapt, score, inter- pret, or use assessment techniques, interviews, tests, or instruments in a manner and for purposes that are appropriate in light of the research on or evidence of the usefulness and proper application of the techniques. (b) Psychologists use assessment instruments whose validity and reliability have been established for use with members of the population tested. When such validity or reliability has not been estab- lished, psychologists describe the strengths and lim- itations of test results and interpretation. (c) Psychologists use assessment methods that are appropriate to an individual’s language prefer- ence and competence, unless the use of an alterna- tive language is relevant to the assessment issues.

9.03 Informed Consent in Assessments

(a) Psychologists obtain informed consent for assessments, evaluations, or diagnostic services, as described in Standard 3.10, Informed Consent, except when (1) testing is mandated by law or gov- ernmental regulations; (2) informed consent is implied because testing is conducted as a routine educational, institutional, or organizational activity (e.g., when participants voluntarily agree to assess- ment when applying for a job); or (3) one purpose of the testing is to evaluate decisional capacity. Informed consent includes an explanation of the nature and purpose of the assessment, fees, involvement of third parties, and limits of confidenti- ality and sufficient opportunity for the client/patient to ask questions and receive answers. (b) Psychologists inform persons with questionable capacity to consent or for whom testing is mandated

by law or governmental regulations about the nature and purpose of the proposed assessment services, using language that is reasonably understandable to the person being assessed. (c) Psychologists using the services of an inter- preter obtain informed consent from the client/ patient to use that interpreter, ensure that confidenti- ality of test results and test security are maintained, and include in their recommendations, reports, and diagnostic or evaluative statements, including foren- sic testimony, discussion of any limitations on the data obtained. (See also Standards 2.05, Delegation of Work to Others; 4.01, Maintaining Confidentiality; 9.01, Bases for Assessments; 9.06, Interpreting Assessment Results; and 9.07, Assessment by Unqualified Persons.)

9.04 Release of Test Data

(a) The term test data refers to raw and scaled scores, client/patient responses to test questions or stimuli, and psychologists’ notes and recordings concerning client/patient statements and behavior during an examination. Those portions of test mate- rials that include client/patient responses are included in the definition of test data. Pursuant to a client/patient release, psychologists provide test data to the client/ patient or other persons identified in the release. Psychologists may refrain from releas- ing test data to protect a client/patient or others from substantial harm or misuse or misrepresenta- tion of the data or the test, recognizing that in many instances release of confidential information under these circumstances is regulated by law. (See also Standard 9.11, Maintaining Test Security.) (b) In the absence of a client/patient release, psy- chologists provide test data only as required by law or court order.

9.05 Test Construction

Psychologists who develop tests and other assess- ment techniques use appropriate psychometric pro- cedures and current scientific or professional knowledge for test design, standardization, valida- tion, reduction or elimination of bias, and recommen- dations for use.

312 APPENDIX B

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

9.06 Interpreting Assessment Results

When interpreting assessment results, including auto- mated interpretations, psychologists take into account the purpose of the assessment as well as the various test factors, test-taking abilities, and other characteristics of the person being assessed, such as situational, personal, linguistic, and cultural differences, that might affect psychologists’ judg- ments or reduce the accuracy of their interpretations. They indicate any significant limitations of their inter- pretations. (See also Standards 2.01b and c, Bound- aries of Competence, and 3.01, Unfair Discrimination.)

9.07 Assessment by Unqualified Persons

Psychologists do not promote the use of psycholog- ical assessment techniques by unqualified persons, except when such use is conducted for training purposes with appropriate supervision. (See also Standard 2.05, Delegation of Work to Others.)

9.08 Obsolete Tests and Outdated Test Results

(a) Psychologists do not base their assessment or intervention decisions or recommendations on data or test results that are outdated for the current purpose. (b) Psychologists do not base such decisions or recommendations on tests and measures that are obsolete and not useful for the current purpose.

9.09 Test Scoring and Interpretation Services

(a) Psychologists who offer assessment or scoring services to other professionals accurately describe the purpose, norms, validity, reliability, and applica- tions of the procedures and any special qualifica- tions applicable to their use. (b) Psychologists select scoring and interpretation services (including automated services) on the basis

of evidence of the validity of the program and pro- cedures as well as on other appropriate considera- tions. (See also Standard 2.01b and c, Boundaries of Competence.) (c) Psychologists retain responsibility for the appro- priate application, interpretation, and use of assess- ment instruments, whether they score and interpret such tests themselves or use automated or other services.

9.10 Explaining Assessment Results

Regardless of whether the scoring and interpretation are done by psychologists, by employees or assis- tants, or by automated or other outside services, psychologists take reasonable steps to ensure that explanations of results are given to the individual or designated representative unless the nature of the relationship precludes provision of an explanation of results (such as in some organizational consulting, preemployment or security screenings, and forensic evaluations), and this fact has been clearly explained to the person being assessed in advance.

9.11. Maintaining Test Security

The term test materials refers to manuals, instru- ments, protocols, and test questions or stimuli and does not include test data as defined in Standard 9.04, Release of Test Data. Psychologists make rea- sonable efforts to maintain the integrity and security of test materials and other assessment techniques consistent with law and contractual obligations, and in a manner that permits adherence to this Ethics Code.

Source: From American Psychologist 57, pp. 1060–1073. Copyright © 2002/2010 by the American Psychological Association. Reprinted with permission.

APPENDIX B 313

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

CA P P E N D I X Code of Fair TestingPractices in Education

PREPARED BY THE JOINT COMMITTEE ON TESTING PRACTICES

The Code of Fair Testing Practices in Education (Code) is a guide for professionals in fulfilling their obligation to provide and use tests that are fair to all test takers regardless of age, gender, disability, race, ethnicity, national origin, religion, sexual orientation, linguistic background, or other personal characteristics. Fairness is a primary consideration in all aspects of testing. Careful standardization of tests and administration conditions helps to ensure that all test takers are given a compara- ble opportunity to demonstrate what they know and how they can perform in the area being tested. Fairness implies that every test taker has the opportunity to pre- pare for the test and is informed about the general nature and content of the test, as appropriate to the purpose of the test. Fairness also extends to the accurate reporting of individual and group test results. Fairness is not an isolated concept, but must be considered in all aspects of the testing process.

The Code applies broadly to testing in education (admissions, educational assessment, educational diagnosis, and student placement) regardless of the mode of presentation, so it is relevant to conventional paper-and-pencil tests, computer- based tests, and performance tests. It is not designed to cover employment testing, licensure or certification testing, or other types of testing outside the field of educa- tion. The Code is directed primarily at professionally developed tests used in for- mally administered testing programs. Although the Code is not intended to cover

314

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

tests made by teachers for use in their own classrooms, teachers are encouraged to use the guidelines to help improve their testing practices.

The Code addresses the roles of test developers and test users separately. Test developers are people and organizations that construct tests, as well as those that set policies for testing programs. Test users are people and agencies that select tests, administer tests, commission test development services, or make decisions on the basis of test scores. Test developer and test user roles may overlap, for exam- ple, when a state or local education agency commissions test development services, sets policies that control the test development process, and makes decisions on the basis of the test scores.

Many of the statements in the Code refer to the selection and use of existing tests. When a new test is developed, when an existing test is modified, or when the administration of a test is modified, the Code is intended to provide guidance for this process.1

The Code provides guidance separately for test developers and test users in four critical areas:

1. Developing and Selecting Appropriate Tests 2. Administering and Scoring Tests 3. Reporting and Interpreting Test Results 4. Informing Test Takers

The Code is intended to be consistent with the relevant parts of the Standards for Educational and Psychological Testing (American Educational Research Associa- tion [AERA], American Psychological Association [APA], and National Council on Measurement in Education [NCME], 1999). The Code is not meant to add new principles over and above those in the Standards or to change their meaning. Rather, the Code is intended to represent the spirit of selected portions of the Stan- dards in a way that is relevant and meaningful to developers and users of tests, as well as to test takers and/or their parents or guardians. States, districts, schools, organizations, and individual professionals are encouraged to commit themselves to fairness in testing and safeguarding the rights of test takers. The Code is intended to assist in carrying out such commitments.

The Code has been prepared by the Joint Committee on Testing Practices, a cooperative effort among several professional organizations. The aim of the Joint

1 The Code is not intended to be mandatory, exhaustive, or definitive, and may not be applicable to every situation. Instead, the Code is intended to be aspirational, and is not intended to take precedence over the judgment of those who have competence in the subjects addressed.

Index terms: assessment, education.

Reference this material as: Code of Fair Testing Practices in Education. (2004). Washington, DC: Joint Committee on Testing Practices.

Copyright 2004 by the Joint Committee on Testing Practices. (Mailing Address: Joint Committee on Testing Practices, Science Directorate, American Psychological Association, 750 First Street, NE, Washington, DC 20002-4242; http://www.apa.org/science/jctpweb.html.) Contact APA for additional copies. This material may be reproduced in whole or in part without fees or permission, provided that acknowledgment is made to the Joint Committee on Testing Practices. Reproduction and dissemination of this document are encouraged.

APPENDIX C 315

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Committee is to act, in the public interest, to advance the quality of testing practices. Members of the Joint Committee include the American Counseling Association (ACA), the American Educational Research Association (AERA), the American Psychological Association (APA), the American Speech-Language- Hearing Association (ASHA), the National Association of School Psychologists (NASP), the National Association of Test Directors (NATD), and the National Council on Measurement in Education (NCME).

A. DEVELOPING AND SELECTING APPROPRIATE TESTS

Test Developers Test Users

Test developers should provide the information and supporting evidence that test users need to select appropriate tests.

Test users should select tests that meet the intended purpose and that are appropriate for the intended test takers.

A-l. Provide evidence of what the test measures, the recommended uses, the intended test takers, and the strengths and limitations of the test, including the level of precision of the test scores.

A-l. Define the purpose for testing, the content and skills to be tested, and the intended test takers. Select and use the most appropriate test based on a thorough review of available information.

A-2. Describe how the content and skills to be tested were selected and how the tests were developed.

A-2. Review and select tests based on the appropriateness of test content, skills tested, and content coverage for the intended purpose of testing.

A-3. Communicate information about a test’s characteristics at a level of detail appropriate to the intended test users.

A-3. Review materials provided by test developers and select tests for which clear, accurate, and complete information is provided.

A-4. Provide guidance on the levels of skills, knowledge, and training necessary for appropriate review, selection, and administration of tests.

A-4. Select tests through a process that includes persons with appropriate knowledge, skills, and training.

A-5. Provide evidence that the technical quality, including reliability and validity, of the test meets its intended purposes.

A-5. Evaluate evidence of the technical quality of the test provided by the test developer and any independent reviewers.

A-6. Provide to qualified test users representative samples of test questions or practice tests, directions, answer sheets, manuals, and score reports.

A-6. Evaluate representative samples of test questions or practice tests, directions, answer sheets, manuals, and score reports before selecting a test.

A-7. Avoid potentially offensive content or language when developing test questions and related materials.

A-7. Evaluate procedures and materials used by test developers, as well as the resulting test, to ensure that potentially offensive content or language is avoided.

A-8. Make appropriately modified forms of tests or administration procedures available for test takers with disabilities who need special accommodations.

A-8. Select tests with appropriately modified forms or administration procedures for test takers with disabilities who need special accommodations.

A-9. Obtain and provide evidence on the performance of test takers of diverse subgroups, making significant efforts to obtain sample sizes that are adequate for subgroup analyses. Evaluate the evidence to ensure that differences in performance are related to the skills being assessed.

A-9. Evaluate the available evidence on the performance of test takers of diverse subgroups. Determine to the extent feasible which performance differences may have been caused by factors unrelated to the skills being assessed.

316 APPENDIX C

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

B. ADMINISTERING AND SCORING TESTS

Test Developers Test Users

Test developers should explain how to administer and score tests correctly and fairly.

Test users should administer and score tests correctly and fairly.

B-l. Provide clear descriptions of detailed procedures for administering tests in a standardized manner.

B-l. Follow established procedures for administering tests in a standardized manner.

B-2. Provide guidelines on reasonable procedures for assessing persons with disabilities who need special accommodations or those with diverse linguistic backgrounds.

B-2. Provide and document appropriate procedures for test takers with disabilities who need special accommodations or those with diverse linguistic backgrounds. Some accommodations may be required by law or regulation.

B-3. Provide information to test takers or test users on test question formats and procedures for answering test questions, including information on the use of any needed materials and equipment.

B-3. Provide test takers with an opportunity to become familiar with test question formats and any materials or equipment that may be used during testing.

B-4. Establish and implement procedures to ensure the security of testing materials during all phases of test development, administration, scoring, and reporting.

B-4. Protect the security of test materials, including respecting copyrights and eliminating opportunities for test takers to obtain scores by fraudulent means.

B-5. Provide procedures, materials and guidelines for scoring the tests, and for monitoring the accuracy of the scoring process. If scoring the test is the responsibility of the test developer, provide adequate training for scorers.

B-5. If test scoring is the responsibility of the test user, provide adequate training to scorers and ensure and monitor the accuracy of the scoring process.

B-6. Correct errors that affect the interpretation of the scores and communicate the corrected results promptly.

B-6. Correct errors that affect the interpretation of the scores and communicate the corrected results promptly.

B-7. Develop and implement procedures for ensuring the confidentiality of scores.

B-7. Develop and implement procedures for ensuring the confidentiality of scores.

APPENDIX C 317

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

C. REPORTING AND INTERPRETING TEST RESULTS

Test Developers Test Users

Test developers should report test results accurately and provide information to help test users interpret test results correctly.

Test users should report and interpret test results accurately and clearly.

C-l. Provide information to support recommended interpretations of the results, including the nature of the content, norms or comparison groups, and other technical evidence. Advise test users of the benefits and limitations of test results and their interpretation. Warn against assigning greater precision than is warranted.

C-l. Interpret the meaning of the test results, taking into account the nature of the content, norms or comparison groups, other technical evidence, and benefits and limitations of test results.

C-2. Provide guidance regarding the interpretations of results for tests administered with modifications. Inform test users of potential problems in interpreting test results when tests or test administration procedures are modified.

C-2. Interpret test results from modified test or test administration procedures in view of the impact those modifications may have had on test results.

C-3. Specify appropriate uses of test results and warn test users of potential misuses.

C-3. Avoid using tests for purposes other than those recommended by the test developer unless there is evidence to support the intended use or interpretation.

C-4. When test developers set standards, provide the rationale, procedures, and evidence for setting performance standards or passing scores. Avoid using stigmatizing labels.

C-4. Review the procedures for setting performance standards or passing scores. Avoid using stigmatizing labels.

C-5. Encourage test users to base decisions about test takers on multiple sources of appropriate information, not on a single test score.

C-5. Avoid using a single test score as the sole determinant of decisions about test takers. Interpret test scores in conjunction with other information about individuals.

C-6. Provide information to enable test users to accurately interpret and report test results for groups of test takers, including information about who were and who were not included in the different groups being compared, and information about factors that might influence the interpretation of results.

C-6. State the intended interpretation and use of test results for groups of test takers. Avoid grouping test results for purposes not specifically recommended by the test developer unless evidence is obtained to support the intended use. Report procedures that were followed in determining who were and who were not included in the groups being compared and describe factors that might influence the interpretation of results.

C-7. Provide test results in a timely fashion and in a manner that is understood by the test taker.

C-7. Communicate test results in a timely fashion and in a manner that is understood by the test taker.

C-8. Provide guidance to test users about how to monitor the extent to which the test is fulfilling its intended purposes.

C-8. Develop and implement procedures for monitoring test use, including consistency with the intended purposes of the text.

318 APPENDIX C

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

D. INFORMING TEST TAKERS Under some circumstances, test developers have direct communication with the test takers and/or control of the tests, testing process, and test results. In other circum- stances the test users have these responsibilities.

Test developers or test users should inform test takers about the nature of the test, test taker rights and responsibilities, the appropriate use of scores, and procedures for resolving challenges to scores.

D-l. Inform test takers in advance of the test administration about the coverage of the test, the types of question formats, the directions, and appropriate test-taking strategies. Make such information available to all test takers.

D-2. When a test is optional, provide test takers or their parents/guardians with information to help them judge whether a test should be taken—including indications of any consequences that may result from not taking the test (e.g., not being eligible to compete for a particular scholarship)—and whether there is an available alternative to the test.

D-3. Provide test takers or their parents/guardians with information about rights test takers may have to obtain copies of tests and completed answer sheets, to retake tests, to have tests rescored, or to have scores declared invalid.

D-4. Provide test takers or their parents/guardians with information about responsibilities test takers have, such as being aware of the intended purpose and uses of the test, performing at capacity, following directions, and not disclosing test items or interfering with other test takers.

D-5. Inform test takers or their parents/guardians how long scores will be kept on file and indicate to whom, under what circumstances, and in what manner test scores and related information will or will not be released. Protect test scores from unauthorized release and access.

D-6. Describe procedures for investigating and resolving circumstances that might result in canceling or withholding scores, such as failure to adhere to specified testing procedures.

D-7. Describe procedures that test takers, parents/guardians, and other interested parties may use to obtain more information about the test, register complaints, and have problems resolved.

APPENDIX C 319

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Note: The membership of the Working Group that developed the Code of Fair Testing Practices in Education and of the Joint Committee on Testing Practices that guided the Working Group is as follows:

Peter Behuniak, PhD Stephanie H. McConaughy, PhD Lloyd Bond, PhD Julie P. Noble, PhD Gwyneth M. Boodoo, PhD Wayne M. Patience, PhD Wayne Camara, PhD Carole L. Perlman, PhD Ray Fenton, PhD Douglas K. Smith, PhD (deceased) John J. Fremer, PhD (Co-Chair) Janet E. Wall, EdD (Co-Chair) Sharon M. Goldsmith, PhD Pat Nellor Wickwire, PhD Bert F. Green, PhD Mary Yakimowski, PhD William G. Harris, PhD Janet E. Helms, PhD Lara Frumkin, PhD, of the

APA served as staff liaison.

The Joint Committee intends that the Code be consistent with and supportive of existing codes of conduct and standards of other professional groups who use tests in educational contexts. Of particular note are the Responsibilities of Users of Standardized Tests (Association for Assessment in Counseling and Education, 2003), APA Test User Qualifications (2000), ASHA Code of Ethics (2001), Ethical Principles of Psychologists and Code of Conduct (1992), NASP Professional Con- duct Manual (2000), NCME Code of Professional Responsibility (1995), and Rights and Responsibilities of Test Takers: Guidelines and Expectations (Joint Committee on Testing Practices, 2000).

320 APPENDIX C

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

DA P P E N D I XSample AssessmentReport*

Demographic Information Name: Eduardo (Ed) Unclear DOB: 01/08/1966

Address: 223 Confused Lane Age: 48 Coconut Creek, Florida Sex: Male

Phone: 954-969-5555 Ethnicity: Hispanic (Cuban-American)

E-mail: [email protected] Date of Interview: 10/22/2014

Name of Interviewer: Sigmund Freud, MD

Presenting Problem or Reason for Referral Eduardo Unclear is a 48-year-old Hispanic male of average stature and build. He was self-referred to counseling due to stress and inability to sleep. The client reported feeling anxious for approximately two years and intermittently depressed for approximately seven or eight years. He states that he feels discontent with his marriage and confused about his future. Mr. Unclear appeared appropriately dressed and was attentive during the session. An assessment was conducted to determine differential diagnosis and the course of treatment.

*Updated and revised by Mandi Hughes.

321

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Family Background Mr. Unclear was raised in Miami, Florida. When he was five years old, his parents fled from Cuba on a fishing boat with him and his two brothers, José, who is two years older, and Juan, who is two years younger. Mr. Unclear comes from an intact family. He reports that his father was a bookkeeper and his mother was a stay-at-home mom. He states that his parents were “loving but strict” and notes that his father was “in charge” of the family and would often “take a belt to me.” He reports that he and his brothers were always close and that both brothers currently live within 1 mile of his home. He states that his younger brother is mar- ried and has two children. He describes his other brother as single and “gay but not out.” He and his brothers went to Catholic school, and he states that he was a good student and had the “normal” number of friends. His father died approxi- mately four years ago of a “heart disorder.” His mother currently resides in a retirement community in North Miami Beach.

Mr. Unclear notes that he met his wife Carla in college when he was 20. They married when he was 21 and quickly had two children, Carlita and Carmen, who are now 27 and 26. Both daughters are college-educated, have professional jobs, and are married. Carlita has two children aged 3 and 4, while Carmen has one child aged 5. He notes that both daughters and their families live close to him, and he maintains positive relationships with them. He states that although his mar- riage was “good” for the first 20 years, in recent years he has found himself feeling unloved and depressed. He wonders if he should remain in the marriage.

Significant Medical/Counseling History Mr. Unclear reports that approximately four years ago he was in a serious car accident that subsequently left him with chronic back pain. Although he is pre- scribed medication for the pain (Flexeril, 5 mg), he prefers not to take it, stating that he mostly tries to “live without drugs.” He notes that he often feels fatigued and has trouble sleeping, usually sleeping around four hours a night. He reports that a recent medical exam revealed no apparent medical reason for his fatigue and sleep difficulties. He notes that in the past two years, he has had obsessive worry related to fears of dying of a heart attack. He describes his eating habits as “normal” and reports no other significant medical history.

Mr. Unclear explained that after the birth of his second child, his wife required surgery to repair vaginal tears. He states that since that time she has experienced pain during intercourse and their level of intimacy has significantly decreased. He notes that he and his wife attended couples counseling for about two months approximately 15 years ago. He feels that counseling did not help, and he reports that it “particularly did nothing to help our sex life.”

Substance Use and Abuse Mr. Unclear states that he does not smoke cigarettes but does occasionally smoke cigars, adding that he “will never smoke a Cuban cigar.” He describes himself as a moderate alcohol user, stating that he has a “couple of beers a day” but rarely

322 APPENDIX D

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

drinks “hard liquor.” He reports taking prescription medication intermittently for chronic back pain, and he denies the use of illegal substances.

Educational and Vocational History Mr. Unclear attended Catholic school in Miami, Florida. He reports that he excelled in math but had difficulty with reading and spelling. After high school, he attended college at the University of Miami, where he majored in business administration. After graduating with his bachelor’s degree, he obtained a job as an accountant at a major tobacco import company, where he worked for 17 years. During that time, he began to work on his master’s in business administra- tion but stated he never finished his degree because it was “boring.” Approxi- mately eight years ago, he changed jobs to “make more money.” He obtained employment as an accountant at a local new car company.

Mr. Unclear states that as an accountant, his “books were always perfect,” although he went on to note that he was embarrassed by his inability to prepare a well-written report. He expresses dissatisfaction with his career path and wants to “do something more meaningful with his life.” He adds, however, that “I am probably too old to change careers now.”

Other Pertinent Information Mr. Unclear states that he is unhappy with his sex life and reports limited intimacy with his wife. He denies an extramarital affair but states “I would have one if I met the right person.” He notes that he is “just making it” financially and that it was difficult to support his two children through college. He denies any problems with the law.

Mental Status Eduardo Unclear appeared for his appointment casually but neatly dressed and groomed. He was able to maintain appropriate eye contact and was oriented to time, place, and person. Visual acuity appeared to be within normal limits; audi- tion and speech were unremarkable. During the interview, he appeared anxious, often rubbing his hands together. Mr. Unclear was cooperative with the examiner and demonstrated satisfactory levels of motivation, interest, and energy. He is cur- rently prescribed pain medication, which he only takes occasionally for chronic back pain.

He stated that he often feels fatigued because he usually sleeps approximately four hours a night. He described himself as feeling intermittently depressed over the past seven or eight years. He appeared to be of above average intelligence and his memory was intact. His judgment seemed fair and his insight fair to good. He stated that he has some suicidal ideation but denies he has a plan or would kill himself, noting that it is “against my religion.” He denied homicidal ideation.

APPENDIX D 323

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Assessment Results Mr. Unclear was administered a battery of objective and projective personality tests, including the Beck Depression Inventory-II (the BDI-II), the Minnesota Multi- phasic Personality Inventory-II (the MMPI-II), the Rorschach Inkblot Test (the Rorschach), the Thematic Apperception Test (TAT), the Kinetic Family Drawing (KFD), the Sentence Completion Test, the Strong Interest Inventory (the Strong), and the Wide Range Achievement Test-4 (the WRAT).

Through self-report of the past two weeks, Mr. Unclear’s score on the BDI-II indicates that he has moderate depression (raw score ¼ 24). His responses showed some evidence of possible suicidal ideation. The BDI-II is not only able to diagnose depression due to its consistency with the DSM diagnostic criteria, but it can also determine the severity of depressive symptoms. The MMPI-II supports this finding of moderate to severe depression and also indicates some mild anxiety. The MMPI-II, which reveals dissatisfaction with one’s life, demonstrates that Mr. Unclear is generally “discontent with the world” and feels a lack of intimacy in his life. It suggests assessing for possible suicidal ideation.

The Rorschach and the TAT are projective assessment tools used to evaluate psychological functioning. Both demonstrated that Mr. Unclear is grounded in reality and open to testing, as evidenced by his willingness to readily respond to initial inkblots and TAT cards, his ability to complete stories in the TAT, and the fact that many of his responses were “common” responses. Feelings of depression and hopelessness are evident in a number of responses, such as not readily seeing color in many of the responses to the “color cards” of the Rorschach and making a number of pessimistic stories that generally had depressive endings to them on the TAT.

When he was administered the KFD, a projective test that asks the client to draw his family all doing something together, Mr. Unclear placed his father as an angel in the sky and included his wife, mother, children, and grandchildren. His mother was standing next to him while his wife was off to the side with the grand- children. He also placed himself in a chair and when describing the picture stated, “I’m sitting because my back hurts.” The picture showed the client and his family at his mother’s house having a Sunday dinner while it rained outside. Rain could be indicative of depressive feelings. A cross was prominent in the background and was larger than most of the people in the picture, which is likely an indication of strong religious beliefs and could also indicate a need to be taken care of.

On the Sentence Completion Test, Mr. Unclear made a number of references to missing his father, such as “The thing I think most about is missing my father.” He also referenced continual back pain. Finally, he noted discontent with his marriage, including the statement, “Sex is nonexistent.”

On the Strong, a self-report assessment tool used to evaluate both personality and career interest, Mr. Unclear’s two highest personality codes were conventional and enterprising, respectively. All the other codes were significantly lower. Indivi- duals of the conventional type are stable, controlled, conservative, sociable, and like to follow instructions. Enterprising individuals are self-confident, adventurous, and sociable. They have good persuasive skills, and prefer positions of leadership.

324 APPENDIX D

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Careers in business and industry where persuasive skills are important are good choices for these individuals.

On the WRAT-4, Mr. Unclear scored at the 86th percentile in math, 75th per- centile in reading, 64th percentile in sentence comprehension, and 42nd percentile in spelling. His reading composite score was at the 69th percentile. These results could indicate a possible learning disorder in spelling, although cross-cultural con- siderations should be taken into account due to the fact that Mr. Unclear was an immigrant to this country at a young age.

Diagnosis

296.22 (F32.1) Major depression, single episode, moderate

309.28 (F43.23) Rule out: Adjustment disorder with mixed anxiety and depressed mood

V62.29 (Z56.9) Problems related to employment

V61.10 (Z63.0) Relationship distress with spouse

722.0 Displacement cervical intervertebral disc (chronic back pain)

Summary and Conclusions Mr. Unclear is a 48-year-old married male who was self-referred due to feelings of depression, anxiety, and discontent with his job and his marriage. Mr. Unclear fled from Cuba to Miami, Florida with his parents and two siblings when he was 5 years old. He describes his family as close, and he continues to live near his children, siblings, and mother. His father died approximately four years ago. He married while in college. He and his wife subsequently raised two girls who are now in their mid-20s, married, and have their own children.

Mr. Unclear finished college with a degree in business and has been working as an accountant for the past 25 years. He reports feeling dissatisfied in his career and states that he wants to “do something more meaningful with his life.” He also reports marital discord, which he attributes partly to medical problems his wife had after the birth of their second child. These problems, he states, resulted in diminishing sexual relations with his wife.

Mr. Unclear was oriented during the session but appeared anxious and talked about feelings of depression. He noted that he often feels fatigued, has difficulty sleeping, and has fleeting thoughts of suicide, which he states he would not act upon. Recently, he has had obsessive worries about having a heart attack, although there is no medical reason to support his concerns. Chronic back pain due to a car accident a few years ago seems to exacerbate his current feelings of depression.

Throughout testing the consistent themes of depression, isolation, and hope- lessness emerged. High scores on the BDI-II and the MMPI-II depression scale evi- denced this. This was additionally indicated by specific responses to the Rorschach, TAT cards, the KFD, and the sentence completion.

APPENDIX D 325

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Dissatisfaction with his marriage, sadness about the loss of his father, and chronic pain were also major themes that arose during testing. Testing also revealed a person whose career is a good match for his personality. However, he might be more challenged if he entered a position requiring additional responsibili- ties and leadership skills. Such a change may be disadvantageous if Mr. Unclear does not receive treatment for his depression. Finally, testing also shows a possible learning disability in spelling, although cross-cultural issues may have affected his score.

On a positive note, testing and the clinical interview showed a man who was neatly dressed and open to collaborating with the examiner. He has worked hard in his life and is proud of the family he has raised. He was grounded in reality, willing to engage interpersonally, and showed fair to good judgment and insight. He seems to be aware of many of his most pressing concerns and showed some willingness to address them.

Recommendations 1. Counseling, 1 hour a week for depression, possible anxiety, marital discord,

and career dissatisfaction. 2. Possible marital counseling with particular focus on sexual relations of the

couple. 3. Referral to a physician/psychiatrist for medication, possibly antidepressants. 4. Possible further assessment for learning problems. 5. Long-term consideration of a career move following alleviation of depressive

feelings and addressing possible learning problems. 6. Possible orthopedic reevaluation of back problems.

Sigmund Freud, MD Sigmund Freud

326 APPENDIX D

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

EA P P E N D I XSupplemental StatisticalEquations

PEARSON PRODUCT-MOMENT CORRELATION

r ¼ N∑XY � ð∑XÞð∑YÞffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi N∑X2 � ð∑XÞ2

ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi N∑Y2 � ð∑YÞ2

qr

Where: N is the number of scores X is the first set of test scores Y is the second set of test scores

KUDER-RICHARDSON

KR20 ¼ n

n � 1 h i" SD2 � ∑pq

SD2

#

Where: n is the number of items on the test SD is the standard deviation p is the proportion of correct items q is the proportion of incorrect items

327

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

COEFFICIENT ALPHA

r ¼ n n � 1 h i" SD2 � ∑SD2i

SD2

#

Where: n is the number of items on the test SD is the standard deviation ∑SD2i is the sum of the variances of item scores

STANDARD DEVIATION—ALTERNATIVE FORMULA

SD ¼ ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ∑X2

N � M2

s

Where: SD is the standard deviation X is the test scores N is the number of scores M is the mean of the test scores

328 APPENDIX E

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

FA P P E N D I XConverting Percentilesfrom z-Scores

The look-up table below can be used to quickly convert a z-score to an approximate percentile rank.

z-score Percentile z-score Percentile z-score Percentile

−6.0 0.0000001% −1.9 2.87% −0.5 30.85%

−5.0 0.00003% −1.8 3.59% −0.4 34.46%

−4.0 0.0032% −1.7 4.46% −0.3 38.21%

−3.0 0.13% −1.6 5.48% −0.2 42.07%

−2.9 0.19% −1.5 6.68% −0.1 46.02%

−2.8 0.26% −1.4 8.08% 0.0 50.00%

−2.7 0.35% −1.3 9.68% 0.1 53.98%

−2.6 0.47% −1.2 11.51% 0.2 57.93%

−2.5 0.62% −1.1 13.57% 0.3 61.79%

−2.4 0.82% −1.0 15.87% 0.4 65.54%

−2.3 1.07% −0.9 18.41% 0.5 69.15%

−2.2 1.39% −0.8 21.19% 0.6 72.57%

−2.1 1.79% −0.7 24.20% 0.7 75.80%

−2.0 2.28% −0.6 27.43% 0.8 78.81%

(Continued)

329

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

z-score Percentile z-score Percentile z-score Percentile

0.9 81.59% 1.7 95.54% 2.5 99.38%

1.0 84.13% 1.8 96.41% 2.6 99.53%

1.1 86.43% 1.9 97.13% 2.7 99.65%

1.2 88.49% 2.0 97.72% 2.8 99.74%

1.3 90.32% 2.1 98.21% 2.9 99.81%

1.4 91.92% 2.2 98.61% 3.0 99.87%

1.5 93.32% 2.3 98.93% 4.0 99.997%

1.6 94.52% 2.4 99.18% 5.0 99.99997%

Due to the fact percentiles are not evenly spaced along the bell curve (by defini- tion), you cannot interpolate between scores on the above table. Consequently, if a more precise percentile is needed that is not listed above, the formula for calcu- lating percentiles can be used.

The formula used to convert a z-score to a percentile rank is:

fðzÞ ¼ 1ffiffiffiffiffiffi 2p

p ez22 where z is the given z-score.

330 APPENDIX F

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Glossary

ability test Test that measures what a person can do in the cognitive realm. Achievement and aptitude tests are types of ability tests.

Accreditation Standards of Professional Associations Professional bodies that oversee graduate schools’ curriculum and training regarding therapists to en- sure standardization. Examples include the American Psychological Association (APA) and the Council for the Accredita- tion of Counseling and Related Educa- tional Programs (CACREP).

achievement test A type of ability test that measures what one has learned. Types of achievement tests include sur- vey battery tests, diagnostic tests, and readiness tests.

ACT score Created by converting a raw score to a standard score that gen- erally uses a mean of 21 and a standard deviation of 5 for college-bound stu- dents. The mean score for all students, including those who are not college- bound, is 18.

ADA (see “Americans with Disabilities Act”)

age comparison scoring A type of stan- dard score calculated by comparing an individual score to the average score of others who are the same age.

alternate, parallel, or equivalent forms reliability A method for determining

reliability by creating two or more alter- nate, parallel, or equivalent forms of the same test. These alternate forms mimic one another yet are different enough to eliminate some of the problems found in test-retest reliability (e.g., looking up an answer). In this case, rather than giving the same test twice, the examiner gives the alternate form the second time.

Americans with Disabilities Act (ADA) (PL 101-336) Law stating that to assure proper test administration, accommodations must be made for individuals with disabilities who are taking tests for employment and that testing must be shown to be relevant to the job in question.

anecdotal information A record of an individual that generally includes beha- viors that are consistent or inconsistent. A type of record and personal document.

aptitude test A type of ability test that measures what one is capable of doing. It includes intelligence tests, neuropsy- chological assessments cognitive ability tests, special aptitude tests, and multi- ple aptitude tests.

Armed Services Vocational Aptitude Battery (ASVAB) A multiple aptitude test that measures many abilities required for military and civilian jobs.

Army Alpha test An instrument cre- ated by Robert Yerkes, Lewis Terman,

and others during World War I to screen recruits for the military. Gener- ally considered the first modern group test.

Army Beta test A test of nonverbal intelligence used to screen recruits dur- ing World War I.

assessment A broad array of evaluative procedures that yield information about a person. An assessment may consist of many procedures, including a clinical interview, personality tests, ability tests, and informal assessment.

assessment of lethality A process of determining the risk one is for harming self or others. Common factors to assess include his or her development of thoughts of harm along a continuum, risk factors, protective factors, and will- ingness to contract for safety.

assessment report The “deliverable” or “end product” of the assessment pro- cess. Its purpose is to synthesize an assortment of assessment techniques so that a deeper understanding of the examinee is obtained and recommended courses of action can be offered.

Association for Assessment and Research in Counseling (AARC) A professional counselors’ organization dedicated to “promoting best practices in assessment, research, and evaluation in counseling.”

331

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

autobiography A procedure that asks an individual to write subjective infor- mation that stands out in his or her life. In some ways, the information highlighted in an individual’s autobiog- raphy is a type of projective test of his or her unconscious mind. A type of record and personal document.

Binet, Alfred Commissioned by the Ministry of Public Education in Paris in 1904 to develop an intelligence test to assist in the integration of “subnor- mal” children into the schools. His work with Theophile Simon led to the development of the first modern-day intelligence test.

Binet and Simon scale One of the first tests of intelligence.

biographical inventories Provides a detailed picture of the individual from birth. They can be obtained by con- ducting an involved, structured inter- view or by having the client answer a series of items on a checklist or respond to a series of questions.

Boston Process Approach A type of flexible battery approach to neuropsy- chological assessment.

breadth Covering all important or rel- evant issues; casting a wide net when assessing a client.

Career Exploration Tools A series of free instruments available on the O*NET Web site that can be used to assess abilities, interest, and values related to work.

Carl Perkins Act (PL 98-524) Origi- nally passed in 1984 and subsequently amended, this law assures that adults or special groups in need of job training have access to vocational assessment, counseling, and placement. These groups include (a) individuals with dis- abilities; (b) individuals from economi- cally disadvantaged families, including foster children; (c) individuals preparing for nontraditional fields; (d) single par- ents, including single pregnant women; (e) displaced homemakers; and (f) indi- viduals with limited English proficiency.

Cattell-Horn-Carroll (CHC) Integrated Model of Intelligence Developed a theory of intelligence that includes 16 broad ability factors, 6 of which are tentative, and over 70 associated tasks that may or may not be related to a g factor.

Cattell, James One of the earliest psy- chologists to use statistical concepts to understand people. His main emphasis became testing mental functions, and he is known for coining the term men- tal test.

Cattell, Raymond Differentiated fluid (innate) intelligence from crystallized (learned) intelligence and attempted to remove cultural bias from intelligence testing.

CHC Theory (see “Cattell-Horn-Carroll (CHC) Integrated Model of Intelligence”)

Civil Rights Acts (1964 and Amend- ments) Laws requiring that any test used for employment or promotion be shown to be suitable and valid for the job in question. If this is not done, alter- native means of assessment must be pro- vided. Differential test cutoffs are not allowed.

class interval Grouping scores from a frequency distribution within a prede- termined range to create a histogram or frequency polygon.

classification method A type of infor- mal assessment procedure where infor- mation is provided about whether an individual possesses certain attributes or characteristics (asking a person to check off those adjectives that seem to best describe him or her). It includes behavior and feeling word checklists.

Clerical Aptitude Test A special apti- tude test that assesses for clerical ability.

clinical assessment The process of assessing the client through multiple methods, including the clinical inter- view, informal assessment techniques, and objective and projective tests.

clinical interview A critical step in the assessment process in which the exam- inee is asked a broad range of back- ground questions to develop a profile of the individual. Interviews can be struc- tured, unstructured, or semi-structured.

clinical neuropsychology The assess- ment and intervention principles related to the central nervous system.

coefficient of determination (shared variance) The underlying commonality that accounts for the relationship between two sets of variables. It is calculated by squaring the correlation coefficient.

cognitive abilities tests Often based on what one has learned in school, these

instruments measure a broad range of cognitive abilities and are useful in mak- ing predictions about the future (e.g., whether an individual is likely to succeed in college). When a cognitive ability test score is much higher than a survey bat- tery achievement test it could indicate problems in learning (e.g., a learning dis- ability, motivation, poor teaching, etc.). This is a type of aptitude test.

competence in the use of tests In accor- dance with most professional codes of ethics, examiners are required to have adequate training and knowledge before using a test. Some test publishers have a tiered system to describe the levels of training required to administer their tests.

Comprehensive Test of Nonverbal Intelligence (CTONI-II) A nonverbal test of intelligence.

computer-driven assessment Allowing acomputerprogramtoassistintheassess- ment process and preparation of reports. Some observers believe computer-assisted questioning is at least as reliable as struc- tured interviews and can provide an accu- rate diagnosis at a low cost.

Conant, James Bryant Harvard presi- dent who conceived the idea of the SAT (formerly Scholastic Aptitude Test), which was developed by the Edu- cational Testing Service after World War II. Conant thought that such tests could identify the ability of indivi- duals and ultimately help to equalize educational opportunities.

concurrent validity Evidence that test scores are related to an external source that can be measured at around the same time the test is being given (“here and now” validity). This is a type of criterion-related validity.

Confidentiality An ethical (not legal) obligation to protect the client’s right to privacy. There are some instances in which confidentiality should be bro- ken, such as when clients are in danger of harming themselves or someone else.

construct validity Evidence that a test measures a specific concept or trait. Construct validity includes an analysis of a test through one or more of the following methods: experimental design, factor analysis, convergence with other instruments, and/or discrim- ination with other measures.

332 GLOSSARY

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

content validity Evidence that the test developer adequately surveyed the domain (field) the test is to cover, that test items match that domain, and that test items are accurately weighted for relative importance.

convergent validity A method of dem- onstrating construct validity by correlat- ing a test with some other well-known measure or instrument.

COPSystem A career measurement package that contains three instru- ments that measure interests, abilities, and values.

Corey, Corey, and Callanan’s ethical decision-making model Eight steps that a practitioner should go through when making complex ethical decisions: identifying the problem, identifying the potential issues involved, reviewing the relevant ethical guidelines, knowing relevant laws and regulations, obtaining consultation, considering possible and probable courses of action, listing the consequences of various decisions, and deciding on what appears to be the best course of action.

correlation coefficient The relationship between two sets of scores. Correlation coefficients range from þ1 to �1 and generally are reported in decimals of one-hundredths. A positive correla- tion shows a tendency for scores to be related in the same direction, while a negative correlation indicates an inverse relationship.

criterion referenced testing A method of scoring in which test scores are com- pared to a predetermined value or a set criterion.

criterion-related validity The relation- ship between a test and a standard (external source) to which the test should be related. The external stan- dard may be in the here-and-now (con- current validity) or a predictor of future criteria (predictive validity).

Cronbach’s coefficient alpha A method of measuring internal consistency by calculating test reliability using all the possible split-half combinations. This is done by correlating the scores for each item on the test with the total score on the test and finding the average correlation for all of the items.

cross-cultural sensitivity Awareness of the potential biases of assessment

procedures when selecting, administer- ing, and interpreting such procedures, as well as acknowledging the potential effects of age, cultural background, dis- ability, ethnicity, gender, religion, sex- ual orientation, and socioeconomic status on test administration and test interpretation.

crystalized intelligence As identified by Raymond Cattell, learned intelligence that tends to increase over time.

cumulative distribution A method of converting a frequency distribution of scores into increasing percentages as a function of the percentage of scores counted. A bar graph is often generated by placing the class interval scores along the x-axis and the cumulative percentages along the y-axis.

cumulative records File containing information about a client’s test scores, grades, behavioral problems, family issues, relationships with others, and other matters. School and workplace records are examples of cumulative records that can add vital information to our understanding of clients. It is a type of record and personal document.

depth Ensuring one is assessing the extent and seriousness of a concern.

derived score Score obtained by com- paring an individual’s score to his or her norm group and converting the individual’s raw score to a percentile or standard score such as z-score, T-score, deviation IQ, stanine, sten score, normal curve equivalent (NCE), college or graduate school entrance exam score (e.g., SAT, GRE, and ACT), or publisher-type score, or by using developmental norms such as age comparison and grade equivalent.

developmental norms Direct compari- son of an individual’s score to the aver- age scores of others at the same age or grade level. Examples include age com- parison and grade equivalent scoring.

deviation IQ Standard score with a mean of 100 and a standard deviation of 15. As the name implies, these scores are generally used in intelligence testing.

Diagnostic and Statistical Manual-5 for Mental Disorders (DSM-5) A compre- hensive diagnostic, single axis system of mental disorders published by the American Psychiatric Association that uses a dimensional assessment model

(mild, moderate, severe, and very severe) when assessing mental disorders.

diagnostic test Test that assesses prob- lem areas of learning. Often used to assess learning disabilities. Generally classified as a type of achievement test.

Differential Aptitude Test (DAT) A multiple aptitude test that measures abilities for career decision-making. An interest inventory is also available.

dimensional assessments (see “DSM-5”)

discriminant validity A method of demonstrating construct validity by correlating a test with other dissimilar instruments to ensure lack of a relationship.

direct observation A type of environ- mental assessment in which the exam- iner visits the classroom, workplace, or other setting in an effort to obtain addi- tional information not usually retrieved in the office.

Division 5 of the American Psychologi- cal Association A professional organi- zation for psychologists that promotes “research and practical application of psychological assessment, evaluation, measurement, and statistics.”

environmental assessment A naturalis- tic and systems approach to assessment in which practitioners collect informa- tion about clients from their home, work, or school environments. It includes direct observation, conducting a situational assessment, applying a sociometric assessment, or using an envi- ronmental assessment instrument. It is a type of informal assessment procedure.

environmental assessment instrument A type of environmental assessment in which any of a number of instruments are used in conjunction with simple observation to obtain additional infor- mation about a client.

equivalent forms reliability (see “alter- nate forms”)

Esquirol, Jean Used language to iden- tify different levels of intelligence while working in the French mental asylums; his work led to the concept of verbal intelligence.

ethical code Professional guidelines for appropriate behavior and guidance on how to respond under certain condi- tions, including when conducting assessment. Each professional group

GLOSSARY 333

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

has its own ethical code, including guidelines for counselors, psycholo- gists, social workers, marriage and family therapists, psychiatrists, human service professionals, and others.

event and time sampling Observing a targeted behavior for a set amount of time. It is a type of observation.

event sampling A type of observation in which one is observing a specific behavior with no regard to time.

experimental design validity Using experimentation to show that a test measures a specific concept or con- struct. It is a type of construct validity.

face validity Superficial observation that a test appears to cover the correct content or domain. This is not an actual form of validity since the appearance of test items may or may not accurately reflect the domain.

factor analysis A method of demon- strating construct validity by statisti- cally examining the relationship between subscales and the larger con- struct. It is frequently used in multiple aptitude tests to help developers deter- mine the purity of subtests.

false positives and negatives When an instrument incorrectly classifies some- one as having an attribute they do not have (false positive) or not having an attribute, which in fact, they do have (false negative).

feeling word checklist A type of classi- fication method that allows an individ- ual to identify those words that best describe kinds of feelings an individual might typically or atypically exhibit.

FERPA Family Educational Rights and Privacy Act of 1974 affirms the right of all individuals to gain access to their school records, including test records.

fixed battery approach to neuropsy- chological assessment Using a uniform (standard) set of instruments for all cli- ents who require a neuropsychological evaluation.

flexible battery approach to neuropsy- chological assessment Tailoring the instruments and techniques specific to each client who require a neuropsycho- logical evaluation.

fluid intelligence As identified by Raymond Cattell, a type of innate intelligence.

forensic evaluations A specialized type of assessment often used in a court of law that requires the author to testify as an expert witness. The assessing examiner frequently collects informa- tion from interviews, testing, and the reviewing of supplemental records. Professional organizations offer partic- ular endorsements for this.

Freedom of Information Act This law assures the right of individuals to access their federal records, including test records. Most states have similar laws that assure access to state records.

frequency distribution A method of understanding test scores by ordering a set of scores from high to low and listing the corresponding frequency of each score across from it.

frequency polygon A method of con- verting a frequency distribution of scores into a line graph. After combining scores by class interval, the class intervals are placed along the x-axis and the fre- quency of scores along the y-axis.

Galton’s board (see “quincunx”)

Galton, Sir Francis Examined relation- ship of sensorimotor responses to intelligence. He hypothesized that individuals who had a quicker reaction time and stronger grip strength were superior intellectually.

Gardner, Howard Vehemently opposed current constructs of intelligence measure- ment and developed his own theory of multiple intelligences asserting that there are eight or nine intelligences: verbal- linguistic, mathematical-logical, musical, visual-spatial, bodily-kinesthetic, interper- sonal, intrapersonal, naturalist, and exis- tential intelligence.

General Aptitude Test Battery (GATB) One of the first tests to measure multi- ple aptitudes. It was develped by the U.S. Employment Service.

General (g) factor of intelligence A belief that there is an underlying, over- all, factor that mediates intelligence. It was popularized by Charles Edward Spearman.

generosity error Error that occurs when an individual rates another per- son inaccurately because he or she identifies with the person being rated.

genogram A map of an individual’s family tree that may include the family

history of illnesses, mental disorders, substance use, expectations, relation- ships, cultural issues, and other con- cerns relevant to counseling. Special symbols may be used to assist in creat- ing this map. It is a type of record and personal document.

Gesell Developmental Observation A readiness test that assesses development of the whole child.

grade equivalent scoring A type of standard score calculated by compar- ing an individual’s score to the average score of others at the same grade level.

graphic scale (see “Likert-Type scale”)

Griggs v. Duke Power Company Asserted that tests used for hiring and advancement at work must show that they can predict job performance for all groups.

Guilford, J. P. Developed a multifactor/ multidimensional model of intelligence based on 180 factors. His three- dimensional model can be represented as a cube and involves three kinds of cognitive ability: operation, content, and product.

Hall, G. S. Worked with Wundt and eventually set up his own experimental lab at John Hopkins University. Became a mentor to other great American psy- chologists and was the founder and first president of the American Psycho- logical Association in 1892.

halo effect Error that occurs when the overall impression of an individual clouds the rating of that person in one or more select areas.

Halstead-Reitan A widely utilized fixed battery neuropsychological assessment consisting of eight core tests.

Health Insurance Portability and Accountability Act (HIPAA) Ensures the privacy of client records, including testing records, and the sharing of such information. In general, HIPAA restricts the amount of information that can be shared without client con- sent and allows clients to have access to their records, except for process notes used in counseling.

high-stakes testing The term used to describe the pressure placed on exam- inees, teachers, administrators, and others as a result of the major decisions that are made from the use of tests,

334 GLOSSARY

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

usually national standardized tests such as the SATs and those used as the result of No Child Left Behind.

Histogram A method of converting a frequency distribution of scores into a bar graph. After combining scores by class interval, the class intervals are placed along the x-axis and the fre- quency of scores along the y-axis.

holistic process The idea that it is important to use multiple measures of assessment to draw adequate conclu- sions about a person.

Individualized Education Plan (IEP) PL94-142 states that children who are identified as having a learning dis- ability will be provided a school team that will create an individualized edu- cation plan to assist the student with his or her learning problem(s).

Individuals with Disabilities Education Act (IDEA) (Expansion of PL 94-142) This legislative act assures the right of students to be tested, at the school sys- tem’s expense, if they are suspected of having a disability that interferes with learning. These students must be given accommodations for their disability and taught within the “least restrictive environment,” which often is a regular classroom.

Informalassessmentinstruments Assess- ment instruments that are often devel- oped by the user and are specific to the testing situation. All of these instruments can be used to assess broad areas of ability and personality attributes in a variety of settings. Types of informal assessment instruments include observa- tion, rating scales, classification methods, environmental assessment, records and personal documents, and performance- based assessment.

informed consent Principle that indivi- duals being assessed should give their permission for the assessment after they are given information concerning such items as the nature and purposes of the assessment, fees, involvement by others in the assessment process (e.g., teachers, therapists), and the limits of confidentiality.

intellectual and cognitive functioning Tests that measure a broad range of cognitive functioning in the following domains: general intelligence, mental retardation, giftedness, and changes

in overall cognitive functioning. It includes intelligence testing and neuro- psychological assessment. They are types of aptitude tests.

intelligence quotient Concept devel- oped by Lewis Terman, who divided mental age by chronological age and multiplied the quotient by 100 to derive a score that he called “IQ.”

intelligence test Individual tests that assess a broad range of cognitive capa- bilities that generally results in an “IQ” score. Often used to identify mental retardation, giftedness, learning dis- abilities, and general cognitive func- tioning. It is a type of aptitude test.

interest inventories Tests that measure likes and dislikes as well as one’s per- sonality orientation toward the world of work. Generally used in career counsel- ing. It is a type of personality test.

internal consistency A method of determining reliability of an instrument by looking within the test itself, or not going “outside of the test” to determine a reliability estimate as is done with test-retest or parallel forms reliability. Some types of internal consistency reli- ability include split-half (or odd-even), Cronbach’s Coefficient Alpha, and Kuder-Richardson.

Interquartile range Provides the range of the middle 50% of scores around the median. Because it eliminates the top and bottom quartiles, the quartile range is most useful with skewed curves because it offers a more repre- sentative picture of where a large per- centage of the scores fall.

interrater reliability When two or more raters observe or rank order a behavior in order to find the level of concordance between the raters. The higher the con- cordance, the higher the reliability.

interval scale A scale of measurement in which there are equal distances between measurements but no absolute zero reference point.

invasion of privacy All tests invade one’s privacy but concerns about inva- sion of privacy are lessened if the client has given informed consent, has the ability to accept or refuse testing, and knows the limits of confidentiality.

Iowa Test of Basic Skills (ITBS) A sur- vey battery achievement test that

measures skills to satisfactorily prog- ress through school.

item response theory An extension of classical test theory that allows for more detailed analyses of individual test items. Tools such as the item char- acteristic curve allow in depth exami- nation of an item’s discrimination and difficulty.

Jaffee v. Redmond In this case, the Supreme Court upheld the right of a licensed social worker to keep her case records confidential. Describing the social worker as a “therapist” and “psychother- apist,” the ruling will likely protect all licensed therapists in federal courts and may affect all licensed therapists who have privileged communication.

journals and diaries Procedures that allow individuals to describe them- selves and can provide valuable insight into self, or to a clinician, by revealing unconscious drives and desires and by uncovering patterns that highlight issues in a client’s life. Types of record and personal documents.

Jung, Carl Developed a list of 156 words (word association) to which sub- jects were asked to respond as quickly as possible. Depending on the response and the answer time, Jung believed he could identify complexes. He also created the personality type construct that is used in the Myers-Briggs Type Inventory.

Kaufman Assessment Battery for Chil- dren (KABC-II) An intelligence test that measures cognitive ability for ages 3–8.

KeyMath3 A comprehensive test to assess for learning disabilities in math.

Kraeplin, Emil Developed a crude word association test to study schizo- phrenia in the 1880s.

Kuder-Richardson A method of calcu- lating internal consistency by using all the possible split-half combinations. This is done by correlating the scores for each item on the test with the total score on the test and finding the aver- age correlation for all of the items.

Likert-type scale Sometimes called graphic scales, this kind of rating scale contains a number of items that are being rated on the same theme and are anchored by both numbers and a state- ment that correspond to the numbers.

GLOSSARY 335

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

major professional associations (see “Association for Assessment and Research in Counseling” and “Division 5 of the American Psychological Association”).

mean A measure of central tendency that is calculated by adding all of the scores and dividing by the total number of scores. It is the arithmetic average of the scores.

measures of central tendency Indica- tors of what is occurring in the mid- range or “center” of a group of scores. Three measures of central ten- dency are the mean, median, and mode.

measures of variability Indicators of how much scores vary in a distribution. Three types of measures of variability are the range, quartile range, and the standard deviation.

median A measure of central tendency that is the middle score, or the score for which 50% of scores fall above and 50% fall below. In a skewed curve or skewed distribution of test scores, the median is generally the most accurate measure of central tendency since it is not affected by unusually high or low scores.

Medical College Admission Test (MCAT) Assesses knowledge of physi- cal sciences, verbal reasoning, biologi- cal sciences; predicts grades in medical school.

Mental Measurements Yearbook by Buros A sourcebook of reviews of more than 2,000 different tests, instru- ments, or screening devices. Most large universities carry this in hardbound copy and/or online.

mental status exam A portion of the assessment and written report that addresses the client’s appearance and behavior, emotional state, thought com- ponents, and cognitive functioning.

Miner, J. B. Developed one of first group interest inventories in 1922 to assist large groups of high school stu- dents in selecting an occupation.

mode A measure of central tendency that is the score that occurs more often than any other.

moral principles of ethical-decision making model Six moral principles often identified as important in making thorny ethical decisions: autonomy,

nonmaleficence, beneficence, justice, fidelity, and veracity.

multiple aptitude tests Tests that mea- sure many aspects of ability. Often use- ful in determining the likelihood of success in a number of vocations. It is a type of aptitude test.

Murray, Henry Developed the The- matic Apperception Test (TAT), which asks a subject to view a number of standard pictures and create a story to explain the situation as he or she best understands it. This test is based on his needs-press theory.

Musical Aptitude Test A special apti- tude test that assesses for musical ability.

National Assessment of Educational Progress (NAEP) Sometimes called “the nation’s report card,” these assess- ment instruments allows states to com- pare progress in achievement to other states around the country.

negatively skewed curve A set of test scores where the majority fall at the upper or positive end. It is said to be a negatively skewed curve or distribu- tion because a few “negative” or low- end scores have stretched or skewed the curve to the left.

neuropsychology A domain of psy- chology that examines brain–behavior relationships.

neuropsychological assessment A spe- cialized form of assessment that evalu- ates the central nervous system for functioning. It is frequently applied after a traumatic brain injury, illnesses that may have caused brain damage, or with elderly that may be experiencing changes in the brain.

No Child Left Behind (NCLB) Act A federal act that requires all states to have a plan in place to show how stu- dents will have obtained proficiency in reading/ language arts and math. The act has caused states to create or iden- tify achievement tests whose purpose is to assess all students and ensure they are meeting the minimum standards.

nominal scale A scale of measurement in which numbers are arbitrarily assigned to represent different catego- ries or variables. The only math that can be applied to these numbers is cal- culation of the mode.

nonverbal intelligence tests Intelli- gence tests that rely on little or no ver- bal expression.

norm referencing A method of scoring in which test scores are compared to a group of other scores called the norm group.

normal curve A bell-shaped curve showing the usual frequency distribu- tion of values of measured human traits and other natural phenomena.

normal curve equivalents (NCE) Fre- quently used in the educational com- munity, this is a form of standard scoring that has 99 equal units along a bell-shaped curve, with a mean of 50 and a standard deviation of 21.06.

numerical scale A type of rating scale in which a statement or question is fol- lowed by a choice of numbers arranged from high to low along a number line.

objective personality testing Multiple- choice or true/false type formats that assess various aspects of personality. Often used to increase client insight, to identify psycho-pathology, and to assist in treatment planning. It is a type of personality test.

observation Observing behaviors of an individual to develop a deeper under- standing of one or more specific behaviors (e.g., observing a student’s acting-out behavior in class or asses- sing a client’s ability to perform eye– hand coordination tasks to determine potential vocational placements). It includes time sampling, event sampling, and time and event sampling. It is a type of informal assessment.

O*NET A free online database main- tained by the U.S. Department of Labor/Employment and Training Administration that contains hundreds of classifications of job descriptions. It replaced the Dictionary of Occupa- tional Titles (DOT).

ordinal scale A scale of measurement in which the magnitude or rank order is implied; however, the distance between measurements is unknown.

Otis-Lennon School Ability Test 8 (OLSAT 8) A cognitive ability test that assesses abstract thinking and reasoning skills via verbal and nonverbal sections.

parallel forms reliability (see “alter- nate forms”)

336 GLOSSARY

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Parsons, Frank Founder of the voca- tional counseling movement.

Peabody Individual Achievement Test (PIAT) A diagnostic test that measures six content areas for screening K–12 students.

Percentiles A method of comparing raw scores to a norm group by calcu- lating the percentage of people falling below an obtained score, with ranges from 1 to 99, and 50 being the mean.

performance-based assessment The evaluation of an individual using a variety of informal assessment proce- dures that are often based on real world responsibilities. This kind of assessment is not highly loaded for cog- nitive skills, and these procedures are sometimes seen as an alternative to standardized testing (e.g., a portfolio). Type of informal assessment.

personality testing Tests in the affective realm used to assess habits, tempera- ment, likes and dislikes, character, and similar behaviors. Types of personality tests include interest inventories, objec- tive personality tests, and projective per- sonality tests.

Piaget, Jean Approached intelligence from a developmental perspective by observing how children’s cognitions were shaped as they grew. He identified four stages of cognitive development: sen- sorimotor, preoperational, concrete oper- ational, and formal operational. He also believed that cognitive development is adaptive, and he conceived the concepts of assimilation and accommodation.

portfolio assessment A type of performance-based assessment in which a number of items are collected and eventually shown to such indivi- duals as potential employer’s or poten- tial colleges as an indication of the applicant’s ability. Sometimes used in lieu of standardized testing.

positively skewed curve A set of test scores in which the majority fall at the lower or negative end. It is said to be a positively skewed curve or distri- bution because a few “positive” or high-end scores have stretched or skewed the curve to the right.

practicality When selecting a test, one of the cornerstones of test worthiness is practicality. Practical concerns include time, cost, format, readability, and

ease of administration, scoring, and interpretation.

predictive validity Evidence that test scores are able to predict a future crite- rion or standard. A type of criterion- related validity.

privileged communication The legal right to maintain privacy of a conversation. The privilege belongs to the client, and only the client can waive that privilege.

projective personality tests Tests that present a stimulus to which individuals can respond. Personality factors are interpreted based on the individual’s response. Often used to identify psycho- pathology and to assist in treatment planning. A type of personality test.

proper diagnosis Due to the delicate nature of diagnoses, ethical codes stress that professionals should be particu- larly careful when deciding which assessment techniques to use in form- ing a diagnosis for a mental disorder.

publisher-type scores Created by test developers who generate their own unique standard score that employs a mean and standard deviation of the publisher’s choice.

quincunx Also known as Galton’s board, it is a board with protruding pins or nails where on balls dropped will fall along the bell or normal curve.

range The simplest measure of vari- ability is the range, which is calculated by subtracting the lowest score from the highest score and adding 1.

rank order A rating scale providing a series of statements that the respondent is asked to place from highest to lowest based on his or her preferences.

rating scales Scales developed to assess any of a number of attributes of the examinee. Can be rated by the exam- inee or someone who knows the exam- inee well. Some commonly used rating scales include numerical scales, Likert- type scales (graphic scales), semantic differential scales, and rank-order scales. A type of informal assessment.

ratio scale A scale of measurement that has a meaningful zero point and equal intervals and thus can be manipulated by all mathematical principles.

raw score An untreated score before manipulation or processing to make it a standard score, as must be done for

all norm-referenced tests. Raw scores alone tell us little, if anything, about how a person has done on a test. We must take an individual’s raw score and do something to it to give it meaning.

readiness tests Tests that measure one’s readiness for moving ahead in school. Often used to assess readiness to enter kindergarten or first grade. It is a type of achievement test.

records and personal documents Asses- sing behaviors, values, and beliefs of an individual by examining such items as autobiographies, diaries, personal jour- nals, genograms, or school records. It is a type of informal assessment.

release of test data Test data should be released to others only if the client has signed a release form. The release of such data is generally granted only to individuals who can adequately interpret the test data, and professionals should assure that those who receive such data do not misuse the information.

reliability The degree to which test scores are free from errors of measure- ment, also, the capacity of an instru- ment to provide consistent results. (see “test-re-test,” “alternative forms,” and “internal consistency”)

Rorschach, Herman A student of Carl Jung, he created the Rorschach Inkblot test by splattering ink onto sheets of paper and folding them in half. He believed the interpretation of an indivi- dual’s reactions to these forms could tell volumes about the individual’s uncon- scious life.

Rorschach Inkblot Test (see “Rorschach, Herman”)

S factors of intelligence The belief that some aspects of intelligence are mediated by very specialized (s) aspects of human functioning and largely not influenced by the generalized (g) factor. Popularized by Charles Edward Spearman.

SAT-type scores Standard score that has a standard deviation of approxi- mately 100 and a mean of 500 for each section of the exam. It is a type of cognitive ability test.

scales of measurement Ways of defin- ing the attributes of numbers and how they can be manipulated. The four types of measurement scales are nomi- nal, ordinal, interval, and ratio.

GLOSSARY 337

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

scatterplot A graph of two sets of test scores used to visually display the rela- tionship or correlation. If the dots are plotted closer together, the correlation is moving toward a positive or negative 1. If the dots are spread out or completely random, the correlation is closer to zero.

Section 504 of the Rehabilitation Act This act applies to all federally funded programs receiving financial assistance and was established to prevent discrim- ination based on disability.

Seguin, Edouard Nineteenth-century scientist who suggested that language ability is related to intelligence.

Self-Directed Search A self-adminis- tered, scored, and interpreted interest inventory that was created by Holland and uses his personality types.

semantic differential A rating scale that provides a statement followed by one or more pairs of words that reflect opposing traits. A number line may or may not be associated with the dichot- omous pairs of words.

semi-structured interview Uses pre- scribed items and thereby allows the examiner to obtain the necessary infor- mation within a relatively short amount of time. However, this kind of interview also gives leeway to the examiner should the client need to “drift” during the interview process.

Sequin, Edouard Worked with indivi- duals with intellectual disabilities and developed the form board to increase his patients’ motor control and sensory discrimination. This was the forerunner to performance IQ.

Simon, Theophile With Alfred Binet, developed the Binet-Simon scale, one of the first instruments to assess intelligence.

single aptitude tests See special apti- tude tests.

Situational assessments A type of envi- ronmental assessment used to examine how an individual is likely to respond in a contrived but natural situation. An example of this type of procedure is when a potential doctoral student counsels a role-play client.

skewed curve A set of test scores that do not fall along the normal curve.

sociometric instruments A kind of environmental assessment that identifies

the relative position of an individual within a group. This type of instrument is often used when one wants to deter- mine the dynamics of individuals within a group, organization, or institution.

Spearman, Charles Edward Believed in a two-factor approach to intelligence that included a general factor (g) and a specific factor (s), both of which he considered important in understanding intelligence.

Spearman-Brown formula A mathe- matical formula that can be used with split-half or odd-even reliability esti- mates to increase the accuracy, which is impaired because of the shortening (splitting in half) of the test.

special aptitude tests Tests that mea- sure one aspect of ability. Often useful in determining the likelihood of success in a vocation (e.g., a mechanical apti- tude test to determine success as a mechanic). A type of aptitude test.

split-half or odd-even reliability This method of internal consistency reliabil- ity splits the test in half and correlates the scores of one half of the test with the other half. Hence, it requires only one form and one administration of the test.

standard deviation A measure of vari- ability that describes how scores vary around the mean. In all normal curves the percentage of scores between stan- dard deviation units are the same; hence, the standard deviation com- bined with the mean can tell us a great deal about a set of test scores.

standard error of estimate The range of scores we would expect a person’s score to fall on an instrument based on his or her score from a previous known test. The equation allows us to use the score on instrument X to pre- dict a range of scores on instrument Y.

standard error of measurement The range of scores where we would expect a person’s score to fall if he or she took the instrument over and over again—in other words, where a “true” score might lie. It is calculated by taking the square root of 1 minus the reliability and multiplying that number by the standard deviation of the desired score.

standards in assessment Any of a num- ber of standards developed to help guide practitioner’s in the proper use

of tests and inform test users of their rights. (see list of standards, Table 2.1)

standard scores Scores derived by con- verting an individual’s raw score to a new score that has a new mean and new standard deviation. Standard scores are generally used to make test results easier for the examinee to interpret.

Stanford-Binet 5 (SB5) An individual intelligence tests that uses basal and ceil- ing levels to determine start and stop points and measures verbal and nonver- bal intelligence across five factors.

stanines Derived from the term “stan- dard nines,” this is a standardized score frequently used in schools. Often used with achievement tests, stanines have a mean of 5 and a standard deviation of 2, and range from 1 to 9.

sten scores Derived from the name “standard 10,” a standard score that is commonly used on personality inven- tories and questionnaires. Stens have a mean of 5.5 and a standard deviation of 2.

Sternberg’s triarchic theory (see “triarchic theory”)

Strong, Edward Led a team of researchers in the 1920s to develop the Strong Vocational Interest Blank. The test is now known as the Strong Interest Inventory and is still one of the most popular interest inventories ever created.

Strong Interest Inventory A popular interest inventory that uses the Holland Codes to assess general occupational themes, basic interests, occupational scales, personal style, and offers a response summary.

Strong Vocational Interest Blank Developed by Edward Strong, one of the first tests to measure interests and used for career counseling.

structured interview An interview in which the examinee is asked to respond to a set of preestablished items. This is often done verbally, although sometimes clients can respond to written items.

suicide assessment (see “assessment of lethality”)

survey battery tests Multiple choice and true/false tests usually given in school set- tings, that measure broad content areas. Often used to assess progress in school. It is a type of achievement test.

338 GLOSSARY

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

T-score A type of standard score that can be easily converted from a z-score. T-scores have a mean of 50 and a standard deviation of 10 and are gen- erally used with personality tests.

Terman, Lewis Professor at Stanford University who analyzed the Binet and Simon scale and made a number of revi- sions to create the Stanford-Binet intelli- gence test that is still used today. Terman was the first to incorporate in his test the ratio of chronological age and mental age, calling it the “intelli- gence quotient” or “IQ.”

test administration Tests should be administered appropriately as defined by the way they were established and standardized. Alterations to this pro- cess should be noted and interpreta- tions of test data adjusted if conditions were not ideal.

test scoring and interpretation The process of examining tests and making judgments about test data. Professionals should reflect on how issues of test wor- thiness, includingthe reliability, validity, cross-cultural fairness, and practicality of the test, might affect the results.

test security Professionals have the responsibility to make reasonable efforts to assure the integrity of test content and the security of the test itself. Professionals should not dupli- cate tests or change test material with- out the permission of the publisher.

test worthiness Determined by an objective analysis of a test in four criti- cal areas: (1) validity: whether or not it measures what it is supposed to mea- sure; (2) reliability: whether or not the score an individual receives on a test is an accurate measure of his or her true score; (3) cross-cultural fairness: whether or not the score the individual obtains is a true reflection of the individual and not a reflection of cul- tural bias inherent in the test, and

(4) practicality: whether or not it makes sense to use a test in a particular situation.

test-retest reliability Giving the test twice to the same group of people and then correlating the scores of the first test with those of the second test to deter- mine the reliability of the instrument.

tests A subset of assessment techniques that yield scores based on the gathering of collected data.

Thematic Apperception Test (TAT) Developed by Henry Murray, one of the first projective tests. Asks individuals to tell stories about pictures they view.

Thorndike, Edward Believed that tests could be given in a format that was more reliable than the previous meth- ods. His work culminated with the development of the Stanford Achieve- ment Test in 1923.

Thurstone, Louis Developed a multi- factor approach or model of intelligence that included seven primary factors: verbal meaning, number ability, word fluency, perception speed, spatial abil- ity, reasoning, and memory.

time sampling A form of observation in which behaviors are noted during a set duration of time.

triarchic theory A theory of intelligence that has three subtheories including componential (analytical), experiential (creative), and (practical) aspects.

Universal Nonverbal Intelligence Test (UNIT) A nonverbal test of intelligence.

unstructured interview An interview in which the examiner does not have a preestablished list of items or questions to which the client can respond; instead, client responses to examiner inquiries establish the direction for follow-up questioning.

validity The degree to which all of the accumulated evidence supports the

intended interpretation of test scores for the intended purpose. Validity is a unitary concept that attempts to answer the question: How well does a test measure what it’s supposed to measure? (see content validity; criterion-related validity, and con- struct validity)

Vernon, Philip Believed that subcom- ponents of intelligence could be added in a hierarchical manner to get a score for a cumulative (g) or general factor. Many of today’s intelligence tests con- tinue to use this concept.

Wechsler Nonverbal Scale of Ability (WNV) A nonverbal test of intelligence.

Wide Range Achievement Test 4 (WRAT4) Assesses basic learning prob- lems in reading, spelling, math, and sentence comprehension.

Woodcock-Johnson A diagnostic test that assesses a broad assessment of ability, ages 2–90.

Woodworth’s Personal Data Sheet An instrument with 116 items developed to screen World War I recruits for their susceptibility to mental health prob- lems. It is considered the precursor of modern-day personality testing.

Wundt, Wilhelm Developed one of the first psychological laboratories. He also set out to create “a new domain of science” that he called physiological psychology, which later became known as psychology.

Yerkes, Robert President of the American Psychological Association during World War I, he chaired a spe- cial committee designed to screen new recruits. The committee developed the Army Alpha test.

z-score The most fundamental stan- dard score, which is created by convert- ing an individual’s raw score to a new score that has a mean of 0 and a stan- dard deviation of 1.

GLOSSARY 339

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Index

A AAC. See Achievement/Ability Comparison AACE. See Association for Assessment in

Counseling and Education Ability tests

definition of, 14 emergence of, 7–12 types of, 7–12

Abusive behaviors, 288 Accommodation, 196 Accreditation standards, 34 Achievement/Ability Comparison (AAC),

175 Achievement tests

definition of, 16 diagnostic, 165–171 Metropolitan, 165 NCLB and, 159, 160 readiness, 171–174 survey battery, 159–165 teachers’ interpretation, 182 vs. aptitude testing, 159

ACT. See American College Testing Assessment

ADA. See Americans with Disabilities Act Addictive and substance-related disorders, 51 ADHD. See Attention-deficit/hyperactivity

disorder Administration, of good test, 103–106 Adverse effects of medications disorders,

other, 52 Age comparison scores, 139 American Board for Forensic Psychology, 35 American College Testing Assessment

(ACT), 174 raw scores, 136–138 scores, 178–179

American Counseling Association (ACA), 22–25, 34

American Psychological Association (APA), 22–25

Americans with Disabilities Act, 29, 32, 100 Analytical facet, 198 Anecdotal information, 295 Anxiety disorders, 50 Apperceptive Personality Test (APT), 268 APT. See Apperceptive Personality Test Aptitude testing achievement tests vs., 159 definition of, 16 multiple, 233–238 special, 238–241

Armed Services Vocational Aptitude Battery (ASVAB), 221, 233, 234–237

Army Alpha administration of, 8 definition of, 10 eugenics and, 11 sample of, 9

Army Beta, 10, 11 Artistic aptitude test, 240 Artistic personality type, 224 Assessment ability vs. developmental maturity, 174 career. See Career assessment clinical. See Clinical assessment cross-cultural issues, 36–37 definition of, 4 diagnosis in. See Diagnosis; Diagnostic

and Statistical Manual, Fourth Edition-Text Revision

environmental. See Environmental assessment

ethical issues, 22–25 history of

ability test, 7–12 ancient, 5 informal procedures, 14 instrument types, 15 modern, 5–6, 14 personality tests, 12–13

informal. See Informal assessment

intelligence. See Intelligence testing legal issues, 28–33 multiple procedures, 5 NAEP, 160–162 neuropsychological. See Neuropsycholog-

ical assessment ongoing nature of, 36 performance-based, 297–298 portfolio, 298 procedure for, 5 purpose of, 158 situational, 291 sociometric, 291 standards for, 26 subset, 4–5 survey battery, 159–165 test category-based definitions, 16–17

Assessment reports breadth, 61 computer-driven, 63–64 depth, 61 information gathering, 60–61 instrument selection, 64 interviews, 61–63 purpose of, 60 sections conclusions, 74–75 demographic information, 65–66 diagnosis, 73–74 educational history, 68 family background, 66–67 medical/counseling history, 67 mental status, 68–72 problem presentation, 66 recommendations, 75 results, 72–73 substance use/abuse, 67 summary, 74–75 vocational history, 68

summary of, 75–77 Assimilation, 196

340

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Association for Assessment in Counseling and Education, 33

Associations accreditation standards of, 34 professional, 34

ASVAB. See Armed Services Vocational Aptitude Battery

ataque de nervios (“attack of nerves”), 54 Attention-deficit/hyperactivity disorder

(ADHD), 44, 264 Autobiography, 296 Autonomy, 27

B BAC. See Blood alcohol content Basic Interest Scale, 225 Beck Depression Inventory-II

development, 256 norm group, 256 validity, 256

Behavior abusive, 288 checklists, 287–288

Behavior Disorder Scale-Revised, 292 Behavior Rating Inventory of Executive

Function-Adult version, 292 Bell-shaped curve. See Normal curve Bender Visual-Motor Gestalt II, 270–272 Beneficence, 27 Big Five personality traits, 262–264 Binet, Alfred, 7, 191 Binet scale, 7 Biographical inventories, 293–295 Bipolar and related disorders, 49 Blood alcohol content (BAC), 146–147 Bodily-kinesthetic intelligence, 197 Boston Process Approach (BPA), 213–214 BPA. See Boston Process Approach Breadth, of assessment, 61 Briggs, Katharine, 257 Buckley Amendment. See Family Education

Rights and Privacy Act Buros Mental Measurements Yearbook, 105

C CACREP. See Council for the Accreditation

of Counseling and Related Educa- tional Programs

Campbell Interest and Skill Survey (CISS), 233

CAPS. See Career Ability Placement Survey Career Ability Placement Survey (CAPS), 230 Career assessment

definition of, 222 helpers’ role, 241–242 interest inventories, 222–233 multiple aptitude tests, 233–238 special aptitude testing, 238–241

Career Assessment Inventory, 233 Career Interest Inventory (CII), 237 Career Occupational Preference System

(COPS), 229 Career Orientation Placement and Evalua-

tion Survey (COPES), 230 Carl Perkins Act (PL 98–524), 100 Carl Perkins Career and Technical Educa-

tion Improvement Act, 33 CASE-IMS. See Comprehensive Assessment

of School Environments Information Management System

CAT. See Children’s Apperception Test Category Test, of Halstead-Reitan battery,

212 Cattell, James McKeen, 7 Cattell, Raymond, 193, 195, 201, 260,

262 Cattell-Horn-Carroll (CHC) integrated

model, 199 CHC integrated model. See Cattell-

Horn-Carroll integrated model Checklists

behavior, 287–288 feeling word, 288

Children’s Apperception Test (CAT), 269 CII. See Career Interest Inventory CISS. See Campbell Interest and Skill

Survey Civil Rights Acts, 29, 30 Civil Rights Acts (1964 and Amendments),

100–101 Classification systems, 287–290 Class intervals, 113 Clerical aptitude tests, 238–239 Clerical Test Battery (CTB2), 238 Clinical assessment, 52

definition of, 248 helpers’ role, 274 labeling issues, 274 objective personality testing, 248–266 projective personality testing, 266–274

Clinical interviews, 61–63 Code of Fair Testing Practices, 26 Coefficient of determination, 86–87 CogAT. See Cognitive ability tests Cognitive ability tests (CogAT), 177–178

college admission, 178–182 definition of, 16 function, 177 OLSAT, 174–177 and survey battery tests, 164

Cognitive development model of intelligence, 195–196, 201

Cognitive domain, 7–12 College admission exams. See specific tests Competence, 23 Componential subtheory, 198 Comprehensive Assessment of School

Environments Information Management System (CASE-IMS), 291

Comprehensive Test of Nonverbal Intelligence (CTONI), 209

Computer-driven assessments, 63–64 Conant, James Bryant, 10–11 Concurrent validity, 89, 264 Conduct disorders, 51 Confidence interval, 143 Confidentiality, 23–24 Conners 3rd Edition, 264–265 Construct validity, 91–93, 94 Content, cognitive, 193 Content validity, 87–89 Contextual subtheory, 198 Conventional personality type, 224 Convergent validity, 92 Co-occurring disorders, 52 COPES. See Career Orientation Placement

and Evaluation Survey COPS. See Career Occupational Preference

System

Correlation coefficient, 84–86 Cost factors, 102 Council for the Accreditation of Counseling

and Related Educational Programs (CACREP), 34

Creative facet, 198 Criterion referencing, 128–129 Criterion-related validity, 89–91 Cronbach’s coefficient reliability, 97 Cross-cultural fairness, 81

description of, 99–101 examining, 105 informal assessment, 301–302 intelligence tests, 101 legal aspects, 100–101

Cross-cultural sensitivity, 24 Crystal intelligence, 193–195, 201 CTB2. See Clerical Test Battery CTONI. See Comprehensive Test of

Nonverbal Intelligence Cultural considerations, 54 Cumulative distributions, 115–116 Cumulative records, 295

D Darwin, Charles, 6 DAT. See Differential Aptitude Test Data release, 24 DAT PCA. See Differential Aptitude Battery

for Personnel and Career Assessment Decision-making, 27–28, 29 Delirium, 51 Departments of motor vehicles (DMV),

128 Depression. See Beck Depression

Inventory-II Depressive disorders, 49 Depth, of assessment, 61 Derived scores. See also Standard scores

age comparisons, 139 college and graduate school entrance

exam scores, 136–138 definition of, 130 deviation IQ, 133–134 grade equivalents, 139–140 NCE scores, 136, 137 percentile, 130 publisher type scores, 138–139 stanines, 135–136 sten, 136 T-scores, 133 z-scores, 131–133

Developmental maturity assessments, 174 Developmental norms, 139–140 Deviation IQ (DIQ), 133–134 Diagnosis

diagnostic categories, 49–52 dimensional, 48 DSM and, 45–46 DSM-5 and, 46–54 importance of, 44–45 making and reporting, 47–49 medical considerations, 52–53 multiaxial, 47 ordering, 47 practice making, 55 principal, 47 proper, 24 provisional, 48 single-axis, 47

INDEX 341

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Diagnostic and Statistical Manual, Fourth Edition-Text Revision (DSM-IV-TR), 2, 46

cultural considerations, 54 diagnostic categories, 49–52 five-axis diagnostic system, 46 history, 45–46 making and reporting diagnosis, 47–49 other medical considerations, 52–53 psychosocial and environmental

considerations, 53–54 single-axis vs. multiaxial diagnosis, 47 specified disorders and unspecified

disorders, 48–49 Diagnostic and Statistical Manual of Mental

Disorders (DSM), 14, 43 Diagnostic categories, 49–52 Diagnostic tests

definition of, 16 history of, 165–166 KeyMath3, 170–171 PIAT-R/NU, 167–170 WIAT-III, 167, 169 Woodcock-Johnson® III, 170 WRAT4, 166–167

Diaries, 296 Dictionary of Holland Occupational Codes

(Holland), 228 Dictionary of Occupational Titles, 230 Differential Aptitude Battery for Personnel

and Career Assessment (DAT PCA), 238

Differential Aptitude Test (DAT), 237–238 Dimensional diagnosis, 48 DIQ. See Deviation IQ Direct observation, 290–291 Discriminant validity, 92–93, 252 Disorders. See specific types Disruptive disorders, 51 Disruptive observations, 283 Dissociative disorders, 50 DMV. See Departments of motor vehicles Documents. See Personal documents Drawing tests, 272 DSM. See Diagnostic and Statistical Manual

of Mental Disorders DSM-IV-TR. See Diagnostic and Statistical

Manual, Fourth Edition-Text Revision

E Eating and feeding disorders, 50 Educational ability. See also Achievement

tests assessment of, 158–159, 183 cognitive ability testing, 174–182 defining assessment of, 158–159 diagnostic testing, 165–171 helpers’ role, 182–183 readiness testing, 171–174 survey battery testing, 159–165

EHR. See Electronic health records Electronic health records (EHR), 63 Elimination disorders, 50 Emotional Disorder Scale-Revised, 292 Enterprising personality type, 224 Entrance exams, 136–138 Environmental assessment, 290–293

definition of, 290 direct observation, 290–291

instruments for, 291–292 situational assessment, 291 sociometric assessment, 291

Environmental/psychosocial factors, 53–54 Erickson, Milton, 290 Esquirol, Jean, 5 Ethical codes

ACA and APA, 22–25 definition of, 22 professional associations, 22

Ethical decision, 37 Ethical issues

administration conditions, 24 choosing assessment instruments, 22 confidentiality, 23–24 cross-cultural sensitivity, 24 decision-making, 27–29 ethical decision, 37 fair testing, 27 informed consent, 24 privacy, 24 proper diagnosis, 24 releasing data, 24–25 RUST, 26 scores, 25 test administration, 25 test security, 25 users, 23

Eugenics, 11 Event sampling, 283, 284 Evolution, theory of, 6 Existential intelligence, 197 Exner’s scoring system, 270 Experiential subtheory, 198 Experimental design validity, 91–92

F Factor analysis, 92, 233 Fairness. See Cross-cultural fairness Family Education Rights and Privacy Act

(FERPA), 29, 65, 100 Feeding and eating disorders, 50 Feeling word checklists, 288 FERPA. See Family Education Rights and

Privacy Act Fidelity, 28 Finger Tapping Test, of Halstead-Reitan

battery, 212–213 Five-axis diagnostic system, 46 Fixed battery approach to

neuropsychological assessment, 211 Flexible battery approach to

neuropsychological assessment, 213–214

Fluid intelligence, 193–195, 201 Forensic evaluations, 34–35 Forensic psychologists, 34–35 Format of test, 102–103 Freedom of Information Act, 29, 30, 65, 101 Frequency distribution, 112–113 Frequency polygon, 113, 114

G GAF. See Global assessment of functioning Galton, Francis, 6–7, 11 Galton’s Board. See Quincunx Gardner, Howard, 196–197 GATB. See General Aptitude Test Battery Gender dysphoria, 51 General Aptitude Test Battery, 11

General Occupational Themes, 223, 227 Generosity error, 284 Genograms, 296–297 Gesell School Readiness Test, 173–174 g factor, 192 Global assessment of functioning (GAF), 46 Good test, 103–106 Grade equivalents, 139–140 Graduate Record Exam (GRE), 174

definition of, 180 general tests, 180 scoring, 136–138 subject tests, 181

Graphic-type scales, 286 GRE. See Graduate Record Exam Griggs v. Duke Power Company, 36 Grip strength, 6 Guilford, J. L., 193, 201

H Hall, G.S., 7 Halo effect, 284 Halstead Neuropsychological Test Battery

for Older Children, 212 Halsted-Reitan, 211–213 Health Insurance Portability and

Accountability Act (HIPAA), 29–30, 65

Hexagon model of personality types, 223–224

Hierarchical model of intelligence, 193, 201 High stakes testing, 31, 129, 160 HIPAA. See Health Insurance Portability

and Accountability Act Histograms, 113, 114 Holistic process, 35–36 Holland, John, 223–224, 228 Home visits, 290 House-Tree-Person (HTP) test, 272 HTP test. See House-Tree-Person test

I IDEIA. See Individuals with Disabilities

Education Improvement Act IEP. See Individualized Education Plan;

Individualized education plan Impulse control disorders, 51 Individual intelligence testing, 16 Individualized Education Plan (IEP), 32, 44,

163 Individuals with Disabilities Education

Improvement Act (IDEIA), 29, 32, 101, 165

Informal assessment advantages, 282 benefits, 302 classification systems, 287–290 cross-cultural fairness, 301–302 definition of, 282 environmental, 290–293 instruments for, 16–17 observations, 283–284 rating scales, 284–287 records and personal documents, 293–297 reliability, 300 validity, 299 worthiness of, 299–302

Information, anecdotal, 295 Informed consent, 24 Inkblot test. See Rorschach Inkblot test

342 INDEX

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Intelligence models CHC, 199, 201 cognitive development, 195–196, 201 fluid and crystal, 193–195, 201 hierarchical, 193, 201 multifactor/multi-dimensional, 193, 201 multiple, 196–197, 201 nonverbal. See Nonverbal intelligence tests Triarchic theory, 197–198, 201 two-factor, 192–193, 201

Intelligence quotients definition of, 7 development of, 7–8 deviation IQ, 133–134 full-scale, 205 population and, 134 precursors of, 5

Intelligence testing brief history of, 191 definition of, 191–192 eugenics and, 11 helpers’ role in, 214–215 individual, 7–8, 10–11 Kaufman Assessment Battery for

Children, 208 models and, 192–199 neuropsychological assessment and,

comparisons between, 211 precursors of, 5 Stanford-Binet, 202–203 Wechsler scales, 203–208

Interest blanks, 12 Interest inventories

characterization, 222–223 CISS, 233 COPSystem, 228–230 definition of, 16 history, 12 O*NET, 230–232 SDS, 228 Strong Interest Inventory, 223–228

Internal consistency, 96–97 Internet, 105 Interpersonal intelligence, 197 Interquartile range, 119, 120–121 Interrater reliability, 300 Interval scale, 146 Interviews. See also specific types

clinical, 61–63 semi-structured, 63 structured, 62 unstructured, 62

Intrapersonal intelligence, 197 Investigative personality type, 224 Iowa Test of Basic Skills (ITBS), 160,

163–165 Iowa Test of Music Literacy, 241 IRT. See Item response theory ITBS. See Iowa Test of Basic Skills Item response theory (IRT), 98–99, 258

J Jaffee v. Redmond, 30 Journal of Psychological Type, 257 Journals, 105, 296 Jung, Carl, 13, 257

K KABC-II. See Kaufman Assessment Battery

for Children

Kaufman Assessment Battery for Children (KABC-II), 208

KeyMath3, 170–171 Keynotes Music Evaluation Software Kit,

241 KFD. See Kinetic Family Drawing KHTP. See Kinetic House-Tree-Person Test Kindergarten Readiness Test (KRT), 172 Kinetic Family Drawing (KFD), 272 Kinetic House-Tree-Person Test (KHTP),

272 Kraeplin, Emil, 12 KRT. See Kindergarten Readiness Test Kuder–Richardson reliability, 97

L Law School Admission Test (LSAT), 174,

181 Least restrictive environment, 165 Legal issues

legislation, 28–33 privileged communication, 30

LEP. See Limited English proficiency Likert-type scales, 286 Limited English proficiency (LEP), 162 LSAT. See Law School Admission Test

M Marital Satisfaction Inventory, Revised

(MSI-R), 266 MAT. See Miller Analogy Test Mathematical-logical intelligence, 197 MBTI. See Myers-Briggs Type Indicator MCAT. See Medical College Admission Test MCMI. See Millon Clinical Multiaxial

Inventory Mean, 118 Measurement

central tendency, 118–119 scales of, 145–147 standard error of, 141–143 variability, 119–124

Measurement and Evaluation in Counseling and Development (MECD), 34

MECD. See Measurement and Evaluation in Counseling and Development

Mechanical aptitude tests, 239–240 Median, 118–119 Medical College Admission Test (MCAT),

174, 182 Medical conditions, 47

diagnosing, 53 Medical considerations, 52–53 Medication-induced movement disorders,

52 Mental disorders, other, 52 Mental Measurements Yearbook (MMY),

105 Mental status exam

appearance, 68 behavior, 68 defined, 68

Mental test, 7 Metropolitan Achievement Test, 165 Metropolitan Readiness Test (MRT6), 173 Miller Analogy Test (MAT), 174, 181 Millon Clinical Multiaxial Inventory

(MCMI), 252–253 Miner’s interest blank, 12 Minnesota Clerical Assessment Battery, 239

Minnesota Multiphasic Personality Inventory-2 (MMPI-2)

development, 249 function, 249–252 history, 12 profile, 250 validity scales, 251

Minnesota Multiphasic Personality Inventory-2-RF (MMPI-2), 252

MMPI-2. See Minnesota Multiphasic Per- sonality Inventory-2

MMY. See Mental Measurements Yearbook Mode, 119 Moral model, 27 MRT6. See Metropolitan Readiness Test MSI-R. See Marital Satisfaction Inventory,

Revised Multiaxial diagnosis, 47 Multi-dimensional approach, 193, 201 Multifactor approach, 193, 201 Multiple aptitude testing

ASVAB, 221, 233, 234–237 DAT, 237–238 definition of, 16, 233 factor analysis, 233

Multiple intelligence, 196–197 Murray, Henry, 13 Musical aptitude tests, 241 Musical intelligence, 197 Music Aptitude Profile, 241 Myers, Isabel Briggs, 257 Myers-Briggs Type Indicator

development, 257 dichotomies, 258 forms, 258–259 manual reports, 260 sample profile, 259 settings, 258

Myers-Briggs Type Indicator (MBTI), 248, 257–260

N NAEP. See National Assessment of Educa-

tional Progress NASP. See National Association of School

Psychologists National Assessment of Educational Prog-

ress (NAEP), 160–162 National Association of School Psycholo-

gists (NASP), 34 National Report Card, 160, 161. See also

National Assessment of Educational Progress

Naturalist intelligence, 197 NCAA, 31 NCD. See Neurocognitive disorders NCEs. See Normal curve, equivalents NCLB. See No Child Left Behind Act Needed services, identifying of, 177 Negative correlations, 84 Negatively skewed curve, 117 NEO-FFI-3. See NEO Five-Factor

Inventory-3 NEO Five-Factor Inventory-3 (NEO-FFI-3),

262–264 NEO Personality Inventory-Revised (NEO

PI-R), 262 NEO PI-R. See NEO Personality

Inventory-Revised Neurocognitive disorders (NCD), 51

INDEX 343

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Neurodevelopmental disorders, 49 Neuropsychological assessment

applications of results, 211 Boston Process Approach, 213–214 defining, 210–211 description of, 7–8, 210 domains assessed, 211 fixed battery approach, 211 flexible battery approach, 213–214 Halsted-Reitan, 211–213 history of, 210 intelligence testing and, comparisons

between, 211 methods of, 211–214 Wechsler Nonverbal Scale of Ability,

209–210 Neuropsychology, 7–8, 210 No Child Left Behind Act (NCLB),

101, 128 achievement tests and, 159, 160

Nominal scale, 146 Nonmaleficence, 27 Nonverbal intelligence tests

characteristics of, 208–209 Comprehensive Test of Nonverbal

Intelligence, 209 definition of, 208 populations assessed using, 208–209 Universal Intelligence Test, 209

Normal curve defined, 117 equivalents (NCEs), 136, 137, 168 functions, 116–117 SEM and, 141–143 standard deviation, 122–124

Norm referencing, 128–129 Norms, 128–129

developmental, 139–140 Numerical scales, 285–286

O Objective personality assessment, 12 Objective personality testing

BDI-II, 256 Conners 3rd Edition, 264–265 definition of, 16, 248 description, 12 MBTI, 257–260 MCMI, 252–253 MMPI-2, 24–252 MSI-R, 266 NEO PI-R, 262–264 PAI, 253–255 16 PF, 260–262 SASSI, 265–266 TJTA, 266

Observations definition of, 16, 283 direct, 290–291 disruptive, 283 event sampling, 283, 284 time sampling, 283, 284

Obsessive-compulsive and related disorders, 50

Occupational assessment. See Career assessment

Occupational Information Network (O*NET), 230–232

Occupational Scales, 225

OCEAN. See Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism

Odd-even reliability, 96–97 Ogive curve. See Cumulative distributions OLSAT. See Otis-Lennon School Ability

Test O*NET. See Occupational Information

Network Openness, Conscientiousness, Extraversion,

Agreeableness, and Neuroticism (OCEAN), 262

Operations, cognitive, 193 Ordinal scale, 146 Other specified disorder, 49 Otis-Lennon School Ability Test (OLSAT),

162, 174–177

P PAI. See Personality Assessment Inventory Paraphilic disorders, 52 Parsons, Frank, 11 Peabody Individual Achievement Test

(PIAT-R/NU), 167–170 Percentile, 130 Perceptual Reasoning Index (PRI), 205 Performance-based assessment, 297–298 Personal documents, 293–297 Personality Assessment Inventory (PAI),

253–255 Personality assessments, 16 Personality testing

definition of, 12–13, 16 helpers’ role, 274 interest inventories, 12, 222–233. See also

Interest Inventories objective, 12–13, 248–266. See also

Objective personality testing projective, 13, 266–274. See also Projec-

tive tests Personality types, 224 Personal Style Scales, 225 Piaget, Jean, 195, 196, 201 PIAT-R/NU. See Peabody Individual

Achievement Test PL 94–142. See Individuals with Disabilities

Education Improvement Act Portfolio assessment, 298 Positive correlations, 84 Positively skewed curve, 117 Posttraumatic stress disorder (PTSD), 255 Practical facet, 198 Practicality, 81

description, 102 examining, 105 factors in testing, 102–103 informal assessment, 302

Predictive validity, 89–91 PRI. See Perceptual Reasoning Index Principal diagnosis, 47 Privacy, invasion of, 24 Privileged communication, 30 Products, cognitive, 193 Professional associations, 34 Professional issues, 33–38 Projective tests

APT, 268 Bender Gestalt II, 270–272 CAT, 269

definition of, 16, 266–267 development, 12–13 drawing, 272 HTP, 272 KHTP, 272 Rorschach Inkblot, 269–270 sentence completion technique, 273–274 SM-TAT, 269 TAT, 267–269

Provisional diagnosis, 48 Psychological Types (Jung), 257 Psychologists, forensic, 34–35 Psychosocial/environmental factors, 53–54 Psychotic disorders, 49 PTSD. See Posttraumatic stress disorder Publisher resource catalogs, 105 Publisher type scores, 138–139

Q Quincunx, 116

R Range, 119, 120 Rank-order scales, 287 Rating scales, 16, 284–287

graphic-type, 286 Likert-type. See Graphic-type scales numerical, 285–286 rank order, 287 semantic differential, 286–287

Ratio scales, 146–147 Raw scores, 111–112

defined, 111 individuals’ position, 130 norm referencing, 128–129 percentile, 130

Reaction times, 6–7 Readability, 103 Readiness testing

definition of, 16, 158 function, 171 Gesell School Readiness Test, 173–174 KRT, 172 MRT6, 173

Realistic personality type, 224 Records

cumulative, 295 definition of, 17 function, 293

Rehabilitation Act, 29, 33 Reitan-Indiana Aphasia Screening Test, 213 Reitan Indiana Neuropsychological Test,

212 Reitan-Klove Sensory-Perceptual Examina-

tion, 213 Reliability, 81, 142

alternate forms, 96 Cronbach’s coefficient alpha, 97 defined, 94–95 equivalent forms, 96 examining, 105 informal assessment, 300 internal consistency, 96–97 interrater, 300 IRT, 98–99 Kuder–Richardson, 97 odd-even, 96–97 parallel forms, 96

344 INDEX

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

split-half, 96–97 test-retest, 95–96 visual representation, 97–98

Reports. See Assessment reports Response Summary, 225 Responsibilities of Users of Standardized

Tests (RUST), 26 Rhythm Test, of Halstead-Reitan battery,

213 Rorschach, Herman, 13, 269 Rorschach Inkblot test, 13, 269–270 Routing test, 202 RUST. See Responsibilities of Users of

Standardized Tests

S SAI. See School Ability Index Sampling

event, 283, 284 time, 284

SASSI. See Substance Abuse Subtle Screening Inventory

SAT. See Scholastic Aptitude Test SAT10. See Stanford Achievement Test Scale(s). See also specific types

Basic Interest, 225 graphic-type, 286 interval, 146 Likert-type, 286 of measurement, 145–147 MMPI-2, 249–252 nominal, 146 numerical, 285–286 occupational, 225 ordinal, 146 personal style, 225 rank order, 287 rating, 284–287 ratio, 146–147 semantic differential, 286–287 Wechsler, 203–208

Scaled scores, 159 Scatterplot, 85–86 Schizophrenia Spectrum, 49 Scholastic Aptitude Test (SAT)

description of, 179–180 development of, 10–11 interval scale and, 146 scoring of, 136–138 student athletes, 31

School Ability Index (SAI), 175 Score(s). See also specific types

curves, 116–117 derived. See Derived scores frequency distributions, 112–113 interpretation, 25 measures of central tendency, 118–119 measures of variability, 119–124 raw. See Raw scores scaled, 159 SEM, 141–143 standard. See Standard scores standard deviation, 119, 122–124

SDS. See Self-directed search Section 504, 33, 39, 101 Security, 25 SEest. See Standard error of estimate Seguin, Edouard, 5 Selection, test, 103–106 Self-directed search (SDS), 228

SEM. See Standard error of measurement Semantic differential scales, 286–287 Semi-structured interview, 63. see also

Interviews Sentence Completion Technique, 273–274 Severity, 48 Sexual dysfunctions, 51 s factor, 192 SIGI3. See System of Integrated Guidance

and Information-3 Simon, Theophile, 7, 191 Single-axis diagnosis, 47 Situational assessment, 291 Sixteen Personality Factors Questionnaire

(16 PF), 260–262 16 PF. See Sixteen Personality Factors

Questionnaire Skewed curves, 117 Sleep-wake disorders, 50 SM-TAT. See Southern Mississippi’s TAT Social personality type, 224 Sociometric assessment, 291 Somatic symptom and related disorders, 50 Source books, 104–105 Southern Mississippi’s TAT (SM-TAT), 268 Spearman, Charles Edward, 192–193 Spearman’s two-factor approach, 192–193,

201 Special aptitude testing

artistic, 240 clerical, 238–239 definition of, 16, 238 mechanical, 239–240 musical, 241

Specifiers, 47–48 Speech Sounds Perception Test, of Halstead-

Reitan battery, 213 Split-half reliability, 96–97 Standard deviation, 119, 122–124

Graduate Record Exams, 180–181 Standard error of estimate (SEest), 143–144

steps for calculating, 145 Standard error of measurement (SEM),

141–143, 166 determining, 143

Standardized testing, 298 Standards

accreditation, 34 user qualifications, 26

Standard scores. See also Derived scores age comparisons, 139 college and graduate school entrance

exam scores, 136–138 definition of, 130 description, 130–131 deviation IQ, 133–134 grade equivalents, 139–140 NCE scores, 136, 137 publisher type scores, 138–139 stanines, 135–136 sten, 136 T-scores, 133 z-scores, 131–133

Stanford Achievement Test (SAT10) definition of, 162 development, 11 interpretive reports, 162 reliability, 162

Stanford-Binet administering of, 202

development of, 202 interpretive report, 204 measures, 202 norm group, 202–203 organization of, 203

Stanford Revision of the Binet and Simon Scale, 7

Stanines, 135–136 Sten scores, 136 Strong, Edward, 12 Strong Interest Inventory

Basic Interest Scales, 225 definition of, 223 General Occupational Themes, 223 Occupational Scales, 225 Personal Style Scales, 225 Response Summary, 225

Strong Vocational Interest Blank, 12, 223 Structured interview, 62. See also Interviews Subjective Unit of Discomfort (SUD) scale,

285 Substance Abuse Subtle Screening Inventory

(SASSI), 248, 265–266 Substance-related and addictive disorders,

51 Subtypes, 47 Successful intelligence, 197–198, 201 SUD. See Subjective Unit of Discomfort scale Survey battery tests

definition of, 16 description, 159–160 ITBS, 163–165 Metropolitan Achievement Test, 165 NAEP, 160–162 SAT10, 162–163

System of Integrated Guidance and Information-3 (SIGI3), 233

T Tactual Performance Test, of Halstead-

Reitan battery, 212 TAT. See Thematic Apperception Test Taylor-Johnson Temperament Analysis

(TJTA), 266 TBI. See Traumatic brain injury Technical Test Battery (TTB2), 240 Technology and engineering literacy (TEL),

160 TEL. See Technology and engineering

literacy Terman, Lewis, 7–8, 11, 191 Test administration, 25 Test-retest reliability, 95–96 Tests, definition of, 4 Tests in Print, 105 Test worthiness, 83–84 Thematic Apperception Test (TAT)

description of, 267–269 development, 13 standardization issues, 267–268

Thorndike, Edward, 11–12, 285 Thurstone, Louis L., 193, 201 Time factors, 102 Time sampling, 283, 284 TJTA. See Taylor-Johnson Temperament

Analysis Trail Making Test, of Halstead-Reitan

battery, 213 Trauma- and stressor-related disorders, 50 Traumatic brain injury (TBI), 214

INDEX 345

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

Triarchic theory, 197–198, 201 “True reasoning,” 11 T-scores, 133 TTB2. See Technical Test Battery

U UNIT. See Universal Nonverbal Intelligence

Test Universal Nonverbal Intelligence Test

(UNIT), 209 Unspecified disorder, 49 Unstructured interview, 62. See also

Interviews U.S. Department of Labor/Employment and

Training Administration (USDOL/ ETA), 230

U.S. Postal Service (USPS), 238, 239 USDOL/ETA. See U.S. Department of

Labor/Employment and Training Administration

USPS. See U.S. Postal Service

V Validity, 81

concurrent, 89, 264 construct, 91–93, 94 content, 87–89 convergent, 92 criterion-related, 89–91 defined, 87 discriminant, 92–93, 252 examining, 105

experimental design, 91–92 factor analysis, 92 informal assessment, 299 predictive, 89–91 visual representation, 93–94

VCI. See Verbal comprehension index Verbal comprehension index (VCI), 205 Verbal-linguistic intelligence, 197 Vernon, Philip, 193, 201 Visual-spatial intelligence, 197 Vocational assessment, 12

W Wechsler Adult Intelligence Scales-Third

Edition, 205 Wechsler Individual Achievement Test-Third

Edition (WIAT-III), 167, 169 Wechsler Intelligence Scale for Children-

Fourth Edition (WISC-IV), 205 composite score indexes, 207 description, 205 mental measurement review, 104 record form, 207 subtests, 206 validity, 208

Wechsler Nonverbal Scale of Ability (WNV), 209–210

Wechsler Preschool and Primary Scale of Intelligence-Third Edition, 205

Wechsler Scales, 203–208 WHODAS 2.0. See World Health Organiza-

tion Disability Assessment Schedule 2.0

WIAT-III. See Wechsler Individual Achievement Test-Third Edition

Wide Range Achievement Test 4 (WRAT4), 166–167

Wiesen Test of Mechanical Aptitude (WTMA), 239

WISC-IV. See Wechsler Intelligence Scale for Children-Fourth Edition

WJ III. See Woodcock-Johnson® III WMI. See Working Memory Index WNV. See Wechsler Nonverbal Scale of

Ability Woodcock-Johnson® III, 170 Woodworth’s Personal Data Sheet, 12, 13 Word lists, 289 Working Memory Index (WMI), 205 World Health Organization Disability

Assessment Schedule 2.0 (WHODAS 2.0), 47

WRAT4. See Wide Range Achievement Test 4

WTMA. See Wiesen Test of Mechanical Aptitude

Wundt, Wilheim, 6

Y Yerkes, Robert, 8, 11

Z z-scores, 131–133

346 INDEX

Copyright 201 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).

Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.

  • Brief Contents
  • Contents
  • Preface
  • Section I: Understanding the Assessment Process: History, Ethical and Professional Issues, Diagnosis, and the Assessment Report
    • Ch 1: History of Testing and Assessment����������������������������������������������
      • Distinguishing between Testing and Assessment����������������������������������������������������
      • The History of Assessment��������������������������������
      • Questions to Consider When Assessing Individuals�������������������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 2: Ethical, Legal, and Professional Issues in Assessment������������������������������������������������������������������
      • Ethical Issues in Assessment�����������������������������������
      • Legal Issues in Assessment���������������������������������
      • Professional Issues��������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 3: Diagnosis in the Assessment Process������������������������������������������������
      • The Importance of Diagnosis����������������������������������
      • The Diagnostic and Statistical Manual (DSM): A Brief History�������������������������������������������������������������������
      • The DSM-5����������������
      • Summary��������������
      • Chapter Review���������������������
      • Answers to Exercise��������������������������
      • References�����������������
    • Ch 4: The Assessment Report Process: Interviewing the Client and Writing the Report������������������������������������������������������������������������������������������
      • Purpose of the Assessment Report���������������������������������������
      • Gathering Information for the Report: Garbage in, Garbage Out
      • Structured, Unstructured, and Semi-Structured Interviews���������������������������������������������������������������
      • Choosing an Appropriate Assessment Instrument����������������������������������������������������
      • Writing the Report�������������������������
      • Summarizing the Writing of an Assessment Report������������������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
  • Section II : Test Worthiness and Test Statistics
    • Ch 5: Test Worthiness: Validity, Reliability, Cross-Cultural Fairness, and Practicality����������������������������������������������������������������������������������������������
      • Correlation Coefficient������������������������������
      • Coefficient of Determination (Shared Variance)�����������������������������������������������������
      • Validity���������������
      • Reliability������������������
      • Cross-Cultural Fairness������������������������������
      • Practicality�������������������
      • Selecting and Administering a Good Test����������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 6: Statistical Concepts: Making Meaning out of Raw Scores�������������������������������������������������������������������
      • Raw Scores�����������������
      • Frequency Distributions������������������������������
      • Histograms and Frequency Polygons����������������������������������������
      • Cumulative Distributions�������������������������������
      • Normal Curves and Skewed Curves��������������������������������������
      • Measures of Central Tendency�����������������������������������
      • Measures of Variability������������������������������
      • Remembering the Person�����������������������������
      • Summary��������������
      • Chapter Review���������������������
      • Answers to Items 5 through 8�����������������������������������
      • Reference����������������
    • Ch 7: Statistical Concepts: Creating New Scores to Interpret Test Data�����������������������������������������������������������������������������
      • Norm Referencing versus Criterion Referencing����������������������������������������������������
      • Normative Comparisons and Derived Scores�����������������������������������������������
      • Putting It All Together
      • Standard Error of Measurement������������������������������������
      • Standard Error of Estimate���������������������������������
      • Scales of Measurement����������������������������
      • Summary��������������
      • Chapter Review���������������������
      • Answers to Items 4 through 10������������������������������������
      • References�����������������
  • Section III: Commonly Used Assessment Techniques
    • Ch 8: Assessment of Educational Ability: Survey Battery, Diagnostic, Readiness, and Cognitive Ability Tests������������������������������������������������������������������������������������������������������������������
      • Defining Assessment of Educational Ability�������������������������������������������������
      • Survey Battery Achievement Testing�����������������������������������������
      • Diagnostic Testing�������������������������
      • Readiness Testing������������������������
      • Cognitive Ability Tests������������������������������
      • The Role of Helpers in the Assessment of Educational Ability�������������������������������������������������������������������
      • Final Thoughts about the Assessment of Educational Ability�����������������������������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 9: Intellectual and Cognitive Functioning: Intelligence Testing and Neuropsychological Assessment�����������������������������������������������������������������������������������������������������������
      • A Brief History of Intelligence Testing����������������������������������������������
      • Defining Intelligence Testing������������������������������������
      • Models of Intelligence�����������������������������
      • Intelligence Testing���������������������������
      • Neuropsychological Assessment������������������������������������
      • The Role of Helpers in the Assessment of Intellectual and Cognitive Functioning��������������������������������������������������������������������������������������
      • Final Thoughts on the Assessment of Intellectual and Cognitive Functioning���������������������������������������������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 10: Career and Occupational Assessment: Interest Inventories, Multiple Aptitude, and Special Aptitude Tests���������������������������������������������������������������������������������������������������������������������
      • Defining Career and Occupational Assessment��������������������������������������������������
      • Interest Inventories���������������������������
      • Multiple Aptitude Testing��������������������������������
      • Special Aptitude Testing�������������������������������
      • The Role of Helpers in Occupational and Career Assessment����������������������������������������������������������������
      • Final Thoughts Concerning Occupational and Career Assessment�������������������������������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 11: Clinical Assessment: Objective and Projective Personality Tests�����������������������������������������������������������������������������
      • Defining Clinical Assessment�����������������������������������
      • Objective Personality Testing������������������������������������
      • Projective Testing�������������������������
      • The Role of Helpers in Clinical Assessment�������������������������������������������������
      • Final Thoughts on Clinical Assessment��������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
    • Ch 12: Informal Assessment: Observation, Rating Scales, Classification Methods, Environmental Assessment, Records and Personal Documents, and Performance-Based Assessment���������������������������������������������������������������������������������������������������������������������������������������������������������������������������������
      • Defining Informal Assessment�����������������������������������
      • Types of Informal Assessment�����������������������������������
      • Test Worthiness of Informal Assessment���������������������������������������������
      • The Role of Helpers in the Use of Informal Assessment������������������������������������������������������������
      • Final Thoughts on Informal Assessment��������������������������������������������
      • Summary��������������
      • Chapter Review���������������������
      • References�����������������
  • Appendix A: Websites of Codes of Ethics of Select Mental Health Professional Associations
  • Appendix B: Assessment Sections of ACA's and APA's Codes of Ethics
  • Appendix C: Code of Fair Testing Practices in Education
  • Appendix D: Sample Assessment Report
  • Appendix E: Supplemental Statistical Equations
  • Appendix F: Converting Percentiles from z-Scores
  • Glossary
  • Index