PLAGIARISM FREE "A" WORK

profileNeNe1994
Chapter4ValidityReliabilityandBias.pdf

FEW IF ANY SIGNIFICANT ASSESSMENT CONCEPTS ARE COMPLETELY UNRELATED TO

a teacher’s instructional decision making. To prove it, in this chapter

I’m going to trot out three of the most important measurement

ideas—validity, reliability, and assessment bias—and then show you

how each of them bears directly on the instructional choices that

teachers must make.

These three measurement concepts are just about as important as

measurement concepts can get. Although teachers need not be meas-

urement experts, basic assessment literacy is really a professional obli-

gation, and this chapter will unpack some key terminology and clar-

ify what you really need to know.

It is impossible to be assessment literate without possessing at

least a rudimentary understanding of validity, reliability, and assess-

ment bias. They are the foundation for trustworthy inferences. As

teachers, we can’t guarantee that any test-based inference we make is

accurate, but a basic understanding of validity, reliability, and assess-

ment bias increases the odds in our favor. And the more we can trust

our test-based inferences, the better our insight into students and the

better our test-based decisions. Also, in this age of accountability, the

fallout from invalid inferences could be invalid conclusions by

4 Validity, Reliability, and Bias

4 2

ch4.qxd 7/10/2003 9:56 AM Page 42

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 3

administrators, by parents, and by politicians that you and your

school are doing a sub-par job.

Validity At the very apex of all measurement concepts is the notion of valid-

ity. Indeed, the concept of validity almost always finds its way into

any conversation about educational testing. Well, in the next few

paragraphs you’ll learn that there is no such thing as a valid test.

The Validity of Inferences The reason that there’s no such thing as a valid test is quite straight-

forward: It’s not the test itself that can be valid or invalid but, rather,

the inference that’s based on a student’s test performance. Is the score-

based inference that a teacher has made a valid one? Or, in contrast,

has the teacher made a score-based inference that’s invalid? All valid-

ity analysis should center on the test-based inference, rather than on

the test itself. Let’s see how this inference-making process works and,

thereafter, how educators can determine if their own score-based

inferences are valid or not.

You’ll remember from Chapter 1 that educators use educational

tests to secure overt evidence about covert variables, such as a stu-

dent’s ability to spell, read, or perform arithmetic operations. Well,

even though teachers can look at students’ overt test scores, they’re

still obliged to come up with the interpretation about what those test

scores mean. If the interpretation is accurate, we say that the teacher

has arrived at a valid test-based inference. If the interpretation is inac-

curate, then the teacher’s test-based inference is invalid.

You might be wondering, why is this author making such a big

fuss about whether it’s the test or the test-based inference that’s valid?

Well, if a test can be labeled as valid or invalid, then it is surely true

that assessment accuracy resides in the test itself. By this logic, that

test would yield unerringly accurate information no matter how it’s

used or to whom it’s administered. For instance, let’s say a group of

ch4.qxd 7/10/2003 9:56 AM Page 43

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R4 4

Ethiopian test-developers have created a brand new science test for

children. Because the test was developed by Ethiopians in Ethiopia,

the test items and text passages are in Amharic, the official language

of that nation. Now, if that test were administered to Ethiopian chil-

dren, the scores would probably yield valid inferences about these

children’s science skills and knowledge. But if it were administered to

English-speaking children in a Kansas school district, any test-based

inferences about those children’s science skills and knowledge would

be altogether inaccurate. The test itself would yield accurate interpre-

tations in one setting with one group of test-takers, yet inaccurate

interpretations in another setting, with another group of test-takers.

It is the score-based inference that is accurate in Ethiopia, but inaccu-

rate in Kansas. It is not the test.

The same risk of invalid inferences applies in any testing situation

where factors interfere with test-takers’ ability to demonstrate what

they know and can do. For example, consider a 14-year-old gifted

writer recently arrived from El Salvador who cannot express herself in

English; a trigonometry student with cerebral palsy who cannot con-

trol a pencil well enough to draw sine curves; a 6th grader who falls

asleep 5 minutes into the test and only completes 10 of the 50 test

items. Even superlative tests, when used in these circumstances or

other settings where extraneous factors can diminish the accuracy of

score-based inferences, will often lead to mistaken interpretations. If

you recognize that educational tests do not possess some sort of

inherent accurate or inaccurate essence, then you will more likely

realize that assessment validity rests on human judgment about the

inferences derived from students’ test performances. And human

judgment is sometimes faulty.

The task of measurement experts who deal professionally with

assessment validity, then, is to assemble evidence that a particular

score-based inference made in a particular context is valid. With tests

used in large-scale assessments (such as a statewide, high school grad-

uation test), the experts’ quest is to assemble a collection of evidence

ch4.qxd 7/10/2003 9:56 AM Page 44

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 5

that will support the validity of the test-based inferences that are

drawn from the test. Rarely does a single “validity study” supply suf-

ficiently compelling evidence regarding the validity of any test-based

inference. In most cases, to determine whether any given type of

score-based inference is on the mark, it is necessary to consider the

collection of validity evidence in a number of studies.

Three Varieties of Validity Evidence There are three kinds of validity evidence sanctioned by the relevant

professional organizations: (1) criterion-related evidence, (2) construct-

related evidence, and (3) content-related evidence. Each type, usually col-

lected via some sort of investigation or analytic effort, contributes to

the conclusion that a test is yielding data that will support valid infer-

ences. Typically, such investigations are funded by test-makers before

they bring their off-the-shelf product to the market. In other in-

stances, validity studies are required by state authorities prior to test-

selection or prior to “launching” a test they’ve commissioned. We’ll

take a brief peek at each evidence type, paying particular attention to

the one that should most concern a classroom teacher.

Before we do that, though, I need to call your attention to an

important distinction to keep in mind when dealing with educa-

tional tests: the difference between an achievement test and an aptitude

test. An achievement test is intended to measure the skills and know-

ledge that a student currently possesses in a particular subject area.

For instance, the social studies test in the Metropolitan Achievement

Tests is intended to supply an idea of students’ social studies skills

and knowledge. Classroom tests that teachers construct to see how

much their students have learned are other examples of achievement

tests. In contrast, an aptitude test is intended to help predict a stu-

dent’s future performance, typically in a subsequent academic set-

ting. The best examples of this test type are the widely used ACT and

SAT, which are supposed to predict how well high school students

will perform when they get to college.

ch4.qxd 7/10/2003 9:56 AM Page 45

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R4 6

As you’ll read in Chapter 9, there are many times when these two

supposedly different kinds of educational tests actually function in an

almost identical manner. Nevertheless, it’s an important terminology

difference that you need to know, and it comes into particular focus

regarding the first type of validity evidence, which we’re going to

examine right now.

Criterion-related evidence of validity. This kind of validity concerns

whether aptitude tests really predict what they were intended to pre-

dict. Investigators directing a study to collect criterion-related evi-

dence of validity simply administer the aptitude test to high school

students and then follow those students during their college careers

to see if the predictions based on the aptitude test were accurately

predictive of the criterion. In the case of the SAT or ACT, the criterion

would be those students’ college grade-point averages. If the rela-

tionship between the predictive test scores and the college grades is

strong, then this finding constitutes criterion-related evidence sup-

porting the validity of score-based inferences about high school stu-

dents’ probable academic success in college.

Clearly, classroom teachers do not have the time to collect

criterion-referenced evidence of validity, especially about aptitude

tests. That’s a task better left to assessment specialists. Few teachers I

know have ever been mildly tempted to create an aptitude test, much

less collect criterion-related validity evidence regarding that test.

Construct-related evidence of validity. Most measurement specialists

regard construct-related evidence as the most comprehensive form of

validity evidence because, in a sense, it covers all forms of validity

evidence. To explain what construct-related evidence of validity is, I

need to provide a short description of how it is collected.

The first step is identifying some sort of hypothetical construct,

another term for the covert variable sought. As we’ve learned, this

can be something as exotic as a person’s “depression potential” or as

straightforward as a student’s skill in written composition. Next,

based on the understanding of the nature of the hypothetical

ch4.qxd 7/10/2003 9:56 AM Page 46

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 7

construct identified, a test to measure that construct is developed,

and a study is designed to help determine whether the test does, in

fact, measure the construct. The form of such studies can vary sub-

stantially, but they are all intended to supply empirical evidence that

the test is behaving in the way it is supposed to behave.

Here’s what I hope is a helpful example. One kind of construct-

related validity study is called a differential population investigation.

The validity investigator identifies two groups of people who are

almost certain to differ in the degree to which they possess the con-

struct being measured. For instance, let’s say a new test is being inves-

tigated that deals with college students’ mathematical skills. The

investigator locates 25 college math majors and another 25 college

students who haven’t taken a math course since 9th grade. The inves-

tigator then predicts that when all 50 students take the new test, the

math whizzes will blow away the math nonwhizzes. The 50 students

take the test, and that’s just the way things turn out. The investiga-

tor’s prediction has been confirmed by empirical evidence, so this

represents construct-related evidence that, yes, the new test really

does measure college students’ math skills.

There are a host of other approaches to the collection of

construct-related validity evidence, and some are quite exotic. As a

practical matter, though, busy classroom teachers don’t have time to

be carrying out such investigations. Still, I hope it’s clear that because

all attempts to collect validity evidence really do relate to the exis-

tence of an unseen educational variable, most measurement special-

ists believe that it’s accurate to characterize every validity study as

some form of a construct-related validity study.

Content-related evidence of validity. This third kind of validity evi-

dence is also (finally) the kind that classroom teachers might wish to

collect. Briefly put, this form of evidence tries to establish that a test’s

items satisfactorily reflect the content the test is supposed to repre-

sent. And the chief method of carrying out such content-related

validity studies is to rely on human judgment.

ch4.qxd 7/10/2003 9:56 AM Page 47

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R4 8

I’ll start with an example that shows how to collect content-

related evidence of validity for a district-built or school-built test.

Let’s say that the English teachers in a high school are trying to build

a new test to assess students’ mastery of four language arts content

standards—the four standards from their state-approved set that they,

as a group, have decided are the most important. They decide to

assess students’ mastery of each of these content standards through a

32-item test, with 8 items devoted to each of the 4 standards.

The team of English teachers creates a draft version of their new

test and then sets off in pursuit of content-related evidence of the

draft test’s validity. They assemble a review panel of a dozen individ-

uals, half of them English teachers from another school and the other

half parents who majored or minored in English in college. The

review panel meets for a couple of hours on a Saturday morning, and

its members make individual judgments about each of the 32 items,

all of which have been designated as primarily assessing one of the 4

content standards. The judgments required of the review panel all

revolve around the four content standards that the items are suppos-

edly measuring. For example, all panelists might be asked to review

the content standards and then respond to the following question for

each of the 32 test items:

Will a student’s response to this item help teachers determine

whether a student has mastered the designated content standard

that the item is intended to assess?

Yes No Uncertain

After the eight items linked to a particular content standard have

been judged by members of the review panel, the panelists might be

asked to make the following sort of judgment:

Considering the complete set of eight items intended to measure

this designated content standard, indicate how accurately you

ch4.qxd 7/10/2003 9:56 AM Page 48

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 9

believe a teacher will be able to judge a student’s content-standard

mastery based on the students’ response to these eight items.

Very Accurately Somewhat Accurately

Not Too Accurately

It’s true that the process I’ve just described represents a consider-

able effort. That’s okay for state-developed, district-developed, or

school-developed tests . . . but how might individual teachers go

about collecting content-related evidence of validity for their own,

individually created classroom tests? First, I’d suggest doing so only

for very important tests, such as midterm or final exams. Collecting

content-related evidence of validity does take time, and it can be a

ton of trouble, so do it judiciously. Second, you can get by with a

small number of validity judges, perhaps a colleague or two or a par-

ent or two. Remember that the essence of the judgments you’ll be

asking others to make revolves around whether your tests satisfacto-

rily represent the content they are supposed to represent. The more

“representative” your tests, the more likely it is that any of your test-

based inferences will be valid.

Ideally, of course, individual teachers should construct classroom

tests from a content-related validity perspective. What this means, in

a practical manner, is that teachers who develop their own tests

should be continually attentive to the curricular aims each test is

supposed to represent. It’s advisable to keep a list of the relevant

objectives on hand and create specific items to address those objec-

tives. If teachers think seriously about the content-representativeness

of the tests they are building, those tests are more likely to yield valid

score-based inferences.

Figure 4.1 presents a summing-up graphic depiction of how stu-

dents’ mastery of a content standard is measured by a test . . . and the

three kinds of validity evidence that can be collected to support the

validity of a test-based inference about students’ mastery of the con-

tent standard.

ch4.qxd 7/10/2003 9:56 AM Page 49

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R5 0

In contrast to measurement specialists, who are often called on to

collect all three varieties of validity evidence (especially for important

tests), classroom teachers rarely collect any validity evidence at all.

But it is possible for teachers to collect content-related evidence of

validity, the type most useful in the classroom, without a Herculean

effort. When it comes to the most important of your classroom

exams, it might be worthwhile for you to do so.

Validity in the Classroom Now, how does the concept of assessment validity relate to a teacher’s

instructional decisions? Well, it’s pretty clear that if you come up

with incorrect conclusions about your students’ status regarding

important educational variables, including their mastery of particular

content standards, you’ll be more likely to make unsound instruc-

tional decisions. The better fix you get on your students’ status

4 . 1 HOW A TEST-BASED INFERENCE IS MADE AND HOW THE VALIDITY OF THIS INFERENCE IS SUPPORTED

ch4.qxd 7/10/2003 9:56 AM Page 50

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 1

(especially regarding such unseen variables as their cognitive skills),

the more defensible your instructional decisions will be. Valid

assessment-based inferences about students don’t always translate

into brilliant instructional decisions; however, invalid assessment-

based inferences about students almost always lead to dim-witted or,

at best, misguided instructional decisions.

Let me illustrate the sorts of unsound instructional decisions that

teachers can make when they base those decisions on invalid test-

based inferences. Unfortunately, I can readily draw on my own class-

room experience to do so. When I first began teaching in that small,

eastern Oregon high school, one of my assigned classes was senior

English. As I thought about that class in the summer months leading

up to my first salaried teaching position, I concluded that I wanted

those 12th grade English students to leave my class “being good writ-

ers.” In other words, I wanted to make sure that if any of my students

(whom I’d not yet met) went on to college, they’d be able to write

decent essays, reports, and so on.

Well, when I taught that English course, I was pretty pleased with

my students’ progress in being able to write. As the year went by, they

were doing better and better on the exams I employed to judge their

writing skills. My instructional decisions seemed to be sensible,

because the 32 students in my English class were scoring well on my

exams. The only trouble was . . . all of my exams contained only

multiple-choice items about the mechanics of writing. With suitable

shame, I now confess that I never assessed my students’ writing skills by

asking them to write anything! What a cluck.

I never altered my instruction, not even a little, because I used my

students’ scores on multiple-choice tests to arrive at an inference that

“they were learning how to write.” Based on my students’ measured

progress, my instruction was pretty spiffy and didn’t need to be

changed. An invalid test-based inference had led me to an unsound

instructional decision, namely, providing mechanics-only writing

instruction. If I had known back then how to garner content-related

ch4.qxd 7/10/2003 9:56 AM Page 51

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R5 2

evidence of validity for my multiple-choice tests about the mechan-

ics of writing, I’d probably have figured out that my selected-response

exams weren’t truly able to help me reach reasonable inferences

about my students’ constructed-response abilities to write essays,

reports, and so on.

You will surely not be as much of an assessment illiterate as I was

during my first few years of teaching, but you do need to look at the

content of your exams to see if they can contribute to the kinds of

inferences about your students that you truly need to make.

Reliability Reliability is another much-talked-about measurement concept. It’s a

major concern to the developers of large-scale tests, who usually

devote substantial energy to its calculation. A test’s reliability refers to

its consistency. In fact, if you were never to utter the word “reliability”

again, preferring to employ “consistency” instead, you could still live

a rich and satisfying life.

Three Kinds of Reliability As was true with validity, assessment reliability also comes in three

flavors. However, the three ways of thinking about an assessment

instrument’s consistency are really quite strikingly different, and it’s

important that educators understand the distinctions.

Stability reliability. This first kind of reliability concerns the con-

sistency with which a test measures something over time. For

instance, if students took a standardized achievement test on the first

day of a month and, without any intervening instruction regarding

what the test measured, took the same test again at the end of the

month, would students’ scores be about the same?

Alternate-form reliability. The crux of this second kind of reliability

is fairly evident from its name. If there are two supposedly equivalent

forms of a test, do those two forms actually yield student scores that

are pretty similar? If Jamal scored well on Form A, will Jamal also

ch4.qxd 7/10/2003 9:56 AM Page 52

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 3

score well on Form B? Clearly, alternate-form reliability only comes

into play when there are two or more forms of a test that are sup-

posed to be doing the same assessment job.

Internal consistency reliability. This third kind of reliability focuses

on the consistency of the items within a test. Do all of the test’s items

appear to be doing the same kind of measurement job? For internal

consistency reliability to make much sense, of course, a test should be

aimed at a single overall variable—for instance, a student’s reading

comprehension. If a language arts test tried to simultaneously meas-

ure a student’s reading comprehension, spelling ability, and punctu-

ation skills, then it wouldn’t make much sense to see if the test’s

items were functioning in a similar manner. After all, because three

distinct things are being measured, the test’s items shouldn’t be func-

tioning in a homogeneous manner.

Do you see how the three forms of reliability, although all related

to aspects of a test’s consistency, are conceptually dissimilar? From

now on, if you’re ever told that a significant educational test has

“high reliability,” you are permitted to ask, ever so suavely, “What

kind or what kinds of reliability are you talking about?” (Most often,

you’ll find that the test developers have computed some type of

internal consistency reliability, because it’s possible to calculate such

reliability coefficients on the basis of only one test administration.

Both of the other two kinds of reliability require at least two test

administrations.) It’s also important to note that reliability in one

area does not ensure reliability in another. Don’t assume, for exam-

ple, that a test with high internal consistency reliability will auto-

matically yield stable scores over time. It just isn’t so.

Reliability in the Classroom So, what sorts of reliability evidence should classroom teachers col-

lect for their own tests? My answer may surprise you. I don’t think

teachers need to assemble any kind of reliability evidence. There’s just

too little payoff for the effort involved.

ch4.qxd 7/10/2003 9:56 AM Page 53

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R5 4

I do believe that classroom teachers ought to understand test reli-

ability, especially that it comes in three quite distinctive forms and

that one type of reliability definitely isn’t equivalent to another. And

this is the point at which even reliability has a relationship to instruc-

tion, although I must confess it’s not a strong relationship. Suppose

your students are taking important external tests—say, statewide

achievement tests assessing their mastery of state-sanctioned content

standards. Well, if the tests are truly important, why not find out

something about their technical characteristics? Is there evidence of

assessment reliability presented? If so, what kind or what kinds of

reliability evidence?

If the statewide test’s reliability evidence is skimpy, then you

should be uneasy about the test’s quality. Here’s why: An unreliable

test will rarely yield scores from which valid inferences can be drawn.

I’ll illustrate this point with an example of stability evidence of

reliability. Suppose your students took an important exam on

Tuesday morning, and there was a fire in the school counselor’s

office on Tuesday afternoon. Your student’s exam papers went up in

smoke. A week later, you re-administer the big test, only to learn

soon after that the original exam papers were saved from the flames

by an intrepid school custodian. When you compare the two sets of

scores on the very same exam administered a week apart, you are

surprised to find that your students’ scores seem to bounce all over

the place. For example, Billy scored high on the first test, yet scored

low on the second. Tristan scored low on the first test, but soared on

the second.

Based on the variable test scores, have these students mastered

the material or haven’t they? Would it be appropriate or inappropri-

ate to send Billy on to more challenging work? Does Tristan need

extra guided practice, or is she ready for independent application? If

a test is unreliable—inconsistent—how can it contribute to accurate

score-based inferences and sound instructional decisions? Answer: It

can’t.

ch4.qxd 7/10/2003 9:56 AM Page 54

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 5

Assessment Bias Bias is something that everyone understands to be a bad thing. Bias

beclouds one’s judgment. Conversely, the absence of bias is a good

thing; it permits better judgment. Assessment bias is one species of

bias and, as you might have already guessed, it’s something educators

need to identify and eliminate, both for moral reasons and in the

interest of promoting better, more defensible instructional decisions.

Let’s see how to go about doing that for large-scale tests and for the

tests that teachers cook up for their own students.

The Nature of Assessment Bias Assessment bias occurs whenever test items offend or unfairly penal-

ize students for reasons related to students’ personal characteristics,

such as their race, gender, ethnicity, religion, or socioeconomic sta-

tus. Notice that there are two elements in this definition of assess-

ment bias. A test can be biased if it offends students or if it unfairly

penalizes students because of students’ personal characteristics.

An example of a test item that would offend students might be

one in which a person, clearly identifiable as a member of a particu-

lar ethnic group, is described in the item itself as displaying patently

unintelligent behavior. Students from that same ethnic group might

(with good reason) be upset at the test item’s implication that persons

of their ethnicity are not all that bright. And that sort of upset often

leads the offended students to perform less well than would other-

wise be the case. Another example of offensive test items would arise

if females were always depicted in a test’s items as being employed in

low-level, undemanding jobs whereas males were always depicted as

holding high-paying, executive positions. Girls taking a test com-

posed of such sexist items might (again, quite properly) be annoyed

and, hence, might perform less well than they would have otherwise.

Turning to unfair penalization, think about a series of mathemat-

ical items set in a context of how to score a football game. If girls par-

ticipate less frequently in football and watch, in general, fewer

ch4.qxd 7/10/2003 9:56 AM Page 55

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R5 6

football games on television, then girls are likely to have more diffi-

culty in answering math items that revolve around how many points

are awarded when a team “scores a safety” or “kicks extra points

rather than running or passing for extra points.” These sorts of math-

ematics items simply shimmer with gender bias.

Not all penalties are unfair, of course. If students don’t study

properly and then perform poorly on a teacher’s test, such penalties

are richly deserved. What’s more, the inference a teacher would

make, based on these low scores, would be a valid one: These students

haven’t mastered the material! But if students’ personal characteris-

tics, such as their socioeconomic status (SES), are a determining fac-

tor in weaker test performance, then assessment bias has clearly

raised its unattractive head. It distorts the accuracy of students’ test

performances and invariably leads to invalid inferences about stu-

dents’ status and to unsound instructional decisions about how best

to teach those students.

Bias Detection in Large-Scale Assessments There was a time, not too long ago, when the creators of large-scale

tests (such as nationally standardized achievement tests) didn’t do a

particularly respectable job of identifying and excising biased items

from their tests. I taught educational measurement courses in the

UCLA Graduate School of Education for many years, and in the late

1970s, one of the nationally standardized achievement tests I rou-

tinely had my students critique employed a bias-detection procedure

that would be considered laughable today. Even back then, it made

me smirk a bit.

Here’s how the test-developers’ bias-detection procedure worked.

A three-person bias review committee was asked to review all of the

items of a test under development. Two of the reviewers “represented”

minority groups. If all three reviewers considered an item to be biased,

the item was eliminated from the test. Otherwise, that item stayed on

the test. If the two minority-representing reviewers considered an

ch4.qxd 7/10/2003 9:56 AM Page 56

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 7

item biased beyond belief, but the third reviewer disagreed, the item

stayed. Today, such a superficial bias-review process would be recog-

nized as absurd.

Developers of large-scale tests now employ far more rigorous bias-

detection procedures, especially with respect to bias based on race

and gender. A typical approach these days calls for the creation of a

bias-review committee, usually of 15–25 members, almost all of

whom are themselves members of minority groups. Bias-review com-

mittees are typically given ample training and practice in how best to

render their item-bias judgments. For each item that might end up

being included in the test, every member of the bias-review commit-

tee would be asked to respond to the following question:

Might this item offend or unfairly penalize students because of

such personal characteristics as gender, ethnicity, religion, or

socioeconomic status?

Yes No Uncertain

Note that the above question asks reviewers to judge whether a stu-

dent might be penalized, not whether the item would (for absolutely

certain) penalize a student. This kind of phrasing makes it clear that

the test’s developers are going the extra mile to eliminate any items

that might possibly offend or unfairly penalize students because of per-

sonal characteristics. If the review question asked whether an item

“would” offend or unfairly penalize students, you can bet that fewer

items would be reported as biased.

If a certain percentage of bias-reviewers believe the item might be

biased, it is eliminated from those to be used in the test. Clearly, the

determination of what proportion of reviewers is needed to delete an

item on potential-bias grounds represents an important issue. With

the three-person bias-review committee my 1970s students read

about at UCLA, the percentage of reviewers necessary for an item’s

elimination was 100 percent. In recent years I have taken part in

ch4.qxd 7/10/2003 9:56 AM Page 57

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R5 8

bias-review procedures where an item would eliminated if more than

five percent of the reviewers rated the item as biased. Times have def-

initely changed.

My experiences suggest that the developers of most of our current

large-scale tests have been attentive to the potential presence of

assessment bias based on students’ gender or membership in a minor-

ity racial or ethnic group. (I still think that today’s major educational

tests feature far too many items that are biased on SES grounds. We’ll

discuss this issue in Chapter 9.) There may be a few items that have

slipped by the rigorous item-review process but, in general, those test

developers get an A for effort. In many instances, such assiduous

attention to gender and minority-group bias eradication stems

directly from the fear that any adverse results of their tests might be

challenged in court. The threat of litigation often proves to be a

potent stimulant!

Assessment Bias in the Classroom Unfortunately, there’s much more bias present in teachers’ classroom

tests than most teachers imagine. The reason is not that today’s class-

room teachers deliberately set out to offend or unfairly penalize cer-

tain of their students; it’s just that too few teachers have systemati-

cally attended to this issue.

The key to unbiasing tests is a simple matter of serious, item-by-

item scrutiny. The same kind of item-review question that bias-

reviewers typically use in their appraisal of large-scale assessments will

work for classroom tests, too. A teacher who is instructing students

from racial/ethnic groups other than the teacher’s own racial/ethnic

group might be wise to ask a colleague (or a parent) from those

racial/ethnic groups to serve as a one-person bias review committee.

This can be very illuminating. Happily, most biased items can be

repaired with only modest effort. Those that can’t should be tossed.

This whole bias-detection business is about being fair to all

students and assessing them in such a way that they are accurately

ch4.qxd 7/10/2003 9:56 AM Page 58

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 9

measured. That’s the only way a teacher’s test-based inferences will be

valid. And, of course, valid inferences about students serve as the

foundation for defensible instructional decisions. Invalid inferences

don’t.

Recommended Resources

American Educational Research Association. (1999). Standards for educational and psychological testing. Washington, DC: Author.

McNeil, L. M. (2000, June). Creating new inequalities: Contradictions of reform. Phi Delta Kappan, 81(10), 728–734.

Popham, W. J. (2000). Modern educational measurement: Practical guidelines for educational leaders (3rd ed.). Boston: Allyn & Bacon.

Popham, W. J. (Program Consultant). (2000). Norm- and criterion-referenced testing: What assessment-literate educators should know [Videotape]. Los Angeles: IOX Assessment Associates.

Popham, W. J. (Program Consultant). (2000). Standardized achievement tests: How to tell what they measure [Videotape]. Los Angeles: IOX Assessment Associates.

INSTRUCTIONALLY FOCUSED TESTING TIPS

• Recognize that validity refers to a test-based inference, not to

the test itself.

• Understand that there are three kinds of validity evidence, all of

which can contribute to the confidence teachers have in the

accuracy of a test-based inference about students.

• Assemble content-related evidence of validity for your most

important classroom tests.

• Know that there are three related, but meaningfully different

kinds of reliability evidence that can be collected for educational

tests.

• Give serious attention to the detection and elimination of

assessment bias in classroom tests.

ch4.qxd 7/10/2003 9:56 AM Page 59

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .