PLAGIARISM FREE "A" WORK
FEW IF ANY SIGNIFICANT ASSESSMENT CONCEPTS ARE COMPLETELY UNRELATED TO
a teacher’s instructional decision making. To prove it, in this chapter
I’m going to trot out three of the most important measurement
ideas—validity, reliability, and assessment bias—and then show you
how each of them bears directly on the instructional choices that
teachers must make.
These three measurement concepts are just about as important as
measurement concepts can get. Although teachers need not be meas-
urement experts, basic assessment literacy is really a professional obli-
gation, and this chapter will unpack some key terminology and clar-
ify what you really need to know.
It is impossible to be assessment literate without possessing at
least a rudimentary understanding of validity, reliability, and assess-
ment bias. They are the foundation for trustworthy inferences. As
teachers, we can’t guarantee that any test-based inference we make is
accurate, but a basic understanding of validity, reliability, and assess-
ment bias increases the odds in our favor. And the more we can trust
our test-based inferences, the better our insight into students and the
better our test-based decisions. Also, in this age of accountability, the
fallout from invalid inferences could be invalid conclusions by
4 Validity, Reliability, and Bias
4 2
ch4.qxd 7/10/2003 9:56 AM Page 42
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 3
administrators, by parents, and by politicians that you and your
school are doing a sub-par job.
Validity At the very apex of all measurement concepts is the notion of valid-
ity. Indeed, the concept of validity almost always finds its way into
any conversation about educational testing. Well, in the next few
paragraphs you’ll learn that there is no such thing as a valid test.
The Validity of Inferences The reason that there’s no such thing as a valid test is quite straight-
forward: It’s not the test itself that can be valid or invalid but, rather,
the inference that’s based on a student’s test performance. Is the score-
based inference that a teacher has made a valid one? Or, in contrast,
has the teacher made a score-based inference that’s invalid? All valid-
ity analysis should center on the test-based inference, rather than on
the test itself. Let’s see how this inference-making process works and,
thereafter, how educators can determine if their own score-based
inferences are valid or not.
You’ll remember from Chapter 1 that educators use educational
tests to secure overt evidence about covert variables, such as a stu-
dent’s ability to spell, read, or perform arithmetic operations. Well,
even though teachers can look at students’ overt test scores, they’re
still obliged to come up with the interpretation about what those test
scores mean. If the interpretation is accurate, we say that the teacher
has arrived at a valid test-based inference. If the interpretation is inac-
curate, then the teacher’s test-based inference is invalid.
You might be wondering, why is this author making such a big
fuss about whether it’s the test or the test-based inference that’s valid?
Well, if a test can be labeled as valid or invalid, then it is surely true
that assessment accuracy resides in the test itself. By this logic, that
test would yield unerringly accurate information no matter how it’s
used or to whom it’s administered. For instance, let’s say a group of
ch4.qxd 7/10/2003 9:56 AM Page 43
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R4 4
Ethiopian test-developers have created a brand new science test for
children. Because the test was developed by Ethiopians in Ethiopia,
the test items and text passages are in Amharic, the official language
of that nation. Now, if that test were administered to Ethiopian chil-
dren, the scores would probably yield valid inferences about these
children’s science skills and knowledge. But if it were administered to
English-speaking children in a Kansas school district, any test-based
inferences about those children’s science skills and knowledge would
be altogether inaccurate. The test itself would yield accurate interpre-
tations in one setting with one group of test-takers, yet inaccurate
interpretations in another setting, with another group of test-takers.
It is the score-based inference that is accurate in Ethiopia, but inaccu-
rate in Kansas. It is not the test.
The same risk of invalid inferences applies in any testing situation
where factors interfere with test-takers’ ability to demonstrate what
they know and can do. For example, consider a 14-year-old gifted
writer recently arrived from El Salvador who cannot express herself in
English; a trigonometry student with cerebral palsy who cannot con-
trol a pencil well enough to draw sine curves; a 6th grader who falls
asleep 5 minutes into the test and only completes 10 of the 50 test
items. Even superlative tests, when used in these circumstances or
other settings where extraneous factors can diminish the accuracy of
score-based inferences, will often lead to mistaken interpretations. If
you recognize that educational tests do not possess some sort of
inherent accurate or inaccurate essence, then you will more likely
realize that assessment validity rests on human judgment about the
inferences derived from students’ test performances. And human
judgment is sometimes faulty.
The task of measurement experts who deal professionally with
assessment validity, then, is to assemble evidence that a particular
score-based inference made in a particular context is valid. With tests
used in large-scale assessments (such as a statewide, high school grad-
uation test), the experts’ quest is to assemble a collection of evidence
ch4.qxd 7/10/2003 9:56 AM Page 44
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 5
that will support the validity of the test-based inferences that are
drawn from the test. Rarely does a single “validity study” supply suf-
ficiently compelling evidence regarding the validity of any test-based
inference. In most cases, to determine whether any given type of
score-based inference is on the mark, it is necessary to consider the
collection of validity evidence in a number of studies.
Three Varieties of Validity Evidence There are three kinds of validity evidence sanctioned by the relevant
professional organizations: (1) criterion-related evidence, (2) construct-
related evidence, and (3) content-related evidence. Each type, usually col-
lected via some sort of investigation or analytic effort, contributes to
the conclusion that a test is yielding data that will support valid infer-
ences. Typically, such investigations are funded by test-makers before
they bring their off-the-shelf product to the market. In other in-
stances, validity studies are required by state authorities prior to test-
selection or prior to “launching” a test they’ve commissioned. We’ll
take a brief peek at each evidence type, paying particular attention to
the one that should most concern a classroom teacher.
Before we do that, though, I need to call your attention to an
important distinction to keep in mind when dealing with educa-
tional tests: the difference between an achievement test and an aptitude
test. An achievement test is intended to measure the skills and know-
ledge that a student currently possesses in a particular subject area.
For instance, the social studies test in the Metropolitan Achievement
Tests is intended to supply an idea of students’ social studies skills
and knowledge. Classroom tests that teachers construct to see how
much their students have learned are other examples of achievement
tests. In contrast, an aptitude test is intended to help predict a stu-
dent’s future performance, typically in a subsequent academic set-
ting. The best examples of this test type are the widely used ACT and
SAT, which are supposed to predict how well high school students
will perform when they get to college.
ch4.qxd 7/10/2003 9:56 AM Page 45
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R4 6
As you’ll read in Chapter 9, there are many times when these two
supposedly different kinds of educational tests actually function in an
almost identical manner. Nevertheless, it’s an important terminology
difference that you need to know, and it comes into particular focus
regarding the first type of validity evidence, which we’re going to
examine right now.
Criterion-related evidence of validity. This kind of validity concerns
whether aptitude tests really predict what they were intended to pre-
dict. Investigators directing a study to collect criterion-related evi-
dence of validity simply administer the aptitude test to high school
students and then follow those students during their college careers
to see if the predictions based on the aptitude test were accurately
predictive of the criterion. In the case of the SAT or ACT, the criterion
would be those students’ college grade-point averages. If the rela-
tionship between the predictive test scores and the college grades is
strong, then this finding constitutes criterion-related evidence sup-
porting the validity of score-based inferences about high school stu-
dents’ probable academic success in college.
Clearly, classroom teachers do not have the time to collect
criterion-referenced evidence of validity, especially about aptitude
tests. That’s a task better left to assessment specialists. Few teachers I
know have ever been mildly tempted to create an aptitude test, much
less collect criterion-related validity evidence regarding that test.
Construct-related evidence of validity. Most measurement specialists
regard construct-related evidence as the most comprehensive form of
validity evidence because, in a sense, it covers all forms of validity
evidence. To explain what construct-related evidence of validity is, I
need to provide a short description of how it is collected.
The first step is identifying some sort of hypothetical construct,
another term for the covert variable sought. As we’ve learned, this
can be something as exotic as a person’s “depression potential” or as
straightforward as a student’s skill in written composition. Next,
based on the understanding of the nature of the hypothetical
ch4.qxd 7/10/2003 9:56 AM Page 46
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 7
construct identified, a test to measure that construct is developed,
and a study is designed to help determine whether the test does, in
fact, measure the construct. The form of such studies can vary sub-
stantially, but they are all intended to supply empirical evidence that
the test is behaving in the way it is supposed to behave.
Here’s what I hope is a helpful example. One kind of construct-
related validity study is called a differential population investigation.
The validity investigator identifies two groups of people who are
almost certain to differ in the degree to which they possess the con-
struct being measured. For instance, let’s say a new test is being inves-
tigated that deals with college students’ mathematical skills. The
investigator locates 25 college math majors and another 25 college
students who haven’t taken a math course since 9th grade. The inves-
tigator then predicts that when all 50 students take the new test, the
math whizzes will blow away the math nonwhizzes. The 50 students
take the test, and that’s just the way things turn out. The investiga-
tor’s prediction has been confirmed by empirical evidence, so this
represents construct-related evidence that, yes, the new test really
does measure college students’ math skills.
There are a host of other approaches to the collection of
construct-related validity evidence, and some are quite exotic. As a
practical matter, though, busy classroom teachers don’t have time to
be carrying out such investigations. Still, I hope it’s clear that because
all attempts to collect validity evidence really do relate to the exis-
tence of an unseen educational variable, most measurement special-
ists believe that it’s accurate to characterize every validity study as
some form of a construct-related validity study.
Content-related evidence of validity. This third kind of validity evi-
dence is also (finally) the kind that classroom teachers might wish to
collect. Briefly put, this form of evidence tries to establish that a test’s
items satisfactorily reflect the content the test is supposed to repre-
sent. And the chief method of carrying out such content-related
validity studies is to rely on human judgment.
ch4.qxd 7/10/2003 9:56 AM Page 47
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R4 8
I’ll start with an example that shows how to collect content-
related evidence of validity for a district-built or school-built test.
Let’s say that the English teachers in a high school are trying to build
a new test to assess students’ mastery of four language arts content
standards—the four standards from their state-approved set that they,
as a group, have decided are the most important. They decide to
assess students’ mastery of each of these content standards through a
32-item test, with 8 items devoted to each of the 4 standards.
The team of English teachers creates a draft version of their new
test and then sets off in pursuit of content-related evidence of the
draft test’s validity. They assemble a review panel of a dozen individ-
uals, half of them English teachers from another school and the other
half parents who majored or minored in English in college. The
review panel meets for a couple of hours on a Saturday morning, and
its members make individual judgments about each of the 32 items,
all of which have been designated as primarily assessing one of the 4
content standards. The judgments required of the review panel all
revolve around the four content standards that the items are suppos-
edly measuring. For example, all panelists might be asked to review
the content standards and then respond to the following question for
each of the 32 test items:
Will a student’s response to this item help teachers determine
whether a student has mastered the designated content standard
that the item is intended to assess?
Yes No Uncertain
After the eight items linked to a particular content standard have
been judged by members of the review panel, the panelists might be
asked to make the following sort of judgment:
Considering the complete set of eight items intended to measure
this designated content standard, indicate how accurately you
ch4.qxd 7/10/2003 9:56 AM Page 48
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 4 9
believe a teacher will be able to judge a student’s content-standard
mastery based on the students’ response to these eight items.
Very Accurately Somewhat Accurately
Not Too Accurately
It’s true that the process I’ve just described represents a consider-
able effort. That’s okay for state-developed, district-developed, or
school-developed tests . . . but how might individual teachers go
about collecting content-related evidence of validity for their own,
individually created classroom tests? First, I’d suggest doing so only
for very important tests, such as midterm or final exams. Collecting
content-related evidence of validity does take time, and it can be a
ton of trouble, so do it judiciously. Second, you can get by with a
small number of validity judges, perhaps a colleague or two or a par-
ent or two. Remember that the essence of the judgments you’ll be
asking others to make revolves around whether your tests satisfacto-
rily represent the content they are supposed to represent. The more
“representative” your tests, the more likely it is that any of your test-
based inferences will be valid.
Ideally, of course, individual teachers should construct classroom
tests from a content-related validity perspective. What this means, in
a practical manner, is that teachers who develop their own tests
should be continually attentive to the curricular aims each test is
supposed to represent. It’s advisable to keep a list of the relevant
objectives on hand and create specific items to address those objec-
tives. If teachers think seriously about the content-representativeness
of the tests they are building, those tests are more likely to yield valid
score-based inferences.
Figure 4.1 presents a summing-up graphic depiction of how stu-
dents’ mastery of a content standard is measured by a test . . . and the
three kinds of validity evidence that can be collected to support the
validity of a test-based inference about students’ mastery of the con-
tent standard.
ch4.qxd 7/10/2003 9:56 AM Page 49
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R5 0
In contrast to measurement specialists, who are often called on to
collect all three varieties of validity evidence (especially for important
tests), classroom teachers rarely collect any validity evidence at all.
But it is possible for teachers to collect content-related evidence of
validity, the type most useful in the classroom, without a Herculean
effort. When it comes to the most important of your classroom
exams, it might be worthwhile for you to do so.
Validity in the Classroom Now, how does the concept of assessment validity relate to a teacher’s
instructional decisions? Well, it’s pretty clear that if you come up
with incorrect conclusions about your students’ status regarding
important educational variables, including their mastery of particular
content standards, you’ll be more likely to make unsound instruc-
tional decisions. The better fix you get on your students’ status
4 . 1 HOW A TEST-BASED INFERENCE IS MADE AND HOW THE VALIDITY OF THIS INFERENCE IS SUPPORTED
ch4.qxd 7/10/2003 9:56 AM Page 50
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 1
(especially regarding such unseen variables as their cognitive skills),
the more defensible your instructional decisions will be. Valid
assessment-based inferences about students don’t always translate
into brilliant instructional decisions; however, invalid assessment-
based inferences about students almost always lead to dim-witted or,
at best, misguided instructional decisions.
Let me illustrate the sorts of unsound instructional decisions that
teachers can make when they base those decisions on invalid test-
based inferences. Unfortunately, I can readily draw on my own class-
room experience to do so. When I first began teaching in that small,
eastern Oregon high school, one of my assigned classes was senior
English. As I thought about that class in the summer months leading
up to my first salaried teaching position, I concluded that I wanted
those 12th grade English students to leave my class “being good writ-
ers.” In other words, I wanted to make sure that if any of my students
(whom I’d not yet met) went on to college, they’d be able to write
decent essays, reports, and so on.
Well, when I taught that English course, I was pretty pleased with
my students’ progress in being able to write. As the year went by, they
were doing better and better on the exams I employed to judge their
writing skills. My instructional decisions seemed to be sensible,
because the 32 students in my English class were scoring well on my
exams. The only trouble was . . . all of my exams contained only
multiple-choice items about the mechanics of writing. With suitable
shame, I now confess that I never assessed my students’ writing skills by
asking them to write anything! What a cluck.
I never altered my instruction, not even a little, because I used my
students’ scores on multiple-choice tests to arrive at an inference that
“they were learning how to write.” Based on my students’ measured
progress, my instruction was pretty spiffy and didn’t need to be
changed. An invalid test-based inference had led me to an unsound
instructional decision, namely, providing mechanics-only writing
instruction. If I had known back then how to garner content-related
ch4.qxd 7/10/2003 9:56 AM Page 51
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R5 2
evidence of validity for my multiple-choice tests about the mechan-
ics of writing, I’d probably have figured out that my selected-response
exams weren’t truly able to help me reach reasonable inferences
about my students’ constructed-response abilities to write essays,
reports, and so on.
You will surely not be as much of an assessment illiterate as I was
during my first few years of teaching, but you do need to look at the
content of your exams to see if they can contribute to the kinds of
inferences about your students that you truly need to make.
Reliability Reliability is another much-talked-about measurement concept. It’s a
major concern to the developers of large-scale tests, who usually
devote substantial energy to its calculation. A test’s reliability refers to
its consistency. In fact, if you were never to utter the word “reliability”
again, preferring to employ “consistency” instead, you could still live
a rich and satisfying life.
Three Kinds of Reliability As was true with validity, assessment reliability also comes in three
flavors. However, the three ways of thinking about an assessment
instrument’s consistency are really quite strikingly different, and it’s
important that educators understand the distinctions.
Stability reliability. This first kind of reliability concerns the con-
sistency with which a test measures something over time. For
instance, if students took a standardized achievement test on the first
day of a month and, without any intervening instruction regarding
what the test measured, took the same test again at the end of the
month, would students’ scores be about the same?
Alternate-form reliability. The crux of this second kind of reliability
is fairly evident from its name. If there are two supposedly equivalent
forms of a test, do those two forms actually yield student scores that
are pretty similar? If Jamal scored well on Form A, will Jamal also
ch4.qxd 7/10/2003 9:56 AM Page 52
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 3
score well on Form B? Clearly, alternate-form reliability only comes
into play when there are two or more forms of a test that are sup-
posed to be doing the same assessment job.
Internal consistency reliability. This third kind of reliability focuses
on the consistency of the items within a test. Do all of the test’s items
appear to be doing the same kind of measurement job? For internal
consistency reliability to make much sense, of course, a test should be
aimed at a single overall variable—for instance, a student’s reading
comprehension. If a language arts test tried to simultaneously meas-
ure a student’s reading comprehension, spelling ability, and punctu-
ation skills, then it wouldn’t make much sense to see if the test’s
items were functioning in a similar manner. After all, because three
distinct things are being measured, the test’s items shouldn’t be func-
tioning in a homogeneous manner.
Do you see how the three forms of reliability, although all related
to aspects of a test’s consistency, are conceptually dissimilar? From
now on, if you’re ever told that a significant educational test has
“high reliability,” you are permitted to ask, ever so suavely, “What
kind or what kinds of reliability are you talking about?” (Most often,
you’ll find that the test developers have computed some type of
internal consistency reliability, because it’s possible to calculate such
reliability coefficients on the basis of only one test administration.
Both of the other two kinds of reliability require at least two test
administrations.) It’s also important to note that reliability in one
area does not ensure reliability in another. Don’t assume, for exam-
ple, that a test with high internal consistency reliability will auto-
matically yield stable scores over time. It just isn’t so.
Reliability in the Classroom So, what sorts of reliability evidence should classroom teachers col-
lect for their own tests? My answer may surprise you. I don’t think
teachers need to assemble any kind of reliability evidence. There’s just
too little payoff for the effort involved.
ch4.qxd 7/10/2003 9:56 AM Page 53
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R5 4
I do believe that classroom teachers ought to understand test reli-
ability, especially that it comes in three quite distinctive forms and
that one type of reliability definitely isn’t equivalent to another. And
this is the point at which even reliability has a relationship to instruc-
tion, although I must confess it’s not a strong relationship. Suppose
your students are taking important external tests—say, statewide
achievement tests assessing their mastery of state-sanctioned content
standards. Well, if the tests are truly important, why not find out
something about their technical characteristics? Is there evidence of
assessment reliability presented? If so, what kind or what kinds of
reliability evidence?
If the statewide test’s reliability evidence is skimpy, then you
should be uneasy about the test’s quality. Here’s why: An unreliable
test will rarely yield scores from which valid inferences can be drawn.
I’ll illustrate this point with an example of stability evidence of
reliability. Suppose your students took an important exam on
Tuesday morning, and there was a fire in the school counselor’s
office on Tuesday afternoon. Your student’s exam papers went up in
smoke. A week later, you re-administer the big test, only to learn
soon after that the original exam papers were saved from the flames
by an intrepid school custodian. When you compare the two sets of
scores on the very same exam administered a week apart, you are
surprised to find that your students’ scores seem to bounce all over
the place. For example, Billy scored high on the first test, yet scored
low on the second. Tristan scored low on the first test, but soared on
the second.
Based on the variable test scores, have these students mastered
the material or haven’t they? Would it be appropriate or inappropri-
ate to send Billy on to more challenging work? Does Tristan need
extra guided practice, or is she ready for independent application? If
a test is unreliable—inconsistent—how can it contribute to accurate
score-based inferences and sound instructional decisions? Answer: It
can’t.
ch4.qxd 7/10/2003 9:56 AM Page 54
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 5
Assessment Bias Bias is something that everyone understands to be a bad thing. Bias
beclouds one’s judgment. Conversely, the absence of bias is a good
thing; it permits better judgment. Assessment bias is one species of
bias and, as you might have already guessed, it’s something educators
need to identify and eliminate, both for moral reasons and in the
interest of promoting better, more defensible instructional decisions.
Let’s see how to go about doing that for large-scale tests and for the
tests that teachers cook up for their own students.
The Nature of Assessment Bias Assessment bias occurs whenever test items offend or unfairly penal-
ize students for reasons related to students’ personal characteristics,
such as their race, gender, ethnicity, religion, or socioeconomic sta-
tus. Notice that there are two elements in this definition of assess-
ment bias. A test can be biased if it offends students or if it unfairly
penalizes students because of students’ personal characteristics.
An example of a test item that would offend students might be
one in which a person, clearly identifiable as a member of a particu-
lar ethnic group, is described in the item itself as displaying patently
unintelligent behavior. Students from that same ethnic group might
(with good reason) be upset at the test item’s implication that persons
of their ethnicity are not all that bright. And that sort of upset often
leads the offended students to perform less well than would other-
wise be the case. Another example of offensive test items would arise
if females were always depicted in a test’s items as being employed in
low-level, undemanding jobs whereas males were always depicted as
holding high-paying, executive positions. Girls taking a test com-
posed of such sexist items might (again, quite properly) be annoyed
and, hence, might perform less well than they would have otherwise.
Turning to unfair penalization, think about a series of mathemat-
ical items set in a context of how to score a football game. If girls par-
ticipate less frequently in football and watch, in general, fewer
ch4.qxd 7/10/2003 9:56 AM Page 55
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R5 6
football games on television, then girls are likely to have more diffi-
culty in answering math items that revolve around how many points
are awarded when a team “scores a safety” or “kicks extra points
rather than running or passing for extra points.” These sorts of math-
ematics items simply shimmer with gender bias.
Not all penalties are unfair, of course. If students don’t study
properly and then perform poorly on a teacher’s test, such penalties
are richly deserved. What’s more, the inference a teacher would
make, based on these low scores, would be a valid one: These students
haven’t mastered the material! But if students’ personal characteris-
tics, such as their socioeconomic status (SES), are a determining fac-
tor in weaker test performance, then assessment bias has clearly
raised its unattractive head. It distorts the accuracy of students’ test
performances and invariably leads to invalid inferences about stu-
dents’ status and to unsound instructional decisions about how best
to teach those students.
Bias Detection in Large-Scale Assessments There was a time, not too long ago, when the creators of large-scale
tests (such as nationally standardized achievement tests) didn’t do a
particularly respectable job of identifying and excising biased items
from their tests. I taught educational measurement courses in the
UCLA Graduate School of Education for many years, and in the late
1970s, one of the nationally standardized achievement tests I rou-
tinely had my students critique employed a bias-detection procedure
that would be considered laughable today. Even back then, it made
me smirk a bit.
Here’s how the test-developers’ bias-detection procedure worked.
A three-person bias review committee was asked to review all of the
items of a test under development. Two of the reviewers “represented”
minority groups. If all three reviewers considered an item to be biased,
the item was eliminated from the test. Otherwise, that item stayed on
the test. If the two minority-representing reviewers considered an
ch4.qxd 7/10/2003 9:56 AM Page 56
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 7
item biased beyond belief, but the third reviewer disagreed, the item
stayed. Today, such a superficial bias-review process would be recog-
nized as absurd.
Developers of large-scale tests now employ far more rigorous bias-
detection procedures, especially with respect to bias based on race
and gender. A typical approach these days calls for the creation of a
bias-review committee, usually of 15–25 members, almost all of
whom are themselves members of minority groups. Bias-review com-
mittees are typically given ample training and practice in how best to
render their item-bias judgments. For each item that might end up
being included in the test, every member of the bias-review commit-
tee would be asked to respond to the following question:
Might this item offend or unfairly penalize students because of
such personal characteristics as gender, ethnicity, religion, or
socioeconomic status?
Yes No Uncertain
Note that the above question asks reviewers to judge whether a stu-
dent might be penalized, not whether the item would (for absolutely
certain) penalize a student. This kind of phrasing makes it clear that
the test’s developers are going the extra mile to eliminate any items
that might possibly offend or unfairly penalize students because of per-
sonal characteristics. If the review question asked whether an item
“would” offend or unfairly penalize students, you can bet that fewer
items would be reported as biased.
If a certain percentage of bias-reviewers believe the item might be
biased, it is eliminated from those to be used in the test. Clearly, the
determination of what proportion of reviewers is needed to delete an
item on potential-bias grounds represents an important issue. With
the three-person bias-review committee my 1970s students read
about at UCLA, the percentage of reviewers necessary for an item’s
elimination was 100 percent. In recent years I have taken part in
ch4.qxd 7/10/2003 9:56 AM Page 57
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R5 8
bias-review procedures where an item would eliminated if more than
five percent of the reviewers rated the item as biased. Times have def-
initely changed.
My experiences suggest that the developers of most of our current
large-scale tests have been attentive to the potential presence of
assessment bias based on students’ gender or membership in a minor-
ity racial or ethnic group. (I still think that today’s major educational
tests feature far too many items that are biased on SES grounds. We’ll
discuss this issue in Chapter 9.) There may be a few items that have
slipped by the rigorous item-review process but, in general, those test
developers get an A for effort. In many instances, such assiduous
attention to gender and minority-group bias eradication stems
directly from the fear that any adverse results of their tests might be
challenged in court. The threat of litigation often proves to be a
potent stimulant!
Assessment Bias in the Classroom Unfortunately, there’s much more bias present in teachers’ classroom
tests than most teachers imagine. The reason is not that today’s class-
room teachers deliberately set out to offend or unfairly penalize cer-
tain of their students; it’s just that too few teachers have systemati-
cally attended to this issue.
The key to unbiasing tests is a simple matter of serious, item-by-
item scrutiny. The same kind of item-review question that bias-
reviewers typically use in their appraisal of large-scale assessments will
work for classroom tests, too. A teacher who is instructing students
from racial/ethnic groups other than the teacher’s own racial/ethnic
group might be wise to ask a colleague (or a parent) from those
racial/ethnic groups to serve as a one-person bias review committee.
This can be very illuminating. Happily, most biased items can be
repaired with only modest effort. Those that can’t should be tossed.
This whole bias-detection business is about being fair to all
students and assessing them in such a way that they are accurately
ch4.qxd 7/10/2003 9:56 AM Page 58
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
V a l i d i t y , R e l i a b i l i t y , a n d B i a s 5 9
measured. That’s the only way a teacher’s test-based inferences will be
valid. And, of course, valid inferences about students serve as the
foundation for defensible instructional decisions. Invalid inferences
don’t.
Recommended Resources
American Educational Research Association. (1999). Standards for educational and psychological testing. Washington, DC: Author.
McNeil, L. M. (2000, June). Creating new inequalities: Contradictions of reform. Phi Delta Kappan, 81(10), 728–734.
Popham, W. J. (2000). Modern educational measurement: Practical guidelines for educational leaders (3rd ed.). Boston: Allyn & Bacon.
Popham, W. J. (Program Consultant). (2000). Norm- and criterion-referenced testing: What assessment-literate educators should know [Videotape]. Los Angeles: IOX Assessment Associates.
Popham, W. J. (Program Consultant). (2000). Standardized achievement tests: How to tell what they measure [Videotape]. Los Angeles: IOX Assessment Associates.
INSTRUCTIONALLY FOCUSED TESTING TIPS
• Recognize that validity refers to a test-based inference, not to
the test itself.
• Understand that there are three kinds of validity evidence, all of
which can contribute to the confidence teachers have in the
accuracy of a test-based inference about students.
• Assemble content-related evidence of validity for your most
important classroom tests.
• Know that there are three related, but meaningfully different
kinds of reliability evidence that can be collected for educational
tests.
• Give serious attention to the detection and elimination of
assessment bias in classroom tests.
ch4.qxd 7/10/2003 9:56 AM Page 59
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-13 03:38:36.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .