PLAGIARISM FREE "A" WORK 18 HOURS or LESS
IN THE SPRING OF MY FIRST FULL YEAR OF TEACHING, I ADMINISTERED MY FIRST
nationally standardized achievement test. We set aside an hour, our
high school’s juniors took the test, and we received the score-reports
the following fall. I looked over the results to see which of my stu-
dents had earned high-percentile scores and which had earned low-
percentile scores. But, because ours was a small high school with only
about 35 students per grade level, I had already discovered which of
my students performed well on tests. The standardized test’s results
yielded no surprises.
Our school district required us to give this test every year, and we
routinely mailed all students’ score-reports to their parents. As far as
I can recall, no parent ever contacted me, the school’s principal, or
any other teachers about the child’s standardized test scores. Why
would they? There was nothing riding on the results. Our low-
scoring students were not held back a grade level, denied diplomas,
or forced to take summer school classes. And no citizen of our rural
Oregon town ever tried to evaluate our school’s success on the basis
of its students’ performances on those standardized achievement
tests. Those tests, in contrast to today’s high-stakes tests, were gen-
uinely no-stakes tests. Things have really changed.
9 Uses and Misuses of Standardized Achievement Tests
1 2 2
ch9.qxd 7/30/2003 12:43 PM Page 122
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 3
Standardized Achievement Tests as the Evaluative Yardstick In the United States today, most citizens regard students’ perform-
ances on standardized achievement tests as the definitive indicator of
school quality. These test scores, published in newspapers, monitored
by district administrators and state departments of education, and
reported to the federal government, mark school staffs as either suc-
cessful or unsuccessful. Schools whose students score well on stan-
dardized achievement tests are often singled out for applause or,
increasingly, given significant monetary rewards. On the flip side of
the evaluative coin, schools whose students score too low on stan-
dardized tests are singled out for intensive staff development. If test
scores do not improve “sufficiently” after substantial staff-develop-
ment efforts, the schools can be taken over by for-profit corporations
or, in some instances, simply closed down altogether.
If you think the consequences of low standardized test scores are
considerable now, just wait until NCLB’s adequate yearly progress
requirements kick in. The NCLB Act requires schools to promote their
students’ adequate yearly progress (AYP) according to a state-deter-
mined time schedule. Schools that fail to get sufficient numbers of
their students to make AYP (as measured by statewide tests tied to
challenging content standards) will be labeled “low performing.”
After two years of low performance, schools and districts that receive
NCLB Title I funds will be subject to a whole series of negative sanc-
tions. For instance, if a school fails to achieve its AYP targets for two
consecutive years, parents of children in the school will be permitted
to transfer their children to a nonfailing district school—with trans-
portation costs picked up by the district. After another year of miss-
ing AYP requirements, the school will be required to supply its stu-
dents with supplemental instruction, such as tutoring sessions.
Unfortunately, the increasing evaluative significance of standardized
achievement tests and the resulting pressure on teachers to raise their
ch9.qxd 7/30/2003 12:43 PM Page 123
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 2 4
students’ test scores are contributing to a number of educationally
indefensible practices now seen with increasing frequency. Perhaps
the most obvious fallout of the score-boosting frenzy is curricular
reductionism, wherein teachers have chosen (or have been directed) to
give short shrift to any content not assessed on standardized achieve-
ment tests. In locales where this kind of curricular shortsightedness is
rampant, students simply aren’t being given an opportunity to learn
the things they should be learning.
A second by-product is a dramatic increase in the amount of
drudgery drilling in classrooms. Students are required to devote sub-
stantial chunks of their school day to relentless, often mind-numbing
practice on items similar to those they will encounter on a standard-
ized test. Such drilling can, of course, rapidly extinguish any joy that
students might derive either from school or from learning itself.
Remember the previous chapter’s discussion of affect? Well, today’s
ubiquitous test-preparation “drill and kill” sessions can quickly
destroy the positive attitude toward school that children really ought
to have.
Finally, as a result of the enormous pressure on educators to
improve students’ test scores, we have seen far too many instances of
improper test preparation or improper test administration. In some cases,
students have been given test-preparation practice sessions based on
the very same items they will encounter on “real” test. In other
instances, students have been given substantially more time to com-
plete a standardized test than is stipulated. (Standardized tests admin-
istered in a nonstandardized manner are, of course, no longer stan-
dardized.) There are even cases in which students’ answer sheets have
been massively “refined” by educators prior to the official submission
of those answer sheets to a state-designated scoring firm. Obviously,
this sort of unethical conduct by educators sends an inappropriate
message to students. Thankfully, such conduct is still relatively rare.
It is because the widespread use of standardized achievement tests
as the dominant evaluative school-quality yardstick has led to
ch9.qxd 7/30/2003 12:43 PM Page 124
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 5
increasingly frequent instances of curricular reductionism, drudgery
drilling, and improper test-preparation or test-administration that I
believe today’s teachers really need to learn more about the uses and
misuses of standardized tests. I’ll be up-front about my point of view:
The primary use of standardized achievement tests today—to evalu-
ate school and teacher quality—is a misuse. It is mistaken. It is just
plain wrong.
The Measurement Mission of Standardized Achievement Tests A standardized test is any assessment device that’s administered and
scored in a standard, predetermined manner. Earlier in this book, I
explained that achievement tests (such as the Stanford Achievement
Tests) attempt to measure students’ skills and knowledge, whereas
aptitude tests (such as the ACT and SAT) attempt to predict students’
success in some subsequent academic setting. Actually, in a bygone
era, educators used to consider aptitude tests “group intelligence”
tests. That interpretation has long since gone by the wayside, as it
conveys the impression that intelligence is an immutable commodi-
ty. Interestingly, even the term “aptitude” has become rather unfash-
ionable. Several years ago, the distributors of the highly esteemed SAT
decided to change the official name of their exam from the
“Scholastic Aptitude Test” to the “Scholastic Assessment Test.” Their
current preference, though, is to use the acronym only: SAT. One sus-
pects that this words-to-letters transformation is an effort to avoid
using the term aptitude. From a marketing perspective, though, the
letters-only approach may have real merit. Consider how Kentucky
Fried Chicken successfully reinvented itself as “KFC.”
At any rate, a standardized achievement test is designed to measure
a student’s relative ability to answer the test’s items. A student’s score
is compared to the scores of a carefully selected group of previous
test-takers known as the test’s norm group. Based on these com-
parisons, we can discover that Sally scored at the 92nd percentile
ch9.qxd 7/30/2003 12:43 PM Page 125
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 2 6
(meaning that Sally out-performed 92 percent of the students in the
norm group), while Billy scored at the 13th percentile (meaning that
Billy’s score only topped 13 percent of the scores earned by students
in the norm group).
Such relative comparisons can be useful to both teachers and par-
ents. If a 4th grade teacher discovers that a student has earned an
87th percentile score on a standardized language arts achievement
test, but only a 23rd percentile score in a standardized mathematics
achievement test, this suggests that some serious instructional atten-
tion should be directed toward boosting the student’s mathematics
moxie. Parents can benefit from such comparative test-based results
because these results do serve to identify a child’s relative strengths
and weaknesses.
This ability to provide accurate, fine-grained comparisons
between the scores of a current test-taker and the scores of those pre-
vious test-takers who constitute the test’s norm group is the corner-
stone of standardized achievement testing and has been since stan-
dardized testing’s origins in the early 20th century. We refer to these
comparative test-based interpretations as norm-referenced interpreta-
tions because we “reference” a student’s test score back to the scores
of the test’s norm group and, thereby, give the student’s score mean-
ing. Raw test scores all by themselves are really quite uninterpretable.
In order for a standardized test to permit fine-grained, norm-ref-
erenced inferences about a given student’s performance, however, it
is necessary for the test to produce sufficient score-spread (technically
referred to as test score variance). If the scores yielded by a standard-
ized test were all bunched together within a few points of each other,
then precise comparisons among students’ scores would be impossi-
ble. The production of adequate score-spread, therefore, is imperative
for the creators of traditional standardized achievement tests. But it is
this quest for score-spread that turns out to render such tests unsuit-
able for the evaluation of school and teacher quality. Let’s see why.
ch9.qxd 7/30/2003 12:43 PM Page 126
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 7
Test Design Features Contributing to Score-Spread . . . and Inhibiting Evaluation of Educational Effectiveness In subsequent sections of this chapter, I am going to be taking a slap
or two at standardized achievement tests when they’re used to evalu-
ate schools. I’ll be disparaging this sort of test use not because the
tests themselves are tawdry. To the contrary, I regard traditional stan-
dardized achievement tests as first-rate assessment tools when they
are used for an appropriate purpose. Nor do I want to imply that the
designers of these tests are malevolent measurement monsters out to
mislabel students or schools. If we discover that a surgical scalpel has
been used as a weapon during an assault, that doesn’t mean the firm
that manufactured the scalpel is at fault. It’s just a case of a tool being
used for the wrong purpose. This is just what’s happening with the
use of standardized tests, created to permit comparisons among stu-
dents but misapplied to assess educational effectiveness.
An Emphasis on Mid-Difficulty Items Because most standardized tests are built to be administered in about
an hour or so (otherwise, students would become restless or, worse,
openly rebellious), the developers of such tests must be very judicious
in the kinds of items they select. Their goal is to get maximum score-
spread from the fewest number of items and still measure all the
required variables.
Statistically, test items that produce the maximum score-spread
are those that will be answered correctly by roughly half of the test-
takers. The testing term p-value indicates the percentage of students
who answer an item correctly. To create the ample score-spread nec-
essary for precise comparison, test developers select the vast majority
of their items so that those items have p-values of between .40 and
.60—that is, the items were answered correctly by between 40 percent
and 60 percent of test-takers when the under-development items
were tried out during early field tests.
ch9.qxd 7/30/2003 12:43 PM Page 127
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 2 8
What the developers of standardized tests resolutely avoid are test
items that have extremely low or extremely high p-values. Such items
are viewed as space-wasters because they don’t “do their share” to
spread out students’ scores. Accordingly, most of these items are jet-
tisoned before a test is released. And, as a test is revised (which typi-
cally takes place every half-dozen years or so), the developers will
look at data based on how real test-takers have actually responded.
They’ll then replace almost all items with p-values that are very high
or very low with items that have mid-range p-values, more friendly to
score-spread.
Here’s the catch: The avoidance of items in the high p-value
ranges (p-values of .80 or .90) tends to reduce the ability of standard-
ized achievement tests to detect truly effective instruction. Think
about it. Items with high p-values indicate that most students possess
the knowledge or have mastered the skills that the items represent.
The skills and knowledge that teachers regard as most important tend
to be the ones that those teachers stress in their instruction. And,
even allowing for plenty of differences in teachers’ instructional
skills, the more that teachers stress certain content, the better their
students will perform on items that measure such teacher-stressed
content. But the better students perform on those items, the more
likely it will be that those very items will be jettisoned when the stan-
dardized test is revised.
In short, the quest for score-spread creates a clearly identifiable
tendency to remove from traditionally constructed standardized
achievement tests those items that measure the most important,
teacher-stressed content. Clearly, a test that deliberately dodges the
most important things teachers try to teach should not be used to
judge teachers’ instructional success.
Items Linked to Test-Takers’ Socioeconomic Status Remember, the traditional measurement mission of standardized
achievement tests is to provide accurate norm-referenced interpretations,
ch9.qxd 7/30/2003 12:43 PM Page 128
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 9
and to do that, the test must produce ample score-spread. Again, due to
limited test-administration time and the need to get maximum score
variance from a minimal number of items, some of the items on stan-
dardized achievement tests are highly related to a student’s socioeco-
nomic status (SES).
Here’s an example taken from a currently used standardized
achievement test. It’s a 6th grade science item, and I’ve modified it
slightly, changing some words to preserve the test’s security. I want to
stress, though, that I have not altered the nature of the original item’s
cognitive demand—what it asks students to do.
AN SES-LINKED ITEM
Because a plant’s fruit always contains seeds, which one of the
following is not a fruit?
a. pumpkin
b. celery
c. orange
d. pear
If you look carefully at this sample item, you’ll realize that children
from more-privileged backgrounds (with parents who can routinely
afford to buy fresh celery at the supermarket and purchase fresh
pumpkins for Halloween carving) will generally do better on it than
will children from less-privileged backgrounds (with parents who are
eking out the family meals on government-issued food stamps). This
is a classic SES-linked item.
It just so happens that socioeconomic status is a nicely spread out
variable, and it doesn’t change all that rapidly. So, by linking test items
to SES, the developers of standardized achievement tests are almost cer-
tain to get the score-spread they need. But SES-linked items measure what
students bring to school, not what they learn there. For this reason, SES-
linked items are not appropriate for evaluating instructional quality.
ch9.qxd 7/30/2003 12:43 PM Page 129
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 3 0
Items Linked to Test-Takers’ Inherited Academic Aptitude Children differ at birth, depending on what transpired during the
gene-pool lottery. Some children are destined to grow up taller, heav-
ier, or more attractive than their age-mates. Children also differ from
birth in certain academic aptitudes, namely, in their verbal, quantita-
tive, or spatial potentials.
From a teacher’s perspective, classroom instruction would be far
simpler if all children were born with identical academic aptitudes.
But that’s not the world we live in. We know, for example, that some
children come into class with inherited quantitative smarts that
exceed those of their classmates. Such children “catch on” quickly to
most mathematical concepts, and they are likely to sail more easily
through most of a school’s mathematical challenges.
Of course, this is not to say that children born without inherit
superior quantitative aptitude should cease their mathematical jour-
ney shortly after mastering 2 + 2 or that they will not go on to high
levels of mathematical prowess. It’s just that children whose inborn
quantitative aptitude is low will probably have to work harder and
longer to do so. That’s the way academic aptitudes work.
I concur with Howard Gardner’s contention that there are multi-
ple intelligences. Kids can be weak in verbal smarts, yet possess superb
aesthetic smarts. I’m pretty good at mathematical stuff, yet I’m a
blockhead when it comes to interpersonal sensitivities. Surely, there
is not just one kind of intelligence. The people who create tradition-
al standardized achievement tests are particularly concerned with
three specific sorts: quantitative, verbal, and spatial aptitudes. You will
find a good many items in standardized achievement tests that pri-
marily assess these three kinds of smarts.
Consider, for example, the following 4th grade mathematics
item. It, too, was drawn from a current standardized achievement test
and is presented with only minor modifications to preserve test
security.
ch9.qxd 7/30/2003 12:43 PM Page 130
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 1
AN INHERITED APTITUDE-LINKED ITEM
Which one of the letters below, when folded in half, can have
exactly two matching parts?
a. Z
b. F
c. Y
d. S
Children who were born with ample spatial smarts will have a far eas-
ier time identifying the correct answer. (It’s choice c.) Yes, this item is
designed to measure a student’s inborn spatial aptitude. It’s certainly
not measuring a skill that teachers promote through instruction.
After all, how often is “mental letter-folding” taught in 4th grade
mathematics classrooms? Answer: Never.
Like socioeconomic status, inherited academic aptitudes are nice-
ly spread out in the population. By linking a test’s items to one of
these aptitudes, test developers have a better chance of creating the
kind of score-spread that traditionally constructed standardized
achievement tests must possess if they’re going to carry out their
comparative measurement mission properly.
But again: inheritance-linked items measure what students bring to
school, not what they learn there. Such items are not appropriate for
evaluating instructional quality. I suppose it could be argued that
inheritance-linked items have a role to play in aptitude tests (espe-
cially if you regard such assessments as some sort of intelligence test).
Still, aptitude-linked items really have no place at all in what is sup-
posed to be an achievement test.
To reiterate, the measurement function of traditionally construct-
ed standardized achievement tests is to permit relative comparisons
among test-takers, usually by contrasting an individual’s score with a
norm-group’s scores on the same test. For these relative (norm-refer-
enced) comparisons to be accurate, the test must create a considerable
ch9.qxd 7/30/2003 12:43 PM Page 131
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 3 2
spread in test-takers’ scores. However, in the pursuit of score-spread,
the developers of standardized achievement test often include items
blatantly unsuitable for evaluating the effectiveness of instruction.
The Prevalence of Inappropriate Items How many such score-spreading items are there in a typical stan-
dardized achievement test? Well, the number surely varies from test
to test, but I recently went through a pair of different standardized
achievement tests, item by item, at two different grade levels. I really
was trying to be objective in my judgments, but if I thought the dom-
inant factor in a student’s coming up with a correct answer was either
socioeconomic status or inherited academic aptitude, I flagged the
item. These are the approximate percentages I found:
• 50 percent of reading items.
• 75 percent of language arts items.
• 15 percent of mathematics items.
• 85 percent of science items.
• 65 percent of social studies items.
Yes, it’s a little shocking. Even if you were to cut my percentages
in half (because, although I was trying to be objective, I may have let
my biases blind me), these tests would still include way too many
items that ought not to be used to evaluate the quality of instruction.
But then, they are absolutely appropriate for a standardized test’s tra-
ditional comparative assessment mission.
I challenge you to spend an hour or two with a copy of a stan-
dardized achievement test and do your own judging about the num-
ber of items in the test that careful analysis will reveal to be strongly
dependent on children’s socioeconomic status or on their inherited
academic aptitudes. If you accept this challenge, let me caution
against judging an item positively because you’d like a test-taker be
able to answer the item correctly. Heck, we’d like all test-takers to
answer every item correctly. Nor should you defer to the technical
ch9.qxd 7/30/2003 12:43 PM Page 132
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 3
expertise of those who originally wrote the item. Remember, the item
is apt to have been written to satisfy a different assessment function
than the evaluation of educators’ instructional effectiveness.
If you’re up to this challenge, your task is to make a Yes, No, or
Uncertain judgment about each test item based on this question:
Will this test item, along with others, be helpful in determining
what students were taught in school?
If your item-by-item scrutiny yields many No or Uncertain judgments
because items are either SES-linked or inheritance-linked, then the
test you’re reviewing should definitely not be used to evaluate teach-
ers’ instructional success.
Another Problem: Standardized Tests’ Ill-Defined Instructional Targets With so much pressure on U.S. teachers to raise students’ scores on
standardized achievement tests, it is not surprising that a vigorous
test-preparation industry has blossomed in this country. “Test-prep”
booklets and computer programs now abound, and many are linked
to a specific standardized achievement test. In some districts, teach-
ers have been directed to devote substantial segments of their regular
classroom time to unabashed preparation for a particular standard-
ized achievement test, either a nationally published test or a test cus-
tomized for their state’s accountability program.
Unfortunately, because many of these state-customized tests were
built by the same firms that distribute the national standardized
achievement tests, they too have been developed according to the
traditional score-spreading measurement model. As a consequence,
these customized tests are often no better for evaluating instruction
than an off-the-shelf, nationally standardized achievement test.
I realize that some of you reading this book may be teaching in
states where a customized statewide test has been constructed so that
ch9.qxd 7/30/2003 12:43 PM Page 133
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 3 4
it is closely aligned with the state’s official content standards. Surely,
you might think, the results of such a test must provide some insight
into the quality of classroom instruction. Sadly, this is rarely the case.
One reason—a serious shortcoming of today’s so-called standards-
based tests and the whole standards-based reform strategy—is that
these tests typically do not supply teachers with a report regarding a
student’s standard-by-standard mastery. How can teachers decide
which aspects of their instruction need to be modified if they are
unable to determine which content standards their students have
mastered and which they have not? Without per-standard reporting,
all that teachers get is a general and potentially misleading report of
students’ overall standards mastery. This information has little
instructional value.
Another instructional shortcoming of most standards-based tests
is that they don’t spell out what they’re actually measuring with suf-
ficient clarity so that a teacher can teach toward the bodies of skills
and knowledge the tests represent. Remember, a test is only supposed
to represent (that is, sample) a body of knowledge and skills. Based on
the student’s score on that test-created representation, the teacher
reaches an inference about the student’s content mastery. But, as we
discussed back in Chapter 2, the teacher should direct the actual
instruction—and all test-preparation activities—toward the body of
knowledge and skills represented by a specific set of test items, not
toward the test itself. I’ve represented this graphically in Figure 9.1.
For purposes of a teacher’s instructional decision making, the dif-
ficulty is that the description of what the standardized achievement
test measures is typically way too skimpy to help a teacher direct
instruction properly. Why then don’t test developers just take the
time to provide instructionally helpful descriptions? Well, remember
that as long as a traditional standardized achievement test provides
satisfactory comparative interpretations, there’s no compelling rea-
son for the developers to do so.
ch9.qxd 7/30/2003 12:43 PM Page 134
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 5
As you’ll see in the next chapter, it is possible to create standard-
ized achievement tests so they actually do define what they assess at
a level suitable for instructional decision making. However, if you
find yourself forced to use a traditional standardized achievement
test, you must be wary of teaching too specifically toward the test’s
actual items. Your litmus test, when you judge your own test-prepa-
ration activities, should be your answer to the following question:
Will this test-preparation activity not only improve students’ test
performance, but also improve their mastery of the skills and
knowledge this test represents?
Oh, it’s all right to give students an hour or two of preparation dealing
with general test-taking tactics, such as how to allocate test-taking time
judiciously or how to make informed guesses. But beyond such brief
one-size-fits-all preparation to help students cope with the trauma of
9 . 1 PROPER AND IMPROPER DIRECTIONS FOR A TEACHER’S INSTRUCTIONAL EFFORTS
ch9.qxd 7/30/2003 12:43 PM Page 135
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 3 6
test taking, your instruction should focus on what the test represents,
not on the test itself.
A Misleading Label? For years, educators have been using the label “standardized achieve-
ment tests” to identify tests such as the Iowa Tests of Basic Skills or
the California Achievement Tests, and the term is in even greater cir-
culation these days as each state is preparing to comply with the new
measurement requirements of the No Child Left Behind Act. But
according to my dictionary, achievement refers to something that has
been accomplished “through great effort.” In fact, that same diction-
ary describes an achievement test as “a test to measure a person’s
knowledge or proficiency in something that can be learned or
taught.” It’s safe to say that most people think of an achievement test
as a measure of what “students have learned in school,” which is one
reason so many educational policymakers automatically believe that
students’ scores on standardized achievement tests provide a defensi-
ble indication of a school’s instructional quality.
What most people don’t know, but you now do, is that the his-
toric mission of standardized testing is at cross-purposes with the
intent of achievement testing. And because of the historic need to
produce score-spread, standardized achievement tests don’t do a very
good job of measuring what students have learned in school through
their efforts and the efforts of their teachers. As I’ve indicated, a sub-
stantial part of a student’s score on a standardized achievement test
is likely to reflect not what was taught in school, but what the stu-
dent brought to school in the first place.
Our educational community is, in my view, partially to blame for
today’s widespread misconception that standardized achievement tests
can be used to determine instructional quality. (I fault myself, too, for
I should personally have been working much harder to help dissuade
both educators and the public from the idea that standardized test
scores accurately reflect educational quality.) But it’s not too late to
start correcting this prevalent and harmful misconception.
ch9.qxd 7/30/2003 12:43 PM Page 136
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 7
I encourage you to spread the word, first among your colleagues
and then to parents and other concerned citizens. There are legitimate
ways to evaluate instructional quality, and we’ll look at some of these
in the next two chapters. But please do what you can to get the word
out that evaluating instructional quality with traditional standard-
ized achievement tests is flat-out wrong.
Recommended Resources
Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Lawrence Erlbaum Associates.
Kohn, A. (2000). The case against standardized testing: Raising the scores, ruining the schools. Westport, CT: Heinemann.
Kohn, A. (Program Consultant). (2000). Beyond the standards movement: Defending quality education in an age of test scores [Videotape]. Port Chester, NY: National Professional Resources, Inc.
INSTRUCTIONALLY FOCUSED TESTING TIPS
• Explain to colleagues and parents why standardized achieve-
ment tests’ traditional function to provide accurate norm-refer-
enced interpretations is dependent on sufficient score-spread
among students’ test performances.
• Describe to colleagues and parents how it is that three types of
score-spreading items (mid-difficulty items, SES-linked items, and
aptitude-linked items) reduce the suitability of traditionally con-
structed standardized achievement tests for evaluating instruc-
tional quality.
• Spend time reviewing an actual standardized achievement
test’s items to determine the proportion of items you regard as
unsuitable for determining what students were taught in school.
• Recognize that the descriptive information supplied with tradi-
tional standardized achievement tests does not describe the skills
and knowledge represented by those tests in a manner adequate
to support teachers’ instructional decision making.
ch9.qxd 7/30/2003 12:43 PM Page 137
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 3 8
Lemann, N. (2002). The big test: The secret history of the American meritocracy. New York: Farrar, Straus and Giroux.
Northwest Regional Educational Laboratory. (1991). Understanding standard- ized tests [Videotape]. Los Angeles: IOX Assessment Associates.
Popham, W. J. (Program Consultant). (2000). Standardized achievement tests: Not to be used in judging school quality [Videotape]. Los Angeles: IOX Assessment Associates.
Popham, W. J. (Program Consultant). (2002). Evaluating schools: Right tasks, wrong tests [Videotape]. Los Angeles: IOX Assessment Associates.
Sacks, P. (1999). Standardized minds: The high price of America’s testing culture and what we can do to change it. Cambridge, MA: Perseus Books.
ch9.qxd 7/30/2003 12:43 PM Page 138
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .