PLAGIARISM FREE "A" WORK 18 HOURS or LESS

profileNeNe1994
Chapter9UsesandMisusesofStandardizedAchievementTests.pdf

IN THE SPRING OF MY FIRST FULL YEAR OF TEACHING, I ADMINISTERED MY FIRST

nationally standardized achievement test. We set aside an hour, our

high school’s juniors took the test, and we received the score-reports

the following fall. I looked over the results to see which of my stu-

dents had earned high-percentile scores and which had earned low-

percentile scores. But, because ours was a small high school with only

about 35 students per grade level, I had already discovered which of

my students performed well on tests. The standardized test’s results

yielded no surprises.

Our school district required us to give this test every year, and we

routinely mailed all students’ score-reports to their parents. As far as

I can recall, no parent ever contacted me, the school’s principal, or

any other teachers about the child’s standardized test scores. Why

would they? There was nothing riding on the results. Our low-

scoring students were not held back a grade level, denied diplomas,

or forced to take summer school classes. And no citizen of our rural

Oregon town ever tried to evaluate our school’s success on the basis

of its students’ performances on those standardized achievement

tests. Those tests, in contrast to today’s high-stakes tests, were gen-

uinely no-stakes tests. Things have really changed.

9 Uses and Misuses of Standardized Achievement Tests

1 2 2

ch9.qxd 7/30/2003 12:43 PM Page 122

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 3

Standardized Achievement Tests as the Evaluative Yardstick In the United States today, most citizens regard students’ perform-

ances on standardized achievement tests as the definitive indicator of

school quality. These test scores, published in newspapers, monitored

by district administrators and state departments of education, and

reported to the federal government, mark school staffs as either suc-

cessful or unsuccessful. Schools whose students score well on stan-

dardized achievement tests are often singled out for applause or,

increasingly, given significant monetary rewards. On the flip side of

the evaluative coin, schools whose students score too low on stan-

dardized tests are singled out for intensive staff development. If test

scores do not improve “sufficiently” after substantial staff-develop-

ment efforts, the schools can be taken over by for-profit corporations

or, in some instances, simply closed down altogether.

If you think the consequences of low standardized test scores are

considerable now, just wait until NCLB’s adequate yearly progress

requirements kick in. The NCLB Act requires schools to promote their

students’ adequate yearly progress (AYP) according to a state-deter-

mined time schedule. Schools that fail to get sufficient numbers of

their students to make AYP (as measured by statewide tests tied to

challenging content standards) will be labeled “low performing.”

After two years of low performance, schools and districts that receive

NCLB Title I funds will be subject to a whole series of negative sanc-

tions. For instance, if a school fails to achieve its AYP targets for two

consecutive years, parents of children in the school will be permitted

to transfer their children to a nonfailing district school—with trans-

portation costs picked up by the district. After another year of miss-

ing AYP requirements, the school will be required to supply its stu-

dents with supplemental instruction, such as tutoring sessions.

Unfortunately, the increasing evaluative significance of standardized

achievement tests and the resulting pressure on teachers to raise their

ch9.qxd 7/30/2003 12:43 PM Page 123

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 2 4

students’ test scores are contributing to a number of educationally

indefensible practices now seen with increasing frequency. Perhaps

the most obvious fallout of the score-boosting frenzy is curricular

reductionism, wherein teachers have chosen (or have been directed) to

give short shrift to any content not assessed on standardized achieve-

ment tests. In locales where this kind of curricular shortsightedness is

rampant, students simply aren’t being given an opportunity to learn

the things they should be learning.

A second by-product is a dramatic increase in the amount of

drudgery drilling in classrooms. Students are required to devote sub-

stantial chunks of their school day to relentless, often mind-numbing

practice on items similar to those they will encounter on a standard-

ized test. Such drilling can, of course, rapidly extinguish any joy that

students might derive either from school or from learning itself.

Remember the previous chapter’s discussion of affect? Well, today’s

ubiquitous test-preparation “drill and kill” sessions can quickly

destroy the positive attitude toward school that children really ought

to have.

Finally, as a result of the enormous pressure on educators to

improve students’ test scores, we have seen far too many instances of

improper test preparation or improper test administration. In some cases,

students have been given test-preparation practice sessions based on

the very same items they will encounter on “real” test. In other

instances, students have been given substantially more time to com-

plete a standardized test than is stipulated. (Standardized tests admin-

istered in a nonstandardized manner are, of course, no longer stan-

dardized.) There are even cases in which students’ answer sheets have

been massively “refined” by educators prior to the official submission

of those answer sheets to a state-designated scoring firm. Obviously,

this sort of unethical conduct by educators sends an inappropriate

message to students. Thankfully, such conduct is still relatively rare.

It is because the widespread use of standardized achievement tests

as the dominant evaluative school-quality yardstick has led to

ch9.qxd 7/30/2003 12:43 PM Page 124

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 5

increasingly frequent instances of curricular reductionism, drudgery

drilling, and improper test-preparation or test-administration that I

believe today’s teachers really need to learn more about the uses and

misuses of standardized tests. I’ll be up-front about my point of view:

The primary use of standardized achievement tests today—to evalu-

ate school and teacher quality—is a misuse. It is mistaken. It is just

plain wrong.

The Measurement Mission of Standardized Achievement Tests A standardized test is any assessment device that’s administered and

scored in a standard, predetermined manner. Earlier in this book, I

explained that achievement tests (such as the Stanford Achievement

Tests) attempt to measure students’ skills and knowledge, whereas

aptitude tests (such as the ACT and SAT) attempt to predict students’

success in some subsequent academic setting. Actually, in a bygone

era, educators used to consider aptitude tests “group intelligence”

tests. That interpretation has long since gone by the wayside, as it

conveys the impression that intelligence is an immutable commodi-

ty. Interestingly, even the term “aptitude” has become rather unfash-

ionable. Several years ago, the distributors of the highly esteemed SAT

decided to change the official name of their exam from the

“Scholastic Aptitude Test” to the “Scholastic Assessment Test.” Their

current preference, though, is to use the acronym only: SAT. One sus-

pects that this words-to-letters transformation is an effort to avoid

using the term aptitude. From a marketing perspective, though, the

letters-only approach may have real merit. Consider how Kentucky

Fried Chicken successfully reinvented itself as “KFC.”

At any rate, a standardized achievement test is designed to measure

a student’s relative ability to answer the test’s items. A student’s score

is compared to the scores of a carefully selected group of previous

test-takers known as the test’s norm group. Based on these com-

parisons, we can discover that Sally scored at the 92nd percentile

ch9.qxd 7/30/2003 12:43 PM Page 125

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 2 6

(meaning that Sally out-performed 92 percent of the students in the

norm group), while Billy scored at the 13th percentile (meaning that

Billy’s score only topped 13 percent of the scores earned by students

in the norm group).

Such relative comparisons can be useful to both teachers and par-

ents. If a 4th grade teacher discovers that a student has earned an

87th percentile score on a standardized language arts achievement

test, but only a 23rd percentile score in a standardized mathematics

achievement test, this suggests that some serious instructional atten-

tion should be directed toward boosting the student’s mathematics

moxie. Parents can benefit from such comparative test-based results

because these results do serve to identify a child’s relative strengths

and weaknesses.

This ability to provide accurate, fine-grained comparisons

between the scores of a current test-taker and the scores of those pre-

vious test-takers who constitute the test’s norm group is the corner-

stone of standardized achievement testing and has been since stan-

dardized testing’s origins in the early 20th century. We refer to these

comparative test-based interpretations as norm-referenced interpreta-

tions because we “reference” a student’s test score back to the scores

of the test’s norm group and, thereby, give the student’s score mean-

ing. Raw test scores all by themselves are really quite uninterpretable.

In order for a standardized test to permit fine-grained, norm-ref-

erenced inferences about a given student’s performance, however, it

is necessary for the test to produce sufficient score-spread (technically

referred to as test score variance). If the scores yielded by a standard-

ized test were all bunched together within a few points of each other,

then precise comparisons among students’ scores would be impossi-

ble. The production of adequate score-spread, therefore, is imperative

for the creators of traditional standardized achievement tests. But it is

this quest for score-spread that turns out to render such tests unsuit-

able for the evaluation of school and teacher quality. Let’s see why.

ch9.qxd 7/30/2003 12:43 PM Page 126

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 7

Test Design Features Contributing to Score-Spread . . . and Inhibiting Evaluation of Educational Effectiveness In subsequent sections of this chapter, I am going to be taking a slap

or two at standardized achievement tests when they’re used to evalu-

ate schools. I’ll be disparaging this sort of test use not because the

tests themselves are tawdry. To the contrary, I regard traditional stan-

dardized achievement tests as first-rate assessment tools when they

are used for an appropriate purpose. Nor do I want to imply that the

designers of these tests are malevolent measurement monsters out to

mislabel students or schools. If we discover that a surgical scalpel has

been used as a weapon during an assault, that doesn’t mean the firm

that manufactured the scalpel is at fault. It’s just a case of a tool being

used for the wrong purpose. This is just what’s happening with the

use of standardized tests, created to permit comparisons among stu-

dents but misapplied to assess educational effectiveness.

An Emphasis on Mid-Difficulty Items Because most standardized tests are built to be administered in about

an hour or so (otherwise, students would become restless or, worse,

openly rebellious), the developers of such tests must be very judicious

in the kinds of items they select. Their goal is to get maximum score-

spread from the fewest number of items and still measure all the

required variables.

Statistically, test items that produce the maximum score-spread

are those that will be answered correctly by roughly half of the test-

takers. The testing term p-value indicates the percentage of students

who answer an item correctly. To create the ample score-spread nec-

essary for precise comparison, test developers select the vast majority

of their items so that those items have p-values of between .40 and

.60—that is, the items were answered correctly by between 40 percent

and 60 percent of test-takers when the under-development items

were tried out during early field tests.

ch9.qxd 7/30/2003 12:43 PM Page 127

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 2 8

What the developers of standardized tests resolutely avoid are test

items that have extremely low or extremely high p-values. Such items

are viewed as space-wasters because they don’t “do their share” to

spread out students’ scores. Accordingly, most of these items are jet-

tisoned before a test is released. And, as a test is revised (which typi-

cally takes place every half-dozen years or so), the developers will

look at data based on how real test-takers have actually responded.

They’ll then replace almost all items with p-values that are very high

or very low with items that have mid-range p-values, more friendly to

score-spread.

Here’s the catch: The avoidance of items in the high p-value

ranges (p-values of .80 or .90) tends to reduce the ability of standard-

ized achievement tests to detect truly effective instruction. Think

about it. Items with high p-values indicate that most students possess

the knowledge or have mastered the skills that the items represent.

The skills and knowledge that teachers regard as most important tend

to be the ones that those teachers stress in their instruction. And,

even allowing for plenty of differences in teachers’ instructional

skills, the more that teachers stress certain content, the better their

students will perform on items that measure such teacher-stressed

content. But the better students perform on those items, the more

likely it will be that those very items will be jettisoned when the stan-

dardized test is revised.

In short, the quest for score-spread creates a clearly identifiable

tendency to remove from traditionally constructed standardized

achievement tests those items that measure the most important,

teacher-stressed content. Clearly, a test that deliberately dodges the

most important things teachers try to teach should not be used to

judge teachers’ instructional success.

Items Linked to Test-Takers’ Socioeconomic Status Remember, the traditional measurement mission of standardized

achievement tests is to provide accurate norm-referenced interpretations,

ch9.qxd 7/30/2003 12:43 PM Page 128

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 2 9

and to do that, the test must produce ample score-spread. Again, due to

limited test-administration time and the need to get maximum score

variance from a minimal number of items, some of the items on stan-

dardized achievement tests are highly related to a student’s socioeco-

nomic status (SES).

Here’s an example taken from a currently used standardized

achievement test. It’s a 6th grade science item, and I’ve modified it

slightly, changing some words to preserve the test’s security. I want to

stress, though, that I have not altered the nature of the original item’s

cognitive demand—what it asks students to do.

AN SES-LINKED ITEM

Because a plant’s fruit always contains seeds, which one of the

following is not a fruit?

a. pumpkin

b. celery

c. orange

d. pear

If you look carefully at this sample item, you’ll realize that children

from more-privileged backgrounds (with parents who can routinely

afford to buy fresh celery at the supermarket and purchase fresh

pumpkins for Halloween carving) will generally do better on it than

will children from less-privileged backgrounds (with parents who are

eking out the family meals on government-issued food stamps). This

is a classic SES-linked item.

It just so happens that socioeconomic status is a nicely spread out

variable, and it doesn’t change all that rapidly. So, by linking test items

to SES, the developers of standardized achievement tests are almost cer-

tain to get the score-spread they need. But SES-linked items measure what

students bring to school, not what they learn there. For this reason, SES-

linked items are not appropriate for evaluating instructional quality.

ch9.qxd 7/30/2003 12:43 PM Page 129

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 3 0

Items Linked to Test-Takers’ Inherited Academic Aptitude Children differ at birth, depending on what transpired during the

gene-pool lottery. Some children are destined to grow up taller, heav-

ier, or more attractive than their age-mates. Children also differ from

birth in certain academic aptitudes, namely, in their verbal, quantita-

tive, or spatial potentials.

From a teacher’s perspective, classroom instruction would be far

simpler if all children were born with identical academic aptitudes.

But that’s not the world we live in. We know, for example, that some

children come into class with inherited quantitative smarts that

exceed those of their classmates. Such children “catch on” quickly to

most mathematical concepts, and they are likely to sail more easily

through most of a school’s mathematical challenges.

Of course, this is not to say that children born without inherit

superior quantitative aptitude should cease their mathematical jour-

ney shortly after mastering 2 + 2 or that they will not go on to high

levels of mathematical prowess. It’s just that children whose inborn

quantitative aptitude is low will probably have to work harder and

longer to do so. That’s the way academic aptitudes work.

I concur with Howard Gardner’s contention that there are multi-

ple intelligences. Kids can be weak in verbal smarts, yet possess superb

aesthetic smarts. I’m pretty good at mathematical stuff, yet I’m a

blockhead when it comes to interpersonal sensitivities. Surely, there

is not just one kind of intelligence. The people who create tradition-

al standardized achievement tests are particularly concerned with

three specific sorts: quantitative, verbal, and spatial aptitudes. You will

find a good many items in standardized achievement tests that pri-

marily assess these three kinds of smarts.

Consider, for example, the following 4th grade mathematics

item. It, too, was drawn from a current standardized achievement test

and is presented with only minor modifications to preserve test

security.

ch9.qxd 7/30/2003 12:43 PM Page 130

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 1

AN INHERITED APTITUDE-LINKED ITEM

Which one of the letters below, when folded in half, can have

exactly two matching parts?

a. Z

b. F

c. Y

d. S

Children who were born with ample spatial smarts will have a far eas-

ier time identifying the correct answer. (It’s choice c.) Yes, this item is

designed to measure a student’s inborn spatial aptitude. It’s certainly

not measuring a skill that teachers promote through instruction.

After all, how often is “mental letter-folding” taught in 4th grade

mathematics classrooms? Answer: Never.

Like socioeconomic status, inherited academic aptitudes are nice-

ly spread out in the population. By linking a test’s items to one of

these aptitudes, test developers have a better chance of creating the

kind of score-spread that traditionally constructed standardized

achievement tests must possess if they’re going to carry out their

comparative measurement mission properly.

But again: inheritance-linked items measure what students bring to

school, not what they learn there. Such items are not appropriate for

evaluating instructional quality. I suppose it could be argued that

inheritance-linked items have a role to play in aptitude tests (espe-

cially if you regard such assessments as some sort of intelligence test).

Still, aptitude-linked items really have no place at all in what is sup-

posed to be an achievement test.

To reiterate, the measurement function of traditionally construct-

ed standardized achievement tests is to permit relative comparisons

among test-takers, usually by contrasting an individual’s score with a

norm-group’s scores on the same test. For these relative (norm-refer-

enced) comparisons to be accurate, the test must create a considerable

ch9.qxd 7/30/2003 12:43 PM Page 131

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 3 2

spread in test-takers’ scores. However, in the pursuit of score-spread,

the developers of standardized achievement test often include items

blatantly unsuitable for evaluating the effectiveness of instruction.

The Prevalence of Inappropriate Items How many such score-spreading items are there in a typical stan-

dardized achievement test? Well, the number surely varies from test

to test, but I recently went through a pair of different standardized

achievement tests, item by item, at two different grade levels. I really

was trying to be objective in my judgments, but if I thought the dom-

inant factor in a student’s coming up with a correct answer was either

socioeconomic status or inherited academic aptitude, I flagged the

item. These are the approximate percentages I found:

• 50 percent of reading items.

• 75 percent of language arts items.

• 15 percent of mathematics items.

• 85 percent of science items.

• 65 percent of social studies items.

Yes, it’s a little shocking. Even if you were to cut my percentages

in half (because, although I was trying to be objective, I may have let

my biases blind me), these tests would still include way too many

items that ought not to be used to evaluate the quality of instruction.

But then, they are absolutely appropriate for a standardized test’s tra-

ditional comparative assessment mission.

I challenge you to spend an hour or two with a copy of a stan-

dardized achievement test and do your own judging about the num-

ber of items in the test that careful analysis will reveal to be strongly

dependent on children’s socioeconomic status or on their inherited

academic aptitudes. If you accept this challenge, let me caution

against judging an item positively because you’d like a test-taker be

able to answer the item correctly. Heck, we’d like all test-takers to

answer every item correctly. Nor should you defer to the technical

ch9.qxd 7/30/2003 12:43 PM Page 132

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 3

expertise of those who originally wrote the item. Remember, the item

is apt to have been written to satisfy a different assessment function

than the evaluation of educators’ instructional effectiveness.

If you’re up to this challenge, your task is to make a Yes, No, or

Uncertain judgment about each test item based on this question:

Will this test item, along with others, be helpful in determining

what students were taught in school?

If your item-by-item scrutiny yields many No or Uncertain judgments

because items are either SES-linked or inheritance-linked, then the

test you’re reviewing should definitely not be used to evaluate teach-

ers’ instructional success.

Another Problem: Standardized Tests’ Ill-Defined Instructional Targets With so much pressure on U.S. teachers to raise students’ scores on

standardized achievement tests, it is not surprising that a vigorous

test-preparation industry has blossomed in this country. “Test-prep”

booklets and computer programs now abound, and many are linked

to a specific standardized achievement test. In some districts, teach-

ers have been directed to devote substantial segments of their regular

classroom time to unabashed preparation for a particular standard-

ized achievement test, either a nationally published test or a test cus-

tomized for their state’s accountability program.

Unfortunately, because many of these state-customized tests were

built by the same firms that distribute the national standardized

achievement tests, they too have been developed according to the

traditional score-spreading measurement model. As a consequence,

these customized tests are often no better for evaluating instruction

than an off-the-shelf, nationally standardized achievement test.

I realize that some of you reading this book may be teaching in

states where a customized statewide test has been constructed so that

ch9.qxd 7/30/2003 12:43 PM Page 133

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 3 4

it is closely aligned with the state’s official content standards. Surely,

you might think, the results of such a test must provide some insight

into the quality of classroom instruction. Sadly, this is rarely the case.

One reason—a serious shortcoming of today’s so-called standards-

based tests and the whole standards-based reform strategy—is that

these tests typically do not supply teachers with a report regarding a

student’s standard-by-standard mastery. How can teachers decide

which aspects of their instruction need to be modified if they are

unable to determine which content standards their students have

mastered and which they have not? Without per-standard reporting,

all that teachers get is a general and potentially misleading report of

students’ overall standards mastery. This information has little

instructional value.

Another instructional shortcoming of most standards-based tests

is that they don’t spell out what they’re actually measuring with suf-

ficient clarity so that a teacher can teach toward the bodies of skills

and knowledge the tests represent. Remember, a test is only supposed

to represent (that is, sample) a body of knowledge and skills. Based on

the student’s score on that test-created representation, the teacher

reaches an inference about the student’s content mastery. But, as we

discussed back in Chapter 2, the teacher should direct the actual

instruction—and all test-preparation activities—toward the body of

knowledge and skills represented by a specific set of test items, not

toward the test itself. I’ve represented this graphically in Figure 9.1.

For purposes of a teacher’s instructional decision making, the dif-

ficulty is that the description of what the standardized achievement

test measures is typically way too skimpy to help a teacher direct

instruction properly. Why then don’t test developers just take the

time to provide instructionally helpful descriptions? Well, remember

that as long as a traditional standardized achievement test provides

satisfactory comparative interpretations, there’s no compelling rea-

son for the developers to do so.

ch9.qxd 7/30/2003 12:43 PM Page 134

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 5

As you’ll see in the next chapter, it is possible to create standard-

ized achievement tests so they actually do define what they assess at

a level suitable for instructional decision making. However, if you

find yourself forced to use a traditional standardized achievement

test, you must be wary of teaching too specifically toward the test’s

actual items. Your litmus test, when you judge your own test-prepa-

ration activities, should be your answer to the following question:

Will this test-preparation activity not only improve students’ test

performance, but also improve their mastery of the skills and

knowledge this test represents?

Oh, it’s all right to give students an hour or two of preparation dealing

with general test-taking tactics, such as how to allocate test-taking time

judiciously or how to make informed guesses. But beyond such brief

one-size-fits-all preparation to help students cope with the trauma of

9 . 1 PROPER AND IMPROPER DIRECTIONS FOR A TEACHER’S INSTRUCTIONAL EFFORTS

ch9.qxd 7/30/2003 12:43 PM Page 135

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 3 6

test taking, your instruction should focus on what the test represents,

not on the test itself.

A Misleading Label? For years, educators have been using the label “standardized achieve-

ment tests” to identify tests such as the Iowa Tests of Basic Skills or

the California Achievement Tests, and the term is in even greater cir-

culation these days as each state is preparing to comply with the new

measurement requirements of the No Child Left Behind Act. But

according to my dictionary, achievement refers to something that has

been accomplished “through great effort.” In fact, that same diction-

ary describes an achievement test as “a test to measure a person’s

knowledge or proficiency in something that can be learned or

taught.” It’s safe to say that most people think of an achievement test

as a measure of what “students have learned in school,” which is one

reason so many educational policymakers automatically believe that

students’ scores on standardized achievement tests provide a defensi-

ble indication of a school’s instructional quality.

What most people don’t know, but you now do, is that the his-

toric mission of standardized testing is at cross-purposes with the

intent of achievement testing. And because of the historic need to

produce score-spread, standardized achievement tests don’t do a very

good job of measuring what students have learned in school through

their efforts and the efforts of their teachers. As I’ve indicated, a sub-

stantial part of a student’s score on a standardized achievement test

is likely to reflect not what was taught in school, but what the stu-

dent brought to school in the first place.

Our educational community is, in my view, partially to blame for

today’s widespread misconception that standardized achievement tests

can be used to determine instructional quality. (I fault myself, too, for

I should personally have been working much harder to help dissuade

both educators and the public from the idea that standardized test

scores accurately reflect educational quality.) But it’s not too late to

start correcting this prevalent and harmful misconception.

ch9.qxd 7/30/2003 12:43 PM Page 136

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

U s e s a n d M i s u s e s o f S t a n d a r d i z e d A c h i e v e m e n t T e s t s 1 3 7

I encourage you to spread the word, first among your colleagues

and then to parents and other concerned citizens. There are legitimate

ways to evaluate instructional quality, and we’ll look at some of these

in the next two chapters. But please do what you can to get the word

out that evaluating instructional quality with traditional standard-

ized achievement tests is flat-out wrong.

Recommended Resources

Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Lawrence Erlbaum Associates.

Kohn, A. (2000). The case against standardized testing: Raising the scores, ruining the schools. Westport, CT: Heinemann.

Kohn, A. (Program Consultant). (2000). Beyond the standards movement: Defending quality education in an age of test scores [Videotape]. Port Chester, NY: National Professional Resources, Inc.

INSTRUCTIONALLY FOCUSED TESTING TIPS

• Explain to colleagues and parents why standardized achieve-

ment tests’ traditional function to provide accurate norm-refer-

enced interpretations is dependent on sufficient score-spread

among students’ test performances.

• Describe to colleagues and parents how it is that three types of

score-spreading items (mid-difficulty items, SES-linked items, and

aptitude-linked items) reduce the suitability of traditionally con-

structed standardized achievement tests for evaluating instruc-

tional quality.

• Spend time reviewing an actual standardized achievement

test’s items to determine the proportion of items you regard as

unsuitable for determining what students were taught in school.

• Recognize that the descriptive information supplied with tradi-

tional standardized achievement tests does not describe the skills

and knowledge represented by those tests in a manner adequate

to support teachers’ instructional decision making.

ch9.qxd 7/30/2003 12:43 PM Page 137

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 3 8

Lemann, N. (2002). The big test: The secret history of the American meritocracy. New York: Farrar, Straus and Giroux.

Northwest Regional Educational Laboratory. (1991). Understanding standard- ized tests [Videotape]. Los Angeles: IOX Assessment Associates.

Popham, W. J. (Program Consultant). (2000). Standardized achievement tests: Not to be used in judging school quality [Videotape]. Los Angeles: IOX Assessment Associates.

Popham, W. J. (Program Consultant). (2002). Evaluating schools: Right tasks, wrong tests [Videotape]. Los Angeles: IOX Assessment Associates.

Sacks, P. (1999). Standardized minds: The high price of America’s testing culture and what we can do to change it. Cambridge, MA: Perseus Books.

ch9.qxd 7/30/2003 12:43 PM Page 138

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:13:57.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .