PLAGIARISM FREE "A" WORK 18 HOURS or LESS

profileNeNe1994
Chapter7ConstructedResponseItems.pdf

IN THIS CHAPTER, I’LL BE DESCRIBING THE TWO MOST COMMON TYPES OF

constructed-response items: short-answer items and essay items. Just as

in the previous chapter, I’ll briefly discuss the advantages and disad-

vantages of these two item-types and then provide a concise set of

item-construction rules for each. From there, I’ll address how to eval-

uate students’ responses to essay items or, for that matter, to any type

of constructed-response item. We’ll take a good look at how rubrics

are employed—not only to evaluate students’ responses, but also to

teach students how they should respond. Finally, we’ll consider per-

formance assessment and portfolio assessment, two currently popular

but fairly atypical forms of constructed-response testing.

Constructed-Response Assessment: Positives and Negatives The chief virtue of any type of constructed-response item is that it

requires students to create their responses rather than select a

prepackaged response from the answer shelf. Clearly, creating a

response represents a more complicated and difficult task. Some stu-

dents who might stumble onto a selected-response item’s correct

answer simply by casting a covetous eye over the available options

7 Constructed-Response Items

8 6

ch7.qxd 7/10/2003 9:58 AM Page 86

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 8 7

would never be able to concoct an original correct answer without

access to such options. Constructed-response items, in a word, are

tougher for test-takers. And, because a student really needs to under-

stand something in order to construct a response based on that

understanding, in many instances (but not all), students’ responses to

these sorts of items will better contribute to valid inferences than will

students’ answers to selected-response items.

On the downside, students’ answers to constructed-response

items take longer to score, and it’s also more of challenge to score

them accurately. In fact, the more complex the task that’s presented

to students in a constructed-response item, the tougher it is for teach-

ers to score. Whereas a teacher might have little difficulty accurately

scoring a student’s one-word response to a short-answer item, a

1,000-word essay is another matter entirely. The teacher might score

the long essay too leniently or too strictly. And if asked to evaluate

the same essay a month later, it’s very possible that the teacher might

give the essay a very different score.

As a side note, a decade or so ago, some states created elaborate

large-scale tests that incorporated lots of constructed-response items;

those states quickly learned how expensive it was to get those tests

scored. That’s because scoring constructed responses involves lots of

human beings, not just electrically powered machines. Due to such

substantial costs, most large-scale tests these days rely primarily on

selected-response items. In most of those tests, however, you will still

find a limited number of constructed-response items.

Clearly, there are pluses and minuses associated with the use of

constructed-response items. As a teacher, however, you can decide

whether you want to have your own tests include many, some, or no

constructed-response items. Yes, the scoring of constructed-response

items requires a greater time investment on your part. Still, I hope to

convince you that constructed-response items will help you arrive at

more valid inferences about your students’ actual levels of mastery.

What you’ll soon see, when we consider rubrics later in the chapter,

ch7.qxd 7/10/2003 9:58 AM Page 87

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R8 8

is that it’s possible to conceptualize constructed-response items and

the scoring for those items so that both you and your students will

reap giant instructional payoffs. If you want your students to master

truly powerful cognitive skills, you almost always need to rely on at

least a certain number of constructed-response items.

Let’s turn now to a closer inspection of the two most popular

forms of constructed-response items: short-answer items and essay

items.

Short-Answer Items A short-answer item requires a student to supply a word, a phrase, or

a sentence or two in response either to a direct question or an incom-

plete statement. A short-answer item’s only real distinction from an

essay item is in the brevity of the response the student is supposed to

supply.

Advantages and Disadvantages One virtue of short-answer items is that, like all constructed-response

items, they require students to generate a response rather than pluck-

ing one from a set of already-presented options. The students’ own

words can offer a good deal of insight into their understanding,

revealing if they are on the mark or conceptualizing something very

differently from how the teacher intended it to be understood. Short-

answer items have the additional advantage of being time-efficient;

students can answer them relatively quickly, and teachers can score

students’ answers relatively easily. Thus, as is true with binary-choice

and matching items, short-answer items allow test-making teachers

to measure a substantial amount of content. A middle school social

studies teacher interested in finding out whether students knew the

meanings of a set of 25 technical terms could present a term’s defini-

tion as the stimulus in each short-answer item, and then ask students

to supply the name of the defined term. (Alternately, that teacher

could present the names of terms and then ask for their full-blown

ch7.qxd 7/10/2003 9:58 AM Page 88

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 8 9

definitions, but then the students’ responses would be substantially

longer than “short”!)

An item asking for a full definition, of course, would still be

regarded as a short-answer item. The longer the short answer that’s

called for, the more class time the test will take up. However, the

more revealing the responses may be. The single-word response, 25-

item vocabulary test might take up 10 minutes of class time; the full-

definition version of all 25 words might take about 30 minutes. But

operationalizing a student’s vocabulary knowledge in the form of a

full-definition test would probably promote deeper understanding of

the vocabulary terms involved. Shorter short-answer items, therefore,

maximize the content-coverage value of this item-type, but also con-

tribute to the item-type’s disadvantages.

One of these disadvantages is that short-answer items tend to fos-

ter the student’s memorization of factual information, which is

another feature they have in common with binary-choice items.

Teachers must be wary of employing many assessment approaches

that fail to nudge students a bit higher on the cognitive ladder. Be

sure to balance your assessment approaches and remember that gains

in time and content coverage can also lead to losses in cognitive

challenge.

In addition, although short-answer items are easier to score than

essay items, they are still more difficult to score accurately than are

selected-response items. Suppose, for example, that one student’s

response uses a similar but different word or phrase than the one the

teacher was looking for. If the student’s response is not identical to

the teacher’s “preferred” short answer, is the student right or wrong?

And what about spelling? If a student’s written response to “Who, in

1492, first set foot on America?” is “Column-bus,” is this right or

wrong? As you see, even in their most elemental form, constructed-

response items can present teachers with some nontrivial scoring

puzzles.

ch7.qxd 7/10/2003 9:58 AM Page 89

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R9 0

Item-Writing Rules Again, my item-writing rules are ultra-short. If you wish to dig deeper

into the nuances of creating good short-answer items, be sure to con-

sult this chapter’s recommended resources section.

▲ Choose direct questions over incomplete statements.

There’s less chance of confusing students, especially the little ones in

grades K–3, if you write your short-answer items as direct questions

rather than as incomplete statements. You’d be surprised how many

times a teacher has one answer in mind for an incomplete statement,

but ambiguities in that incomplete statement lead students to come

up with answers that are dramatically different.

▲ Structure an item so that it seeks a brief, unique

response. There’s really a fair amount of verbal artistry associated

with the construction of short-answer items, for the item should elicit

a truly distinctive response that is also quite terse. I have a pair of

examples for you to consider. Whereas the first item might elicit a

considerable variety of responses from students, the second is far

more likely to elicit the term that the teacher had in mind.

AN EXCESSIVELY OPEN CONSTRUCTION

What one word can be used to represent a general truth or

principle?

(Answer: Maxim . . . or truism? Aphorism? Axiom? Adage?)

A CHARMINGLY CONSTRAINED CONSTRUCTION, OPAQUELY PHRASED

What one word describes an expression of a general truth

or principle, especially an aphoristic or sententious one?

(Answer: Maxim)

Incidentally, please observe that the second item, despite its worthy

incorporation of constraints, flat-out violates Chapter 5’s general

ch7.qxd 7/10/2003 9:58 AM Page 90

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 9 1

item-writing roadblock concerning difficult vocabulary. When was

the last time you used or heard anyone use the words “sententious”

or “aphoristic”? By following one rule, item-writers sometimes risk

breaking another. The trick, of course, is to figure out how to violate

no rules at all. Occasionally, this is almost impossible. We may not

live in a sin-free world, but we can at least strive to sin in moderation.

Similarly, try to break as few item-writing rules as you can.

▲ Place response-blanks at the end of incomplete state-

ments or, for direct questions, in the margins. If any of your

short-answer items are based on incomplete statements, try to place

your response-blanks near the end of the sentence. Early-on blanks

tend to confuse students. For direct questions, the virtue of placing

response-blanks in the margins is that it allows you to score the stu-

dents’ answers more efficiently: a quick scan, rather than a line-by-

line hunt.

▲ For incomplete statements, restrict the number of

blanks to one or two. If you use too many blanks in an incom-

plete statement, your students may find it absolutely incomprehensi-

ble. For obvious reasons, test-developers describe such multiblank

items as “Swiss Cheese” items. Think of the variety of responses stu-

dents could supply for the following, bizarrely ambiguous example:

A SWISS CHEESE ITEM

Following a decade of Herculean struggles, in the year

, and his partner, ,

finally discovered how to make .

(Answer: ???)

▲ Make all response-blanks equal in length. Novice item-

writers often vary the lengths of their response-blanks to better

match the correct answers they’re seeking. It just looks nicer that

way. To avoid supplying your students with unintended clues, always

stick with response blanks of uniform length.

ch7.qxd 7/10/2003 9:58 AM Page 91

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R9 2

▲ Supply sufficient answer space. Yes, the idea is to solicit

short answers, but be sure you give your students adequate room to

provide the sought-for word or phrase. Don’t forget to factor in the

size of students’ handwriting. Your average 2nd grader, for example,

will need more space per word than your average 9th grader.

Essay Items Essay items have been around for so long that it’s a good bet Socrates

once asked Plato to describe Greek-style democracy in an essay of at

least 500 words. I suspect that Plato’s essay would have satisfied any

of today’s content standards dealing with written communication,

not to mention a few content standards dealing with social studies.

And that’s one reason essay items have endured for so long in our

classrooms: They are very flexible and can be used to measure stu-

dents’ mastery of all sorts of truly worthwhile curricular aims.

Advantages and Disadvantages As an assessment tactic to measure truly sophisticated types of stu-

dent learning, essay items do a terrific job. If you teach high school

biology and you want your students to carry out a carefully reasoned

analysis of creationism versus evolutionism as alternative explana-

tions for the origin of life, asking students to respond to an essay item

can put them, at least cognitively, smack in the middle of that com-

plex issue.

Essay items also give students a great opportunity to display their

composition skills. Indeed, for more than two decades, most of the

statewide tests taken by U.S. students have required the composition

of original essays on assigned topics such as “a persuasive essay on

Topic X” or “a narrative essay on Topic Y.” Referred to as writing sam-

ples, these essay-writing tasks have been instrumental in altering the

way that many teachers provide composition instruction.

Essay items have two real shortcomings: (1) the time required to

score students’ essays and (2) the potential inaccuracies associated with

ch7.qxd 7/10/2003 9:58 AM Page 92

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 9 3

that scoring. There’s really no way for classroom teachers to get around

the time-requirement problem, but recent research in the electronic

scanning and scoring of students’ essays for large-scale assessments

suggests that we are getting quite close to computer-based scoring of

essays. With respect to the second problem, inaccurate scoring, we

have learned from efforts to evaluate essays in statewide and national

tests that it is possible to do so with a remarkable degree of accuracy,

provided that sufficient resources are committed to the effort and

well-trained scoring personnel are in place. And thanks to the use of

rubrics, which I’ll get to in a page or two, even classroom teachers can

go a long way toward making their own scoring of essays more pre-

cise. It can be tricky, though, as students’ compositional abilities can

sometimes obscure their mastery of the curricular aim the item is

intended to reveal. For example, Shelley knows the content, but she

cannot express herself, or her knowledge, very well in writing; Dan,

on the other hand, doesn’t know the content but can fake it through

very skillful writing. Rubrics are one way to keep a student’s ability or

inability to spin out a compelling composition from deluding you

into an incorrect inference about that student’s knowledge or skill.

Item-Writing Rules There are five rules that classroom teachers really need to follow

when constructing essay items.

▲ Structure items so that the student’s task is explicitly

circumscribed. Phrase your essay items so that students will have no

doubt about the response you’re seeking. Don’t hesitate to add details

to eliminate ambiguity. Consider the following two items. One of them

is likely to leave far too much uncertainty in the test-taker’s mind and,

as a result, is unlikely to provide evidence about the particular educa-

tional variables the item-writer is (clumsily) looking to uncover.

A DREADFULLY AMBIGUOUS ITEM

Discuss youth groups in Europe.

ch7.qxd 7/10/2003 9:58 AM Page 93

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R9 4

A DELIGHTFULLY CIRCUMSCRIBED ITEM

Describe, in 400–600 words, how the rulers of Nazi Germany in

the 1930s used the Hitler Youth Movement to solidify the Nazi

political position.

▲ For each question, specify the point value, an accept-

able response-length, and a recommended time allocation.

What this second rule tries to do is give students the information

they need to respond appropriately to an essay item. The less guess-

ing that your students are obliged to do about how they’re supposed

to respond, the less likely it is that you’ll get lots of off-the-wall essays

that don’t give you the evidence you need. Here’s an example of how

you might tie down some otherwise imponderables and give your

students a better chance to show you what you’re interested in find-

ing out.

SUITABLE SPECIFICITY

What do you believe is an appropriate U.S. policy regarding the

control of global warming? Include in your response at least

two concrete activities that would occur if your recommended

policy were implemented. (Value: 20 points; Length: 200–300

words; Recommended Response Time: 20 minutes)

▲ Employ more questions requiring shorter answers

rather than fewer questions requiring longer answers. This

rule is intended to foster better content sampling in a test’s essay

items. With only one or two items on a test, chances are awfully good

that your items may miss your students’ areas of content mastery or

nonmastery. For instance, if a student hasn’t properly studied the

content for one item in a three-item essay test, that’s one-third of the

exam already scuttled.

ch7.qxd 7/10/2003 9:58 AM Page 94

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 9 5

▲ Don’t employ optional questions. I know that teachers

love to give students options because choice tends to support positive

student engagement. Yet, when students can choose their essay items

from several options, you really end up with different tests, unsuitable

for comparison. For this reason, I recommend sticking with non-

optionality. It is always theoretically possible for teachers to create two

or more “equidifficult” items to assess mastery of the same curricular

aim. But in the real word of schooling, a teacher who can whip up two

or more constructed-response items that are equivalent in difficulty is

most likely an extraterrestrial in hiding. Items that aren’t really equi-

difficult lead to students taking on noncomparable challenges.

▲ Gauge a question’s quality by creating a trial response

to the item. A great way to determine if your essay items are really

going to get at the responses you want is to actually try writing a

response to the item, much as a student might do. And you can

expect that kind of payoff even if you make your trial response only

“in your head” instead of on paper.

Rubrics: Not All Are Created Equal I have indicated that a significant shortcoming of constructed-

response tests, and of essay items in particular, is the potential inac-

curacy of scoring. A properly fashioned rubric can go a long way

toward remedying that deficit. And even more importantly, a properly

fashioned rubric can help teachers teach much more effectively and

help students learn much more effectively, too. In short, the instruc-

tional payoffs of properly fashioned rubrics may actually exceed their

assessment dividends. However, and this is an important point for you

to remember, not all rubrics are properly fashioned. Some are wonderful,

and some are worthless. You must be able to tell the difference.

What Is a Rubric, Anyway? A rubric is a scoring guide that’s intended to help those who must

score students’ responses to constructed-response items. You might

ch7.qxd 7/10/2003 9:58 AM Page 95

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R9 6

be wondering why it is that measurement people chose to use a cryp-

tic word such as “rubric” instead of the more intuitively understand-

able “scoring guide.” My answer is that all specialist fields love to

have their own special terminology, and “rubric” is sufficiently

opaque to be a terrific turn-on to the measurement community.

The most important component in any rubric is the rubric’s eval-

uative criteria, which are the factors a scorer considers when deter-

mining the quality of a student’s response. For example, one evalua-

tive criterion in an essay-scoring rubric might be the organization of

the essay. Evaluative criteria are truly the guts of any rubric, because

they lay out what it is that distinguishes students’ winning responses

from their weak ones.

A second key ingredient in a rubric is a set of quality definitions

that accompany each evaluative criterion. These quality definitions

spell out what is needed, with respect to each evaluative criterion, for

a student’s response to receive a high rating versus a low rating on

that criterion. A quality definition for an evaluative criterion like

“essay organization” might identify the types of organizational struc-

tures that are acceptable and those that aren’t. The mission of a

rubric’s quality definitions is to reduce the likelihood that the rubric’s

evaluative criteria will be misunderstood.

Finally, a rubric should set forth a scoring strategy that indicates

how a scorer should use the evaluative criteria and their quality defi-

nitions. There are two main contenders, as far as scoring strategy is

concerned. A holistic scoring strategy signifies that the scorer must

attend to how well a student’s response satisfies all the evaluative cri-

teria in the interest of forming a general, overall evaluation of the

response based on all criteria considered in concert. In contrast, an

analytic approach to scoring requires a scorer to make a criterion-by-

criterion judgment for each of the evaluative criteria, and then amal-

gamate those per-criterion ratings into a final score (this is often done

via a set of predetermined, possibly numerical, rules).

ch7.qxd 7/10/2003 9:58 AM Page 96

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 9 7

The most widespread use of rubrics occurs with respect to the scor-

ing of students’ writing samples. In large-scale scoring operations, holis-

tic scoring is usually the preferred approach because it is faster and,

hence, less expensive. However, analytic scoring is clearly more diag-

nostic, as it permits the calculation of a student’s per-criterion perform-

ance. Some states keep costs down by scoring all student writing sam-

ples holistically first, then scoring only “failing” responses analytically.

Early rubrics used by educators in the United States were intended

to help score students’ writing samples. Although there were surely

some qualitative differences among those rubrics, for the most part

they all worked remarkably well. Not only could different scorers

score the same writing samples and come up with similar appraisals,

but the evaluative criteria in those early rubrics were instructionally

addressable. Thus, teachers could use each evaluative criterion to

teach their students precisely how their writing samples would be

judged. Indeed, students’ familiarity with the evaluative criteria con-

tained in those writing sample rubrics meant they could critique their

own compositions and those of their classmates as a component of

instruction.

To this day, most writing sample rubrics contribute wonderfully

to accurate scoring and successful instruction. Unfortunately, many

teachers have come to believe that any rubric will lead to those twin

dividends. As you’re about to see, that just isn’t so.

Reprehensible Rubrics I’ve identified three kinds of rubrics that are far from properly fash-

ioned if the objective is both evaluative accuracy and instructional

benefit. They are (1) task-specific rubrics, (2) hypergeneral rubrics, and

(3) dysfunctionally detailed rubrics; each of these reprehensible rubrics

takes its name chiefly from the type of evaluative criteria it contains.

Let’s take a closer look to see where the failings lie.

Task-specific rubrics. This type of rubric contains evaluative crite-

ria that refer to the particular task that the student has been asked to

ch7.qxd 7/10/2003 9:58 AM Page 97

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R9 8

perform, not the skill that task was created to represent. For example,

one pivotal skill in science class might be the ability to understand

and then explain how certain scientifically rooted processes function.

If an essay item in a science class asked students to explain how a vac-

uum thermos works, a task-specific rubric’s evaluative criteria would

help only in the scoring of responses to that specific vacuum thermos–

explaining task. The rubric’s evaluative criteria, focused exclusively

on the particular task, would have no relevance to the evaluation of

students’ responses to other tasks representing the cognitive skill

being measured. Task-specific rubrics don’t help teachers teach and

they don’t help students learn. They are way too specific!

Hypergeneral rubrics. A second type of reprehensible rubric is one

whose evaluative criteria are so very, very general that the only thing

we really know after consulting the rubric is something such as “a

good response is . . . good” and “a bad response is the opposite”! To

illustrate, in recent years, I have seen many hypergeneral rubrics that

describe “distinguished” student responses in a manner similar to

this: “A complete and fully accurate response to the task posed, pre-

sented in a particularly clear and well-written manner.” Hypergeneral

rubrics provide little more clarity to teachers and students than

what’s generally implied by an A through F grading system. Hyper-

general rubrics don’t help teachers teach, and they don’t help stu-

dents learn. They are way too general!

Dysfunctionally detailed rubrics. Finally, there are the rubrics that

have loads and loads of evaluative criteria, each with its own

extremely detailed set of quality descriptions. These rubrics are just

too blinking long. Because few teachers and students have the

patience or time to wade through rubrics of this type, these rubrics

don’t help teachers teach or students learn. They are, as you have

already guessed, way too detailed!

A Skill-Focused Rubric: The Kind You Want For accurate evaluation of constructed-response tests and to provide

instructional benefits, you need a rubric that deals with the skill being

ch7.qxd 7/10/2003 9:58 AM Page 98

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 9 9

tapped by that test. These skill-focused rubrics are exemplified by the

early rubrics used to score U.S. students’ writing samples. Here’s what

you’ll find in a properly designed skill-focused rubric:

▲ It includes only a handful of evaluative criteria. Skill-

focused rubrics incorporate only the most important evaluative factors

to be used in appraising a student’s performance. This way, teachers and

students can remain attentive to a modest number of super-important

evaluative criteria, rather than being overwhelmed or sidetracked.

▲ It reflects something teachable. Teachers need to be able

to teach students to master an important skill. Thus, each evaluative

criterion on a good skill-focused rubric should be included only after

an affirmative answer is given to this question: “Can I get my stu-

dents to be able to use this evaluative criterion in judging their own

mastery of the skill I’m teaching them?”

▲ All the evaluative criteria are applicable to any skill-

reflective task. If this condition is satisfied, of course, it is impos-

sible for the rubric to be task-specific. For example, if a student’s

impromptu speech is being evaluated with a skill-focused rubric, the

evaluative criteria of “good eye contact” will apply to any topic the

student is asked to speak about, whether it’s “My Fun Summer” or

“How to Contact the City Council.”

▲ It is concise. Brevity is the best way to ensure that a rubric is

read and used by both teachers and students. Ideally, each evaluative

criterion should have a brief, descriptive label. And, as long as we’re

going for ideal, it usually makes sense for the teacher to conjure up a

short, plain-talk student version of any decent skill-focused rubric.

Illustrative Evaluative Criteria If you’ve put two and two together, you’ve come to the conclusion

that a rubric’s instructional quality usually hinges chiefly on the way

that its evaluative criteria are given. Figure 7.1 clarifies how it is that

those criteria can help or hinder a teacher’s instructional thinking.

In it, you’ll see three different types of evaluative criteria: one crite-

rion each from a hypergeneral rubric, a task-specific rubric, and a

ch7.qxd 7/10/2003 9:58 AM Page 99

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 0 0

skill-focused rubric. All three focus on the same variable: students’

organization of a narrative essay.

Notice that the hypergeneral evaluative criterion offers little

insight into what the organization of the student’s essay must be like

other than “good.” Instructionally, that’s not much to go on. The

task-specific evaluative criterion would work well for evaluating

students’ narrative essays about this particular firefighters’ visit to

class, but how many times would a student need to write a narrative

7 . 1 ILLUSTRATIVE EVALUATIVE CRITERIA FOR THREE KINDS OF RUBRICS

Introduction: A 6th grade class (fictional) has been visited by a group of local fire- fighters. The teacher asks the students to compose a narrative essay recounting the previous day’s visit. Organization is one of the key evaluative criteria in the rubric that the teacher routinely uses to appraise students’ compositions. Presented below are three ways of describing that criterion:

A Hypergeneral Rubric Organization: “Superior” essays are those in which the essay’s content has been arranged in a genuinely excellent manner, whereas “inferior” essays are those that display altogether inadequate organization. An “adequate” essay is one that repre- sents a lower organizational quality than a superior essay, but a higher organiza- tional quality than an inferior essay.

A Task-Specific Rubric Organization: “Superior” essays will (1) commence with a recounting of the partic- ular rationale for home fire-escape plans that the local firefighters presented, then (2) follow up with a description of the six elements in home-safety plans in the order that those elements were described, and (3) conclude by citing at least three of the life-death safety statistics the firefighters provided at the close of their class- room visit. Departures from these three organizational elements will result in lower evaluations of essays.

A Skill-Focused Rubric Organization: Two aspects of organization will be employed in the appraisal of stu- dents’ narrative essays, namely, overall structure and sequence. To earn maximum credit, an essay must embody an overall structure containing an introduction, a body, and a conclusion. The content of the body of the essay must be sequenced in a reasonable manner, for instance, in a chronological, logical, or order-of- importance sequence.

ch7.qxd 7/10/2003 9:58 AM Page 100

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 1 0 1

essay on that specific topic? Finally, the skill-focused evaluative crite-

rion isolates two organizational features: a reasonable sequence and

the need for an “introduction-body-conclusion” structure. Those two

organizational elements are not only teachable, but they are also

applicable to a host of narrative essays.

In short, you’ll find that skill-focused rubrics can be used not only

to score your students’ responses to constructed-response tests, but

also to guide both you and your students toward students’ mastery of

the skill being assessed. Skill-focused rubrics help teachers teach; they

help students learn. They go a long way toward neutralizing the

drawbacks of constructed-response assessments while maximizing

the advantages.

Special Variants of Constructed-Response Assessment Before we close the primer on constructed-response tests, I want to

address a couple of alternative forms of constructive response evalu-

ation, both of which have become increasingly prominent in recent

years: performance assessment and portfolio assessment.

Performance Assessment Educators differ in their definitions of what a performance assess-

ment is. Some educators believe that any kind of constructed response

constitutes a performance test. Those folks are in the minority. Most

educators regard performance assessment as an attempt to measure a

student’s mastery of a high-level, sometimes quite sophisticated skill

through the use of fairly elaborate constructed-response items and a

rubric. Performance tests always present a task to students. In the case

of evaluating students’ written communication skills, that task might

be to “write a 500-word narrative essay.” (Assessment folks call the

tasks on writing tests prompts.) And the more that the student’s assess-

ment task resembles the tasks to be performed by people in real life,

the more likely it is that the test will be labeled a performance

assessment.

ch7.qxd 7/10/2003 9:58 AM Page 101

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 0 2

An example of a performance assessment in social studies might

be for each student to (1) come up with a current-day problem that

could be solved by drawing directly on historical lessons; (2) prepare

a 3,000-word written report describing the applicability of the histor-

ical lessons to the current-day problem; and finally, (3) present a 10-

minute oral report in class describing the student’s key conclusions.

No one would dispute that this sort of a problem presents students

with a much more difficult task to tackle than choosing between true

and false responses in a test composed of binary-choice items.

There are a couple of drawbacks to performance assessments, one

obvious and the other more subtle. The obvious one is time. This

kind of assessment requires a significant time investment from teach-

ers (designing a multistep performance task, developing an evaluative

rubric, and applying the evaluative rubric to the performance) and

from students (preparing for the performance and the performance

itself). A less obvious but more serious obstacle for those who advo-

cate the use of performance assessment is the problem of generaliz-

ability. How many performance tests do students need to complete

before the teacher can come up with valid inferences about their gen-

eralizable skill-mastery? Will one well-written persuasive essay be

enough to determine the degree to which a particular student has

mastered the skill of writing persuasive essays? How many reports

about the use of history-based lessons are needed before a teacher can

say that a student really is able to employ lessons from past historical

events to deal with current-day problems?

The research evidence related to this issue is not encouraging.

Empirical investigations suggest that teachers would be obliged to

give students many performance assessments before arriving at a truly

accurate interpretation about a student’s skill-mastery (Linn &

Burton, 1994). “Many performance assessments” translates into even

more planning and classroom time. On practical grounds, therefore,

teachers need to be very selective in their use of this exciting, but

time-consuming assessment approach. I recommend reserving

ch7.qxd 7/10/2003 9:58 AM Page 102

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 1 0 3

performance assessments for only the most significant of your high-

priority curricular aims.

Portfolio Assessment Let’s close out the chapter with a quick look at another special sort of

constructed-response measurement: portfolio assessment. A portfolio is

a collection of one’s work. Essentially, portfolio assessment requires

students to continually collect and evaluate their ongoing work for

the purpose of improving the skills they need to create such work.

Language arts portfolios, usually writing portfolios or journals, are

the example cited most often; however, portfolio assessment can be

used to assess the student’s evolving mastery of many kinds of skills

in many different subject areas.

Suppose Mr. Miller, a 5th grade teacher, is targeting his students’

abilities to write brief multiparagraph compositions. Mr. Miller sets

aside a place in the classroom (perhaps a file cabinet) where students,

using teacher-supplied folders, collect and critique the compositions

they create throughout the entire school year. Working with his 5th

graders to decide what the suitable qualities of good compositions are,

Mr. Miller creates a skill-focused rubric students will use to critique

their own writing. When students evaluate a composition, they attach

a dated, filled-in rubric form. Students also engage in a fair amount of

peer critique, using the same skill-focused rubric to review each other’s

compositions. Every month, Mr. Miller holds a brief portfolio confer-

ence with each student to review the student’s current work and to

decide collaboratively on directions for improvement. Mr. Miller also

calls on parents to participate in those portfolio conferences at least

once per semester. The emphasis is always on improving students’ abil-

ities to evaluate their own multiparagraph compositions in these

“working” portfolios. At the end of the year, each student selects a

best-work set of compositions, and these “showcase” portfolios are

then taken home to students’ parents.

ch7.qxd 7/10/2003 9:58 AM Page 103

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 0 4

As you can see, given such a continuing emphasis on the

enhancement of students’ self-appraisal skills, there is a considerable

likelihood that portfolio assessment will yield major payoffs for stu-

dents. And teachers who have used portfolios agree with their power

as a combined assessment and instruction strategy.

But those same teachers, if they are honest, will also point out

that staying abreast of students’ portfolios and carrying out periodic

portfolio conferences is really time-consuming. Be sure to consider

the time-consumption drawback carefully before diving headfirst

into the portfolio pool.

Recommended Resources

Andrade, H. G. (2000, February). Using rubrics to promote thinking and learning. Educational Leadership, 57(5), 13–18.

Eisner, E. W. (1999, May). The uses and limits of performance assessment. Phi Delta Kappan, 80(9), 658–660.

Letts, N., Kallick, B., Davis, H. B., & Martin-Kniep, G. (Presenters). (1999). Portfolios: A guide for students and teachers [Audiotape]. Alexandria, VA: Association for Supervision and Curriculum Development.

INSTRUCTIONALLY FOCUSED TESTING TIPS

• Understand the relative strengths of constructed-response

items and selected-response items.

• Become familiar with the advantages and disadvantages of

short-answer and essay items.

• Make certain to follow experience-based rules for creating short-

answer and essay items.

• When scoring students’ responses to constructed-response

tests, employ skill-focused scoring rubrics rather than rubrics that

are task specific, hypergeneral, or dysfunctionally detailed.

• Recognize both the pros and cons associated with performance

testing and portfolio assessment.

ch7.qxd 7/10/2003 9:58 AM Page 104

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o n s t r u c t e d - R e s p o n s e I t e m s 1 0 5

Linn, R. L., & Burton, E. (1994). Performance-based assessment: Implications of task specificity. Educational measurement: Issues and practice, 13(1), 5–8, 15.

Mabry, L. (1999, May). Writing to the rubric: Lingering effects of traditional standardized testing on direct writing assessment. Phi Delta Kappan, 80(9), 673–679.

McMillan, J. H. (2001). Classroom assessment: Principles and practice for effective instruction (2nd ed.). Boston: Allyn & Bacon.

Northwest Regional Educational Laboratory. (1991). Developing assessments based on observation and judgment [Videotape]. Los Angeles: IOX Assessment Associates.

Pollock, J. (Presenter). (1996). Designing authentic tasks and scoring rubrics [Audiotape]. Alexandria, VA: Association for Supervision and Curriculum Development.

Popham, W. J. (Program Consultant). (1996). Creating challenging classroom tests: When students CONSTRUCT their responses [Videotape]. Los Angeles: IOX Assessment Associates.

Popham, W. J. (Program Consultant). (1998). The role of rubrics in classroom assessment [Videotape]. Los Angeles: IOX Assessment Associates.

Stiggins, R. J. (Program Consultant). (1996). Assessing reasoning in the classroom [Videotape]. Portland, OR: Assessment Training Institute.

Wiggins, G., Stiggins, R. J., Moses, M., & LeMahieu, P. (Program Consultants). (1991). Redesigning assessment: Portfolios [Videotape]. Alexandria, VA: Association for Supervision and Curriculum Development.

ch7.qxd 7/10/2003 9:58 AM Page 105

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:02:36.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .