PLAGIARISM FREE "A" WORK 18 HOURS or LESS

profileNeNe1994
Chapter11CollectingCredibleClassroomEvidenceofInstructional.pdf

1 4 9

AS I’VE NOTED, THE CHIEF REASON THAT THE DESIGNERS OF EDUCATIONAL

accountability systems rely on standardized tests is that those design-

ers don’t trust educators to provide truly accurate evidence regarding

their own classroom successes or failures. And let’s be honest: That’s

not an absurd reason for incredulity. No one really likes to be evalu-

ated and found wanting. The architects of educational accountability

systems typically make students’ standardized test scores the center-

piece of their evaluative activities because it’s accepted that such tests

provide more believable evidence.

However, because no single source of data should ever be

employed to make significant decisions about schools, teachers, or

students, even standardized test scores should be supplemented by

other evidence of teachers’ instructional success. This supplemental

evidence of a teacher’s instructional impact must clearly be credible. If

evidence of effective instruction is not truly credible, well, no one

will believe it . . . and it won’t do any good.

A Post-Test–Focused Data-Gathering Design You’ll recall from Chapter 1 that one of the major drawbacks of using

data from once-a-year standardized achievement test (even an

11 Collecting Credible

Classroom Evidence of Instructional Impact

ch11.qxd 7/30/2003 2:06 PM Page 149

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 5 0

instructionally supportive one) to see how well a teacher has taught

is that the quality of students in a given teacher’s class varies from

year to year. If a 5th grade teacher’s class this year is brimming with

bright, motivated children, those students are apt to perform pretty

well on this year’s standardized test. But if a year later, the same

teacher draws a markedly less able collection of 5th graders, the stu-

dents’ test scores are apt to be much lower. Does this year-to-year

drop in 5th graders’ scores indicate that the teacher’s instructional

effectiveness had decreased? Of course not. The students were different.

The key to getting a more accurate fix on a teacher’s true instruc-

tional impact is to collect evidence from the same students while they

remain with the same teacher. One way to do this, represented in

Figure 11.1, is to test the teacher’s students after instruction is fin-

ished. This is called a post-test–only design and its chief weakness, for

teachers who would employ it, is that it doesn’t factor in the stu-

dents’ pre-instruction status. If your students score superbly on the

post-instruction test, is it because you taught them well or because

the tested material was something they all learned years ago? With no

basis for comparison, the post-test–only design flops if you’re trying

to isolate the impact of your instruction.

1 1 . 1 A POST-TEST–O NLY DATA-GATHERING DESIGN

ch11.qxd 7/30/2003 2:06 PM Page 150

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 1

The Pretest/Post-Test Data-Gathering Design The most common data-gathering model teachers use to get a picture

of how well they have taught their students is a straightforward

pretest/post-test model such as the one presented graphically in Figure

11.2. (You also saw this data-gathering model in Chapter 1.) Most of

us are familiar with how this works: a teacher gives students the same

test, once before instruction and again after. Elementary teachers, for

example, might administer a 20-item test at the start of a school year

and the same test again at year’s end. Secondary teachers in schools

on a semester system might administer a test at the semester’s outset

and then again at its conclusion.

The virtue of the classic pretest/post-test evaluative model is that,

for the most part, it does measure the same group of students before

instruction and after, meaning that a comparative analysis of the two

sets of test data provides a clearer picture of the teacher’s instruction-

al impact on student mastery levels than do post-test data alone. In

classrooms with high student mobility, it is usually sensible to com-

pare pretest and post-test performances of only those students who

have been in the class long enough for the relevant instruction to

“work.” Thus, whereas all students might take a post-instruction

1 1 . 2 A PRETEST/POST-TEST DATA-GATHERING DESIGN

ch11.qxd 7/30/2003 2:06 PM Page 151

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 5 2

assessment, for purposes of appraising instructional effectiveness, you

would analyze the post-tests scores of just the students who had been

enrolled in the class for, say, eight weeks or whatever instructional

period seems sufficient.

There are problems with the pretest/post-test model that reduces

its accuracy. One major drawback that occurs when a teacher uses the

same test before and after instruction is referred to as pretest reactivi-

ty. What this means is that students’ experience taking a pretest will

often make them react differently to the post-test. They have been

sensitized to “what’s coming.” A skeptic might conclude that even if

post-test scores went up, students didn’t necessary learn anything;

perhaps they simply figured out how to take the test the second time

through. Unfortunately, the skeptic might be right.

Well, how about using different assessments for the pretest and

the post-test? It sounds like a good idea, and it definitely prevents

pretest reactivity, but the truth is that it is much more difficult than

most people realize to create two genuinely equidifficult tests. And

unless the tests are identical in difficulty, all sorts of evaluative crazi-

ness can ensue. If the post-test is the tougher of the two tests, then

the post-test scores have a good chance of being lower, and the

teacher’s instruction will look ineffective no matter what. On the

other hand, if the pretest is much more difficult than the post-test,

then the teacher will appear to be successful even when that’s not

true. No, there are definite difficulties associated with a teacher’s use

of the rudimentary kind of pretest/post-test model seen in Figure

11.2. But this model can be improved.

A Better Source of Evidence: The Split-and-Switch Data-Gathering Design Let me describe a way to collect instruction data for a pretest versus

post-test comparison that skirts the difficulties associated with the

standard pretest/post-test design. I call it the split-and-switch model,

and here’s how to use it in your classroom.

ch11.qxd 7/30/2003 2:06 PM Page 152

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 3

▲ Choose a skill. Select an important skill that your students

should acquire as a consequence of your instruction over a substan-

tial period of time (a semester or school year). The more important

the skill being taught and tested, the more impressive will be any evi-

dence of your instructional impact.

▲ Choose your assessment type. I recommend you opt for a

constructed-response test, which is almost always the best choice to

assess important skills. Performance tests are often ideal. Don’t forget,

though, that if you choose a constructed-response format, you’ll also

need to design a rubric to evaluate students’ responses.

▲ Create two assessments. Now, instead of designing one

assessment to measure your students’ mastery of the skill, design

two—two versions that measure the same variable. Let’s call them

Form A and Form B. The two forms should require approximately the

same administration time and, although they should be similar in

difficulty, they need not be identical in difficulty.

▲ Create two test groups. Randomly divide your class into

two halves. You could simply use your alphabetical class list and split

the class into two half-classes: A–L last names and M–Z last names.

Let’s call these Half-Class 1 and Half-Class 2.

▲ Pretest, teach, and post-test. Finally, you’ll come to the

testing. As a pretest, administer one test form to one test group, and

the other test form to the other test group (for example, Half-Class 1

takes Form A and Half-Class 2 takes Form B). As you might have

already guessed, at the end of instruction, you will post-test by switch-

ing the two test forms so that each test group gets the test form it

didn’t get at pretest-time.

▲ Compare data sets. You’ll end up with two sets of compara-

tive data to consider: the Form A pretests versus the Form A post-tests

and the Form B pretests versus the Form B post-tests. The split-and-

switch design is depicted graphically in Figure 11.3, where you can

see the two appropriate comparisons to make.

ch11.qxd 7/30/2003 2:06 PM Page 153

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 5 4

What you want to see, of course, are results indicating that the

vast majority of high test performances were produced at post-test-

time. There will be no pretest reactively because each half-class gets

what is to them a brand new test form as a post-test. And because this

data-gathering design yields two separate indications of a teacher’s

instructional impact, the results (positive or negative) ought to con-

firm one another. Remember, even if Form A is very tough and Form

B is not, you’re comparing the tough Form A pretests with the tough

Form A post-tests, and the easy Form B post-tests with the easy Form

B post-tests.

Credibility-Enhancing Measures Now, although you could use this model as is, without any refine-

ments, I don’t think you should. Here’s what I recommend.

1 1 . 3 THE SPLIT-AND-SWITCH DATA-GATHERING DESIGN

ch11.qxd 7/30/2003 2:06 PM Page 154

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 5

▲ “Blind score” the pretests and post-tests. The rationale

behind blind scoring is to counteract every teacher’s understandable

desire to see improvement even if it may not be in evidence. When

creating your two test forms, also create identical response-booklets

that you’ll pass out to all students, both at pretest and post-test time.

Tell students not to put dates on their response booklets. When you

collect the completed response booklets at pretest time, code them on

the reverse side of the last sheet so that it’s possible, but not easy, to

tell that the responses are pretests. You might use a short string of

numbers, and make the second number an odd number (1, 3, 5, 7, or

9). Go ahead and look at your students’ pretest responses if you’d like

to get an idea of your students’ skill-levels at the beginning of instruc-

tion, but do not make any marks on their test forms. File the forms

away.

At the end of the semester or school year, conduct a similar cod-

ing procedure when you collect the post-tests, but make the second

number in your code string an even number. Again, you can use these

post-tests as you would any other (to make inferences about what

your students have learned), but if you use post-tests for grading pur-

poses, be sure to take into consideration any difficulty-level differ-

ences you may have detected between Form A and Form B.

Finally, when it comes time to compare the two sets of data, mix

all the Form A responses together (Form A pretests and post-tests). Do

the same for the Form B responses. After you’ve scored them, use the

codes to separate out your pretests and post-tests, and then see what

the data have to say about your instructional effectiveness.

▲ Bring in nonpartisan judges. You can enhance the credi-

bility of your evidence even further by turning over pretest/post-test

evaluations to someone else, thereby ensuring undistorted judgment.

(If you’ve spent some time with your students’ pretests to help you

determine entry-level skills or with the post-tests for grading purpos-

es, you’ll probably recognize which are which, despite your best

efforts.)

ch11.qxd 7/30/2003 2:06 PM Page 155

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

T E S T B E T T E R , T E A C H B E T T E R1 5 6

Parent volunteers make great nonpartisan blind-scorers. Provide

them with an evaluation rubric and a brief orientation to its use. The

judges will get all the Form A responses (both pretests and post-tests

thoroughly mixed together) and then score the responses. After all

the Form A responses have been scored and returned to you, use your

coding system to sort the responses into pretests and post-tests. Then,

it’s just a matter of comparing your students’ pretest responses with

their post-test responses. Use the same process for blind-scoring the

Form B tests.

If you’ve used nonpartisan judges and they have blind-scored

your students’ papers properly, the results ought to be truly credible

to the accountability minded. The results should help you, and the

external world, know if your instruction is working.

Caveats for Using the Model Of course, even the best-laid evaluative schemes can be foiled. The

major weakness in the split-and-switch model is that some teachers

might teach toward the test forms themselves, thereby eroding the

validity of any test-based inferences about students’ true skill mastery.

Always teach only toward the skill being promoted; never teach

toward the tests being used to measure skill-mastery.

Another potential problem is that, in order to produce stable

data, the split-and-switch model requires a relatively large group of

students. A class of around 30 is ideal. If your class is much smaller

(say, 20 students or less), it’s best to revert to a standard pretest/post-

test model, even with its drawbacks. You can still employ the kind of

nonpartisan blind-scoring process I’ve described, which will increase

the likelihood of your pretest/post-test data being regarded as

credible.

Though not flawless, the split-and-switch design can help you

buttress the evidence of your instructional effectiveness in a more

believable manner. And, of course, if this test-based evidence suggests

that your instruction isn’t working as well as you wish, you will need

ch11.qxd 7/30/2003 2:06 PM Page 156

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 7

to address that shortcoming through the use of alternative instruc-

tional strategies.

I urge you to consider employing a split-and-switch design to

evaluate your effectiveness in promoting your students’ mastery of a

very limited number of high-import outcomes. Let’s be honest.

Implementation of this data-gathering design takes time and is clear-

ly more trouble than conventional testing. Use it judiciously, but do

use it. I think you’ll like it.

Comparisons of “Pre-” and “Post-” Affective Data Pretest/post-test contrasts of students’ attitudes, interests, or values as

measured on self-report inventories are a final source of useful credi-

ble evidence regarding a teacher’s own instructional effectiveness.

Remember, if you will, the confidence inventories described in

Chapter 8 and the positive relations between students’ expressed con-

fidence in being able to use a skill and their actual possession of that

skill. If you can assemble pretest/post-test evidence, collected anony-

mously, indicating that your students’ confidence has grown appre-

ciably with respect to their ability to perform significant cognitive

skills, these results will reflect very favorably on your instruction.

If you do decide to collect pretest/post-test affective evidence as

an indication of your own instructional success, remember that you

need to collect it in a manner that even nonbelievers will regard as

credible. Just spend a few mental moments, casting yourself in the

role of a doubter. Pretend you are a person who has serious doubts

about whether the evidence you’ll be presenting about your own

instructional effectiveness is believable. Do whatever you can, in the

collection of any evaluative evidence, to allay the concerns that such

a doubter might have. Happily, a split-and-switch design, especially if

students’ responses are blind-scored by nonpartisans, will take care of

most of a disbeliever’s concerns.

If you are worried about pretest reactivity—and such reactivity

often is a concern when assessing affect—you can employ a split-and-

ch11.qxd 7/30/2003 2:06 PM Page 157

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

switch design, simply chopping your affective inventories into two

half-inventories so that students will get different inventories as

pretests and post-tests.

Final Thoughts There are a number of ways that properly constructed educational

tests can contribute to the quality of a teacher’s instructional deci-

sions. In my opinion, however, there is no more important contribu-

tion than helping a teacher supply an answer to the question “How

effective was my instruction?”

There are also a number of ways to judge a teacher’s success.

Many factors ought to be considered, some of which are not test-

based in any way. We can gain insights into a school staff’s effective-

ness by considering the school’s statistics regarding truancy, tardi-

ness, and dropouts. At the secondary school level, we can look at how

many students decide to pursue post-secondary education. And sure-

ly, any sort of sensible teacher-evaluation model must consider a

teacher’s classroom conduct.

Few would argue, however, the most important evidence regard-

ing instructional effectiveness must revolve around what students

learn. And tests can help teachers by allowing them not only to deter-

mine for themselves what students have learned, but also to tell that

story to the world. In this chapter, two test-based sources of evalua-

tive evidence were considered: pretest/post-test cognitive assessments

and pretest/post-test affective assessments. The evidence you can col-

lect from both of those approaches, ideally buttressing your students’

performances on the kinds of instructionally supportive standards-

based tests described in Chapter 10, will provide you with a defensi-

ble notion of how well you’ve been teaching. If the use of those kinds

of evidence can help you decide which parts of your instruction are

working and which parts aren’t, then you can keep the winning-parts

and overhaul the losing-parts. Your students deserve no less.

T E S T B E T T E R , T E A C H B E T T E R1 5 8

ch11.qxd 7/30/2003 2:06 PM Page 158

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .

C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 9

Recommended Resources

Bridges, E. M., & Groves, B. R. (2000). The Macro-and micropolitics of person- nel evaluation: A framework. Journal of Personnel Evaluation in Education, 13(4), 321–337.

Guskey, T., & Johnson, D. (Presenters). (1996). Alternative ways to document and communicate student learning [Audiotape]. Alexandria, VA: Association for Supervision and Curriculum Development.

Popham, W. J. (Program Consultant). (2000). Evidence of school quality: How to collect it! [Videotape]. Los Angeles: IOX Assessment Associates.

Popham, W. J. (2001). The truth about testing: An educator’s call to action. Alexandria, VA: Association for Supervision and Curriculum Development.

Popham, W. J. (Program Consultant). (2002). How to evaluate schools [Videotape]. Los Angeles: IOX Assessment Associates.

Shepard, L. (2000, October). The role of assessment in a learning culture. Educational Researcher, 29(7), 4–14.

Stiggins, R. J. (2001). Student-involved classroom assessment (4th ed.). Upper Saddle River, NJ: Prentice Hall.

Stiggins, R. J., & Davies, A. (Program Consultants). (1996). Student-involved conferences: A professional development video [Videotape]. Portland, OR: Assessment Training Institute.

INSTRUCTIONALLY FOCUSED TESTING TIPS

• Collect truly credible pretest/post-test evidence of instructional

impact regarding students’ mastery of important cognitive skills.

• Use the split-and-switch design judiciously in your own classes,

applying it only to the pretest/post-test appraisal of a limited num-

ber of high-import outcomes.

• Use pretest/post-test affective inventories to determine students’

affective changes.

ch11.qxd 7/30/2003 2:06 PM Page 159

Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.

C o p yr

ig h t ©

2 0 0 3 . A

ss o ci

a tio

n f o r

S u p e rv

is io

n &

C u rr

ic u lu

m D

e ve

lo p m

e n t. A

ll ri g h ts

r e se

rv e d .