PLAGIARISM FREE "A" WORK 18 HOURS or LESS
1 4 9
AS I’VE NOTED, THE CHIEF REASON THAT THE DESIGNERS OF EDUCATIONAL
accountability systems rely on standardized tests is that those design-
ers don’t trust educators to provide truly accurate evidence regarding
their own classroom successes or failures. And let’s be honest: That’s
not an absurd reason for incredulity. No one really likes to be evalu-
ated and found wanting. The architects of educational accountability
systems typically make students’ standardized test scores the center-
piece of their evaluative activities because it’s accepted that such tests
provide more believable evidence.
However, because no single source of data should ever be
employed to make significant decisions about schools, teachers, or
students, even standardized test scores should be supplemented by
other evidence of teachers’ instructional success. This supplemental
evidence of a teacher’s instructional impact must clearly be credible. If
evidence of effective instruction is not truly credible, well, no one
will believe it . . . and it won’t do any good.
A Post-Test–Focused Data-Gathering Design You’ll recall from Chapter 1 that one of the major drawbacks of using
data from once-a-year standardized achievement test (even an
11 Collecting Credible
Classroom Evidence of Instructional Impact
ch11.qxd 7/30/2003 2:06 PM Page 149
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 5 0
instructionally supportive one) to see how well a teacher has taught
is that the quality of students in a given teacher’s class varies from
year to year. If a 5th grade teacher’s class this year is brimming with
bright, motivated children, those students are apt to perform pretty
well on this year’s standardized test. But if a year later, the same
teacher draws a markedly less able collection of 5th graders, the stu-
dents’ test scores are apt to be much lower. Does this year-to-year
drop in 5th graders’ scores indicate that the teacher’s instructional
effectiveness had decreased? Of course not. The students were different.
The key to getting a more accurate fix on a teacher’s true instruc-
tional impact is to collect evidence from the same students while they
remain with the same teacher. One way to do this, represented in
Figure 11.1, is to test the teacher’s students after instruction is fin-
ished. This is called a post-test–only design and its chief weakness, for
teachers who would employ it, is that it doesn’t factor in the stu-
dents’ pre-instruction status. If your students score superbly on the
post-instruction test, is it because you taught them well or because
the tested material was something they all learned years ago? With no
basis for comparison, the post-test–only design flops if you’re trying
to isolate the impact of your instruction.
1 1 . 1 A POST-TEST–O NLY DATA-GATHERING DESIGN
ch11.qxd 7/30/2003 2:06 PM Page 150
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 1
The Pretest/Post-Test Data-Gathering Design The most common data-gathering model teachers use to get a picture
of how well they have taught their students is a straightforward
pretest/post-test model such as the one presented graphically in Figure
11.2. (You also saw this data-gathering model in Chapter 1.) Most of
us are familiar with how this works: a teacher gives students the same
test, once before instruction and again after. Elementary teachers, for
example, might administer a 20-item test at the start of a school year
and the same test again at year’s end. Secondary teachers in schools
on a semester system might administer a test at the semester’s outset
and then again at its conclusion.
The virtue of the classic pretest/post-test evaluative model is that,
for the most part, it does measure the same group of students before
instruction and after, meaning that a comparative analysis of the two
sets of test data provides a clearer picture of the teacher’s instruction-
al impact on student mastery levels than do post-test data alone. In
classrooms with high student mobility, it is usually sensible to com-
pare pretest and post-test performances of only those students who
have been in the class long enough for the relevant instruction to
“work.” Thus, whereas all students might take a post-instruction
1 1 . 2 A PRETEST/POST-TEST DATA-GATHERING DESIGN
ch11.qxd 7/30/2003 2:06 PM Page 151
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 5 2
assessment, for purposes of appraising instructional effectiveness, you
would analyze the post-tests scores of just the students who had been
enrolled in the class for, say, eight weeks or whatever instructional
period seems sufficient.
There are problems with the pretest/post-test model that reduces
its accuracy. One major drawback that occurs when a teacher uses the
same test before and after instruction is referred to as pretest reactivi-
ty. What this means is that students’ experience taking a pretest will
often make them react differently to the post-test. They have been
sensitized to “what’s coming.” A skeptic might conclude that even if
post-test scores went up, students didn’t necessary learn anything;
perhaps they simply figured out how to take the test the second time
through. Unfortunately, the skeptic might be right.
Well, how about using different assessments for the pretest and
the post-test? It sounds like a good idea, and it definitely prevents
pretest reactivity, but the truth is that it is much more difficult than
most people realize to create two genuinely equidifficult tests. And
unless the tests are identical in difficulty, all sorts of evaluative crazi-
ness can ensue. If the post-test is the tougher of the two tests, then
the post-test scores have a good chance of being lower, and the
teacher’s instruction will look ineffective no matter what. On the
other hand, if the pretest is much more difficult than the post-test,
then the teacher will appear to be successful even when that’s not
true. No, there are definite difficulties associated with a teacher’s use
of the rudimentary kind of pretest/post-test model seen in Figure
11.2. But this model can be improved.
A Better Source of Evidence: The Split-and-Switch Data-Gathering Design Let me describe a way to collect instruction data for a pretest versus
post-test comparison that skirts the difficulties associated with the
standard pretest/post-test design. I call it the split-and-switch model,
and here’s how to use it in your classroom.
ch11.qxd 7/30/2003 2:06 PM Page 152
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 3
▲ Choose a skill. Select an important skill that your students
should acquire as a consequence of your instruction over a substan-
tial period of time (a semester or school year). The more important
the skill being taught and tested, the more impressive will be any evi-
dence of your instructional impact.
▲ Choose your assessment type. I recommend you opt for a
constructed-response test, which is almost always the best choice to
assess important skills. Performance tests are often ideal. Don’t forget,
though, that if you choose a constructed-response format, you’ll also
need to design a rubric to evaluate students’ responses.
▲ Create two assessments. Now, instead of designing one
assessment to measure your students’ mastery of the skill, design
two—two versions that measure the same variable. Let’s call them
Form A and Form B. The two forms should require approximately the
same administration time and, although they should be similar in
difficulty, they need not be identical in difficulty.
▲ Create two test groups. Randomly divide your class into
two halves. You could simply use your alphabetical class list and split
the class into two half-classes: A–L last names and M–Z last names.
Let’s call these Half-Class 1 and Half-Class 2.
▲ Pretest, teach, and post-test. Finally, you’ll come to the
testing. As a pretest, administer one test form to one test group, and
the other test form to the other test group (for example, Half-Class 1
takes Form A and Half-Class 2 takes Form B). As you might have
already guessed, at the end of instruction, you will post-test by switch-
ing the two test forms so that each test group gets the test form it
didn’t get at pretest-time.
▲ Compare data sets. You’ll end up with two sets of compara-
tive data to consider: the Form A pretests versus the Form A post-tests
and the Form B pretests versus the Form B post-tests. The split-and-
switch design is depicted graphically in Figure 11.3, where you can
see the two appropriate comparisons to make.
ch11.qxd 7/30/2003 2:06 PM Page 153
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 5 4
What you want to see, of course, are results indicating that the
vast majority of high test performances were produced at post-test-
time. There will be no pretest reactively because each half-class gets
what is to them a brand new test form as a post-test. And because this
data-gathering design yields two separate indications of a teacher’s
instructional impact, the results (positive or negative) ought to con-
firm one another. Remember, even if Form A is very tough and Form
B is not, you’re comparing the tough Form A pretests with the tough
Form A post-tests, and the easy Form B post-tests with the easy Form
B post-tests.
Credibility-Enhancing Measures Now, although you could use this model as is, without any refine-
ments, I don’t think you should. Here’s what I recommend.
1 1 . 3 THE SPLIT-AND-SWITCH DATA-GATHERING DESIGN
ch11.qxd 7/30/2003 2:06 PM Page 154
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 5
▲ “Blind score” the pretests and post-tests. The rationale
behind blind scoring is to counteract every teacher’s understandable
desire to see improvement even if it may not be in evidence. When
creating your two test forms, also create identical response-booklets
that you’ll pass out to all students, both at pretest and post-test time.
Tell students not to put dates on their response booklets. When you
collect the completed response booklets at pretest time, code them on
the reverse side of the last sheet so that it’s possible, but not easy, to
tell that the responses are pretests. You might use a short string of
numbers, and make the second number an odd number (1, 3, 5, 7, or
9). Go ahead and look at your students’ pretest responses if you’d like
to get an idea of your students’ skill-levels at the beginning of instruc-
tion, but do not make any marks on their test forms. File the forms
away.
At the end of the semester or school year, conduct a similar cod-
ing procedure when you collect the post-tests, but make the second
number in your code string an even number. Again, you can use these
post-tests as you would any other (to make inferences about what
your students have learned), but if you use post-tests for grading pur-
poses, be sure to take into consideration any difficulty-level differ-
ences you may have detected between Form A and Form B.
Finally, when it comes time to compare the two sets of data, mix
all the Form A responses together (Form A pretests and post-tests). Do
the same for the Form B responses. After you’ve scored them, use the
codes to separate out your pretests and post-tests, and then see what
the data have to say about your instructional effectiveness.
▲ Bring in nonpartisan judges. You can enhance the credi-
bility of your evidence even further by turning over pretest/post-test
evaluations to someone else, thereby ensuring undistorted judgment.
(If you’ve spent some time with your students’ pretests to help you
determine entry-level skills or with the post-tests for grading purpos-
es, you’ll probably recognize which are which, despite your best
efforts.)
ch11.qxd 7/30/2003 2:06 PM Page 155
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
T E S T B E T T E R , T E A C H B E T T E R1 5 6
Parent volunteers make great nonpartisan blind-scorers. Provide
them with an evaluation rubric and a brief orientation to its use. The
judges will get all the Form A responses (both pretests and post-tests
thoroughly mixed together) and then score the responses. After all
the Form A responses have been scored and returned to you, use your
coding system to sort the responses into pretests and post-tests. Then,
it’s just a matter of comparing your students’ pretest responses with
their post-test responses. Use the same process for blind-scoring the
Form B tests.
If you’ve used nonpartisan judges and they have blind-scored
your students’ papers properly, the results ought to be truly credible
to the accountability minded. The results should help you, and the
external world, know if your instruction is working.
Caveats for Using the Model Of course, even the best-laid evaluative schemes can be foiled. The
major weakness in the split-and-switch model is that some teachers
might teach toward the test forms themselves, thereby eroding the
validity of any test-based inferences about students’ true skill mastery.
Always teach only toward the skill being promoted; never teach
toward the tests being used to measure skill-mastery.
Another potential problem is that, in order to produce stable
data, the split-and-switch model requires a relatively large group of
students. A class of around 30 is ideal. If your class is much smaller
(say, 20 students or less), it’s best to revert to a standard pretest/post-
test model, even with its drawbacks. You can still employ the kind of
nonpartisan blind-scoring process I’ve described, which will increase
the likelihood of your pretest/post-test data being regarded as
credible.
Though not flawless, the split-and-switch design can help you
buttress the evidence of your instructional effectiveness in a more
believable manner. And, of course, if this test-based evidence suggests
that your instruction isn’t working as well as you wish, you will need
ch11.qxd 7/30/2003 2:06 PM Page 156
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 7
to address that shortcoming through the use of alternative instruc-
tional strategies.
I urge you to consider employing a split-and-switch design to
evaluate your effectiveness in promoting your students’ mastery of a
very limited number of high-import outcomes. Let’s be honest.
Implementation of this data-gathering design takes time and is clear-
ly more trouble than conventional testing. Use it judiciously, but do
use it. I think you’ll like it.
Comparisons of “Pre-” and “Post-” Affective Data Pretest/post-test contrasts of students’ attitudes, interests, or values as
measured on self-report inventories are a final source of useful credi-
ble evidence regarding a teacher’s own instructional effectiveness.
Remember, if you will, the confidence inventories described in
Chapter 8 and the positive relations between students’ expressed con-
fidence in being able to use a skill and their actual possession of that
skill. If you can assemble pretest/post-test evidence, collected anony-
mously, indicating that your students’ confidence has grown appre-
ciably with respect to their ability to perform significant cognitive
skills, these results will reflect very favorably on your instruction.
If you do decide to collect pretest/post-test affective evidence as
an indication of your own instructional success, remember that you
need to collect it in a manner that even nonbelievers will regard as
credible. Just spend a few mental moments, casting yourself in the
role of a doubter. Pretend you are a person who has serious doubts
about whether the evidence you’ll be presenting about your own
instructional effectiveness is believable. Do whatever you can, in the
collection of any evaluative evidence, to allay the concerns that such
a doubter might have. Happily, a split-and-switch design, especially if
students’ responses are blind-scored by nonpartisans, will take care of
most of a disbeliever’s concerns.
If you are worried about pretest reactivity—and such reactivity
often is a concern when assessing affect—you can employ a split-and-
ch11.qxd 7/30/2003 2:06 PM Page 157
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
switch design, simply chopping your affective inventories into two
half-inventories so that students will get different inventories as
pretests and post-tests.
Final Thoughts There are a number of ways that properly constructed educational
tests can contribute to the quality of a teacher’s instructional deci-
sions. In my opinion, however, there is no more important contribu-
tion than helping a teacher supply an answer to the question “How
effective was my instruction?”
There are also a number of ways to judge a teacher’s success.
Many factors ought to be considered, some of which are not test-
based in any way. We can gain insights into a school staff’s effective-
ness by considering the school’s statistics regarding truancy, tardi-
ness, and dropouts. At the secondary school level, we can look at how
many students decide to pursue post-secondary education. And sure-
ly, any sort of sensible teacher-evaluation model must consider a
teacher’s classroom conduct.
Few would argue, however, the most important evidence regard-
ing instructional effectiveness must revolve around what students
learn. And tests can help teachers by allowing them not only to deter-
mine for themselves what students have learned, but also to tell that
story to the world. In this chapter, two test-based sources of evalua-
tive evidence were considered: pretest/post-test cognitive assessments
and pretest/post-test affective assessments. The evidence you can col-
lect from both of those approaches, ideally buttressing your students’
performances on the kinds of instructionally supportive standards-
based tests described in Chapter 10, will provide you with a defensi-
ble notion of how well you’ve been teaching. If the use of those kinds
of evidence can help you decide which parts of your instruction are
working and which parts aren’t, then you can keep the winning-parts
and overhaul the losing-parts. Your students deserve no less.
T E S T B E T T E R , T E A C H B E T T E R1 5 8
ch11.qxd 7/30/2003 2:06 PM Page 158
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .
C o l l e c t i n g C r e d i b l e C l a s s r o o m E v i d e n c e o f I n s t r u c t i o n a l I m p a c t 1 5 9
Recommended Resources
Bridges, E. M., & Groves, B. R. (2000). The Macro-and micropolitics of person- nel evaluation: A framework. Journal of Personnel Evaluation in Education, 13(4), 321–337.
Guskey, T., & Johnson, D. (Presenters). (1996). Alternative ways to document and communicate student learning [Audiotape]. Alexandria, VA: Association for Supervision and Curriculum Development.
Popham, W. J. (Program Consultant). (2000). Evidence of school quality: How to collect it! [Videotape]. Los Angeles: IOX Assessment Associates.
Popham, W. J. (2001). The truth about testing: An educator’s call to action. Alexandria, VA: Association for Supervision and Curriculum Development.
Popham, W. J. (Program Consultant). (2002). How to evaluate schools [Videotape]. Los Angeles: IOX Assessment Associates.
Shepard, L. (2000, October). The role of assessment in a learning culture. Educational Researcher, 29(7), 4–14.
Stiggins, R. J. (2001). Student-involved classroom assessment (4th ed.). Upper Saddle River, NJ: Prentice Hall.
Stiggins, R. J., & Davies, A. (Program Consultants). (1996). Student-involved conferences: A professional development video [Videotape]. Portland, OR: Assessment Training Institute.
INSTRUCTIONALLY FOCUSED TESTING TIPS
• Collect truly credible pretest/post-test evidence of instructional
impact regarding students’ mastery of important cognitive skills.
• Use the split-and-switch design judiciously in your own classes,
applying it only to the pretest/post-test appraisal of a limited num-
ber of high-import outcomes.
• Use pretest/post-test affective inventories to determine students’
affective changes.
ch11.qxd 7/30/2003 2:06 PM Page 159
Popham, W. James. Test Better, Teach Better : The Instructional Role of Assessment, Association for Supervision & Curriculum Development, 2003. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/amridge/detail.action?docID=5704436. Created from amridge on 2022-01-22 03:17:01.
C o p yr
ig h t ©
2 0 0 3 . A
ss o ci
a tio
n f o r
S u p e rv
is io
n &
C u rr
ic u lu
m D
e ve
lo p m
e n t. A
ll ri g h ts
r e se
rv e d .