Edu 645 wk 6

profileonefan1
06ch_lefrancois_learning.pdf

Summative Assessment

Focus Questions

After reading this chapter, you should be able to answer the following questions:

1. What is summative assessment?

2. What are the main differences between norm-referenced and criterion-referenced interpretations of assessment data?

3. What are some instructional systems based on criterion-referenced comparisons?

4. What is curriculum-based measurement?

5. What are benchmark assessments?

6

Ableimages/Digital Vision/Thinkstock

If you live among wolves you have to howl like a wolf. —Russian proverb

Not to howl like a wolf if you live among them—say, to cluck like a chicken instead—might invite unwanted attention. On the other hand, howling like a wolf if you plan to live among chickens might win you very few friends. The point is that certain characteristics and behav- iors ensure success and survival in a given environment, but they might be totally useless and even dangerous in another—like making bird noises among wolves.

For example, in another book, I describe a jungle kingdom, Ochawa, where survival depends on an unexpected set of behaviors (Lefrançois, 2011).

“Ochawa,” I write, “is a small kingdom hidden in a steamy jungle somewhere” (p. 326). I describe how one of this kingdom’s borders is a river and the other is a row of rugged moun- tains. The inhabitants are trapped between river and mountain: They cannot cross the river because it’s too swift and deep; and they cannot cross the mountains because the other side is an unbroken wall of sheer cliff.

Summative Assessment Chapter 6

Chapter Outline 6.1 The Nature of Summative Assessment

Comparisons in Summative Assessment

What Is Summative Assessment?

Main Purposes of Summative Assessment

6.2 Norm-Referenced Interpretations

Uses of Norm-Referenced Interpretations

Advantages and Disadvantages of Norm-Referenced Approaches

6.3 Criterion- and Self-Referenced Approaches

Criteria and Standards in Education

Criterion-Referenced Interpretations of Assessment in the Classroom

Self-Referenced Approaches to Assessment

Comparisons of Norm-, Criterion-, and Self-Referenced Interpretations

6.4 Instructional Systems Based on Criterion-Referenced Interpretations

Bloom’s Mastery Learning

Keller’s Personalized System of Instruction (PSI)

Research and Applications of Competency-Based Approaches

6.5 Applications and Examples of Summative Assessment

Planning for Summative Assessment

Curriculum-Based Measurement (CBM)

Benchmark Assessment

A Reminder: The Main Purpose of Educational Assessment

Chapter 6 Themes and Questions

Section Summaries

Applied Questions

Key Terms

The Nature of Summative Assessment Chapter 6

These circumstances present an important survival test to all the people of Ochawa every day of their lives. You see, during the day, they must venture into the jungle for food; but every night, they need to climb up to their caves on the mountainside before the nocturnal preda- tors that roam the jungle awaken. That is their test. Those who pass the test survive; those who don’t, well. . .

The test given to citizens of Ochawa is quite different from the tests that are common in most schools. To pass the Ochawa test, you don’t have to be the first to reach safety; you don’t even have to be among the first 90% to do so. Nor do you have to climb higher than the oth- ers. In fact, you will pass just as surely if you are the very last to reach safety. And you might be far less hungry.

6.1 The Nature of Summative Assessment Much like daily life in Ochawa, schools present our learners with various tests—although these are not normally a matter of life and death. Still, there is sometimes a very close parallel between assessment in Ochawa and summative assessment in schools. Recall that summa- tive assessment is the type of assessment that normally occurs at the end of an instructional sequence and that is designed mainly to provide a grade.

Comparisons in Summative Assessment

Assume, for example, that students are expected to reach a certain level of competence—that is, to learn certain identifiable concepts and develop a repertoire of specific skills. Learning these concepts and skills can be viewed as defining the criteria of success—criteria we can denote as X. In a sense, X is analogous to escaping from the beasts: you either reach X or you don’t (you pass or you fail).

Say, for example, that the criterion for success in Ms. Espinaco’s third-grade arithmetic class is that learners understand parentheses and that, as a result, they be able to solve four prob- lems of the kind shown in Figure 6.1. Furthermore, they must be able to do so within five minutes and with no more than one error. Doing so successfully is the criterion of success for this arithmetic unit. Assessment that determines whether each third-grader in Ms. Espinaco’s class has reached this level of competence would be an example of a criterion-referenced approach. It is important to note that a criterion-referenced approach does not involve the use of a different kind of test; rather, it involves a different sort of interpretation of test scores. Hence it is more accurate to refer to criterion-referenced interpretations or comparisons rather than to criterion-referenced tests or assessments. As we see shortly, criteria in education are often based on predefined state- or system-wide standards to which the performance of each learner is compared.

Figure 6.1: Example of a brief criterion-referenced test in third-grade arithmetic

A criterion-referenced test is typically a pass-fail test. In this case, third-graders are deemed to have passed (to have reached the passing criterion) when they cor- rectly solve four of the following five problems. Criteria for criterion-referenced tests are often based on state- mandated standards.

f06.01_EDU645.ai

24 – (8 + 4) =

30 + (15 – 5) =

5 – (22 – 20) =

18 + (3 + 22) =

34 – (7 + 7) =

The Nature of Summative Assessment Chapter 6

In many assessment situations, however, the performance of each learner isn’t compared to a standard (or criterion); instead, it is compared to the performance of other learners. As a result, students who have reached X but who fall below the average performance of the class may be assigned mediocre marks. And in lower performing classes, students who have not reached X might be given exceptionally high grades. This approach illustrates a norm- referenced interpretation.

A third option is also available: it compares each learner’s performance not to a standard (or criterion), as in criterion-referenced interpretations, nor to the average performance of other comparable learners (as in norm-referenced comparisons). Instead, learners are compared to themselves. A self-referenced approach to assessment compares the learner’s current performance with earlier performance or with expected performance based on ability, experi- ence, and other personal factors. Teacher comments such as “Elvira is not working up to her ability” or “Renaldo should be able to complete three more laps after eight weeks of training” are examples of self-referenced assessment.

In brief, summative assessments typically involve one or more of three principal kinds of comparisons:

1. Comparisons with the performance of other comparable learners (norm-referenced interpretations)

2. Comparisons with a predefined criterion or standard of acceptable or expected perfor- mance (criterion- or standards-based interpretations)

3. Comparisons of each learner with him or herself, often reflecting expectations based on measured or assumed ability, background, and other personal factors such as motiva- tion, parental encouragement, and so on (self-referenced approach to assessment). (See Figure 6.2).

Each of these three approaches to interpreting test scores—criterion-referenced, norm- referenced, and self-referenced—can be used for summative purposes.

What Is Summative Assessment?

Earlier, we distinguished among four kinds of assessment:

1. Formative assessment: An integral part of instruction designed mainly to provide immedi- ate, ongoing feedback to assist learners and teachers in improving the teaching–learning process.

2. Placement assessment: Closely related to formative assessment, and used for making selec- tion and placement decisions.

3. Diagnostic assessment: Directed at uncovering learner strengths and weaknesses, allowing for differentiated assessment and differentiated instruction. Differentiation in education refers to assessment procedures that identify important differences among learners and that lead to instructional procedures that accommodate these differences. Accommodations in approaches to teaching and learning are a key feature of formative assessment.

4. Summative assessment: Assessment that typically occurs at the end of an instructional sequence and that is used to summarize student progress and achievement and to provide a grade.

The Nature of Summative Assessment Chapter 6

Note that these types of assessment differ mainly in terms of their primary purposes—to assist, to place, to diagnose, or to summarize. However, although their main purposes might differ, they are not always very distinct in practice. Not only might the same formal and infor- mal assessment procedures be used for all four purposes, but the purposes themselves are all directed toward the same end. Simply put, that end is to provide the most effective possible educational experience for all learners. The point is worth repeating: The main purpose of educational assessment is to foster learning in all its forms.

Moreover, any single assessment, no matter its primary use, might serve all four purposes. For example, when Miss Aloysius gives her third-graders a quiz before starting her first arithmetic unit in September, she might use the results to

• Divide the class into learning groups based on their individual performance—a place- ment function

• Look for weaknesses and strengths in individual learners’ understanding of basic grouping concepts—a diagnostic function

• Use the results to guide instructional and learning activities for each group—a forma- tive function

• Grade each learner to obtain information, some of which can be shared with parents during the meet-the-teacher evening at the start of term

Figure 6.2: Comparison bases for summative assessment

The group that the learner is compared to for assessment determines the broad approach to assessment.

f06.02_EDU645.ai

The average score of the group, no matter

what it may be, is assigned a “C” grade and all other scores

are distributed around this mark.

Learners are assigned a “pass-fail” grade

according to whether they achieve at a predetermined

level.

Learners are evaluated

in terms of their improvement.

Learner’s performance is compared

with that of peers.

Learner’s performance is

judged relative to some standard of

acceptable performance.

Learner’s performance

is assessed relative to previous or expected

performance.

Norm- Referenced

Criterion- Referenced

Self- Referenced

Other Learners

Pre- determined Standards

Self

Comparison Group

Type of Comparison Explanation Example

The Nature of Summative Assessment Chapter 6

Main Purposes of Summative Assessment

It’s important to note that the main purpose of summative assessment is somewhat differ- ent from the purposes of the other three categories of assessment. The common purpose of placement, formative, and diagnostic assessment is basically formative: specifically, to opti- mize instructional and learning experiences for the learner. Summative assessment, on the other hand, has its own main purpose: to measure the outcomes of learning and instructional experiences. Put simply, its purpose is summative: It provides an index of the sum or total of the effects of schooling. As Black (1998) explains, when the cook tastes the soup, she is involved in formative assessment: She can still add herbs, thicken the broth, simmer the ingre- dients, and put in new spices, new vegetables, new meats, and new starches. But when the customer tastes the soup, he is engaged in summative assessment: The time to change and improve the broth is past: Now is the time for the final grade.

As mentioned earlier, a useful way of distinguishing between summative assessment and other approaches to assessment is implicit in the observation that formative, diagnostic, and placement assessment are, in a sense, assessment for learning. In contrast, summative assess- ment is assessment of learning (see Figure 6.3).

Figure 6.3: Some distinctions between summative and formative assessment

Despite their differences in principal purposes, timing, and uses, a single assessment might serve both formative and summative functions.

f06.03_EDU645.ai

Formative Assessment

Assessment for

learning

Occurs during learning

Used to improve

teaching and learning

Learner involved with teacher in interpreting

and using results of assessment

Summative Assessment

Assessment of

learning

Occurs after

learning

Used to summarize effects of

teaching and learning

Learner less involved

Norm-Referenced Interpretations Chapter 6

In summary, the main purpose of summative assessment is to provide a mark or a grade that reflects progress and achievement. At the same time, however, summative assessments can be used to draw conclusions about the effectiveness of school programs and of instruc- tional strategies. Summative assessments also say something about the appropriateness of curriculum offerings, about the readiness of learners, and perhaps about learner and teacher characteristics.

6.2 Norm-Referenced Interpretations In norm-referenced approaches, the student’s performance is interpreted in terms of how all students perform. In a sense, the performance of other students sets the norm or the stan- dard. This type of norm is quite different from the predetermined standards that character- ize criterion-referenced assessment. Norm-referenced interpretations compare a test score to scores obtained by other test takers; criterion-referenced interpretations compare a test taker’s score to a predetermined standard.

It’s important to note, however, that standards (the criteria) for both criterion-referenced and norm-referenced comparisons might have very similar sources. For example, both might be based on state-prescribed standards for performance at a specific grade level where the stan- dards have been determined by looking at the performance of a large, representative group on relevant assessments. The distinction is that criterion-referenced assessment evaluates the learner in terms of a pre- determined standard; norm-referenced assessment evaluates the learner in rela- tion to other learners.

Norm-referenced interpretations lend themselves well to competitive approaches to teaching but are poorly suited to more cooperative approaches. Norm-referenced comparison, says Kingston (2008), is the dominant form of assessment in most schools.

Norm-referenced interpretations often rely on assessments based on objective tests. These are tests that consist of items that require short, factual, unambigu- ous answers. Objective tests typically can be scored by anyone who either has an answer key or who happens to know the correct answers. There tends to be high agreement among scores from different markers.

The most common objective tests are multiple-choice tests. Many are commercially devel- oped. Often, these are based on national norms rather than on local curricula. They are usually designed to rank students relative to the performance of a representative sample of similar students—referred to as a norming group. Their emphasis is less on covering

Digital Vision/Thinkstock

▲ The most common norm-referenced tests are objective, multiple-choice tests like the one these students are writing. Such tests provide an easy basis for comparing these learners to each other. If the test is a standardized test, test takers can also be compared to students in other schools, in the entire state, or in the nation. But the validity and fairness of com- parisons might be suspect if cheating occurs.

Norm-Referenced Interpretations Chapter 6

curriculum content that might have been mastered than on selecting items that most clearly differentiate among examinees. As a result, standardized tests do not usually ask questions that everyone might answer correctly. Instead, they ask questions that only a percentage of test takers can answer. (Teacher-made objective tests are discussed in Chapter 8; standardized tests are the subject of Chapter 10.)

Uses of Norm-Referenced Interpretations

Much of the assessment that occurs in the classroom is not based on commercially developed, standardized tests, but on teacher-made assessments, many of which are used for norm- referenced interpretations. Based on these assessments, students in a class may be ranked according to how well they perform relative to each other. Ranking is one way of comparing an individual’s performance to that of others exposed to the same assessment. That Benjamin is ranked first following a third-grade arithmetic quiz simply means that he has performed better than all his third-grade classmates. The norm that his performance is being compared to has been established by the performance of his entire class.

Ranking is sometimes expressed as the percentage of students below a given point. For exam- ple, Benjamin’s performance is at the 99th percentile: 99% of his classmates score at or below his score. A percentile is simply the point at or below which a given percentage of cases fall (percentiles and other important statistics are discussed in Chapter 9).

Another approach to reporting norm-referenced grades is to compare each test taker’s score to the mean (average) of the entire group. Say, for example, that the average of all scores obtained by Benjamin’s class on this arithmetic test was 15 out of 40. Benjamin’s performance is clearly better than average: He ranked first. So, if the teacher is using a letter grading system—say A, B, C, D, and F—Benjamin would probably be given an A.

Nevertheless, for a variety of reasons that we look at in Chapter 9, Benjamin’s A, all by itself, might be highly misleading. If his classmates did extraordinarily poorly, he might have answered only 19 of the 40 test items correctly and done better than everyone else. Conversely, if they did exceptionally well, he might have had to answer 39 items correctly to be at the 99th per- centile. If he answered only 19 correctly, he would have done rather poorly in that class.

Advantages and Disadvantages of Norm-Referenced Approaches

The previous example highlights a significant disadvantage of norm-referenced grading when norms are based on the performance of a small group such as a single class. As Figure 6.4 shows, the group you belong to can make an enormous difference under these circumstances.

Of course, when norms are based on the performance of large, comparable samples—which is the case for standardized tests—these disadvantages disappear. Now all examinees are being compared to the same highly representative norms. Performing at the 99th percen- tile—or at the 10th, meaning that approximately 90% of the norming group did better—has consistent meaning no matter how the rest of Benjamin’s class might have done.

For several reasons, norm-referenced interpretations are widely used in schools:

• They provide an effective means of comparing students. If, for example, there are limited resources for programs for gifted and talented learners, a valuable and fair way

Norm-Referenced Interpretations Chapter 6

of selecting candidates for admission to these programs is to rank them. Ranks and percentiles are especially meaningful if they compare learners to state- or nationwide norms rather than only to their classmates.

• Norm-referenced comparisons based on standardized tests provide quick, inexpensive, and useful information about the extent to which students have learned what they were expected to learn.

• Norm-referenced interpretations are often more easily developed by teachers than are criterion-based comparisons.

It’s important to recognize that although a score such as a rank or a percentile indicates how well or how poorly Benjamin has performed relative to his classmates, it says nothing about any specific competencies he might have acquired—or have failed to acquire. True, we can assume with some confidence that if Benjamin is at the 99th percentile on a nationally standardized achievement test, he will have acquired all relevant knowledge and skills for his grade level. But that he scores at the 99th percentile relative to his small class tells us much less about what he has learned, especially if that class happens to be disadvantaged—or advantaged—in one or more ways.

Figure 6.4: Interpreting norm-referenced scores on teacher-made tests

Hypothetical distributions of identical norm-referenced tests given to two classes. If Sally were in class A, her rank would be at a respectable 64th percentile. But if she were in class B, the same score would place her just above the 40th percentile.

f06.04_EDU645.ai

10 20 30 40 50 60 70 80 90 99

10 20 30 40 50 60 70 80 90 99

(percentiles in class A)

(percentiles in class B)

Sally’s mark

Class A distribution

Class B distribution

Criterion- and Self-Referenced Approaches Chapter 6

6.3 Criterion- and Self-Referenced Approaches Criterion-referenced interpretations of assessment provide us with information that is quite different from that provided by norm-referenced interpretations. Whereas norm-referenced interpretations reveal how the learner has performed relative to the performance of others, a criterion-referenced interpretation reveals how the learner is performing relative to some criterion or to a set of criteria. In the jungle example, assessment is criterion referenced. The criterion is obvious: Did you climb high enough and early enough? The grade is equally clear: life or death.

Criteria and Standards in Education

Criterion-based interpretations of assessment depend on having clear, measurable criteria. These might be derived from lists of objectives and goals that can serve as guides for teach- ers and school systems as they develop learning targets. Bloom’s taxonomy of educational objectives, described in Chapter 4, provides one approach. Recall that one part of the tax- onomy describes a series of activities that underlie our cognitive processing—activities like

remembering, understanding, analyzing, evaluating, and creating. Figure 4.4, for example, lists a large number of verbs for each of these activities—verbs such as indicate, demonstrate, draw, calculate, appraise, produce, and so on. Each of these verbs suggests different instructional and assessment procedures.

The No Child Left Behind legislation of 2001 required every state to create educational standards to guide edu- cators and test developers. Ten years later, the Common Core State Standards (CCSS) were created by the National Governors Association and the Council of Chief State School Officers to “provide a consistent, clear understanding of what students are expected to learn, so teachers and parents know what they need to do to help them” (from http://www.corestandards.org/). The CCSS, which have been adopted by most states, are statements of instructional objectives describing the skills and knowl- edge that define minimum competency in core subjects at each grade level. All schools are encouraged to adopt and apply these standards.

The development of state standards in education under- lies what is termed standards-based education. The term describes educational systems where instructional and assessment methods are directed toward reaching explicit academic and performance levels. As we saw in Chapter 4, standards-based education, and the related practice of standards-based grading, are very much part of ongo- ing educational reforms. Standards-based grading is an approach to grading that looks specifically at the learn- er’s performance with respect to clearly defined state- wide standards denoting proficiency or mastery.

iStockphoto/Thinkstock

▲ Common core standards are statements of basic levels of performance and achievement that are expected in core subjects at each grade level and that are common to more than one school. For example, high school stu- dents should be able to apply trigonometry to general triangles.

Criterion- and Self-Referenced Approaches Chapter 6

Aligning Standards and Assessments An ongoing issue in education has to do with what is termed educational alignment—the need to align local curriculum with state-mandated standards and to adjust instructional and assessment standards accordingly. Degree of alignment between assessment procedures and content standards serves as an indication of the validity of the assessments.

Alignment between assessments and educational standards is usually achieved in one of three ways, as Case, Jorgensen, and Zucker (2008) explain:

1. Standards are developed first and are then used as blueprints for developing assessments.

2. Experts are called in to review the alignment between assessments and content after both have been developed.

3. Content standards and assessments are analyzed, and alignment models are used to quan- tify the match between them.

The last of these approaches, the application of models following analysis of assessments and content, is widely used by state educational authorities in an effort to ensure that the standardized statewide achievement tests they use reflect core standards (see Chapter 10). One commonly used model is similar to what might be based on a taxonomy such as Bloom’s (Porter, 2002). It looks at alignment with respect to the content of both standards and assessments, and it looks at the cognitive activities required by each. Degree of alignment is reflected in how closely the content and cognitive activity requirements parallel each other on both assessments and standards. Thus, if an educational standard stresses being able to apply knowledge in new situations, a well-aligned assessment procedure would require that the learner demonstrate applications.

Criterion-Referenced Interpretations of Assessment in the Classroom

Using common core standards as a basis for establishing proficiency and competency criteria for learners gives teachers a basis for criterion-referenced comparisons. Because criteria are generally expressed as specific competencies (skills, understanding, performance levels, and so on), criterion-referenced approaches to assessment and grading provide the teacher with very specific information about what the individual has learned. Also, specifying criteria in the first place may do much to clarify the teacher’s role. Few instructional guides are more useful than those that provide a clear view of learning objectives defined in terms of specific performance criteria.

Both norm-referenced and criterion-referenced comparisons can be based on teacher-made tests or on commercially prepared standardized tests. A clear example of a criterion-referenced approach based on a standardized test is the written test most of us had to take to obtain a driver’s permit. In most such tests, the criterion is a specific score. In California, for example, the test for a regular driver’s class C license consists of 36 multiple-choice questions. The criterion for passing is 31 correct responses. Accordingly, unlike the case for a norm-referenced assess- ment, the relative performances of all those who take the test in a given month is irrelevant to whether a test taker passes or fails. It makes no difference how well—or how poorly—others perform: As long as you meet the criterion, you will pass. (And if you fail, you can take the test again, although it might be a different test because there are five different forms of it.)

By definition, criterion-referenced approaches depend on a clear understanding of instruc- tional objectives and on knowing what an acceptable level of proficiency and performance is.

Criterion- and Self-Referenced Approaches Chapter 6

The test maker needs to answer questions like these: How many items must drivers answer correctly to demonstrate that they are qualified to operate a motor vehicle? Or, How rapidly and accurately must the student be able to translate the Spanish poem to demonstrate profi- ciency at a fourth-grade level?

Teachers need to describe instructional objectives in terms of measurable competencies and achievements. These descriptions are the criteria that the teacher will then use in criterion- referenced grading. Without a clear statement of learning criteria, it would not be possible to devise assessments to determine whether instructional objectives have been reached. Accordingly, criterion-referenced approaches to assessment have these characteristics:

• They are typically based on tests that include items that are directly relevant to instruc- tional objectives rather than items that are useful for discriminating among students. For example, whereas norm-referenced approaches are often based on tests that include a range of easy to more difficult items and so tend to produce a wide range of scores, a criterion-referenced approach might be based on a test that produces only two grades (pass or fail; satisfactory or unsatisfactory).

• They allow instructors and parents to determine and describe relatively precisely the learner’s achievement and progress. For example, they might permit the teacher to describe the specific learning tasks that have been mastered, or to establish whether a predetermined standard of acceptable performance has been reached.

• In many instances, they are based on the sorts of assessments on which all learners will succeed. In contrast with norm-referenced interpretations, discriminating among learners is not an objective; ascertaining that all—or most—learners have reached predetermined criteria is their defining purpose. As explained later in this chapter, in the discussion of Bloom’s mastery learning, criterion-referenced interpretations are especially well suited to competency-oriented classrooms.

Self-Referenced Approaches to Assessment

Criterion-referenced interpretations compare the learner’s performance to some predeter- mined standard; norm-referenced approaches compare the learner’s performance to that of a group of peers. A third option, a self-referenced approach to assessment, compares the learner’s current performance to the same individual’s performance at some earlier time, or it compares the learner’s performance with performance expected on the basis of ability. That is, the comparison might be based on the learner’s progress (or lack thereof), or on the extent to which the learner is achieving as might be expected on the basis of factors such as ability and experience.

Teacher comments such as “Marcia is doing better than expected” or “Samuel can do a lot better” are examples of self-referenced comparisons. These comments reflect the teacher’s expectations of possible and probable achievement based on ability. An ability-based, self- referenced approach to assessment usually means that students who achieve according to their ability or who do even better than what might be expected receive high grades. The term overachiever is often used to describe these learners.

In contrast, learners for whom there are high expectations based on ability but who do not perform up to expectations are given lower grades. These learners are sometimes described as underachievers.

Criterion- and Self-Referenced Approaches Chapter 6

Problems With Self-Referenced Interpretations One problem associated with self-referenced comparisons based on ability is that teachers are not always very good judges of learners’ abilities. Nor can they easily identify and differentiate among all the various abilities and characteristics that contribute to school achievement. For example, research conducted by Sommer, Fink, and Neubauer (2008) suggests that teach- ers’ estimates of social and creative abilities are especially unreliable. Similarly, Lee and Reeve (2012) found that teachers are often poor judges of students’ motivation and effort. There is evidence, too, that they tend to overestimate the progress of their lower ability students (Graney, 2008).

A second problem related to self-referenced approaches to assessment is that evaluations based on either ability or progress are typically vague and confusing for both parents and learners. For example, when a student of lower ability who is an overachiever receives the same grade as a student with higher ability who is underachieving, does this mean that they have achieved at the same level? And the comment that “Mario is improving” says nothing specific about Mario’s performance relative to instructional objectives or relative to other students.

A third problem with assessing students on the basis of expected performance is that teacher expectations are often strongly affected by a variety of factors that are not always related to ability. Expectations, as Damber, Samuelsson, and Taube (2012) point out, are often tied to socioeconomic and language factors. They may also be profoundly influenced by race, class, and gender (Riley & Ungerleider, 2012). And although these factors may sometimes be related to achievement, they are less likely to be related to innate ability.

An associated problem with assessment based on expected achievement is that a growing body of evidence suggests that teacher expectations themselves affect learner outcomes (e.g., Workman, 2012). This is partly because, as Workman notes, teachers’ expectations influence the sorts of learning opportunities they provide their learners. Unfortunately, however, expec- tations for some students are inappropriate. As a result, the learning opportunities provided for them may be inadequate and unfair. Or the overly optimistic expectations teachers have for some learners may place too great a burden on them and might even result in judgments of underachievement.

Uses of Self-Referenced Interpretations In light of these problems and weaknesses, self-referenced approaches are not generally rec- ommended as a basis for grading, especially in the higher grades. However, as we see in Chapter 11, evaluative comments relating to progress and even to expected achievement are highly common on report cards, especially in early grades.

Even though self-referenced comparisons are often vague and confusing and are uncommon as a grading system, they can be very useful for a variety of purposes in the classroom. First, a self-referenced approach can be highly beneficial for teachers to encourage learners to set goals that are challenging but within with their capabilities. They can then be encouraged to monitor and control their learning activities with a view to helping them become self- regulated learners—learners who have learned how to learn. These are learners who have developed strong self-assessment skills and are therefore able to set their own learning goals as well as evaluate and direct their activities to maximize the attainment of these goals. Self- regulated learning is autonomous, self-directed learning. Becoming a self-regulated learner

Criterion- and Self-Referenced Approaches Chapter 6

is highly dependent on accurate monitoring of learning progress—that is, on self-referenced evaluations (Bercher, 2012).

Second, self-referenced approaches to assessment can serve an important motivational func- tion. Informing learners and parents that Maximilian is doing better than expected might be highly rewarding for little Max, and he might be encouraged to redouble his efforts. Similarly, the evaluative comments that “Janet is capable of doing much better in arithmetic” or “Sammy is progressing well in language literacy” might motivate Janet and Sammy to work even harder than they have been.

Third, self-referenced comparisons can be helpful when the teacher is interested in a student’s learning progressions—in the learner’s mastery of a given sequence of cognitive skills and understanding. Self-referenced interpretations, whether they are based on comparisons with expected performance or with student improvement, can suggest critical interventions to the teacher. They might indicate to the learner that certain cognitive strategies are better than others for different kinds of learning.

Also, in certain endeavors, self-referenced comparisons are the most logical and informative approach. For example, a student athlete preparing for the annual track meet might, along with her coach, begin by recording her usual fastest times. Their ultimate target—say, win- ning the 100-meter distance—is based on the expected winning time calculated from previ- ous years’ performances. Now the athlete sets sequential targets and monitors her progress toward these targets over the training period. This is a clear example of a self-referenced approach based on individual progress.

Comparisons of Norm-, Criterion-, and Self-Referenced Interpretations

The primary difference among the three approaches to summative assessment—norm- referenced, criterion-referenced, and self-referenced—is that each approach compares the learner’s performance to a different reference group. Norm-referenced approaches, compare a learner’s performance to that of others in the same or a similar group; criterion-referenced approaches refer to a predetermined standard; and self-referenced approaches are based on comparisons with the self. However, the same tests might be used for each of these comparisons.

As an example of different uses of a single measuring instrument, Mr. Leacock administers a standardized mathematics achievement test to his sixth-graders. Knowing that there is space for three students in the math enrichment club, he then uses the results to rank his students so that he can select the top three for admission into this club. This is an example of norm- referenced interpretation. The level of achievement required to be selected for entry into the math enrichment club depends on the performance of all members of this sixth-grade class.

At the same time, however, Mr. Leacock might use the test to identify those of his students who have not yet mastered an understanding of concepts dealing with percentages, ratio, rate, and rounding. Those he deems to have met the criteria for understanding these concepts have answered at least four of test items 5, 7, 9, 12, 43, and 67 correctly. Thus the standard- ized test that had been part of a norm-referenced interpretation is now used for a criterion- referenced approach.

Furthermore, Mr. Leacock compares the performance of each of his students to that student’s performance on the same test written three months earlier. This step, which allows him to gauge individual progress during the term, illustrates the use of this standardized test for self- referenced observations. Figure 6.5 summarizes these concepts.

Criterion- and Self-Referenced Approaches Chapter 6

Which Approach Is Best? No single approach to assessment—norm-referenced, criterion-referenced, or self- referenced—is always best. Each one has advantages and disadvantages (see Applications: What Kind of Interpretation of Assessment Data is needed?). In the final analysis, the best approach clearly depends on how the results will be used.

Still, some educators strongly advocate one approach over the other. Advocates of criterion- referenced comparisons point to the inherent justice of their approach. No student need fail simply because others learn faster or have less to learn in the first place: All who reach the criterion will be equally successful.

Criterion-referenced approaches present a powerful argument for individualizing instruction. And there is strong evidence that when instruction pays attention to individual differences and tailors offerings in response to those differences, it can be highly effective. The most individualized of instructional approaches is one-on-one tutoring. After an extensive review of the literature, Bloom (1984) claimed that an average student paired with a good tutor can be expected to achieve at somewhere around the 98th percentile when compared with a group of similar students taught in a more conventional manner. This is a strong argument for the use of computer-based intelligent tutoring systems (discussed in Chapter 5).

But criterion-referenced comparisons do have their limitations. For one thing, it is sometimes very difficult to specify the criteria that define acceptable performance or high proficiency.

Figure 6.5: Three assessment approaches

The three approaches are distinguished mainly by the comparison group used. The same assess- ment procedures might be used for all three.

f06.05_EDU645.ai

Criterion- Referenced

Compares learner’s performance to

an external, predetermined

standard

Based on standardized or teacher-devised

assessments

A principal aspect of mastery learning

approaches (Bloom & Keller) and

curriculum-based measurement

Self- Referenced

Compares learner’s performance to own

past performance

Based on a variety of informal and formal

assessment procedures

Has important motivational

effects; central in individualized

instruction approaches

Norm- Referenced

Compares learner’s performance to that of

others in same (or a similar) group

Often based on standardized tests, but also relies on

teacher-made tests

Highly useful for comparing students

getting quick snapshots of how well learning targets are being met

Criterion- and Self-Referenced Approaches Chapter 6

A P P L I C A T I O N S :

What Kind of Interpretation of Assessment Data Is Needed?

Señora Sánchez teaches ninth-grade Spanish 1 in an urban high school with an enrollment of 3,000 students. At the end of the year, she gives a final exam to her class. She explains how she interprets the scores:

My final exam covers everything we learned in Spanish 1. It has five sections: vocabulary, read- ing, listening, writing, and speaking. From this test, I can assign a student a grade for the second semester, I can decide what to write in the narrative portion of the grade book, and I can make a recommendation about placement for 10th-grade Spanish (whether or not a student should go to Spanish 2 or Spanish 2 Honors).

Here’s how it works: There are 10 questions in each of the five parts of the test for a total of 50 questions. The first 5 questions in each section are from the first semester, and the second 5 are based only on material we covered in the second semester.

Semester Grade: To calculate the second semester grade, I use a norm-referenced interpretation of the second 5 items in each section. I rank the students based on their test score and give an A to the students in the top 10%, a B to those in the next 20%, and so on.

Narrative Comments: When deciding what to write in the narrative part of the grade book, I use a self-referenced interpretation. To do so, I compare how the student did on the first half of the test (the material from first semester) to performance on the second half. This provides information about each person’s progress over the course of the year.

Placement Recommendation: Lastly, I have to decide who is going into regular Spanish 2 and who is ready for Honors—a criterion-referenced interpretation. For that, I look at the whole test. Because there is usually room for only 30 people in the Honors section, the bar is set really high. As a department, we have decided that unless a student can answer at least 47 of the items correctly, he or she probably isn’t ready for Honors Spanish. Of course we take other factors into consideration, such as the student’s effort and past performance, but the score on that final exam is a big part of the decision.

Test Section

Item Numbers Used for Norm-Referenced Interpretation of Semester Grade

Item Numbers Used for Self-Referenced Interpretations

Item Numbers Used for Criterion- Referenced Interpretation of Placement

Vocabulary 6–10 1–5 compared to 6–10 1–10

Reading 16–20 11–15 compared to 16–20 11–20

Listening 26–30 21–25 compared to 26–30 21–30

Writing 36–40 31–35 compared to 36–40 31–40

Speaking 46–50 41–45 compared to 46–50 41–50

Criterion- and Self-Referenced Approaches Chapter 6

For example, while it might be relatively simple to detail exactly what fourth-graders should know and be able to do following a unit on multiplication, it is quite another thing to spell out what they should understand and be able to do following a unit in social studies or religion.

A second limitation of criterion-referenced approaches is that some students are easily able to go beyond established criteria. As we see in the discussion of competency-based learning (in the next section), some educators fear that this situation might reduce students’ motivation and stifle their initiative and creativity.

One advantage of norm-referenced approaches is that they provide useful information about the likelihood of success in a competitive, dog-eat-dog kind of world. It isn’t especially surpris- ing that postsecondary admissions are often based largely on the individual’s relative ranking on standardized assessments.

Self-referenced interpretations, as we noted, can have important motivational consequences. They can also be useful for setting goals and assessing progress toward those goals. For those who would improve themselves, it is essential to know how they are doing now relative to how they have done in the past (see In the Classroom: A Normative Grading Conundrum).

I N T H E C L A S S R O O M :

A Normative Grading Conundrum

Suppose you are every bit as fantastic a teacher as you have sometimes fantasized you might be. As a result, your students are exceptionally motivated, enormously successful, and deliriously happy. The grades you submit to central office for posting and dissemination are heavily weighted toward the top end of the school’s grading system; absolutely no one deserves any less.

Wonderful situation, isn’t it?

But no: Administration says, “Listen, about your grades . . . there’s no spread, you see. We can’t really have that. It’s not like you have only five students, for heaven’s sake! You have 34 in that class. Spread them out. Give us a more reasonable average.”

Do you think the opposite might have happened if, despite your unparalleled skills and dedication, all of your students remained pathetically uninterested and appallingly unsuccessful? Would admin- istration now say, “Hey, listen, about those grades . . . it’s not only that there’s no spread but, well what’s going on in that class? Do you need some help with your instructional strategies? With your formative assessments? We’re going to have to spread and raise these grades a bit. Give them a more reasonable average.”

These scenarios illustrate the conundrum that normative grading can sometimes pose. If assess- ments were based on definite criteria, and if these were clearly aligned with educational goals and with assessments, should you disagree with administration, maintaining that your grades are cor- rect and appropriate? Would your grades be less defensible if your assessments compared learners to one another rather than to predetermined criteria?

Instructional Systems Based on Criterion-Referenced Approaches Chapter 6

6.4 Instructional Systems Based on Criterion- Referenced Approaches

Instructional systems based on criterion-referenced approaches are generally referred to as competency-based learning because they are designed to develop competency relative to specific criteria. The best known and most useful of these systems are Benjamin Bloom’s mas- tery learning and Fred Keller’s Personalized System of Instruction (PSI).

Bloom’s Mastery Learning

There are faster learners, Bloom (1976) informs us, and there are slower learners—which is a far cry from saying there are highly gifted and less gifted learners. As Bloom (1968) explains, most students—perhaps as many as 90% of them—can master what schools teach. The main task of educators is threefold:

1. Determine what is meant by mastery. Determining what is meant by mastery is, in effect, the process of establishing instructional objectives. But it also requires that instructional objectives be described in terms of actual performances and measurable competencies— in other words, in terms of criteria.

2. Devise instructional strategies that will ensure that most learners achieve mastery.

3. Devise assessment instruments that will assist learners in mastering instructional objectives (formative assessments) and that will reveal to both learners and teachers that mastery has been achieved (summative assessment).

Basic Assumptions of Mastery Learning That most learners can achieve the criteria that define mas- tery (competence) is one of two basic assumptions underly- ing Bloom’s model of mastery learning. The second is that achieving mastery requires constant formative assessment— assessment designed not for grading, but directed instead toward helping the student learn.

Reflecting the first of these two assumptions—namely, that there are faster and slower learners, but that most, given enough time, can master course objectives—Bloom’s model of mastery learning holds that degree of learning is a func- tion of the amount of time spent learning relative to the amount of time required. And the amount of time required is a function not only of ability but also of quality of instruc- tion (Figure 6.6). As a result, the model attempts to provide all learners with two things:

1. Sufficient time to master all important concepts and skills

2. The best instructional approaches available

Ingram Publishing/Thinkstock

▲ Determining what is meant by mas- tery is not always a simple task. For this Buddhist monk in Yunnan Province, China, the criterion of mastery is enlightenment. Our learning targets are less ambitious, although not always a whole lot clearer.

Instructional Systems Based on Criterion-Referenced Approaches Chapter 6

Characteristics of Mastery Learning The instructional method that Bloom developed based on these considerations has a number of important characteristics. First, it is directed toward mastering course or unit objectives. Bloom believed that any learning sequence can be broken down into sequential steps and clear objectives so that most, if not all, learners can reach them.

Second, Bloom’s approach to instruction makes extensive use of formative evaluation. At every step in the learning sequence, the learner takes one or more mastery quizzes. Progress to the next step in the sequence requires mastery of earlier steps. Mastery quizzes sometimes take the form of computer-adaptive testing (CAT. As we see in Chapter 10, computer-adaptive testing uses a computer to administer a sequence of items that are selected on the basis of the learner’s responses. In essence, the computer estimates the candidate’s ability based on responses to preceding items and selects the next item accordingly (Shapiro & Gebhardt, 2012). As a result, a good computer-adaptive test can yield a great deal of information with far fewer items than a more conventional test. When CAT is used in mastery learning, the computer is programmed to determine whether the learner has achieved a criterion of mas- tery. Weiss and Kingsbury (1984) cite evidence that computer-adaptive mastery testing can sometimes be more efficient and effective than more conventional quizzes.

Third, mastery learning requires the use of a wide variety of corrective procedures guided by ongoing formative evaluation. While many of these corrective procedures will make use of traditional approaches to teaching, a variety of other classroom activities are highly recom- mended. These include study sessions, individualized tutoring, reviewing, cooperative group approaches, and a wide variety of instructional materials and different media.

Figure 6.6: Bloom’s view of the determinants of degree of learning

In Bloom’s model, how much the child learns is a function of time spent versus time needed. Time needed is, in turn, a function of ability, of the effectiveness of instruction, and of other factors such as motivation. The examples shown here are hypothetical.

f06.06_EDU645.ai

Degree of Learning = ƒ )( time spenttime needed

Degree of Learning = ƒ )( time spenttime needed

Degree of Learning = ƒ )( time spenttime needed

Examples: If Arnold needs 4 hours to reach criterion X, but only spends 2 hours studying, his estimated degree of learning is:

4 = 50%

If Sarah needs only 2 hours to reach X, and spends 2 hours learning, her degree of learning is:

2 = 100%

=

=

2

2

Instructional Systems Based on Criterion-Referenced Approaches Chapter 6

Fourth, Bloom suggests that all learners should progress from one unit to the next as a group. This is accomplished by providing additional enrichment for those who are first to master course or unit objectives. As a result, progress occurs at the rate of the slowest learners. All students who master course objectives are given A’s; those who don’t are given I’s (for Incomplete). Figure 6.7 summarizes the main assumptions and features of Bloom’s mastery learning.

Keller’s Personalized System of Instruction (PSI)

Keller’s Personalized System of Instruction (PSI) is another competency-based approach to teaching and learning (Keller, 1968). Originally designed as an approach to classroom teach- ing, PSI has since been used extensively in distance education and computer-based courses. Unlike Bloom’s mastery learning, which is designed for use with groups of learners, PSI was intended for individual instruction. Keller first developed PSI for teaching introductory psychol- ogy, but it has since been used in a variety of other courses (e.g., Ironsmith & Eppler, 2007).

Characteristics of PSI Like Bloom’s mastery learning, PSI requires that (a) course material be broken down into small units, (b) students be given as much time as necessary to master each unit, and (c) criterion- referenced tests be developed to gauge learner progress. At the end of the unit, learners are given a test that covers all the material. In PSI, units are often entire chapters of textbooks.

Unlike Bloom’s mastery learning, PSI does not use traditional teaching methods such as lec- turing or presenting lessons. Instead, the emphasis is on the written word. As Grant and Spencer (2003) explain, in PSI, instructors typically prepare a written study guide containing

Figure 6.7: Basic assumptions and characteristics of mastery learning

The basic elements of Bloom’s mastery learning emphasize its reliance on criterion-referenced for- mative and summative assessment.

f06.07_EDU645.ai

Main Assumptions of

Mastery Learning

Practical Implications

Broad Characteristics of Teaching Methods

Formative evaluation is accompanied by

various instructional procedures including

individual study sessions, cooperative approaches, tutoring, reteaching, and so on.

Teachers need to develop criterion- referenced quizzes and tests to gauge

learning progressions and assist learners.

Learning requires constant formative evaluation to guide

the teaching– learning process.

Content is divided into small, sequential

units. Instruction is directed toward specific targets

describable in terms of measurable criteria.

Learners must be given as much time as

required to master sequentially ordered

competencies.

There are faster learners and slower

learners.

Instructional Systems Based on Criterion-Referenced Approaches Chapter 6

objectives and questions. The guide also outlines precisely what students should do and what they should learn. Nor does PSI make extensive use of remedial materials. Instead, the onus is placed largely on the learner: Students are responsible for their own learning.

In many applications of PSI, computers are used extensively. In addition, student (peer) tutors are closely involved. In many cases, tutors are responsible for administering and marking unit quizzes. The main elements of PSI are summarized in Table 6.1.

Table 6.1 Main features of Keller’s Personalized System of Instruction

Instruction is directed toward mastery.

Objectives to be mastered are clearly specified.

Instructional systems are self-paced and largely individual.

Instruction is based mainly on the written word (chapters rather than lectures).

Material is carefully sequenced in small subunits.

Unit mastery is required before moving on.

Repeated testing is employed.

Proctors (tutors) are used to help learners and to administer tests.

The emphasis is on reward for success rather than penalties for errors.

Lectures are used for motivation, as a reward.

Research and Applications of Competency-Based Approaches

There is evidence that competency-based approaches to education can lead to highly posi- tive attitudes among learners (e.g., Colquitt, Pritchard, & McCollum, 2011). There is evidence, too, that such approaches can be very effective in helping learners reach course goals. Based on several meta-analyses of competency-based approaches to learning (basically, a form of research that summarizes a large group of related studies), Zimmerman and Dibenedetto (2008) conclude: “The effectiveness of mastery learning approaches has been documented in numerous studies” (p. 209). For example, Cracolice and Roth (1996) report evidence that PSI increases student scores by as much as 20 percentile points. As they put it, “What would you do if you discovered an instructional strategy that raised the scores of your students from the 50th percentile to the 70th percentile?” (p. 1). Their answer: “Even given this evidence, few educators employ this instructional strategy today” (p. 1).

There are several reasons for this lack of interest in competency-based approaches: One is that these systems require a dramatic change in the teacher’s role. Teachers become facilitators rather than primary sources of instruction and information. And they become administrators who are charged with recruiting, organizing, and overseeing proctors—which is not always an easy task (Martin, Pear, & Martin, 2002).

Another reason competency-based approaches are not highly popular is the fear that they might lead to boredom among the faster learners. Critics argue that emphasis on objectives that all learners can master might destroy motivation and render the assignment of grades meaning- less and non-reinforcing. If all who work long enough get A’s, there is perhaps little incentive to work hard and finish first. Furthermore, claim the critics, restricting instructional emphases to simple, identifiable objectives might prevent the occurrence of important incidental learning.

Applications and Examples of Summative Assessment Chapter 6

But perhaps the most important reason competency-based approaches are not more popular, despite their demonstrated effectiveness, is that they require an enormous expenditure of time and effort by the teacher—and a fair allotment of intelligence as well. Sadly, these are not always available.

6.5 Applications and Examples of Summative Assessment During a preschool, school, and postsecondary educational career, student achievement will be gauged and summed in a huge variety of ways and at many different times. Here are just a few examples:

• End-of-unit or end-of-chapter tests

• End-of-term or end-of-year tests

• State assessments

• District-wide benchmark or achievement tests

• All assessment procedures used for calculating grades

• Assessment procedures used as measures of school and teacher accountability

• A variety of professionally prepared high-stakes achievement tests

Summative assessments can take many different forms. Often assessments will involve the evaluation of written, keyboard, or touch-screen responses to teacher-made tests, to stan- dardized tests, or to other district, state, or even national assessments. Some assessments may require the learner to do or perform in ways that do not require written responses—termed performance-based assessment. Other assessments might involve rubrics or various kinds of checklists, portfolios, demonstrations, or other indicators of development and achievement. Chapter 7 focuses on these different approaches to performance-based assessment.

Planning for Summative Assessment

Planning for summative assessment should begin well before the start of instruction. Developing an appropriate and systematic plan for summative assessment requires that teach- ers answer a sequence of questions similar to those provided here.

1. What Should My Students Learn? It all originates with a clear statement of instructional objectives: What will students learn? What skills, values, attitudes, and understanding will they develop?

Clearly, instructional objectives are not entirely the prerogative of the classroom teacher. Mr. Smithers, a third-grade teacher, is not likely to be told, “Well, Smithers, why don’t you go ahead and decide what your third-graders should learn in arithmetic this year?” No. There are certain curriculum requirements, certain state-determined standards that specify what third-graders should know. For example, among the 2010 common core state standards for mathematics developed by the Council of Chief State School Officers is the statement that third-graders should be able to “Interpret products of whole numbers, e.g., interpret 5 × 7 as the total number of objects in 5 groups of 7 objects each” (Common Core State Standards: Implementation Tools and Resources, 2013). Statements of these standards also include sam- ple test items (Figure 6.8). These can be used as a basis for generating items for both summa- tive and formative assessments.

Applications and Examples of Summative Assessment Chapter 6

Unfortunately, statements of statewide (common) standards don’t always simplify matters very much. For example, the sample standard given in Figure 6.8 is just one of dozens of dif- ferent standards intended to serve as guides for instructional objectives. In addition, these many standards are each linked to a variety of specific skills. As Marzano (2000) notes, lists of educational standards and instructional objectives have many problems. Here are some of them:

• There are often too many standards listed. Interpreting and using them becomes con- fusing and difficult.

• Many standards are overly ambitious: They represent lofty but unrealistic ideals.

• Over-adherence to standards can be restrictive for teachers, constraining their instruc- tional approaches and limiting the spontaneity and creativity of their efforts.

• Some prescribed standards make excessive and difficult demands on learners.

• Standards-based curricula and assessment have led to what many consider to be the overuse of testing and, indirectly, to some of the negative implications of high-stakes testing (see Chapter 10).

Figure 6.8: Sample items for one common core standard, third-grade mathematics

The common core standards developed by the National Governors Association Center for Best Practices and Council of Chief State School Officers lists dozens of standards for each grade level. These define academic expectations in English-language arts and mathematics. In many states, these are the basis for standardized summative assessments. (Based on Common Core State Standards: Implementation Tools and Resources: National Governors Association Center for Best Practices and Council of Chief State School Officers. Retrieved March 22, 2013, from http://www. ccsso.org/documents/2012/common_core_resources.pdf).

f06.08_EDU645.ai

Standard 3, Third-Grade Mathematics: Multiply one-digit whole numbers by multiples of 10 in the range 10–90 using strategies based on place value and properties of operations.

Sample Items:

Which number’s underlined digit is worth 8? ____ 834 ____ 384 ____ 438

In 698, which number is in the hundreds place? ____ 6 ____ 9 ____ 8

Multiply:

7 × 50 =

Source: Reprinted by permission of the NGA Center for Best Practices (NGA Center) and the Council of Chief State School Officers.

Applications and Examples of Summative Assessment Chapter 6

Clearly, however, these criticisms and problems don’t apply to all uses of educational stan- dards. There is little doubt that clear instructional objectives are absolutely fundamental to good educational assessment. And, as we see in Chapter 10, statewide standards can go a long way toward making schools and teachers more accountable and, ultimately, toward improving teaching and learning.

2. How Will I Structure Instruction? As we saw in Chapter 1, recognition of instructional objectives should then be followed by careful planning designed to bring instructional methods into alignment with instructional objectives. Hence, the statement of instructional targets should be followed by answers to the questions: How can I help learners achieve these targets? What teaching and learning approaches are most appropriate for this content? For these instructional objectives? For this group of learners? For each learner in my class?

3. How Will I Know Learning Has Occurred? Along with planning for instruction, the teacher must also plan for assessment. Might place- ment or diagnostic tests be useful before instruction? What sorts of informal and formal assessment will I use for formative purposes during instruction?

Some important considerations should be kept in mind when planning for assessment:

• Most important, emphasis should be placed on assessments that are most likely to enhance learning. That is, the formative rather than the summative functions of assess- ment are most important.

• Care should be taken to use approaches to assessment that are interesting, meaning- ful, and motivating for the 21st-century learner. Recall that, as we saw in Chapter 3, this learner is a digital native, reared with computers, video games, the Internet, social media, and wireless technology. Some of the educationally important identifying characteristics of many—though clearly not all—of today’s learners are summarized in Figure 6.9.

• Assessments should be closely aligned with instructional objectives. For example, if instructional objectives have more to do with higher order thinking skills than with the learning of specific facts, assessments should not require learners to reproduce verba- tim chunks of information.

4. How Will Assessment Information Be Gathered and Used? Planning for assessment requires that the teacher decide beforehand what sorts of assessment methods to use as well as how assessment information will be gathered, how often assess- ments will occur, and when they will take place. Ideally, the teacher should now prepare a writ- ten assessment plan for the unit, term, or course that is to be assessed. This plan should contain information that reflects answers to many of the questions listed in this section, relating to

• The most important instructional objectives for the assessment period

• The nature of the evidence that will be used to determine progress and achievement relative to these targets

• How assessment might be used to improve teaching and learning

An example of a general assessment plan for a phonics unit in early reading instruction is shown in Figure 6.10.

Applications and Examples of Summative Assessment Chapter 6

5. How Can I Ensure That Assessment Procedures Are Valid, Fair, and Reliable? Recall from Chapter 2 that an assessment is valid when it measures what it is intended to measure; it is reliable when it measures consistently and as accurately as possible; and it is fair when it is free from biases, it assesses instructional objectives that all learners have had an opportunity to reach, it gives candidates sufficient time, and it makes accommodations for special needs as required.

As we saw in Chapter 2, teachers can increase the validity and reliability of their measures in several ways. We know, for example, that using test blueprints (tables of specifications) of the kind described and illustrated in Chapter 2 can do much to ensure that assessment items are aligned with instructional objectives. Assessments that are not well aligned with instructional objectives are less likely to be valid measures of those targets. (Chapters 4 and 8 also present more information about test blueprints, along with some illustrative samples.)

Figure 6.9: Some educationally relevant common characteristics of today’s learner

Although these descriptions do not apply to all of today’s learners, they are common enough to be tremendously important for those teachers who want to excel. Increasingly, effective teaching depends on the skilled use and understanding of educational technology.

f06.09_EDU645.ai

. . . have a relatively shorter attention span.

. . . be more independent

and less authority oriented.

. . . be more visually oriented than textbook

oriented.

. . . multitask with relative ease.. . . be highly social.

. . . have little patience for traditional

lectures and lessons.

. . . believe that learning

should be fun.

Today’s learners receive

information quite rapidly from a variety of

high- impact, colorful, multimodal, high-

interest sources. As a result, they

may. . .

Applications and Examples of Summative Assessment Chapter 6

Test reliability is affected by several factors:

• The length of tests: In general, tests consisting of more items tend to be more reliable than those with fewer items.

• The stability of a characteristic: Measures of characteristics that fluctuate are often inconsistent and therefore, by definition, unreliable.

• Item difficulty: Tests comprising moderately difficult items tend to be more reliable than tests that contain very difficult or extraordinarily easy items.

6. How Will I Calculate and Report the Results of My Summative Assessments? Finally, the teacher needs to answer questions relating to how the results of summative assess- ment are to be calculated and reported: What role will the teacher’s professional judgment play in the final assessment? What provision is there for reconciling dramatic differences between end-of-unit or end-of-term results and performance during the term? In what form will results be communicated to those who have a right to know (usually students, parents or guardians, and educational personnel)? These are questions that we look at in detail in Chapter 11.

Figure 6.10: Sample assessment plan

Assessment plans might cover a single unit of instruction, as shown here. Or they might be devel- oped for an entire term or course. Where a variety of tests and quizzes are used, and where assessments include other information and indicators such as class participation, improvement, homework, class assignments, and so on, the nature and weighting of these contributors would also be included in the assessment plan.

f06.10_EDU645.ai

Unit 1 Learning Target: Basic Phonics:

Letter and Word Recognition

Instructional Time Period

Formative Assessment

Summative Assessment

Grading

• Recognizing letters and sounds of alphabet • Learning initial, medial, and final sounds in one-syllable words; first

10 pages of Basic Blue Reader • Translating letter patterns into spoken language; Basic Blue Reader

word list 1

• Target 1: 2 weeks • Target 2: 2 weeks • Target 3: 1 week

• Daily one-on-one interaction with each learner; observation of progress; guidance and questioning to identify problems; integrated checks for learning

• A criterion-referenced performance evaluation at the end of five-week period. Criterion: Reads with no more than 3 errors summative reading quiz 1, Page 9, Teacher's guide for Basic Blue Reader.

• “Satisfactory” if criterion is met; “Progressing well” if not met. If appropriate, more specific descriptors will be used at the end of second grading period.

First-Grade Reading Program: Building Reading Skills: Phonics Target:

Applications and Examples of Summative Assessment Chapter 6

Figure 6.11 summarizes the decision-making sequence involved in planning for summative assessment. Given the often significant and long-term implications of the results of summa- tive assessments, asking and answering each of the questions shown in Figure 6.11 is hugely important.

But asking and answering the questions may not be the most difficult part of the task: Implementing the answers might pose the biggest challenges.

In addition to asking these questions, planning for assessment often involves generating or finding appropriate test items and instruments for the more formal aspects of assessment. Some of these will be teacher constructed (discussed in Chapter 8); others will be commer- cially prepared. Some commercially prepared tests consist of items produced to accompany and be closely aligned with a specific text or curriculum; others might be standardized tests developed to gauge the extent to which learners have reached state-defined standards. Standardized tests are discussed in Chapter 10.

Figure 6.11: Planning for summative assessment

Planning for summative assessment should begin before instruction and requires not only answer- ing each of these questions—and others—but implementing the answers.

f06.11_EDU645.ai

3. How will I know learning

has occurred?

2. How will I structure

instruction?

4. How will assessment

information be gathered and

used?

5. How can I ensure that assessment

procedures are fair, valid, and

reliable?

1. What should my students

learn?

6. How will I calculate and

report the results of summative assessments?

Questions to answer when

planning summative assessment

Applications and Examples of Summative Assessment Chapter 6

The next two sections look at two other approaches to educational assessment: curriculum- based measurement and benchmark testing. These approaches are especially useful for gaug- ing progress through a curriculum, and for monitoring the extent to which course content is being mastered.

Curriculum-Based Measurement (CBM)

Closely related to the criterion-referenced mastery testing advocated by Bloom and Keller is an approach labeled curriculum-based measurement (CBM; also referred to as curriculum-based assessment or CBA). As the label implies, curriculum-based measurement attempts to measure basic skills and knowledge that are tied directly to the curriculum. In CBM, the assessments are formative as well as summative: They occur during an instructional sequence and are intended to foster learning.

Curriculum-based measurement is most common in core subject areas such as lan- guage arts and mathematics. It is designed mainly to provide information about how students are progressing. Curriculum- based assessments generally consist of a small selection of tasks or problems that are actually part of the curriculum. The learner’s performance on these serves as a direct indicator of progress. For example, a CBM measure in arithmetic might pre- sent the learner with four or five problems that closely reflect the curriculum being taught. In an early reading program, a CBM assessment might count the number of words the learner reads correctly in a fixed period of time. Following a series of such assessments, the teacher can plot the results to provide a graphic represen- tation of learner progress.

CBM Probes CBM measures are sometimes developed by classroom teachers, but they are also widely available commercially in the form of what are often termed CBM probes. In their commer- cialized form, these assessments are given under standardized conditions: They are timed, they use consistent instructions, and they are marked according to predetermined procedures. Performance is typically scored for speed and accuracy. CBM measures are usually given fre- quently and serve as an indicator of student progress and as a source of guidance both for instruction and for learning.

CBM probes are usually very brief series of questions, test items, or performances that can be administered in a few minutes. The most common CBM probes have been developed for curriculum areas such as arithmetic, spelling, reading, and writing. For example, CBM probes in mathematics might consist of worksheets that tap single skills (such as adding two-digit numbers), or they might be probes that look at a variety of skills (such as addition and multi- plication). Figure 6.12 presents an example of a single-skill mathematics CBM probe.

iStockphoto/Thinkstock

▲ Curriculum-based measurement (CBM) focuses on basic skills and knowledge tied directly to the curriculum. It is most common in core subject areas. Some fear that with increas- ing emphasis on common core standards, subjects such as art, drama, physical education, and—much to these children’s dis- appointment—music will increasingly be ignored.

Applications and Examples of Summative Assessment Chapter 6

Uses and Advantages of CBM Measures An important advantage of CBM measures over other standardized norm-referenced assess- ments is that they are tied more closely to the school’s curriculum. Because test makers want the largest possible number of users, most standardized tests are normed using national samples of individuals and of curriculum content. As a result, some standardized tests might be a poor match for a district or state curriculum.

Other advantages of CBM measures are that they can be administered quickly, often in just a few minutes. And because they can be given often, they are highly sensitive to short-term gains or declines. Accordingly, they can provide useful information about changes in learners as well as about the effectiveness of instructional procedures. Not surprisingly, CBM probes have been widely used by teachers of children with special needs (Shinn, 2008). They have proven very useful for assessing learning problems and for helping students who have been performing less well than expected (Anderson, Lai, Alonzo, & Tindal, 2011).

Benchmark Assessment

CBM measurement is typically a recurring, classroom-based form of assessment that occurs during instruction and that is designed to provide a picture of student progress toward speci- fied goals that are often explicit in statewide standards. They are basically a form of what is widely referred to as benchmark assessment.

Literally, a benchmark is a standard of excellence or achievement that provides a basis for comparing and judging something. Accordingly, benchmark assessment in education is a way of comparing achievement to some standard. But benchmark assessments are quite different from annual statewide tests, whose main purpose is determining to what extent mandated standards have been met. Benchmark tests are typically administered at regular intervals dur- ing a course or a term; statewide tests are typically given at the end.

Unlike formative assessments, benchmark assessments are not an intrinsic part of ongoing instruction: They are not designed specifically to provide teachers and learners with informa- tion that is immediately useful for modifying and improving teaching and learning. Instead,

Figure 6.12: A teacher-made CBM math probe

Curriculum-based measurement (CBM) presents brief, commercially prepared or teacher-made probes designed to assess the extent to which learners are developing skills and understanding related directly to the curriculum. In addition to its use as an indication of progress (a summa- tive function), CBM provides feedback that can be useful for teaching and learning (a formative purpose).

f06.12_EDU645.ai

Example of a Single-Skill, Teacher-Made CBM Math Probe

Second-grade students are asked to add as many of the following pairs of single-digit numbers as they can in one minute:

7 + 6 = _____ 5 + 8 = _____ 2 + 9 = _____ 9 + 5 = _____

6 + 8 = _____ 4 + 9 = _____ 5 + 6 = _____ 7 + 9 = _____

Applications and Examples of Summative Assessment Chapter 6

they are periodic summative assessments. In many jurisdictions, they are uniform across grades and schools. Simply put, their main purpose is to show progress by evaluating or checking achievement in comparison with a standard (a benchmark).

A large number of benchmark assessments are available, many of them commercially pre- pared; others are developed locally by teachers or by personnel at the district or state level. One example is Dynamic Indicators of Basic Early Literacy Skills—abbreviated as DIBELS (Good, Kaminski, Simmons, & Kame’enui, 2001). DIBELS is a collection of benchmark measures that looks at sequential growth and development of early literacy skills. Its main purpose is to iden- tify children who need intervention and to assess the adequacy of intervention strategies. To accomplish this, it looks at benchmarks in developing reading skills, such as the child’s aware- ness of the sounds of letters (phonological awareness), and the ability to recognize syllables. One of the benchmark DIBELS tests, for example, asks students to read nonsense words. Doing so correctly reflects knowledge of letter–sound correspondences (rather than simple recognition of a word pattern learned by rote).

As is the case for all educational assessment instruments, the most useful benchmark assess- ments are characterized by the qualities of good measuring instruments described in Chapter 2: test fairness, validity, reliability, and alignment with instruction, curriculum, and instruc- tional objectives. Alignment with state standards is especially important for commercially prepared benchmark assessments, because not all of them reflect the goals of all state edu- cational systems.

Accordingly, teachers who want to prepare their own benchmark tests might profit from the guidelines summarized in Figure 6.13.

Although their main purpose is to monitor progress toward final goals, benchmark assess- ments serve a number of other purposes (Herman, Osmundson, & Dietel, 2010):

1. They help teachers recognize learners who are in need of help.

2. They communicate expectations not only to learners but also to teachers and parents. Implicit in the content of benchmark assessments is information about what is important to learn.

3. As a way of monitoring student progress, they also provide teachers and school systems with information about the effectiveness and appropriateness of instructional materials and strategies.

4. By providing ongoing evaluation and monitoring of student progress, they yield informa- tion that can guide instruction.

Interestingly, there also is evidence that benchmark testing can improve student learning (Shrago & Smith, 2006). Hence, it serves formative as well as summative purposes. The iden- tifying characteristics of benchmark assessment are summarized in Figure 6.14.)

A Reminder: The Main Purpose of Educational Assessment

As stated in Chapter 1, “The most important purpose of all the different approaches to edu- cational assessment is to foster learning.”

The point bears repeating. True, the specific purpose of summative assessment is to summa- rize learner progress and achievement and to provide learning-relevant information that can

Applications and Examples of Summative Assessment Chapter 6

Figure 6.13: Guidelines for teacher-prepped benchmark assessment

Suggestions for preparing and using benchmark assessments.

f06.13_EDU645.ai

Guidelines for Preparing Benchmark Assessment Instruments

5. Consider Feasibility and Utility It’s important to consider whether benchmark testing improves student

learning, and the extent to which test results are meaningful and useful. Is the testing program worth the effort, time, and money expended?

4. Attend to Characteristics of Good Measuring Instruments Benchmark assessments need to be fair, valid, and reliable.

3. Focus on Big Ideas Assessments should be built around the key ideas that define curriculum.

2. Map Content To ensure alignment, curriculum content should be “mapped” in terms of

specific content and skills learners need to acquire.

1. Align with State Standards Tests must reflect state standards and statewide assessments.

Source: Based on Herman, J. L., & Baker, E.L. (2005). Making benchmark testing work. Educational Leadership, 63(3), 48–54.

Figure 6.14: Characteristics of benchmark assessments

Benchmark assessments are designed to assess the learner’s progress toward course standards. They are intended mainly as periodic summative assessments.

f06.14_EDU645.ai

• Monitor learner progress toward final goals.

• Are often commercially prepared by test-preparation firms. May be part of published text-related instructional material for a course.

• Are given periodically, usually at least three times a year, but sometimes as often as monthly.

• Focus mainly on core subjects such as reading and mathematics.

• Reflect statewide standards or school district standards.

• Provide a “benchmark” measure of learner progress

relative to state standards.

Defining Characteristics of Benchmark Assessments

Section Summaries Chapter 6

then be communicated to learners (and others). Similarly, the defining purpose of diagnostic assessment is to identify (diagnose) problems and weaknesses; the main purpose of place- ment assessment is to provide a basis for selection and placement; and the chief purpose of formative assessment is to provide information that is immediately useful to teachers and learners.

But ultimately, the goal of all these assessments is to help learners.

It’s also important to keep in mind that test items, informal questions, and the variety of tasks that might be used for assessment are not, in and of themselves, summative, formative, diag- nostic, or what have you. Rather, we label them that way depending on how they are used.

In Chapter 7, we look at a large variety of tasks that can be used for summative, formative, diagnostic, or placement purposes. These tasks fall under the general label of performance assessment.

Chapter 6 Themes and Questions

Section Summaries 6.1 The Nature of Summative Assessment Although educational assessments can be dis- tinguished in terms of their main purposes, the overall goal of all forms of educational assess- ment is to provide learners with the best educational experience possible, maximizing the attainment of instructional targets. The main purpose of summative assessment is to provide an indication of achievement after a unit of study, a term, or at year-end. Given its importance and implications, summative assessment should be carefully planned. Comparisons in sum- mative assessment can be to the performance of other learners (norm-referenced); to prede- termined criteria (criterion-referenced); or to the individual’s prior or expected performance (self-referenced).

6.2 Norm-Referenced Interpretations In norm-referenced approaches, the performance of other students sets the norm or standard to which the individual is compared. This inter- pretation is highly compatible with competitive approaches to education. It permits relative ranking of learners and is therefore useful for comparing learners.

6.3 Criterion- and Self-Referenced Approaches Comparing individual performance to predetermined criteria is the basis of criterion-referenced interpretations. Criteria in educa- tion are often set by states and may be common across states (common core state stan- dards). Adherence to these standards defines standards-based education. Where educational goals and instructional objectives are determined by state standards, it’s important that there be alignment between the standards, instructional objectives, approaches to teaching, and assessments. Self-referenced interpretations involve comparisons to the individual’s own prior or expected performance. This approach to assessment is useful for gauging learner progress and often has important motivational consequences.

6.4 Instructional Systems Based on Criterion-Referenced Approaches Instructional systems based on criterion-referenced interpretations are designed to develop competency relative to specific criteria and are known as competency-based systems. Among the bet- ter known competency-based systems is Bloom’s mastery learning, which is based on the assumption that all learners can master course content if given sufficient time and support.

Key Terms Chapter 6

It makes extensive use of formative assessment as learners progress through content that has been divided into small, sequential steps. Keller’s Personalized System of Instruction (PSI) breaks textual material into small units with accompanying objectives, study guide, and quiz- zes. There is evidence that these approaches can be highly effective. However, they are time consuming and sometimes difficult to organize and implement.

6.5 Applications and Examples of Summative Assessment Planning for summative assess- ment requires a clear view of instructional objectives as well as decisions about approaches to instruction, assessment, and the use of assessment results. Curriculum-based measure- ment (CBM) is a form of assessment that uses probes that are embedded in the curriculum to determine learner progress. Benchmark assessments are periodic summative assessments of learner progress administered at regular intervals during a term rather than at the end. The main purpose of all educational assessment is to foster learning.

Applied Questions 1. What is summative assessment? Explain the identifying distinctions between summative

and formative assessment.

2. What are the main differences between norm-referenced and criterion-referenced inter- pretations of assessment data? Design a graphic that compares norm-referenced, criterion- referenced, and self-referenced comparisons.

3. What are some instructional systems based on criterion-referenced comparisons? Using library and online resources, write a brief description and evaluation of competency-based instructional and assessment systems.

4. What is curriculum-based measurement? Design sample items that might be used as CBM probes in your area of expertise.

5. What are benchmark assessments? Define benchmark assessment and illustrate its usefulness.

Key Terms benchmark assessment Periodic assessments undertaken during instruction and used to monitor learner progress toward learning targets.

CBM probes Standardized sets of test items and quizzes used in curriculum-based mea- surement; typically available for mathematics, spelling, reading, and writing.

common core standards Statements of basic, minimum requirements describing the skills and knowledge students are expected to develop in core subjects at each grade level and that are common to more than one school or school system.

Common Core State Standards (CCSS) Common core standards that states are encour- aged to adopt and apply to all state schools in an effort to bring into alignment different state curricula. See common core standards.

criterion-referenced An assessment procedure in which the student is judged rela- tive to a criterion rather than relative to the performance of other students. See also norm-referenced.

Key Terms Chapter 6

curriculum-based measurement (CBM) Assessment of skills and knowledge related directly to curriculum content; used as an indicator of learning progress and to assist learn- ing and instruction (formative function) rather than as an indicator of achievement (summa- tive function).

mean The arithmetic average of a set of scores. See also central tendency, median, mode.

norming group A group (sample), judged to be representative of a target population, that serves to set the standards for commercially prepared tests designed for that population.

norm-referenced Assessment where the student competes against the performance of other students rather than in relation to a preestablished criterion of acceptable perfor- mance. See also criterion-referenced.

objective tests A test consisting of questions for which responses are typically brief, fac- tual, and unambiguous. Multiple-choice, matching, and fill-in-the-blanks tests are typically objective tests.

overachiever Describes a person who is ambitious, driven, highly motivated to achieve, and whoconsequently achieve at a higher level than expected based on their presumed abili- ties and skills.

percentile The point at or below which a specified proportion of cases in a distribution lie. For example, the 70th percentile is the point (score) at or below which 70% of all scores in the sample fall.

Personalized System of Instruction (PSI) A mastery learning instructional approach in which course material is broken down into small units, study is largely individual, a variety of study material is available, and progress depends on performance on criterion-referenced unit tests.

portfolio A collection of actual samples of students’ performances and achievements used for assessment.

Self-referenced approach to assessment Assessment that compares the learner’s cur- rent performance with performance at some earlier time or with expected performance based on ability, experience, and other personal factors.

self-regulated learners Learners who have mastered the art of learning how to learn, which includes effective self-assessment skills and the desire to take control of personal learning. See self-regulated learning.

standards-based grading Grading based on assessments that look at the learner’s per- formance with respect to clearly defined standards of proficiency or mastery.

underachiever Describes a person whose achievement and progress is less than what might be expected based on presumed abilities.