Creation of Summative Assessment

profilenjordan_23
FOCUS.docx

FOCUS QUESTIONHow do you go about developing a rubric for performance tasks?

Among the pervasive problems in educational assessment is the fact that many of the things we need to measure are very complex.

Determining whether first graders know the alphabet or whether second graders have learned how to add three numbers is relatively simple: We can request that they recite their ABCs, or we can ask them, "Hey, what is 2 plus 6 plus 3?" But assessing the extent to which they can read fluently or evaluating the clarity of their thought processes is quite another matter.

For both of these educational assessments—and many others—we are likely to require students to perform or demonstrate what they can do.

But unless we are careful, the evaluations that result from these performances might be so subjective as to be of little value.

And if that is the case, as Trace, Janssen, and Meier (2017) found when they looked at the ratings four different examiners gave for 524 second-language essays, we will often disagree with each other when we evaluate our learners. In measurement terms, our assessments will lack validity and reliability. Moreover, they may be biased and unfair and of little worth if we want to generalize from them.

Analytic and Holistic Rubrics

One way to make assessment of performances more objective, reliable, and valid is to use carefully developed guidelines on which we can agree. Rubrics, such as the one shown in Figure 7.5, are a way of expressing these guidelines.

Rubrics specify the criteria and the standards that might be used to assess a performance.

In addition, most rubrics include a scale that denotes different levels of achievement or competence. For example, the criteria that make up a rubric might be scaled from poor to excellent or from does not meet standard to exceeds standard, with a number of steps in between.

The rubric thus communicates to the learner precisely what is required for performance at successive levels of excellence.

A distinction is sometimes made between two different kinds of rubrics in terms of the scores and information they yield. The holistic rubric is designed to yield a single, global assessment; the analytic rubric may also provide an overall score, but it is purposely designed to yield scores or ratings for different aspects of the performance.

A holistic rubric often consists of a single criterion with which to assess a performance, thus resulting in a global score or rating. As a result, holistic rubrics often take the form of a paragraph or even just one sentence, describing a standard of acceptable performance. An example of a simple holistic rubric would be "The student will be able to recite the alphabet correctly within 55 seconds."

Analytic rubrics typically provide a list of criteria by which a performance can be assessed as well as detailed guidelines for evaluating performance. They typically take the form of a two-dimensional grid, as shown in Figure 7.7, where rows represent criteria and columns are used to describe levels of performance relative to each criterion.

A two-dimensional, grid-type rubric can also be used in a holistic sense whereby, instead of giving a rating for each criterion, the examiner weighs the relative ratings on each to arrive at a single score or rating. When using the rubric in an analytic sense, the examiner assigns a rating relative to each criterion and might, or might not, determine an overall rating. For example, a sixth-grade teacher using the essay scoring rubric shown in Figure 7.7 will evaluate each learner's essay on all six criteria. Every criterion will be assigned a numerical score between 1 and 4, reflecting the student's performance. When using the rubric in an analytic sense, the teacher will communicate a separate rating for each of the six criteria shown. When using the same rubric in a holistic manner, the teacher will calculate a single rating or score determined by assigning a weighting to each criterion. If focusing on the main topic is more important than having a good introduction, it is given more weighting in arriving at an overall assessment. (For an example of a ready-to-use, free, public rubric, see https://www.rcampus.com/rubricshowc.cfm?code=A36CAW&sp=yes&.)

Developing Rubrics

Figure 7.6 summarizes the four steps just described for devising a performance-based approach to assessment.

Summative Assessment

FOCUS QUESTIONSWhat is the purpose of summative assessment? What are the main differences between norm-referenced and criterion-referenced interpretations of assessment data?

Much like daily life in Ochawa, schools present our learners with various tests—although these are not normally a matter of life and death. Still, there is sometimes a very close parallel between assessment in Ochawa and summative assessment in schools. Recall that summative assessment is the type of assessment that normally occurs at the end of an instructional sequence and is designed mainly to provide a grade.

What Is Summative Assessment?

In previous chapters, we distinguished among four kinds of assessment:

· Formative assessment: An integral part of instruction designed mainly to provide immediate, ongoing feedback to assist learners and teachers in improving the teaching–learning process. Occurs during instruction.

· Placement assessment: Preinstruction assessment used for making selection and placement decisions.

· Diagnostic assessment: Assessment that often occurs prior to instruction and is then used for placement purposes. Directed at uncovering learner strengths and weaknesses, allowing for differentiated assessment and differentiated instruction. Differentiation in education refers to assessment procedures that identify important differences among learners and that lead to placements and instructional procedures that accommodate these differences.

· Summative assessment: Assessment that typically occurs at the end of an instructional sequence and is used to summarize student progress and achievement and to provide a grade.

LEARN MORE

Main Purposes of Summative Assessment

The main purpose of summative assessment is somewhat different from the purposes of the other three categories of assessment.

Diagnostic and placement assessment are generally preinstruction assessments designed to provide information for selecting learners and placing them in appropriate groups, classes, or programs and for selecting appropriate interventions as required; formative assessment occurs during learning and is specifically directed toward improving learning and instruction.

Summative assessment, on the other hand, occurs after instruction and is a measure of the outcomes of instructional experiences. In other words, its purpose is summative: It provides an indication of the sum or total of the effects of schooling.

As Black (1998) explains, when the chef tastes the soup, she is involved in formative assessment: She can still add herbs, thicken the broth, simmer the ingredients, and put in new spices, new vegetables, new meats, and new starches. But when the customer tastes the soup, he is engaged in summative assessment: The time to change and improve the broth is past. Now is the time for the final grade.

As mentioned earlier, a useful way of distinguishing between summative assessment and other approaches to assessment is implicit in the observation that formative, diagnostic, and placement assessment are in a sense assessment for learning. In contrast, summative assessment is assessment of learning (see Figure 6.1).

Figure 6.1 Some distinctions between summative and formative assessment

Despite their differences in principal purposes, timing, and uses, a single assessment might serve both formative and summative functions.

Some of the differences between formative assessment and summative assessment are that formative assessment is assessment for learning, occurs during learning, is used to improve teaching and learning, and is where the learner is involved with the teacher in interpreting and using the results of assessment. Summative assessment is the assessment of learning that occurs after learning, is used to summarize the effects of teaching and learning, and the learner is less involved.

In summary, the main purpose of summative assessment is to provide a mark or a grade that reflects progress and achievement. At the same time, however, summative assessments can be used to draw conclusions about the effectiveness of school programs and of instructional strategies. Summative assessments also say something about the appropriateness of curriculum offerings, the readiness of learners, and perhaps the characteristics of learners and teachers.

Comparisons in Summative Assessment

Summative assessment almost invariably involves comparisons. The most common approach in many schools is to compare the performance of each learner to the performance of other learners.

What this would mean in Ochawa is that even if all inhabitants succeeded in climbing beyond the reach of the predators, those who came in last might be sacrificed; they would have failed this norm-referenced comparison. And those who climbed first and highest might receive high praise to make up for their greater hunger.

In some assessment situations, however, the performance of each learner is not compared to that of other students but instead is compared to a standard (a criterion).

Assume, for example, that students are expected to reach a certain level of competence—that is, to learn certain identifiable concepts and develop a repertoire of specific skills. In this case, summative evaluation might involve comparing their performance to criteria that denote attainment of these concepts and skills. Learning these concepts and skills can be viewed as defining the criteria of success—criteria we can denote as X. In a sense, X is analogous to escaping from the beasts in Ochawa: Providing they escape the beasts, those who arrive last succeed just as surely as those who climb first. This is referred to as a criterion- or standards-referenced comparison.

A third option is also available: it compares each learner's performance not to a standard (as in criterion-referenced interpretations), nor to the average performance of other comparable learners (as in norm-referenced comparisons).

Instead, learners are compared to themselves in what is termed a self-referenced approach to assessment.

Self-referenced interpretations compare the learner's current performance with earlier performances or with performance that is expected based on ability, experience, and other personal factors. Teacher comments such as "Elvira is not working up to her ability" or "With his long legs, Renaldo should be able to complete three more laps after 8 weeks of training" are examples of self-referenced assessment.

To summarize, summative assessments typically involve one or more of three principal kinds of comparisons:

1. Norm-referenced interpretations: Comparisons with the performance of other similar learners.

2. Criterion- or standards-based interpretations: Comparisons with a predefined criterion or standard of acceptable or expected performance.

3. Self-referenced approach to assessment: Comparisons of each learner with him- or herself, often reflecting expectations based on measured or assumed ability, background, and personal factors such as motivation and parental encouragement (see Table 6.1).

Table 6.1 Comparison bases for summative assessment

Comparison group

Type of comparison

Explanation

Example

Other learners

Norm-referenced

Learner's performance is compared with that of peers.

The average score of the group, no matter what it is, is assigned a C grade, and all other scores are distributed around this mark.

Predetermined standards

Criterion-referenced

Learner is judged relative to some standard of acceptable performance.

Learners are assigned a "pass–fail" grade according to whether they achieve at a predetermined level.

Self

Self-referenced

Learner's performance is assessed relative to previous or expected performance.

Learners are evaluated in terms of their improvement.

Each of these three approaches to interpreting test scores—criterion-referenced, norm-referenced, and self-referenced—can be used for summative purposes.

Figure 7.6 Developing performance-based assessment

Teacher-constructed performance-based assessment tasks may well require teachers to progress sequentially through these four steps, paying attention to the cautions and criteria summarized in the text. But in many cases ready-made and entirely appropriate performance-based assessments are available from a variety of sources, including commercially prepared instructional and assessment guides that often accompany courses or course material, system or state assessment guides and standards, and web-based material.

Four boxes stacked with an arrow moving from one box to the next one below. The top box is labeled "Step 1. Deciding on the main purpose of the assessment." The second box is labeled "Step 2. Establishing learning targets." The third box is labeled "Step 3. Selecting and developing performance tasks." The bottommost box is labeled "Step 4. Developing a scoring system."

Note that the final step in this process is developing a scoring system. Basically, this step involves identifying and refining criteria that can be used to assess the performance of a task or the quality of a product.

Identifying Criteria

Say, for example, that you want to devise a performance-based assessment task for essay writing in a sixth-grade class. First you need to know the criteria for a good essay. A variety of approaches are possible: You might begin by researching published material, online or hard copy, and ferreting around for information about how your state, your school, or other authorities define the qualities of a good sixth-grade essay. You might also look at the pile of sixth-grade essays you or others have collected over the years and analyze them, trying to discover what makes some better than others. Or, better yet, why not spearhead the formation of a PLC so that you can consult and collaborate with other educators and perhaps even with parents and students? What do others think the characteristics of a good essay are? Why?

In the end, you might come up with a list of criteria that looks something like this:

A good sixth grade essay:

· focuses on its main topic

· is interesting and clearly written

· uses appropriate vocabulary and grammar

· has a good introduction

· is logically sequenced

· contains material that is demonstrably factual or opinions that are identified as such

Note that these simplified criteria reflect important common core standards that might be adopted by a state or a school jurisdiction.

Chapter 8

Planning for Teacher-Made Tests

FOCUS QUESTIONWhat are some important steps in planning for assessment?

Reading the chapter might have improved Mr. Moskal's construction of teacher-made tests (as opposed to commercially prepared standardized tests, discussed in Chapter 10).

He would know that he should not rely solely on his memory and intuition when constructing a test but should begin with a clear notion of his educational goals. He then needs to decide on the best ways of determining whether his learners have reached these goals. If his assessments are to be useful for determining how well his students have learned (summative function of tests) and for improving their learning (formative function of tests), he will need some detailed test blueprints, and perhaps some rubrics and checklists, to help him evaluate student performances.

Identify Goals and Learning Objectives

Educational goals are the nation's, state's, school district's, or teacher's general statements of the broad intended outcomes of the educational process. Learning, or instructional, objectives are more specific statements of intended learning outcomes relative to a lesson, unit, or even course. Whereas educational goals are often somewhat vague and idealistic, the most useful learning objectives for the classroom tend to be very explicit. Most are phrased in terms of behaviors that can be taught and learned and that can be assessed.

National Educational Standards (Goals)

The nation's educational goals (or standards), for example, are often detailed in legislation and regulations. As we saw in Chapter 1, in the United States the legislation governing K–12 education, ESSA, requires all states to develop standards and to devise or select appropriate assessments to determine the extent to which these standards are being met.

Virtually all states have now published descriptions of educational standards and criteria that can be used to assess the extent to which educational goals are being met. As we saw, following a nationwide education initiative involving a consortium of educators, Common Core State Standards (Council of Chief State School Officers, 2018) were developed. These standards describe what students should know at each grade level for each subject. One intended result of adopting common core standards is to bring about a realignment of curricula across different states.

LEARN MORE

National Science Standards

Another example of common core standards are the Next Generation Science Standards (NGSS) (Next Generation Science Standards, 2018). These standards are the product of a collaboration among 26 states and various national groups, including the National Research Council, the American Association for the Advancement of Science, and the National Science Teachers Association. They are designed to improve the teaching of science in U.S. schools and are built around the notion that there are three distinct dimensions involved in learning science:

1. core ideas in each of the main domains of science (physical science, earth and space science, life science, and engineering design);

2. practices, meaning what it is that scientists do and the range of skills and knowledge required to engage in these practices; and

3. crosscutting, which refers to making connections across the four domains of science.

Accordingly, each of the many standards is described in terms of all three of these dimensions. An illustration of a standard is shown in Table 8.1. More than two thirds of American states have now adopted the NGSS standards or have used them as a basis for developing their own science standards (National Science Teachers Association, 2018).

Table 8.1: NGSS sample standard

Middle school chemical reactions

Students who demonstrate understanding can:

1. Analyze and interpret data on the properties of substances before and after the substances interact to determine if a chemical reaction has occurred

2. Develop and use a model to describe how the total number of atoms does not change in a chemical reaction and thus mass is conserved

3. Undertake a design project to construct, test, and modify a device that either releases or absorbs thermal energy by chemical processes

Science and engineering practices e.g., developing and using models

Disciplinary core ideas e.g., each pure substance has characteristic and physical and chemical properties that can be used to identify it

Crosscutting concepts e.g., matter is conserved because atoms are conserved in physical and chemical processes

Source: Based on "Middle School Chemical Reactions," by National Science Teachers Association, 2018 (http://ngss.nsta.org/DisplayStandard.aspx?view=topic&id=25).

Create Test Blueprints

The best way of ensuring that assessments are directed toward instructional objectives is to use test blueprints. These are basically tables of specifications for developing assessment instruments. They are typically based closely on the instructional objectives for a course or a unit. They may also reflect a list or a hierarchical arrangement of relevant intellectual or motor activities such as those provided by revisions of Bloom's taxonomy (described in Chapter 4).

Many states provide blueprints for large-scale testing. For example, the Ohio Department of Education provides detailed, multigrade test blueprints for mathematics, English language arts, science, and social studies (Ohio Department of Education, 2018; Johnstone & Thurlow, 2012).

Examples of Test Blueprints

Suppose you are teaching sixth-grade mathematics in California. California core standards list detailed objectives at that grade level for five different areas: ratios and proportional relationships, the number system, expressions and equations, geometry, and statistics and probability (Sacramento County Office of Education, 2018a). The first of six core standards for geometry reads as follows:

Find the area of right triangles, other triangles, special quadrilaterals, and polygons by composing into rectangles or decomposing into triangles and other shapes; apply these techniques in the context of solving real-world and mathematical problems. (Sacramento County Office of Education, 2018a, p. 27)

Part of a test blueprint reflecting related learning objectives, based on Bloom's revised taxonomy, might look something like that in Table 8.2. Numbers in the grid indicate the number of test items for each category. Questions in parentheses are examples of the sorts of items that might be used to assess a specific cognitive process with respect to a given topic. Test blueprints of this kind might also include the value assigned to each type of test item.

Table 8.2 Part of a sample test blueprint for a single geometry objective reflecting Bloom's revised taxonomy, cognitive domain

Topic

Remembering

Understanding

Higher processes (applying, analyzing, evaluating, creating)

Right triangles

4 items (e.g., What is the formula for finding the area of a right triangle?)

1 item (e.g., If you were building a house and could have a total of only 80 feet of perimeter wall, which of the following shapes would give you the largest area? Quadrilateral; polygon; square; right-angle triangle; other shape. Prove that your answer is correct.)

Quadrilaterals

3 items

Other triangles

3 items

2 items (e.g., Illustrate how you would find the area of an isosceles triangle by sketching a solution.)

1 item

There are several other approaches to devising test blueprints. For example, the blueprint might list what learners are expected to understand, remember, or be able to do. In addition, the most useful blueprints will include an indication of how many items or questions there might be for each entry in the list and the test value for each. Figure 8.2 gives an example of a checklist that can be used as a blueprint for a unit covering part of the content of Chapter 2 in this text. (For other examples of test blueprints, see Tables 4.4 and 4.5 in Chapter 4.)

Figure 8.2 Checklist test blueprint

A teacher might use a checklist test blueprint such as this as an assessment rubric; a student might use it as a study guide and self-assessment guide.

Example of checklist test blueprint for a unit on "Characteristics of Good Testing Instruments." The box is divided into three areas: fairness (for example, know what test fairness means), validity (for example, understand how test validity can be improved), and reliability (for example, know how reliability is calculated).

Uses and Limitations of Test Blueprints

A blueprint such as that shown in Figure 8.2 is useful for more than simply organizing and writing items for a test. It can also be adapted as a rubric that might be used for assessing learner progress, as well as for guiding learner efforts. Perhaps most important, it directs the attention of both teachers and learners toward the higher levels of mental activity.

In this connection, it is worth noting that despite teachers' best intentions and their most carefully prepared test blueprints, assessments do not always reflect instructional objectives. For a variety of reasons, including that they are much easier to assess, the lowest levels of cognitive activity in Bloom's taxonomy (knowledge and comprehension) are often far more likely to be tapped by school assessments than are the higher levels. As Adams (2015) notes with respect to health education, in spite of the importance of higher level cognitive skills that would foster critical thinking and evaluative judgment, educators "focus overwhelmingly on the lower levels of the taxonomy, knowledge and comprehension" (p. 153). When Broman, Bernholt, and Parchmann (2015) looked at problems used to determine student interest and knowledge in science, they found that almost all test items were assessing lower rather than higher order thinking. Similarly, Jideani and Jideani (2012) report that knowledge- and comprehension-based assessments predominated in their investigation of assessment in food sciences classes. And this was true even though instructors intended that their students go beyond remembering and understanding—that they also learn to apply, analyze, evaluate, and create.

Create Rubrics

As we saw in Chapter 7, another important tool for assessment is the rubric. A rubric is a written guide for assessment. Rubrics are used extensively in performance assessments—in which, without such guides, evaluations are often highly subjective, unpredictable, and unfair. Inconsistent assessments are the hallmark of a lack of test reliability. And measures that are unreliable are also invalid.

Rubrics, like test blueprints, are a guide not only for assessment but also for instruction. Also like blueprints, they are typically given to the learner before instruction begins. They often tell the student what is important and expected far more clearly than might be expressed verbally.

LEARN MORE

Approaches to Classroom Assessment

As we saw earlier, assessment can serve at least three different functions in schools.

1. Assessment might be used for placement purposes before instruction (placement assessment). Diagnostic assessment, which is a form of preassessment used to identify strengths and weakness and to detect learning difficulties, is generally considered a form of placement assessment.

2. Assessment might assume a helping role when feedback from ongoing assessments is given to learners to help them improve their learning and when ongoing assessments suggest to the teacher how instructional strategies might be modified (formative assessment).

3. School assessments often serve to provide a summary of the learner's performance and achievements. These unit- or year-end assessments are usually the basis for grades (summative assessment).

Two students perform a chemistry experiment as their teacher observes.

Monkeybusinessimages/iStock/Getty Images Plus

Because they are closer to real-life situations, performance-based assessments are often described as more authentic assessments. Some of the most important learning targets associated with the chemistry class to which these students belong cannot easily be assessed with a selected-response test. The test is in the performance.

Teacher-made assessments, no matter to which of these uses they are put, can take any one of several forms. Among them are performance-based assessments, selected-response assessments, and constructed-response assessments.