4 EDUCATION DISCUSSIONS DUE IN 72 HOURS
Chapter 11
Understanding Measurement and Testing
Learning Objectives
After reading this chapter, you will be able to:
• Define measurement, assessment, and evaluation and explain how school leaders can use assessment results to drive instruction.
• Discuss the differences between formative and summative assessment.
• Describe various methods of assessment, including traditional and alternative means.
• Explain how school leaders collect and apply growth model metrics.
11
© monkeybusinessimages/iStock/Thinkstock
Believing we can improve schooling with more tests is like believing you can make yourself grow taller by measuring your height.
—Robert Schaeffer of FairTest
Introduction Chapter 11
Introduction So how do we know if a student has mastered a topic, concept, or skill? Ben Franklin once wrote, “Tim was so learned that he could name a horse in nine languages; so ignorant that he bought a cow to ride on.” In today’s educational climate, mastery-based learning is premised on the notion that learning is dependent on proficiency as opposed to the amount of seat time spent on academic work. Most schools currently operate around time-based structures for credit—a 6-hour school day, 9-month school year, with students grouped by age, attend- ing subject matter specific courses, and being assessed by standardized exams. The current system forces students to move forward to stay with their norm (age) groups, rather than focusing on actual competencies. This system is not necessarily designed to support differen- tiated instruction or personalized learning to a level sufficient to allow all students to reach their full potential. And, standardized assessments don’t necessarily measure how well learn- ers are able to apply their knowledge and skills.
In most cases, assessment is used to document deficiencies. Imagine if, instead, assessment were used to gather information for making instructional decisions. Data-driven decision mak- ing is the process by which administrators and teachers collect and analyze data to guide their next steps of instructional practice (Ikemoto & Marsh, 2007). These decisions are based, in part, on what we know about a specific audience of students being taught and their per- formance. According to O’Connor (2002), “Too often, educational tests, grades, and report cards are treated by teachers as autopsies when they should be viewed as physicals” (p. 112). Frequent and varied formative assessment provides important data that can be moved into useable knowledge to change practice. Standardized test score data and student work allow educators to differentiate instruction and target it toward students’ individual needs (Mandinach, Honey, & Light, 2006). However, according to (Hubbard, Datnow, & Pruyn, 2013, in press), data literacy among educators remains a persistent concern.
To determine if students have reached mastery of particular content or concepts, we have to set goals. Much as a hiker knows he has reached his goal when he conquers the summit, students need to know where their summit is, and teachers and school leaders need to know when their students have conquered it.
This chapter will focus on four basic questions:
1. What defines what students should know, understand, or be able to do?
2. How will students demonstrate their mastery?
3. How will teachers or leaders know what level of mastery students have reached?
4. How can school leaders use assessment to inform change?
Once teachers have determined what they are supposed to teach and have then determined what students will do to demonstrate that they understand the material, we must look at how school leaders interpret achievement scores to reflect a school’s progress. We must also observe the extent to which particular targets are being met and what school leaders can do to innovate.
Accountability is a central thread running through any change process. In today’s assessment- based educational climate, schools have to provide information about their performance across a range of issues. This chapter concludes with a look at an accountability cycle of monitoring, evaluation, and reflection. Monitoring is the process of collecting and presenting information
Introduction Chapter 11
in relation to specific objectives on a systematic basis; evaluation ana- lyzes the information so that value judgments are made; and reflection uses evaluation data to inform deci- sions for strategic planning.
V O I C E S F R O M T H E F I E L D
From classrooms to boardrooms, plans and planning processes abound in educational organiza- tions. Anita worked on a plan with the leadership team at Benavides High School to reduce their student dropout rate as a primary objective of their redesign efforts. At one meeting, leadership team members came to realize the importance of thinking beyond a plan for what might go wrong and consider the consequences of success. In some ways trying to imagine what happens if things “go right” is a challenge.
I attended a leadership team meeting at a high school working to decrease their dropout rate of over 50%. Fernando took the initiative in facilitating this meeting. Fernando was one of two school improvement facilitators (SIFs) for the school. SIFs were new positions created to help with imple- menting the various elements of high school redesign. The school’s redesign project was in the planning year, with implementation scheduled for the next school year. I felt privileged to be invited to observe and chronicle this school’s change efforts.
Previously, the leadership and the faculty had agreed on creating smaller learning communities within the school as a major part of their redesign. Today’s meeting was to decide the primary objectives for the coming year, such as what changes they wanted to occur as a result of the rede- sign effort.
Fernando asked each team member (DeLisa, Cynthia, Rolando, and Berta) to bring to the meeting the ideas that had been generated during their small-group sessions with faculty members. It didn’t take long to identify the priority—reducing the dropout rate. The leadership team was pleased about the consensus across the faculty. With that done, the team members began to talk about the “what if’s” to include in the implementation plan.
After a while, I interrupted with a few questions: “What happens if you succeed? I know there will be celebrations, pats on the back, high fives, right? But, you are already overcrowded here—half of the grounds have portable classrooms, and I have even heard some rumors about running double shifts next year. And what about the faculty you will need to teach more kids? Your recruitment last year only attracted one math teacher for the coming year. If you all succeed—and I surely hope you do—who is planning what will be needed to accommodate the consequences of your good work?“
Suddenly all five team members were looking at me, but no words were uttered. After what seemed a long time, DeLisa said, “We’ve never talked about anything like that.” The other four nodded their heads, but didn’t speak. I filled the silence with a suggestion: “Maybe you all should spend a few minutes talking about what may be needed to support the changes so that your suc- cesses do not turn into problems?” Rolando replied with enthusiasm, “Good idea! And, maybe we should invite the principals to help us with this part of our plan.”
As a follow-up to this meeting, Fernando first talked with the other SIF for the school. They decided that meeting with the principal would be a good place to begin. It was obvious that the kind of changes needed to address these issues of success went far beyond anything they could do on their own. They walked together to the main office of the school to make an appointment with Dr. Easton, the BHS principal.
Think About It
What would an educational system look like that truly customizes students’ learning based on assessment and doesn’t move them forward in grade levels until they actually master objectives?
Measurement, Assessment, and Evaluation Chapter 11
Pre-Test 1. “Students will be able to appreciate art” would be an example of a(n)
a. objective.
b. goal.
c. self-fulfilling prophecy.
d. assessment.
2. Which of the following is not a good method to provide formative feedback?
a. Help students address areas they need to improve on.
b. Provide feedback in small increments.
c. Give timely and routine feedback.
d. Compare a student’s performance with that of other students.
3. A 30-question true-or-false test on Chapters 6–8 of the textbook would be a(n)
a. nonobjective traditional assessment.
b. authentic alternative assessment.
c. nonauthentic alternative assessment.
d. objective traditional assessment.
4. Pablo got 82% correct on his physics test. This value is a
a. scaled score.
b. standardized score.
c. true score.
d. raw score.
Answers 1. b. goal. The answer can be found in Section 11.1.
2. d. Compare a student’s performance with that of other students. The answer can be found in Section 11.2.
3. d. objective traditional assessment. The answer can be found in Section 11.3.
4. d. raw score. The answer can be found in Section 11.4.
11.1 Measurement, Assessment, and Evaluation Understanding the difference between measurement, assessment, and evaluation is funda- mental knowledge for school leaders, because these terms are used for very different tasks. Measurement, in general, is a critical part of all our lives. Never before has testing been so prominent in schooling and in the lives of students and teachers as it has been in the past decade. Schools, superintendents, teachers, students, and parents have been dramatically affected. Simply stated, measurement is defined as assigning numbers to some character- istic. We measure things with a ruler that has been given a numerical scale called inches. We step on a scale and measure weight in pounds. In the same way, we measure achievement by
Measurement, Assessment, and Evaluation Chapter 11
giving a test and seeing how many items a student answers correctly. However, as we move away from more objective physical measurement, we move into the more abstract realm.
Measurement refers to the process of determining the attributes of some physical object, attitude, or preference. When we measure, we generally use some instrument (e.g., yard- stick, exam, performance criteria) to obtain quantitative data. We can then compare this raw data to an answer sheet, established norms, or standards. The accuracy of the measurement depends on the precision of the instrument being used to take the measurement and the skill of the person using the instrument or device. We measure how big a classroom is in terms of square feet using a ruler. We are not assessing anything; we are simply collecting information relative to some established rule or standard. Some of the basic measurements in education are raw scores, percentile ranks, derived scores, standard scores, or test scores.
Assessment is a process of documenting knowledge, skills, attitudes, or beliefs, usually in measurable terms. A test is a special form of assessment because it goes beyond simply measuring knowledge; rather, it is interpreting information about learning. Assessment also includes methods such as observations, interviews, or behavior monitoring. From an assess- ment, feedback can be given. We assess progress at the end of a school year through stan- dardized testing, which establishes the adequate yearly progress (AYP), and we assess verbal and quantitative skills through the SAT and GRE, which can be interpreted in terms of meeting a goal for college or graduate school entrance.
According to Rogers and Badham (2004), “Evaluation is the process of systematically col- lecting and analyzing information in order to form value judgments based on firm evidence” (p. 3). When we evaluate, we are determining the worthiness, appropriateness, goodness, validity, or legality of something for which a reliable measurement or assessment has been made. For example, we can repeatedly measure the temperature of a classroom using a ther- mometer to determine an average temperature, such as 75 degrees Fahrenheit. That action is simple measuring. We can then poll students to determine if this is the ideal temperature for a learning environment. It is the context of the temperature for a particular purpose (learning) that provides the criteria for evaluation (Huitt, Hummel, & Kaeck, 2001). Evaluation deter- mines whether the subject (i.e., students) meet a preset criteria, such as a passing score on an exit exam. Evaluations use assessment to make a determination of qualification in accordance with predetermined criteria.
Evaluation has two main purposes: (a) accountability to prove quality and (b) development to improve quality (Rogers & Badham, 2004). School-based evaluations have the highest rates of success when the evaluation does not take up too much time, effort, or resources. Constraints on schools to conduct evaluations include a shortage of time, lack of expertise in evaluation, and reluctance of staff to embrace evaluation as a part of normal practice. To address these concerns, Rogers and Badham (2004) suggested the following tasks:
• Limit the evaluation to a few specific targets by establishing priorities that are achiev- able in the short term and that are easily measured.
• Only collect essential information and measures that are necessary for the purposes of evaluation. Keep it short and simple.
• Make the maximum use of information already available before rushing in to collect more.
• Ensure that the process is sustainable by making it cost-effective in terms of time and resources.
Measurement, Assessment, and Evaluation Chapter 11
Setting Goals and Measurable Objectives
Before going further in depth with measurement, assessment, and evaluation, let’s first con- sider what defines what students should know, understand, or be able to do. Student mastery is often framed as a goal or an objective. What is the difference between these two terms? Table 11.1 provides an overview of the differences.
Table 11.1: Goals vs. objectives
Goals Objectives Broad statements Specific
Measurable Measurable
Abstract Concrete
Long-term Short-term
Focuses on what learner will experience Focuses on defining the learning outcome of the experience
With the publication of the landmark report A Nation at Risk in 1983, the national educa- tion standards movement was underway and continues today. The report suggested that American schools were failing, and it touched off a wave of local, state, and federal reform efforts. In the ensuing years, however, many objections have been raised about the common standards movement. Kober and Rentner (2012) reported that most states are a long way from turning the learning goals into real classroom practice. A key concern is how well aligned tests are with the standards they are designed to measure (validity). Another concern is the slowness with which change occurs in revising curriculum to reflect standards and changing classroom practices to teach the standards.
The reality is that standards are being translated into exams (assessment), and test results (measurement) are being used to hold schools accountable (evaluation), all of which makes these exams “high stakes.” Because of the significance attached to some of these standard- ized exams, Standards for Educational and Psychological Testing have been created by the American Psychological Association, the American Educational Research Association, and the National Council on Measurement in Education. These standards present a number of prin- ciples designed to promote fairness in testing and avoid unintended consequences. They include the following:
• The results of a single test should not be the only data used for making decisions about a student’s continued education, such as retention, tracking, or graduation.
• For high-stakes exams that determine a student’s eligibility for promotion to the next grade or for high school graduation, there should be multiple opportunities for learn- ers to demonstrate mastery of materials through equivalent testing procedures.
• School districts or states that mandate standard- ized exams also need to take responsibility for monitoring the impact on ethnic minority or low socioeconomic status students and to identify and minimize potential negative consequences of such testing.
• Special accommodations for students with lim- ited English proficiency or students with disabil- ity may be necessary to obtain valid test scores.
Think About It
Are high-stakes exams a valid measure of con- tent knowledge for English language learners? Why or why not?
Measurement, Assessment, and Evaluation Chapter 11
Determining Objectives
Many school districts and universities are currently providing training on how to develop measurable student-learning outcomes and outcome-based assessments (Chard, Cook, & Tankersley, 2012). Learning outcomes can be written as measurable objectives. An objective is a clear, specific, concise statement that outlines in precise terms a short-term outcome that is relevant to achieving the ultimate learning goal, which is fairly broad. Objectives play a fundamental role in designing instruction because they identify what students should master, know, or be able to do at the conclusion of an instructional activity. Objectives create a level of accountability for student performance and become the yardstick used to describe learning outcomes. Behaviors for educational objectives fall into three categories or domains, which were defined in 1956 when Benjamin Bloom headed a group of educational psychologists:
• Cognitive: mental skills or knowledge (80% of educational objectives are in this domain.)
• Affective: growth in feelings or emotional areas (attitude or self)
• Psychomotor: manual or physical skills (objectives for hands-on activities)
Bloom believed that after a learning segment, learners should have acquired new knowledge (cognitive domain), appropriate dispositions (affective domain), or new skills (psychomotor domain). Bloom’s taxonomy for all three of these categories is hierarchical (see Figure 11.1), starting from the simplest type of thinking or behavior and progressing to the most complex.
f11.01_EDU675.ai
Creating Create, invent,
compose, predict, plan, imagine,
construct, design
Bloom’s Taxonomy
Evaluating Judge, select, decide, justify,
debate, discuss, recommend, rate
Analyzing Analyze, explain, investigate, distinguish,
compare, separate
Applying Solve, show, use, illustrate, complete, classify,
compare, design
Understanding Explain, interpret, compare, discuss, predict, describe, give an example
Remembering State, name, list describe, label, relate, find
Figure 11.1: Revised cognitive domain of Bloom’s Taxonomy with measurable verbs
The measureable verbs in Bloom’s revised cognitive domain are useful for designing learning objectives.
Source: From Lorin Anderson et al., A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom's Taxonomy of Educational Objectives, Completed Edi- tion, 1st ed. Copyright © 2001. Printed and Electronically reproduced by permission of Pearson Education, Inc., Upper Saddle River, New Jersey.
Measurement, Assessment, and Evaluation Chapter 11
Mastery at the lower levels is assumed, because each level builds on itself. In the cognitive domain, the hierarchical levels were revised in 2001 to reorder the top cognitive levels and to clarify verb usage. The verbs are important because they are measurable and therefore pro- vide a source of measurable verbs for designing learning objectives.
In the lowest level, or remembering, of Bloom’s taxonomy, thinking is confined to recall, or locating knowledge in long-term memory that is consistent with presented material. The cog- nitive processes used to do this are recognizing or identifying factual knowledge. Moving up a level in Bloom’s taxonomy, the cognitive processes advance to a demonstration of under- standing based on interpreting, classifying, summarizing, inferring, comparing and contrast- ing, or explaining. The action verbs associated with these cognitive processes include clarify, paraphrase, illustrate, map, predict, and generalize. These types of verbs require that the learner reveal conceptual knowledge, which goes beyond providing the factual knowledge required at the lowest level. At the next level, application, learners have the ability to use a new concept to solve or address a problem.
When applying Bloom’s taxonomy in designing assessment items or writing measurable objectives, we need to keep several points in mind:
• The categories are not meant to be mutually exclusive.
• To build a learner’s critical thinking skills, we have to move beyond the lowest cognitive levels of remembering and understanding.
Tips for Designing Learning Objectives Well-written objectives are SMART: specific, measurable, attainable, realistic, and timebound (Association of College and Research Libraries, 2013). Objectives are not the activities con- ducted in class or the assignment given after class. Rather, objectives represent important learning outcomes that are measurable via assessment. A good, measurable, SMART objective has four components:
• Content: The content tells what the student will learn by the conclusion of the lesson. The content stems from the standard. Begin with nouns, or the things you want stu- dents to learn (e.g., steps of the research process).
• Behavior: The behavior tells what students will do to show that they have learned. This component describes an observable, measurable action. Select a verb from the level of knowledge that is observable or measurable (e.g., describe the steps of the research process).
• Criterion: The criterion is the level of acceptable performance, the standard of mastery, or the proficiency level expected. How well will students have to perform before you can say they have met the objective? (e.g., speed, accuracy, quality, quantity; accu- rately describe seven steps of the research method).
• Condition (optional): Describes the circumstances, situation, or setting under which the student will perform the behavior and be evaluated (as opposed to learning condition). This is the method or activity in which students will demonstrate their understanding (e.g., describe the seven steps of the research process by writing an essay). This com- ponent of the objective can be put in a “by” clause at the end of the objective.
Here is an example of a good, SMART learning objective:
Think About It
Generate a measurable objective using the provided verbs: list, justify, analyze, explain, predict. How does the verb used in the objec- tive change the level of learning outcome?
Measurement, Assessment, and Evaluation Chapter 11
OBJECTIVE: Explain five phases (evaporation, condensation, transpiration, percolation, pre- cipitation) of the water cycle in your own words by creating a digital storytelling document.
Common Pitfalls When Writing Objectives Following are some of the common mistakes educators make when writing objectives:
• Not using a measureable action verb. Too often, educators will use verbs such as “understand,” “gain knowledge of,” or “learn” as the verb of the educational objec- tive. (For instance, “The learner will understand the water cycle.”) Understand is just too vague, and it is very difficult to measure or discriminate between levels of understand- ing. Instead, use a verb from Bloom’s taxonomy or the six facets to anchor the objective.
• Not listing the degree required for mastery. How will you know when the objective has been met if a criterion is not set for mastery? What is “good enough”? What level does the learner have to reach to succeed in attaining an objective?
Bloom’s Taxonomy for Affective Learning and Teaching
Teachers also need to know who must be encouraged to speak in class and who does not need encouragement; who is interested in science but not in social studies; and so on. To examine these behaviors, Bloom and his colleagues included a taxonomy for an affective domain. In these categories, Bloom grouped how we deal with things emotionally, again from the simplest behavior to the most complex. The affective domain includes feeling, values, appreciation, enthusiasm, motivation, and attitudes. Although the indicators in this domain are rarely formally assessed in classrooms, educators constantly assess affective behaviors informally through interactions with students. Teachers use this domain to examine how a student approaches learning: With confidence? With a can-do attitude? Table 11.2 outlines Bloom’s affective domain and provides key words for each category.
Table 11.2: Affective domain of Bloom’s taxonomy
Category Example and Key Words (verbs) Receiving phenomena: Awareness, will- ingness to hear, selected attention.
Examples: Listen to others with respect. Listen for and remember the names of newly introduced people.
Key words: asks, chooses, describes, follows, gives, holds, identi- fies, locates, names, points to, selects, sits, erects, replies, uses
Responding to phenomena: Active participation on the part of the learners. Learners attend and react to a particular phenomenon. Learning outcomes may emphasize compliance in responding, willingness to respond, or satisfaction in responding (motivation).
Examples: Participates in class discussions. Gives a presentation. Questions new ideals, concepts, or models in order to fully under- stand them. Knows the safety rules and practices them.
Key words: answers, assists, aids, complies, conforms, discusses, greets, helps, labels, performs, practices, presents, reads, recites, reports, selects, tells, writes
Valuing: The worth or value a person attaches to a particular object, phenom- enon, or behavior. This ranges from simple acceptance to the more complex state of commitment. Valuing is based on the internalization of a set of specified values, while clues to these values are expressed in the learner’s overt behavior and are often identifiable.
Examples: Demonstrates belief in the democratic process. Is sensitive toward individual and cultural differences (value diversity). Shows the ability to solve problems. Proposes a plan to social improvement and follows through with commitment. Informs management on matters that one feels strongly about.
Key words: completes, demonstrates, differentiates, explains, follows, forms, initiates, invites, joins, justifies, proposes, reads, reports, selects, shares, studies, works
(continued)
Measurement, Assessment, and Evaluation Chapter 11
Category Example and Key Words (verbs) Organization: Organizes values into priorities by contrasting different values, resolving conflicts between them, and creating a unique value system. The emphasis is on comparing, relating, and synthesizing values.
Examples: Recognizes the need for balance between freedom and responsible behavior. Accepts responsibility for one’s behavior. Explains the role of systematic planning in solving problems. Accepts professional ethical standards. Creates a life plan in harmony with abilities, interests, and beliefs. Prioritizes time effectively to meet the needs of the organization, family, and self.
Key words: adheres, alters, arranges, combines, compares, completes, defends, explains, formulates, generalizes, identifies, inte- grates, modifies, orders, organizes, prepares, relates, synthesizes
Internalizing values (characteriza- tion): Has a value system that controls behavior. The behavior is pervasive, consis- tent, predictable, and, most important, characteristic of the learner. Instructional objectives are concerned with the student’s general patterns of adjustment (personal, social, emotional).
Examples: Shows self-reliance when working independently. Cooperates in group activities (displays teamwork). Uses an objective approach in problem solving. Displays a professional commitment to ethical practice on a daily basis. Revises judgments and changes behavior in light of new evidence. Values people for what they are, not how they look.
Key words: acts, discriminates, displays, influences, listens, modifies, performs, practices, proposes, qualifies, questions, revises, serves, solves, verifies
Source: Clark, D. (2013). Bloom’s taxonomy of learning domains. Retrieved from http://www.nwlink.com/~donclark/hrd/bloom.html. Reprinted with permission.
Think About It
Online courses can include affective components by providing students with a place to post questions, get feedback, complete self-assessment checklists, and hear encouraging messages from peers and the instruc- tor. How can encouraging students to set reasonable goals for them- selves enhance affective learning?
Bloom’s Taxonomy for Psychomotor Learning
This domain focuses on developing skills to a specified level of accuracy, speed, smoothness, or force. This development is particularly applicable to science lab courses, vocational train- ing, physical education, music, or performing-arts classes. The stages (from simplest to most complex) of the psychomotor domain have been described as follows:
• Action (elementary movement)
• Coordination (synchronized movement)
• Formation (bodily movement)
• Production (combined verbal and nonverbal movement)
There is a cognitive component in the psychomotor domain. Students who are new to a con- tent area will benefit from a hands-on approach. As students become more familiar with the content, videos and images can be used to teach the skill.
Six Facets of Understanding
Wiggins and McTighe (2005) proposed an alternative way to measure the level of a student’s learning. Their book, Understanding by Design (UbD), suggests a backward curriculum
Measurement, Assessment, and Evaluation Chapter 11
design methodology that focuses on “teaching for understanding” and suggests that edu- cators concentrate on learning outcomes when designing the curriculum. The learning out- comes are based on the big ideas of the content area, and a learner’s “understanding” of the concept is deconstructed into six facets: explain, apply, interpret, perspective, empathy, and self-knowledge. Unlike Bloom’s taxonomy, the facets are not hierarchical. Rather, learn- ers gain understanding through their own personal journey, which is certain to be unique for each individual, and the teacher facilitates the development of these six facets according to the teaching context. Understanding of a topic is based on the cultivation of six equal aspects of learning that emerge from the learning context, as opposed to being dictated to the learner at the beginning of a unit of study.
The following points describe these six facets of understanding in more detail:
• Explanation goes beyond the simple recitation of facts and figures and provides evi- dence and reasons to support learning insights. Asking learners to show their work, or respond to “because” prompts, pushes them to provide a rationale for their thinking. If students understand a topic, then they can explain concepts, principles, and pro- cesses using their own words; teach the topic to others; justify their answers; or show their reasoning.
• Interpretation concerns a learner’s meaning-making ability. If learners “understand” a topic, then they are able to make sense of data, text, and experiences through images, analogies, stories, and models. Learners demonstrate their understanding by providing apt translations; revealing a historical or personal dimension to the topic; or relating to ideas and events through images, anecdotes, analogies, models, and metaphors.
• Application means that learners can effectively use and adapt what they know in new and complex contexts; it is similar to what Bloom intended with his description.
• Perspective requires a level of maturity to not necessarily believe in another’s view- point, but to be able to speak from a different social, racial, or religious context. Students demonstrate their understanding of perspective by being able to see the big picture and recognize different points of view.
• Empathy is more subjective than perspective when it comes to seeing something from another’s viewpoint. It challenges the learner to look past his or her assumptions about a given topic and use creative reasoning to “walk in someone else’s shoes.” To move from the view of “they” to the view of “I” is a powerful shift in understanding. The empathy facet addresses the social concern that students are somehow unable to look beyond their own sense of self.
• Self-knowledge is the manner by which students show meta-cognitive awareness, use productive habit of mind, and reflect on the meaning of their learning and experi- ence. It is the practice of questioning their long-standing beliefs in an objective way. It demands that students uncover what understanding looks like beyond themselves (Wiggins & McTighe, 2005).
In their entirety, these six facets construct a holistic picture of understanding. The facets serve as indicators of how understanding is revealed and provide guidance about the kinds of assessments needed to determine the extent of student understanding. However, Wiggins and McTighe (2005) cautioned:
• All six facets of understanding need not be used all of the time in assessment.
• Performance tasks based on one or more facets are not intended for use in daily lessons. Rather, these tasks should be seen as culminating performances for a unit
Measurement, Assessment, and Evaluation Chapter 11
of study. Just as practices in athletics prepare a team for an upcoming game, daily lessons develop the affiliated knowledge and skills needed for the understanding performances.
By deconstructing “understanding” into facets, learners can be challenged to accomplish more intellectual work and the critical thinking stressed as a 21st century skill. Newmann, Bryk, and Nagaoka (2001) examined the relationship of the nature of classroom assignments to standardized test performance and concluded that students who received assignments requiring more challenging intellectual work also achieved greater-than-average gains on the Iowa Tests of Basic Skills in reading and mathematics and demonstrated higher performance in reading.
Think About It
What are the similarities and differences between Bloom’s taxonomy and the six facets of understanding? How do these categories help dis- criminate depths of student understanding?
Performance Indicators
Performance indicators (PIs) are the signals of success that would show that an individual has achieved the objective. This indicator is not a checklist or an opinion; rather it is a qualita- tive or quantitative measure that indicates the extent of progress made in a certain area. For example, a performance indicator describing what a leader would do in relation to improving students’ learning might read as follows:
A leader uses multiple sources of information and analyzes data about current practices and outcomes to shape a vision, mission, and goals with high, measurable expectations for all students and educators.
From this PI, the following specific measurable objectives for leaders can be identified:
• Identify and use multiple sources of information.
• Analyze data about current practices.
• Write a mission statement with high, measurable expectations. (Sanders & Kearney, 2008)
PIs are collected consistently over time using various evaluation instruments as relevant to the stated objectives. For example, one of a vocational school’s objectives was to strengthen and increase the number of connections between the world of work and the school’s curriculum. A PI for this objective might be the number of department teams using links with the world of work in their teaching. To provide evidence for the PI, a survey is designed as the evalua- tion instrument. In the survey, department chairs are asked to indicate how many students visited or had projects involving local businesses or other workplaces, how many teach- ers had guest speakers from businesses, whether teachers used students’ work-placement experiences in their curriculum, and so on. These performance indicators can be expressed as rates, ratios, or percentages. One performance indicator used extensively to serve as an actionable measure of how well schools are educating their students is the adequate yearly progress gauge.
Measurement, Assessment, and Evaluation Chapter 11
Adequate Yearly Progress
The adequate yearly progress (AYP) measure was developed under Title I of the No Child Left Behind (NCLB) Act, the current version of the Elementary and Secondary Education Act, to determine if schools are successfully educating their students. The law requires states to use this single accountability system for public schools to determine whether all students, as well as individual subgroups of students, are making progress toward meeting state academic con- tent standards. The goal is to have all students reaching proficient levels in reading and math by 2014. Progress on those standards is tested yearly in grades 3 through 8 and then in one grade in high school. The results are then compared to the previous years to determine if the school has made adequate progress toward the proficiency goal (Department of Education, 2001).
The Department of Education (2001) stated:
According to the law, states have the flexibility to define this yearly progress, but it must include the following elements:
• State tests must be the primary factor in the state’s measure of AYP, but the use of at least one other academic indicator of school performance is required, and additional indicators are permitted;
• For secondary schools, the other academic indicator must be the high school graduation rate;
• States must set a baseline for measuring students’ performance toward the goal of 100 percent proficiency by spring 2014. The baseline is based on data from the 2001–02 school year;
• States must also create benchmarks for how students will progress each year to meet the goal of 100 percent proficiency by spring 2014;
• A state’s AYP must include separate measures for both reading/language arts and math. In addition, the measures must apply not only to students on average, but also to students in subgroups, including economically disadvantaged students, students with disabilities, English-language learners, African-American students, Asian-American stu- dents, Caucasian students, Hispanic students, and Native American students.
• To make AYP, at least 95 percent of students in each of the subgroups, as well as 95 percent of students in a school as a whole, must take the state tests, and each subgroup of students must meet or exceed the measurable annual objectives set by the state for each year
If a school or district fails to make AYP for two consecutive years, it is then identified for school improvement. A school so identified must notify parents, provide students with an option of transfer to another public school, or offer supplemental services (tutoring). Additional sanc- tions (including ordering restructuring of the school) are added if the school continuously fails to make AYP. As AYP requirements have increased, the number of schools failing to make AYP has also increased. According to McNeil (2011), in 2007, 28% of schools failed to make AYP. By 2010, that number had risen to 38%, and by 2013, the failure rate hit nearly 50%. “Because the policies underlying which schools fail to make AYP vary from state to state, it is not the reliable yardstick that advocates had hoped it would be when the NCLB law was passed in 2001” (McNeil, 2011).
The extent to which these flaws have undermined AYP as an effective performance indicator for measuring academic growth has led many states to ask for and receive waivers from AYP
The Purpose of Assessment Chapter 11
reporting. A state must agree to changes in three main areas in exchange for elimination of the AYP designation: (a) college- and career-ready standards and assessments that measure stu- dent achievement and growth; (b) a differentiated accountability system that both recognizes high-achieving, high progress schools (reward schools) and supports chronically low-achieving
schools (priority and focus schools); and (c) teacher and principal evaluation and support systems to improve instruction. A team of peer reviewers, along with Department of Education staff, study waiver proposals and offer suggestions to help states win approval. Since February 2012, 34 states plus Washington, D.C., have been granted waiv- ers. These waivers are in effect until the end of the 2013–14 school year, when states can renew the waiver for another 2 years. States without waivers are still under the mandates of NCLB.
11.2 The Purpose of Assessment For the most part, assessment is designed to help teachers and students build a shared under- standing of the academic achievement the student has accomplished so that feedback for further development is possible. Assessments are designed differently to ensure their fitness for different purposes. One purpose is “institutional monitoring,” which comprises a huge number of uses for assessment data. Some examples are school-by-school performance com- parisons, judgments on whether schools have reached their student-achievement targets, performance pay, and assessments of teachers’ qualifications for promotion. Assessment information has become a proxy measure that is supposed to facilitate evaluation on the qual- ity of most elements in our education system: students, teachers, schools, school districts, and even the federal government itself. For the purposes of further discussion, we will divide the purpose of assessment into two broad categories: formative and summative. These dis- tinctions are not labels for different types or forms of assessment; instead, they describe how assessments are used.
Formative Assessments
Formative assessments are changing the entire assessment paradigm in 21st century schools. Formative assessment is the use of day-to-day, often informal assessments to explore students’ understanding so that instruction can be tailored to develop that understanding. According to Moss and Brookhart (2009), “Formative assessment is an active and intentional learning process that partners the teacher and the students to continuously and systematically gather evidence of learning with the express goal of improving student achievement” (p. 6). This definition is congruent with the view that “formative classroom assessment occurs only when evidence is used to make a needed change” (Schneider & Randel, 2010, p. 251). Teachers who gather and examine various types of evidence of a student’s learning and use that informa- tion to either adapt instruction or provide feedback to students are using formative classroom assessment (Brookhart, Moss, & Long, 2008). An essential feature of formative assessment is a purposeful link between the results of an assessment and succeeding actions that lead to gains in student learning (Nichols, Meyers, & Burling, 2009; Wiliam, 2010). The advantages of formative assessment are greatly reduced if teachers are not able to use evidence of student learning to establish succeeding instructional steps and help students move their own learning forward (Heritage, Kim, Vendlinski, & Herman, 2009).
Think About It
Additional performance indicators would help guide decision making about school improve- ment. What additional performance indicators do you think would help schools more accu- rately measure the challenges the school faces, direct school improvement efforts, and recog- nize progress toward established goals?
The Purpose of Assessment Chapter 11
For the most part, formative assessments are ongoing, varied, and dynamic. Feedback may be motivational, informative, or corrective; teachers may provide it immediately or delay it, present it in writing or express it verbally. Formative assessment can also have a diagnos- tic purpose. When given before instruction begins, it serves as a pretest, or benchmark, to determine prior knowledge or mastery of prerequisite knowledge and skills (e.g., a reading- readiness test). When curriculum is designed with the end in mind, formative assessment is a central part of pedagogy. Table 11.3 shows 10 things to remember about using formative feedback (Buczynski, 2009).
Table 11.3: Key points regarding formative feedback
Characteristic Explanation
It probes the status of learners’ knowledge.
By using formative feedback, you can inquire into the student’s current knowledge and experience. Use feedback to explore the what, how, and why of the student’s thinking.
It challenges the critical-thinking process.
Target feedback to encourage learners to conduct the error analysis them- selves. Ask them questions like, “Did you make any assumptions when
?” and “Would you consider your to be strong or weak?”
It aligns with self-assessment. Ask students to assess their own performance and then help them diagnose any areas that need improvement. Give sufficient information so that students can address those areas. Be sure to focus your feedback on the task, not the learner.
It provides nonevaluative input. Make your comments constructive to guide learners to the next level of understanding. For example, you could say, “You might consider measuring
. What would happen if ?” Avoid comparisons with other students—directly or indirectly—or providing an overall grade at this stage of assessment.
Use formative feedback to make observations.
Consider how a student’s thinking has changed over time, and point this out to the student to reinforce his or her thought process. Feedback like this supports students’ autonomy when they are completing assignments.
Make your comments simple and specific.
Rather than giving many suggestions for improvement, work with the learner to set a single goal for the next assignment. Give only enough information to initiate new thinking, remove uncertainties, or clarify objectives. Provide feedback in small chunks so that it is not overwhelming or ignored.
Offer formative feedback frequently.
Build in multiple checkpoints for feedback. Provide timely responses to students’ work. Use immediate feedback to help students retain procedural or conceptual knowledge. To promote transfer of learning, consider using delayed feedback.
Differentiate formative feedback. Tailor your comments based on the learner’s characteristics. There is no “one size fits all” formative feedback. Provide early, structured, and corrective support for low-achieving students by offering explicit guidance or directive feedback. For high-achieving students, provide verification and facilitative feedback in the form of accuracy checks, hints, cues, and prompts. Consider the nature of the task and instructional goals, and customize your comments to ensure that your feedback is valid, objective, focused, and clear to the learner.
Use multimedia formats to provide formative feedback.
Explore the potential of multiple modes for feedback, including written, verbal, graphic, video, and electronic means.
Do not allow formative feedback to interrupt active learning.
Minimize the use of formative feedback during implementation of an assignment. Comments at this time may take control or overly influence the direction of the task. They may also distract the learner from focusing on the task at hand.
The Purpose of Assessment Chapter 11
As a classroom practice, formative assessment can have a significant positive impact on stu- dent learning. Falk (2012) reported on a review by Black and Wiliam (1998a) of 250 articles and chapters that examined the effects of formative assessment on student achievement. Black and Wiliam provided substantial evidence of this impact. They found that improved formative assessment produced effect sizes between 0.4 and 0.7, “amongst the largest ever reported for educational interventions” (Black & Wiliam, 1998a, p. 61). Other studies have also provided strong evidence that formative assessment leads to increased learning (Bell & Cowie, 2001; Black & Wiliam, 1998a). According to Ainsworth & Viegut (2006), teachers who use formative assessments are more likely to (a) determine what standards students already know and to what degree; (b) decide what minor modifications or major changes in instruc- tion they need to make so that all students can succeed in upcoming instruction and on subsequent assessments; (c) create appropriate lessons and activities for groups of learners or individual students; and (d) inform students about their current progress in order to help them set goals for improvement (p. 23).
However positive the research support is for formative assessment, other studies point to the difficulties teachers encounter with both interpreting evidence of student learning from formative assessment processes and providing students with feedback that moves learning forward. A study by Ruiz-Primo and Li (cited in Schneider & Andrade, 2013) noted that only 14% of the feedback comments that teachers provided students in a particular science unit met the authors’ criteria for having the potential to move student learning forward. Similarly, Schneider and Gowan (cited in Schneider & Andrade, 2013) stated that teachers found it dif- ficult to provide feedback to move students forward:
Schneider and Gowan conjectured that the low quality of feedback found in their study was attributable to the teachers’ general inability to analyze the cause of student confu- sion. Teachers tended to take a one-size-fits-all approach when determining what instruc- tional adaptations to make or feedback to provide to a student based on the student’s response. (Schneider & Andrade, 2013, p. 161)
Finally, Hoover and Abrams (cited in Schneider & Andrade, 2013) found that 64% of teach- ers reported that the pacing of instruction prohibited reteaching of concepts. These studies suggest that teachers need support in (a) designing instructional sequences based on state standards, (b) effectively engaging in formative assessment practices, and (c) moving beyond a one-size-fits-all approach to reteaching concepts (Schneider & Andrade, 2013, p. 160). The latter two items also require “that teachers develop skills related to interpreting evidence of student learning, and in targeting feedback and instructional adaptations directly to a stu-
dent’s present level of performance” (p. 161).
Summative Assessment
Summative assessment comes at the end of a learning segment and is the more formal “summing up” of a student’s cumulative progress that can be used for purposes ranging from providing a letter
grade for the course to providing information to parents. Reliability and validity of summative assessment are critical if the data from the assessments will be used for evaluation. Reliability is about the extent to which an assessment can be trusted to give consistent information on a student’s progress; validity is about whether the assessment actually measures what it was meant to measure. Many have heard the admonition to “measure twice, cut once” as applied
Think About It
What are some solid reasons for using forma- tive assessments in the classroom?
The Purpose of Assessment Chapter 11
to sewing or cutting wood. But this axiom reflects the two major concepts of measurement: reliability and validity.
Reliability and validity refer to two characteristics of measurement. When anything is mea- sured, the results must be repeatable and consistent (reliable) and must truly measure what was intended to be measured (validity). The more concrete and well defined the characteristic that is being measured, the more likely it will be reliable and valid. In addition, a measure must be reliable; otherwise, it cannot be valid (you cannot say you have measured something unless you can consistently measure something). Conversely, reliability is necessary but does not ensure validity.
Think of a clock that is always 10 minutes fast. If it is always 10 minutes fast, you can say the measure (time) is reliable. However, it is not valid (unless you consistently subtract 10 minutes). Or consider target shooting. You look down the sight and shoot 10 rounds at the center. If all the hits are up and to the right, your rifle may be reliable (consistently up and to the right) but not valid (aiming at, but missing, the center). You then might adjust to the sight.
When attempting to measure achievement, you write items directly aimed at the content to be measured. For example, assume you are trying to measure the addition of two-digit numbers without carrying (e.g., 10 + 12). You would write a minimum of 7–10 items that represent the concept. After giv- ing the test to a sample of respondents, there are ways to check the reliability and validity of your test.
A teacher-leader examining a summative assessment given at a school site may encounter these various forms of reliability mentioned in the test-administra- tor booklet. To understand the reliability discussion in the booklet, the teacher-leader will encounter dif- ferent ways to assess reliability. Test-retest reliability refers to giving a test to a group and then giving the test to the same group again. If the test is reliable (consistent), then the high scorers in the group should score equally high on both—the same being true for the lower scorers.
Alternative forms of reliability refer to giving two forms of the same content (two tests) to the same group of students. Again, if the test is reliable (consistent), then the high scorers in the group should score equally high on both forms of the test—the same being true for the lower scorers. Split-half reliability consists of essentially two forms of the same test with one form being the odd items and one form being the even items. If the original test is reliable (consistent), then the high scorers in the group on the odd items should score equally high on the even items of the test—the same being true for the lower scorers. Yet another form of reliability is referred to as coefficient alpha, which measures reliability on an item-by-item basis. All reliability measures yield a score from 0 to 1, with 0.6 or higher seen as a minimum value to be considered reasonably reliable (Nunnaly, 1967).
Once a test is found to be reliable, validity is assessed. For validity, content or face validity is the assessment used most often. A group of content experts writes items on different top- ics; then, another group of experts examines those items to affirm that they indeed measure the topic as defined. In classrooms, an individual teacher typically develops a test based on content judgment. Classroom tests are usually then combined to reflect a grade in a particular subject. District-level test development usually consists of a group of teachers in a content
Think About It
For many graduate students, the graduate records exam (GRE) is given as an admission requirement to many graduate schools. The GRE is an incredibly reliable test. However, its validity as a predictor of success in graduate school is open to question. Why is this so?
Methods of Assessment Chapter 11
area. On the larger scale of state and national tests, a panel of experts across a state or the country (teachers, professors, etc.) is usually convened.
How tests are developed and shown to be reliable and valid are critical for the teacher- leader. Ensuring that you are measuring something and that something is what you intend to measure provides the leader with a sense of confidence that correct judgments are being made with respect to students, instruction, and curriculum. If a teacher-leader is charged with selecting an existing test, then a basic understanding of measurement is a requirement so that the test is both reliable and valid for particular students. In addition, it is critical that the content the test assesses matches the curriculum being taught. Too often, a mismatch results in low test scores.
11.3 Methods of Assessment Now that we have explored what defines what students should know, understand, or be able to do, let’s turn our attention to how students will demonstrate their mastery: assess- ment methods. Assessment instruments can be divided into what are considered “traditional” assessments—in which there are two main categories of questions: objective and nonobjective —and “alternative” assessment—which is more constructive in nature and might include evi- dence of student mastery via portfolios, project work, performances, or interviews. Usually, these alternative assessments are graded with a rubric. Within the alternative category of assessment is the “authentic” assessment, which mirrors “real world” expectations or tasks (e.g., lab practical, oral history project). Bergen (1993) identified three qualities of authen- tic assessment: (a) it measures many facets simultaneously, (b) it is applied in a way that is realistic and that reflects the complex roles of the real world, and (c) it is often group-based and requires individual contribution for success. Let’s take a closer look at each assessment methodology.
Traditional Assessment
Traditional assessments are usually paper-and-pencil exams in which the student selects or composes a response to a prompt. Considerable skill is required to develop multiple-choice test items that go beyond measuring factual knowledge to measuring higher-order thinking and problem-solving skills. Students usually produce responses on demand at a specified time and within a certain amount of time. These constraints contribute to standardization of testing conditions, contribute to the efficiency of test taking, and increase the comparabil- ity of the results. However, these constraints can also create “test anxiety” and promote a certain degree of answer guessing. In a traditional assessment, the multiple-choice portion of the exam can be formatted to be scored by a Scantron machine (optical mark sensors). With these machines, student responses can be scored and reported extremely quickly and inexpensively.
Objective assessment questions are items that are generally not open to interpretation; there is a right answer and a wrong answer. Nonobjective questions are more open-ended and allow greater room for constructed responses. An example is an open-ended question that requires a short, written answer that consists of a word or phrase or a longer response, such as an explanation of how to apply particular knowledge or skills to a particular context. Constructed response questions can be used to measure more complex reasoning, logical thinking, interpretation, or analysis.
Methods of Assessment Chapter 11
Each type of questions is appropriate for different purposes, as shown in Table 11.4.
Table 11.4: Assessment question types and their uses
Question Type Category Purpose Example Multiple choice Objective Discriminate between
options, make simple judgments, apply vocabulary
What is the role of the leader in the “forming” stage of group development?
a. Organizer
b. Negotiator
c. Participant
d. Cheerleader
Matching, sequencing
Objective Identifying relationships, classifying items, charting cause and effect
What is the sequence of stages in Tuckman’s group development theory?
a. Forming norming performing storming
b. Forming performing storming norming
c. Forming norming storming performing
d. Forming storming norming performing
True/false or yes/no Objective Knowledge of generaliza- tions, relationships, and examples; predicting and evaluating
Conflict can be good for developing a team’s problem-solving skills.
a. Misconception
b. Truth
Factual short answer, fill in the blanks, computational
Objective Recall; classifying facts, terms, or concepts; solving simple problems
Draw a diagram explaining the action research cycle.
Higher-order short answer
Nonobjective Summarizing, applying, concluding, evaluating, predicting, analyzing
What would John Dewey say about today’s educational system if he were alive today?
Short or long essay Nonobjective Organizing ideas, devel- oping a logical argument, comparing concepts, evaluating a position, communicating thoughts
Are high-stakes exams a valid measure of content knowledge for English language learners? Why or why not?
Source: Adapted from Steven Farr and Teach for America, Teaching As Leadership: The Highly Effective Teacher's Guide to Closing the Achievement Gap. Reprinted by permission of John Wiley & Sons, Inc.
With the new smarter, balanced assessment for the Common Core State Standards, stu- dents take the entire test on the computer. Computerized assessment tools are being used to simulate interactive, real-world problems or to provide immediate feedback so that learn- ers participate in the construction of the exam itself. The standardized questions on these high-stakes exams are written to assess specific points from the standards. Supovitz’s 2010 review of research on high-stakes testing provides four effects for improving the education system:
1. High-stakes testing does motivate educators, but responses are often superficial, focusing on literacy and numeracy skills.
2. Test-based accountability fosters alignment of curriculum, standards, and assessments.
3. High-stakes testing regimes are useful to policymakers for assessing school- and system- level performance but insufficient for individual-level accountability.
Methods of Assessment Chapter 11
4. Test-based accountability is an appealing political strategy to demonstrate to the public that tax dollars are being spent judiciously
Alternative Assessment
Many alternatives to traditional assessments offer a variety of ways to measure student under- standing. The teacher usually designs these alternative assessments, which are described as authentic, comprehensive, holistic, or performance assessments. Examples of these types of measures include oral presentations, model building, projects, experiments, and portfolios of student work. Alternative assessments are designed to align with the content of the instruc- tion and to be evaluated based on progress or mastery. As long as these assessments maintain their status as nonmainstream or “alternative,” their use for large-scale and high-stakes evalu- ation remains minimal. Alternative assessments have a number of advantages and disadvan- tages, as shown in Table 11.5.
Table 11.5: Advantages and disadvantages of alternative assessment
Advantages of Alternative Assessment Disadvantages of Alternative Assessment Provides a more realistic setting for student performance.
Can be costly in terms of time, effort, equipment, materials, facilities, or funds.
Focuses on quality of work. Rating process is more subjective.
Can be easily aligned with learning outcomes.
Humans must score the assessment, not machines.
Although one of the most significant advantages of alternative assessment is its flexibility and allowance for diversity, alternative assessment will not reach its full potential in edu- cational settings unless criteria are introduced for how the assessment will meet learning objectives.
We also need to consider the issue of accommodations to testing. The Individuals with Disabilities Education Act (IDEA) requires that an individualized education program (IEP) be developed for each learner with a disability so that learning differences can be accommo- dated. Leaders need to be keenly aware of state- and district-approved policies for approv- ing and making accommodations. Sometimes, for individuals with identified special needs or those with second-language issues, accommodation involves very different practices. For example, some learners may need extended time for completion of tests or a reduction of paper-and-pencil tasks. Learners with disabilities can be accommodated through presentation of assessment material in small steps, having subject matter read to them or paraphrased. When testing, teachers might consider fewer repetitive test items or a test format that allows more space for answering questions. Also, leaders may consider providing test directions that are read aloud and explained thoroughly or using oral, short-answer, or modified tests for those with different learning styles. Setting accommodations, such as testing in an alternative location, individual testing, small-group testing, preferential seating, or using a bilingual test administrator, are appropriate to improve accessibility. For example, a student with a diag- nosed attention deficit disorder may be allowed to take a test alone in a private library test carrel to assist in focusing concentration. Making these types of modifications in assessment administration yields a better insight regarding a student’s true ability and level of learning outcomes. Another way to gain a truer picture of a student’s learning is to look for progress over time, as is accomplished in portfolio-based assessment.
Interpreting Results Chapter 11
A portfolio, hard copy or electronic, is a collection of student work—some teacher selected and some student selected—that the learner assembles to demonstrate academic growth over time. A portfolio is not a scrapbook of assignments, and it is not a showcase of only the best work. If students simply collected their test results, journal entries, homework, essays, or products of student activities without reflecting on their learning progression, then the collection of work would simply be a keepsake. In a portfolio, however, learners provide a collection of artifacts in a systematic, organized manner to provide evidence that highlights their learning journey. The portfolio includes student reflections on their learning progression and rationales for why they chose particular pieces. This method of assessment is particularly apt for differentiated classes because it is individualized and results in a longitudinal, or “big picture,” look at a student’s mastery of a topic. The portfolio is not a single snapshot of learn- ing, as you might get from an exam or essay; rather, it is a mirror for students to see their own development and take charge of their learning (Wormeli, 2006).
A critical element in portfolio assessment involves record keeping. According to Thomas et al. (2005), “Goals and objectives for the portfolios must be set a priori and the contents negoti- ated by individual teachers and students rather than set by committee, board, or administra- tors” (p. 3). Periodic reviews of the portfolios should be conducted to collect data for the evaluation by conferencing notes; anecdotal records; field-note nar- ratives; rating scales ranking each item according to ability, frequency, extent, and so on; and checklists to condense student performance across time. Teachers are especially concerned with the amount of time this approach involves as an alternative to traditional, objective testing. Proponents of portfolio assess- ment contend that this type of assessment provides a means for those students at risk for academic failure to demonstrate progress within a format that is less restrictive and inflexible. This assessment method also allows students to demonstrate specific skills within the context in which the skills were taught, rather than in an artificial context determined by test developers.
11.4 Interpreting Results Now we turn to our final question for this chapter: How will teachers know what level of mastery students have reached? Using test data to make decisions about students, instruction, and curriculum is a critical responsibility of a teacher-leader. As noted previously, teachers who have been recognized by their profession have expressed concern about the effects of testing in their school and district (Dozier, 2007). One critical aspect for which they need more infor- mation is how to understand test scores and their interpretation for use in decision making.
Scaled Scores
Comprehension of scaled scores in standardized testing is critical for the teacher-leader. Everyone has had numerous experiences with tests. From the beginnings of life, babies are tested for developmental aspects at birth and then at various ages as they grow. They are scored to see how they compare to known behaviors of the average baby at different ages.
Think About It
• How can portfolio assessment promote more positive attitudes toward learning among students?
• Can a portfolio accurately measure a stu- dent’s academic achievement? Why or why not?
• How does a portfolio evaluate a student’s performance in terms of meeting instruc- tional objectives?
• Would you consider a portfolio an evidence- based assessment?
Interpreting Results Chapter 11
Some would argue that the earliest forms of understanding and seeing scores as objective measures came when a child was trying not to get out of bed on a school day, claiming to be sick. The adult immediately went for the ultimate measure—the thermometer! If the child’s temperature did not exceed 98.6 degrees Fahrenheit, then the child was not sick. A score suddenly had meaning.
In school, one is tested on reading, math, and myriad other subject areas. Typically, students take a test and are then told how many questions they got correct or the percentage they got correct. Such scores are often called raw scores. These scores then are given meaning by the teacher, who declares how many correct or what percentage correct (a criterion) determines a particular grade. This is a fairly straightforward presentation of test results by teachers.
On so-called high-stakes tests (tests that are used to evaluate schools, teachers, or stu- dents), the original raw score is often mathematically transformed into a scaled score, though the exact mathematical process is not necessary to understand scaled scores. For example, a raw score of 35 correct (out of 45) on a test might be transformed to a scaled score of 480. If the raw score of 35 was the average score (the mean) of all students who took the test, then 480 becomes the mean of all students for the scaled score. The mathematical process of transformation does nothing to manipulate where every child scored; it just changes the numbering system.
One might immediately ask, Why is this mathematical process done? The main reason is to adjust for multiple forms of a test measuring the same content but with potentially different difficulty levels. Testing companies often develop multiple forms of the same test to avoid overexposure of a test. Thus, second graders in one school in a district may get a different form of a math test than other schools, even though both tests cover essentially the same material. If you have ever taught high school, you know the issue of your second-period class revealing the test questions to your fifth-period class! Test makers understand the importance of test scores to educators and want to ensure test integrity.
In a process called test equating, the test maker can write a second-grade math test (cover- ing the concepts taught in second grade) and embed for consistency some anchor items that would also appear on the alternative form. In this way, a scaled score is equated, no matter what form of the test the student took. Some standardized tests have what are called vertical scaled scores—scores that allow comparisons across grades. If you think about it, the whole idea of scaled scores is to consider difficulty levels. Thus, if a third grader takes a math test and gets a certain scaled score and then takes a more difficult math test in fourth grade, the fourth-grade scaled score (adjusted for difficulty) can be compared to the third-grade score to see if absolute growth has occurred.
Norm and Criterion-based Referencing
Understanding norm groups is important for interpreting test scores. A score only has mean- ing if a reference point has been established. If someone is 4 feet in height, we cannot say that the person is short or tall unless we have a reference point; for instance, is the person 10 years old or 30 years old (different norm groups)? Or, if a person is 3 feet 8 inches in height (44 inches), the critical reference point might be the minimum height of 48 inches to ride the roller coaster (criterion). Thus, understanding or interpreting a test score is often logical, if one understands how the score was developed and defined.
In criterion referencing, a standard is set (usually a number) beyond which a score is given meaning. Back to the GRE: Some universities establish a minimum score (criterion) on the GRE
Interpreting Results Chapter 11
before they will consider admission. Likewise, some high school or trade school typing courses have minimum words-per-minute requirements before they will issue a certificate of typing competency. In these two examples, the criterion was most likely established through studies of successful graduate students or office workers, respectively.
State departments of education will often set a criterion score on the standardized test that is used to determine if a particular district, school, or student is meeting expectations. Experts might set this criterion based on what concepts should be learned at which grade level. There might also be multiple criterion scores set depending on levels of achievement (on grade, below grade, above grade, exemplary, etc.).
In norm referencing, scores gain meaning by comparison to other scores in a group (the norm group). In schools, one often finds students given a percentile rank score. That score reflects where the particular student scored among all students. Thus, if a student’s percentile rank on the third-grade math test was the 78th percentile, we would know that the student scored higher than 78% of the other students in the norm group.
As noted earlier, it is critical that the teacher-leader understand how a particular score was developed and how that score was given meaning. A good example comes from a particular state’s use of a norm group. For example, a school that scores in the 25th percentile or below on the state achievement test (reading or math) is then compared to other schools in that same range the next year to see if growth has occurred, rather than comparing that school to the state as a whole. The confusion comes in differentiating relative growth (norm referenced) from absolute growth (criterion referenced). A simple example might serve to clarify.
Josephine is an 8-year-old girl who is 4 feet 1 inch tall. For illustrative purposes, let’s say this happens to be at the 50th percentile of 8-year-old girls (half of 8-year-old girls are taller, half are shorter). If during the next year, Josephine grows 2 inches to 4 feet 3 inches and the median height of 9-year-old girls is 4 feet 3 inches, then again, she is at the 50th percentile. Did she grow? Absolutely—she grew 2 inches. Relatively, though, she stayed in the middle of the pack for girls her age. If a note were sent home to her parents saying that she had grown 2 inches, her parents would know that she had grown not only because the school said so, but also because they can objectively see the growth. Conversely, if that note home said she was at the 50th percentile on both measurements, her parents would not know how much she had grown. Would they understand that she stayed at the same place among her age peers?
Now, let’s replace the reference to physical height with math achievement scores and any reference to age with grade level.
Josephine is a third grader who had a score of 450 on the math achievement state test. For illustrative purposes, let’s say this happens to be at the 50th percentile of third graders (half of third graders scored higher, half scored lower). If, at the end of the year, Josephine scored 470 (20 points higher) on the same test and the median score for end-of-the-year third graders is 470, then again, she is at the 50th percentile.
Did she grow? Absolutely—she grew 20 points. But relatively, she stayed in the middle of the pack for end-of-the-year third graders. If a note were sent home to her parents saying that she had gained 20 points, her parents would know that she had grown in achievement because the school said so. If that note home said she had remained at the 50th percentile, her parents would not know how much she had grown, if at all. Would they understand that she had grown at the same pace among her age peers, or would they conclude that she had not grown at all? Math achievement is not always objectively observable.
Interpreting Results Chapter 11
Now, consider the implications for those in the third grade who start out below average—for example, those in the 20th percentile. If they grow at the same rate as their peers every year, they will stay at the 20th percentile forever. They may gain in content knowledge, but they stay at the same place as their peers—that is, all their peers would gain the same.
So how can a child or a class move to a higher percentile rank? To do so would require achievement growth faster than that of the norm group to which the child or class is being compared. Think of what that means for teaching. The third-grade teacher must exceed the rate of growth as compared with other third-grade teachers. Having students learn at the same rate will not be good enough. And worse, if the particular class slows a little, they will look even worse.
The pressure on teachers to have children grow academically is immense. As a society, we are comfortable having our children grow in height while maintaining their relative position as average in height. But we will not accept average results over the years in achievement (relative), even if absolute growth is occurring. Everyone has to be above average, which is a mathematical dilemma depending on the norm group. The only way everyone can be above average in a class is if the norm group is broader than the particular class.
Making Sense of the Metrics
States across the country are addressing growth models as part of the Race to the Top initia- tives. Previously, states focused on a system that collected snapshots of student performances; from these snapshots, they drew inferences about students’ progress. The major assumption in this system is that passing one grade level means a student is on track to pass the next grade level. However, very little data support this assumption. The new national trend is to collect more direct evidence from longitudinal student data (O’Malley, McClarty, Magda, & Burling, 2011).
One of the most confusing issues in testing and measurement is that of growth. Some of the confusion comes from the terminology: student growth, value-added models, and teacher effectiveness. According to O’Malley et al. (2011), student growth focuses on the performances of individual students and measures how much a student has progressed or whether the student is on track. Student growth measures produce a label and a score; the label indicates whether the student is on track, and the score informs the amount of gain. For example, a student growth model reports that Troy has made a sufficient performance-level transition (e.g., from high below basic to middle basic), which indicates that he is on track to reach the proficiency performance level within three years. This can be reported as a growth percentile or an actual scale score gain to indicate how a student’s growth compares to the growth of students with similar score history. Student growth measures calculate the growth a student is expected to make. However, student characteristics (e.g., gender, ethnicity) are not included in the model. Thus, the same high expectations are expected of all students, regardless of gender, socioeconomic status, or ethnicity.
Value-added models focus on the effects of teachers and leaders within schools on student score gains. These models use student test scores, student demographics, and other teacher or school-level variables to estimate the value-added measure for a teacher or leaders using three steps:
1. Determine the amount of growth expected for the teacher’s class or leader’s school.
2. Calculate the amount of growth the teacher’s class (leader’s school) actually made.
Interpreting Results Chapter 11
3. Define the difference as the “value” that the teacher or leader added. (O’Malley et al., 2011, p. 2)
For example, after accounting for differences in student (class composition) and teacher or leader characteristics (education, level of experience, demographics), the score gains of stu- dents in Mr. Quezada’s class are compared with score gains in other comparable teacher’s classes or leader’s schools.
Teacher effectiveness measures whether the teacher is successful in improving student out- comes. The measures include subject matter knowledge, communication skills, value-added scores, and pedagogical content knowledge. These measures are obtained through multiple methods (e.g., classroom observations, surveys, portfolios, assessments) that produce an over- all composite effectiveness rating or a separate score on each measure used. A teacher might be rated as exceeds expectation, meets expectations, below expectations, or as satisfactory, needs improvement or unsatisfactory.
Moving from static snapshots of student performance to longitudinal measures of student growth, combined with an evaluation of teachers and leaders, certainly provides a more accurate picture of a school’s or district’s progress toward its educational goals. This under- standing can guide professional development for teachers and leaders and appropriate ways to improve student learning.
Ethical Dilemmas and Integrity of the Testing Process
Since early in 2013, the federal government has issued waivers to 33 states to lift the 2014 deadline of math and reading proficiency requirements as mandated by No Child Left Behind. The government also began to allow states to design new school accountability standards that deemphasize testing and introduce other performance indicators to measure student achievement. However, the era of high-stakes testing is not over. States are still required to administer tests. And now, many school districts are developing and implementing plans to tie teacher evaluations (and, thus, teacher pay) to student achievement as measured by these high-stakes tests.
On July 19, 2011, The Washington Post quoted U.S. Secretary of Education Arnie Duncan saying,
Recent news reports of widespread or suspected cheating on standardized tests in several school districts around the country have been taken by some as evidence that we must reduce reliance on testing to measure student growth and achievement. Others have gone even farther, claiming that cheating is an inevitable consequence of “high-stakes testing” and a threat so we should abandon testing altogether. To be sure, there are lessons to be learned from these jarring incidents, but the existence of cheating says nothing about the merits of testing. Instead, cheating reflects a willingness to lie at children’s expense to avoid accountability—an approach I reject entirely.
The availability of test data is important to improve instruction, identify the needs of individual students, implement targeted interventions, and help students reach high levels of achieve- ment. M. B. Pell (2012) reported a testing consultant commenting with a broad and disturbing statement: “If you think there’s cheating now under school accountability, wait until what you see under teacher accountability.” After the state of Georgia implemented a new teacher evaluation system that depends largely on student standardized test scores, professors from
Interpreting Results Chapter 11
colleges and universities in Georgia wrote an open letter of concern to the governor outlining why teacher accountability models are ineffective (Strauss, 2012).
Cheating can be defined as any action that violates the rules for administering a test. Testing irregularities, such as breaches of test security and student cheating, undermine efforts to use those data to improve student achievement. Amrein-Beardsley (2013) developed a taxonomy that divides cheating into three different degrees, ranging from involuntary and accidental to willful and premeditated, as shown in Table 11.6.
Table 11.6: Definitions of cheating according to severity
Cheating in the First Degree (Willful and premeditated acts)
Cheating in the Second Degree (Subtle forms of misconduct)
Cheating in the Third Degree (Unintentional or accidental)
Erasing and changing student answers
Filling in answers left blank by students
Overtly or covertly providing correct answers on tests
Falsifying student test identification or tracking numbers
Suspending or otherwise excluding students with poor academic perfor- mance on testing days, so that they are not tested.
Cueing students on incorrect answers (e.g., tapping on the desk, head shake, or nudging)
Distributing “cheat sheets;” talking students through processes and definitions
Giving extra time on tests during recess or before or after school
Failure to act when whistleblowers step forward
Poor or haphazard oversight of test security
Not standardizing the administration of the exams to deter cheating
Source: Adapted from Amrein-Beardsley (2013).
For as long as there have been tests, there has been cheating. According to Cizek (2001), “Critics of accountability view cheating as the natural, and not so reprehensible, result of placing undue emphasis on the results of a single test. Some even view cheating as a kind of civil disobedience.”
What can be done to address the problem of cheating? Cizek (2001) suggested the following pragmatic actions:
• Clearly worded guidelines: When caught cheating, school leaders protest that they did not know that they weren’t following the rules. Therefore, every high-stakes test should come with clearly worded guidelines for all who handle testing materials.
• Procedural changes: If bar-coding were used to identify test materials and a tracking system implemented (similar to Fed Ex’s tracking system, which can tell you where a package is at any given time), the ability of educators to cheat would be hindered.
• No “truth in testing”: Some states have laws that require the content of state- mandated tests to be disclosed following the administration of a test. This requirement allows teachers to use previous versions of a test for classroom practice, further nar- rowing instruction. The unintended consequence of this law is a “teach to the test” mentality.
• Scale back: The exclusive use of objective questions for accountability systems tempts educators to alter a bubbled-in response or to provide the key to multiple-choice items. Constructed response formats, though more costly to score, would be less prone to corruption.
Case Study in Educational Leadership Chapter 11
The difficulty with these solutions is that they fail to address the core issues. Redecker and Johannessen (2013) suggested a dramatic shift of the assessment paradigm.
Toward a New Assessment Paradigm
In the past, the use of technology-assisted assessment has improved the validity and reliability of test scores; however, it has still been grounded in the traditional assessment paradigm. Against the background of 21st century skill requirements, Redecker and Johannessen (2013) believe that assessment strategies need to refocus on fostering more holistic key compe- tences. To seize these opportunities, the education community needs to develop new strate- gies for embedded, authentic, and holistic assessment through technological solutions. The authors’ ideas for current and future e-assessments are outlined in Figure 11.2.
Think About It
What features of personalized assessment through learning Web 2.0 tools make this paradigm a good solution to the current ethical dilemmas facing today’s standardized testing paradigm? Is it a good solution?
11.5 Case Study in Educational Leadership School District 16 (SD16) is at a point of transition. The district superintendent has retired after serving 15 years. Five of the six school board members have been voted out of office this past year. Two new principals will be hired over the summer, along with the new superintendent. The district consists of three elementary, two middle, and one high school. The test scores for the district have been less than stellar for the past few years. In frustration over the outcomes of their students, the community voted down a tax referendum earlier in the year as a com- mentary on their dissatisfaction with the way the district is operating.
f11.02_EDU675.ai
Computer-Based Assessment (CBA)
Efficient testing by using technology for test administration. Examples include multiple choice and short answer assessments.
1990 1995 2000 2005 2010 2015 2020 2025
Embedded Assessment
Personalized learning by using technology to provide a collaborative learning environment. Examples include learning analytics, intelligent feedback, virtual worlds, simulations, and peer assessment.
Figure 11.2: Current and future e-assessment strategies
In the past, technology was used for automated administration and scoring for assessments. There is now a shift toward more personalized learning. Technology can be used to provide a collaborative and interactive learning environment. In the future, e-assessment will be embedded, providing personalized feedback to learners.
Source: Adapted from Redecker & Johannessen (2013, p. 82)
Case Study in Educational Leadership Chapter 11
The assistant superintendent, Mr. Brun, is a former teacher in SD16. The board has asked him to step into the interim superintendent position to create some stability and look at possible changes for improvement. Mr. Brun knows how hard the teachers have worked and how difficult it had been to get anything changed in the system with the old superintendent and school board. He is well aware of the tenuous relationship between the teachers and the com- munity at this time. Teachers feel unsupported by the community, while the community feels let down by the teachers and the system. He also knows that in the past, the students have had little say in what they think would improve their learning.
Mr. Brun set out a plan for bringing all of the involved parties together for some serious dia- logue to direct change. He then took the plan, with a summary of recent research as support, to the board for approval. He wanted a critical mass of participants to work on a two-year transformation project that would identify the needs of students, the school, and the com- munity. Then he wanted them to set five-year goals for improving student outcomes, realign- ing organizational structures as appropriate, and crystalizing a vision and mission statement for the district. After the board gave its approval, he spoke with key business organizations and community groups to share the basic idea and ask for financial or other kinds of support for meetings (e.g., food for meetings, rooms away from school at no cost to the group). He gained support from the community when he promised complete transparency on the project through frequent reports to the community, open house opportunities, and question-and- answer sessions.
Mr. Brun held separate meetings with teachers, staff, and administrators; community mem- bers; students and graduates from the district; parents; and business representatives and college representatives to determine who would be on key teams for the two-year transfor- mation project. He clearly outlined his expectations that each person would (a) be involved for the full two years, (b) participate fully in their roles, (c) be open to others’ ideas, (d) bring honest challenges and solutions to the table, and (e) create a risk-free environment for creativ- ity and progress. He informed the teachers, staff, and students that they would have special support if meetings required them to miss their normal work or class time.
Once he had commitment to the project, he held a three-day initiative session. His expected outcomes from this session were for the group to
1. create cross-sectional teams to research, develop, and create enhancements for the dis- trict’s strengths and improvements for areas of weakness;
2. develop specific research projects for the critical areas identified by the teams;
3. draft the vision and mission statements for the district; and
4. write a clear summary of the plans for the two-year project to be published on the district’s website and in the local paper for all community members to review.
During that time frame, a facilitator led the group through an inclusive envisioning process to set specific reform objectives and clear program goals. The facilitator and principal worked closely together to provide support and guidance in the process. The process involved a thor- ough discussion of what was expected of graduates for the district over the next 20 years. Using that information as a base, the group then looked at the current student outcomes for the graduates, including drop-out rates and postsecondary transition success based on how many went to a technical, two- or four-year school, the military, or employment. Furthermore,
Case Study in Educational Leadership Chapter 11
they reviewed the test scores of their students in relationship to the state averages and to outcomes of other state schools with similar demographics.
The teams used these data to identify strengths and gaps in the district’s expectations and outcomes. Once this was completed, they created smaller mixed teams to conduct research on reform initiatives that could enhance their strengths and improve their gaps in performance. Mr. Brun had given the directions that each team’s plan should include, at a minimum, the following:
1. A set of goals
2. A timeline
3. Budget expectations
4. Materials
5. Organizational support
6. Measureable outcomes
7. Training component for others affected by the plan
8. Research and reporting component
9. Method of communicating progress and receiving feedback from the larger community
Using all of the information collected during the first two days of the reform initiative session, the facilitator led the large group in the process of crafting the vision not only of the group for the two-year project but also the overall mission for the district. These statements were to be reviewed and approved by the board and the community.
At the end of the third day, the large group set specific calendar dates for meeting again to review each other’s progress, fine-tune ideas, and report outcomes. In addition, they discussed plans for open dialogue with the community, including creating a website conversation board, offering open house meetings with the community to share and answer questions, and iden- tifying a leader from each group as the contact for more information.
At the closing of the third day, Mr. Brun handed each team member a wristband that had DREAM 16 on it. He wanted each person to wear this to remember how important they were to District 16 and the dreams of current and future students.
Critical Thinking Questions 1. Mr. Brun put together a complex plan. Name the different components of the plan and why they
would be critical to guarantee change.
2. Research shows that for true change to occur, there must be commitment from all involved par- ties. Identify the parties that Mr. Brun thought were important to change. Why do you think he wanted them involved?
3. What would be the next steps for the school district?
4. How would you keep the public informed on the progress of teachers and administrators? What kinds of information would you share?
Post-Test Chapter 11
Summary Never before has assessment been so prominent in schooling and the lives of students as it has been in the past decade due to the national standards movement. The academic standards provide the content that educators use to construct learning outcomes, as defined by objec- tives. Objectives define what students should know, understand, or be able to do because of instruction. Bloom’s cognitive taxonomy and Wiggins and McTighe’s facets of understand- ing provide a framework for developing measurable objectives. Students demonstrate their mastery of the objectives through formative and summative assessments. These assessments take many forms, including traditional paper-and-pencil objective tests and alternative assess- ments, such as project-based learning, portfolios, and authentic tasks. Teachers know what level of mastery students have reached by interpreting raw scores for use in decision making, which we will discuss in more detail in the next chapter.
Post-Test 1. “Students will be able to appreciate art” would be an example of a(n)
a. objective.
b. goal.
c. self-fulfilling prophecy.
d. assessment.
2. Which of the following is not a good method to provide formative feedback?
a. Help students address areas they need to improve on.
b. Provide feedback in small increments.
c. Give timely and routine feedback.
d. Compare a student’s performance with that of other students.
3. A 30-question true-or-false test on Chapters 6–8 of the textbook would be a(n)
a. nonobjective traditional assessment.
b. authentic alternative assessment.
c. nonauthentic alternative assessment.
d. objective traditional assessment.
4. Pablo got 82% correct on his physics test. This value is a
a. scaled score.
b. standardized score.
c. true score.
d. raw score.
Post-Test Chapter 11
5. Ms. Pumpernickel has an objective that states, “Students list major battles from the Civil War.” She then teaches students the major battles, has them make study aids with lists of the battles, and tests them by having them orally list the battles to her individually. This skill would fall within which of Bloom’s (1956) taxonomy levels?
a. Remembering
b. Evaluating
c. Analyzing
d. Understanding
6. Another term for performance indicator would be
a. evidence.
b. objective.
c. target.
d. process.
7. A portfolio
a. should include a reflective component.
b. is the same as a scrapbook of academic work.
c. should have goals and objectives determined as it is assembled.
d. tend to be more restrictive and inflexible than traditional assessments.
8. Toni’s score at the 74th percentile means that she
a. got 74% of the items correct.
b. got a score lower than 26% of those who took the same test.
c. got the same score as 74% of those who took the same test.
d. is below average.
Answers 1. b. goal. The answer can be found in Section 11.1.
2. d. Compare a student’s performance with that of other students. The answer can be found in Section 11.2.
3. d. objective traditional assessment. The answer can be found in Section 11.3.
4. d. raw score. The answer can be found in Section 11.4.
5. a. Remembering. The answer can be found in Section 11.1.
6. a. evidence. The answer can be found in Section 11.2.
7. a. should include a reflective component. The answer can be found in Section 11.3.
8. b. got a score lower than 26% of those who took the same test. The answer can be found in Section 11.4.
Key Terms Chapter 11
Key Ideas • Any assessment must be reliable and valid.
• Scores may start out as the number of items answered correctly on a particular test; for standardized tests, however, they are often mathematically transformed into scaled scores.
• Criteria are established to define the norm group and are used for accurate assessment of test scores.
• Systemic cheating gives educators a false understanding of student achievement.
Critical Thinking Questions 1. What constitutes evidence of understanding? How can we promote understanding more
by design than by good fortune or native ability?
2. Why is it important to understand norm groups for educational testing purposes?
3. Why is measuring student growth so difficult?
4. How has high-stakes testing changed the educational environment over the past decade?
5. What are the ethical dilemmas of the testing process?
Key Terms affective That which appeals to the emotional or social side.
assess To gather data to make informed decisions.
backward curriculum design Starting instruction with the end (and evidence of real understanding or transfer) in mind.
Bloom’s taxonomy A hierarchical yardstick or measure used to describe learning out- comes. It uses behaviors for educational objectives.
criterion referencing A standard set (usually a number) beyond which a score is given meaning.
evaluate To judge the worthiness of something or to compare it to a set of criteria.
feedback Provides information to learners to help them improve their product. It has no evaluative component.
formative assessment Frequent and ongoing “checkpoints” on students’ progress and the basis for feedback.
high-stakes tests Tests that are used to evaluate schools, teachers, or students.
measurement Assigning numbers to some characteristic.
norm referencing Scores gain meaning by comparison to other scores in a group.
raw score Original data that has not been aggregated or transformed into scaled scores.
Additional Resources Chapter 11
reliability Results are repeatable and consistent—that is, they are reliable.
rubrics Provide a guideline for measuring a student’s work.
scaled score The transformation of one score (the raw score) through a mathematical process.
student growth Measures how much the student progresses and whether the student is on track.
summative assessment Usually completed at the end of a learning segment, where stu- dents demonstrate mastery of key essential understandings; it is gradable.
teacher effectiveness Measures the overall effectiveness (satisfactory, needs improve- ment, unsatisfactory) of the teacher in improving student outcomes.
test equating A statistical process of determining comparable scores on different form of an exam covering the same content.
validity Asks whether measures truly measure what was intended to be measured.
value-added models Measure how score gains of the students of one teacher or leader compare with average score gains to indicate if students grow more or less than expected.
vertical scaled score Converts the raw score of the number of questions answered cor- rectly into a common unit of measurement to allow comparisons across grades.
Additional Resources Further Reading
• Improving school accountability measures: http://www.nber.org/papers/w8156. pdf?new_window=1
Videos
• High-stakes test’s cheating dilemma: http://www.youtube.com/watch?v=nUNvzaIQjng
• Norm- vs. criterion-referenced scoring: Advantages & disadvantages: http://www.youtube.com/watch?v=nUNvzaIQjng
• Validity and reliability: How to assess the quality of a research study: http://education -portal.com/academy/lesson/validity-and-reliability-how-to-assess-the-quality-of-a -research-study.html
• Formative Assessment
– Formative assessment and differentiated learning: http://www.youtube.com/ watch?v=gFXbuE-21I4
– Rick Wormeli on formative assessment: http://www.youtube.com/ watch?v=rJxFXjfB_B4
– Examples of formative assessment: http://wvde.state.wv.us/teach21/ ExamplesofFormativeAssessment.html)
Additional Resources Chapter 11
Weblinks
• Applying technology to Bloom’s taxonomy: http://www.schrockguide.net/bloomin -apps.html
• Educational goals and objectives: http://www.ineedce.com/courses/1561/PDF/ed_ goals_objctvs.pdf
• Elementary & Secondary Education Reauthorization: A blueprint for reform: http:// www2.ed.gov/policy/elsec/leg/blueprint/index.html
• Project-based learning for the 21st century: Rubrics: http://www.bie.org/tools/freebies/ cat/rubrics
• Revised Bloom’s taxonomy: http://www.utar.edu.my/fegt/file/Revised_Blooms_Info.pdf