Paper due for American National Government class on Every Student Succeeds Act.
Educational Researcher, Vol. 46 No. 7, pp. 378 –396 DOI: 10.3102/0013189X17726752 © 2017 AERA. http://edr.aera.net
378 EDUCATIONAL RESEARCHER
I n recent years, policy reforms at the federal, state, and local levels have dramatically changed the ways that educators are evaluated (Donaldson & Papay, 2015). These reforms, along
with growing public scrutiny, arose from widespread recogni- tion that traditional teacher evaluation systems neither differen- tiated between low- and high-performing teachers (Donaldson, 2009; Toch & Rothman, 2008; Tucker, 1997; Weisberg et al., 2009) nor provided teachers with meaningful feedback about their practice (Almy, 2011; Sartain, Stoelinga, & Brown, 2011; Sinnema & Robinson, 2007; Stronge & Tucker, 2003). By the 2015–2016 school year, 88% of both states and the largest 25 districts and the District of Columbia had revised and imple- mented new teacher evaluation systems (Steinberg & Donaldson, 2016).
Traditional systems of teacher evaluation tended to be per- functory exercises, relying on a single measure of teacher perfor- mance (typically a cursory observation of classroom practice) and binary summative ratings (i.e., proficient or not), and few if any consequences were tied to teachers’ summative ratings (Weisberg et al., 2009). In contrast, teachers’ evaluation ratings under newly implemented evaluation systems have become high stakes for both individual teachers and for districts as a whole. Policymakers and the public are increasingly asking districts to reconcile teachers’ evaluation ratings with the performance of their students. This is in light of evidence that, under both
traditional evaluation systems and many newly implemented systems, nearly all teachers continue to be rated professionally proficient (Anderson, 2013; Kraft & Gilmour, 2017; Steinberg & Sartain, 2015).1
Efforts to reform teacher evaluation systems have been focused on three primary system design features: the incorpora- tion of multiple measures of teacher performance, the use of multiple performance ratings categories, and the creation of pro- fessional support and incentive structures tied to teachers’ rat- ings. District policymakers have been deeply engaged in decisions about which performance metrics should be incorporated into their evaluation systems, including test-based performance mea- sures, such as value-added measures (VAMs) or student growth percentiles (SGPs), as well as rubric-based observation ratings of a teacher’s instructional practice. Further, nearly all new systems have expanded the range of performance ratings to include at least four categories defining a teacher’s summative performance. Teachers who receive low ratings—typically the bottom two rat- ings categories—are now overwhelmingly required to participate in additional targeted professional development and are increas- ingly at risk of being terminated or nonrenewed during the ten- ure review process (Steinberg & Donaldson, 2016).2 Teachers
726752EDRXXX10.3102/0013189X17726752Educational ResearcherMONTH XXXX research-article2017
1University of Pennsylvania, Philadelphia, PA 2Brown University, Providence, RI
The Sensitivity of Teacher Performance Ratings to the Design of Teacher Evaluation Systems Matthew P. Steinberg1 and Matthew A. Kraft2
In recent years, states and districts have responded to federal incentives and pressure to institute major reforms to their teacher evaluation systems. The passage of the Every Student Succeeds Act in 2015 now provides state policymakers with even greater autonomy to redesign existing evaluation systems. Yet, little evidence exists to inform decisions about two key system design features: teacher performance measure weights and performance ratings thresholds. Using data from the Measures of Effective Teaching study, we conduct simulation-based analyses that illustrate the critical role that performance measure weights and ratings thresholds play in determining teachers’ summative evaluation ratings and the distribution of teacher proficiency rates. These findings offer insights to policymakers and administrators as they refine and possibly remake teacher evaluation systems.
Keywords: descriptive analysis; educational policy; policy analysis; school/teacher effectiveness; teacher assessment
FEATURE ARTICLES
OCTObER 2017 379
with exemplary ratings may be rewarded with merit pay or pro- moted to new positions on a career ladder (Donaldson & Papay, 2015; Steinberg & Donaldson, 2016).
Research on teacher evaluation reforms has mirrored these patterns. Most existing studies focus on the reliability and valid- ity of performance measures—VAMs (e.g., Chetty, Friedman, & Rockoff, 2014; Kane, McCaffrey, Miller, & Staiger, 2013), class- room observation rubrics (e.g., Garrett & Steinberg, 2015; Hill, Charalambous, & Kraft, 2012; Kane & Staiger, 2012), and stu- dent surveys (e.g., Kane & Cantrell, 2010; Wallace, Kelcey, & Ruzek, 2016). A related line of research evaluates how these new high-stakes systems affect teacher performance, student achieve- ment, and teacher turnover (Cullen, Koedel, & Parsons, 2016; Dee & Wyckoff, 2015; Steinberg & Sartain, 2015; Sartain & Steinberg, 2016). Even practitioner-facing guidebooks and edited volumes have primarily focused on how to design evalua- tion systems to evaluate teachers more reliably and/or leverage the evaluation process to promote teacher development (Darling- Hammond, 2013; Grissom & Youngs, 2015; Kane, Kerr, & Pianta, 2014; Marzano & Toth, 2013).
With this article, we illustrate the central role that two equally important system design features play in shaping teachers’ sum- mative evaluation ratings but that have received far less policy and research attention: performance measure weights and summative evaluation ratings thresholds. In comparison to decisions about which measures to choose and how to design consequential incen- tives, state and local policymakers have almost no empirically based evidence to inform their decision process about how to combine scores across multiple performance measures and then how to map these summative evaluation scores onto performance ratings categories. Informal conversations with administrators and researchers involved in the design process suggest that decisions about weights and performance ratings thresholds are often made through a somewhat arbitrary and iterative process, one that is shaped by political considerations in place of empirical evidence. As we demonstrate in this article, these decisions can have impor- tant consequences for both individual teachers’ ratings and the share of teachers deemed to be professionally proficient.
The passage of the Every Student Succeeds Act (ESSA) in December 2015 makes research that can inform the evaluation system design process more important now than ever before. ESSA has ushered in a new phase in the teacher evaluation reform movement by granting states and districts considerable autonomy to redesign and implement teacher evaluations sys- tems independent of federal influence. Research that informs the evaluation system design process is especially important given the existence of what Richard Elmore (2002) termed the “capac- ity gap” in state departments of education: the gap between what they are expected to do and what they are staffed to accomplish (Le Floch, Boyle, & Therriault, 2008). Several recent studies have found that limited technical expertise in state departments of education has constrained their ability to fully attend to all important design features of teacher evaluation systems (Herlihy et al., 2014; McGuinn, 2012).
We address this need by conducting simulation-based analyses to examine how teachers’ summative evaluation ratings are affected by the decisions district administrators make about the weights they assign to multiple performance measures and the ratings
thresholds that they choose. Specifically, we investigate how the proportion of teachers deemed professionally proficient changes under different weighting and ratings thresholds schemes. We examine how teacher proficiency rates change when we vary the performance weights (holding the ratings thresholds scheme fixed), when we vary the ratings thresholds scheme (holding the performance weights fixed), and how these design decisions inter- act with one another (i.e., when we jointly vary performance weights and ratings thresholds). Our analyses also allow us to pro- vide additional empirical insights into how the properties of teacher evaluation measures—specifically, the mean, variance, and cross-measure correlations—shape the distribution of teacher pro- ficiency ratings under these different system design parameters.
It is straightforward to infer that teacher proficiency rates will improve as, for example, greater weight is given to performance measures with higher average scores and/or the minimum threshold required to receive a proficient rating is lowered. Ours is the first article, to our knowledge, to more precisely illustrate the degree to which marginal changes in the weights assigned to performance measures and the placement of ratings thresholds can shift the distribution of teacher ratings and, ultimately, affect the proportion of teachers deemed professionally proficient. Although our findings are not intended to provide specific rec- ommendations about what weights and ratings to select—such decisions are fundamentally subject to local district priorities and preferences—they do offer important insights about how these decisions will affect the distribution of teacher perfor- mance ratings as policymakers and administrators continue to refine and possibly remake teacher evaluation systems.
We accomplish this by drawing on rich data collected by the Measures of Effective Teaching (MET) Project. The MET Project affords a unique opportunity to examine the sensitivity of teacher performance ratings. In particular, the MET data contain a wide range of performance measures that are common to more than 1,000 teachers and that are currently being incorporated into new teacher evaluation systems. We use these data to illustrate how the summative performance ratings for the same set of MET teachers change as we impose different evaluation design param- eters based on existing evaluation systems across a range of large urban school districts. To do this, we first construct teacher evalu- ation scores from combinations of three performance measures found in many new teacher evaluation systems: scores from class- room observation rubrics, estimates of teachers’ contributions to student achievement, and student survey responses capturing their perceptions of teacher performance in the classroom. We then combine these data with publicly available information on the performance ratings thresholds used across eight large and geographically diverse urban school districts. Together, these data allow us to conduct a range of simulation analyses that illustrate the consequences of different weighting regimes and ratings thresholds. Although our analyses focus on teachers in tested grades and subjects, we also discuss how our findings relate to the evaluation ratings received by the majority of teachers who teach in nontested grades/subjects.
We first describe the considerable variation across districts in both the weights they assign to different performance measures and the percentage of available evaluation points required to earn a given summative evaluation rating. We then show how teachers
380 EDUCATIONAL RESEARCHER
can receive substantially different summative ratings, with the same underlying scores on individual performance measures, depending on how weights are assigned to individual performance measures and how summative performance scores map on to sum- mative rating categories. Our findings also reveal the important role that the properties of teacher performance measures play in determining teacher proficiency rates. First, if all performance measures are assigned equal weight, then the measure with the highest cross-measure correlation will contribute the most to a teacher’s summative evaluation score. Second, teacher perfor- mance measures that are weakly correlated with the other mea- sures will contribute less to a teacher’s summative score than would be expected based on the weight that the evaluation system assigns to it. And third, teacher proficiency rates depend not just on the properties of performance measures but also on the location of the proficiency threshold relative to the actual distribution of teachers’ summative evaluation scores. In evaluation systems where the pro- ficiency threshold is located near the center of the distribution of teachers’ summative evaluation scores, proficiency rates will be more sensitive to marginal changes in performance measure weights than in evaluation systems where the proficiency thresh- old is located at the upper end of the score distribution (where the density of teachers is lower). We conclude by discussing the impli- cations of these findings for policy and practice.
The Anatomy of a Teacher Evaluation System
The process of assigning a summative evaluation rating to teach- ers is shaped by four primary design features of a teacher evalua- tion system: (a) the teacher performance measures used, (b) the approach used to place performance measures on a common scale, (c) the weights assigned to teacher performance measures, and (d) the performance ratings thresholds. We describe each of these design features in detail below.
Teacher Performance Measures
A key feature of newly implemented evaluation systems is the incorporation of multiple measures of teacher performance. In this article, we focus on three distinct and widely used measures: observations of a teacher’s classroom instruction, a teacher’s con- tribution to student achievement growth, and students’ percep- tions of teacher effectiveness. According to a recent analysis documenting the extent of teacher evaluation reform, all 46 states and 23 districts (of the largest 25 school districts and the District of Columbia) that have, or plan to have, implemented new teacher evaluation systems no later than the 2016–2017 school year incorporate classroom observation as a measure of teacher performance. Further, 80% of these states and districts use one or more measures of teacher performance based on student achieve- ment.3 Finally, 17% of these states (eight) and districts (four) incorporate student surveys capturing students’ perceptions of teacher performance (Steinberg & Donaldson, 2016).
Classroom observations. Observation rubrics provide scales for criterion-based assessments of a teacher’s classroom instruc- tion and professional practice. Evaluation system designers first select among classroom observation protocols (Framework for
Teaching [FFT], CLASS, PLATO, etc.) and decide whether to incorporate the full protocol (i.e., all observation components across multiple domains of teacher practice) or a subset of the domains. Next, designers decide on the number of formal/ informal classroom observations that each teacher is subject to and who (e.g., principal, assistant principal, master teacher) is responsible for conducting the classroom observation and rating a teacher’s instructional and professional practices on the chosen observation rubric. Evaluation scores from multiple observations are then combined to construct a final teacher practice score.4
Contributions to student achievement. Measures of teacher perfor- mance based on student achievement rely on student test scores and aim to capture student growth attributable to the teacher’s instructional performance. Evaluation system designers choose a particular statistical approach for calculating a teacher’s contri- bution to student learning based on state-administered standard- ized exams; such approaches include teacher-level VAMs and/or SGPs.5 These norm-referenced measures capture teachers’ con- tributions relative to their peers rather than on an absolute scale. Because upward of 70% of teachers nationwide do not teach in grades and/or subjects in which state-administered exams are available (Watson, Kraemer, & Thorn, 2009), many systems also incorporate criterion-based student learning objectives (SLOs) to evaluate a teacher’s contribution to student learning.
Student surveys. Student feedback on teacher performance is captured by student perspective surveys. These surveys ask stu- dents to report about their teacher’s performance and objective occurrences of specific instructional practices. Designers select among a variety of surveys (such as the Tripod survey) and then determine how to construct scores based on students’ responses to create these criterion-based measures.
Placing Teacher Performance Measures on a Common Scale
Once a teacher has been evaluated on multiple performance measures, consideration must be given to how to place these dif- ferent measures, which typically vary in how they are scored, onto the same scale. For example, classroom observations that use the FFT observation rubric are scored on an integer scale from 1 to 4 (i.e., a range of 3). In contrast, VAM scores have no theoretical minimum or maximum value (i.e., an infinite range), and the mean score is centered at zero. By placing teacher perfor- mance measures on a common scale, weights can be applied to each performance measure to construct a teacher’s summative evaluation score. Then, ratings thresholds can be applied to the summative evaluation score to determine a teacher’s summative evaluation rating.
District policymakers therefore determine (a) the point range for the common scale and (b) the mapping of points from differ- ent measures onto a common scale. In practice, there exists con- siderable variation in the point range assigned to a teacher’s summative evaluation score. Table 1 provides the range of avail- able evaluation points across a purposeful sample of eight large and geographically diverse districts that have newly implemented teacher evaluation systems. For example, available evaluation points in Chicago Public Schools range from 100 to 400; in New
OCTObER 2017 381
York City, available evaluation points range from 0 to 100; in Philadelphia, available evaluation points range from 0 to 3. Importantly, the distribution of teacher proficiency—which will depend on the performance measure weights and ratings thresholds— will be invariant to the choice of the range of a common point
scale. District policymakers typically assign points to each perfor- mance measure on a one-to-one basis, because the weight applied to different performance measures allows for local preferences to guide decisions about which measure should have more (or less) influence on a teacher’s summative evaluation rating.
Table 1 Overview of Teacher Evaluation System Designs
Summative Performance Ratings
District (State) Evaluation System
(Year of Study)
District Ranking
(Size) Teacher Performance
Measure (Weight)
Scale Range (Available
Points)
Threshold (% of Available
Points) Category (Level)
Chicago Public Schools (Illinois)
Recognizing Educators Advancing Chicago’s Students (2014–2015)
3rd Observation (70%) 100–400 (300 points) 0%–36% Unsatisfactory (1) Student achievement (30%) 36%–61% Developing (2) Student survey (0%) 61%–80% Proficient (3) Other (0%) 80%–100% Excellent (4)
Clark County School District (Nevada)
Nevada Educator Performance Framework (2014–2015)
5th Observation (100%) 1–4 (3 points) 0%–30% Ineffective (1) Student achievement (0%) 30%–60% Minimally effective (2) Student survey (0%) 60%–86% Effective (3) Other (0%) 86%–100% Highly effective (4)
Denver Public Schools (Colorado)
Leading Effective Academic Practice (2014–2015)
34th Observation (40%) 0–50 (50 points) 0%–47% Not meeting (1) Student achievement (50%) 47%–64% Approaching (2) Student survey (10%) 64%–81% Effective (3) Other (0%) 81%–100% Distinguished (4)
Fairfax County Public Schools (Virginia)
Teacher Performance Evaluation Program (2012–2013)
11th Observation (60%) 10–40 (30 points) 0%–30% Ineffective (1) Student achievement (40%) 30%–50% Developing/needs
improvement (2) Student survey (0%) 50%–80% Effective (3) Other (0%) 80%–100% Highly effective (4)
Gwinnett County Public Schools (Georgia)
Gwinnett Teacher Effectiveness System (2014–2015)
13th Uses matrix rather than weights for observation and student achievement
0–30 (30 points) 0%–20% Ineffective (1) 20%–53% Needs development (2) 53%–87% Proficient (3) 87%–100% Exemplary (4)
Miami-Dade County Public Schools (Florida)
Instructional Performance Evaluation and Growth System (2014–2015)
4th Observation (50%) 0–100 (100 points) 0%–36% Unsatisfactory (1) Student achievement (35%) 36%–73% Developing (2) Student survey (0%) 73%–88% Effective (3) Other (15%) 88%–100% Highly effective (4)
New York City Department of Education (New York)
NYC Advance (2013– 2014)
1st Observation (60%) 0–100 (100 points) 0%–64% Ineffective (1) Student achievement (40%) 64%–74% Developing (2) Student survey (0%) 74%–90% Effective (3) Other (0%) 90%–100% Highly effective (4)
School District of Philadelphia (Pennsylvania)
Educator Effectiveness System (2012–2013)
19th Observation (50%) 0–3 (3 points) 0%–16% Failing (1) Student achievement (50%) 16%–50% Needs improvement (2) Student survey (0%) 50%–83% Proficient (3) Other (0% ) 83%–100% Distinguished (4)
Note. See the appendix for full details about data sources. District ranking (size) is based on district enrollment (National Center on Teacher Quality, n.d.). Teacher performance measures (and associated weights) are for teachers in tested grades/subjects. For evaluation systems that evaluate teachers on professional responsibility (Clark County, Nevada; Denver, Colorado; Fairfax County, Virginia; Miami, Florida), we have included the weight assigned to professional responsibility as part of the weight assigned to a teacher’s observation score. Some systems incorporate multiple measures of student achievement (in addition to value-added measures) into teachers’ summative evaluation scores (e.g., Chicago, Illinois; New York, New York); we have aggregated all student achievement-based measures of teacher performance into the weight assigned to student achievement. In some evaluation systems, performance measures other than observation, student achievement, or student surveys are incorporated into teachers’ summative evaluation scores (e.g., professional development plan, as in Miami, Florida).
382 EDUCATIONAL RESEARCHER
Performance Measure Weights
After multiple performance measures have been selected and scores have been placed on a common scale, designers must decide how to combine scores into a single summative evalua- tion score. In the vast majority of systems, this is done by assign- ing weights (relative proportions of a teacher’s summative evaluation score) to each performance measure. For example, if we were to randomly select a teacher teaching in a tested grade/ subject across the nation’s largest school districts with newly implemented evaluation systems (i.e., the typical teacher in a tested grade/subject nationwide), 82% of his or her evaluation score (and subsequent summative evaluation rating) will be based on the three performance measures we use in our analyses: classroom observations of teacher practice (52%), student per- formance on state-administered exams (28%), and student sur- veys (2%). The balance of this teacher’s evaluation will depend on other measures of teacher performance, including SLOs, schoolwide achievement, professional conduct, and/or parent/ caregiver surveys (Steinberg & Donaldson, 2016). In some eval- uation systems, scores are not aggregated into a single summative evaluation score but instead are mapped from multiple perfor- mance measures onto a rating category using a ratings matrix (e.g., Gwinnett County Public Schools in Table 1).
Performance Rating Thresholds
Given a teacher’s summative evaluation score, a teacher’s perfor- mance rating in a given school year is most often determined by an evaluation system’s ratings thresholds.6 These thresholds are based on the percentage of available points that a teacher earns for his or her performance across multiple evaluation measures. The percentage of available evaluation system points that a teacher earns may be calculated as
Percent of Available Points Earned =
Summative Evaluation Scorre-Minimum Score
Maximum Score-Minimum Score ( )
( ) .
(1)
For example, if the evaluation system point scale ranges from a minimum of 1 to a maximum of 4 points (i.e., a scale range of 3) and a teacher’s summative evaluation score is 2.5, then a teacher has earned 50% of available evaluation system points (i.e., [2.5 – 1]/[4 – 1]). An important implication of this evalua- tion system design feature is that once a teacher has been evalu- ated and has earned his or her evaluation points (and, by extension, the percentage of available points in a given evalua- tion system), the same teacher may be rated differently depend- ing on where the system sets its ratings thresholds.
As shown in Table 1 (and accompanying Figure 1), newly implemented teacher evaluation systems assign quite different thresholds to determine teachers’ summative ratings. For exam- ple, a teacher who earns 60% of available (district-specific) evalu- ation points on his or her summative evaluation score would be rated the lowest possible rating level based on New York City’s evaluation system; the second-lowest rating level in Chicago, Denver, and Miami-Dade; but proficient (Level 3) in Clark County, Fairfax County, Gwinnett County, and Philadelphia. A
teacher in New York City must earn at least 74% of available evaluation points to be rated proficient/effective (Level 3 in each district), whereas a teacher in Philadelphia must earn 50% of available evaluation points to receive the same rating. Such varia- tion suggests that districts may differ in both their views concern- ing what it means for teachers to meet proficiency standards and the degree of difficulty in earning evaluation score points across different evaluation systems. Districts, may, for example, adjust their ratings thresholds to correspond to the degree of difficulty of earning points on the district-specific evaluation measures. This may result in districts with vastly different ratings thresholds having quite similar distributions of teacher performance. In our simulation-based analyses described below, we hold constant the set of performance measures used to better illustrate how differ- ent ratings thresholds shape teacher proficiency rates.
Data and Sample
We use data from the MET study, which was carried out over 2 school years (2009–2010 and 2010–2011) and across six dis- tricts.7 The MET study is among the most ambitious efforts to date to systematically measure teacher effectiveness and affords a unique opportunity to examine the sensitivity of teacher perfor- mance ratings. In particular, the MET data contain a wide range of performance measures that are common to more than 1,000 teachers and that are currently being incorporated into new teacher evaluations systems, including measures based on teacher practice, student achievement, and student reports. Teacher practice is measured using multiple classroom observation pro- tocols; we use scores from Danielson’s (1996) FFT protocol
FIGURE 1. Evaluation system ratings thresholds Stacked bars represent the range of evaluation score points associated with each of the four different evaluation rating categories across eight districts. Level 1 corresponds with a district’s lowest (of four) evaluation rating; Level 3 corresponds with a proficient or effective rating; Level 4 corresponds with a district’s highest evaluation rating. Performance ratings thresholds for each district’s evaluation system are captured by the vertical lines separating ratings categories. All scoring systems have been rescaled so that they map onto a common point scale ranging from 1 to 4 total available evaluation system points. See Table 1 for more detail on each district’s evaluation system.
OCTObER 2017 383
given its widespread adoption across newly implemented teacher evaluation systems (Garrett & Steinberg, 2015; Steinberg & Donaldson, 2016). Teacher performance based on student achievement is measured by VAM scores calculated by MET Project researchers. Teacher performance based on students’ reports is measured using student responses on the Tripod sur- vey. In the next section, we discuss how we construct scores for each teacher from the FFT, VAM, and Tripod survey data.
The teacher sample includes 1,275 teachers in Grades 4 to 8 who participated in the MET study during the 2009–2010 school year. We focus on the 1st year of the MET study because all teachers were assigned to classes by the normal, within-school process, in contrast to the 2nd year, when many MET teachers were randomly assigned to classes just prior to the start of the 2010–2011 school year. We are interested in how teacher ratings may be sensitive to weighting schema under conditions in which teachers are assigned to their classes in the typical manner (that is, nonrandomly). Table 2 summarizes the characteristics of teachers included in the sample. All simulations are based on the full sample of teachers (N = 1,275).
Empirical Approach
Constructing Performance Measure Scores
We begin by constructing a score for each teacher on each of the three performance measures—teacher practice, teacher contri- butions to student achievement, and students’ reports of teacher practice—as described below.
Classroom observations of instructional practice. The MET Project used an abbreviated version of the Danielson FFT observation protocol, including eight components across two domains—Domain 2 (the classroom environment) and Domain 3 (instruction)—with each component rated on a 1-to-4 (unsatis- factory to distinguished) integer scale. Scores for each of the eight FFT components were generated by MET raters from videos of subject-specific (e.g., math or English language arts [ELA]) lessons that MET teachers conducted on multiple occasions dur- ing the 2009–2010 school year.8 We average across FFT compo- nents within lesson observations and then average across lesson observations (within a teacher) to generate a teacher’s practice score. We create both a subject-specific practice score (FFTis) for teacher i observed teaching lesson subject s (math or ELA) as well as an aggregate practice score (FFTi,Aggregate) across all les- sons and subjects.9 This simple approach is used by most school districts to construct classroom observation scores, and prior research using the MET data has found this to be an appropriate approach for aggregating teacher effectiveness measures based on classroom observation scores (Garrett & Steinberg, 2015; Kane et al., 2013; Mihaly, McCaffrey, Staiger, & Lockwood, 2013).10
We find that MET teachers received an average FFT score of 2.5 (see Table 3), approximately half a point (and more than one standard deviation) lower than mean FFT scores received in newly implemented teacher evaluation systems (e.g., Chicago Public Schools; Jiang & Sporte, 2016; Steinberg & Jiang, 2016; and Pennsylvania; Lipscomb, Terziev, & Chaplin, 2015). This is likely the result of several factors. First, MET raters were not physically present in the classroom, as is the case with school-based
Table 2 Teacher Characteristics
Teacher Characteristic All Teachers Math Teachers ELA Teachers
Female 0.83 0.82 0.87 White 0.58 0.56 0.58 Black 0.35 0.37 0.36 Hispanic 0.05 0.05 0.05 Other 0.02 0.02 0.01 Grade 4 0.22 0.28 0.28 Grade 5 0.23 0.30 0.31 Grade 6 0.21 0.16 0.16 Grade 7 0.17 0.13 0.12 Grade 8 0.17 0.13 0.13 Experience (total) 10.7
(8.89) 10.9
(9.56) 10.4
(8.43) Experience (district) 7.6
(6.85) 7.2
(6.63) 7.3
(6.62) Master’s or higher 0.36 0.42 0.40 Generalist 0.34 0.52 0.49 Teachers 1,275 833 874 Schools 207 189 196 Districts 6 6 6
Note. Proportions reported for all characteristics except experience, which reports mean with standard deviation in parentheses. Data are from the 2009–2010 school year. Generalist teachers are included in both the math and ELA teacher samples. For the full teacher sample, 1,237 teachers reported gender, 1,235 reported race, 580 reported years of experience (total), 992 reported years of experience (in district), and 993 reported educational attainment (master’s or higher). ELA = English language arts.
384 EDUCATIONAL RESEARCHER
evaluators. Second, MET raters had no personal connections to the teachers they rated remotely and did not participate in either pre- or postobservation meetings with teachers, as is the practice in many newly implemented evaluation systems. Third, MET rat- ings of teacher practice were not tied to consequential, high-stakes personnel decisions. Research has found that evaluators systemati- cally assign higher summative ratings to teachers relative to forma- tive ratings that are decoupled from high-stakes consequences (Kraft & Gilmour, 2017). Fourth, MET raters received extensive training and were required to pass certification tests in order to conduct remote observations. Finally, MET raters used an abbre- viated version of the FFT instrument, which did not require them to evaluate teachers on domains such as Planning and Preparation or Professional Responsibilities. Recent evidence from Baltimore Public Schools indicates that school-based evaluators (i.e., princi- pals and assistant principals) rate teacher practice, based on class- room observations scores, higher than evaluators who are external to the teacher’s school (Jackson & Steinberg, 2017).
To more closely approximate the consequential ratings teach- ers receive in the context of newly implemented evaluation sys- tems, we adjust MET teachers’ FFT scores by adding 0.5 points (which we refer to as FFTA). By shifting the mean of the FFT scores, we are able to better approximate (though not replicate)
the distribution of teachers’ summative evaluation ratings in newly implemented teacher evaluation systems and to more clearly illustrate how performance measure weights and ratings thresholds shape the distribution of teacher effectiveness. We do not adjust the variance of the underlying FFT scores given by external MET raters, because we find that the variance in MET FFT scores is no greater (and in some cases is lower) than the variance of observation scores found in newly implemented sys- tems in Chicago (Jiang & Sporte, 2016; Steinberg & Jiang, 2016) and Pennsylvania (Lipscomb et al., 2015). By not upwardly adjusting the variance of MET FFT scores, we avoid overstating the effective weight—which is increasing in the vari- ance of the underlying teacher performance measure—given to observation scores by external raters in the MET data (Schochet, 2008). Our substantive findings presented below are not sensi- tive to adjusting the mean of the FFT scores.
Student reports of teacher practice. Students’ reports of their teachers’ practices were captured using the Tripod Elementary and Secondary surveys developed by Ron Ferguson (Kane & Cantrell, 2010). Both versions of the Tripod survey are orga- nized around seven domains—the “7Cs”—of a teacher’s class- room practice (i.e., care, control, clarify, challenge, captivate,
Table 3 Teacher Performance Measures
Performance Measure All Teachers Math Teachers ELA Teachers Generalist Teachers
Panel A: Classroom observations FFT (aggregate) 2.52
(0.316) 2.46
(0.329) 2.49
(0.364) 2.61
(0.217) FFT (math) 2.52
(0.304) 2.46
(0.329) — 2.59
(0.261) FFT (ELA) 2.56
(0.317) — 2.49
(0.364) 2.63
(0.239) FFTA (aggregate) 3.02
(0.316) 2.96
(0.329) 2.99
(0.364) 3.11
(0.217) FFTA (math) 3.02
(0.034) 2.96
(0.329) — 3.09
(0.261) FFTA (ELA) 3.06
(0.317) — 2.99
(0.364) 3.13
(0.239) Panel B: Student achievement VAM (aggregate) 2.41
(0.354) 2.51
(0.341) 2.31
(0.293) 2.42
(0.392) VAM (math) 2.52
(0.416) 2.51
(0.341) — 2.52
(0.475) VAM (ELA) 2.31
(0.363) — 2.31
(0.293) 2.32
(0.422) Panel C: Student perceptions Survey 3.12
(0.287) 3.01
(0.294) 3.06
(0.282) 3.29
(0.202) Teachers 1,275 401 442 432
Note. Mean (standard deviation) of teacher performance measures reported. Data are from the 2009–2010 school year. VAM and Survey measures have been transformed so that they are on a 1-to-4 continuous scale. FFTA is the teacher’s adjusted FFT score, adjusted by adding 0.5 points to FFT score. For subject specialists teaching multiple sections of the same subject (i.e., math or ELA teachers), aggregate FFT and VAM scores are a weighted average (weighted by section enrollment) for a single subject. For subject-matter generalists, aggregate FFT and VAM scores represent the average of math and ELA scores for the same section (class) of students. Among the full sample of 1,275 teachers, 833 teachers had VAM (math), 874 teachers had VAM (ELA), 817 had FFT (math), and 867 had FFT (ELA) scores. ELA = English language arts; FFT = Framework for Teaching; VAM = value-added measure.
OCTObER 2017 385
confer, and consolidate). Students respond to 36 items on a 5-point Likert scale ranging from no, never to yes, always (ele- mentary) or from totally true to totally untrue (secondary). Fol- lowing the practices of the Tripod project and the MET Project, we constructed an overall measure of students’ assessments of their teachers’ instructional practices by assigning point values of 1 to 5 to Likert scale responses, reverse-coding items with nega- tive valence, averaging responses across the 36 items for each stu- dent, and averaging students’ overall scores to the teacher level (Surveyi; for further details, see Kane & Cantrell, 2010).
11
Teacher contributions to student achievement. VAM scores were created for the MET sample of teachers using student achieve- ment data from state-administered accountability exams. MET researchers estimated VAMs by grade and district for a single achievement outcome (ELA or math). Student achievement was modeled as a function of student background characteristics and prior-year achievement, in addition to average class background characteristics and prior-year achievement. Residuals from these models were then averaged to generate subject-specific teacher VAM scores (VAMis).
12 For subject-matter specialists teaching more than one section of the same subject, we created a weighted average VAM score, weighted by the number of students tested in each of the teacher’s sections.
Placing Measures on a Common Scale
We rescaled teachers’ VAM and Survey scores so that they share the same theoretical and continuous 4-point scale (i.e., 1 to 4) as the FFT (with corresponding 3-point range). As described above, this is a necessary step for applying weights, but the choice of a common scale does not affect the distribution of teacher proficiency.
A simple linear transformation allows us to rescale the Survey measure as follows:
Survey Survey Survey
Rescale TheoreticalRangei i=
* 3
+ −
1
3 Survey
Survey TheoreticalMin
TheoreticalRange*
.
(2)
In Equation (2), Surveyi is the overall score from student sur- veys for teacher i based on students’ responses to the 34-item Tripod survey. The value of SurveyTheoreticalRange equals 4 (between 1 and 5) and of SurveyTheoreticalMin equals 1, reflecting the mini- mum value on the 5-point Likert response scale. The first term on the right side of the equation rescales the range of all teacher survey scores to equal 3, and the second term shifts the score range upward so that the minimum value of the rescaled survey score (Surveyi
Rescale) equals 1. This approach preserves the relative position of teachers’ empirical survey scores within the full theo- retically possible range.
The approach taken in Equation (2) is not possible for teach- ers’ VAM scores because VAM is a relative measure with no true theoretical range or minimum value. Thus, we substitute empiri- cal analogues into Equation (2) as follows:
VAM VAM VAM
VA
Rescale ObservedRangeis is s
=
+ −
* 3
1 MM VAM
ObservedMin ObservedRanges s
* . 3
(3)
In Equation (3), VAMis is teacher i’s VAM score in subject s (math or ELA). The variable VAMs
ObservedRange is the observed range of VAM scores in subject s among all teachers in the sam- ple. The term VAMs
ObservedMin is the observed minimum value of VAM scores in subject s among all teachers in the sample. As in Equation (1), the first term on the right side of the equation rescales the range of all teacher VAM scores for subject s to equal 3, and the second term shifts the score range upward so that the minimum value of the rescaled VAM score for subject s (VAMis
Rescale) equals 1. We rescale math and ELA VAM scores separately, and for generalist teachers, we create an aggregate VAM score ( ),VAMi Aggregate
Rescale by averaging teacher i’s rescaled VAM math
( ),VAMi math Rescale
and ELA ( ),VAMi ELA Rescale
scores.13 Figure 2 shows the score distribution of the three teacher per-
formance measures. Teacher performance was judged to be better by students (based on survey reports) than by external evaluators’ observations of teacher practice or student achievement (see Table 3 and Figure 2, Panel A). Interestingly, the distribution of teacher performance based on unadjusted classroom observation scores (FFT) is similar to teacher performance based on student achieve- ment measures (VAM). However, after shifting the distribution of observation scores upward (FFTA) to more closely reflect how teachers may be rated in the context of newly implemented evalu- ation systems, we find that the distribution of observation scores is nearly identical to that of scores based on student survey responses (see Table 3 and Figure 2, Panel B).
Assigning Weights to Performance Measure Scores
We construct a summative evaluation score for each teacher i as a weighted average of the performance measure scores, as follows:
Score FFT
VAM
Aggregate
Aggregate Rescale
VAM
i j
i A
FFT j
i
W
W
=( ) +
,
,
*
* jj
i jW
( ) +( )SurveyRescale Survey* ,
(4)
where Scorei j is teacher i’s summative evaluation score based on
weighting scheme j, FFT is teacher i’s practice score based on class- room observations, VAM is the score that captures teacher i’s con- tribution to student achievement, and Survey is teacher i’s score based on students’ reports of teacher practice. Each of the three performance scores has an associated nominal weight (W) that corresponds to weighting scheme j. We use the adjusted FFT score (FFTA) in the calculation of all summative evaluation scores.
In practice, teacher evaluation systems may assign any feasi- ble set of nominal (i.e., policy) weights to each of the three per- formance measure scores, so long as they sum to 100%. The assignment of nominal weights to performance measures has
386 EDUCATIONAL RESEARCHER
been shown (using MET data) to yield statistically more reliable summative evaluation scores than empirically determined weights (e.g., optimal prediction weights, which are used to pre- dict student test scores); as a result, nominal weights both are better suited for high-stakes evaluation systems and better reflect the relative value that policymakers place on different measures of teacher performance (Martinez, Schweig, & Goldschmidt, 2016). We assign nominal weights to each of the three perfor- mance measure scores based on the following approach. First, we allow W jFFT to vary from 0% to 100% along an integer scale, such that W jFFT = [0,100]. Next, we construct the weight associated with a teacher’s contribution to student achievement as follows:
W Wj
j
VAM FFT
Survey VAMRatio =
− +( )
100
1 / , where RatioSurvey/VAM =
W
W
j
j Survey
VAM
,
or the ratio of the weights assigned to the student survey and VAM scores for weighting scheme j. Ratio allows for varia- tion in the value that an evaluation system places on stu- dents’ classroom experiences relative to student achievement as measures of teacher performance. From this, we construct the weight associated with student survey scores as follows:
W jSurvey = (100 – (W j FFT + W
j VAM)) .
Following this approach, we generate four summative perfor- mance scores for each teacher. For the first performance score (Scorei
j ,1), we set Ratio = 1/10. This value of RatioSurvey/VAM is
motivated by the fact that among the largest school districts with newly implemented teacher evaluation systems, the average weight assigned to VAM is approximately 10 times the average weight assigned to student survey scores for teachers teaching in tested grades/subjects (Steinberg & Donaldson, 2016).14 For example, if the entirety of a teacher’s summative evaluation score depends on classroom observations (i.e., W jFFT = 100), then zero
FIGURE 2. Distribution of teacher performance scores Data are from the 2009–2010 school year. Sample includes all teachers (N = 1,275). In Panels A and B, VAM is the VAM (aggregate) measure from Table 3; Survey is the Survey measure from Table 3. In Panel A, FFT is the FFT (aggregate) measure from Table 3. In Panel B, FFT (Adjusted) is the FFTA (aggregate) from Table 3. VAM = value-added measure; FFT = Framework for Teaching.
weight will be assigned to the VAM and student survey scores. If, however, none of a teacher’s summative evaluation score depends on classroom observations (i.e., W jFFT = 0), and the evaluation system assigns 10 times as much weight to VAM as it does to the student survey measure, then
W Wj
j
VAM FFT
Ratio =
− +( )
= −
+
100
1
100 0
1 1
10
= 90.9%, and W j Survey = 100 –
(W jFFT + W j VAM ) = 100 – (0 + 90.9) = 9.1%.
To allow for student surveys (and, by extension, students’ reports of teacher performance) to play a more prominent role in teachers’ summative evaluation scores (relative to student achievement), we construct a second performance score (Scorei
j ,2 )
by setting Ratio = 1/5. This value of RatioSurvey/VAM is motivated by evidence that the weight assigned to VAM is approximately 5 times the weight assigned to student survey scores, on average, across newly implemented systems (in the largest school dis- tricts) that give nonzero weight to student surveys and VAM (Steinberg & Donaldson, 2016). Further, in some evaluation systems, student surveys contribute even more to a teacher’s summative evaluation score. Indeed, in some districts, student surveys are assigned approximately half the weight that is assigned to teacher performance based on student achieve- ment.15 We therefore construct a third performance score (Scorei
j ,3 ) by setting Ratio = 1/2. Finally, many new evaluation
systems do not incorporate student surveys into teachers’ sum- mative evaluation scores.16 We construct a fourth performance score that is composed of only observation and VAM scores (Scorei
j ,4 ) by setting Ratio = 0. The incorporation of multiple
ratios for teacher performance measures into our analysis allows for greater insight into how the distribution of teacher ratings responds dynamically to the interaction between score construc- tion and the two key system design features: performance mea- sure weights and ratings thresholds.
OCTObER 2017 387
A performance measure’s contribution to a teacher’s summa- tive score depends not only on the weight system designers assign to the measure (i.e., nominal weight) but also on the underlying variance of the measure and its correlation with the other measures used to construct the summative score (i.e., effec- tive weight; Schochet, 2008). As Schochet (2008) notes, equal weight assigned to performance measures does not imply that each performance measure will contribute equally to the overall variance of a teacher’s summative score. Specifically, the effective weight for each performance measure will depend on its average correlation with the other performance measures; if the average correlations are similar across measures, then the effective and nominal weights should also be similar (Schochet, 2008). Simply put, measures with lower correlations with other performance measures will have lower effective weights than their nominal (i.e. assigned) weights suggest.
In our analytic sample of MET teachers, we observe that VAM is relatively weakly correlated with both the FFT score (.11) and the student survey score (.17). In contrast, the FFT score is more highly correlated with the student survey score (.41; see Table 4). For example, suppose equal weights are assigned to each of the three performance measures (i.e., 33.3% assigned to observation scores, VAM, and student survey scores). Based on the perfor- mance measures in the MET data used in this article, the effective weights will be as follows: 34.8% to observation scores, 29.2% to VAM, and 36.0% to student surveys (Schochet, 2008).17 This example illustrates that VAM, which has the lowest correlation with the other performance measures, will also have the lowest effective weight. In practical terms, the measure with the lowest effective weight will contribute the least to a teacher’s summative evaluation score when equal nominal weight is assigned to each teacher performance measure.
Examining the Sensitivity of Ratings to System Design
To examine the sensitivity of teachers’ evaluation ratings to eval- uation system design parameters, we conduct two sets of simulation- based analyses. For the first analysis, we examine how, under a fixed evaluation ratings system (i.e., ratings thresholds employed in one of the eight teacher evaluation systems), the distribution of teacher ratings may be sensitive to the underlying weights assigned to performance measures. On the basis of a given rat- ings system, we vary the weights assigned to the three perfor- mance measures and calculate the proportion of teachers who
would be rated proficient under each weighting scheme. Teachers are deemed proficient if the evaluation points that they earn are sufficient for them to receive one of the two highest ratings— Level 3 or Level 4—which are based on the fixed ratings thresh- olds of each district’s evaluation system. Teachers who achieve at least a Level 3 summative evaluation rating are deemed profi- cient in each of the eight districts included in our analysis (see Table 1).
For the second analysis, we examine how, under a fixed weighting scheme, the distribution of teacher ratings may be sensitive to different ratings thresholds schema found across our sample of eight district evaluation systems. We do so by calculat- ing the proportion of teachers who would be rated proficient when only the ratings thresholds vary. This analysis allows us to demonstrate the extent to which teachers who receive the same summative evaluation score (and, by extension, the same per- centage of total evaluation points available) may be rated differ- ently as a consequence of policy-determined ratings thresholds. These complementary analyses also allow us to examine how the properties of teacher evaluation measures influence teacher pro- ficiency rates under different system design parameters.
Results
Table 5 summarizes our primary results. These simulated find- ings do not (nor are they intended to) replicate the actual ratings distributions in the eight districts from which we draw our rat- ings thresholds. Indeed, the simulated proficiency rates reported in Table 5 are substantially lower for some districts than the actual proficiency rates found in new evaluation systems (Anderson, 2013; Kraft & Gilmour, 2017). This is likely due to a number of factors, including the following: the specific perfor- mance measures used by each district, the norms across districts about what constitutes proficient practice, the exclusion of other types of measures and observation domains, and the conse- quences and rewards attached to teacher ratings.
To examine how variation in teacher performance measure weights shape the distribution of teacher proficiency, we look within a given teacher evaluation system, allowing us to hold constant the performance ratings thresholds while varying the performance measure weights. First, we find that teacher profi- ciency rates change substantially as the weights assigned to teacher performance measures change. Looking down a given column, or evaluation system (within a panel of Table 5), we see how the proportion of proficient teachers differs under different component weight schemes. Take the rates of teacher proficiency based on the ratings thresholds of Fairfax County Public Schools’ evaluation system (see Table 5, Panel A, which is based on Score1). Under a component weight scheme where FFT
A receives zero weight (VAM contributes 90.9% and student survey con- tributes 9.1% to a teacher’s summative evaluation score), 45% of teachers in our sample would be rated proficient. If we change only the component weights—say, to 50% FFTA (and 45.5% VAM and 4.5% student survey)—then teacher proficiency increases to 85%, an increase of 40 percentage points.18
Our findings in Table 5 reveal two important facts with respect to the weight assigned to performance measures with higher mean values. First, the more weight assigned to measures
Table 4 Correlation Matrix of Teacher Performance
Measures
Variable FFTA
(Aggregate) VAM
(Aggregate) Survey
FFTA (aggregate) — VAM (aggregate) .11 — Survey .41 .17 —
Note. All correlations are statistically significant at the .001 level. There are 1,275 teachers in the sample. FFT = Framework for Teaching; VAM = value-added measure.
388 EDUCATIONAL RESEARCHER
Table 5 Simulated Teacher Proficiency Rates, by Performance Measure Weights and District Ratings Thresholds
FFT (%) VAM (%) Survey (%) Chicago System
Clark County System
Denver System
Fairfax County System
Gwinnett County System
Miami-Dade System
New York City System
Philadelphia System
Panel A: Score1 (Survey/VAM = 1/10) 0 90.9 9.1 .12 .15 .09 .45 .32 .02 .02 .47 10 81.8 8.2 .14 .17 .09 .53 .38 .02 .02 .54 20 72.7 7.3 .17 .21 .10 .62 .46 .02 .02 .64 30 63.6 6.4 .20 .26 .13 .72 .57 .02 .02 .73 40 54.5 5.5 .26 .34 .17 .80 .66 .02 .02 .81 50 45.5 4.5 .37 .45 .24 .85 .75 .03 .03 .86 60 36.4 3.6 .48 .56 .33 .89 .80 .04 .03 .90 70 27.3 2.7 .58 .64 .44 .91 .85 .07 .05 .91 80 18.2 1.8 .66 .70 .55 .93 .87 .14 .11 .92 90 9.1 0.9 .70 .76 .62 .94 .88 .22 .19 .93 100 0.0 0.0 .76 .79 .67 .93 .89 .31 .27 .94 Panel B: Score2 (Survey/VAM = 1/5) 0 83.3 16.7 .14 .17 .09 .52 .37 .03 .02 .54 10 75.0 15.0 .16 .20 .10 .61 .45 .03 .02 .62 20 66.7 13.3 .19 .24 .13 .69 .54 .02 .02 .71 30 58.3 11.7 .24 .31 .15 .77 .63 .03 .02 .79 40 50.0 10.0 .32 .40 .20 .83 .71 .03 .02 .84 50 41.7 8.3 .41 .50 .28 .88 .78 .04 .03 .89 60 33.3 6.7 .52 .59 .36 .90 .83 .05 .04 .91 70 25.0 5.0 .61 .67 .47 .91 .86 .08 .06 .92 80 16.7 3.3 .67 .72 .56 .93 .88 .15 .11 .93 90 8.3 1.7 .71 .76 .63 .93 .88 .23 .20 .94 100 0.0 0.0 .76 .79 .67 .93 .89 .31 .27 .94 Panel C: Score3 (Survey/VAM = 1/2) 0 66.7 33.3 .21 .26 .14 .71 .56 .03 .03 .73 10 60.0 30.0 .25 .31 .16 .78 .63 .03 .02 .80 20 53.3 26.7 .29 .38 .19 .84 .69 .03 .02 .85 30 46.7 23.3 .37 .47 .23 .87 .77 .03 .02 .88 40 40.0 20.0 .45 .53 .30 .90 .82 .04 .03 .91 50 33.3 16.7 .53 .61 .39 .92 .85 .05 .04 .92 60 26.7 13.3 .60 .67 .47 .93 .87 .07 .05 .93 70 20.0 10.0 .65 .71 .53 .93 .88 .11 .08 .94 80 13.3 6.7 .69 .74 .59 .94 .88 .19 .14 .94 90 6.7 3.3 .72 .77 .64 .94 .89 .25 .21 .94 100 0.0 0.0 .76 .79 .67 .93 .89 .31 .27 .94 Panel D: Score4 (Survey/VAM = 0) 0 100 0.0 .10 .13 .07 .37 .27 .02 .02 .38 10 90 0.0 .12 .15 .08 .44 .32 .02 .02 .47 20 80 0.0 .13 .17 .09 .54 .39 .02 .02 .55 30 70 0.0 .17 .22 .11 .64 .48 .02 .02 .66 40 60 0.0 .22 .28 .15 .73 .60 .02 .02 .75 50 50 0.0 .31 .39 .20 .82 .69 .03 .02 .83 60 40 0.0 .43 .51 .29 .87 .78 .04 .03 .88 70 30 0.0 .55 .61 .40 .90 .83 .06 .05 .91 80 20 0.0 .64 .69 .52 .92 .87 .12 .09 .92 90 10 0.0 .70 .75 .61 .93 .88 .21 .18 .93 100 0 0.0 .76 .79 .67 .93 .89 .31 .27 .94
Note. Each cell reports the proportion of teachers rated Level 3 or Level 4 for a given set of performance measure weights and based on the ratings thresholds of a given evaluation system. All scores are based on FFTA (aggregate). Results in each column (within a panel) include teacher data pooled across the six Measures of Effective Teaching study districts and are based on simulations of teachers’ performance scores and the ratings thresholds of district-specific teacher evaluation systems. Panel A presents the distribution of simulated teacher ratings based on a teacher’s summative performance score where the ratio of the Survey/VAM weights is set to 1/10 (i.e., Score1), Panel B presents the distribution of simulated teacher ratings based on a teacher’s summative performance score where the ratio of the Survey/VAM weights is set to 1/5 (i.e., Score2), Panel C presents the distribution of simulated teacher ratings based on a teacher’s summative performance score where the ratio of the Survey/VAM weights is set to 1/2 (i.e., Score3), and Panel D presents the distribution of simulated teacher ratings based on a teacher’s summative performance score where the ratio of the Survey/VAM weights is set to 0 (i.e., Score4). See Figure 1 and Table 1 for each district’s rating thresholds. There are 1,275 teachers in the sample. FFT = Framework for Teaching; VAM = value-added measure.
OCTObER 2017 389
with higher relative means, the greater the rate of teacher profi- ciency. This can be seen by looking within a given evaluation system (i.e., within a column of Table 5) as the weight for obser- vations scores (FFT) increases relative to VAM scores across all four Survey/VAM ratios (i.e., within a panel of Table 5). Second, when greater relative weight is assigned to measures with lower means—for example, by reducing the Survey/VAM ratio from 1/2 (in Panel C) to 0 (in Panel D)—assigning more weight to a third measure (FFT) with a higher mean value will produce larger incremental changes in teacher proficiency rates. Specifically, focusing on teacher proficiency rates based on Chicago’s system, when the Survey/VAM ratio is the greatest (at 1/2; see Panel C), increasing the weight assigned to FFT from 50% to 100% increases teacher proficiency rates by 23 percent- age points, from 53% to 76%. In contrast, when the Survey/ VAM ratio is the lowest (at 0; see Panel D), increasing the weight assigned to FFT from 50% to 100% has a much larger effect on the change in proficiency rates, increasing teacher proficiency this time by 45 percentage points, from 31% to 76%. This empirical fact bears out across each of the other seven systems with different ratings thresholds.
We further illustrate these results with a series of heat maps in Figure 3. For each of the eight evaluation systems, these figures illustrate how the distribution of teacher ratings changes as the weight assigned to the adjusted observation score (FFTA) increases from 0% to 100%. These figures clearly show how, in a single evaluation system with fixed ratings thresholds, the per- centage of teachers assigned to each rating category substantively changes across different weighting schemes.
Figure 3 also demonstrates how performance weights and rat- ings thresholds interact differently across evaluation systems to determine the distribution of teacher effectiveness (i.e., the pro- portion of teachers in each performance rating category). Evidence from Figure 3 reveals how changes to the distribution of teacher ratings depend on the specific rating threshold system with which a given set of performance measure weights is combined. Specifically, based on Miami’s and New York City’s evaluation sys- tems, teacher proficiency rates among our sample of teachers remain relatively constant until FFTA contributes (approximately) at least 70% of the weight to a teacher’s summative evaluation score, after which teacher proficiency increases at a relatively con- stant rate (see Figure 3, Panels F and G). Based on Denver’s evalu- ation system, teacher proficiency rates remain relatively constant until FFTA contributes (approximately) at least 20% of the weight to a teacher’s summative evaluation score (see Figure 3, Panel C). In contrast, teacher proficiency rates based on the ratings thresh- olds in evaluation systems located in Chicago, Clark County, Fairfax County, Gwinnett County, and Philadelphia, increase at a relatively constant rate as the FFT weight increases across the full range of the FFT weight distribution.
Second, we find that teacher proficiency rates change sub- stantially when, holding constant the performance measure weights, the same teachers are evaluated using different perfor- mance ratings thresholds. By looking across a given row, or weight scheme (within a panel of Table 5), we see how the pro- portion of proficient teachers differs across evaluation systems. Based on a weight scheme where FFTA contributes 50% to a teacher’s summative evaluation score, our simulated teacher
proficiency rates range from 3% and 4%, based on the ratings thresholds in Miami’s and New York City’s evaluation systems, respectively, to approximately 90%, based on the ratings thresh- olds in Fairfax County’s and Philadelphia’s systems (see Table 5, Panel B). Figure 4 illustrates these results graphically by captur- ing the full distribution of teacher ratings across the eight evalu- ation systems (and across the four score constructions) under a weight scheme where FFT contributes 50% to a teacher’s sum- mative evaluation score. Here we see that the proportion of teachers rated in all four categories, and in particular, the lowest two rating categories (i.e., Levels 1 and 2), vary substantially due to differences across districts’ ratings thresholds.
Third, we show that the relative weights teacher evaluation systems place on student information—student survey responses relative to student achievement exams—in the construction of a teacher’s summative performance score will have real conse- quences for the distribution of teacher proficiency rates. Table 6 summarizes the range of teacher proficiency rates, within and across teacher evaluation systems, for different constructions of a teacher’s performance score (fixing the weight assigned to obser- vation scores at 50%). For example, we find that lowering the Survey/VAM ratio from 1/2 to 0 reduces teacher proficiency rates by up to 22 percentage points, from 53% to 31% (based on the ratings thresholds in Chicago’s system) and from 61% to 39% (based on the ratings thresholds in Clark County’s system). These findings further reveal that teacher proficiency rates are lowest across all systems when norm-referenced teacher perfor- mance measures, such as VAM, are given greater relative weight than criterion-based measures, such as student surveys. This result is not surprising given that teachers’ VAM scores are, on average, lower than teacher scores based on student surveys (see Table 3 and Figure 2) and have lower correlations with the other teacher performance measures in the MET data (see Table 4).
Discussion
Recent policy reforms have spurred a major overhaul of teacher evaluation systems in the United States, highlighted by the incor- poration of multiple measures of teacher performance and the expansion of teacher ratings categories in an effort to better mea- sure and differentiate teacher effectiveness. The designs of these new systems also incorporate equally important choices that poli- cymakers have made concerning the weights assigned to multiple performance measures and the placement of teachers’ summative scores into discrete performance categories. Yet, little guidance has been available to inform policymakers about the consequences these design decisions may have on the distribution of teacher rat- ings and the proportion of teachers deemed proficient. The absence of empirically based guidance to inform these decisions is particularly notable given that teachers’ summative ratings are increasingly being used to make high-stakes personnel decisions.
We find not only that both the weighting schemes assigned to performance measures and the ratings thresholds set by evalua- tion systems play a critical role in determining teacher profi- ciency rates, but also that the properties of performance measures directly influence the distribution of teacher proficiency rates. First, if teacher performance measures are assigned the same weight, then measures with stronger correlations will effectively
390 EDUCATIONAL RESEARCHER
FIGURE 3. Distribution of teacher ratings, by performance measure weights and system ratings thresholds Each panel shows the distribution of teacher ratings (Levels 1–4) based on a given evaluation system’s performance rating thresholds and across different weighting schemes assigned to a teacher’s summative evaluation score with the Survey-to-VAM ratio of 1/5. Teacher summative evaluation scores based on weights assigned to observation score (FFTA [aggregate]), VAM score (VAM [aggregate]) and survey score (Survey; see Table 3). Sample includes all teachers (N = 1,275). The vertical line indicates the weight assigned to observation scores in each district’s evaluation system (see Table 1). VAM = value-added measure; FFT = Framework for Teaching.
OCTObER 2017 391
contribute more to a teacher’s summative rating than measures that are weakly correlated. Second, teacher proficiency rates are increasing in the weight assigned to teacher performance mea- sures with highest mean values.
Further, we show that teacher proficiency rates depend on the relative value a system places on student information—student surveys relative to student achievement—in the construction of a teacher’s summative performance score. We also demonstrate that teacher proficiency rates are much more sensitive to chang- ing teacher performance weights when the proficiency ratings threshold is located near the center of the distribution of actual ratings (e.g., Chicago and Clark County) than in systems where the proficiency threshold is located at the upper end of the score distribution (e.g., New York City and Miami).
These results provide new evidence to inform policymakers about how design decisions related to teacher performance measure weights and ratings thresholds affect the distribution of teacher
proficiency rates. Our analysis also provides empirical insights into how the properties of teacher evaluation measures shape the distri- bution of teacher proficiency rates. In doing so, our results point to the fact that variation in teacher proficiency rates can be predictable and quantifiable based on both the observed properties of teacher performance measures and system design decisions related to per- formance measure weights and ratings thresholds.
Implications for Policy
Our analyses illustrate several important findings that are par- ticularly salient for policymakers. First, differences in the relative weights assigned to norm-referenced versus criterion-based mea- sures of teacher performance can result in substantially different distributions of teacher performance ratings. Norm-referenced measures, such as VAM, are relative scores that are normalized within a given group of teachers; by construction, the mean of
FIGURE 4. Distribution of teacher ratings, by system ratings thresholds and teacher performance score construction (Survey/VAM weight ratios) Each panel shows the distribution of teacher ratings (Levels 1–4) based on a given evaluation system’s performance rating thresholds and one (of four) score constructions. A teacher’s summative evaluation score is based on weights assigned to observation score (FFTA [aggregate]), VAM score (VAM [aggregate]) and survey score (Survey; see Table 3). In Panel A, a teacher’s summative evaluation score (Score1) is based on RatioSurvey/VAM = 1/10 and the following performance measure weights: FFT = 50%, VAM = 45.5%, and Survey = 4.5%. In Panel B, a teacher’s summative evaluation score (Score2) is based on RatioSurvey/VAM = 1/5 and the following performance measure weights: FFT = 50%, VAM = 41.7%, and Survey = 8.3%. In Panel C, a teacher’s summative evaluation score (Score3) is based on RatioSurvey/VAM = 1/2 and the following performance measure weights: FFT = 50%, VAM = 33.3%, and Survey = 16.7%. In Panel D, a teacher’s summative evaluation score (Score4) is based on RatioSurvey/VAM = 0 and the following performance measure weights: FFT = 50% and VAM = 50%. Sample includes all teachers (N = 1,275).
392 EDUCATIONAL RESEARCHER
VAM scores will be centered at the middle of the score distribu- tion. In contrast, the mean of criterion-based measures, such as observation and survey scores, can be located anywhere along the score distribution. In practice, criterion-based measures used in teacher evaluation systems are often centered well above the middle of the score range (see Figure 2, Panel B). Therefore, if scores on criterion-based measures are systematically skewed toward the upper end of the score distribution, then the relative weights assigned to norm-referenced versus criterion-based per- formance measures will have real consequences for the distribu- tion of teachers’ summative evaluation ratings. Indeed, we show that teacher proficiency rates reach a minimum when VAM scores receive the greatest nominal weight (see Tables 5 and 6). Moreover, we also show that performance measures with lower average correlations with the other performance measures, such as VAM scores, will contribute less to teachers’ summative evalu- ation scores—they will have lower effective weights—than their nominal weights would suggest.
These findings suggest that summative evaluation ratings for the nearly 70% of teachers in nontested grades/subjects (Watson et al., 2009) are likely to differ in systematic ways from the rat- ings of teachers for whom VAM scores can be calculated. We remind readers that VAM scores were available for all MET study teachers included in our simulation-based analyses. In the MET data, we observe that VAM has a substantially lower mean and greater variance than both observation scores (based on the adjusted FFT ratings of teacher performance) and student sur- vey scores (see Table 3). Further evidence of this pattern is found, for example, in Tennessee’s teacher evaluation system (Tennessee Department of Education, 2015). Because greater weight is con- sistently assigned to observation scores for teachers in nontested grades and subjects, we would expect these teachers’ evaluation ratings to be systematically higher than those for teachers with VAM scores. Data provided to us from a large (anonymous) urban school district in the Midwest bore out this prediction. In the absence of student achievement on state accountability exams, replacing VAMs with measures based on locally devel- oped tests of student progress (i.e., SLOs) is unlikely to resolve
this ratings disparity. SLOs are criterion-based measures of teacher performance that are often systematically and upwardly skewed relative to VAM scores. This also helps to further explain why our simulated teacher proficiency rates understate the share of teachers who, in a typical school district, are deemed to be proficient.
Second, states’ and districts’ decisions about where to place the summative ratings thresholds have direct implications for the distribution of teacher effectiveness. In systems that apply abso- lute thresholds—as do the eight evaluation systems included in our analysis—all teachers, in principle, may be rated proficient. In contrast, a system that imposes a target distribution of teacher effectiveness, such as the one used by the Dallas Independent School District, defines the proportion of summative evaluation scores (and, by extension, the proportion of teachers) assigned to a summative evaluation rating.19 Despite the important differ- ences in teacher performance measures used across districts and the weights assigned to these measures, there is no clear justifica- tion for the variation in ratings thresholds that districts set for teachers to meet proficiency standards (as shown in Figure 1). As a result, these differences in ratings thresholds complicate com- parisons of teacher effectiveness across districts in the same way that state-specific differences in performance standards limit the comparability of student proficiency rates across states.
Further, both the meaning attached to and the consequences of a teacher’s summative rating also differ across districts. For example, a Level 2 rating of “developing” (as in Chicago, Miami, and New York City) may imply that teachers are making prog- ress toward proficiency, whereas a Level 2 rating of “minimally effective” (as in Clark County) may imply that they are not. As a result, the district-specific distribution of teacher proficiency will depend, in part, on how different districts assign different meaning and, ultimately, different high-stakes consequences to teacher ratings. Our analysis provides insight into how a district’s design decisions may influence the desired distribution of teacher ratings.
Third, the design and consequences of evaluation systems may be shaped in important ways by political considerations as
Table 6 Simulated Teacher Proficiency Rates, by Teacher Performance Score
Construction (Survey/VAM Weight Ratios)
Teacher Performance Score (Survey/VAM Weight Ratio)
Chicago System
Clark County System
Denver System
Fairfax County System
Gwinnett County System
Miami-Dade System
New York City System
Philadelphia System
Score1 (Survey/VAM =1/10) .37 .45 .24 .85 .75 .03 .03 .86 Score2 (Survey/VAM = 1/5) .41 .50 .28 .88 .78 .04 .03 .89 Score3 (Survey/VAM = 1/2) .53 .61 .39 .92 .85 .05 .04 .92 Score4 (Survey/VAM = 0) .31 .39 .20 .82 .69 .03 .02 .83 Minimum .31 .39 .20 .82 .69 .03 .02 .83 Maximum .53 .61 .39 .92 .85 .05 .04 .92 Range .22 .22 .19 .10 .16 .02 .02 .09
Note. Each cell reports the proportion of teachers rated Level 3 or Level 4 for a given construction of teacher’s summative performance scores. Results are based on weights that fix FFTA (aggregate) at 50% and allow Survey and VAM weights to vary. Score1 is based on a Survey/VAM weights ratio of 1/10, Score2 is based on a Survey/ VAM weights ratio of 1/5, Score3 is based on a Survey/VAM weights ratio of 1/2, and Score4 is based on a Survey/VAM weights ratio of 0. There are 1,275 teachers in the sample. VAM = value-added measure; FFT = Framework for Teaching.
OCTObER 2017 393
well as local implementation practices. Although our simulation- based results reveal that the same teachers may be assigned to different performance categories depending on district-specific ratings thresholds, teacher proficiency rates do not differ in prac- tice nearly as much. This is because system designers may start by selecting ratings thresholds based on objective criteria but then, upon inspection of the distribution of teacher proficiency they generate, revise the ratings thresholds to produce a distribu- tion of teacher proficiency that is both professionally as well as politically feasible. Further, school-based evaluators can respond to a system established by state and district administrators by adjusting scores on subjective teacher performance measures (such as observation scores) in order to produce a desired ratings distribution. Such policy and practice decisions may be made to avoid negative externalities (e.g., teacher exits among high-per- forming, lower-rated teachers and lower staff moral) that can result when teacher ratings are not consistent with perceptions of effectiveness among educators, even in systems designed to gen- erate greater variation in teachers’ summative ratings than what existed under traditional evaluation systems.
Finally, our work calls for greater attention to and transpar- ency around the policymaking process for selecting perfor- mance measure weights and determining ratings thresholds. The empirical findings of this article reveal the sensitivity of teacher ratings to these design features of newly implemented evaluation systems. These findings demonstrate why it is impor- tant that policymakers both understand the consequences of their design decisions and more clearly communicate how these decisions are made. Such considerations are critical in light of policy efforts to improve teacher quality and, ultimately, stu- dent achievement.
NOTES
The authors thank Allison Atteberry, Cory Koedel, and Eric Taylor for feedback on previous versions of this article and participants at the 2016 Association for Education Finance and Policy conference for valu- able comments and discussions. The authors thank Filippo Bulgarelli and Mariela Mannion for excellent research assistance and Jennifer Moore for editorial assistance.
1Under the district’s traditional evaluation system, nearly all Chicago Public Schools teachers (93%) were rated proficient (i.e., “superior” or “excellent”), whereas only 34% of Chicago Public Schools met state proficiency standards (Steinberg & Sartain, 2015).
2Specifically, 83% of states with newly implemented evaluation systems link teacher summative ratings to required professional devel- opment for low-rated teachers; among the largest districts and the District of Columbia, 74% require professional development for low- rated teachers. Further, 61% of newly implemented state evaluation systems (and 39% of newly implemented systems in the largest school districts and the District of Columbia) tie teacher ratings to employ- ment termination, whereas 48% of new state systems and 22% of new district systems tie teacher ratings to tenure granting/revocation deci- sions (Steinberg & Donaldson, 2016).
3Approximately 61% (14) of the largest districts and 30% (14) of states use value-added measures (VAMs), whereas 35% (eight) of the largest districts and 59% (27) of states use student growth percentiles (Steinberg & Donaldson, 2016).
4In practice, many states and districts construct a final teacher practice score by averaging across protocol components (within a given
observation) and then averaging across multiple observations. Indeed, prior research suggests that a simple average is an appropriate approach for aggregating teacher practice scores across multiple classroom obser- vations (Garrett & Steinberg, 2015; Kane, McCaffrey, Miller, & Staiger, 2013; Mihaly, McCaffrey, Staiger, & Lockwood, 2013). Alternatively, some evaluation systems weight observation components differently (e.g., Chicago Public Schools’ REACH system gives greater weight to observation components related to aspects of a teacher’s instructional performance [such as engaging students in learning] than to aspects of a teacher’s practice related to managing the classroom environment [such as managing student behavior]).
5Although the statistical models used to produce teacher VAM scores differ, in general, such models provide an estimate of a teach- er’s contribution to student learning by controlling for prior student achievement and other observable student and classroom characteris- tics. We further note that the necessity of a student’s baseline test score to construct VAM often precludes teachers in early elementary grades (i.e., Grades K–3) from receiving VAM scores. Student growth percen- tiles provide an estimate of a teacher’s contribution to student learning by comparing student achievement growth to students’ peers with simi- lar prior test score histories. Student learning objectives are subject- and grade-specific learning goals, which tend to be based on locally selected (i.e., school and/or district) measures of student achievement and are used to estimate a teacher’s contribution to student learning in grades/ subjects that are not tested via the state’s accountability exam (Steinberg & Donaldson, 2016).
6Some districts allow evaluators to use their professional judgment to assign a summative rating based on their own synthesis of multiple performance measures available for each teacher (e.g., Boston Public Schools).
7The six participating districts were Charlotte-Mecklenburg Schools (North Carolina), Dallas Independent School District (Texas), Denver Public Schools (Colorado), Hillsborough County Public Schools (Florida), Memphis City Schools (Tennessee), and the New York City Department of Education (New York).
8Classroom generalist teachers were videoed, on average, on 4 separate days throughout the year, with each day producing one English language arts (ELA) and one math lesson video. Subject-specific (i.e., departmentalized) teachers were videoed on 2 separate days, capturing two different sections of the same subject taught by the teacher.
9Aggregate practice scores for departmentalized teachers’ (i.e., teachers teaching multiple sections of the same subject) will equal their subject-specific practice scores.
10We also pursued measurement-based approaches—empirical Bayes and principal components analysis—as alternative ways of construct- ing teacher performance scores using classroom observation scores. The scores generated by these measurement-based approaches are all very highly correlated with scores constructed by averaging across observa- tion components and domains.
11We also pursued alternative approaches to constructing teacher performance scores based on students’ perceptions. These include (a) averaging within the “7C” domains first and then averaging across these domains so that each domain is weighted equally, (b) averaging individual items across students for each teacher and then averaging across items, and (c) a measurement-based approach where items are weighted based on the loadings from the first eigenvector of a principal components analysis. All three alternative approaches produce perfor- mance scores that are correlated at 0.98 or above. Further sensitivity analyses, including dropping students who “straight line” (i.e., fill in the same answer for every single item—less than 0.0001% of students) or students whose rating of a teacher is beyond 2 standard deviations (approximately 3% of students) from the class mean, do not change a teacher’s performance score.
394 EDUCATIONAL RESEARCHER
12For more details on the construction of teacher VAM scores, see White and Rowan (2012).
13For departmentalized teachers (i.e., teachers teaching multiple sections of the same subject), their subject-specific VAM scores equal their aggregate VAM scores. Averaging across ELA and math VAM scores for a generalist teacher is a practice used in many school districts (e.g., Chicago Public Schools).
14For the typical teacher teaching in a tested grade/subject, 21.7% and 1.9% of a teacher’s summative evaluation score is based on VAM and student survey scores, respectively (Steinberg & Donaldson, 2016).
15For example, for most Grade 3-to-12 teachers in the Dallas Independent School District, student surveys (based on students’ class- room experiences) and student-achievement based measures account for 15% and 35%, respectively, of a teacher’s summative evaluation score (Dallas Independent School District, n.d.).
16Some of the largest school districts, including New York City and Chicago Public Schools, do not incorporate student surveys into teachers’ summative evaluation scores (see Table 1).
17Following Schochet (2008), the effective weight—the contribu- tion of performance measure j to the variance of the composite teacher evaluation score—may be calculated as
w w w wj j j k
N
j k jk effective nominal nominal nominal= ( ) +
≠ ∑
2 ρ ,
where nominal indicates the nominal weight assigned to performance measure j and ρjk is the correlation between performance measure j and performance measure k (e.g., the correlation between the Framework for Teaching [FFT] score and the VAM score).
18Under an evaluation system with very different ratings thresh- olds, changing the underlying component weights may have little conse- quence for teacher proficiency. That is, as we vary the component weight scheme, teacher proficiency changes differently depending on where a given evaluation system sets its performance ratings thresholds. Contrast the results based on Fairfax County’s system with New York City’s system (see Figure 1 and Table 1 for the contrast in ratings thresholds across districts). Under a component weight scheme where FFTA receives zero weight, 2% of teachers in our sample would be rated proficient based on New York City’s ratings thresholds. If we change only the component weights to 50% FFTA (and 45.5% VAM and 4.5% Survey), teacher pro- ficiency among our sample of teachers would increase to just 3%, based on the ratings thresholds under New York City’s evaluation system.
19For the 2014–2015 and 2015–2016 school years, Dallas Independent School District’s (2015, 2016) Teacher Excellence Initiative determined teachers’ summative evaluation ratings by arranging scores to follow a target distribution, as follows: 3% of teachers were rated unsatisfactory; 37% progressing; 58% proficient; and 2% exemplary.
REFERENCES
Almy, S. (2011). Fair to everyone: Building the balanced teacher eval- uations that educators and students deserve. Washington, DC: Education Trust.
Anderson, J. (2013, March 30). Curious grade for teachers: Nearly all pass. The New York Times. Retrieved from http://www.nytimes .com/2013/03/31/education/curious-grade-for-teachers-nearly- all-pass.html
Chetty, R., Friedman, J., & Rockoff, J. (2014). Measuring the impacts of teachers I: Evaluating bias in teacher value-added estimates. American Economic Review, 104(9), 2593–2632.
Cullen, J. B., Koedel, C., & Parsons, E. (2016). The compositional effect of rigorous teacher evaluation on workforce quality (No. w22805). Cambridge, MA: National Bureau of Economic Research.
Dallas Independent School District. (2015). Rules and procedures for calculating achievement statistics, evaluation scores, and effectiveness levels for Dallas ISD’s Teacher Excellence Initiative. Retrieved June 19, 2015, from http://www.dallasisd.org/Page/28269
Dallas Independent School District. (2016). Resources. Retrieved March 10, 2016, from http://tei.dallasisd.org/home-2/resources/
Dallas Independent School District. (n.d.). Defining excellence. Retrieved from http://tei.dallasisd.org/home-2/defining-excellence/
Danielson, C. (1996). Enhancing professional practice: A framework for teaching. Alexandria, VA: Association for Supervision and Curriculum Development.
Darling-Hammond, L. (2013). Getting teacher evaluation right: What really matters for effectiveness and improvement. New York, NY: Teachers College Press.
Dee, T., & Wyckoff, J. (2015). Incentives, selection and teacher per- formance: Evidence from IMPACT. Journal of Policy Analysis and Management, 34(2), 267–297.
Donaldson, M. L. (2009). So long, Lake Wobegon? Using teacher evaluation to raise teacher quality. Washington, DC: Center for American Progress. Retrieved from https://cdn.american- progress.org/wp-content/uploads/issues/2009/06/pdf/teacher_ evaluation.pdf
Donaldson, M. L., & Papay, J. P. (2015). Teacher evaluation for accountability and development. In H. F. Ladd & M. E. Goertz (Eds.), Handbook of research in education finance and policy (pp. 174–193). New York, NY: Routledge.
Elmore, R. F. (2002). Unwarranted intrusion. Education Next, 2(1). Garrett, R., & Steinberg, M. P. (2015). Examining teacher effectiveness
using classroom observation scores: Evidence from the random- ization of teachers to students. Educational Evaluation and Policy Analysis, 37(2), 224–242.
Grissom, J. A., & Youngs, P. (Eds.). (2015). Improving teacher evalua- tion systems: Making the most of multiple measures. New York, NY: Teachers College Press.
Herlihy, C., Karger, E., Pollard, C., Hill, H. C., Kraft, M. A., Williams, M., & Howard, S. (2014). State and local efforts to investigate the validity and reliability of scores from teacher evaluation systems. Teachers College Record, 116(1), 1–28.
Hill, H. C., Charalambous, C. Y., & Kraft, M. A. (2012). When rater reliability is not enough teacher observation systems and a case for the generalizability study. Educational Researcher, 41(2), 56–64.
Jackson, C., & Steinberg, M. P. (2017). Does teacher effectiveness depend on who rates classroom practice? Evidence from an urban teacher prep- aration program. Working paper.
Jiang, J. Y., & Sporte, S. (2016). Teacher evaluation in Chicago: Differences in observation and value-added scores by teacher, stu- dent, and school characteristics. Chicago, IL: University of Chicago Consortium on School Research.
Kane, T., & Cantrell, S. (2010). Learning about teaching: Initial findings from the measures of effective teaching project. MET Project research paper. Seattle, WA: Bill & Melinda Gates Foundation.
Kane, T., Kerr, K., & Pianta, R. (2014). Designing teacher evalua- tion systems: New guidance from the Measures of Effective Teaching Project. New York, NY: Wiley.
Kane, T., McCaffrey, D., Miller, T., & Staiger, D. (2013). Have we identified effective teachers? Validating measures of effective teach- ing using random assignment. MET Project research paper. Seattle, WA: Bill & Melinda Gates Foundation.
Kane, T. J., & Staiger, D. O. (2012). Gathering feedback for teach- ing: Combining high-quality observations with student surveys and achievement gains. MET Project research paper. Seattle, WA: Bill & Melinda Gates Foundation.
OCTObER 2017 395
Kraft, M. A., & Gilmour, A. (2017). Revisiting the widget effect: Teacher evaluation reforms and distribution of teacher effective- ness ratings. Educational Researcher, 46(5), 234–249.
Le Floch, K. C., Boyle, A., & Therriault, S. B. (2008). Help wanted: State capacity for school improvement. AIR research brief. Washington, DC: American Institutes for Research.
Lipscomb, S., Terziev, J., & Chaplin, D. (2015). Measuring teach- ers’ effectiveness: A report from Phase 3 of Pennsylvania’s pilot of the Framework for Teaching. Princeton, NJ: Mathematica Policy Research.
Martinez, J. F., Schweig, J., & Goldschmidt, P. (2016). Approaches for combining multiple measures of teacher performance: Reliability, validity, and implications for evaluation policy. Educational Evaluation and Policy Analysis, 38(4), 738–756.
Marzano, R. J., & Toth, M. D. (2013). Teacher evaluation that makes a difference: A new model for teacher growth and student achievement. Alexandria, VA: ASCD.
McGuinn, P. (2012). Stimulating reform: Race to the Top, competi- tive grants, and the Obama education agenda. Educational Policy, 16(1), 136–159.
Mihaly, K., McCaffrey, D. F., Staiger, D., & Lockwood, J. R. (2013). A composite estimator of effective teaching. MET Project technical report. Seattle, WA: Bill & Melinda Gates Foundation.
National Center on Teacher Quality. (n.d.). NCTQ district policy. Washington, DC: Author. Retrieved from http://www.nctq.org/ districtPolicy/contractDatabase/customReport.do#criteria
Sartain, L., Stoelinga, S. R., & Brown, E. R. (2011). Rethinking teacher evaluation: Lessons learned from observations, principal-teacher con- ferences, and district implementation. Chicago, IL: Consortium on Chicago School Research.
Sartain, L., & Steinberg, M. P. (2016). Teachers’ labor market responses to performance evaluation reform: Experimental evi- dence from Chicago public schools. Journal of Human Resources, 51(3), 615–655.
Schochet, P. Z. (2008). Technical methods report: Guidelines for multiple testing in impact evaluations (NCEE 2008-4018). Washington, DC: National Center for Education Evaluation and Regional Assistance, Institute of Education Sciences, U.S. Department of Education.
Sinnema, C., & Robinson, V. (2007): The leadership of teaching and learning: Implications for teacher evaluation. Leadership and Policy in Schools, 6(4), 319–343.
Steinberg, M. P., & Donaldson, M. (2016). The new educational accountability: Understanding the landscape of teacher evaluation in the post-NCLB era. Education Finance and Policy, 11(3), 340–359.
Steinberg, M. P., & Jiang, J. (2016). Rater bias or teacher sorting? Examining the causes and consequences of racial gaps in teacher per- formance ratings. Working paper.
Steinberg, M. P., & Sartain, L. (2015). Does teacher evaluation improve school performance? Experimental evidence from Chicago’s Excellence in Teaching Project. Education Finance and Policy, 10(4), 535–572.
Stronge, J. H., & Tucker, P. D. (2003). Teacher evaluation. Assessing and improving performance. Larchmont, NY: Eye on Education.
Tennessee Department of Education. (2015). Teacher and adminis- trator evaluation in Tennessee: A report on Year 3 implementation. Retrieved from http://team-tn.org/wp-content/uploads/2013/08/ rpt_teacher_evaluation_year_31.pdf
Toch, T., & Rothman, R. (2008). Rush to judgment: Teacher evaluation in public education. Washington, DC: Education Sector.
Tucker, P. D. (1997). Lake Wobegon: Where all teachers are compe- tent (or, have we come to terms with the problem of incompe- tent teachers?). Journal of Personnel Evaluation in Education, 11, 103–126.
Wallace, T. L., Kelcey, B., & Ruzek, E. (2016). What can student perception surveys tell us about teaching? Empirically testing the underlying structure of the Tripod student perception survey. American Educational Research Journal, 53(6), 1834–1868.
Watson, J. G., Kraemer, S. B., & Thorn, C. A. (2009). The other 69 percent. Washington, DC: Center for Educator Compensation Reform, U.S. Department of Education, Office of Elementary and Secondary Education.
Weisberg, D., Sexton, S., Mulhern, J., Keeling, D., Schunck, J., Palcisco, A., & Morgan, K. (2009). The widget effect: Our national failure to acknowledge and act on differences in teacher effectiveness. Brooklyn, NY: New Teacher Project.
White, M., & Rowan, B. (2012). A user guide to the “core study” data files available to MET early career grantees. Ann Arbor: Inter-University Consortium for Political and Social Research, University of Michigan.
AuThORS
MATTHEW P. STEINBERG, PhD, is an assistant professor of edu- cation policy at the University of Pennsylvania Graduate School of Education, 3700 Walnut Street, Philadelphia, PA 19104; steima@ upenn.edu. His research focuses on teacher evaluation and human capital, school discipline and safety, urban school reform, and school finance.
MATTHEW A. KRAFT, EdD, is an assistant professor of education and economics at Brown University, P.O. Box 1938, Providence, RI 02912; [email protected]. His research focuses on efforts to improve educator and organizational effectiveness in K–12 urban public schools.
Manuscript received April 25, 2016 Revisions received November 30, 2016,
March 30, 2017, and July 14, 2017 Accepted July 20, 2017
396 EDUCATIONAL RESEARCHER
Appendix
Table 1 Data Sources
Chicago, IL Chicago Public Schools. (2014). REACH students: Educator evalu-
ation handbook 2014–2015 (p. 61). Retrieved from http://www .ctunet.com/rights-at-work/teacher-evaluation/text/CPS- REACH-Educator-Evaluation-Handbook-FINAL.pdf
Clark County, NV Nevada State Board of Education. (2015). NEPF Educator Performance
Framework (NEPF): Statewide evaluation system (p. 17). Retrieved from http://www.doe.nv.gov/Educator_Effectiveness/Educator_ Develop_Support/NEPF/Tools_and_Protocols/
After following the link, you will be directed to a page on the Nevada Department of Education website that gives an overview of the Nevada Educator Performance Framework (NEPF). To access the NEPF document, click on the hyperlink that says “Protocols” under the headline “NEPF Protocols” to download the document.
Denver, CO Denver Public Schools. (n.d.). LEAP handbook 2014–2015 (p. 5).
Retrieved from http://www.nctq.org/docs/denver_2014-15-LEAP- handbook-master_1.pdf
Fairfax County, VA Fairfax County Public Schools. (2015). Teacher Performance Evaluation
Program handbook (p. 16). Retrieved from http://www.nctq.org/ docs/TEHandbook.pdf
Gwinnett County, GA Gwinnett County Public Schools. (n.d.). A primer for teachers 2015–
2016 (p. 6). Retrieved from http://www.nctq.org/docs/2015-16- GTES-Primer_FINAL_June25.pdf
Miami-Dade, FL Miami-Dade County Public Schools. (2015). IPEGS procedural hand-
book 2015 edition (p. 92). Retrieved from http://ipegs.dadeschools .net/pdfs/2015_IPEGS_Procedural_Handbook.pdf
New York City, NY New York City Department of Education. (2014). Advance overall rat-
ings guide (p. 12). Retrieved from http://www.uft.org/files/attach ments/advance-ratings-guide-2013-14.pdf
Philadelphia, PA Pennsylvania Department of Education. (2014). Educator effectiveness
administrative manual (p. 19). Retrieved from http://www.nctq .org/docs/Educator_Effectiveness_Administrative_Manual.pdf