help with 401 article review
Learning Disabilities Research
Learning Disabilities Research & Practice, 25(2), 60–75 C© 2010 The Division for Learning Disabilities of the Council for Exceptional Children
Creating a Progress-Monitoring System in Reading for Middle-School Students: Tracking Progress Toward Meeting High-Stakes Standards
Christine Espin, Teri Wallace, Erica Lembke, Heather Campbell, and Jeffrey D. Long University of Minnesota
In this study, we examined the reliability and validity of curriculum-based measures (CBM) in reading for indexing the performance of secondary-school students. Participants were 236 eighth-grade students (134 females and 102 males) in the classrooms of 17 English teachers. Students completed 1-, 2-, and 3-minute reading aloud and 2-, 3-, and 4-minute maze selection tasks. The relation between performance on the CBMs and the state reading test were examined. Results revealed that both reading aloud and maze selection were reliable and valid predictors of performance on the state standards tests, with validity coefficients above .70. An exploratory follow-up study was conducted in which the growth curves produced by the reading-aloud and maze-selection measures were compared for a subset of 31 students from the original study. For these 31 students, maze selection reflected change over time whereas reading aloud did not. This pattern of results was found for both lower- and higher-performing students. Results suggest that it is important to consider both performance and progress when examining the technical adequacy of CBMs. Implications for the use of measures with secondary-level students for progress monitoring are discussed.
In recent years, much attention has been directed to early intervention and prevention in reading. An alternative to a singular focus on early intervention is an approach in which early intervention is combined with continuous, long-term, intensive interventions for struggling readers. “Long term” in this approach refers to reading instruction that extends into the high school years. The goal of such an approach would be to diminish the magnitude of reading difficulties experienced by struggling readers and increase the likelihood of postgraduation success. Supporting the notion that long- term, intensive reading interventions may be needed for a select group of students are two sources of data: (1) results of early intervention studies and (2) results of secondary- school studies for students with learning disabilities.
Need for Long-Term, Intensive Intervention Efforts
Recent research on the effects of early identification and intervention programs have produced promising outcomes and demonstrated reductions in the magnitude and preva- lence of reading failure (O’Connor, Fulmer, Harty, & Bell, 2005; O’Connor, Harty, & Fulmer, 2005; Vaughn, Linan- Thompson, & Hickman, 2003). However, these studies also
Requests for reprints should be sent to Christine Espin, Wassenaarseweg 52, PO Box 9555, 2300 RB Leiden, The Netherlands. Electronic inquiries should be sent to [email protected].
have uncovered a small group of children who “fail to thrive” (Vaughn et al., 2003), even when given intensive and poten- tially powerful interventions. Such children either do not reach a level of performance that warrants placement into a typical instructional setting or do not maintain satisfactory levels of performance without continued intensive interven- tions. These students have reading difficulties that seem to be especially resistant to change (see Torgesen, 2000) and are often considered to have learning disabilities (LD).
Research at the secondary-school level reveals that stu- dents with LD continue to experience reading difficulties well into their high school years. Secondary-school students with LD experience difficulties with phonological, language comprehension, and reading fluency skills (Fuchs, Fuchs, Mathes, & Lipsey, 2000; Vellutino, Fletcher, Snowling, & Scanlon, 2004; Vellutino, Scanlon, & Tanzman, 1994; Vel- lutino, Tunmer, Jaccard, & Chen, 2007). They typically per- form at levels 4–6 years behind non-LD peers in reading and score in the lowest decile on reading achievement tests (Deshler, Schumaker, Alley, Warner, & Clark, 1982; Levin, Zigmond, & Birch, 1985; Warner, Schumaker, Alley, & Desh- ler, 1980). For example, on the 2007 National Assessment of Educational Progress (Lee, Grigg, & Donahue, 2007), 66 percent of students with disabilities in public schools scored below a Basic Level, compared to only 24 percent of students without disabilities. (A Basic Level implies partial mastery of the knowledge and skills needed for proficient work at a given grade level.)
LEARNING DISABILITIES RESEARCH 61
Taken together, research on younger and older children with reading difficulties produces a picture of students whose reading difficulties begin early and persist throughout their school career. For such students a program of intervention that begins early—and then continues throughout their school careers—is needed.
Reading Interventions at the Secondary-School Level
Two questions arise when considering reading interventions for secondary-school students with LD. The first is: At what level do students need to read to be successful after high school graduation? In recent years, this question often has been addressed through the development of state standards tests in reading. Such tests define, by design or default, the level of reading considered to be necessary for students to be successful at the secondary-school level—this despite the fact that the extent to which many state tests reflect the type of reading necessary for success either in school or in postsec- ondary settings is unknown. However, given the high-stakes nature of state tests for schools in terms of meeting No Child Left Behind standards, and for students who are required to pass reading tests to graduate (as is the case in 23 states; Center on Education Policy, 2008), the tests are an important outcome for students and schools at the secondary-school level.
The second question is: How can we determine whether our reading interventions are effective? The reading progress of secondary-school students with LD might prove to be slow and incremental—but not necessarily unimportant. For example, improvement of even one grade level (to use a typical metric) in reading over the course of 4 years in high school might translate into large advantages in post–high school settings. Yet are there instruments that are sensitive to such slow and incremental growth? Are those instruments reliable and valid, and can they be tied to success on tasks of importance, such as performance on state reading tests or performance in postsecondary educational settings? One instrument that might potentially fulfill these requirements is curriculum-based measurement (CBM).
CBM
CBM is a system of measurement designed to allow teachers to monitor student progress and evaluate the effectiveness of instructional programs (Deno, 1985). The success of CBM relies on two key characteristics: practicality and technical adequacy (Deno, 1985). With respect to practicality, if the measures are to be given on a frequent basis, they must be time efficient and easy to develop, administer, and score and must allow for the creation of multiple equivalent forms. With respect to technical adequacy, if the measures are to provide educationally useful information, they must be valid and re- liable indicators of performance in an academic area. For a measure to be considered a valid indicator of performance, evidence must demonstrate that performance on the measure relates to performance in the academic domain more broadly.
In reading, the number of words read correctly in 1 minute is often used as a CBM indicator of general reading perfor- mance at the elementary-school level (Wayman, Wallace, Wiley, Ticha, & Espin, 2007). One-minute reading-aloud measures are time efficient and easy to develop, adminis- ter, and score, and they allow for the creation of multiple equivalent forms. Further, a large body of research supports the relation between the number of words read aloud in 1 minute and other measures of reading proficiency, including reading comprehension (see reviews by Marston, 1989; Way- man et al., 2007). Although most CBM reading research has focused on a reading-aloud measure, support also has been found for the technical adequacy of a maze-selection mea- sure (see Wayman et al., 2007). In a maze-selection measure, every seventh word of a passage is deleted and replaced with a multiple-choice item consisting of the correct word plus two distracters. Students read through the text and choose the correct word for each multiple-choice item. Specific to the present study, both reading-aloud (Crawford, Tindal, & Stieber, 2001; Hintze & Silberglitt, 2005; McGlinchey & Hixson, 2004; Silberglitt & Hintze, 2005; Stage & Jacobsen, 2001) and maze-selection measures (Wiley & Deno, 2005) have been shown to predict performance on state standards tests.
Although research supports the technical adequacy of both reading aloud and maze selection, the majority of that research has been done at the elementary-school level (Wayman et al., 2007). Far less research has been conducted in reading at the secondary-school level, even though the re- sults of cross-age studies suggest that the nature and type of CBM in reading might need to change as students become older and more proficient readers (Jenkins & Jewell, 1993; MacMillan, 2000; Yovanoff, Duesbery, Alonzo, & Tindal, 2005). Many of the studies that have been conducted in read- ing at the secondary-school level have focused on reading as it relates to learning in the content areas (e.g., Espin & Deno, 1993a, 1993b; Espin & Deno, 1994–1995; Fewster & MacMillan, 2002) rather than on the development of general reading proficiency. However, a small group of studies has focused on general reading proficiency.
Fuchs, Fuchs, and Maxwell (1988) examined the va- lidity of reading aloud for students with mild disabili- ties across grades 4–8. Across-grade correlations between words read correctly (WRC) in 1 minute and scores on comprehension and word study subtests of a stan- dardized achievement test were .91 and .80, respectively; however, because the study was not specifically focused on the secondary-school level, correlations were not re- ported separately for the secondary-school students in the study.
Three subsequent studies focused specifically on secondary-school students. Espin and Foegen (1996) ex- amined the validity of three CBMs—reading aloud, maze selection, and vocabulary matching—on the comprehension, acquisition, and retention of expository text for students in grades 6–8. Comprehension, acquisi- tion, and retention were measured with researcher-designed, multiple-choice questions given immediately after reading (comprehension), immediately after instruction on the text (acquisition), and a week or more following instruction
62 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
(retention). Correlations ranged from .54 to .65 and were similar for comprehension, acquisition, and retention mea- sures. Brown-Chidsey, Davis, and Maya (2003) examined the reliability and validity of a 10-minute maze task—a somewhat long task by CBM standards—as an indicator of reading for students in grades 5–8. They found that scores generally differentiated students by grade level and special education status. Rasinski et al. (2005), in discussing the importance of reading fluency for high school students, re- ported correlations between WRC in 1 minute and scores on a state standards test of .53 for ninth-grade students. Descriptive data and methods were not reported in the article.
In sum, little research has been conducted at the secondary-school level on the development of CBM read- ing measures as indicators of general reading proficiency, and that which has been done has been limited in terms of measures and methodology, or has not focused specifically on secondary-school students. What is more, the research to date has focused on the characteristics of the measures as performance or static measures, not as progress or growth measures. The validity and reliability of the measures may differ based on their intended use.
In this article, we examine the technical adequacy of CBM reading measures for secondary-school students. Specif- ically, the reliability and validity of CBMs as predic- tors of performance on a state standards test in reading is examined. Differences related to time frame and scor- ing procedure are examined. Reading-aloud and maze- selection measures are selected because of previous re- search demonstrating their practical and technical adequacy at the elementary-school level and their potential promise at the secondary-school level. Time frames are examined because longer samples of work might be needed at the middle-school level to obtain a distribution of student scores. For example, reading-aloud scores might bunch together at 1 minute but spread out at 3 minutes. Finally, scoring pro- cedures are examined to determine the influence of er- rors on the reliability and validity of students’ scores. For example, counting the number of correct selections on a maze task is less time consuming than counting the number of correct minus incorrect selections, but using a correct minus incorrect score may help to control for guessing.
Two research questions are addressed in the study:
(1) What are the reliability and validity of reading aloud and maze selection for predicting performance on a state standards test in reading?
(2) Do reliability and validity vary with time frame and scoring procedures?
Our primary focus was on the technical adequacy of CBMs as static measures or indicators of performance at a single point in time. However, we were also able to col- lect progress measures on a small subsample of the orig- inal sample. Thus, we conducted an exploratory study in which we compared the growth rates produced by reading- aloud and maze-selection measures for this subsample of students.
STUDY 1: READING ALOUD AND MAZE SELECTION AS PERFORMANCE INDICATORS
Method
Setting and Participants
The study took place in two middle schools in an urban dis- trict of a large, midwestern metropolitan area. The district enrolled over 47,000 students. Seventy-five percent of the students were from diverse cultural backgrounds, 24 percent received ESL services, 67 percent were eligible for free and reduced lunches, and 13 percent were in special education. The first school had 669 students in grades 6–8. Eighty- three percent of the students were from diverse cultural back- grounds, 35 percent received ESL services, 83 percent were eligible for free and reduced lunches, and 15 percent were in special education. The second school had 778 students in grades 6–8. Sixty-two percent of the students were from diverse cultural backgrounds, 18 percent received ESL ser- vices, 56 percent were eligible for free or reduced lunches, and 16 percent were in special education.
All eighth-grade students were invited to participate in the study to ensure a range of student performance levels. Participants were 236 eighth-grade students (134 females and 102 males) in the classrooms of 17 English teachers from the two schools. Fifty-eight percent of the participants were eligible for free or reduced lunches. Students were Caucasian (34 percent), Asian American (24 percent), African American (20 percent), Hispanic (19 percent), and Native American (3 percent). Nine percent of the students were receiving special education services for learning disabilities or mild disabilities (4 percent), speech and language (3 percent), emotional and behavior disorders (1 percent), or other health impaired (1 percent). Fifty-eight percent of the students spoke English at home. The rest spoke Spanish (18.5 percent), Hmong (16 percent), Laotian (4 percent), Vietnamese (1 percent), Cambodian (1 percent), Amharic (.5 percent), Chinese (0.5 percent), and Somali (0.5 percent). The mean standard score on the state standards reading test for Sample 1 was 626.9. This compared to a state-wide mean score of 640.6 and a district-wide mean score of 607.3.
Note that the sample did not consist of struggling read- ers only, even though the primary purpose of the study was to identify performance and progress measures for strug- gling readers. To establish the reliability and validity of CBM, it was necessary to have a sample that represented a range of student ability levels, because validity and relia- bility coefficients could be negatively affected by a truncated distribution of scores. We had two options. One was to se- lect students who were struggling readers across a range of grade levels, similar to the approach taken by Fuchs et al. (1988). A second was to work within one grade level, but to include students across a range of performance lev- els within that grade. Given that the purpose of the study was to tie the CBM to performance on a state standards test, and given that the state standards test was given in only one grade, we chose the latter approach. This approach is not unique. In a review of the CBM research in reading,
LEARNING DISABILITIES RESEARCH 63
(Wayman et al., 2007), 28 of the 29 technical adequacy stud- ies conducted at the elementary-school level used general education samples (13 studies) or mixed samples of general and special education (15 studies). Only 1 used an exclusively special education sample.
Measures
Predictor variables. Predictor variables were scores on two CBM tasks: reading aloud and maze selection. The reading- aloud and maze-selection tasks were drawn from human- interest stories published in the local daily newspaper and were selected on the basis of content, readability level, length, and scores on a pilot test conducted with four students who were not involved in the study. Passages whose content was determined to be too technical or culturally specific were not used. To ensure that students would not complete the CBM tasks before time was expired, only passages that were longer than 800 words were selected. Readability was calculated us- ing the Flesch-Kincaid formula (Kincaid, Fishburne, Rogers, & Chissom, 1975) via Microsoft Word, and the Degrees of Reading Power (DRP; Touchstone Applied Science and As- sociates, 2006). Readability levels for the selected passages ranged from fifth to seventh grade and DRP levels ranged from 51 to 61. Means (number of words read aloud in 3 min- utes) and standard deviations from the pilot study for selected passages were: 421.5 (SD = 80.5), 489.5 (SD = 117), 432.5 (SD = 140.5), and 401.7 (SD = 75).
The reading-aloud task was administered to students on an individual basis using standardized administration proce- dures. Students read aloud from the passage while the exam- iner followed along on a numbered copy of the same passage, making a slash through words read incorrectly or words sup- plied for the student. The examiner timed for 3 minutes using a stopwatch, marking progress at 1, 2, and 3 minutes. Read- ing aloud was scored for total words read (TWR) and WRC at 1, 2, and 3 minutes.
Maze-selection passages were created from the same sto- ries used for reading aloud. Every seventh word was deleted and replaced by the correct choice and two distracters. The distracters were within one letter in length of the correct word but started with different letters of the alphabet and comprised different parts of speech (see Fuchs, Fuchs, Ham- lett, & Ferguson, 1992, for maze-construction procedures). The three word choices were underlined in bold print and were not split at the end of the sentence in order to preserve continuity for the reader.
The maze selection task was administered to students in a group setting using standardized administration procedures. Students read silently for 4 minutes, making selections for each multiple-choice item. Examiners timed for 4 minutes and instructed students to mark their progress with a slash at 2, 3, and 4 minutes. Examiners monitored to ensure that students made the slashes. Maze selection was scored for cor- rect maze choices (CMC) and correct minus incorrect choices (CMI) in 2, 3, and 4 minutes. As a control for guessing, and following the procedures used in previous research on maze selection (Espin, Deno, Maruyama, & Cohen, 1989; Fuchs et al, 1992), maze scoring was stopped when three consec-
utive incorrect choices were made. A recent investigation comparing different maze-selection scoring procedures re- vealed no differences in criterion-related validity associated with using a two-in-a-row versus three-in-a-row incorrect rule (Wayman et al., 2009).
Criterion variables. The criterion variable in this study was performance on the Minnesota Basic Standards Test (MBST) in reading, a high-stakes test required for gradu- ation. The MBST was designed by the state of Minnesota to test the minimum level of reading skills needed for sur- vival (MN Department of Education, 2001) and, at the time of the study, was administered annually in the winter to all eighth-grade students in Minnesota.1 The untimed test com- prised four or more passages of 500 words or more selected from newspaper and magazine articles. Passages were both narrative and expository and had average DRP levels ranging from 64 to 67. Each passage was followed by multiple-choice questions, with approximately 40 questions per test. The test was constructed so that 60 percent of the questions on the test were literal, 30 percent inferential, and 10 percent could be either. The test was machine-scored on a scale from 0 to 40, and then the raw score was converted to a scale score between 375 and 750. A passing scale score was 600, which corre- sponded to 75 percent correct (MN Department of Education, 2001). Students who did not pass the test were permitted to retake it two times each year. Students had to pass the test in order to graduate from high school.
The MBST Technical Manual (MN Department of Educa- tion, 2001), reported reliability and validity information for the MBST Reading test. Internal consistency measures for reliability were based on the Rasch model index of person separation. The Kuder–Richardson 20 internal consistency reliability estimate was .90. No alternate-form reliability was calculated. Content validity, according to the manual, was determined by the relationship of the reading test items to statewide content standards as verified by educators, item developers, and experts in the field. Construct validity was measured by item point-biserial correlations (the correlation between students’ raw scores on the MBST and their scores on individual test items). The mean point biserial correlation was .38. There were no criterion-related validity statistics noted.
Procedures
In the fall, students completed two maze passages in a group setting in their classrooms. On a subsequent day in the same week, students completed two reading-aloud passages indi- vidually. Type of measure (reading aloud vs. maze selection) and passage were counterbalanced across students, as was the order in which the students completed the passages within reading aloud or maze selection. Examples of each task were given to students prior to administration. The MBST was administered by teachers to students in February.
Sixteen graduate students administered and scored the reading-aloud and maze-selection measures. Prior to data collection, the graduate students were interviewed by mem- bers of the research team to ascertain their ability to work
64 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
with students and to accurately score reading samples. Fol- lowing this initial screening, the graduate students partic- ipated in two 2-hour training sessions on administration and scoring. During training, the graduate students admin- istered and scored three samples. Inter-scorer agreement on the three passages between the data collectors and the trainer was calculated by dividing the smaller by the larger score and multiplying by 100. Inter-scorer agreement ex- ceeded 95 percent on maze selection and 90 percent on read- ing aloud for all scorers. During data collection and scor- ing, 33 percent of the reading-aloud and 10 percent of the maze-selection probes were randomly selected to be checked for accuracy of scoring. Inter-scorer agreement exceeded 90 percent for all measures.
Results
Means and standard deviations for reading-aloud and maze- selection scores for each time frame are reported in Table 1. Examination of mean scores reveals that students worked at a steady pace across the duration of the passages. Students read aloud approximately 125 words with 6 errors per minute across the 3 minutes and made approximately 6 correct maze choices with 0.5 errors per minute across the 4 minutes of maze. The mean score for study participants on the MBST in reading was a standard score of 626.90 (SD = 65.66), with a range of 475–750.
To determine alternate-form reliability, correlations be- tween scores on the two forms of the maze-selection and reading-aloud measures were calculated for each time frame and scoring procedure (see Table 2). Reliabilities for both reading aloud and maze were generally above .80. Relia- bilities for reading aloud ranged from .93 to .96, and were similar across scoring method and sample duration. Relia- bilities for maze ranged from .79 to .96, and were generally similar for scoring method, but increased somewhat with time frame. The highest obtained reliability coefficient was for the 4-minute maze passages scored for CMI (r = .96); however, reliabilities for the 3-minute maze selection were above .85, regardless of scoring method.
TABLE 1 Means and Standard Deviations for Reading Aloud and Maze
Selection by Scoring Procedure and Time Frame
Curriculum-Based Measurements and scoring procedure Time
Reading aloud 1 minute 2 minutes 3 minutes Total words read 125.88 250.46 373.31
(43.75) (85.05) (125.95) Words read correct 119.82 238.54 355.27
(47.29) (92.14) (136.92) Maze selection 2 minutes 3 minutes 4 minutes
Correct choices 12.33 18.76 25.24 (7.12) (10.87) (14.53)
Correct minus 11.18 17.17 23.10 incorrect choices (7.53) (11.40) (15.17)
Note: Standard deviations are in parentheses.
TABLE 2 Alternate-Form Reliability for Reading Aloud and Maze Selection
by Scoring Procedure and Time Frame
Curriculum-Based Measurements and scoring procedure Time
Reading aloud 1 minute 2 minutes 3 minutes Total words read .93 .96 .95 Words read correct .94 .96 .94
Maze selection 2 minutes 3 minutes 4 minutes Correct choices .80 .86 .88 Correct minus .79 .86 .96
incorrect choices
Note: All correlations significant at p < .01.
TABLE 3 Predictive Validity Coefficients for Reading Aloud and Maze
Selection with MBST by Scoring Procedure and Time Frame
Curriculum-Based Measurements and scoring procedure Time
Reading aloud 1 minute 2 minutes 3 minutes Total words read .76 .77 .76 Words read correct .78 .79 .78
Maze selection 2 minutes 3 minutes 4 minutes Correct choices .75 .77 .80 Correct minus .77 .78 .81
incorrect choices
Note: All correlations significant at p < .01. MBST : Minnesota Basic Standards Test.
To examine the predictive validity of the measures, cor- relations between mean scores on the two forms of reading- aloud and maze-selection measures and scores on the MBST were calculated (see Table 3). Correlations ranged from .75 to .81. The magnitude of the correlations was similar across type of measure (reading aloud and maze) and method of scoring. For reading aloud, correlations for 1, 2, and 3 min- utes were virtually identical. For maze selection, a consis- tent but small increase in correlations was seen across time frames, with correlations of .75 (CMC) and .77 (CMI) for the 2-minute measure and .80 (CMC) and .81 (CMI) for the 4-minute measure.
In summary, results revealed that both maze selection and reading aloud produced respectable alternate-form reliabil- ities, although reading aloud yielded consistently larger re- liability coefficients than maze. Few differences in reliabili- ties were seen for scoring procedure or time frame with the exception that reliabilities for the maze selection increased somewhat with time. Predictive validity coefficients were similar for the two types of measures. Correlations were sim- ilar across scoring procedures for both measures. With regard to time frame, small but consistent increases in correlations were seen for maze selection.
Discussion
In this study, we examined the reliability and validity of read- ing aloud and maze selection as indictors of performance on
LEARNING DISABILITIES RESEARCH 65
a state standards test. Difference in technical characteristics related to time frame and scoring procedure were examined.
Both reading aloud and maze selection showed reason- able alternate-form reliabilities at all time frames, with most coefficients at or above .80. In general, reading aloud re- sulted in higher alternate-form reliability coefficients (rang- ing from .93 to .96) than did maze selection (ranging from .79 to .96), but reliability for maze selection was in the range typical for CBM. Time frame did not influence re- liability coefficients for reading aloud but had some influ- ence on maze selection. Obtained reliability coefficients for maze increased with time frame, with coefficients for the 2-minute time frame hovering around .80, but increasing for 3-minute (r’s = .86) and 4-minute (r = .88 and .96) time frames. Finally, scoring procedure had little effect on relia- bility, with the exception that when 4-minute maze selection was scored for CMI, reliability was somewhat larger (r = .96) than when it was scored for CMC (r = .88).
Like reliability coefficients, validity coefficients were quite similar across type of measure, time frame, and scoring procedure. Validity coefficients for reading aloud ranged be- tween .76 and .79 and were similar across scoring procedure and time frames. Maze-selection coefficients ranged between .75 and .81 and also were similar across scoring procedure. A systematic increase in validity coefficients was seen with an increase in time for maze, but differences were small.
We wish to make two observations regarding the magni- tude of the validity coefficients found in the performance study. First, the correlations obtained in our study were larger than those found in previous research at the middle- school level. For example, Yovanoff et al. (2005) reported correlations of .51 and .52 between WRC in 1 minute and scores on a reading comprehension task for eighth-grade stu- dents. Espin and Foegen (1996) reported correlations of .57 and .56, respectively, between WRC in 1 minute and CMC in 2 minutes and scores on a reading comprehension task.
One might hypothesize that the differences in correla- tions are related to the materials used to develop the CBMs, although no consistent pattern of differences can be seen across studies. Yovanoff et al. (2005) used grade-level prose material, Espin and Foegen (1996) used fifth-grade level ex- pository material, and we used fifth- to seventh-grade human- interest stories from the newspaper—material that might be considered to be both narrative and expository. Moreover, previous research conducted at the elementary-school level has revealed few differences in reliability and validity for CBMs drawn from material of different difficulty levels or from various sources (see Wayman et al., 2007, for a review).
It is possible that differences are related to the criterion variable used. Both Yovanoff et al. (2005) and Espin and Foegen (1996) used a limited number of researcher-designed multiple-choice questions as an outcome, whereas in our study we used a broad-based measure of comprehension de- signed to scale student performance across a range of levels. Supporting this hypothesis are data from two studies demon- strating nearly identical correlations (in the .70s) to those we found between the CBM reading-aloud and maze-selection measures and the MBST (Muyskens & Marston, 2006; Ticha, Espin, & Wayman, 2009). In addition, Ticha et al. (2009)
found high correlations between maze-selection scores and a standardized achievement test.
Second, the state standards test used in the current study was designed to test the minimal reading competency for students in eighth grade. Thus, one might question whether the CBMs would predict reading competence as well if the criterion measures were measures of broader reading compe- tence. Results of Ticha et al. (2009) indicate that the reading measures predict performance on a standardized reading test as well as (or better than) they predict performance on the state standards test. Perhaps the nature of the state test serves to reduce the overall variability in scores and thus serves to reduce the correlations. Replication of the current study with other outcome measures of reading proficiency is in order.
In summary, the results supported the reliability and va- lidity of both reading aloud and maze selection as indictors of performance on a state standards reading test for middle- school students. For reading aloud, our data, combined with practical considerations, would suggest use of a 1-minute sample scored for TWR or WRC as a valid and reliable indi- cator of performance. Little was gained in technical adequacy by increasing the reading time. Given that reading aloud is typically scored for WRC, and given that this scoring pro- cedure is no more time consuming than scoring TWR, we would recommend scoring the sample for WRC rather than TWR.
For maze selection, our data, combined with practical considerations, would suggest use of a 3-minute selec- tion task scored for CMC as a valid and reliable indicator of performance. Although reliability and validity coeffi- cients were the strongest for 4 minutes, the differences be- tween 3- and 4-minute coefficients were small in magnitude, and both data collectors and teachers reported anecdotally that a 4-minute maze task was tedious for the students to complete.
Although our data support the use of both WRC in 1 minute and CMC in 3 minutes as predictors of performance on a state standard test, one might ask how teachers can use such data in their decision making. A common approach is to create a district-wide cutoff score on the CBM that is associated with a high probability of passing the state standards test. For example, district-wide data may show that, of students who read 145 WRC in 1 minute, 80 percent pass the state standards test. Teachers might then set a goal of 145 WRC in 1 minute for their students. The disadvantage of a cutoff score for students who struggle in reading is that these students often perform well below the cutoff score. An alternative approach is to present the relationship between performance on the CBM measures and the likelihood of passing the state standards test along the entire performance continuum. For example, district-wide data may show that, of students who read 100 WRC in 1 minute, 26 percent pass the state standards test, but of students who read 126 WRC in 1 minute, 57 percent pass. Teachers may choose to set an annual goal of 126 WRC for a student who begins the year reading only 100 WRC. This goal would move the student closer to a level of likely success. A method that can be used to create these Tables of Probable Success using CBM data is explained and illustrated in Espin et al. (2008).
66 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
In conclusion, our study supports the technical adequacy of both reading aloud and maze selection as indicators of performance on a state standards reading test. In the past, technical adequacy research would stop here, with the as- sumption that, if both measures were shown to be valid and reliable with respect to performance on a criterion measure, then both measures would reflect growth as progress mea- sures. However, development of more advanced statistical techniques such as Hierarchical Linear Modeling now allow for examination of the characteristics of CBM measures as progress as well as performance measures (e.g., see Shin, Es- pin, Deno, & McConnell, 2004). From our original sample, we had access to one classroom with 31 students for weekly progress monitoring. Although this sample size was too small to produce generalizable results about typical growth rates on the measures, it was large enough to conduct an exploratory, within-subject comparison of the growth rates produced by the two measures for a sample of students. Specifically, we examined differences in the sensitivity of the two measures to growth and their relation to performance on the state stan- dards test. Results of this exploratory study could help us to generate hypotheses for future research.
STUDY 2: EXPLORATORY PROGRESS STUDY
Method
Participants
Participants in exploratory progress study were selected from the original sample and were 31 (10 male; 21 female) stu- dents from one classroom in the first school described above. Fifty-five percent of the students were eligible for free or re- duced lunches. Students were Caucasian (42 percent), Asian American (26 percent), African American (16 percent), His- panic (10 percent), and Native American (6 percent). Ten percent of the students received special education services for emotional disturbance or speech-language difficulty. Six- teen percent of the students were identified as English lan- guage learners (ELL) but did not receive ESL services. The mean standard score on the state standards reading test for the students was 646.17.
Procedures
Students were monitored weekly on both a maze-selection task, administered in a group setting by the classroom teacher, and a reading-aloud task, administered on an individual basis
TABLE 4 Alternate-Form Reliability for Reading-Aloud and Maze-Selection Progress-Monitoring Passages
Reading aloud, words correct, 1 minute Passages 1 and 2 2 and 3 3 and 4 4 and 5 5 and 6 6 and 7 7 and 8 8 and 9 9 and 10
.92 .91 .85 .88 .88 .86 .79 .84 .83
Maze, correct choices, 3 minutes .72 .84 .69 .80 .80 .85 .90 .83 .74
n = 25 to 31. Note: All correlations significant at p < .01.
by a member of the research team. The maze-selection and reading-aloud tasks were created from the same passages each week. Sixteen passages were selected from human- interest stories from the newspaper. Passages that required specific background knowledge (e.g., knowledge of the game of baseball) were eliminated from consideration. For the re- maining passages, readability levels were calculated using both the DRP (Touchstone Applied Science and Associates, 2006) and Flesch-Kincaid (Kincaid et al., 1975). In addi- tion, teachers were consulted regarding appropriateness of the passages for secondary-school students. A final set of 10 passages was selected based on readability formula and teacher input. DRP scores ranged from 51 to 61, representing approximately a sixth-grade level, and Flesch-Kincaid read- ability level was between the fifth and seventh grade levels. Passages were on average 750 words long. Alternate-form reliabilities between adjacent pairs of passages are reported in Table 4. All reliabilities were statistically significant, all but one were above .70, and all but three were above .80. For reading aloud, reliabilities ranged from .79 to .92, and for maze selection from .69 to .90.
Maze selection was administered first, usually on a Mon- day, and reading aloud was administered on a subsequent day within the same week, usually on a Friday. Progress data were collected over a period of approximately 3 months, yielding an average of 10 data points per student (note that during vacation weeks, no data were collected).
Scoring
Maze selection was administered by the classroom teacher using a standard script. Maze-selection probes were scored by graduate students. Prior to administering the measure the first time, the teacher observed one of the members of the re- search team administering the maze task to her class. Fidelity of treatment checks were conducted at equal intervals three times during the course of the study to assess accuracy of the administration and timing of the maze. For each fidelity check, the teacher was found to read the directions and com- plete the timings correctly. Reading aloud was administered and scored by 11 of the data collectors from the original study. Every week, 10 of the reading-aloud samples were tape-recorded and checked for fidelity and reliability, and 10 of the maze-selection passages were checked for accuracy of scoring. On all occasions, data collectors read the directions and timed correctly for the reading-aloud samples. Accu- racy of scoring for reading aloud and maze selection was checked by the two graduate students involved in the study.
LEARNING DISABILITIES RESEARCH 67
Percentage agreement between the graduate students and scorers was calculated by dividing agreements by agree- ments plus disagreements. Accuracy of scoring across the study for both maze selection and reading aloud ranged from 97 percent to 100 percent.
Results
The sensitivity of the measures to growth and the validity of the growth rates produced by the measures were examined. Growth curve analyses were carried out using the MIXED procedure of SAS 9.1. Linear mixed effects growth curves were used (see Fitzmaurice, Laird, & Ware, 2004, ch. 8) with different models specified for each measure. Similar patterns of results were found for each time frame and scoring pro- cedure within both reading aloud and maze selection, thus we report results selectively. First, we report on the mea- sures with the best reliability, validity, and efficiency from the performance study—WRC in a 1-minute reading-aloud and CMC in a 3-minute maze-selection task. Second, to pro- vide a direct comparison of reading aloud and maze selection with time held constant, we also report results for WRC in a 3-minute reading-aloud task. Finally, taking into consid- eration practicality of use, we report results for CMC in a 2-minute maze selection. A 2-minute maze selection is more practical than a 3-minute maze selection for ongoing and fre- quent progress monitoring, and we considered the reliability and validity coefficients for 2-minute maze selection to be within an acceptable range for progress monitoring.
Sensitivity to growth for reading aloud. The growth curve model for reading aloud was a simple linear growth curve,
Yi j = β0 + β1ti j + εi j , (1)
where i is the participant subscript, i = 1, . . , N , and j is the wave subscript, j = 1, . . ., ni. In Equation (1), β 0 is the intercept (status at wave 1), and β 1 is the linear slope with tij = j − 1, and ε ij = b0i + b1itij + eij, which is the random effects structure with b0i being the deviation of an individual’s individual intercept from the mean intercept, b1i
being the deviation of an individual’s slope from the mean slope, and eij is random error (Fitzmaurice et al., 2004, ch. 8). Restricted maximum likelihood was used for parameter estimation, and degrees of freedom (df ) for the t tests of the parameter estimates were estimated using the method of Kenward and Roger (1997).
The observed means and predicted means based on the linear model for 1-minute reading aloud are presented at the top of Figure 1. Detailed results of all the growth curve anal- yses are in Table 5. Results reveal that the linear slope was significant, β̂1 = .84, t(245) = 2.63, p = .009, and the in- tercept was significant, β̂0 = 139.01, t(28.8) = 24.87, p < .0001. The variance of the linear slopes was estimated to be zero, so the df were based on the sum of all the time points over all participant, � ini, rather than the number of partic- ipants (N). This resulted in relatively high df (i.e., 245) for the t test of the slopes. However, the result is still significant
when based on the same df used for testing the intercept (i.e., t(28.8) = 2.63, p = .014). The mean linear slope indicates that the number of WRC in 1 minute tended to increase at a rate of .84 per wave.
The observed means and predicted means based on the linear model for 3-minute reading aloud are presented at the bottom of Figure 1. The linear slope for WRC3 was not significant, β̂1 = −0.41, t(29.4) = −0.55, p = 0.58, but the intercept was significant, β̂0 = 420.04, t(29) = 26.71, p < 0.001. The linear slope indicates that WRC in 3 minutes did not show a significant increase over time. (Although not reported here, the 2-minute reading-aloud measure scored for WRC also showed no significant change over time.)
Sensitivity to growth for maze selection. The observed and predicted means for 3- and 2-minute maze selection scored for CMC are presented in Figure 2. As illustrated in Figure 2, there was a change of direction at wave 8 for the maze-selection scores. This presented a problem for the analyses as the major goal was to estimate linear growth over time. After a close examination of the data and discussions with the teacher, it was determined that this shift might be due to a passage effect. To address this problem, a piecewise or spline model was used to fit an additional linear predictor starting at wave 8 to account for the observed nonlinearity (see Ruppert, Wand, & Carroll, 2003, ch. 3). That is, we decided to model the near-linear growth apparent before wave 8 without deleting any of the data. The spline growth curve model was
Yi j = β0 + β1ti j + β2t∗ i j + εi j . (2)
In Equation (2), β 0 is the intercept (constant across all 10 waves), β 1 is the linear slope over wave 1 to wave 7 (with tij = j − 1), and β 2 is the linear slope starting at wave 8, with tij
∗ = 0 for wave 1 through wave 7 and tij ∗ = (tij − 7)
starting at wave 8. Random effects terms were specified for each linear slope and the intercept, that is, ε ij = b0i + b1itij
+ b2itij ∗ + eij. Primary interest was on β 1 as this was the
linear slope for waves 1–7. The observed and predicted means based on the spline
model for 3-minute maze selection are shown at the top of Figure 2. The results show that each parameter estimate of the spline model was significant, β̂0 = 21.22, t(25.3) = 17.22, p < .0001, β̂1 = 2.88, t(32.7) = 12.79, p < .0001, and β̂2 = −7.26, t(227) = −10.82, p < .0001. (The random effects component of the linear slope starting at wave 8 was estimated to be zero accounting for the higher df for the test of H0: β 2 = 0, also see Table 5.) The latter two estimates indicate there was an overall rate of increase of 2.88 CMC in 3 minutes per wave, but a decrease of 7.26 per wave at wave 8.
The observed and predicted means based on the spline model for 2-minute maze selection scored for CMC are pre- sented at the bottom of Figure 2. Each parameter estimate of the spline model was significant, β̂0 = 13.58, t(24.00) = 16.93, p < .0001, β̂1 = 2.17, t(32.7) = 12.85, p < .0001, and β̂2 = −5.75, t(225) = −10.19, p < .0001. Thus, for the 2-minute maze, there was an overall rate of increase of 2.17 CMC in 2 minutes per wave, but a decrease of 5.75 per wave
68 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
Linear Model: Words Read Correct 1 Minute
0
20
40
60
80
100
120
140
160
180
10987654321
Wave
WRC1 Obs.
WRC1 Pred.
Linear Model: Words Read Correct 3 Minutes
280
300
320
340
360
380
400
420
440
460
10987654321
Wave
WRC3 Obs. WRC3 Pred
FIGURE 1 One- and 3-minute reading aloud (words read correctly) observed and predicted means by wave.
LEARNING DISABILITIES RESEARCH 69
TABLE 5 Detailed Results of the Growth Curve Analysis
Parameter Estimate SE df t value p value
WRC1 β 0 138.1 5.5538 28.8 24.87 <.0001 β 1 0.8391 0.319 245 2.63 .0091
WRC3 β 0 420.04 15.7232 29 26.71 <.0001 β 1 −0.4057 0.7333 29.4 −0.55 0.5842
CMC2 β 0 13.5786 0.8022 24 16.93 <.0001 β 1 2.1674 0.1686 32.7 12.85 <.0001 β 2 −5.747 0.5642 225 −10.19 <.0001
CMC3 β 0 21.22 1.2321 25.3 17.22 <.0001 β 1 2.8808 0.2253 32.7 12.79 <.0001 β 2 −7.2639 0.6716 227 −10.82 <.0001
Note: WRC1, WRC3: Words read correctly in 1 and 3 minutes. CMC2, CMC3: Correct maze choices in 2 and 3 minutes.
at wave 8. (Although not reported here, a similar pattern of results was found for 4-minute maze-selection measures.)
In summary, the results of the growth curve analysis for the reading-aloud and maze-selection measures demon- strated that reading aloud showed minimal or no growth over time, whereas maze selection showed significant and substan- tial growth except after week 7. This pattern of results held generally across scoring procedure and time frame. Thus, there were no differences in patterns of growth found for the different scoring methods for either reading aloud or maze selection. With regard to time frame, maze selection demon- strated significant and substantial growth for both 2-minute (2.17 CMC) and 3-minute (2.88 CMC per week) time frames (although recall that these growth rates were obtained fol- lowing a correction for a score shift). However, results for reading aloud revealed a statistically significant but minimal growth rate for the 1-minute (.84 words per week) reading- aloud measure, but no significant growth for the 3-minute (or 2-minute) measure.
One might conjecture that reading aloud might be more sensitive to growth for lower-performing than for higher- performing students. However, examination of Figure 3, which presents individual student data across time, reveals that growth across time for reading aloud (top graph) was fairly flat for those at both the lower and higher levels of CBM performance. In contrast, growth across time for maze selection (bottom graph) reflects a fanning out of scores over time, with students at the higher levels of CBM performance reflecting larger gains than those at the lower levels of perfor- mance. The meaningfulness of the growth rates produced by the maze-selection measures were examined in the follow- ing analysis. Reading aloud was not included in this analysis due to the lack in interindividual variability in growth rates, which would mean that slopes would not be correlated with other variables.
Relation between growth rates and performance on the MBST. To examine the validity of the growth rates, we
investigated the extent to which the growth rates produced by the 3-minute maze-selection (CMC) measures were related to performance on the MBST. Specifically, we used MBST scores as predictors of linear slopes and intercepts but were primarily interested in the former. A random effects was associated with each fixed effects in the same manner as the growth curves above (i.e., ε ij = b0i + b1itij + eij, or ε ij = b0i + b1itij + b2it∗ij + eij). The mean score on the MBST for participants in Study 2 was a standard scored of 646.17 (SD = 38.11), with a range of 587 to 750. Let mi = the MBST score for the ith participant, which is a static predictor (not varying over time). For the CMC, the statistic predictor was incorporated into the spline model,
Yi j = β0 + β1ti j + β2t∗ i j + β3mi + β4mi ti j + εi j . (3)
In Equation (3), β 4 represents the association between the MBST and the slopes for waves 1–7.
The number of CMC in 3 minutes resulted in significant growth over time, and this growth was related to performance on the MBST, with students passing the MBST obtaining higher rates of growth over time than those not passing the MBST (β̂4 = 0.009, t(31.3) = 2.35, p = 0.025). Figure 4 shows the predicted curves based on the estimated parameters of Equation (3) for MBST scores of 500 (not passing) and 700 (passing). Note the students passing the MBST start higher and increase at a faster rate of change. Although not reported here, this same pattern of results also was found for the 2-minute maze-selection measure.
Discussion
The characteristics of the measures as progress measures were compared for a small subset of our original partici- pant sample. For these students, only maze selection resulted in substantial and significant growth over time (2.88 selec- tions per week for a 3-minute sample), while the 1-minute reading-aloud measure revealed statistically significant, but minimal growth over time (.84 WRC per week). Controlling for time frame and using a 3-minute reading-aloud task did not increase the amount of growth. In fact, a 3-minute sam- ple of reading aloud resulted in no significant growth over time. Further, the growth rates produced by maze selection were significantly related to performance on the MBST— students with higher scores on the MBST also grew more on the maze-selection measures. Growth on the 1-minute reading-aloud measure was not related to performance on the MBST.
The differences in growth for reading-aloud and maze- selection measures are surprising and difficult to explain, especially in light of the good technical adequacy of both measures as performance measures. One obvious explana- tion might be differences in the materials used to construct the measures—but recall that the reading-aloud and maze- selection tasks were created from the same passages each week. Another obvious explanation might be a bunching of scores over time because of a ceiling effect on the reading aloud; however, inspection of the data reveals no ceiling ef- fect for the reading-aloud scores in the original study. In
70 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
Spline Model: Correct Maze Choices 3 Minutes
0
5
10
15
20
25
30
35
40
45
50
10987654321
Wave
CMC3 Obs.
CMC3 Pred.
Spline Model: Correct Maze Choices 2 Minutes
0
5
10
15
20
25
30
35
40
45
50
10987654321
Wave
CMC2 Obs CMC2 Pred
FIGURE 2 Three- and 2-minute maze selection (correct maze choices) observed and predicted means by wave.
LEARNING DISABILITIES RESEARCH 71
Words Read Correctly in 1 Minute
50
70
90
110
130
150
170
190
210
230
10987654321
Wave
Mean
Correct Maze Choices in 3 M inutes
0
10
20
30
40
50
60
70
80
10987654321
W ave
C o
rr ec
t ch
o ic
es
Mean
FIGURE 3 Individual growth rates for 1-minute reading aloud (words read correctly) and 3-minute maze selection (correct maze choices).
addition, examination of the individual growth rates over time, as illustrated in Figure 3, reveals no bunching of scores over time. Both higher- and lower-performing students (in terms of CBM scores) tended to maintain their relative levels of performance over time. When viewing a similar picture of the maze selection 3-minute task (Figure 3), one sees a fairly steady increase in scores for all students, with a fan- ning out of the scores over time. This fanning out is due to higher-performing students showing greater growth than the lower-performing students, an observation confirmed by the subsequent analysis with the state standards test. A fi-
nal possible explanation is that the maze-selection task was relatively novel to the students while the reading-aloud task was not and that growth on the maze selection was due to practice on the task, rather than improvements in reading performance. However, both the reading-aloud and maze- selection tasks were novel to the students (i.e., the district did not regularly collect reading-aloud data on the students), and the relation between growth on the maze and the criterion variable would argue against simple practice effects.
A more plausible reason for differences in growth rates produced by the measures might be the order in which the
72 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
0
5
10
15
20
25
30
35
40
45
50
10987654321 Wave
MBST=500
MBST=700
FIGURE 4 Relationship between 3-minute maze selection (correct maze choices) and Minnesota Basic Standards Test (MBST) scores.
measures were administered. In the current study, maze al- ways was administered first and reading aloud second. Per- haps completing the maze task diminished the sensitivity of the reading-aloud measure to change over time. Although the effect of order must be considered, we would note that, in a follow-up study in which the reading-aloud measure was administered first and maze selection second (Ticha et al., 2009), a similar pattern of results was obtained, with maze selection reflecting growth over time but reading aloud not.2
Another plausible reason for differences in the measures may be related to the small convenience sample used in this exploratory study. Compared to the larger sample, students in the exploratory study had relatively higher mean MBST scores (646.17 vs. 626.90). Correlations between the CBM and MBST scores for this subsample were lower than for the larger sample (see Table 6), ranging from .30 to .34 for read- ing aloud, and from .48 to .57 for maze selection (despite no restriction in the range of scores on the predictor and crite- rion variables). In addition, correlations for reading aloud are consistently and substantially lower than for maze selection. Perhaps for this particular sample of students, reading aloud did not function as a reasonable indicator of growth, but maze selection did. However, in the follow-up study referred to earlier (Ticha et al., 2009), similar results were obtained regarding growth rates produced by reading aloud and maze selection. It is important to note that, in the follow-up study, performance-related correlations between the CBM and the MBST were virtually identical to those obtained in this study.
Our results suggest that there may be differences in the characteristics of the measures when used as predictors of performance versus measures of progress, and that both should be considered when examining the technical ade- quacy of the measures. Our sample is too small to draw conclusions regarding the best measure for monitoring the progress of secondary-school students, but it does suggest the need for further research examining the characteristics of the two measures for reflecting growth over time for older students. It may be that, although students as a group do not reach a ceiling in scores, each individual student reaches a “natural” level of reading fluency that, when compared to
TABLE 6 Correlations for Reading Aloud and Maze Selection with MBST for
Study 2 Participants at Time of MBST
CBM measure and scoring procedure Time
Reading aloud 1 minute 2 minutes 3 minutes Total words read .30 .32 .33 Words read correct .32 .33 .34
Maze selection 2 minutes 3 minutes 4 minutes Correct choices .52∗∗ .50∗∗ .48∗∗ Correct minus .57∗∗ .55∗∗ .51∗∗
incorrect choices
n = 31. ∗∗p < .01. Note: MBST : Minnesota Basic Standards Test.
others, reveals a general level of reading proficiency but does not change with time. In addition, despite the fact that in our study neither lower- nor higher-performing students showed growth on the reading-aloud measure, it would be important to replicate the findings with a large, cross-grade sample of just struggling readers.
Although not the original intent of the study, this ex- ploratory study also provides us with data regarding between- passage variability. The order in which the measures were administered was not counterbalanced across students, pre- venting us from drawing conclusions regarding typical growth rates for lower- and higher-performing readers. How- ever, the design does allow us to examine characteristics of the passages themselves as growth measures. As is evident in Figure 2, students displayed a steady rate of growth on the maze-selection measure until week 8, when there was a sudden spike in scores for nearly all students, followed by a fairly low score in week 9 for nearly all students. Even more interesting, if one examines the reading-aloud graph in Figure 2, this same spike in scores on week 8 is not evi- dent (recall that the same passages were used for reading aloud and maze selection each week). Perhaps the simplest explanation for this pattern would be administration error on the particular days that maze passages 8 and 9 were given. Although fidelity checks on teacher administration of maze selection revealed that the teacher administered the passages correctly, those checks were conducted only three times dur- ing the course of the study and were not done on the days that passages 8 and 9 were administered. However, when asked, the teacher reported no particular problems with ad- ministration on those days (although one must still consider administration error a potential explanation for the pattern of results).
There is, however, a potential, somewhat troubling ex- planation for the pattern seen at points 8 and 9, related to between-passage variability. As discussed in Wayman et al. (2007), determination of passage difficulty is an important, but complicated, task for progress monitoring. The task is important because passages must be equivalent if we are to attribute growth over time to change in student performance as opposed to passage variability. The task is complicated be- cause it is difficult to predict what may make a passage easy
LEARNING DISABILITIES RESEARCH 73
or difficult. The most common technique for determining passage equivalence, use of readability formulas, is not reli- able (see Ardoin, Suldo, Witt, Aldrich, & McDonald, 2005; Compton, Appleton, & Hosp, 2004). Especially troubling with respect to our data is that whatever affected the scores in week 8 on maze selection did not affect the scores on read- ing aloud. Thus, it would not be merely the characteristics of the passage itself, but the characteristic of the passage as a maze passage that produced the variability for this particular sample.3
We suggest further examination of the effects of passage variability on growth and emphasize the need for careful and systematic approaches to developing equivalent passages for CBM progress monitoring, especially when that progress monitoring is to be used as a part of a high-stakes decision- making process. Perhaps the best guarantee of passage equiv- alence is to administer the passages to a group of students and examine whether the group as a whole obtains higher or lower scores on particular passages or to use the same passage for repeated testing (see Griffiths, VanDerHeyden, Skokut, & Lilles, 2009, for a discussion of these approaches). If using parallel forms, it would be wise to counterbalance the order in which the passages are administered, especially if the goal is to establish normative growth rates on the measures.
CONCLUSION
We examined the technical characteristics of two CBM read- ing measures as indicators of performance for middle-school students. We also conducted an exploratory study to examine the characteristics of the measures as progress measures. Our goal was to develop measures that could be used to moni- tor the progress of students with reading difficulties; how- ever, to examine technical characteristics of the measures, we needed to include students across a range of performance levels. Our results supported the use of both reading aloud and maze selection as indicators of performance on a state standards test representing survival levels of reading perfor- mance. Reliability and validity were good for both measures, and within the range of levels found in previous research at the elementary-school level. Further, few differences were found related to scoring procedure or time frame, although reliability did increase somewhat for maze with an increase in time frame. Given the results of the performance study, and taking into account practical considerations, we would recommend use of WRC on a 1-minute reading-aloud task or CMC on a 3-minute maze-selection task as indicators of performance.
Results of the exploratory progress study implied that it is important to consider technical adequacy of the measures as both performance and progress measures. In our study, only maze selection revealed growth over time, reading aloud did not. Further, growth on the maze-selection measure was related to performance on the MBST.
The studies represent only the first step in the development of CBM for monitoring the progress of students with reading difficulties at the secondary-school level. First, our results ap-
ply only to the sample used in the study, and replication with other samples is necessary. Second, our research addressed students at the middle-school level. There is a need for re- search at the high school level. We do not know whether our results would generalize to older students. Third, our progress study was a pilot study. Results must be replicated with a larger, representative sample. Fourth, once reliable and valid measures are developed, it will be important to examine whether teacher use of the measures leads to improvement for struggling readers. That, of course, is the ultimate goal of the research program and one which will need to be examined directly at both the middle- and high school level because one cannot assume that the positive results for use of the measures at the elementary-school level (see Stecker, Fuchs, & Fuchs, 2005, for a review) will necessarily replicate at the secondary-school level. Finally, although our results support the use of the CBM reading measures as indicators of per- formance in reading, we would support the use of multiple measures for determination of students’ need for additional intensive reading instruction.
ACKNOWLEDGMENTS
The research reported here was funded by the Office of Special Education Programs, U.S. Department of Education, Field Initiated Research Projects, CFDA 84.324C. We thank Mary Pickart for her contributions to this research and Stan- ley Deno for his insights. We also thank the Netherlands Institute for Advanced Study in the Humanities and Social Sciences for its support in the preparation of this manuscript.
NOTES
1. The MBST in reading is being replaced by the Minnesota Comprehensive Assessment, a more broad-based reading test that is given annually in 3rd through 8th grade and again in 10th grade.
2. We would like to note that there was an error in the Ticha et al. (2009) article. In the methods section, the maze selection is said to be given first, and the reading aloud second. Later in the discussion section, the reading aloud is said to be given first and the maze selection second. In fact, the reading-aloud measures were given first, and the maze-selection measures second.
3. We hypothesize that the difficulty of the passage in week 8 had to do with a fairly infrequent word appearing in the very first maze selection item that created difficulties for all students. In the reading-aloud measure, this word was supplied after 3 seconds and, thus, may have had less of an effect on the overall score of the students.
REFERENCES
Ardoin, S. P., Suldo, S. M., Witt, J., Aldrich, S., & McDonald, E. (2005). Accuracy of readability estimates’ predictions of CBM performance. School Psychology Review, 20, 1–22.
Brown-Chidsey, R., Davis, L., & Maya, C. (2003). Sources of variance in curriculum-based measures of silent reading. Psychology in the Schools, 40, 363–377.
74 ESPIN ET AL.: CREATING A READING PROGRESS MEASUREMENT SYSTEM
Center on Education Policy. (2008). State high school exit exams: A move to- ward end-of-course exams. Washington, D.C.: U.S. Government Print- ing Office. Retrieved on March 25, 2009, from www.cep-dc.org
Compton, D. L., Appleton, A., & Hosp, M. K. (2004). Exploring the relation- ship between text-leveling systems and reading accuracy and fluency in second grade students who are average and poor decoders. Learning Disabilities Research & Practice, 19, 176–184.
Crawford, L., Tindal, G., & Stieber, S. (2001). Using oral reading rate to pre- dict student performance on statewide achievement tests. Educational Assessment, 7, 303–323.
Deno, S. L. (1985). Curriculum-based measurement: The emerging alterna- tive. Exceptional Children, 52, 219–232.
Deshler, D. D., Schumaker, J. B., Alley, G. B., Warner, M. M., & Clark, F. L. (1982). Learning disabilities in adolescent and young adult populations: Research implications. Focus on Exceptional Children, 15(1), 1–12.
Espin, C. A., Deno, S. L., Maruyama, G., & Cohen, C. (1989). The Basic Academic Skills Samples (BASS): An instrument for the screening and identification of children at risk for failure in regular education class- rooms. Paper presented at the National Convention of the American Educational Research Association, March.
Espin, C. A., & Deno, S. L. (1993a). Performance in reading from con- tent area text as an indicator of achievement. Remedial and Special Education, 14, 47–59.
Espin, C. A., & Deno, S. L. (1993b). Content-specific and general reading disabilities of secondary-level students: Identification and educational relevance. The Journal of Special Education, 27, 321–337.
Espin, C. A., & Deno, S. L. (1994–95). Curriculum-based measures for secondary students: Utility and task specificity of text-based reading and vocabulary measures for predicting performance on content-area tasks. Diagnostique, 20, 121–142.
Espin, C. A., & Foegen, A. (1996). Validity of general outcome measures for predicting secondary students’ performance on content-area tasks. Exceptional Children, 62, 497–514.
Espin, C. A., Wallace, T., Campbell, H., Lembke, E. S., Long, J. D., & Ticha, R. (2008). Curriculum-based measurement in writing: Predicting the success of high-school students on state standards tests. Exceptional Children, 74, 174–193.
Fewster, A., & MacMillan, P. D. (2002). School-based evidence for the valid- ity of curriculum-based measurement of reading and writing. Remedial and Special Education, 23, 149–156.
Fitzmaurice, G. M., Laird, N. M., & Ware, J. H. (2004). Applied longitudinal analysis. New York: Wiley.
Fuchs, D., Fuchs, L. S., Mathes, P. G., & Lipsey, M. W. (2000). Reading differences between low-achieving students with and without learning disabilities: A meta-analysis. In R. Gersten, E. Schiller, & S. Vaughn (Eds.), Research syntheses in special education (pp. 81–104). Mahwah, NJ: Erlbaum.
Fuchs, L. S., Fuchs, D., Hamlett, C. L., & Ferguson, C. (1992). Effects of expert system consultation within curriculum-based measurement, using a reading maze task. Exceptional Children, 58, 436–450.
Fuchs, L. S., Fuchs, D., & Maxwell, L. (1988). The validity of informal reading measures. Remedial and Special Education, 9, 20–28.
Griffiths, A. J., VanDerHeydeyn, A. M., Skokut, M., & Lilles, E. (2009). Progress monitoring in oral reading fluency within the context of RTI. School Psychology Quarterly, 24, 13–23.
Hintze, J. M., & Silberglitt, B. (2005). A longitudinal examination of the diagnostic accuracy and predictive validity of R-CBM and high-stakes testing. School Psychology Review, 34, 372–386.
Jenkins, J. R., & Jewell, M. (1993). Examining the validity of two measures for formative teaching: Reading aloud and maze. Exceptional Children, 59, 429–432.
Kenward, M. G., & Roger, J. H. (1997). Small sample inference for fixed effects from restricted maximum likelihood. Biometrics, 53, 983–997.
Kincaid, J. P., Fishburne, R. P., Rogers, R. L., & Chissom, B. S. (1975). Derivation of new readability formulas (Automated Readability In- dex, Fog Count, and Flesch Reading Ease Formula) for Navy enlisted personnel (Research Branch report 8–75). Memphis, TN: Naval Air Station.
Lee, J., Grigg, W., & Donahue, P. (2007). The Nation’s Report Card: Reading 2007 (NCES 2007–496). National Center for Education Statistics, Institute of Education Sciences, U.S. Department of Education, Washington, DC. Retrieved January 16, 2008, from: http://nces.ed.gov/pubsearch/pubsinfo.asp?pubid=2007496.
Levin, E. K., Zigmond, N., & Birch, J. W. (1985). A followup study of 52 learning disabled adolescents. Journal of Learning Disabilities, 18, 2–7.
MacMillan, P. (2000). Simultaneous measurement of reading growth, gen- der, and relative-age effects: Many-faceted Rasch applied to CBM reading scores. Journal of Applied Measurement, 1, 393–408.
Marston, D. (1989). A curriculum-based measurement approach to assessing academic performance: What it is and why do it. In M. Shinn (Ed.), Curriculum-based measurement: Assessing special children (pp. 18– 78). New York: Guilford.
McGlinchey, M. T., & Hixson, M. D. (2004). Using curriculum-based measurement to predict performance on state assessments in reading. School Psychology Review, 33, 193–203.
Minnesota Department of Education. (2001). Minnesota Basic Skills Test Technical Manual. Accountability_Programs/Assessment_and_ Testing/Assessments/BST/BST_Technical_Reports/index.html
Muyskens, P., & Marston, D. (2006). The relationship between Curriculum- Based Measurement and outcomes on high-stakes tests with secondary students. Minneapolis Public Schools. Unpublished manuscript.
O’Connor, R. E., Fulmer, D., Harty, K. R., & Bell, K. M. (2005). Layers of reading intervention in kindergarten through third grade: Changes in teaching and student outcomes. Journal of Learning Disabilities, 38, 440–445.
O’Connor, R. E., Harty, K. R., & Fulmer, D. (2005). Tiers of intervention in kindergarten through third grade. Journal of Learning Disabilities, 38, 532–538.
Rasinski, T. V., Padak, N. D., McKeon, C. A., Wilfong, L. G., Friedauer, J. A., & Heim, P. (2005). Is reading fluency a key for successful high school reading? Journal of Adolescent and Adult Literacy, 48, 22–27.
Ruppert, D., Wand, M. P., & Carroll, R. J. (2003). Semiparametric regression. New York: Cambridge University Press.
Shin, J., Espin, C. A., Deno, S. L., & McConnell, S. (2004). Use of hierarchi- cal linear modeling and curriculum-based measurement for assessing academic growth and instructional factors for students with learning difficulties. Asia Pacific Education Review, 5, 136–148.
Silberglitt, B., & Hintze, J. (2005). Formative assessment using CBM-R cut scores to track progress toward success on state-mandated achieve- ment tests: A comparison of methods. Journal of Psychoeducational Assessment, 23, 304–325.
Stage, S. A., & Jacobsen, M. A. (2001). Predicting student success on a state- mandated performance-based assessment using oral reading fluency. School Psychology Review, 30, 407–419.
Stecker, P. M., Fuchs, L. S., & Fuchs, D. (2005). Using curriculum-based measurement to improve student achievement: Review of research. Psychology in the Schools, 42, 795–819.
Ticha, R., Espin, C. A., & Wayman, M. M. (2009). Reading progress moni- toring for secondary-school students: Reliability, validity, and sensitiv- ity to growth of reading aloud and maze selection measures. Learning Disabilities Research & Practice, 24, 132–142.
Torgesen, J. K. (2000). Individual differences in response to early interven- tions in reading: The lingering problem of treatment resisters. Learning Disabilities Research & Practice, 15, 55–64.
Touchstone Applied Science and Associates. (2006). Degrees of reading power. Brewster, NY: Author.
Vaughn, S., Linan-Thompson, S., & Hickman, P. (2003). Response to in- struction as a means of identifying students with reading/learning dis- abilities. Exceptional Children, 69, 391–409.
Vellutino, F. R., Fletcher, J. M., Snowling, M. J., & Scanlon, D. (2004). Specific reading disability (dyslexia): What have we learned in the past four decades? Journal of Child Psychology and Psychiatry, 45, 2–40.
Vellutino, F. R., Scanlon, D. M., & Tanzman, M. S. (1994). Components of reading ability: Issues and problems in operationalizing word identifi- cation, phonological coding, and orthographic coding. In G. R. Lyon (Ed.), Frames of references for the assessment of learning disabilities: new views on measurement issues (pp. 279–332), Baltimore: Brookes Publishing.
Vellutino, F. R., Tunmer, W. E., Jaccard, J. J., & Chen, R. (2007). Components of reading ability: Multivariate evidence for a convergent skills model of reading development. Scientific Studies of Reading, 11, 3–32.
Warner, M. M., Schumaker, J. B., Alley, G. R., & Deshler, D. D. (1980). Learning disabled adolescents in the public schools: Are they different from other low achievers? Exceptional Education Quarterly, 1(2), 27– 36.
LEARNING DISABILITIES RESEARCH 75
Wayman, M. M., Ticha, R., Wallace, T., Espin, C. A., Wiley, H. I., Du, X., & Long, J. (2009). Comparison of different scoring procedures for CBM maze selection measures. (Technical Report No. 10). Minneapolis, MN: University of Minnesota, Research Institute on Progress Monitoring.
Wayman, M. M., Wallace, T., Wiley, H. I., Ticha, R., & Espin, C. A. (2007). Literature synthesis on curriculum-based measurement in read- ing. Journal of Special Education, 41, 85–120.
Wiley, H. I., & Deno, S. L. (2005). Oral reading and maze measures as pre- dictors of success for English learners on a state standards assessment. Remedial and Special Education, 26, 207–214.
Yovanoff, P., Duesbery, L., Alonzo, J., & Tindal, G. (2005). Grade-level invariance of a theoretical causal structure predicting reading compre- hension with vocabulary and oral reading fluency. Educational Mea- surement: Issues and Practice, 24, 4–12.
About the Authors
Christine Espin is a professor in Education and Child Studies at Leiden University, Leiden, the Netherlands. She is also an adjunct professor in Cognitive Sciences at the University of Minnesota. Her research interests focus on the development of curriculum-based measurement (CBM) procedures in reading, written expression, and content-area learning for secondary students with learning disabilities and teachers’ use of CBM data.
Teri Wallace is an associate professor of Special Education at Minnesota State University in Mankato. Her research focuses on the development of general outcome measures for students with significant cognitive disabilities, implementation of response to intervention and utilization of data in decision making at the student, classroom, school, and district level.
Heather Campbell is an assistant professor of Education at St. Olaf College in Northfield, Minnesota. She works with educational opportunity programs at St. Olaf, and her research interests include the development of CBM procedures in written expression for English language learners.
Erica Lembke is an associate professor in the Department of Special Education at the University of Missouri. Her research interests focus on the development of CBM procedures in reading, written expression, and mathematics for students in early elementary grades as well as implementation of response to intervention in classrooms.
Jeffrey D. Long is an associate professor of Educational Statistics in the Quantitative Methods in Education program in the Department of Educational Psychology, University of Minnesota. His interest is longitudinal data analysis.
Copyright of Learning Disabilities Research & Practice (Blackwell Publishing Limited) is the property of
Wiley-Blackwell and its content may not be copied or emailed to multiple sites or posted to a listserv without
the copyright holder's express written permission. However, users may print, download, or email articles for
individual use.