05978 -2- Pages within 10hrs
Modules chapter 3 wk2 p655
C H A P T E R 3
A Statistics Refresher
From the red-pencil number circled at the top of your first spelling test to the computer printout of your college entrance examination scores, tests and test scores touch your life. They seem to reach out from the paper and shake your hand when you do well and punch you in the face when you do poorly. They can point you toward or away from a particular school or curriculum. They can help you to identify strengths and weaknesses in your physical and mental abilities. They can accompany you on job interviews and influence a job or career choice.
JUST THINK . . .
For most people, test scores are an important fact of life. But what makes those numbers so meaningful? In general terms, what information, ideally, should be conveyed by a test score?
In your role as a student, you have probably found that your relationship to tests has been primarily that of a testtaker. But as a psychologist, teacher, researcher, or employer, you may find that your relationship with tests is primarily that of a test user—the person who breathes life and meaning into test scores by applying the knowledge and skill to interpret them appropriately. You may one day create a test, whether in an academic or a business setting, and then have the responsibility for scoring and interpreting it. In that situation, or even from the perspective of one who would take that test, it’s essential to understand the theory underlying test use and the principles of test-score interpretation.
Test scores are frequently expressed as numbers, and statistical tools are used to describe, make inferences from, and draw conclusions about numbers. 1 In this statistics refresher, we cover scales of measurement, tabular and graphic presentations of data, measures of central tendency, measures of variability, aspects of the normal curve, and standard scores. If these statistics-related terms look painfully familiar to you, we ask your indulgence and ask you to remember that overlearning is the key to retention. Of course, if any of these terms appear unfamiliar, we urge you to learn more about them. Feel free to supplement the discussion here with a review of these and related terms in any good elementary statistics text. The brief review of statistical concepts that follows can in no way replace a sound grounding in basic statistics gained through an introductory course in that subject.Page 76
Scales of Measurement
We may formally define measurement as the act of assigning numbers or symbols to characteristics of things (people, events, whatever) according to rules. The rules used in assigning numbers are guidelines for representing the magnitude (or some other characteristic) of the object being measured. Here is an example of a measurement rule: Assign the number 12 to all lengths that are exactly the same length as a 12-inch ruler. A scale is a set of numbers (or other symbols) whose properties model empirical properties of the objects to which the numbers are assigned. 2
JUST THINK . . .
What is another example of a measurement rule?
There are various ways in which a scale can be categorized. One way of categorizing a scale is according to the type of variable being measured. Thus, a scale used to measure a continuous variable might be referred to as a continuous scale, whereas a scale used to measure a discrete variable might be referred to as a discrete scale. A continuous scale exists when it is theoretically possible to divide any of the values of the scale. A distinction must be made, however, between what is theoretically possible and what is practically desirable. The units into which a continuous scale will actually be divided may depend on such factors as the purpose of the measurement and practicality. In measurement to install venetian blinds, for example, it is theoretically possible to measure by the millimeter or even by the micrometer. But is such precision necessary? Most installers do just fine with measurement by the inch.
As an example of measurement using a discrete scale, consider mental health research that presorted subjects into one of two discrete groups: (1) previously hospitalized and (2) never hospitalized. Such a, categorization scale would be characterized as discrete because it would not be accurate or meaningful to categorize any of the subjects in the study as anything other than “previously hospitalized” or “not previously hospitalized.”
JUST THINK . . .
The scale with which we are all perhaps most familiar is the common bathroom scale. How are a psychological test and a bathroom scale alike? How are they different? Your answer may change as you read on.
JUST THINK . . .
Assume the role of a test creator. Now write some instructions to users of your test that are designed to reduce to the absolute minimum any error associated with test scores. Be sure to include instructions regarding the preparation of the site where the test will be administered.
Measurement always involves error. In the language of assessment, error refers to the collective influence of all of the factors on a test score or measurement beyond those specifically measured by the test or measurement. As we will see, there are many different sources of error in measurement. Consider, for example, the score someone received on a test in American history. We might conceive of part of the score as reflecting the testtaker’s knowledge of American history and part of the score as reflecting error. The error part of the test score may be due to many different factors. One source of error might have been a distracting thunderstorm going on outside at the time the test was administered. Another source of error was the particular selection of test items the instructor chose to use for the test. Had a different item or two been used in the test, the testtaker’s score on the test might have been higher or lower. Error is very much an element of all measurement, and it is an element for which any theory of measurement must surely account.
Measurement using continuous scales always involves error. To illustrate why, let’s go back to the scenario involving venetian Page 77blinds. The length of the window measured to be 35.5 inches could, in reality, be 35.7 inches. The measuring scale is conveniently marked off in grosser gradations of measurement. Most scales used in psychological and educational assessment are continuous and therefore can be expected to contain this sort of error. The number or score used to characterize the trait being measured on a continuous scale should be thought of as an approximation of the “real” number. Thus, for example, a score of 25 on some test of anxiety should not be thought of as a precise measure of anxiety. Rather, it should be thought of as an approximation of the real anxiety score had the measuring instrument been calibrated to yield such a score. In such a case, perhaps the score of 25 is an approximation of a real score of, say, 24.7 or 25.44.
It is generally agreed that there are four different levels or scales of measurement. Within these levels or scales of measurement, assigned numbers convey different kinds of information. Accordingly, certain statistical manipulations may or may not be appropriate, depending upon the level or scale of measurement. 3
JUST THINK . . .
Acronyms like noir are useful memory aids. As you continue in your study of psychological testing and assessment, create your own acronyms to help remember related groups of information. Hey, you may even learn some French in the process.
The French word for black is noir (pronounced “‘nwăre”). We bring this up here only to call attention to the fact that this word is a useful acronym for remembering the four levels or scales of measurement. Each letter in noir is the first letter of the succeedingly more rigorous levels: N stands for nominal, ofor ordinal, i for interval, and r for ratio scales.
Nominal Scales
Nominal scales are the simplest form of measurement. These scales involve classification or categorization based on one or more distinguishing characteristics, where all things measured must be placed into mutually exclusive and exhaustive categories. For example, in the specialty area of clinical psychology, a nominal scale in use for many years is the Diagnostic and Statistical Manual of Mental Disorders. Each disorder listed in that manual is assigned its own number. In a past version of that manual, the version really does not matter for the purposes of this example, the number 303.00 identified alcohol intoxication, and the number 307.00 identified stuttering. But these numbers were used exclusively for classification purposes and could not be meaningfully added, subtracted, ranked, or averaged. Hence, the middle number between these two diagnostic codes, 305.00, did not identify an intoxicated stutterer.
Individual test items may also employ nominal scaling, including yes/no responses. For example, consider the following test items:
Instructions: Answer either yes or no.
Are you actively contemplating suicide? __________
Are you currently under professional care for a psychiatric disorder? _______
Have you ever been convicted of a felony? _______
JUST THINK . . .
What are some other examples of nominal scales?
In each case, a yes or no response results in the placement into one of a set of mutually exclusive groups: suicidal or not, under care for psychiatric disorder or not, and felon or not. Arithmetic operations that can legitimately be performed with Page 78nominal data include counting for the purpose of determining how many cases fall into each category and a resulting determination of proportion or percentages. 4
JUST THINK . . .
What are some other examples of interval scales?
Ordinal Scales
Like nominal scales, ordinal scales permit classification. However, in addition to classification, rank ordering on some characteristic is also permissible with ordinal scales. In business and organizational settings, job applicants may be rank-ordered according to their desirability for a position. In clinical settings, people on a waiting list for psychotherapy may be rank-ordered according to their need for treatment. In these examples, individuals are compared with others and assigned a rank (perhaps 1 to the best applicant or the most needy wait-listed client, 2 to the next, and so forth).
Although he may have never used the term ordinal scale, Alfred Binet, a developer of the intelligence test that today bears his name, believed strongly that the data derived from an intelligence test are ordinal in nature. He emphasized that what he tried to do with his test was not to measure people (as one might measure a person’s height), but merely to classify (and rank) people on the basis of their performance on the tasks. He wrote:
I have not sought . . . to sketch a method of measuring, in the physical sense of the word, but only a method of classification of individuals. The procedures which I have indicated will, if perfected, come to classify a person before or after such another person, or such another series of persons; but I do not believe that one may measure one of the intellectual aptitudes in the sense that one measures a length or a capacity. Thus, when a person studied can retain seven figures after a single audition, one can class him, from the point of his memory for figures, after the individual who retains eight figures under the same conditions, and before those who retain six. It is a classification, not a measurement . . . we do not measure, we classify. (Binet, cited in Varon, 1936, p. 41)
Assessment instruments applied to the individual subject may also use an ordinal form of measurement. The Rokeach Value Survey uses such an approach. In that test, a list of personal values—such as freedom, happiness, and wisdom—are put in order according to their perceived importance to the testtaker (Rokeach, 1973). If a set of 10 values is rank ordered, then the testtaker would assign a value of “1” to the most important and “10” to the least important.
Ordinal scales imply nothing about how much greater one ranking is than another. Even though ordinal scales may employ numbers or “scores” to represent the rank ordering, the numbers do not indicate units of measurement. So, for example, the performance difference between the first-ranked job applicant and the second-ranked applicant may be small while the difference between the second- and third-ranked applicants may be large. On the Rokeach Value Survey, the value ranked “1” may be handily the most important in the mind of the testtaker. However, ordering the values that follow may be difficult to the point of being almost arbitrary.
JUST THINK . . .
What are some other examples of ordinal scales?
Ordinal scales have no absolute zero point. In the case of a test of job performance ability, every testtaker, regardless of standing on the test, is presumed to have some ability. No testtaker is presumed to have zero ability. Zero is without meaning in such a test because the number of units that separate one testtaker’s score from another’s is simply not known. The scores are ranked, but the actual number of units separating one score from the next may be many, just a few, or practically none. Because there is no zero point on an ordinal scale, the ways in which data from such scales can be analyzed statistically are limited. One cannot average the qualifications of the Page 79first- and third-ranked job applicants, for example, and expect to come out with the qualifications of the second-ranked applicant.
Interval Scales
In addition to the features of nominal and ordinal scales, interval scales contain equal intervals between numbers. Each unit on the scale is exactly equal to any other unit on the scale. But like ordinal scales, interval scales contain no absolute zero point. With interval scales, we have reached a level of measurement at which it is possible to average a set of measurements and obtain a meaningful result.
Scores on many tests, such as tests of intelligence, are analyzed statistically in ways appropriate for data at the interval level of measurement. The difference in intellectual ability represented by IQs of 80 and 100, for example, is thought to be similar to that existing between IQs of 100 and 120. However, if an individual were to achieve an IQ of 0 (something that is not even possible, given the way most intelligence tests are structured), that would not be an indication of zero (the total absence of) intelligence. Because interval scales contain no absolute zero point, a presumption inherent in their use is that no testtaker possesses none of the ability or trait (or whatever) being measured.
Ratio Scales
In addition to all the properties of nominal, ordinal, and interval measurement, a ratio scale has a true zero point. All mathematical operations can meaningfully be performed because there exist equal intervals between the numbers on the scale as well as a true or absolute zero point.
In psychology, ratio-level measurement is employed in some types of tests and test items, perhaps most notably those involving assessment of neurological functioning. One example is a test of hand grip, where the variable measured is the amount of pressure a person can exert with one hand (see Figure 3–1 ). Another example is a timed test of perceptual-motor ability that requires the testtaker to assemble a jigsaw-like puzzle. In such an instance, the time taken to successfully complete the puzzle is the measure that is recorded. Because there is a true zero point on this scale (or, 0 seconds), it is meaningful to say that a testtaker who completes the assembly in 30 seconds has taken half the time of a testtaker who completed it in 60 seconds. In this example, it is meaningful to speak of a true zero point on the scale—but in theory only. Why? Just think . . .
Figure 3–1 Ratio-Level Measurement in the Palm of One’s Hand Pictured above is a dynamometer, an instrument used to measure strength of hand grip. The examinee is instructed to squeeze the grips as hard as possible. The squeezing of the grips causes the gauge needle to move and reflect the number of pounds of pressure exerted. The highest point reached by the needle is the score. This is an example of ratio-level measurement. Someone who can exert 10 pounds of pressure (and earns a score of 10) exerts twice as much pressure as a person who exerts 5 pounds of pressure (and earns a score of 5). On this test it is possible to achieve a score of 0, indicating a complete lack of exerted pressure. Although it is meaningful to speak of a score of 0 on this test, we have to wonder about its significance. How might a score of 0 result? One way would be if the testtaker genuinely had paralysis of the hand. Another way would be if the testtaker was uncooperative and unwilling to comply with the demands of the task. Yet another way would be if the testtaker was attempting to malinger or “fake bad” on the test. Ratio scales may provide us “solid” numbers to work with, but some interpretation of the test data yielded may still be required before drawing any “solid” conclusions.© BanksPhotos/Getty Images RF
JUST THINK . . .
What are some other examples of ratio scales?
No testtaker could ever obtain a score of zero on this assembly task. Stated another way, no testtaker, not even The Flash (a comic-book superhero whose power is the ability to move at superhuman speed), could assemble the puzzle in zero seconds.
Measurement Scales in Psychology
The ordinal level of measurement is most frequently used in psychology. As Kerlinger (1973, p. 439) put it: “Intelligence, aptitude, and personality test scores are, basically and strictly speaking, ordinal. These tests indicate with more or less accuracy not the amount of intelligence, aptitude, and personality traits of individuals, but rather the rank-order positions of the individuals.” Kerlinger allowed that “most psychological and educational scales approximate interval equality fairly well,” though he cautioned that if ordinal measurements are treated as if they were interval measurements, then the test user must “be constantly alert to the possibility of gross inequality of intervals” (pp. 440–441).Page 80
Why would psychologists want to treat their assessment data as interval when those data would be better described as ordinal? Why not just say that they are ordinal? The attraction of interval measurement for users of psychological tests is the flexibility with which such data can be manipulated statistically. “What kinds of statistical manipulation?” you may ask.
In this chapter we discuss the various ways in which test data can be described or converted to make those data more manageable and understandable. Some of the techniques we’ll describe, such as the computation of an average, can be used if data are assumed to be interval- or ratio-level in nature but not if they are ordinal- or nominal-level. Other techniques, such as those involving the creation of graphs or tables, may be used with ordinal- or even nominal-level data.Page 81
Describing Data
Suppose you have magically changed places with the professor teaching this course and that you have just administered an examination that consists of 100 multiple-choice items (where 1 point is awarded for each correct answer). The distribution of scores for the 25 students enrolled in your class could theoretically range from 0 (none correct) to 100 (all correct). A distribution may be defined as a set of test scores arrayed for recording or study. The 25 scores in this distribution are referred to as raw scores. As its name implies, a raw score is a straightforward, unmodified accounting of performance that is usually numerical. A raw score may reflect a simple tally, as in number of items responded to correctly on an achievement test. As we will see later in this chapter, raw scores can be converted into other types of scores. For now, let’s assume it’s the day after the examination and that you are sitting in your office looking at the raw scores listed in Table 3–1 . What do you do next?
|
Student |
Score (number correct) |
|
Judy |
78 |
|
Joe |
67 |
|
Lee-Wu |
69 |
|
Miriam |
63 |
|
Valerie |
85 |
|
Diane |
72 |
|
Henry |
92 |
|
Esperanza |
67 |
|
Paula |
94 |
|
Martha |
62 |
|
Bill |
61 |
|
Homer |
44 |
|
Robert |
66 |
|
Michael |
87 |
|
Jorge |
76 |
|
Mary |
83 |
|
“Mousey” |
42 |
|
Barbara |
82 |
|
John |
84 |
|
Donna |
51 |
|
Uriah |
69 |
|
Leroy |
61 |
|
Ronald |
96 |
|
Vinnie |
73 |
|
Bianca |
79 |
|
Table 3–1 Data from Your Measurement Course Test |
JUST THINK . . .
In what way do most of your instructors convey test-related feedback to students? Is there a better way they could do this?
One task at hand is to communicate the test results to your class. You want to do that in a way that will help students understand how their performance on the test compared to the performance of other students. Perhaps the first step is to organize the data by transforming it from a random listing of raw scores into something that immediately conveys a bit more information. Later, as we will see, you may wish to transform the data in other ways.
Frequency Distributions
The data from the test could be organized into a distribution of the raw scores. One way the scores could be distributed is by the frequency with which they occur. In a frequency distribution , all scores are listed alongside the number of times each score occurred. The scores might be listed in tabular or graphic form. Table 3–2 lists the frequency of occurrence of each score in one column and the score itself in the other column.
|
Score |
f (frequency) |
|
96 |
1 |
|
94 |
1 |
|
92 |
1 |
|
87 |
1 |
|
85 |
1 |
|
84 |
1 |
|
83 |
1 |
|
82 |
1 |
|
79 |
1 |
|
78 |
1 |
|
76 |
1 |
|
73 |
1 |
|
72 |
1 |
|
69 |
2 |
|
67 |
2 |
|
66 |
1 |
|
63 |
1 |
|
62 |
1 |
|
61 |
2 |
|
51 |
1 |
|
44 |
1 |
|
42 |
1 |
|
Table 3–2 Frequency Distribution of Scores from Your Test |
Often, a frequency distribution is referred to as a simple frequency distribution to indicate that individual scores have been used and the data have not been grouped. Another kind of Page 82frequency distribution used to summarize data is a grouped frequency distribution. In a grouped frequency distribution , test-score intervals, also called class intervals, replace the actual test scores. The number of class intervals used and the size or width of each class interval (or, the range of test scores contained in each class interval) are for the test user to decide. But how?
In most instances, a decision about the size of a class interval in a grouped frequency distribution is made on the basis of convenience. Of course, virtually any decision will represent a trade-off of sorts. A convenient, easy-to-read summary of the data is the trade-off for the loss of detail. To what extent must the data be summarized? How important is detail? These types of questions must be considered. In the grouped frequency distribution in Table 3–3 , the test scores have been grouped into 12 class intervals, where each class interval is equal to 5 points. 5 The highest class interval (95–99) and the lowest class interval (40–44) are referred to, respectively, as the upper and lower limits of the distribution. Here, the need for convenience in reading the data outweighs the need for great detail, so such groupings of data seem logical.
|
Class Interval |
f (frequency) |
|
95–99 |
1 |
|
90–94 |
2 |
|
85–89 |
2 |
|
80–84 |
3 |
|
75–79 |
3 |
|
70–74 |
2 |
|
65–69 |
5 |
|
60–64 |
4 |
|
55–59 |
0 |
|
50–54 |
1 |
|
45–49 |
0 |
|
40–44 |
2 |
|
Table 3–3 A Grouped Frequency Distribution |
Frequency distributions of test scores can also be illustrated graphically. A graph is a diagram or chart composed of lines, points, bars, or other symbols that describe and illustrate Page 83data. With a good graph, the place of a single score in relation to a distribution of test scores can be understood easily. Three kinds of graphs used to illustrate frequency distributions are the histogram, the bar graph, and the frequency polygon ( Figure 3–2 ). A histogram is a graph Page 84with vertical lines drawn at the true limits of each test score (or class interval), forming a series of contiguous rectangles. It is customary for the test scores (either the single scores or the midpoints of the class intervals) to be placed along the graph’s horizontal axis (also referred to as the abscissa or X-axis) and for numbers indicative of the frequency of occurrence to be placed along the graph’s vertical axis (also referred to as the ordinate or Y-axis). In a bar graph , numbers indicative of frequency also appear on the Y-axis, and reference to some categorization (e.g., yes/no/maybe, male/female) appears on the X-axis. Here the rectangular bars typically are not contiguous. Data illustrated in a frequency polygon are expressed by a continuous line connecting the points where test scores or class intervals (as indicated on the X-axis) meet frequencies (as indicated on the Y-axis).
Figure 3–2 Graphic Illustrations of Data from Table 3–3 A histogram (a), a bar graph (b), and a frequency polygon (c) all may be used to graphically convey information about test performance. Of course, the labeling of the bar graph and the specific nature of the data conveyed by it depend on the variables of interest. In (b), the variable of interest is the number of students who passed the test (assuming, for the purpose of this illustration, that a raw score of 65 or higher had been arbitrarily designated in advance as a passing grade). Returning to the question posed earlier—the one in which you play the role of instructor and must communicate the test results to your students—which type of graph would best serve your purpose? Why? As we continue our review of descriptive statistics, you may wish to return to your role of professor and formulate your response to challenging related questions, such as “Which measure(s) of central tendency shall I use to convey this information?” and “Which measure(s) of variability would convey the information best?”
Graphic representations of frequency distributions may assume any of a number of different shapes ( Figure 3–3 ). Regardless of the shape of graphed data, it is a good idea for the consumer of the information contained in the graph to examine it carefully—and, if need be, critically. Consider, in this context, this chapter’s Everyday Psychometrics.
Figure 3–3 Shapes That Frequency Distributions Can Take
As we discuss in detail later in this chapter, one graphic representation of data of particular interest to measurement professionals is the normal or bell-shaped curve. Before getting to that, however, let’s return to the subject of distributions and how we can describe and characterize them. One way to describe a distribution of test scores is by a measure of central tendency.
Measures of Central Tendency
A measure of central tendency is a statistic that indicates the average or midmost score between the extreme scores in a distribution. The center of a distribution can be defined in different ways. Perhaps the most commonly used measure of central tendency is the arithmetic mean (or, more simply, mean ), which is referred to in everyday language as the “average.” The mean takes into account the actual numerical value of every score. In special instances, such as when there are only a few scores and one or two of the scores are extreme in relation to the remaining ones, a measure of central tendency other than the mean may be desirable. Other measures of central tendency we review include the median and the mode. Note that, in the formulas to follow, the standard statistical shorthand called “summation notation” (summation meaning “the sum of”) is used. The Greek uppercase letter sigma, Σ, is the symbol used to signify “sum”; if X represents a test score, then the expression Σ X means “add all the test scores.”
The arithmetic mean
The arithmetic mean , denoted by the symbol (and pronounced “X bar”), is equal to the sum of the observations (or test scores, in this case) divided by the number of observations. Symbolically written, the formula for the arithmetic mean is where n equals the number of observations or test scores. The arithmetic mean is typically the most appropriate measure of central tendency for interval or ratio data when the distributions are believed to be approximately normal. An arithmetic mean can also be computed from a frequency distribution. The formula for doing this is
where Σ(fX) means “multiply the frequency of each score by its corresponding score and then sum.” An estimate of the arithmetic mean may also be obtained from a grouped frequency distribution using the same formula, where X is equal to the midpoint of the class interval. Table 3–4 illustrates a calculation of the mean from a grouped frequency distribution. After doing the math you will find that, using the grouped data, a mean of 71.8 (which may be rounded to 72) is calculated. Using the raw scores, a mean of 72.12 (which also may be rounded to 72) is calculated. Frequently, the choice of statistic will depend on the required degree of precision in measurement.
|
Class Interval |
f |
X (midpoint of class interval) |
fX |
|
95–99 |
1 |
97 |
97 |
|
90–94 |
2 |
92 |
184 |
|
85–89 |
2 |
87 |
174 |
|
80–84 |
3 |
82 |
246 |
|
75–79 |
3 |
77 |
231 |
|
70–74 |
2 |
72 |
144 |
|
65–69 |
5 |
67 |
335 |
|
60–64 |
4 |
62 |
248 |
|
55–59 |
0 |
57 |
000 |
|
50–54 |
1 |
52 |
52 |
|
45–49 |
0 |
47 |
000 |
|
40–44 |
2 |
42 |
84 |
|
|
Σ f = 25 |
|
Σ (fX) = 1,795 |
|
Table 3–4 Calculating the Arithmetic Mean from a Grouped Frequency Distribution |
To estimate the arithmetic mean of this grouped frequency distribution,
To calculate the mean of this distribution using raw scores,
JUST THINK . . .
Imagine that a thousand or so engineers took an extremely difficult pre-employment test. A handful of the engineers earned very high scores but the vast majority did poorly, earning extremely low scores. Given this scenario, what are the pros and cons of using the mean as a measure of central tendency for this test?
Page 85
EVERYDAY PSYCHOMETRICS
Consumer (of Graphed Data), Beware!
One picture is worth a thousand words, and one purpose of representing data in graphic form is to convey information at a glance. However, although two graphs may be accurate with respect to the data they represent, their pictures—and the impression drawn from a glance at them—may be vastly different. As an example, consider the following hypothetical scenario involving a hamburger restaurant chain we’ll call “The Charred House.”
The Charred House chain serves very charbroiled, microscopically thin hamburgers formed in the shape of little triangular houses. In the 10-year period since its founding in 1993, the company has sold, on average, 100 million burgers per year. On the chain’s tenth anniversary, The Charred House distributes a press release proudly announcing “Over a Billion Served.”
Reporters from two business publications set out to research and write a feature article on this hamburger restaurant chain. Working solely from sales figures as compiled from annual reports to the shareholders, Reporter 1 focuses her story on the differences in yearly sales. Her article is entitled “A Billion Served—But Charred House Sales Fluctuate from Year to Year,” and its graphic illustration is reprinted here.
Quite a different picture of the company emerges from Reporter 2’s story, entitled “A Billion Served—And Charred House Sales Are as Steady as Ever,” and its accompanying graph. The latter story is based on a diligent analysis of comparable data for the same number of hamburger chains in the same areas of the country over the same time period. While researching the story, Reporter 2 learned that yearly fluctuations in sales are common to the entire industry and that the annual fluctuations observed in the Charred House figures were—relative to other chains—insignificant.
Compare the graphs that accompanied each story. Although both are accurate insofar as they are based on the correct numbers, the impressions they are likely to leave are quite different.
Incidentally, custom dictates that the intersection of the two axes of a graph be at 0 and that all the points on the Y-axis be in equal and proportional intervals from 0. This custom is followed in Reporter 2’s story, where the first point on the ordinate is 10 units more than 0, and each succeeding point is also 10 more units away from 0. However, the custom is violated in Reporter 1’s story, where the first point on the ordinate is 95 units more than 0, and each succeeding point increases only by 1. The fact that the custom is violated in Reporter 1’s story should serve as a warning to evaluate pictorial representations of data all the more critically.
Page 87
The median
The median , defined as the middle score in a distribution, is another commonly used measure of central tendency. We determine the median of a distribution of scores by ordering the scores in a list by magnitude, in either ascending or descending order. If the total number of scores ordered is an odd number, then the median will be the score that is exactly in the middle, with one-half of the remaining scores lying above it and the other half of the remaining scores lying below it. When the total number of scores ordered is an even number, then the median can be calculated by determining the arithmetic mean of the two middle scores. For example, suppose that 10 people took a preemployment word-processing test at The Rochester Wrenchworks (TRW) Corporation. They obtained the following scores, presented here in descending order:
66 65 61 59 53 52 41 36 35 32
Page 88
The median of these data would be calculated by obtaining the average (or, the arithmetic mean) of the two middle scores, 53 and 52 (which would be equal to 52.5). The median is an appropriate measure of central tendency for ordinal, interval, and ratio data. The median may be a particularly useful measure of central tendency in cases where relatively few scores fall at the high end of the distribution or relatively few scores fall at the low end of the distribution.
Suppose not 10 but rather tens of thousands of people had applied for jobs at The Rochester Wrenchworks. It would be impractical to find the median by simply ordering the data and finding the midmost scores, so how would the median score be identified? For our purposes, the answer is simply that there are advanced methods for doing so. There are also techniques for identifying the median in other sorts of distributions, such as a grouped frequency distribution and a distribution wherein various scores are identical. However, instead of delving into such new and complex territory, let’s resume our discussion of central tendency and consider another such measure.
The mode
The most frequently occurring score in a distribution of scores is the mode . 6 As an example, determine the mode for the following scores obtained by another TRW job applicant, Bruce. The scores reflect the number of words Bruce word-processed in seven 1-minute trials:
43 34 45 51 42 31 51
It is TRW policy that new hires must be able to word-process at least 50 words per minute. Now, place yourself in the role of the corporate personnel officer. Would you hire Bruce? The most frequently occurring score in this distribution of scores is 51. If hiring guidelines gave you the freedom to use any measure of central tendency in your personnel decision making, then it would be your choice as to whether or not Bruce is hired. You could hire him and justify this decision on the basis of his modal score (51). You also could not hire him and justify this decision on the basis of his mean score (below the required 50 words per minute). Ultimately, whether Rochester Wrenchworks will be Bruce’s new home away from home will depend on other job-related factors, such as the nature of the job market in Rochester and the qualifications of competing applicants. Of course, if company guidelines dictate that only the mean score be used in hiring decisions, then a career at TRW is not in Bruce’s immediate future.
Distributions that contain a tie for the designation “most frequently occurring score” can have more than one mode. Consider the following scores—arranged in no particular order—obtained by 20 students on the final exam of a new trade school called the Home Study School of Elvis Presley Impersonators:
51 49 51 50 66 52 53 38 17 66 33 44 73 13 21 91 87 92 47 3
These scores are said to have a bimodal distribution because there are two scores (51 and 66) that occur with the highest frequency (of two). Except with nominal data, the mode tends not to be a very commonly used measure of central tendency. Unlike the arithmetic mean, which has to be calculated, the value of the modal score is not calculated; one simply counts and determines which score occurs most frequently. Because the mode is arrived at in this manner, the modal score may be totally atypical—for instance, one at an extreme end of the distribution—which nonetheless occurs with the greatest frequency. In fact, it is theoretically possible for a bimodal distribution to have two modes, each of which falls at the high or the low end of the distribution—thus violating the expectation that a measure of central tendency should be . . . well, central (or indicative of a point at the middle of the distribution).Page 89
Even though the mode is not calculated in the sense that the mean is calculated, and even though the mode is not necessarily a unique point in a distribution (a distribution can have two, three, or even more modes), the mode can still be useful in conveying certain types of information. The mode is useful in analyses of a qualitative or verbal nature. For example, when assessing consumers’ recall of a commercial by means of interviews, a researcher might be interested in which word or words were mentioned most by interviewees.
The mode can convey a wealth of information in addition to the mean. As an example, suppose you wanted an estimate of the number of journal articles published by clinical psychologists in the United States in the past year. To arrive at this figure, you might total the number of journal articles accepted for publication written by each clinical psychologist in the United States, divide by the number of psychologists, and arrive at the arithmetic mean. This calculation would yield an indication of the average number of journal articles published. Whatever that number would be, we can say with certainty that it would be more than the mode. It is well known that most clinical psychologists do not write journal articles. The mode for publications by clinical psychologists in any given year is zero. In this example, the arithmetic mean would provide us with a precise measure of the average number of articles published by clinicians. However, what might be lost in that measure of central tendency is that, proportionately, very few of all clinicians do most of the publishing. The mode (in this case, a mode of zero) would provide us with a great deal of information at a glance. It would tell us that, regardless of the mean, most clinicians do not publish.
Because the mode is not calculated in a true sense, it is a nominal statistic and cannot legitimately be used in further calculations. The median is a statistic that takes into account the order of scores and is itself ordinal in nature. The mean, an interval-level statistic, is generally the most stable and useful measure of central tendency.
JUST THINK . . .
Devise your own example to illustrate how the mode, and not the mean, can be the most useful measure of central tendency.
Measures of Variability
Variability is an indication of how scores in a distribution are scattered or dispersed. As Figure 3–4 illustrates, two or more distributions of test scores can have the same mean even though differences in the dispersion of scores around the mean can be wide. In both distributions A and B, test scores could range from 0 to 100. In distribution A, we see that the mean score was 50 and the remaining scores were widely distributed around the mean. In distribution B, the mean was also 50 but few people scored higher than 60 or lower than 40.
Figure 3–4 Two Distributions with Differences in Variability
Statistics that describe the amount of variation in a distribution are referred to as measures of variability . Some measures of variability include the range, the interquartile range, the semi-interquartile range, the average deviation, the standard deviation, and the variance.Page 90
The range
The range of a distribution is equal to the difference between the highest and the lowest scores. We could describe distribution B of Figure 3–3 , for example, as having a range of 20 if we knew that the highest score in this distribution was 60 and the lowest score was 40 (60 − 40 = 20). With respect to distribution A, if we knew that the lowest score was 0 and the highest score was 100, the range would be equal to 100 − 0, or 100. The range is the simplest measure of variability to calculate, but its potential use is limited. Because the range is based entirely on the values of the lowest and highest scores, one extreme score (if it happens to be the lowest or the highest) can radically alter the value of the range. For example, suppose distribution B included a score of 90. The range of this distribution would now be equal to 90 − 40, or 50. Yet, in looking at the data in the graph for distribution B, it is clear that the vast majority of scores tend to be between 40 and 60.
JUST THINK . . .
Devise two distributions of test scores to illustrate how the range can overstate or understate the degree of variability in the scores.
As a descriptive statistic of variation, the range provides a quick but gross description of the spread of scores. When its value is based on extreme scores in a distribution, the resulting description of variation may be understated or overstated. Better measures of variation include the interquartile range and the semi-interquartile range.
The interquartile and semi-interquartile ranges
A distribution of test scores (or any other data, for that matter) can be divided into four parts such that 25% of the test scores occur in each quarter. As illustrated in Figure 3–5 , the dividing points between the four quarters in the distribution are the quartiles . There are three of them, respectively labeled Q1, Q2, and Q3. Note that quartile refers to a specific point whereas quarter refers to an interval. An individual score may, for example, fall at the third quartile or in the third quarter (but not “in” the third quartile or “at” the third quarter). It should come as no surprise to you that Q2 and the median are exactly the same. And just as the median is the midpoint in a distribution of scores, so are quartiles Q1 and Q3 the quarter-points in a distribution of scores. Formulas may be employed to determine the exact value of these points.
Figure 3–5 A Quartered DistributionPage 91
The interquartile range is a measure of variability equal to the difference between Q3 and Q1. Like the median, it is an ordinal statistic. A related measure of variability is the semi-interquartile range , which is equal to the interquartile range divided by 2. Knowledge of the relative distances of Q1 and Q3 from Q2 (the median) provides the seasoned test interpreter with immediate information as to the shape of the distribution of scores. In a perfectly symmetrical distribution, Q1 and Q3 will be exactly the same distance from the median. If these distances are unequal then there is a lack of symmetry. This lack of symmetry is referred to as skewness, and we will have more to say about that shortly.
The average deviation
Another tool that could be used to describe the amount of variability in a distribution is the average deviation , or AD for short. Its formula is
The lowercase italic x in the formula signifies a score’s deviation from the mean. The value of x is obtained by subtracting the mean from the score (X − mean = x). The bars on each side of x indicate that it is the absolute value of the deviation score (ignoring the positive or negative sign and treating all deviation scores as positive). All the deviation scores are then summed and divided by the total number of scores (n) to arrive at the average deviation. As an exercise, calculate the average deviation for the following distribution of test scores:
85 100 90 95 80
Begin by calculating the arithmetic mean. Next, obtain the absolute value of each of the five deviation scores and sum them. As you sum them, note what would happen if you did not ignore the plus or minus signs: All the deviation scores would then sum to 0. Divide the sum of the deviation scores by the number of measurements (5). Did you obtain an AD of 6? The AD tells us that the five scores in this distribution varied, on average, 6 points from the mean.
JUST THINK . . .
After reading about the standard deviation, explain in your own words how an understanding of the average deviation can provide a “stepping-stone” to better understanding the concept of a standard deviation.
The average deviation is rarely used. Perhaps this is so because the deletion of algebraic signs renders it a useless measure for purposes of any further operations. Why, then, discuss it here? The reason is that a clear understanding of what an average deviation measures provides a solid foundation for understanding the conceptual basis of another, more widely used measure: the standard deviation. Keeping in mind what an average deviation is, what it tells us, and how it is derived, let’s consider its more frequently used “cousin,” the standard deviation.
The standard deviation
Recall that, when we calculated the average deviation, the problem of the sum of all deviation scores around the mean equaling zero was solved by employing only the absolute value of the deviation scores. In calculating the standard deviation, the same problem must be dealt with, but we do so in a different way. Instead of using the absolute value of each deviation score, we use the square of each score. With each score squared, the sign of any negative deviation becomes positive. Because all the deviation scores are squared, we know that our calculations won’t be complete until we go back and obtain the square root of whatever value we reach.
We may define the standard deviation as a measure of variability equal to the square root of the average squared deviations about the mean. More succinctly, it is equal to the square root of the variance. The variance is equal to the arithmetic mean of the squares of the differences between the scores in a distribution and their mean. The formula used to calculate the variance (s2) using deviation scores is
Page 92
Simply stated, the variance is calculated by squaring and summing all the deviation scores and then dividing by the total number of scores. The variance can also be calculated in other ways. For example: From raw scores, first calculate the summation of the raw scores squared, divide by the number of scores, and then subtract the mean squared. The result is
The variance is a widely used measure in psychological research. To make meaningful interpretations, the test-score distribution should be approximately normal. We’ll have more to say about “normal” distributions later in the chapter. At this point, think of a normal distribution as a distribution with the greatest frequency of scores occurring near the arithmetic mean. Correspondingly fewer and fewer scores relative to the mean occur on both sides of it.
For some hands-on experience with—and to develop a sense of mastery of—the concepts of variance and standard deviation, why not allot the next 10 or 15 minutes to calculating the standard deviation for the test scores shown in Table 3–1 ? Use both formulas to verify that they produce the same results. Using deviation scores, your calculations should look similar to these:
Using the raw-scores formula, your calculations should look similar to these:
In both cases, the standard deviation is the square root of the variance (s2). According to our calculations, the standard deviation of the test scores is 14.10. If s = 14.10, then 1 standard deviation unit is approximately equal to 14 units of measurement or (with reference to our example and rounded to a whole number) to 14 test-score points. The test data did not provide a good normal curve approximation. Test professionals would describe these data as “positively skewed.” Skewness, as well as related terms such as negatively skewed and positively skewed, are covered in the next section. Once you are “positively familiar” with terms like positively skewed, you’ll appreciate all the more the section later in this chapter entitled “The Area Under the Normal Curve.” There you will find a wealth of information about test-score interpretation Page 93in the case when the scores are not skewed—that is, when the test scores are approximately normal in distribution.
The symbol for standard deviation has variously been represented as s, S, SD, and the lowercase Greek letter sigma (σ). One custom (the one we adhere to) has it that s refers to the sample standard deviation and σ refers to the population standard deviation. The number of observations in the sample is n, and the denominator n − 1 is sometimes used to calculate what is referred to as an “unbiased estimate” of the population value (though it’s actually only lessbiased; see Hopkins & Glass, 1978). Unless n is 10 or less, the use of n or n − 1 tends not to make a meaningful difference.
Whether the denominator is more properly n or n − 1 has been a matter of debate. Lindgren (1983) has argued for the use of n − 1, in part because this denominator tends to make correlation formulas simpler. By contrast, most texts recommend the use of n − 1 only when the data constitute a sample; when the data constitute a population, n is preferable. For Lindgren (1983), it doesn’t matter whether the data are from a sample or a population. Perhaps the most reasonable convention is to use n either when the entire population has been assessed or when no inferences to the population are intended. So, when considering the examination scores of one class of students—including all the people about whom we’re going to make inferences—it seems appropriate to use n.
Having stated our position on the n versus n − 1 controversy, our formula for the population standard deviation follows. In this formula, represents a sample mean and M a population mean:
The standard deviation is a very useful measure of variation because each individual score’s distance from the mean of the distribution is factored into its computation. You will come across this measure of variation frequently in the study and practice of measurement in psychology.
Skewness
Distributions can be characterized by their skewness , or the nature and extent to which symmetry is absent. Skewness is an indication of how the measurements in a distribution are distributed. A distribution has a positive skew when relatively few of the scores fall at the high end of the distribution. Positively skewed examination results may indicate that the test was too difficult. More items that were easier would have been desirable in order to better discriminate at the lower end of the distribution of test scores. A distribution has a negative skew when relatively few of the scores fall at the low end of the distribution. Negatively skewed examination results may indicate that the test was too easy. In this case, more items of a higher level of difficulty would make it possible to better discriminate between scores at the upper end of the distribution. (Refer to Figure 3–3 for graphic examples of skewed distributions.)
The term skewed carries with it negative implications for many students. We suspect that skewed is associated with abnormal, perhaps because the skewed distribution deviates from the symmetrical or so-called normal distribution. However, the presence or absence of symmetry in a distribution (skewness) is simply one characteristic by which a distribution can be described. Consider in this context a hypothetical Marine Corps Ability and Endurance Screening Test administered to all civilians seeking to enlist in the U.S. Marines. Now look again at the graphs in Figure 3–3 . Which graph do you think would best describe the resulting distribution of test scores? (No peeking at the next paragraph before you respond.)
No one can say with certainty, but if we had to guess, then we would say that the Marine Corps Ability and Endurance Screening Test data would look like graph C, the positively skewed distribution in Figure 3–3 . We say this assuming that a level of difficulty would have been built into the test to ensure that relatively few assessees would score at the high end of Page 94the distribution. Most of the applicants would probably score at the low end of the distribution. All of this is quite consistent with the advertised objective of the Marines, who are only looking for a few good men. You know: the few, the proud. Now, a question regarding this positively skewed distribution: Is the skewness a good thing? A bad thing? An abnormal thing? In truth, it is probably none of these things—it just is. By the way, although they may not advertise it as much, the Marines are also looking for (an unknown quantity of) good women. But here we are straying a bit too far from skewness.
Various formulas exist for measuring skewness. One way of gauging the skewness of a distribution is through examination of the relative distances of quartiles from the median. In a positively skewed distribution, Q3 − Q2 will be greater than the distance of Q2 − Q1. In a negatively skewed distribution, Q3− Q2 will be less than the distance of Q2 − Q1. In a distribution that is symmetrical, the distances from Q1 and Q3 to the median are the same.
Kurtosis
The term testing professionals use to refer to the steepness of a distribution in its center is kurtosis . To the root kurtic is added to one of the prefixes platy-, lepto-, or meso- to describe the peakedness/flatness of three general types of curves ( Figure 3–6 ). Distributions are generally described as platykurtic (relatively flat), leptokurtic (relatively peaked), or—somewhere in the middle— mesokurtic . Distributions that have high kurtosis are characterized by a high peak and “fatter” tails compared to a normal distribution. In contrast, lower kurtosis values indicate a distribution with a rounded peak and thinner tails. Many methods exist for measuring kurtosis. According to the original definition, the normal bell-shaped curve (see graph A from Figure 3–3 ) would have a kurtosis value of 3. In other methods of computing kurtosis, a normal distribution would have kurtosis of 0, with positive values indicating higher kurtosis and negative values indicating lower kurtosis. It is important to keep the different methods of calculating kurtosis in mind when examining the values reported by researchers or computer programs. So, given that this can quickly become an advanced-level topic and that this book is of a more introductory nature, let’s move on. It’s time to focus on a type of distribution that happens to be the standard against which all other distributions (including all of the kurtic ones) are compared: the normal distribution.
Figure 3–6 The Kurtosis of Curves
JUST THINK . . .
Like skewness, reference to the kurtosis of a distribution can provide a kind of “shorthand” description of a distribution of test scores. Imagine and describe the kind of test that might yield a distribution of scores that form a platykurtic curve.
Page 95
The Normal Curve
Before delving into the statistical, a little bit of the historical is in order. Development of the concept of a normal curve began in the middle of the eighteenth century with the work of Abraham DeMoivre and, later, the Marquis de Laplace. At the beginning of the nineteenth century, Karl Friedrich Gauss made some substantial contributions. Through the early nineteenth century, scientists referred to it as the “Laplace-Gaussian curve.” Karl Pearson is credited with being the first to refer to the curve as the normal curve, perhaps in an effort to be diplomatic to all of the people who helped develop it. Somehow the term normal curve stuck—but don’t be surprised if you’re sitting at some scientific meeting one day and you hear this distribution or curve referred to as Gaussian.
Theoretically, the normal curve is a bell-shaped, smooth, mathematically defined curve that is highest at its center. From the center it tapers on both sides approaching the X-axis asymptotically (meaning that it approaches, but never touches, the axis). In theory, the distribution of the normal curve ranges from negative infinity to positive infinity. The curve is perfectly symmetrical, with no skewness. If you folded it in half at the mean, one side would lie exactly on top of the other. Because it is symmetrical, the mean, the median, and the mode all have the same exact value.
Why is the normal curve important in understanding the characteristics of psychological tests? Our Close-Up provides some answers.
The Area Under the Normal Curve
The normal curve can be conveniently divided into areas defined in units of standard deviation. A hypothetical distribution of National Spelling Test scores with a mean of 50 and a standard deviation of 15 is illustrated in Figure 3–7 . In this example, a score equal to 1 standard deviation above the mean would be equal to 65 (+ 1s = 50 + 15 = 65).
Figure 3–7 The Area Under the Normal CurvePage 96
CLOSE-UP
The Normal Curve and Psychological Tests
Scores on many psychological tests are often approximately normally distributed, particularly when the tests are administered to large numbers of subjects. Few, if any, psychological tests yield precisely normal distributions of test scores (Micceri, 1989). As a general rule (with ample exceptions), the larger the sample size and the wider the range of abilities measured by a particular test, the more the graph of the test scores will approximate the normal curve. A classic illustration of this was provided by E. L. Thorndike and his colleagues (1927). They compiled intelligence test scores from several large samples of students. As you can see in Figure 1 , the distribution of scores closely approximated the normal curve.
Figure 1 Graphic Representation of Thorndike et al. Data The solid line outlines the distribution of intelligence test scores of sixth-grade students (N = 15,138). The dotted line is the theoretical normal curve (Thorndike et al., 1927).
Following is a sample of more varied examples of the wide range of characteristics that psychologists have found to be approximately normal in distribution.
· The strength of handedness in right-handed individuals, as measured by the Waterloo Handedness Questionnaire (Tan, 1993).
· Scores on the Women’s Health Questionnaire, a scale measuring a variety of health problems in women across a wide age range (Hunter, 1992).
· Responses of both college students and working adults to a measure of intrinsic and extrinsic work motivation (Amabile et al., 1994).
· The intelligence-scale scores of girls and women with eating disorders, as measured by the Wechsler Adult Intelligence Scale–Revised and the Wechsler Intelligence Scale for Children–Revised (Ranseen & Humphries, 1992).
· The intellectual functioning of children and adolescents with cystic fibrosis (Thompson et al., 1992).
· Decline in cognitive abilities over a one-year period in people with Alzheimer’s disease (Burns et al., 1991).
· The rate of motor-skill development in developmentally delayed preschoolers, as measured by the Vineland Adaptive Behavior Scale (Davies & Gavin, 1994).Page 97
· Scores on the Swedish translation of the Positive and Negative Syndrome Scale, which assesses the presence of positive and negative symptoms in people with schizophrenia (von Knorring & Lindstrom, 1992).
· Scores of psychiatrists on the Scale for Treatment Integration of the Dually Diagnosed (people with both a drug problem and another mental disorder); the scale examines opinions about drug treatment for this group of patients (Adelman et al., 1991).
· Responses to the Tridimensional Personality Questionnaire, a measure of three distinct personality features (Cloninger et al., 1991).
· Scores on a self-esteem measure among undergraduates (Addeo et al., 1994).
In each case, the researchers made a special point of stating that the scale under investigation yielded something close to a normal distribution of scores. Why? One benefit of a normal distribution of scores is that it simplifies the interpretation of individual scores on the test. In a normal distribution, the mean, the median, and the mode take on the same value. For example, if we know that the average score for intellectual ability of children with cystic fibrosis is a particular value and that the scores are normally distributed, then we know quite a bit more. We know that the average is the most common score and the score below and above which half of all the scores fall. Knowing the mean and the standard deviation of a scale and that it is approximately normally distributed tells us that (1) approximately two-thirds of all testtakers’ scores are within a standard deviation of the mean and (2) approximately 95% of the scores fall within 2 standard deviations of the mean.
The characteristics of the normal curve provide a ready model for score interpretation that can be applied to a wide range of test results.
Before reading on, take a minute or two to calculate what a score exactly at 3 standard deviations below the mean would be equal to. How about a score exactly at 3 standard deviations above the mean? Were your answers 5 and 95, respectively? The graph tells us that 99.74% of all scores in these normally distributed spelling-test data lie between ±3 standard deviations. Stated another way, 99.74% of all spelling test scores lie between 5 and 95. This graph also illustrates the following characteristics of all normal distributions.
· 50% of the scores occur above the mean and 50% of the scores occur below the mean.
· Approximately 34% of all scores occur between the mean and 1 standard deviation above the mean.
· Approximately 34% of all scores occur between the mean and 1 standard deviation below the mean.
· Approximately 68% of all scores occur between the mean and ±1 standard deviation.
· Approximately 95% of all scores occur between the mean and ±2 standard deviations.
A normal curve has two tails. The area on the normal curve between 2 and 3 standard deviations above the mean is referred to as a tail . The area between −2 and −3 standard deviations below the mean is also referred to as a tail. Let’s digress here momentarily for a “real-life” tale of the tails to consider along with our rather abstract discussion of statistical concepts.
As observed in a thought-provoking article entitled “Two Tails of the Normal Curve,” an intelligence test score that falls within the limits of either tail can have momentous consequences in terms of the tale of one’s life:
Individuals who are mentally retarded or gifted share the burden of deviance from the norm, in both a developmental and a statistical sense. In terms of mental ability as operationalized by tests of intelligence, performance that is approximately two standard deviations from the mean (or, IQ of 70–75 or lower or IQ of 125–130 or higher) is one key element in identification. Success at life’s tasks, or its absence, also plays a defining role, but the primary classifying feature of both gifted and retarded groups is intellectual deviance. These individuals are out of sync with more average people, simply by their difference from what is expected for their age Page 98and circumstance. This asynchrony results in highly significant consequences for them and for those who share their lives. None of the familiar norms apply, and substantial adjustments are needed in parental expectations, educational settings, and social and leisure activities. (Robinson et al., 2000, p. 1413)
Robinson et al. (2000) convincingly demonstrated that knowledge of the areas under the normal curve can be quite useful to the interpreter of test data. This knowledge can tell us not only something about where the score falls among a distribution of scores but also something about a person and perhaps even something about the people who share that person’s life. This knowledge might also convey something about how impressive, average, or lackluster the individual is with respect to a particular discipline or ability. For example, consider a high-school student whose score on a national, well-respected spelling test is close to 3 standard deviations above the mean. It’s a good bet that this student would know how to spell words like asymptotic and leptokurtic.
Just as knowledge of the areas under the normal curve can instantly convey useful information about a test score in relation to other test scores, so can knowledge of standard scores.
Standard Scores
Simply stated, a standard score is a raw score that has been converted from one scale to another scale, where the latter scale has some arbitrarily set mean and standard deviation. Why convert raw scores to standard scores?
Raw scores may be converted to standard scores because standard scores are more easily interpretable than raw scores. With a standard score, the position of a testtaker’s performance relative to other testtakers is readily apparent.
Different systems for standard scores exist, each unique in terms of its respective mean and standard deviations. We will briefly describe z scores, Tscores, stanines, and some other standard scores. First for consideration is the type of standard score scale that may be thought of as the zero plus or minus one scale. This is so because it has a mean set at 0 and a standard deviation set at 1. Raw scores converted into standard scores on this scale are more popularly referred to as z scores.
z Scores
A z score results from the conversion of a raw score into a number indicating how many standard deviation units the raw score is below or above the mean of the distribution. Let’s use an example from the normally distributed “National Spelling Test” data in Figure 3–7 to demonstrate how a raw score is converted to a z score. We’ll convert a raw score of 65 to a z score by using the formula
In essence, a z score is equal to the difference between a particular raw score and the mean divided by the standard deviation. In the preceding example, a raw score of 65 was found to be equal to a z score of +1. Knowing that someone obtained a z score of 1 on a spelling test provides context and meaning for the score. Drawing on our knowledge of areas under the normal curve, for example, we would know that only about 16% of the other testtakers obtained higher scores. By contrast, knowing simply that someone obtained a raw score of 65 on a spelling test conveys virtually no usable information because information about the context of this score is lacking.Page 99
In addition to providing a convenient context for comparing scores on the same test, standard scores provide a convenient context for comparing scores on different tests. As an example, consider that Crystal’s raw score on the hypothetical Main Street Reading Test was 24 and that her raw score on the (equally hypothetical) Main Street Arithmetic Test was 42. Without knowing anything other than these raw scores, one might conclude that Crystal did better on the arithmetic test than on the reading test. Yet more informative than the two raw scores would be the two z scores.
Converting Crystal’s raw scores to z scores based on the performance of other students in her class, suppose we find that her z score on the reading test was 1.32 and that her z score on the arithmetic test was −0.75. Thus, although her raw score in arithmetic was higher than in reading, the z scores paint a different picture. The z scores tell us that, relative to the other students in her class (and assuming that the distribution of scores is relatively normal), Crystal performed above average on the reading test and below average on the arithmetic test. An interpretation of exactly how much better she performed could be obtained by reference to tables detailing distances under the normal curve as well as the resulting percentage of cases that could be expected to fall above or below a particular standard deviation point (or z score).
T Scores
If the scale used in the computation of z scores is called a zero plus or minus one scale, then the scale used in the computation of T scores can be called a fifty plus or minus ten scale; that is, a scale with a mean set at 50 and a standard deviation set at 10. Devised by W. A. McCall (1922, 1939) and named a Tscore in honor of his professor E. L. Thorndike, this standard score system is composed of a scale that ranges from 5 standard deviations below the mean to 5 standard deviations above the mean. Thus, for example, a raw score that fell exactly at 5 standard deviations below the mean would be equal to a T score of 0, a raw score that fell at the mean would be equal to a T of 50, and a raw score 5 standard deviations above the mean would be equal to a T of 100. One advantage in using T scores is that none of the scores is negative. By contrast, in a z score distribution, scores can be positive and negative; this can make further computation cumbersome in some instances.
Other Standard Scores
Numerous other standard scoring systems exist. Researchers during World War II developed a standard score with a mean of 5 and a standard deviation of approximately 2. Divided into nine units, the scale was christened a stanine , a term that was a contraction of the words standard and nine.
Stanine scoring may be familiar to many students from achievement tests administered in elementary and secondary school, where test scores are often represented as stanines. Stanines are different from other standard scores in that they take on whole values from 1 to 9, which represent a range of performance that is half of a standard deviation in width ( Figure 3–8 ). The 5th stanine indicates performance in the average range, from 1/4 standard deviation below the mean to 1/4 standard deviation above the mean, and captures the middle 20% of the scores in a normal distribution. The 4th and 6th stanines are also 1/2 standard deviation wide and capture the 17% of cases below and above (respectively) the 5th stanine.
Figure 3–8 Stanines and the Normal Curve
Another type of standard score is employed on tests such as the Scholastic Aptitude Test (SAT) and the Graduate Record Examination (GRE). Raw scores on those tests are converted to standard scores such that the resulting distribution has a mean of 500 and a standard deviation of 100. If the letter A is used to represent a standard score from a college or graduate school admissions test whose distribution has a mean of 500 and a standard deviation of 100, then the following is true:
Page 100
Have you ever heard the term IQ used as a synonym for one’s score on an intelligence test? Of course you have. What you may not know is that what is referred to variously as IQ, deviation IQ, or deviation intelligence quotient is yet another kind of standard score. For most IQ tests, the distribution of raw scores is converted to IQ scores, whose distribution typically has a mean set at 100 and a standard deviation set at 15. Let’s emphasize typically because there is some variation in standard scoring systems, depending on the test used. The typical mean and standard deviation for IQ tests results in approximately 95% of deviation IQs ranging from 70 to 130, which is 2 standard deviations below and above the mean. In the context of a normal distribution, the relationship of deviation IQ scores to the other standard scores we have discussed so far (z, T, and A scores) is illustrated in Figure 3–9 .
Figure 3–9 Some Standard Score Equivalents Note that the values presented here for the IQ scores assume that the intelligence test scores have a mean of 100 and a standard deviation of 15. This is true for many, but not all, intelligence tests. If a particular test of intelligence yielded scores with a mean other than 100 and/or a standard deviation other than 15, then the values shown for IQ scores would have to be adjusted accordingly.
Standard scores converted from raw scores may involve either linear or nonlinear transformations. A standard score obtained by a linear transformation is one that retains a direct numerical relationship to the original raw score. The magnitude of differences between such standard scores exactly parallels the differences between corresponding raw scores. Sometimes scores may undergo more than one transformation. For example, the creators of the SAT did a second linear transformation on their data to convert z scores into a new scale that has a mean of 500 and a standard deviation of 100.
A nonlinear transformation may be required when the data under consideration are not normally distributed yet comparisons with normal distributions need to be made. In a nonlinear transformation, the resulting standard score does not necessarily have a direct numerical relationship to the original, raw score. As the result of a nonlinear transformation, the original distribution is said to have been normalized.
Normalized standard scores
Many test developers hope that the test they are working on will yield a normal distribution of scores. Yet even after very large samples have been tested with the instrument under development, skewed distributions result. What should be done?
One alternative available to the test developer is to normalize the distribution. Conceptually, normalizing a distribution involves “stretching” the skewed curve into the shape of a normal curve and creating a corresponding scale of standard scores, a scale that is technically referred to as a normalized standard score scale .
Normalization of a skewed distribution of scores may also be desirable for purposes of comparability. One of the primary advantages of a standard score on one test is that it can readily be compared with a standard score on another test. However, such comparisons are appropriate only when the distributions from which they derived are the same. In most instances, Page 101they are the same because the two distributions are approximately normal. But if, for example, distribution A were normal and distribution B were highly skewed, then z scores in these respective distributions would represent different amounts of area subsumed under the curve. A z score of −1 with respect to normally distributed data tells us, among other things, that about 84% of the scores in this distribution were higher than this score. A z score of −1 with respect to data that were very positively skewed might mean, for example, that only 62% of the scores were higher.
JUST THINK . . .
Apply what you have learned about frequency distributions, graphing frequency distributions, measures of central tendency, measures of variability, and the normal curve and standard scores to the question of the data listed in Table 3–1 . How would you communicate the data from Table 3–1 to the class? Which type of frequency distribution might you use? Which type of graph? Which measure of central tendency? Which measure of variability? Might reference to a normal curve or to standard scores be helpful? Why or why not?
For test developers intent on creating tests that yield normally distributed measurements, it is generally preferable to fine-tune the test according to difficulty or other relevant variables so that the resulting distribution will approximate the normal curve. That usually is a better bet than attempting to normalize skewed distributions. This is so because there are technical cautions to be observed before attempting normalization. For example, transformations should be made only when there is good reason to believe that the test sample was large enough and representative enough and that the failure to obtain normally distributed scores was due to the measuring instrument.Page 102
Correlation and Inference
Central to psychological testing and assessment are inferences (deduced conclusions) about how some things (such as traits, abilities, or interests) are related to other things (such as behavior). A coefficient of correlation (or correlation coefficient ) is a number that provides us with an index of the strength of the relationship between two things. An understanding of the concept of correlation and an ability to compute a coefficient of correlation is therefore central to the study of tests and measurement.
The Concept of Correlation
Simply stated, correlation is an expression of the degree and direction of correspondence between two things. A coefficient of correlation (r) expresses a linear relationship between two (and only two) variables, usually continuous in nature. It reflects the degree of concomitant variation between variable X and variable Y. The coefficient of correlation is the numerical index that expresses this relationship: It tells us the extent to which X and Y are “co-related.”
The meaning of a correlation coefficient is interpreted by its sign and magnitude. If a correlation coefficient were a person asked “What’s your sign?,” it wouldn’t answer anything like “Leo” or “Pisces.” It would answer “plus” (for a positive correlation), “minus” (for a negative correlation), or “none” (in the rare instance that the correlation coefficient was exactly equal to zero). If asked to supply information about its magnitude, it would respond with a number anywhere at all between −1 and +1. And here is a rather intriguing fact about the magnitude of a correlation coefficient: It is judged by its absolute value. This means that to the extent that we are impressed by correlation coefficients, a correlation of −.99 is every bit as impressive as a correlation of +.99. To understand why, you need to know a bit more about correlation.
“Ahh . . . a perfect correlation! Let me count the ways.” Well, actually there are only two ways. The two ways to describe a perfect correlation between two variables are as either +1 or −1. If a correlation coefficient has a value of +1 or −1, then the relationship between the two variables being correlated is perfect—without error in the statistical sense. And just as perfection in almost anything is difficult to find, so too are perfect correlations. It’s challenging to try to think of any two variables in psychological work that are perfectly correlated. Perhaps that is why, if you look in the margin, you are asked to “just think” about it.
JUST THINK . . .
Can you name two variables that are perfectly correlated? How about two psychological variables that are perfectly correlated?
If two variables simultaneously increase or simultaneously decrease, then those two variables are said to be positively (or directly) correlated. The height and weight of normal, healthy children ranging in age from birth to 10 years tend to be positively or directly correlated. As children get older, their height and their weight generally increase simultaneously. A positive correlation also exists when two variables simultaneously decrease. For example, the less a student prepares for an examination, the lower that student’s score on the examination. A negative (or inverse) correlation occurs when one variable increases while the other variable decreases. For example, there tends to be an inverse relationship between the number of miles on your car’s odometer (mileage indicator) and the number of dollars a car dealer is willing to give you on a trade-in allowance; all other things being equal, as the mileage increases, the number of dollars offered on trade-in decreases. And by the way, we all know students who use cell phones during class to text, tweet, check e-mail, or otherwise be engaged with their phone at a questionably appropriate time and place. What would you estimate the correlation to be between such daily, in-class cell phone use and test grades? See Figure 3–10 for one such estimate (and kindly refrain from sharing the findings on Facebook during class).Page 103
Figure 3–10 Cell Phone Use in Class and Class Grade This may be the “wired” generation, but some college students are clearly more wired than others. They seem to be on their cell phones constantly, even during class. Their gaze may be fixed on Mech Commander when it should more appropriately be on Class Instructor. Over the course of two semesters, Chris Bjornsen and Kellie Archer (2015) studied 218 college students, each of whom completed a questionnaire on their cell phone usage right after class. Correlating the questionnaire data with grades, the researchers reported that cell phone usage during class was significantly, negatively correlated with grades.© Caia Image/Glow Images RF
If a correlation is zero, then absolutely no relationship exists between the two variables. And some might consider “perfectly no correlation” to be a third variety of perfect correlation; that is, a perfect noncorrelation. After all, just as it is nearly impossible in psychological work to identify two variables that have a perfect correlation, so it is nearly impossible to identify two variables that have a zero correlation. Most of the time, two variables will be fractionally correlated. The fractional correlation may be extremely small but seldom “perfectly” zero.
JUST THINK . . .
Bjornsen & Archer (2015) discussed the implications of their cell phone study in terms of the effect of cell phone usage on student learning, student achievement, and post-college success. What would you anticipate those implications to be?
JUST THINK . . .
Could a correlation of zero between two variables also be considered a “perfect” correlation? Can you name two variables that have a correlation that is exactly zero?
As we stated in our introduction to this topic, correlation is often confused with causation. It must be emphasized that a correlation coefficient is merely an index of the relationship between two variables, not an index of the causal relationship between two variables. If you were told, for example, that from birth to age 9 there is a high positive correlation between hat size and spelling ability, would it be appropriate to conclude that hat size causes spelling ability? Of course not. The period Page 104from birth to age 9 is a time of maturation in all areas, including physical size and cognitive abilities such as spelling. Intellectual development parallels physical development during these years, and a relationship clearly exists between physical and mental growth. Still, this doesn’t mean that the relationship between hat size and spelling ability is causal.
Although correlation does not imply causation, there is an implication of prediction. Stated another way, if we know that there is a high correlation between X and Y, then we should be able to predict—with various degrees of accuracy, depending on other factors—the value of one of these variables if we know the value of the other.
The Pearson r
Many techniques have been devised to measure correlation. The most widely used of all is the Pearson r , also known as the Pearson correlation coefficientand the Pearson product-moment coefficient of correlation. Devised by Karl Pearson ( Figure 3–11 ), r can be the statistical tool of choice when the relationship between the variables is linear and when the two variables being correlated are continuous (or, they can theoretically take any value). Other correlational techniques can be employed with data that are discontinuous and where the relationship is nonlinear. The formula for the Pearson r takes into account the relative position of each test score or measurement with respect to the mean of the distribution.
Figure 3–11 Karl Pearson (1857–1936) Karl Pearson’s name has become synonymous with correlation. History records, however, that it was actually Sir Francis Galton who should be credited with developing the concept of correlation (Magnello & Spies, 1984). Galton experimented with many formulas to measure correlation, including one he labeled r. Pearson, a contemporary of Galton’s, modified Galton’sr, and the rest, as they say, is history. The Pearson r eventually became the most widely used measure of correlation.© TopFoto/Fotomas/The Image Works
A number of formulas can be used to calculate a Pearson r. One formula requires that we convert each raw score to a standard score and then multiply each pair of standard scores. A mean for the sum of the products is calculated, and that mean is the value of the Pearson r. Even from this simple verbal conceptualization of the Pearson r, it can be seen that the sign of the resulting r would be a function of the sign and the magnitude of the standard scores used. If, for example, negative Page 105standard score values for measurements of X always corresponded with negative standard score values for Y scores, the resulting r would be positive (because the product of two negative values is positive). Similarly, if positive standard score values on X always corresponded with positive standard score values on Y, the resulting correlation would also be positive. However, if positive standard score values for Xcorresponded with negative standard score values for Y and vice versa, then an inverse relationship would exist and so a negative correlation would result. A zero or near-zero correlation could result when some products are positive and some are negative.
The formula used to calculate a Pearson r from raw scores is
This formula has been simplified for shortcut purposes. One such shortcut is a deviation formula employing “little x,” or x in place of and “little y,” or y in place of
Another formula for calculating a Pearson r is
Although this formula looks more complicated than the previous deviation formula, it is easier to use. Here N represents the number of paired scores; Σ XY is the sum of the product of the paired X and Y scores; Σ X is the sum of the X scores; Σ Y is the sum of the Y scores; Σ X2 is the sum of the squared Xscores; and Σ Y2 is the sum of the squared Y scores. Similar results are obtained with the use of each formula.
The next logical question concerns what to do with the number obtained for the value of r. The answer is that you ask even more questions, such as “Is this number statistically significant, given the size and nature of the sample?” or “Could this result have occurred by chance?” At this point, you will need to consult tables of significance for Pearson r—tables that are probably in the back of your old statistics textbook. In those tables you will find, for example, that a Pearson r of .899 with an N = 10 is significant at the .01 level (using a two-tailed test). You will recall from your statistics course that significance at the .01 level tells you, with reference to these data, that a correlation such as this could have been expected to occur merely by chance only one time or less in a hundred if X and Y are not correlated in the population. You will also recall that significance at either the .01 level or the (somewhat less rigorous) .05 level provides a basis for concluding that a correlation does indeed exist. Significance at the .05 level means that the result could have been expected to occur by chance alone five times or less in a hundred.
The value obtained for the coefficient of correlation can be further interpreted by deriving from it what is called a coefficient of determination , or r2. The coefficient of determination is an indication of how much variance is shared by the X- and the Y-variables. The calculation of r2 is quite straightforward. Simply square the correlation coefficient and multiply by 100; the result is equal to the percentage of the variance accounted for. If, for example, you calculated r to be .9, then r2 would be equal to .81. The number .81 tells us that 81% of the variance is accounted for by the X- and Y-variables. The remaining variance, equal to 100(1 − r2), or 19%, could presumably be accounted for by chance, error, or otherwise unmeasured or unexplainable factors. 7 Page 106
Before moving on to consider another index of correlation, let’s address a logical question sometimes raised by students when they hear the Pearson r referred to as the product-moment coefficient of correlation. Why is it called that? The answer is a little complicated, but here goes.
In the language of psychometrics, a moment describes a deviation about a mean of a distribution. Individual deviations about the mean of a distribution are referred to as deviates. Deviates are referred to as the first moments of the distribution. The second moments of the distribution are the moments squared. The third moments of the distribution are the moments cubed, and so forth. The computation of the Pearson r in one of its many formulas entails multiplying corresponding standard scores on two measures. One way of conceptualizing standard scores is as the first moments of a distribution. This is because standard scores are deviates about a mean of zero. A formula that entails the multiplication of two corresponding standard scores can therefore be conceptualized as one that entails the computation of the product of corresponding moments. And there you have the reason r is called product-moment correlation. It’s probably all more a matter of psychometric trivia than anything else, but we think it’s cool to know. Further, you can now understand the rather “high-end” humor contained in the cartoon (below).
The Spearman Rho
The Pearson r enjoys such widespread use and acceptance as an index of correlation that if for some reason it is not used to compute a correlation coefficient, mention is made of the statistic that was used. There are many alternative ways to derive a coefficient of correlation. One commonly used alternative statistic is variously called a rank-order correlation coefficient , a rank-difference correlation coefficient , or simply Spearman’s rho .Developed by Charles Spearman, a British psychologist ( Figure 3–12 ), this coefficient of correlation is frequently used when the sample size is small (fewer than 30 pairs of measurements) and especially when both sets of measurements are in ordinal (or rank-order) form. Special tables are used to determine whether an obtained rho coefficient is or is not significant.
Figure 3–12 Charles Spearman (1863–1945) Charles Spearman is best known as the developer of the Spearman rho statistic and the Spearman-Brown prophecy formula, which is used to “prophesize” the accuracy of tests of different sizes. Spearman is also credited with being the father of a statistical method called factor analysis, discussed later in this text.© Atlas Archive/The Image Works Copyright 2016 by Ronald Jay Cohen. All rights reserved.Page 107
Graphic Representations of Correlation
One type of graphic representation of correlation is referred to by many names, including a bivariate distribution , a scatter diagram , a scattergram , or—our favorite—a scatterplot . A scatterplot is a simple graphing of the coordinate points for values of the X-variable (placed along the graph’s horizontal axis) and the Y-variable (placed along the graph’s vertical axis). Scatterplots are useful because they provide a quick indication of the direction and magnitude of the relationship, if any, between the two variables. Figures 3–13 and 3–14 offer a quick course in eyeballing the nature and degree of correlation by means of scatterplots. To distinguish positive from negative correlations, note the direction of the curve. And to estimate the strength of magnitude of the correlation, note the degree to which the points form a straight line.
Figure 3–13 Scatterplots and Correlations for Positive Values of r Figure 3–14 Scatterplots and Correlations for Negative Values of r
Scatterplots are useful in revealing the presence of curvilinearity in a relationship. As you may have guessed, curvilinearity in this context refers to an “eyeball gauge” of how curved a graph is. Remember that a Pearson r should be used only if the relationship between the variables is linear. If the graph does not appear to take the form of a straight line, the chances are good that the relationship is not linear ( Figure 3–15 ). When the relationship is nonlinear, other statistical tools and techniques may be employed. 8
Page 110Figure 3–15 Scatterplot Showing a Nonlinear Correlation
A graph also makes the spotting of outliers relatively easy. An outlier is an extremely atypical point located at a relatively long distance—an outlying distance—from the rest of the coordinate points in a scatterplot ( Figure 3–16 ). Outliers stimulate interpreters of test data to speculate about the reason for the atypical score. For example, consider an outlier on a scatterplot that reflects a correlation between hours each member of a fifth-grade class spent studying and their grades on a 20-item spelling test. And let’s say that one student studied for 10 hours and received a failing grade. This outlier on the scatterplot might raise a red flag and compel the test user to raise some important questions, such as “How effective are this student’s study skills and habits?” or “What was this student’s state of mind during the test?”
Figure 3–16 Scatterplot Showing an Outlier
In some cases, outliers are simply the result of administering a test to a very small sample of testtakers. In the example just cited, if the test were given statewide to fifth-graders and the sample size were much larger, perhaps many more low scorers who put in large amounts of study time would be identified.
As is the case with very low raw scores or raw scores of zero, outliers can sometimes help identify a testtaker who did not understand the instructions, was not able to follow the instructions, or was simply oppositional and did not follow the instructions. In other cases, an outlier can provide a hint of some deficiency in the testing or scoring procedures.
People who have occasion to use or make interpretations from graphed data need to know if the range of scores has been restricted in any way. To understand why this is so necessary to know, consider Figure 3–17 . Let’s say that graph A describes the relationship between Public University entrance test scores for 600 applicants (all of whom were later admitted) and their grade point averages at the end of the first semester. The scatterplot indicates that the relationship between entrance test scores and grade point average is both linear and positive. But what if the admissions officer had accepted only the applications of the students who scored within the top half or so on the entrance exam? To a trained eye, this scatterplot (graph B) appears to indicate a weaker correlation than that indicated in graph A—an effect attributable exclusively to the restriction of range. Graph B is less a straight line than graph A, and its direction is not as obvious.
Figure 3–17 Two Scatterplots Illustrating Unrestricted and Restricted Ranges
Meta-Analysis
Generally, the best estimate of the correlation between two variables is most likely to come not from a single study alone but from analysis of the data from several studies. One option to Page 111facilitate understanding of the research across a number of studies is to present the range of statistical values calculated from a number of different studies of the same phenomenon. Viewing all of the data from a number of studies that attempted to determine the correlation between variable X and variable Y, for example, might lead the researcher to conclude that “The correlation between variable Xand variable Y ranges from .73 to .91.” Another option might be to combine statistically the information across the various studies; that is what is done using a statistical technique called meta-analysis. Using this technique, researchers raise (and strive to answer) the question: “Combined, what do all of these studies tell us about the matter under study?” For example, Imtiaz et al. (2016) used meta-analysis to draw some conclusions regarding the relationship between cannabis use and physical health. Colin (2015) used meta-analysis to study the correlations of use-of-force decisions among American police officers.
Meta-analysis may be defined as a family of techniques used to statistically combine information across studies to produce single estimates of the data under study. The estimates derived, referred to as effect size , may take several different forms. In most meta-analytic studies, effect size is typically expressed as a correlation coefficient. 9 Meta-analysis facilitates the drawing of conclusions and the making of statements like, “the typical therapy client is better off than 75% of untreated individuals” (Smith & Glass, 1977, p. 752), there is “about 10% increased risk for antisocial behavior among children with incarcerated parents, compared to peers” (Murray et al., 2012), and “GRE and UGPA [undergraduate grade point average] are generalizably valid predictors of graduate grade point average, 1st-year graduate grade point average, comprehensive examination scores, publication citation counts, and faculty ratings” (Kuncel et al., 2001, p. 162).Page 112
MEET AN ASSESSMENT PROFESSIONAL
Meet Dr. Joni L. Mihura
Hi, my name is Joni Mihura, and my research expertise is in psychological assessment, with a special focus on the Rorschach. To tell you a little about me, I was the only woman* to serve on the Research Council for John E. Exner’s Rorschach Comprehensive System (CS) until he passed away in 2006. Due to the controversy around the Rorschach’s validity, I began reviewing the research literature to ensure I was teaching my doctoral students valid measures to assess their clients. That is, the controversy about the Rorschach has not been that it is a completely invalid test—the critics have endorsed several Rorschach scales as valid for their intended purpose—the main problem that they have highlighted is that only a small proportion of its scales had been subjected to “meta-analysis,” a systematic technique for summarizing the research literature. To make a long story short, I eventually published my review of the Rorschach literature in the top scientific review journal in psychology (Psychological Bulletin) in the form of systematic reviews and meta-analyses of the 65 main Rorschach CS variables (Mihura et al., 2013), therefore making the Rorschach the psychological test with the most construct validity meta-analyses for its scales!
My meta-analyses also resulted in two other pivotal events. They formed the backbone for a new scientifically based Rorschach system of which I am a codeveloper—the Rorschach Performance Assessment System (R-PAS; Meyer et al., 2011), and they resulted in the Rorschach critics removing the “moratorium” they had recommended for the Rorschach (or, Garb, 1999) for the scales they deemed had solid support in our meta-analyses (Wood et al., 2015; also see our reply, Mihura et al., 2015).
I’m very excited to talk with you about meta-analysis. First, to set the stage, let’s take a step back and look at what you might have experienced so far when reading about psychology. When students take their first psychology course, they are often surprised how much of the field is based on research findings rather than just “common sense.” Even so, because undergraduate textbooks have numerous topics about which they cannot cite all of the research, it can appear that the textbook is relying on just one or two studies as the “proof.” Therefore, you might be surprised just how many psychological research studies actually exist! Conducting a quick search in the PsycINFO database shows that over a million psychology journal articles are classified as empirical studies—and that excludes chapters, theses, dissertations, and many other studies not listed in PsycINFO.
Joni L. Mihura, Ph.D. is Associate Professor of Psychology at the University of Toledo in Toledo, Ohio © Joni L. Mihura, Ph.D.
But, good news or bad news, a significant challenge with many research studies is how to summarize results. The classic example of such a dilemma and the eventual solution is a fascinating one that comes from the psychotherapy literature. In 1952, Hans Eysenck published a classic article entitled Page 113“The Effects of Psychotherapy: An Evaluation,” in which he summarized the results of a few studies and concluded that psychotherapy doesn’t work! Wow! This finding had the potential to shake the foundation of psychotherapy and even ban its existence. After all, Eysenck had cited research that suggested that the longer a person was in therapy, the worse-off they became. Notwithstanding the psychotherapists and the psychotherapy enterprise, Eysenck’s publication had sobering implications for people who had sought help through psychotherapy. Had they done so in vain? Was there really no hope for the future? Were psychotherapists truly ill-equipped to do things like reduce emotional suffering and improve peoples’ lives through psychotherapy?
In the wake of this potentially damning article, several psychologists—and in particular Hans H. Strupp—responded by pointing out problems with Eysenck’s methodology. Other psychologists conducted their own reviews of the psychotherapy literature. Somewhat surprisingly, after reviewing the same body of research literature on psychotherapy, various psychologists drew widely different conclusions. Some researchers found strong support for the efficacy of psychotherapy. Other researchers found only modest support for the efficacy of psychotherapy. Yet other researchers found no support for it at all.
How can such different conclusions be drawn when the researchers are reviewing the same body of literature? A comprehensive answer to this important question could fill the pages of this book. Certainly, one key element of the answer to this question had to do with a lack of systematic rules for making decisions about including studies, as well as lack of a widely acceptable protocol for statistically summarizing the findings of the various studies. With such rules and protocols absent, it would be all too easy for researchers to let their preexisting biases run amok. The result was that many researchers “found” in their analyses of the literature what they believed to be true in the first place.
A fortuitous bi-product of such turmoil in the research community was the emergence of a research technique called “meta-analysis.” Literally, “an analysis of analyses,” meta-analysis is a tool used to systematically review and statistically summarize the research findings for a particular topic. In 1977, Mary Lee Smith and Gene V. Glass published the first meta-analysis of psychotherapy outcomes. They found strong support for the efficacy of psychotherapy. Subsequently, others tried to challenge Smith and Glass’ findings. However, the systematic rigor of their meta-analytic technique produced findings that were consistently replicated by others. Today there are thousands of psychotherapy studies, and many meta-analysts ready to research specific, therapy-related questions (like “What type of psychotherapy is best for what type of problem?”).
What does all of this mean for psychological testing and assessment? Meta-analytic methodology can be used to glean insights about specific tools of assessment, and testing and assessment procedures. However, meta-analyses of information related to psychological tests brings new challenges owing, for example, to the sheer number of articles to be analyzed, the many variables on which tests differ, and the specific methodology of the meta-analysis. Consider, for example, that multiscale personality tests may contain over 50, and sometimes over 100, scales that each need to be evaluated separately. Furthermore, some popular multiscale personality tests, like the MMPI-2 and Rorschach, have had over a thousand research studies published on them. The studies typically report findings that focus on varied aspects of the test (such as the utility of specific test scales, or other indices of test reliability or validity). In order to make the meta-analytic task manageable, meta-analyses for multiscale tests will typically focus on one or another of these characteristics or indices.
In sum, a thoughtful meta-analysis of research on a specific topic can yield important insights of both theoretical and applied value. A meta-analytic review of the literature on a particular psychological test can even be instrumental in the formulation of revised ways to score the test and interpret the findings (just ask Meyer et al., 2011). So, the next time a question about psychological research arises, students are advised to respond to that question with their own question, namely “Is there a meta-analysis on that?”
Used with permission of Joni L. Mihura.
*I have also edited the Handbook of Gender and Sexuality in Psychological Assessment (Brabender & Mihura, 2016).
Page 114
A key advantage of meta-analysis over simply reporting a range of findings is that, in meta-analysis, more weight can be given to studies that have larger numbers of subjects. This weighting process results in more accurate estimates (Hunter & Schmidt, 1990). Some advantages to meta-analyses are: (1) meta-analyses can be replicated; (2) the conclusions of meta-analyses tend to be more reliable and precise than the conclusions from single studies; (3) there is more focus on effect size rather than statistical significance alone; and (4) meta-analysis promotes evidence-based practice , which may be defined as professional practice that is based on clinical and research findings (Sánchez-Meca & Marin-Martinez, 2010). Despite these and other advantages, meta-analysis is, at least to some degree, art as well as science (Hall & Rosenthal, 1995). The value of any meta-analytic investigation is very much a matter of the skill and ability of the meta-analyst (Kavale, 1995), and use of an inappropriate meta-analytic method can lead to misleading conclusions (Kisamore & Brannick, 2008).
It may be helpful at this time to review this statistics refresher to make certain that you indeed feel “refreshed” and ready to continue. We will build on your knowledge of basic statistical principles in the chapters to come, and it is important to build on a rock-solid foundation.
Self-Assessment
Test your understanding of elements of this chapter by seeing if you can explain each of the following terms, expressions, and abbreviations:
· coefficient of determination
· error
· graph
· grouped frequency distribution
· kurtosis
· mean
· median
· mode
· normalized standard score scale
· outlier
· quartile
· range
· rank-order/rank-difference correlation coefficient
· scale
· skewness
· stanine
· T score
· tail
· variance
· z score