Variables, Values, and Levels of Measurement - Discussion Question
147
Level of Measurement
Once Over Again
EDGAR F. BORGATTA Graduate Center, City University of New York
GEORGE W. BOHRNSTEDT Indiana University—Bloomington
The distinctions between nominal, ordinal, interval, and ratio measurement were
popularized by S. S. Stevens. Unfortunately, positions taken by Stevens have often been disseminated without criticism. One problem is the common assumption that "ordinal" statistics are the best statistics to use for presumed noninterval continuous social variables, when, in fact, they use addition, subtraction, and division, which make the measurements interval by definition. Additionally, the relationship of the normal distribution to interval measurement is commonly misunderstood, the latter existing by definition if a normal distribution exists. Concern with levels of measurements may mislead persons into attending to issues other than maximizing (utility, given) the particular limits of the state of the measurement art in the social sciences.
f f those who already believe do not need to be convinced, and those who are skeptics will not believe, then a lot of
preaching goes on without purpose. Therefore, we hope what follows is less a matter of preaching and more a matter of pur- suing the consistency and appropriateness of one position on the &dquo;level of measurement&dquo; issue in the social sciences. The position taken here is that the question of measurement assumptions in statistical applications often has been treated inappropriately, and, more to the point, erroneously. As we look back, we see resulting misinformation that can be described as a hindrance to the development of the social and psychological sciences. Our objective here is to emphasize the misinformation and the errors, with some historical reference. We hope to ease the minds of those researchers who have been &dquo;mindlessly&dquo; applying parametric
SOCIOLOGICAL METHODS & RESEARCH, Vol 9 No 2, November 1980 147-160 @ 1980 Sage Publications, Inc
148
statistics to social variables. As we will show, this practice has been a wise one.
Most discussions of the level of measurement issue can be traced back to the work of Stevens (1966). While many people have written about the ideas of measurement, the variety of approaches that are involved is still impressive. Essentially, one’s objective determines, to some extent, how one enters the argument. To begin with, no one enters this arena without some experience with numbers and ideas of measurement; therefore, there are few neutral persons involved in the debate.
As one works more and more with Stevens’ concepts, aspects of
circularity involved in definitions and refinements of concepts become evident. In at least part of our presentation, we shall follow Stevens not only as a matter of convenience, but also to maintain continuity with presentations that have been made earlier by other writers. Following common convention, Stevens defines measurement thusly: &dquo;A rule for the assignment of numerals (numbers) to aspects of objects or events creates a scale [of measurement]&dquo; (1966: 22). More simply measurement is the assignment of numbers to objects according to rules. Stevens ( 1966: 23) proceeds directly to a discussion of types of scales, a set of distinctions that has received much attention:
The type of scale achieved when we deputize the numerals to serve as representatives for a state of affairs in nature depends upon the character of the basic empirical operations performed on nature. These operations are limited ordinarily by the peculiarities of the thing being scaled and by our choice of concrete procedures, but once selected the procedures determine that there will eventuate one or another of four types of scale: nominal, ordinal, interval or ratio. Each of these classes of scales is best characterized by its range of invariance-by the kinds of transformations that leave the &dquo;structure&dquo; of the scale undistorted. And the nature of the invariance sets limits to the kinds of statistical manipulation that can legitimately be applied to the scaled data. This question of the applicability of the various statistics is of great practical concern to several of the sciences.
149
Stevens presents the four types of scales as ordered in a cumulative sense. The nominal scale requires the determination of equality for placement in the classes implied; the ordinal scale additionally requires a determination of &dquo;greater than&dquo; or &dquo;less than&dquo; for objects; the interval scale in addition requires deter- mination of equality of differences between scale intervals; and the ratio scale further requires determination of a true zero point. The ratio scale, of course, has all the qualities of the three previously named scales. It is appropriate to go over the meanings associated with these scales in terms of the elaboration by Stevens.
NOMINAL SCALES
Stevens sees two types of nominal scales: the type used for
purposes such as numbering individuals for identification as one, and a scale used for classification into types where those who are
placed within the type are given the same number as the other. The former is a special case of the latter. Nominal scales are seen as a primitive type of scale, one in which the basic rule is that the same numeral is not assigned to different classes. A class is defined as or based on the demonstration of equality in respect to some characteristic of the object. We shall bypass here the question of the logical definition of equality, although this is an important issue in itself.
According to Stevens (1966) &dquo;the ordinal scale arises from the operation of rank ordering.&dquo; It is at this point that Stevens appears to have gone astray, possibly because of the direction from which he approached the subject matter. And, indeed, the most serious misdirection that has developed in the area of measurement resulted from this statement. Stevens (1966: 26) asserts:
As a matter of fact, most of the scales used widely and effectively by psychologists are ordinal scales. In the strictest propriety, the ordinary statistics involving means and standard deviations ought
150
not to be used with those scales, for these statistics imply a knowledge of something more than the relative rank order of data.
The quarrel here is simple and direct. Most of the scales used widely and effectively by psychologists (and other social scien- tists) simply are not ordinal scales; the procedures used by social scientists usually fit badly to an interval scale, but they certainly are not ordinal scales. In other words, we measure latent continuous variables with error at the manifest level.
Stevens appears aware of the limitations associated with the idea of ordinal scales, but unfortunately was distracted from pursuing this matter:
In earlier discussions (e.g., Stevens, 1946) I expressed the opinion that rank-order correlation does not apply to ordinal scales because the derivation of the formula for this correlation involves the assumption that the differences between successive ranks are equal. My colleague, Frederick Mosteller convinces me that this conservative view can be liberalized, provided that the resultant coefficient (e.g., Spearman’s P or Kendall’s T ) is interpreted only as a test function for a hypothesis about order [ 1966: 26].
If only Stevens had not been convinced, the enormous wasted energy and misdirection of the fads and fashions of so-called
nonparametric or distribution-free statistics might possibly have been avoided. This extremely important point will receive specific and special attention below.
INTERVAL SCALES
An interval scale is identified by Stevens as what we ordinarily think of when we consider quantitative procedures. An interval scale is subject to a linear transformation with invariance, and is represented by many common measures, such as the Fahrenheit and Celsius temperature scales. With minor limitations, these are subject to all the algebraic operations with which we are ordinarily concerned in science. The transition to a ratio scale is easily handled by the fact that operations on the interval scale can conveniently fix a true zero.
151
Stevens (1966: 28) goes astray in discussing interval measure- ment relative to the work of psychologists:
The variability of a psychological measure is itself sometimes used to equalize the units of a scale. This process smacks of a kind of magic-a rope trick for climbing the hierarchy of scales. The rope in this case is the assumption that in the sample of individuals tested the trait in question has a canonical distribution (e.g., &dquo;normal&dquo;). Then it is a simple matter to adjust the units of the scale so that the assumed distribution is recovered when the individuals are measured. But this procedure is obviously no better than the gratuitous postulate behind it.
Stevens recognizes that &dquo;the fact remains that the assumption of normality has the advocacy of a certain pragmatic usefulness in the measurement of many human traits,&dquo; but he does not pursue this point, we shall do later in this essay.
STEVENS MORE RECENT POSITION
In another treatment of measurement, Stevens (1975) takes a position that is emphasized here and leads to quite a different global interpretation of the measurement process as appropriate for the social and psychological sciences. Stevens notes that &dquo;it is helpful ... to regard measurement as a two-part endeavor, consisting on the one hand of manipulations and on the other of models. You do something, you perform a sequence of opera- tions, and afterwards you invoke a model or a schema to stand for what you have done.&dquo; He observes that there are really two aspects to the idea of measurement: the everyday operations in- volved, and the model of measurement used to rationalize these
operations. What is most important is a kind of circularity. That is, scientists go through operations and invoke models, then go through additional operations and invoke additional models. We eventually enter an area in which there are many models and many sets of possible operations. Stevens (1975: 47) continues: &dquo;The model used in measurement is usually the system of numbers bequeathed to us by mathematics. The empirical operations may vary enormously, however, depending on what is to be measured.&dquo;
152
It should be noted here that this emphasis on models is somewhat in contrast to the negative comment relative to earlier assumptions made by Stevens (1966: 28) and noted here.
In dealing with measurement, then, one may begin wherever one pleases, but one must give attention, as has been suggested by Stevens and many others, to the correspondence between what one does and an underlying model. Of course, however, there is no reason that one should not begin with a model, then design operations which conform to it within given bounds of error. It is exactly this approach that we take here.
A SUGGESTED MODEL
The variables of greatest interest to social scientists are latent unobserved constructs rather than constructs that are opera- tionally defined. And most of these constructs are conceptualized to be continuous at the latent level, even though they are usually manifestly measured as discrete variables. Examples of this include the constructs of industrialization, social status, power, authoritarianism, and self-esteem. If the constructs are contin- uous, they must also be interval.
Importantly, the use of inferential regression analysis, factor analysis, and other multivariate techniques require that one’s dependent variables be continuous and distributed normally (multivariate normality for multiple dependent variables) for each outcome associated with the independent variable(s). It is worth emphasizing here that level of measurement is not a requirement for the use of parametric statistics, as was suggested by Stevens. For a recent discussion of this point and a brief history of this issue, see Gaito (1980).
Because most of the variables that interest us as researchers are
continuous at the conceptual level and are reasonably close to normally distributed in the population of interest, there is no reason to eschew the use of parametric statistics. And why not also assume that most of these latent variables are
roughly bell-shaped in the population? For most constructs, does
153
it not make sense to assume that the bulk of observations in the
population lie close to the mean, with relatively few cases falling at the extremes? It should be pointed out that Chebycheff’s well- known theorem guarantees that the probability of observing a given outcome increases the closer that outcome is to the population mean. While the theorem in no sense guarantees that the distribution will be normal, it is well known that regression and regression-like procedures tend to be relatively robust even if the assumption of normality is violated (Bohrnstedt and Carter, 1971). To summarize this argument, most of the central constructs in
the social sciences are conceptualized as continuous, and their distributions are such that the application of parametric statistics to their analyses will not result in seriously biased estimates. And if the variables are continuous, they must also by definition, be interval.
IMPERFECT INTERVAL-LEVEL MANIFEST SCALES
At the manifest, observed level, our measures are likely to be imperfect interval-level scales. Social measurement is crude. Unlike the physicist, who can measure with high precision through the use of pointer and meter readings which reflect agreed-upon international standards, social scientists are likely to generate measures which are (1) unquestionably discrete, rather than continuous, and (2) likely to result in measures with intervals which are not truly equal. However, if we have been careful in developing our measures (e.g., item and indicator development procedures that are extensive and systematic) we should be able to assume that there is a monotonic relationship between the manifest scale and the underlying latent construct. That is, a positive difference between two points on the manifest scale reflects a positive difference on the latent scale as well, even though the intervals on the manifest scale may not be isomorphic.
This is another way of saying that our measures contain error. Measurement error is defined, of course, as a function of the fit between the manifest scale and the latent construct. Importantly,
154
measurement error will have an effect on parameter estimates. But the fact that our observations do not correspond perfectly to the underlying model in no way implies that ordinal level statistics are required to analyze the data.
ORDINALITY?
We think the emphasis on using ordinal statistics in the social sciences is misplaced for two reasons. First, we do not visualize a model in which each and every individual in the population is ordered (ranked) relative to each and every other individual, nor do we think most social scientists visualize such a model. Indeed, the notion that ranking is a &dquo;natural&dquo; and appropriate way of thinking about most variables in social and psychological science is simply nonsense.
Second, let us examine how a manifest, well-ordered scale is modified in order to arrive at an ordinal scale. In practice, this occurs by the allocation of unit distances between individuals, independently of how they might be distributed on a variable. For example, if on an interval scale 3 persons have scores of 1, 3, and 14, the procedure of converting the scale to a set of ordered ranks is to revalue them as 1, 2, and 3. Ordinal scaling, then, may be thought of as a form of scaling in which the interval information is lost. Presumably, because the ordinal scale does not have appropriate allocation of numbers relative to the interval scale, it cannot be handled with all the convenient properties that would otherwise accrue. Nunnally (1967: 18) states this without ambi- guity : &dquo;With ordinal scales, none of the fundamental operations of algebra may be applied. In the use of descriptive statistics, it makes no sense to add, subtract, divide, or multiply ranks.&dquo; Blalock (1979: 17) puts it as follows: &dquo;When we translate order relations into mathematical operations, we cannot, in general, use the usual operations of addition, subtraction, multiplication, and division.&dquo;
Here we will make some immediately relevant comments, the first of which may be extremely disturbing to researchers who have been persuaded they are making no assumptions in using
155
nonparametric statistics or distribution-free statistics. According to the statements by Nunnally and Blalock quoted above, computing a tau, a rank correlation statistic, or a Wilcoxon Test makes no sense at all if one thinks one is dealing with ordinal-level measurement. However, in computing these statistics, the ranks are added-and more. This means that least one important error is involved: the researcher has erroneously labeled the scale as ordinal. The other obvious error may be that the person is
applying the operations of algebra where they should not be applied. If the operations are examined, it is clear that a form of interval measurement is being utilized that presumably distorts the true interval measurement, because the ordinary integer number system is usually applied directly to the so-called ranks. In particular, a unit distance is being placed between each of the elements or observations.
Let us illustrate the point simply. Assume that we have 3 individuals of different heights, and by observing them we can order them into ranks 1, 2, and 3. Furthermore say that 3 is greater than 2 is greater than 1. That is, we observe that 3 is some height greater than 2, which is some height greater than 1. However, information about distance either is not recorded or is discarded, and only the information concerning greater than and less than is retained. If anything is to be done with this set of ranks, numbers are allocated, and these numbers are in terms of unit differences between the ranks. The operations most people carry out, and all ordinal scales, should not be thought of as having only some intrinsic quality of greater than and less than. Often, in fact, they are simply very poorly devised interval scales which result from the method of collecting or analyzing the information. Does the fact that people are given rank order in height in any way vitiate the idea that height is to be measured in the population as an interval scale? Emphatically, no. Skeptics should try to compute ordinal statistics using the alphabet, which can also be ordered, but resists addition and subtration. If greater than and less than are all that is involved, the substitution should be simple.
156
SOME OPERA TIONS AND
THE MODEL OF MEASUREMENT
The issue of an appropriate model of measurement has been dealt with briefly, and now we can progress to some operations that may be seen to generate measures that, for various reasons, will be convenient. Suppose that we think of a latent variable X on which we wish to order individuals. This variable presumably can be tapped in a number of ways. But let us assume that the items will be answered as simple dichotomies-yes or no. Assume also that we can find dichotomies that measure this variable with reasonable reliability. In addition, for the sake of convenience, the variables are all chosen so that each divides the population exactly in half on the variable X. (It is clear that we have a special case with perfectly reliable data, none of which has ever been known, and that some other observations could be detailed here.)
Because the variables are not perfect measures of X, if we assume that each is answered independently of each other, then we could expect a certain progression of events. First, those persons who are exactly at the middle point of the population would have a hard time deciding whether they are yeses or nos. Thus, for that particular group, one might expect that yeses and nos might be given randomly. Indeed, the expectation would be that over a long series of questions, half would be answered yes and half would be answered no. On the no side, the closer that one is to the decision point on the
dimension-that is, the amount of the dimension which is measured that divides the population in half-the closer the individual would be in a long series of responses of giving half yeses and half nos, but one would expect more nos than from those who are exactly at the division point. It then follows immediately that the further a person is in the direction of no on the latent variable X, the higher the proportion of no answers he would be expected to give. Indeed, if it is assumed that individuals are distributed in some way along the latent X axis, a smooth curve that will unavoidably remind one of the normal curve will manifest itself. In fact, even if the questions were totally unreliable, the Central Limit Theorem guarantees that as the
157
number of items increases, the distribution of their sum ap- proaches normality.
The point here is that if we deal with the assumption of a large number of dichotomous items assumed to be equally reliable, for a variable X, we expect a subsequent distribution that has characteristics approaching those of a normal curve. We use the weak language at the end of the sentence because it is appropriate to generalize this immediately. Suppose all the variables are not at exactly one-half the cutting point of the population-what would happen? The answer, relatively simply, is that whatever was operating in the original example would continue to operate, but because of the different cutting points and the possibility that quite different combinations of items might be included, the resulting curve could be spread, flattened, made irregular, or distorted in other ways. The distortion, however, is not likely to be enormous if there are many items and if the items do not deviate radically from dividing the population in half. And, if we approximate such normally distributed scores, it is difficult to argue that the assumptions for parametric analyses are not approximately satisfied. It is granted without question that the measures are not exactly isomorphic with the normal model, and therefore efficiency will be lost to the extent that there is a lack of correspondence. But, again, this may be thought of as a form of measurement error, and in no way suggests that ordinal level statistics are more appropriate to analyze the data.
Measurement error may be increased by using fewer items. For example, assuming 20 equally reliable items, the measure must be better than if there are 10. It follows that 5 items would be worse, 2 even worse, and 1 item would represent the worst possible condition. But what is absolutely clear is that what was true of 20 items, of 10 to a lesser extent, of 5 to a still lesser extent, must still be true, although to an even lesser extent, of the single item. It is still a measure corresponding to a model which would produce a normal distribution, but it is a degenerate case in which the measurement has become as poor as possible-the single item dichotomy. The underlying model has not changed; all that has occurred is that in developing the scale, the researcher has elected, either out of ignorance or by design, to use a single item, and
158
therefore the score that is generated is less good that it could have been.
It is now possible to generalize the above results. If trichoto- mies or fourfold response categories are utilized, using large numbers of items, the distribution generated by their sum will also be roughly normal. But because more precision is possible by using the additional cutting points, one can expect a more efficient (more reliable) score X than we would expect if using an equal number of dichotomies (assuming that question content is held constant and only the response categories are changed). The important point here is that a bell-shaped distribution will be generated by the arbitrarily scored dichotomies as well as with items with more ordered categories.
In building scales, it is rather a ridiculous question to ask why one should arbitrarily allocate the value 0 to no and 1 to yes, then raise the question of why one should not be able to allocate the values of 0 to no, 1 to maybe, and 2 to yes, on the grounds that ostensibly equal distances do not exist between the midpoints of those categories of response. The question is not whether or not there is error in such an allocation of numbers, but whether or not using that allocation contributes to the unreliability of the resulting measure. In other words, persons who had been following the common sense procedure should not have felt guilty about doing so.
INAPPROPRIATE CRITICISMS OF INTERVAL MEASURES
The problem of making an appropriate distinction between ordinal and interval measurement is commonplace, and unfor- tunately, among some scholars, it may be highly visible and influential in presentation. For example, Blalock (1979) appar- ently draws his distinction between ordinal and interval from Cohen and Nagel (1934). The latter writers, however, draw a distinction between what they call intensive and extensive qualities, and in the former designation they include temperature and density, while in the latter they include lengths, time intervals, areas, angles, electric current, and electric resistance. The distinc- tion made by Cohen and Nagel, which is not discussed here, is the
159
distinction between nonratio (less than ratio) and ratio measure- ment. Confusion of this distinction with definitions of ordinal and interval is visible in Blalock’s thinking (1979: 18) when he states that &dquo;this means that it is possible to add or subtract scores in an analogous manner to the way we can add weights on a balance or subtract 6 inches from a board by sawing it in two [Refer- ence to Cohen and Nagel].&dquo; The examples are the kind associated with ratio measurement, rather than interval measurement.
Blalock also perpetuates a common fallacious criticism of a common score encountered in social science; namely, the IQ score. This is done, however, in a manner that can only be said to reflect the lack of attention to the distinction he outlines between theoretical definitions and operational definitions. Accordingly, in one sentence ( 1979: 18) he states that &dquo;there are no such interval units of intelligence, authoritarianism, or prestige,&dquo; but directly following that he refers to the IQ score as a specific operational measure: &dquo;Similarly, we can add the incomes of husband and wife, whereas it makes no sense to add their IQ scores.&dquo; Note here that income happens to be a ratio scale, as implied, and IQ scores are an operational definition, presumably of intelligence. If it does not make sense to add IQ scores, does it make sense to add other properties of the husband and wife that can be measured at the ratio level? For example, does adding their weights make one larger, heavier person, presumably analogous to one larger total income? Does adding their ages make one older age? This simply is not the question that should be asked to determine whether weight, age, temperature, and IQ scores are interval measures.
Another point that Blalock (1979: 23) makes is interesting, but may be misleading. He states that &dquo;ideally, one should make use of a data-gathering technique that permits the lowest levels of measurement, if these are all the data will yield, rather than using techniques which force a scale on the data.&dquo; The problem is raised by the word &dquo;ideally,&dquo; as research never seems to be carried out under ideal conditions. More to the point, the data-gathering approach should maximize the amount of information that can be gathered given the limiting circumstances under which mea- surement will be carried out.
This, then, requires pragmatic decisions. For example, a method of paired comparisons may be discarded because of
160
cumbersome procedures-procedures which, of course, may become virtually impossible if large numbers of comparisons are to be carried out. What is needed is some implicit notion of efficiency. The question is whether the more cumbersome procedure is worth the presumed increment of accuracy in ranking over the direct allocation of a set of ranks, balanced against other possible losses. However, choice of the paired comparison procedure, (or of rankings) would make sense only if more information could not be obtained by some other proce- dure, such as the allocation of scores implying some notion of distance (with error, of course).
SUMMARY
In our opinion, the following is worth stating: what makes an appropriate ordinal scale is not merely the assignment of ranks to observations. An ordinal scale is appropriate if we assume that only the properties of &dquo;greater than&dquo; and &dquo;less than&dquo; define an underlying latent construct. We doubt this is the case for most variables of interest to social scientists. As it seems to us that most constructs are conceptualized as continuous and can be thought of as reasonably distributed in the population using a bell-shaped curve as a model, we see no reason not to analyze the manifest data using parametric statistics, even though they are imperfect interval-level scales.
REFERENCES
BLALOCK, H. M., Jr. (1979) Social Statistics. New York: McGraw-Hill. BOHRNSTEDT, G. W. and T. M. CARTER (1971) Robustness in regression analysis,"
pp. 118-146 in G. W. Bohrnstedt and E. F. Borgatta (eds.) Sociological Methodology. San Francisco: Jossey-Bass.
COHEN, M. R. and E. NAGEL (1934) An Introduction to Logic and Scientific Method. New York: Harcourt Brace Jovanovich.
GAITO, J. ( 1980) "Measurement scales and statistics: resurgence of an old misconcep- tion." Psych. Bull. 87: 564-567.
NUNNALLY, J. C. (1967) Psychometric Theory. New York: McGraw-Hill. STEVENS, S. S. (1975) Psychophysics: Introduction to Its Perceptual, Neural and
Social Prospects. New York: John Wiley. ——— [ed.] (1966) Handbook of Experimental Psychology. New York: John Wiley.