2-3 pages APA format, masters level, review attachments

profileashley772
effectsize.pdf

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

Law and Human Behavior, Vol. 29, No. 5, October 2005 ( C© 2005) DOI: 10.1007/s10979-005-6832-7

Comparing Effect Sizes in Follow-Up Studies: ROC Area, Cohen’s d, and r

Marnie E. Rice1,2 and Grant T. Harris1

In order to facilitate comparisons across follow-up studies that have used different measures of effect size, we provide a table of effect size equivalencies for the three most common measures: ROC area (AUC), Cohen’s d, and r. We outline why AUC is the preferred measure of predictive or diagnostic accuracy in forensic psychology or psychiatry, and we urge researchers and practitioners to use numbers rather than verbal labels to characterize effect sizes.

KEY WORDS: effect size; ROC area; risk assessment; predictive accuracy.

In studies of prediction, the key issue is how well different methods perform. How should investigators quantify predictive accuracy, and what is the best way to communicate about accuracy? In the field of risk assessment, for example, how well do different methods predict who will recidivate within a given follow-up time? Many investigators in the field of law and human behavior, and especially in the area of risk assessment, are heeding recent advice about improving the accuracy of deci- sions (Mossman, 1994; Rice & Harris, 1995) by reporting effect size in terms of area under the receiver operating characteristic (ROC area or AUC). To facilitate com- parisons among the most common current measures of effect size—AUC, Cohen’s d, and r—we compiled information from various sources into Table 1. The table ap- plies to any ordinal or continuous predictor variable and a dichotomous outcome.3

Thus, the r provided in the table is rpb, the point-biserial correlation, and assumes a 50% base rate for the outcome.

Cohen (1969, 1988, 1992) provided widely cited (and easily understood by non- statisticians) discussions of effect sizes for what were until recently its most common measures: r (the Pearson product-moment correlation coefficient) and d (known

1Mental Health Centre, Penetanguishene, Ontario, Canada. 2To whom correspondence should be addressed at Mental Health Centre, 500 Church St., Penetan- guishene, Ontario, Canada L9M 1G3; e-mail: [email protected].

3Strictly speaking, d values pertain only to variables scored on an interval scale. When the nondichoto- mous variable is ordinally scaled, r or AUC should be used. Nevertheless, the values in Table 1 allow one to compare the relative magnitudes across studies that have reported any of the three effect size measures.

615

0147-7307/05/1000-0615/1 C© 2005American Psychology-Law Society/Division 41 of the American Psychological Association

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

616 Rice and Harris

Table 1. The Relationships Among AUC, and Two More Longstanding Measures of Effect Size, Cohen’s d, and the Point-Biserial Correlation Coefficient, rpb

z(AUC)a AUC d rpb z(AUC) AUC d rpb

.00 .500 .0 .0 .67 .749 .948 .428

.01 .504 .014 .007 .68 .752 .962 .433

.02 .508 .028 .014 .69 .755 .976 .438

.03 .512 .042 .021 .70 .758 .990 .444

.04 .516 .057 .028 .71 .761 1.00 .449

.05 .520 .071 .035 .73 .767 1.03 .459

.06 .524 .085 .042 .75 .773 1.06 .469

.07 .528 .099 .049 .76 .776 1.07 .473

.08 .532 .113 .056 .77 .779 1.09 .478

.09 .536 .127 .064 .79 .785 1.12 .488

.10 .540 .141 .071 .81 .791 1.15 .497

.11 .544 .156 .078 .82 .794 1.16 .502

.12 .548 .170 .085 .83 .797 1.17 .506

.13 .552 .184 .092 .85 .802 1.20 .515

.14 .556 .198 .099 .86 .805 1.22 .520

.141 .556 .200 .100 .88 .811 1.24 .528

.15 .560 .212 .105 .89 .813 1.26 .533

.16 .564 .226 .112 .90 .816 1.27 .537

.17 .568 .240 .119 .92 .821 1.30 .545

.18 .571 .255 .126 .94 .826 1.33 .554

.19 .575 .269 .133 .95 .829 1.34 .558

.20 .579 .283 .140 .96 .831 1.36 .562

.21 .583 .297 .147 .97 .834 1.37 .566

.22 .587 .311 .154 .98 .836 1.39 .570

.23 .591 .325 .161 .99 .839 1.40 .573

.24 .595 .339 .167 1.00 .841 1.41 .577

.25 .599 .354 .174 1.02 .846 1.44 .585

.26 .603 .368 .181 1.03 .848 1.46 .589

.27 .606 .382 .188 1.05 .853 1.48 .596

.28 .610 .396 .194 1.06 .855 1.50 .600

.29 .614 .410 .201 1.08 .860 1.53 .607

.30 .618 .424 .208 1.10 .864 1.56 .614

.31 .622 .438 .214 1.11 .867 1.57 .617

.32 .626 .453 .221 1.14 .873 1.61 .628

.33 .629 .467 .227 1.16 .877 1.64 .634

.34 .633 .481 .234 1.17 .879 1.65 .637

.35 .637 .495 .240 1.20 .885 1.70 .647

.354 .639 .500 .243 1.21 .887 1.71 .650

.36 .641 .509 .247 1.23 .891 1.74 .656

.37 .644 .523 .253 1.24 .893 1.75 .659

.38 .648 .537 .259 1.26 .896 1.78 .665

.39 .652 .552 .266 1.30 .903 1.84 .677

.40 .655 .566 .272 1.31 .905 1.85 .680

.41 .659 .580 .278 1.33 .908 1.88 .685

.42 .663 .594 .285 1.37 .915 1.94 .696

.43 .666 .608 .291 1.38 .916 1.95 .698

.44 .670 .622 .297 1.41 .921 1.99 .706

.45 .674 .636 .303 1.44 .925 2.04 .713

.46 .677 .651 .309 1.45 .926 2.05 .716

.47 .681 .665 .315 1.49 .932 2.11 .725

.48 .684 .679 .321 1.52 .936 2.15 .732

.49 .688 .693 .327 1.53 .937 2.16 .734

.50 .691 .707 .333 1.54 .938 2.18 .737

.51 .695 .721 .339 1.58 .943 2.23 .745

.52 .698 .735 .345 1.60 .945 2.26 .749

.53 .702 .750 .351 1.63 .948 2.31 .755

.54 .705 .764 .357 1.67 .953 2.36 .763

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

Comparing Effect Sizes 617

Table 1. Continued

z(AUC)a AUC d rpb z(AUC) AUC d rpb

.55 .709 .778 .362 1.68 .954 2.38 .765

.56 .712 .792 .368 1.70 .955 2.40 .769 .566 .714 .800 .371 1.74 .959 2.46 .776 .57 .716 .806 .374 1.80 .964 2.55 .786 .58 .719 .820 .379 1.81 .965 2.56 .788 .59 .722 .834 .385 1.82 .966 2.57 .790 .60 .726 .849 .391 1.86 .969 2.63 .796 .61 .729 .863 .396 1.88 .970 2.66 .799 .62 .732 .877 .402 1.92 .973 2.72 .805 .63 .736 .891 .407 1.95 .974 2.76 .810 .64 .739 .905 .412 1.96 .975 2.77 .811 .65 .742 .919 .418 1.99 .977 2.81 .815 .66 .745 .933 .423

Note. Because the values of all three indices change more slowly beyond what Cohen calls a large effect, and to save space, for values beyond large we include only those rows for which the value of at least one of the three effect size indices changes when rounded to two significant digits. The rows in bold correspond to minimum values for small, medium, and large effects. aThe first column of the table is z(AUC), the normal deviate of AUC based on the assumption that the predictor variable is normally distributed within each of the outcome groups. Under this assumption, the values of z(AUC) and AUC may be obtained from the z and F(z) columns respectively of Pearson and Hartley (1954) and reprinted in many statistical textbooks.

as Cohen’s d). Cohen stated that the values of d for small, medium, and large ef- fects, respectively, are .2, .5, and .8, and that the corresponding values for r are .1, .3, and .5. However, neither we (Rice & Harris, 1995) nor others (e.g., Hemphill, 2003) realized that Cohen first derived different corresponding values for compar- ing r to d (e.g., Cohen, 1988, p. 82) using point-biserial correlations (rpb), for the case where there is one ordinal or continuous variable and one dichotomous vari- able. As shown in the table, the relevant values of rpb corresponding to the d values for small, medium, and large effects are .10, .243, and .371, respectively.4 Because correlations between two continuously distributed variables are more common and because he wanted to characterize effect sizes in general, Cohen corrected these point-biserial correlations (multiplying by the required 1.253 correction factor for a 50% base rate) to approximate the familiar .1, .3, and .5 guideline for rs calculated from two continuous variables (Cohen, 1969, 1988). However, the smaller values of r in the table are correct for the d and r relationship for dichotomous outcomes (in- cluding recidivism). It is also important to note that for base rates other than 50%, the rpb values for small, medium, and large effect are even smaller than those in the table. For example, using the formulae provided in the Appendix, readers can calculate that for a 25% base rate (perhaps more typical of violent and sexual re- cidivism studies), the rpb values for small, medium, and large are .086, .212, and .327, respectively.

Most readers will be familiar with the suggestion that an index of predictive ac- curacy is Pearson r2, the amount of variance accounted for (e.g., Berlin, Galbreath,

4Assuming a 50% base rate. For base rates other than 50%, the formulae in the Appendix can be used to calculate rpb values corresponding to values of d and AUC presented in the table. Note also that the difference between the values of r and rpb required to equate to values of d also reveal the power lost when continuous variables are dichotomized unnecessarily.

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

618 Rice and Harris

Geary, & McGlone, 2003). Rosenthal and Rubin (1982) emphasize that, in many circumstances, r2 gives a gross misrepresentation of the importance of a finding, particularly when one of the variables is dichotomous. Rosenthal (1990) cited the example of a study of aspirin and heart attack in which an experiment was termi- nated prematurely because it had become so clear that aspirin was effective that it would have been unethical to continue to administer placebos to the control par- ticipants (who, interestingly, were physicians). The Pearson r in that study was only .034, and r2 only .001. Rosenthal and Rubin (1982) pointed out that this reflected a 3.4% decrease in heart attacks between treated versus untreated physicians. Given the importance of the problem and the low cost of the intervention, this result has led many of us to take a daily dose of aspirin. Attending only to the squared corre- lation coefficient would have led to this finding being ignored. Were such a squared correlation to be observed in a study of violent recidivism, a potentially useful risk factor might be neglected (if either the factor occurred at a low rate or the study obtained a low base rate of recidivism). Such an error would be less likely, however, with the use of d or AUC; in this case, .30 and .58, respectively.5

Among the most commonly used effect size measures (or, perhaps more pre- cisely here, indices of detection performance), which measure do we recommend? Cohen’s d was designed for use where scores of the two populations being compared are continuous and normally distributed, a condition seldom met in such applied ar- eas as risk assessment. The size of a correlation coefficient as an indication of effect size depends upon both the particular correlation coefficient and the base rate. Thus, we, as have others (Mossman, 1994; Swets, Dawes, & Monahan, 2000) recommend the use of AUC as the preferred measure of predictive or diagnostic accuracy in forensic psychology and psychiatry. AUC also has a plain language interpretation because it is conceptually and mathematically similar to the common language ef- fect size or CL (McGraw & Wong, 1992; see also Delaney & Vargha, 2002). AUC equals the probability that a score (on an ordinal or continuous measure such as a risk-assessment instrument) drawn at random from one sample or population (e.g., recidivists’ scores) is higher than that drawn at random from a second sample or population (e.g., nonrecidivists’ scores). Use of a consistent measure would facili- tate understanding not only among expert witnesses, but also among other criminal justice officials.

Cohen (1969) provided his rule of thumb tentatively, suggesting that a barely perceptible (but real) difference was “small.” Cohen’s examples of a small effect are the difference in average IQ between twins and nontwins and the difference in av- erage height (about half an inch) between 15- and 16-year-old girls (Cohen, 1988). Cohen described a “medium” effect size as one “large enough to be visible to the naked eye” (Cohen, 1988, p. 26); as an example, he offers the difference in average height (about 1 inch) between 14- and 18-year-old girls. Cohen described a “large” effect as grossly perceptible, as, for example, the difference in average height be- tween 13- and 18-year-old girls (about two inches), or the average difference in IQ between persons with a Ph.D. and first year college students. Cohen also stated

5Assuming a continuously distributed risk factor and a base rate equal to the aspirin and heart attack study (.013%).

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

Comparing Effect Sizes 619

that in applied psychology, effect sizes of d = .8 are “about as high as they come” (Cohen, 1988, p. 81). Cohen also pointed out that labels depend on empirical and social contexts. Early in the study of a phenomenon when methodological control is nascent, standards may be lower than when instrumentation and control have ad- vanced. This seems quite true of the field of risk assessment research. We regarded (Rice & Harris, 1995) the predictive validity of the Violence Risk Appraisal Guide as “large” because d values greater than .80, or their equivalents, have been com- monly reported (www.mhcp-research/ragreps.htm), but further research has sug- gested how even larger effects could be achieved (Harris & Rice, 2003).

The verbal characterizations proposed by Cohen could be very useful when communicating to laypersons such as court officials, for example, if they were com- monly understood and adhered to by researchers and practitioners. However, we suggest the field of risk assessment place little reliance on plain language verbal la- bels because of the considerable disagreement about what they mean. For example, our colleagues (Hilton, Carter, Harris, & Bryans, 2005) found that, among a small group of forensic clinicians, some risks (i.e., probabilities between .40 and .55 of violent recidivism in 10 years of opportunity) were simultaneously characterized as “high,” “moderate,” and “low” by different clinicians. Clarity is best reflected by nu- merical characterization and we offer the accompanying table to aid communication among researchers and practitioners.

APPENDIX

The conversion formulae (after Rosenthal, 1991; Swets, 1986):

r = d√ d2 + (1/pq)

,

where p = the base rate, q = 1 − p, and r is the point-biserial r d = r√

pq(1 − r2) ,

where p = the base rate, q = 1 − p, and r is the point-biserial r d =

√ 2 × z(AUC),

where z(AUC) is the z-transform or normal deviate of AUC, and where variances of the two populations are equal.

REFERENCES

Berlin, F. S., Galbreath, N. W., Geary, B., & McGlone, G. (2003). The use of actuarials at civil com- mitment hearings to predict the likelihood of future sexual violence. Sexual Abuse: A Journal of Research and Treatment, 15, 377–382.

Cohen, J. (1969). Statistical power analysis for the behavioral sciences. New York: Academic Press. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Erlbaum. Cohen, J. (1992). A power primer. Psychological Bulletin, 122, 155–159. Delaney, H. D., & Vargha, A. (2002). Comparing several robust tests of stochastic equality with ordinally

scaled variables and small to moderate sized samples. Psychological Methods, 7, 485–503.

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

620 Rice and Harris

Harris, G. T., & Rice, M. E. (2003). Actuarial assessment of risk among sex offenders. In R. A. Prentky, E. S. Janus, & M. C. Seto (Eds.), Understanding and managing sexually coercive behavior, Vol. 989 (pp. 198–210). New York: Annals of the New York Academy of Sciences.

Hemphill, J. F. (2003). Interpreting the magnitudes of correlation coefficients. American Psychologist, 58, 78–80.

Hilton, N. Z., Carter, A. M., Harris, G. T., & Bryans, A. (2005). Using categorical judgments to commu- nicate risk of violence. Unpublished manuscript.

McGraw, K. O., & Wong, S. P. (1992). A common language effect size statistic. Psychological Bulletin, 111, 361–365.

Mossman, D. (1994). Assessing predictions of violence being accurate about accuracy. Journal of Con- sulting and Clinical Psychology, 62, 783–792.

Pearson, E. S., & Hartley, H. O. (Eds.). (1954). Biometrika tables for statisticians, Vol. 1 (1st ed.). Cambridge: Cambridge University Press.

Rice, M. E., & Harris, G. T. (1995). Violent recidivism: Assessing predictive validity. Journal of Consult- ing and Clinical Psychology, 63, 737–748.

Rosenthal, R. (1990). How are we doing in soft psychology? American Psychologist, 45, 775–777. Rosenthal, R. (1991). Meta-analytic procedures for social research. Newbury Park, CA: Sage. Rosenthal, R., & Rubin, D. B. (1982). A simple, general purpose display of magnitude of experimental

effect. Journal of Educational Psychology, 74, 166–169. Swets, J. A. (1986). Indices of discrimination or diagnostic accuracy: Their ROCs and implied models.

Psychological Bulletin, 99, 100–117. Swets, J. A., Dawes, R. M., & Monahan, J. (2000). Psychological science can improve diagnostic decisions.

Psychological Science in the Public Interest: A Journal of the American Psychological Society, 1, 1– 26.

  • Untitled
  • Untitled
  • Untitled
  • Untitled
  • Untitled
  • Untitled
  • Untitled