2-3 pages APA format, masters level, review attachments

profileashley772
quantitative.pdf

O R I G I N A L A R T I C L E

Quantifying the Accuracy of Forensic Examiners in the Absence of a ‘‘Gold Standard’’

Douglas Mossman Æ Michael D. Bowen Æ David J. Vanness Æ David Bienenfeld Æ Terry Correll Æ Jerald Kay Æ William M. Klykylo Æ Douglas S. Lehrer

Published online: 22 September 2009

� American Psychology-Law Society/Division 41 of the American Psychological Association 2009

Abstract This study asked whether latent class modeling

methods and multiple ratings of the same cases might per-

mit quantification of the accuracy of forensic assessments.

Five evaluators examined 156 redacted court reports

concerning criminal defendants who had undergone

hospitalization for evaluation or restoration of their adju-

dicative competence. Evaluators rated each defendant’s

Dusky-defined competence to stand trial on a five-point

scale as well as each defendant’s understanding of, appre-

ciation of, and reasoning about criminal proceedings.

Having multiple ratings per defendant made it possible to

estimate accuracy parameters using maximum likelihood

and Bayesian approaches, despite the absence of any ‘‘gold

standard’’ for the defendants’ true competence status.

Evaluators appeared to be very accurate, though this finding

should be viewed with caution.

Keywords Competence to stand trial � Adjudicative competence � ROC analysis � Diagnostic accuracy � Maximum likelihood � Bayesian � Gold standard

Daubert v. Merrell Dow Pharmaceuticals (1993) directs

judges to evaluate proffered scientific testimony based on

factors such as ‘‘the known or potential rate of error’’ of the

‘‘particular scientific technique’’ and whether the technique

has been tested (pp. 592–593). A subsequent U.S. Supreme

Court decision, Kumho Tire Co. v. Carmichael (1999),

extended trial courts’ gate-keeping role and the applica-

bility of Daubert factors to ‘‘other experts who are not

scientists’’ (p. 137). Thus, in federal courts and other U.S.

jurisdictions that follow Daubert-like evidentiary rules,

testifying mental health professionals may be asked,

‘‘Doctor, has your method been tested?’’ and ‘‘How accu-

rate is it?’’

Practitioners of most medical specialties use diagnostic

methods for which accuracy statistics such as sensitivity

and specificity are available. Also, for many diagnostic

modalities—e.g., mammography for breast cancer detec-

tion (Berg et al., 2008; Lehman et al., 2007)—long-term

Portions of this work were presented at (1) Annual Meeting of the

American Academy of Psychiatry and the Law, Miami Beach,

Florida, October 18, 2007; (2) American Psychology-Law Society

Conference, Jacksonville, Florida, March 7, 2008; (3) Annual

Meeting of the Midwest Chapter of the American Academy of

Psychiatry and the Law, Renaissance Hotel, Cleveland, Ohio, March

29, 2008; (4) Cincinnati Psychiatric Society, June 17, 2008; and

(5) Summit Behavioral Healthcare, June 11, 2009.

D. Mossman (&) Glenn M. Weaver Institute of Law and Psychiatry, University of

Cincinnati College of Law, Clifton Avenue & Calhoun Street,

PO Box 210040, Cincinnati, OH 45221-0040, USA

e-mail: [email protected]

D. Mossman

Department of Psychiatry, University of Cincinnati College of

Medicine, Cincinnati, USA

D. Mossman � M. D. Bowen � D. Bienenfeld � T. Correll � J. Kay � W. M. Klykylo � D. S. Lehrer Department of Psychiatry, Wright State University, Boonshoft

School of Medicine, Dayton, USA

D. J. Vanness

Department of Population Health Sciences, University of

Wisconsin School of Medicine and Public Health, Madison,

USA

D. J. Vanness

Center for Health Economics and Science Policy, United

BioSource Corporation, Bethesda, USA

123

Law Hum Behav (2010) 34:402–417

DOI 10.1007/s10979-009-9197-5

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

follow-up or biopsy results provide an independent, virtu-

ally infallible criterion for true disease status. By contrast,

psychiatric classifications usually have no independent

‘‘gold standard’’ (Faraone & Tsuang, 1994), and diagnostic

criteria consist entirely of clinical findings.

Where psycholegal determinations are concerned, the

absence of an indubitable truth criterion is both common

and arguably more troubling than in clinical mental health

practice. Ethical principles of forensic psychiatrists and

forensic psychologists encourage adherence to high stan-

dards of competence, impartiality, and scientific rigor

(American Academy of Psychiatry and the Law, 2005;

Committee on the Revision of the Specialty Guidelines for

Forensic Psychology, 2006; Weissman & DeBow, 2003).

Yet popular works (Hagen, 1997), appellate-level legal

opinions (Mossman, 1999), and evidence texts (Faigman,

Saks, Sanders, & Cheng, 2008) suggest that psycholegal

experts seem like ‘‘whores’’ and ‘‘hired guns’’ who ‘‘are

merely selling their testimony to the highest bidder’’

(Melton et al., 2007, p. 577). When opposing experts dis-

agree, courtroom cross-examination often becomes an

intensive effort to question the integrity of psychiatric

diagnoses and to discredit all mental health expertise.

Several instruments relevant to forensic assessment

offer bases for expert opinion that appear more reliable and

systematic than unaided clinical judgment (Dawes, Faust,

& Meehl, 1989; Gardner, Lidz, Mulvey, & Shaw, 1996;

Harris, Rice, & Cormier, 2002). Over the last two decades,

many publications have described the accuracy of such

instruments, often using receiver operating characteristic

(ROC) methods (Douglas, Ogloff, Nicholls, & Grant, 1999;

Rice & Harris, 1995). For tools used to assess violence risk,

subjects’ arrest or conviction records supplemented with

interview and collateral data have served as external cri-

teria for gauging accuracy (Steadman et al., 1998). In many

circumstances, however, the closest approximation to truth

is a professional’s well-considered opinion. In several

studies that examine accuracy of assessment tools or pre-

dictions (e.g., Akinkunmi, 2002; Kim et al., 2007;

Mossman, 2007), judgments of experienced clinicians have

provided the ‘‘truth’’ or ultimately ‘‘right’’ conclusion. Yet

the wisest experts err in their clinical and forensic assess-

ments, and though judges or juries render ultimately

binding decisions in courtrooms, legal fact-finders are

humanly fallible. Thus, for most psycholegal determina-

tions, all we can hope for are various individuals’

conclusions.

Faced with this epistemological barrier, several studies

(e.g., Cooper & Zapf, 2003; Jacobs, Ryba, & Zapf, 2008)

have looked for correlates of or factors related to forensic

assessment tools or examiners’ opinions. Other studies

(e.g., Boccaccini, Turner, & Murrie, 2008; Murrie, Boc-

caccini, Zapf, Warren, & Henderson, 2008; Murrie et al.,

2009) have quantified absolute diagnostic agreement

between evaluators and have examined possible sources of

disagreement. Though several of these studies recognize

and discuss consequences of rater bias and disagreement,

they have been agnostic (so to speak) about whether a

question such as ‘‘Is this defendant competent to stand

trial?’’ has a ‘‘right’’ answer. Yet the existence of a right

answer is a condition of the possibility of thinking—as any

experienced forensic examiner occasionally does—that a

court has made an error in finding a particular criminal

defendant competent or incompetent. And if we acknowl-

edge that judgments about psycholegal matters can be right

or wrong, we can also wonder how accurate such judg-

ments are, despite our having no gold standard to establish

absolute truth in any particular case.

Over the last two decades, latent class modeling (LCM)

(Uebersax & Grove, 1990) has shown promise in permit-

ting ROC analyses without gold standards in subject areas

as diverse as imaging liver metastases (Henkelman, Kay, &

Bronskill, 1990) and detecting infections in dairy cattle

(Choi, Johnson, Collins, & Gardner, 2006). Implementing

the LCM approach involves evaluating the same cases with

multiple diagnostic modalities, which often permits statis-

tical identification of models that includes accuracy

parameters for those modalities. To learn whether LCM

could characterize and quantify the accuracy of forensic

examiners, the present study considers the most-often-

performed (Melton et al., 2007; Mossman et al., 2007)

criminal forensic evaluation: assessing competence to

stand trial.

BACKGROUND

Legal Criteria

All U.S. jurisdictions (Bennett, 1985) define competence to

stand trial (CST) consistent with Dusky v. United States,

which states that a defendant is CST if ‘‘he has sufficient

present ability to consult with his lawyer with a reasonable

degree of rational understanding’’ and ‘‘has a rational as

well as factual understanding of the proceedings against

him’’ (Dusky v. United States, 1960, p. 402). Whenever ‘‘a

bona fide doubt’’ about a criminal defendant’s fitness

arises, the trial court must hold a hearing concerning his

CST (Pate v. Robinson, 1966). All U.S. jurisdictions permit

courts to order mental health evaluations of criminal

defendants for use in hearings on CST (Mossman et al.,

2007). Following findings of incompetence, courts usually

order defendants to undergo ‘‘restoration’’—treatment,

usually at a public sector hospital, aimed at rendering the

defendants competent—if such treatment has a substantial

probability of being successful (Jackson v. Indiana, 1972;

Law Hum Behav (2010) 34:402–417 403

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

Mossman et al., 2007; State v. Sullivan, 2001). Around

60,000 U.S. criminal defendants undergo CST evaluations

each year, and roughly 4,000 U.S. hospital beds are

occupied by defendants who are undergoing CST restora-

tion (Mossman, 2005). Hospitals send periodic reports to

referring courts concerning defendants’ progress toward

achieving competence (Parry & Drogin, 2007).

Conceptualizing Experts’ Assessments of CST

When forensic examiners assess CST, they (explicitly or

implicitly) consider many mental faculties and several

dimensions of social and interpersonal functioning,

including defendants’ abilities to comprehend social situ-

ations, to project themselves into hypothetical situations, to

function in collaborative relationships, to recognize what

things are relevant in complex social situations, to com-

municate logically, and to maintain self-control (Mossman,

2008). For this reason, writers have characterized adjudi-

cative competence as an abstract, ‘‘open-textured’’

construct (Bonnie, 1990) intended ‘‘to apply to an infinite

number of fact situations’’ (Golding, Roesch, & Schreiber,

1984, p. 323), or as some actual though hypostatized fea-

ture of defendants (Grisso, 2003). On this view,

adjudicative competence is what Cronbach and Meehl call

a ‘‘postulated attribute’’ (1955, p. 283)—an imperfectly

defined but real property of defendants—and judgments

about adjudicative competence (whether made by exam-

iners or courts) reflect beliefs about the degree to which

defendants exhibit this property. Alternatively (says this

view), competent and incompetent defendants may com-

prise two natural, distinct groups or ‘‘taxa,’’ and a decision

about adjudicative competence is a decision about whether

a defendant falls into one taxon or the other.

Following Mossman (2008), however, we regard CST

evaluations as contextual assessments that ask whether a

defendant can do something—meet a standard—rather than

whether the defendant has a property or belongs to one

natural category or another. On our view, to ask the

question ‘‘Is Defendant Jones competent to stand trial?’’ is

like asking whether Jones can jump over a hurdle of a

given height. The features and qualities that allow people

to clear hurdles (their height, weight, leg strength, coor-

dination, etc.) are various and fall along continua, but

whether Jones can clear a particular hurdle is a yes-or-no

matter. If the hurdle is 15 cm (6 in.) high, experience lets

us be very confident that a randomly selected, able-bodied,

middle-aged adult should have no trouble jumping over it,

while toddlers and frail elderly people will not succeed.

The hurdle homology is not perfect, of course. If we

agree on an operational definition of ‘‘able to clear a 15-cm

hurdle’’ (e.g., ‘‘you get ten tries, and you have to succeed

just once’’), we can evaluate Jones against the criterion and

determine definitively whether he can clear the hurdle. By

contrast, evaluating CST requires a forensic examiner to

assess many hard-to-quantify personal qualities of a

defendant against the also-hard-to-quantify demands of a

particular criminal case. In many cases, examiners justifi-

ably feel high confidence about a defendant’s adjudicative

competence or lack thereof. But we have no way to verify

for certain whether a defendant who faces one or more

specific charges understands the nature and objective of the

proceedings against him and can assist his attorney in

preparing a defense. The characterizations of CST found in

statutes and case law guide examiners’ beliefs and judges’

opinions about adjudicative competence, but statutes

and case law do not operationally define adjudicative

competence.

Yet, just as we know that most able-bodied middle-aged

adults can clear 15-cm hurdles, we know that most defen-

dants—indeed, almost all adults who do not have serious

psychopathology or cognitive impairment—are competent

to stand trial. A major portion of a CST assessment involves

determining whether a serious psychiatric disorder or cog-

nitive impairment prevents a particular defendant from

doing what most individuals could easily do. Just as general

experience informs everyone about the personal character-

istics that might preclude someone from jumping over a low

hurdle, training and experience inform mental health pro-

fessionals about the kinds of impairments that would keep a

defendant from doing the basic (though harder to measure)

mental tasks needed to stand trial. For example, verbal

incoherence is hard to quantify, but its adverse impact on

communication—and on adjudicative competence—is easy

to apprehend. This suggests that if forensic evaluators have

adequate information available, they should be able to make

very good (though not perfect) judgments about adjudica-

tive competence.

METHOD

Data Collection and Development

Our study received approval from the Institutional Review

Board of Wright State University and from the Office of

Program Evaluation and Research of the Ohio Department

of Mental Health. We obtained 156 reports on CST that a

public sector hospital had submitted to Ohio criminal

courts in 1994–2001. The reports described criminal

defendants who had undergone court-ordered hospitaliza-

tions either for evaluation of CST or for restoration. In

accordance with state statutory requirements (Ohio

Revised Code §2945.38(F)), each report included the

examiner’s ‘‘penultimate’’ opinion on CST (i.e., a state-

ment about whether the defendant could understand the

404 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

pending proceedings and assist defense counsel). We

selected reports largely at random, though we had to reject

some reports that were poorly written or did not include

enough detail about defendants’ condition, treatment

course, current mental status, and responses to CST-spe-

cific inquiries to allow formation of independent opinions

about adjudicative competence. Previous studies on quan-

tifying accuracy without gold standards (Albert, 2007;

Henkelman et al., 1990; Zhou, Castelluccio, & Zhou, 2005)

suggested that our raters should examine at least 150

reports to facilitate convergence of the mathematical

algorithms used to identify accuracy parameters; including

six additional reports provided a margin of safety in case

raters found some reports unusable.

Two authors prepared ‘‘sanitized’’ versions of each

original report by (1) removing the original examiners’

diagnoses and forensic opinions about CST, (2) substitut-

ing pseudonyms for original names, and (3) approximating,

disguising, paraphrasing, or deleting other identifying data

(e.g., precise ages, ethnicity, dates and locales of previous

hospitalizations). Each sanitized report retained disguised

background information (including medical, legal, and

psychiatric history), though we removed material that was

irrelevant to CST or too personal (e.g., childhood sexual

abuse). Defendants’ criminal charges and sexes were not

changed because these items often were needed to assess

CST or to interpret background information. Sanitized

reports also retained the original examiners’ descriptions of

evaluees’ hospital course, current functioning, mental sta-

tus, responses to questions specifically related to CST (e.g.,

‘‘What are you charged with?’’), results from psychological

testing (e.g., MMPI-2 or intelligence scales), and/or scores

from structured assessment instruments (e.g., the Georgia

Court Competency Test).

Serving as raters were five board-certified, experienced

(12–33 years post-residency) psychiatrists (three also

board-certified in forensic psychiatry) who had not helped

prepare the sanitized reports. After reading each of the 156

reports, raters assigned scores on five-point scales con-

cerning the defendant’s understanding of information

relevant to, ability to reason about, and appreciation of his

current legal situation, applying definitions of these con-

cepts used in the MacArthur studies on adjudicative

competence (Poythress, Monahan, Bonnie, Otto, & Hoge,

2002). Each rater also provided ordinal scale scores

reflecting his overall rating of CST as defined under Dusky

v. United States (1960). Table 1 shows portions of a rater’s

data sheet.

Conceptualizing the Data

Having raters provide impressions using five-point scales

differs from usual forensic experts’ practice and from

courts’ usual expectation that experts will render binary

opinions (either competent or incompetent) with reason-

able medical or scientific certainty. Buchanan (2006) notes,

however, that experts reach yes-or-no opinions with vary-

ing degrees of assuredness; judgments about CST

incorporate subordinate judgments about defendants’

mental functioning, potential penalties, case-specific

demands (e.g., complexity of information that a defendant

must process), and consequences of errors (e.g., possibly

having a marginally incompetent defendant stand trial for a

serious crime). Some evidence suggests that examiners

may require a higher level of competence for defendants

charged with more serious offenses (Rosenfeld & Ritchie,

1998). An expert’s conclusion about a defendant’s CST

thus reflects an estimate of the defendant’s relevant abili-

ties coupled with the expert’s judgments about benefits and

costs of correct and incorrect outcomes. Also, recent pub-

lications suggest that forensic opinions display evidence of

‘‘adversarial allegiance’’ to retaining parties (Murrie et al.,

2009), and that yes-or-no judgments about CST reflect

experts’ professional backgrounds and views about the

impact of psychosis (Murrie et al., 2008). To properly

gauge an expert’s ability to discern competent from

incompetent defendants, one should use a mathematical

method that can tease out intrinsic discriminatory capacity

from biasing factors and the expert’s concerns about con-

sequences of errors (Mossman, 2008).

Over the last four decades, investigators in medicine and

psychology have increasingly turned to ROC analysis

to address this type of problem. ROC analysis separates

Table 1 Portions of the Rater Scoring Sheet, with numbers assigned to each rating category

Understanding

Defendant’s cognitive apprehension, at a descriptive level, of basic

legal functions; his ability to comprehend information relevant

to the adjudicative process

h 1 = definitely satisfactory

h 2 = probably satisfactory

h 3 = uncertain

h 4 = probably unsatisfactory

h 5 = definitely unsatisfactory

Competence to Stand Trial

Overall rating of the defendant’s rational and factual understanding

of the proceedings against him and his ability to consult rationally

with an attorney

h 1 = very likely competent

h 2 = probably competent

h 3 = uncertain

h 4 = probably incompetent

h 5 = very likely incompetent

Law Hum Behav (2010) 34:402–417 405

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

trade-offs between diagnostic sensitivity and specificity in a

diagnostic method from cost–benefit judgments that influ-

ence decisions based on diagnostic data (Obuchowski, 2003;

Swets, 1995; Zweig & Campbell, 1993). A ROC graph plots

sensitivity (the test’s true positive rate, tpr) as a function of a

test’s false positive rate (fpr, equal to 1—specificity); area

under the ROC curve (AUC) is a summary index of overall

diagnostic accuracy (Mossman & Somoza, 1991).

Statistical Model

If five raters provide ratings about 156 defendants’ adju-

dicative competence, one can set out the raters’ scores for

each defendant in a five-element row, or vector. Also, one

can array scores for all defendants in a matrix with

5 9 156 = 780 elements arranged in five columns (one

column for each rater) and 156 rows (one row vector for

each defendant).

Earlier, we said that asking whether a defendant is CST

is homologous to asking whether a person can do some-

thing, such as clear a particular hurdle. We also noted that

even when abilities needed to do various tasks lie along

continua, whether a particular person can do a particular

task may well be answerable with ‘‘yes’’ or ‘‘no.’’ For

purposes of exposition, we now concretize inability to do a

task—here, being incompetent to stand trial—as ‘‘having’’

a condition or disorder D.

Suppose, for a moment, that some ‘‘gold standard’’—

perhaps the declaration of an Omniscient Being—told us

whether a defendant (or ‘‘subject’’) had D, and we wanted

to describe how accurately a particular rater could detect D.

The raters have provided scores expressing their confidence

about presence or absence of D along a five-point scale;

this means that four (fpr, tpr) pairs, or four ‘‘cut-off’’

points, describe each rater’s accuracy in detecting D.

Table 2 contains results for a hypothetical rater and shows

how the rater’s scores translate into four (fpr, tpr) pairs.

Figure 1 illustrates how these four (fpr, tpr) pairs define a

ROC curve that summarizes the rater’s accuracy.

Note that, if an Omniscient Being declared which sub-

jects had D, we would know the prevalence of D in our

particular subject pool. We could also quantify any con-

ditional dependence in raters’ scores that arose from

peculiarities—such as whether the cases in that particular

pool are especially easy or hard—of the competent and

incompetent subjects. Thus, if we knew whether each

subject had D, we could calculate 43 parameters—four

(fpr, tpr) pairs for five raters (4 9 2 9 5 = 40 parame-

ters), plus the prevalence (one parameter), plus values

expressing conditional dependence of the competent and

incompetent subjects (two parameters)—directly from the

780-element matrix.

In reality, we have no truth criterion for the presence or

absence of D, only imperfect human opinion. One response

to this situation, reflected in previously published studies

(e.g., Cooper & Zapf, 2003; Murrie et al., 2008; Skeem,

Golding, Cohn, & Berge, 1998) would be to examine

agreement in and correlates of experts’ opinions about

CST. These studies have used statistics such as kappa and

the intraclass correlation coefficients (ICCs) to quantify

(dis)agreement between experts and to explore factors that

cause them to disagree.

The statistical approach we employ here aims at some-

thing different: quantifying the accuracy of experts, or,

more specifically, gauging how well they can distinguish

competent from incompetent defendants. If one pre-

sumes—despite the absence of a ‘‘gold standard’’—that

having or lacking D is something that experts can get right

or wrong, and that experts actually can distinguish com-

petent from incompetent defendants fairly well, one might

ask, ‘‘What combination of values for the 43 accuracy

parameters would maximize the probability of producing

this 780-element ratings matrix?’’ An answer to this

Table 2 Calculation of accuracy indices based on 156 hypothetical rating results if the actual CST status of evaluees were known

Rating Actual CST status fpr tpr

Competent Not

competent

1 = very likely competent 60 1

2 = probably competent 30 4 0.406 0.982

3 = uncertain 5 5 0.109 0.889

4 = probably incompetent 3 10 0.059 0.818

5 = very likely incompetent 3 35 0.030 0.636

Total 101 55

0

0.5

1

10.50

false positive rate

tr u

e p

o s

it iv

e r

a te

5

4

3 2

1

Fig. 1 ROC graph depicting a rater’s performance in detecting adjudicative competence, based on hypothetical data in Table 2.

Numbers correspond to rating categories in Table 1

406 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

question would provide the most plausible estimates of the

raters’ accuracy parameters, given the data available.

A Pictorial Explanation

We created Fig. 2a–c to help readers get an informal,

intuitive feel of when and why our statistical approach

might succeed. Suppose two fairly accurate forensic

examiners examine 150 defendants, half of whom—though

this is known only to an Omniscient Being—actually lack

adjudicative competence. Each examiner can make finely

graded judgments about CST that can be represented along

a continuous (or at least finely grained) mental decision

scale. In Fig. 2a, the judgments of Rater 2 are plotted as a

function of the judgments of Rater 1. For both examiners,

the impact of judging CST is to shift the distribution of

actually incompetent defendants by two standard devia-

tions along the mental decision axes. For Rater 1, the shift

takes place rightward (in the direction of the arrow) along

the horizontal axis in Fig. 2a, and for Rater 2, upward (in

the direction of the arrow) along the vertical axis. Another

way to think about these shifts is to say that the raters’

judgments about presence and absence of D have an effect

size of 2, which implies a ROC area of 0.92.

Figure 2a suggests that each examiner’s mental decision

scale might allow for virtually continuous judgments about

competence. However, the examiners have summarized

(per instructions) their judgments about adjudicative com-

petence as five-category ratings. Using a five-category

rating scale implies that their judgments have four non-

trivial dichotomous thresholds (= 5, C4, C3, and C2), and

each dichotomous threshold has associated with it a true

positive and false positive rate. Therefore, four (fpr, tpr)

pairs—eight parameters in all—summarize each exam-

iner’s accuracy. These (fpr, tpr) pairs appear numerically in

Fig. 2a and are also represented by the vertical and hori-

zontal dashed lines.

If the Omniscient Being provided the truth about each

defendant’s status, one could simply calculate the four (fpr,

tpr) pairs for Rater 1 or Rater 2 very easily, without ref-

erence to the other rater’s judgments. In reality, however,

we have no such truth. If one tried to calculate eight

parameters for a single examiner from just five ratings plus

a ninth parameter representing the fraction of D defendants

in the whole 150-subject group, a huge number of possible

(fpr, tpr) pairs would be possible. One would have no way

to distinguish the ‘‘best’’ combination of parameters

because there are more degrees of freedom (nine) than

rating categories (five).

If, however, one looks at both raters’ judgments simul-

taneously, one sees that their joint ratings produce 25

categories. The total number of accuracy parameters—four

(fpr, tpr) pairs for each rater—equals 16. If one adds

additional parameters for correlations in the D and non-D

subjects and for the fraction of D defendants in the entire

group, one obtains 19 parameters in all, implying 19

degrees of freedom. This number is smaller than the 25

joint categories formed by combining both sets of ratings.

Fig. 2 Two hypothetical raters’ continuous-scale judgments about adjudicative competence, with cut-offs (dashed lines) that result when raters group their judgments in five categories. Open circles Actually competent defendants; filled squares actually incompetent defendants. a Uncorrelated judgments, effect size = 2. b Uncorrelated judgments, effect size = 1.2. c Correlated judgments (r = 0.7), effect size = 2

Law Hum Behav (2010) 34:402–417 407

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

In theory, a mathematical search algorithm could ‘‘find’’ a

single set of values for 19 parameters with the maximum

likelihood of producing the 25 empirical category counts—

or alternatively, using a Bayesian procedure, a distribution

of values for the 19 parameters that was most plausible,

given the 25-category data.

To understand how the maximum likelihood estimation

(MLE) algorithm works, imagine asking a two-dimensional

creature to locate the highest elevation or ‘‘peak’’ of the

highest ‘‘mountain’’ on an irregular three-dimensional

surface. The creature cannot get outside the surface to look

and simply ‘‘see’’ where the peak is, but before moving to a

new location on the surface, the creature can take a ten-

tative step of any size in any direction and get information

about which step would place it at the highest elevation

relative to its current location. By a series of successive

elevation-maximizing steps, the creature would eventually

find a point such that a small step in any direction would

place it at an elevation that was lower than the present

location. This location would be the highest point on the

surface, unless the creature had unintentionally found a

local maximum (the ‘‘top’’ of a ‘‘hill,’’ but not the highest

mountain’s ‘‘peak’’).

For the situation depicted in Fig. 2a, the actual mathe-

matical search task involves looking for a peak on a

probability surface that lies above a 19-dimensional plane.

Yet looking at Fig. 2a, we sense that ‘‘finding’’ this prob-

ability peak might be relatively ‘‘easy’’ because the D and

non-D populations are relatively separated. A relatively

‘‘obvious’’ transition zone exists, which suggests that the

algorithm would find it relatively easy to ‘‘locate’’ the

transition zone, which it does by finding the set of (fpr, tpr)

pairs that best represents that zone. The algorithm would

not get ‘‘confused’’ by local maxima and would proceed

relatively directly to the set of values that represent the

locus of the probability peak.

If a transition zone is less obvious, however, an opti-

mization algorithm might have more difficulty locating it—

that is, finding the (fpr, tpr) pairs that best represent the

zone. Figure 2b and c depict two situations reasons why a

transition zone might be ambiguous. In Fig. 2b, the dis-

criminative ability of Raters 1 and 2 is equivalent to an

effect size of 1.2, or a ROC area of 0.80. In Fig. 2c, both

raters have the same discriminatory power as in Fig. 2a

(i.e., effect size = 2, ROC area = 0.92), but their judg-

ments are highly correlated (r = 0.7). In both cases,

distributions of the D and non-D subjects overlap much

more, the transition zone is less obvious, and an algorithm

might have more trouble locating the optimal (probability

maximizing or most plausible) values for the accuracy

parameters.

Our Bayesian approach to locating accuracy parameters

used Gibbs sampling implemented by WinBUGS (Lunn,

Thomas, Best, & Spiegelhalter, 2000) to find a posterior

distribution from which we could make inferences about

the parameters’ values. To understand the process, picture

a drunk individual who, though able to walk, is entirely

unable to walk in a continuously straight line. Instead, the

drunkard takes a random number of steps ahead or back,

then stops to rest; then he walks left or right for a random

number of steps before resting again. He alternates between

forward/backward and left/right random walking over and

over again. As the drunkard staggers, he is affected by

gravity, such that he is more likely to step downhill than

up, and his steps are likely to be smaller as he moves uphill

and larger as he stumbles downhill. Now imagine setting

the drunkard loose in a large, irregularly shaped ravine

with an uneven bottom. Every other time the drunkard

stops to rest, we record his location in the (x, y) plane.

When we are done, we will have a bivariate dot-plot of the

ravine, where the highest density of dots corresponds to the

lowest point in the ravine. This odd ‘‘topographical map’’ is

a sampling from the two-dimensional posterior distribution

for the plausible though unknown (x, y) coordinates of the

bottom of the ravine.

For the situation depicted in Fig. 2a, the Bayesian

drunkard staggers in 19 dimensions rather than just two,

and we record his locations as a sampling from a 19-

dimensional posterior distribution of the model parameters,

which are analogous to the two-dimensional coordinates

for the bottom of the ravine. In Fig. 2a, the D and non-D

populations are relatively separated and distinguishable.

This is like sending the drunkard out to stagger in a steeply

sloped ravine; the steep slope, coupled with gravity, will

help the drunkard quickly get to the low point. In the case

of the 19-dimensional problem posed in Fig. 2a, the

‘‘drunkard’’—the Gibbs sampling process—will quickly

stagger to a clear posterior distribution for the (fpr, tpr)

pairs and other parameters that locate the transition zone.

As was the case with the MLE algorithm, however, situa-

tions in which the transition zone is less obvious—e.g.,

those shown in Fig. 2b and c—may give the Gibbs sam-

pling algorithm a less clearly sloped surface on which to

stagger toward the (fpr, tpr) pairs that best define the

transition zone.

The preceding informal description used two raters

(rather than five) because doing so allows for easy repre-

sentation in two-dimensional figures. The task in our study

was to use a 780-category matrix to identify 43 parameters.

The Appendix formally describes the statistical model and

methods we used to make inferences about empirical class

membership (i.e., about being competent or incompetent)

based on our five raters’ judgments. We implemented the

model under conditional dependence (CD) and conditional

independence (CI) assumptions. We then evaluated accu-

racy in assessing CST with both MLE and Bayesian

408 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

methods based on the Dusky competence scores that raters’

assigned after examining the 156 reports. We also evalu-

ated raters’ scores on understanding, reasoning, and

appreciation (URA) by treating these values as diagnostic

‘‘tests’’ for incompetence. In addition, we examined the

sum of URA ratings as a diagnostic test under the CI

model. Finally, we transformed URA scores (x0 = 5 - x) so that 0 implied highest likelihood of incompetence, and

evaluated the diagnostic accuracy of the product of these

transformed scores under the CI model. (Note that the

transformation causes a ‘‘5’’ score on any URA rating to

yield a ‘‘0’’ product.)

RESULTS

Evaluee Characteristics

Table 3 summarizes the demographic and diagnostic

characteristics of the 156 defendants whose court docu-

ments served as sources for our sanitized reports. Eighty-

six (55%) of the defendants had primary or co-morbid

substance use disorders, but because the patients had been

confined for substantial periods before evaluation, current

or recent intoxication did not influence their clinical pre-

sentation when they underwent evaluation.

Raters’ Performance

Results of the accuracy analyses appear in Table 4. Here,

AUC equals the probability that a rater, examining one

randomly chosen competent defendant and one randomly

chosen incompetent defendant, would assign a higher score

(i.e., a score more indicative of incompetence) to the

actually incompetent subject. An AUC of 1.0 would imply

perfect sorting, and an AUC of 0.5 would imply no-better-

than-chance discrimination between competent and incom-

petent defendants.

The average accuracy of raters’ assessments of Dusky

competence, calculated with MLE or Bayesian methods

under CD or CI assumptions, was at least 0.967, which

suggests they could correctly distinguish a randomly

selected competent defendant from a randomly selected

incompetent defendant in 29 out of 30 attempts. Raters’

ability to evaluate adjudicative competence from the san-

itized reports thus was comparable to accuracy achieved in

using advanced positron emission imaging methods to

detect Alzheimer’s disease (Small et al., 2006). We had

hypothesized that the sum or product of URA scores might

yield greater accuracy than global judgments of compe-

tence. This did not happen, both because raters’ URA

scores tended to yield lower accuracy than their

assessments of Dusky competence and because raters’

assessment using Dusky criteria left little room for

improvement.

One can compare results from various methods and

models using the Akaike (1974) information criterion,

calculated as AIC = - 2 ln L ? 2p, where p is the number

of model parameters. In calculating the AIC, the superior

fit (measured by the -2 ln L term) expected from adducing

additional parameters is offset by a penalty (the 2p term).

Minimum AIC thus can serve as a basis for choosing, from

among several models with different numbers of parame-

ters, a model that the data best support. In Table 4, AIC

values are substantially smaller for CD models than for CI

models, and substantially smaller for Bayesian estimates

than the MLE estimates. Though overall accuracy (mea-

sured by AUCs) is comparable, the MLE and Bayesian

Table 3 Demographic and diagnostic characteristics of the original 156 defendants

Age (years)

Mean ± SD 37.4 ± 12.6

Range 18.2–84.9

Race

African-American 79

Caucasian 77

Sex

Female 19

Male 137

Original opinion on CST a

Competent 101

Not competent 55

Primary diagnoses b

Schizophrenia 53

Bipolar disorder 23

Depressive disorders 7

Substance use disorders 8

Malingering 12

Delusional disorders 3

Adjustments disorders 4

Schizoaffective disorders 24

Psychotic disorder NOS 14

Mental retardation 2

Neurocognitive disorders 2

Impulse control disorders 2

Pedophilia 2

Comorbid conditions

Substance use disorders 78

Mental retardation 10

a Opinion concerning competence to stand trial provided by original

report’s author b

Hospital’s final DSM-IV diagnosis, based on all available clinical information, that best accounted for defendant’s hospital stay

Law Hum Behav (2010) 34:402–417 409

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

estimates of the prevalence of incompetent individuals

often differed. The Bayesian prevalence estimates for

incompetence also tended to be lower than the 35% of

defendants so identified by the original evaluators.

Though AUC is a useful summary index of accuracy,

individual operating points can also be informative.

Figures 3 and 4 depict MLE and Bayesian results for raters’

performance in detecting Dusky incompetence under the

CD model. Comparing these figures suggests that a different

ROC for Rater 4 explains some of the differences found in

Table 4. Using the MLE-CD results from our study, ratings

of ‘‘5’’ on Dusky competence were associated with an

average sensitivity (in detecting incompetence) of 0.828

and an average specificity of 0.977. Using the Bayesian-CD

Table 4 ROC areas for five raters based on MLE and Bayesian estimates of accuracy parameters

Estimation

method

Criterion, model ROC area (MLE standard error or Bayesian posterior standard deviation) prev logL AIC

Rater 1 Rater 2 Rater 3 Rater 4 Rater 5

MLE Dusky, CI 0.974 (0.016) 0.966 (0.018) 0.941 (0.023) 0.958 (0.020) 1.000 (0.000) 0.333 -816.6 1715.3

Dusky, CD 0.984 (0.012) 0.965 (0.018) 0.955 (0.021) 0.942 (0.023) 0.993 (0.008) 0.339 -739.2 1564.3

Understanding, CI 0.925 (0.025) 0.959 (0.019) 0.945 (0.022) 0.957 (0.019) 0.972 (0.016) 0.375 -816.6 1715.3

Understanding, CD 0.801 (0.038) 0.917 (0.025) 0.783 (0.039) 0.853 (0.033) 0.907 (0.027) 0.405 -727.2 1540.4

Reasoning, CI 0.974 (0.016) 0.954 (0.021) 0.967 (0.018) 0.977 (0.015) 0.997 (0.006) 0.314 -859.0 1799.9

Reasoning, CD 0.978 (0.015) 0.962 (0.019) 0.959 (0.020) 0.985 (0.012) 0.998 (0.004) 0.323 -785.4 1656.8

Appreciation, CI 0.970 (0.017) 0.957 (0.020) 0.946 (0.023) 0.973 (0.016) 0.990 (0.010) 0.319 -863.8 1809.5

Appreciation, CD 0.902 (0.031) 0.947 (0.024) 0.930 (0.027) 0.950 (0.023) 0.997 (0.006) 0.297 -800.5 1687.0

Sum, CI 0.962 (0.018) 0.973 (0.015) 0.958 (0.019) 0.977 (0.014) 0.982 (0.012) 0.378 -1373.2 2988.3

Product, CI 0.956 (0.019) 0.943 (0.021) 0.957 (0.019) 0.964 (0.017) 0.985 (0.011) 0.398 -1311.3 2944.6

WinBUGS Dusky, CI 0.967 (0.015) 0.960 (0.016) 0.937 (0.022) 0.950 (0.019) 0.991 (0.009) 0.334 -737.5 1557.0

Dusky, CD 0.975 (0.017) 0.975 (0.016) 0.963 (0.022) 0.968 (0.017) 0.985 (0.014) 0.276 -601.0 1288.0

Understanding, CI 0.948 (0.023) 0.968 (0.016) 0.932 (0.024) 0.957 (0.016) 0.960 (0.018) 0.330 -734.5 1555.0

Understanding, CD 0.971 (0.019) 0.976 (0.016) 0.940 (0.028) 0.970 (0.019) 0.978 (0.018) 0.267 -594.5 1275.0

Reasoning, CI 0.963 (0.017) 0.949 (0.019) 0.951 (0.019) 0.966 (0.017) 0.989 (0.009) 0.320 -782.0 1646.0

Reasoning, CD 0.973 (0.017) 0.962 (0.021) 0.955 (0.021) 0.978 (0.015) 0.987 (0.011) 0.286 -641.5 1369.0

Appreciation, CI 0.962 (0.016) 0.950 (0.019) 0.936 (0.024) 0.961 (0.019) 0.980 (0.012) 0.328 -786.5 1655.0

Appreciation, CD 0.952 (0.021) 0.960 (0.022) 0.950 (0.025) 0.980 (0.017) 0.980 (0.015) 0.276 -606.5 1389.0

prev sample prevalence, logL loge likelihood, AIC Akaike Information Criterion, MLE maximum likelihood estimation, CI conditional inde- pendence assumption, CD condition dependence assumption

0

0.5

1

10.50 false positive rate

tr u

e p

o s it

iv e r

a te

Rater 1

Rater 2

Rater 3

Rater 4

Rater 5

Fig. 3 ROC graph depicting maximum likelihood estimates of raters’ performance in detecting adjudicative incompetence (applying the

Dusky standard), under conditional dependence assumptions

0

0.5

1

10.50

false positive rate

tr u

e p

o s it

iv e r

a te

Rater 1

Rater 2

Rater 3

Rater 4

Rater 5

Fig. 4 ROC graph depicting Bayesian (WinBUGS) estimates of raters’ performance in detecting adjudicative incompetence (applying

the Dusky standard), under conditional dependence assumptions

410 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

results, the average sensitivity for scores of ‘‘5’’ was 0.761,

and the average specificity was 0.991.

As the high accuracy of raters suggests, the raters’

results were consistent with and highly correlated with

each other. On Dusky competence, for example, correla-

tions for the ten possible pairs of raters ranged between

0.757 and 0.859. Table 5 reports values of ICC(A,1) (i.e.,

single rater, absolute agreement; see McGraw & Wong,

1996) and ICC(C,1) (single rater, consistency) for all four

sets of ratings. These findings offer numerical expressions

of what Figs. 3 and 4 illustrate. In these two figures, we see

that each rater’s ROC bends at a different location in the

ROC square, implying that raters’ operating points are not

identical. Figure 3, for example, shows that the boundaries

between Dusky competence scores of ‘‘4’’ (‘‘probably

incompetent’’), ‘‘3’’ (‘‘uncertain’’), and ‘‘2’’ (‘‘probably

competent’’) for Rater 3 are roughly the same location as

the boundary between scores of ‘‘5’’ (‘‘very likely incom-

petent’’) and 4 (‘‘probably incompetent’’) for Rater 2.

Raters 2 and 3 earned similar MLE estimates of overall

accuracy—their ROC areas are 0.965 and 0.955, respec-

tively (see Table 4)—but Rater 3 gave scores of ‘‘4’’ and

‘‘3’’ to several defendants whom Rater 2 scored ‘‘5.’’

For the MLE results, the estimated dependence param-

eters for the random effects model were r̂0 ¼ 1:53 (for the non-D subgroup) and r̂1 ¼ 1:01 (for the D subgroup); in the Bayesian-estimated random effects model, the non-D

and D dependence parameters were r̂0 ¼ 1:73 and r̂1 ¼ 0:406, respectively. Both pairs of results suggest substantial conditional dependence. Following Albert

(2007) and Qu, Tan, and Kutner (1996), we examined the

‘‘observed minus expected correlations’’ for each of the ten

rater pairs, that is, the differences between the actual

pairwise correlations of ratings minus the correlation

that would be expected based on the model’s parameters.

Figure 5 shows a ‘‘diagnostic plot’’ comparing these dif-

ferences for the MLE (dashed lines) and Bayesian (solid

lines) estimates under conditional independence and con-

ditional dependence (random effects) models. Generally,

the differences for the random effects models (open sym-

bols) are lower (closer to zero) than the differences for the

conditional independence models, which suggests (as do

the AIC values in Table 4) that the random effects models

provide better descriptions of the pairwise correlations than

do the conditional independence models.

DISCUSSION

Over the last two decades, ROC analysis has become a

popular technique for quantifying the accuracy of forensic

mental health assessments. In such applications, humans’

judgments based on available data—e.g., about whether

violence occurred (Steadman et al., 1998; Monahan et al.,

2006), about whether malingering has occurred (Miller,

2005), or about whether an evaluee is competent (Kim et al.,

2007)—provide the truth criteria or ‘‘gold standards’’ against

which investigators have judged accuracy. But given the

central role that psychiatrists and psychologists play in psy-

cholegal determinations, the accuracy of the professionals

themselves is a matter of major legal and social significance.

To our knowledge, ours is the first study to use latent

class models and the ROC analytic concepts discussed by

Mossman (2008) to quantify the accuracy of forensic

experts. Our study shows that raters can assess CST on a

graded scale (rather than simply providing binary opinions,

as mental health experts customarily do) and that such

ratings can lead to estimated ROC parameters of detection

accuracy despite there being no ‘‘gold standard’’ for CST.

Insofar as our chief aims were to find out whether clini-

cians could provide graded judgments and whether our

statistical approach was feasible, our study met its goals.

Our findings also suggest that mental health experts’

intrinsic ability to discriminate between competent and

Table 5 Intraclass correlation coefficients for five raters’ judgments concerning understanding, reasoning, appreciation, and Dusky competence

Rating ICC(A,1) ICC(C,1)

Dusky competence 0.7946 0.9508

Understanding 0.8015 0.9528

Reasoning 0.7931 0.9504

Appreciation 0.7664 0.9425

ICC(A,1) single rater, absolute agreement ICC, ICC(C,1) single rater, consistency

0

0.02

0.04

0.06

0.08

0.1

0.12

0.14

0.16

1,2 1,3 1,4 1,5 2,3 2,4 2,5 3,4 3,5 4,5

Rater pairs

o b

s e rv

e d

- e

x p

e c te

d c

o rr

e la

ti o

n

MLE Conditional Independence

MLE random effects

WinBUGS Conditional Independence

WinBUGS random effects

Fig. 5 Observed minus expected correlations under conditional independence and conditional dependence assumptions

Law Hum Behav (2010) 34:402–417 411

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

incompetent defendants is high (though not perfect).

Readers should recognize, however, that several features of

our study limit the generalizability of this conclusion.

First, our raters based their judgments on written reports

rather than on personal examinations—a typical require-

ment for a forensic assessment (American Academy of

Psychiatry and the Law, 2005). We have no reason to

believe that we introduced systematic bias by excluding

poorly written reports from our sample, but we cannot be

sure that the exclusions altered our ultimate sample so as to

make the diagnostic task easier (or harder) for our raters.

Using ‘‘sanitized’’ reports was time-efficient for raters

and protected defendants’ confidentiality. However, written

reports condense information that examiners obtain during

personal interviews (which may have the effect of reducing

rater accuracy). The original reports’ text may have

selectively included data consistent with original author-

examiners’ CST opinions, while excluding data that con-

flicted with their opinions (which could have the statistical

effect of inflating our raters’ apparent accuracy). Also,

having raters base judgments on data provided by someone

else does not permit evaluation of the raters’ ability to

gather data and recognize its importance, which is an

essential feature of evaluative accuracy.

Of course, one could never replicate typical circum-

stances of adjudicative competence evaluations for a study

such as ours. Attempting to do this would require having

multiple examiners conduct in-person, legally superfluous,

closely spaced evaluations of dozens of defendants. Even if

doing this were practicable, multiple evaluations would

themselves affect the evaluees and distort the resulting

findings. The process would also confront serious ethical

questions concerning evaluees who would not be compe-

tent to consent to the research, but whose participation

would be necessary to reach a meaningful judgment about

the accuracy of CST determinations.

However, because our study has shown that our statis-

tical approaches are workable, it may be reasonable to

conduct future studies that better replicate the data obtained

in actual CST evaluations. For example, a future study

might evaluate raters’ accuracy when they base their CST

judgments on documents prepared using highly systema-

tized formats for data collection and written presentation.

Also, the likely prospect of generating findings with high

social and legal significance might make it easier to justify

a study like ours in which multiple raters based their

judgments on actual, videotaped CST assessments (though

obtaining consent from defendants with severe mental

impairments would remain problematic).

A second limit on generalizability stems from the ori-

ginal reports’ context—evaluations following court-

ordered hospitalizations for CST assessment or restoration.

Such hospitalizations usually give evaluators substantial

time to assemble background information and generate

copious observational data that help to clarify psychiatric

diagnosis. Acute effects of intoxicants abate during hos-

pitalization, and for those patient-defendants who accept

treatment, examiners can take into account how medication

and psychotherapy have affected evaluees’ functioning and

clinical presentation. All these circumstances probably

made our raters’ tasks easier—and their assessments more

accurate—than would be the case when an examiner

encounters a defendant during an initial court-ordered

evaluation of CST.

A third limitation relates to how our findings apply to

the binary, competent-or-incompetent opinions often

required by courts or statutes. Our study’s five-category

scorings helped us focus specifically on raters’ intrinsic

ability to assess adjudicative competence. When examiners

provide yes-or-no opinions, however, those opinions

incorporate examiner biases, including examiners’ feelings

about the relative undesirability of false-positive and false-

negative errors. If these value judgments translate into

disagreement about the ultimate yes-or-no conclusion, they

will lower apparent accuracy.

To see why this is the case, consider what Fig. 4 sug-

gests about defendants evaluated by Raters 1, 2, and 4. In

Table 4, the Bayesian CD-model ROC areas for these

raters on Dusky competence are very similar (implying

near-identical diagnostic accuracy), and for all three raters,

Fig. 4 locates an operating point at roughly (fpr,

tpr) = (0.1, 0.95). For Rater 1, however, this area marks

the boundary between ratings of ‘‘3’’ (‘‘uncertain’’) and

‘‘2’’ (‘‘probably competent’’), while for Raters 2 and 4, the

boundary is between 4 (‘‘probably incompetent’’) and 3

(‘‘uncertain’’). This implies that Rater 1 will score a few

defendants ‘‘uncertain’’ whom Raters 2 and 4 think are

probably incompetent. If Rater 1 believes ‘‘when uncertain,

presume competent,’’ then he will disagree with Raters 2

and 4 if they believe that probably incompetence should

generate a yes-no opinion of incompetence. In such a case,

at least one rater would be ‘‘wrong’’ about the defendant,

and his apparent accuracy based on just the binary rating

would fall. To understand another source of inaccuracy

caused by binary ratings, imagine that all five of our raters

thought a particular report described a ‘‘probably compe-

tent’’ defendant, but—because they were required to

provide binary opinions—they disagreed in their value

judgment about whether that marginally competent

defendant should face trial. Here, at least one rater would

have been ‘‘wrong’’ about the defendant, and his apparent

accuracy would have fallen.

What this suggests is that experts may well be better

discriminators of competence than one would conclude

from CST evaluations undertaken in ‘‘real life’’ criminal

cases, because some disagreements may arise from

412 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

different value judgments rather than from different opin-

ions about clinical findings relevant to CST. Indeed, recent

studies by Murrie and colleagues (Boccaccini et al., 2008;

Murrie et al., 2008, 2009) suggest that individual exam-

iners’ biases explain a large amount of the variance in their

ultimate opinions about CST or the scores they report when

using actuarial risk assessment instruments. Findings from

these studies, which use ICC(A,1) (the absolute, single-

rater intraclass correlation coefficient) as their chief ana-

lytical statistic, complement this study’s ROC-based

findings. The difference is that these previous studies are

agnostic about whether forensic evaluees have any true

diagnostic status—that is, these studies have examined

degrees of absolute diagnostic agreement between evalu-

ators and possible sources of disagreement, but have not

asked how well forensic examiners perform or how often

they might get the answer ‘‘right.’’ In contrast, our study

assumes—as do mental health experts and courts—that

while defendants display varying degrees of the abilities

that underlie adjudicative competence, they ultimately are

either competent to stand trial or not. Our study then

attempts to quantify and characterize individual evaluators’

ability to distinguish between these two (presumptively

mutually exclusive) forensic options, despite having no

gold standard for defendants’ true status. Our study

accomplishes this by focusing on and quantifying evaluator

accuracy using appropriate statistics (i.e., ROC indices),

rather than on measuring inter-rater agreement using sta-

tistics (e.g., ICCs) that are appropriate to that task.

Our findings may reflect statistical limitations arising

from sampling error; we report here on the performance of

just five psychiatrists who read a particular set of redacted

materials based on reports generated at just one institution.

Though our findings reflect a reasonable statistical model

for our data, we recognize that other plausible models exist

and might generate somewhat different outcomes. Also, as

Uebersax (1988) notes, latent class modeling provides

upper bounds for accuracy under certain conditions. LCM

chooses underlying classes that minimize error rates

defined within the model, but these error-minimizing,

empirically generated latent classes can differ from the true

classes when probabilities of the empirical classes depend

on covariates. Knowing whether this actually has occurred

is difficult to ascertain (Spencer, 2008), but it is a limitation

that we must acknowledge.

Our findings may also reflect limitations due to mis-

taking reliability for validity and to related problems of

undefined ontology. As the ‘‘Background’’ section

explains, we assumed that assessing CST involves evalu-

ating abilities needed to perform a task; we also assumed

that raters’ Dusky-guided notions about CST reflected valid

conceptions of those abilities. A problem with our statis-

tical methods is that if raters all used a very reliable but

irrelevant method for assessing competence (e.g., the

length of a subject’s last name), they might appear very

accurate despite their really having no-better-than-chance

accuracy. Our response is, simply, that our raters applied

the same definitions and ideas about CST that mental

health experts and courts regularly use in legal decision

making—that is, the same definitions and ideas that experts

and courts regularly accept as being valid.

We also did not explore whether CST is a dimensional

construct or whether it admits of a valid dichotomy, which

are typical uses of latent class methods when diagnostic

validity is questionable. We think that forensic psychiatry

and psychology might well benefit from explorations of

empirically based, natural taxa that are relevant to CST,

but our study leaves such explorations to others. We note,

however, that even if the results of such efforts were

known, natural taxa might not correctly track or mean-

ingfully distinguish between individuals who are and are

not competent to stand trial; it is just as reasonable to

suppose that some taxa would overlap the competent-

incompetent boundary. Whatever a CST taxonomy might

tell us, we think it is reasonable to gather data that lets us

characterize raters’ accuracy about a distinction (between

competent and incompetent) that is imperfectly understood,

but that everyone usually takes to be genuine.

A final limitation arises from our raters’ neutrality.

Numerous studies show that framing and anchoring of

information affect decision-makers’ judgments, even when

the decision-makers receive instruction about such effects

or have incentives to be accurate (Cain & Detsky, 2008).

Despite attempts by forensic consultants to remain objec-

tive, a variety of conscious or unconscious factors—

identification with or desire to please the retaining party, or

sympathy for or antipathy toward an evaluee, or being

prosecution- or defense-oriented—influence forensic opin-

ions (Boccaccini et al., 2008; Gutheil, 2004; Murrie et al.,

2008, 2009). In the United States, most mental health

opinions about CST are accepted by courts without dispute

(Zapf, Hubbard, Cooper, Wheeles, & Ronan, 2004). When

second opinions are sought, however, it is often because the

defense or the prosecution has disagreed with the first

opinion’s conclusion. This non-random referral pattern may

inflate apparent expert disagreement above what one would

find if experts and cases were chosen at random. Such

selection patterns may encourage or induce experts to dis-

agree with each other (Murrie et al., 2009), and ironically,

acknowledging biasing potential may actually exacerbate

the problem (Cain, Loewenstein, & Moore, 2005). In our

study, raters were ignorant of their collaborators’ opinions,

and because they could offer graded ratings about CST, they

did not have to reach yes-or-no conclusions about ambig-

uous cases. Arguably, our raters performed their evaluation

tasks in circumstances more conducive to being accurate

Law Hum Behav (2010) 34:402–417 413

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

than the circumstances in which opposing forensic evalua-

tors usually find themselves.

Though our study’s primary aim was to test whether

forensic opinions could be conceptualized so as to allow

accuracy characterization using latent class methods, we

believe our endeavor’s success and our numerical findings

have some immediate practical implications.

Having shown that forensic examiners differ greatly in how

often they think criminal defendants are incompetent, Murrie

and colleagues suggest that future research might aim to

identify the ‘‘thresholds’’ at which clinicians consider

a defendant to be incompetent to stand trial (IST).

Although the legal determination regarding compe-

tence is dichotomous (i.e., competent or not

competent), it seems reasonable to think of the

capacities underlying trial competence as dimen-

sional; in other words, some defendants are ‘‘more

competent’’ and others ‘‘less competent.’’ If these

capacities are dimensional, we might not be surprised

to find that clinicians sometimes draw the distinction

between competent and incompetent at different

points along the continuum. We also might not be

surprised to find some disagreement among clinicians

regarding cases that fall toward the midpoint of this

continuum. (Murrie et al., 2008, p. 190)

As expressed by Murrie and colleagues, the continuum

(or continua) along which clinicians’ thresholds might lie is

undefined. Our study, however, used the raters’ judgments

about CST itself as the continuum. An implication of our

approach is the finding, illustrated in Figs. 3 and 4, that

raters’ thresholds for deeming defendants ‘‘probably

competent,’’ of ‘‘uncertain’’ competence, or ‘‘probably

incompetent’’ can vary enough to generate practical dis-

agreement, even though raters are very accurate. Our study

also suggests that disagreement in a dichotomous judgment

about CST may not necessarily stem from examiners’ dif-

ferences in judgments about defendants’ abilities. Even

when examiners agree about what specific defendants can

and cannot do relative to standing trial, they may disagree

about whether marginally capable defendants are really fit

to face serious criminal charges.

A second implication is that forensic practitioners

should be modest. Our study looked at raters making

‘‘fairly simple determinations’’ (Murrie et al., 2008, p. 180)

about adjudicative competence in a non-adversarial con-

text. Our raters also worked under conditions that gave

them optimal access to background data, and the data raters

used included information about treatment response fol-

lowing extensive inpatient treatment episodes. Though our

raters were very accurate, they still had varying levels of

clarity about their judgments and sometimes disagreed with

each other. This finding should tell us forensic practitioners

who work and reach conclusions in less ideal circum-

stances that we can be quite accurate yet fallible, and that

our colleagues will disagree with us for reasons that do not

undermine the legitimacy of their or our conclusions.

A final implication is for fellow investigators. Our study

demonstrates a practicable method for using ‘‘ROC anal-

ysis without truth’’ (Henkelman et al., 1990) to estimate the

accuracy of forensic assessments. The ideas and techniques

we describe may be applied to a host of other psycholegal

determinations where no diagnostic gold standard exists,

but where quantification of accuracy would have scientific

and evidentiary value. Many readers of this article can

create or already have available data that would be ame-

nable to the methods described in this article. We hope

other investigators will view our work as something to

improve upon and as inspiration for asking—and answer-

ing—many other questions about forensic assessments and

the quality of mental health expertise.

APPENDIX

Adopting the notation used by Albert (2007), suppose I

subjects (i = 1, 2, …, I) undergo assessment by J raters (j = 1, 2, …, J), who assign ordinal ratings k = 1, 2, …, K to each subject. Without loss of generality, let rating k = 1

indicate lowest confidence and k = K indicate highest

confidence that a subject has the condition or disorder D of

interest (here, incompetence to stand trial). Let Yi = (Yi1,

Yi2, …, YiJ)0 be a vector representing ratings made by the J raters for the ith subject. Because J raters could each assign

one of K ratings to each subject, each Yi has J 9 K possible

combinations of elements. The joint distribution of Yi,

expressed as P(Yi), the probability of Yi, is

PðYiÞ¼ PðYijdi ¼ 1ÞPðdi ¼ 1Þþ PðYijdi ¼ 0ÞPðdi ¼ 0Þ; ð1Þ

where di = 1 means the ith subject has condition D, di = 0

means the ith subject does not have D, P(di = 1) is

the probability or prevalence of D, and P(di = 0) =

1 - P(di = 1).

We would like to model P(Yi|di) so as to include pos-

sible conditional dependence (CD) of ratings, i.e.,

similarity in raters’ responses attributable to specific

characteristics of subjects besides their membership in the

D or non-D subgroups that affect how easy or hard their

particular cases are. Following Albert, we utilize a probit

link function for the parameterization,

U�1 P Yij �kjdi; bdi;i � �� �

¼ Cdi;k;j þ bdi;i ð2Þ

where U is the cumulative standard normal distribution function and U-1 is its inverse, Cdi;k;j are monotonically

414 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

increasing cut-offs for the jth rater, and bdi;i is a random

effect attributable to each subject that characterizes

conditional dependence in multiple ratings of that

subject.

Notice that bdi;i depends on the latent class of each

subject—i.e., whether the individual does or does not have

D. Following Albert (2007) and Qu et al. (1996), we used

the random effect model bdi;i ¼ rdi bi, where bi has a standard normal distribution. Equation 2 thus says that cut-

off points demarcating each rater’s classification thresholds

reflect the presence (di = 1) or absence (di = 0) of D.

However, the probability that the jth rater will assign rating

k to the ith subject reflects the ith subject’s state (di = 0 or

di = 1), locations of the rater’s particular cut-offs, and

peculiarities of the ith subject (which act in common across

all raters). We characterize the random effects of the D and

non-D populations separately because their cut-offs are not

linked (as they would be under the ‘‘binormal’’ ROC

model; see Somoza & Mossman, 1991).

In our data set, I = 156, J = 5, and K = 5. Thus, in our

CD model, I 9 J ratings (J raters evaluating I subjects)

arise from 2(K - 1)J ? 3 = 43 parameters: K - 1 cut-

offs for the D subgroup and K - 1 cut-offs for the non-D

subgroup for each rater, plus a random effect attributable to

each subgroup, plus the prevalence P(di = 1) in the rating

set. We sought values for the 43 parameters that would, in

combination, be most likely to have generated the 5 9 156

rating matrix. We could then construct individual raters’

ROC graphs using (fpr, tpr) coordinates computed as

follows:

fprj;k ¼ 1 � U C0;kffiffiffiffiffiffiffiffiffiffiffiffiffi 1 þ r20

p

!

; tprj;k ¼ 1 � U C1;kffiffiffiffiffiffiffiffiffiffiffiffiffi 1 þ r21

p

!

ð3Þ

Under a conditional independence (CI) assumption (equiv-

alent to setting bdi;i ¼ 0), ROC graphs for the five raters could be constructed from estimates of 2(K - 1)J ? 1 =

41 parameters.

We estimated the model’s accuracy parameters in two

ways. The first approach, standard maximum likelihood

estimation (MLE), used GAUSS 3.6 code kindly fur-

nished by Albert and modified for our data. (The modified

code, which calls the GAUSS 4.0 maxlik library’s quasi-

Newton BFGS optimization algorithm, is available from

the first author.) The natural logarithm of the likelihood

function, ln L ¼ PI

i¼1 ln Li, is (slightly modifying Albert’s notation)

ln L ¼ XK

i1¼1

XK

i2¼1 . . . XK

iJ¼1 IfYi¼ði1;i2;...;iJÞg

� ln P Yi ¼ i1; i2; . . .; iJð Þð Þf g; ð4Þ

with P(Yi) given by Eq. 1. Knowing (fpr, tpr) coordinates

for each rater permitted computation of ‘‘trapezoidal’’

AUCs as overall measures of rater accuracy, with standard

errors computed using the method of Hanley and McNeil

(1982).

In contrast to MLE, which provides point estimates of

the parameter values most likely to have generated the

observed data, Bayesian estimation summarizes knowledge

of unknown parameters using ‘‘posterior’’ distributions

representing the probability that a parameter has a particular

value, given the observed data. According to Bayes’ Rule,

the posterior probability of a parameter’s value is propor-

tional to the likelihood of observing the data given that

parameter value, multiplied by a ‘‘prior’’ probability of the

parameter’s value. The likelihood function is dictated by

statistical model choice, and is the same construct as in

MLE. When a prior is ‘‘non-informative’’ (e.g., P(h) = c for all h [ [a,b], where [a,b] is an arbitrarily large bounded interval, c is a constant, and h is a parameter), Bayesian and MLE methods yield similar inferences (Carlin & Louis,

2000). However, in Bayesian estimation, inference is con-

ducted directly on the unknown parameters (or functions

thereof, such as AUC), while in MLE, inference is con-

ducted on the data. Hence, only Bayesian estimation allows

direct probability statements such as ‘‘the probability that

the AUC for rater j is between .955 and .973 is 95%.’’

Markov chain Monte Carlo (MCMC) methods (Gelfand

& Smith, 1990; Geman & Geman, 1984; Metropolis,

Rosenbluth, Rosenbluth, Teller, & Teller, 1953) are used to

make inferences on posterior distributions for which lack

of analytic methods would make Bayes’ Rule intractable.

Under mild regularity conditions, a Markov chain con-

verges to a unique invariant or ‘‘target’’ distribution. To use

MCMC methods for Bayesian analysis, one constructs the

transition kernel so that the target distribution of the

resulting Markov chain will be the joint posterior distri-

bution of interest. After discarding input from initial ‘‘burn-

in’’ iterations, one can use the remaining draws to make

inferences about model parameters. WinBUGS is a free

software package that allows specification of a Bayesian

model, determines the transition kernel for the Markov

chain, and produces draws from the joint posterior distri-

bution of unknown parameters (Lunn et al., 2000).

For our Bayesian analyses, we used minimally infor-

mative priors and the same statistical models as in our

MLE approach. WinBUGS 1.4.3 ran five parallel MCMC

chains; the Brooks–Gelman–Rubin diagnostic (Brooks &

Gelman, 1998) indicated convergence after 2000–5000

iterations. We ran each chain for 15,000 iterations and

treated each chain’s first 10,000 iterations as ‘‘burn-in’’

values to be discarded, leaving 5 9 5000 = 25,000 draws

for inference. Because our WinBUGS code (available from

Law Hum Behav (2010) 34:402–417 415

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

the first author upon request) calculated (fpr, tpr) coordi-

nates and trapezoidal AUCs directly from the MCMC

parameter draws, we obtained samples of and made

inferences about our accuracy statistics directly from the

posterior distributions.

REFERENCES

Akaike, H. (1974). A new look at the statistical model identification.

IEEE Transactions on Automatic Control, 19, 716–723. Akinkunmi, A. A. (2002). The MacArthur Competence Assessment

Tool Fitness to Plead: A preliminary evaluation of a research

instrument for assessing fitness to plead in England and Wales.

Journal of the American Academy of Psychiatry and the Law, 30, 476–482.

Albert, P. S. (2007). Random effects modeling approaches for

estimating ROC curves from repeated ordinal tests without a

gold standard. Biometrics, 63, 593–602. American Academy of Psychiatry and the Law. (May 2005). Ethics

guidelines for the practice of forensic psychiatry. http://www. aapl.org/ethics.htm. Accessed 19 Sept 2008.

Bennett, G. (1985). A guided tour through selected ABA standards

relating to incompetence to stand trial: Incompetence to stand

trial. George Washington Law Review, 53, 375–413. Berg, W. A., Blume, J. D., Cormack, J. B., Mendelson, E. B., Lehrer,

D., Böhm-Vélez, M., et al. (2008). Combined screening with

ultrasound and mammography vs mammography alone in

women at elevated risk of breast cancer. Journal of the American Medical Association, 299, 2151–2163.

Boccaccini, M. T., Turner, D., & Murrie, D. C. (2008). Do some

evaluators report consistently higher or lower psychopathy

scores than others? Findings from a statewide sample of sexually

violent predator evaluations. Psychology, Public Policy, and Law, 14, 262–283.

Bonnie, R. J. (1990). The competence of criminal defendants with

mental retardation to participate in their own defense. Journal of Criminal Law and Criminology, 81, 419–446.

Brooks, S. P., & Gelman, A. (1998). Alternative methods for

monitoring convergence of iterative simulations. Journal of Computational and Graphical Statistics, 7, 434–455.

Buchanan, A. (2006). Competency to stand trial and the seriousness

of the charge. Journal of the American Academy of Psychiatry and the Law, 34, 458–465.

Cain, D. M., & Detsky, A. S. (2008). Everyone’s a little bit biased

(even physicians). Journal of the American Medical Association, 299, 2893–2895.

Cain, D. M., Loewenstein, G., & Moore, D. A. (2005). The dirt on

coming clean: Perverse effects of disclosing conflicts of interest.

Journal of Legal Studies, 34, 1–25. Carlin, B. P., & Louis, T. A. (2000). Bayes and empirical Bayes

methods for data analysis (2nd ed.). London: Chapman & Hall. Choi, Y. K., Johnson, W. O., Collins, M. T., & Gardner, I. A. (2006).

Bayesian inferences for receiver operating characteristic curves

in the absence of a gold standard. Journal of Agricultural, Biological, and Environmental Statistics, 11, 210–229.

Committee on the Revision of the Specialty Guidelines for Forensic

Psychology. (11 January 2006). Specialty guidelines for forensic psychology, second official draft. http://www.ap-ls.org/links/. Accessed 19 Sept 2008.

Cooper, V. G., & Zapf, P. A. (2003). Predictor variables in

competency to stand trial decisions. Law and Human Behavior, 27, 423–436.

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in

psychological tests. Psychological Bulletin, 52, 281–302. Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579 (1993). Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus

actuarial judgment. Science, 243, 1668–1674. Douglas, K. S., Ogloff, J. R., Nicholls, T. L., & Grant, I. (1999).

Assessing risk for violence among psychiatric patients: The

HCR-20 violence risk assessment scheme and the Psychopathy

Checklist: Screening Version. Journal of Consulting and Clinical Psychology, 67, 917–930.

Dusky v. United States, 362 U.S. 402 (1960). Faigman, D. L., Saks, M. J., Sanders, J., & Cheng, E. K. (2008).

Modern scientific evidence: Standards, statistics, and research methods, student ed.. Eagan, MN: Thomson West.

Faraone, S. V., & Tsuang, M. T. (1994). Measuring diagnostic

accuracy in the absence of a ‘‘gold standard.’’ American Journal of Psychiatry, 151, 650–657.

Gardner, W., Lidz, C. W., Mulvey, E. P., & Shaw, E. C. (1996).

Clinical versus actuarial predictions of violence of patients with

mental illnesses. Journal of Consulting and Clinical Psychology, 64, 602–609.

Gelfand, A. E., & Smith, A. F. M. (1990). Sampling-based

approaches to calculating marginal densities. Journal of the American Statistical Association, 85, 389–409.

Geman, S., & Geman, D. (1984). Stochastic relaxation, Gibbs

distributions, and the Bayesian restoration of images. IEEE Transactions on Pattern Analysis and Machine Intelligence, 6, 721–741.

Golding, S. L., Roesch, R., & Schreiber, J. (1984). Assessment and

conceptualization of competency to stand trial: Preliminary data

on the Interdisciplinary Fitness Interview. Law and Human Behavior, 8, 321–334.

Grisso, T. (2003). Legally relevant assessments for legal competen-

cies. In T. Grisso (Ed.), Evaluating competencies: Forensic assessments, instruments (2nd ed., pp. 21–40). New York: Kluwer Academic/Plenum Publishers.

Gutheil, T. G. (2004). The expert witness. In R. I. Simon & L. H.

Gold (Eds.), The American Psychiatric Publishing textbook of forensic psychiatry (pp. 75–89). Arlington, VA: American Psychiatric Publishing.

Hagen, M. A. (1997). Whores of the court: The fraud of psychiatric testimony and the rape of American justice. New York: ReganBooks.

Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the

area under the receiver operating characteristic (ROC) curve.

Radiology, 143, 29–36. Harris, G. T., Rice, M. E., & Cormier, C. A. (2002). Prospective

replication of the Violence Risk Appraisal Guide in predicting

violent recidivism among forensic patients. Law and Human Behavior, 26, 377–394.

Henkelman, R. M., Kay, I., & Bronskill, M. J. (1990). Receiver

operator characteristic (ROC) analysis without truth. Medical Decision Making, 10, 24–29.

Jackson v. Indiana, 406 U.S. 715 (1972). Jacobs, M. S., Ryba, N. L., & Zapf, P. A. (2008). Competence-related

abilities and psychiatric symptoms: An analysis of the under-

lying structure and correlates of the MacCAT-CA and the BPRS.

Law and Human Behavior, 32, 64–77. Kim, S. Y. H., Appelbaum, P. S., Swan, J., Stroup, T. S., McEvoy, J. P.,

Goff, D. C., et al. (2007). Determining when impairment

constitutes incapacity for informed consent in schizophrenia

research. British Journal of Psychiatry, 191, 38–43. Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999). Lehman, C. D., Gatsonis, C., Kuhl, C. K., Hendrick, R. E., Pisano, E. D.,

Hanna, L., et al. (2007). MRI evaluation of the contralateral breast

416 Law Hum Behav (2010) 34:402–417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

in women with recently diagnosed breast cancer. New England Journal of Medicine, 356, 1295–1303.

Lunn, D. J., Thomas, A., Best, N., & Spiegelhalter, D. (2000).

WinBUGS—A Bayesian modeling framework: Concepts, struc-

ture, and extensibility. Statistics and Computing, 10, 325–337. McGraw, K. O., & Wong, S. P. (1996). Forming inferences about

some intraclass correlations. Psychological Methods, 1, 30–46. Melton, G. B., Petrila, J., Poythress, N., Slobogin, C., Lyons, P., &

Otto, R. K. (2007). Psychological evaluations for the courts: A handbook for mental health professionals and lawyers (3rd ed.). New York: Guilford.

Metropolis, N., Rosenbluth, A., Rosenbluth, M., Teller, A., & Teller,

E. (1953). Equations of state calculations by fast computing

machines. Journal of Chemical Physics, 21, 1087–1091. Miller, H. A. (2005). The Miller-Forensic Assessment of Symptoms

Test (M-Fast): Test generalizability and utility across race

literacy, and clinical opinion. Criminal Justice and Behavior, 32, 591–611.

Monahan, J., Steadman, H. J., Appelbaum, P. S., Grisso, T., Mulvey,

E. P., Roth, L. H., et al. (2006). The classification of violence

risk. Behavioral Science and the Law, 24, 721–730. Mossman, D. (1999). ‘‘Hired guns’’, ‘‘whores’’, and ‘‘prostitutes’’:

Case law references to clinicians of ill repute. Journal of the American Academy of Psychiatry and the Law, 27, 414–425.

Mossman, D. (2005). Is prosecution ‘‘medically appropriate’’? New England Journal on Criminal and Civil Confinement, 31, 15–80.

Mossman, D. (2007). Predicting restorability of incompetent criminal

defendants. Journal of the American Academy of Psychiatry and the Law, 35, 34–43.

Mossman, D. (2008). Conceptualizing and characterizing accuracy in

assessments of competence to stand trial. Journal of the American Academy of Psychiatry and the Law, 36, 340–351.

Mossman, D., Noffsinger, S. G., Ash, P., Frierson, R. L., Gerbasi, J.,

Hackett, M., et al. (2007). AAPL practice guideline for the

forensic psychiatric evaluation of competence to stand trial.

Journal of the American Academy of Psychiatry and the Law, 35(Suppl 4), S3–S72.

Mossman, D., & Somoza, E. (1991). ROC curves, test accuracy, and

the description of diagnostic tests. Journal of Neuropsychiatry and Clinical Neurosciences, 3, 330–333.

Murrie, D. C., Boccaccini, M. T., Turner, D., Meeks, M., Woods, C.,

& Tussey, C. (2009). Rater (dis)agreement on risk assessment

measures in sexually violent predator proceedings: Evidence of

adversarial allegiance in forensic evaluation? Psychology, Public Policy, and Law, 15, 19–53.

Murrie, D. C., Boccaccini, M., Zapf, P. A., Warren, J. I., & Henderson,

C. E. (2008). Clinician variation in findings of competence to

stand trial. Psychology, Public Policy, and Law, 14, 177–193. Obuchowski, N. A. (2003). Receiver operating characteristic curves

and their use in radiology. Radiology, 229, 3–8. Parry, J., & Drogin, E. Y. (2007). Mental disability law, evidence and

testimony: A comprehensive reference manual for lawyers, judges, and mental disability professionals. Washington, DC: American Bar Association.

Pate v. Robinson, 383 U.S. 375 (1966).

Poythress, N., Monahan, J., Bonnie, R., Otto, R. K., & Hoge, S. K.

(2002). Adjudicative competence: The MacArthur studies. New York: Kluwer/Plenum.

Qu, Y., Tan, M., & Kutner, M. H. (1996). Random effects models in

latent class analysis for evaluating accuracy of diagnostic tests.

Biometrics, 53, 797–810. Rice, M. E., & Harris, G. T. (1995). Violent recidivism: Assessing

predictive validity. Journal of Consulting and Clinical Psychol- ogy, 63, 737–748.

Rosenfeld, B., & Ritchie, K. (1998). Competence to stand trial:

Clinician reliability and the role of offense severity. Journal of Forensic Sciences, 43, 151–159.

Skeem, J., Golding, S., Cohn, N., & Berge, G. (1998). The logic and

reliability of expert opinion on competence to stand trial. Law and Human Behavior, 22, 519–547.

Small, G. W., Kepe, V., Ercoli, L. M., Siddarth, P., Bookheimer, S. Y.,

Miller, K. J., et al. (2006). PET of brain amyloid and tau in mild

cognitive impairment. New England Journal of Medicine, 355, 2652–2663.

Somoza, E., & Mossman, D. (1991). ROC curves and the binormal

assumption. Journal of Neuropsychiatry and Clinical Neuros- ciences, 3, 436–439.

Spencer, B. D. (2008). When do latent class models overstate accuracy for binary classifiers? With applications to jury accuracy, survey response error, and diagnostic error. Institute for Policy Research, Northwestern University, Working Paper

Series WP-08-10.

State v. Sullivan, 739 N.E.2d 788 (Ohio 2001). Steadman, H. J., Mulvey, E. P., Monahan, J., Robbins, P. C.,

Appelbaum, P. S., Grisso, T., et al. (1998). Violence by people

discharged from acute psychiatric inpatient facilities and by

others in the same neighborhoods. Archives of General Psychiatry, 55, 393–401.

Swets, J. A. (1995). Signal detection theory and ROC analysis in psychology and diagnostics: Collected papers. Mahwah, NJ: Lawrence Erlbaum Associates.

Uebersax, J. S. (1988). Validity inferences from interobserver

agreement. Psychological Bulletin, 104, 405–416. Uebersax, J. S., & Grove, W. M. (1990). Latent class analysis of

diagnostic agreement. Statistics in Medicine, 9, 559–572. Weissman, H. N., & DeBow, D. M. (2003). Ethical principles and

professional competencies. In I. B. Weiner (Series Ed.) & A. M.

Goldstein (Vol. Ed.), Handbook of psychology: Vol. 11. Forensic psychology (pp. 33–53). New York: Wiley.

Zapf, P. A., Hubbard, K. L., Cooper, V. G., Wheeles, M. C., & Ronan,

K. A. (2004). Have the courts abdicated their responsibility for

determination of competency to stand trial to clinicians? Journal of Forensic Psychology Practice, 4, 27–44.

Zhou, X. H., Castelluccio, P., & Zhou, C. (2005). Nonparametric

estimation of ROC curves in the absence of a gold standard.

Biometrics, 61, 600–609. Zweig, M. H., & Campbell, G. (1993). Receiver operating character-

istic (ROC) plots: A fundamental evaluation tool in clinical

medicine. Clinical Chemistry, 39, 561–577.

Law Hum Behav (2010) 34:402–417 417

123

T hi

s do

cu m

en t i

s co

py ri

gh te

d by

th e

A m

er ic

an P

sy ch

ol og

ic al

A ss

oc ia

tio n

or o

ne o

f i ts

a lli

ed p

ub lis

he rs

. T

hi s

ar tic

le is

in te

nd ed

s ol

el y

fo r t

he p

er so

na l u

se o

f t he

in di

vi du

al u

se r a

nd is

n ot

to b

e di

ss em

in at

ed b

ro ad

ly .

<< /ASCII85EncodePages false /AllowTransparency false /AutoPositionEPSFiles true /AutoRotatePages /None /Binding /Left /CalGrayProfile (Gray Gamma 2.2) /CalRGBProfile (sRGB IEC61966-2.1) /CalCMYKProfile (ISO Coated v2 300% \050ECI\051) /sRGBProfile (sRGB IEC61966-2.1) /CannotEmbedFontPolicy /Error /CompatibilityLevel 1.3 /CompressObjects /Off /CompressPages true /ConvertImagesToIndexed true /PassThroughJPEGImages true /CreateJobTicket false /DefaultRenderingIntent /Perceptual /DetectBlends true /DetectCurves 0.1000 /ColorConversionStrategy /sRGB /DoThumbnails true /EmbedAllFonts true /EmbedOpenType false /ParseICCProfilesInComments true /EmbedJobOptions true /DSCReportingLevel 0 /EmitDSCWarnings false /EndPage -1 /ImageMemory 1048576 /LockDistillerParams true /MaxSubsetPct 100 /Optimize true /OPM 1 /ParseDSCComments true /ParseDSCCommentsForDocInfo true /PreserveCopyPage true /PreserveDICMYKValues true /PreserveEPSInfo true /PreserveFlatness true /PreserveHalftoneInfo false /PreserveOPIComments false /PreserveOverprintSettings true /StartPage 1 /SubsetFonts false /TransferFunctionInfo /Apply /UCRandBGInfo /Preserve /UsePrologue false /ColorSettingsFile () /AlwaysEmbed [ true ] /NeverEmbed [ true ] /AntiAliasColorImages false /CropColorImages true /ColorImageMinResolution 149 /ColorImageMinResolutionPolicy /Warning /DownsampleColorImages true /ColorImageDownsampleType /Bicubic /ColorImageResolution 150 /ColorImageDepth -1 /ColorImageMinDownsampleDepth 1 /ColorImageDownsampleThreshold 1.50000 /EncodeColorImages true /ColorImageFilter /DCTEncode /AutoFilterColorImages true /ColorImageAutoFilterStrategy /JPEG /ColorACSImageDict << /QFactor 0.40 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /ColorImageDict << /QFactor 0.15 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /JPEG2000ColorACSImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /JPEG2000ColorImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /AntiAliasGrayImages false /CropGrayImages true /GrayImageMinResolution 149 /GrayImageMinResolutionPolicy /Warning /DownsampleGrayImages true /GrayImageDownsampleType /Bicubic /GrayImageResolution 150 /GrayImageDepth -1 /GrayImageMinDownsampleDepth 2 /GrayImageDownsampleThreshold 1.50000 /EncodeGrayImages true /GrayImageFilter /DCTEncode /AutoFilterGrayImages true /GrayImageAutoFilterStrategy /JPEG /GrayACSImageDict << /QFactor 0.40 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /GrayImageDict << /QFactor 0.15 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /JPEG2000GrayACSImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /JPEG2000GrayImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /AntiAliasMonoImages false /CropMonoImages true /MonoImageMinResolution 599 /MonoImageMinResolutionPolicy /Warning /DownsampleMonoImages true /MonoImageDownsampleType /Bicubic /MonoImageResolution 600 /MonoImageDepth -1 /MonoImageDownsampleThreshold 1.50000 /EncodeMonoImages true /MonoImageFilter /CCITTFaxEncode /MonoImageDict << /K -1 >> /AllowPSXObjects false /CheckCompliance [ /None ] /PDFX1aCheck false /PDFX3Check false /PDFXCompliantPDFOnly false /PDFXNoTrimBoxError true /PDFXTrimBoxToMediaBoxOffset [ 0.00000 0.00000 0.00000 0.00000 ] /PDFXSetBleedBoxToMediaBox true /PDFXBleedBoxToTrimBoxOffset [ 0.00000 0.00000 0.00000 0.00000 ] /PDFXOutputIntentProfile (None) /PDFXOutputConditionIdentifier () /PDFXOutputCondition () /PDFXRegistryName () /PDFXTrapped /False /CreateJDFFile false /Description << /ARA <FEFF06270633062A062E062F0645002006470630064700200627064406250639062F0627062F0627062A002006440625064606340627062100200648062B062706260642002000410064006F00620065002000500044004600200645062A064806270641064206290020064406440637062806270639062900200641064A00200627064406450637062706280639002006300627062A0020062F0631062C0627062A002006270644062C0648062F0629002006270644063906270644064A0629061B0020064A06450643064600200641062A062D00200648062B0627062606420020005000440046002006270644064506460634062306290020062806270633062A062E062F062706450020004100630072006F0062006100740020064800410064006F006200650020005200650061006400650072002006250635062F0627063100200035002E0030002006480627064406250635062F062706310627062A0020062706440623062D062F062B002E0635062F0627063100200035002E0030002006480627064406250635062F062706310627062A0020062706440623062D062F062B002E> /BGR <FEFF04180437043f043e043b043704320430043904420435002004420435043704380020043d0430044104420440043e0439043a0438002c00200437043000200434043000200441044a0437043404300432043004420435002000410064006f00620065002000500044004600200434043e043a0443043c0435043d04420438002c0020043c0430043a04410438043c0430043b043d043e0020043f044004380433043e04340435043d04380020043704300020043204380441043e043a043e043a0430044704350441044204320435043d0020043f04350447043004420020043704300020043f044004350434043f0435044704300442043d04300020043f043e04340433043e0442043e0432043a0430002e002000200421044a04370434043004340435043d043804420435002000500044004600200434043e043a0443043c0435043d044204380020043c043e0433043004420020043404300020044104350020043e0442043204300440044f0442002004410020004100630072006f00620061007400200438002000410064006f00620065002000520065006100640065007200200035002e00300020043800200441043b0435043404320430044904380020043204350440044104380438002e> /CHS <FEFF4f7f75288fd94e9b8bbe5b9a521b5efa7684002000410064006f006200650020005000440046002065876863900275284e8e9ad88d2891cf76845370524d53705237300260a853ef4ee54f7f75280020004100630072006f0062006100740020548c002000410064006f00620065002000520065006100640065007200200035002e003000204ee553ca66f49ad87248672c676562535f00521b5efa768400200050004400460020658768633002> /CHT <FEFF4f7f752890194e9b8a2d7f6e5efa7acb7684002000410064006f006200650020005000440046002065874ef69069752865bc9ad854c18cea76845370524d5370523786557406300260a853ef4ee54f7f75280020004100630072006f0062006100740020548c002000410064006f00620065002000520065006100640065007200200035002e003000204ee553ca66f49ad87248672c4f86958b555f5df25efa7acb76840020005000440046002065874ef63002> /CZE <FEFF005400610074006f0020006e006100730074006100760065006e00ed00200070006f0075017e0069006a007400650020006b0020007600790074007600e101590065006e00ed00200064006f006b0075006d0065006e0074016f002000410064006f006200650020005000440046002c0020006b00740065007200e90020007300650020006e0065006a006c00e90070006500200068006f006400ed002000700072006f0020006b00760061006c00690074006e00ed0020007400690073006b00200061002000700072006500700072006500730073002e002000200056007900740076006f01590065006e00e900200064006f006b0075006d0065006e007400790020005000440046002000620075006400650020006d006f017e006e00e90020006f007400650076015900ed007400200076002000700072006f006700720061006d0065006300680020004100630072006f00620061007400200061002000410064006f00620065002000520065006100640065007200200035002e0030002000610020006e006f0076011b006a016100ed00630068002e> /DAN <FEFF004200720075006700200069006e0064007300740069006c006c0069006e006700650072006e0065002000740069006c0020006100740020006f007000720065007400740065002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e007400650072002c0020006400650072002000620065006400730074002000650067006e006500720020007300690067002000740069006c002000700072006500700072006500730073002d007500640073006b007200690076006e0069006e00670020006100660020006800f8006a0020006b00760061006c0069007400650074002e0020004400650020006f007000720065007400740065006400650020005000440046002d0064006f006b0075006d0065006e0074006500720020006b0061006e002000e50062006e00650073002000690020004100630072006f00620061007400200065006c006c006500720020004100630072006f006200610074002000520065006100640065007200200035002e00300020006f00670020006e0079006500720065002e> /ESP <FEFF005500740069006c0069006300650020006500730074006100200063006f006e0066006900670075007200610063006900f3006e0020007000610072006100200063007200650061007200200064006f00630075006d0065006e0074006f00730020005000440046002000640065002000410064006f0062006500200061006400650063007500610064006f00730020007000610072006100200069006d0070007200650073006900f3006e0020007000720065002d0065006400690074006f007200690061006c00200064006500200061006c00740061002000630061006c0069006400610064002e002000530065002000700075006500640065006e00200061006200720069007200200064006f00630075006d0065006e0074006f00730020005000440046002000630072006500610064006f007300200063006f006e0020004100630072006f006200610074002c002000410064006f00620065002000520065006100640065007200200035002e003000200079002000760065007200730069006f006e0065007300200070006f00730074006500720069006f007200650073002e> /ETI <FEFF004b00610073007500740061006700650020006e0065006900640020007300e4007400740065006900640020006b00760061006c006900740065006500740073006500200074007200fc006b006900650065006c007300650020007000720069006e00740069006d0069007300650020006a0061006f006b007300200073006f00620069006c0069006b0065002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e00740069006400650020006c006f006f006d006900730065006b0073002e00200020004c006f006f0064007500640020005000440046002d0064006f006b0075006d0065006e00740065002000730061006100740065002000610076006100640061002000700072006f006700720061006d006d006900640065006700610020004100630072006f0062006100740020006e0069006e0067002000410064006f00620065002000520065006100640065007200200035002e00300020006a00610020007500750065006d006100740065002000760065007200730069006f006f006e00690064006500670061002e000d000a> /FRA <FEFF005500740069006c006900730065007a00200063006500730020006f007000740069006f006e00730020006100660069006e00200064006500200063007200e900650072002000640065007300200064006f00630075006d0065006e00740073002000410064006f00620065002000500044004600200070006f0075007200200075006e00650020007100750061006c0069007400e90020006400270069006d007000720065007300730069006f006e00200070007200e9007000720065007300730065002e0020004c0065007300200064006f00630075006d0065006e00740073002000500044004600200063007200e900e90073002000700065007500760065006e0074002000ea0074007200650020006f007500760065007200740073002000640061006e00730020004100630072006f006200610074002c002000610069006e00730069002000710075002700410064006f00620065002000520065006100640065007200200035002e0030002000650074002000760065007200730069006f006e007300200075006c007400e90072006900650075007200650073002e> /GRE <FEFF03a703c103b703c303b903bc03bf03c003bf03b903ae03c303c403b5002003b103c503c403ad03c2002003c403b903c2002003c103c503b803bc03af03c303b503b903c2002003b303b903b1002003bd03b1002003b403b703bc03b903bf03c503c103b303ae03c303b503c403b5002003ad03b303b303c103b103c603b1002000410064006f006200650020005000440046002003c003bf03c5002003b503af03bd03b103b9002003ba03b103c42019002003b503be03bf03c703ae03bd002003ba03b103c403ac03bb03bb03b703bb03b1002003b303b903b1002003c003c103bf002d03b503ba03c403c503c003c903c403b903ba03ad03c2002003b503c103b303b103c303af03b503c2002003c503c803b703bb03ae03c2002003c003bf03b903cc03c403b703c403b103c2002e0020002003a403b10020005000440046002003ad03b303b303c103b103c603b1002003c003bf03c5002003ad03c703b503c403b5002003b403b703bc03b903bf03c503c103b303ae03c303b503b9002003bc03c003bf03c103bf03cd03bd002003bd03b1002003b103bd03bf03b903c703c403bf03cd03bd002003bc03b5002003c403bf0020004100630072006f006200610074002c002003c403bf002000410064006f00620065002000520065006100640065007200200035002e0030002003ba03b103b9002003bc03b503c403b103b303b503bd03ad03c303c403b503c103b503c2002003b503ba03b403cc03c303b503b903c2002e> /HEB <FEFF05D405E905EA05DE05E905D5002005D105D405D205D305E805D505EA002005D005DC05D4002005DB05D305D9002005DC05D905E605D505E8002005DE05E105DE05DB05D9002000410064006F006200650020005000440046002005D405DE05D505EA05D005DE05D905DD002005DC05D405D305E405E105EA002005E705D305DD002D05D305E405D505E1002005D005D905DB05D505EA05D905EA002E002005DE05E105DE05DB05D90020005000440046002005E905E005D505E605E805D5002005E005D905EA05E005D905DD002005DC05E405EA05D905D705D4002005D105D005DE05E605E205D505EA0020004100630072006F006200610074002005D5002D00410064006F00620065002000520065006100640065007200200035002E0030002005D505D205E805E105D005D505EA002005DE05EA05E705D305DE05D505EA002005D905D505EA05E8002E05D005DE05D905DD002005DC002D005000440046002F0058002D0033002C002005E205D905D905E005D5002005D105DE05D305E805D905DA002005DC05DE05E905EA05DE05E9002005E905DC0020004100630072006F006200610074002E002005DE05E105DE05DB05D90020005000440046002005E905E005D505E605E805D5002005E005D905EA05E005D905DD002005DC05E405EA05D905D705D4002005D105D005DE05E605E205D505EA0020004100630072006F006200610074002005D5002D00410064006F00620065002000520065006100640065007200200035002E0030002005D505D205E805E105D005D505EA002005DE05EA05E705D305DE05D505EA002005D905D505EA05E8002E> /HRV (Za stvaranje Adobe PDF dokumenata najpogodnijih za visokokvalitetni ispis prije tiskanja koristite ove postavke. Stvoreni PDF dokumenti mogu se otvoriti Acrobat i Adobe Reader 5.0 i kasnijim verzijama.) /HUN <FEFF004b0069007600e1006c00f30020006d0069006e0151007300e9006701710020006e0079006f006d00640061006900200065006c0151006b00e90073007a00ed007401510020006e0079006f006d00740061007400e100730068006f007a0020006c006500670069006e006b00e1006200620020006d0065006700660065006c0065006c0151002000410064006f00620065002000500044004600200064006f006b0075006d0065006e00740075006d006f006b0061007400200065007a0065006b006b0065006c0020006100200062006500e1006c006c00ed007400e10073006f006b006b0061006c0020006b00e90073007a00ed0074006800650074002e0020002000410020006c00e90074007200650068006f007a006f00740074002000500044004600200064006f006b0075006d0065006e00740075006d006f006b00200061007a0020004100630072006f006200610074002000e9007300200061007a002000410064006f00620065002000520065006100640065007200200035002e0030002c0020007600610067007900200061007a002000610074007400f3006c0020006b00e9007301510062006200690020007600650072007a006900f3006b006b0061006c0020006e00790069007400680061007400f3006b0020006d00650067002e> /ITA <FEFF005500740069006c0069007a007a006100720065002000710075006500730074006500200069006d0070006f007300740061007a0069006f006e00690020007000650072002000630072006500610072006500200064006f00630075006d0065006e00740069002000410064006f00620065002000500044004600200070006900f900200061006400610074007400690020006100200075006e00610020007000720065007300740061006d0070006100200064006900200061006c007400610020007100750061006c0069007400e0002e0020004900200064006f00630075006d0065006e007400690020005000440046002000630072006500610074006900200070006f00730073006f006e006f0020006500730073006500720065002000610070006500720074006900200063006f006e0020004100630072006f00620061007400200065002000410064006f00620065002000520065006100640065007200200035002e003000200065002000760065007200730069006f006e006900200073007500630063006500730073006900760065002e> /JPN <FEFF9ad854c18cea306a30d730ea30d730ec30b951fa529b7528002000410064006f0062006500200050004400460020658766f8306e4f5c6210306b4f7f75283057307e305930023053306e8a2d5b9a30674f5c62103055308c305f0020005000440046002030d530a130a430eb306f3001004100630072006f0062006100740020304a30883073002000410064006f00620065002000520065006100640065007200200035002e003000204ee5964d3067958b304f30533068304c3067304d307e305930023053306e8a2d5b9a306b306f30d530a930f330c8306e57cb30818fbc307f304c5fc59808306730593002> /KOR <FEFFc7740020c124c815c7440020c0acc6a9d558c5ec0020ace0d488c9c80020c2dcd5d80020c778c1c4c5d00020ac00c7a50020c801d569d55c002000410064006f0062006500200050004400460020bb38c11cb97c0020c791c131d569b2c8b2e4002e0020c774b807ac8c0020c791c131b41c00200050004400460020bb38c11cb2940020004100630072006f0062006100740020bc0f002000410064006f00620065002000520065006100640065007200200035002e00300020c774c0c1c5d0c11c0020c5f40020c2180020c788c2b5b2c8b2e4002e> /LTH <FEFF004e006100750064006f006b0069007400650020016100690075006f007300200070006100720061006d006500740072007500730020006e006f0072011700640061006d00690020006b0075007200740069002000410064006f00620065002000500044004600200064006f006b0075006d0065006e007400750073002c0020006b00750072006900650020006c0061006200690061007500730069006100690020007000720069007400610069006b007900740069002000610075006b01610074006f00730020006b006f006b007900620117007300200070006100720065006e006700740069006e00690061006d00200073007000610075007300640069006e0069006d00750069002e0020002000530075006b0075007200740069002000500044004600200064006f006b0075006d0065006e007400610069002000670061006c006900200062016b007400690020006100740069006400610072006f006d00690020004100630072006f006200610074002000690072002000410064006f00620065002000520065006100640065007200200035002e0030002000610072002000760117006c00650073006e0117006d00690073002000760065007200730069006a006f006d00690073002e> /LVI <FEFF0049007a006d0061006e0074006f006a00690065007400200161006f00730020006900650073007400610074012b006a0075006d00750073002c0020006c0061006900200076006500690064006f00740075002000410064006f00620065002000500044004600200064006f006b0075006d0065006e007400750073002c0020006b006100730020006900720020012b00700061016100690020007000690065006d01130072006f00740069002000610075006700730074006100730020006b00760061006c0069007401010074006500730020007000690072006d007300690065007300700069006501610061006e006100730020006400720075006b00610069002e00200049007a0076006500690064006f006a006900650074002000500044004600200064006f006b0075006d0065006e007400750073002c0020006b006f002000760061007200200061007400760113007200740020006100720020004100630072006f00620061007400200075006e002000410064006f00620065002000520065006100640065007200200035002e0030002c0020006b0101002000610072012b00200074006f0020006a00610075006e0101006b0101006d002000760065007200730069006a0101006d002e> /NLD (Gebruik deze instellingen om Adobe PDF-documenten te maken die zijn geoptimaliseerd voor prepress-afdrukken van hoge kwaliteit. De gemaakte PDF-documenten kunnen worden geopend met Acrobat en Adobe Reader 5.0 en hoger.) /NOR <FEFF004200720075006b00200064006900730073006500200069006e006e007300740069006c006c0069006e00670065006e0065002000740069006c002000e50020006f0070007000720065007400740065002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e00740065007200200073006f006d00200065007200200062006500730074002000650067006e0065007400200066006f00720020006600f80072007400720079006b006b0073007500740073006b00720069006600740020006100760020006800f800790020006b00760061006c0069007400650074002e0020005000440046002d0064006f006b0075006d0065006e00740065006e00650020006b0061006e002000e50070006e00650073002000690020004100630072006f00620061007400200065006c006c00650072002000410064006f00620065002000520065006100640065007200200035002e003000200065006c006c00650072002000730065006e006500720065002e> /POL <FEFF0055007300740061007700690065006e0069006100200064006f002000740077006f0072007a0065006e0069006100200064006f006b0075006d0065006e007400f300770020005000440046002000700072007a0065007a006e00610063007a006f006e00790063006800200064006f002000770079006400720075006b00f30077002000770020007700790073006f006b00690065006a0020006a0061006b006f015b00630069002e002000200044006f006b0075006d0065006e0074007900200050004400460020006d006f017c006e00610020006f007400770069006500720061010700200077002000700072006f006700720061006d006900650020004100630072006f00620061007400200069002000410064006f00620065002000520065006100640065007200200035002e0030002000690020006e006f00770073007a0079006d002e> /PTB <FEFF005500740069006c0069007a006500200065007300730061007300200063006f006e00660069006700750072006100e700f50065007300200064006500200066006f0072006d00610020006100200063007200690061007200200064006f00630075006d0065006e0074006f0073002000410064006f0062006500200050004400460020006d00610069007300200061006400650071007500610064006f00730020007000610072006100200070007200e9002d0069006d0070007200650073007300f50065007300200064006500200061006c007400610020007100750061006c00690064006100640065002e0020004f007300200064006f00630075006d0065006e0074006f00730020005000440046002000630072006900610064006f007300200070006f00640065006d0020007300650072002000610062006500720074006f007300200063006f006d0020006f0020004100630072006f006200610074002000650020006f002000410064006f00620065002000520065006100640065007200200035002e0030002000650020007600650072007300f50065007300200070006f00730074006500720069006f007200650073002e> /RUM <FEFF005500740069006c0069007a00610163006900200061006300650073007400650020007300650074010300720069002000700065006e007400720075002000610020006300720065006100200064006f00630075006d0065006e00740065002000410064006f006200650020005000440046002000610064006500630076006100740065002000700065006e0074007200750020007400690070010300720069007200650061002000700072006500700072006500730073002000640065002000630061006c006900740061007400650020007300750070006500720069006f006100720103002e002000200044006f00630075006d0065006e00740065006c00650020005000440046002000630072006500610074006500200070006f00740020006600690020006400650073006300680069007300650020006300750020004100630072006f006200610074002c002000410064006f00620065002000520065006100640065007200200035002e00300020015f00690020007600650072007300690075006e0069006c006500200075006c0074006500720069006f006100720065002e> /RUS <FEFF04180441043f043e043b044c04370443043904420435002004340430043d043d044b04350020043d0430044104420440043e0439043a043800200434043b044f00200441043e043704340430043d0438044f00200434043e043a0443043c0435043d0442043e0432002000410064006f006200650020005000440046002c0020043c0430043a04410438043c0430043b044c043d043e0020043f043e04340445043e0434044f04490438044500200434043b044f00200432044b0441043e043a043e043a0430044704350441044204320435043d043d043e0433043e00200434043e043f0435044704300442043d043e0433043e00200432044b0432043e04340430002e002000200421043e043704340430043d043d044b04350020005000440046002d0434043e043a0443043c0435043d0442044b0020043c043e0436043d043e0020043e0442043a0440044b043204300442044c002004410020043f043e043c043e0449044c044e0020004100630072006f00620061007400200438002000410064006f00620065002000520065006100640065007200200035002e00300020043800200431043e043b043504350020043f043e04370434043d043804450020043204350440044104380439002e> /SKY <FEFF0054006900650074006f0020006e006100730074006100760065006e0069006100200070006f0075017e0069007400650020006e00610020007600790074007600e100720061006e0069006500200064006f006b0075006d0065006e0074006f0076002000410064006f006200650020005000440046002c0020006b0074006f007200e90020007300610020006e0061006a006c0065007001610069006500200068006f0064006900610020006e00610020006b00760061006c00690074006e00fa00200074006c0061010d00200061002000700072006500700072006500730073002e00200056007900740076006f00720065006e00e900200064006f006b0075006d0065006e007400790020005000440046002000620075006400650020006d006f017e006e00e90020006f00740076006f00720069016500200076002000700072006f006700720061006d006f006300680020004100630072006f00620061007400200061002000410064006f00620065002000520065006100640065007200200035002e0030002000610020006e006f0076016100ed00630068002e> /SLV <FEFF005400650020006e006100730074006100760069007400760065002000750070006f0072006100620069007400650020007a00610020007500730074007600610072006a0061006e006a006500200064006f006b0075006d0065006e0074006f0076002000410064006f006200650020005000440046002c0020006b006900200073006f0020006e0061006a007000720069006d00650072006e0065006a016100690020007a00610020006b0061006b006f0076006f00730074006e006f0020007400690073006b0061006e006a00650020007300200070007200690070007200610076006f0020006e00610020007400690073006b002e00200020005500730074007600610072006a0065006e006500200064006f006b0075006d0065006e0074006500200050004400460020006a00650020006d006f0067006f010d00650020006f0064007000720065007400690020007a0020004100630072006f00620061007400200069006e002000410064006f00620065002000520065006100640065007200200035002e003000200069006e0020006e006f00760065006a01610069006d002e> /SUO <FEFF004b00e40079007400e40020006e00e40069007400e4002000610073006500740075006b007300690061002c0020006b0075006e0020006c0075006f00740020006c00e400680069006e006e00e4002000760061006100740069007600610061006e0020007000610069006e006100740075006b00730065006e002000760061006c006d0069007300740065006c00750074007900f6006800f6006e00200073006f00700069007600690061002000410064006f0062006500200050004400460020002d0064006f006b0075006d0065006e007400740065006a0061002e0020004c0075006f0064007500740020005000440046002d0064006f006b0075006d0065006e00740069007400200076006f0069006400610061006e0020006100760061007400610020004100630072006f0062006100740069006c006c00610020006a0061002000410064006f00620065002000520065006100640065007200200035002e0030003a006c006c00610020006a006100200075007500640065006d006d0069006c006c0061002e> /SVE <FEFF0041006e007600e4006e00640020006400650020006800e4007200200069006e0073007400e4006c006c006e0069006e006700610072006e00610020006f006d002000640075002000760069006c006c00200073006b006100700061002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e007400200073006f006d002000e400720020006c00e4006d0070006c0069006700610020006600f60072002000700072006500700072006500730073002d007500740073006b00720069006600740020006d006500640020006800f600670020006b00760061006c0069007400650074002e002000200053006b006100700061006400650020005000440046002d0064006f006b0075006d0065006e00740020006b0061006e002000f600700070006e00610073002000690020004100630072006f0062006100740020006f00630068002000410064006f00620065002000520065006100640065007200200035002e00300020006f00630068002000730065006e006100720065002e> /TUR <FEFF005900fc006b00730065006b0020006b0061006c006900740065006c0069002000f6006e002000790061007a006401310072006d00610020006200610073006b013100730131006e006100200065006e0020006900790069002000750079006100620069006c006500630065006b002000410064006f006200650020005000440046002000620065006c00670065006c0065007200690020006f006c0075015f007400750072006d0061006b0020006900e70069006e00200062007500200061007900610072006c0061007201310020006b0075006c006c0061006e0131006e002e00200020004f006c0075015f0074007500720075006c0061006e0020005000440046002000620065006c00670065006c0065007200690020004100630072006f006200610074002000760065002000410064006f00620065002000520065006100640065007200200035002e003000200076006500200073006f006e0072006100730131006e00640061006b00690020007300fc007200fc006d006c00650072006c00650020006100e70131006c006100620069006c00690072002e> /UKR <FEFF04120438043a043e0440043804410442043e043204430439044204350020044604560020043f043004400430043c043504420440043800200434043b044f0020044104420432043e04400435043d043d044f00200434043e043a0443043c0435043d044204560432002000410064006f006200650020005000440046002c0020044f043a04560020043d04300439043a04400430044904350020043f045604340445043e0434044f0442044c00200434043b044f0020043204380441043e043a043e044f043a04560441043d043e0433043e0020043f0435044004350434043404400443043a043e0432043e0433043e0020043404400443043a0443002e00200020042104420432043e04400435043d045600200434043e043a0443043c0435043d0442043800200050004400460020043c043e0436043d04300020043204560434043a0440043804420438002004430020004100630072006f006200610074002004420430002000410064006f00620065002000520065006100640065007200200035002e0030002004300431043e0020043f04560437043d04560448043e04570020043204350440044104560457002e> /ENU (Use these settings to create Adobe PDF documents best suited for high-quality prepress printing. Created PDF documents can be opened with Acrobat and Adobe Reader 5.0 and later.) /DEU <FEFF004a006f0062006f007000740069006f006e007300200066006f00720020004100630072006f006200610074002000440069007300740069006c006c0065007200200038002000280038002e0032002e00310029000d00500072006f006400750063006500730020005000440046002000660069006c0065007300200077006800690063006800200061007200650020007500730065006400200066006f00720020006f006e006c0069006e0065002e000d0028006300290020003200300031003000200053007000720069006e006700650072002d005600650072006c0061006700200047006d006200480020000d000d0054006800650020006c00610074006500730074002000760065007200730069006f006e002000630061006e00200062006500200064006f0077006e006c006f0061006400650064002000610074002000680074007400700073003a002f002f0070006f007200740061006c002d0064006f0072006400720065006300680074002e0073007000720069006e006700650072002d00730062006d002e0063006f006d002f00500072006f00640075006300740069006f006e002f0046006c006f0077002f00740065006300680064006f0063002f00640065006600610075006c0074002e0061007300700078000d0054006800650072006500200079006f0075002000630061006e00200061006c0073006f002000660069006e0064002000610020007300750069007400610062006c006500200045006e0066006f0063007500730020005000440046002000500072006f00660069006c006500200066006f0072002000500069007400530074006f0070002000500072006f00660065007300730069006f006e0061006c00200030003800200061006e0064002000500069007400530074006f0070002000530065007200760065007200200030003800200066006f007200200070007200650066006c00690067006800740069006e006700200079006f007500720020005000440046002000660069006c006500730020006200650066006f007200650020006a006f00620020007300750062006d0069007300730069006f006e002e000d> >> /Namespace [ (Adobe) (Common) (1.0) ] /OtherNamespaces [ << /AsReaderSpreads false /CropImagesToFrames true /ErrorControl /WarnAndContinue /FlattenerIgnoreSpreadOverrides false /IncludeGuidesGrids false /IncludeNonPrinting false /IncludeSlug false /Namespace [ (Adobe) (InDesign) (4.0) ] /OmitPlacedBitmaps false /OmitPlacedEPS false /OmitPlacedPDF false /SimulateOverprint /Legacy >> << /AddBleedMarks false /AddColorBars false /AddCropMarks false /AddPageInfo false /AddRegMarks false /ConvertColors /ConvertToCMYK /DestinationProfileName () /DestinationProfileSelector /DocumentCMYK /Downsample16BitImages true /FlattenerPreset << /PresetSelector /MediumResolution >> /FormElements false /GenerateStructure false /IncludeBookmarks false /IncludeHyperlinks false /IncludeInteractive false /IncludeLayers false /IncludeProfiles false /MultimediaHandling /UseObjectSettings /Namespace [ (Adobe) (CreativeSuite) (2.0) ] /PDFXOutputIntentProfileSelector /DocumentCMYK /PreserveEditing true /UntaggedCMYKHandling /LeaveUntagged /UntaggedRGBHandling /UseDocumentProfile /UseDocumentBleed false >> ] >> setdistillerparams << /HWResolution [2400 2400] /PageSize [595.276 841.890] >> setpagedevice