2-3 pages APA format, masters level, review attachments
O R I G I N A L A R T I C L E
Quantifying the Accuracy of Forensic Examiners in the Absence of a ‘‘Gold Standard’’
Douglas Mossman Æ Michael D. Bowen Æ David J. Vanness Æ David Bienenfeld Æ Terry Correll Æ Jerald Kay Æ William M. Klykylo Æ Douglas S. Lehrer
Published online: 22 September 2009
� American Psychology-Law Society/Division 41 of the American Psychological Association 2009
Abstract This study asked whether latent class modeling
methods and multiple ratings of the same cases might per-
mit quantification of the accuracy of forensic assessments.
Five evaluators examined 156 redacted court reports
concerning criminal defendants who had undergone
hospitalization for evaluation or restoration of their adju-
dicative competence. Evaluators rated each defendant’s
Dusky-defined competence to stand trial on a five-point
scale as well as each defendant’s understanding of, appre-
ciation of, and reasoning about criminal proceedings.
Having multiple ratings per defendant made it possible to
estimate accuracy parameters using maximum likelihood
and Bayesian approaches, despite the absence of any ‘‘gold
standard’’ for the defendants’ true competence status.
Evaluators appeared to be very accurate, though this finding
should be viewed with caution.
Keywords Competence to stand trial � Adjudicative competence � ROC analysis � Diagnostic accuracy � Maximum likelihood � Bayesian � Gold standard
Daubert v. Merrell Dow Pharmaceuticals (1993) directs
judges to evaluate proffered scientific testimony based on
factors such as ‘‘the known or potential rate of error’’ of the
‘‘particular scientific technique’’ and whether the technique
has been tested (pp. 592–593). A subsequent U.S. Supreme
Court decision, Kumho Tire Co. v. Carmichael (1999),
extended trial courts’ gate-keeping role and the applica-
bility of Daubert factors to ‘‘other experts who are not
scientists’’ (p. 137). Thus, in federal courts and other U.S.
jurisdictions that follow Daubert-like evidentiary rules,
testifying mental health professionals may be asked,
‘‘Doctor, has your method been tested?’’ and ‘‘How accu-
rate is it?’’
Practitioners of most medical specialties use diagnostic
methods for which accuracy statistics such as sensitivity
and specificity are available. Also, for many diagnostic
modalities—e.g., mammography for breast cancer detec-
tion (Berg et al., 2008; Lehman et al., 2007)—long-term
Portions of this work were presented at (1) Annual Meeting of the
American Academy of Psychiatry and the Law, Miami Beach,
Florida, October 18, 2007; (2) American Psychology-Law Society
Conference, Jacksonville, Florida, March 7, 2008; (3) Annual
Meeting of the Midwest Chapter of the American Academy of
Psychiatry and the Law, Renaissance Hotel, Cleveland, Ohio, March
29, 2008; (4) Cincinnati Psychiatric Society, June 17, 2008; and
(5) Summit Behavioral Healthcare, June 11, 2009.
D. Mossman (&) Glenn M. Weaver Institute of Law and Psychiatry, University of
Cincinnati College of Law, Clifton Avenue & Calhoun Street,
PO Box 210040, Cincinnati, OH 45221-0040, USA
e-mail: [email protected]
D. Mossman
Department of Psychiatry, University of Cincinnati College of
Medicine, Cincinnati, USA
D. Mossman � M. D. Bowen � D. Bienenfeld � T. Correll � J. Kay � W. M. Klykylo � D. S. Lehrer Department of Psychiatry, Wright State University, Boonshoft
School of Medicine, Dayton, USA
D. J. Vanness
Department of Population Health Sciences, University of
Wisconsin School of Medicine and Public Health, Madison,
USA
D. J. Vanness
Center for Health Economics and Science Policy, United
BioSource Corporation, Bethesda, USA
123
Law Hum Behav (2010) 34:402–417
DOI 10.1007/s10979-009-9197-5
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
follow-up or biopsy results provide an independent, virtu-
ally infallible criterion for true disease status. By contrast,
psychiatric classifications usually have no independent
‘‘gold standard’’ (Faraone & Tsuang, 1994), and diagnostic
criteria consist entirely of clinical findings.
Where psycholegal determinations are concerned, the
absence of an indubitable truth criterion is both common
and arguably more troubling than in clinical mental health
practice. Ethical principles of forensic psychiatrists and
forensic psychologists encourage adherence to high stan-
dards of competence, impartiality, and scientific rigor
(American Academy of Psychiatry and the Law, 2005;
Committee on the Revision of the Specialty Guidelines for
Forensic Psychology, 2006; Weissman & DeBow, 2003).
Yet popular works (Hagen, 1997), appellate-level legal
opinions (Mossman, 1999), and evidence texts (Faigman,
Saks, Sanders, & Cheng, 2008) suggest that psycholegal
experts seem like ‘‘whores’’ and ‘‘hired guns’’ who ‘‘are
merely selling their testimony to the highest bidder’’
(Melton et al., 2007, p. 577). When opposing experts dis-
agree, courtroom cross-examination often becomes an
intensive effort to question the integrity of psychiatric
diagnoses and to discredit all mental health expertise.
Several instruments relevant to forensic assessment
offer bases for expert opinion that appear more reliable and
systematic than unaided clinical judgment (Dawes, Faust,
& Meehl, 1989; Gardner, Lidz, Mulvey, & Shaw, 1996;
Harris, Rice, & Cormier, 2002). Over the last two decades,
many publications have described the accuracy of such
instruments, often using receiver operating characteristic
(ROC) methods (Douglas, Ogloff, Nicholls, & Grant, 1999;
Rice & Harris, 1995). For tools used to assess violence risk,
subjects’ arrest or conviction records supplemented with
interview and collateral data have served as external cri-
teria for gauging accuracy (Steadman et al., 1998). In many
circumstances, however, the closest approximation to truth
is a professional’s well-considered opinion. In several
studies that examine accuracy of assessment tools or pre-
dictions (e.g., Akinkunmi, 2002; Kim et al., 2007;
Mossman, 2007), judgments of experienced clinicians have
provided the ‘‘truth’’ or ultimately ‘‘right’’ conclusion. Yet
the wisest experts err in their clinical and forensic assess-
ments, and though judges or juries render ultimately
binding decisions in courtrooms, legal fact-finders are
humanly fallible. Thus, for most psycholegal determina-
tions, all we can hope for are various individuals’
conclusions.
Faced with this epistemological barrier, several studies
(e.g., Cooper & Zapf, 2003; Jacobs, Ryba, & Zapf, 2008)
have looked for correlates of or factors related to forensic
assessment tools or examiners’ opinions. Other studies
(e.g., Boccaccini, Turner, & Murrie, 2008; Murrie, Boc-
caccini, Zapf, Warren, & Henderson, 2008; Murrie et al.,
2009) have quantified absolute diagnostic agreement
between evaluators and have examined possible sources of
disagreement. Though several of these studies recognize
and discuss consequences of rater bias and disagreement,
they have been agnostic (so to speak) about whether a
question such as ‘‘Is this defendant competent to stand
trial?’’ has a ‘‘right’’ answer. Yet the existence of a right
answer is a condition of the possibility of thinking—as any
experienced forensic examiner occasionally does—that a
court has made an error in finding a particular criminal
defendant competent or incompetent. And if we acknowl-
edge that judgments about psycholegal matters can be right
or wrong, we can also wonder how accurate such judg-
ments are, despite our having no gold standard to establish
absolute truth in any particular case.
Over the last two decades, latent class modeling (LCM)
(Uebersax & Grove, 1990) has shown promise in permit-
ting ROC analyses without gold standards in subject areas
as diverse as imaging liver metastases (Henkelman, Kay, &
Bronskill, 1990) and detecting infections in dairy cattle
(Choi, Johnson, Collins, & Gardner, 2006). Implementing
the LCM approach involves evaluating the same cases with
multiple diagnostic modalities, which often permits statis-
tical identification of models that includes accuracy
parameters for those modalities. To learn whether LCM
could characterize and quantify the accuracy of forensic
examiners, the present study considers the most-often-
performed (Melton et al., 2007; Mossman et al., 2007)
criminal forensic evaluation: assessing competence to
stand trial.
BACKGROUND
Legal Criteria
All U.S. jurisdictions (Bennett, 1985) define competence to
stand trial (CST) consistent with Dusky v. United States,
which states that a defendant is CST if ‘‘he has sufficient
present ability to consult with his lawyer with a reasonable
degree of rational understanding’’ and ‘‘has a rational as
well as factual understanding of the proceedings against
him’’ (Dusky v. United States, 1960, p. 402). Whenever ‘‘a
bona fide doubt’’ about a criminal defendant’s fitness
arises, the trial court must hold a hearing concerning his
CST (Pate v. Robinson, 1966). All U.S. jurisdictions permit
courts to order mental health evaluations of criminal
defendants for use in hearings on CST (Mossman et al.,
2007). Following findings of incompetence, courts usually
order defendants to undergo ‘‘restoration’’—treatment,
usually at a public sector hospital, aimed at rendering the
defendants competent—if such treatment has a substantial
probability of being successful (Jackson v. Indiana, 1972;
Law Hum Behav (2010) 34:402–417 403
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
Mossman et al., 2007; State v. Sullivan, 2001). Around
60,000 U.S. criminal defendants undergo CST evaluations
each year, and roughly 4,000 U.S. hospital beds are
occupied by defendants who are undergoing CST restora-
tion (Mossman, 2005). Hospitals send periodic reports to
referring courts concerning defendants’ progress toward
achieving competence (Parry & Drogin, 2007).
Conceptualizing Experts’ Assessments of CST
When forensic examiners assess CST, they (explicitly or
implicitly) consider many mental faculties and several
dimensions of social and interpersonal functioning,
including defendants’ abilities to comprehend social situ-
ations, to project themselves into hypothetical situations, to
function in collaborative relationships, to recognize what
things are relevant in complex social situations, to com-
municate logically, and to maintain self-control (Mossman,
2008). For this reason, writers have characterized adjudi-
cative competence as an abstract, ‘‘open-textured’’
construct (Bonnie, 1990) intended ‘‘to apply to an infinite
number of fact situations’’ (Golding, Roesch, & Schreiber,
1984, p. 323), or as some actual though hypostatized fea-
ture of defendants (Grisso, 2003). On this view,
adjudicative competence is what Cronbach and Meehl call
a ‘‘postulated attribute’’ (1955, p. 283)—an imperfectly
defined but real property of defendants—and judgments
about adjudicative competence (whether made by exam-
iners or courts) reflect beliefs about the degree to which
defendants exhibit this property. Alternatively (says this
view), competent and incompetent defendants may com-
prise two natural, distinct groups or ‘‘taxa,’’ and a decision
about adjudicative competence is a decision about whether
a defendant falls into one taxon or the other.
Following Mossman (2008), however, we regard CST
evaluations as contextual assessments that ask whether a
defendant can do something—meet a standard—rather than
whether the defendant has a property or belongs to one
natural category or another. On our view, to ask the
question ‘‘Is Defendant Jones competent to stand trial?’’ is
like asking whether Jones can jump over a hurdle of a
given height. The features and qualities that allow people
to clear hurdles (their height, weight, leg strength, coor-
dination, etc.) are various and fall along continua, but
whether Jones can clear a particular hurdle is a yes-or-no
matter. If the hurdle is 15 cm (6 in.) high, experience lets
us be very confident that a randomly selected, able-bodied,
middle-aged adult should have no trouble jumping over it,
while toddlers and frail elderly people will not succeed.
The hurdle homology is not perfect, of course. If we
agree on an operational definition of ‘‘able to clear a 15-cm
hurdle’’ (e.g., ‘‘you get ten tries, and you have to succeed
just once’’), we can evaluate Jones against the criterion and
determine definitively whether he can clear the hurdle. By
contrast, evaluating CST requires a forensic examiner to
assess many hard-to-quantify personal qualities of a
defendant against the also-hard-to-quantify demands of a
particular criminal case. In many cases, examiners justifi-
ably feel high confidence about a defendant’s adjudicative
competence or lack thereof. But we have no way to verify
for certain whether a defendant who faces one or more
specific charges understands the nature and objective of the
proceedings against him and can assist his attorney in
preparing a defense. The characterizations of CST found in
statutes and case law guide examiners’ beliefs and judges’
opinions about adjudicative competence, but statutes
and case law do not operationally define adjudicative
competence.
Yet, just as we know that most able-bodied middle-aged
adults can clear 15-cm hurdles, we know that most defen-
dants—indeed, almost all adults who do not have serious
psychopathology or cognitive impairment—are competent
to stand trial. A major portion of a CST assessment involves
determining whether a serious psychiatric disorder or cog-
nitive impairment prevents a particular defendant from
doing what most individuals could easily do. Just as general
experience informs everyone about the personal character-
istics that might preclude someone from jumping over a low
hurdle, training and experience inform mental health pro-
fessionals about the kinds of impairments that would keep a
defendant from doing the basic (though harder to measure)
mental tasks needed to stand trial. For example, verbal
incoherence is hard to quantify, but its adverse impact on
communication—and on adjudicative competence—is easy
to apprehend. This suggests that if forensic evaluators have
adequate information available, they should be able to make
very good (though not perfect) judgments about adjudica-
tive competence.
METHOD
Data Collection and Development
Our study received approval from the Institutional Review
Board of Wright State University and from the Office of
Program Evaluation and Research of the Ohio Department
of Mental Health. We obtained 156 reports on CST that a
public sector hospital had submitted to Ohio criminal
courts in 1994–2001. The reports described criminal
defendants who had undergone court-ordered hospitaliza-
tions either for evaluation of CST or for restoration. In
accordance with state statutory requirements (Ohio
Revised Code §2945.38(F)), each report included the
examiner’s ‘‘penultimate’’ opinion on CST (i.e., a state-
ment about whether the defendant could understand the
404 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
pending proceedings and assist defense counsel). We
selected reports largely at random, though we had to reject
some reports that were poorly written or did not include
enough detail about defendants’ condition, treatment
course, current mental status, and responses to CST-spe-
cific inquiries to allow formation of independent opinions
about adjudicative competence. Previous studies on quan-
tifying accuracy without gold standards (Albert, 2007;
Henkelman et al., 1990; Zhou, Castelluccio, & Zhou, 2005)
suggested that our raters should examine at least 150
reports to facilitate convergence of the mathematical
algorithms used to identify accuracy parameters; including
six additional reports provided a margin of safety in case
raters found some reports unusable.
Two authors prepared ‘‘sanitized’’ versions of each
original report by (1) removing the original examiners’
diagnoses and forensic opinions about CST, (2) substitut-
ing pseudonyms for original names, and (3) approximating,
disguising, paraphrasing, or deleting other identifying data
(e.g., precise ages, ethnicity, dates and locales of previous
hospitalizations). Each sanitized report retained disguised
background information (including medical, legal, and
psychiatric history), though we removed material that was
irrelevant to CST or too personal (e.g., childhood sexual
abuse). Defendants’ criminal charges and sexes were not
changed because these items often were needed to assess
CST or to interpret background information. Sanitized
reports also retained the original examiners’ descriptions of
evaluees’ hospital course, current functioning, mental sta-
tus, responses to questions specifically related to CST (e.g.,
‘‘What are you charged with?’’), results from psychological
testing (e.g., MMPI-2 or intelligence scales), and/or scores
from structured assessment instruments (e.g., the Georgia
Court Competency Test).
Serving as raters were five board-certified, experienced
(12–33 years post-residency) psychiatrists (three also
board-certified in forensic psychiatry) who had not helped
prepare the sanitized reports. After reading each of the 156
reports, raters assigned scores on five-point scales con-
cerning the defendant’s understanding of information
relevant to, ability to reason about, and appreciation of his
current legal situation, applying definitions of these con-
cepts used in the MacArthur studies on adjudicative
competence (Poythress, Monahan, Bonnie, Otto, & Hoge,
2002). Each rater also provided ordinal scale scores
reflecting his overall rating of CST as defined under Dusky
v. United States (1960). Table 1 shows portions of a rater’s
data sheet.
Conceptualizing the Data
Having raters provide impressions using five-point scales
differs from usual forensic experts’ practice and from
courts’ usual expectation that experts will render binary
opinions (either competent or incompetent) with reason-
able medical or scientific certainty. Buchanan (2006) notes,
however, that experts reach yes-or-no opinions with vary-
ing degrees of assuredness; judgments about CST
incorporate subordinate judgments about defendants’
mental functioning, potential penalties, case-specific
demands (e.g., complexity of information that a defendant
must process), and consequences of errors (e.g., possibly
having a marginally incompetent defendant stand trial for a
serious crime). Some evidence suggests that examiners
may require a higher level of competence for defendants
charged with more serious offenses (Rosenfeld & Ritchie,
1998). An expert’s conclusion about a defendant’s CST
thus reflects an estimate of the defendant’s relevant abili-
ties coupled with the expert’s judgments about benefits and
costs of correct and incorrect outcomes. Also, recent pub-
lications suggest that forensic opinions display evidence of
‘‘adversarial allegiance’’ to retaining parties (Murrie et al.,
2009), and that yes-or-no judgments about CST reflect
experts’ professional backgrounds and views about the
impact of psychosis (Murrie et al., 2008). To properly
gauge an expert’s ability to discern competent from
incompetent defendants, one should use a mathematical
method that can tease out intrinsic discriminatory capacity
from biasing factors and the expert’s concerns about con-
sequences of errors (Mossman, 2008).
Over the last four decades, investigators in medicine and
psychology have increasingly turned to ROC analysis
to address this type of problem. ROC analysis separates
Table 1 Portions of the Rater Scoring Sheet, with numbers assigned to each rating category
Understanding
Defendant’s cognitive apprehension, at a descriptive level, of basic
legal functions; his ability to comprehend information relevant
to the adjudicative process
h 1 = definitely satisfactory
h 2 = probably satisfactory
h 3 = uncertain
h 4 = probably unsatisfactory
h 5 = definitely unsatisfactory
Competence to Stand Trial
Overall rating of the defendant’s rational and factual understanding
of the proceedings against him and his ability to consult rationally
with an attorney
h 1 = very likely competent
h 2 = probably competent
h 3 = uncertain
h 4 = probably incompetent
h 5 = very likely incompetent
Law Hum Behav (2010) 34:402–417 405
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
trade-offs between diagnostic sensitivity and specificity in a
diagnostic method from cost–benefit judgments that influ-
ence decisions based on diagnostic data (Obuchowski, 2003;
Swets, 1995; Zweig & Campbell, 1993). A ROC graph plots
sensitivity (the test’s true positive rate, tpr) as a function of a
test’s false positive rate (fpr, equal to 1—specificity); area
under the ROC curve (AUC) is a summary index of overall
diagnostic accuracy (Mossman & Somoza, 1991).
Statistical Model
If five raters provide ratings about 156 defendants’ adju-
dicative competence, one can set out the raters’ scores for
each defendant in a five-element row, or vector. Also, one
can array scores for all defendants in a matrix with
5 9 156 = 780 elements arranged in five columns (one
column for each rater) and 156 rows (one row vector for
each defendant).
Earlier, we said that asking whether a defendant is CST
is homologous to asking whether a person can do some-
thing, such as clear a particular hurdle. We also noted that
even when abilities needed to do various tasks lie along
continua, whether a particular person can do a particular
task may well be answerable with ‘‘yes’’ or ‘‘no.’’ For
purposes of exposition, we now concretize inability to do a
task—here, being incompetent to stand trial—as ‘‘having’’
a condition or disorder D.
Suppose, for a moment, that some ‘‘gold standard’’—
perhaps the declaration of an Omniscient Being—told us
whether a defendant (or ‘‘subject’’) had D, and we wanted
to describe how accurately a particular rater could detect D.
The raters have provided scores expressing their confidence
about presence or absence of D along a five-point scale;
this means that four (fpr, tpr) pairs, or four ‘‘cut-off’’
points, describe each rater’s accuracy in detecting D.
Table 2 contains results for a hypothetical rater and shows
how the rater’s scores translate into four (fpr, tpr) pairs.
Figure 1 illustrates how these four (fpr, tpr) pairs define a
ROC curve that summarizes the rater’s accuracy.
Note that, if an Omniscient Being declared which sub-
jects had D, we would know the prevalence of D in our
particular subject pool. We could also quantify any con-
ditional dependence in raters’ scores that arose from
peculiarities—such as whether the cases in that particular
pool are especially easy or hard—of the competent and
incompetent subjects. Thus, if we knew whether each
subject had D, we could calculate 43 parameters—four
(fpr, tpr) pairs for five raters (4 9 2 9 5 = 40 parame-
ters), plus the prevalence (one parameter), plus values
expressing conditional dependence of the competent and
incompetent subjects (two parameters)—directly from the
780-element matrix.
In reality, we have no truth criterion for the presence or
absence of D, only imperfect human opinion. One response
to this situation, reflected in previously published studies
(e.g., Cooper & Zapf, 2003; Murrie et al., 2008; Skeem,
Golding, Cohn, & Berge, 1998) would be to examine
agreement in and correlates of experts’ opinions about
CST. These studies have used statistics such as kappa and
the intraclass correlation coefficients (ICCs) to quantify
(dis)agreement between experts and to explore factors that
cause them to disagree.
The statistical approach we employ here aims at some-
thing different: quantifying the accuracy of experts, or,
more specifically, gauging how well they can distinguish
competent from incompetent defendants. If one pre-
sumes—despite the absence of a ‘‘gold standard’’—that
having or lacking D is something that experts can get right
or wrong, and that experts actually can distinguish com-
petent from incompetent defendants fairly well, one might
ask, ‘‘What combination of values for the 43 accuracy
parameters would maximize the probability of producing
this 780-element ratings matrix?’’ An answer to this
Table 2 Calculation of accuracy indices based on 156 hypothetical rating results if the actual CST status of evaluees were known
Rating Actual CST status fpr tpr
Competent Not
competent
1 = very likely competent 60 1
2 = probably competent 30 4 0.406 0.982
3 = uncertain 5 5 0.109 0.889
4 = probably incompetent 3 10 0.059 0.818
5 = very likely incompetent 3 35 0.030 0.636
Total 101 55
0
0.5
1
10.50
false positive rate
tr u
e p
o s
it iv
e r
a te
5
4
3 2
1
Fig. 1 ROC graph depicting a rater’s performance in detecting adjudicative competence, based on hypothetical data in Table 2.
Numbers correspond to rating categories in Table 1
406 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
question would provide the most plausible estimates of the
raters’ accuracy parameters, given the data available.
A Pictorial Explanation
We created Fig. 2a–c to help readers get an informal,
intuitive feel of when and why our statistical approach
might succeed. Suppose two fairly accurate forensic
examiners examine 150 defendants, half of whom—though
this is known only to an Omniscient Being—actually lack
adjudicative competence. Each examiner can make finely
graded judgments about CST that can be represented along
a continuous (or at least finely grained) mental decision
scale. In Fig. 2a, the judgments of Rater 2 are plotted as a
function of the judgments of Rater 1. For both examiners,
the impact of judging CST is to shift the distribution of
actually incompetent defendants by two standard devia-
tions along the mental decision axes. For Rater 1, the shift
takes place rightward (in the direction of the arrow) along
the horizontal axis in Fig. 2a, and for Rater 2, upward (in
the direction of the arrow) along the vertical axis. Another
way to think about these shifts is to say that the raters’
judgments about presence and absence of D have an effect
size of 2, which implies a ROC area of 0.92.
Figure 2a suggests that each examiner’s mental decision
scale might allow for virtually continuous judgments about
competence. However, the examiners have summarized
(per instructions) their judgments about adjudicative com-
petence as five-category ratings. Using a five-category
rating scale implies that their judgments have four non-
trivial dichotomous thresholds (= 5, C4, C3, and C2), and
each dichotomous threshold has associated with it a true
positive and false positive rate. Therefore, four (fpr, tpr)
pairs—eight parameters in all—summarize each exam-
iner’s accuracy. These (fpr, tpr) pairs appear numerically in
Fig. 2a and are also represented by the vertical and hori-
zontal dashed lines.
If the Omniscient Being provided the truth about each
defendant’s status, one could simply calculate the four (fpr,
tpr) pairs for Rater 1 or Rater 2 very easily, without ref-
erence to the other rater’s judgments. In reality, however,
we have no such truth. If one tried to calculate eight
parameters for a single examiner from just five ratings plus
a ninth parameter representing the fraction of D defendants
in the whole 150-subject group, a huge number of possible
(fpr, tpr) pairs would be possible. One would have no way
to distinguish the ‘‘best’’ combination of parameters
because there are more degrees of freedom (nine) than
rating categories (five).
If, however, one looks at both raters’ judgments simul-
taneously, one sees that their joint ratings produce 25
categories. The total number of accuracy parameters—four
(fpr, tpr) pairs for each rater—equals 16. If one adds
additional parameters for correlations in the D and non-D
subjects and for the fraction of D defendants in the entire
group, one obtains 19 parameters in all, implying 19
degrees of freedom. This number is smaller than the 25
joint categories formed by combining both sets of ratings.
Fig. 2 Two hypothetical raters’ continuous-scale judgments about adjudicative competence, with cut-offs (dashed lines) that result when raters group their judgments in five categories. Open circles Actually competent defendants; filled squares actually incompetent defendants. a Uncorrelated judgments, effect size = 2. b Uncorrelated judgments, effect size = 1.2. c Correlated judgments (r = 0.7), effect size = 2
Law Hum Behav (2010) 34:402–417 407
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
In theory, a mathematical search algorithm could ‘‘find’’ a
single set of values for 19 parameters with the maximum
likelihood of producing the 25 empirical category counts—
or alternatively, using a Bayesian procedure, a distribution
of values for the 19 parameters that was most plausible,
given the 25-category data.
To understand how the maximum likelihood estimation
(MLE) algorithm works, imagine asking a two-dimensional
creature to locate the highest elevation or ‘‘peak’’ of the
highest ‘‘mountain’’ on an irregular three-dimensional
surface. The creature cannot get outside the surface to look
and simply ‘‘see’’ where the peak is, but before moving to a
new location on the surface, the creature can take a ten-
tative step of any size in any direction and get information
about which step would place it at the highest elevation
relative to its current location. By a series of successive
elevation-maximizing steps, the creature would eventually
find a point such that a small step in any direction would
place it at an elevation that was lower than the present
location. This location would be the highest point on the
surface, unless the creature had unintentionally found a
local maximum (the ‘‘top’’ of a ‘‘hill,’’ but not the highest
mountain’s ‘‘peak’’).
For the situation depicted in Fig. 2a, the actual mathe-
matical search task involves looking for a peak on a
probability surface that lies above a 19-dimensional plane.
Yet looking at Fig. 2a, we sense that ‘‘finding’’ this prob-
ability peak might be relatively ‘‘easy’’ because the D and
non-D populations are relatively separated. A relatively
‘‘obvious’’ transition zone exists, which suggests that the
algorithm would find it relatively easy to ‘‘locate’’ the
transition zone, which it does by finding the set of (fpr, tpr)
pairs that best represents that zone. The algorithm would
not get ‘‘confused’’ by local maxima and would proceed
relatively directly to the set of values that represent the
locus of the probability peak.
If a transition zone is less obvious, however, an opti-
mization algorithm might have more difficulty locating it—
that is, finding the (fpr, tpr) pairs that best represent the
zone. Figure 2b and c depict two situations reasons why a
transition zone might be ambiguous. In Fig. 2b, the dis-
criminative ability of Raters 1 and 2 is equivalent to an
effect size of 1.2, or a ROC area of 0.80. In Fig. 2c, both
raters have the same discriminatory power as in Fig. 2a
(i.e., effect size = 2, ROC area = 0.92), but their judg-
ments are highly correlated (r = 0.7). In both cases,
distributions of the D and non-D subjects overlap much
more, the transition zone is less obvious, and an algorithm
might have more trouble locating the optimal (probability
maximizing or most plausible) values for the accuracy
parameters.
Our Bayesian approach to locating accuracy parameters
used Gibbs sampling implemented by WinBUGS (Lunn,
Thomas, Best, & Spiegelhalter, 2000) to find a posterior
distribution from which we could make inferences about
the parameters’ values. To understand the process, picture
a drunk individual who, though able to walk, is entirely
unable to walk in a continuously straight line. Instead, the
drunkard takes a random number of steps ahead or back,
then stops to rest; then he walks left or right for a random
number of steps before resting again. He alternates between
forward/backward and left/right random walking over and
over again. As the drunkard staggers, he is affected by
gravity, such that he is more likely to step downhill than
up, and his steps are likely to be smaller as he moves uphill
and larger as he stumbles downhill. Now imagine setting
the drunkard loose in a large, irregularly shaped ravine
with an uneven bottom. Every other time the drunkard
stops to rest, we record his location in the (x, y) plane.
When we are done, we will have a bivariate dot-plot of the
ravine, where the highest density of dots corresponds to the
lowest point in the ravine. This odd ‘‘topographical map’’ is
a sampling from the two-dimensional posterior distribution
for the plausible though unknown (x, y) coordinates of the
bottom of the ravine.
For the situation depicted in Fig. 2a, the Bayesian
drunkard staggers in 19 dimensions rather than just two,
and we record his locations as a sampling from a 19-
dimensional posterior distribution of the model parameters,
which are analogous to the two-dimensional coordinates
for the bottom of the ravine. In Fig. 2a, the D and non-D
populations are relatively separated and distinguishable.
This is like sending the drunkard out to stagger in a steeply
sloped ravine; the steep slope, coupled with gravity, will
help the drunkard quickly get to the low point. In the case
of the 19-dimensional problem posed in Fig. 2a, the
‘‘drunkard’’—the Gibbs sampling process—will quickly
stagger to a clear posterior distribution for the (fpr, tpr)
pairs and other parameters that locate the transition zone.
As was the case with the MLE algorithm, however, situa-
tions in which the transition zone is less obvious—e.g.,
those shown in Fig. 2b and c—may give the Gibbs sam-
pling algorithm a less clearly sloped surface on which to
stagger toward the (fpr, tpr) pairs that best define the
transition zone.
The preceding informal description used two raters
(rather than five) because doing so allows for easy repre-
sentation in two-dimensional figures. The task in our study
was to use a 780-category matrix to identify 43 parameters.
The Appendix formally describes the statistical model and
methods we used to make inferences about empirical class
membership (i.e., about being competent or incompetent)
based on our five raters’ judgments. We implemented the
model under conditional dependence (CD) and conditional
independence (CI) assumptions. We then evaluated accu-
racy in assessing CST with both MLE and Bayesian
408 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
methods based on the Dusky competence scores that raters’
assigned after examining the 156 reports. We also evalu-
ated raters’ scores on understanding, reasoning, and
appreciation (URA) by treating these values as diagnostic
‘‘tests’’ for incompetence. In addition, we examined the
sum of URA ratings as a diagnostic test under the CI
model. Finally, we transformed URA scores (x0 = 5 - x) so that 0 implied highest likelihood of incompetence, and
evaluated the diagnostic accuracy of the product of these
transformed scores under the CI model. (Note that the
transformation causes a ‘‘5’’ score on any URA rating to
yield a ‘‘0’’ product.)
RESULTS
Evaluee Characteristics
Table 3 summarizes the demographic and diagnostic
characteristics of the 156 defendants whose court docu-
ments served as sources for our sanitized reports. Eighty-
six (55%) of the defendants had primary or co-morbid
substance use disorders, but because the patients had been
confined for substantial periods before evaluation, current
or recent intoxication did not influence their clinical pre-
sentation when they underwent evaluation.
Raters’ Performance
Results of the accuracy analyses appear in Table 4. Here,
AUC equals the probability that a rater, examining one
randomly chosen competent defendant and one randomly
chosen incompetent defendant, would assign a higher score
(i.e., a score more indicative of incompetence) to the
actually incompetent subject. An AUC of 1.0 would imply
perfect sorting, and an AUC of 0.5 would imply no-better-
than-chance discrimination between competent and incom-
petent defendants.
The average accuracy of raters’ assessments of Dusky
competence, calculated with MLE or Bayesian methods
under CD or CI assumptions, was at least 0.967, which
suggests they could correctly distinguish a randomly
selected competent defendant from a randomly selected
incompetent defendant in 29 out of 30 attempts. Raters’
ability to evaluate adjudicative competence from the san-
itized reports thus was comparable to accuracy achieved in
using advanced positron emission imaging methods to
detect Alzheimer’s disease (Small et al., 2006). We had
hypothesized that the sum or product of URA scores might
yield greater accuracy than global judgments of compe-
tence. This did not happen, both because raters’ URA
scores tended to yield lower accuracy than their
assessments of Dusky competence and because raters’
assessment using Dusky criteria left little room for
improvement.
One can compare results from various methods and
models using the Akaike (1974) information criterion,
calculated as AIC = - 2 ln L ? 2p, where p is the number
of model parameters. In calculating the AIC, the superior
fit (measured by the -2 ln L term) expected from adducing
additional parameters is offset by a penalty (the 2p term).
Minimum AIC thus can serve as a basis for choosing, from
among several models with different numbers of parame-
ters, a model that the data best support. In Table 4, AIC
values are substantially smaller for CD models than for CI
models, and substantially smaller for Bayesian estimates
than the MLE estimates. Though overall accuracy (mea-
sured by AUCs) is comparable, the MLE and Bayesian
Table 3 Demographic and diagnostic characteristics of the original 156 defendants
Age (years)
Mean ± SD 37.4 ± 12.6
Range 18.2–84.9
Race
African-American 79
Caucasian 77
Sex
Female 19
Male 137
Original opinion on CST a
Competent 101
Not competent 55
Primary diagnoses b
Schizophrenia 53
Bipolar disorder 23
Depressive disorders 7
Substance use disorders 8
Malingering 12
Delusional disorders 3
Adjustments disorders 4
Schizoaffective disorders 24
Psychotic disorder NOS 14
Mental retardation 2
Neurocognitive disorders 2
Impulse control disorders 2
Pedophilia 2
Comorbid conditions
Substance use disorders 78
Mental retardation 10
a Opinion concerning competence to stand trial provided by original
report’s author b
Hospital’s final DSM-IV diagnosis, based on all available clinical information, that best accounted for defendant’s hospital stay
Law Hum Behav (2010) 34:402–417 409
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
estimates of the prevalence of incompetent individuals
often differed. The Bayesian prevalence estimates for
incompetence also tended to be lower than the 35% of
defendants so identified by the original evaluators.
Though AUC is a useful summary index of accuracy,
individual operating points can also be informative.
Figures 3 and 4 depict MLE and Bayesian results for raters’
performance in detecting Dusky incompetence under the
CD model. Comparing these figures suggests that a different
ROC for Rater 4 explains some of the differences found in
Table 4. Using the MLE-CD results from our study, ratings
of ‘‘5’’ on Dusky competence were associated with an
average sensitivity (in detecting incompetence) of 0.828
and an average specificity of 0.977. Using the Bayesian-CD
Table 4 ROC areas for five raters based on MLE and Bayesian estimates of accuracy parameters
Estimation
method
Criterion, model ROC area (MLE standard error or Bayesian posterior standard deviation) prev logL AIC
Rater 1 Rater 2 Rater 3 Rater 4 Rater 5
MLE Dusky, CI 0.974 (0.016) 0.966 (0.018) 0.941 (0.023) 0.958 (0.020) 1.000 (0.000) 0.333 -816.6 1715.3
Dusky, CD 0.984 (0.012) 0.965 (0.018) 0.955 (0.021) 0.942 (0.023) 0.993 (0.008) 0.339 -739.2 1564.3
Understanding, CI 0.925 (0.025) 0.959 (0.019) 0.945 (0.022) 0.957 (0.019) 0.972 (0.016) 0.375 -816.6 1715.3
Understanding, CD 0.801 (0.038) 0.917 (0.025) 0.783 (0.039) 0.853 (0.033) 0.907 (0.027) 0.405 -727.2 1540.4
Reasoning, CI 0.974 (0.016) 0.954 (0.021) 0.967 (0.018) 0.977 (0.015) 0.997 (0.006) 0.314 -859.0 1799.9
Reasoning, CD 0.978 (0.015) 0.962 (0.019) 0.959 (0.020) 0.985 (0.012) 0.998 (0.004) 0.323 -785.4 1656.8
Appreciation, CI 0.970 (0.017) 0.957 (0.020) 0.946 (0.023) 0.973 (0.016) 0.990 (0.010) 0.319 -863.8 1809.5
Appreciation, CD 0.902 (0.031) 0.947 (0.024) 0.930 (0.027) 0.950 (0.023) 0.997 (0.006) 0.297 -800.5 1687.0
Sum, CI 0.962 (0.018) 0.973 (0.015) 0.958 (0.019) 0.977 (0.014) 0.982 (0.012) 0.378 -1373.2 2988.3
Product, CI 0.956 (0.019) 0.943 (0.021) 0.957 (0.019) 0.964 (0.017) 0.985 (0.011) 0.398 -1311.3 2944.6
WinBUGS Dusky, CI 0.967 (0.015) 0.960 (0.016) 0.937 (0.022) 0.950 (0.019) 0.991 (0.009) 0.334 -737.5 1557.0
Dusky, CD 0.975 (0.017) 0.975 (0.016) 0.963 (0.022) 0.968 (0.017) 0.985 (0.014) 0.276 -601.0 1288.0
Understanding, CI 0.948 (0.023) 0.968 (0.016) 0.932 (0.024) 0.957 (0.016) 0.960 (0.018) 0.330 -734.5 1555.0
Understanding, CD 0.971 (0.019) 0.976 (0.016) 0.940 (0.028) 0.970 (0.019) 0.978 (0.018) 0.267 -594.5 1275.0
Reasoning, CI 0.963 (0.017) 0.949 (0.019) 0.951 (0.019) 0.966 (0.017) 0.989 (0.009) 0.320 -782.0 1646.0
Reasoning, CD 0.973 (0.017) 0.962 (0.021) 0.955 (0.021) 0.978 (0.015) 0.987 (0.011) 0.286 -641.5 1369.0
Appreciation, CI 0.962 (0.016) 0.950 (0.019) 0.936 (0.024) 0.961 (0.019) 0.980 (0.012) 0.328 -786.5 1655.0
Appreciation, CD 0.952 (0.021) 0.960 (0.022) 0.950 (0.025) 0.980 (0.017) 0.980 (0.015) 0.276 -606.5 1389.0
prev sample prevalence, logL loge likelihood, AIC Akaike Information Criterion, MLE maximum likelihood estimation, CI conditional inde- pendence assumption, CD condition dependence assumption
0
0.5
1
10.50 false positive rate
tr u
e p
o s it
iv e r
a te
Rater 1
Rater 2
Rater 3
Rater 4
Rater 5
Fig. 3 ROC graph depicting maximum likelihood estimates of raters’ performance in detecting adjudicative incompetence (applying the
Dusky standard), under conditional dependence assumptions
0
0.5
1
10.50
false positive rate
tr u
e p
o s it
iv e r
a te
Rater 1
Rater 2
Rater 3
Rater 4
Rater 5
Fig. 4 ROC graph depicting Bayesian (WinBUGS) estimates of raters’ performance in detecting adjudicative incompetence (applying
the Dusky standard), under conditional dependence assumptions
410 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
results, the average sensitivity for scores of ‘‘5’’ was 0.761,
and the average specificity was 0.991.
As the high accuracy of raters suggests, the raters’
results were consistent with and highly correlated with
each other. On Dusky competence, for example, correla-
tions for the ten possible pairs of raters ranged between
0.757 and 0.859. Table 5 reports values of ICC(A,1) (i.e.,
single rater, absolute agreement; see McGraw & Wong,
1996) and ICC(C,1) (single rater, consistency) for all four
sets of ratings. These findings offer numerical expressions
of what Figs. 3 and 4 illustrate. In these two figures, we see
that each rater’s ROC bends at a different location in the
ROC square, implying that raters’ operating points are not
identical. Figure 3, for example, shows that the boundaries
between Dusky competence scores of ‘‘4’’ (‘‘probably
incompetent’’), ‘‘3’’ (‘‘uncertain’’), and ‘‘2’’ (‘‘probably
competent’’) for Rater 3 are roughly the same location as
the boundary between scores of ‘‘5’’ (‘‘very likely incom-
petent’’) and 4 (‘‘probably incompetent’’) for Rater 2.
Raters 2 and 3 earned similar MLE estimates of overall
accuracy—their ROC areas are 0.965 and 0.955, respec-
tively (see Table 4)—but Rater 3 gave scores of ‘‘4’’ and
‘‘3’’ to several defendants whom Rater 2 scored ‘‘5.’’
For the MLE results, the estimated dependence param-
eters for the random effects model were r̂0 ¼ 1:53 (for the non-D subgroup) and r̂1 ¼ 1:01 (for the D subgroup); in the Bayesian-estimated random effects model, the non-D
and D dependence parameters were r̂0 ¼ 1:73 and r̂1 ¼ 0:406, respectively. Both pairs of results suggest substantial conditional dependence. Following Albert
(2007) and Qu, Tan, and Kutner (1996), we examined the
‘‘observed minus expected correlations’’ for each of the ten
rater pairs, that is, the differences between the actual
pairwise correlations of ratings minus the correlation
that would be expected based on the model’s parameters.
Figure 5 shows a ‘‘diagnostic plot’’ comparing these dif-
ferences for the MLE (dashed lines) and Bayesian (solid
lines) estimates under conditional independence and con-
ditional dependence (random effects) models. Generally,
the differences for the random effects models (open sym-
bols) are lower (closer to zero) than the differences for the
conditional independence models, which suggests (as do
the AIC values in Table 4) that the random effects models
provide better descriptions of the pairwise correlations than
do the conditional independence models.
DISCUSSION
Over the last two decades, ROC analysis has become a
popular technique for quantifying the accuracy of forensic
mental health assessments. In such applications, humans’
judgments based on available data—e.g., about whether
violence occurred (Steadman et al., 1998; Monahan et al.,
2006), about whether malingering has occurred (Miller,
2005), or about whether an evaluee is competent (Kim et al.,
2007)—provide the truth criteria or ‘‘gold standards’’ against
which investigators have judged accuracy. But given the
central role that psychiatrists and psychologists play in psy-
cholegal determinations, the accuracy of the professionals
themselves is a matter of major legal and social significance.
To our knowledge, ours is the first study to use latent
class models and the ROC analytic concepts discussed by
Mossman (2008) to quantify the accuracy of forensic
experts. Our study shows that raters can assess CST on a
graded scale (rather than simply providing binary opinions,
as mental health experts customarily do) and that such
ratings can lead to estimated ROC parameters of detection
accuracy despite there being no ‘‘gold standard’’ for CST.
Insofar as our chief aims were to find out whether clini-
cians could provide graded judgments and whether our
statistical approach was feasible, our study met its goals.
Our findings also suggest that mental health experts’
intrinsic ability to discriminate between competent and
Table 5 Intraclass correlation coefficients for five raters’ judgments concerning understanding, reasoning, appreciation, and Dusky competence
Rating ICC(A,1) ICC(C,1)
Dusky competence 0.7946 0.9508
Understanding 0.8015 0.9528
Reasoning 0.7931 0.9504
Appreciation 0.7664 0.9425
ICC(A,1) single rater, absolute agreement ICC, ICC(C,1) single rater, consistency
0
0.02
0.04
0.06
0.08
0.1
0.12
0.14
0.16
1,2 1,3 1,4 1,5 2,3 2,4 2,5 3,4 3,5 4,5
Rater pairs
o b
s e rv
e d
- e
x p
e c te
d c
o rr
e la
ti o
n
MLE Conditional Independence
MLE random effects
WinBUGS Conditional Independence
WinBUGS random effects
Fig. 5 Observed minus expected correlations under conditional independence and conditional dependence assumptions
Law Hum Behav (2010) 34:402–417 411
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
incompetent defendants is high (though not perfect).
Readers should recognize, however, that several features of
our study limit the generalizability of this conclusion.
First, our raters based their judgments on written reports
rather than on personal examinations—a typical require-
ment for a forensic assessment (American Academy of
Psychiatry and the Law, 2005). We have no reason to
believe that we introduced systematic bias by excluding
poorly written reports from our sample, but we cannot be
sure that the exclusions altered our ultimate sample so as to
make the diagnostic task easier (or harder) for our raters.
Using ‘‘sanitized’’ reports was time-efficient for raters
and protected defendants’ confidentiality. However, written
reports condense information that examiners obtain during
personal interviews (which may have the effect of reducing
rater accuracy). The original reports’ text may have
selectively included data consistent with original author-
examiners’ CST opinions, while excluding data that con-
flicted with their opinions (which could have the statistical
effect of inflating our raters’ apparent accuracy). Also,
having raters base judgments on data provided by someone
else does not permit evaluation of the raters’ ability to
gather data and recognize its importance, which is an
essential feature of evaluative accuracy.
Of course, one could never replicate typical circum-
stances of adjudicative competence evaluations for a study
such as ours. Attempting to do this would require having
multiple examiners conduct in-person, legally superfluous,
closely spaced evaluations of dozens of defendants. Even if
doing this were practicable, multiple evaluations would
themselves affect the evaluees and distort the resulting
findings. The process would also confront serious ethical
questions concerning evaluees who would not be compe-
tent to consent to the research, but whose participation
would be necessary to reach a meaningful judgment about
the accuracy of CST determinations.
However, because our study has shown that our statis-
tical approaches are workable, it may be reasonable to
conduct future studies that better replicate the data obtained
in actual CST evaluations. For example, a future study
might evaluate raters’ accuracy when they base their CST
judgments on documents prepared using highly systema-
tized formats for data collection and written presentation.
Also, the likely prospect of generating findings with high
social and legal significance might make it easier to justify
a study like ours in which multiple raters based their
judgments on actual, videotaped CST assessments (though
obtaining consent from defendants with severe mental
impairments would remain problematic).
A second limit on generalizability stems from the ori-
ginal reports’ context—evaluations following court-
ordered hospitalizations for CST assessment or restoration.
Such hospitalizations usually give evaluators substantial
time to assemble background information and generate
copious observational data that help to clarify psychiatric
diagnosis. Acute effects of intoxicants abate during hos-
pitalization, and for those patient-defendants who accept
treatment, examiners can take into account how medication
and psychotherapy have affected evaluees’ functioning and
clinical presentation. All these circumstances probably
made our raters’ tasks easier—and their assessments more
accurate—than would be the case when an examiner
encounters a defendant during an initial court-ordered
evaluation of CST.
A third limitation relates to how our findings apply to
the binary, competent-or-incompetent opinions often
required by courts or statutes. Our study’s five-category
scorings helped us focus specifically on raters’ intrinsic
ability to assess adjudicative competence. When examiners
provide yes-or-no opinions, however, those opinions
incorporate examiner biases, including examiners’ feelings
about the relative undesirability of false-positive and false-
negative errors. If these value judgments translate into
disagreement about the ultimate yes-or-no conclusion, they
will lower apparent accuracy.
To see why this is the case, consider what Fig. 4 sug-
gests about defendants evaluated by Raters 1, 2, and 4. In
Table 4, the Bayesian CD-model ROC areas for these
raters on Dusky competence are very similar (implying
near-identical diagnostic accuracy), and for all three raters,
Fig. 4 locates an operating point at roughly (fpr,
tpr) = (0.1, 0.95). For Rater 1, however, this area marks
the boundary between ratings of ‘‘3’’ (‘‘uncertain’’) and
‘‘2’’ (‘‘probably competent’’), while for Raters 2 and 4, the
boundary is between 4 (‘‘probably incompetent’’) and 3
(‘‘uncertain’’). This implies that Rater 1 will score a few
defendants ‘‘uncertain’’ whom Raters 2 and 4 think are
probably incompetent. If Rater 1 believes ‘‘when uncertain,
presume competent,’’ then he will disagree with Raters 2
and 4 if they believe that probably incompetence should
generate a yes-no opinion of incompetence. In such a case,
at least one rater would be ‘‘wrong’’ about the defendant,
and his apparent accuracy based on just the binary rating
would fall. To understand another source of inaccuracy
caused by binary ratings, imagine that all five of our raters
thought a particular report described a ‘‘probably compe-
tent’’ defendant, but—because they were required to
provide binary opinions—they disagreed in their value
judgment about whether that marginally competent
defendant should face trial. Here, at least one rater would
have been ‘‘wrong’’ about the defendant, and his apparent
accuracy would have fallen.
What this suggests is that experts may well be better
discriminators of competence than one would conclude
from CST evaluations undertaken in ‘‘real life’’ criminal
cases, because some disagreements may arise from
412 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
different value judgments rather than from different opin-
ions about clinical findings relevant to CST. Indeed, recent
studies by Murrie and colleagues (Boccaccini et al., 2008;
Murrie et al., 2008, 2009) suggest that individual exam-
iners’ biases explain a large amount of the variance in their
ultimate opinions about CST or the scores they report when
using actuarial risk assessment instruments. Findings from
these studies, which use ICC(A,1) (the absolute, single-
rater intraclass correlation coefficient) as their chief ana-
lytical statistic, complement this study’s ROC-based
findings. The difference is that these previous studies are
agnostic about whether forensic evaluees have any true
diagnostic status—that is, these studies have examined
degrees of absolute diagnostic agreement between evalu-
ators and possible sources of disagreement, but have not
asked how well forensic examiners perform or how often
they might get the answer ‘‘right.’’ In contrast, our study
assumes—as do mental health experts and courts—that
while defendants display varying degrees of the abilities
that underlie adjudicative competence, they ultimately are
either competent to stand trial or not. Our study then
attempts to quantify and characterize individual evaluators’
ability to distinguish between these two (presumptively
mutually exclusive) forensic options, despite having no
gold standard for defendants’ true status. Our study
accomplishes this by focusing on and quantifying evaluator
accuracy using appropriate statistics (i.e., ROC indices),
rather than on measuring inter-rater agreement using sta-
tistics (e.g., ICCs) that are appropriate to that task.
Our findings may reflect statistical limitations arising
from sampling error; we report here on the performance of
just five psychiatrists who read a particular set of redacted
materials based on reports generated at just one institution.
Though our findings reflect a reasonable statistical model
for our data, we recognize that other plausible models exist
and might generate somewhat different outcomes. Also, as
Uebersax (1988) notes, latent class modeling provides
upper bounds for accuracy under certain conditions. LCM
chooses underlying classes that minimize error rates
defined within the model, but these error-minimizing,
empirically generated latent classes can differ from the true
classes when probabilities of the empirical classes depend
on covariates. Knowing whether this actually has occurred
is difficult to ascertain (Spencer, 2008), but it is a limitation
that we must acknowledge.
Our findings may also reflect limitations due to mis-
taking reliability for validity and to related problems of
undefined ontology. As the ‘‘Background’’ section
explains, we assumed that assessing CST involves evalu-
ating abilities needed to perform a task; we also assumed
that raters’ Dusky-guided notions about CST reflected valid
conceptions of those abilities. A problem with our statis-
tical methods is that if raters all used a very reliable but
irrelevant method for assessing competence (e.g., the
length of a subject’s last name), they might appear very
accurate despite their really having no-better-than-chance
accuracy. Our response is, simply, that our raters applied
the same definitions and ideas about CST that mental
health experts and courts regularly use in legal decision
making—that is, the same definitions and ideas that experts
and courts regularly accept as being valid.
We also did not explore whether CST is a dimensional
construct or whether it admits of a valid dichotomy, which
are typical uses of latent class methods when diagnostic
validity is questionable. We think that forensic psychiatry
and psychology might well benefit from explorations of
empirically based, natural taxa that are relevant to CST,
but our study leaves such explorations to others. We note,
however, that even if the results of such efforts were
known, natural taxa might not correctly track or mean-
ingfully distinguish between individuals who are and are
not competent to stand trial; it is just as reasonable to
suppose that some taxa would overlap the competent-
incompetent boundary. Whatever a CST taxonomy might
tell us, we think it is reasonable to gather data that lets us
characterize raters’ accuracy about a distinction (between
competent and incompetent) that is imperfectly understood,
but that everyone usually takes to be genuine.
A final limitation arises from our raters’ neutrality.
Numerous studies show that framing and anchoring of
information affect decision-makers’ judgments, even when
the decision-makers receive instruction about such effects
or have incentives to be accurate (Cain & Detsky, 2008).
Despite attempts by forensic consultants to remain objec-
tive, a variety of conscious or unconscious factors—
identification with or desire to please the retaining party, or
sympathy for or antipathy toward an evaluee, or being
prosecution- or defense-oriented—influence forensic opin-
ions (Boccaccini et al., 2008; Gutheil, 2004; Murrie et al.,
2008, 2009). In the United States, most mental health
opinions about CST are accepted by courts without dispute
(Zapf, Hubbard, Cooper, Wheeles, & Ronan, 2004). When
second opinions are sought, however, it is often because the
defense or the prosecution has disagreed with the first
opinion’s conclusion. This non-random referral pattern may
inflate apparent expert disagreement above what one would
find if experts and cases were chosen at random. Such
selection patterns may encourage or induce experts to dis-
agree with each other (Murrie et al., 2009), and ironically,
acknowledging biasing potential may actually exacerbate
the problem (Cain, Loewenstein, & Moore, 2005). In our
study, raters were ignorant of their collaborators’ opinions,
and because they could offer graded ratings about CST, they
did not have to reach yes-or-no conclusions about ambig-
uous cases. Arguably, our raters performed their evaluation
tasks in circumstances more conducive to being accurate
Law Hum Behav (2010) 34:402–417 413
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
than the circumstances in which opposing forensic evalua-
tors usually find themselves.
Though our study’s primary aim was to test whether
forensic opinions could be conceptualized so as to allow
accuracy characterization using latent class methods, we
believe our endeavor’s success and our numerical findings
have some immediate practical implications.
Having shown that forensic examiners differ greatly in how
often they think criminal defendants are incompetent, Murrie
and colleagues suggest that future research might aim to
identify the ‘‘thresholds’’ at which clinicians consider
a defendant to be incompetent to stand trial (IST).
Although the legal determination regarding compe-
tence is dichotomous (i.e., competent or not
competent), it seems reasonable to think of the
capacities underlying trial competence as dimen-
sional; in other words, some defendants are ‘‘more
competent’’ and others ‘‘less competent.’’ If these
capacities are dimensional, we might not be surprised
to find that clinicians sometimes draw the distinction
between competent and incompetent at different
points along the continuum. We also might not be
surprised to find some disagreement among clinicians
regarding cases that fall toward the midpoint of this
continuum. (Murrie et al., 2008, p. 190)
As expressed by Murrie and colleagues, the continuum
(or continua) along which clinicians’ thresholds might lie is
undefined. Our study, however, used the raters’ judgments
about CST itself as the continuum. An implication of our
approach is the finding, illustrated in Figs. 3 and 4, that
raters’ thresholds for deeming defendants ‘‘probably
competent,’’ of ‘‘uncertain’’ competence, or ‘‘probably
incompetent’’ can vary enough to generate practical dis-
agreement, even though raters are very accurate. Our study
also suggests that disagreement in a dichotomous judgment
about CST may not necessarily stem from examiners’ dif-
ferences in judgments about defendants’ abilities. Even
when examiners agree about what specific defendants can
and cannot do relative to standing trial, they may disagree
about whether marginally capable defendants are really fit
to face serious criminal charges.
A second implication is that forensic practitioners
should be modest. Our study looked at raters making
‘‘fairly simple determinations’’ (Murrie et al., 2008, p. 180)
about adjudicative competence in a non-adversarial con-
text. Our raters also worked under conditions that gave
them optimal access to background data, and the data raters
used included information about treatment response fol-
lowing extensive inpatient treatment episodes. Though our
raters were very accurate, they still had varying levels of
clarity about their judgments and sometimes disagreed with
each other. This finding should tell us forensic practitioners
who work and reach conclusions in less ideal circum-
stances that we can be quite accurate yet fallible, and that
our colleagues will disagree with us for reasons that do not
undermine the legitimacy of their or our conclusions.
A final implication is for fellow investigators. Our study
demonstrates a practicable method for using ‘‘ROC anal-
ysis without truth’’ (Henkelman et al., 1990) to estimate the
accuracy of forensic assessments. The ideas and techniques
we describe may be applied to a host of other psycholegal
determinations where no diagnostic gold standard exists,
but where quantification of accuracy would have scientific
and evidentiary value. Many readers of this article can
create or already have available data that would be ame-
nable to the methods described in this article. We hope
other investigators will view our work as something to
improve upon and as inspiration for asking—and answer-
ing—many other questions about forensic assessments and
the quality of mental health expertise.
APPENDIX
Adopting the notation used by Albert (2007), suppose I
subjects (i = 1, 2, …, I) undergo assessment by J raters (j = 1, 2, …, J), who assign ordinal ratings k = 1, 2, …, K to each subject. Without loss of generality, let rating k = 1
indicate lowest confidence and k = K indicate highest
confidence that a subject has the condition or disorder D of
interest (here, incompetence to stand trial). Let Yi = (Yi1,
Yi2, …, YiJ)0 be a vector representing ratings made by the J raters for the ith subject. Because J raters could each assign
one of K ratings to each subject, each Yi has J 9 K possible
combinations of elements. The joint distribution of Yi,
expressed as P(Yi), the probability of Yi, is
PðYiÞ¼ PðYijdi ¼ 1ÞPðdi ¼ 1Þþ PðYijdi ¼ 0ÞPðdi ¼ 0Þ; ð1Þ
where di = 1 means the ith subject has condition D, di = 0
means the ith subject does not have D, P(di = 1) is
the probability or prevalence of D, and P(di = 0) =
1 - P(di = 1).
We would like to model P(Yi|di) so as to include pos-
sible conditional dependence (CD) of ratings, i.e.,
similarity in raters’ responses attributable to specific
characteristics of subjects besides their membership in the
D or non-D subgroups that affect how easy or hard their
particular cases are. Following Albert, we utilize a probit
link function for the parameterization,
U�1 P Yij �kjdi; bdi;i � �� �
¼ Cdi;k;j þ bdi;i ð2Þ
where U is the cumulative standard normal distribution function and U-1 is its inverse, Cdi;k;j are monotonically
414 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
increasing cut-offs for the jth rater, and bdi;i is a random
effect attributable to each subject that characterizes
conditional dependence in multiple ratings of that
subject.
Notice that bdi;i depends on the latent class of each
subject—i.e., whether the individual does or does not have
D. Following Albert (2007) and Qu et al. (1996), we used
the random effect model bdi;i ¼ rdi bi, where bi has a standard normal distribution. Equation 2 thus says that cut-
off points demarcating each rater’s classification thresholds
reflect the presence (di = 1) or absence (di = 0) of D.
However, the probability that the jth rater will assign rating
k to the ith subject reflects the ith subject’s state (di = 0 or
di = 1), locations of the rater’s particular cut-offs, and
peculiarities of the ith subject (which act in common across
all raters). We characterize the random effects of the D and
non-D populations separately because their cut-offs are not
linked (as they would be under the ‘‘binormal’’ ROC
model; see Somoza & Mossman, 1991).
In our data set, I = 156, J = 5, and K = 5. Thus, in our
CD model, I 9 J ratings (J raters evaluating I subjects)
arise from 2(K - 1)J ? 3 = 43 parameters: K - 1 cut-
offs for the D subgroup and K - 1 cut-offs for the non-D
subgroup for each rater, plus a random effect attributable to
each subgroup, plus the prevalence P(di = 1) in the rating
set. We sought values for the 43 parameters that would, in
combination, be most likely to have generated the 5 9 156
rating matrix. We could then construct individual raters’
ROC graphs using (fpr, tpr) coordinates computed as
follows:
fprj;k ¼ 1 � U C0;kffiffiffiffiffiffiffiffiffiffiffiffiffi 1 þ r20
p
!
; tprj;k ¼ 1 � U C1;kffiffiffiffiffiffiffiffiffiffiffiffiffi 1 þ r21
p
!
ð3Þ
Under a conditional independence (CI) assumption (equiv-
alent to setting bdi;i ¼ 0), ROC graphs for the five raters could be constructed from estimates of 2(K - 1)J ? 1 =
41 parameters.
We estimated the model’s accuracy parameters in two
ways. The first approach, standard maximum likelihood
estimation (MLE), used GAUSS 3.6 code kindly fur-
nished by Albert and modified for our data. (The modified
code, which calls the GAUSS 4.0 maxlik library’s quasi-
Newton BFGS optimization algorithm, is available from
the first author.) The natural logarithm of the likelihood
function, ln L ¼ PI
i¼1 ln Li, is (slightly modifying Albert’s notation)
ln L ¼ XK
i1¼1
XK
i2¼1 . . . XK
iJ¼1 IfYi¼ði1;i2;...;iJÞg
� ln P Yi ¼ i1; i2; . . .; iJð Þð Þf g; ð4Þ
with P(Yi) given by Eq. 1. Knowing (fpr, tpr) coordinates
for each rater permitted computation of ‘‘trapezoidal’’
AUCs as overall measures of rater accuracy, with standard
errors computed using the method of Hanley and McNeil
(1982).
In contrast to MLE, which provides point estimates of
the parameter values most likely to have generated the
observed data, Bayesian estimation summarizes knowledge
of unknown parameters using ‘‘posterior’’ distributions
representing the probability that a parameter has a particular
value, given the observed data. According to Bayes’ Rule,
the posterior probability of a parameter’s value is propor-
tional to the likelihood of observing the data given that
parameter value, multiplied by a ‘‘prior’’ probability of the
parameter’s value. The likelihood function is dictated by
statistical model choice, and is the same construct as in
MLE. When a prior is ‘‘non-informative’’ (e.g., P(h) = c for all h [ [a,b], where [a,b] is an arbitrarily large bounded interval, c is a constant, and h is a parameter), Bayesian and MLE methods yield similar inferences (Carlin & Louis,
2000). However, in Bayesian estimation, inference is con-
ducted directly on the unknown parameters (or functions
thereof, such as AUC), while in MLE, inference is con-
ducted on the data. Hence, only Bayesian estimation allows
direct probability statements such as ‘‘the probability that
the AUC for rater j is between .955 and .973 is 95%.’’
Markov chain Monte Carlo (MCMC) methods (Gelfand
& Smith, 1990; Geman & Geman, 1984; Metropolis,
Rosenbluth, Rosenbluth, Teller, & Teller, 1953) are used to
make inferences on posterior distributions for which lack
of analytic methods would make Bayes’ Rule intractable.
Under mild regularity conditions, a Markov chain con-
verges to a unique invariant or ‘‘target’’ distribution. To use
MCMC methods for Bayesian analysis, one constructs the
transition kernel so that the target distribution of the
resulting Markov chain will be the joint posterior distri-
bution of interest. After discarding input from initial ‘‘burn-
in’’ iterations, one can use the remaining draws to make
inferences about model parameters. WinBUGS is a free
software package that allows specification of a Bayesian
model, determines the transition kernel for the Markov
chain, and produces draws from the joint posterior distri-
bution of unknown parameters (Lunn et al., 2000).
For our Bayesian analyses, we used minimally infor-
mative priors and the same statistical models as in our
MLE approach. WinBUGS 1.4.3 ran five parallel MCMC
chains; the Brooks–Gelman–Rubin diagnostic (Brooks &
Gelman, 1998) indicated convergence after 2000–5000
iterations. We ran each chain for 15,000 iterations and
treated each chain’s first 10,000 iterations as ‘‘burn-in’’
values to be discarded, leaving 5 9 5000 = 25,000 draws
for inference. Because our WinBUGS code (available from
Law Hum Behav (2010) 34:402–417 415
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
the first author upon request) calculated (fpr, tpr) coordi-
nates and trapezoidal AUCs directly from the MCMC
parameter draws, we obtained samples of and made
inferences about our accuracy statistics directly from the
posterior distributions.
REFERENCES
Akaike, H. (1974). A new look at the statistical model identification.
IEEE Transactions on Automatic Control, 19, 716–723. Akinkunmi, A. A. (2002). The MacArthur Competence Assessment
Tool Fitness to Plead: A preliminary evaluation of a research
instrument for assessing fitness to plead in England and Wales.
Journal of the American Academy of Psychiatry and the Law, 30, 476–482.
Albert, P. S. (2007). Random effects modeling approaches for
estimating ROC curves from repeated ordinal tests without a
gold standard. Biometrics, 63, 593–602. American Academy of Psychiatry and the Law. (May 2005). Ethics
guidelines for the practice of forensic psychiatry. http://www. aapl.org/ethics.htm. Accessed 19 Sept 2008.
Bennett, G. (1985). A guided tour through selected ABA standards
relating to incompetence to stand trial: Incompetence to stand
trial. George Washington Law Review, 53, 375–413. Berg, W. A., Blume, J. D., Cormack, J. B., Mendelson, E. B., Lehrer,
D., Böhm-Vélez, M., et al. (2008). Combined screening with
ultrasound and mammography vs mammography alone in
women at elevated risk of breast cancer. Journal of the American Medical Association, 299, 2151–2163.
Boccaccini, M. T., Turner, D., & Murrie, D. C. (2008). Do some
evaluators report consistently higher or lower psychopathy
scores than others? Findings from a statewide sample of sexually
violent predator evaluations. Psychology, Public Policy, and Law, 14, 262–283.
Bonnie, R. J. (1990). The competence of criminal defendants with
mental retardation to participate in their own defense. Journal of Criminal Law and Criminology, 81, 419–446.
Brooks, S. P., & Gelman, A. (1998). Alternative methods for
monitoring convergence of iterative simulations. Journal of Computational and Graphical Statistics, 7, 434–455.
Buchanan, A. (2006). Competency to stand trial and the seriousness
of the charge. Journal of the American Academy of Psychiatry and the Law, 34, 458–465.
Cain, D. M., & Detsky, A. S. (2008). Everyone’s a little bit biased
(even physicians). Journal of the American Medical Association, 299, 2893–2895.
Cain, D. M., Loewenstein, G., & Moore, D. A. (2005). The dirt on
coming clean: Perverse effects of disclosing conflicts of interest.
Journal of Legal Studies, 34, 1–25. Carlin, B. P., & Louis, T. A. (2000). Bayes and empirical Bayes
methods for data analysis (2nd ed.). London: Chapman & Hall. Choi, Y. K., Johnson, W. O., Collins, M. T., & Gardner, I. A. (2006).
Bayesian inferences for receiver operating characteristic curves
in the absence of a gold standard. Journal of Agricultural, Biological, and Environmental Statistics, 11, 210–229.
Committee on the Revision of the Specialty Guidelines for Forensic
Psychology. (11 January 2006). Specialty guidelines for forensic psychology, second official draft. http://www.ap-ls.org/links/. Accessed 19 Sept 2008.
Cooper, V. G., & Zapf, P. A. (2003). Predictor variables in
competency to stand trial decisions. Law and Human Behavior, 27, 423–436.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in
psychological tests. Psychological Bulletin, 52, 281–302. Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579 (1993). Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus
actuarial judgment. Science, 243, 1668–1674. Douglas, K. S., Ogloff, J. R., Nicholls, T. L., & Grant, I. (1999).
Assessing risk for violence among psychiatric patients: The
HCR-20 violence risk assessment scheme and the Psychopathy
Checklist: Screening Version. Journal of Consulting and Clinical Psychology, 67, 917–930.
Dusky v. United States, 362 U.S. 402 (1960). Faigman, D. L., Saks, M. J., Sanders, J., & Cheng, E. K. (2008).
Modern scientific evidence: Standards, statistics, and research methods, student ed.. Eagan, MN: Thomson West.
Faraone, S. V., & Tsuang, M. T. (1994). Measuring diagnostic
accuracy in the absence of a ‘‘gold standard.’’ American Journal of Psychiatry, 151, 650–657.
Gardner, W., Lidz, C. W., Mulvey, E. P., & Shaw, E. C. (1996).
Clinical versus actuarial predictions of violence of patients with
mental illnesses. Journal of Consulting and Clinical Psychology, 64, 602–609.
Gelfand, A. E., & Smith, A. F. M. (1990). Sampling-based
approaches to calculating marginal densities. Journal of the American Statistical Association, 85, 389–409.
Geman, S., & Geman, D. (1984). Stochastic relaxation, Gibbs
distributions, and the Bayesian restoration of images. IEEE Transactions on Pattern Analysis and Machine Intelligence, 6, 721–741.
Golding, S. L., Roesch, R., & Schreiber, J. (1984). Assessment and
conceptualization of competency to stand trial: Preliminary data
on the Interdisciplinary Fitness Interview. Law and Human Behavior, 8, 321–334.
Grisso, T. (2003). Legally relevant assessments for legal competen-
cies. In T. Grisso (Ed.), Evaluating competencies: Forensic assessments, instruments (2nd ed., pp. 21–40). New York: Kluwer Academic/Plenum Publishers.
Gutheil, T. G. (2004). The expert witness. In R. I. Simon & L. H.
Gold (Eds.), The American Psychiatric Publishing textbook of forensic psychiatry (pp. 75–89). Arlington, VA: American Psychiatric Publishing.
Hagen, M. A. (1997). Whores of the court: The fraud of psychiatric testimony and the rape of American justice. New York: ReganBooks.
Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the
area under the receiver operating characteristic (ROC) curve.
Radiology, 143, 29–36. Harris, G. T., Rice, M. E., & Cormier, C. A. (2002). Prospective
replication of the Violence Risk Appraisal Guide in predicting
violent recidivism among forensic patients. Law and Human Behavior, 26, 377–394.
Henkelman, R. M., Kay, I., & Bronskill, M. J. (1990). Receiver
operator characteristic (ROC) analysis without truth. Medical Decision Making, 10, 24–29.
Jackson v. Indiana, 406 U.S. 715 (1972). Jacobs, M. S., Ryba, N. L., & Zapf, P. A. (2008). Competence-related
abilities and psychiatric symptoms: An analysis of the under-
lying structure and correlates of the MacCAT-CA and the BPRS.
Law and Human Behavior, 32, 64–77. Kim, S. Y. H., Appelbaum, P. S., Swan, J., Stroup, T. S., McEvoy, J. P.,
Goff, D. C., et al. (2007). Determining when impairment
constitutes incapacity for informed consent in schizophrenia
research. British Journal of Psychiatry, 191, 38–43. Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999). Lehman, C. D., Gatsonis, C., Kuhl, C. K., Hendrick, R. E., Pisano, E. D.,
Hanna, L., et al. (2007). MRI evaluation of the contralateral breast
416 Law Hum Behav (2010) 34:402–417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
in women with recently diagnosed breast cancer. New England Journal of Medicine, 356, 1295–1303.
Lunn, D. J., Thomas, A., Best, N., & Spiegelhalter, D. (2000).
WinBUGS—A Bayesian modeling framework: Concepts, struc-
ture, and extensibility. Statistics and Computing, 10, 325–337. McGraw, K. O., & Wong, S. P. (1996). Forming inferences about
some intraclass correlations. Psychological Methods, 1, 30–46. Melton, G. B., Petrila, J., Poythress, N., Slobogin, C., Lyons, P., &
Otto, R. K. (2007). Psychological evaluations for the courts: A handbook for mental health professionals and lawyers (3rd ed.). New York: Guilford.
Metropolis, N., Rosenbluth, A., Rosenbluth, M., Teller, A., & Teller,
E. (1953). Equations of state calculations by fast computing
machines. Journal of Chemical Physics, 21, 1087–1091. Miller, H. A. (2005). The Miller-Forensic Assessment of Symptoms
Test (M-Fast): Test generalizability and utility across race
literacy, and clinical opinion. Criminal Justice and Behavior, 32, 591–611.
Monahan, J., Steadman, H. J., Appelbaum, P. S., Grisso, T., Mulvey,
E. P., Roth, L. H., et al. (2006). The classification of violence
risk. Behavioral Science and the Law, 24, 721–730. Mossman, D. (1999). ‘‘Hired guns’’, ‘‘whores’’, and ‘‘prostitutes’’:
Case law references to clinicians of ill repute. Journal of the American Academy of Psychiatry and the Law, 27, 414–425.
Mossman, D. (2005). Is prosecution ‘‘medically appropriate’’? New England Journal on Criminal and Civil Confinement, 31, 15–80.
Mossman, D. (2007). Predicting restorability of incompetent criminal
defendants. Journal of the American Academy of Psychiatry and the Law, 35, 34–43.
Mossman, D. (2008). Conceptualizing and characterizing accuracy in
assessments of competence to stand trial. Journal of the American Academy of Psychiatry and the Law, 36, 340–351.
Mossman, D., Noffsinger, S. G., Ash, P., Frierson, R. L., Gerbasi, J.,
Hackett, M., et al. (2007). AAPL practice guideline for the
forensic psychiatric evaluation of competence to stand trial.
Journal of the American Academy of Psychiatry and the Law, 35(Suppl 4), S3–S72.
Mossman, D., & Somoza, E. (1991). ROC curves, test accuracy, and
the description of diagnostic tests. Journal of Neuropsychiatry and Clinical Neurosciences, 3, 330–333.
Murrie, D. C., Boccaccini, M. T., Turner, D., Meeks, M., Woods, C.,
& Tussey, C. (2009). Rater (dis)agreement on risk assessment
measures in sexually violent predator proceedings: Evidence of
adversarial allegiance in forensic evaluation? Psychology, Public Policy, and Law, 15, 19–53.
Murrie, D. C., Boccaccini, M., Zapf, P. A., Warren, J. I., & Henderson,
C. E. (2008). Clinician variation in findings of competence to
stand trial. Psychology, Public Policy, and Law, 14, 177–193. Obuchowski, N. A. (2003). Receiver operating characteristic curves
and their use in radiology. Radiology, 229, 3–8. Parry, J., & Drogin, E. Y. (2007). Mental disability law, evidence and
testimony: A comprehensive reference manual for lawyers, judges, and mental disability professionals. Washington, DC: American Bar Association.
Pate v. Robinson, 383 U.S. 375 (1966).
Poythress, N., Monahan, J., Bonnie, R., Otto, R. K., & Hoge, S. K.
(2002). Adjudicative competence: The MacArthur studies. New York: Kluwer/Plenum.
Qu, Y., Tan, M., & Kutner, M. H. (1996). Random effects models in
latent class analysis for evaluating accuracy of diagnostic tests.
Biometrics, 53, 797–810. Rice, M. E., & Harris, G. T. (1995). Violent recidivism: Assessing
predictive validity. Journal of Consulting and Clinical Psychol- ogy, 63, 737–748.
Rosenfeld, B., & Ritchie, K. (1998). Competence to stand trial:
Clinician reliability and the role of offense severity. Journal of Forensic Sciences, 43, 151–159.
Skeem, J., Golding, S., Cohn, N., & Berge, G. (1998). The logic and
reliability of expert opinion on competence to stand trial. Law and Human Behavior, 22, 519–547.
Small, G. W., Kepe, V., Ercoli, L. M., Siddarth, P., Bookheimer, S. Y.,
Miller, K. J., et al. (2006). PET of brain amyloid and tau in mild
cognitive impairment. New England Journal of Medicine, 355, 2652–2663.
Somoza, E., & Mossman, D. (1991). ROC curves and the binormal
assumption. Journal of Neuropsychiatry and Clinical Neuros- ciences, 3, 436–439.
Spencer, B. D. (2008). When do latent class models overstate accuracy for binary classifiers? With applications to jury accuracy, survey response error, and diagnostic error. Institute for Policy Research, Northwestern University, Working Paper
Series WP-08-10.
State v. Sullivan, 739 N.E.2d 788 (Ohio 2001). Steadman, H. J., Mulvey, E. P., Monahan, J., Robbins, P. C.,
Appelbaum, P. S., Grisso, T., et al. (1998). Violence by people
discharged from acute psychiatric inpatient facilities and by
others in the same neighborhoods. Archives of General Psychiatry, 55, 393–401.
Swets, J. A. (1995). Signal detection theory and ROC analysis in psychology and diagnostics: Collected papers. Mahwah, NJ: Lawrence Erlbaum Associates.
Uebersax, J. S. (1988). Validity inferences from interobserver
agreement. Psychological Bulletin, 104, 405–416. Uebersax, J. S., & Grove, W. M. (1990). Latent class analysis of
diagnostic agreement. Statistics in Medicine, 9, 559–572. Weissman, H. N., & DeBow, D. M. (2003). Ethical principles and
professional competencies. In I. B. Weiner (Series Ed.) & A. M.
Goldstein (Vol. Ed.), Handbook of psychology: Vol. 11. Forensic psychology (pp. 33–53). New York: Wiley.
Zapf, P. A., Hubbard, K. L., Cooper, V. G., Wheeles, M. C., & Ronan,
K. A. (2004). Have the courts abdicated their responsibility for
determination of competency to stand trial to clinicians? Journal of Forensic Psychology Practice, 4, 27–44.
Zhou, X. H., Castelluccio, P., & Zhou, C. (2005). Nonparametric
estimation of ROC curves in the absence of a gold standard.
Biometrics, 61, 600–609. Zweig, M. H., & Campbell, G. (1993). Receiver operating character-
istic (ROC) plots: A fundamental evaluation tool in clinical
medicine. Clinical Chemistry, 39, 561–577.
Law Hum Behav (2010) 34:402–417 417
123
T hi
s do
cu m
en t i
s co
py ri
gh te
d by
th e
A m
er ic
an P
sy ch
ol og
ic al
A ss
oc ia
tio n
or o
ne o
f i ts
a lli
ed p
ub lis
he rs
. T
hi s
ar tic
le is
in te
nd ed
s ol
el y
fo r t
he p
er so
na l u
se o
f t he
in di
vi du
al u
se r a
nd is
n ot
to b
e di
ss em
in at
ed b
ro ad
ly .
<< /ASCII85EncodePages false /AllowTransparency false /AutoPositionEPSFiles true /AutoRotatePages /None /Binding /Left /CalGrayProfile (Gray Gamma 2.2) /CalRGBProfile (sRGB IEC61966-2.1) /CalCMYKProfile (ISO Coated v2 300% \050ECI\051) /sRGBProfile (sRGB IEC61966-2.1) /CannotEmbedFontPolicy /Error /CompatibilityLevel 1.3 /CompressObjects /Off /CompressPages true /ConvertImagesToIndexed true /PassThroughJPEGImages true /CreateJobTicket false /DefaultRenderingIntent /Perceptual /DetectBlends true /DetectCurves 0.1000 /ColorConversionStrategy /sRGB /DoThumbnails true /EmbedAllFonts true /EmbedOpenType false /ParseICCProfilesInComments true /EmbedJobOptions true /DSCReportingLevel 0 /EmitDSCWarnings false /EndPage -1 /ImageMemory 1048576 /LockDistillerParams true /MaxSubsetPct 100 /Optimize true /OPM 1 /ParseDSCComments true /ParseDSCCommentsForDocInfo true /PreserveCopyPage true /PreserveDICMYKValues true /PreserveEPSInfo true /PreserveFlatness true /PreserveHalftoneInfo false /PreserveOPIComments false /PreserveOverprintSettings true /StartPage 1 /SubsetFonts false /TransferFunctionInfo /Apply /UCRandBGInfo /Preserve /UsePrologue false /ColorSettingsFile () /AlwaysEmbed [ true ] /NeverEmbed [ true ] /AntiAliasColorImages false /CropColorImages true /ColorImageMinResolution 149 /ColorImageMinResolutionPolicy /Warning /DownsampleColorImages true /ColorImageDownsampleType /Bicubic /ColorImageResolution 150 /ColorImageDepth -1 /ColorImageMinDownsampleDepth 1 /ColorImageDownsampleThreshold 1.50000 /EncodeColorImages true /ColorImageFilter /DCTEncode /AutoFilterColorImages true /ColorImageAutoFilterStrategy /JPEG /ColorACSImageDict << /QFactor 0.40 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /ColorImageDict << /QFactor 0.15 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /JPEG2000ColorACSImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /JPEG2000ColorImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /AntiAliasGrayImages false /CropGrayImages true /GrayImageMinResolution 149 /GrayImageMinResolutionPolicy /Warning /DownsampleGrayImages true /GrayImageDownsampleType /Bicubic /GrayImageResolution 150 /GrayImageDepth -1 /GrayImageMinDownsampleDepth 2 /GrayImageDownsampleThreshold 1.50000 /EncodeGrayImages true /GrayImageFilter /DCTEncode /AutoFilterGrayImages true /GrayImageAutoFilterStrategy /JPEG /GrayACSImageDict << /QFactor 0.40 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /GrayImageDict << /QFactor 0.15 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /JPEG2000GrayACSImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /JPEG2000GrayImageDict << /TileWidth 256 /TileHeight 256 /Quality 30 >> /AntiAliasMonoImages false /CropMonoImages true /MonoImageMinResolution 599 /MonoImageMinResolutionPolicy /Warning /DownsampleMonoImages true /MonoImageDownsampleType /Bicubic /MonoImageResolution 600 /MonoImageDepth -1 /MonoImageDownsampleThreshold 1.50000 /EncodeMonoImages true /MonoImageFilter /CCITTFaxEncode /MonoImageDict << /K -1 >> /AllowPSXObjects false /CheckCompliance [ /None ] /PDFX1aCheck false /PDFX3Check false /PDFXCompliantPDFOnly false /PDFXNoTrimBoxError true /PDFXTrimBoxToMediaBoxOffset [ 0.00000 0.00000 0.00000 0.00000 ] /PDFXSetBleedBoxToMediaBox true /PDFXBleedBoxToTrimBoxOffset [ 0.00000 0.00000 0.00000 0.00000 ] /PDFXOutputIntentProfile (None) /PDFXOutputConditionIdentifier () /PDFXOutputCondition () /PDFXRegistryName () /PDFXTrapped /False /CreateJDFFile false /Description << /ARA <FEFF06270633062A062E062F0645002006470630064700200627064406250639062F0627062F0627062A002006440625064606340627062100200648062B062706260642002000410064006F00620065002000500044004600200645062A064806270641064206290020064406440637062806270639062900200641064A00200627064406450637062706280639002006300627062A0020062F0631062C0627062A002006270644062C0648062F0629002006270644063906270644064A0629061B0020064A06450643064600200641062A062D00200648062B0627062606420020005000440046002006270644064506460634062306290020062806270633062A062E062F062706450020004100630072006F0062006100740020064800410064006F006200650020005200650061006400650072002006250635062F0627063100200035002E0030002006480627064406250635062F062706310627062A0020062706440623062D062F062B002E0635062F0627063100200035002E0030002006480627064406250635062F062706310627062A0020062706440623062D062F062B002E> /BGR <FEFF04180437043f043e043b043704320430043904420435002004420435043704380020043d0430044104420440043e0439043a0438002c00200437043000200434043000200441044a0437043404300432043004420435002000410064006f00620065002000500044004600200434043e043a0443043c0435043d04420438002c0020043c0430043a04410438043c0430043b043d043e0020043f044004380433043e04340435043d04380020043704300020043204380441043e043a043e043a0430044704350441044204320435043d0020043f04350447043004420020043704300020043f044004350434043f0435044704300442043d04300020043f043e04340433043e0442043e0432043a0430002e002000200421044a04370434043004340435043d043804420435002000500044004600200434043e043a0443043c0435043d044204380020043c043e0433043004420020043404300020044104350020043e0442043204300440044f0442002004410020004100630072006f00620061007400200438002000410064006f00620065002000520065006100640065007200200035002e00300020043800200441043b0435043404320430044904380020043204350440044104380438002e> /CHS <FEFF4f7f75288fd94e9b8bbe5b9a521b5efa7684002000410064006f006200650020005000440046002065876863900275284e8e9ad88d2891cf76845370524d53705237300260a853ef4ee54f7f75280020004100630072006f0062006100740020548c002000410064006f00620065002000520065006100640065007200200035002e003000204ee553ca66f49ad87248672c676562535f00521b5efa768400200050004400460020658768633002> /CHT <FEFF4f7f752890194e9b8a2d7f6e5efa7acb7684002000410064006f006200650020005000440046002065874ef69069752865bc9ad854c18cea76845370524d5370523786557406300260a853ef4ee54f7f75280020004100630072006f0062006100740020548c002000410064006f00620065002000520065006100640065007200200035002e003000204ee553ca66f49ad87248672c4f86958b555f5df25efa7acb76840020005000440046002065874ef63002> /CZE <FEFF005400610074006f0020006e006100730074006100760065006e00ed00200070006f0075017e0069006a007400650020006b0020007600790074007600e101590065006e00ed00200064006f006b0075006d0065006e0074016f002000410064006f006200650020005000440046002c0020006b00740065007200e90020007300650020006e0065006a006c00e90070006500200068006f006400ed002000700072006f0020006b00760061006c00690074006e00ed0020007400690073006b00200061002000700072006500700072006500730073002e002000200056007900740076006f01590065006e00e900200064006f006b0075006d0065006e007400790020005000440046002000620075006400650020006d006f017e006e00e90020006f007400650076015900ed007400200076002000700072006f006700720061006d0065006300680020004100630072006f00620061007400200061002000410064006f00620065002000520065006100640065007200200035002e0030002000610020006e006f0076011b006a016100ed00630068002e> /DAN <FEFF004200720075006700200069006e0064007300740069006c006c0069006e006700650072006e0065002000740069006c0020006100740020006f007000720065007400740065002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e007400650072002c0020006400650072002000620065006400730074002000650067006e006500720020007300690067002000740069006c002000700072006500700072006500730073002d007500640073006b007200690076006e0069006e00670020006100660020006800f8006a0020006b00760061006c0069007400650074002e0020004400650020006f007000720065007400740065006400650020005000440046002d0064006f006b0075006d0065006e0074006500720020006b0061006e002000e50062006e00650073002000690020004100630072006f00620061007400200065006c006c006500720020004100630072006f006200610074002000520065006100640065007200200035002e00300020006f00670020006e0079006500720065002e> /ESP <FEFF005500740069006c0069006300650020006500730074006100200063006f006e0066006900670075007200610063006900f3006e0020007000610072006100200063007200650061007200200064006f00630075006d0065006e0074006f00730020005000440046002000640065002000410064006f0062006500200061006400650063007500610064006f00730020007000610072006100200069006d0070007200650073006900f3006e0020007000720065002d0065006400690074006f007200690061006c00200064006500200061006c00740061002000630061006c0069006400610064002e002000530065002000700075006500640065006e00200061006200720069007200200064006f00630075006d0065006e0074006f00730020005000440046002000630072006500610064006f007300200063006f006e0020004100630072006f006200610074002c002000410064006f00620065002000520065006100640065007200200035002e003000200079002000760065007200730069006f006e0065007300200070006f00730074006500720069006f007200650073002e> /ETI <FEFF004b00610073007500740061006700650020006e0065006900640020007300e4007400740065006900640020006b00760061006c006900740065006500740073006500200074007200fc006b006900650065006c007300650020007000720069006e00740069006d0069007300650020006a0061006f006b007300200073006f00620069006c0069006b0065002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e00740069006400650020006c006f006f006d006900730065006b0073002e00200020004c006f006f0064007500640020005000440046002d0064006f006b0075006d0065006e00740065002000730061006100740065002000610076006100640061002000700072006f006700720061006d006d006900640065006700610020004100630072006f0062006100740020006e0069006e0067002000410064006f00620065002000520065006100640065007200200035002e00300020006a00610020007500750065006d006100740065002000760065007200730069006f006f006e00690064006500670061002e000d000a> /FRA <FEFF005500740069006c006900730065007a00200063006500730020006f007000740069006f006e00730020006100660069006e00200064006500200063007200e900650072002000640065007300200064006f00630075006d0065006e00740073002000410064006f00620065002000500044004600200070006f0075007200200075006e00650020007100750061006c0069007400e90020006400270069006d007000720065007300730069006f006e00200070007200e9007000720065007300730065002e0020004c0065007300200064006f00630075006d0065006e00740073002000500044004600200063007200e900e90073002000700065007500760065006e0074002000ea0074007200650020006f007500760065007200740073002000640061006e00730020004100630072006f006200610074002c002000610069006e00730069002000710075002700410064006f00620065002000520065006100640065007200200035002e0030002000650074002000760065007200730069006f006e007300200075006c007400e90072006900650075007200650073002e> /GRE <FEFF03a703c103b703c303b903bc03bf03c003bf03b903ae03c303c403b5002003b103c503c403ad03c2002003c403b903c2002003c103c503b803bc03af03c303b503b903c2002003b303b903b1002003bd03b1002003b403b703bc03b903bf03c503c103b303ae03c303b503c403b5002003ad03b303b303c103b103c603b1002000410064006f006200650020005000440046002003c003bf03c5002003b503af03bd03b103b9002003ba03b103c42019002003b503be03bf03c703ae03bd002003ba03b103c403ac03bb03bb03b703bb03b1002003b303b903b1002003c003c103bf002d03b503ba03c403c503c003c903c403b903ba03ad03c2002003b503c103b303b103c303af03b503c2002003c503c803b703bb03ae03c2002003c003bf03b903cc03c403b703c403b103c2002e0020002003a403b10020005000440046002003ad03b303b303c103b103c603b1002003c003bf03c5002003ad03c703b503c403b5002003b403b703bc03b903bf03c503c103b303ae03c303b503b9002003bc03c003bf03c103bf03cd03bd002003bd03b1002003b103bd03bf03b903c703c403bf03cd03bd002003bc03b5002003c403bf0020004100630072006f006200610074002c002003c403bf002000410064006f00620065002000520065006100640065007200200035002e0030002003ba03b103b9002003bc03b503c403b103b303b503bd03ad03c303c403b503c103b503c2002003b503ba03b403cc03c303b503b903c2002e> /HEB <FEFF05D405E905EA05DE05E905D5002005D105D405D205D305E805D505EA002005D005DC05D4002005DB05D305D9002005DC05D905E605D505E8002005DE05E105DE05DB05D9002000410064006F006200650020005000440046002005D405DE05D505EA05D005DE05D905DD002005DC05D405D305E405E105EA002005E705D305DD002D05D305E405D505E1002005D005D905DB05D505EA05D905EA002E002005DE05E105DE05DB05D90020005000440046002005E905E005D505E605E805D5002005E005D905EA05E005D905DD002005DC05E405EA05D905D705D4002005D105D005DE05E605E205D505EA0020004100630072006F006200610074002005D5002D00410064006F00620065002000520065006100640065007200200035002E0030002005D505D205E805E105D005D505EA002005DE05EA05E705D305DE05D505EA002005D905D505EA05E8002E05D005DE05D905DD002005DC002D005000440046002F0058002D0033002C002005E205D905D905E005D5002005D105DE05D305E805D905DA002005DC05DE05E905EA05DE05E9002005E905DC0020004100630072006F006200610074002E002005DE05E105DE05DB05D90020005000440046002005E905E005D505E605E805D5002005E005D905EA05E005D905DD002005DC05E405EA05D905D705D4002005D105D005DE05E605E205D505EA0020004100630072006F006200610074002005D5002D00410064006F00620065002000520065006100640065007200200035002E0030002005D505D205E805E105D005D505EA002005DE05EA05E705D305DE05D505EA002005D905D505EA05E8002E> /HRV (Za stvaranje Adobe PDF dokumenata najpogodnijih za visokokvalitetni ispis prije tiskanja koristite ove postavke. Stvoreni PDF dokumenti mogu se otvoriti Acrobat i Adobe Reader 5.0 i kasnijim verzijama.) /HUN <FEFF004b0069007600e1006c00f30020006d0069006e0151007300e9006701710020006e0079006f006d00640061006900200065006c0151006b00e90073007a00ed007401510020006e0079006f006d00740061007400e100730068006f007a0020006c006500670069006e006b00e1006200620020006d0065006700660065006c0065006c0151002000410064006f00620065002000500044004600200064006f006b0075006d0065006e00740075006d006f006b0061007400200065007a0065006b006b0065006c0020006100200062006500e1006c006c00ed007400e10073006f006b006b0061006c0020006b00e90073007a00ed0074006800650074002e0020002000410020006c00e90074007200650068006f007a006f00740074002000500044004600200064006f006b0075006d0065006e00740075006d006f006b00200061007a0020004100630072006f006200610074002000e9007300200061007a002000410064006f00620065002000520065006100640065007200200035002e0030002c0020007600610067007900200061007a002000610074007400f3006c0020006b00e9007301510062006200690020007600650072007a006900f3006b006b0061006c0020006e00790069007400680061007400f3006b0020006d00650067002e> /ITA <FEFF005500740069006c0069007a007a006100720065002000710075006500730074006500200069006d0070006f007300740061007a0069006f006e00690020007000650072002000630072006500610072006500200064006f00630075006d0065006e00740069002000410064006f00620065002000500044004600200070006900f900200061006400610074007400690020006100200075006e00610020007000720065007300740061006d0070006100200064006900200061006c007400610020007100750061006c0069007400e0002e0020004900200064006f00630075006d0065006e007400690020005000440046002000630072006500610074006900200070006f00730073006f006e006f0020006500730073006500720065002000610070006500720074006900200063006f006e0020004100630072006f00620061007400200065002000410064006f00620065002000520065006100640065007200200035002e003000200065002000760065007200730069006f006e006900200073007500630063006500730073006900760065002e> /JPN <FEFF9ad854c18cea306a30d730ea30d730ec30b951fa529b7528002000410064006f0062006500200050004400460020658766f8306e4f5c6210306b4f7f75283057307e305930023053306e8a2d5b9a30674f5c62103055308c305f0020005000440046002030d530a130a430eb306f3001004100630072006f0062006100740020304a30883073002000410064006f00620065002000520065006100640065007200200035002e003000204ee5964d3067958b304f30533068304c3067304d307e305930023053306e8a2d5b9a306b306f30d530a930f330c8306e57cb30818fbc307f304c5fc59808306730593002> /KOR <FEFFc7740020c124c815c7440020c0acc6a9d558c5ec0020ace0d488c9c80020c2dcd5d80020c778c1c4c5d00020ac00c7a50020c801d569d55c002000410064006f0062006500200050004400460020bb38c11cb97c0020c791c131d569b2c8b2e4002e0020c774b807ac8c0020c791c131b41c00200050004400460020bb38c11cb2940020004100630072006f0062006100740020bc0f002000410064006f00620065002000520065006100640065007200200035002e00300020c774c0c1c5d0c11c0020c5f40020c2180020c788c2b5b2c8b2e4002e> /LTH <FEFF004e006100750064006f006b0069007400650020016100690075006f007300200070006100720061006d006500740072007500730020006e006f0072011700640061006d00690020006b0075007200740069002000410064006f00620065002000500044004600200064006f006b0075006d0065006e007400750073002c0020006b00750072006900650020006c0061006200690061007500730069006100690020007000720069007400610069006b007900740069002000610075006b01610074006f00730020006b006f006b007900620117007300200070006100720065006e006700740069006e00690061006d00200073007000610075007300640069006e0069006d00750069002e0020002000530075006b0075007200740069002000500044004600200064006f006b0075006d0065006e007400610069002000670061006c006900200062016b007400690020006100740069006400610072006f006d00690020004100630072006f006200610074002000690072002000410064006f00620065002000520065006100640065007200200035002e0030002000610072002000760117006c00650073006e0117006d00690073002000760065007200730069006a006f006d00690073002e> /LVI <FEFF0049007a006d0061006e0074006f006a00690065007400200161006f00730020006900650073007400610074012b006a0075006d00750073002c0020006c0061006900200076006500690064006f00740075002000410064006f00620065002000500044004600200064006f006b0075006d0065006e007400750073002c0020006b006100730020006900720020012b00700061016100690020007000690065006d01130072006f00740069002000610075006700730074006100730020006b00760061006c0069007401010074006500730020007000690072006d007300690065007300700069006501610061006e006100730020006400720075006b00610069002e00200049007a0076006500690064006f006a006900650074002000500044004600200064006f006b0075006d0065006e007400750073002c0020006b006f002000760061007200200061007400760113007200740020006100720020004100630072006f00620061007400200075006e002000410064006f00620065002000520065006100640065007200200035002e0030002c0020006b0101002000610072012b00200074006f0020006a00610075006e0101006b0101006d002000760065007200730069006a0101006d002e> /NLD (Gebruik deze instellingen om Adobe PDF-documenten te maken die zijn geoptimaliseerd voor prepress-afdrukken van hoge kwaliteit. De gemaakte PDF-documenten kunnen worden geopend met Acrobat en Adobe Reader 5.0 en hoger.) /NOR <FEFF004200720075006b00200064006900730073006500200069006e006e007300740069006c006c0069006e00670065006e0065002000740069006c002000e50020006f0070007000720065007400740065002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e00740065007200200073006f006d00200065007200200062006500730074002000650067006e0065007400200066006f00720020006600f80072007400720079006b006b0073007500740073006b00720069006600740020006100760020006800f800790020006b00760061006c0069007400650074002e0020005000440046002d0064006f006b0075006d0065006e00740065006e00650020006b0061006e002000e50070006e00650073002000690020004100630072006f00620061007400200065006c006c00650072002000410064006f00620065002000520065006100640065007200200035002e003000200065006c006c00650072002000730065006e006500720065002e> /POL <FEFF0055007300740061007700690065006e0069006100200064006f002000740077006f0072007a0065006e0069006100200064006f006b0075006d0065006e007400f300770020005000440046002000700072007a0065007a006e00610063007a006f006e00790063006800200064006f002000770079006400720075006b00f30077002000770020007700790073006f006b00690065006a0020006a0061006b006f015b00630069002e002000200044006f006b0075006d0065006e0074007900200050004400460020006d006f017c006e00610020006f007400770069006500720061010700200077002000700072006f006700720061006d006900650020004100630072006f00620061007400200069002000410064006f00620065002000520065006100640065007200200035002e0030002000690020006e006f00770073007a0079006d002e> /PTB <FEFF005500740069006c0069007a006500200065007300730061007300200063006f006e00660069006700750072006100e700f50065007300200064006500200066006f0072006d00610020006100200063007200690061007200200064006f00630075006d0065006e0074006f0073002000410064006f0062006500200050004400460020006d00610069007300200061006400650071007500610064006f00730020007000610072006100200070007200e9002d0069006d0070007200650073007300f50065007300200064006500200061006c007400610020007100750061006c00690064006100640065002e0020004f007300200064006f00630075006d0065006e0074006f00730020005000440046002000630072006900610064006f007300200070006f00640065006d0020007300650072002000610062006500720074006f007300200063006f006d0020006f0020004100630072006f006200610074002000650020006f002000410064006f00620065002000520065006100640065007200200035002e0030002000650020007600650072007300f50065007300200070006f00730074006500720069006f007200650073002e> /RUM <FEFF005500740069006c0069007a00610163006900200061006300650073007400650020007300650074010300720069002000700065006e007400720075002000610020006300720065006100200064006f00630075006d0065006e00740065002000410064006f006200650020005000440046002000610064006500630076006100740065002000700065006e0074007200750020007400690070010300720069007200650061002000700072006500700072006500730073002000640065002000630061006c006900740061007400650020007300750070006500720069006f006100720103002e002000200044006f00630075006d0065006e00740065006c00650020005000440046002000630072006500610074006500200070006f00740020006600690020006400650073006300680069007300650020006300750020004100630072006f006200610074002c002000410064006f00620065002000520065006100640065007200200035002e00300020015f00690020007600650072007300690075006e0069006c006500200075006c0074006500720069006f006100720065002e> /RUS <FEFF04180441043f043e043b044c04370443043904420435002004340430043d043d044b04350020043d0430044104420440043e0439043a043800200434043b044f00200441043e043704340430043d0438044f00200434043e043a0443043c0435043d0442043e0432002000410064006f006200650020005000440046002c0020043c0430043a04410438043c0430043b044c043d043e0020043f043e04340445043e0434044f04490438044500200434043b044f00200432044b0441043e043a043e043a0430044704350441044204320435043d043d043e0433043e00200434043e043f0435044704300442043d043e0433043e00200432044b0432043e04340430002e002000200421043e043704340430043d043d044b04350020005000440046002d0434043e043a0443043c0435043d0442044b0020043c043e0436043d043e0020043e0442043a0440044b043204300442044c002004410020043f043e043c043e0449044c044e0020004100630072006f00620061007400200438002000410064006f00620065002000520065006100640065007200200035002e00300020043800200431043e043b043504350020043f043e04370434043d043804450020043204350440044104380439002e> /SKY <FEFF0054006900650074006f0020006e006100730074006100760065006e0069006100200070006f0075017e0069007400650020006e00610020007600790074007600e100720061006e0069006500200064006f006b0075006d0065006e0074006f0076002000410064006f006200650020005000440046002c0020006b0074006f007200e90020007300610020006e0061006a006c0065007001610069006500200068006f0064006900610020006e00610020006b00760061006c00690074006e00fa00200074006c0061010d00200061002000700072006500700072006500730073002e00200056007900740076006f00720065006e00e900200064006f006b0075006d0065006e007400790020005000440046002000620075006400650020006d006f017e006e00e90020006f00740076006f00720069016500200076002000700072006f006700720061006d006f006300680020004100630072006f00620061007400200061002000410064006f00620065002000520065006100640065007200200035002e0030002000610020006e006f0076016100ed00630068002e> /SLV <FEFF005400650020006e006100730074006100760069007400760065002000750070006f0072006100620069007400650020007a00610020007500730074007600610072006a0061006e006a006500200064006f006b0075006d0065006e0074006f0076002000410064006f006200650020005000440046002c0020006b006900200073006f0020006e0061006a007000720069006d00650072006e0065006a016100690020007a00610020006b0061006b006f0076006f00730074006e006f0020007400690073006b0061006e006a00650020007300200070007200690070007200610076006f0020006e00610020007400690073006b002e00200020005500730074007600610072006a0065006e006500200064006f006b0075006d0065006e0074006500200050004400460020006a00650020006d006f0067006f010d00650020006f0064007000720065007400690020007a0020004100630072006f00620061007400200069006e002000410064006f00620065002000520065006100640065007200200035002e003000200069006e0020006e006f00760065006a01610069006d002e> /SUO <FEFF004b00e40079007400e40020006e00e40069007400e4002000610073006500740075006b007300690061002c0020006b0075006e0020006c0075006f00740020006c00e400680069006e006e00e4002000760061006100740069007600610061006e0020007000610069006e006100740075006b00730065006e002000760061006c006d0069007300740065006c00750074007900f6006800f6006e00200073006f00700069007600690061002000410064006f0062006500200050004400460020002d0064006f006b0075006d0065006e007400740065006a0061002e0020004c0075006f0064007500740020005000440046002d0064006f006b0075006d0065006e00740069007400200076006f0069006400610061006e0020006100760061007400610020004100630072006f0062006100740069006c006c00610020006a0061002000410064006f00620065002000520065006100640065007200200035002e0030003a006c006c00610020006a006100200075007500640065006d006d0069006c006c0061002e> /SVE <FEFF0041006e007600e4006e00640020006400650020006800e4007200200069006e0073007400e4006c006c006e0069006e006700610072006e00610020006f006d002000640075002000760069006c006c00200073006b006100700061002000410064006f006200650020005000440046002d0064006f006b0075006d0065006e007400200073006f006d002000e400720020006c00e4006d0070006c0069006700610020006600f60072002000700072006500700072006500730073002d007500740073006b00720069006600740020006d006500640020006800f600670020006b00760061006c0069007400650074002e002000200053006b006100700061006400650020005000440046002d0064006f006b0075006d0065006e00740020006b0061006e002000f600700070006e00610073002000690020004100630072006f0062006100740020006f00630068002000410064006f00620065002000520065006100640065007200200035002e00300020006f00630068002000730065006e006100720065002e> /TUR <FEFF005900fc006b00730065006b0020006b0061006c006900740065006c0069002000f6006e002000790061007a006401310072006d00610020006200610073006b013100730131006e006100200065006e0020006900790069002000750079006100620069006c006500630065006b002000410064006f006200650020005000440046002000620065006c00670065006c0065007200690020006f006c0075015f007400750072006d0061006b0020006900e70069006e00200062007500200061007900610072006c0061007201310020006b0075006c006c0061006e0131006e002e00200020004f006c0075015f0074007500720075006c0061006e0020005000440046002000620065006c00670065006c0065007200690020004100630072006f006200610074002000760065002000410064006f00620065002000520065006100640065007200200035002e003000200076006500200073006f006e0072006100730131006e00640061006b00690020007300fc007200fc006d006c00650072006c00650020006100e70131006c006100620069006c00690072002e> /UKR <FEFF04120438043a043e0440043804410442043e043204430439044204350020044604560020043f043004400430043c043504420440043800200434043b044f0020044104420432043e04400435043d043d044f00200434043e043a0443043c0435043d044204560432002000410064006f006200650020005000440046002c0020044f043a04560020043d04300439043a04400430044904350020043f045604340445043e0434044f0442044c00200434043b044f0020043204380441043e043a043e044f043a04560441043d043e0433043e0020043f0435044004350434043404400443043a043e0432043e0433043e0020043404400443043a0443002e00200020042104420432043e04400435043d045600200434043e043a0443043c0435043d0442043800200050004400460020043c043e0436043d04300020043204560434043a0440043804420438002004430020004100630072006f006200610074002004420430002000410064006f00620065002000520065006100640065007200200035002e0030002004300431043e0020043f04560437043d04560448043e04570020043204350440044104560457002e> /ENU (Use these settings to create Adobe PDF documents best suited for high-quality prepress printing. Created PDF documents can be opened with Acrobat and Adobe Reader 5.0 and later.) /DEU <FEFF004a006f0062006f007000740069006f006e007300200066006f00720020004100630072006f006200610074002000440069007300740069006c006c0065007200200038002000280038002e0032002e00310029000d00500072006f006400750063006500730020005000440046002000660069006c0065007300200077006800690063006800200061007200650020007500730065006400200066006f00720020006f006e006c0069006e0065002e000d0028006300290020003200300031003000200053007000720069006e006700650072002d005600650072006c0061006700200047006d006200480020000d000d0054006800650020006c00610074006500730074002000760065007200730069006f006e002000630061006e00200062006500200064006f0077006e006c006f0061006400650064002000610074002000680074007400700073003a002f002f0070006f007200740061006c002d0064006f0072006400720065006300680074002e0073007000720069006e006700650072002d00730062006d002e0063006f006d002f00500072006f00640075006300740069006f006e002f0046006c006f0077002f00740065006300680064006f0063002f00640065006600610075006c0074002e0061007300700078000d0054006800650072006500200079006f0075002000630061006e00200061006c0073006f002000660069006e0064002000610020007300750069007400610062006c006500200045006e0066006f0063007500730020005000440046002000500072006f00660069006c006500200066006f0072002000500069007400530074006f0070002000500072006f00660065007300730069006f006e0061006c00200030003800200061006e0064002000500069007400530074006f0070002000530065007200760065007200200030003800200066006f007200200070007200650066006c00690067006800740069006e006700200079006f007500720020005000440046002000660069006c006500730020006200650066006f007200650020006a006f00620020007300750062006d0069007300730069006f006e002e000d> >> /Namespace [ (Adobe) (Common) (1.0) ] /OtherNamespaces [ << /AsReaderSpreads false /CropImagesToFrames true /ErrorControl /WarnAndContinue /FlattenerIgnoreSpreadOverrides false /IncludeGuidesGrids false /IncludeNonPrinting false /IncludeSlug false /Namespace [ (Adobe) (InDesign) (4.0) ] /OmitPlacedBitmaps false /OmitPlacedEPS false /OmitPlacedPDF false /SimulateOverprint /Legacy >> << /AddBleedMarks false /AddColorBars false /AddCropMarks false /AddPageInfo false /AddRegMarks false /ConvertColors /ConvertToCMYK /DestinationProfileName () /DestinationProfileSelector /DocumentCMYK /Downsample16BitImages true /FlattenerPreset << /PresetSelector /MediumResolution >> /FormElements false /GenerateStructure false /IncludeBookmarks false /IncludeHyperlinks false /IncludeInteractive false /IncludeLayers false /IncludeProfiles false /MultimediaHandling /UseObjectSettings /Namespace [ (Adobe) (CreativeSuite) (2.0) ] /PDFXOutputIntentProfileSelector /DocumentCMYK /PreserveEditing true /UntaggedCMYKHandling /LeaveUntagged /UntaggedRGBHandling /UseDocumentProfile /UseDocumentBleed false >> ] >> setdistillerparams << /HWResolution [2400 2400] /PageSize [595.276 841.890] >> setpagedevice