Creating a Professional Resumé

profiletashrifie
Psychopathology_History_Diagnosis_and_Empirical_Fo..._----_Pages_37_to_72.pdf

CRITICISMS OF THE CURRENT CLASSIFICATION SYSTEM 21

Whatever its causes, extensive comorbidity is potentially problematic for the DSM, because an ideal classification system yields largely mutually exclusive categories with few overlapping cases (Lilienfeld, VanValkenberg, Larntz, & Akiskal, 1986; Sullivan & Kendler, 1998). As a consequence, such comorbidity may suggest that the current classification system is attaching multiple labels to differing manifestations of the same underlying condition. Defenders of the current classification system are quick to point out that high levels of comorbidity are also prevalent in organic medicine, and often indicate that certain conditions (e.g., diabetes) increase individuals’ risk for other conditions (e.g., blindness), a phenomenon that Kaplan and Feinstein (1974) termed pathogenetic comorbidity. Nevertheless, in stark contrast to organic medicine, in which the causal pathways contributing to pathogenetic comorbidity are often well understood, the causal pathways contributing to pathogenetic comorbidity in the domain of psychopathology generally remain unknown.

MEDICALIZATION OF NORMALITY

A number of critics have raised concerns that recent DSMs, DSM-5 in particular, have overmedicalized normality (Sommers & Satel, 2005). They have done so, these authors contend, in two ways: (1) increasing the number of diagnoses and (2) lowering the threshold for a number of extant diagnoses. In this way, recent DSMs, including DSM-5, may risk opening the floodgates to a pathologizing of largely normative behaviors, emotions, and thoughts. Probably the most vocal critic in this regard has been psychiatrist Allen Frances, who was the principal architect of DSM-IV. In a number of publications, Frances and others have decried DSM-5’s apparently lowered diagnostic thresholds for a number of conditions, as well as its introduction of new and largely unvalidated disorders (Batstra & Frances, 2012a).

Historically, one dramatic change from DSM-I to DSM-IV was the massive increase in the sheer number of diagnoses, a trend potentially reversed by DSM-5. Some critics have argued that this increase reflects the tendency for successive editions of the DSM to expand their range of coverage into new and largely uncharted waters (Houts, 2001). Many of these novel diagnoses, which describe relatively mild problems, may be of questionable validity. For example, the new DSM-5 diagnosis of disruptive mood dysregulation disorder, which is intended to capture many cases of what some authors believe to be pediatric bipolar disorder, has been harshly criticized by Frances (2012) and others for “turn[ing] temper tantrums into a mental disorder.” Another potential example is the new DSM-5 category of minor neurocognitive disorder, which some authors contend may unduly pathologize mild forgetfulness and other largely normative cognitive problems often associated with aging.

As Wakefield (2001) noted, however, there is little evidence that DSM actually expanded its range of coverage from DSM-III to DSM-IV. Although it is unclear at present, the same conclusion may hold for DSM-5. As Wakefield observed, most increases in the number of diagnoses across previous DSMs, since possibly stabilized by DSM-5, reflect an increased splitting of broader diagnoses into progressively narrower subtypes.

The distinction between splitting and lumping derives from biological taxonomy (Mayr, 1982) and refers to the difference between two classificatory styles: the ten- dency to subdivide broad and potentially heterogeneous categories into narrower and presumably more homogeneous categories (splitting) or the tendency to combine narrow and presumably more homogeneous categories into broad and potentially heterogeneous

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

22 ISSUES IN DIAGNOSIS

categories (lumping). For example, given evidence that bipolar I disorder and bipolar II disorder are related (although by no means identical) conditions with relatively similar family histories, laboratory correlates, prognoses, and treatment response, should we keep these diagnoses separate or combine them into a more encompassing, albeit more heterogeneous, category? In the case of autism and allied conditions, the developers of DSM-5 elected to embrace a lumping approach, combining several conditions, such as autistic disorder, Asperger’s disorder, and childhood disintegrative disorder, into the broader domain of what is now termed autism spectrum disorder. Nevertheless, some authors have argued that this change will incorrectly exclude children with milder forms of autism spectrum conditions from the DSM (McPartland, Reichow, & Volkmar, 2012).

The splitting preferences of the architects of DSM-III, DSM-III-R, and DSM-IV in particular have been widely maligned (Houts, 2001). Herman van Praag (2000) even humorously “diagnosed” the DSM’s predilection for splitting as the disorder of nosologomania (also see Ghaemi, 2003). Nevertheless, a preference for splitting is defensible from the standpoint of research and nosological revision. A key point is that the relation between splitting and lumping is asymmetrical: If we begin by splitting diagnostic categories, we can always lump them later if research demonstrates that they are essentially identical according to the Robins and Guze (1970) criteria for validity. Yet, if we begin by lumping it would often be difficult or impossible to split later. As a consequence, we may overlook potentially crucial distinctions among etiologically separable subtypes that bear differing implications for treatment and prevention.

At the same time, it is unclear whether DSM-5’s new diagnoses, such as disruptive mood regulation disorder, reflect a splitting of the diagnostic pie into narrower slices or an enlargement of the pie. If the latter, DSM-5 may indeed risk extending the umbrella of pathology to relatively mild and normative problems.

As noted earlier, a second way in which recent DSMs, including DSM-5, may overmedicalize normality is by lowering the threshold for a number of conditions (Batstra & Frances, 2012a; Frances & Widiger, 2012). For example, by increasing the age of onset from 7 to 12 years of age and decreasing the proportion of symptoms for the diagnosis, DSM-5 appears to have made it easier to meet criteria for ADHD (Batstra & Frances, 2012b). Even more controversial was the decision in DSM-5 to remove the bereavement criterion for major depression, allowing individuals to be diagnosed with this condition as soon as 2 weeks following the death of a loved one (APA, 2013).

We believe the concerns of Frances and others regarding the potential overmedicaliza- tion of normality are important and worth raising. Nevertheless, the ultimate question is whether the changes in DSM-5 increase or decrease the construct validity of the resultant disorders, not whether they alter the prevalence of individuals diagnosed with these disorders (cf. Batstra & Frances, 2012a). The answer to the latter question will surely vary by disorder, and at present awaits clarification in light of future research.

NEGLECT OF THE ATTENUATION PARADOX

Much of the impetus behind DSM-III was the laudable attempt to increase the reliability of psychiatric diagnosis and, thereby, place the fields of psychiatry and clinical psychology on firmer scientific footing. The importance of the reliability of diagnostic categories continues to be recognized in DSM-5. Nevertheless, reliability is only a means to an end,

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

CRITICISMS OF THE CURRENT CLASSIFICATION SYSTEM 23

namely validity; moreover, as noted earlier, validity is limited not by reliability per se, but by its square root (Meehl, 1986). Therefore, diagnoses of even modest reliability can, in principle, achieve high levels of validity.

Ironically, efforts to achieve higher reliability, especially internal consistency, can sometimes produce decreases in validity, a phenomenon that Loevinger (1957) referred to as the attenuation paradox (also see Clark & Watson, 1995). This paradox can result when an investigator uses a narrowly circumscribed pool of items to capture a broad and multifaceted construct. In such a case, the measure of the construct may exhibit high internal consistency yet low validity, because it does not adequately tap the full breadth and richness of the construct.

Some authors have argued that this state of affairs occurred with several DSM diagnoses. Putting it a bit differently, they have suggested that DSM-III and its descendants sacrificed validity at the altar of reliability (Vaillant, 1984). For example, the current DSM diagnosis of antisocial personality disorder (ASPD) is intended to assess the core interpersonal and affective features of psychopathic personality (psychopathy) delineated by Cleckley (1941), Karpman (1948), and others. Indeed, the accompanying text of DSM-IV even referred misleadingly to ASPD as synonymous with psychopathy (APA, 2000, p. 702). Because the developers of DSM-III (APA, 1980) were concerned that the personality features of psychopathy—such as guiltlessness, callousness, and self-centeredness—were difficult to assess reliably, they opted for a diagnosis emphasizing overt and easily agreed on antisocial behaviors—such as vandalism, stealing, and physical aggression (Hare, 2003; Lilienfeld, 1994). These changes may have resulted in a diagnosis with greater internal consistency and interrater reliability than the more traditional construct of psychopathy (although evidence for this possibility is lacking). Nevertheless, they may have also resulted in a diagnosis with lower validity, because the DSM diagnosis of ASPD largely fails to assess the personality features central to psychopathy (Lykken, 1995; Skeem, Polaschek, Patrick, & Lilienfeld, 2011). Indeed, accumulating evidence suggests that measures of ASPD are less valid for predicting a number of theoretically meaningful variables—including laboratory indicators—than are measures of psychopathy (Hare, 2003; also see Vaillant, 1984, for a discussion of the reliability trade-off in the case of the DSM-III diagnosis of schizophrenia).

UNSUPPORTED RETENTION OF A CATEGORICAL MODEL

Technically, the DSM is agnostic on the question of whether psychiatric diagnoses are truly categories in nature, or what Meehl (Meehl & Golden, 1982) termed taxa, as opposed to continua or dimensions. Taxa differ from normality in kind, whereas dimensions differ in degree. Pregnancy is a taxon, as a woman cannot be slightly pregnant; in contrast, height is almost always a dimension (although certain rare taxonic conditions, like hormonal abnormalities, can lead to heights that differ qualitatively from the general population). The opening pages of DSM-IV state: “There is no assumption that each category of mental disorder is a completely discrete entity with absolute boundaries dividing it from other mental disorders or from no mental disorder” (p. xxxi). Yet at the measurement level, the DSM embraces an exclusively categorical model, classifying individuals as either meeting criteria for a disorder or not meeting them. In a highly contentious move, DSM-5 passed on an opportunity to embrace a dimensional model of personality disorders,

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

24 ISSUES IN DIAGNOSIS

leaving the current categorical model in place and relegating a proposed dimensional alternative to Section III of the manual (devoted to provisional criterion sets meriting future consideration).

The DSM categorical model is problematic for at least two reasons. First, there is grow- ing evidence from taxometric analyses (Meehl & Golden, 1982)—namely, those that allow researchers to ascertain whether a single observed distribution is decomposable into multiple independent distributions—that many or even most DSM diagnoses are under- pinned by dimensions rather than taxa (Kendell & Jablensky, 2003), with schizophrenia and schizophrenia-spectrum disorders being notable probable exceptions (Lenzenweger & Korfine, 1992). This is particularly true for most personality disorders (Cloninger, 2009; Trull & Durrett, 2005), including antisocial personality disorder (Marcus, Lilienfeld, Edens, & Poythress, 2006). Even many or most other mental disorders—such as major depression (Slade & Andrews, 2005); social anxiety disorder (Kollman, Brown, Liver- ant, & Hofmann, 2006); and ADHD (Marcus, Norris, & Coccaro, 2012)—appear to be dimensional as opposed to taxonic in structure.

Second, setting aside the ontological issue of taxonicity versus dimensionality, there is good evidence that measuring most disorders (especially personality disorders) dimen- sionally by using the full range of scores almost always results in higher correlations with external validating variables than does measuring them categorically in an all-or-none fashion (Craighead, Sheets, Craighead, & Madsen, 2011; Markon, Chmielewski, & Miller, 2011; Ullrich, Borkenau, & Marneros, 2001). Such findings are not surprising given that artificial dichotomization of variables almost always results in a loss of information and, hence, statistical power (Cohen, 1983; MacCallum, Zhang, Preacher, & Rucker, 2002).

The DSM: Quo Vadis?

In some respects, DSM-III-R and DSM-IV were disappointments, as they did not resolve many of the serious problems endemic to DSM-III (Ghaemi, 2003). If anything, comorbid- ity in DSM-III-R and DSM-IV mushroomed due to the dismantling of many hierarchical exclusion rules (Lilienfeld & Waldman, 2004). Moreover, some diagnostic categories (e.g., dependent personality disorder) of questionable validity remained. It is too early to tell whether DSM-5 will help to resolve these and other problems. DSM-5 presents both challenges and opportunities: challenges because many conceptual and method- ological quandaries regarding psychiatric diagnosis remain unresolved, and opportunities because a new manual opens the door for novel approaches to the classification of psychopathology.

With these considerations in mind, we sketch out two promising future directions for psychiatric diagnosis: adoption of a dimensional approach and the incorporation of endophenotypic markers into psychiatric diagnosis (see Widiger & Clark, 2000, for other proposals for DSM-5 and future DSMs).

A DIMENSIONAL APPROACH

The accumulating evidence for the dimensionality of many psychiatric conditions, particularly personality disorders, has led many authors to suggest replacing or at least supplementing the DSM with a set of dimensions derived from the basic science of personality (Krueger et al., 2011; Widiger & Clark, 2000). One early candidate for a

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE DSM: QUO VADIS? 25

dimensional model is the five-factor model (FFM; Goldberg, 1993), which consists of five major dimensions that have emerged repeatedly in factor analyses of omnibus (broad) measures of personality: extraversion, neuroticism, agreeableness, conscientiousness, and openness to experience (the FFM, incidentally, can easily be recalled using the waterlogged mnemonics of OCEAN or CANOE). These five dimensions also contain lower-order facets that provide a fine-grained description of personality; for example, the FFM dimension of extraversion contains facets of warmth, gregariousness, assertiveness, excitement seeking, and so on (Costa & McCrae, 1992).

The framers of DSM-5 considered a dimensional model for personality disorders influenced substantially by the work of Harkness (see Harkness & McNulty, 1994) and others. In this model, five broad dimensions of antagonism, detachment, negative affectivity (similar to but broader than neuroticism), disinhibition, and psychoticism would have been used to describe all personality variation in the abnormal range. Nevertheless, this bold proposal was ultimately vetoed by the American Psychiatric Association board of trustees, in part because its clinical feasibility was deemed to be insufficiently demonstrated. As noted earlier, however, these dimensions appear in Section III of the current manual in an effort to encourage further research with an eye toward DSM-6.

In addition to clinical feasibility, there are other potential objections to a dimensional model. For example, there is disagreement regarding both the precise nature and number of the personality dimensions to be used, with some authors advocating for alternative (e.g., three-dimensional) models. Another objection derives from the often neglected distinction between basic tendencies and characteristic adaptations in personality psy- chology (Harkness & Lilienfeld, 1997; McCrae & Costa, 1995). Basic tendencies are core personality traits, whereas characteristic adaptations are the behavioral manifestations of these traits. A large body of personality research suggests that basic tendencies can often be expressed in a wide variety of different characteristic adaptations depending on the upbringing, interests, cognitive skills, and other personality traits of the individual. For example, the scores of firefighters on a well-validated measure of the personality trait of sensation seeking (a construct closely related to, although broader than, risk taking) are significantly higher than those of college students, but comparable to those of incarcerated prisoners (Zuckerman, 1994). This finding dovetails with the notion that the same basic tendency, in this case sensation seeking, can be expressed in either socially constructive or destructive outlets, depending on yet unidentified moderating influences.

The distinction between basic tendencies and characteristic adaptations implies that personality dimensions may never be sufficient to capture the full variance in personality disorders. This is because these dimensions (basic tendencies) do not adequately assess many key aspects of psychopathological functioning, many of which can be viewed as maladaptive characteristic adaptations (Sheets & Craighead, 2007). This theoretical conjecture is corroborated by findings that the FFM dimensions do not account for a sizable chunk of variance in many DSM personality disorders. For example, in one study the correlations between FFM prototype scores of DSM personality disorders (derived from expert ratings of the FFM facets most closely associated with each disorder) and structured interview-based measures of these disorders were high for some disorders (e.g., avoidant personality disorder, r = .67) and modest and even negligible for others (e.g., obsessive-compulsive disorder, r = .13; Miller, Reynolds, & Pilkonis, 2004). The latter finding may reflect the fact that some obsessive-compulsive traits—such as perfectionism—may be adaptive in certain settings and, therefore, may not lead inevitably to personality pathology. Moreover, Skodol et al. (2005) reported that the

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

26 ISSUES IN DIAGNOSIS

dimensions of the Schedule for Nonadaptive and Adaptive Personality (SNAP; Clark, 1993)—a measure that assesses many pathological behaviors associated with personality disorders—displayed incremental validity above and beyond the FFM dimensions in distinguishing among DSM personality disorders (also see Reynolds & Clark, 2001). This finding suggests that the FFM overlooks crucial distinctions captured by the SNAP, perhaps in part because the SNAP assesses not only basic tendencies but also the maladaptive characteristic adaptations of many personality disorders (Lilienfeld, 2005).

The findings reviewed here imply that a dimensional model may be useful in capturing core features of many DSM personality disorders. Nevertheless, they raise the possibility that personality dimensions may not be sufficient by themselves to capture personality pathology, because they cannot tell us whether individuals’ behavioral adaptations to these dimensions are adaptive or maladaptive, nor the phenotypic (behavioral) manifestations these adaptations have assumed.

ENDOPHENOTYPIC MARKERS

As noted earlier, considerable recent interest has focused on the use of endophenotypes in the validation of psychiatric diagnoses (Andreasen, 1995; Waldman, 2005). Nevertheless, endophenotypic markers have thus far been excluded from DSM diagnostic criterion sets, which consist almost entirely of the classical signs and symptoms of disorders (exophenotypes). This omission is noteworthy, because endophenotypes may lie closer to the etiology of many disorders than exophenotypes do.

This situation may change in coming years with accumulating evidence from studies of biochemistry, brain imaging, and performance on laboratory tasks; this evidence holds the promise of identifying more valid markers of certain mental disorders (Widiger & Clark, 2000). To take just two examples, many impulse control disorders (e.g., pathological gambling, intermittent explosive disorder) appear to be associated with low levels of serotonin metabolites (Moeller, Barratt, Dougherty, Schmitz, & Swann, 2001) and major depression is frequently associated with left frontal hypoactivation (Henriques & Davidson, 1991).

Nevertheless, at least two potential obstacles confront the use of endophenotypic markers in psychiatric diagnosis, the first conceptual and the second empirical. First, the widespread assumption that endophenotypic markers are more closely linked to underlying etiological processes than exophenotypic markers (Kihlstrom, 2002) is just that: an assumption. For example, the well-replicated finding that diminished amplitude of the P300 (a brain event–related potential appearing approximately 300 milliseconds following stimulus onset) is dependably associated with externalizing disorders—such as conduct disorder and substance dependence (Patrick et al., 2006)—could reflect the fact that P300 is merely a sensitive indicator of attention. As a consequence, diminished P300 amplitude could be a downstream consequence of the inattention and low levels of motivation often associated with externalizing disorders. This possibility would not necessarily negate the incorporation of P300 amplitude into diagnostic criterion sets, although it could raise questions concerning its specificity to externalizing disorders, let alone specific externalizing disorders.

Second, no endophenotypic markers yet identified are close to serving as inclusion tests for their respective disorders. Even smooth pursuit eye movement dysfunction, which is perhaps the most dependable biological marker of schizophrenia, is present only

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

SUMMARY AND FUTURE DIRECTIONS 27

in anywhere from 40% to 80% of patients with schizophrenia, so it would miss many individuals with the disorder. It may come closer, however, to serving as a good exclusion test, as it is present in only about 10% of normal individuals (Clementz & Sweeney, 1990; Keri & Janka, 2004). Thus, although endophenotypic markers may eventually add to the predictive efficiency of some diagnostic criteria sets, they are likely to be fallible indicators, just like traditional signs and symptoms. These markers also hold out the hope of assisting in the identification of more etiologically pure subtypes of disorders; for example, schizophrenia patients with abnormal smooth pursuit eye movements may prove to be separable in important ways from other patients with this disorder.

Until recently, most of the proposals to implement endophenotypic markers were limited to supplementing the diagnosis of existing psychiatric categories, such as major depression or bipolar I disorder. A more radical proposal emanates from the recent initiative supported by the National Institute of Mental Health (NIMH) to develop Research Domain Criteria (RDoC) as a full-fledged alternative to the DSM and similar diagnostic manuals. At this point in time, RDoC is more of an envisioned research approach than a proposed system. Nevertheless, its goal is to identify well-established psychobiological systems that undergird psychopathology (Morris & Cuthbert, 2012), along with promising markers of these systems. Examples of such systems might include reward systems, fear systems, impulse control systems, and working memory. In turn, each of these systems could be measured using indicators at different levels of analysis, including observable behavior, self-report measures, laboratory measures, and brain imaging findings (Insel et al., 2010; Sanislow et al., 2010), Ultimately, such a system could supplement or even supplant the extant DSM system, but as of this writing progress along these lines remains preliminary.

Summary and Future Directions

We conclude the chapter with 10 take-home messages:

1. A systematic system of psychiatric classification is a prerequisite for psychiatric diagnosis.

2. Psychiatric diagnoses serve important, even essential, communicative functions.

3. A valid psychiatric diagnosis gives us new information—for example, it tells about the diagnosed individual’s probable family history, performance on laboratory tests, natural history, and perhaps response to treatment—and it also distinguishes that person’s diagnosis from other, related diagnoses.

4. The claim that mental illness is a myth rests on a misunderstanding of the role of lesions in medical disorders.

5. Prevalent claims to the contrary, psychiatric diagnoses often achieve adequate levels of reliability and validity, and do not typically pigeonhole or stigmatize individuals when correctly applied.

6. There is no clear consensus on the correct definition of mental disorder, and some authors have suggested that the higher-order concept of mental disorder is intrinsically undefinable. Even if true, this should have no effect on the sci- entific investigation, assessment, or treatment of specific mental disorders (e.g., schizophrenia, panic disorder), which undeniably exist.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

28 ISSUES IN DIAGNOSIS

7. Early versions of the diagnostic manual (DSM-I and DSM-II) were problematic because they provided clinicians and researchers with minimal guidance for establishing diagnoses and required high levels of subjective judgment and clinical inference.

8. DSM-III, which appeared in 1980, helped to alleviate this problem by providing diagnosticians with explicit diagnostic criteria, algorithms (decision rules), and hierarchical exclusion criteria, leading to increases in the reliability of many psychiatric diagnoses.

9. The current classification system, DSM-5, is a clear advance over DSM-I and DSM-II. Nevertheless, from initial reports, DSM-5 continues to be plagued by a variety of problems, especially extensive comorbidity, reliable diagnoses that are nevertheless of questionable validity, and retention of a categorical model in the absence of compelling scientific evidence.

10. Fruitful potential directions for psychiatric classification include a dimensional model of personality to replace or supplement the existing categorical system of personality disorders and the adoption of endophenotypic markers for diagnostic purposes.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

References

Alliance of Psychoanalytic Organizations. (2006). Psy- chodynamic diagnostic manual. Silver Spring, MD: Alliance of Psychoanalytic Organizations.

Amador, X. F., & Paul-Odouard, R. (2000). Defend-

ing the Unabomber: Anosognosia in schizophrenia.

Psychiatric Quarterly, 71, 363–371. American Psychiatric Association. (1952). Diagnostic

and statistical manual of mental disorders. Washing- ton, DC: Author.

American Psychiatric Association. (1968). Diagnostic and statistical manual of mental disorders (2nd ed.). Washington, DC: Author.

American Psychiatric Association. (1980). Diagnostic and statistical manual of mental disorders (3rd ed.). Washington, DC: Author.

American Psychiatric Association. (1987). Diagnostic and statistical manual of mental disorders (3rd ed., rev.). Washington, DC: Author.

American Psychiatric Association. (1994). Diagnostic and statistical manual of mental disorders (4th ed.). Washington, DC: Author.

American Psychiatric Association. (2000). Diagnostic and statistical manual of mental disorders (4th ed., text rev.). Washington, DC: Author.

American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). Washington, DC: Author.

Andreasen, N. C. (1995). The validation of psychiatric

diagnosis: New models and approaches. American Journal of Psychiatry, 152, 161–162.

Barlow, D. H. (2001). Anxiety and its disorders: The nature and treatment of anxiety and panic (2nd ed.). New York, NY: Guilford Press.

Batstra, L., & Frances, A. (2012a). Diagnostic inflation:

Causes and a suggested cure. Journal of Nervous and Mental Disease, 200, 474–479.

Batstra, L., & Frances, A. (2012b). Holding the

line against diagnostic inflation in psychiatry. Psy- chotherapy and Psychosomatics, 81, 5–10.

Bayer R., & Spitzer, R. L. (1982). Edited correspon-

dence on the status of homosexuality in DSM-III. Journal of the History of the Behavioral Sciences, 18, 32–52.

Benton, A. L. (1992). Gerstmann’s syndrome. Archives of Neurology, 49, 445–447.

Berkson, J. (1946). Limitations of the application of

fourfold table analysis to hospital data. Biometrics Bulletin, 2, 47–53.

Blashfield, R., & Burgess, D. (2007). Classification

provides an essential basis for organizing mental dis-

orders. In S. O. Lilienfeld & W. T. O’Donohue (Eds.),

The great ideas of clinical science: 17 principles that every mental professional should understand (pp. 93–118). New York, NY: Routledge.

Borsboom, D., Cramer, A. O., Kievit, R. A., Zand

Scholten, A., & Franic, S. (2009). The end of con-

struct validity. In R. W. Lissitz (Ed.), The concept of validity: Revisions, new directions, and applica- tions (pp. 135–170). Charlotte, NC: Information Age Publishing.

Caplan, P. J. (1995). They say you’re crazy: How the world’s most powerful psychiatrists decide who’s normal. Reading, MA: Addison-Wesley.

Caron, C., & Rutter, M. (1991). Comorbidity in

child psychopathology: Concepts, issues and research

strategies. Journal of Child Psychology and Psychi- atry, 32, 1063–1080.

Clark, L. A. (1993). Manual for the schedule for non-adaptive and adaptive personality (SNAP). Min- neapolis: University of Minnesota Press.

Clark, L. A., & Watson, D. (1995). Constructing valid-

ity: Basic issues in objective scale development.

Psychological Assessment, 7, 309–319. Cleckley, H. (1941). The mask of sanity. St. Louis, MO:

Mosby.

Clementz, B. A., & Sweeney, J. A. (1990). Is eye move-

ment dysfunction a biological marker for schizophre-

nia? A methodological review. Psychological Bul- letin, 108, 77–92.

Cloninger, C. R. (2009). Foreword. In W. O’Donohue,

K. A. Fowler, & S. O. Lilienfeld (Eds.), Personal- ity disorders: Toward the DSM-V (pp. vii–xv). Los Angeles, CA: Sage.

Cohen, H. (1981). The evolution of the concept of

disease. In A. L. Caplan, H. T. Engelhardt, Jr., &

J. J. McCartney (Eds.), Concepts of health and dis- ease: Interdisciplinary perspectives (pp. 209–220). Reading, MA: Addison-Wesley.

Cohen, J. (1983). The cost of dichotomization. Applied Psychological Measurement, 7, 249–253.

Compton, W. M., & Guze, S. B. (1995). The neo-

Kraepelinian revolution in psychiatric diagnosis.

European Archives of Psychiatry and Clinical Neu- roscience, 245, 196–201.

Cooper, R. (2011). Mental health and disorder. In H. T.

Have, R. Chadwick, & E. M. Meslin (Eds.), The

29

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

30 ISSUES IN DIAGNOSIS

SAGE handbook of health care ethics (pp. 251–260). London, England: Sage Publications.

Cornez-Ruiz, S., & Hendricks, B. (1993). Effects of

labeling and ADHD behaviors on peer and teacher

judgments. Journal of Educational Research, 86, 349–355.

Costa, P. T., Jr., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO-PI-R) and NEO Five- Factor Inventory (NEO-FFI) professional manual. Odessa, FL: Psychological Assessment Resources.

Craighead, W. E., Sheets, E. S., Craighead, L. W., &

Madsen, J. W. (2011). Recurrence of MDD: A

prospective study of personality pathology and cog-

nitive distortions. Personality Disorders: Theory, Research, and Treatment, 2, 83–97.

Cramer, A. O., Waldorp, L. J., van der Maas, H. L., &

Borsboom, D. (2010). Comorbidity: A network

perspective. Behavioral and Brain Sciences, 33, 137–150.

Cronbach, L. J. (1951). Coefficient alpha and the inter-

nal structure of tests. Psychometrika, 16, 297–335. Cronbach, L. J., & Meehl, P. E. (1955). Construct valid-

ity in psychological tests. Psychological Bulletin, 52, 281–302.

Draguns, J. G., & Tanaka-Matsumi, J. (2003). Assess-

ment of psychopathology across and within cultures:

Issues and findings. Behaviour Research and Ther- apy, 41, 755–776.

Drake, R. E., & Wallach, M. A. (2007). Is comorbidity a

psychological science? Clinical Psychology: Science and Practice, 14, 20–22.

du Fort, G. G., Newman, S. C., & Bland, R. C.

(1993). Psychiatric comorbidity and treatment seek-

ing: Sources of selection bias in the study of clinical

populations. Journal of Nervous and Mental Disease, 181, 467–474.

Eysenck, H. J., Wakefield, J., & Friedman, A. (1983).

Diagnosis and clinical assessment: The DSM-III. Annual Review of Psychology, 34, 167–193.

Faust, D., & Miner, R. A. (1986). The empiricist in

his new clothes: DSM-III in perspective. American Journal of Psychiatry, 143, 962–967.

Feighner, J., Robins, E., Guze, S., Woodruff, R.,

Winokur, G., & Munoz, R. (1972). Diagnostic criteria

for use in psychiatric research. Archives of General Psychiatry, 26, 57–63.

First, M. B., Spitzer, R. L., Gibbon, M., & Williams,

J. B. W. (2002). Structured clinical interview for DSM-IV-TR Axis I disorders, research version, patient edition. New York, NY: Biometrics Research, New York State Psychiatric Institute.

Frances, A. J. (2012, December 2). The DSM-5 is not a bible: Ignore its ten worst changes. Psychology Today. Retrieved from http://www.psychologytoday .com/blog/dsm5-in-distress/201212/dsm-5-is-guide-

not-bible-ignore-its-ten-worst-changes

Frances, A. J., & Widiger, T. (2012). Psychiatric diag-

nosis: Lessons from the DSM-IV past and cautions for the DSM-5 future. Annual Review of Clinical Psychology, 8, 109–130.

Garb, H. N. (1998). Studying the clinician: Judgment research and psychological assessment. Washington, DC: American Psychological Association.

Garber, J., & Strassberg, Z. (1991). Construct valid-

ity: History and application to developmental psy-

chopathology. In W. M. Grove & D. Cicchetti (Eds.),

Personality and psychopathology (pp. 218–258). Minneapolis: University of Minnesota Press.

Ghaemi, N. (2003). The concepts of psychiatry: A plu- ralistic approach to the mind and mental illness. Baltimore, MD: Johns Hopkins University Press.

Goldberg, L. R. (1993). The structure of phenotypic per-

sonality traits. American Psychologist, 48, 266–275. Goodwin, D. W., & Guze, S. B. (1996). Psychiatric

diagnosis (5th ed.). New York, NY: Oxford Univer- sity Press.

Gorenstein, E. E. (1992). The science of mental illness. San Diego, CA: Academic Press.

Gottesman, I. I., & Gould, T. D. (2003). The endopheno-

type concept in psychiatry: Etymology and strategic

intentions. American Journal of Psychiatry, 160, 636–645.

Gough, H. (1971). Some reflections on the mean-

ing of psychodiagnosis. American Psychologist, 26, 160–167.

Grob, G. N. (1991). Origins of DSM-I: A study in appearance and reality. American Journal of Psychi- atry, 148, 421–431.

Guilford, J. P. (1936). Psychometric methods. New York, NY: McGraw-Hill.

Hare, R. D. (2003). Manual for the Revised Psychopathy Checklist (2nd ed.). Toronto, Canada: Multi-Health Systems.

Harkness, A. R., & Lilienfeld, S. O. (1997). Individual

differences science for treatment planning: Personal-

ity traits. Psychological Assessment, 9, 349–360. Harkness, A. R., & McNulty, J. L. (1994). The Person-

ality Psychopathology Five (PSY-5): Issue from the

pages of a diagnostic manual instead of a dictionary.

In S. Strack & M. Lorr (Eds.), Differentiating normal and abnormal personality (pp. 291–315). New York, NY: Springer.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

REFERENCES 31

Harris, M. J., Milich, R., Corbitt, E. M., Hoover,

D. W., & Brady, M. (1992). Self-fulfilling effects of

stigmatizing information on children’s social interac-

tions. Journal of Personality and Social Psychology, 63, 41–50.

Henriques, J. B., & Davidson, R. J. (1991). Left frontal

hypoactivation in depression. Journal of Abnormal Psychology, 100, 535–545.

Houts, A. C. (2001). The diagnostic and statistical

manual’s new white coat and circularity of plau-

sible dysfunctions: Response to Wakefield, part 1.

Behavior Research and Therapy, 39, 315–345. Insel, T., Cuthbert, B., Garvey, M., Heinssen, R.,

Kozak, M., Pine, D. S., . . . Wang, P. (2010). Research

Domain Criteria (RDoC): Developing a valid diag-

nostic framework for research on mental disorders.

American Journal of Psychiatry, 167, 748–751. Joiner, T. (2006). Why people die by suicide. Cam-

bridge, MA: Harvard University Press.

Kaplan, M. H., & Feinstein, A. R. (1974). The impor-

tance of classifying initial co-morbidity in evaluating

the outcome of diabetes mellitus. Journal of Chronic Diseases, 27, 387–404.

Karpman, B. (1948). The myth of the psychopathic

personality. American Journal of Psychiatry, 104, 523–524.

Kazdin, A. E. (1983). Psychiatric diagnosis, dimensions

of dysfunction, and child behavior therapy. Behavior Therapy, 14, 73–99.

Kendell, R., & Jablensky, A. (2003). Distinguishing

between the validity and utility of psychiatric diag-

noses. American Journal of Psychiatry, 160, 4–12. Kendell, R. E. (1975). The concept of disease and

its implications for psychiatry. British Journal of Psychiatry, 127, 305–315.

Kendler, K. S. (1980). The nosologic validity of para-

noia (simple delusional disorder): A review. Archives of General Psychiatry, 37, 699–706.

Keri, S., & Janka, Z. (2004). Critical evaluation of cog-

nitive dysfunctions as endophenotypes of schizophre-

nia. Acta Psychiatrica Scandinavica, 110, 83–91. Kihlstrom, J. F. (2002). To honor Kraepelin . . . :

From symptoms to pathology in the diagnosis of

mental illness. In L. Beutler & M. Malik (Eds.),

Rethinking the DSM: A psychological perspective (pp. 279–303). Washington, DC: American Psycho-

logical Association.

Kim, J., Park, S., & Blake, R. (2011). Perception of

biological motion in schizophrenia and healthy indi-

viduals: A behavioral and fMRI study. PloS One, 6(5), e19971.

Kirk, S. A., & Kutchins, H. (1992). The selling of DSM: The rhetoric of science in psychiatry. New York, NY: Aldine de Gruy.

Klein, D., & Riso, L. P. (1993). Psychiatric disor-

ders: Problems of boundaries and comorbidity. In

C. G. Costello (Ed.), Basic issues in psychopathology (pp. 19–66). New York, NY: Guilford Press.

Kleinknecht, R. A., Dinnel, D. L., Tanouye-Wilson,

S., & Lonner, W. J. (1994). Cultural variation in

social anxiety and phobia: A study of taijin kyofusho. Behavioral Therapist, 17(8), 175–178.

Klerman, G. (1984). The advantages of DSM-III. Amer- ican Journal of Psychiatry, 141, 539–542.

Kollman, D. M., Brown, T. A., Liverant, G. I., & Hof-

mann, S. G. (2006). A taxometric investigation of the

latent structure of social anxiety disorder in outpa-

tients with anxiety and mood disorders. Depression and Anxiety, 23, 190–199.

Kraupl Taylor, F. (1971). A logical analysis of medi-

cophysiological concept of disease. Psychological Medicine, 1, 356–364.

Krueger, R. F., Eaton, N. R., Clark, L. A., Watson, D.,

Markon, K. E., Derringer, J., . . . Livesley, W. J.

(2011). Deriving an empirical structure of person-

ality pathology for DSM-5. Journal of Personality Disorders, 25, 170–191.

Lenzenweger, M. F., & Korfine, L. (1992). Confirming

the latent structure and base rate of schizotypy: A tax-

ometric analysis. Journal of Abnormal Psychology, 101, 567–571.

Lief, A. A. (Ed.). (1948). The commonsense psychiatry of Dr. Adolf Meyer: Fifty-two selected papers. New York, NY: McGraw-Hill.

Lilienfeld, S. O. (1994). Conceptual problems in the

assessment of psychopathy. Clinical Psychology Review, 14, 17–38.

Lilienfeld, S. O. (1995). Seeing both sides: Classic con- troversies in abnormal psychology. Pacific Grove, CA: Brooks/Cole.

Lilienfeld. S. O. (2003). Comorbidity between and

within childhood externalizing and internalizing dis-

orders: Reflections and directions. Journal of Abnor- mal Child Psychology, 31, 285–291.

Lilienfeld, S. O. (2005). Longitudinal studies of per-

sonality disorders: Four lessons from personality

psychology. Journal of Personality Disorders, 19, 547–556.

Lilienfeld, S. O. (2013). Is psychopathy a syndrome?

Commentary on Marcus, Fulton, and Edens. Per- sonality Disorders: Theory, Practice, and Research, 4(1), 85–86.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

32 ISSUES IN DIAGNOSIS

Lilienfeld, S. O., & Marino, L. (1995). Mental disor-

der as a Rochian concept: A critique of Wakefield’s

“harmful dysfunction” analysis. Journal of Abnormal Psychology, 104, 411–420.

Lilienfeld, S. O., & Marino, L. (1999). Essentialism

revisited: Evolutionary theory and the concept of a

mental disorder. Journal of Abnormal Psychology, 108, 400–411.

Lilienfeld, S. O., VanValkenburg, C., Larntz, K., &

Akiskal, H. S. (1986). The relationship of histrionic

personality to antisocial personality and somatiza-

tion disorders. American Journal of Psychiatry, 143, 718–722.

Lilienfeld, S. O., & Waldman, I. D. (2004). Comorbidity

and Chairman Mao. World Psychiatry, 3, 26–27. Lilienfeld, S. O., Waldman, I. D., & Israel, A. C. (1994).

A critical note on the use of the term and concept of

“comorbidity” in psychopathology research. Clinical Psychology: Science and Practice, 1, 71–83.

Link, B. G., & Cullen, F. T. (1990). The labeling

theory of mental disorder: A review of the evi-

dence. Research in Community and Mental Health, 6, 75–105.

Lobbestael, J., Leurgans, M., & Arntz, A. (2011). Inter-

rater reliability of the Structured Clinical Interview

for DSM-IV Axis I disorders (SCID I) and Axis II disorders (SCID II). Clinical Psychology & Psy- chotherapy, 18, 75–79.

Loevinger, J. (1957). Objective tests as instruments

of psychological theory. Psychological Reports, 3, 635–694.

Lykken, D. T. (1995). The antisocial personalities. Hillsdale, NJ: Erlbaum.

Lynam, D. R., & Miller, J. D. (2012). Fearless dom-

inance and psychopathy: A response to Lilienfeld

et al. Personality Disorders: Theory, Research, and Treatment, 3, 341–353.

MacCallum, R. C., Zhang, S., Preacher, K. J., & Rucker,

D. D. (2002). On the practice of dichotomization

of quantitative variables. Psychological Methods, 7, 19–40.

Maffei, C., Fossati, A., Agostoni, I., Barraco, A.,

Bagnato, M., Namia, C., . . . Petrachi, M. (1997).

Interrater reliability and internal consistency of the

Structured Clinical Interview for DSM-IV Axis II personality disorders (SCID-II), version 2.0. Journal of Personality Disorders, 11, 279–284.

Marcus, D. K., Lilienfeld, S. O., Edens, J. F., &

Poythress, N. G. (2006). Is antisocial personality

disorder continuous or categorical? A taxometric

analysis. Psychological Medicine, 36, 1571–1581.

Marcus, D. K., Norris, A. L., & Coccaro, E. F. (2012).

The latent structure of attention deficit/hyperactivity

disorder in an adult sample. Journal of Psychiatric Research, 46, 782–789.

Markon, K. E., Chmielewski, M., & Miller, C. J. (2011).

The reliability and validity of discrete and continuous

measures of psychopathology: A quantitative review.

Psychological Bulletin, 137, 856–879. Martel, M. M., Von Eye, A., & Nigg, J. T. (2010).

Revisiting the latent structure of ADHD: Is there a

“g” factor? Journal of Child Psychology and Psychi- atry, 51, 905–914.

Matarazzo, J. D. (1983). The reliability of psychiatric

and psychological diagnosis. Clinical Psychology Review, 3, 103–145.

Mayes, R., & Horwitz, A. V. (2005). DSM-III and the revolution in the classification of mental illness.

Journal of the History of the Behavioral Sciences, 41, 249–267.

Mayr, E. (1982). The growth of biological thought: Diversity, evolution, and inheritance. Cambridge, MA: Belknap Press.

McCann, J. T., Shindler, K. L., & Hammond, T. R.

(2003). The science and pseudoscience of expert tes-

timony. In S. O. Lilienfeld, J. M. Lohr, & S. J. Lynn

(Eds.), Science and pseudoscience in contemporary clinical psychology (pp. 77–108). New York, NY: Guilford Press.

McCrae, R. R., & Costa, P. T. (1995). Trait explana-

tions in personality psychology. European Journal of Personality, 9, 231–252.

McHugh, P. R., & Slavney, P. R. (1998). The perspec- tives of psychiatry (2nd ed.). Baltimore, MD: Johns Hopkins University Press.

McPartland, J. C., Reichow, B., & Volkmar, F. R.

(2012). Sensitivity and specificity of proposed

DSM-5 diagnostic criteria for autism spectrum dis- order. Journal of the American Academy of Child & Adolescent Psychiatry, 51, 368–383.

Meehl, P. E. (1973). Why I do not attend case con-

ferences. In P. E. Meehl (Ed.), Psychodiagnosis: Selected papers (pp. 225–302). Minneapolis: Uni- versity of Minnesota Press.

Meehl, P. E. (1977). Specific etiology and other forms

of strong influence: Some quantitative meanings.

Journal of Medicine and Philosophy, 2, 33–53. Meehl, P. E. (1986). Diagnostic taxa as open concepts:

Metatheoretical and statistical questions about relia-

bility and construct validity in the grand strategy of

nosological revision. In T. Millon & G. L. Klerman

(Eds.), Contemporary directions in psychopathology:

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

REFERENCES 33

Toward the DSM-IV (pp. 215–231). New York, NY: Guilford Press.

Meehl, P. E. (1990). Schizotaxia as an open concept.

In A. I. Rabin, R. Zucker, R. Emmons, & S. Frank

(Eds.), Studying persons and lives (pp. 248–303). New York, NY: Springer.

Meehl, P. E., & Golden, R. (1982). Taxometric methods.

In P. C. Kendall & J. N. Butcher (Eds.), Hand- book of research methods in clinical psychology (pp. 127–181). New York, NY: Wiley.

Messick, S. (1995). Validity of psychological assess-

ment: Validation of inferences from persons’

responses and performances as scientific inquiry into

score meaning. American Psychologist, 50, 741–749. Meyer, G. J. (1997). Assessing reliability: Critical cor-

rections for a critical examination of the Rorschach

Comprehensive System. Psychological Assessment, 9, 480–489.

Michels, R. (1984). A debate on DSM-III: First rebuttal. American Journal of Psychiatry, 141, 548–553.

Milich, R., McAninich, C. B., & Harris, M. J. (1992).

Effects of stigmatizing information on children’s peer

relations: Believing is seeing. School Psychology Review, 21, 399–408.

Miller, J. D., Reynolds, S. K., & Pilkonis, P. A. (2004).

The validity of the five-factor model prototypes for

personality disorders in two clinical samples. Psy- chological Assessment, 16, 310–322.

Millon, T. (1975). Reflections on Rosenhan’s “On being

sane in insane places.” Journal of Abnormal Psychol- ogy, 84, 456–461.

Moeller, F. G., Barratt, E. S., Dougherty, D. M.,

Schmitz, J. M., & Swann, A. C. (2001). Psychi-

atric aspects of impulsivity. American Journal of Psychiatry, 158, 1783–1793.

Morey, L. C. (1991). Classification of mental disorders

as a collection of hypothetical constructs. Journal of Abnormal Psychology, 100, 289–293.

Morris, S. E., & Cuthbert, B. N. (2012). Research

Domain Criteria: Cognitive systems, neural circuits,

and dimensions of behavior. Dialogues in Clinical Neuroscience, 14, 29–37.

Morrison, J. (1997). When psychological problems mask medical disorders: A guide for psychothera- pists. New York, NY: Guilford Press.

Neese, R., & Williams, G. (1994). Why we get sick. New York, NY: Vintage.

Patrick, C. J., Bernat, E., Malone, S. M., Iacono, W. G.,

Krueger, R. F., & McGue, M. K. (2006). P300 ampli-

tude as an indicator of externalizing in adolescent

males. Psychophysiology, 43, 84–92.

Patrick, C. J., Fowles, D. C., & Krueger, R. F.

(2009). Triarchic conceptualization of psychopathy:

Developmental origins of disinhibition, boldness,

and meanness. Development and Psychopathology, 21(3), 913.

Pelham, W. E., & Bender, M. E. (1982). Peer rela-

tionships in hyperactive children: Description and

treatment. In K. D. Gadow & I. Bialer (Eds.),

Advances in learning and behavioral disabilities (Vol. 1, pp. 365–436). Greenwich, CT: JAI Press.

Pincus, H. A., Tew, J. D., & First, M. B. (2004). Psychi-

atric comorbidity: Is more less? World Psychiatry, 3, 18–23.

Pouissant, A. F. (2002). Is extreme racism a men-

tal illness? Point-counterpoint. Western Journal of Medicine, 176, 4.

Reynolds, S. K., & Clark, L. A. (2001). Predicting

personality disorder dimensions from domains and

facets of the five-factor model. Journal of Personal- ity, 69, 199–222.

Robins, E., & Guze, S. B. (1970). Establishment of diag-

nostic validity in psychiatric illness: Its application

to schizophrenia. American Journal of Psychiatry, 126, 983–987.

Rosch, E. R. (1973). Natural categories. Cognitive Psy- chology, 4, 328–350.

Rosch, E. R., & Mervis, C. B. (1975). Family resem-

blances: Studies in the internal structure of categories.

Cognitive Psychology, 7, 573–605. Rosenhan, D. L. (1973). On being sane in insane places.

Science, 179, 250–258. Rosenhan, D. L., & Seligman, M. E. (1995). Abnormal

psychology. New York, NY: W.W. Norton. Ross, C., & Pam, A. (1996). Pseudoscience in biolog-

ical psychiatry: Blaming the body. New York, NY: Wiley.

Ruscio, J. (2004). Diagnoses and the behaviors they

denote: A critical evaluation of the labeling theory

of mental illness. Scientific Review of Mental Health Practice, 3(1), 5–22.

Sanislow, C. A., Pine, D. S., Quinn, K. J., Kozak, M. J.,

Garvey, M. A., Heinssen, R. K., . . . Cuthbert, B.

N. (2010). Developing constructs for psychopathol-

ogy research: Research Domain Criteria. Journal of Abnormal Psychology, 119, 631–639.

Sarbin, T. R. (1969). On the distinction between social

roles and social types, with special reference to

the hippie. American Journal of Psychiatry, 125, 1024–1031.

Scadding, J. G. (1996). Essentialism and nominalism in

medicine: Logic of diagnosis in disease terminology.

Lancet, 348, 594–596.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

34 ISSUES IN DIAGNOSIS

Schaler, J. A. (Ed.). (2004). Szasz under fire: The psy- chiatric abolitionist faces his critics. Chicago, IL: Open Court.

Scheff, T. (Ed.). (1975). Labeling madness. Englewood Cliffs, NJ: Prentice Hall.

Schmidt, F. L., Le, H., & Ilies, R. (2003). Beyond alpha:

An empirical examination of the effects of different

sources of measurement error on reliability estimates

for measures of individual differences constructs.

Psychological Methods, 8, 206–224. Seitz, S., & Geske, D. (1976). Mothers’ and gradu-

ate trainees’ judgments of children: Some effects of

labeling. American Journal of Mental Deficiency, 81, 362–370.

Selkoe, D. J. (1992). Aging brain, aging mind. Scientific American, 267, 134–142.

Sheets, E., & Craighead, W. E. (2007). Toward an

empirically based classification of personality pathol-

ogy. Clinical Psychology: Science and Practice, 14, 77–93.

Skeem, J. L., & Cooke, D. J. (2010). One measure

does not a construct make: Directions toward rein-

vigorating psychopathy research—Reply to Hare

and Neumann (2010). Psychological Assessment, 22, 455–459.

Skeem, J. L., Polaschek, D. L., Patrick, C. J., & Lilien-

feld, S. O. (2011). Psychopathic personality bridging

the gap between scientific evidence and public pol-

icy. Psychological Science in the Public Interest, 12, 95–162.

Skinner, H. A. (1981). Toward the integration of clas-

sification theory and methods. Journal of Abnormal Psychology, 90, 68–87.

Skinner, H. A. (1986). Construct validation approach

to psychiatric classification. In T. Millon & G. L.

Klerman (Eds.), Contemporary directions in psy- chopathology: Toward the DSM-IV (pp. 307–330). New York, NY: Guilford Press.

Skodol, A. E., Oldham, J. M., Bender, D. S., Dyck,

I. R., Stout, R. L., Morey, L. C., . . . Gunderson,

J. C. (2005). Dimensional representations of DSM-IV

personality disorders: Relationships to functional

impairment. American Journal of Psychiatry, 162(10), 1919–1925.

Slade, T., & Andrews, G. (2005). Latent structure of

depression in a community sample: A taxometric

analysis. Psychological Medicine, 35, 489–497. Slater, L. (2004). Opening Skinner’s box: Great psycho-

logical experiments of the 20th century. New York, NY: W.W. Norton.

Sommers, C. H., & Satel, S. (2005). One nation under therapy: How the helping culture is eroding self- reliance. New York, NY: St. Martin’s.

Spitzer, R. L. (1975). On pseudoscience, logic in

remission, and psychiatric diagnosis: A critique of

Rosenhan’s “On being sane in insane places.” Jour- nal of Abnormal Psychology, 84, 442–452.

Spitzer, R. L., Endicott, J., & Robins, E. (1978).

Research Diagnostic Criteria: Rationale and reliabil-

ity. Archives of General Psychiatry, 35(6), 773–782. Spitzer, R. L., Foreman, J. B. W., and Nee, J. (1979).

DSM-III field trials. American Journal of Psychiatry, 136, 815–820.

Spitzer, R. L., Lilienfeld, S. O., & Miller, M. B. (2005).

Rosenhan revisited: The scientific credibility of Lau-

ren Slater’s pseudopatient diagnosis study 1. Journal of Nervous and Mental Disease, 193, 734–739.

Stuart, S., Pfohl, B., Battaglia, M., Bellodi, L., Grove,

W., & Cadoret, R. (1998). The cooccurrence of

DSM-III-R personality disorders. Journal of Person- ality Disorders, 12, 302–315.

Subotnik, K. L., Nuechterlein, K. H., Ventura, J.,

Gitlin, M. J., Marder, S., Mintz, J., . . . Singh, I. R.

(2011). Risperidone nonadherence and return of pos-

itive symptoms in the early course of schizophrenia.

American Journal of Psychiatry, 168(3), 286–292. Sullivan, P. F., & Kendler, K. S. (1998). The genetic

epidemiology of smoking. Nicotine and Tobacco Research, 1, S51–S57.

Sutton, E. H. (1980). An introduction to human genetics. Philadelphia, PA: Saunders College.

Szasz, T. (1960). The myth of mental illness. American Psychologist, 15, 113–118.

Trull, T. J., & Durrett, C. A. (2005). Categorical and

dimensional models of personality disorder. Annual Review of Clinical Psychology, 1, 355–380.

Ullrich, S., Borkenau, P., & Marneros, A. (2001). Per-

sonality disorders in offenders: Categorical versus

dimensional approaches. Journal of Personality Dis- orders, 15, 442–449.

Vaillant, G. E. (1984). The disadvantages of DSM- III outweigh its advantages. American Journal of Psychiatry, 14, 542–545.

van Praag, H. M. (2000). Nosologomania: A disorder

of psychiatry. World of Biological Psychiatry, 1, 151–158.

Wakefield, J. C. (1992). The concept of mental disor-

der: On the boundary between biological facts and

social values. American Psychologist, 47, 373–388. Wakefield, J. C. (1998). The DSM’s theory-neutral

nosology is scientifically progressive: Response to

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

REFERENCES 35

Follette and Houts. Journal of Consulting and Clini- cal Psychology, 66, 846–852.

Wakefield, J. C. (1999). Evolutionary versus proto-

type analyses of the concept of disorder. Journal of Abnormal Psychology, 108, 374–399.

Wakefield, J. C. (2001). The myth of DSM’s invention of new categories of disorder: Houts’s diagnostic dis-

continuity thesis disconfirmed. Behaviour Research and Therapy, 39, 575–624.

Waldman, I. D. (2005). Statistical approaches to

complex phenotypes: Evaluating neuropsychologi-

cal endophenotypes of attention-deficit/hyperactivity

disorder. Biological Psychiatry, 57, 1347–1356. Waldman, I. D., Lilienfeld, S. O., & Lahey, B. B. (1995).

Toward construct validity in the childhood disrup-

tive behavior disorders: Classification and diagnosis

in DSM-IV and beyond. In T. H. Ollendick & R. J. Prinz (Eds.), Advances in clinical child psychology (Vol. 17, pp. 323–363). New York, NY: Plenum

Press.

Widiger, T. A. (1997). The construct of mental disor-

der. Clinical Psychology: Science and Practice, 4, 262–266.

Widiger, T. A. (2007). Alternatives to DSM-IV: Axis II. In W. O’Donohue, K. A. Fowler, & S. O. Lilienfeld

(Eds.), Personality disorders: Toward the DSM-V (pp. 21–40). Thousand Oaks, CA: Sage.

Widiger, T. A., & Clark, L. A. (2000). Toward DSM-V and the classification of psychopathology. Psycho- logical Bulletin, 126, 946–963.

Widiger, T. A., Frances, A. J., Pincus, H. A., Ross, R.,

First, M. B., & Davis, W. W. (Eds.). (1998). DSM-

IV sourcebook (Vol. 4). Washington, DC: American Psychiatric Press.

Widiger, T. A., Frances, A. J., Spitzer, R. L., &

Williams, J. B. W. (1991). The DSM-III-R person- ality disorders: An overview. American Journal of Psychiatry, 145, 786–795.

Widiger, T. A., & Rogers, J. H. (1989). Prevalence

and comorbidity of personality disorders. Psychiatric Annals, 19, 132–136.

Widiger, T. A., & Trull, T. (1985). The empty debate

over the existence of mental illness: Comments on

Gorenstein. American Psychologist, 40, 468–470. Zimmerman, M. (1994). Diagnosing personality disor-

ders. Archives of General Psychiatry, 51, 225–245. Zimmerman, M., & Mattia, J. I. (2000). Principal and

additional DSM-IV disorders for which outpatients seek treatment. Psychiatric Services, 51, 1299–1304.

Zuckerman, M. (1994). Behavioral expressions and biosocial bases of sensation seeking. New York, NY: Cambridge University Press.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

Chapter 2

Strategies for Evidence-Based Assessment of Children

and Adolescents

Measuring Prediction, Prescription, and Process

ERIC A. YOUNGSTROM AND THOMAS W. FRAZIER

I often say that when you can measure what you are speaking about and express it in numbers you know something about it; but when you cannot measure it, when you cannot express it in numbers, your knowledge is of a meagre and unsatisfactory kind: it may be the beginning of knowledge, but you have scarcely, in your thoughts, advanced to the stage of science, whatever the matter may be.

—Lord Kelvin, Sir William Thomson, “Electrical Units of Measurement” (1883), in Popular Lectures and Addresses (1891), Vol. 1, pp. 80–81

Our goal is to evaluate critically psychological assessment as it applies to develop-mental psychopathology. The litmus test for assessment methods is the extent to which they succeed in answering one of the Three Ps of assessment: Do they predict important criteria? Do they prescribe specific treatments? Do they inform our under- standing of processes in developmental psychopathology? If the method does not address one of these purposes, then it is not clear why we would add it to either a research or a clinical assessment battery. Advances in quantitative methods can contribute much toward bridging the science-practice gap by helping make better use of existing tools.

Background

Assessment represents a paradox in the field of psychology. On the one hand, assess- ment has a long history, and it is arguably the activity most uniquely the province of psychology. Whereas multiple professional disciplines can offer therapy or counseling, or can conduct basic research into social behavior or neuroscience, psychometric and behavioral assessment typically has remained within the guild of professional psycholog- ical activities. Measurement of constructs as diverse as personality, cognitive abilities,

36

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

BACKGROUND 37

academic achievement, behavior problems, family functioning, or quality of life has been a distinctly psychological enterprise.

The paradox is that psychological practice and graduate training have increasingly become disconnected from this core professional function (Archer & Newsom, 2000; Stedman, Hatch, & Schoenfeld, 2001). There are several forces contributing to the current state of torpor in psychological assessment. One major issue is a lack of clear linkage between measurement tools and clinical practices. Well-developed instruments often have painstakingly honed psychometric properties, yet have had unclear validity in terms of guiding treatment or improving outcomes—central questions to the practicing clinician. Cognitive ability tests are an excellent case in point. They have the longest pedigree and most extensive corpus of research of any assessment tool, yet many experts question their treatment utility (Flanagan, McGrew, & Ortiz, 2000).

A second issue has to do with the emergence of managed care. The erosion of service reimbursement, and the challenge to demonstrate cost effectiveness of psychological assessments, has lifted the potential barriers between tool and practice from an intellectual issue to an iceberg against which clinical assessment practices have foundered (Cashel, 2002; Eisman et al., 1998; Piotrowski, 1999).

The history of psychological assessment has also contributed much inertia to current practices. Our current library of assessment devices is based more on convention and habit than on any clear sense that these are the best tools for specific purposes. Much validity evidence for contemporary tests is also recursive: The Wechsler Intelligence Scale for Children—Fourth Edition (Wechsler, 2003), was validated against the Third Edition, which was validated against the WISC-R, which in turn was validated against the WISC, the Binet, and ultimately against the Raven’s Progressive Matrices, the Army Alpha, and other older ability measures (Sattler, 2001). Fewer studies shine a light forward instead of back into the prior validating lineage, showing that the test accomplishes some external criterion task better than competing measures. In the cognitive ability literature, fewer studies show predictive or ecological validity than other criterion correlations (Neisser et al., 1996), yet cognitive ability probably represents the place where more work has been done than anywhere else in the field of developmental psychopathology assessment.

Some degree of conservatism and inertia is appropriate in clinical practice. To use a computer software analogy, the largest portion of the test user “installed base” is familiar with the well-established measures. There also are economies of scale for professionals and institutions when certain tests become industry standards. The most popular tests often (though not always) have the most research available. However, the correlations between popularity and research quantity are imperfect, and even smaller when correlating popularity with the amount of high-quality, clinically relevant research.

Assessment of developmental psychopathology faces further obstacles. One is the inheritance of nosological systems that were developed for adult clinical presentations, and then later adapted for use with adolescents or children (Kazdin & Kagan, 1994). Another disconnect between the field of developmental psychopathology and typical clinical assessment is that developmental psychopathology focuses on understanding processes and mechanisms, whereas adult nosology has consciously adopted a noncausal approach to diagnosis, using phenomenological description instead of mechanisms as a way of organizing observations and classifying presentations (Carson, 1997). Developmental psychopathology recognizes that people grow and change within the context of other systems, such as the family and the broader cultural environment. Understanding growth

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

38 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

requires tools that are developmentally appropriate, deployed at multiple levels of analysis to capture an accurate picture of change.

The following sections provide a brief overview of current practices in assessment and training, and then develop the concepts of the Three Ps of assessment (prediction, prescription, and process). Subsequent sections examine the added considerations imposed by studying developmental phenomena, including the issues of norms, developmentally appropriate modalities and content, continuity and discontinuity of phenomena, and the concepts of equipotentiality and multifinality (Cicchetti & Cohen, 1995; Kazdin & Kagan, 1994). The chapter concludes with recommendations about priority areas for research and also for “technology transfer” from existing research into meaningful changes in clinical practice and clinical training.

A Snapshot of Current Assessment Training and Practice

There is considerable inertia in the choice of assessment tools taught in graduate training programs or internships, and also used in clinical practice. Table 2.1 lists tests used for the assessment of personality and psychopathology in both children and adults, in descending order of use by practicing clinical psychologists. The table also reports the rank for usage by practicing neuropsychologists responding to the same survey (Camara, Nathan, & Puente, 1998), along with a ranking of what percentage of assessment courses cover a particular instrument (based on syllabi and questionnaire responses from 84 doctoral programs) (Childs & Eyde, 2002). The test rankings are highly consistent across levels of training, across disciplines, and over time, although there has been a historical trend for the emphasis on projective techniques to decrease at the predoctoral training level.

One striking feature of the list of instruments is the persistence of the incumbent measures. Although many other measures of cognitive ability are available, and many arguably have technical advantages over the Wechslers, the Wechslers remain the industry standard for training and practice (Sattler, 2001). Projective measures remain high on the list, despite questions about the reliability of the administration or scoring procedures or incremental validity over more easily acquired measures (cf. Meyer & Handler, 1997; Wood, Nezworski, & Stejskal, 1996). Table 2.1 suggests that projectives remain alive and well in clinical practice.

In contrast, many other well-validated instruments have not permeated far into clinical training or practice. The Big Three or Big Five models of personality have amassed considerable research, but none of the personality measures appears in the top dozen assessment tools, despite their demonstrated validity as measures of personality and their clinical relevance (Barnett et al., 2011; Harkness & Lilienfeld, 1997). Similarly, the extensive research on temperament, using both parent report and laboratory measures (e.g., Carey, 1998; Derryberry & Rothbart, 1997; Kagan, 1997b), is not reflected in the modal assessment batteries of practitioners. The Child Behavior Checklist (CBCL; Achenbach & Rescorla, 2001), which has been used in literally thousands of research studies and translated into dozens of languages, appears at number 17 on the list of assessment devices used by clinical psychologists (Camara et al., 1998); and neither the CBCL nor behavior checklists more generally appear in the top 20 tools or techniques covered in graduate assessment courses (Stedman et al., 2001). It is possible that there have been changes in training or usage in the decade since these surveys were conducted, but it is sobering to consider that the CBCL was initially published in 1983 and thus

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

T A

B L

E 2.

1 C

om m

on ly

U se

d A

ss es

sm en

t M

et h

od s

fo r

P er

so n

al it

y an

d P

sy ch

op at

h ol

og y

A ss

es sm

en t

T es

t

C li

n ic

a l

P sy

ch o

lo g

y

R a

n k

(n )

(C a

m a

ra et

a l.

, 1

9 9

8 )

N eu

ro p

sy ch

o lo

g y

R a

n k

(n )

(C a

m a

ra et

a l.

, 1

9 9

8 )

T ra

in in

g R

a n

k (P

er ce

n ta

g e

o f

P ro

g ra

m s

T ea

ch in

g )

(C h

il d

s &

E yd

e, 2

0 0

2 )

R a n k

b y

M ed

ia n

N u m

b er

o f

R ep

o rt

s W

ri tt

en (M

ed ia

n N

u m

b er

o f

R ep

o rt

s) —

C li

n ic

a l

P ro

g ra

m s

(N =

1 1 1 )

(S te

d m

a n

et a l.

, 2 0 0 1 )

M M

P I

1 (n

= 1 3 1 )

1 (n

= 3

1 0

) 3

(8 6

% )

1 (7

.5 re

p o

rt s)

R o

rs c h

a c h

2 (n

= 1 2 2 )

3 (n

= 1 4 4 )

4 (8

1 %

) 4

(4 .3

)

T A

T 3

(n =

1 0 0 )

5 (n

= 8

4 )

5 (7

1 %

) 7

(1 .5

)

H o

u se

-T re

e -P

e rs

o n

4 (n

= 6 0 )

6 (n

= 7

2 )

1 7

(2 4

% )

9 .5

(1 .0

)

B e n d e r

G e st

a lt

5 (n

= 5 4 )

1 6

(n =

2 7 )

7 (4

6 %

) 1 2

(0 .5

)

B e c k

D e p

re ss

io n

In v

e n

to ry

6 (n

= 5 2 )

2 (n

= 1 6 3 )

— 5

(4 .0

)

M il

lo n

C li

n ic

a l

M u

lt ia

x ia

l In

v e n

to ry

7 (n

= 5 1 )

4 (n

= 9 6 )

8 (3

8 %

) 1 3

(0 .3

)

W A

IS -R

/- II

I 8

(n =

4 8 )

9 (n

= 4

2 )

1 (9

3 %

) 2

(6 .2

)

H u m

a n

F ig

u re

D ra

w in

g 9

(n =

4 6 )

8 (n

= 4

7 )

1 7

(2 4

% )

9 .5

(1 .0

)

R o tt

e r

In c o m

p le

te S

e n te

n c e s

1 0

(n =

4 3 )

1 0

(n =

3 6 )

— —

S e n te

n c e

C o m

p le

ti o n

1 1

(n =

4 0 )

7 (n

= 4

9 )

1 2

(2 9

% )

6 (2

.0 )

M A

C I

1 2

(n =

3 8 )

1 1 .5

(n =

3 3 )

— —

W IS

C -I

II 1

3 (n

= 3 5 )

1 9

(n =

2 3

) 2

(8 8

% )

3 (6

.0 )

C B

C L

1 7

(n =

2 2 )

1 1 .5

(n =

3 3 )

— —

N o

te :

M M

P I =

M in

n e so

ta M

u lt

ip h

a si

c P

e rs

o n

a li

ty In

v e n

to ry

; T

A T

= T

h e m

a ti

c A

p p e rc

e p ti

o n

T e st

; W

A IS

-R /I

II =

W e c h sl

e r

A d u lt

In te

ll ig

e n c e

S c a le

s– R

e v is

e d

a n d

3 rd

E d it

io n ;

M A

C I =

M il

lo n

A d

o le

sc e n

t C

li n

ic a l

In v

e n

to ry

; W

IS C

-I II

= W

e c h

sl e r

In te

ll ig

e n

c e

S c a le

s fo

r C

h il

d re

n –

3 rd

E d

it io

n ;

C B

C L

= A

c h

e n

b a c h

C h

il d

B e h

a v

io r

C h

e c k

li st

.

39

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

40 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

had 15 years of research and dissemination before the survey results that are reflected in Table 2.1. The stability in the ranking of tests when surveys have been repeated (Archer & Newsom, 2000; Stedman et al., 2001) also indicates that there is more reason to expect inertia rather than innovation.

A third discouraging observation is that the order of popularity does not correspond closely with evidence of validity. Established measures of psychopathology, such as the CBCL or the Beck Depression Inventory (Beck & Steer, 1987), rank below measures such as the Bender Gestalt Test and the Wechslers for psychopathology assessment. This order not only ignores the validity evidence that has accumulated for the rating scales and checklists (e.g., Achenbach, 1999; Beck, Steer, & Garbin, 1988), but it also fails to take into account the lack of evidence for the validity of the Bender Gestalt as a measure of personality or psychopathology besides visual-motor integration difficulties (Sattler, 2002), or the demonstrated invalidity of subtest analysis from cognitive ability batteries as a measure of personality or psychopathology (Glutting, Youngstrom, Oakland, & Watkins, 1996; Watkins, Glutting, & Youngstrom, 2005). Unfortunately, the lack of connection between evidence and practice is a phenomenon that has been observed throughout medicine (Guyatt & Rennie, 2002), and the enduring popularity of tests that are devoid of evidence has been true for decades. Almost 50 years ago, the editor of the Mental Measurements Yearbooks opined that “bad tests will always be with us” (Buros, 1965), an observation that remains apt today.

The First P: Prediction

What assessment tools should be used to measure the development of psychopathology? How can users compare the myriad tests that are available and make an informed choice among them? What evidence would persuade a user to adopt a different test rather than continuing to rely on an entrenched approach to assessment? One of the first heuristics for evaluating an assessment strategy is determining whether it predicts criteria of interest. Prediction could mean demonstrating concurrent correlations as well as showing associations with criteria that are separated by time.

CONCURRENT CRITERION VALIDITY

Assessment tools can be useful by virtue of correlating with other meaningful criteria. Screening measures are useful because they correlate with diagnosis. Diagnoses are valuable in part because they correlate with associated features of illness, as well as courses and outcomes. A nomothetic network of correlations also helps validate a diagnosis as a construct, by showing associations with family history, biological processes, or experimental psychopathology task performance (Cantwell, 1996; Robins & Guze, 1970).

An important part of the research enterprise is establishing the validity of a construct or diagnosis. This can be thought of as an iterative process that involves both elaboration and consolidation. Elaboration can involve expanding the network of correlations to include not just validators but also correlates in terms of typical treatment response, outcome, quality of life, and functioning. The elaborative process is crucial to contextualizing the construct and understanding its connections to various aspects of development.

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE FIRST P: PREDICTION 41

The consolidation aspect of research involves clarifying when different measures are assessing the same underlying latent variables, and also creating models about the relation- ships among these underlying processes. The consolidation aspect becomes increasingly valuable as the measures proliferate in number. To what extent are different self-report measures that claim to assess depression really measuring the same thing? Instruments may include scales named “Externalizing” (Achenbach & Rescorla, 2001), “Undercon- trolled Behavior” (McDermott, 1994), and “Aggression,” but despite the different names, these are likely to be tapping similar phenomena and underlying latent constructs.

The consolidation process can be expanded to include theoretical or conceptual linkages across constructs. For example, Gray has developed a model focusing on three major systems: the Behavioral Inhibition System (BIS), focused on cues of threat or punishment; the Behavioral Activation System (BAS), focused on cues of reward; and the Fight/Flight System (FFS) (Gray & McNaughton, 1996). Other scholars have investigated the relationship between these basic motivational systems and either underlying neurophysiological systems (Panksepp, 2000) or markers of temperament, personality, and psychopathology (Depue & Lenzenweger, 2001; Depue, Luciana, Arbisi, Collins, & Leon, 1994; Quay, 1993, 1997). The BIS/BAS example also underscores two other facets along which meaningful consolidation can occur: (1) across systems of functioning, as has been done with the construct of emotions, which in turn organize behavior at a physiological, cognitive, facial, and behavioral level (Lazarus, 1991), and (2) across different informants who are reporting about the same behaviors (Achenbach, McConaughy, & Howell, 1987). More recently, the Research Domain Criteria (RDoC) initiative of the National Institute of Mental Health has advocated for the integration of assessment across levels of analysis, moving from gene to cell function at the molecular level, up to circuits, systems, behavior, and interpersonal functioning (Cuthbert, 2005; Insel et al., 2010). The RDoC domains include constructs such as negative and positive affect, and attempt to connect them with animal models, experimental methods, and ultimately clinically measurement that would be informative with patients.

STATISTICAL METHODS

There are a variety of statistical models that help to examine different aspects of the validity of an assessment tool. The following section reviews some widely used methods, as well as some newer approaches that have great potential relevance, often modeling aspects of development or change more directly than older statistical methods.

Correlations

The correlation coefficient is probably the most widely reported measure of predictive association in developmental psychopathology. It measures effect size, and it can be compared readily across studies and across measures within studies. The main drawback of the correlation coefficient as an index of prognosis is that it is difficult to apply to individual cases (Wiggins, 1973). A second limitation is that it can consider only one correlate at a time. It also can be unduly influenced by extreme scores, particularly in smaller samples.

Comparing Correlations

It is possible to formally test whether correlation coefficients between the same variables but drawn from different samples are close enough in size to attribute any differences to

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

42 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

sampling error. This can be done with a z-test (Cohen, Cohen, West, & Aiken, 2003). Conceptually, it is a test of whether differences between the two samples moderate the correlation between the variables. An example of this would be to compare the correlation between externalizing and internalizing problem scores in a new sample of data with the correlation published in the standardization sample. If the z-test rejected the null hypothesis that the two coefficients were sampled from the same population, then investigators would be alerted to look for aspects of the sample or design that might have changed the relationship between the two dimensions of behavior problems.

It also is possible to test whether the correlations between two different predictors and the same criterion are different from each other in the same sample. For example, one could test whether parent report or teacher report of depressive symptoms shows a stronger correlation with youth self-report within the same sample. The formula for this test is more complicated, as it needs to account for the nuisance correlation between the two predictors (in this case, the correlation between parent and teacher report) (see Cohen & Cohen, 1983, pp. 56–57; note that this formula is not included in the newer edition). Conceptually, this formal comparison of correlations addresses a pragmatic assessment question: Is one of these scores a significantly better predictor of the criterion than the other is? In assessment situations where both tools are available, then the one with the higher correlation would typically be the first-choice assessment strategy.

The formal test is much better than the more common practice of examining the statistical significance of both correlations. Both correlations could be significantly different from zero (the null hypothesis), yet one could still be a better predictor than the other. Worse, one could be statistically significant (i.e., different from zero) and the other not achieve significance, yet the two coefficients might not differ reliably. For example, with an N of 100, a correlation of .22 would be significant (p < .05, two-tailed), but a correlation of .18 would not be. However, it would be a mistake to conclude that the predictor yielding a correlation of .22 had reliably outperformed the other predictor. Unfortunately, this happens when people focus only on whether particular variables achieve statistical significance and tally up significant versus nonsignificant variables. The field would benefit from more widespread adoption of direct comparison of validity correlations for different measures.

Regression

Regression analyses offer several additional refinements beyond what is possible with bivariate correlations (Crawford, Garthwaite, Denham, & Chelune, 2012). These include (a) preserving the actual units of measurement in the unstandardized regression weights, (b) providing formal tests of whether combinations of predictors provide incremental improvement over a single predictor, (c) making predictions about an individual’s score on the criterion variable, and (d) creating a framework where it is possible to test statistical mediation or moderation of relationships. A limitation of regression, like correlation, is that it models linear associations. If the developmental phenomenon does not follow a linear pattern (for example, the acquisition of language), then care needs to be taken with variable transformations or alternate statistical models used to quantify relationships.

Raw Is Good

As a field, developmental psychopathology has tended to ignore the unstandardized coefficients and focus instead on standardized correlations. However, there are major advantages to working with the variables in their actual metric when it comes to

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE FIRST P: PREDICTION 43

assessment of an individual, as has long been recognized in industrial-organizational psychology (Guion, 1998; Wiggins, 1973). Working in the raw units brings the focus back to how clinicians would actually use the instrument. Instead of relatively abstract concepts such as the correlation being .3 (or the predictor explaining 9% of the variance in the criterion), the unstandardized regression coefficients would indicate the actual scores that would be predicted on the dependent variable (e.g., every point increase in depression is associated with a third of a point decrement in average scores on the quality of life measure, or each 10-point increase in attention problems is linked with a .25 decrease in grade point average).

Incremental Validity

The test of incremental validity has been widely used in research (Haynes & Lench, 2003). If the additional variable can provide a statistically significant improvement in prediction of the criterion (often quantified as a significant increase in the R2 of the regression model, or as a significant regression weight for the new variable entered in the model), then it has demonstrated incremental validity at a statistical level. This would provide statistical justification for exploring a more complex assessment battery that included both sources of information in order to provide a more accurate prediction about the individual.

Observed Versus Predicted Performance

It is possible to use regression formulas to compare an individual’s observed scores to what would be predicted based on the general relationship of the variables in the population. This technique has been most articulated in the areas of comparing cognitive ability and academic achievement (Psychological Corporation, 1992) and in neuropsychological testing, where it is often used to examine outcomes in pre- and posttesting designs (Sawrie, Chelune, Naugle, & Luders, 1996). The tool could be applied productively to other areas of psychopathology research (Crawford et al., 2012).

For example, clinicians often are struck by the relatively low levels of concerns reported by teachers or youths when referrals are initiated by parents. Applying the regression framework would remind clinicians that given the typical amount of agreement between parents and youths (r = .22 in a meta-analysis; Achenbach et al., 1987), the level of concern we intuitively might expect to see across informants would actually represent an exceptionally high level of agreement (Youngstrom, Meyers, Youngstrom, Calabrese, & Findling, 2006b). This method also could identify instances where cross-informant agreement was actually significantly worse than expected based on normative data, triggering more detailed examination of the factors contributing to disagreement in the specific case (De Los Reyes, Henry, Tolan, & Wakschlag, 2009). This method has been applied to parent-youth and parent-teacher agreement about symptoms of mood disorder, emphasizing that substantial differences in opinion are par for the course given the relatively modest levels of typical agreement across informants (Youngstrom et al., 2006b).

Challenges to Using Regression Equations Clinically

Obstacles to applying regression models in clinical practice have included the lack of sufficient published statistics (e.g., when researchers publish correlations or regression weights without including the intercept statistic), the increased computational burden on the clinician, and the fact that the regression weights are dependent on the specific

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

44 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

combination of instruments used. Sometime predictions generalize better when weighting is ignored and unit weights are used instead (Perloff & Persons, 1988).

All of these challenges are solvable (Crawford et al., 2012). Researchers can publish more complete details of the regression results. The computational burden has sometimes been managed by creating tables where clinicians can look up the predicted value (e.g., Psychological Corporation, 1992), although this works best for simple bivariate models. Web-based or mobile device applications could perform the necessary calculations quickly, conveniently, and accurately. Meta-analyses can formally test the question of sample dependence. If sample characteristics change the performance of the measure, then separate regression models can be used as appropriate (much as separate norms are used for many tests). In short, the technical challenges are no longer a major impediment to applying regression models to individual cases. The benefits could be considerable: Even relatively simple models often provide predictions that are much more accurate than the results achieved via nonactuarial decision making (Grove, Zald, Lebow, Snitz, & Nelson, 2000; Jenkins, Youngstrom, Washburn, & Youngstrom, 2011; Meehl, 1954).

Statistical Methods for Consolidation of Scores

Principal components analysis, exploratory factor analysis, and confirmatory factor analysis are all methods for synthesizing correlations among sets of variables. Within an assessment framework, these techniques have been widely used to test the dimensionality of assessment instruments. There are a variety of different methods available for determining the number of dimensions underlying a battery. Three have consistently performed well in simulation studies (Velicer, Eaton, & Fava, 2000), yet are relatively rarely used in developmental psychopathology research. These are the Scree Test, Minimum Average Partials (Velicer et al., 2000), and Parallel Analysis (Glorfeld, 1995; Horn, 1965). The latter two are not included as options in the most popular statistical packages, but free code is available to run these techniques on many platforms (O’Connor, 2000). All three of these methods tend to converge on more parsimonious factor structures than maximum-likelihood-based techniques identify, which form the basis of confirmatory factor analytic approaches. Parsimony offers considerable advantages. Retaining fewer factors means that there will be fewer scales to interpret, reducing the risk of Type I (false positive) errors in clinical assessment (Silverstein, 1993). The smaller number of factors also tends to include a larger number of items per factor, yielding greater internal consistency reliability, smaller standard errors of measurement, and more accurate description of an individual (Brown, 2006).

Experts who have suggested that overfactoring is preferable to underfactoring (e.g., Fabrigar, Wegener, MacCallum, & Strahan, 1999) view the issue from the perspective of a statistician and not a clinician or researcher who needs to apply measurement to a specific individual. Greater reliance on approaches such as Parallel Analysis would accelerate the consolidation of measures and also promote the development of scales with psychometric properties better suited for making decisions about individuals. By putting factor structures on firmer footing using exploratory methods early in the cycle of scale development, it is likely that the field would produce more robust scales for validation with confirmatory methods.

Covariance modeling approaches can also quantify the overlap between related con- structs. To what extent are anxiety and depression distinct phenomena, as opposed to sharing a common general internalizing component? Are the more than 200 extant measures of depression and internalizing problems all measuring the same underlying

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE FIRST P: PREDICTION 45

latent variable (Nezu, Ronan, Meadows, & McClure, 2000)? Within a developmental framework, covariance models are also important methods to clarify the points of contact between assessment tools deployed at different ages, such as measures of temperament versus measures of personality (Shiner, 1998). All of these can be viewed as examples of conceptual consolidation, where covariance modeling is used to integrate and synthesize different sources of information, different tests, and different age epochs to identify underlying constructs. However, the same tools can also be helpful in the context of elaboration, as in the case of “structural equation models,” where path models link latent constructs that are quantified by means of a “measurement model” that includes a factor analysis for each of the latent variables (Bollen, 1989). Latent variable models now make it possible to test measurement invariance to see the extent to which items or measures change their performance across different demographic or clinical groups (Borsboom, Romeijn, & Wicherts, 2008; Zumbo, 2007).

Grouping Individuals Instead of Items or Scales

An alternate approach to descriptive assessment would be to group individuals together based on similarity, rather than grouping variables together. These have been described as “q-methods,” as distinct from the “r-methods” of describing correlation among variables (Thompson, 2000). One specific statistical method would be to use common factor analysis, but to aggregate people with similar profiles of scores across variables, rather than the more widespread approach of grouping variables with similar scores across people. Cluster analysis is another method that has a long history of use to aggregate similar cases in biology as well as the social sciences (Achenbach, 1993; Aldenderfer & Blashfield, 1984). Latent class analysis (LCA) (McCutcheon, 1987) is a related methodology for clustering similar cases on the basis of observed categorical indicator variables (e.g., Hudziak, Althoff, Derks, Faraone, & Boomsma, 2005; Hudziak et al., 1998). Mixture modeling, and the factor mixture model in particular, is a related statistical method that explicitly takes into account residual relationships among indicators. This is helpful in developmental psychopathology, where symptoms often correlate strongly, even within relatively homogeneous classes, producing a severity gradient (Lubke et al., 2007). The use of mixture and factor mixture models not only is consistent with the complex (categorical and dimensional) nature of many psychopathological constructs (Frazier et al., 2012), but also meshes with DSM-5 conceptualizations of psychopathological categories as having dimensional attributes.

A third family of methods is the “coherent cut kinetics” or “taxometric” methods developed by Meehl, Waller, and colleagues (Schmidt, Kotov, & Joiner, 2004; Waller & Meehl, 1998). Taxometric methods are well suited for testing whether data are distributed along a continuum, or whether they reflect two naturally occurring underlying categories; but they will not perform well in situations where there are three or more underlying categories. Latent class analysis and cluster analyses, on the other hand, are better when there are multiple underlying groups, but they will provide inaccurate clustering solutions in situations where the underlying data actually are dimensional. Perhaps both methods should be used in tandem (Beauchaine & Beauchaine, 2002; Solomon, Haaga, & Arnow, 2001).

Q-methods have been employed several times in assessment models in developmental psychopathology. Taxometric methods have been used in more than 100 studies across a variety of disorders and age groups (Haslam, Holland, & Kuppens, 2012), and latent class analyses have been applied to attention problems, aggression, and other behavior

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

46 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

problems, as mentioned earlier. There also are intriguing examples of using clustering methods to define normative profiles within standardization samples of both clinical syndrome scales (e.g., Achenbach, 1993; Kamphaus, Petoskey, Cody, Rowe, & Huberty, 1999; McDermott, 1994) and cognitive ability tests (e.g., Konold, Glutting, & McDermott, 1997).

Q-methods provide an empirical approach for creating a nosology of common patterns of behavior problems. Clustering methods allow the data to dictate how many core profiles emerge, and what the relative prevalence is of each profile. If the same indicators are used at multiple settings or time points, then it becomes possible to identify secular trends in prevalence and group composition. For example, by clustering the entire standardization sample together (pooling ages, sex groups, and ethnicity), it becomes evident that profiles characterized by elevations in attention problems and aggressive behavior are more common in males, but decrease in prevalence with age, whereas anxious and depressed profiles are more common in females, with the depressed scores rising even higher in females after the transition to adolescence (McDermott & Weiss, 1995). Q-methods also provide a parsimonious response to the problem of comorbidity: Most core profiles involve elevations on multiple scales, reflecting that symptoms and behaviors described as separate disorders in DSM have a strong tendency to co-occur (Angold, Costello, & Erkanli, 1999). Again, factor-mixture modeling may be an optimal statistical approach for modeling complex associations among variables and classifying individuals.

Q-methods can also classify individual cases, assigning them to group membership based on the similarity of their scores to the average scores for each core profile. In the Achenbach approach, similarity was indexed as a set of q-correlations between the individual’s profile of scores and the average scores for each of the core profiles, with the individual classified as belonging to the group that showed the highest correlation coefficient (Achenbach, 1993). In another approach, the similarity of an individual to each core profile was quantified using “generalized distance,” the sum of the squared discrepancies between the individual’s score and the cluster average on each measure (McDermott, 1998).

An extra fillip offered by McDermott’s method was the inclusion of a “maximum distance,” after which an individual’s scores would be considered unique and not closely matching any of the core profiles from the standardization sample (McDermott, 1998). This threshold was empirically determined by examining the frequency distribution of discrepancy scores and finding what generalized distance score was so extreme that 95% of cases in the standardization sample scored at or below that point. This approach offers a statistically based system for determining when a person shows an unusual multivariate profile. Advantages of this include that the clustering method accommodates the fact that the average profile is not “flat” (i.e., on average youths do not have all of their cognitive abilities equally developed, nor do they typically display a wide range of behavior problems to the same degree), and it also avoids making multiple univariate decisions about an individual when looking at correlated measures (McDermott & Weiss, 1995). Thus, these methods would help to cut down on the rate of Type I errors, or false positive inferences, in clinical decision making (Silverstein, 1993).

Another advantage of multivariate profiles is that they can incorporate information from multiple informants (e.g., De Los Reyes et al., 2011) or different domains of functioning. For example, some investigators have clustered intelligence and academic achievement data at the same time to identify profiles of the functioning, including commonly occurring patterns where academic achievement is markedly different than the level of cognitive

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE FIRST P: PREDICTION 47

abilities (e.g., Konold et al., 1997). An intriguing extension of this methodology would be to cluster behavior problems and positive aspects of social functioning or quality of life at the same time. Epidemiological studies have shown that there are many people who present with symptoms of psychopathology, but without impairment or decrement in quality of life (Bird, 1996; Bird et al., 1990). It would be interesting to document how common these patterns of functioning might be that show elevated symptom levels without a corresponding deficit in functioning. These individuals have sometimes been described as the “worried well,” but an alternate conceptualization might be that these are people who are resilient despite having some symptoms. These individuals also might be important to include in investigations of endophenotypic markers or heritability studies (Gottesman & Gould, 2003; Hasler, Drevets, Gould, Gottesman, & Manji, 2006).

PROGNOSIS: PREDICTION INTO THE FUTURE

Clinical prognostication has been more obvious in the area of developmental psy- chopathology than in many other areas of psychology. Prognosis refers to the course of illness or the longitudinal outcomes that are likely for individuals affected by a condition or showing a particular marker or trait. Prognostic data help validate diagnoses, and they also inform decisions about whether to intervene. Anxiety problems in childhood frequently have been dismissed as simple shyness or as a behavior pattern that youths are likely to outgrow; however, longitudinal data show that children meeting criteria for anxiety disorders are at higher risk for substance use, depression, peer rejection, and continued dependence on parents or the welfare system as young adults (Silverman & Ollendick, 2005). Thus, the prognostic value of an anxiety disorder diagnosis contributes to the justification for clinical intervention. Similarly, the higher rates of car accidents, substance use, arrest, and other poor outcomes associated with a diagnosis of attention- deficit/hyperactivity disorder (ADHD) provide grounds for initiating treatment (Barkley, 2002). Prognostic value can apply to test results or to clinical signs as well, and could be used to predict treatment response.

Addressing questions of development or change over time is more an issue of research design than analysis, in the sense that the underlying matrix algebra is the same regardless of whether the covariances are drawn from a single panel or multiple time points. All the analytic methods described in the previous section are special cases of more general analytic methods, where both the independent and dependent variables happened to have been gathered at the same time. Thus, correlation, regression, and other covariance structure-based methods continue to be helpful in describing and quantifying longitudinal relationships. The strength of the inference comes from the research design and from having the same participants followed for multiple time points.

Individual Trajectories (Growth on Continuous Measures)

A variety of statistical refinements are available for modeling longitudinal data with even greater sophistication. These include mixed effect models (also referred to as “hierarchical linear models”), where repeated measures are treated as being nested within the individual participant, and other stable differences between participants (such as sex or ethnicity) are modeled in a second, higher-level regression equation (Raudenbush & Bryk, 2002; Tabachnick & Fidell, 2007); growth curve models, where observed variables are treated as indicators of change for indirectly observed “latent variables” (Duncan, Duncan, Strycker,

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

48 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

Li, & Alpert, 1999); and growth mixture models or latent class growth models where unique trajectories of indicators are empirically identified. Each has specific advantages and disadvantages. A technical treatment of these methods is outside the scope of a chapter on assessment, and the interested reader is referred to the more comprehensive treatments cited earlier, or to more conceptual overviews offered elsewhere (Grimm & Yarnold, 1995; Nylund, Asparouhov, & Muthén, 2007).

For our purposes, an important generalization to bear in mind is that these techniques concentrate on change in scores on a continuous dependent measure, such as academic achievement, or level of anxious symptoms, over time. In principle, each of these methods could also be used to generate a predicted score on the dependent measure for an individual, as discussed earlier in the regression section. These predictions could be incorporated into a clinical decision-making framework (Straus, Glasziou, Richardson, & Haynes, 2011), or they could be used to generate more refined hypotheses for future research or policy work (Kaplan & Elliott, 1997). However, there is a trade-off, such that as the models or analytic methods become more complex, the application to an individual case becomes more difficult. Such an exercise is technically possible; for example, we very recently found increased rates of conversion to bipolar disorder in youths with a particular pattern of elevated symptoms of mania using growth mixture model classifications (Findling et al., under review). As these techniques become more widely accessible and can be powerfully applied to longitudinal data sets, it is possible that what was previously viewed as a limitation to the clinical utility of these approaches (Oh, Glutting, Watkins, Youngstrom, & McDermott, 2004) will become a strength.

Time Until an Event Occurs

A second family of methods, less used in developmental psychopathology research to date, is event history analysis. This family of techniques includes survival analysis and Cox regression—methods for looking at the length of time until a particular event of interest happens (Tabachnick & Fidell, 2007). The event is categorical; examples could include pregnancy, dropping out of school, graduation, arrest, or relapse of an illness. A major strength is that event history methods can model data where the outcome is “censored,” or not directly observed in all participants at the time of analysis. For example, when investigating time until arrest, survival analysis can accommodate the fact that many participants have not been arrested at the time of the conclusion of the study, without needing to assume that these people would never get arrested. Instead, the survival analysis weights these “censored” cases differently than cases with directly observed outcomes, producing less biased estimates (Tabachnick & Fidell, 2007). Cox regression can include both continuous and categorical predictors as covariates that may account for differences in the time until the occurrence of the event (Willett & Singer, 1993).

Event history methods are uniquely suited to investigations of onset, cessation, relapse, and recovery (Willett & Singer, 1993). There are even models that allow for repeated recurrences of the event of interest (as would happen with multiple nonfatal heart attacks or relapses of mood disorder). Event history techniques also produce “hazard ratios” or regression weights that could be applied to an individual case to make predictions about risk or outcome. Although event history models are beginning to appear in the research literature, particularly in forensic assessment contexts (e.g., Richards, Casey, & Lucente, 2003; Tengstroem, Grann, Langstroem, & Kullgren, 2000), there have been fewer attempts to generate individual prediction models. This has been due to the combination

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE SECOND P: PRESCRIPTION 49

of unfamiliarity with Cox regression in the field of psychopathology assessment, and also because of the computational burden involved in applying the regression weights (which represent changes in the log odds of the event happening). However, the greater availability of inexpensive computing power in the form of mobile applications or web-based programs reduces the computational barriers to use. Both Cox regression and survival analysis are now available in popular statistical software packages such as SAS and SPSS, and more specialized software such as MPlus (Muthen & Muthen, 2004) makes it possible to test models that mix together growth curve and event history analyses, enabling the flexible specification of models that closely approximate the developmental processes of interest.

SUMMARY: PREDICTION

Prediction is one of the cornerstones of both clinical assessment and also developmental psychopathology. There is a wealth of published data containing correlations between published measures of a myriad of constructs. However, most articles and test manuals have concentrated on the statistical significance of associations, sometimes also reporting the correlation coefficient. The extant literature on almost all measures falls short of its potential to inform choices about test selection or application to decisions about individuals. From a research standpoint, the agenda for improving the predictive value of assessment tools would include (a) publishing studies that consolidate existing measures into more parsimonious dimensions; (b) conducting studies that elaborate the connections between measures and constructs across development, such as linking measures of temperament to personality; (c) directly comparing the predictive value of multiple tests under the same conditions (versus the current convention of focusing only on null-hypothesis significance testing) so that test selection can be guided by empirical results; and (d) supplementing or supplanting correlational analyses with multivariate regression analyses, reported in enough detail to allow application to individual cases (with appropriate confidence bands). More speculative research endeavors could include development of q-method approaches to assessment, creating empirical, multivariate taxonomies of individuals.

The Second P: Prescription

Assessment should guide choices about treatment. Assessment findings can lead to a prescription of a type of treatment, sometimes referred to as “treatment matching” between a diagnosis and an intervention strategy.

THRESHOLDS FOR ASSESSMENT AND TREATMENT

The likelihood of a person having a particular condition can be thought of as a probability, ranging somewhere between 0% (when the person definitely does not have the condition) up to 100% (when the person definitely does have the condition). In clinical practice, we are never absolutely positive about the presence or absence of a condition. Instead, our assessment of the probability will range somewhere between these two extremes. Within this conceptual framework, assessment is helpful to the extent that it changes

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

50 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

our probability estimate. Negative results on a valid test decrease the probability of a condition, and positive results increase the probability.

When we become sufficiently confident in a diagnosis or a case formulation, then we go ahead and initiate treatment. Not all cases receiving treatment actually have the condition—clinical diagnosis is imperfectly reliable (Garb, 1998; Kraemer, 1992). Similarly, we cannot be 100% certain that any specific child actually has a given disorder. However, once our estimate of the probability exceeds a certain point, we are confident enough to go ahead and begin treatment. This has been called the “treatment threshold” (Straus et al., 2011). If our estimate of the probability falls below this threshold, then we would do more assessment until either we achieved sufficient confidence that the condition was present to warrant intervention or else the probability estimate became so low that we considered the diagnosis ruled out. There is a second threshold, the “assessment threshold,” below which the diagnosis is considered so unlikely that further testing is not needed. The range of probabilities falling between the assessment threshold and the treatment threshold represents the situation in which additional assessment is clinically indicated, as we are neither confident enough to begin treatment nor certain that the condition is absent. Figure 2.1 illustrates the threshold concept.

Typically, the assessment and treatment thresholds are not explicitly defined by clinicians. Instead, we make intuitive decisions about when testing is needed. Writing down our probability estimate and our thresholds would immediately make the decision- making process more conscious and transparent. This framework also can be empowering for the patient, as it makes it possible to weigh the costs and benefits associated with testing and with treatment, and to negotiate where to set the assessment and treatment thresholds (Straus et al., 2011; Youngstrom, 2013).

Combining the threshold model of decision making with the idea of levels of inter- vention (Mechanic, 1989) produces an integrated model of assessment and treatment. Primary intervention, or universal prevention, treats everyone, regardless of risk. This approach is rational when treatment’s costs and risks are low and the benefits decisively outweigh them. Primary interventions include such techniques as fluoridating drinking supplies to prevent dental problems, mandating a comprehensive vaccination program, or advertising the benefits of exercise as a protective factor against heart disease. In terms of the threshold model of decision making, primary intervention sets the treatment threshold at a probability of zero: Even if it is highly unlikely that people have the target condition, they will still receive the treatment. Assessment is not needed as a gatekeeper to determine eligibility in a primary intervention model; in fact, assessment may add unnecessary expense.

Secondary interventions concentrate on those who are at risk of developing a condition, but who have not yet fully manifested the syndrome (Mechanic, 1989; Youngstrom, 2013). Risk could be defined by exposure to a risk factor, such as a traumatic event, or it could be defined as a prodromal expression of the condition or an intermediate probability of developing a bad outcome. The advantages of secondary intervention include that fewer cases are treated, permitting the dose or expense of the treatment to be greater than in a primary prevention model. Limiting treatment to those at probable risk also often avoids the ethical concerns inherent in universal treatment, where informed consent may not exist.

Secondary interventions may be appropriate for cases in the midrange of probability of having a diagnosis. In this view, secondary interventions might be deployed for the same people who would also fall between the assessment and treatment thresholds—cases

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE SECOND P: PRESCRIPTION 51

100% High:

Treat with Tertiary

Intervention

Moderate:

Test and

Use Secondary

Intervention

Low:

Wait Zone or

Primary Prevention

Treatment:

Acute Interventions (intensive

therapy, medication,

hospitalization)

Assessment:

Supplemental testing to gather

enough data to confirm or disconfirm

diagnosis

More diagnostically specific measures

Treatment:

Target diagnosis ruled out

Treat any other conditions

Assessment:

No further assessment for disorder

unless there is a new risk factor or

change in status

Treatment:

Secondary interventions

Nonspecific and low-risk treatments

0%

Assessment:

Potential treatment moderators

Monitor process and adherence

Probability

Wait-Test Threshold

Test-Treat Threshold

FIGURE 2.1 Decision Thresholds for Assessment and Treatment, Combined With the “Levels of Intervention” Model. Note: The assessment process combines information about the presenting problem, risk factors, and assessment results into a revised estimate of the probability of a diagnosis. The revised

probability in turn guides next actions about assessment and treatment. Adapted from Straus et al. (2011)

and Youngstrom (2013).

in which the probability is high enough to justify taking treatment precautions as well as continuing assessment, but not yet so diagnostically clear as to warrant bringing the heavy artillery of tertiary interventions to bear. For example, after gathering test data and clinical history, a practitioner using an evidence-based assessment approach might conclude that a person has a 60% probability of having schizophrenia. This is above the assessment threshold, indicating continued evaluation. At the same time, the probability is likely below the treatment threshold at which the clinician and patient would be ready

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

52 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

to start antipsychotic medication. In addition to recommending continued evaluation, the clinician might also recommend initiating secondary interventions at this point, such as reducing stressors, curtailing substance use, or seeking supportive interactions with friends (Youngstrom, 2013).

An alternate way of integrating secondary interventions with the assessment threshold model would be to make clear definitions of risk status that are triggers for intervention. In cardiology, elevated blood pressure has become conceptualized as a risk factor or prodrome of heart disease that is sufficient to trigger a variety of interventions, including changes in diet, exercise, or the use of medications specifically intended to regulate blood pressure. Similarly, the schizotaxic construct (or schizotypy) appears to be a constellation of behaviors and biological markers that signal a diathesis for developing schizophrenia (Blanchard, Gangestad, Brown, & Horan, 2000). Screening or early identification programs could identify individuals with schizotaxia, and then initiate psychoeducation and prevention programs to lessen the risk of progression into schizophrenia. Other temperamental or personality dimensions are well established as correlates of different forms of pathology. High trait neuroticism, shyness, or behavioral inhibition are not exactly the same thing as a DSM-defined psychiatric disorder of anxiety or social phobia, but all of these variables have demonstrated robust associations with anxiety disorders as well as mood disorders (Harkness & Lilienfeld, 1997; Kagan, 1997a; Lonigan, Vasey, Phillips, & Hazen, 2004). Analogous to blood pressure, elevations of these factors could be conceptualized as triggers for secondary intervention to prevent exacerbation into an impairing disorder.

Tertiary interventions are the most intense treatments. They also are often the most expensive, and may have the most serious risks of side effects. The increased costs (both fiscal and risk of harm) militate against deploying tertiary interventions more broadly. Assessment serves as the gatekeeper determining when to start tertiary interventions. In the evidence-based medicine (EBM) threshold model, the treatment threshold is set based on the costs and benefits associated with treatment; and treatment begins when the diagnostic probability of having the condition rises above the treatment threshold (Straus et al., 2011). One of the main sources of dissatisfaction with tertiary intervention models is that the illness is often quite advanced by the time it is recognized and treated. The severe progression of the condition often makes it more difficult to treat, with even greater expense and lower rates of success. At an individual level, it is much less costly to treat high blood pressure than it is to do open heart surgery, and the patient has a higher probability of survival. Assessment also protects people from unnecessary interventions, which convey all of the risks associated with the treatment, but few or none of the potential benefits. At the tertiary intervention level, the assessment question is usually not if there is a problem, but rather what the specific nature of the problem is, so that optimal treatment can be prescribed. Figure 2.1 maps the levels of intervention model onto the test and treat threshold model (see also Youngstrom, 2013), showing how individual assessment findings are integrated into a revised estimate of the probability of diagnosis. The probability in turn guides the choice of next clinical actions about assessment and treatment.

DIAGNOSIS AS A SHORTHAND

There have been extensive critiques of the validity of the DSM and the International Classification of Diseases (ICD) diagnostic systems (e.g., Carson, 1997; Insel et al., 2010;

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE SECOND P: PRESCRIPTION 53

Kihlstrom, 2001; Wakefield, 1997). The number of diagnoses has proliferated with each revision of the nosology, with DSM-IV-TR containing more than 350 distinct diagnoses. DSM-5 collapses some disorders into a single category with dimensional substructure (e.g., autism spectrum disorder). However, it is unclear whether each of the DSM-5 categories actually represents a distinct categorical entity, versus being a somewhat arbitrary label that is imposed on extreme levels of an underlying trait. In fact, accumulating evidence suggests that most conditions involve core features that vary along a continuum, rather than representing distinct categories (Frazier, Youngstrom, & Naugle, 2007; Markon & Krueger, 2005; Ruscio & Ruscio, 2002; Schmidt et al., 2004). The extremely high rate of comorbidity among putative psychiatric diagnoses also strongly suggests that the current classification systems are splitting too much—imposing artificial distinctions that do not have an underlying basis (Angold et al., 1999). For example, the high rates of overlap between ADHD and bipolar disorder (Galanter & Leibenluft, 2008; Youngstrom, Arnold, & Frazier, 2010) or between generalized anxiety and major depression (Cuthbert, 2005; Mineka, Watson, & Clark, 1998) suggest that shared mechanisms contribute to the apparent comorbidity.

However, there are major advantages to categorical classification. There usually is a dichotomous choice between treating and not treating something, as well as the practical issues of billing and the conceptual advantages of a label as a shorthand for a constellation of related variables. A key future direction is determining how to consolidate diagnoses into more parsimonious groupings that could guide treatment. Once these groups are defined, then assessment tools can be reevaluated in terms of how well they differentiate between groups with distinct etiologic patterns or that would benefit from a different intervention approach. Although the subsequent examples will concentrate on current diagnoses, bear in mind that (a) current diagnoses need further validation and will likely be modified during the validation process, (b) the assessment framework being advocated here would still apply to any categorical decision, and (c) it will be incumbent on developmental psychopathology researchers to revisit the validity of extant measures with regard to any new diagnostic definitions.

ASSESSMENT AS AID IN DIAGNOSIS

Assessment can greatly aid practitioners in their decision about whether to initiate treatment, or which treatment to select. However, psychology as a field has largely under- utilized this capacity: Interpretation of tests has largely been intuitive and impressionistic (Garb, 1998). Unstructured clinical approaches to diagnosis consistently demonstrate low reliability. A recent meta-analysis found that the average kappa between semistruc- tured diagnoses and diagnosis as usual was around .27 (Rettew, Lynch, Achenbach, Dumenci, & Ivanova, 2009), indicating that there is room for assessment to help improve accuracy. Concordance between research and clinical diagnoses also is associated with better treatment outcomes (Jensen-Doss & Weisz, 2008). Fortunately, there is a highly refined framework for evaluating the contributions of a test to diagnosis, based on signal detection theory or Bayes’ theorem (Kraemer, 1992). The performance of a test can be described in terms of its sensitivity and specificity to a diagnosis, where sensitivity refers to the percentage of cases with the diagnosis that would be classified correctly by the test. Specificity quantifies the percentage of cases without the diagnosis that would be classified correctly by the test. The specificity of a test can always be improved by using a

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

54 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

more stringent threshold. However, there is almost always a trade-off between sensitivity and specificity.

Receiver operating characteristic (ROC) analysis is a method for quantifying the relationship between the sensitivity and the specificity of an assessment tool for a particular diagnosis. ROC curves plot the sensitivity as a function of the specificity (or the false alarm rate—the complement of specificity, arithmetically equal to 1 minus specificity), moving across the entire range of possible test scores. It is possible to measure the area under the curve (AUC) for the ROC plot of test performance, yielding an index that ranges from 1.00 (for a perfectly discriminating test, achieving 100% sensitivity and 0% false alarms, or 100% specificity) to .50 (reflecting chance performance). The AUC can be interpreted as the probability that a randomly selected case with the diagnosis would have a higher score on the test than would a case without the disorder (McFall & Treat, 1999). The AUC provides a single, global index of how well a test can aid in the differentiation of a diagnosis.

There are statistical methods for comparing the AUCs for the same test evaluated in different samples (Hanley & McNeil, 1983). This procedure is helpful because it is possible for the performance of the test to change depending on sample characteristics. A common study design that exaggerates test performance compared to what it would actually deliver under clinically realistic conditions would be to limit the sample to affected versus healthy normal controls (e.g., Steer, Cavalieri, Leonard, & Beck, 1999; Tillman & Geller, 2005). Such designs magnify the apparent discriminating power in two ways: (a) they create groups that differ based on global impairment, as well as specific features of the illness, making the groups easier to tell apart (increasing the sensitivity of the test to the target condition); and (b) the samples exclude other illnesses that might share some of the features of the target condition (decreasing the number of false alarms for the test, and thus raising its apparent specificity). Youngstrom, Meyers, Youngstrom, Calabrese, and Findling (2006a) demonstrated both artifacts by analyzing the same mood rating scales under two different sampling designs, one taking all comers and producing more modest estimates of diagnostic efficiency, and the other design excluding 30% or more of the complex cases and yielding inflated estimates.

It is possible to test whether one instrument is performing significantly better than another in the same sample, adjusting for the correlation between the two assessment measures (Hanley & McNeil, 1983). Because same study is comparing the two tests, the design holds constant most of the factors that might change test performance across studies, such as definitions of the illness, the base rate of the illness, the severity of the presentation, or the amount of comorbidity present in the sample. Such “horse race” designs simultaneously evaluate multiple measures under the same conditions, offering persuasive evidence if there are differences between the measures’ diagnostic performances. Because two candidate measures may be highly correlated, tests of the difference between their discriminative performances will have high power to detect clinically meaningful differences.

Clinical Implications

Clinicians should keep an eye open for “diagnostic horse race” papers, and if a measure outperforms the tool that they are currently using, then they should change instruments. These studies also will be greatly helpful in winnowing the field of assessment measures. Consider the example of assessing depression. There are more than 200 published depression measures available (Nezu et al., 2000). It is unlikely they are equally good

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

THE SECOND P: PRESCRIPTION 55

at discriminating depression from other presenting problems or diagnoses. The sheer number of published measures creates an obstacle to identifying the instruments with the best validity for the purpose of detecting depression. It would take a prohibitive amount of time for a practitioner to obtain and evaluate all 200 tests, let alone survey all of the published literature on each. As reviewed earlier, neither graduate training nor internship supervision was likely to provide thorough exposure to the different tests available for this purpose (only the Beck Depression Inventory and the Child Behavior Checklist appear in the top 20 frequently taught or administered tests based on multiple surveys). If clinicians consulted a handbook or assessment text, they would receive no guidance on the comparison of tests in terms of diagnostic efficiency.

Evidence-based medicine recommends turning to online search engines to answer clinical questions (Straus et al., 2011). In this case, the question might be formulated as: “What is the best assessment tool to improve diagnosis of depression in children and adolescents?” Searching MedLine using the term “Depression” in any field yields an unmanageable number of hits, with 281,347 records (all searches conducted in May 2007). However, limiting the search to “children OR adolescents,” AND “depression and diagnosis” AND “sensitivity and specificity,” which is the recommended Medical Sub-Heading (MeSH) term, to identify studies of diagnostic methods (Straus et al., 2011) produces 1,168 hits, of which 34 are reviews. A quick scan of the titles of the reviews eliminates several as being peripherally related (depression as correlate of cancer or recurrent abdominal pain), and finds a review that synthesizes 160 studies covering 33 different diagnostic and symptom assessment measures (Brooks & Kutcher, 2001).

Based on this quick search and review, the list of candidate instruments dropped from several hundred to a half dozen. Having identified these as the contenders, the clinician could then concentrate on the more recent studies and attend only to articles reporting AUCs of ∼.8 or higher in outpatient samples (i.e., performing as well as or better than the tests covered in the Brooks & Kutcher review). If an instrument demonstrated higher diagnostic efficiency, then the article would warrant more careful scrutiny to determine if the results were valid (Bossuyt et al., 2003) and if they were more applicable to the specific patient in question (Jaeschke, Guyatt, & Sackett, 1994). This brief exercise indicates several important general conclusions: (a) both typical clinical training and review materials (textbooks, handbooks) often fail to provide the information necessary to answer clinically relevant questions, (b) relevant information often exists in the published literature, (c) rapid online searches using appropriate keywords can find clinically relevant evidence, and (d) online searches can both identify high-quality reviews and also augment them with more current clinical evidence. With regard to the specific question of assessing depression in youths, the evidence indicates that of the plethora of available tests, only a handful have been investigated in sufficient detail to consider adoption as an evidence-based tool for the purpose of aiding diagnosis. Thus, this exercise simultaneously illuminates important next steps to take in research as well as guiding evidence-based approaches to clinical assessment.

Applying Test Results to Individual Cases

Both clinicians and consumers are primarily concerned with the status of an individual child, not with abstract properties of instruments applied to large groups. It is possible to take information about the probability of a diagnosis and combine it with the test result to estimate a new probability (“posterior probability”) that the person has the diagnosis in question. The algebra needed to integrate these two pieces of information has been

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .

56 STRATEGIES FOR EVIDENCE-BASED ASSESSMENT OF CHILDREN AND ADOLESCENTS

known for several centuries in the form of Bayes’ theorem. However, recent approaches have attempted to simplify the application so that the clinician need not formally do computations to estimate the posterior probability (Jenkins, Youngstrom, Washburn, et al., 2011; Straus et al., 2011). Figure 2.2 provides a nomogram designed to aid the clinical use of test results. The left-hand column charts the prior probability. The middle column contains the likelihood ratio, which is simply the ratio of the percentage of cases with the diagnosis that would exceed the cutoff (i.e., the sensitivity of the test) divided by the percentage of cases without the diagnosis that would also exceed the cutoff (the false alarm rate, or the complement of specificity). A practitioner would simply put a dot on the left-hand line to indicate the prior probability, then put a dot on the middle line corresponding to the likelihood ratio (the sensitivity divided by false alarm rate), connect the two dots with a straight line, and finally extend the line across the right-hand line

.1

1,000 500

200 100 50

20

10 5

2

1 .50

.20

.10

.05

.02

.01

.005

.002

.001

99

95

90

80

70

60

50

40

30

20

10

5

2

1

.5

.2

.1

.2

.5

1

2

5

10

20

30

40

50

60

70

80

90

95

99 Pretest Probability

% %

Likelihood Ratio Posttest Probability

FIGURE 2.2 Nomogram for Combining Probability With Likelihood Ratios

Psychopathology : History, Diagnosis, and Empirical Foundations, edited by Linda W. Craighead, et al., John Wiley & Sons, Incorporated, 2013. ProQuest Ebook Central, http://ebookcentral.proquest.com/lib/ashford-ebooks/detail.action?docID=1380177. Created from ashford-ebooks on 2021-10-31 19:43:13.

C o p yr

ig h t ©

2 0 1 3 . Jo

h n W

ile y

& S

o n s,

I n co

rp o ra

te d . A

ll ri g h ts

r e se

rv e d .