1 / 50100%
YES, IT HAS A MANUAL: THE ASSOCIATIONS BETWEEN THE RISB-2 AND
PSYCHOLOGICAL SYMPTOMS, DIAGNOSES, AND QUALITY OF LIFE
CHAPTER I
INTRODUCTION AND LITERATURE REVIEW
Rationale for Research
Though the necessity and appropriateness of diagnoses in the mental health professions remains
contested (Sartorious, 2015), the current state of evidence-based practice applies specific treatment
approaches to patients exhibiting specific diagnoses (e.g. exposure and response prevention for
obsessive-compulsive disorder); which treatment is applied in a given case is, ideally, based upon
the research evidence accumulated in support of a given approach’s effectiveness in ameliorating
symptoms of a given diagnosis (Division 12, 2016). As evidencebased practice hinges on
accurately matching a treatment approach with a given patient’s presenting problem, the
importance of accurate assessment cannot be understated. Accurate assessment, however, can be
difficult, particularly at the outset of therapy when the clinician has not yet established an
understanding of a patient’s behavior across time and situations.
Personality disorders, in particular, can be difficult to diagnose at the outset of treatment due to
their conceptualization as enduring and inflexible behavioral patterns, as well as the “apparent
competence” that individuals presenting with personality pathology may show (Koerner &
Linehan, 2012; Ward, 2004).
Though established instruments assessing for the presence of personality pathology exist, they can
be time-intensive to administer and impractical for standard screening in practice. By contrast,
sentence completion tests tend to be quick to administer and are already among the most
commonly-administered personality assessment instruments (Holaday et al., 2000). The Rotter
Incomplete Sentences Blank (RISB; Rotter, 1949), a well-established sentence completion test,
was informally observed to yield higher scores in patients with diagnosed personality pathology
decades ago (Fulkerson & Gettys, 1965), though it is only recently that this observation was
statistically validated in a study involving the second edition of the RISB (Torstrick et al., 2015).
As the RISB has repeatedly demonstrated strong associations with both clinical status (Church and
Crandall, 1955; Lah, 1989) and specific symptom presentations (Atchinson, 1968; Cohen, 1966),
the present research sought to replicate, by a different methodology, Torstrick and colleagues’
(2015) findings that the RISB-2 is significantly associated with the presence of personality
pathology, thus establishing it as an effective and efficient screening device. The present research
also examined the RISB-2’s association with
other clinically-relevant constructs such as depression, anxiety, and quality of life.
Objective and Projective Assessment
Assessment is a key domain of clinical psychology (Benjamin, 2005). One of the tools
available to the clinical psychologist conducting an assessment is testing (APA, 2013). Testing
refers to the use of a formal instrument, or “test,” to evaluate a client’s thoughts, attitudes,
actions, or other tendencies or behaviors (APA, 2013). Tests are used in addition to less formal
or structured methods, such as a clinical interview (APA, 2013).
Psychological tests can be classified in many different ways, including what they purport
to measure (e.g., personality traits, cognitive ability, academic achievement, etc.), whether they
measure typical behavior or maximal performance, and the nature of the test stimulus (Institute of
Medicine, 2015). In terms of the nature of the test stimulus, though many different formats of test
exist, a broad dichotomy that is often used for classification purposes is that of “objective” and
“projective” tests (Lichtenberg, 1985). Though Lichtenberg (1985) argued that the distinction
between objective and projective tests may not be as clear as it is generally assumed to be, one
way to distinguish between these two types of tests is in looking at the examinee’s freedom of
response. Specifically, objective tests tend to allow examinees to respond to structured stimuli with
limited options (e.g. true/false and Likert-scale responses to statements or questions), whereas
projective tests tend to allow examinees to respond to an ambiguous stimulus with near-infinite
variation and flexibility (e.g. completing a sentence, interpreting an inkblot). An underlying
assumption of projective tests is that the projective test encourages the client to respond to the
ambiguous stimulus in a way that reveals important information about his or her cognitive style,
attitudes, or personality more generally (Clark, 1995). By contrast, objective tests simply compare
a client’s pattern of responses to an established criterion or reference norms in order to establish
where the examinee stands on a trait compared to a larger sample (Clark, 1995).
Though the underlying logic of projective tests makes sense at first glance, the requirement
that a clinician interprets a client’s unstructured response has resulted in debate over the reliability
and validity of any information gained (Wiederman, 1999). Indeed, it is impossible to calculate a
coefficient for interrater reliability when qualitative impressions are compared. Similarly, efforts
to correlate such qualitative data with outcome measures or other criteria, in an attempt to establish
validity evidence, would be rife with methodological problems.
In an effort to address the concerns about reliability and validity discussed above, several
test developers, as well as independent researchers, have developed “objective,” or otherwise
quantifiable, scoring systems for traditionally projective tests. Perhaps the most famous of these
is Exner’s s (1974) method of interpreting the Rorschach (WPS, 2018). However, the present
research will focus on the scoring system described in the manual of the Rotter Incomplete
Sentences Blank, Second Edition (RISB-2; Rotter et al., 1992).
The RISB-2
Background
The RISB-2 is one of several sentence completion tests in use by clinicians today. Extant sentence
completion tasks differ in terms of theory (e.g. psychodynamic, Murray’s Theory of Needs,
atheoretical, etc.), target population, and scoring procedures (Holaday et al., 2000); however, as
implied by their name, each test in this broad category requires the examinee to respond to a neutral
prompt. Responses, in turn, are then interpreted by the clinician, whether purely subjectively or
according to a scoring system. Holaday and colleagues (2000) found that sentence completion tests
are commonly used by clinicians but are rarely scored/interpreted in in a systematic way as dictated
by the associated manual.
Description of the Test
The RISB-2 (Rotter et al., 1992) is the latest iteration of Rotter’s (1949) original test
(RISB). The RISB was, in turn, derived from earlier work by Shor (1946), Hutt (1945), and
Holzberg and colleagues (1947). Like its predecessor, the RISB-2 consists of 40 sentence stems
(e.g. “I like…”), each of which the examinee completes in order to make a full, meaningful
sentence (e.g. “I like… ice cream.”). The RISB-2 (Rotter et al., 1992) can be, and often is,
interpreted solely in terms of the clinician’s qualitative, subjective appraisals of the patient’s
responses. It, along with other sentence completion measures, is often used, either in isolation or
as part of a larger personality assessment battery, in this manner (Holaday et al., 2000).
Scoring
Despite the RISB-2’s frequent use as a qualitative instrument, the quantitative scoring
system, originally developed by Rotter and colleagues (1949) based on the psychodynamic
conceptualization of conflict, allows a clinician to rate a given sentence on a 7-point scale. For
each sentence stem, the RISB-2’s (Rotter et al., 1992) manual provides sample responses for
each possible rating (i.e. 0-6). Scoring is performed through comparison of the examinee’s
response to the given examples, in a manner not unlike the scoring of the Vocabulary and
Similarities subtests of the Wechsler Adult Intelligence Scale (WAIS-IV; Wechsler, 2008).
Scores assigned to individual sentences are then summed, and the overall score may be
compared to a cut score, with higher scores generally assumed to represent higher experienced
conflict (Lah, 1989; Rotter et al., 1949).
Reliability Evidence
Numerous studies have investigated the reliability of the RISB and the RISB-2. Though reliability
may be assessed through multiple metrics, as Wiederman (1999) points out, with projective
measures, the primary point of contention between proponents and opponents is the consistency
of interpretations across clinicians. As such, inter-rater reliability stands out as a key concern in
the case of the RISB-2. In Rotter and colleagues’ (1949) original study, inter-rater reliability
coefficients of .96 and .91 were found for females and males, respectively. The RISB-
2’s manual (Rotter et al., 1992) cites multiple studies investigating the inter-rater reliability of
the RISB-2’s scoring system, reporting coefficients ranging from .72 (Feher et al., 1983) to .99
(Snow, 1972; Vernallis et al., 1970), with a reported median coefficient of .93. Moreover,
estimates of internal consistency of .83 and .84 (females and males respectively) were found in
the original Rotter’s original validation study (1949). More recently, Torstrick and colleagues
(2015) found similar numbers, with a combined internal consistency coefficient of .81, and
interrater reliability coefficients ranging from .85-.95 among pairs of graduate-level raters.
These numbers are consistent with generally-accepted guidelines for minimum standards in
clinical use instruments (Cook & Beckman, 2006); moreover, they are not much lower than
those of goldstandard psychometric instruments, such as the WAIS-IV, which boasts index
internal consistency estimates of .87-.98, subtest coefficients of.71-.96, and, on subtests
involving examiner judgment, interscorer-reliability estimates ranging from .91 to .97
(Wechsler, 2008). In other words, when scored in accordance with the manual, the RISB-2 has
consistently demonstrated reliability estimates that are well within the acceptable range for
clinical application.
Validity Evidence- General
The RISB and RISB-2 have also accumulated substantial validity evidence. Rotter
initially sought to validate the RISB by administering it to 124 undergraduate students, who were
recruited from either classes on campus or from clinicians on campus (Rotter et al., 1949). In
either case, the referring professor or clinician was asked to declare each referred participant as
either “adjusted” or “maladjusted.” Rotter and colleagues (1949) then proceeded to compute
biserial correlation coefficients to determine the correlation between group membership and
RISB score. Perhaps unsurprisingly given their questionable methodology, which assumed a
given participant’s professor would have knowledge of that participant’s degree of adjustment,
their results were inconsistent; specifically, while they found a difference of means between the
male students identified as maladjusted by their professors and the male counseling clients
reaching the .01 level significance, the differences between the students labelled as
“maladjusted” and “adjusted” by their professors not consistently significant across classes
(Rotter et al., 1949).
While Rotter and colleagues’ original validation study was very flawed, methodologically
speaking (Rotter et al., 1992), due to its assumption that professors were knowledgeable raters of
their student’s mental health status, subsequent studies have looked at the RISB-2’s scoring system
in a variety of contexts. In 1955, Church and Crandall attempted to validate the RISB using a
different methodology; they drew their sample from both a small college, where the RISB was
routinely administered to incoming freshmen, and an existing sample of adult women participating
in a longitudinal study of human development. In the sample of 344 college students (the adult
women were used solely for inter-rater and test-retest reliability coefficient calculation), they used
subsequent requests for counseling services as the criterion variable, and calculated biserial
correlations that were moderate in strength, and significant at the .01 level, between RISB score
and group membership (r=.42 and .37 for women and men, respectively; Church & Crandall,
1955).
Three decades later, in an attempt to update the RISB’s cut scores, Lah (1989) conducted
a series of two studies. In the first study, Lah (1989) recruited 116 participants from two
fraternities and two sororities at a university. He administered to each participant the RISB and a
sociometric measure which asked the participant to nominate peers who were present on several
continua between negative and positive personality traits (e.g. happy-unhappy; Lah, 1989), as
well as on the continuum of neat to sloppy, which was considered a control item not theoretically
associated with adjustment. In addition, participants were asked to rate the members of the group
with whom they were most and least friendly (Lah, 1989). Data analysis consisted of checking
the distribution of ratings within each group, testing the homogeneity of the four groups of
participants, testing the internal consistency of the sociometric measure (alpha coefficients on
trait ratings ranged from .70-93, with the exception of one outlier, easygoing-driving in group 3,
which had a coefficient of .48), and testing the inter-rater reliability of RISB with a sample of
20% of the protocols, by comparing the author’s ratings to those of a graduate student who used
the manual with minimal training (r=.94; Lah, 1989). Lah (1989) then correlated participants’
overall rankings on the positive-negative trait continua with their maladjustment scores, and found
that ratings of personality traits were significantly negatively correlated with the RISB
maladjustment scores (i.e. higher happiness was associated with lower maladjustment). The
correlations between traits and maladjustment ranged from -.12 to -.40, and were significant, with
one meeting significance at the .05 level (independent-dependent), two meeting significance at the
.01 level (open-defensive & easygoing-driving), three meeting significance at the .001 level (calm-
nervous, self-accepting-nonself-accepting, & humor-humorless), and one meeting significance at
the .0001 level (happy-unhappy; Lah, 1989). The second study’s methodology was more consistent
with Church in Crandall’s (1955), in that Lah (1989) compared the RISBs of 120 university
students who had been self-referred to campus counseling with those of 120 controls who were
recruited from introductory classes and fraternities and sororities associated with the university. It
is worth noting that the author plainly states that he did not attempt to ensure the control group had
not sought counseling other than comparing the groups to ensure no participant was included in
both (Lah, 1989). As such, it is possible that some members of the control group actually were in
counseling and, as such, Lah (1989) cautions that the control group be seen as “average,” as
opposed to “well-adjusted.” Data analysis of the second study included testing inter-rater
reliability between the investigator and a graduate assistant (r=.90) and comparing the mean
maladjustment scores of the clinic (160.7 and 162.7 for males and females respectively) and
control (130.1 and 129.0 for males and females respectively) samples. The difference between
means was significant at the .0001 level, and Pearson Biserial coefficients of .67 and .72 (males
and females respectively) were found between group membership and maladjustment scores on
the RISB. Both correlations were significant at the .0001 level (Lah, 1989).
Validity Evidence- Specific Constructs
In terms of specific symptoms, both Atchinson (1968) and Cohen (1966) found significant
correlations between the RISB adjustment score and measures of anxiety. Most recently, Torstrick
and colleagues (2015) compared a clinical group of 72 adults, who either were currently in
psychotherapy at an outpatient clinic or previously received psychiatric treatment previously, to a
group of 69 undergraduate controls on a battery of personality and outcome measures, including
the RISB-2. All participants completed the self-report measures in a random order, and clinical
participants also underwent two structured interviews (Torstrick et al., 2015). In addition to
computing inter-rater reliability (.85-.95, depending on rater pairing), split-half (.77), and internal-
consistency (.81) statistics for the RISB-2, they performed several analyses of the RISB-2’s ability
to discriminate between groups and contribute incrementally to variance accounted for by other
measures (Torstrick et al., 2015). In terms of discriminant validity, they found a significant
difference between the two groups on mean maladjustment score, as well as between clinical
participants with and without an active Axis I diagnosis; both findings were significant at the .001
level, and they exhibited medium (d=.58) and large (d=.94) effect sizes, respectively (Torstrick et
al., 2015). In terms of criterion validity, they found significant negative correlations between the
RISB-2 maladjustment score, self-reported life satisfaction (r=-.50), and clinician-rated global
assessment of functioning (GAF; r=-.29), with the former finding significant at the .001 level, and
the latter significant at the .05 level (Torstrick et al, 2015). In addition, they found a significant
difference in mean maladjustment scores between clinical participants with (M=161.36) and
without (M=145.52) a personality disorder diagnosis. In terms of construct validity, the RISB-2’s
maladjustment score correlated in the expected direction with several constructs measured by
various inventories of personality, including, at the .001 significance level, neuroticism (r=.50),
self-downing (r=.34), need for achievement (r=.28), need for approval (r=.27), need for comfort
(r=.29), cold/distant (r=.35), socially inhibited (r=.40), nonassertiveness (r=.35), and overly
accommodating (r=.31). In addition, Torstrick and colleagues found that, even after accounting for
general psychological distress (GPD), participant scores on the RISB-2 explained significant
variance in several criterion variables, including satisfaction with life. In addition to these
correlational analyses, they compared the results of various cut scores for different diagnoses, and
found that a cut score of 130 provided the best overall correct classification (OCC) for any DSM
diagnosis, whereas 155 provided the best OCC for Axis II diagnoses specifically, and had a better
OCC than a measure specifically designed to assess Axis II pathology (Torstrick et al., 2015),
indicating that the RISB-2 may have some utility as a screener for personality pathology. Finally,
they found that the RISB-2 added incremental validity to a model for three of four criterion
variables, including life satisfaction, after general psychological distress (GPD) was included as a
covariate (Torstrick et al., 2015), indicating that the RISB-2 accounts for variance in these criterion
variables over and above
GPD. Though Torstrick and colleagues’ (2015) study is not without its flaws (e.g. structured
interviews were only administered to clinical participants, clinical participants mean age was
significantly older than that of controls), it nevertheless provides a wealth of data that supports
the utility of the RISB-2 and establishes directions for future research.
Validity Evidence- Diverse Samples
Recently, in an effort to assess the validity of the RISB-2 with a non-white population,
Logan and Waehler (2001) administered the RISB-2 to 100 white and 94 African American
students at a midwestern university. Participants were members of fraternities and sororities
associated with the university, and were recruited through the Office of Greek Affairs. The
groups were comparable in terms of age, mean number of years in attendance, and gender
composition (Logan & Waehler, 2001). In addition to the RISB-2, the authors (Logan &
Waehler, 2001) administered a test designed to assess response bias tendencies, specifically
SelfDeceptive-Enhancement (SDE) and Impression Management (IM). Data analysis included
interrater reliability of the RISB-2 (coefficients of .89, .90, .96 among pairs of graduate students
who used the manual, t-tests comparing scores of protocols scored by multiple raters were
nonsignificant; Logan & Waehler, 2001), as well as comparison of RISB-2 scores across sexes and
racial groups. The authors found that scores did not differ between sexes within racial groups, and
so they collapsed males and females within racial groups into a single group for the purposes of
racial comparison. In comparing racial groups, they found no significant difference in mean scores
between white and African American students, however, when comparing classification rates
based on different cutoffs, they found that the cutoff of 135 suggested in the manual resulted in
significantly more African American students being classified as maladjusted (but not the other
cutoffs of 125, 130, and 140; Logan & Waehler, 2001). In addition, in analyzing the difference
between the two groups on the response bias test, they found that the RISB-2 maladjustment score
was significantly negatively correlated with SDE, with the correlation significantly stronger for
whites than African Americans (r= -.335 for the total sample, r=-.292 for the African American
group, r=-.457 for the white group; Logan & Waehler,
2001). Moreover, they found that the white group engaged in significantly more Impression
Management than the African American group (r=-.237 for whites vs. r=-.152 for African
Americans; Logan & Waehler, 2001). They concluded that, as a result of this finding, the greater
classification of African Americans by the standard RISB-2 cutoff may be reflective of their
responses more accurately portraying their experiences (Logan & Waehler, 2001).
Additional Scoring/Interpretive Systems
In addition to Rotter and colleagues’ (1949) original scoring system, at least two
alternative scoring systems have been developed for the RISB family of tests, independent of
Rotter and colleagues. The first, developed by Renner, Maher, and Campbell in 1962, is designed
to assess an examinee in terms of anxiety, dependency, and hostility. To do this, they
operationalized the three personality traits in question by developing scoring keys which
contained behaviors associated with the target traits (e.g. for anxiety, nightmares, phobias,
inability to relax, and being easily startled are considered referent behaviors; Renner et al., 1962).
They then drew fifty RISB protocols from a pool of 206 obtained from students in an introduction
to psychology class, and used these protocols to establish examples of 0-point, 1point, and 2-point
responses (Renner et al., 1962). Their scoring system was designed such that direct references to
the target behaviors are given a score of “2,” whereas indirect or qualified references are given a
score of “1,” and a response with no reference to any of the behavioral referents in the scoring
keys is given a score of “0 (Renner at al., 1962).” These guidelines, examples, and scoring keys
were then compiled into a score-by-example manual, not unlike that of the RISB manual itself
(Rotter, 1949), or the verbal subtests of a Wechsler-series cognitive test (Wechsler, 2008; Renner
et al., 1962). In testing the inter-rater reliability of this system on additional protocols pulled from
the initial development pool of 206, the authors found coefficients of .88, .83, and .94 for anxiety,
dependency, and hostility respectively (Renner et al., 1962). To test the validity of the system,
they administered to 44 women in the dorms at one university and 48 men in the dorms at another
university both the RISB and a sociometric measure asking the participants to rate themselves and
their dorm-mates on 27 variables, including anxiety, dependency, and hostility (Renner et al.,
1962). In the female sample, they found a significant correlation between peer-rated anxiety,
(r=.43, significant at the .01 level), self-rated anxiety (r=.34, significant at the .05 level) and self-
rated hostility (r=.48, significant at the .01 level) and their respective traits scores on the RISB.
In the male sample, they found significant correlations between peer-rated hostility (r=.34,
significant at the .05 level) and selfrated dependency (r=.37, significant at the .01 level) and their
respective trait scores on the RISB
(Renner at al., 1962). In an effort to establish discriminant validity, they also compared correlations
between the two raters of the same traits (e.g. rater-1-scored anxiety x rater-2-scored anxiety) and
different traits (e.g. rater-1-scored anxiety x rater-2-scored hostility) and found that different traits
correlated much lower than they did to different rater’s ratings of the same trait (Renner at al.,
1962). Finally, the authors also calculated coefficient alphas for each trait, separated by sex; these
values were less impressive, with the highest being .74 (male dependency) and the lowest being -
.02 (female hostility). As such, while some of Renner and colleagues’ (1962) are promising, their
system clearly was not sufficiently valid and reliable for
clinical use as of publication.
The second alternative scoring system for the RISB family of tests, the Cognitive Rating
Form (CRF) is based on the cognitive theory of depression (Clark & Steer, 1996; Lehnert et al.,
1996). The CRF is based on responses to 25 of the original 40 stems of the RISB-2, and, in
keeping with the theory of cognitive therapy (Clark & Steer, 1996), analyzes responses in terms
of 25 types of cognitions, including: complex/abstract thinking, simple/concrete thinking,
relativistic thinking, blaming, and extreme thinking (Lehnert et al., 1996). The assessed
cognitions also include the response’s focus (e.g. self/other), whether it is active or passive, and
whether it is positive or negative in outlook or emotional valence (Lehnert et al., 1996). Lehnert
and colleagues (1996) attempted to validate this alternative scoring system using a sample of 56
adolescent psychiatric inpatients and 102 high school controls. They found that, of the 25
categories of cognitions assessed, 10 met their stated inter-rater reliability criterion of .70 or
better (Lehnert et al., 1996). These 10 categories were used for subsequent analysis. In
examining internal consistency, they performed a factor analysis, and decided on a four-factor
solution using the “Kaiser rule” and scree plot analysis (Lehnert et al., 1996). These factors
included “cognitive complexity,” “negative cognitions,” “negative others,” and “positive view of
the future (Lehnert et al., 1996).” Analysis of the CRF’s ability to discriminate between the
study’s two participant groups revealed that negative cognitions correlated with group
membership at the .00001 level of significance, though the other three factors did not correlate
significantly with group membership (Lehnert et al., 1996). In comparing diagnostic accuracy
between the CRF and a battery of other measures comprised of the Children’s Depression
Inventory, Hopelessness Scale for Children, and the Rosenberg Self Esteem Scale (CDI, HSC,
and RSES, respectively; Lehnert, Overholser, & Adams, 1996), they found that the CRF had a
greater sensitivity rate than the CDI, which, in their model, was the only one of the three possible
predictors to explain significant variance (50% vs. 44%; Lehnert et al., 1996). Although Lehnert
and colleagues’ (1996) study does not boast the same impressive reliability coefficients that
Rotter’s (1949) scoring system has established over the past decades, it appears to at least have
potential, and warrant future investigation.
Administration
According to the publisher, the RISB-2 requires approximately 20-40 minutes to administer
(Pearson, 2021). When this is taken into consideration, along with the reliability and validity
evidence discussed above, a strong case is made for using the RISB-2 having potential as a screener
in clinical settings (Renner et al., 1962). Specifically, it appears to be capable of screening for
personality disorders, as, when a cutoff of 155 was applied in Torstrick and colleagues’ (2015)
sample, the RISB-2 demonstrated sensitivity of .50 and specificity of .83, resulting in an overall
correct classification (OCC) of .76. Given the RISB-2’s already widespread use as a qualitative
assessment tool (Torstrick et al., 2015), as well as its relatively short time to administer compared
to many other measurements commonly used to assess for serious psychopathology, such as the
MMPI-2-RF (35-50 minutes; Pearson, 2021) and the SCID5 PD (30-120 minutes; APA, 2021), it
seems a simple, practical, and efficient way for clinicians to gather both valuable qualitative data
and simultaneously screen for more serious personality pathology.
Shortcomings of the Present Literature
Despite the argument made above, and though the RISB-2’s sensitivity to overall
psychological distress is well established, Torstrick and colleagues’ (2015) data on the RISB-2
and personality disorders represents only a single study. The present author is only aware of one
other study which specifically looked at the relationship between the RISB’s score and personality
disorders, and while that study noted that patients with personality disorders generally had higher
maladjustment scores than those with psychotic disorders and other disorders, they reported that
this trend only approached significance, not quite reaching the .05 level (Fulkerson & Gettys,
1965); as such, this observation cannot in good faith be considered empirical evidence of an
association between the RISBV maladjustment score and the presence of personality pathology.
Additionally, the present author is unaware of any research other than that of Torstrick and
colleagues (2015) to look at the RISB-2’s ability to predict outcomes above and beyond GPD. As
such, the present study aims to investigate the RISB-2’s relationship with several other constructs
in a way methodologically distinct from that of previous studies.
Anxiety, Depression, and General Psychological Distress
Anxiety disorders are a diverse class of mental disorders in the Diagnostic and Statistical Manual
of Mental Disorders (DSM-5) that are marked by the common feature of worry or fear that is
reasonably considered excessive or out of proportion in a given context (APA, 2013).
Depressive disorders represent a subset of the DSM-5’s mood disorders, and are marked by
sadness, lack of energy, a lack of interest in previously enjoyable activities, and disturbances in
appetite and sleep (APA, 2013). Both anxiety and depression are extremely common, with
estimated prevalence rates ranging from 6% to 23.4% (Leon et al., 1995; NIMH, 2017) for anxiety
and 1.5-19% (Kessler, 2012) for depression. Moreover, both anxiety and depression are associated
with substantial social burden, including disrupted education, healthcare utilization, disrupted
relationships, and disability (Kessler, 2012; Leon et al., 1995). Anxiety and depression, though
currently conceptualized as independent disorders, commonly co-occur, with comorbidity
estimates as high as 60% (Salcedo, 2018). As a result, while standalone models of anxiety (e.g.
Wells, 1999) and depression (e.g. Lewinsohn, 1986) exist, some have proposed models that are
combined or progressive, where depression results from anxiety (Barlow, 1991; Clark & Steer,
1996).
In keeping with models that view anxiety and depression as interrelated constructs, the concept of
general psychological distress (GPD) stems from work studying the possibility that many negative
emotions and subjective experiences can be represented by a single higher-order construct,
originally proposed as “demoralization (Veit & Ware, 1983).” This concept of undifferentiated
negative experience has since become commonly referred to as GPD (Veit & Ware, 1983). While
GPD is conceptualized differently by different investigators (e.g. Gulliver et al., 2012; Torstrick et
al., 2015), depression and anxiety are two components that are almost always included, with other
factors such as somatization, or an external locus of control, included at the discretion of
investigators. Studies investigating GPD have found that GPD significantly correlates with
variables such as group membership (clinical vs. nonclinical), and even accounts for enough
variance that previously significant associations are rendered nonsignificant after GPD is entered
into structural models (Torstrick et al., 2015). As a result, interventions for GPD, as opposed to
specific diagnoses such as Major Depressive Disorder or Generalized Anxiety Disorder, have
recently begun to emerge in the literature (APA, 2013;
Blainey et al., 2017).
Personality Disorders
In previous editions of the DSM (APA, 1980; APA, 2000) a multiaxial system of
diagnosis was used. Fives axes existed under this system, with Axis I representing most mental
illness (e.g. major depressive disorder), Axis II representing personality disorders (e.g. borderline
personality disorder) and developmental disorders (e.g. mental retardation), Axis III representing
comorbid medical conditions (e.g. diabetes), Axis IV representing other psychosocial conditions
relevant to treatment (e.g. exposure to war), and Axis V representing the Global Assessment of
Functioning, a value on a 100-point scale which represented an overall impression of the client’s
current functioning (APA, 1980; APA, 2000). Conceptually, Axes I and II, while both representing
mental disorders, were separated because the disorders they represented were thought to have
different degrees of persistence; specifically, personality disorders and developmental disorders
were placed together on Axis II because they were believed to be more enduring, whereas the
disorders placed on Axis I were believed to be more transient in nature (Widiger & Shea, 1991).
This system has since been abandoned (APA, 2013), and research has called into question the
separation of at least some personality disorders (e.g. borderline personality disorder) from other
disorders, such as mood disorders (Links & Eynan, 2013; Parker, 2014). Still, the strong
association of personality disorders with poor outcomes (e.g. inpatient treatment, suicide; Paris,
2019) as compared to other disorders (e.g. mood disorders; Angst et al., 1999; Isometsä, 2014), as
well as the relative difficulty of treatment (1-3 years of treatment vs. 3-4 months of treatment for
depression; Division 12, 2016) supports the clinical importance of being able to effectively identify
clients who are suffering from personality pathology, especially in the case of clients who may
appear more competent and functional than they actually are (Koerner & Linehan, 2012).
Negative Cognitions
Negative cognitions, whether referred to as such, as cognitive distortions, or as irrational
beliefs, are illogical, irrational, or otherwise unhelpful thoughts or schemata which are theorized
by cognitive therapists to play a role in anxiety, depression, and general psychological distress
(David et al., 2004). Consistent with this theory, much research has demonstrated the association
between negative cognitions and various forms of psychopathology and distress (McDermut et
al., 1997; Tecuta et al., 2019). Moreover, Cognitive Behavioral Therapy (CBT; Beck Institute,
2020), an empirically-supported treatment for anxiety, depression, and other focuses of clinical
attention (Division 12, 2016) is structured around the assumption that negative cognitions play a
role in maintaining a variety of psychological disturbance.
In addition to their role in more common anxiety and mood disorders, cognitive distortions
are believed to be associated with more severe and enduring psychopathology. Specifically,
negative cognitions are hypothesized to have a role in borderline personality disorder (Baer et
al., 2012) and narcissistic personality disorder (Zeigler-Hill et al., 2011). As such, negative
cognitions appear to be generally associated with psychological disturbance.
Purpose
The aim of the present study is to add to the body of literature investigating the reliability of the
RISB-2, as well as to attempt to replicate, by a different method, findings that the RISB-2
maladjustment score is associated with several clinically relevant constructs. There is a
considerable body of literature (e.g. Church & Crandall, 1955; Lah, 1989) demonstrating that
Rotter’s scoring system, based on the psychodynamic conceptualization of conflict, has utility in
differentiating the patient population from the non-patient population. In addition, there is evidence
that the maladjustment score of the RISB and RISB-2 is associated with specific symptom
presentations such as anxiety (Atchinson, 1968; Cohen, 1966) and, at high levels of maladjustment,
personality pathology (Torstrick et al., 2015). The confirmation and clarification of these
relationships by additional research would add support for the use of the RISB-2 as a screening
instrument in clinical practice. Given this, the present study attempted to establish that the RISB-
2 maladjustment score is positively correlated with anxiety, as in Atchinson (1968) and Cohen
(1966), as well as depression and dysfunctional attitudes. The present study also attempted to
replicate the finding that the RISB-2 maladjustment score is negatively correlated with quality of
life, as well as that it differs significantly between diagnostic groups (e.g. healthy, no personality
disorder, personality disorder; Torstrick et al., 2015). Finally, the present study attempted to
corroborate that the significant correlation between the RISB-2 and quality of life remains
significant when general psychological distress is taken into account (Torstrick et al., 2015.)
The present study will add to the existing literature on the RISB-2’s reliability and validity, and
will attempt to replicate Torstrick and colleagues’ (2015) findings that the RISB-2 may be useful
in screening for personality pathology.
CHAPTER II
METHOD
Participants
Participants were recruited via requests for participation posted by the present author on
social media. The present author requested those connected to him on social media share these
posts in order to reach as many potential participants as possible. Prospective participants were
directed via a link to Qualtrics, where they were presented with information regarding their rights
and a request for informed consent. The author intended to have a sample of at least 270,
consistent with G*Power (HHU, 2020; IDRE, 2021) required sample estimates for one-way,
three level ANOVAs (both omnibus, fixed effects and fixed effects, special, main effects, and
interactions, effect size= .25, alpha level= .05, power= .8, numerator df= 10, 3 groups). Target
sample size is based on the requirements for a one-way ANOVA because the correlational
analyses required to test most of the present study’s hypotheses require significantly fewer
participants (Brysbaert, 2019), however, reducing the number of participants to these numbers
would render the ANOVA required to test that hypothesis underpowered.
Materials
The above hypotheses were tested by analyzing participant responses to a demographics
questionnaire, a mental health history questionnaire, the RISB-2, the Big Five Inventory (BFI),
the Patient Health Questionnaire (PHQ-9; Kroenke et al., 2001), the Generalized Anxiety
Disorder-7 (GAD-7; Spitzer et al., 2006), the short form of the Dysfunctional Attitude Scale (DAS-
SF1), and the Abbreviated World Health Organization Quality of Life (WHOQOL-BREF;
WHO, 1996).
Demographics Questionnaire- Participants are asked to indicate their age, gender, sexual
orientation, education level, household income, political affiliation, religious affiliation, and
ethnicity. See appendix A for the proposed measure.
Mental Health History Questionnaire- Participants are asked to disclose information on both
their current mental health and their lifetime mental health history. For both current mental health
and lifetime mental health history, participants are asked if they have one or more official
diagnoses from a credentialed healthcare provider. Those that report having an official diagnosis
in either category will be asked to list all diagnoses that apply. See appendix B for the measure.
RISB-2- Described in detail above.
Big Five Inventory- The BFI was developed as a brief, self-report measure of the five broad
domains of personality, as well as the lower-order facets within each domain (John et al., 1991).
The BFI is a 44-item instrument which asks the examinee to rate themselves on a 5-point scale on
a series of statements (Berkeley Personality Lab, 2009). As an example, item one states “I see
myself as someone who is talkative,” and offers response options of “1- disagree strongly,” “2-
disagree a little,” “3- neither agree nor disagree,” “4- agree a little,” and “5- agree strongly.” The
BFI boasts impressive psychometric properties given its brief nature, including a test-retest
coefficient of .84 (Rammstedt & John, 2007). Moreover, overall intercorrelation between scales is
.21 (Rammstedt & John, 2007), suggesting that, consistent with the five-factor model, the scales
of the BFI are distinct personality constructs that do not highly correlate with one another. Patient
Health Questionnaire- The PHQ-9 is a brief, self-report measure of symptoms consistent with
depression. The PHQ-9 consists of nine items which the participant will rate on a 4-point scale
ranging from 0 to 3, with higher numbers indicating greater symptom severity
(Kroenke et al., 2001). As an example, item 1 asks the examinee to rate how much they have
been bothered by “little interest or pleasure in doing things” over the previous two weeks and
offers the response choices “0- not at all,” “1- several days,” “2- more than half the days,” and
“3- nearly every day (MD+Calc, 2024b).” The PHQ-9 is a proven instrument in clinical practice
and boasts a coefficient alpha of .89 and test-retest reliability of .84 (Kroenke et al., 2001). In
addition, it is highly related to both major depression diagnostic status and functional impairment
(Kroenke et al., 2001).
Generalized Anxiety Disorder-7- The GAD-7 is a brief, self-report measure of symptoms of
anxiety. The GAD-7 consists of seven items which the participant will rate on a 4-point scale
ranging from 0 to 3, with higher numbers indicating greater symptom severity (Spitzer et al.,
2006). As an example, item 1 asks participants to rate the extent of which they have experienced
“feeling nervous, anxious, or on edge” over the past two weeks, and offers the response choices
“0- not at all,” “1- several days,” “2-more than half the days,” and “3- nearly every day
(MD+Calc, 2024a).” Spitzer and colleagues (2006) report Cronbach’s alpha of .92 and test-retest
estimates of .83, as well as sensitivity and specificity of 89% and 82%, respectively.
Dysfunctional Attitudes Scale- Short Form 1- The DAS-SF1 is a 9-item short form of the DAS,
originally authored by Weissman in 1979. The original DAS is a self-report measure which consists
of two forty-item alternate forms (A and B) derived from an initial 100-item question bank
(Beevers et al., 1979). Items on the DAS and its short forms focus on assessing irrational,
inflexible, negative, and otherwise unhelpful attitudes. As an example, item 1 states “if
I don’t set the highest standards for myself, I am likely to end up a second-rate person,” and
offers “totally agree,” “agree,” “disagree,” and “totally disagree” as response options. Internal
consistency estimates for the DAS-A’s two factors, perfectionism and need for approval, are good
(.91 and .82, respectively; Imber, et al., 1990) and, in a massive study (N=8,960) study of a
shortened version of the DAS-A (DAS-A-17; de Graaf et al., 2009), both the factors and the total
score (internal consistency estimates of .90, .81, and .91, respectively) showed significant
correlations with depressed status. Moreover, in the same sample, the scales and total score of the
DAS-A-17 together accounted for a significant 31% of the variance in depression severity, as
measured by the Diagnostic Interview for Depression (de Graaf et al., 2009). The shorter
DASSF1 was derived by applying Item Response Theory (IRT) techniques to DAS-A data from
two separate randomized clinical trials (RCT), with a combined sample of 367 participants
(Beevers et al., 2007). Specifically, the authors of this study removed items that failed to make
multiple discriminations between the 5th and 95th percentiles (Beevers et al., 2007). They
reported that the resulting short forms (SF1 and SF2) were both internally consistent (coefficients
of .84 and .83, respectively) and highly correlated with the DAS-A at both pretreatment and
posttreatment (.91 and .93 respectively at pretreatment; Beevers et al., 2007). The DAS-SF1 was
also found to correlate significantly with both the Cognitive Beliefs Questionnaire (CBQ, r=.53 ;
Beevers et al., 2007) and the BDI (r=.3). As a result, for the sake of brevity, the present study will
include the DAS-SF1, rather than its longer counterparts.
Abbreviated World Health Organization Quality of Life- The WHOQOL-BREF is the 26item
short form of the WHOQOL-100, a 100-item, self-report measure of four domains of quality of
life (physical, psychological, social, and environment; WHO, 1995; WHOQOL Group,
1998). As an example, item 1 asks “how would you rate your quality of life,” and offers response
options of “1- very poor,” “2- poor,” “3- neither good nor poor,” “4- good,” and “5- very good.”
The original instrument was derived from an initial bank of 300 items, which were tested on a
sample of 4,500 participants at 15 institutions worldwide (WHO, 1995). The WHOQOL-BREF
was in turn derived from data from the WHOQOL-100 field trials. Its domains correlate highly
with their respective counterparts from the larger instrument (.89 or higher), and multiple sources
report both good reliability estimates and validity data (Skevington et al., 2004; WHOQOL
Group, 1998). Skevington and colleagues (2004) specifically report a validation study of 11,830
sick and well participants across 23 countries. Since its initial development, the WHOQOLBREF
has been translated and validated in many different languages and populations (Berlim et al.,
2005; Min et al., 2002).
Procedure
Upon clicking the link advertised on social media, participants were directed to Qualtrics.
After reading and agreeing to the informed consent document, which was presented first,
participants completed the measures listed above in the following order: first, they completed the
RISB-2, followed by the WHOQOL-BREF and the BFI. Participants were then presented with
the DAS-SF1, the GAD-7, and the PHQ-9 in a randomized order. Finally, participants were
presented with the demographics questionnaire and the mental health history questionnaire.
The above procedure was chosen to minimize the risk of biasing participant responses
through priming. It is established in the field of psychology that recall is rife with potential for
errors, even when recalling one’s own experienced symptoms (Wells & Horwood, 2004), and
priming is a known source of bias in cognitions (Molden, 2014). By putting the RISB-2 first, the
author hoped to eliminate impacts on the participants’ responses that might occur through the
activation of thoughts associated with the other measures of the study (e.g. quality of life,
depressive/anxious symptoms). Similarly, by putting the WHOQOL-BREF and BFI immediately
after the RISB-2, the author hoped that the participants’ responses to this measure were unaffected
by possible activations by the PHQ-9, GAD-7, and DAS-SF1 of thoughts associated with
symptomatology. Finally, by putting the demographic and mental health questionnaires last, after
the three measures of symptomatology and/or negative cognitions presented in a randomized order,
the author hoped to avoid recall of symptoms being influenced by the increased availability of a
diagnostic label (e.g. I have depression, so I must have experienced crying at some point). Though
there is always the risk that the content of one measure will unduly influence a participant’s
response on another, the author expected that the order of
presentation discussed above would minimize such effects to the greatest extent possible.
An investigator blind to the participant’s mental health history and responses to items
scored the completed RISB-2 protocols. Twenty percent of the RISB-2 protocols received were
selected at random for scoring by two other investigators, also blind to participant mental health
history and item responses, for the purposes of calculating inter-rater reliability. All additional
raters have received brief training on how to use the manual to score responses from the primary
investigator.
CHAPTER III
RESULTS
IBM Statistical Package for Social Sciences (SPSS; IBM, n.d.) was used for the purposes of data
analysis. The author conducted multiple Pearson Product-Moment, Biserial, and Intraclass
correlations, as well as weighted kappas, Cronbach’s alpha, Analyses of Variance, and Multiple
Regressions. In the paragraphs that follow, these analyses will be discussed in detail, as will the
hypotheses they were intended to test, and the results obtained from these analyses.
Observed Sample
The Qualtrics survey received 191 responses. When responses that did not consent to participation
or did not complete enough of the measures to contribute meaningfully to data analysis (e.g.
stopping before the end of the RISB-2) were removed (52.36% of responses), the observed
sample contained 91 participants. The observed sample consisted of 16 (17.4%) men, 67 (72.8%)
women, one (1.1%) individual who identified as nonbinary, and one (1.1%) individual who
declined to specify. 76 participants (82.6%) identified as Non-Hispanic White, while the African
American, Asian/Pacific Islander, and Mixed-Race categories each had one (1.1%) participant
identify. Three (3.3%) participants identified as “other,” while three (3.3%) of participants
declined to specify their race. Reported age ranged from 22 to 75, with a mean age of 48.32 and
a standard deviation 16.205 for the sample. 67 participants (72.8%) of the sample identified as
heterosexual, seven (7.6%) identified as homosexual, five (5.4%) identified as bisexual, and one
(1.1%) identified as asexual. One (1.1%) participant identified as “other” sexual orientation, and
three (3.3%) declined to specify.
Hypotheses
In keeping with the above stated purpose, the present study collected data to test the following
six hypotheses. First, the authors hypothesized that the RISB-2 would demonstrate both good
internal consistency and good inter-rater reliability. Second, the authors hypothesized that the
maladjustment score on the RISB-2 would have significant, positive correlations with both
selfreported depression and self-reported anxiety. Third, the authors hypothesized that the
maladjustment score on the RISB-2 would have a significant, positive correlation with measured
negative cognitions. Fourth, the authors hypothesized that the RISB-2 would have a significant
negative correlation with quality of life. Fifth, the authors hypothesized that RISB-2 mean scores
would differ significantly between participants diagnosed with mental illness but not personality
disorders, participants diagnosed with personality disorders, and participants with no previous
psychiatric diagnosis. Finally, the authors hypothesized that the RISB-2 maladjustment score
would account for variance in self-reported outcomes over and above a combination of
selfreported depression, self-reported anxiety, and negative cognitions, which would serve as an
approximation of general psychological distress. It is worth noting that, due both to a lack of
individuals with self-reported personality disorders in the sample and unexpected findings once
data analysis began, several of these original hypotheses were rendered impossible or
unnecessary to test in their original forms. In addition, due to both these unexpected findings and
the lack of self-reported personality disorders among participants, the authors tested additional,
post-hoc hypotheses that seemed warranted in light of initial results. Specifically, the authors
observed that the PHQ-9 was at least as strong in its associations as the RISB-2 in most analyses,
as was the GAD-7 in some analyses; as a result, the authors decided to compare the RISB-2,
PHQ-9, GAD-7, and DAS-SF-1 in several analyses, rather than only conduct those analyses
using the RISB-2 as originally planned.
Findings
Interrater Reliability
In order to test the hypothesis that the RISB-2 maladjustment score would display good interrater
reliability, the authors ran a series of three Pearson Product-Moment (PPM, table 1) correlations,
consistent with previous literature. Consistent with previous literature, correlations between rater
scores of 20% of RISB-2 protocols were excellent, with all coefficients exceeding
.8.
Table 1 Pearson Product Moment Correlations Between Raters
Rater 1
Rater 2
Rater 3
Rater 1
Pearson
Correlation
x
p (2-tailed)
N
x
Rater 2
Pearson
Correlation
.947**
p (2-tailed)
<.001
x
N
22
Rater 3
Pearson
Correlation
.819**
.844**
p (2-tailed)
<.001
<.001
N
21
21
Note. ** = correlation is significant at the 0.01 level (2-tailed)
Because the PPM, a measure of association, rather than agreement, is considered less-than-ideal
as a measure of interrater reliability (D. DeRose, personal communication, September 5th, 2023),
the authors ran a series of three Intraclass Correlations (ICC, table 2), consistent with Torstrick and
colleagues (2015) methodology. Similar to the results from the PPM correlations, the ICC
coefficients were consistently high, with each over .8.
Table 2 Intraclass Correlations Between Raters
Rater 1
Rater 2
Rater 3
Rater 1
Pearson
Correlation
x
p df
x
Rater 2
Pearson
Correlation
.942**
p
<.001
x
df
21
Rater 3
Pearson
Correlation
.818**
.840**
p
<.001
<.001
df
20
20
Note. ** = correlation is significant at the 0.01 level
Additionally, in order to assess the interrater reliability of each individual item of the RISB-2,
the authors ran a series of forty weighted kappas. Though fully reproducing the results of this
series of analyses would be impractical, the frequency distribution of the average of weighted
kappas among the three rater pairs is reproduced in table 3. The observed range of weighted
kappas, averaged across the three pairs of raters for each item, was .286 to .824, with a mean of
.549 and a standard deviation of .121. Finally, the observed Cronbach’s alpha for the sample of
RISB-2 responses was .844, consistent with the hypothesis that the RISB-2 would demonstrate
good internal consistency.
Table 3
RISB-2 Item Weighted Kappa Average Among Rater Pairs
Average Weighted Kappa Among Number
Rater Pairs of
Items
.0-.1
.1-.2
.2-.3 1
.3-.4 4
.4-.5 7
.5-.6 12
.6-.7 14
.7-.8 1
.8-.9 1
.9-1
Relationships Between Measures of Psychological Symptoms
The authors ran a series of PPM (table 4) correlations to establish the degree of association
between the RISB-2 maladjustment score and/or the total scores of the measures of symptoms
administered in the present study, consistent with the hypotheses that the maladjustment score on
the RISB-2 would have significant, positive correlations with both self-reported depression and
self-reported anxiety and with measured negative cognitions. The RISB-2, PHQ-9, GAD-7, and
DAS-SF-1 were all highly correlated; all observed correlations between these measures were
significant at the .001 level or better, with the highest observed correlation coefficient (r=.689)
between the PHQ-9 and GAD-7, and the lowest (r=.348) observed between the RISB-2 and the
DAS-SF-1.
Table 4 Pearson Product Moment Correlations Between Symptom Measures
RISB-2 GAD-7 PHQ-9 DAS-SF-1
RISB-2 Pearson X
Correlation
p (2-tailed)
N
GAD-7 Pearson .497** x
Correlation
p (2-tailed) <.001
N 85
PHQ-9 Pearson .617** .689** x
Correlation
p (2-tailed) <.001 <.001
N 85 85 85
DAS-SF-1 Pearson .348** .555** .359** x
Correlation
p (2-tailed) .001 <.001 <.001
N 85 85 85 85
Note. ** = correlation is significant at the 0.01 level (2-tailed)
Relationships Between Diagnostic Status and the RISB-2, PHQ-9, and GAD-7
Consistent with the present study’s original hypotheses, the authors intended to divide the sample
into participants with personality disorders, participants with a mental health diagnosis but no
personality disorders, and participants with no mental health diagnosis. These categories would
then be used for the purposes of conducting a series of ANOVAs. However, due to the lack of
personality disorders reported in the sample (there was only one observed case of borderline
personality disorder), this plan for data analysis was unfeasible. As such, given that it has been
suggested that personality disorders can be conceptualized as the far end of a spectrum of
possible reactions to trauma (Bounoua et al., 2022; Cloitre, 2016), the author instead divided the
sample into participants with trauma-related disorders (n=12), participants with mental health
diagnoses without trauma-related disorders (n=30), and participants with no known mental health
diagnosis (n=49). These categories were used for several of the following analyses.
Consistent with previous literature on the RISB-2, the authors ran biserial correlations between
diagnostic status (diagnosis of any kind or no diagnosis) and RISB-2 maladjustment scores,
PHQ-9 scores, and GAD-7 scores (table 5), hypothesizing that each measure would correlate
significantly with diagnostic status. The correlation between diagnostic status and the RISB-2
resulted in a correlation coefficient of .357. This correlation was significant at the .001 level.
Similarly, the PHQ-9 and GAD-7 each correlated significantly with diagnostic status, with the
correlation between status and the PHQ-9 generating a coefficient of .405, and the correlation
between diagnostic status and the GAD-7 generating a coefficient of .372. Both correlations were
significant at the .000 level.
Table 5 Biserial Correlation Between Diagnostic Status and RISB-2 Maladjustment Score, PHQ-
9, and GAD-7 Scores
PHQ-9
GAD-7
Diagnostic status
Pearson
Correlation
.405**
.372**
p (2-tailed).
.000
.000
N
85
85
Note. ** = correlation is significant at the 0.01 level
After seeing that several participants reported multiple psychiatric diagnoses over the course of
their lives (this observation was made after scoring was complete, as not to jeopardize rater
blindness), the authors ran a post-hoc PPM (table 6) between the number of mental health
diagnoses reported and RISB-2 maladjustment score, hypothesizing that the RISB-2 would
correlate positively and significantly with the number of diagnoses a participant reported. This
analysis resulted in a correlation coefficient of .462 and was significant at the .000 level.
Additionally, the authors ran PPMs between both the PHQ-9 total score and the GAD-7 total
score and the number of diagnoses reported by participants. The coefficient resulting from the
correlation of the PHQ-9 and the number of reported diagnoses was .629 and was significant at
the .000 level. The correlation between the GAD-7 and the number of reported diagnoses was
.445 and significant at the .000 level.
Table 6 PPM Correlation Between Number of Reported Diagnoses and Measures
RISB-2
Maladjustment
Score
PHQ-9
Score
GAD-7
Score
Number of Diagnoses
Reported
Pearson
Correlation
.462**
.629**
.445**
p (2-tailed).
.000
.000
.000
N
85
85
85
Note. ** = correlation is significant at the 0.01 level
Analysis of Variance
In order to test the modified hypothesis that RISB-2 mean scores would differ significantly
between participants diagnosed with mental illness but not trauma-related disorders, participants
diagnosed with trauma disorders, and participants with no previous psychiatric diagnosis, the
authors ran a series of one-way ANOVAs in order to examine how RISB-2 maladjustment scores
(table 7), PHQ-9 scores (table 8), and GAD-7 scores (table 9) differ between the three diagnostic
groups. For the RISB-2, the ANOVA was significant, with an eta-squared of .165, F(84)= 8.090,
p=.001. Moreover, a Tukey post-hoc test determined significant differences between no diagnosis
(Mean= 131.65) and both diagnosis without a trauma-related diagnosis (Mean= 141.30, p= .043,
Cohen’s d=.588) and trauma-related diagnosis (Mean= 154.5, p= .001, Cohen’s d=1.201).
However, no significant difference was found between the two groups with a diagnosis
(p= .142, Cohen’s d=.645). For the PHQ-9, the ANOVA was significant, with an eta-squared of
.207, F(84)= 10.715, p=.000. The Tukey post-hoc test determined significant differences between
no diagnosis (Mean= 3.60) and both diagnosis without a trauma-related diagnosis
(Mean= 7.20, p= .014, Cohen’s d=.72) and trauma-related diagnosis (Mean= 11.00, p= .000,
Cohen’s d=1.255). No significant difference was found between the two groups with a diagnosis
(p= .092, Cohen’s d=.609). Finally, for the GAD-7, the ANOVA was significant, with an
etasquared of .139, F(84)= 6.629, p=.002. A Tukey post-hoc test determined significant
differences between no diagnosis (Mean= 3.91) and both diagnosis without a trauma-related
diagnosis
(Mean= 8.00, p= .006, Cohen’s d=.724) and trauma-related diagnosis (Mean= 8.58, p= .026,
Cohen’s d=.977). However, as with the RISB-2 and PHQ-9, no significant difference was found
between the two groups with a diagnosis (p= .947, Cohen’s d=.104).
Table 7 RISB-2 Score by Diagnostic Category
N
Mean
Difference
Std.
Error
Lower
Bound
Upper
Bound
p
1-2
43-30
-11.765
4.813
-23.25
-.28
.043
1-3
43-12
-24.965
6.605
-40.73
-9.20
.001
2-3
30-12
-13.200
6.910
-29.69
3.29
.142
Note. CI = confidence interval for mean. 1= No diagnosis (Mean= 131.65), 2= diagnosis without
trauma-related diagnosis (Mean= 141.30), 3 = trauma related diagnosis (Mean= 154.5)
Table 8 PHQ-9 Score by Diagnostic Category
N
Mean
Difference
Std.
Error
Lower
Bound
Upper
Bound
p
1-2
43-30
-3.595
1.247
-6.57
-.62
.014
1-3
43-12
-7.395
1.712
-11.48
-3.31
.000
2-
3
30-12
-3.800
1.791
-8.08
.48
.092
Note. CI = confidence interval for mean. 1= No diagnosis (Mean= 3.60), 2= diagnosis without
trauma-related diagnosis (Mean= 7.20), 3 = trauma related diagnosis (Mean = 11.00)
% CI
95
% CI
95
1-2 43-30 -4.093 .006
2-3 30-12 -.583 1.849 -5.00 3.83 .947
Note. CI = confidence interval for mean. 1= No diagnosis (Mean= 3.91), 2= diagnosis without
trauma-related diagnosis (Mean= 8.00), 3 = trauma related diagnosis (Mean= 8.58)
Correlations Between RISB-2, PHQ-9, GAD-7, and Quality of Life
In order to test the hypothesis that the RISB-2 would have a significant negative correlation with
quality of life, the authors ran a series of PPMs in order to correlate the RISB-2, PHQ-9, and GAD-
7 with the domain scores of the WHOQOL-BREF (tables 10-13). With the exception of the GAD-
7 and the WHOQOL-BREF social domain, all measures were significant with all domains at the
.001 level or higher. The PHQ-9 had the strongest associations with the physical (r=-.681, p= .000)
and psychological (r=-.754, p= .000) domains, followed by the RISB-2 (physical r=-.621, p=.000;
psych r=-.718, p= .000) and the GAD-7 (physical r=-.459, p= .000; psych r=-.587, p= .000). The
RISB-2 had the strongest association with the social domain (r=.486, p= .000), with the PHQ-9
having the only other significant association (r=-.359, p=.001) with this domain. The GAD-7 had
the strongest association with the environmental domain (r=.558, p= .000), followed closely by the
PHQ-9 (r=-.554, p=.000), with the RISB-2 showing the weakest, though still significant,
association (r=-.458, p=.000).
Table
9
GAD
-
7
Score by Diagnostic Category
% CI
95
N
Mean
Difference
Std.
Error
Lower
Bound
Upper
Bound
p
1.288
-
7.17
-
1.02
1
-
3
43
-
12
-
4.676
1.767
-
8.89
-
.46
.026
Table 10 PPM Correlation Between WHOQOL-BREF Physical Domain and Measures
RISB-2
Maladjustment
Score
PHQ-9
Score
GAD-7
Score
WHOQOL-BREF Physical
Domain Score
Pearson
Correlation
-.621**
-
.681**
-.459**
p (2-tailed).
.000
.000
.000
N
85
85
85
Note. ** = correlation is significant at the 0.01 level
Table 11 PPM Correlation Between WHOQOL-BREF Psychological Domain and Measures
RISB-2
Maladjustment
Score
PHQ-9
Score
GAD-7
Score
WHOQOL-BREF Pearson -.718**
Psychological Domain Score Correlation
-
.754**
-.587**
p (2-tailed). .000
.000
.000
N 85
85
85
Note. ** = correlation is significant at the 0.01 level
Table 12
PPM Correlation Between WHOQOL-BREF Social Domain and Measures
RISB-2
Maladjustment
Score
PHQ-9
Score
GAD-7
Score
WHOQOL-BREF Social Pearson -.486** Domain Score
Correlation
-
.359**
-.204
p (2-tailed). .000
.001
.061
N 85
85
85
Note. ** = correlation is significant at the 0.01 level
Table 13
PPM Correlation Between WHOQOL-BREF Environmental Domain and
Measures
RISB-2
Maladjustment
Score
PHQ-9
Score
GAD-7
Score
WHOQOL-BREF
Environmental Domain
Score
Pearson
Correlation
-.458**
-
.554**
-.558**
p (2-tailed).
.000
.000
.000
N
85
85
85
Note. ** = correlation is significant at the 0.01 level
As the RISB-2 was not generally superior in its association with quality of life than the other
measures in the study which were meant to, when combined, approximate general psychological
distress, the authors felt it did not make sense to test the hypothesis that the RISB-2 maladjustment
score would account for variance in self-reported outcomes over and above a combination of self-
reported depression, self-reported anxiety, and negative cognitions. Instead, the authors ran a
series of Linear Regressions in order to determine the relative contributions of each measure to a
given domain of quality of life when all were entered together as separate variables. For the
physical domain (table 14), the RISB-2 (B=-.090, p= .000), PHQ-9 (B= -.472, p= .000), and DAS-
SF-1 (B=.197, p= .048) each contributed significantly to the variance, while the GAD-7 (B=-.036,
p= .747) did not, with the final model demonstrating a R of .744 and R2 of
.554. The physical domain’s resulting regression equation was WHOQOL-BREF Physical
Domain score= 38.461- (.090x RISB-2 maladjustment score) - (.036x GAD-7 score)- (.472x
PHQ-9 score)+ (.197xDAS-SF-1 score). For the psychological domain (table 15), the RISB-2
(B=-.070, p= .000), PHQ-9 (B=-.349, p= .000), and DAS-SF-1 (B= -.202, p= .001) all
contributed significantly, while the GAD-7 (B=.047, p=.486) did not, with the final model
demonstrating a R of .846 and R2 of .716. The resulting regression equation for the psychological
domain was WHOQOL-BREF Psychological Domain score= 35.918- (.070x RISB-2
maladjustment score) + (.047x GAD-7 score)- (.349x PHQ-9 score)- (.202xDAS-SF-1 score).
For the social domain (table 16), the RISB-2 (B=-.047, p=.001) was the only variable that
maintained significance when entered into the Multiple Regression, with the final model
demonstrating a R of .524 and R2 of .274. The resulting regression equation for the social domain
was WHOQOL-BREF Social Domain score= 18.627- (.047x RISB-2 maladjustment score) +
(.105x GAD-7 score)- (.087x PHQ-9 score)- (.087xDAS-SF-1 score). Finally, in terms of the
environmental domain (table 17), the GAD-7 (B=-.244, p=.043) was the only variable that
maintained significance when entered into the multiple regression, with the final model
demonstrating a R of .619 and R2 of .383. The resulting regression equation for the
environmental domain was WHOQOL-BREF Environmental Domain score= 40.183- (.033x
RISB-2 maladjustment score)- (.244x GAD-7 score)- (.217x PHQ-9 score)- (.063xDAS-SF-1
score).
Table 14 Multiple Regression, Physical Domain Dependent Variable
95% CI
Unstandardized
Coefficient B
Std.
Error
Standardized
Coefficient
Beta
Lower
Bound
Upper
Bound
p
(Constant)
38.461
3.146
-
32.201
44.721
.000
RISB-2
-.090
.024
-.358
-.138
-.042
.000
GAD-7
PHQ-9
DAS-SF-1
-.036
-.472
.197
.110
.109
.098
-.037
-.500
.182
-.255
-.688
.002
.184
-.256
.392
.747
.000
.048
Note. CI = confidence interval for mean.
Table 15 Multiple Regression, Psychological Domain Dependent Variable
95% CI
Unstandardized
Coefficient B
Std.
Error
Standardized
Coefficient
Beta
Lower
Bound
Upper
Bound
p
(Constant)
35.918
1.910
-
32.116
39.719
.000
RISB-2
-.070
.015
-.365
-.099
-.041
.000
GAD-7
PHQ-9
DAS-SF-1
.047
-.349
-.202
.067
.066
.060
.065
-.486
-.246
-.086
-.480
-.321
.180
-.218
-.084
.486
.000
.001
Note. CI = confidence interval
Table 16 Multiple Regression, Social Domain Dependent Variable
95% CI
Unstandardized
Coefficient B
Std.
Error
Standardized
Coefficient
Beta
Lower
Bound
Upper
Bound
p
(Constant)
18.627
1.804
-
15.037
22.217
.000
RISB-2
-.047
.014
-.418
-.075
-.020
.001
GAD-7
PHQ-9
DAS-SF-1
.105
-.087
-.087
.063
.062
.056
.245
-.206
-.178
-.021
-.211
-.199
.231
.036
.025
.101
.164
.126
Note. CI = confidence interval
Table 17 Multiple Regression, Environmental Domain Dependent Variable
95% CI
Unstandardized
Coefficient B
Std.
Error
Standardized
Coefficient
Beta
Lower
Bound
Upper
Bound
p
(Constant)
40.183
3.390
-
33.436
46.930
.000
RISB-2
-.033
.026
-.142
-.085
.019
.214
GAD-7
PHQ-9
DAS-SF-1
-.244
-.217
-.063
.119
.117
.106
-.280
-.250
-.064
-.481
-.450
-.274
-.008
.016
.147
.043
.068
.552
Note. CI = confidence interval
CHAPTER IV
DISCUSSION
The present study aimed to contribute to the literature on the RISB-2, particularly as it relates to
the RISB-2’s association with personality pathology. This was undertaken in order to more firmly
establish the RISB-2 as an instrument with clinical utility in assessing personality pathology, in
addition to general maladjustment. Though a lack of personality disorders in the obtained sample
rendered analyses related to personality disorders impossible (see limitations below), the present
study was nevertheless able to examine both the RISB-2’s psychometric properties and its
associations with multiple constructs related both to psychopathology and quality of life. These
findings will be discussed below.
Interpretation of Findings
Consistent with previous research, the RISB-2 was found to be highly reliable in terms of both
interrater reliability and internal consistency. Moreover, as expected, the RISB-2 maladjustment
score was strongly associated with both other psychological instruments measuring symptoms (i.e.
the PHQ-9, GAD-7, and DAS-SF-1) and diagnostic status of the participants in this sample; this
finding further contributes to an existing body of literature suggesting that Rotter’s (1992) concept
of maladjustment has construct validity. Finally, the RISB-2 was found to be highly associated
with all measures of quality of life administered in the present study, which is further evidence of
the maladjustment score’s construct validity.
However, despite the impressive associations found between the RISB-2 and multiple relevant
constructs, it is important to consider the implications of these associations for the RISB-2’s use
as a clinical instrument. With this in mind, it must be noted that the PHQ-9 and GAD-7 were
both at least as strong as the RISB-2 in their associations with diagnostic status, both in terms of
correlation and comparing means through ANOVA. Additionally, the PHQ-9 demonstrated an
association with the number of diagnoses reported by participants that was at least as strong as
that of the RISB-2’s maladjustment score. Taken together with the fact that the RISB-2 inherently
takes more time to administer and score than either the PHQ-9 and GAD-7, the present study’s
results seem to indicate that the RISB-2 is less efficient than these brief measures for the
purposes of general psychological assessment.
Despite this apparent relative weakness of the RISB-2 in comparison to the GAD-7 and the PHQ-
9, the present study found an unexpected, unique strength of the instrument. When examining the
associations between the RISB-2, GAD-7, and PHQ-9 and the various domains of quality of life
assessed by the WHOQOL-BREF, it was found that the RISB-2 exhibited the strongest
association with social quality of life. Moreover, when the PHQ-9, GAD-7, DAS-SF-1, and
RISB-2 were entered into a multiple regression with the social domain of the WHOQOLBREF as
the dependent variable, the RISB-2 emerged as the only measure that retained significance. As
such, it seems that the RISB-2 is uniquely sensitive to an examinee’s social functioning and, as
such, may serve a unique role within a larger psychological assessment battery. Moreover, it is
worth reiterating that, though the RISB-2 was not the only predictor that retained significance
when entered into multiple regressions with the WHOQOL-BREF physical and psychological
domains as the dependent variables, it did retain significance in both of these analyses, while the
GAD-7 did not. As such, the RISB-2 contributed unique variance in predicting three of the four
domains of quality of life measured by the WHOQOL-BREF.
Limitations of the Study
The present study is limited in several ways by its design. First, In the study by Torstrick and
colleagues (2015), diagnoses were assigned by assessments performed as part of data collection,
rather than relying on patient report. Utilizing patient report as a way of ascertaining diagnostic
status introduces multiple sources of potential error, including reluctance to disclose diagnoses
on the part of the participant, ignorance of a diagnosis on the part of the participant, and
misdiagnosis on the part of participant’s past/present clinicians. This latter possibility is
particularly salient when studying personality disorder diagnoses, as borderline personality
disorder is known to be frequently misdiagnosed as bipolar disorder (Fruzzetti, 2017). As such,
asking for participant report of previous diagnoses is inherently less reliable than diagnosis by a
standardized interview or assessment battery.
The second limitation of the present study is the sampling method used. As discussed above, the
investigator was unsuccessful in recruiting a sample of participants diagnosed with personality
disorders large enough to conduct the planned analyses. In addition to this, the sample obtained
was not representative of the US population in terms of either gender or ethnicity. As such, though
the present study’s design and obtained sample did allow for several interesting findings, the
original objective of attempting to replicate Torstrick and colleagues’ (2015) findings through a
different method was not possible. Finally, the final sample obtained (91) was far smaller than the
minimum sample required for adequate power when running a one-way ANOVA (270). As such,
while several significant effects emerged from the ANOVAs and multiple regressions run during
the present study, as indicated by observed effect sizes, it is possible that other significant effects
(e.g. differences between scores in patients with and without trauma-related diagnoses) were not
found because the tests conducted were underpowered.
Strengths of the Present Study
Despite its limitations, the present study has a number of strengths. First, the present study was
able to compare RISB-2 scores, both in terms of overall maladjustment and at the level of
individual items, between three independent raters. These analyses allow the present study to add
to the extant literature demonstrating the RISB-2’s inter-rater reliability, but also set the
groundwork for looking at the psychometric properties of the individual items of the RISB-2 and,
perhaps, establishing subscales or an abbreviated version.
In addition to allowing for multiple types of inter-rater reliability analyses, the present study
looked at the RISB-2’s association, not just with a dichotomous clinical status, but also with
different categories of diagnosis (i.e. no diagnosis, diagnosis other than trauma, and
traumarelated diagnosis), as well as multiple indicators of both psychological symptoms and
quality of life. As such, the present study contributes to a more nuanced understanding of the
RISB-2 maladjustment score and its relationship to both other, more specific clinical constructs,
and outcomes.
Finally, and paradoxically, the authors would argue that, though the present study’s small sample
renders certain analyses underpowered (e.g. ANOVA), this limitation also makes the many
significant findings that did emerge all the more compelling. As underpowered samples reduce
the likelihood of finding a significant effect, the fact that most of the analyses conducted,
including underpowered analyses, found significant effects speaks to, in the author’s opinion,
how robust the relationships between the constructs assessed in the present study actually are. It
is the author’s sincere hope that more nuanced relationships (e.g. whether there is a significant
difference in symptoms severity in non-trauma diagnoses and trauma-related diagnoses) can be
more fully examined in the future, when a sample better suited towards such analyses is
available.
Future Directions and Implications
As discussed above, Torstrick and colleagues’ (2015) finding that the RISB-2 maladjustment
score was uniquely associated with personality pathology could not be replicated in the present
study due to problems with the present study’s obtained sample. As such, it falls to future
investigators to examine the RISB-2 maladjustment score’s demonstrated association with
personality pathology in order to better establish whether or not the RISB-2 may serve as a
reliable alternative method of assessing personality pathology. However, despite this setback, it
is possible that the RISB-2’s demonstrated unique association with social quality of life in the
present study may explain its strong association with personality pathology in the Torstrick and
colleagues’ (2015) study. As disruptions in interpersonal functioning have been established as a
core feature of personality pathology (Wilson et al., 2017), it is possible that the RISB-2’s
apparent sensitivity to social quality of life may also cause it to more strongly associate with
disorders characterized by disruptions in social quality of life. Future research should thus
examine whether or not any association between the RISB-2 and personality pathology is
mediated by associations between the RISB-2 and social quality of life and/or interpersonal
functioning.
Additionally, the present study found significant variability in the weighted kappa coefficients
associated with each of the RISB-2’s items. It is possible that these coefficients, rather than
being the result of simple chance, reflect items and/or scoring criteria that are simply better at
pulling for responses that are more easily scored with consistency between raters. As such, and
especially in light of the RISB-2’s relative inefficiency as a diagnostic instrument compared to
the brief measures also administered in this study, it may be worthwhile to develop a brief
version of the RISB-2 using only the items with the demonstrated highest weighted kappa
coefficients in the present study. It is possible that, if such a brief version were developed, its
construction exclusively from items with higher kappa coefficients could result in a measure
that is both more efficient to administer and more psychometrically sound.
The majority of past research on the RISB-2’s psychometric properties, to include the present
study, has focused on the general population, often by recruiting college students (e.g. Church &
Crandall, 1955; Lah, 1989). Though certain studies have focused on more specific populations,
such as adolescents (Weis et al., 2008) and African-Americans (Logan & Waehler, 2001), further
research examining the psychometric properties of the RISB-2 in diverse populations, as well as
other, special populations (e.g. forensic and residential populations) would help establish the
RISB-2’s overall reliability, validity, and utility in both clinical and research contexts. Though it
is always important to establish the reliability and validity of a measure across populations, the
authors would argue it is especially important in the case of the RISB-2; unlike many objective
measures, which give limited options for response to any given item, the open format of the RISB-
2 allows respondents to complete sentences in a way that is influenced by, not just their own
experiences, but also cultural norms. As norms of expression and ways of meaning-making vary
across cultures (HHSOMH, n.d.), the authors would argue that it is incumbent upon future
researchers to establish that the RISB-2’s instructions and scoring criteria, which past research
(e.g. Lah, 1989; Torstrick et al., 2015) and the present study have demonstrated result in scores
which strongly correlate with clinically-significant constructs for those raised in mainstream
American culture, function equivalently for those who hail from different cultures, with different
norms of expression. As an example, the rule that one point is added to any response longer than
ten words may serve a different function when score responses by individuals hailing from more
expressive cultures. Thus, future research should continue to study the RISB-2 in diverse and
special populations.
Finally, the authors would be remiss if they did not discuss the larger implications of the RISB2’s
contribution of unique variance to linear regressions predicting multiple domains of quality of life,
as well as the finding that it was the only measure that retained significance in the linear regression
predicting social quality of life. As social quality of life is strongly associated with many areas of
functioning (Umberson & Montez, 2010), it is worth discussing why the RISB-2’s maladjustment
score may have been uniquely associated with it. While the other measures included in the study
were designed to assess specific constructs of clinical importance (i.e. depressive symptoms,
anxiety symptoms, and negative cognitions), the RISB-2’s scoring criteria are designed to assess
overall maladjustment. As maladjustment, like general psychological distress, is not a single
symptom, nor a construct specific to any given class of diagnoses, it seems likely that it serves as
a higher-order construct that captures other clinically important constructs that are either
inadequately assessed by the measures included in the present study (e.g. shame), or altogether
excluded from those measures (e.g. anger), all of which likely impact multiple domains of quality
of life, including the social domain. As such, it is possible that the RISB-2 contributes unique
variance to these outcome variables by accounting for these constructs, which more specialized
measures, by virtue of their design, do not assess. The author suspects that an individual’s social
quality of life is particularly sensitive to high levels of these constructs in combination, as
neuroticism, a higher-order personality construct marked by multiple clinically relevant symptoms,
is known to impact an individual’s relationships (Widiger & Oltmanns, 2017).
The finding that the RISB-2’s maladjustment score contributed unique variance to linear
regressions predicting multiple domains of quality of life is also significant because it contributes
to a growing body of literature (e.g. Exner, 1974; Torstrick et al., 2015) indicating that, when
interpreted in an empirically-validated manner, projective measures can add unique, quantitative
data to a psychological assessment battery. This is especially important in light of criticism of
projective techniques which hinges its argument on limited empirical evidence supporting their
validity and reliability (Lilienfeld et al., 2000). It is the author’s hope that the present study’s results
will support a renewed interest in the development of both novel methods of assessment of
psychological constructs and new, empirically-supported methods of interpretation for existing
projective methods. Not only would such developments have the potential to help researchers and
clinicians account for constructs not easily addressed by more conventional, objective measures,
the validation of projective and other novel assessment techniques as reliable sources of unique,
quantitative data would also be consistent with the spirit of multifactored assessment, which, in
itself, has been put forth as a means of reducing bias in psychological assessment and increasing
confidence in assessment results (Gresham, 1983).
Conclusion
In conclusion, despite the limitations of the current study’s methodology and sample discussed
above, several important findings emerged. First, the current study is the latest in a long line of
research supporting the RISB-2’s interrater reliability, internal consistency, and strong
association with both general psychopathology and quality of life. Second, the present study
found that, for the purposes of assessing general psychopathology, the RISB-2, while effective,
is likely less efficient than brief screeners such as the PHQ-9. Despite this, the RISB-2 seems to
have a unique association with social quality of life that remains significant even when other
measures, such as the PHQ-9, are entered into the regression.
In addition to these findings, multiple significant directions of future research emerged from the
present study’s results. Given the RISB-2’s previously established association with personality
disorders (Torstrick et al., 2015), and the known association of personality disorders with
impairment in interpersonal functioning (Wilson et al., 2017), it is possible that the present
study’s finding that the RISB-2 is uniquely associated with social quality of life may explain the
association between the RISB-2 and personality pathology. Additionally, the observed variability
in the weighted kappa coefficients of the RISB-2’s items suggests that, were a brief version of the
RISB-2 derived from only the most psychometrically-sound items, the result may be a measure
that is both more efficient to administer and more psychometrically sound. Finally, as always,
there is room for additional research in more diverse and specialized populations.
Students also viewed