1 / 8100%
Assignment: Comparison of IRT, CFA, and CTT
Tests and Measurements
Professor Richard Thompson
1
Comparison of IRT, CFA, and CTT
Clinical psychologists are advised to assess clinical and statistical significance when
assessing change in individual patients. Individual change assessment can be conducted using
either the methodologies of classical test theory (CTT), Confirmatory Factor Analysis or item
response theory (IRT). The following section will compare and contrast the three measurement
approaches (IRT, CFA, and CTT). The advantages and disadvantage of both method are
discussed (Courville, 2004).
Classical Test Theory
CTT is a psychometric theory concerned with the difficulty of items and the capacity of
test-takers, and it is used to predict the results of psychological testing. The goal of classical test
theory is to increase the reliability of psychological tests by clarifying them. CTT is a model that
is linked to true score theory (Kim et al., 2015). CTT is based on the true score model, which is
based on an examinee's aggregate score in a test and hence does not provide a consideration of
examinees' replies to any specific item, offering no foundation to forecast how a given examinee
will perform on a specific test question. A theory presupposes that each person has a genuine
score that would be attained if there were no measurement mistakes. The predicted number
correct score over an unlimited number of separate test administrations is defined as a person's
accurate score (Meredith, 1993).
Item Response Theory (IRT)
Item response theory (IRT), also known as latent trait theory, is a psychometric theory
developed to understand better how people react to specific items on psychological and
educational examinations. The fundamental theory is based on mathematical formulas with
parameters that must be calculated using sophisticated statistical methods (Courville, 2004).
These variables have to do with the features of particular items and the characteristics of
individual respondents. IRT is referred to as a latent trait since unique features cannot be directly
observed; instead, they must be inferred using certain assumptions about the response process
that aid in estimating these parameters (Fan, 1998).
2
Any model relating the probability of an examinee's response to a test item to an
underlying ability is referred to as item response theory. It is known as latent trait theory, and it
aims to predict observations based on latent variables. Compared to conventional test theory,
item response theory has had a considerable impact on psychology by providing more exact
techniques for measuring test features (Fan, 1998). In addition, IRT has had a significant impact
on psychology by enabling the development of several instruments that would have been
difficult to develop without it. The discovery of IRT has substantially enhanced psychometric
applications such as computerized adaptive testing, detecting item bias, equating tests, and
identifying aberrant individuals. Computerized adaptive testing, in particular, deserves more
attention (Meredith, 1993).
IRT item parameters are not dependent on the sample used to create the parameters, and
they are supposed to be invariant (within a linear transformation) between divergent groups
within a research population and across populations (Courville, 2004). According to Cappelleri
et al. 2014). The two essential postulates of IRT area. A set of traits, latent traits, or abilities can
be used to predict or explain an examinee's performance on a test item; b. A monotonically
increasing function called an item characteristic function or item characteristic curve can
describe the relationship between examinees' item performance and the set of traits underlying
item performance (ICC). According to this function, the chance of a correct response to an item
grows as the amount of the attribute increases.
Confirmatory Factor Analysis (CFA)
CFA and IRT are sophisticated scales quality evaluation methods. Each adds to the
argument for the validity of scale scores and their application in research and practice by
providing specific information regarding item and scale quality. CFA is frequently used in social
work research to study and establish the psychometric properties of scores derived from
collections of items assessing common constructs (Kim et al., 2015). CFA is a latent variable
modeling technique in which scores on related questionnaire questions are assumed to reflect an
everyday complicated, unobservable reality (construct). CFA researchers frequently start with
exploratory factor analysis (EFA) to learn more about a set of items' factor structure, such as the
number of factors and the pattern of interactions between items and factors. Then, with
theoretical or empirical evidence, they employ CFA to assess those models and potential
3
competing factor structures. CFA allows for more factor and error structure testing than EFA and
more comprehensive model quality assessments (Kim et al., 2015).
The majority of scales used in social work research and practice have items with
dichotomous or ordinal response options that do not meet normality assumptions, so researchers
conducting CFAs can choose between estimation options based on the measurement level and
distributional characteristics of observed variables an important feature given that the majority of
scales used in social work research and practice have items with dichotomous or ordinal response
options that do not meet normality assumptions (Meredith, 1993). Factor analysis reveals
information on item quality, construct validity and dependability. The overarching purpose is to
determine whether and to what extent items on a scale reflect an underlying hypothetical
construct or constructs, referred to as factors an analytical method with high sensitivity for
detecting problematic items and determining the number of components (Meredith, 1993).
A Comparison of CTT, IRT and CFA
In comparison to CTT, IRT emphasizes assumptions, findings, and error characteristics
more strongly and its model-based nature; it also has several advantages over equivalent CTT
discoveries, which lead to critical practical outcomes. CTT scores are simple to compute and
elaborate, but IRT scores require complex computation processes (Courville, 2004). Still, they
have the advantage of comparing the difficulty of an item and a person's ability on a test
meaningfully. Furthermore, because the parameters of IRT models are not sample or test
dependent, IRT can give greater flexibility in circumstances involving a variety of samples or
test formats, and its findings are essential for computerized adaptive testing (Cappelleri et al.,
2014).
When the model fit is present, IRT produces person parameter invariance (test scores are
not dependent on the particular choice of test items), and test information functions provide the
amount of information or "measurement precision" captured by the test on the scale measuring
the construct of interest, which has several advantages over CTT (Courville, 2004).
Classical test theory (CTT), the most common psychometric theory taught in
undergraduate and graduate programs, is supplemented and contrasted by item response theory.
Classical test theory differs from IRT in several areas, described later in this article. On the other
4
hand, IRT can be compared to an electron microscope for item analysis, while CTT is more akin
to a standard optical microscope (Meredith, 1993). Both strategies are beneficial in their own
right. Like the electron microscope, IRT provides robust measurement analysis; IRT is useful
when specialized, precise analysis is required. However, CTT can be as valuable as IRT when
the research questions are imprecise and generic. The optical microscope is sometimes favored
over the electron microscope in medical research. CTT may also be preferable in some
circumstances (Fan, 1998).
An explanation of the advantages and disadvantages of IRT, CFA, and CTT.
According to Cappelleri et al., (2014), the benefits of applying classical test models to
measurement problems include:
1. It can be done with smaller samples of test-takers as representative of a
population.
2. It uses relatively simple mathematical procedures and conceptually
straightforward model parameter estimations.
3. It is considered a weak model because its assumptions are easily met by
traditional testing procedures.
Even though CTT is useful in test development, it still has certain drawbacks. The main
elements used in this theory, item difficulty and item discrimination, are sample dependent,
meaning that the derived conclusions are heavily reliant on samples for interpretation. CTT can
also be defined as test dependent or test-based. The complexity of the exams impacts the test
scores obtained and the actual score model CTT on which it is built, and it leaves no room for
examinees' reactions to specific items (Courville, 2004). As a result, it is impossible to forecast
how an examinee will fare in a particular test item. In addition to the shortcomings listed above,
there is no distinction in CTT between test-takers and test characteristics, and they must be
interpreted in the context of one another. Furthermore, the concept of reliability is problematic
because it is described as the correlation between test scores on parallel forms of a test the
difficulty is that different people have different ideas about what parallel tests (Costa et al. 2017).
Person and item statistics are recognized to be test and sample dependent in the CTT
paradigm. With IRT, this is not the case. As a result, when estimating human and item
parameters, the IRT framework is thought to be theoretically superior to the CTT framework. In
5
earlier simulation experiments, IRT models have been employed as both generating and fitting
models. As a result, the results favoring the IRT framework could be explained because IRT is
the data-generation framework. On the other hand, CFA allows researchers to specify which
variables load on which factors ahead of time. It enables researchers to determine which
variables are linked and which are not. CFA is frequently used to demonstrate measurement
validity. CFA can be used to calculate estimates of dependability (coefficient omega) (Fan,
1998).
6
References
Cappelleri, J.C., Lundy, J.J., & Hays, R. D. (2014). Overview of classical test theory and item
response theory for quantitative assessment of items in developing patient-reported
outcome measures. Doi: 10.1016/j.clinthera.2014.04.006
Costa, J. A., Marôco, J., & Pinto‐Gouveia, J. (2017). Validation of the psychometric properties
of cognitive fusion questionnaire. A study of the factorial validity and factorial invariance
of the measure among osteoarticular disease, diabetes mellitus, obesity, depressive
disorder, and general populations. Clinical psychology & psychotherapy, 24(5), 1121-
1129.
Courville, T. G. (2004). An empirical comparison of item response theory and classical test
theory item/person statistics (Doctoral dissertation, Texas AandM University).
Fan, X. (1998). Item response theory and classical test theory: An empirical comparison of their
item/person statistics. Educational and psychological measurement, 58(3), 357-381. Doi:
10.1177/0013164498058003001
Kim, E. S., Yoon, M., Wen, Y., Luo, W., & Kwok, O. M. (2015). Within-level group factorial
invariance with multilevel data: multilevel factor mixture and multilevel mimic models.
Structural Equation Modeling: A Multidisciplinary Journal, 22(4), 603-616. Doi:
10.1080/10705511.2014.938217
7
Meredith, W. (1993). Measurement invariance, factor analysis, and factorial invariance.
Psychometrika, 58(4), 525-543. Doi: 10.1007/bf02294825
8
Students also viewed