1
Chapter 1: Introduction to the Study
Introduction
The volume of data in health care is exploding exponentially. This trend is
supported by digitization of health care data from different sources along with advances
in imaging technology and molecular diagnostic methods (Murdoch & Detsky, 2013;
Raghupathi & Raghupathi, 2014). The overwhelming volume of data in health care
contributes to chaos and uncertainty that can lead to diagnostic errors with devastating
consequences such as chronic pain, disability, or death (Gupta et al., 2017; Saber-Tehrani
et al., 2013). The radiologist is rapidly becoming one of the most important curators and
gatekeepers of big data. Despite this important role, the complexity and volume of
available data during the interpretive stage of radiology workflow outpaces the individual
radiologist’s capacity to make fully informed and timely decisions (Thrall et al., 2016).
The rate of diagnostic error associated with the interpretation of abnormal radiology
studies has been reported as high as 30% with retrospective review of diagnostic imaging
studies yielding even higher error rates (Berlin, 2007; Brady et al., 2012; Donald &
Barnard, 2012).
The causes of errors during the interpretive stage of radiology workflow include
exposure to complex data from new imaging technology, limited access to relevant non-
imaging data, increasing workload, physiological fatigue, and human bias (Brady et al.,
2012; Donald & Barnard, 2012; C. S. Lee, Nagy, Weaver, & Newman-Toker, 2012).
Without adequate technological assistance, the human interpretive process within
radiology workflow will become progressively more inaccurate, inefficient, and untimely
2
(Croskerry 2013; Manrai et al., 2014; Obermeyer & Emanuel, 2016; Ragupathi &
Ragupathi, 2014). The solution requires the adoption and implementation of new data
management and decision support solutions such as artificial intelligence (AI).
AI offers a variety of approaches to assist in problem solving and prediction
(Pannu, 2015). The process encompasses different computational methods such as
machine learning, deep learning, and cognitive computing. Using any of these AI
solutions during radiology workflow might improve diagnostic accuracy and timing. AI
is capable of revealing biological variability, heterogeneity of pathology, and
comorbidities, all of which support more personalized and precise diagnosis (Gillies,
Kinahan, & Hricka, 2016; Limkin et al., 2017; Pannu, 2015; Yip & Aerts, 2017). AI will
fundamentally alter the field of radiology by facilitating more predictive, preventive, and
participatory health care (Ghasemi et al., 2016; Hillman & Goldsmith, 2011; Jha &
Topol, 2016). The potential impact of AI in radiology must be better understood to be
adopted, implemented, and supported. Research will contribute to this process and
contribute to positive social change by introducing decision support solutions capable of
improving diagnostic accuracy and reducing health care costs.
My primary goal for this qualitative exploratory research study was to identify the
potential impact of AI on spine imaging interpretation and diagnosis. Chapter 1 includes
the fundamental elements of the study. These include the background, problem statement,
and research questions. The chapter also includes an introduction to relevant terms and
definitions to enhance the reader’s understanding of the research topic. I present the
theoretical perspectives and conceptual framework used for the design of the study. I also
3
discuss relevant assumptions, delimitations, and limitations in the study. Chapter 1 ends
with a discussion of the social significance of the research study along with a transitional
summary to the next chapter.
Background
The success of diagnostic imaging is highly dependent on technology and
protocol. Technological advances in radiology result in a perpetual cycle of discovery,
disruption, opportunity, and adaptation. Current innovations are rapidly transforming
radiology from a qualitative to a quantitative science with a growing capacity to obtain
molecular and physiologic measures (Jha & Topol, 2016; Quer et al., 2017). This
evolutionary process has led to unprecedented growth in the volume of multidimensional
data which has and will continue to have an impact on radiology workflow and the type
of decisions which have to be made. Exposure to an increasing quantity of complex data
increases the likelihood of diagnostic errors (Obermeyer & Ezekierl, 2016; Ragupathi &
Ragupathi, 2014). The growing demand to read imaging studies faster also influences
interpretive accuracy. For example, the average emergency radiologist may read up to
50,000 to 100,000 images per day, resulting in an average of 2 to 3 seconds spent on each
image (Syeda-Mahmood, 2015). Sokolovskaya et al. (2015) reported that radiologists
pushed to read complex diagnostic imaging studies at a faster speed made over twice as
many interpretive errors as those who read the same studies at a normal rate. Emerging
health care system productivity goals often drive the radiologist to interpret imaging
studies in a shorter period of time. The time constraints combined with the growing
complexity of studies adds to the burden of decision making.
4
The error rate in radiology has been a concern for decades. Pioneering research
performed by Garland (1949) over 60 years ago revealed an estimated 30% level of
inaccuracy regarding the interpretation of abnormal chest radiographs. This landmark
study prompted follow-up research. The error rate of radiologists on the interpretation of
abnormal imaging studies between 1949 and 1992 was estimated at 30%, consistent with
Garland’s work decades prior (Berlin, 1994; Renfrew et al., 1992). Additional studies
have revealed error rates for the interpretation of different types of abnormal imaging
studies in the range of 15–35% (Elmore, et al., 1994; Lehr et al., 1976; Janjul et al.,
1998). Interpretive errors in the current health care environment are potentially more
influential than they have been in the past because of widely distributed results within
electronic medical record (EMR) systems. Interpretive errors embedded within EMR can
lead to exposure at multiple levels along the health care chain, setting the stage for
additional errors in subsequent testing, clinical application, and judgment.
Paralleling the burden of big data in radiology there is a growing demand for
more precise imaging interpretation and concise reporting to support personalized care.
To successfully meet this demand, new decision support solutions have to be embedded
within the interpretive stage of radiology workflow. Success will require the integration
of human and machine attributes, a form of collective intelligence (CI). Radiomics, a
rapidly emerging AI supported process, is capable of performing high-throughput
automated analysis of imaging data through a series of sequential steps such as pathology
detection, feature extraction, feature characterization, and analysis (Lambin et al., 2017;
Lee et al., 2017; Limkin et al., 2017). Pathology features include physiological,
5
molecular, morphological, statistical, and textual attributes (Aerts, 2017). Successful use
of radiomics has the potential to help detect, characterize, and classify disease. It is
capable of revealing new signatures of pathology not available through visual
interpretation and contributing to probability-based decision support.
IBM developed a cognitive computing system in 2007 referred to as Watson,
which introduced to the public in 2011 as a potential AI solution for decision support in
the presence of complex data and uncertainty (Hoyt, Snider, Thompson, & Mantravadi,
2016). Watson was designed to augment the role of the radiologist during the interpretive
stage of radiology workflow and during the differential diagnostic process (Brink et al.,
2017; Jha & Topol, 2016; Kharat & Singhal, 2017; Jiang et al., 2017; Ranschaert, 2016).
AI solutions such as IBM Watson have the potential to improve diagnostic accuracy and
timing through the use of automated pattern detection, evidence analysis, and probability-
weighted scoring (Doyle-Lindrud, 2015; Ferrucci, 2012; Kohn et al., 2014). Some AI
systems are capable of learning with exposure to structured and unstructured data. This
includes exposure to data from radiomic methods, prior imaging studies, electronic health
records, disease databases, computational disease models, and peer-reviewed literature.
Interpretive and diagnostic errors in radiology occur due to many conditions
including technological limitations, system inefficiencies, and human error (Berlin, 2013;
Thammasitboon, Thammasitboon, & Singhal, 2013). Human perceptual errors occur
more frequently than cognitive errors (Berlin, 2013; Bruno, Walker, & Abujudeh, 2015;
Donald & Barnard, 2012). The causes of perceptual errors include reader fatigue,
distractions, and the presence of hidden or subtle disease patterns overlooked with visual
6
analysis methods (Bruno et al., 2015; C. S. Lee et al., 2012). Potential solutions for
reducing errors include redesign of the interpretive stage of radiology workflow to
include AI solutions, better integration of radiology-pathology data, 3D viewing options,
and structured reporting (C. S. Lee et al., 2012). The use of AI has already demonstrated
that it can improve the accuracy and quality of some decisions in radiology and in other
areas of health care (Brink et al., 2017; Jiang et al., 2017). The potential role of AI in
radiology is relatively new, subsequently; research and development will be influenced
by needs analysis and meaningful use cases. Research on the use of AI in spine care, and
more specifically spine imaging is extremely limited, despite its potential for having a
favorable impact on one of the most common causes of pain and disability.
The volume of personalized data acquired with advanced diagnostic imaging is
growing as the result of more detailed tissue interrogation, expanded fields of view,
thinner slice thickness, and the use of multimodality and multiparametric methods (Aerts,
2017; Hillman & Goldsmith, 2011; Jha & Topol, 2016; Quer et al., 2017). This trend will
influence all areas of radiology including spine imaging. The increased volume of
accessible personalized data contributes to the complexity of decision making
surrounding individual patient care and to inconsistencies and variability in radiology
reports.
According to Lundstrom, Gilmore, and Ros (2017), the interpretive stage of
radiology workflow may serve as the primary hub for the convergence and analysis of
imaging data, pathology data, and genomic data. This would empower the radiologist as a
curator of integrated patient data and as a clinical consultant. New decision support
7
solutions accessible during image interpretation will have a significant impact on the
timing and precision of the diagnosis process and subsequently decisions at the point-of
care. Research has demonstrated that AI use offers “higher than clinician-grade accuracy”
in numerous fields such as in dermatology, ophthalmology, and radiology (Syeda-
Mahmood, 2018, p. 573). AI has the potential to support more accurate, efficient, and
timely decision making in radiology. Its role in spine imaging needs to be addressed. AI
applications could help overcome errors associated with human bias while democratizing
expert decision support. I designed this study to identify how AI could improve
personalized spine care and help direct further investigation on what is required to
achieve this goal.
Problem Statement
The unprecedented growth of data in radiology is driven by new imaging
technology and methods capable of performing whole body surveys and evaluating
tissues at multiple length scales such as anatomic, physiologic, cellular, and molecular
levels (Aerts, 2017; Hillman & Goldsmith, 2011, Murdoch & Detsky, 2013; Raghupathi
& Raghupathi, 2014). Radiology is rapidly evolving into a dynamic, whole systems
diagnostic discipline capable of interrogating multiple dimensions of pathology in vivo
(D. Y. Lee & Li, 2009). The growing complexity and volume of imaging and non-
imaging data available during the interpretive stage of radiology workflow is exceeding
the individual radiologist’s ability to make fully informed decisions (Thrall et al., 2016).
Radiologists are finding it increasingly difficult to determine what imaging findings are
clinically significant and meaningful. Big data in radiology has become a burden which
8
can be overcome to create new opportunities for patient care. The primary problem
addressed by this research was defining how using AI during the interpretive stage of
spine imaging could augment the role of the radiologist and improve diagnostic precision.
As previously stated, the rate of error on abnormal radiology studies has been
reported as high as 30%; retrospective review of diagnostic imaging studies yielding even
higher error rates (Berlin, 2007, 2013; Brady et al., 2012, Donald & Barnard, 2012). An
overwhelming amount of data contributes to missed patterns and relationships resulting
in diagnostic errors with deleterious consequences (Gupta et al., 2017; Saber-Tehrani et
al., 2013). The causes of error in radiology include the burden of complex data,
increasing workloads; limited time, physiological fatigue, and human bias (Brady et al.,
2012; Donald & Barnard, 2012; C. S. Lee et al., 2012). The high level of inconsistency
and variability in radiology workflow is influenced by what data are acquired, how the
data are analyzed, and how the results are reported (Aerts et al., 2013). Several authors
have suggested that without adequate technological assistance, the human interpretive
process within radiology workflow will become progressively more inaccurate,
inefficient, and untimely (Croskerry 2013; Manrai et al., 2014; Obermeyer & Emanuel,
2016; Ragupathi & Ragupathi, 2014). Few studies have addressed how to reduce
interpretive errors in radiology (Fitzgerald, 2005; Lee, 2017).
I was unable to identify any peer-reviewed articles in which the authors
specifically address how to reduce interpretive errors associated with big data in spine
imaging. I was also unable to locate peer-reviewed publications in which the authors
address the use of AI methods such as radiomics combined with text analysis to reduce
9
interpretive errors in spine imaging. The spine-related complaints represent one of the
most common causes of chronic pain and disability in the United States (Hurwitz et al.,
2018; Loney & Stratford, 1999; Ricca et al., 2006). The American College of Radiology
(2016) recently acknowledged the need for further research to help identify how AI could
be used to access meaningful data, enhance the interpretive phase of radiology workflow,
improve diagnostic accuracy, and reduce errors.
The use of decision support solutions such as AI during the interpretive stage of
spine imaging workflow could lead to more accurate detection, characterization, and
diagnosis of pathology. Despite heightened awareness of the need for new decision
support solutions and knowledge of the potential benefits of AI use in radiology adequate
research on the potential impact on spine imaging has not been performed. Qualitative
exploratory research is required to lay the foundation for further research surrounding the
use of AI and to identify how benefits can be achieved.
Purpose of the Study
There is a considerable amount of meaningful data and insight embedded within
medical images, undetectable through routine visual analysis, and therefore not
considered in the diagnostic process (Gillies et al., 2016; V. S. Lee, 2017). Such data
embedded within non-imaging records could be made available during the interpretive
stage of radiology workflow for correlation with imaging findings. Missed data or
misinterpretation of data may lead to an error in diagnosis. Overlooked data and data
patterns can also lead to missed opportunities in care. The primary purpose of this
research study was to explore the potential of AI to favorably impact diagnostic accuracy
10
and precision during the interpretive stage of spine imaging workflow. The approach
includes the potential role of natural language processing and radiomic methods to assist
in the detection, characterization, and monitoring of pathology.
Research Questions
The primary research question for this study was: What are the opinions of
experts regarding the potential use and impact of AI intelligence during the interpretive
stage of spine imaging workflow?
The subquestions for this study follow:
1. How could the use of AI-supported methods (auto detection, segmentation,
radiomics, natural language processing) during the interpretive stage of spine
imaging influence the differential diagnostic process?
2. How could the use of AI-supported methods during the interpretive stage of
spine imaging influence disease classification and staging?
3. What AI solutions could be used to create interpretive priority in spine
imaging?
4. What will the future of spine imaging interpretation workflow look like?
5. Could AI-supported solutions such as radiomics be used to interrogate spinal
tissue in vivo and eventually lead to a virtual (digital) biopsy?
6. What are some potential advantages of an in vivo virtual “digital” biopsy over
a traditional needle biopsy in spine care?
7. How could the use of AI-supported augmented reality (AR) or virtual reality
(VR) enhance the evaluation of pathology in spine imaging?
11
8. What are some potentially “meaningful use” applications of AI in spine
imaging?
9. Which construct of the technology acceptance model (TAM) will likely have
a greater impact on AI adoption during the interpretive stage of spine imaging:
perceived benefits or perceived ease-of-use?
10. Which characteristics of innovations proposed by the diffusion of innovation
theory (DOI) will likely have the greatest impact on AI adoption during the
interpretive stage of spine imaging: complexity, compatibility and
interoperability, observed effects or trialability?
Theoretical Perspectives and Conceptual Framework
Adopting and using new technology is influenced by many factors, including
awareness of technology’s potential, ease of use, and interoperability with existing
technology and workflow (Khan & Woosley, 2011; Oye, Iahad, & Abrahim, 2012). The
potential impact of new technology must be studied in-depth using a flexible and iterative
approach. A rigid theoretical framework would not allow for adequate exploration of this
topic of study. Exploratory research offers a descriptive and inductive approach which
supports discovery, and helps reveal potential advantages and disadvantages surrounding
the use of new technology (Creswell, 2009; Patton, 1990; Yin, 2014). Effective
qualitative research design requires a conceptual framework to help align theoretical
perspectives, the research purpose, and research strategies (Creswell, 2007; Marshall &
Rossman, 2011; Patton, 1990). Conceptual boundaries can be developed through the
integration of ontological, epistemological, methodological, and structural perspectives.
12
This research study was inductive, beginning with a set of assumptions and
research questions leading to data acquisition and analysis. This study would have been
restricted by the rigid application of an a priori theory or hypothesis. I did not use a
deductive approach. A qualitative exploratory approach supports the investigation of
possibilities within the framework of different contexts and realities (Creswell, 2013).
Because of the potential depth and breadth of this topic of study, it would have been
unfeasible to perform the exploratory study through a single theoretical lens or
worldview.
AI represents a complex technology associated with numerous processes;
therefore, its potential impact must be evaluated from a pluralistic perspective. No single
theory can be used to fully explore the potential role and impact of AI during the
interpretive stage of spine imaging workflow. I used an interpretivist-constructivist
epistemological view to address the topic of study. The interpretivist perspective helped
operationalize how and why questions surrounding the use of AI. A constructivist
perspective supported the development of a proposed model for using AI solutions during
the interpretive stage of spine imaging workflow. Consistent with the work of Creswell
(2013) and Patton (1990), I used theoretical perspectives to help develop the conceptual
framework of the study and to develop research questions which aligned with the
research methodology and purpose.
The diffusion of innovations theory (DOI) proposed by Rogers (1995) represents
one of the most widely accepted theories for exploring how the attributes of innovative
technology influence its adoption, use, and distribution (Kahn & Woosley, 2011).
13
Furthermore, DOI offers insight about the stages of adoption and adopter categories
(Dillon & Morris, 1996; Rogers, 2003). The characteristics of innovations that influence
adoption proposed by Rogers (2003) include relative advantage, interoperability with
existing systems, the potential for technology evolution, reinvention, complexity, and
trialability. Additional factors that influence adoption include awareness of benefits
(knowledge), knowledge of advantages and disadvantages (decision), method of use
(implementation), and clinical utility (confirmation). DOI includes five adopter
categories: innovators, early adopters, early majority, late majority, and laggards (Rogers,
2003). The literature review summarized in Chapter 2 includes information regarding the
use of theoretical constructs and related conceptual perspectives introduced in this
treatise. I designed this study to obtain consensus-based perspectives from published
white papers and insights from purposively selected innovators and early adopters in
focus group sessions.
The technology acceptance model (TAM) proposed by Venkatesh (2008) is used
to address behavior and perceptions that influence willingness to use new technology.
The two primary constructs of TAM are perceived usefulness (PU) and perceived ease-
of-use (PEOU). The synthesis of DOI and TAM constructs have been previously used to
frame exploratory research of computer and information technology (Carter & Belanger,
2005; Y. Lee, Hsieh, & Hsu, 2011; Legris, Ingram, & Colerette, 2003). DOI represents an
effective theory for addressing variables surrounding technology adoption, whereas TAM
offers insight about post-adoption use and support (Hameed & Arachchilage, 2016).
Select constructs from DOI were combined with constructs from TAM to guide the
14
development of research questions to be used in focus group sessions and to help develop
an initial a priori list of thematic coding categories for data analysis.
A well-developed conceptual framework provides boundaries that help direct the
acquisition, management, and analysis of data during research. A conceptual framework
can also be used to assess the impact of new technology on professional behavior and
workflow (Bogdan & Bilken, 1982; Patton, 1999). The conceptual framework of this
study supported a contextual inductive and iterative process for exploring the potential
impact of AI on the interpretive stage of spine imaging workflow. I used the conceptual
framework for this study to choose data analysis strategies and to help guide the literature
search summarized in Chapter 2. I focused on literature that revealed AI technology
characteristics and the goals of early adopters of AI in radiology. Research questions
were developed to address some of the perspectives offered by DOI and TAM constructs.
The chosen method of inquiry in this study required answers to how and why questions
surrounding the potential benefits and challenges associated with AI use during the
interpretive stage of spine imaging workflow.
Nature of the Study
I performed this qualitative exploratory case study with an in-depth, inductive
approach to acquire and analyze data to address the potential impact of AI on spine
imaging interpretation and diagnosis. Qualitative exploratory case study research offers
an effective method for acquiring contextual knowledge surrounding the use of new
technology (Creswell, 2013; Ponelis, 2015; Yin, 2014). The approach is also effective at
revealing themes surrounding the potential impact and role of new technology (Bradley,
15
Curry, & Devers, 2007; Patton, 1999). Qualitative exploratory research studies have been
successfully used in radiology (Sanberg et al., 2012). The approach in this case offered
the level of holistic investigation required to reveal previously unknown variables
surrounding the use of AI during the interpretive stage of radiology workflow.
A well-defined unit of study helps direct qualitative research. It also helps frame
and inform research strategies. The boundaries of a qualitative case study may consist of
a relationship, situation, process or a culture (Miles, Huberman, & Saldana, 2014). The
primary unit of study for this research is a process defined as the interpretive stage of
spine imaging workflow. This includes the workstation, display technology, embedded
AI solutions, available data, and the role of the radiologist. I chose the unit of study to
identify how and why AI solutions should be used during the interpretation of advanced
spine imaging. I also designed the study to evaluate the potential impact of AI on the
current state of the unit of analysis.
Qualitative exploratory research is capable of revealing contextual complexities
not adequately addressed by explanatory or quantitative research methods (Ponelis, 2015;
Yin, 1984). More specifically, qualitative research can be used to reveal themes and
patterns, not revealed with restrictive quantitative methods (Dubin, 1969; Patton, 2002;
Ryan & Bernard; 2003). This study required the acquisition of data from multiple sources
including expert documents, reflective journaling, and focus group sessions. The focus
groups comprised radiologists and AI experts who met study inclusion and exclusion
criteria. I used numerous methods to improve the trustworthiness of the study.
16
Definitions of Relevant Terms
I have listed the operational definitions of selected terms and phrases to help the
reader understand the research topic, findings, and conclusions.
Algorithm: An algorithm represents a mathematical formula or program, used as a
set of instructions for computational analysis (Kakhani et al., 2017).
Artificial intelligence (AI): Artificial intelligence refers to the use of technology to
identify patterns, solve problems, and make predictions using biological-like approaches
(Jiang et al., 2017).
Augmented intelligence (AI): Augmented intelligence refers to the use of
technology to enhance human capability and to provide decision support (Liew, 2018).
Big data: Big data refers to a large volume of complex data that is difficult to
analyze utilizing traditional methods (Farooki, Almeida, & Saltz, 2016).
Collective intelligence (CI): Collective intelligence refers to the combined use of
AI solutions and human intelligence to identify a pattern, solve a problem, perform a
task, or make a prediction.
Computer-aided detection (CADe): Computer-aided detection refers to the use of
a computer system to identify a pattern within structured data, unstructured data, or on a
set of diagnostic images (Hillman & Goldsmith, 2011).
Computer-aided diagnosis (CADx): Computer-aided diagnosis refers to the use of
a computer system and computational methods of analysis to provide a list of differential
diagnostic possibilities or a specific diagnosis (Hillman & Goldsmith, 2011).
17
Data science: Data science represents an interdisciplinary field that uses
computational methods to analyze data, to reveal patterns, study processes or to make
predictions (Aerts, 2017).
Deep learning (DL): Deep learning represents a subset of machine learning
capable of becoming more accurate and capable with exposure to data through the use of
multiple hidden processing layers (Tang et al., 2018).
Differential diagnostic process: A differential diagnostic process is a series of
interrelated steps that use probability-based logic or reasoning to differentiate a disease or
disorder from others that may have a similar presentation.
Ground truth: Ground truth refers to data assumed or proven to be true (Choy et
al., 2018)
Innovation: Innovation is defined as a new or improved method of performing
tasks or solving problems (George et al., 2005).
In vivo: In vivo refers to the evaluation of biological elements, processes or
systems within a living organism (Lambin et al., 2012; Rizzo et al., 2018).
Machine learning (ML): Machine learning represents a subset of AI that uses a
computer system and a set of algorithms to reveal patterns in data with or without
handcrafted explicit instructions (Giger, 2018).
Natural language processing (NLP): Natural language processing refers to the use
of automated computational methods to analyze and interpret spoken or written language
(J. Y. Chen et al., 2017).
18
Neural networks: The phrase neural networks refer to a sophisticated network of
layered interconnected data paths and nodes, which process signals in a manner similar to
neurons in a biological system (Tang et al., 2018).
Precision medicine (PM): Precision medicine refers to an objective approach to
the diagnosis and delivery of health care which takes into account the heterogeneity of
disease along with an individual’s unique biology, disease risk, and variable response to
treatment (Mesko, 2017).
Radiomics: Radiomics refers to the science of high-throughput data analysis used
for identifying, extracting, characterizing, quantifying, and classifying nonvisible
elements of pathology within medical imaging data sets (Napel & Giger, 2015).
Segmentation: The definition of the term segmentation in radiology refers to the
use of manual, semiautomated or automated methods to identify and outline a region of
interest or an abnormality on images (Gillies et al., 2016).
Systems medicine: Systems medicine also referred to as network medicine refers
to the science associated with the relationships and interconnections between molecular
structures, cells tissues, and organs (McCue & McCue, 2017).
Virtual biopsy: The use of the phrase virtual biopsy in radiology refers to the use
of various computational and radiomic methods to reveal, extract, and quantify in vivo
characteristics of pathology derived from imaging data guided by a defined region of
interest (Echgaray et al., 2016; Lambin et al., 2012; Thrall, 2016).
19
Assumptions
Assumptions that are true or likely to be true represent beliefs or perspectives.
Assumptions influence research designs, methods, and conclusions. It is therefore
important to acknowledge assumptions that may influence this research process. I had to
consider the potential influence of numerous assumptions in this research study. I
assumed that by choosing consensus-based white papers (research documents) published
by reputable radiology organizations that I would acquire expert insights regarding the
potential role of AI in radiology now and in the future. I assumed that four to six
purposively selected participants (experts) placed into one of two homogenous focus
group sessions would be enough to reach topic saturation. In addition, I assumed that the
research participants would have sufficient knowledge and experience to answer focus
group questions and productively contribute to focus group discussions. I also assumed
that the research participants would adequately represent the levels of expertise held by
others in their relevant fields.
I assumed based on the inclusion and exclusion criteria of this study that the
participants would have a comfortable understanding of the research topic and goals. My
use of a purposive sampling method increased the likelihood that each research
participant had the experience and knowledge required to address the research topic
questions. Each of the participants works in the health care field, thus improving the
likelihood that they are aware of the importance of an honest and ethical approach to
research.
20
I assumed that participating radiologists had limited knowledge of the potential
role of AI in spine care and that AI experts had limited knowledge of the potential impact
of AI on spine imaging interpretation. The presumed knowledge and limitations of AI
specialists and radiologists opened the door for creative discussions and discovery in each
session. The level of dedication required of individuals within each group to reach expert
status infers a high level of passion and dedication to the subject; therefore, I assumed
that each of the participants could contribute valuable insight and opinions during the
focus group sessions. I assumed that, given the preparatory steps taken prior to data
acquisition, each of the research participants fully understood their rights and
responsibilities in the research project. The research assumptions helped guide me during
data collection from focus group sessions and from other sources.
Scope and Delimitations
The topic of delimitations refers to anticipated constraints that may arise during
the research design or while conducting research. Several delimitations arose in this case.
Research participants voluntarily participated in the focus group sessions; therefore, it
was likely that they were interested in the role of AI in radiology and spine care. The
defined unit of study and boundaries surrounding this study served as relative constraints
to data acquisition and interpretation. Participant inclusion and exclusion criteria also
represented delimitations. The moderator guide and questions developed for the focus
group sessions served as adaptable delimiters in the study.
The primary purpose of this research study was to explore the potential impacts of
AI use during the interpretive stage of spine imaging workflow. I did not investigate the
21
impact of AI on interprofessional relationships within the radiology department or its
influence on economic issues surrounding the use of AI in spine care. I also did not
address how AI might specifically impact radiologists of different backgrounds, training,
and skill levels. I did not design the research project to address the administrative
challenges or costs associated with developing, implementing, or supporting AI solutions.
Limitations
Research limitations refer to potential barriers or boundaries that investigators
cannot control. Limitations in this study included the available knowledge and expertise
from a small study population. Furthermore, the survey sample used in this study was
purposive and small, resulting in a data analysis process which is more descriptive than
inferential. The focus group participants’ levels of expertise influenced research study
conclusions. Limitations associated with lack of familiarity with the research topic were
relatively low given the highly qualified nature of the experts who participated. The
biases associated with the use of a single examiner posed a potential limitation. I reduced
this risk with the use of validation methods such as member checking, within and
between group analyses, reflective journaling, and triangulation of qualitative data
acquired from numerous sources.
This research study consisted of a small purposively chosen population of experts;
thereby, limiting the ability to generalize results. Determining whether sample size was
adequate was influenced by many factors such as the outcome of the research validation
methods. The type of methods used such as triangulation of data, member checking, and
peer review must be carefully considered when determining sample size (Patton, 2002).
22
The use of two research participant categories and two focus group sessions limited the
scope of this portion of the study. The potential role of AI in spine imaging is broad;
therefore, a single case study cannot address all relevant dimensions of the topic. In
addition, there was limited time and resources that restricted the duration of this study,
rendering it difficult to study the perceived benefits of AI use during the interpretive
stage of spine imaging over time.
Significance of the Study
Medical errors represent the third most common cause of death in the United
States (Makary & Daniel, 2016) and one of the most costly and avoidable health care
expenses (Saber-Tehrani et al., 2013). Diagnostic errors represent one of the most
common reasons for malpractice claims (Andel et al., 2012). Diagnostic imaging
represents a widely used method for detecting, characterizing, and monitoring disease.
Subsequently, many of the important decisions made in health care arise from diagnostic
imaging studies (D. Y. Lee & Li, 2009). The critical role and relevance of diagnostic
imaging in health care is increasing. This study addresses a gap in the research literature
associated with the potential impacts of AI on spine imaging workflow and the diagnostic
process. Using AI during the interpretive stage of spine imaging workflow could help
reduce human errors and favorably contribute to a more accurate diagnostic and reporting
process, leading to improved patient care. Research in other health care and radiologic
specialties has demonstrated that using AI can contribute more consistent, quantitative,
and actionable information to the final report, thus influencing point of care decisions
(Augimeri et al., 2016; Boone et al., 2015; Mohebian et al., 2017). AI solutions combined
23
with effective data governance and data management in spine imaging could result in
fewer interpretive errors and improved diagnostic precision.
Traditionally, radiology has relied upon the visual perception and interpretation of
the radiologist (Pinto & Brunese, 2010). Perceptual factors surrounding visual
interpretation represent a common cause of error. Factors that contribute to perceptual
errors include reader fatigue, distractions, the presence of hidden or subtle disease
patterns, and various form of human bias (Bruno et al., 2015; C. S. Lee et al., 2012). A
specialized application of AI, referred to as radiomics, has been successfully used by
radiologists to reveal the heterogeneity and in vivo features of pathology not detectable
by the radiologist during the visual inspection of a study (Aerts, 2017; et al., 2016; H. W.
Wu et al., 2012; Yip & Aerts, 2016). Radiomic methods improve the ability to classify
disease and stratify treatment approaches, all leading to more precise and personalized
patient care. In addition, AI solutions have the potential to improve the diagnostic
imaging process through the ability to detect subvisual pathology and to correlate data
from other imaging and non-imaging sources (Dreyer & Geis, 2017). This includes data
from electronic health records, published literature, genetic databases, as well as from
prognostic and predictive disease models. Digitization of in vivo and in vitro pathology
supports image sharing and access to remote data analysis and expert consultations.
Sharing of digital information also allows for consensus-based decision support.
Advances in medical technology have historically contributed to improved
methods of detecting and treating disease (Clinton, 2000). Integrated co-evolution of AI
supported methods such as radiomics will contribute to the discovery of new molecular
24
signatures and biomarkers of disease. This will lead to new standards for evidence-based
care in all fields including spine care. Successful integration of human and machine
intelligence during the interpretive stage of spine imaging workflow has the potential to
help control unsustainable health care costs, and support personalized care (Hillman &
Goldsmith, 2011; Jha & Topol, 2016; Kressel, 2017; Lee, 2017). In addition, AI is
capable of democratizing decision support that will aid underserved facilities,
underserved regions, and inexperienced radiologists.
New levels of expectations and knowledge surrounding the use of AI will
influence standards of care, which will ultimately benefit individuals, families, and
society. Greater use of AI-supported diagnostic methods such as radiomics will expand
classifications of disease and improve personalized care. The results of this research
study contribute to positive social change by identifying how AI use during the
interpretive stage of radiology workflow could favorably shape the future spine care.
Successful use of AI during spine imaging interpretation could result in early detection,
early intervention, better treatment outcome, and reduced direct, as well as indirect costs
associated with chronic pain and disability. Moreover, AI could facilitate collaboration
between spine care providers of all disciplines by exposing the fundamental basis for
disease, democratizing expert decision support, and by providing evidence-based
measures of treatment outcome. Widespread use of AI decision support will help
overcome some of the barriers to collaboration associated with human ignorance and
biases. I designed this study to reveal potential applications of AI in spine imaging, as
well as in other fields and specialties. This study introduces the concept of the digital
25
(virtual) biopsy that could have a profound influence on further investigation of this
concept and its application in all fields of health care.
Summary
Given the variability in people and their spine disorders, spine care delivery needs
to be more precise and personalized. Better integration of machine and human
capabilities during the interpretive stage of radiology workflow will contribute to earlier
detection of pathology, better characterization of pathology, and a more precise and
timely diagnosis. AI and related solutions used in other fields such as oncology and
cardiology can be adapted and used during the interpretive stage of spine imaging. The
primary goal of diagnostic imaging in spine care is to provide insight and knowledge that
can be used by a provider to deliver care for the right patient, for the right condition, at
the right time.
Imaging represents one of the most commonly performed and revealing elements
of the diagnostic workup in spine care. The combined burden associated with the growing
volume of imaging and non-imaging data available during the interpretive stage of
radiology workflow is increasing the complexity of decision making for the radiologist.
The missed opportunity and error rate associated with the interpretation of abnormal
imaging studies is too high. The growing complexity of data acquired through more
advanced imaging technology will only contribute to more complex decisions and higher
incidence of error. Applying AI solutions during the interpretive stage of spine imaging
workflow has the potential to detect early-stage pathology, offer decision support, and
facilitate personalized care.
26
The primary purpose of this research study was to identify whether the use of AI
solutions during the interpretive stage of spine imaging could reduce the risk for
diagnostic error and improve the precision of the final diagnosis. The potential AI
solutions introduced in this chapter and investigated in the next chapter include natural
language processing, radiomics, disease modeling, and computational analysis of
structured and unstructured data. The success of the interpretive stage of radiology
workflow is dependent on integrated solutions. It is becoming increasingly important to
use quantitative measures in diagnostic imaging. The integration of various decision
support solutions will likely change the landscape of radiology and spine care.
An accurate and precise diagnosis is dependent on heightened awareness of
possibilities and the ability to investigate the possibilities by analyzing available data.
Too often knowledge of differential diagnostic possibilities is limited by human
experience and limited access to technologies required to identify subtle or hidden
patterns within medical records or in the imaging data. This leads to errors in diagnosis
and clinical judgment. AI technologies such as natural language processing and radiomics
offer potential solutions. In Chapter 2, I address how current applications and
contributions of AI in radiology could be adapted or further developed for use during the
interpretive stage of spine imaging.
In this chapter, I introduced the research problem, research purpose, and research
significance, to help guide my literature search and review addressed in Chapter 2. The
subsequent chapter acknowledges current and potential applications of AI in radiology. In
Chapter 2, I synthesized the research purpose and goals from this chapter with the results
27
of an extensive literature search. This helped inform the research design and
methodology introduced in Chapter 3.
28
Chapter 2: Literature Review
Introduction
Exposure to large volumes of complex data in radiology increases the level of
uncertainty and complex decision making. This challenge combined with human
limitations and bias increases the risk for errors and oversights. The field of radiology has
always been dependent on the use of technology and has long served as a pillar in health
care for the acquisition, analysis, and management of complex data (Thrall et al., 2016).
Imaging has become one of the most important sources of data and diagnostic
information in health care (Aerts et al., 2013; Gillies et al., 2016: Kim et al., 2013;
Kinahan, & Hricak, 2016). The radiologist is poised to become a gatekeeper of big data
and a disease consultant (di Piro et al., 2017). This role will be augmented with the
convergence and integration of imaging, laboratory, genetic, and pathology data at the
radiology workstation.
The field of radiology is rapidly transforming from a predominantly qualitative to
a quantitative science supporting the rising demand for a more personalized diagnoses
(Jha & Topol, 2016; Quer et al., 2017). In addition to the demand for a more precise
diagnostic process, the radiologist is burdened with a growing quantity of complex data
arising from advances in whole body imaging, molecular diagnostic methods, and the
integration of multimodality and multiparametric imaging approaches (Murdoch &
Detsky, 2013; Raghupathi & Raghupathi, 2014). New forms of decision support are
required during radiology workflow to reduce the potential for interpretive and diagnostic
errors. The results of an extensive literature search revealed how AI solutions have begun
29
to have a significant impact on disease detection, characterization, and surveillance in
health care.
In this chapter, I discuss the literature search strategy that I used to address the
topic of study. I also address the fundamental concepts and research findings that I
identified in the literature that support qualitative exploration of the potential role of AI
use during the interpretative stage of spine imaging. Spine and spine-related research
involving the use of computational decision support such as AI were limited in number
and scope. The majority of published research and reviews addressed the role of AI
solutions in other fields such as oncology and neuroimaging. The literature review
identified prior methods of inquiry and research used to perform the studies. It also
revealed gaps in the literature regarding the use of AI during the interpretive stage of
non-spinal and spinal imaging. Scholarly publications provided the rationale for refining
the research problem, as well as developing the research questions and methodology for
this study. Given the limited number of publications referencing the use of AI in spine
imaging, the literature search was expanded to include research on the use of AI during
the interpretive stage of imaging other regions of the body. I chose and organized the
topics of this chapter to present relevant findings from the literature. The literature search
established a scholarly foundation for the research design and methodology covered in
Chapter 3.
Literature Search Strategy
I performed an extensive literature search to explore existing perspectives and
facts surrounding the topic of study. My helped to identify relevant theories, conceptual
30
frameworks, and research strategies that were applied in this study. The literature search
revealed current expectations and standards associated with the use of AI in radiology.
My literature search concentrated on articles that were published within 5 years of
the anticipated completion date of this work. My search was primarily limited to peer-
reviewed scholarly publications. I retrieved journal articles from the following research
databases; PubMed, EBSCO, ProQuest, and Google Scholar. I used key terms and
phrases along with different techniques to perform and refine the search process.
Common search terms and phrases included artificial intelligence, deep learning,
machine learning, radiomics, natural language processing, interpretive radiology, spine,
spine imaging, virtual biopsy, in vivo, voxel-wise detection, computer-aided detection,
computer-aided diagnosis, and biometric analysis. I used independent and combined
search terms to improve the investigative process.
The literature search was highly iterative to achieve adequate saturation of the
topic. I used single and combined Boolean operators, truncation, and wild card symbols
in the search process. I reviewed publication abstracts to ascertain article relevance. All
relevant research and review articles were printed, read, and saved. I evaluated the
bibliographic reference lists of seminal articles for additional works. My literature search
identified research studies, as well as consensus-based documents and position papers
published by respected institutions and organizations. During the literature review
process I was able to identify key experts on different research topics. I performed
author-based searches to look for relevant material they may have authored or contributed
31
to. Research librarians provided access to published articles I was unable to find through
more traditional methods. Thus, the literature search on the topic was exhaustive.
Applied Theoretical and Conceptual Framework
Successful adoption and use of new technology are influenced by knowledge of
its utility and its role within existing workflow (Khan & Woosley, 2011; Oye, Iahad, &
Abrahim, 2012). This knowledge is acquired through reading published works, listening
to colleagues or through hands-on experience. A rigid theoretical framework cannot be
successfully used to study the role of AI in radiology due to its numerous components,
rapid evolution and the complexity of its impact on the spectrum of workflow. As
previously stated, AI represents a complex technology associated with numerous
processes; therefore, its potential impact must be evaluated from a pluralistic perspective.
Exploration using a conceptual framework developed from the synthesis of constructs
and perspectives from different theories and models supported an adaptive and iterative
approach to addressing the potential role of AI in this study. This approach is consistent
with the work of Creswell (2009), Patton (1990), and Yin (2014).
The synthesis of DOI and TAM constructs have been used in numerous research
studies to explore the potential applications and benefits of new technology (Carter &
Belanger, 2005; Y. Lee et al, 2011; Legris, Ingram, & Colerette, 2003). In this case, the
synthesis of constructs from DOI and TAM led to the use of practical descriptive
categories such as perceived usefulness, perceived ease-of-use, clinical utility, diagnostic
accuracy, interoperability, and workflow compatibility. Each of these perspectives is
applied to the role of AI in radiology and is addressed through a variety of headings in
32
this chapter. In summary, my use of DOI and TAM perspectives supported the
development of a guiding conceptual framework used to help direct the literature review,
create focus group questions, and analyze acquired data. These conceptual perspectives
were also used to help frame the results of the literature search presented in this chapter.
The DOI proposed by Rogers (1995) acknowledges various categories of
technology adopters. These categories include the innovator, early adopter, early
majority, late majority, and laggards. Scholarly investigation of publications by
innovators and early adopters of AI use in non-spine related specialty fields of radiology
provided some of the insights required to address the topic of this research study.
Technology-Induced Transitions
The field of radiology has and will continue to be shaped by technology
development and its evolution (Hillman & Goldsmith, 2011). Advances in imaging
technology influence the timing and type of decisions, which have to be made during the
interpretive stage of radiology workflow. Innovations also restructure the realm of
expectations surrounding the analysis and flow of imaging data. Historically,
transformative technological advances in radiology have included radiology information
systems (RIS), picture archiving communication systems (PACs), natural language
processing (NLP), and more recently AI solutions (Hillman & Goldsmith, 2011). Each of
these technologies has improved some aspect of radiology workflow including the
detection, characterization, and reporting of disease (Sardanelli, 2017). The rapid
evolution of technology and data management in radiology is supported by the
33
miniaturization of materials, increased computer processing power, improved data
storage, enhanced connectivity, as well as greater access to disease models and registries.
Technological advances that offer decision support such as AI has and will
continue to influence every stage of radiology workflow. Computational decision support
influences the roles and responsibilities of radiologists. Widespread adoption and use of
AI-based decision support will require efficient integration with existing workflow and
legacy systems (Bauer, 2017; Dreyer & Geis, 2017). Learning how AI might improve
spine imaging interpretation initially requires exploratory research.
Computer systems and related AI solutions are often associated with numerous
components and processes, each of which contributes to new insights and technological
advances. This process is often referred to as co-evolution. Arthur (2009) introduced the
supporting premise that “existing technologies beget further technologies” (p. 21). The
impact of new technologies and their spinoffs is often underestimated, whereas the speed
of implementation is often overestimated. Knowledge and appreciation for AI technology
co-evolution is required to predict its impact on interpretive workflow.
Technological advances in radiology lead to perpetual cycles of discovery,
disruption, and adaptation. This unstable process introduces threats to established
protocols and standards (Lai, 2017). Each time new technology emerges, its potential
impact on the delivery of care has to be evaluated, and its clinical utility determined
(Kressel, 2017). In addition, the benefits and risks associated with an innovation have to
be disclosed and discussed. The attributes and influences of new technology such as AI
34
must be considered when evaluating its potential impact during the interpretive stage of
spine imaging workflow.
Potential AI solutions in radiology cannot be studied in isolation. They must be
studied in the context of other technologies and processes embedded into workflow. This
includes assessing its potential impact on diagnostic accuracy and efficiency. Zhang et al.
(2004) introduced a hierarchical system of human and technological relationships that
can contribute to or amplify medical error. The hierarchy includes the “individual,
individual-technology interaction, distributed systems, organizational structures,
institutional functions and overarching national regulations” (Zhang et al., 2004, p. 194).
The impact of new technology and decision support can have both favorable and
unfavorable outcomes. For this reason, the potential advantages and disadvantages of AI
use during the interpretive stage of radiology workflow must be considered within the
context of hierarchical professional and technological relationships.
Big Data Attributes and Related Burdens in Radiology
Diagnostic images are more than pictures. They contain massive quantities of
minable and potentially meaningful data, often not considered during the interpretive
process. Medical imaging is estimated to represent as much as 90% of all stored medical
data, contained within billions of images (Lambin et al., 2017). Imaging data is present in
many forms including symbols, words, images, and binary digits. Individual datum and
aggregations of data have unique characteristics that influence how it is acquired,
analyzed, and applied. The primary attributes of datum and data include value,
variability, veracity, velocity, and volume (Gandomi & Haiderm, 2015; Nendaz &
35
Perrier, 2012). The primary sources of data available during the interpretive stage of
radiology workflow are the study requisition, electronic medical records, diagnostic
imaging, and disease registries (Brown, 2014). In the near future there will be greater
access to relevant pathomic and genomic data at the radiology workstation. Human
interpretation of diagnostic images is often limited to visual analysis and qualitative
descriptive reporting (Thrall et al., 2016). Additional methods are required to analyze
textual and nonvisible data in radiology.
Data are commonly classified as structured or unstructured. Structured data have
well-defined form and context, whereas unstructured data tend to have inconsistent form,
rendering it more difficult to analyze. The most common type of unstructured data is
language, whether printed or verbal (Raghupathi & Ragupathi, 2014). The majority of
personalized data in health care are unstructured, in the form of text, represented in
medical records and reports.
Emerging computational methods are being used in radiology to help transform
data and related patterns to meaningful information or knowledge that has clinical utility.
Determining data relevancy in radiology requires knowledge of its validity and clinical
utility. The majority of acquired data in radiology are considered noise and are not
relevant to clinical care. This perspective is influenced by current limitations in pattern
detection and limited human capacity to analyze nonvisual imaging data (Kohn et al.,
2014). The large volume of data acquired during current imaging studies only represents
a fraction of what will be acquired, mined, and transformed into action in the near future
36
(Kohn et al., 2014). What are considered irrelevant and meaningless data now may
represent actionable data in the future.
A considerable amount of data and insight are embedded within medical images,
but remain undetectable through traditional visual analysis (Gillies et al., 2016; Lee et al.,
2017). As a result, radiologists face immense challenges during the interpretive stage of
radiology workflow. Despite these current challenges advances in diagnostic imaging
continues to create progressively larger and more complex data sets. Without adequate
technological assistance human interpretation of complex imaging data will become
progressively more inefficient, inaccurate, and untimely (Crosskery 2013; Manrai et al.,
2014; Murdoch & Detsky, 2013; Obermeyer & Ezekierl, 2016; Ragupathi & Ragupathi,
2014; Weber et al., 2017). As previously stated, new solutions are required. In the future,
whoever has access to the best data and best interpretive solutions will likely provide the
best care.
Decision-Making Processes
Clinical decision making in health care including radiology is often challenging,
and associated with complex nonlinear information (Hussain & Oestreicher, 2017).
Available methods to simplify the process include the use of published guidelines,
professional collaboration, crowd-sourcing, and computational decision support (Nendaz
& Perrier, 2012; Phua & Tan, 2013). Many health care providers, including radiologists,
are often confronted with decisions they are unprepared or unqualified to make.
Radiologists like other health care providers tend to look for what they know, identify
what they are familiar with, and render decisions based on experience. Variables such as
37
the quantity and quality of data, experience, professional knowledge, incentives, and
access to technological support influence human decisions (Gandomi & Haider; Weber &
El-Kareh, 2017; H. W. Wu et al., 2012). Accurate and timely decisions require awareness
of a well-defined problem, knowledge of options and alternatives, and the capacity to
evaluate the problem (Croskerry, 2013; Kohn et al., 2014). In addition to addressing
diagnostic variables, a radiologist’s decisions must meet the standard of care and be
consistent with a patient’s needs, values, and expectations. The diagnosis offered by a
radiologist after interpreting a set of images should be delivered in a manner that provides
adequate decision support for the referring health care provider at the point of care.
Decision Support Solutions
Decision support refers to a process or technique used to help determine the right
course or courses of action. The primary purpose of clinical decision support (CDS) in
radiology is to avoid errors, improve diagnostic accuracy, and improve the quality of care
delivered to the patient. A successful decision support system requires various attributes
such as availability, ease-of-use, accuracy, and consistency (Kahn, 1994). Decisions
made during image interpretation, the reporting process, and at the-point-of care
(Raghupathi & Raghupathi, 2014). The rate of discovery and knowledge creation
generally outpaces the individual health care provider’s ability to keep up to date and
make fully informed decisions (Kohn et al., 2014; Thrall et al., 2016). This phenomenon
applies to data intensive specialties such as radiology; therefore, the radiologist must
remain aware of decision support options.
38
Numerous solutions have been proposed to address complex decision making in
radiology. One of the more recent solutions is the computerized decision support system
(CDSS). CDSS solutions are placed into one of two categories; knowledge-based systems
or non-knowledge-based systems (Stivaros et al., 2010). Knowledge-based systems
contain programmed rules, an inference engine, and a well-defined communication
mechanism. Non-knowledge-based systems often use machine-learning techniques or
algorithms that learn from the ground up through the exposure, assimilation, and analysis
of available data. Effective use of decision support in radiology will help deliver the right
information, in the right format, at the right level of workflow to the right person. In
summary, access to decision support can augment the role of the radiologist.
Heuristics
Heuristics refers to the application of rules or processes to simplify decision
making. Radiologists often rely on heuristic methods such as mental shortcuts to
minimize delay, reduce task complexity, and to simplify decisions (Itri & Patel, 2018;
(Tversky & Kahneman, 1974). Heuristic methods are used to improve the accuracy, as
well as the efficiency of the interpretive process in radiology. Inappropriate use of
heuristics or the use of inaccurate heuristic methods can result in errors that adversely
influence patient care.
The three principal categories of heuristics are anchoring, availability, and
representative (Tversky & Kahneman, 1974). Anchoring heuristics refers to the limitation
of further considerations due to perceived truth. In contrast, availability heuristics refers
to the assignment of value based upon an individual’s memory or recall. This form of
39
heuristics can result in errors secondary to limited or selective memory of prior events or
outcomes. Representative heuristics is characterized by the use of categories to simplify
data, information, and knowledge. This approach may be used to simplify a process
leading to oversight and conjunction fallacy.
Dual Processes: Intuition and Analytics
The dual process theory of decision making consists of intuitive processing (type
I) and analytic (type II) approaches (Croskerry, Petrie, Reily, & Tait, 2014). An intuitive
response is characterized by a low cognitive demand and rapid application, whereas an
analytic approach is characterized by a high cognitive demand, a slow process, and
greater reliance on working memory. There are risks associated with isolated application
of one decision-making process over another (Phua & Tan, 2013). Croskerry (2014)
acknowledged the importance of discriminant use of both methods during complex
decision making. The combined use of problem-solving methods increases the likelihood
of a good decision and a good outcome. This perspective applies to human and machine-
based approaches.
Sources of Interpretive and Diagnostic Errors in Radiology
Diagnostic errors are often underreported and underappreciated due to a lack of
standards in defining, recognizing, and acknowledging their presence (C. S. Lee et al.,
2012). It is difficult to estimate the impact of radiology errors due to the reasons
mentioned and the limited capacity to measure their short and long-term impact on
overall health (C. S. Lee et al., 2012). Errors can occur anywhere along the path of
radiology workflow (Huassian & Oestreicher, 2017; Kassier & Kopeman, 1989). Errors
40
that occur during the interpretive stage of radiology workflow are likely to have the
greatest impact on the final diagnosis and report.
Health care providers too often make point of care decisions with irrelevant,
incomplete or incorrect information (Kelly & Hamm, 2013). In addition, many physicians
have limited experience with uncommon diseases or complex presentations associated
with coexistent pathology (Manrai et al., 2014). This condition can influence the
radiologist during image interpretation and can influence the referring physician who
reads the report at the point of care. According to Latts (2016). it is impossible for a
single health care provider to stay up-to-date in their field and to remain aware of all of
the relevant data and variables associated with any particular disease process or state.
This position supports the need for better decision support along the path of care.
Types of Errors
A medical error represents a deviation from a consensus opinion or a standard of
care. Errors may occur secondary to missed presentations, oversights, or mistakes of
judgment, all of which could lead to failure to implement a process or a plan of action
(Andel et al., 2012; Makary & Daniel, 2016). One of the most common forms of error in
radiology is diagnostic error, surfacing as a missed diagnosis, wrong diagnosis, or an
untimely diagnosis. Diagnostic errors in radiology have generally been classified as
perceptual or cognitive (Berlin, 1996, 2013). Several researchers have acknowledged
perceptual error as the most common cause of diagnostic errors in radiology (Berlin,
2013; Bruno et al., 2015). In support of this premise, a large radiographic research study
revealed that 80% of diagnostic errors were perceptual and 20% were interpretive
41
(Donald & Barnard, 2012). Causes of nondiagnostic errors in radiology include failure to
recommend or perform an indicated test or test protocol and failure to address clinical
concerns or reported patient presentations on imaging test requisitions.
Y. W. Kim and Mansfield (2014) proposed one of the most widely accepted
classification of errors in health care. The approach represented an expansion of
classifications previously proposed by Renfrew (1992), a few decades earlier. According
to Kim and Mansfield (2014) common causes of error in radiology include faulty
reasoning, lack of knowledge, satisfaction of search bias, miscommunication, and an
acquired inaccurate or incomplete history. Cognitive based-errors include limited skills,
stress, bias, faulty heuristics, memory loss, and inattention (Zhang et al., 2004). All types
of errors are amplified in the presence of large volumes of complex data. In addition to
the reasons given, diagnostic errors may occur as the result of technological limitations,
restricted access to data, and system failure (Berlin, 2013; Thammasitboon et al., 2013).
Errors in radiology may also occur secondary to cognitive bias rather than the result of a
perceptual error or lack of knowledge (Hussain & Oestreicher, 2017; Nendaz & Perrier,
2012). The cause of interpretive error in radiology is often the result of more than one
factor.
Expanding domain knowledge combined with human variables such as limited
time and preoccupation with the care of complex patients, contributes to a high incidence
of diagnostic errors within any specialty field (Weber & El-Kareh, 2017). Radiologists
are more likely to make wrong decisions in the presence of signs, symptoms or
conditions, which they have, limited experience with or knowledge of (Weber & El-
42
Kareh, 2017). Overwhelming data combined with limited knowledge and competing
environmental pressures contribute to chaos and confusion, which can lead to oversights
and diagnostic errors (Gupta et al., 2017; Saber-Tehrani et al., 2013). Access to
computational decision support such AI increases the potential for problem solving in the
presence of human limitations, as well as cultural and environmental pressures.
Human Limitations
Human decisions made in the presence of high degrees of complexity and
uncertainty increase the likelihood of error (Stivaros et al., 2010). The ability to process
information is limited by physiological mental capacity, a phenomena often referred to as
the cognitive threshold. Radiologists interpret images using a variety of cognitive
methods such as visual detection, pattern recognition, memory, and reasoning. A
radiologist’s cognitive performance is influenced by personal attributes, physiology, and
professional skills. It is also influenced by environmental factors such as ambient noise,
workload intensity, and workflow distractions. Stress influences interpretive accuracy.
Stress may be associated with uneven work distribution, poor reimbursement, limited
time, increasing liability pressures, along with heightened awareness of the complexity
and heterogeneity of pathology.
Humans are flawed in their capacity to process large volumes of multidimensional
or deeply nested data (Hatt et al., 2017; Wolf et al., 2015). The human brain is also
limited in its capacity to perform highly scalable functions that involve voluminous or
unrecognized confounding variables (Wolf et al., 2015; H. W. Wu et al., 2012). Humans
are limited in their capacity to perform accurate complex data analysis in a relatively
43
short period of time. Factors that contribute to these limitations include physiological
fatigue, cognitive bias, limited knowledge, and distractibility (H. W. Wu et al., 2012). A
conscious or unconscious response to self-limitations may result in the use of heuristic
methods that introduce bias to a decision-making process.
Increased complexity in a work environment increases the likelihood of human
error compounded by deficiencies of the system (Institute of Medicine, 2000). For this
reason, health care specialists responsible for analyzing large quantities of complex data
such as radiologists are often exposed to high cognitive demands and subsequently high
rates of diagnostic error (Crosskery 2013; Nendaz & Perrier, 2012; Obermeyer &
Ezekierl, 2016; Ragupathi & Ragupathi, 2014; Weber et al., 2017). Technological
solutions can be implemented to reduce and simplify the differential diagnostic process
by performing pre-analytic functions prior to human interpretation of images.
Inconsistency and Variability
Conventional imaging interpretation and reporting methods are highly subjective;
and subsequently, associated with a high degree of variability (Bosmans, Weyler,
DeSchepper, & Parizel, 2011; Bruno et al., 2015). Most radiology reports primarily
consist of subjective narrative descriptions of normal and abnormal findings (J. Y. Chen
et al., 2017). Moreover, radiologists vary in their use of interpretive descriptors and
reporting structures (Napel & Giger, 2015). This personalized approach to reporting
contributes to inconsistencies within and between radiology reports. In support of this
premise, radiology reports have been described by numerous researchers as incomplete,
inconsistent, and inconclusive (Bosmans et al., 2011; J. Y. Chen, Sippel-Schmidt, Carr, &
44
Kahn, 2017). In one particular example, variability of the interpretation of spine imaging
by different radiologists was reflected by a misinterpretation rate of 43.6% plus or -11.7
(Herzog et al., 2017). The study was based on the interpretation of 10 MRI studies
performed on the same patient at 10 MRI centers, read by 10 independent radiologists.
The degree of inconsistency and variability occurring during the interpretive stage of
radiology workflow in all fields including spine care must be improved.
In addition to variability of the reporting process, there is also a high level of data
management and data access variability during interpretive workflow (Aerts et. al.,
2013). Factors, which adversely influence the flow of data and the pattern of data access
during radiology workflow, include incomplete access to medical records and prior
imaging reports, inadequate imaging protocols, human error, and technical workstation
deficiencies. Image interpretation and the description of pathology often vary between
radiologists. The degree of interpretive variability is influenced by a radiologist’s level of
experience, the time allowed for the interpretive process, the complexity of the study, and
the presence of human bias (Napel & Giger, 2015). Interpretive variability is also
influenced by reporting requirements. AI has the potential to improve patient care by
improving the flow of data and access to data across the spectrum of radiology workflow
(Augimeri et al., 2016; Boone et al., 2015). AI can also influence how radiology reports
are structured.
Cognitive Bias
Cognitive bias represents an error in reasoning. Over 100 different forms of
cognitive bias have been identified in health care (Croskerry, 2017). Common forms of
45
bias in radiology include availability bias, alliterative bias, anchoring bias, framing bias,
satisfaction of search bias, and pro-innovation bias. Availability bias occurs when a
decision is influenced by experiences, whereas alliterative bias occurs when an
individual’s judgment is influenced by another. Anchoring bias refers to limiting the
search for additional possibilities due to the belief that a prior assumption is correct or
that a current diagnosis fully explains a patients presentation (Tversky & Kahneman,
1974). Framing bias refers to the use of a limited perspective. Satisfaction of search bias
refers to the assumption that the diagnostic process is complete due to lack of knowledge
of other differential diagnostic possibilities. Pro-innovation bias refers to assigning value
to the role of new technology or the data it provides without consideration for potential
inaccuracies or inconsistencies (Bauman & Martigoni, 2012; Rogers, 2003). Radiologists
may not be aware of their own cognitive bias under different circumstances. Machine-
based decision support is not subject to most forms of human bias, and therefore can be
used to help avoid or overcome adverse consequences of human bias.
Research has revealed that radiologists, like experts in other fields, are subject to
a phenomenon referred to as inattentional blindness, characterized by missing what
should be obvious due to search bias (Drew et al., 2013; Memmert, 2006). For example,
Drew et al., (2013) revealed that 83% of radiologists asked to review computed
tomography lung scans for nodules and other abnormalities failed to identify a gorilla
image located in the lung field. The gorilla image was more than 40 times larger than the
average nodule. This acclaimed research study confirmed that a prioritized search for
specific pathology could lead to blinding of other significant findings.
46
Cognitive bias can influence data acquisition, data analysis, and data
interpretation during the course of radiology workflow. The primary factors that
influence cognitive bias include poor training, lack of experience, stress, uncertainty,
physiological fatigue, and incomplete information (H. W. Wu et al., 2012). Different
forms of bias may overlap or coexist within the same decision-making process. For
example, anchoring bias may be amplified by confirmation bias leading to premature
diagnostic closure (Hussain & Oestreicher, 2017). The use of detrimental heuristics may
complicate cognitive bias and result in higher risk for diagnostic error.
Process and Workflow Error
Any situation that disrupts or interrupts the interpretive process during radiology
workflow can lead to human distraction and diagnostic error. Examples include slow data
access, difficulty accessing prior imaging studies or records, and lack of familiarity with
complicated workstation technology. Additional distractions include phone calls,
interventional procedures, and conversations with health care providers (Schemmel et al.,
2016). Numerous disruptive factors commonly occur simultaneously or within a short
time frame in the radiology setting.
The absence of technological decision support in the presence of complex data
can result in an inefficient, inaccurate, and untimely diagnostic process. Potential
solutions include the use of embedded AI decision support, access to integrated radiology
and pathology data, physical workflow modification, and simplified workstation
interfaces (Bruno et al., 2015; C. S. Lee et al., 2012). It is important to embed
47
technological solutions within radiology workflow to simplify the interpretive process
and to reduce the risk for diagnostic errors.
Artificial Intelligence: An Introduction
Artificial intelligence (AI) refers to technology that exhibits biological-like
properties to assist, augment or replace human processes or actions (Pannu, 2015). The
Turing test developed decades ago offered an operational definition of AI, which required
that a machine possess certain human-like attributes and capabilities that could be used
for problem solving (Turing, 1950). The two principal forms of AI are machine learning
and natural language processing (Jiang et al., 2017). Natural language processing (NLP)
is used for textual analysis, whereas the use of machine learning (ML) in radiology
supports computer-aided disease detection, characterization, and monitoring. ML can be
used to assist in the differential diagnostic process.
The three primary forms of AI are assisted intelligence, augmented intelligence,
and autonomous intelligence (Bothum & Lancefield, 2017). Rapid advances in computer
technology, software programming, and algorithm development have accelerated the
evolution of AI in the direction of autonomy for some tasks. Assisted intelligence refers
to the use of technology to improve a process a human is capable of performing. In
contrast, augmented intelligence refers to the use of technology to enhance human
potential. Autonomous intelligence refers to the use of technology to perform a task that
exceeds human capability. One of the most important attributes of AI is speed. It can
perform most tasks much faster than humans can. Some AI solutions will evolve from
48
offering assistance to becoming autonomous. The primary elements of an intelligent
system include infrastructure, algorithms, data, software, and an ecosystem.
The broad topic of AI encompasses different computational methods such as
machine learning (ML), deep learning (DL), and cognitive computing (CC). The
numerous subcategories of AI, each of have different potential applications and response
characteristics (Figure 1). The basic elements of expert machine systems include a
knowledge base, an inference engine, and established rules operationalized by algorithms
(Salem, 2017). Algorithms represent digital rules used to perform an automated task or
operation. DL algorithms have many applications in radiology. For example, they can be
used to reveal new features of disease, not previously identified. Conventional ML
algorithms are linear, whereas DL and CC algorithms are more abstract, characterized by
a hierarchy of increasing complexity. Leading radiology companies have adopted
different terminologies for their AI solutions. For example, Phillips refers to its AI
solution as adaptive intelligence, General Electric refers to its solution as applied
intelligence, and IBM refers to its solution as cognitive computing (Freiherr, 2018).
Despite the use of different terms, the goal is to use technology to enhance or replace
human performance for improving a process and/or an outcome.
Traditional computer programming requires the use of explicit rules to perform
tasks. In contrast, ML uses statistical techniques and algorithms that do not require
explicit rules (Cai et al., 2016). The processes associated with ML performance can be
classified into three primary categories, which are supervised learning, unsupervised
learning, and semi-supervised learning. Supervised learning refers to the use of expert
49
derived handcrafted rules that serve as map for the flow of data between input, output,
and ground truth. Unsupervised learning refers to the use of one or more algorithms
designed to reveal patterns within data without a priori rules or human intervention.
Semi-supervised learning represents a combination of both approaches. Supervised
learning is often used to train a model to make a prediction, whereas unsupervised
learning methods are often used to explore data without a preconceived determination.
Fundamental ML data analysis methods include classification, regression, clustering,
pattern matching, density estimation, and dimensionality reduction (Kotsiantis, 2007).
ML systems can be used in radiology to detect patterns, aggregate data, and classify
disease features.
Figure 1. There are different forms and applications of artificial intelligence. Each
subtype in the above figure is less dependent on handcrafted rules and more capable of
detecting patterns in complex data and learning with exposure to data.
50
Deep learning (DL) represents a subset of machine learning, modeled after
neurological signal transmission in the nervous system (Zaharchuk et al., 2018). One of
the most common forms of DL is referred to as convolution neural networks (CNN),
characterized by an advanced network of connectivity capable of performing complex
parallel and serial processing (Cascianelli et al., 2017; Dreyer & Geis, 2017; Zaharchuk
et al., 2018). DL methods are able to detect patterns within high-dimensional data sets
using layered processes that impart logic (Lakhani & Sundaram, 2017). Neural networks
have generally outperformed individual algorithms in the analysis of complex data
(Lerner et al., 2018). DL solutions tend to become more accurate with exposure to data
due to their ability to detect patterns and create supportive algorithms (Zaharchuk et al.,
2018). DL methods have been successfully used in different health care fields including
neuroradiology. For example, Gao et al., (2017) used CNN for the automatic
classification of 285 non-contrast brain CT examinations into one of three categories of
pathology. The autonomous capacity of DL systems to learn and to construct pattern
detection algorithms renders them capable of detecting previously unrecognized patterns
within complex imaging datasets.
Algorithms represent a programmed set of rules delivered through a sequence of
operations to manipulate and analyze data to achieve a desired outcome (Obermeyer &
Emanuel, 2016). They may be handcrafted or created by a computer. Algorithms
essentially represent mathematical formulas that are used as instructions for a digital
process (Mapoka, Masebu, & Zuva, 2013). They are used to detect patterns, sample
variables, register instances, highlight structures, and create reference maps. Algorithms
51
may be deterministic, logical, or recursive. Their success is dependent upon their
relevance, accuracy, and speed. Algorithms render computer systems capable of
augmenting the role of the radiologist.
The most common types of algorithms used in radiology are k-nearest neighbors,
convolutional neural networks, fuzzy logic, support vector machines, decision trees, and
Naïve Bayes algorithms (Adduru et al., 2017; Erickson, Korfiatis, Akkus, & Kline, 2017;
Fan, Lin, & Tang, 2017). Algorithmic outcomes can be expressed in many forms, which
include statistics, natural language, flowcharts, diagrams, and a list of probability-based
differential diagnostic possibilities. The handcrafted algorithm sometimes serves as a
computational bridge between human and machine processes to transform complex data
into practical information (Beam & Kohane, 2017; Obermeyer & Ezekiel, 2016).
Algorithms can be used to sort through the millions of variables and patterns within
advanced imaging datasets. They can also be used to correlate structured and
unstructured data within radiology workflow.
AI uses two primary types of algorithms during computational analysis. These
algorithms are referred to as classifiers and controllers. Classifiers are used for pattern
detection and pattern matching, whereas controllers are used to assign an action or task to
computational outcomes. Special digital filters combined with algorithmic functions have
been used to mine and interrogate data (Obermeyer & Emanuel, 2016). Algorithms are
not subject to cognitive bias but have the potential to introduce machine bias during their
creation and evolution. Machine bias can lead to false negative or false positive
outcomes.
52
Introduction to IBM Watson: A Cognitive Computing System
IBM Watson is a cognitive computing system that was introduced in 2015 to offer
solutions for problem solving and decision making in the presence of complexity and
uncertainty (Hoyt, Snider, Thompson, & Mantravadi, 2016). The system is capable of
generating “ descriptive, predictive, and visual analytics, thus, reducing the risk for error”
(Hoyt et al., 2016, p. e165.). Cognitive computing is capable of analyzing structured and
unstructured data; thereby, increasing its potential utility in radiology.
IBM Watson represents the first large-scale integrated clinical diagnostic support
system (CDSS) capable of analyzing structured and unstructured data, and rendering a
differential diagnostic list based on pattern detection and probabilistic calculations. The
system uses a series of interdependent computational methods to respond to inquiries.
The process generally begins with a question followed by hypothesis development,
evidence analysis, and probability assignment (Deloitte, 2015; Ferrucci, 2012). With
further development, IBM Watson may become capable of augmenting the role of the
radiologist during image interpretation (Brink et al., 2017; Jha & Topol, 2016; Kharat &
Singhal, 2017; Jiang et al., 2017; Ranschaert, 2016). In summary, IBM Watson has the
potential to auto detect abnormalities, filter irrelevant or normal imaging data,
characterize disease, and assist in the differential diagnostic process with probability
assignment. Watson is capable of offering data and knowledge-driven decision support
that can be used to assist, augment or in some cases replace the role of the radiologist
(Kohn et al., 2014). Furthermore, IBM Watson is a pioneering solution capable of
integrating ML and NLP applications during the interpretive stage of radiology workflow
53
(Jiang et al., 2017). IBM technology simply serves as an example of what AI is capable
of achieving in health care. Many products will be developed to perform the same or
similar tasks. It is too early to tell which technologies will meet the reproducibility and
validation demands of future research and regulatory requirements.
Successful adoption and use of AI solutions such as IBM Watson in radiology
requires validation studies, heightened awareness of its potential value, and an adequate
state of readiness. AI solutions have the potential to reveal biological variability and the
features of pathology in vivo in a manner which exceeds human capabilities (Gillies et
al., 2016; Limkin et al., 2017; Pannu, 2015; Yip & Aerts, 2016). AI solutions also have
the potential to calculate and assign probability to differential diagnostic possibilities
based on analysis of structured and unstructured data. For these reasons, AI solutions
should be developed to improve the precision and personalization of the diagnostic
process.
Collective Intelligence: A Collaborative Approach
Data can be analyzed and health care decisions can be made using collective
intelligence (CI), a process referring to the integration of human expertise and computer
analytics. The capabilities of CI include reducing the complexity of data, revealing
patterns within the data, assigning value to classifications of data, and offering decision
support (Hoyt, Snider, Thompson, & Mantravadi, 2016). The future impact of human-
machine collaboration is difficult to predict because human performance tends to
improve in a linear and relatively predictable fashion, whereas the capabilities of AI grow
exponentially. Advanced AI systems can learn from mistakes and successes and are
54
therefore less likely to repeat mistakes than humans are. Humans add an element of
creativity and intuition to AI output.
Unlike computers, humans have a limited capacity to analyze complex data.
Humans also have difficulty correlating multidimensional nonlinear variables,
performing syntactic transformations, and revealing patterns within variable high velocity
data (Nendaz & Perrier, 2012). Humans are prone to physiologic limitations and fatigue,
whereas computational technology is stable and consistent (El-Kareh et al., 2013; Russel
& Norvig, 2010). Human intelligence is highly dependent on experience, analytic skills,
intuition, and motivations (El-Kareh et al., 2013). In contrast, AI systems are highly
dependent upon access to annotated training data, ground truth, and validation methods
(Russel & Norvig, 2010). Human intelligence has many characteristics not offered by AI
such as adaptability, intuition, creativity, flexibility, and the ability to plan.
AI systems are not capable of replacing the full breadth and depth of human
reasoning and judgment. For example, the human role in radiology offers the benefits of
unique experiential insight, empathy, and thoughtful decisions. The combination of
human and artificial intelligence offers a collaborative approach with a greater chance of
success then either isolated approach in some settings. A collaborative relationship
supports what the radiologist does best and combines it with what AI does best.
Radiology Workflow Defined
The radiology workflow environment is complex. Each stage of radiology
workflow has unique elements or processes (Figure 2). Radiology workflow has been
divided into two primary stages referred to as interpretive workflow and non-interpretive
55
workflow (Lee et al., 2017; Schemmel et al., (2016). I divided the non-interpretive stage
into two subsets referred to as the pre-interpretive and post-interpretive stages. AI is
poised to play a significant role within all stages of radiology workflow, although its
greatest potential is likely within the interpretive stage, the focus of this research study.
AI solutions can be used during the interpretive stage of radiology workflow to detect,
characterize, classify, and monitor disease. Reasoning during the diagnostic process will
involve various types of diagnostic inference and decision support. AI support during the
interpretive stage of radiology workflow can also be used to provide a probability-based
differential diagnostic list for the attending radiologist.
Figure 2. The three stages of radiology workflow that are pre-interpretive stage,
interpretive stage and post interpretive stage. The figure also highlights elements of the
interpretive stage of workflow, which can benefit from AI support.
56
Many radiologists spend a greater amount of time performing non-interpretive
tasks than interpretive tasks during the course of radiology workflow (Dhanoa et al.,
2013). This pattern of labor contributes to inefficient and occasionally inaccurate
interpretive outcomes. Limited time combined with exposure to large volumes of
complex data exposes the radiologist to a higher degree of uncertainty and subsequently
to greater potential for the influence of human bias and error during the interpretive stage
of radiology workflow. There are numerous forms of human bias, which can take place
during the interpretive stage of radiology workflow (Figure 3). Quite often more than one
form of bias will be present. AI-based decision support can reduce the impact of human
bias.
Non-interpretive human tasks during radiology workflow include setting image
protocols, supervising studies, directly caring for patients, accessing analytic tools,
performing image-guided intervention, and consulting with health care providers. Human
tasks typically performed during the interpretive and post-interpretive stages of radiology
workflow include visual evaluation of images and qualitative report generation. The use
of AI methods during the interpretive stage of radiology workflow could help reduce the
incidence and prevalence of interpretive and diagnostic errors (Busby, Coutier, &
Glastonbury, 2018; Itri & Patel, 2018). Radiomic methods are increasingly becoming a
more important quantitative measure in diagnostic imaging, one that may change the
landscape of health care. The success of AI use in radiology will be dependent on its
ease-of-use, utility, and ability to be integrated into existing workflow.
57
Figure 3. Types of human bias that can occur during the interpretive stage of radiology
workflow.
Emerging Options for Interpretive Radiology Workflow
The current tasks performed during the interpretive stage of radiology workflow
must be adapted or replaced to support a more accurate and timely diagnostic process in
the presence of complex and high velocity data. The potential for reducing the incidence
of diagnostic errors in radiology must be addressed at all stages of radiology workflow,
including the final report, to have maximum impact on patient care (Bruno et al., 2015).
AI has the potential to assist and augment the radiologist during interpretation through its
capacity to access, analyze, and correlate data acquired from numerous sources (Dreyer
& Geis, 2017). This includes prior imaging studies, electronic health records, laboratory
studies, genetic profiles, published literature, and disease registries. AI can be used to
pre-analyze data, flag abnormal presentations, and subsequently prioritize the order of
58
what needs to be interpreted by the radiologist (Kahn, 2017). A single health care
provider such as a radiologist or pathologist cannot keep track of all relevant information
on a particular patient or on a disease process. AI solutions can help overcome human
limitations such as lack of experience or unawareness of possibilities.
Radiology workstations of the future will be required to meet new standards based
upon performance, reliability, and scalability. Effective performance will require easy
access to an AI menu of options, technological interoperability, and seamless networking
with external resources (Kahn, 2017). Adaptive workstations will be used by radiologists
to leverage the experience of peers from remote locations and to democratize expert
decision support with the help of AI. It will also be used by the radiologist to acquire
knowledge-based assistance from databases and published resources almost
simultaneously. In summary, for radiology workflow to be successful it must support the
role of the radiologist and lead to better patient care.
Natural Language Processing
The science of linguistics encompasses the form, meaning, and context of sounds
and symbols used to communicate (Zipf, 2012). Natural language processing (NLP)
represents a solution for linguistic analysis. It can be used to assign meaning or value to
unstructured (textual) data allowing it to be analyzed with computational methods (Pons,
Braun, Hunick, & Kors, 2016). The majority of digitized health care data within a
patient’s medical records and within the realm of population databases is unstructured.
The presence of pathology often reported semantically in radiology reports (Acharya et
al., 2017). NLP systems can access prior pathology and radiology reports to reveal prior
59
evidence and characteristics of pathology. NLP can therefore be used to leverage
information within electronic medical record systems to generate an active problem list.
It can also be used to provide access to prior imaging study findings for comparative
analysis and to provide decision support during radiology workflow (Massat, 2018).
Biomedical text mining with NLP leads to knowledge discovery that can aid imaging
interpretation.
NLP uses different methods to analyze unstructured data. The fundamental
mechanisms include pattern matching, parsing, and statistical approaches (Cai et al.,
2016). Hassanpour and Langlotz (2016) demonstrated that ML combined with NLP could
be effectively used to analyze textual data acquired from medical records including prior
radiology reports to assist in the differential diagnostic process. NLP represents an
important data management tool within the interpretive stage of radiology workflow. AI-
supported NLP can only reveal what was previously described, not the current state of
pathology evaluated through imaging. Characterization of an individual’s current state of
pathology requires in vivo imaging and quantitative measures using various techniques
such as radiomics. The combined use of NLP and radiomic methods will improve the
radiologists’ access to relevant information.
Radiomics and Radiogenomics
Radiomics refers to the science associated with high-throughput extraction and
quantitative analysis of non-visible data acquired from images to characterize pathology
(Lambin et al., 2017; Lee et al., 2017; Limkin et al., 2017; Griethuysen et al., 2017; Yip
& Aerts, 2017). Radiomics, first reported in 2012, has evolved quickly during the last
60
couple of years (Verma et al., 2017). Automated radiomic functions include pathology
detection, segmentation, feature extraction, and feature analysis (Lambin et al., 2017;
Limkin et al., 2017; Zaharchuk et al., 2018). Radiomic methods can be used to expand
available data within three-dimensional space on imaging studies to characterize
pathology and subsequently refine the diagnostic process (Gillies et al., 2015; D. Kumar
et al., 2012; Lee et al., 2017; Lambin et al., 2017; Limkin et al., 2017; Peeken et al.,
2018). Expanded dimensionality amplifies spatial heterogeneity (Cook et al., 2014). One
of the goals in radiomics is to identify disease features and biomarkers that have greater
causal rather than correlative relationships (Sanduleanu et al., 2018). Radiomic methods
have the potential to expand sub classifications of disease and support more personalized
care.
Radiomic measures can be used to quantify characteristics of pathology that are
not visible to the radiologist. This form of analysis addresses features of pathology such
as shape, volume, edge characteristics, texture, and other statistical measures (Aerts,
2016; Court et al., 2016; Gillies et al., 2016; V. Kumar et al., 2012; Lambin et al., 2012).
Radiomics is a discovery process rather than a validation process performed using one of
two approaches. The first approach focuses on mining images for predetermined patterns
of disease. The second approach uses deep learning methods to discover and learn disease
features not currently known. Radiomic methods like other analytic approaches tested for
reliability and validity prior to clinical use. Testing should include evaluation of
sensitivity and specificity along with the reliability of predicting positive and negative
values. The testing must be disease-specific and take into account different cohorts. The
61
primary goals of radiomic development are to help detect, characterize, and monitor
pathology. This includes classifying and staging disease.
Radiomic methods have been successfully used to detect and characterize some
diseases using different imaging methods such as CT, MRI, and PET (Acharya et al.,
2018; Parekh & Jacobs, 2016). Radiomics has been successfully used in different
research settings to evaluate breast cancer (Li et al., 2016), brain tumors (Li et al., 2017),
lung disease (Bak et al., 2018; Vallieres et al., 2015), liver disease (Naganawa et al.,
2018), brain metastasis (Ortiz-Ramon et al., 2017), and prostate cancer (Tanadin-Lang,
2018). Most of the important contributions to radiomics have come from the field of
oncology (Gillies et al., 2016). Radiomic methods show promise in the field of breast
imaging for the differentiation of benign versus malignant tumors (Hui et al., 2016).
Specific radiomic features such as enhancing tumor volume and texture features are
emerging as discriminatory factors in the differential diagnostic workup of breast cancer
(Drukker et al., 2018; Hui et al., 2016). Ongoing research is required to identify the
potential applications and benefits of radiomics in different fields. Future research will
reveal whether some of the methods used in other fields may be adapted and used in
spine care.
Radiomic methods are not limited to the initial diagnostic process. They can also
be used to help monitor disease progression and the response of pathology to treatment
(Vargas et al., 2017; Zhou et al., 2017). Radiomic methods can be used for the
surveillance of early-stage pathology prior to intervention. In addition, radiomic
signatures can ”…be used as precision biomarkers for the prognosis of individual
62
patients” (Castiglioni & Gilardi, 2018, p.412). The clinical role of radiomics is based on
the premise that imaging of disease requires a large volume of data that reflects multi-
scale pathological mechanisms, not detectable through routine visual assessment of
images.
The specialized field of radiogenomics sometimes referred to as imaging
genomics refers to the science associated with the correlation of anatomic characteristics
(phenotype) of pathology with genetic (genotype) data (Pinker et al., 2017).
Radiogenomic methods have been successfully used in the evaluation of lung cancer,
glioblastoma multiforme, kidney cancer, prostate cancer, and liver cancer (Incoronato et
al., 2018). Radiomic measures have proven useful for revealing phenotypic
manifestations of genetic expression (Giger, 2018; Panth et al., 2015). Phenotypic
characterization is important for there are many non-genetic determinants in pathology,
especially involving age-related diseases (Oakden-Rayner et al., 2017). Radiogenomics
has the potential to support precise classification of disease and subsequently the
personalized delivery of care (Bai et al., 2016). This conclusion is based on the premise
that alterations of genetic expression influence pathology represented by phenotypes
revealed through diagnostic imaging methods.
Volumetrics: New Realities and 3D Perspectives
Radiologists have a unique opportunity to combine the use of AI and volumetric
datasets to display pathology in three dimensions (3D) to help inform the diagnostic
process and therapeutic planning (Farahani et al., 2017). A few years ago, Denzel et al.
(2014) acknowledged the importance of in vivo interrogation of pathology within the
63
framework of volumetric imaging and display. A study limited to planar or 2-D imaging
can limit the diagnostic process.
Advanced imaging methods such as MRI and CT are capable of acquiring high-
resolution volumetric data sets that can be formatted to create 3D perspectives (Douglas
et al., 2016). The use of augmented reality and 3D viewing can improve the conspicuity
of pathology features (Douglas et al., 2016; Hamacher et al., 2016). Segmented pathology
can be manipulated and interrogated using virtual cut plane technology.
Multidimensional imaging data sets can be used to create virtual reality (VR) and
augmented reality (AR) based representations of pathology. Virtual reality offers an
immersive experience that can be used for planning and training. In contrast, augmented
reality offers digital images or prompts in the physical world. AR can be used to project
nonvisible perspectives of pathology during an exploratory biopsy or over a surgical
field, thus, limiting attention shifts between available imaging and the patient. The
augmented virtual or interactive display of pathology may soon be used to guide the
digital (virtual) biopsy and acquisition of radiomic measures.
Disease Modeling and Computational Diagnosis
Mathematical formulas can be used to help study normal biological and disease
processes (Mapoka et al., 2013). In silico modeling sometimes referred to as ex vivo
modeling refers to the use of quantitative imaging data to create comprehensive
mathematical models which can be used within the context of a computer system to study
the characteristics and behavior of pathology (Jeanquartier, Jean-Quartier, Cemernek, &
Holzinger, 2016; Louis et al., 2016; Mapoka et al., 2013). Mathematical models support
64
the use of simulated variables to predict the behavior of disease under different
circumstances (Louis et al., 2016). Radiomic features acquired in vivo can be used to
help develop in silico models (Echegaray et al., 2016). The use of mathematical disease
models built from data acquired over time through monitoring of a disease process could
help support predictive modeling.
The benefits associated with in silico modeling include the ability to study
disease, zoom in on pathology subsystems, and manipulate time scales, all of which
reduce the need for laborious and costly biological experimentation (Jeanquartier, et al.,
2016). Computational disease models can be used to test assumptions and hypotheses.
They can also be used to help predict outcomes and to reduce the frequency and
magnitude of uncertainties surrounding disease behaviors (Mapoka et al., 2013). Access
to mathematical disease models from the radiology workstation can augment the
differential diagnostic process. In the future, mathematical models will be used to help
identify ground truth and to train AI systems.
Computational diagnostics is the science associated with the combined use of
algorithms, mathematical formulas, and computer systems to detect and study disease.
Computational pathology refers to the use of a digital disease model to evaluate
assumptions and to study disease behavior under simulated circumstances. The
convergence of interoperable technologies offers an unprecedented opportunity to
correlate pathomic, radiomic, and genomic features to create computational disease
models. In the near future, computational analysis may become as important to radiology
as the microscope is to pathology (Louis et al., 2016). Access to AI supported
65
computational diagnostic tools during the interpretive stage of radiology workflow will
enable radiologists to address the complexities of pathology.
Workflow Optimization and Data Integration
The differential diagnostic process in radiology typically requires correlation of
imaging findings with data from other sources such as prior imaging studies, medical
records, peer-reviewed literature, and disease registries. Radiology workstations will
eventually consist of embedded AI solutions and networked functions to perform
interpretive functions, to obtain second opinions, and to support multidimensional
viewing (Morgan, Branstetter, Mates, & Chang, 2006). Radiomic workflow includes data
acquisition, pathology segmentation, feature extraction, feature analysis, and correlation
with computational disease models. The radiologist will serve as the expert on the in vivo
diagnostic process and as a gatekeeper to the flow of data and access to knowledge from
diverse data sources.
The radiology workstation of the near future will likely represent the principle
hub of converging data from the fields of radiology, pathology, and genetics (Gillies et
al., 2016; Jha & Topol, 2016). Computational methods will be required to analyze data
from these diverse sources. Unlike the radiologist, AI solutions do not represent a single
layer of interpretation. It represents multilayered approaches, consisting of logical
analysis and problem solving, not subject to the variability and inconsistencies of a
biological system.
Data acquired from nonimaging sources need to be annotated or tagged in a
manner that prepares it for sorting, aggregation, and analysis. Complex data often has to
66
be conditioned and prepared for exposure to algorithms and computational analysis. The
curating process requires well-defined steps to identify inconsistencies, to standardize
symbolic representations, and to transform data to a uniform interoperable and
interpretable language. This process can be augmented through the identification of
common data elements (CDEs). CDEs help integrate disparate clinical, phenotypic, and
genotypic data. The consistent use of CDEs supports the development and use of disease
registries (Rubin & Kahn, 2017). Widespread application of CDEs will support the
development of integrated knowledge databases accessible by AI for use during the
interpretive stage of imaging. Curated data sets are required to train AI.
Access to Disease Databases and Registries
The increasing ability to extract pathology features during imaging procedures
will result in new forms of data storage and disease registries (Rastegar-Mojarad et al.,
2017). Digital tissue samples can be stored in a manner that will maintain accurate spatial
and relational data. Unlike in vitro tissue samples, in vivo digital representation of
pathology will not degrade. Digital representation of in vivo characteristics can be
acquired and stored in a manner that maintains three-dimensional relationships. The
radiologist will have increasing access to digital pathology registries and mathematical
disease model libraries from their workstation to augment the differential diagnostic
process.
Structured Reporting
The radiology report reflects what took place during the interpretive stage of
radiology workflow. The decisions made influence the content and structure of the final
67
report. Structured reporting refers to the use of a standardized format, quantitative
measures, and consistent descriptions. Structured radiology reports can be
operationalized with quantitative measures derived from AI applications (J. Y. Chen et
al., 2017; Dreyer & Geis, 2017). Growing use of widely accepted terminology and
standardized disease nomenclature will support more consistent diagnostic reporting and
reliable NLP applications. Using keywords and phrases on reports can serve as common
data elements used to trigger prospective or retrospective interpretive or comparative
processes. For example, describing focal pathology in a prior radiology report may direct
auto detection and quantitative radiomic measures within the same volume of interest in a
subsequent study. Structured reporting provides more concise decision support at the
point of care (Alkasab et al., 2017). According to Schwartz et al. (2011), “Referring
clinicians and radiologists found that highly structured reports had better content and
greater clarity than conventional reports” (p. 174).
In the near future, AI will offer a full spectrum of solutions that will begin with
pathology detection and end with automated reporting within a structured format
(Zaharchuk et al., 2018). The combined application of ML techniques and NLP to extract
patterns from current imaging studies and prior radiology reports will improve imaging
interpretive accuracy (J. Y. Chen et al., 2017). Text clues within a prior radiology report
can be used to trigger access to current imaging and prior non-imaging data to support the
interpretive process.
68
Digital Exploration: The Virtual Biopsy
Growing use of AI combined with radiomic methods will expose new features of
disease and expand the spectrum of pathology. This process will lead to a longer list of
disease subtypes and differential diagnostic possibilities during the interpretative stage of
imaging, thereby, increasing the complexity of decision making. Advances in technology
will support the exploration and interrogation of in vivo pathology with visual,
augmented, and subvisual techniques. Augmented visual approaches include the use of
interactive 3-D displays of pathology.
Subvisual Tissue Interrogation
Radiomic methods can be used to detect and characterize features of pathology
from imaging studies in a manner undetectable by traditional visual interpretation (Aerts,
2017; Gillies et al., 2016). Advanced algorithms can be used to correlate subvisual
pathology features with other sources of data such as histological features from pathology
and genetic profiles (Bucking et al., 2017). Echegaray et al. (2016) introduced the
concept of the “digital biopsy” which refers to the targeted non-invasive acquisition of
pathology features in vivo” (p. 283). In another field, Mancini et al. (2018) introduced the
concept of the “digital liver biopsy” referring to the use of multimodality and
multiparametric imaging of the liver (p. 3). The use of a digital or virtual biopsy approach
has the potential to interrogate and map the entire landscape or volume of a pathological
state, whereas the traditional needle biopsy offers limited characterization of pathology
(Echegaray et al., 2016; Lambin et al., 2012; Thrall et al., 2016). It is important to
evaluate and characterize a whole region of pathology to reduce or eliminate sampling
69
bias. Under sampling may poorly represent the entire state or spectrum of pathology
present and subsequently misdirect treatment. Traditional and virtual biopsy results can
be combined to better characterize pathology. The virtual biopsy can be also used to
augment or direct the traditional needle biopsy.
Radiomics is not limited to the detection and extraction of pathology features. It
can be used to create or reveal new quantitative descriptors of pathology. Subsequently,
radiomic methods will continue to contribute to the development of new molecular
biomarkers and imaging signatures of pathology (J. Wu et al., 2018). For this reason,
radiomic methods are being considered for the expansion of the traditional tumor-node-
metastasis (TNM) staging process in oncology (Lai-Kwon, Siva, & Lewin, 2018). Digital
exploration of pathology can be enhanced through partitioning of an image and by
digitally extracting tissues or structures from the field of view, which can enhance
regions of interest. Special digital filters can also be used to increase the dimensionality
of a volume of interest (VOI), thereby improving tissue feature analysis and the ability to
sub-classify disease (Parekh & Jacobs, 2016). Edge detection algorithms is used to help
identify the boundaries between healthy and diseased tissues. Contour analysis is used to
segment and highlight the borders of pathology. Knowledge of the true border of
pathology supports more accurate volumetric measures and feature mapping of
pathology. Radiomics is capable of providing an automated solution for the assessment of
disease characteristics and evolution.
Radiomic methods do not always serve as the final or conclusive method of
pathology assessment. They can be used to help identify the signature for an aggressive
70
region of pathology that can serve as a high-risk target for a traditional needle biopsy
within a well-defined region of pathology (Sala et al., 2017). Radiomic feature analysis
using multiparametric or hybrid imaging technologies such as PET/CT and PET/MRI can
be used to interrogate large volumes of tissue in vivo with subcentimeter resolution,
which enhances the accuracy of the virtual biopsy (Kressel, 2017). Integrated data from
these approaches can also be used to better guide traditional biopsy methods.
Pathology Feature Extraction and Analysis
Radiomic methods are effective for revealing geometric, statistical, and textual
features of pathology from imaging data. Over 440 radiomic features of pathology have
been acknowledged in the literature (Wu et al., 2018). Common categorical features of
pathology include texture, edge-attributes, volume measures, molecular relationships,
contrast kinetics, uptake values, and subvisual pathology boundaries (Bai et al., 2016).
Individual voxels and three-dimensional matrices comprised of an array of voxels
provide the volumes of interest (VOI) for radiomic assessment. Texture analysis refers to
the quantitative evaluation of grayscale intensities and their relationships within and
between well-defined three-dimensional space referred to as voxels (Davnall et al., 2012;
Lubner et al., 2017). Texture may correspond to various pathological processes.
Specialized digital filters can be used to help reveal unique features of pathology such as
entropy and uniformity (Nandu, Wen, & Huang, 2018). The greater the number of
clinically relevant pathology features, the greater the potential for differentiating and
classifying pathology.
71
Diagnostic images are comprised of first, second, and third order data which when
exposed to advanced computational methods can detect patterns, relationships, and trends
that support health care decisions (Gillies et al, 2016; Huang et al., 2017). First-order
statistics include mean gray-level intensity, standard deviations, entropy, skewness,
kurtosis and uniformity. Second-order statistics include local homogeneity, dissimilarity,
correlation, and angular second moment energy (Gillies et al., 2016). Higher order
statistics include coarseness, contrast, and complexity (Davnall et al., 2012). Data can be
acquired and analyzed during the interpretive stage of radiology workflow with
handcrafted and/or deep learning radiomic methods.
Multidimensional 3D Exploration
A focal region of pathology such as a tumor often has a high degree of spatial and
temporal heterogeneity, which limits the usefulness of conventional structural imaging
and traditional biopsy results. A digital (virtual) biopsy can be performed using an in vivo
voxel by voxel (voxel-wise) interrogation process across the entire volume of pathology
in multiple dimensions. Voxel-wise classifiers can be used to help reveal varying stages
and subtypes of pathology within a single lesion (Bucking et al., 2017; Ng et al., 2013).
Radiomic methods have been successfully used to help differentiate benign and
malignant characteristics of tumors and to provide prognostic insights (Sala et al., 2017).
Further research will expand knowledge of the molecular attributes and radiomic
signatures of aggressive pathology within different biological systems and tissues.
Comprehensive in vivo assessment of pathology requires analysis across the
three-dimensional volume of pathology rather than from a two-dimensional plane
72
(Echegaray et al., 2016; Nandu, Wen, & Huang, 2018). As data acquisition and
computational methods become faster, the trend will be to auto interrogate larger
volumes of pathology and auto map pathology features using in vivo data (Zhao et al.,
2016). Virtual exploration supports interrogation of pathology within an augmented or
immersive 3-D environment.
Pathology does not exist in isolation. It functions and interacts within a biological
ecosystem that includes the surrounding microenvironment. Within this ecosystem, there
is both phenotypic and genotypic plasticity that contributes to the evolution of pathology.
The use of multiparametric imaging methods and radiomic measures can help expand and
reveal feature data sets from the whole region of pathology including the surrounding
microenvironment. This represents a significant advantage of multidimensional
interrogation of pathology. It also offers an advantage over lab (blood) studies that
represent circulating biomarkers, which do not localize pathology within an organ or
tissue. In the near future, whole pathology slide mounts will be matched with image slice
acquisitions and voxel by voxel radiomic measures. This approach will be used to help
build computational disease models and disease detection algorithms.
Disease Screening: Discovery Radiomics
Early detection and characterization of pathology influences patient care and
treatment outcome. Handcrafted algorithms and rules used with AI are often limited in
their capacity to reveal subtle or unknown characteristics of healthy and disease states. D.
Kumar et al. (2017) introduced the concept of “discovery radiomics” which refers to the
use of high-throughput analysis of imaging data to detect early-stage pathology and to
73
help predict the outcome of an asymptomatic disease process. A specialized approach to
data analysis referred to as “evolutionary deep radiomic sequencing” offers a rapidly
evolving method for detecting pathology, which is not dependent on a priori knowledge
of disease criterion (Shafiee et al. 2017). In summary, various forms of radiomic
applications can be used to provide a low-cost, fast, and potentially reliable method of
screening for pathology.
Reimagining the Differential Diagnostic Process
The differential diagnostic process is comprised of a series of interrelated steps
applied to a patient’s presentation, which uses probability-based logic or reasoning to
differentiate a disease or disorder from others that may have a similar presentation. The
differential diagnostic process is often dependent upon different sources of information
such as the history, physical examination, laboratory evaluation, imaging, and other
specialized forms of testing. In general, the more complicated a patients’ presentation, the
greater the list of possible causes and contributing factors. Advances in diagnostic
imaging have contributed to a growing appreciation for the complexity and heterogeneity
of disease at anatomic, cellular, molecular, and genetic levels (Aerts, et al., 2013; Davnall
et al., 2012; Lubner et al., 2017; Sala et al., 2017; Yip & Aerts, 2017). Radiomic
measures can be used to “ bridge evidence across different biological scales” in a manner
which can inform the differential diagnostic process (Hsu, Markey, & Wang, p. 1010). It
is important to discover new dimensions and features of pathology that have meaningful
impact on patient care.
74
Heightened awareness of the complexity of disease expands the list of differential
diagnostic considerations. Advances in diagnostic imaging will continue to reveal unique
heterogeneic features that can be used to classify and subtype pathology (McCue &
McCue, 2017). The heterogeneity of pathology is associated with evolutionary changes at
the cellular level associated with genotypic and phenotypic plasticity. Heightened
awareness of this process improves diagnostic and prognostic accuracy. For example,
high levels of tumor heterogeneity have been associated with greater variability of
treatment outcome and generally a poorer prognosis (Lundstrom, Gilmore & Ros, 2017;
Sala et al., 2017). Successful delivery of more precise and personalized health care
requires knowledge of biological differences and pathological variability along with
relevant decision support at the radiology workstation.
The current diagnostic process is very nuanced, and influenced by a provider’s
familiarity with disease and related testing. In addition to assisting with disease detection
and characterization, AI can be used to identify appropriate diagnostic tests and related
protocols to help achieve a precise diagnosis (Baldwin, Guo, & Syeda-Mahmoood,
2017). There are differential diagnostic possibilities for every patient presentation
(Hussain & Oestreicher, 2017). The individual health care provider typically relies on an
intuitive diagnostic approach limited to their familiarity of two to six diseases during the
initial differential diagnostic process (Phua & Tan, 2013). This level of awareness is
often insufficient for unusual or complex conditions. Limited awareness of differential
diagnostic possibilities leads to errors and missed treatment opportunities.
75
Traditionally, radiology has relied upon the visual perception of the radiologist,
which limits their capacity to consider microscopic and molecular level differential
diagnostic considerations (Pinto & Brunese, 2010; Yip & Aerts, 2016). Radiomic
methods can be used to overcome visual limitations and expose measurable
characteristics of subvisual pathologic states, thereby, adding value to the differential
diagnostic process. Improvement of the differential diagnostic process is required in all
areas of diagnostic imaging including spine care.
The application of new quantitative imaging (QI) measures and standards will
refine the diagnostic process and lead to expanded criterion and classifications of disease
(Farooki et al., 2016; Gillies et al., 2016; Kharat & Singhal, 2017; V. Kumar et al., 2012;
Parekh & Jacobs, 2016). The expansion of objective measures of pathology will support
AI solutions. Disease states and their evolution vary between individuals because of
molecular, genetic (genotypic), structural (phenotypic), and biological diversity. For this
reason, a one-size-fits-all approach to diagnosis or treatment does not work for everyone.
Radiomic analysis, a specialized application of QI, can be used to characterize biological
variability and pathological heterogeneity at subvisual levels independent of the
radiologist’s interpretation (Larue et al., 2016; Yip & Aerts, 2017). Radiomic data will
assist the radiologist in the differential diagnostic process. The adoption of QI will also
support further development of AI by providing objective labels used to annotate training
and validation data.
The term probability in the diagnostic process refers to measures of the likelihood
of a disease or pathologic process being present. AI can expose radiologists to differential
76
diagnostic considerations (diseases) for which they have limited knowledge and
experience. Probability is assigned to disease patterns and differential diagnostic
possibilities associated with imaging presentations derived from hundreds of thousands of
patients within a database or from case-based publications.
In summary, treatment success is dependent upon an accurate and efficient
differential diagnostic process. The current diagnostic process remains too imprecise and
inconsistent to adequately identify and subtype disease in many situations. This often
results in a “one-size-fits-all” treatment approach. AI solutions have the potential to use
probability calculations to help identify and prioritize diagnostic possibilities. AI will
continue to influence the steps, as well as the sequence of steps used during image
interpretation and the differential diagnostic process. The combined use of AI-supported
data management solutions such as natural language processing, radiomic methods,
machine learning, and advanced computational analysis at the radiology workstation
might support improved accuracy and precision of the differential diagnostic process.
The Radiologist: New Roles and Responsibilities
The principal role of the radiologist is to detect, characterize, and report on
disease processes in a manner that provides decision support at the point of care. For this
reason, there is a growing demand for radiologists to become more involved in
consultation and patient care (Ranschaert, 2016). Historically, radiologists are skilled in
the evaluation of pathology associated with structural changes and are less skilled in the
early detection of pathology based on subvisual criteria revealed by deep learning and
radiomic methods (Malone & Newton, 2018). This inadequacy reinforces the need for
77
computational assistance. AI has the potential to augment the role of the radiologist
during the interpretive stage of radiology workflow and with final reporting. More
specifically, AI may improve the performance of the radiologist by enabling earlier
disease detection, offering more precise disease characterization, providing probability
based differential diagnostic considerations, and reducing diagnostic error rates
(Zherhoni, 2017). Radiologists are rapidly becoming among the most important data
managers, knowledge brokers, and primary gatekeepers of big data and curators of
knowledge for treatment planning and disease surveillance in health care (Hillman &
Goldsmith, 2011; Jha & Topol, 2016). Using AI during the interpretive stage of radiology
workflow can help identify meaningful incorporated into structured reports.
Despite these important roles, the interpretation of diagnostic images is still
primarily limited to the visual detection and characterization of pathology (Pinto &
Brunese, 2010). This approach is no longer adequate. In some areas of radiology AI has
proven it has the potential to augment the role of the radiologist in the detection and
interpretation of subvisual data (Farooki et al., 2016; Larue et al., 2016). AI has the
potential to empower the radiologist to work better, faster, and smarter.
AI can augment the role of the radiologist and empower their role as a clinical
consultant in numerous ways. For example, AI can help access, analyze, and correlate
non-imaging data with imaging findings during the interpretive stage of radiology
workflow. AI can also be used to flag subtle or visually hidden pathology and address
mundane and redundant work, thus, freeing radiologists up to interpret pathology and
better communicate with referring physicians and other members of the health care team
78
(Recht & Bryan, 2017). AI may overcome human limitations such as fatigue, lack of
experience, unawareness of possibilities, and bias. To maintain relevance the radiologist
must be willing to embrace AI, adapt to its use, and contribute to its development.
Additional research is required to address how AI applications augment the role of the
radiologist and improve the delivery of more precise and personalized care (Sutton et al.,
2017).
Artificial Intelligence in Radiology: Current Applications
The individual radiologist is often overwhelmed with the growing burden of high
volume complex data. This dilemma has led to the pursuit of different forms of decision
support. This includes AI solutions. Most of the research surrounding the use of AI in
radiology has been limited to highly specialized fields. There has been little research
surrounding its role in spine imaging. Successful non-spinal applications will pave the
way for use in spine care.
Non-Spinal Imaging
Because of the complexity of available data and the criticality of decisions, most
research on how AI is used in radiology is limited to cardiology, neurology, and
oncology. AI solutions are used in other specialties of radiology albeit to a limited
degree. Deep learning methods have been successfully used to auto detect pulmonary
tuberculosis on chest radiographs (Lakhani & Sundaram, 2017) and to auto detect disease
states in neuroimaging such as intracerebral hemorrhage, stroke, and mass effects (Maier
et al., 2015; Prevedello et al., 2017; Scherer, 2016). Computer-aided diagnostic systems
help detect and characterize some neurodegenerative disorders (Cascianelli et al., 2016).
79
AI has evolved to a level where it has outperformed the human expert in radiology in
some settings (Augimeri et al., 2016; Boone et al., 2015; Mohebian et al., 2017).
Research addressing the use of AI has demonstrated its potential to augment the role of
the radiologist.
The utility of AI is not limited to radiology. For example, deep learning
algorithms have demonstrated greater accuracy than a panel of pathologists in the
detection of lymph node metastasis in women with breast cancer (Bejnordi, Veta, & van
Diest, 2017). Extensive research is also underway to develop methods for automatically
extracting relevant information from unstructured reports in breast imaging (Gupta,
Banerjee, & Rubin, 2018). In another study, convolutional neural networks outperformed
cardiologist’s interpretation of echocardiographic images with 98% accuracy (Mandani et
al., 2017). Successful use of AI solutions in one field of radiology may be adapted to
meet the needs in another field such as spine imaging. The adoption of AI solutions
requires adequate research to identify meaningful applications, as well as to confirm
clinical utility and validity.
Spine Imaging
Diagnostic imaging is often a fundamental and influential component of spine
care. Imaging findings influence decision making at all levels of care. Subsequently,
interpretive imaging errors and missed opportunities have a profound impact on treatment
planning and treatment outcomes. An extensive literature search revealed limited
research and real-world applications of AI in spine imaging and spine care. It is common
for early applications of AI to be applied to basic steps such as anatomic localization and
80
labeling prior to implementation of more detailed applications. For example, machine-
learning algorithms with different imaging modalities can auto-identify vertebral levels
(Daenzer et al., 2014; Hetherington et al., 2017). Another research study demonstrated
that AI could be used to segment and label vertebral bodies (El-Helo et al., 2013). Auto
segmentation and labeling of anatomic regions has to be perfected before tissue features
can be auto extracted and characterized.
AI use in spine care has not been entirely limited to anatomic localization.
Machine learning methods have also been used to determine bone density, as well as to
detect and categorize vertebral compression fractures on computerized tomography
(Burns, Yao, & Summers, 2017; Doi, 2007; Hetherington et al., 2017). In another study
El-Helo et al. (2013) demonstrated that AI performed with greater than 90% accuracy in
the detection of vertebral compression deformities. This condition is relatively common
and places a significant financial burden on the health care system. Vertebral
compression deformities often missed on non-spinal imaging studies could be detected
with AI supported methods. For example, N. Kim et al. (2004) reported that
approximately 50% of vertebral compression fractures that presented on routine lateral
chest radiographic studies were either missed or underreported.
A few research studies have exposed the potential utility of molecular imaging in
spine care. For example, in vivo quantitative voxel-based mapping is used to evaluate the
microstructural and molecular attributes of degenerative intervertebral discs and of the
spinal cord in cervical spondylotic myelopathy (Grabhar et al. 2015; Grunert et al., 2014).
81
AI could be used to help identify the presence of subtle spine pathology that may be
missed during non-spinal imaging studies, which include spine data in the background.
Most diseases progress through an asymptomatic (subclinical) period. For
example, early spinal cord compromise (myelopathy) secondary to degenerative stenosis
and compression often results in asymptomatic changes in regional biochemistry, blood
flow, and tissue architecture preceding the onset of clinical signs and symptoms (Durrant
& True, 2012). These changes are often not evident on routine imaging studies. The
visual presence of pathology within the spinal cord on advanced imaging studies is often
associated with end-stage pathology and permanent neurological deficits. Successful
detection of early stage myelopathy will require non-visual analysis of data acquired form
the spinal cord with molecular imaging and AI-supported radiomics. Researchers have
begun to address the possibilities. For example, a specialized form of MRI referred to as
diffusion tensor imaging (DTI) combined with machine learning classifiers has been
successfully used to detect early stage spinal cord compromise secondary to degenerative
narrowing of the central spinal canal in the neck, a condition referred to as cervical
spondylotic myelopathy (Wang et al., 2015; Wang, Hu, Shen, & Li, 2018). The concept
of discovery (screening) radiomics can applied to any bodily region including the spine
and spinal cord.
Computational AI approaches have been used to auto classify intervertebral discs
as either normal or degenerative based upon the analysis of tissue features such as signal
intensity, texture, and shape (Ghosh & Chaudhary, 2014; Oktay, Albayrak, & Akgul,
2014). In another study, AI assisted the automated detection and characterization of
82
lumbar neuroforaminal stenosis on MRI studies with greater than 90% accuracy (Han et
al., 2018). Further research is required to determine how AI-supported methods could
augment the role of the radiologist during the interpretive stage of spine imaging. Further
research on the use of AI in radiology will shape the future spine care.
Spine Pathology and AI: Meaningful Use Considerations
The capabilities of deep learning AI systems are highly dependent on exposure to
adequate levels of annotated training data and the establishment of ground truth. The
process is often expensive, tedious, and time-consuming. Due to this level of
commitment, it is imperative for stakeholders, including clinicians and radiologists, to
help identify meaningful use applications prior to investing in the process. Meaningful
use in this context refers to the application of an AI solution to address a condition or
disease which is prevalent, not adequately assessed with normal methods and which has a
significant impact on the individual and society. The institutional definition of
meaningful use applications in this context may include economic value assigned to new
solutions. An example of successful AI application with a favorable outcome is referred
to as a meaningful use case.
Some spine disorders, if left undetected and untreated, can lead to devastating
consequences that place an unnecessary burden on the individual, their family, and
society. The evaluation of AI use in spine imaging should start with the most devastating,
prevalent, and costly spine disorders. Examples include the 550,000 to 700,000 vertebral
compression fractures which occur annually in the United States secondary to
osteoporosis (Kondo, 2008) and the two-thirds of patients with cancer who will develop
83
bone metastasis, with the spine representing the most common location (Maccauro et al.,
2011). Intervertebral disc degeneration represents one of the most common causes of
back pain, which afflicts approximately 80% of the adult population throughout their
lifetime (Suthar et al., 2015). The annual prevalence of spinal cord compromise
(myelopathy) secondary to degenerative changes is estimated to be about 196,000 in
North America (Nouri et al., 2015). Each of the above conditions occurs in well-defined
anatomic regions of the spine, which can be interrogated in vivo at multidimensional
levels using emerging AI-supported methods
Some of the most easily segmented and labeled anatomic regions of the spine
house some of the most devastating types of pathology. These regions include the bone
marrow microenvironment within the vertebral body and the spinal cord
microenvironment within the spinal canal. In each case with the exception of trauma,
severe pathology begins with nonvisible, asymptomatic, tissue changes. Early detection
and intervention may lead to better patient outcome. AI-supported interrogation of
vulnerable anatomic regions could help detect and characterize aggressive pathology and
lead to early-personalized intervention.
The bone marrow microenvironment within the vertebral body is involved in
many different disease processes. For example, changes within the bone marrow often
precede the development of degenerative disc disease, as well as the development of
compression deformities and fractures. The bone marrow space is also a common
location for metastatic disease. Approximately two thirds of individuals with cancer will
develop bone metastasis with the spine representing the most common site (Maccauro et
84
al., 2011). Micro metastasis or early metastatic disease is often subclinical and difficult to
detect with traditional anatomic imaging methods. Approximately 36% of vertebral
lesions associated with spine metastasis are asymptomatic and discovered incidentally on
spine imaging for other disorders (Maccuaro et al., 2011). In some cases evidence of
metastasis to vertebral bone marrow may represent the first indicator that cancer is
present somewhere else.
The vertebral bodies and bone marrow are well visualized on routine spine MRI
and CT studies; therefore, rendering it possible for AI supported screening methods to be
applied to detect subtle or early stage pathology. Prior to implementing AI supported
solutions such as radiomics ground truth must be established for normal and abnormal
states. Discovery radiomics and in vivo interrogation of vertebral body
microenvironments may help reveal early stage osteoporosis, micro metastatic disease,
and subchondral degenerative changes that often precede the development of
intervertebral disc disease. Early detection may lead to early intervention.
Some advanced imaging methods and protocols have demonstrated the ability to
reveal micro pathology within the bone marrow of the spine (Long, Yablon, & Eisenberg,
2010; Park et al., 2015). AI-supported radiomic methods have demonstrated the ability to
detect nonvisible evidence of pathology in other fields. Its potential role in spine care
must be investigated further. In the future, radiomic methods will likely detect and
characterize vertebral bone marrow pathology such as myeloproliferative disorders,
subchondral degeneration, osteonecrosis, infection, and tumor (Long et al., 2010). It is
85
imperative to diagnose pathology within the vertebral bone marrow environment at the
earliest stage possible.
The first step in pathology detection using AI methods is auto assessment and
labeling of anatomic structures. The second step is to segment a region of interest. The
third step is to apply specialized AI applications such as radiomic methods to detect and
characterize pathology. AI has already been successfully used to perform automated
vertebral boundary detection and surface texture analysis using advanced context-
encoding features (Mirzaalian et al., 2013). This provides the digital framing and
segmentation required to isolate bone marrow. With the exception of one study, I was
unable to identify any significant research surrounding the use of radiomic methods to
evaluate tissues of the spine or spine pathology. In the referenced study, quantitative
voxel-based feature detection and analysis was performed using computed tomography
studies of the spine to assess true marrow space, fat composition, and mineral-based
marrow density (Pena et al., 2016). Pena et al. (2016) found that radiomic methods offer
advanced tissue differentiation and characterization that could be applied to the spine.
The same spine disease or disorder within different individuals may vary in
presentation, severity, and progression due to molecular, genetic, and structural diversity.
For this reason, a one-size-fits-all approach to spine imaging diagnosis or to treatment
may not work. AI-supported methods such a radiomics can help classify and stratify
disease and therefore help identify the best-personalized treatment plan. This research
study may represent one of the first to address the potential impact of AI in spine imaging
86
along with the potential role of radiomics and the virtual biopsy for the in vivo
interrogation of spine pathology.
Combinatorial Evolution of Artificial Intelligence
The pace of technological development and the rapid evolution of AI will
continue to increase. As technology advances the duration between innovations and new
applications often becomes shorter. In radiology, the phenomenon is amplified by co-
evolution of integrated technologies such as computer processing, imaging modalities,
database networking, and AI workflow solutions. As AI provides better analytic insights
and decision support there will be a greater push for more advanced imaging technology,
further complicating the decision-making process. Once human limitations are overcome,
there will a rising demand for the acquisition and analysis of more complex data to
support more precise and personalized care. Limited decision support for complex data
was previously a barrier to technology development and technology adoption in
radiology.
Perpetual revisions of disease criterion and classifications combined with ongoing
disease biomarker discovery will result in recurrent cycles of disruption and adaptation in
radiology and related clinical workflow. The three principal influences of an innovation
are differentiation, precision, and speed. Each of these factors will be pushed to the limits
in radiology by AI. There is a recursive relationship between technology, data, and
human contributions, all which collectively influence diagnostic decisions and the
delivery of care. The co-evolution of technologies at the radiology workstation
surrounding the management of high-throughput data will lead to unprecedented methods
87
of pathology assessment. AI and related technology support will lead to expanded
knowledge of disease pathophysiology and new standards of evidence-based care.
Exploratory research methods will help reveal the potential use of AI in radiology and
heighten awareness of and readiness for what is to come.
Imaging With AI: The Key to Precision and Collaborative Spine Care
Using big data and AI will transform the entire field of health care (Obermeyer &
Emanuel, 2016), including spine care. The use of AI solutions in radiology will forever
change how decisions are made in all health care specialties (Brink et al., 2017; Jha &
Topol, 2016; Jiang et al., 2017; Ranschaert, 2016). The unprecedented paradigm shift
associated with molecular level diagnostics and AI decision support will include spine
care. Personalized spine care requires identification of biological differences and
pathological variability, topics AI can help address.
Various forms of AI have the ability to improve the precision of the diagnostic
process in radiology by revealing unique attributes of disease (Syed-Mahmood, 2018;
Wang et al., 2015; Wang, et al., 2018; Zherhoni, 2017). For this reason, successful
applications of AI in radiology can have a favorable impact on the delivery of predictive,
pre-emptive and personalized health care (Augimeri et al., 2016; Brink et al., 2017; Jha &
Topol, 2016; Jiang et al., 2017; Ranschaert, 2016). The spine is intricate and complex.
The individual spine has many unique structural and biomechanical attributes that
influence the impact of disease. Greater knowledge of the unique attributes of an
individual’s spine pathology will inform all members of the spine care team. It will
88
support more precise and personalized spine care. Fundamental shared knowledge will
also facilitate more effective collaborative multidisciplinary care.
Growing use of AI in spine care will result in greater appreciation for the
spectrum of pathology, and for the need for a more timely and precise diagnosis. Future
AI applications in spine care will also expose common ground, redefine expert
boundaries, dampen existing professional turf wars, and lead to a few new ones.
Successful application of AI during interpretive stage of radiology workflow in any
specialty field of health care has the potential to transform the delivery of care while
supporting a more precise and personalized diagnostic process (Acharya et al., 2018;
Ghasemi et al., 2016; Hillman & Goldsmith, 2011; Jha & Topol, 2016). In addition to
improving the accuracy of the diagnostic process, AI may serve as a catalyst to new
levels of multidisciplinary collaboration in spine care. Some of the current applications of
AI in oncology, neuroradiology, and breast imaging will likely be adapted for use in
spine imaging.
A better understanding of the molecular basis of spine pathology using deep
learning and radiomic methods will lead to earlier detection and more precise subtyping
of pathology. This process will have an impact on the standard of care and will
perpetually reset clinical expectations. With heightened awareness of available decision
support radiologists, spine care providers, and the public will become less tolerant of
complacency and errors made during the interpretation of spine imaging. New
expectations and standards in spine imaging interpretation and reporting will have a
significant impact on evidence-based spine care.
89
The potential benefits of multidisciplinary care are well established. The success
of intervention is influenced by greater awareness of the heterogeneity of pathology,
acceptance of new disease criterion and classifications, and evolving standards in
decision support (Aerts, 2016; Collins & Varmus, 2015; Hood & Aufrey, 2013).
Collaborative spine care will become progressively more dependent on the central role of
the radiologist as a gatekeeper of big data and as a clinical consultant. The radiologist’s
role and impact is strengthened by improved diagnostic methods and better decision
support (Castenda, et al., 2015; Kressel, 2017). A more precise diagnostic process during
imaging workflow will help expose the fundamental basis of disease and subsequently
help bridge-the-gap between disciplines working at different points along the spectrum of
pathology (Brink et al., 2017; Kressel, 2017; Jiang et al., 2017). The collective use of
human and machine intelligence during the interpretive stage of spine imaging workflow
may help resolve multidisciplinary discordance by democratizing decision support and by
exposing standardized terminology and disease criterion.
Ethical Issues Surrounding the Use of AI in Radiology
Numerous ethical considerations surround the use of AI and big data in radiology
that could influence patient care (Mittelstadt & Floridi, 2016). Expanded use of AI will
subsequently require the creation of new ethical standards surrounding (Mesko, 2017).
Lee et al. (2017) raised concerns about potential challenges associated with deep learning
in radiology. The challenges included a growing dependency on large volumes of training
data and the black box nature of the technology. The later refers to the lack of
transparency of computational methods used to provide decision-support. The steps
90
associated with deep learning analysis of complex nested layers of computational
processes hidden. The lack of transparency limits validation and may lead to
unprecedented categories of liability.
Challenges associated with rapid adoption and evolution of AI in radiology
includes the possibility of over fitting results due to innovation bias, unnecessary hype,
and unrealistic expectations. Misdirected hype surrounding AI use in radiology could
have many direct and indirect adverse consequences on a health care system, such as
draining capital, unproductive reassignment of expertise, and the generation of false
expectations. An AI supported diagnostic process restricted to data mining can lead to
unproductive steps in workflow. Thus, Mayo et al. (2016) introduced the concept of
“farming of data,” in contrast to mining of data (p. 261). Data farming is characterized by
actively harvesting necessary data, locating missing data, picking the best data, and
weeding out unnecessary data. Ethical matters must be addressed to improve meaningful
use and the clinical utility of AI.
Widespread adoption of AI could result in deskilling of health care providers
including radiologists. In addition, the use of AI may result in a growing level of
technology codependency, thus, diminishing human influence in decision making.
Widespread AI adoption could also introduce automation and/or technology bias. Earlier
detection of pathology could result in unnecessary testing and treatment exposure. Access
to AI enhances authority and expertise and will afford the user with an advantage that can
have a significant impact on leadership roles, collaborative efforts, and the equality of
care. For the reasons stated, it is necessary to perpetually explore and address ethical
91
considerations that may arise because of the development and use of AI during the
interpretive stage of spine imaging workflow.
The Research Approaches
Researchers in the disciplines of computer science and radiology have approached
the potential role of AI in radiology from numerous perspectives. The approaches have
identified what may be missed with traditional anatomic imaging and qualitative
interpretation. Additional strengths are related to the investigation of narrow (disease
feature specific) applications of AI using methods such as radiomics and natural language
processing in isolated fields such as oncology. The weakness of this approach is the
limited knowledge acquired surrounding proposed AI, its interoperability with current
workflow, clinical utility, and ease-of-use.
There have been a limited number of qualitative exploratory studies on the
potential role of AI during the interpretive stage of radiology. The absence of qualitative
insights has resulted in numerous reductionist approaches to research on the topic with
limited capacity for generalization and practical clinical application. Many of the research
studies referenced in this work have not led to the development of an adequate concept
map or blueprint of the sequence of research required to address the potential role and
impact of AI on the interpretation of imaging studies and on the final reporting process.
Exploratory research will help identify needs and potential and reveal the sequence of
research studies and methodologies required to achieve desired results.
92
The Literature Gap
The literature clearly establishes numerous variables in radiology, including spine
imaging, which complicate the interpretive and diagnostic process. This includes the
growing burden of complex data, human bias, individual biological variability,
heterogeneity of pathology, and the multifocal nature of spine pathology. Additional
challenges associated with the use of AI and radiomic methods in spine care include the
intricacy of structures, proximity of anatomic elements, and limited access to relevant
databases and computational disease models. Technological variables include the lack of
standards and wide range of differences between imaging modalities and protocols. The
literature review revealed an absence of gold standards in radiomics. The potential role
and impacts of radiomics is underexplored in all imaging specialties (Oakden-Rayner et
al., 2017). An exhaustive literature search revealed a growing number of research studies
designed to address the role of AI decision support in radiology. The search revealed
some of the challenges associated with imaging interpretation. Many of the published
research studies address narrow applications of AI and therefore do not address the
challenges associated with its adoption, consistent use, and support.
The diagnostic process in spine care is primarily limited to the history, physical
examination, electrodiagnostic testing, and diagnostic imaging. The literature search
established the absence of reliable serum biomarkers for confirming the presence or
progression of spine disorders. The search also confirmed that traditional needle biopsies
are rarely performed on spine pathology. Neurologic deficits associated with a spine
disorder often represent end-stage pathology and a certain amount of permanency is
93
likely. For the reasons stated spine care providers of all disciplines are highly dependent
on diagnostic imaging reports for decision support at the point of care. Improved delivery
of spine care will subsequently require the use of advanced imaging and related decision
support for the radiologist resulting in more precise and personalized reporting.
The literature search and review performed for this study supported the need to
improve the differential diagnostic process during the interpretive stage of spine imaging
workflow. Image interpretation and reporting in other fields has also been described as
incomplete, inconsistent, and inconclusive (Bosmans et al., 2011; J. Y. Chen et al., 2017).
Wu et al. (2018) discussed the “unmet need for methods that allow more comprehensive
disease characterization and reliable prediction or early assessment of treatment response
and prognosis toward the goal of personalized or precision medicine” (p. 125). Spine
disorders represent one of the most common causes of pain and disability; therefore,
radiologist’s should do what is necessary during the interpretive stage of radiology
workflow to improve the accuracy of the diagnostic process and help ensure successful
delivery of personalized spine care.
AI offers potential solutions for spine imaging, although, little attention has been
paid to its potential. The use of AI in radiology has generally been limited and slow due
to the challenges associated with its development and validation (Aerts, 2016). I was
unable to identify any scholarly research articles that addressed the potential applications
or impacts of the digital (virtual) biopsy in spine care. Knowledge about how to
implement radiomic measures into routine radiology practice is also limited (Vallieres et
al., 2017). The potential role of radiomics needs to be addressed within the context of
94
spine imaging. The literature reveals many of the challenges associated with its use in
other specialties including lack of standardized imaging protocol, limited access to
annotation training datasets, defined meaningful use applications, clinical utility, and
contouring regions of interest (Lai-Kwon, Siva, Lewin, 2018). The literature review
revealed successful use of AI in many areas of non-spine imaging. Established success,
although limited, has involved the use of natural language processing, radiomics, and
computational diagnostics. The success of AI solutions in radiology requires scalable
applications, clinical utility, and seamless integration into radiology workflow (Court et
al., 2016; Syeda-Mahmood, 2018). There is a gap in the literature surrounding the use
and potential use of AI and AI-supported methods during spine imaging workflow. This
includes the topics of NLP and radiomics.
The role of AI in some fields of radiology such as oncology has advanced more
than in spine imaging. Published research surrounding the role of AI use in other
subspecialty fields of radiology provide the foundation for the discussion of its potential
role in spine care and the design of this research study. The limited research on spine
imaging has been associated primarily with automated identification of normal anatomy
rather than addressing the detection and characterization of disease states. A published
letter in 2016 representing the position of leaders of the American College of Radiology,
a cultural authority in the field of radiology addressed the general gaps in knowledge and
research surrounding the use of AI in radiology. Written to the U.S. Office of Science and
Technology Policy, the authors summarized the gaps in the literature with
acknowledgment of the need for further research to help identify how AI can be used to
95
access meaningful data, enhance the interpretive phase of radiology workflow, improve
diagnostic accuracy, and reduce errors. This request applies to all areas of radiology
including spine imaging. As noted prior, the results of the literature search confirm the
paucity of studies addressing gaps in knowledge regarding the role of AI in spine
imaging.
Summary
The growing appreciation for the heterogeneity and complexity of pathology has
led to the realization that more precise personalized care is possible with the right
decision support. To achieve this goal, the diagnostic process must be more
comprehensive and classifications of pathology expanded. The standards in radiology
must change. The interpretive approach can no longer be limited to one individual’s
visual assessment of overwhelming volumes of two-dimensional images and the
generation of highly variable qualitative reports. The data are too complex and the stakes
are too high. Automated methods of disease detection and characterization are required to
augment the role of the radiologist. The clinical utility associated with the adoption of AI
solutions in radiology, and, more specifically spine care, must be determined. Timely
personalized patient care should take priority.
An extensive scholarly literature review revealed that AI systems are capable of
integrating and analyzing structured and unstructured data to refine the diagnostic
process. The AI methods required to accomplish this goal include quantitative imaging
with feature analysis (radiomics), acquisition analysis of qualitative imaging features
from records (natural language processing), unsupervised feature learning (text and
96
images), and scaling of AI solutions to accommodate multiplatform data (distributed
computational models). Successful integration of AI solutions during the interpretive
stage of spine imaging workflow will reduce errors and result in unprecedented
diagnostic capabilities. AI can provide new perspectives of disease that will lead to more
efficient and effective care. Success requires that gatekeepers of big data, such as
radiologists, must accept new responsibilities, assume new roles, and embrace AI
decision support.
This study focused on the potential role and impacts of AI applications during the
interpretive stage of spine imaging. The results of the literature review served as critical
determinants of the potential applications of AI in spine care. The literature review
established the benefits of using AI supported method such as radiomics, natural
language processing, and computational disease modeling in other fields of radiology to
achieve a more precise probability-based diagnosis. It is evident based upon an extensive
review of the literature that in the near future, the interpretive stage of spine imaging will
likely rely on the use of collective intelligence derived from the integration of human and
machine intelligence. This research study was designed to help determine how and when
this might occur.
The published research suggests the radiology workstation will become a hub of
convergent information from other patient diagnostic procedures and databases, including
laboratory, genetic and pathology test results. Databases will include accessible
computational disease models and disease registries. Integrating AI and radiomic
methods will fundamentally alter how disease is diagnosed, classified, and treated (Langs
97
et al., 2018; Shaikh et al., 2018). Further development of radiomic methods will support
applications for the assessment of non-cancer-related spine pathology.
Current research acknowledges that emerging AI solutions are capable of
providing radiologists with contextually relevant and probability-based differential
diagnostic considerations during the interpretive stage of imaging workflow. This process
improves diagnostic precision and therefore can help overcome human limitations and
bias. In the future, whoever has access to the best data and the best decision support will
likely provide the best care.
Ongoing advances in diagnostic imaging will continue to challenge and expand
our current understanding of disease and related diagnostic criterion. The differential
diagnostic process will soon no longer be limited to the expertise and skills of an
individual; instead, a whole systems process will involve collective intelligence derived
from the contributions of humans and machines. For this outcome to be possible, many
unknown factors must be addressed. This includes what defines meaningful use and
adequate training of an AI system. Heightened awareness of improved accuracy and
efficiency associated with AI decision support will drive computational analytics and
radiomic methods to the forefront of radiology, supporting their eventual role as routine
procedures. Multidisciplinary research will lay the foundation and pave the way for the
transformative process.
Heightened awareness of AI potential and clinical utility in spine imaging is
required to further the research and development process. The endeavor will require the
insights and participation of numerous experts such as AI developers, physicists,
98
radiologists, pathologists, key influencers, and early adopters of AI in radiology. The
establishment of expert opinions and predictions can help direct further discussion and
research of the topic. A well-designed exploratory case studies can provide this necessary
foundation.
Chapter 3 introduces the research design and methodology used to explore the
potential impacts of AI on the interpretive stage of spine imaging and on the differential
diagnostic process. The results of the extensive literature search reported in this chapter
are used to support the chosen research methodology and strategies acknowledged in
Chapter 3. The chapter addresses my role as a researcher, as well as the data acquisition
and data analysis strategies used in the study. Special attention is placed on the
implementation of steps to protect research participants and to improve the credibility and
trustworthiness of the study.
99
Chapter 3: Research Method
Introduction
The primary goal of this study was to establish the potential impact of AI
solutions on spine imaging interpretation and diagnosis. I placed special emphasis on the
potential role of radiomics. The unit of study was the interpretive stage of spine imaging
workflow used to detect, characterize, and monitor pathology. The sources of data
included document review, reflective journaling, and focus group sessions. Focus groups
are an effective method for exploring attitudes, expectations, and potential applications
associated with emerging technologies and related processes (Kitzinger, 1995).
In Chapter 3 I introduce the research design and methods used to address the topic
of study. In this discourse I address my role as a researcher and the role of research
participants. I also address the methods used for data acquisition and analysis. The
chapter indicates the research steps implemented to improve the trustworthiness of the
study, as well as the processes used to help ensure the ethical treatment of participants
and the ethical management of data.
Research Design and Rationale
I used a qualitative exploratory case study design to investigate the potential
impacts of AI on the interpretive stage of spine imaging workflow. This approach offered
a flexible and inductive method for acquiring holistic and in-depth insight. The primary
purpose of qualitative exploratory research is to reveal the potential contributions and
influence of an emerging technology or process (Baxter & Jack, 2008). My chosen
research design was used to address how and why questions surrounding the potential use
100
AI applications such as radiomics, NLP, and diagnostic inference methods using deep
learning approaches. Exploratory approaches are often used to lay the foundation for
additional methods of inquiry such as quantitative and mixed method research (Creswell,
2013; Patton, 1990). AI represents a bridging technology comprised of an assemblage of
evolving elements and processes that are sometimes difficult to identify and assess. This
scenario contributes to complexity and uncertainty in research. A qualitative exploratory
approach is able to reveal contextual relationships not adequately addressed by more
restrictive explanatory or quantitative research methods (Ponelis, 2015; Yin, 1984). It
was necessary for me to offer an inductive contextual perspective of potential AI
applications which might be of value to radiologists and other stakeholders in the field.
The chosen study design supported the triangulation of qualitative data acquired
from numerous sources, which helped to improve study validity and trustworthiness. The
qualitative research approach supported purposive sampling, the acquisition of expert
insight from different sources, inductive investigation, and the ability to formulate a
contextual narrative summary. The primary research question was: What are the opinions
of experts regarding the potential use and impact of AI during the interpretive stage of
spine imaging workflow? I added supportive research questions to address determinants
of AI adoption and various applications of AI such as radiomics and natural language
processing.
I acquired qualitative data from expert documents in the form of white papers
published by thought leaders and radiology organizations which addressed the evolution
of AI in radiology. The documents I used represented consensus opinions on the use and
101
potential use of AI. I framed the method of inquiry with insight acquired from an
extensive literature search and from perspectives offered by a consensus-based
prospective document prepared by the Spinecare Data Science Committee of the
American Academy of Spine Physicians (AASP). The committee provided a list of high-
priority topics (needs analysis) related to the potential use of AI during the interpretive
stage of spine imaging. I considered the proposed topics during my development of
research strategies and in my preparation for the focus group sessions.
Qualitative data acquisition occurred in the following sequential stages: an
extensive literature review, review of a prospective document from the AASP Spinecare
Data Science Committee, review of expert documents, and the use of two focus group
sessions, one consisting of radiologists and the other AI experts. I performed reflective
journaling during the entire data acquisition and data analysis process. I used focus
groups to acquire expert knowledge, opinions, perceptions, and predictions relevant to the
topic of study. The research process was designed to reveal themes, noteworthy quotes,
and new perspectives surrounding the use of AI during spine imaging interpretation. Prior
research had established that exploratory focus groups can be used to identify attitudes,
discover opportunities, generate ideas, and frame new questions for future inquiry (Breen,
2006). I used focus group sessions to facilitate creative discussion and to expose the
potential benefits associated with AI use at the radiology workstation in spine care. The
process revealed new insights and exposed culturally formed attitudes and opinions
surrounding the potential applications of AI. The insights I acquired from the focus group
sessions helped me predict the type of synergies, controversies, and debates which may
102
arise surrounding this topic in other research settings. The design of this research study
supported the transition from shared experiences and insights to higher levels of
abstraction and application.
Role of the Researcher
The role of the researcher is important in any research study but particularly
important in qualitative exploratory studies because of the subjective nature of the
process. Researcher experience, motives, and bias can all have a significant impact on the
research process including the analysis and interpretation of acquired data (Durdella,
2019). A qualitative researcher often assumes a primary role in the acquisition and
analysis of data. I assumed the role of the sole researcher in this study. As the sole
researcher, I represented the primary instrument for the collection and analysis of data.
My clinical experience combined with my desire to help identify new forms of decision
support in spine imaging motivated me to pursue the topic of study. It also increased the
risk for professional bias during the research. I subsequently implemented numerous
steps in the research process to reduce my potential for introducing bias into the study.
The steps taken to reduce my personal or professional (researcher) bias and to improve
the trustworthiness of the study included the use of independent experts to review the
focus group moderator guide, the use of member checking (respondent validation), the
application of within group and between group analysis, and triangulation of data. Field
testing of focus group protocols and related research questions was performed to help
establish credibility and the relevance of the approach. I used reflective journaling to help
103
reveal my biases, thoughts, and opinions throughout the research process and provide the
basis for some of the decisions I made during the research process.
Prior to this study I had extensive clinical experience in diagnostic neurology,
neuroradiology, and spine care. I also had extensive academic training in AI and
radiomics. My professional experience as a clinician combined with my familiarity with
AI provided me with the insights required to develop and implement effective
exploratory and analytic strategies. Prior to and during the study I prioritized conducting
myself and the research process in accordance with acceptable scientific methods and in
accordance with the Walden University Institutional Review Board (IRB) guidelines. I
implemented numerous methods to help support scientific data analysis the disclosure of
research conclusions in an unbiased and objective manner. I implemented the previously
disclosed strategies to help ensure that I was reflective and transparent throughout the
research process.
Personal and Professional Relationships
Prior to or during the course of this research study I did not have any formal
business relationship with IBM, any other AI-related company, or professionals who
participated in the focus groups sessions. I did not pursue or accept research participants
with whom I had any prior business relationship. During the proposal stage of the
dissertation process I participated in numerous conference calls with IBM staff including
data scientists to discuss gaps in the research, research strategies, and the management of
research related data. During the early stage of the dissertation process I used a key
contact from IBM to help identify a few renowned AI experts who met the study
104
inclusion and exclusion criteria. I chose to approach IBM due to their performance
record, current market position, and their potential for developing AI solutions for
radiology.
Researcher Bias
Among the many potential sources of bias in a qualitative exploratory research
study, some involve the researcher. Bias can occur in many forms and can influence
different phases of the research process such as the development of research questions,
participant recruitment, expert interviews, data acquisition, and data analysis (Creswell,
2013). Potential sources of bias in this study include the effects of the researcher on the
study and the effects of the research process on the researcher. Researcher bias is possible
whenever research relationships could lead to future business opportunities. For this
reason, I did not accept any proposals or entertain discussions about potential future
relationships.
I was not offered any position with research participants and/or companies or
institutions they were been affiliated with prior to or during the course of the research
study. I pursued the research topic and study with bias in favor of the eventual use of AI
solutions to improve diagnostic accuracy and the interpretive diagnostic process in spine
imaging. I fully disclose that I do not fully understand how AI could or should be used.
As a practicing neurologist, I acknowledge that the accuracy of the diagnostic process in
spine imaging must be improved. Approximately two years ago I sat at a prototypical
IBM Watson AI workstation and experience its potential contribution to the differential
diagnostic process in radiology. With the exception of the isolated experience with IBM
105
Watson, I have no other practical hands-on experience with the use of AI in radiology or
spine care. I am therefore not biased toward the adoption of AI in radiology based on
personal hands-on experience.
Ethical Issues Surrounding the Researcher
As the researcher in the study, I anticipated and addressed potential ethical issues
that may have arisen prior to, during, or after the research process. Prior to designing the
study, I became familiar with the Belmont Report (1979) published by the U.S.
Department of Health, Education, and Welfare, which acknowledged ethical principles
and guidelines which can be used to protect human subjects while conducting research. In
addition, during the course of my PhD studies at Walden University, I completed the
National Institutes of Health (NIH) web-based training program titled “Protection of
Human Research Participants” (Certificate # 2872343). I assumed the duty as the primary
researcher in this study to handle myself in a scholarly fashion and to treat all research
participants in a professional and ethical manner. This required the implementation of
steps to ensure participants well-being while protecting their rights and minimizing their
exposure to potential harm.
Research Methods
The dissertation proposal was accepted and Walden University IRB approval was
obtained prior to beginning the research process. The IRB reviewed the research plan and
the research methodology, as well as all pertinent documents. The approval process was
completed prior to research participant recruitment and the acquisition of data. I
implemented steps to disclose and address researcher bias, potential conflicts of interest,
106
competing motives, and potentially detrimental power relationships that could have
developed during the course of the study.
Study Population and Sampling
The study population consisted of stakeholders involved in or influenced by the
interpretation of spine imaging. This includes data scientists, device manufacturers, AI
programmers, radiologists, spine care providers, and other health care providers. It was
important to identify the subpopulation of stakeholders most capable of addressing the
exploratory research topic within focus group settings. The success of this study
depended on my ability to recruit research participants experienced and knowledgeable
on topics related to the potential impact of AI used during the interpretive stage of spine
imaging workflow. The focus group study population was limited to radiologists and AI
experts. The members of each category of participants were intricately involved in
processes which took place within the parameters of the unit of study, which was the
interpretive stage of spine imaging. I implemented steps that required that all research
participants met strict research inclusion and exclusion criteria.
A certain degree of homogeneity or similarity within a focus group session
combined with purposive sampling of professional participants has proven to enhance the
potential for exploring new technology (Kitzinger, 1995). Kitzinger (1995) also
demonstrated that a group of professionals with common knowledge along with similar
training and experience are more likely to engage in in-depth discussions. To facilitate
this approach, I placed AI experts in one focus group session and radiologists in a distinct
and separate focus group.
107
Population sampling for the research study was convenient and purposive so that
small groups of confirmed experts could discuss and explore the research topic. Experts
recommend purposive sampling strategy to help ensure that participants have the level of
experience and expertise required to contribute to the topic of study in group sessions
(Creswell, 2012; Patton, 2002). Subsequently, I used purposive sampling in this study to
help ensure that all of the research participants had an adequate level of expertise,
interest, and experience surrounding the development or application of AI solutions
during the interpretive stage of diagnostic imaging. The use of convenience sampling
combined with purposeful sampling brought together like-minded professionals who
were familiar with the topic of study and categorically with role of AI experts and
radiologists in spine care.
The population sampling strategy included estimation of the sample size required
to achieve the level of representation and topic saturation required to explore the potential
impact of AI on spine imaging interpretation and diagnosis. A renowned key AI contact
was used to help identify qualified AI professionals to participate in the study. I used key
radiology contacts to help identify radiologists qualified for participation in this study. I
provided each of my contacts with the purpose of the research study, as well as
participant inclusion and exclusion participation criterion prior to asking for their
recommendations. I confirmed that all potential participants met inclusion and exclusion
criterion prior to their acceptance into the study.
The ideal size of a focus group is often five to eight participants, unless more are
required to address a complex topic (Krueger & Casey, 2015). Qualitative research
108
experts have acknowledged that four to 10 individuals are often adequate for
homogenous sampling of professionals (Creswell, 2013; Krueger & Casey, 2015). I
limited the number of expert participants to eight to 12, or four to six in each of the two
focus group sessions, to facilitate creative and in-depth discussion of a complex and
contemporary topic. The small focus group sizes helped me as the moderator better
manage research topics and the flow of discussion and ensure that each participant had
adequate opportunities to share their expertise and insights. The use of two homogenous
focus groups comprised of confirmed experts was large enough for this study to provide a
diversity of perspectives and opinions.
The research participants accepted for participation in this study were from
different clinical and professional settings. I recruited research participants through
personal contacts. I followed up with potential participants by of email and phone calls. I
did not offer material or monetary incentives during the recruitment process. Research
subjects who agreed to participate in a focus group session were asked to complete a brief
survey prior to the focus group session (Appendix B). I used a survey to acquire
participant demographic information such as their current position, background, and
experience.
Prior to convening for focus group sessions each research participant received an
acceptance letter which included an introduction to the research project, a focus group
agenda, a list of what was expected of them and a research participation consent form to
review, sign and return. Each research participant received email notification of the
scheduled focus group session along with an invitation to participate in person or via a
109
prearranged teleconference option. I arranged for the teleconference option through
Zoom, a highly respected and secure service. Email notifications included the focus
group facility address, focus group directives, and access information for the
teleconference option. Each participant received periodic email reminders of their
scheduled focus group session. I requested that research participants register online in
advance of the focus group sessions.
Sample size in qualitative research is often not as important as the chosen
methods of data acquisition, data analysis, and data validation (Njie & Asimiran, 2014).
The potential use of AI during the interpretive stage of spine imaging is both a new and
complex topic. Subsequently, I chose small sample sizes to help achieve expert in-depth
discussions and topic saturation during the focus group sessions. Data saturation is
reached in a qualitative study when a coherent and consistent perspective is reached
(Guest, Bounce, & Johnson, 2006). I established the criteria for data saturation in this
study prior to the focus group sessions and included the definition on in the moderator
guide (Appendix D). I developed open-ended and probing research questions to help
achieve data saturation during the focus group sessions.
Research Participant Inclusion and Exclusion Criteria
In qualitative research, research participants must be capable of contributing to
the study with the chosen methods of inquiry (Creswell, 2013), particularly when
addressing a complex and rapidly evolving technology such as AI. I used convenient and
purposeful sampling methods to select research participants for this study. The potential
research participants were subjected to explicit study exclusion and inclusion criteria. The
110
inclusion criterion addressed the potential participant’s current professional status,
background, and experience. The inclusion criterion for AI expert participation was a
minimum of five years of experience in health care AI development or applications. In
addition, each AI participant was required to have a minimum of a bachelor’s degree in a
related field such as informatics, data science, computer science, or AI. AI participants
were also required to be actively working in an AI field.
The inclusion criteria for radiologists were a minimum of 10 years of experience
in spine imaging interpretation. In addition, each participant was required to hold a
doctoral degree and to be actively working in a diagnostic radiology capacity. Each
radiologist was required to be board-certified in a field related to the topic of the study.
Study exclusion criteria for the AI expert and the radiologist included a history of or
current employment with the Chicago Neuroscience Institute (CNI) or the American
Academy of Spine Physicians (AASP), both of which I am affiliated with. I prohibited
key contacts for research participant recruitment from participating in the study. I also
prohibited professionals who played a role in field testing of focus group questions and
strategies from participation in the research study.
Data Collection Instruments and Processes
The use of predetermined or validated data collection instruments can help direct
the data acquisition and data management processes. I acquired and developed numerous
data collection instruments for use in this research study. This included the use of
qualitative research software, a brief qualitative survey, and a focus group moderator
guide. I developed a brief survey and gave it to each participant to complete prior to
111
participation in a focus group session (Appendix B). This document and the consent form
were used to confirm that the various experts met the criterion for participation. I
developed a focus group moderator guide to help manage time, topic discussions, and the
method of inquiry during the focus group sessions (Appendix F). Published research has
demonstrated that the use of a focus group moderator guide helps ensure efficient and
systematic in-depth coverage of research topics and related questions (Fraenkel &
Wallen, 2003). The moderator guide I used in this study consisted of carefully crafted
open-ended questions and probing semistructured questions.
I developed questions for the focus groups with the assistance of insight acquired
from an extensive literature search and from the needs analysis document provided by the
AASP Spinecare Data Science Committee (Appendix F). Each question was placed in the
focus group moderator guide. My development of the moderator guide was overseen by
independent expert prior to and after field testing. I made all necessary changes to the
guide. This iterative process of assessment helped me reduce the risk for an inappropriate
or biased approach to research question development and delivery. It also helped me
refine the methods of inquiry I used during each focus group session. The focus group
moderator guide consisted of an agenda, an introduction, along with a list of PowerPoint
concept slides, and research questions followed by closing remarks. (Appendix D). The
moderator guide identified the order of topic presentation and inquiry. I used the guide to
help set the tone for each focus group session and for guiding the order of the process.
Consistent with the recommendations of C. L. Lee et al. (2015) I developed data
coding guidelines to help ensure analytical and categorical consistency during content
112
and thematic analysis (Appendix C). My coding guidelines consisted of predetermined
codes capable of being adapted, modified or replaced during data analysis. I developed a
priori codes consistent with the conceptual framework of the study utilizing theoretical
perspectives from DOI and TAM. The a priori codes aligned with the research topic,
research purpose, and research questions. During the course of the entire research
process, I made regular entries in a reflective journal to memorialize my biases,
impressions, and insights.
I used numerous expert documents (white papers) to help identify current
consensus-based opinions and positions surrounding the use of AI in radiology and
oncology. At the time of this study, there were no white papers published on the potential
role or impacts of AI in spine care. I used a few published papers, which addressed
narrow applications of AI in spine care. The consensus-based “white papers” I used for
this study were published by nationally and internationally recognized organizations: the
Canadian Association of Radiology, the American College of Radiology, the French
Radiology Community, and the European Society of Radiology. I also used seminal
publications of leading experts in the field. I analyzed, thematically coded, and
triangulated the content of the expert documents with data from other sources to improve
the consistency and relevancy of the study’s conclusions. This methodical and transparent
analysis process helped improve the trustworthiness and validity of the study. The
document provided by the AASP Spinecare Data Science Committee offered a list of
potentially meaningful applications of AI during the interpretative stage of spine imaging
workflow. The AASP document was developed by a multidisciplinary group of spine
113
care experts independent of the research process. I used this document along with insight
acquired from an extensive literature search to guide the development of research
questions I used in the focus group sessions.
The research process was not be limited by fixed guidelines or rules, subsequently
allowing for inductive assessment of emerging topics and trends. As stated previously,
focus group research represents a well-established and disciplined scientific method for
acquiring in-depth insight surrounding the use of new technology and related processes
(Krueger & Casey, 2015). I applied the concept of data saturation during focus group
sessions and during thematic data analysis of expert documents. I used an introductory
PowerPoint slide program at the beginning of each focus group session (Appendix E). I
conducted topic-specific discussions during the focus group sessions until reasonable
topic saturation was achieved. I arranged for a recording of all of the contributions during
each focus group session. I also arranged for verbatim transcription of the recorded
sessions to avoid misinterpretation or misrepresentation. I performed document analysis
until I achieved topic saturation. I triangulated the data from the different research
sources to improve the internal, as well as external validity of the study. A concise and
comprehensive informed consent form was developed and used to protect the rights of
participants and to encourage unfettered contribution to the research process.
Document review and analysis offers a unique and often critical contribution to
qualitative exploratory research (Creswell, 2013; Patton, 1999). I used position papers
and consensus-based summaries published by reputable organizations and highly
regarded experts in this study to help address the potential impact of AI during spine
114
imaging workflow. I compared and contrasted the data acquired through expert
documents with data acquired through other expert sources such as the focus group
sessions and reflective journaling to improve the internal validity and transferability of
the results.
Data Acquisition
I initiated the data collection and analysis process after I received dissertation
proposal approval and Institutional Review Board (IRB) approval from Walden
University. I created a data acquisition flow diagram to help guide me in the research
process. The method of inquiry I used in the focus group session was field tested with
two independent experts, one meeting AI expert participation criteria and the other
radiologist criteria. I did not accept the experts who assisted me with field testing as
research participants. I used field testing to evaluate the focus group protocols and
strategies I used in the Focus Group Moderators Guide. I made minimal modifications, as
a result of the field testing.
The focus group sessions each lasted approximately 90 minutes. I achieved an
acceptable degree of data and topic saturation in each session. I arranged for each focus
group session to be digitally recorded, transcribed verbatim, and stored securely.
Emotional responses, body language, and nonverbal forms of communication between
the participants in a focus group setting can be important (Bunnick et al., 2017). I
subsequently recorded any participant behavior during the focus group sessions I felt was
relevant to the study purpose. I led each focus group session with the assistance of the
moderator guide. I developed my focus group approach guided by Krueger’s categorical
115
strategies that included the use of an opening question, introductory questions,
transitional questions, and probing questions (Krueger, 2000). I used probing questions to
help actively engage each of the participants in topic discussions.
Scholarly publications have established that reflective journaling offers the
researcher with a powerful inductive method for recording observations ideas, and
insights using an active voice (Janesick, 2011). I performed reflective journaling during
the course of the research process. Journaling served many purposes. It allowed me to
identify my initial and evolving perspectives, expectations, and biases associated with the
research topic and the research process. Review of journal entries gave me the
opportunity to engage a higher level of critical thinking and implement methods to reduce
my personal influence on the research process and outcome. Journaling included my
impressions of verbal, as well as nonverbal communication during focus group sessions.
This included recoding of body language and expressions. I also recorded the level and
nature of agreements or disagreements that occurred during each focus group session.
Research Participant Debriefing and Follow-Up
At the completion of each focus group session, I reminded participants of the
purpose of the research study and informed them how I would manage and analyze the
data acquired. In informed the research participants that they would receive an overview
of the focus group data analysis in the form of a thematic summary and a list of
supportive quotes for review, a process referred to as member checking or respondent
validation. In addition, I informed each research participant that the records of the
research study including their consent forms would be stored in a secure location for a
116
minimum of 5 years, after which time they would be properly destroyed. I informed each
participant they would receive notice when the dissertation was published. I also assured
each research participants that they would be provided with access to the published work
when it was available.
Data Analysis
I used a multistep process to analyze acquired data. I implemented an inductive
and iterative data analysis process as soon as data was acquired. My analysis process
continued throughout the entire research study. I recorded a chain of evidence to
memorialize the process and analyzed the focus group transcripts with an exhaustive,
inductive, and iterative process of coding for themes. Descriptive codes were clearly
established and defined consistent with the work of Glaser and Laudel (2013). I used a
hybrid approach to coding, allowing for aggregation, subtraction, combining, and
expanding of code categories when necessary. Qualitative data coding offers an effective
method for revealing emergent ideas, themes, and relationships (Rubin & Rubin, 1995;
Strauss & Corbin, 1998). My evaluation of the focus group transcripts included content
analysis and thematic coding. Content analysis is used in the social sciences and in
qualitative research to structure information (Krippendorf, 2004). The analysis of focus
group data should include identification of noteworthy quotes, as well as identification of
outlying factors and unexpected findings consistent with published works (Breen, 2006).
I identified and labeled all noteworthy findings, comments, and quotes derived from my
assessment of the various sources of research data. I performed the data coding process
until topic saturation was achieved.
117
The first step in the coding process was to become familiar with the data and
systematically reduce its complexity. I developed a few provisional (a priori) codes to
initiate axial coding. I developed the provisional codes with insights acquired from my
extensive literature search along with the influence of theoretical constructs from DOI
and TAM. Some of the provisional codes aligned with theoretical constructs of DOI and
TAM, such as relative advantage, interoperability, complexity, ease-of-use, and perceived
usefulness. I replaced, revised or modified many of the provisional codes during data
analysis to better describe and label acquired data. I expanded, contracted, replaced, and
modified the coding categories many times throughout my analysis process. Thematic
coding arose from the integrated applications of provisional coding, open coding, in vivo
coding, axial coding, and selective coding.
My analyses of focus group data included within and between group analysis. I
took into account the unique experiences and backgrounds of the participants in each
focus group session revealed by their demographic surveys. I displayed the results of my
analysis of the acquired research data in many different ways, including a contrast table.
Contrast tables offer an effective method for looking at relationships between exemplars,
extremes, and outliers (Miles, Huberman, & Saldana, 2014). Combining content analysis
and thematic coding supported the development of a concept map depicting the potential
relationships between processes and technologies, used during the interpretive stage of
radiology workflow. I created a concept map to reveal relationships between the flow of
data and processes during the interpretive stage of radiology workflow. I used the concept
118
map to help transform tacit knowledge into a practical resource and a foundation for
further discussion and research.
Concept Mapping
Data analysis in qualitative research involves many steps that include data
reduction, data organization, data interpretation, and data display. Concept mapping has
been successfully used to graphically organize and depict relationships between elements
of a system or process (Baugh, McNallen, & Frazelle, 2014). This includes the
relationships between data, individuals, and technology (Baugh et al., 2014). Concept
mapping has also been used in qualitative research to reveal themes and to depict
workflow (Daley, 2004; Novak, 1998). One of my goals in this research study was to
identify themes that could be used to create one or more concept maps depicting the
potential role of AI solutions during the interpretive stage of spine imaging workflow. A
concept map helps depict stages of a complex process and reveal technological
relationships to achieve desired goals (Daley, 2004; Novak, 1998; Wheeldon & Faubert,
2009). I subsequently developed a concept map to reveal the flow of data and role of
potential AI applications during the differential diagnostic process associated with
interpreting spine images.
I used concept mapping in this study to facilitate a shared vision, to help direct
subsequent research, and to inform further technology development. I also used it to help
determine how to embed AI technology into existing radiology workflow. A concept map
can be augmented with the use of numerous elements such as linked tasks, labeled
processes, communication pathways, and hierarchies of priority (Wheeldon & Faubert,
119
2009). In my concept map I used map lines to link mapped elements and I used
directional arrows to reflect the flow of data and/or the implementation of a process. I
developed a concept map in this study to complement and enhance textual conclusions.
This helped me present research findings in an accurate, concise, and effective manner.
Trustworthiness of the Study
The trustworthiness of a research process is influenced by the role of the
researcher, the source of the data, the management of the data, and the approaches used to
improve study validity and reproducibility (Connelly, 2016; Mays & Pope, 2000;
Shenton, 2004). Qualitative exploratory case study research is inductive and subjective
and therefore requires high levels of trustworthiness, reliability, and validity to be
influential (Creswell, 2013). The attributes of validity and reliability are operationalized
in different ways in qualitative versus quantitative research studies (Mays & Pope, 2000).
The primary risks associated with qualitative case study research include over
generalization of results, researcher bias, inadequate interpretation of data, poor
integration of data, and research question mismatch with methodology. I took extra
precautions and implemented steps throughout this research process to improve the
trustworthiness of the study and its conclusions (Figure 5).
Qualitative studies that include the use of focus group sessions must meet
extremely high standards to be reliable and valid. Research study trustworthiness is
determined by its credibility, confirmability, dependability, and transferability (Lincoln &
Guba, 1985; Shenton, 2004). I addressed each of these elements in this study along with
120
reliability. I implemented numerous steps d to reduce the risk for interjecting personal
bias and to support evidence-based conclusions.
Credibility
I implemented numerous steps to improve the credibility of this research study. I
used respondent validation also referred to as member checking to help confirm the
accuracy of my thematic conclusions. I also provided research participants with a list of
the supportive quotes I acquired from the focus group transcripts. Member checking is an
important step in qualitative research, because it provides participants with an
opportunity to affirm the accuracy of focus group data acquisition, analysis, and
interpretation (Creswell, 2013). Member checking in this study served as an effective
method for establishing interpretive and descriptive validity. It also helped reduce the
impact of my personal bias as the sole researcher.
To help further reduce personal bias during data acquisition and analysis, each
focus group session was recorded and transcribed verbatim. I used a few open-ended
questions during the focus group discussions to help reduce the risk of framing bias. I
read the focus group transcripts numerous times to ensure comprehension of the material
prior to initiating descriptive labeling and coding of data. I used an inductive and iterative
process of hierarchical coding to avoid rigid misclassification of data.
I used a well-defined unit of analysis to help direct the research process and the
flow of data. This approach improved study credibility. Consistent with the work of Mays
and Pope (2000) I performed reflective journaling to expose how my role as the sole
researcher may have influenced the research process and outcomes. I used reflective
121
journaling to record my thoughts and opinions during the entire research process. The
journaling process helped reveal my perspective and beliefs and how they may have had
an impact on data analysis, research design, and research conclusions.
Triangulation of data acquired from different sources improved the credibility and
internal validity of this research study. Published work has demonstrated that the
triangulation of data acquired from a diverse set of expert sources contributes to the
authenticity, plausibility, and validity of qualitative research (Greenlaugh & Singlehurst,
2011). Triangulation of acquired data from the focus group sessions, reflective
journaling, and from consensus-based white papers in this study improved the credibility
of the research conclusions. My use of theoretical constructs from DOI and TAM
combined with insights acquired from an exhaustive literature search helped reduce
personal bias during my formulation of research questions and the interpretation of the
responses. In summary, the methods I used to improve the internal validity of this study
included reflective journaling, respondent validation (member checking), inductive
coding, and triangulation of data from diverse sources.
Transferability
Transferability refers to the ability of a reader to apply the research process or the
research results to another situation or setting. The success of this process is dependent
on transparency and adequate description of research boundaries, parameters, and
processes (Connelly, 2016; Lincoln & Guba, 1985; Shenton, 2004). In contrast to
transferability, generalizability refers to the ability to apply the results from a research
sample to a broader population. In order for research to be transferable, the results must
122
be reproducible in different cultural settings with common variables (Shenton, 2004). In
this study, I used sequential steps, thick descriptions, a data acquisition flow chart, a
focus group moderator guide, and qualitative coding software, all of which contributed to
a high degree of transparency, which supports transferability.
I used numerous redundant and overlapping methods to improve external study
validity and transferability. I disclosed unexpected and conflicting results along with
unforeseen challenges in the research. My research conclusions include alternative and
rival explanations surrounding the potential impact of AI use during the interpretive stage
of spine imaging workflow. The use of two focus group sessions each comprised of
homogenous groups of experts from two related fields supported the detection of patterns
and themes within and across groups. I compared the findings of the focus group sessions
to the themes that emerged from consensus-based white papers, which served as research
documents in this study.
Dependability and Confirmability
The attributes of dependability and confirmability are important elements of
trustworthiness that influence the ability to replicate the research process and to test
related assumptions. The dependability of qualitative research improves with transparent
strategies such as thematic coding, content analysis, and the generation of thick
descriptions (Shenton, 2004). My use and disclosure of a field-tested focus group
moderator guide and data-coding guide offered the level of transparency required to
facilitate accurate interpretation and/or replication of this research. I used valid tools and
measures available on the Atlas.ti, Version 8 software, to analyze and manage the
123
research data in this study. I made journal entries of influential issues surrounding the
integrity or quality of data used for analysis.
I used clear and concise descriptions of the research process and the flow of data
to improve the ability for others to critique or replicate the study. I also recorded the
chain of evidence and performed content analysis and thematic coding, which I
acknowledged in detail with the help of a hierarchical coding table. I developed a data-
coding guide that includes a list of code categories and their definitions and the criteria
and methods I used to achieve and define data saturation during the analysis process. I
performed within and between case analyses with the help of computational methods to
reduce my potential bias.
Ethical Procedures
The basic ethical principles acknowledged by the Belmont Report (1979) are
respect for individuals, beneficence, and justice. Beneficence refers to the treatment of
individuals in an ethical manner by respecting their decisions, securing their safety,
prioritizing their well-being, and protecting them from harm. Justice refers to the equal
and fair management of research participants. To confirm adherence to these basic
principles I treated each research participant equally. I provided each research participant
with the same documents, had them sign the same consent forms, and exposed them to
the same data collection processes. The research protocol and conduct in this study
conformed to the tenants of the Belmont Report and to Walden University IRB
requirements.
124
Documents and Agreements
I performed the research in an ethical and honest manner to ensure the integrity of
the study and to minimize any potential harm or risk to research participants. I used
predetermined protocols and preapproved documents helped to ensure that an appropriate
and ethical approach was used throughout the entire research process. Research in the
Walden University doctoral program requires oversight by the IRB to help ensure the
integrity of the research process, as well as the safety and privacy of all research
participants. The Walden University IRB approval number assigned to this study was 02-
13-19-0129405.
Prior to collecting the data, I provided each research participant with preapproved
documents, which included an overview of the study, a focus group agenda, participant
expectations, and an informed consent form, which included confidentiality terms. The
documents safeguarded the consistent and ethical treatment of each research participant
and the ethical management of research data. The agreements disclosed any anticipated
or potential exposure to risk. The documents also acknowledged the voluntary nature of
study participation.
All of the research participants were required to sign an IRB approved consent
form prior to participating in the study. The form included confidentiality agreements.
Consent included my responsibility as the researcher to keep confidential the personal
identities and contributions of all research participants. I informed all participants of the
purpose and scope of the study, as well as the expectations for their participation. In
125
addition, I informed the research participants that they would each receive a nominal
stipend of $25 for participating in the study.
Treatment of Research Participants
I treated all of the research participants with the utmost respect. Research
participants should also be treated as autonomous agents (Kaiser, 2009). The safety and
rights of research participants must be prioritized at all times (Belmont, 1979). In
qualitative research, the researcher assumes a unique responsibility for protecting each
research participant (Orb, Eisenhauer, & Wynaden, 2001). Compliance with well-
established ethical principles such as autonomy, beneficence, and justice helps ensure
proper care of research participants (Lorell et al., 2015; Orb et al., 2001). Consistent with
the recommendations of Kaiser (2009), the methods and forms I used to obtain informed
consent in this study were adapted to the type of research and type of research
participants required. My use of field testing and an independent review of the focus
group moderator guide helped guarantee appropriate treatment of research participants in
this study.
I informed all of the research participants of their right to withdraw from the
research study at any time and for any reason. In addition, I provided participants with
the option to withdraw verbally or in writing from the study without any repercussions. I
informed all of the research participants that their names would be kept confidential and
their identities would remain anonymous. I removed research participant names from
focus group transcripts and replaced them with unique and anonymous identifiers to help
126
ensure confidentiality of their names and contributions. Examples include Participant 1
(P1), Participant 2 (P2), and so forth.
Treatment of Data
To help safeguard trust and to protect the privacy and contributions of each
research participant during the focus group sessions, participants were informed that they
are not to disclose the names or contrition of participants outside the research setting. I
deleted all names from the final focus group transcripts. The research records will be
securely stored for 5 years from the time of completion of the research study to protect
the rights of all participants. Proper storage of research records will support authorized
access for auditing or for review by qualified individuals. I will take proper steps to
discard all participant records after 5 years.
The Potential for Research Impact on Social Change
The primary purpose of this research study is to explore the potential impacts of
AI on the interpretive stage of spine imaging and to reveal its social implications. The
study addressed the potential influence on standards of care, technology development,
systems applications, and public expectations. In Chapter 5, I expand the discussion of
the research results to include it potential impacts on various levels of society.
Meaningful use of AI during the interpretive stage of spine imaging will augment
the role of the radiologist by reducing data complexity, characterizing pathology, and
offering decision support. The process will empower the radiologist as a gatekeeper of big
data and facilitate their leadership role. Improved availability of AI decision support will
increase the demand for remote access teleradiology services. The process has the
127
potential to support improved democratization of decision support surrounding diagnostic
image interpretation. It will subsequently offer a potential solution for underserved
professionals, facilities, institutions, and geographic locations.
Widespread use of AI use during radiology workflow may alter the health care
landscape, especially in the areas of image analysis, disease characterization, disease
monitoring, decision support, and final report generation. AI could offer radiologists new
solutions, capable of improving their ability to detect early stage pathology. Most
diseases are recognized at advanced stages, thus, resulting in high costs and poor
treatment outcomes. Early disease detection would contribute to more efficient care at
lower costs. These outcomes would all have a favorable social impact at many levels.
Successful use of AI during the interpretive stage of radiology workflow will
require redesign and co-evolution of supportive technologies to benefit the broader field of
health care and society. Supportive solutions will include new levels of interoperability
between databases and management systems (Tang et al., 2018). A successful co-
evolutionary process will require the development of unifying platforms, which facilitate
sharing of data and support more consistent use of disease criteria, disease classifications,
and computational disease models. The summary discussion addresses the potential
relationship between AI and relevant emerging technologies. For example, block chain
technology has the potential to provide proof-of-work validation while recording the flow
of data and computational steps across an AI-based network (Kuo, Kim, & Ohno-
Machado, 2017; Mamoshina et al., 2018). The eventual convergence of AI and block
128
chain technology has the potential to decentralize intelligent decision support and offer
increased access to computational disease models.
In summary, AI decision support could overcome human bias, reduce interpretive
error, and enhance the potential for a more precise and timely diagnosis in spine imaging.
Meaningful use of AI during the interpretive stage of spine imaging and at other levels of
radiology workflow could have a favorable impact on the role of radiologists, as well as
on the co-evolution of decision support technology, standards of care, delivery of care, and
public expectations. I address the potential social consequences of AI development and
use in spine imaging in the conclusion of this study.
Summary
I designed this research study to identify how the use of various AI solutions
could impact data management and the differential diagnostic process during the
interpretive stage of spine imaging. AI involves a rapidly evolving set of technologies
associated with numerous processes. Its role in radiology is difficult to define because of
its wide range of potential applications. It was necessary to address the potential impact
of AI with a qualitative exploratory case study approach to identify possibilities worthy
of further discussion and investigation. The design of this study supported the acquisition
of expert insights and data from different sources. This research design also supports the
development of a concept map representing potential AI applications and contributions
during the interpretive stage of spine imaging workflow. The research design and
methods I detailed in this chapter were supported by expert sources, an extensive
129
literature search, and a review of consensus-based documents surrounding the use of AI
in radiology.
An individual radiologist can no longer be required to function with precision
accuracy in the face of overwhelming, high velocity, and complex data. Radiologists
require new data analysis and decision support systems during the interpretive stage of
imaging in all fields including spine care. Successful use of AI will likely result in earlier
disease detection, better disease characterization, less diagnostic errors, and shorter
lengths of care (Kohn et al., 2014; Lee, 2017).
The research methods introduced in this chapter provide a trustworthy approach
and a scholarly foundation for further discussion and research surrounding the
development and use of AI during the interpretive stage of spine imaging. I designed this
research study to introduce the role of radiomics and the concept of the digital (virtual)
biopsy, in a manner that could be applied in spine care, as well as in other fields of health
care.
There is a growing demand for health care to become more predictive and
preemptive. Success requires a more deliberate approach to the comprehensive and
objective analysis of actionable data. My acquisition and analysis of data in this research
study revealed concepts and themes that can be used to develop additional research
strategies to pursue the role of AI solutions during the interpretive stage of spine imaging
workflow. Discoveries associated with this research can be applied to other areas of
radiology.
130
Chapter 4 presents the results of the research study. I further discuss how I
acquired and analyzed the data and how I improved the trustworthiness of the study and
related data.
131
Chapter 4: Results
Introduction
The primary purpose of this qualitative, exploratory case study was to explore the
potential impacts of artificial intelligence on spine imaging interpretation and diagnosis. I
designed the study to acquire and analyze expert opinions from different sources. My
goals for the study included identifying how AI solutions might improve the accuracy
and efficiency of interpretive workflow and the differential diagnosis process in spine
imaging. I implemented qualitative research methods to explore the possibilities
associated with computational decision support and to establish a thematic basis for
further discussion and research on the topic.
This study is one of the first to address the potential role of AI in spine care and
the concept of the digital (virtual) biopsy characterized by multiscale in vivo
interrogation of pathology. I initiated this study with the fundamental belief that images
are rich in metadata and that diagnostic imaging represents a core diagnostic process in
spine care. Select constructs of the TAM and DOI were used to guide the process of data
acquisition and analysis. During focus group sessions, I asked open-ended and probing
questions to address radiomics, interpretive workflow, the differential diagnostic process,
clinical utility, and determinants of AI adoption and use.
The volume and complexity of data acquired with advanced diagnostic imaging
methods has created a burden and exposed unprecedented opportunities for radiologists.
Big data has exceeded the ability of a radiologist to make fully informed decisions (Aerts,
2017; Gilles et al., 2016). Radiologists require augmentation of their role to reduce errors
132
and to take advantage of new opportunities for rendering a more precise and personalized
diagnosis. Without adequate technological assistance, the human interpretive process
within radiology workflow will become progressively more inaccurate, inefficient, and
untimely (Croskerry 2013; Manrai et al., 2014; Obermeyer & Emanuel, 2016; Ragupathi
& Ragupathi, 2014). Diagnostic imaging represents one of the single most important
methods for detecting and characterizing pathology in all biological systems including
the spine. This study addresses the potential for AI to reveal actionable data from
imaging studies while augmenting the role of the radiologist. My primary motivation for
performing this exploratory study was to acquire insight and provide direction for the
development of decision support solutions to support better spine care.
In this chapter, I address numerous topics such as the research purpose, the
research setting, field testing, research participant demographics, data collection, data
analysis, trustworthiness of the study, and the research results. I laid out this chapter in a
manner consistent with the chronological stages of the research process. In this chapter, I
reveal the themes and subthemes which emerged from triangulation of data and data
analysis. I also provide supportive evidence for the iterative process. This chapter
indicates how the research design and the use of strategic methods improved the
trustworthiness of the research results. In addition, I provide an overview of my reflective
journaling, which includes disclosure of its impact on the research process and results.
The conclusion provides a summary of the research findings along with a transition to
Chapter 5.
133
Field Testing
I field tested the moderator guide to help establish the required level of
appropriateness, clarity, and relevance of the approach used during focus group sessions.
I achieved these goals through independent review of the strategies and resources
outlined in the focus group moderator guide, including the agenda, topic introduction,
ground rules for participation, open-ended questions, and a script for the conclusion of
the session. I also field tested the appropriateness and clarity of my introductory
PowerPoint slides I used during focus group sessions.
I performed the field testing, on separate occasions, with one radiologist and one
AI expert. The experts who participated in the testing met study inclusion criteria but did
not serve as research participants. The field testing process allowed for peer-review of the
focus group protocols and resources. The experts who participated in the field testing
were not asked to answer or respond to any research questions. I did not have an
employment or consulting relationship with the field testing experts.
I used field-testing to help establish appropriate and relevant focus group
protocol. The process included assessment of the methods of data acquisition and the
pattern of inquiry. I made no significant changes as a result of field testing, with the
exception of the order and clustering of questions to be used during the focus group
sessions. Nor did I make contextual revisions to the primary focus group questions. I field
tested which PowerPoint concept slides were the most neutral and concise to help guide
focus group discussions on complex topics. I removed a few slides from the presentation.
I made no content changes to the remaining PowerPoint slides.
134
My field testing helped identify the order of the focus groups. I determined that
the radiology focus group session should take place first, followed by the AI expert focus
group session. The independent experts believed this order would support more
progressive, in-depth coverage of the research topic. I thought it was necessary to utilize
field testing to reduce bias, ensure professionalism, and to efficiently operationalize
available multimedia and data acquisition methods.
Research Setting
I held the focus group sessions at an independent and professional location. The
setting supported physical participation and teleconference access for all participants. I
provided each of the research participants the option to be physically present during the
focus group session or to access the session using Zoom, an independent well-established
teleconferencing solution. Each research participant chose to access the focus group
session with the teleconferencing solution.
During the live focus group sessions, each participant using Zoom had access to
an online image gallery for intimate real-time viewing and interaction with all
participants. Each participant also had the independent option of engaging a speaker
highlight function that prioritized the participant actively contributing. I provided each
participant with access to a dynamic online gallery to facilitate efficient communication. I
served as the sole moderator for each focus group session. I moderated each session from
a conference room designed for focus groups. The room consisted of a boardroom table
and chairs, professional audio system, and a large format wall-mounted screen with a
camera.