Information Technology Managers’ Strategies
for Implementing Data Governance
Section 1: Foundation of the Study
Background of the Problem
Data plays a crucial role in the information economy by providing the knowledge
needed for business decisions and scientific and technical processes. In the financial
sector, data-driven transformation is evident, highlighting the central and critical role of
data (Fryczak, 2020). Information technology (IT) managers in businesses and
organizations rely on data quality, including availability, reliability, accuracy, validity,
and security (Cichy & Rass, 2019). Meeting these requirements necessitates the
establishment of a data governance process, which ensures data availability, usability,
integrity, and security based on standards, policies, and control (Abraham et al., 2019).
The 21st century is characterized by the era of big data in which businesses aim
to harness and utilize vast amounts of data to drive decisions and establish strong
customer relationships (Fan et al., 2019). Artificial intelligence (AI) is playing an
increasingly significant role in organizing and leveraging big data, particularly in the
retail sector (Duan et al., 2019). In the realm of AI and big data, effective data
governance enhances the consistency and trustworthiness of data for big data algorithmic
systems (Janssen et al., 2020). Regulated industries such as health care require data
governance (DG) to safeguard patient records from misuse, especially with the
exponential growth of digital health care data worldwide. Tiffin et al. (2019) proposed a
DG framework to address this challenge, particularly in developing countries where data
security and privacy are in their infancy.
The discussed use cases clearly demonstrated that business leaders and IT
managers have a vested interest in the quality of business data. DG is essential not only
for making data-driven decisions but also for protecting data against misuse. The current
study aimed to gain an in-depth understanding of the strategies employed by IT
managers to implement DG. The findings may reveal the driving factors, challenges, and
opportunities associated with DG implementation.
Problem Statement
In the era of Industry 4.0, data plays a crucial role in enterprise software
architecture, functioning as a valuable business asset and a key component of data-as-
aservice. However, DG has often been informal, fragmented, and inadequate when solely
driven by the IT department (Al-Ruithe et al., 2019). Insufficient DG affects data quality
at the organizational level, which is a significant concern in the information economy
(Abraham et al., 2019). Studies indicated that nearly half of new data sets are error prone
due to the lack of stringent DG, emphasizing the need for organizations to prioritize data
quality improvement (Nagle et al., 2020).
The general IT problem was that many companies fail to fully integrate DG with
IT governance. This oversight can result in inconsistencies in data availability,
compromises in data integrity and security, and various costs such as data loss, legal
actions, and compliance penalties. The specific IT problem was the lack of strategies
among some IT managers to effectively implement DG.
Purpose Statement
This qualitative pragmatic study aimed to investigate the strategies employed by
IT managers in effectively implementing DG. By exploring the factors that facilitate or
impede the implementation of DG within organizations, the study sought to enhance
knowledge in this area. The targeted population consisted of IT managers in the United
States, specifically those affiliated with companies that have achieved successful
implementation of a comprehensive DG process. The potential positive outcomes of this
research include advancements in data quality and security for businesses and
individuals, as well as the promotion of accessible DG solutions through affordable
applications, leading to its democratization.
Nature of the Study
The study aimed to explore the strategies used by IT managers to implement DG.
Research methods include quantitative, qualitative, and mixed methods. Qualitative
research is inductive and often involves interviews to gain a deeper understanding of
cultural aspects or individuals’ perceptions (Small, 2021). On the other hand,
quantitative research is deductive and relies on numerical data to test hypotheses (Lima
& NewellMcLymont, 2021). Mixed-methods research (MMR) combines qualitative and
quantitative approaches, which was not suitable for the current study due to time and
resource constraints. MMR arose in the 1980s from the need to mitigate bias and
weaknesses inherent in quantitative and qualitative methods by using triangulation of
data sources to provide a more balanced view of the phenomenon (Timans et al., 2018).
MMR has gained momentum in the social scientific circles. MMR could be appropriate
for providing a quantitative follow-up to validate driving factors for the implementation
of DG by IT managers.
In the current study, the qualitative method was selected to explore the “what,”
“why,” and “how” of IT managers’ DG strategies. Qualitative methodology allowed for
a comprehensive exploration of the phenomenon and the generation of new knowledge.
Although MMR could have provided quantitative validation, it was not feasible for this
study’s scope.
The pragmatic qualitative design was deemed most suitable for this study because
it helped me gain an in-depth understanding of IT managers’ DG strategies. The case
study design, under the pragmatic approach, enables detailed examination through data
collection methods such as observations, interviews, and secondary data analysis. The
case study design allows for studying a bounded case or multiple cases without
interference (Takahashi & Araujo, 2020). Ethnographic and phenomenological designs
were not chosen because they did not align with the focus on IT managers’ strategies or
data engineering topics. The strength of the ethnographic design is in understanding the
shared patterns of a cultural group (Gherardi, 2019). This was not the focus of the current
study. The phenomenological design was not appropriate because it focuses on the life
experiences of individuals (Greening, 2019). Therefore, this design was not suitable for a
data engineering topic such as DG.
To answer the research questions, I conducted a qualitative study involving
interviewing IT managers. This approach provided insights into the strategies used for
implementing DG. A pragmatic design was justified because it facilitated individual
interviews with IT managers from various institutions without requiring partner
organizations to be approved by the institutional review board (IRB).
Recruiting IT managers for individual interviews and collecting qualitative data
were essential for the planned research design. Interview questions and protocols were
developed to align with the problem statement and purpose of the study. The data
collected included IT managers’ responses regarding their strategies for successful DG
implementation.
Interviewing IT managers at institutions with DG programs may have required
approval from both Walden University’s IRB and the company’s IRB. However, Walden
University offers a pragmatic approach in which doctoral students can interview selected
participants with IRB approval without necessitating partner organization approval. The
main challenge was finding organizations that had successfully implemented DG
programs and IT managers willing to share their insights.
Research Question
The central research question was the following: What strategies do IT managers
use for implementing DG programs?
Interview Questions
Interview questions are among the typical data collection mechanisms for a
qualitative pragmatic study. The questions must be designed to collect the data needed to
answer the main research questions and must be in strict alignment with the conceptual
framework of the study. For the current study, the research question addressed the
strategies that IT managers use for implementing successful DG programs. By exploring
strategies that IT managers use to drive successful DG programs, this study aimed to
provide insights others can use to implement similar programs. The conceptual
frameworks guiding this study were, Abraham et al.’s (2019) DG conceptual framework
and Khatri and Brown’s (2010) unified framework for DG. Table 1 shows an alignment
of the two frameworks by grouping elements that belong together; this helped me
simplify the formulation of interview questions to cover the two frameworks at the same
time.
Table 1
Alignment of Abraham et al.. (2019) Conceptual Framework for Data Governance and
Khatri and Brown (2010) Data Governance Framework
Abraham et al.
(2019) dimension
Description Khatri and Brown
(2010) dimension
Antecedents
The contingency factors, which
impact the adoption and
implementation of data governance;
internal and external antecedents.
Data principles
Data scope Pertains to the data asset an
organization needs to
govern;traditional data and big data
Cloud vs
noncloud
Organizational
scope
Determines the organizational
expansiveness of data governance and
roughly corresponds to the unit of
analysis; intra-organizational and the
inter-organizational scope.
Domain scope Data decision domains, to which
governance mechanisms are applied:
data quality, data security, data
architecture, data lifecycle, meta data,
data storage and infrastructure
Data quality
Data access
Metadata
Governance
mechanisms
Core dimension of the framework and
encompass structural, procedural, and
relational mechanisms
Data lifecycle
Data access
Consequences The effects of data governance;
intermediate performance effects and
risk management
Data quality Data
access
This study had three categories of interview questions. First were
frameworkrelated questions that addressed why IT managers chose to implement DG
programs at the select participant organizations. Elements of these frameworks were
used to formulate interview questions for capturing insights into the principles,
antecedents, and life cycle of DG. The second category of interview questions addressed
the critical success factors (CSFs) that contributed to successful DG programs at the
selected institutions. Alhassan et al.’s (2019) CSF matrix was used for this purpose. The
third category of interview questions addressed the challenges that lead to DG
implementation failures, including but not limited to lack of executive support, costs,
data quality, complexity of business cases, IT focus, lack of clearly defined roles and
responsibilities, and technical skills gap (AlRuithe et al., 2019; Alsousi & Shah, 2022;
Janssen et al,, 2020; Rasheed et al., 2022). It was important for the current study to
capture this third category for lessons learned that may inform IT managers wanting to
implement DG programs.
Framework-Related Questions
1. What data principles (i.e., contingent factors, for example, data quality,
regulatory compliance) caused your organization to decide to implement a
data governance program?
2. What is the nature of data assets (traditional or big data) that are the main
target for data governance?
3. A data governance program is often primarily concerned with internal
organization data assets. Can you please comment on any third-party data
assets that your organization uses and needs governed?
4. Data quality, data access, or metadata management capability are often needs
or challenges for many organizations like yours. Can you please comment on
data quality, data access, or metadata management capability to your
organization and their classification as needs or challenges?
5. Some organizations establish a data life cycle process prior to implementing
the data governance program. What was the case for your organization?
6. A company culture that places data as a corporate asset often facilitates the
adoption and acceptance of data governance. Is data treated as an asset within
your organization’s culture? Can you please comment on the culture and
attitudes toward data governance within your organization?
7. Does your organization treat data governance as a company-wide initiative
backed up by the executive leadership, or an effort driven solely by certain
departments?
CSF-Related Questions
1. Prior to starting the data governance programs, did your organization have in
place a clear data strategy, including data management processes, procedures,
and policies? If so, what were they? Are they organization wide or
departmental?
2. Can you please summarize the various data management activities within
your organization?
3. What individuals are responsible for the data-related activities in your
organization? Who defines data policies and processes? What are data roles
in your organization?
4. What specific tools, applications, or technologies does your organization use
to manage data assets?
Challenges
1. What is the nature of challenges that your organization faced during the
implementation of the data governance program?
2. How did the data governance help (or not) address those challenges?
3. What strategies does your organization understake to address those
challenges?
4. As an organization, if you had to do it again, what would you do differently
and where would you start?
5. Any advice you would give to other companies wanting to build a data
governance program?
Conceptual Framework
DG plays a crucial role in optimizing the management of data within an
organization. DG is one of the seven complementary areas of the data management
framework, which also includes data architecture, metadata, data quality, data life cycle,
analytics, and data privacy. During the 1980s, a significant shift occurred in the field of
IT, moving from sequential storage methods such as punch cards and magnetic tapes to
random access storage through hard disks. This transformation revolutionized computing
and gave rise to the concept of data management (DM; Dagnaw & Tsigie, 2019). DM
encompasses an administrative process that spans the entire life cycle of data, including
its creation, ingestion, processing, storage, utilization, and disposal. The objective of DM
is to ensure the accessibility, reliability, and timely delivery of data to users. By
implementing effective DM practices, organizations can minimize risks and costs
associated with data loss, regulatory noncompliance, misuse, and data breaches.
However, within the broader framework of DM, DG focuses on establishing and
enforcing policies, standards, and procedures related to DM. DG encompasses the rules,
roles, responsibilities, and processes that guide the proper handling, protection, and
utilization of data assets. DG aims to address various aspects of data management
(Abraham et al. 2019):
•Data ownership: Clearly defining who is accountable for data assets within
the organization, ensuring that the right people have the authority and
responsibility to make decisions about data.
•Data dtewardship: Assigning data stewards who are responsible for managing
and maintaining data quality, integrity, and security throughout its life cycle.
•Data policies and standards: Developing and enforcing policies and standards
that govern data management practices, ensuring consistency and adherence
to regulatory requirements.
•Data quality management: Implementing processes and controls to monitor,
assess, and improve the quality of data, ensuring its accuracy, completeness,
consistency, and relevance.
•Data security and privacy: Establishing measures to protect sensitive data
from unauthorized access, ensuring compliance with relevant data protection
regulations and safeguarding individual privacy.
•Data compliance and risk management: Ensuring adherence to regulatory and
legal requirements related to data management and mitigating risks associated
with data breaches, noncompliance, and other potential issues.
•Data integration and interoperability: Facilitating the seamless integration and
sharing of data across different systems, applications, and departments within
the organization.
By optimizing DG practices, organizations can foster a data-driven culture,
enhance decision-making processes, improve operational efficiency, and gain a
competitive edge in the rapidly evolving digital landscape. The current study included
two conceptual frameworks: Abraham, et al.’s (2019) DG conceptual framework and
Khatri and Brown’s (2010) unified framework for DG. The DG conceptual framework
by
Abraham et al. offers a comprehensive approach to comprehend the implementation of
DG, including challenges and successes, through six dimensions of analysis: governance
mechanisms, organizational scope, data scope, domain scope, antecedents, and
consequences of data governance. The logical connection between this framework and
the nature of my study was using the six criteria domains to gain insights into the
motivations and strategies of IT managers when implementing DG programs.
Khatri and Brown’s (2010) unified framework for DG provides a comprehensive
guide for researchers and practitioners seeking to develop effective data governance
programs. This framework encompasses five key decision domains that serve as the
foundation for successful implementation:
•Data principles: This domain focuses on aligning data governance with the
business needs and establishing the expected behaviors of information
systems professionals and users. By defining clear principles, organizations
can ensure that DG efforts are in line with strategic objectives.
•Data quality: Within this domain, the emphasis is assessing the value of data
assets based on dimensions such as accuracy, timeliness, completeness, and
credibility. Understanding and maintaining data quality is crucial for reliable
decision making and efficient operations.
•Metadata: The metadata domain involves capturing and managing data
descriptions. Metadata provides essential context and understanding of data
elements, facilitating effective DG practices and enabling efficient data
discovery and utilization.
•Data access: This domain deals with defining levels of data access, such as
confidential or restricted access, to meet internal and regulatory requirements.
Establishing appropriate access controls ensures data security and compliance
with relevant policies and regulations.
•Data life cycle: The final decision domain focuses on managing the entire life
cycle of data, including its creation, storage, use, archiving, and disposal. By
establishing comprehensive data life cycle processes, organizations can
ensure data integrity, accessibility, and compliance throughout its lifespan.
To investigate strategies employed by IT managers for implementing DG
programs, I employed these five decision domains as a basis for designing interview
questions. By exploring these domains in-depth, I was able to provide insights into the
approaches, challenges, and successes of DG initiatives within different organizational
contexts. The Abraham et al. (2019) DG conceptual framework (see Figure 1) and Khatri
and Brown’s (2010) unified framework for DG (see Figure 2) share common elements,
particularly in terms of data quality and data principles. Table 1 displays the mapping
between Abraham et al.’s framework and Khatri and Brown’s framework. A single
framework, such as Abraham et al.’s (2019), would have sufficed. However, I decided to
incorporate the framework by Khatri and Brown as well because it offered a more
detailed analysis compared to the former. The framework presented by Abraham et al.
focuses on the cause (antecedents), process, and effect (consequences) of a DG program,
while Khatri and Brown’s framework emphasize attributes.
Figure 1
Khatri and Brown (2010) Unified Framework for Data Governance
Figure 2
Abraham, R., Schneider, J., & Vom Brocke, J. (2019) Conceptual Framework Data
Governance
Furthermore, Alhassan et al.’s (2019) CSFs (see Table 2) provided additional
perspectives to gain deeper insights into the effective implementation of data governance
by IT managers. Theories, frameworks, primitives, constructs, and design patterns are
valuable guides for design and implementation. However, the industry’s practical and
tested knowledge, in the form of best practices, provided the most valuable insights.
These best practices are instrumental in assessing the success or failure of an
implementation. Similarly, CSFs played a vital role in this evaluation. To gauge IT
managers’ adherence to industry best practices when implementing DG programs, I
employed CSFs as a metric to measure their response.
Table 2
Appplication of Abraham et al. (2019) Conceptual Framework for Data Governance and
Khatri and Brown (2010) Data Governance Framework to Research
Reference Title Khatri and Brown
(2010)
Abraham et al.
(2019)
Abraham et al. (2019) Data governance: A conceptual
framework, structured review, and research
agenda
x
Al-Badi et al. (2018) Exploring big data governance frameworks x
Alhassan et al. (2017) Data governance activities: A comparison
between scientific and practice-oriented
literature
x
Alhassan et al. (2019) Critical success factors for data governance:
A theory building approach
x
Alsousi & Shah (2022) Data governance for SME: Systematic
literature review
x
Bento et al. (2022) How data governance frameworks can
leverage data-driven decision making
x
D’Hauwers et al. (2022) Business models value and control in data
ecosystems x
Derakhshannia et al. (2020) Data lake governance: Towards a systemic
and natural ecosystem analogy x
Fernandes (2022) Trajectories and challenges for data
governance implementation in the public
sector: A systematic review
x
Janssen et al. (2020) Data governance: Organizing data for
trustworthy artificial intelligence
x x
Jimenez et al. (2019) Oververview of data governance in business
contexts
x
Juddoo et al (2018) Data governance in the health
industry:Investigating data quality
dimensions within a big data context
x
Karkošková (2022) Data governance model to enhance data
quality in financial institutions
x x
Machado et al. (2022) Sustainable data governance: A systematic
review and a conceptual framework x
Mao et al. (2022) Government data governance framework
based on a data middle platform
x x
Nadal et al. (2022) Operationalizing and automating data
governance
x
Nielsen (2017) A comprehensive review of data
governance literature
x
Vilminko-Heikkinen & Pekkola
(2019)
Changes in roles, responsibilities and
ownership in organizing master data
management
x
Vojvodic & Hitz (2022) Relation of data governance, customer
centricity and data processing compliance
x
Walsh et al. (2022) Grounding data governance motivations: A
review of the literature
x
Yebenes & Zorilla (2019) Towards a data governance framework for
third generation platforms
x
Definition of Terms
Antecedents of data governance: Internal (e.g. organization strategy, IT strategy,
corporate decision-making authority, IT infrastructure) or external (e.g. legal and regulatory)
factors that influence the adoption and implementation of data governance (Abraham et al.,
2019; Schlackl et al., 2022).
Big data: Data sets with the following characteristics (volume, variety, velocity,
veracity, value): very large volume, complexity, variety to include structured and
nonstructured data, streaming, flat files; this mix of files is usually stored in a data lake
cloud storage (Janssen et al., 2020; Ramachandran et al., 2022; Yebenes & Zorrilla,
2019).
Big data governance: DG as applied to big data.
Critical success factors (CSFs): Factors that impact the success of DG
implementation; ranked in order of importance they are (a) employee data competencies,
(b) clear data process and procedures, (c) flexible data tools and technologies, (d)
standardized easy to follow data policies, (e) established data roles and responsibilities,
(f) clear inclusive data requirements, and (g) focused and tangible data strategies
(Alhassan et al., 2019).
Data governance (DG): Policies and procedures to maintain authority and control
over the management of an organization’s data assets for the purpose of increasing the
value of data and minimizing data-related cost and risk (Abraham et al., 2019;
Hikmawati et al., 2021). The four pillars of data governance are data quality, data
stewardship, data protection and compliance, and data management.
Data life cycle: Processes for master data management, data production, storage,
retention, retirement, traceability (Khatri & Brown, 2010).
Data quality (DQ): Value of data assets along dimensions such as accuracy,
timeliness, completeness, credibility, reliability, relevancy (Khatri & Brown, 2010;
Hikmawati et al., 2021).
Data principles: DG objectives intended to fulfill business needs including
definining behaviors of information professionals and users; common data principles are
value add, compliance, decision rights, future proofing, performance, process (Khatri &
Brown, 2010).
Decision domains: Five core DG areas of the Khatri and Brown (2010) unified
framework for DG: (a) data principles (to link data governance to the business needs, and
define behaviors of IS professionals and users), (b) data quality (value of data assets
along dimensions such as accuracy, timeliness, completeness, credibility), (c) metadata
(data descriptions), (d) data access (levels, e.g. confidential, restricted for internal and
regulatory purposes), and (e) data life cycle.
Information governance: A decision and accountability framework that dictates
allowed behaviors for the information life cycle (creation, valuation, use, sharing,
storage, archiving, and deletion); information governance includes the policies,
standards, processes, metrics, and roles governing the efficient and effective use of
information in alignment with the organization’s strategy and objectives (Al-Ruithe et
al., 2018; Mikalef et al., 2020).
Master Data Management (MDM): Core mission is to provide method and
processes for maintaining, integrating, and harmonizing master data to ensure consistent
system information across the organization; MDM controls master data by keeping it
consistent, accurate, current, relevant, and contextual to serve business needs
companywide; master data (customer, supplier, and material data, emloyee data) is the
version of the truth and reference data across the company’s business functions (sales,
services, order management, purchasing, manufacturing, billing, accounts receivable,
and accounts payable) to be used by business intelligence and applications (Haneem et
al., 2019;
Hikmawati et al., 2021).
Assumptions, Limitations, and Delimitations
Assumptions
As immutable requirements, assumptions play a crucial role in any academic
research, and the current study was no exception. As the researcher, I needed to state
these assumptions to ensure they were understood. Misguided assumptions can have a
detrimental impact on the analysis of findings and the results of a study. One assumption
was that IT managers interviewed were knowledgeable in DG, had implemented a DG
program, and were able to speak to the implementation of such. Additionally, I assumeed
a working DG program in lieu of a successful DG program. Success means different
things depending on audience and criteria. To declare a DG program successful implies a
prior evaluation. Another assumption was that organization documentation about DG
processes was up to date with the technology and industry.
Limitations
Due to purposeful sampling and a smaller sample size, qualitative research often
faces challenges in terms of reliability and validity, which are crucial for ensuring the
quality of the study (Malakar, 2022; Small, 2021). To address these challenges,
researchers often employ triangulation to enhance the reliability and validity of
qualitative results. However, qualitative studies also have inherent limitations, such as
difficulties in accessing secondary data and the potential for inadequately designed
questionnaires, that may not capture sufficient information to address the research
question. To mitigate these challenges, I used triangulation techniques and designed
interview questions rigorously. Unlike quantitative studies, in which the sample size can
be large, qualitative studies, including case studies, generally recommend a sample size
of 10 to 20 participants, at which point data saturation is typically observed (Malakar,
2022). Considering budget constraints, time limitations, and the availability of IT
managers willing to participate in interviews, I strived to limit my sample size as much
as possible, ideally targeting 12 participants.
This study was based on two conceptual frameworks, Abraham et al.’s (2019)
DG conceptual framework and Khatri and Brown’s (2010) unified framework for DG,
which have been widely referenced in the literature on DG. However, these frameworks
have not been validated through practical implementation, which presented an inherit
weaknesses associated with them. Abraham et al. acknowledged limitations of their
framework, including differences in applicability based on data scope (e.g., traditional
data vs. big data) and the lack of validation against real-world use cases by practitioners.
Delimitations
The current study aimed to address specific research questions. Therefore, it was
essential for me to determine and communicate the scope of the study. Delimitations
play a crucial role in defining the scope by establishing boundaries that determine what
is included and excluded from the study (Malakar, 2022). In the current study, there were
three main delimitations: program delimitation, nature of the study delimitation, and
framework delimitation.
First, the focus was on applying existing research to bridge gaps in knowledge
rather than creating new theories, inventions, or technical solutions. The study did not
propose a novel DG solution, architecture, theory, or framework. Instead, it sought to
understand the strategies employed by IT managers in implementing DG, with the aim of
sharing lessons learned and best practices with other practitioners undertaking similar
programs.
The nature of the study aligned with the characteristics of a pragmatic study.
Qualitative research designs often employ purposeful sampling, such as criterion
sampling, to identify information-rich cases for the case study (Takahashi & Araujo,
2020). For the current study, the participant population was limited to IT managers in
select private sector companies in the United States, with specific selection criteria
outlined in the Research Method and Design section and aligned with the conceptual
frameworks used.
This study included two conceptual frameworks to understand the
implementation of DG: Abraham et al.’s (2019) DG conceptual framework and Khatri
and Brown’s (2010) unified framework for DG. Within Abraham et al.’s framework, the
data scope dimension differentiated between traditional data and big data. Traditional
data is stored and managed on a centralized architecture such as a database or
datawarehouse; big data, on the other hand, is used on distributed architecture and
storage such as a data lake. Although the literature review included big data to ensure
completeness and relevance to the current data economy, the current study’s focus, data
collection, data analysis, and conclusions were limited to traditional data rather than big
data. This distinction was made due to the increased interest in big data management and
governance, which presented a significant research question of its own. By clearly
establishing these delimitations, I could narrow the study’s focus, provide valuable
insights within its defined scope, and contribute to the understanding and application of
DG strategies for IT managers in the private sector.
Significance of the Study
Contribution to IT Practice
The primary objective of IT is to facilitate business functions, thereby leading to
the achievement of an organization’s strategic objectives and profitability. DG plays a
crucial role in maintaining authority and control over an organization’s data assets,
aiming to increase their value while minimizing costs and risks (Abraham et al., 2019;
Hikmawati et al., 2021). The success or failure of DG impacts IT practices.
The findings of the current study have both theoretical and practical implications
that may benefit companies and data governance practitioners. First, the study may
provide valuable insights into the factors and challenges associated with implementing
DG, thereby enabling organizations to make informed decisions. Second, the study may
highlight emerging trends in the DG field, including aspects such as data quality and
trustworthiness, which are essential for enabling advanced analytics and AI in the
context of Industry 4.0 (see Yebenes & Zorrilla, 2019).
Successful DG is a comprehensive initiative that involves the collaboration of
multiple stakeholders encompassing various roles and business units within the
company. However, many organizations tend to conduct DG in isolated silos, often
driven solely by IT (Khatri & Brown, 2010). By recognizing and addressing this issue,
companies can foster synergy between strategy and technology, thereby enhancing their
DG practices.
The positive outcomes of effective DG implementation include improved data
quality, enhanced data security, increased trustworthiness, and overall value. These
attributes amplify the value of data assets, promoting innovation and generating revenue
for the enterprise. Additionally, robust DG helps mitigate risks such as mishandling
sensitive customer data, data breaches, and noncompliance with regulations, which can
cost companies millions of dollars annually (Schlackl et al., 2022). By maximizing the
value and minimizing the risks associated with data assets through robust data
governance practices, companies can achieve a healthy bottom line. This translates to
their ability to sustain jobs by retaining current employees, hiring new talent, and
avoiding significant layoffs during adverse economic situations.
Implications for Social Change
Technological innovation should prioritize the best interests of society, fostering
positive social change. This is achieved through improved data quality, security, and privacy,
as well as the democratization of DG. Effective DG enhances the quality, trustworthiness and
exchange of data (Hikmawati et al., 2021), enabling companies to leverage advanced
analytics and AI to unlock value from big data. Consequently, job creation, technological
development, and economic growth benefit society as a whole. However, the implementation
of DG is often costly and prohibitive for smaller companies, impeding its democratization.
To address this issue, the commoditization and self-serve nature of DG would significantly
reduce costs, enabling smaller companies and individuals to access higher data quality and
security. Currently, DG applications incur significant expenses, ranging from tens to
hundreds of thousands of dollars annually. Such costs are often unaffordable for small and
medium-size companies. Future trends indicate that independent software vendors such as
Tableau and Microsoft plan to integrate DG capabilities into their platforms and applications,
providing businesses with accessible solutions. This becomes even more crucial in the era of
big data, in which DG remains a considerable challenge. The availability of self-serve DG
will cater to the needs of the masses, extending beyond companies alone.
A Review of the Professional and Academic Literature
To conduct a study effectively, researchers should devise a strategy for selecting
academic literature that aligns with the topic and research questions and draws from the
existing body of knowledge. In the current study, the literature review aimed to equip me
and the reader with a theoretical foundation on the DG topic (definition, conceptual
frameworks, challenges, best practices, research topics). Such a premise was critical for
gaining insights into the strategies employed by IT managers to implementing DG
programs, which was the focus of this study. The structure of the literature review must
align not only with the research topic but also with the conceptual frameworks identified
and main themes from the literature.
This literature review focuses on DG and the chosen conceptual frameworks,
aiming to align with the topic. The review follows a structured approach based on the
themes derived from two conceptual frameworks (see Abraham et al., 2019; Khatri &
Brown, 2010). The literature review is structured into three sections. In the initial
section, I explore the use of conceptual frameworks in previous studies and research
contexts. The subsequent section centers on the application of these frameworks in the
context of the current study. The third section addresses alternative DG frameworks,
identifies existing gaps, and outlines potential future research avenues in the field of DG.
The review gives an overview of big data governance and identifies potential future
research topics in DG. The comprehensive understanding of challenges, success factors,
antecedents, and governance mechanisms gained from this review aided me in
formulating interview questions that elicited key insights regarding strategies and
motivating factors that drive companies to invest in DG programs.
The relevance of big data governance in this study was justified for several
reasons. First, it was crucial to ensure the comprehensiveness of the study by not focusing
solely on traditional DG. Second, big data has become pervasive in cloud computing and
Industry 4.0, making it imperative to discuss. Failing to address big data governance
would have been a significant omission in this study. Moreover, as more companies
migrate their IT and data infrastructures to the cloud, the subject of big data and its
governance becomes increasingly relevant. When interviewing IT managers regarding
their DG strategies, I needed to be aware that some have already experienced or were
planning to undergo a complete cloud migration for their companies. Therefore, big data
governance could serve as a potential topic for future research.
A methodological search strategy is fundamental to a good literature review. For
this literature review, I gathered 100 peer-reviewed papers published within the last 5
years (2018 onward) from reliable sources such as Google Scholar, Walden University
Library, and UC Santa Barbara Ulrich Periodicals Directory. To assemble a
comprehensive list of relevant articles, I initially used keywords such as “data
governance,” “big data governance,” “data,” and “governance.” I then filtered the list to
exclude articles unrelated to DG. I narrowed the results further by using the keyword
“framework” to identify frameworks suitable for my study. Two frameworks surfaced
that were comprehensive enough to be used as a foundation for my study: Abraham et
al.’s (2019) conceptual framework for DG and Khatri and Brown’s (2010) DG
framework. Using the dimensions of DG from these frameworks, I further refined my
search by employing a combination query of “data governance” AND keywords such as
“antecedents,” “data principles,” “data scope,” “domain scope,” “governance
mechanisms,” “data access,” “data lifecycle,” and “metadata.” By following this
systematic approach, I expected to provide a solid foundation for the study. Furthermore,
the literature review enabled me to develop a research question related to the core
definition for DG and the existing body of knowledge on the topic.
Application of Khatri and Brown’s (2010) and Abraham et al.’s (2019) DG
Frameworks
The aim of this section is to explore the practical applications of two key data
governance frameworks, the Abraham et al. (2019) conceptual framework for DG and
the Khatri and Brown (2010) DG framework, within the broader academic literature. To
achieve this, I conducted a thorough examination of references to these frameworks
within the literature review, as presented in Table 3. This analysis involved assessing
how these frameworks have been employed in various studies, identifying their
applications, pinpointing any gaps or limitations in their use, and drawing relevant
conclusions. Khatri and Brown (2010) played a foundational role in establishing a
reference point for DG frameworks, serving as a precursor tp subsequent research in this
domain. The current study highlighted a significant gap in the field of DG: the absence
of a guiding framework for practitioners and researchers. To address this issue, I
included five key decision domains that are essential for comprehending DG: data
principles, data quality, data access, data lifecycle, and metadata.
Table 3
Unified Critical Success Factors for Data Governance (Alhassan et al., 2019; Bento et
al., 2022; Yatya et al., 2021; Daneshmandnia, 2019)
Author Category
Employee data
competencies
Data governance activities that involve human
action. Directly impact the defining, implementing,
and monitoring of data processes and procedures,
as well as data policies and data requirements.
[1],[2] skills
Clear data processes
and procedures
Data policies in place which should be detailed
and operationalized during activities related to data
processes and procedures.
[1],[2] management
Flexible data tools and
technologies
All the activities related to software and hardware
that affect the data in an organization, including
the presentation and storage of data.
[1],[2] technology
Standardized easy-
tofollow data policies
Data policies are short statements that provide the
high-level guidelines and rules necessary for
dealing with data. These relate to data principles.
[1],[2],[4] management
Established data roles and
responsibilities
Individual(s) responsible for the data-related
activities in the organization, who defines the
policies and processes for the data as well as
assigning the duties for the actions related todata.
[1],[2],[5] management
Clear inclusive data
requirements
Define all aspects of data implementation, suchas
data flows and integration.
[1],[2] management
Focused and tangible data
strategies
Include planning fordata governance in order to
achieve its goals, as well as the main activities
related to considering data as assets.
[1],[2] management
Phased implementation
,
Start with small steps [5] management
Support from senior
management
.
Get buy-in and support from top leadership;
stakeholder support.
[2],[3],[5] organization
DG implementation and
maturity assessment
Measure goals with metrics [4] management
Communication Prioritize early communication, communicate
vision,strategy, and trust across the organization to
socialize the value of DG, and a culture that
embraces DG and its benefits. Strategic
communications.
[2], [4],[3],[5] management
Data quality Maintain high standards of data quality [4] technology
Team work Collaboration [4] organization
Critical success factor Detail
[1] Alhassan et al. (2019)
[2] Bento et al. (2022)
[3] Daneshmandnia (2019)
[4] Ibrahim et al. (2021)
[5] Yatya et al. (2021)
In building on the work of Khatri and Brown (2010), Abraham et al. (2019)
expanded the existing conceptual framework. This expansion included dimensions that
enhance the understanding of DG motivations and programs. Notably, these dimensions
align with the five decision domains proposed by Khatri and Brown (2010). To provide
practical context, the following sections offer a concise overview of how this study’s
conceptual frameworks have been applied to selected references from Table 3.
Motivations for DG Programmes
This study sought to gain insights into the strategies employed by IT managers
when implementing DG within different organizations in the United States. I aimed to
explore the motivations driving these strategies and the challenges encountered during
the process. This aligned with the broader scope of research in the field, which has
focused on unraveling the underlying motivations behind the adoption of DG practices.
Traditionally, the four pillars of DG have been data quality, data stewardship, data
protection and compliance, and data management (see Figure 3). In a related study
conducted by Walsh et al. (2022), the primary objective was to understand the rationale
behind companies initiating DG programs. Walsh et al. employed a structured
framework developed by Khatri and Brown (2010), encompassing five key decision
domains: data principles, data quality, data access, metadata, and data lifecycle. The
findings were organized within these domains. Notably, the conclusions drawn from the
analysis of data principles, data access, and data quality domains pointed toward
motivations rooted in operational and technological aspects. However, the study included
limitations in revealing motivations within the metadata and data life cycle domains.
This gap in understanding underscored the need for further research in these areas.
Figure 3
Four Pillars of Data Governance: Data Quality, Data Stewardship, Data Management,
Data Protection and Compliance
One significant contribution of the current study was its recognition of the Khatri
and Brown (2010) framework, which serves as a model and a template for constructing
and evaluating DG programs. The framework’s comprehensive attributes within each
domain offer valuable insights for researchers and practitioners. The framework’s
attributes provide a deeper understanding of the motivations driving DG initiatives, as
D
G
D Quality
D
S
D
M
C
D Protection
D
T
well as the organizational and technical factors influencing them, while addressing the
associated challenges. This framework emerged as a valuable tool for advancing the field
of DG.
Sustainability of DG Programs
DG’s primary goal is to enhance the value of data assets while minimizing
organizational risks. To achieve this, companies must ensure that such initiatives are
sustainable and ongoing. Bento et al. (2022) emphasized the significance of DG in
governing data assets, deriving insights, defining roles and accountabilities, and
formulating strategies. Although their research extended the foundations laid by Khatri
and Brown (2010) and DG frameworks proposed by Abraham et al. (2019) and others,
Bento et al. revealed the deficiency in addressing the sustainability aspect necessary for
continual program review, improvement, and value assessment within organizations.
Bento et al. proposed a framework of 12 CSFs for evaluating the sustainability of a DG
program. These factors include employee data competencies, well-defined data processes
and procedures, adaptable data tools and technologies, dtandardized user-friendly data
policies, established data roles and responsibilities, clear and inclusive data requirements,
focused and tangible data strategies, securing stakeholder buy-in, effective and strategic
communications, assessment of data governance status, and definition of sustaining
requirements. In conjunction with the conceptual frameworks, these CSFs offered
valuable perspectives for understanding IT managers’ strategies when implementing DG
programs. Alhassan et al. (2019) also proposed their set of CSFs for DG, further
contributing to the field’s knowledge base.
To address a gap in the field of DG, Alhassan et al. (2019) conducted a
comprehensive investigation using a single case study approach and incorporated
semistructured interviews with Al Rajhi Bank in Saudi Arabia. The primary objective
was to pinpoint the CSFs that underpin effective DG. Findings indicated that positioning
data as a valuable asset serves as the primary motivation for implementing DG. This
foundation drew upon the framework of Khatri and Brown (2010), encompassing five
key decision domains: data principles, data quality, data access, and data lifecycle. The
study uncovered seven core categories of CSFs, each prioritized by its level of
significance: (a) employee data competencies, (b) clear data process and procedures, (c)
flexible data tools and technologies, (d) standardized easy to follow data policies, (e)
established data roles and responsibilities, (f) clear inclusive data requirements, and (g)
focused and tangible data strategies. Notably, the success of DG programs hinged on
employee data competencies and established data roles and responsibilities, making them
critical factors to consider. Although this study’s findings may not be broadly applicable
(i.e., not easily generalized), they provide valuable insights that informed the current
study, particularly leveraging these CSFs in the context of conducting interviews with IT
managers.
Machado et al. (2022) conducted a systematic literature review to investigate the
potential role of DG in sustainable development. Their systematic literature review
commenced with the definition of DG as a set of controls over data management aimed
at maximizing data value and minimizing associated risks. Sustainable development,
conversely, encompasses economic, environmental, and social sustainability components
and strives to balance present consumption with future needs. The study revealed a gap
in existing DG frameworks, which focus on addressing regulatory antecedents (e.g., the
general data protection regulation [GDPR]) to ensure data consistency, trustworthiness,
decision-making accountability, and user privacy during implementation. Building on
Khatri and Brown’s (2010) data life cycle domain, Machado et al. proposed a
conceptual framework for sustainable DG, positioning DG within the product life cycle.
However, the proposed framework was conceptual and relied on a limited number of
articles. Future research is necessary to expand the existing body of literature, develop
models, and apply them in real-world use cases.
Data Ecosystem Model of DG
D’Hauwers et al. (2022) conducted three comprehensive case studies on Smart
Cities, focusing on the Rotterdam Digital Twin, the Helsinki Digital Twin, and the Smart
Retail Dashboard. Their primary objective was to develop a robust business framework
model that facilitates the efficient exchange of data assets among various stakeholders,
be they private or public entities, within or engaged with a data ecosystem. This endeavor
resulted in the formulation of the Data Ecosystem Model Framework, which hinges on
two pivotal factors: (1) the value of data assets (creation, revenue and cost model), and
(2) control (value network for relationships, and data governance) to meet the needs of
organizations. The significance of this Data Ecosystem Model Framework to our current
study lies in its ability to augment our understanding of strategies employed by IT
managers for data governance. Unlike conventional approaches that primarily address
data management within an organization, this framework broadens the scope to
encompass data exchanges that transpire across different organizations. Furthermore, it is
worth noting that the Data Ecosystem Model Framework effectively aligns with the
fundamental domains outlined in the Khatri and Brown (2010) framework, namely: data
principle, data quality, and data lifecycle.
Application of Khatri and Brown (2010) and Abraham et al. (2019) Governance
Frameworks to the Current Study
The aim of this section is to demonstrate how the two conceptual frameworks can
be applied as lenses to understand the IT managers motivations, challenges, and
strategies to implementing data governance. The following sections are included: data
governance definition and scope, historical overview, challenges, success factors,
frameworks, antecedents and data principles, governance scope (data and domain),
governance mechanisms, consequences, Big Data governance, and future research topics
in data governance. Firstly, the review begins by establishing a shared understanding of
data governance by defining its types and setting a common foundation for IT managers’
strategies in implementing data governance. It acknowledges that the term “data
governance” can have different meanings for different audiences. Secondly, a historical
overview of data governance is provided to offer insights into its evolution within the
academic body of knowledge. This section aims to equip readers with a comprehensive
understanding of how the subject has developed over time, up to the present day.
Thirdly, the challenges associated with data governance are addressed. This section
delves into the difficulties companies face when implementing data governance
programs, providing valuable background information. Fourthly, the review discusses
success factors, which serve as lessons learned and best practices in the implementation
of data governance programs. It recognizes that practitioners often combine industry best
practices to theoretical knowledge and design patterns to achieve successful outcomes.
Fifthly, data governance frameworks in the existing literature are examined. This section
explores how researchers approach the design, implementation, and evaluation of data
governance, offering insights into different theoretical perspectives. Sixthly, the review
explores the antecedents and data principles that justify why organizations engage in
implementing data governance in the first place. It aims to shed light on the reasons and
motivating factors behind organizations’ decision to invest in data governance initiatives.
Seventhly, governance domains are discussed. These domains encompass areas such as
data quality, data security, and compliance, where data governance is applied. This
section provides justifications for implementing data governance and highlights the
specific areas where it is crucial. Eighthly, data governance mechanisms and their
consequences are explored. Organizational aspects, roles and responsibilities, and
processes and procedures related to governance implementation are examined.
Additionally, the outcomes and benefits of data governance, such as improved data
quality and compliance, are discussed.
Data Governance Definition and Scope
As an organization asset data needs to be governed and properly managed to
boost competitiveness and help achieve strategic objectives. In recent years since 2010,
data has emerged as a valuable asset in the modern information economy. As both
private and public sector organizations have generated and consumed vast amounts of
data, the exponential growth of data and the increasing use of data-driven decision-
making
through business intelligence, companies have begun to recognize the importance of
managing and governing data as a valuable asset (Al-Ruithe et al., 2018). Data
governance (DG) plays a crucial role in managing data assets, alongside IT governance
and cloud governance, as part of the broader framework of IT information governance
and corporate governance (see Figure 4). As this study focuses on Digital Governance
(DG) as its primary area of investigation, it is crucial to establish a well-defined
framework for the literature review by clearly defining and outlining the scope of DG.
The Data Governance Institute (DGI) defines data governance as a system of
decision rights and accountabilities for managing data assets, determining who can
access and manipulate data and for what purposes (Al-Ruithe et al., 2019). Most scholars
agree that data governance involves the management and control of data assets (Alsousi
&
Shah, 2022; Mao et al., 2022; Al-Ruithe et al., 2018; Janssen et al., 2020). However,
Alsousi and Shah (2022) and Fernandes (2022) propose an expanded definition of data
governance that includes value creation and risk mitigation as additional dimensions.
This broader perspective implies that data governance should not only enhance the value
of data assets, such as driving innovation, but also reduce costs associated with data
assets, such as compliance failures and risks related to data handling (e.g., security
breaches, privacy violations, data leaks, and theft). These dimensions align with the
compliance, data quality, and data access aspects of Khatri and Brown’s (2010)
framework, as well as the antecedents and domain scope of Abraham, Schneider, and
Vom Brocke’s (2019) framework. Hikmawati et al. (2021) emphasized an often
overlooked but critical dimension of data governance, which is the need for
organizationwide initiatives to establish standardized processes and procedures for
managing and controlling data assets. This standardization and control enable alignment
between data management and business strategy, including regulatory compliance and
risk mitigation. In the context of the medical field, McCaig and Rezania (2021) defined
data governance as encompassing all policies, processes, and principles related to the
management, use, and security of data. This definition underscores the significance of
security considerations within data governance programs, similar to Khatri and Brown’s
(2010) data access dimension or Abraham et al.’s (2019) data domain dimension. Data
governance, as defined by various authors, revolves around several key themes that can
drive organizational success. It aims to enhance the value of data through improved
quality and trustworthiness, while also mitigating risks and reducing associated costs.
Compliance with laws and regulations is emphasized to avoid legal complications.
Importantly, data governance should be treated as an organization-wide initiative,
ensuring standardization of practices and promoting collaboration. Lastly, a key focus
lies in ensuring security and privacy for end users and data consumers, fostering trust and
loyalty.
In addition to the definition, it is crucial to distinguish between different
taxonomies of data governance. Abraham et al. (2019) proposed a framework that
includes a Data Scope dimension, which pertains to the type of data asset an organization
needs to govern: traditional data and big data. Similarly, Al-Ruithe et al. (2018) provided
a taxonomy of data governance, differentiating between traditional data governance for
non-cloud computing and cloud data governance for cloud computing. Traditional data
governance primarily focuses on data stored on-premises in structured data stores, such
as databases and data warehouses. It encompasses areas like Master Data Management
(Hikmawati et al., 2021; Haneem et al., 2019), which establish high-quality and trusted
reference data across various business functions within a company, including sales,
services, order management, purchasing, manufacturing, billing, accounts receivable, and
accounts payable. On the other hand, cloud data governance deals with Big Data stored
in flexible cloud stores, such as data lakes (Ramachandran et al., 2022; Yebenes &
Zorrilla,
2019). Traditional data governance is often associated with siloed implementation and an
IT focus, making it challenging to trust data outside of the organization (Al-Ruithe et al.,
2019; Kariotis et al., 2020). In contrast, cloud data governance faces complexities and
challenges inherent to the nature of Big Data, including volume, velocity, variety, and
veracity (Al-Sai et al., 2019). This study specifically concentrates on traditional DG
rather than Big Data DG. Consequently, the main emphasis of the literature review is on
traditional DG. The following paragraph offers a concise history of DG, providing
readers with a contextual understanding of the evolution of this topic within the realm of
research literature.
Figure 4
Interrelations Between Governance Domains (Al-Rhuite et al., 2018)
Brief History of Data Governance
The evolution of DG as a research topic mirrors that of data technologies starting
from relational databases in the 1970s, then datawarehousing and business intelligence in
the 1990s and 2000s, an finally the data explosion with Big Data and advanced analytis
since 2010. Data governance traces its early origins back to the emergence of relational
database management systems (RDBMS) in the 1970s (Al-Ruithe et al., 2019). The need
for structured data management practices became apparent with the introduction of data
models, dictionaries, and integrity enforcement rules. The mid-1990s saw a rise of
regulatory compliance requirements with federal data protection laws (Mulliganet al.,
2019) in regulated sectors such as healthcare, finance, and telecommunications. This
trend led organizations to develop data policies and procedures to ensure data privacy,
security, and reporting (Al-Ruithe et al., 2019). The next period is that of
datawarehousing and business intelligence.
During the late 1990s, as data warehousing and business intelligence gained
prominence, organizations began recognizing the strategic value of data. Consequently,
there arose a need to govern data assets, including master data, to ensure the quality
required for effective decision-making, such as accuracy, consistency, and reliability
(Khatri & Brown, 2010; Abraham et al., 2019). In the early 2000s, data governance
emerged as a distinct discipline within organizations, marked by the establishment of the
Data Governance Institute (DGI) in 2004. The DGI aimed to promote best practices and
frameworks like the Data Management Body of Knowledge (DMBOK), which offered
authoritative guidance in the field (Aisyah & Ruldeviyani, 2018). The digital economy’s
present phase is characterized by an exponential surge in data starting from 2010,
necessitating enhanced data asset governance.
The period since 2010 is marked by an unprecedented data growth worldwide
which is accompanied by increased concerns for data quality, privacy, and security and
justifying the development of additional data governance frameworks. The frameworks
proposed by Khatri and Brown (2010) and a decade later by Abraham et al. (2019) have
been a response to the gap and need. Moreover, the adoption and maturation of cloud
computing, along with the proliferation of Big Data generated by social media and the
Internet of Things (IoT), coupled with the automation and exploitation capabilities of AI,
have presented new challenges for data governance (Al-Sai & Abdullah, 2019). To
address these challenges and safeguard data quality, privacy, and security, organizations
must embrace Big Data and AI governance and research must fill the gap with Big Data
governance frameworks (Yebenes & Zorrilla, 2019; Zorrilla & Yebenes, 2022). This
ensures that both innovation and customer interests are protected (Al-Badi et al., 2018;
Al-Ruithe et al., 2018; Al-Ruithe et al., 2019). Data governance continues to evolve to
keep pace with challenges stemming from advancements in technology, changes in
regulations, and the ever-expanding range of complex use cases.
Data Governance Challenges
The mission of Data Governance (DG) is to effectively manage and control data
assets to maximize their value while minimizing associated costs and risks. This entails
addressing organizational, process, technological, and human challenges inherent in
managing and controlling data assets. The problem statement which is the genesis for
this study is that many companies fail to implement working data governance programs
due to challenges impacting IT managers strategies. This study offers valuable insights
into data governance implementation, benefiting companies and data governance
practitioners in multiple ways. Firstly, it sheds light on the factors and challenges
associated with data governance, providing organizations with a deeper understanding of
the key elements involved. This knowledge empowers companies to make informed
decisions regarding data governance strategies and implementation. Furthermore, the
study highlights emerging trends in the data governance field, enabling organizations to
stay up-to-date with the latest developments. By being aware of these trends, companies
and practitioners can proactively adapt their data governance practices to align with
industry advancements and best practices.
This literature review uncovered four categories of challenges for traditional DG
faces several challenges, and will be discussed in the following sub-sections:
•Methodological challenges: There is a lack of a comprehensive framework to
guide the implementation of DG practices.
•Technical challenges: DG often suffers from siloed or inadequate
implementation, technology limitations, data quality issues, and concerns
related to data security and privacy.
•Organizational challenges: DG efforts tend to be IT-focused and may lack
clear roles and responsibilities, well-defined data policies, and alignment with
the organization’s strategies.
•Compliance and regulatory challenges: DG must navigate the complexities of
compliance and regulatory requirements, ensuring that data management
practices meet legal and industry standards.
Methodological Challenges. Traditional data governance (DG) was initially
suffering a gap in research with a lack of frameworks to guide both research and
practitioners. Although traditional data governance (DG) has reached a certain level of
maturity, there are still gaps in the existing research that warrant further investigation.
Khatri and Brown (2010) were early contributors to the field of data governance
research. They noticed a lack of a framework to assist with DG implementation and thus
proposed a unified conceptual framework (referred to as the Khatri and Brown (2010)
Data Governance Framework in this study, depicted in Figure 1) to guide researchers and
practitioners in developing effective programs. The Khatri and Brown DG framework
encompasses five decision domains crucial for understanding and guiding DG
implementation: data principles, data quality, data access, data lifecycle, and metadata.
Over time, this framework has become a point of reference for future research. Nearly a
decade later, Abraham et al. (2019) found that despite the maturity of DG practices in
various industries, including healthcare, there was still a lack of a comprehensive
framework for data governance that could guide practitioners and researchers. In
response, they developed a conceptual framework (referred to as the Abraham,
Schneider, and Vom Brocke (2019) Conceptual Framework for Data Governance in this
study, illustrated in Figure 2) to evaluate data governance implementation based on six
dimensions of analysis: governance mechanisms, organizational scope, data scope,
domain scope, antecedents, and consequences of data governance. In the quest to enable
data-driven decision making in the evolving information economy and the emerging
Industry 4.0 era, Bento et al. (2022) explored how DG frameworks could best facilitate
this process. Similarly the study identified research gaps in data governance frameworks,
including the absence of a holistic implementation guide as challenges that cause DGs to
be unsustainable. To address this, the authors proposed a matrix of critical success
factors (CSF) that could enhance the success of DG programs: employee data
competencies, clear data processes and procedures, flexible tools and technologies,
standardized data roles and responsibilities, clear and inclusive data requirements,
focused and tangible strategies, stakeholder buy-in, effective and strategic
communications, and assessment of data governance maturity.
Over the past two decades, the body of knowledge on DG (Data Governance) has
significantly expanded, and several conceptual frameworks have been adopted as
references. However, it is widely acknowledged among researchers that more research
on DG is necessary. Specifically, there is a need for authoritative reference architectures
and best practices to guide the implementation of DG programs for DG and IT
practitioners. Implementing DG solutions presents challenges due to their technology-
stack dependency, high costs, and variations across organizations based on their unique
antecedents. This literature review identified Critical Success Factors (Alhassan et al.,
2019; Bento et al., 2022; Ibrahim et al., 2021; Yatya et al., 2021) that hold the potential
to provide a solid foundation for implementing DG. To leverage these factors effectively,
further quantitative research is required to validate their applicability across various DG
practices, industries, and the broader community of researchers. Once validated, these
findings could empower practitioners with valuable insights for successful DG
implementation. Enhancing the effectiveness of DG requires a comprehensive approach
that not only incorporates methodological frameworks with guidance and best practices
but also emphasizes resolving the inherent technical and organizational challenges that
frequently lead to the failure of DG implementations.
Technical Challenges. Traditionally, organizations, both business and
nonbusiness, have relied on their IT departments to manage their computing, networking,
applications, and data infrastructures. This has led to a strong IT-focused approach to
data management and data governance, often implemented in silos. Khatri and Brown
(2010) differentiated between IT assets (computers, networking, database systems) and
information assets (data), where IT assets facilitate the management of information
assets, and governance being concerned with the efficient use of IT. Al-Ruithe et al.
(2018) emphasized the interdependence of data governance and IT governance within the
broader concept of Information Governance (see Figure 1). They also identified IT-focus
and siloed implementation as significant factors contributing to the failure of data
governance initiatives. In contrast to Khatri and Brown (2010), Al-Ruithe et al. (2019)
added the complexity of business use cases for data as another factor affecting data
governance implementation. In healthcare, Kariotis et al. (2020) found that data
governance is often implemented in silos by different institutions, making it challenging
to establish trust across organizations. Similarly, Ngesimani et al. (2022) identified
technological challenges in the healthcare sector in South Africa, such as data quality
issues and limited competencies of IT professionals in operating electronic medical
record systems (EHR) and healthcare information systems (HIS), hindering the success of
data governance programs. In the public sector of China, Mao et al. (2022) highlighted
data quality issues caused by a lack of effective control over the data lifecycle (data entry,
processing, and reporting) as a barrier to successful data governance. Scope et al. (2022)
discussed compliance challenges faced by data governance due to technological
limitations in storage systems to enforce data privacy and security throughout the entire
data lifecycle (creation, storage, usage, archival, destruction). They suggested the need
for further research and development to integrate data governance into platforms natively.
Cloud Service Providers like Microsoft are playing a role in democratizing data
governance by offering tools such as Azure Purview, which require less financial
investment compared to traditional enterprise data governance applications. However,
companies still face challenges in complying with data governance regulations, such as
GDPR, which necessitate a strong understanding of how data is stored, used, and
managed. These challenges arise from issues related to data quality (inaccuracy and
incompleteness), IT-siloed environments (fragmented enterprise architecture and legacy
systems), and compliance with regulations (Abraham et al., 2019).
To address the IT-focus and siloed implementation challenges, companies are
encouraged to make data governance a company-wide initiative, supported by top
leadership, which recognizes data as a valuable business asset (Yatya et al., 2021).
Furthermore, data-driven transformation (Smith & Heffernan, 2019) requires considering
data as a strategic asset within the company and fostering a culture that values high data
quality and trustworthiness. This approach maximizes the value of data and reduces costs
associated with risks, both of which are outcomes of effective data governance.
However, a practical challenge lies in transforming a company’s culture and processes to
prioritize data as an asset, as simply producing transactional data does not equate to
recognizing its value. It requires executive direction to embed this mindset throughout
the organization.
Implementing data governance involves establishing organizational standards and
processes, which can be subject to various organizational challenges, such as IT-focus
and siloed implementation, lack of clearly defined roles and responsibilities, inadequate
data policies, and misalignment with organization-wide business strategies. The issues of
IT-focus and siloed implementation were discussed in the previous paragraph.
Organizational Challenges. As a vital component of Information Governance,
Data Governance (DG) is not an isolated process. It is heavily influenced by various
factors within an organization, such as its culture, strategy, roles and responsibilities, and
policies. Some of these factors can hinder the implementation and success of DG. Yatya
et al. (2021) emphasized the increasing importance of DG in the digital economy and its
crucial role in establishing a foundation for data architecture, security, and quality. By
enhancing the value of data assets and mitigating associated risks, DG contributes
significantly to overall organizational success. However, the authors highlight cultural,
political, and organizational resistance as key factors that can negatively impact the
effectiveness of a DG program. To ensure successful DG implementation, several best
practices have been identified. These include a phased approach to implementation,
support from senior management, clearly defined roles and responsibilities, goal
measurement using metrics, and early communication prioritization. Yatya et al. (2021)
also suggest two additional best practices: providing non-monetary internal rewards,
such as career growth opportunities, to encourage participation in the DG program, and
periodic reminders from a Data Steward about policies and new guidelines.
Aisyah and Ruldeviyani (2018) conducted a study on DG design and
implementation for the Indonesia Deposit Insurance Corporation (IDIC). They identified
challenges such as siloed processes, scattered data assets across business units, and a lack
of defined roles for data management within the organization. Another study by Alsousi
and Shah (2022) focused on Small and Medium-sized Enterprises (SMEs) and evaluated
the value of DG for these organizations. The researchers found that SMEs face specific
challenges in implementing DG, including poorly implemented data governance
frameworks, unclear definitions of duties and responsibilities, a lack of leadership and
organizational support due to perceived cost-benefit concerns, and external factors
beyond SMEs’ control. Several other studies, including those by Alhassan et al. (2019),
Almeida et al. (2019), Alsousi and Shah (2022), and Bento et al. (2022), highlight critical
success factors (CSFs) for DG programs. These studies emphasize the need for DG to be
a company-wide initiative and identify common organizational challenges such as an
ITfocused approach, siloed implementation, lack of clear roles and responsibilities, and
insufficient executive and company-wide support. The discussion of these CSFs will be
further explored in the Data Governance Success Factors section. Companies still
encounter challenges in achieving compliance and effectively implementing DG
practices across industries including healthcare.
Regulatory Compliance Challenges. In regulated industries such as healthcare
and finance, Data Governance (DG) plays a crucial role in protecting patient or
consumer records from misuse and abuse. Tiffin et al. (2019) recognized the benefits of
digital health in developing countries, but also highlighted the prevalent security and
privacy challenges surrounding consumer health data in those regions. The healthcare
sector in developing countries often faces weak or non-existent privacy laws and limited
security implementation. To address this gap, Tiffin et al. (2019) proposed a
comprehensive DG framework with four dimensions (Figure 5): (1) Ethics and consent,
(2) Data access protection, (3) Documentation, backups, accessibility, (4) Legal
framework.
Mahanti (2022) emphasized compliance as a significant driver for data
governance, extending beyond the healthcare sector discussed by Tiffin et al. (2019).
Non-compliance to governance regulations comes at high cost;therefore, the essential
role of data governance in mitigating non-compliance risks. With the exponential growth
of data since 2010, the potential for malicious data breaches has also increased. In 2021
alone, businesses across industries were fined a total of $957.3 million for compliance
violations, with the largest fine of €746 million imposed on Amazon under GDPR.
Technological challenges, such as data breaches, contribute to compliance difficulties,
along with the cumbersome and costly requirements of consumer privacy protection
regulations like GDPR. Mahanti (2022) recommended that compliance should be a
collaborative effort involving all stakeholders, including business, IT, and other
departments with a vested interest in data compliance. Mulligan et al. (2019) also
examined consumer data protection and privacy, providing an overview of US federal
data protection laws and comparing them to the EU’s General Data Protection
Regulation (GDPR) and the California Consumer Privacy Act (CCPA) of 2018. These
regulations apply to companies involved in collecting, storing, distributing, or selling
information and grant consumers three main rights: the right to know, the right to opt
out, and the right to delete. While the CCPA mirrors some aspects of GDPR, the latter
provides more extensive protection for consumer personal data. The study also covered
federal data protection laws such as the Gramm-Leach-Bliley Act (GLBA) of 1999,
Consumer
Financial Protection Bureau (CFPB), Health Insurance Portability and Accountability
Act
(HIPAA), Fair Credit Reporting Act (FCRA), and Consumer Financial Protection Act
(CFPA). These laws establish governance standards for how consumer personal data
should be handled by both data processors and controllers.
Acetedenents and Data Principles
Motivating factors often lead organizations to invest in new information systems
(IS) and implement policies or procedures. These motivations are often driven by the
need to support business functions, ensure regulatory compliance, and meet industry
standards. Antecedents and data principles work the same way when it comes to DG.
The antecedents outlined in Abraham, Schneider, and Vom Brocke’s (2019) Conceptual
Framework for Data Governance align with the data principles in Khatri and Brown’s
Unified Framework for Data Governance (2010). Both frameworks identify the internal
and external factors that influence the adoption and implementation of data governance
(DG). Implementing DG programs has traditionally been motivated by regulatory
compliance and data-related issues such as poor data quality, limited access, and lack of
trustworthiness (Vojvodic & Hitz, C. 2022). Mahanti (2022) emphasized the cost of
noncompliance to governance regulations and highlights the crucial role of DG in
mitigating non-compliance risks. Since 2010, the exponential growth of data volumes
and velocity has increased the potential for malicious data breaches. In 2021 alone,
businesses across industries faced fines totaling $957.3 million for compliance
violations, including the record-breaking €746 million GDPR fine imposed on Amazon.
This study aims to understand IT managers strategies for implementing DG. Those
strategies stem from specific motivations, let it be regulatory compliance with respect to
data, data quality, efficiency, or other organizational needs. Additionally, this study
underscores that regulatory compliance is a significant driver and antecedent for DG,
necessitating the establishment of good data quality practices. Achieving compliance
should be a collaborative effort involving all stakeholders, including business, IT, and
other departments with a vested interest in data compliance. By fostering good
governance on specific domain scope and maintaining high data quality standards,
organizations can streamline the attainment of compliance objectives.
Domain Scope
Data Governance aims to effectively control and manage data assets in specific
domains. DG decisions are applied to various aspects of data assets, such as data quality
and security. Abraham, Schneider, and Vom Brocke (2019) proposed a conceptual
framework for Data Governance that includes three dimensions for data assets: Data
Scope, Organization Scope, and Domain Scope. Data Scope refers to the different types
of data assets, distinguishing between traditional data assets stored on-premises in
relational databases and data warehouses, and cloud-based data assets including
relational, non-relational, unstructured, semi-structured, and Big Data. Organization
Scope determines the unit of analysis for DG, distinguishing between intra-
organizational scope (within a single organization) and inter-organizational scope (across
multiple organizations). Domain Scope corresponds to specific domains of decision-
making that are subject to DG mechanisms, including data quality, data security, data
architecture, data lifecycle, metadata, data storage, and infrastructure. These decision
domains align with the dimensions of Data Quality, Data Access, and Metadata in the
Khatri and Brown (2010) Data Governance Framework. Data quality and data security
are essential pillars of DG. They are crucial for decision-making through data analytics,
as well as for regulatory compliance and risk mitigation. IT managers motivations to
implement DG may stem from the need to improve data quality, trustworthiness,
security, or data infrastructure and processes. Companies employ a combination of
governance mechanisms to control and manage activities within each decision domain of
DG. These mechanisms help ensure that data assets are effectively governed and utilized
in accordance with organizational goals and regulatory requirements.
Governance Mechanisms
Governance mechanisms refer to the formal structures, processes, and procedures
employed to manage a data governance program. Scholars have categorized governance
mechanisms into three groups: structural mechanisms, procedural mechanisms, and
relational mechanisms (Abraham et al., 2019; Khatri and Brown, 2010). These
mechanisms encompass various aspects such as data management, decision-making,
roles and responsibilities, monitoring, and collaboration among stakeholders. Structural
mechanisms encompass reporting structures, governance bodies, and defined roles and
responsibilities. Typical roles within a data governance program include executive
sponsor, data governance leader, data owner, data steward, data governance council, data
governance office, data producer, and data consumer (refer to Table 4). Procedural
mechanisms primarily focus on effective data management, including data acquisition,
security, and utilization. They involve developing data strategies, policies, standards,
processes, procedures, and handling data-related issues. Relational mechanisms aim to
foster collaboration among stakeholders throughout the daily operations of the data
governance program. These mechanisms include communication, training, and
coordination of decision-making. Communication plays a crucial role in raising
awareness of the program among stakeholders and the entire organization. Training
ensures that stakeholders possess the necessary skills and knowledge for successful
program implementation. Analyzing governance mechanisms in a specific DG
implementation offers valuable insights into both the strategies employed by IT
managers and the challenges they encounter during the process. By implementing and
effectively executing governance mechanisms, positive outcomes or consequences can
be achieved for data governance, such as improved data quality and controlled access to
data assets.
Table 4
Data Governance Roles and Responsibilities
Role Responsibility
Executive sponsor Strategic direction, business prioritization and alignment,
funding
Data governance
leader
Day to day management of the DG program, coordinates tasks
of data stewards, reports on program performance
Data owner Line-of-business executives accountable for data assets of
their respective business units
Data steward Business leaders or subject matter experts with detailed
knowledge of business and data requirements; translate
requirements into technical specifications
Data governance
council
Hierarchy with cross-functional body, establishes the strategic
vision of the entire governance program and aligns it with
organization goals and strategy
Data governance
office
Staff that supports governance and decision-making activities
as well as the data steward teams and governance councils
Data producer Creates data or aggregates data from third-party
Data consumer User of the data
Consequences of Data Governance
The primary outcomes sought out for DG are to add value to data assets while
reducing risks associated with their use. Abraham, Schneider, and Vom Brocke (2019)
proposed a Data Governance Conceptual Framework that includes various dimensions,
such as antecedents and consequences. On one hand antecedents, analogous to data
principles in Khatri and Brown’s Unified Framework for Data Governance (2010), refer
to the factors that influence the adoption of DG. Regulatory compliance and data-related
issues (e.g., poor quality, limited access, lack of trustworthiness) have traditionally been
key motivations for implementing DG programs. On the other hand, the consequences of
DG are the positive outcomes sought through its implementation, such as improved data
quality, enhanced data access, and strengthened data security (Yatya et al., 2021;
Hikmawati et al., 2021; Khatri & Brown, 2010). For instance, master data management
(MDM) serves as a traditional form of DG aimed at providing high-quality and
trustworthy reference data utilized across the organization (Ibrahim et al., 2021). Such
reference data mitigates the risks of errors in business intelligence and analytics,
ultimately leading to better decision-making. The primary rationale for data governance
(DG) is to establish policies and procedures that enable the exertion of authority and
control over data assets management. The ultimate goal is to maximize value while
minimizing risks. DG can be conceptualized using an input/output model, comprising
influential factors and desired outcomes. Understanding the consequences of Data
Governance (DG) is crucial in evaluating the success of its implementation and
determining if it aligns with the motivations that drove IT managers to adopt these
programs initially. By comprehending the outcomes of DG, one can gain insights into its
effectiveness and whether it fulfills the original objectives set by IT managers. Numerous
researchers in the field have identified success factors for DG, offering guidance to
practitioners during program implementation.
Data Governance Success Factors
Researchers and practitioners in the fields of Corporate Governance, Information
Governance, and IT Governance not only share a theoretical foundation but also
exchange best practices to guide the design and implementation of various governance
programs, including Data Governance. Alhassan et al. (2019) identified seven CSFs for
successful Data Governance programs, including employee data competencies, clear data
processes and procedures, flexible data tools and technologies, standardized data
policies, established data roles and responsibilities, clear data requirements, and focused
data strategies. Bento et al. (2022) found that organizational Data Governance programs
often lacked a holistic framework for sustainable implementation. As a solution, they
proposed a set of CSFs, which overlapped with the CSFs identified by Alhassan et al.
(2019) and added three additional factors: stakeholder buy-in, effective and strategic
communications, and regular assessment of data governance maturity. Daneshmandnia
(2019) focused on the influence of organizational culture on the effectiveness of
Information Governance in higher education, with Data Governance being a subset of it.
The study highlighted the positive impact of a culture of competition/result-orientation,
control/hierarchy, and trust, while information silos were found to have an adverse effect
on Information Governance. This aligns with the stakeholder support and strategic
communications aspects emphasized by Bento et al. (2022). Ibrahim et al. (2021)
conducted research on critical factors affecting master data quality, which serves as a
single reference for core data across an organization. They identified 19 factors grouped
into five categories: organizational, managerial, stakeholder, technological, and external.
Several of these factors overlapped with the CSFs identified by Bento et al. (2022) and
Alhassan et al. (2019).
In addition to Critical Success Factors (CSFs), researchers have explored various
best practices that significantly contribute to the success of DG (Data Governance)
programs.Yatya et al. (2021) conducted literature reviews on best practices for
implementing successful Data Governance programs. They emphasized the importance
of full stakeholder support, organization-wide cooperation, and adapting strategies to suit
different organizational cultures. The study also acknowledged internal factors such as
cultural, political, and organizational resistance that can hinder the success of a Data
Governance program. The five main best practices they identified were: starting with
small steps, acquiring support from senior management, defining clear roles and
responsibilities, measuring goals with metrics, and prioritizing early communication.
This aligns with the findings of Daneshmandnia (2019), Bento et al. (2022), and
Alhassan et al. (2019) regarding internal organizational factors, senior management
support, clear roles and responsibilities, maturity assessment, and communication.
Considering the overlap in the CSFs identified by different studies, practitioners
including IT managers can leverage this body of literature to implement effective Data
Governance programs. Conversely, the Critical Success Factors (CSFs) could aid in
comprehending DG strategies and challenges. Table 5 provides a unified synthesis of the
CSFs, which can serve as a practical guide for practitioners. The next section will focus
on discussing Big
Data Governance based on the selected literature from this study.
Table 5
Alhassan et al. (2019) Critical Success Factors for Data Governance
Critical success factor
Employee data
competencies
Data governance activities that involve human
action. Directly impact the defining, implementing,
and monitoring of data processes and procedures,
as well as data policies and data requirements.
Clear data processes and
procedures
Data policies in place which should be detailed and
operationalized during activities related to data
processes and procedures.
Flexible data tools and
technologies
All the activities related to software and hardware
that affect the data in an organization, including the
presentation and storage of data.
Standardized easy-tofollow
data policies
Data policies are short statements that provide the
high-level guidelines and rules necessary for
dealing with data. These relate to data principles.
Established data roles and
responsibilities
Individual(s) responsible for the data-related
activities in the organization, who defines the
policies and processes for the data as well as
assigning the duties for the actions related todata.
Clear inclusive data
requirements
Define all aspects of data implementation, suchas
data flows and integration.
Focused and tangible data
strategies
Include planning fordata governance in order to
achieve its goals, as well as the main activities
related to considering data as assets.
Detail
Other Data Governance Models and Future Topics
Since 2010, DG research has undergone significant development, transitioning
from a significant gap to a mature body of knowledge. Numerous conceptual frameworks
have emerged in this field, providing valuable references for subsequent studies. Among
these frameworks, Khatri and Brown’s foundational framework (2010) has served as a
reference for scholars, including Abraham, R., Schneider, J., & Vom Brocke, J. (2019).
In my study, these two frameworks have been extensively discussed within the
Theoretical or Conceptual Framework section. Alsousi and Shah (2022) conducted a
comprehensive literature review to identify factors influencing the adoption of DG by
small and medium-sized businesses (SMEs). Their study proposed a conceptual
framework that encompasses six domains: Technology Context, Organizational Context,
External Task Environment Context, Technology Use, Data Quality, Organization Scope
and Domain Scope dimensions from the framework proposed by Abraham, R.,
Schneider, J., & Vom Brocke, J. (2019). Furthermore, Al-Ruithe et al. (2018) introduced
a conceptual framework aimed at establishing a unified perspective for comprehending
data governance in both traditional and cloud environments. They outlined a taxonomy
consisting of two dimensions: (1) traditional (non-cloud) data governance, which
encompasses people and organizational bodies, policy and processes, and technology;
and (2) cloud data governance, which encompasses data governance structure, policy and
processes, cloud deployment model, service delivery model, cloud actors, organizational
and technological aspects, service level agreements, monitoring matrices, and legal
considerations. Given the increasing prevalence of cloud computing and big data,
companies are progressively migrating their data assets, infrastructure, and processes to
cloud platforms, thereby elevating the significance of cloud DG as a highly compelling
topic. Overall, these conceptual frameworks, along with the insights gained from
research conducted by Alsousi and Shah (2022) and Al-Ruithe et al. (2018), contribute to
a comprehensive understanding of data governance in various contexts, including SMEs,
traditional environments, and cloud-based settings across industries and sectors.
Several DG frameworks have been proposed in various domains, such as
government (Girard, 2020; Mao et al., 2022) and master data management (Hikmawati et
al., 2021; Haneem et al., 2019). Jimenez and Duarte (2019) focused on Data Governance
implementations for business organizations, identifying the a crucial role of DG in
managing data assets to increase profitability, reduce costs, and mitigate risks. The
complexity of data use cases necessitates effective data governance.
Focusing specifically on Canada, Girard (2020) proposed a data governance
framework to assist business leaders in aligning their organizations with digitization and
AI. This framework comprises twelve dimensions, which share similarities with Khatri
and Brown’s (2010) framework. The dimensions include data governance objectives,
scope (historical vs. archived), roles and accountabilities, data ownership rights, data
collection, data access and lifecycle, data analytics use cases, data residency, privacy,
ethics and trust, approval and implementation, as well as compliance and certification.
Other frameworks have also addressed data quality, particularly in the context of master
data management (MDM).
The primary objective of Master Data Management (MDM) is to establish a
centralized repository for master data, encompassing customer, supplier, material, and
employee data, which can be utilized by all organizational processes, including Business
Intelligence. Hikmawati et al. (2021) conducted a literature review to explore the role of
MDM in enhancing organizational data quality (DQ) and data governance (DG). MDM
is particularly sensitive to data quality, as data quality issues are inherent within
organizations due to the diverse sources and varying standards through which data is
generated and managed. Poor data quality can lead to detrimental consequences such as
duplication, inaccuracy, and inconsistency of information. The MDM process primarily
involves data profiling, standardization, consolidating master data into a unified
repository, integrating it with existing applications, and finally synchronizing it with
business processes and applications. DG aims to establish ownership, define roles and
responsibilities, and establish rules for data management. The ultimate objective is to
ensure that data is reliable, trustworthy, and delivered in a timely manner. MDM and DG
are inter-related as MDM facilitates enhanced data management and quality through the
roles and responsibilities defined by DG.
MDM use cases for government sector call for more research. Haneem et al.
(2019) focused on identifying the key factors influencing the adoption of MDM by local
governments in Malaysia. Utilizing the Technology-Organization-Environment (TOE)
framework, they developed a conceptual model that highlighted factors such as
complexity, top management support, technological competence, and citizen demand as
influential in the adoption of MDM. Additionally, Mao et al. (2022) recognized the
increasing need for DG within China’s public sector and the lack of effective controls
over data entry, processing, and reporting in data asset management. To address these
issues, they proposed a robust data governance framework (DGF) tailored for the public
sector. The government DGF encompasses technical aspects of data governance as well
as organizational challenges such as defining roles and responsibilities, establishing data
policies, and aligning with organizational strategies. Using a design science approach,
the study proposed a tiered structure for the government DGF, consisting of public
services departments as the business frontend, the administrative department as the
backend, and a data middle platform as a bridge between the two, enabling reusable data
governance.
Implementing a government DGF ensures control and expansion of core organizational
data assets associated with administrative processes and systems. In the following
paragraphs I am discussing other DG models such as ecosystem data governance,
sustainable data governance, self-governance through intermediaries, trust-based data
governance, and security and privacy-focused data governance.
Ecosystem Data Governance Model
With the rapid expansion of data, its significance as a valuable business asset has
grown, resulting in the necessity for effective data governance. Lis and Otto (2021)
aimed to investigate the key characteristics and attributes of data governance within
ecosystem settings. They defined data ecosystems as independent organizations that
engage in data sharing to drive business innovation and implement use cases. As data
ecosystems continue to proliferate, it becomes crucial to establish data governance
frameworks to bridge existing gaps. However, implementing data governance in such
ecosystems can be challenging due to varying practices and standards across different
parties. The study proposed a three-layer taxonomy consisting of eight dimensions: the
data layer (data ownership, decision rights), the governance layer (configuration,
structure, mechanism), and the interaction layer (purpose, scope, phase). Limitations of
the study are notably the scarcity of research in this area and the absence of a universally
accepted definition of a data ecosystem.
Sustainable Data Governance Model
The explosive growth of data, commonly referred to as Big Data, presents
significant challenges in terms of acquisition, processing, storage, and secure
management, while also ensuring user privacy. In a recent study by Machado et al.
(2022), a systematic literature review was conducted to explore the role of Data
Governance in sustainable development. Data governance was defined as the
implementation of controls over data management to maximize its value and mitigate
associated risks. Sustainable development, on the other hand, aims to promote current
consumption without compromising future utilization, encompassing economic,
environmental, and social sustainability. The literature review revealed a gap in existing
data governance frameworks, which primarily focus on addressing regulatory aspects
such as GDPR to ensure data consistency, trustworthiness, decision-making
accountability, and user privacy. To bridge this gap, the study proposed a conceptual
Data Governance framework for Sustainable Data Governance, integrating data
governance practices within the product lifecycle. It is important to note that the
proposed framework is conceptual and relies on a limited number of articles,
highlighting the need for future research to expand the existing body of literature.
Furthermore, the development of models and their application in practical use cases
should be considered to enhance the framework’s effectiveness and applicability.
Self-Governance Through Intermediaries Model
Traditional DG within an organization is primarily self-governance whereby the
organization establishes its own standards and controls over data to achieve given
compliance requirements. According to Khatri and Brown (2020), data quality is
essential and can be ensured through core requirements such as accuracy, credibility, fit
for purpose, quantifiability, and relevancy. Medzini and Levi-Faur (2023) investigated
the Self-governance approach to Data Governance (DG) by focusing on credibility, using
a proxy in three use cases. The findings of their research indicate that credibility can be
established through third-party data validation or certification, which serves as a robust
foundation for regulatory compliance.
Trust-Based Data Governance Models
Most traditional data governance models in the existing body of research are
control-based as opposed to trust-based. Van der Sloot and Keymolen (2022) extensively
explored the concept of trust-based data governance models as a viable alternative to
control-based models. They examined two main existing models of data regulation
present in the literature: the control model and the legal standards model. The control
model grants individuals full rights to their own data, but its enforcement is challenging,
and it can impede business operations and hinder innovation. On the other hand, the legal
standards model relies on government enforcement of data regulations, with an
independent organization responsible for administering sanctions and enforcing rules.
However, this model places a significant burden on the government and often focuses
primarily on larger data processing companies.
In contrast to these two models, the authors proposed a third model based on
trust, which involves a trustor and a trustee. In this trust-based model, a third-party entity
is entrusted with assisting individuals in making decisions about their data. Three
trustbased data management models were outlined: information fiduciaries, data curators
and stewards, and data trusts. Each of these models has its own advantages and
disadvantages.
In the case of information fiduciaries, organizations are entrusted with managing
individuals’ data, and they must earn and maintain loyalty and trustworthiness. However,
a challenge arises in determining who monitors the fiduciary and ensures that they act in
the best interests of the citizens. The data curator, data custodian, and data steward model
involve assigning specific roles within an organization to enforce and monitor data
quality and privacy. For example, a data protection officer (DPO) may be employed by
the data processing organization, which acts as the trustee. While this model offers some
benefits, it also raises questions about potential conflicts of interest and independence.
Lastly, the data trust model entails the establishment of a dedicated institution
responsible for governing data on behalf of customers and acting in their best interests.
This model aligns well with a data-driven culture, but its implementation requires careful
consideration and planning. Most countries currently employ a combination of regulatory
models for data governance, typically involving informed consent and governmental
regulation. However, the trust-based models proposed in their study offer an alternative
approach, with the data trust model being particularly well-suited to a data-driven
culture.
Consumer Health Data Security and Privacy
Digital health benefits in developing countries have been well-documented in the
health literature; however, it is important to acknowledge the associated risks. Tiffin et
al. (2019) conducted a study focusing on the challenges of consumer health data security
and privacy in developing countries’ healthcare sectors. These challenges arise due to the
nascent stage of security implementation and the absence of robust privacy laws in these
regions. To address the existing data security and privacy gap, the authors proposed a
comprehensive DG framework, as illustrated in Figure 5. The framework is built upon
four essential pillars:
•Ethics and Consent: This pillar emphasizes the need to prioritize ethics and
consent in digital health initiatives. It involves identifying vulnerable
populations, obtaining informed consent, and transparently disclosing the
procedures involved.
•Data Access Protection: Procedural oversight and structural control are
critical aspects of safeguarding data access. This pillar focuses on
implementing mechanisms to protect data from unauthorized access. It
encompasses establishing oversight processes and maintaining structural
controls to ensure data privacy and security.
•Documentation, Backups, Accessibility: This pillar highlights the importance
of maintaining comprehensive documentation, backups, and ensuring
accessibility to health data. Proper documentation facilitates accountability,
while backups ensure data resilience. Additionally, ensuring accessibility
enables effective utilization of health data for improved healthcare outcomes.
•Legal Framework: The final pillar emphasizes the need for a robust legal
framework to protect individuals’ right to privacy and ensure data security. It
also addresses the oversight of third-party access to health data, promoting
responsible data handling practices.
By incorporating these four pillars, the proposed DG framework aims to bridge
the existing gaps in data security and privacy in developing countries’ healthcare sectors.
Implementation of this framework can help enhance the benefits of digital health while
mitigating potential risks and ensuring the protection of individuals’ privacy and
security.
Figure 5
Pillars of Digital Data Governance to Ensure Protection of Individuals (Tiffin et al.,
2019)
Big Data Governance
Although the primary focus of this study is traditional data governance, the
inclusion of Big Data governance in the literature search is essential for several reasons:
1) ensuring completeness, 2) acknowledging the growing use of Big Data assets across
industries, and 3) recognizing the expanding body of knowledge on Big Data and its
need of governance. These aspects informed the literature search strategy for this study.
In the subsequent sections, I will present a concise overview of Big Data, encompassing
its definition, challenges, opportunities, and governance frameworks.
Big Data Definition. Over the past decade, there has been an unprecedented data
explosion driven by social media, the Internet of Things (IoT), and transactional business
processes giving rise to Big Data. This surge in data has necessitated the development of
specialized frameworks for Big Data governance. Big Data refers to extremely large and
complex data sets that require unique management, processing, and storage capabilities
(Jansen et al., 2020). These data sets are characterized by three key attributes: Volume
( vast amounts of information, often measured in terabytes, petabytes, or even exabytes),
velocity (generated at high speeds, e.g., streaming data, requiring real-time or near-
realtime processing to enable continuous processing and storage), variety ( various
formats and types, including structured data, semi-structured data, and unstructured data.
The literature review revealed a lag in Big Data governance research in general
and the lack of data governance frameworks in particular. Yebenes and Zorilla (2019)
identified a lack of conceptual frameworks to guide Big Data governance research and
implementation. To effectively leverage Big Data for advanced analytics, such as
business intelligence, machine learning, and AI, it is crucial that the data is timely,
consistent, reliable, trustworthy, and of high quality. Therefore, the field of Big Data
requires governance frameworks to guide researchers and practitioners. Although Big
Data is being applied across various sectors, including healthcare, retail, education, and
government, there is still a research gap in the area of Big Data governance, primarily
due to the unique nature of these data sets. Yebenes and Zorilla (2019) proposed a
conceptual framework for Big Data governance, and a reference framework to guide Big
Data governance practice (Zorilla & Yebenes, 2022). However, challenges related to
privacy and security persist in protecting the consumers who generate this data and
contribute it to various platforms. In the following section, we will explore the benefits
and challenges associated with Big Data.
Big Data Benefits and Challenges. By its nature and complexity, Big Data sets
cannot be managed, process, or stored by traditional data processing systems. Numerous
authors covered in this literature review discussed opportunities and challenges of Big
Data including governance. Big Data brings the following benefits:
•Improved decision-making: Big Data analytics can help organizations extract
valuable insights to take informed and data-driven decisions, leading to
improved efficiency and competitiveness (Maniam & Singh, 2020; Micheli et
al., 2020; Smith & Heffernan, 2019; Bento et al., 2022; Donald et al., 2023;
Hummel et al., 2021).
•Enhanced customer understanding: Big Data analytics processing for
example wearable IoT data or social media data can gain a better
understanding of customer behavior and preferences, therefore, better target
marketing strategies and product/services sales.
•Innovation and new opportunities: Big Data opens up innovation for example
in healthcare industry, and AI.
•Operational efficiency and cost reduction: Big Data analytics and applied AI
can detect inefficiencies in processes and product defects to enhance both
efficiency and productivity.
•Fraud detection and risk mitigation: Big Data analytics can play a good role
in detecting and preventing fraud. For example, analyzing massive volumes
of network traffic data using AI can help detect fraudulent patterns, data
breaches, and cyber attacks.
This literature review revealed that the the biggest challenges of Big Data are 1)
the lack of data governance frameworks and best practices, 2) data quality (accuracy,
trustworthiness, reliability), 3) data security and privacy. Big Data use cases are ever
growing and penetrating many industries, health care (Chen 2020; Juddoo et al., 2018),
data lakes (Derakhshannia et. Al, 2020; Hlupić et al., 2022; Ramachandran et al., 2022),
corporate sectors like retail and telecommunications (Kastouni et al., 2022; Rasheed et
al., 2022), public sector and education (Fernandes, 2020). Big Data governance
challenges range from a variety of aspects:
•Data governance challenges: lack of governance frameworks and research to
guide practitionners (Yebenes & Zorrilla, 2019; Zorrilla & Yebenes, 2022).
•Data privacy and security: collecting and processing vast amounts of data,
from users, on social media, IoT (e.g. health wearables) or retail platforms
raise concerns for out privacy and security. Also protecting against sharing
sensitive and unauthorized access, breaches, and potential misuse are not
bullet proof for Big Data sets stored on a data lake, unless strong governance
and security are put in place (Achar et al., 2022; Al-Badi et al, 2018; Almeida
Teixeira et al., 2019; Chen, 2020; Cui, 2019; Donaldet al,2023; Hirsch et al.,
2020; Kariotis, et al, 2020; Machado et al., 2022; Maniam & Singh, 2020;
Mulligan, et al., 2019; Tiffin et al., 2019; van der Sloot & Keymolen, 2022).
•Data quality and reliability: Big Data is complex and if not processed
correctly can result in poor quality (incomplete, inaccurate, or inconsistent
data). Ensuring data quality and reliability is challenging and must be
addressed (Hikmawati et al., 2021; Hirsch et al., 2020; Ibrahim et al., 2021;
Juddoo et al., 2018; Karkošková, 2023; Khatri & Brown, 2010; Mahanti,
2022; Mao et al.,2022; Ngesimaniet al., 2022; Smith & Heffernan, 2019; van
der Sloot & Keymolen, 2022; Vojvodic & Hitz, 2022; Walsh et al., 2022).
•Data integration and compatibility: integration of data from various sources
and formats is challenging and requires savvy data engineering and
datascience skills.
•Skill gap and talent shortage: need for skills and expertise in data science,
statistics, programming, domain knowledge, and AI to process Big Data, and
extract godlen data.
•Ethical considerations: ethical questions arise from the collection of massive
user data without their consent and without transparency. Business decisions
made from biased data impact users.
Big Data Governance Frameworks. While the body of literature on Big Data
Governance is growing from non-existent, there is still a research gap in terms of data
governance frameworks that can effectively address challenges related to data quality,
security, and privacy. While several conceptual dimensions have been proposed to tackle
these issues, such as those suggested by Al-Badi et al. (2018), Al-Sai et al. (2019),
Derakhshannia et al. (2020), Janssen et al. (2020), Juddoo et al. (2018), Kastouni et al.
(2022), Maniam & Singh (2020), Osu & Navarra (2022), and Rasheed et al. (2022), two
comprehensive frameworks have been put forth by Yebenes and Zorrilla (2019) and
Zorrilla and Yebenes (2022). Yebenes and Zorrilla conducted a systematic literature
review to address the lack of a Data Governance framework for Industry 4.0. In the
context of Industry 4.0, which represents the fourth industrial revolution, data serves as a
valuable business asset and a fundamental component of software architecture with
dataas-a-service (DaaS) capabilities. Their literature review focused on third-generation
platforms (3GP) and provided them with a model to propose a conceptual framework of
Data Governance for Industry 4.0. In the context of 3GP, Data Governance is defined as
“a companywide framework for assigning decision-related rights and duties to
adequately handle data as a company asset.” While traditional Data Governance
primarily deals with structured data managed and stored on-premises within IT
structures, Big Data Governance framework must account for data-centric architecture
and the requirements of the Big Data era, including cloud platforms and Big Data
technologies.
The proposed Data Governance framework by Yebenes and Zorrilla aligns with
the Industry 4.0 environment and 3GP and consists of four pillars (see Figure 6): (1)
planning, which includes strategies, objectives, principles, policies, standards, and rules,
similar to traditional Data Governance; (2) organization, encompassing decision-making
bodies, decision rights, authority, and responsibility; (3) implementation or operation,
covering processes, people, and technologies; and (4) monitoring. Building on their
previous proposal of a Data Governance conceptual framework for Industry 4.0 based on
3GP (Yebenes & Zorrilla, 2019), the authors provide a detailed design and definition of a
Data Governance reference framework for implementing Data Governance systems in
the context of Industry 4.0 (see Figure 7). This reference framework is based on TOGAF,
ISO-42010, and RAS standards, reinforcing the principles of treating data as a business
asset that requires proper governance. The new Data Governance reference framework
includes a maturity model, a reference architecture, a standards Information Base, and an
architecture development method. Yebenes and Zorrilla’s (2022) reference framework
provides a detailed foundation for implementing Big Data Governance in the context of
Industry 4.0 for enterprises.
Figure 6
Data Governance Framework for Industry 4.0 (Yebenes & Zorrilla, 2019)
Figure 7
Data Governance Reference Framework for Industry 4.0 (Zorrilla & Yebenes, 2020)
Future Research Topics in Data Governance
This study focuses on investigating the strategies utilized by IT managers in
implementing traditional data governance (DG) and emphasizes the significance of Big
Data governance as a complementary aspect. The literature review places primary
emphasis on traditional DG while also acknowledging the growing importance of Big
Data governance due to the widespread adoption of Big Data in various industries and
the emergence of the Industry 4.0 era (Yebenes & Zorilla, 2019; Al-Badi et al., 2018;
Janssen et al., 2020). Therefore, understanding IT managers strategies for Big Data
governance would be a logical extension to the current study in the field of data
governance. This literature review has identified several emerging topics and potential
areas for future research. These include datafication, data markets, data sovereignty, data
cooperatives, data sharing pools, public data trusts, data ecosystems, governance of
Artificial Intelligence, and digital platform ecosystems governance. The subsequent sub-
sections offer an overview of these topics.
Datafication. The current era of the digital economy and the upcoming Industry
4.0 are characterized by a massive influx of data and the widespread adoption of Big
Data technologies across various industries, including healthcare and retail. In their
survey, Donald et al. (2023) explored the phenomenon of datafication, referring to the
explosion of Big Data and the increasing emphasis on data-driven transformation in the
digital economy and Industry 4.0. Datafication presents both opportunities and
challenges. On the one hand, it enables innovation through AI and advanced analytics,
leading to improved revenues and profitability for companies, as well as creating new
use cases for customers. On the other hand, it brings forth challenges such as the need for
effective data governance. Without proper data governance, the potential of data-driven
technologies like AI, machine learning, and the Internet of Things (IoT) can negatively
impact user privacy and security, accentuate biases, or result in job losses due to AI
automation. Building data exchanges and data markets are key components of data-
driven transformation. These frameworks facilitate the exchange of data between
different parties, enabling them to share, collaborate, and derive insights from diverse
datasets.
Data Markets. In the data economy, data is considered a valuable asset that
exists within a market ecosystem consisting of producers, consumers, and intermediaries.
Driessen et al. (2022) conducted a comprehensive literature review to address three
primary research questions. Firstly, they aimed to provide an overview of the existing
research on the design of data markets. Secondly, they sought to identify the application
domains where data markets can be utilized. Lastly, they aimed to establish a taxonomy
of use cases for data markets along with best practices. The ultimate goal of their study
was to contribute knowledge that can be utilized by software architects and engineers in
designing data markets. A data market, in essence, is a data platform that encompasses
both infrastructure and services to facilitate the exchange of data between providers and
consumers operating in diverse environments. The study identified five primary types of
data markets that are commonly implemented in various industries: generalist, specialist,
industry exchange, enabler, and aggregator. Governance and sovereignty are key
considerations when it comes to the exchange and utilization of data produced by an
organization as an asset, especially when involving third parties and end users.
Data Sovereignty. In the digital era of unprecedented data growth, increased data
exchanges among third parties, and the rise of data-driven technologies, addressing the
question of data sovereignty becomes crucial. Hummel et al. (2021) conducted a
literature review with the primary research goal of defining data sovereignty and its
characteristics. Data sovereignty pertains to the control of data within a specific
jurisdiction. It encompasses aspects such as control, ownership, and various claims on
data assets and infrastructures. Agents can include individual consumers, organizations,
entire societies, and even countries. Data storage and processing technologies such as
data mesh have an impact on the practical aspect of data sovereignty. As a decentralized
approach to emphasize domain-oriented ownership and governance (as opposed to
traditional centralized data systems like a data lake), a data mesh has the potential to
impact data sovereignty; however, it does not inherently determine who should be
responsible for the overall governance of the data shared in the mesh. Hummel et al.
(2021) stressed the need for further research to establish a comprehensive understanding
of data sovereignty, especially in relation to emerging technologies and changing data
storage paradigms. Similar to a data marketplace, a data ecosystem offers an optimized
framework for facilitating controlled data exchanges between multiple parties.
Data Ecosystem. Building upon the exponential growth of data, this research
emphasizes the significance of data as an asset and the necessity of data governance to
effectively manage data as an ecosystem asset. Data governance involves managing data
assets within an organization and facilitating data exchanges between organizations in an
inter-organizational context. Lis and Otto (2021) aimed to identify the key characteristics
and attributes of ecosystem data governance through their research. They defined data
ecosystems as autonomous organizations that engage in data sharing to implement
business use cases and drive innovation. With the increasing volume of data and data
exchanges, data ecosystems are emerging as a prominent phenomenon. Traditional data
governance practices have predominantly focused on intra-organizational aspects. The
rise of data ecosystems necessitates the establishment of data governance frameworks to
bridge this gap. Implementing data governance in such ecosystems poses challenges due
to the inherent uncertainty arising from varying data governance practices and standards
across different parties. The study proposed a three-layer taxonomy with eight
dimensions: the data layer (including data ownership and decision rights), the
governance layer (comprising configuration, structure, and mechanism), and the
interaction layer (encompassing purpose, scope, and phase).
D’Hauwers et al. (2022) also focused on the data ecosystem model but for a
specific use case, Smart Cities, to establish a data governance framework specifically
tailored to producers and consumers of data assets within that environment. The
researchers proposed the Data Ecosystem Model Framework, which revolves around two
key factors: (1) the value of data assets, including aspects of creation, revenue, and cost
models, and (2) control, encompassing the value network for relationships and data
governance to cater to the organization’ needs. While a data ecosystem framework
encourages the perception of data as a product, it does not necessarily establish a
foundation of trust among the participating entities.
Data Cooperatives, Data Sharing Pools, Public Data Trusts. Data governance
is a significant challenge in the age of datafication, and further research is required to
address this issue. In a literature review conducted by Micheli et al. (2020), four
emerging models of data governance were identified to regulate the access, control,
sharing, and utilization of data on various platforms:
1. Data sharing pools (DSPs) where data is treated as a commodity and shared
among multiple parties to promote data-driven innovation, create new
services, and provide value-added benefits to all stakeholders.
2. Data cooperatives (DCs) where data and associated rights are distributed
among various actors, similar to DSPs.
3. Public data trusts (PDTs) involves a public entity that accesses, aggregates,
and utilizes data about its citizens, including data held by commercial entities.
PDTs establish a relationship of trust with these entities to access the
necessary data.
4. Personal data sovereignty (PDS) provides data subjects with greater control
over their data in terms of privacy management and data portability compared
to the prevailing dominant model. In PDS, data subjects are considered key
stakeholders alongside digital service providers.
The existing data governance model is primarily controlled by large high-tech
corporations. Just like the General Data Protection Regulation (GDPR) (Mulligan et al.,
2019), the Personal Data Store (PDS) model enables individuals to have authority over
the utilization, sharing, and administration of their data. This model could serve as a
foundation for addressing the governance of artificial intelligence (AI) concerning its
exploitation of end-user data.
Governing Artificial Intelligence. In the era of a maturing data economy and the
advent of Industry 4.0, companies are focusing on harnessing the potential of Big Data
and AI innovation. However, the lack of governance in this field is becoming
increasingly evident. Taeihagh’s (2021) article addressed the pressing need and
challenges of governing Artificial Intelligence (AI). While AI holds the promises of
economic efficiency and improved quality, it also poses risks that can negatively impact
individuals and society as a whole. Concerns such as privacy, bias, unemployment
resulting from automation, and the allocation of responsibility and legal liability (in the
context of human control) being transferred to robots and AI systems, are among the key
issues at hand.
Digital Platform Ecosystems Governance. In today’s information age, the
remarkable rapid rise of digital social media platforms, such as Facebook, YouTube, and
Twitter, has enabled hundreds of millions of users worldwide to produce and exchange
content. However, these platforms lack systematic data governance, leading to
significant security and privacy concerns for their billions of users. While these
platforms effectively connect users and facilitate content creation, exchange, and storage,
their exponential growth has resulted in a lack of clear governance mechanisms. To
address the data governance gap, Hein et al. (2020) proposed a definition of a digital
platform ecosystem consisting of three components: the platform owner, an ecosystem of
autonomous consumers and commentators, and value creation mechanisms. Within this
framework, Une Lee et al. (2017) identified governance challenges related to data
platforms, such as unclear content ownership between the platform owner and end users,
insufficient transparency of data flow, and ethical considerations. The field of research
on data governance for digital platform ecosystems closely aligns with other areas like
Big Data governance (Janssen et al., 2020) and cloud governance (Al-Ruithe et al.,
2018). Une Lee et al. (2017) proposed a decentralized model based on blockchain
technologies as a new approach to designing data governance for platform ecosystems.
This decentralized data governance model has the potential to address the existing
challenges and provide enhanced security, transparency, and user control.
Transition and Summary
The Foundation of the Study section laid the ground for this study by specifying
the problem statement and main research question, identifying conceptual frameworks to
guide the research design, data collection, analysis, and interpretation, and finally a
review of academic literature to support the study. Two conceptual frameworks were
retained for this study: Abraham, R., Schneider, J., & Vom Brocke, J. (2019) Data
Governance Conceptual Framework, and Khatri and Brown (2010) Unified Framework
for Data Governance. The literature review has uncovered several factors influencing the
adoption and implementation of data governance. These include considerations related to
data quality, data security, data validity, data availability, regulatory compliance,
associated costs, and the necessary skillset. Additionally, it also explored various aspects
of governance strategies such as standards, best practices, and policies.The next section,
the Project, details the research design including population and sampling criteria, data
collection and analysis, and a discussion of reliability and validy of the study.
Section 2: The Project
Purpose Statement
The purpose of this qualitative pragmatic study was to explore strategies used by
IT managers to implement DG. The study aimed at understanding the factors that help or
hinder the implementation of DG by companies. The population included IT managers in
the United States at select companies that had successfully implemented a company-wide
DG process. The implications for positive social change are improved data quality and
data security for businesses and individuals, and the democratization of DG because of
the availability of applications at a low cost.
Role of the Researcher
To ensure the integrity of this study, I needed to acknowledge and address
potential sources of personal bias that could have influenced the research process. My
role in shaping the study’s design, data collection, and data analysis was paramount. Two
crucial considerations for researchers were addressed. First, I adhered to ethical
principles outlined in the Belmont Report Protocol (see Nagai et al., 2022; U.S.
Department of Health and Human Services, n.d.) to ensure the responsible conduct of the
study. Second, I identified and mitigated personal biases to prevent their negative
influence on the research process and conclusions.
Abiding by the Belmont Report Ethical Principles
In this study, I adhered to the ethical guidelines outlined in the Belmont Report.
The primary goal of this study was to expand knowledge and address important issues.
However, research can carry risks for individuals and society if researchers do not follow
essential ethical principles. The Belmont Report, which emerged from the 1976 Belmont
Conference, established three fundamental ethical principles that all researchers must
adhere to to safeguard individuals and society from any adverse consequences of research
(Nagai et al., 2022): beneficence, respect of autonomy, and justice. The beneficence
principle emphasizes the commitment to conducting research with kindness and without
causing harm. The respect for autonomy principle ensures that participants are free from
coercion and have the autonomy to choose to participate in or withdraw from research.
Finally, the justice principle upholds the principles of justice by guaranteeing equal
protection, opportunities, and burdens for all research participants.
In the current study, I was committed to upholding ethical principles. The
primary goal of this study was to gain valuable insights into the strategies used by IT
managers when implementing DG. The overarching aim was to provide valuable
knowledge to professionals in the field of DG. This commitment aligned with the ethical
principle of benevolence guiding this study. To honor the autonomy principle, I
integrated a consent form for each participant, granting them the autonomy to freely
engage in interviews and the freedom to withdraw from the study at any point. The
concept of informed consent in research is legal and regulatory in nature (Bazzano et al.,
2021). A consent form is a requirement for any study that engages subjects.
The principle of justice was upheld in the current study through meticulous
design measures, which prevented any form of discrimination among participants. In the
United States, federal and state antidiscrimination laws prohibit any discrimination based
on race, gender, religion, age, occupation (FTC.Gov, n.d.); therefore, avoidance of
discrimination is a legal and regulatory requirement for researchers. Responsible
research with human subjects must remain ethical by design and be conducted in such a
way that it avoids discrimination of participants. These measures ensure equal
opportunities for all individuals to participate in a study. Furthermore, the principle of
justice demands a biasfree study, one that does not adversely affect either the participants
or the study’s conclusions. Survey research is particularly vulnerable to researcher bias,
stemming from the framing of questions and interpretation of responses (Story & Tait,
2019). Interviewbased research shares the same researcher bias potential. I needed to
identify personal bias early and devise strategies to address it prior to and during the
study. I identify potential personal biases in the following sections and discuss strategies
aimed at mitigating their influence throughout the research process, including the
formulation of findings and conclusions.
Addressing My Personal Bias
Personal bias, stemming from my unique perspective, was identified and
mitigated to safeguard the quality and objectivity of the study’s outcomes, particularly in
the context of qualitative research in which I was involved as the primary data collection
instrument (see Jager et al., 2020, Mazhar et al., 2021). My individual background,
encompassing cultural, experimental, and personal dimensions, included biases, values,
and perceptions that may have impacted the interpretation of collected data. In this
pragmatic study, data collection took the form of interviews conducted with IT managers
to gain insights into their strategies for implementing DG. Throughout this process,
several potential sources of personal bias needed to be considered and managed. First,
my professional background had the potential to influence my perspective. For instance,
a technical background might have inclined me to view DG as a technical endeavor
confined to the realm of IT, rather than recognizing it as an organization-wide initiative
with carefully crafted strategies. Second, my personal familiarity with the topic of DG
may have introduced bias. Preconceived notions or a deep understanding of the subject
matter may have inadvertently shaped the interview protocol and data collection process,
potentially skewing the findings.
To address these potential biases, I implemented mitigation strategies. These
strategies included transparency, reflexivity, and methodological rigor. Transparency
involved openly acknowledging my background and potential sources of bias.
Reflexivity encouraged me to continually self-assess my role and biases throughout the
research process, making adjustments as needed to maintain objectivity (see Dodgson,
2019). Methodological rigor ensured that the research design and data collection
techniques were well grounded, reducing the influence of personal bias. Recognizing and
addressing personal bias was fundamental to the success and validity of this study. By
implementing robust mitigation strategies, I was able to navigate the challenges posed by
their unique perspective, ensuring that the study yielded unbiased and reliable
conclusions.
Background Bias
A researcher’s background, such as being a senior data engineer/architect with
over a decade of experience delivering data platform solutions to Fortune 500
companies, can shape their perspective. In this study, my extensive IT background might
have led me to view DG primarily through a technical lens, focusing on technical
implementation using existing software applications or custom solutions. To address and
mitigate this potential bias, I developed a multifaceted strategy relying on adherence to
conceptual frameworks, interview protocol, member checking, and data triangulation.
First, to prevent undue technical bias, I adhered to established conceptual
frameworks: Abraham et al.’s (2019) DG conceptual framework and Khatri and Brown’s
(2010) unified framework for DG. These frameworks provided a structured and
comprehensive approach to studying DG. Second, recognizing that I played a central role
in data collection, steps were taken to minimize interviewer bias during interviews. The
interview protocol was designed to be neutral and impartial. It was designed not to lead
interviewees toward technical aspects but instead encourage responses that addressed
business and organizational dimensions.
Third, to validate data collection, I employed member-checking interviews.
Member checking, a crucial step in qualitative research, involves verifying
participantsupplied data (from interviews or surveys) to ensure the accuracy of
transcriptions (Motulsky, 2021). Lastly, to enhance the credibility of the findings,
multiple data sources and methods were used. Triangulation (combining data from
interviews, surveys, and document analysis) provided a more comprehensive and
balanced view of the research topic. Relying solely on the interpretation of interview
data may have introduced researcher bias due to the data collection protocol and the
application of conceptual frameworks (see Jager et al., 2020). To enhance the credibility
of the current study findings, I supplemented the analysis with external sources such as
public or industry reports. These additional sources offered valuable validation from the
perspectives of DG practitioners and researchers. By implementing these strategies, I
aimed to conduct a rigorous and unbiased study on DG, ensuring that my IT background
did not influence the research outcomes.
Interview Protocol
I gathered data through interviews with participants, using well-designed
questions centered on the research question and two conceptual frameworks. Bias could
have been introduced through the interview questions and my manner. According to
Jager et al. (2020), many types of biases can be introduced during data collection, which
can have an adverse impact on the internal validy of the study, including selection bias
(sampling bias, evidence prevalence bias) and information bias (interviewer bias,
observer bias, lead time bias). To mitigate bias, I adhered to best practices in question
construction and interview delivery by using a well-established and vetted interview
protocol. The primary strategy was to employ an open-minded approach during
interviews, allowing IT managers to freely express their DG strategies without
influencing their responses. First, I crafted questions that aligned with the research
question and conceptual frameworks. Second, I avoided closed-ended questions that
would have yielded yes/no answers. Finally, I refrained from using leading questions.
This strategy was encapsulated in a rigorous interview protocol that I applied uniformly
to all participants (see Appendix B).
Researcher’s World View of the Topic of DG
My perspective may have shaped my approach to the research question and the
use of existing conceptual frameworks from the body of literature. I demonstrated a
strong understanding of IT managers’ strategies for DG and the chosen conceptual
frameworks. However, there was a potential bias in favor of emphasizing the data quality
dimension of the Khatri and Brown (2010) unified framework for DG, potentially
neglecting the broader framework. To address this potential bias, I incorporated
interview questions that covered all six data governance domains and underlying
attributes within the Khatri and
Brown (2010) framework, as well as those within the Abraham et al. (2019) framework.
This approach ensured a more holistic exploration of data governance. Furthermore, I
was mindful not to oversimplify traditional DG by conflating it with MDM. Although
MDM and DG contribute to improved data quality, MDM’s primary focus is on creating
high-quality and reliable reference data for the entire organization, whereas effective DG
may encompass MDM practices in certain instances but not vice versa. To ensure I
remained objective and free from potential researcher bias, I conducted the data
collection and analysis phases with meticulous rigor. This approach helped me maintain
the integrity of the study and its findings.
Data Collection and Analysis
The interview protocol served as the tool enabling me to gather input from IT
managers regarding their DG strategies. During the data collection and analysis phase, I
realized that researcher bias may have adversely impacted the internal validity of the
study and the study’s outcomes (see Jager et al., 2020). To mitigate this, I employed
various strategies such as participant member checking, transcript validation, and data
saturation checks. The trustworthiness, auditability, credibility, and transferability
(TACT) framework for rigor in qualitative research provides guidance on how to boost
trustworthiness, auditability, credibility and transferability of a study (Daniel, 2019).
Member checking (participant or respondent validation) involved transcript validation
and review in which I sent participants their transcript and allowed them to confirm or
correct it (see Motulsky, 2021). This method is highly effective in preventing
misinterpretation or inaccurate recording during interviews. However, it may extend the
overall time required, and participants might be unwilling to allocate extra time.
Additionally, biases or errors may arise during coding and analysis. To maintain the
integrity of the study, I ensured that coded data aligned with the original raw data
obtained from the interviews. In pragmatic studies relying on interviews, participants
play a pivotal role as the primary sources of interview response data.
Participants
Effective participant selection is vital in preparing for data collection in a
qualitative study employing an interview protocol. Poor selection of participants can
introduce selection bias that will negatively impacts the internal validity of the study and
conclusions (Jager et al., 2020). I prevented selection bias by establishing well-defined
eligibility criteria for participant inclusion and outlining strategies for recruiting
participants and nurturing meaningful relationships with them.
Eligibility Criteria
This study focused on IT managers at Fortune 500 companies in the United States
with an established DG program lasting 18 months or more. To qualify, IT managers
needed to have a minimum of 1 year of employment at the company, possess expertise in
DG, and be familiar with the company’s DG program. This ensured participants’
thorough understanding of the program’s implementation within the organization.
Strategies for Gaining Access to Participants
Gaining access to IT managers within companies can be challenging, but my
strategy was designed to efficiently connect with potential participants. The approach
consisted of two key steps:
1. Identify prospects: First, I focused on identifying potential participants within
professional networks and associations. This involved searching platforms
such as LinkedIn and relevant DG groups. The keywords I used for searches
primarily related to “data governance” in job titles or job functions.
2. Solicitation outreach: Once I identified these networks and potential
participants, I initiated contact by sending a well-crafted solicitation message
explaining the purpose of the study and outlining the selection criteria. This
outreach continued until an initial sample of 36 prospects was obtained.
The strategy in this study is purposive sampling selection. For the initial phase,
the goal is not to limit the sample to a specific number but rather to collect as many
potential candidates as possible. It should be understood that the sample size might vary,
and I remained flexible in this regard. The approach is be pragmatic, and I continued
gathering candidates until a point where data saturation was achieved.
In summary, my strategy involved identifying potential participants through
professional networks and associations, utilizing relevant keywords, and initiating
outreach to build a diverse initial sample. The precise sample size was determined
moving along, with a focus on achieving data saturation or a manageable sample size,
possibly around 12 participants.
Strategies for Establishing a Working Relationship With Participants
Before commencing the interview process, it is essential to establish a productive
rapport with potential participants. To accomplish this, I employed effective networking
techniques using platforms such as LinkedIn or send an email containing a concise
overview of the study, participant requirements, eligibility criteria, a request for
informed consent, and a commitment to upholding ethical standards (Johnson et al.,
2020) to ensure the well-being of both participants and their respective organizations.
Following the initial outreach, I sent an email to individuals who had expressed their
intent to participate in the study, along with the consent form. Once the prospective
participant signed and returned the consent form, indicating their willingness to
participate in the research, I collaborated with them to schedule interview dates and
times. Subsequently, I sent a scheduling invitation and awaited confirmation, while also
providing additional follow-ups and reminders as necessary. In cases where a
prospective participant did not responded, I gently remind them with one to three
reminders. When there was still no response, I discontinued contact. These follow-ups
were conducted via email or through LinkedIn chat. Another strategy to maintain a
working relationship with the participant was to extend a connection invitation on
LinkedIn and engage in discussions related to data governance to establish a shared
rapport on the subject. This interaction continued until after the member-check follow-up
interview was complete.
Research Method and Design
Method
In the “Nature of the Study” section, I chose a qualitative research method,
specifically a pragmatic inquiry, as the most suitable approach for this study. Unlike
quantitative research, which relies on systematic empirical methods to examine
phenomena through numerical data collection and analysis, qualitative research takes an
inductive approach, offering the advantage of delving deeply into the subject matter,
thereby facilitating the generation of fresh insights and theories (Lima &
NewellMcLymont, 2021). This research method is exceptionally well-suited for gaining
a comprehensive understanding of the data governance strategies utilized by IT
managers. In this study, we have opted for a qualitative approach with an exploratory
nature to investigate the “what,” “why,” and “how” behind the data governance strategies
implemented by IT managers.
While mixed-method research, which combines both quantitative and qualitative
elements, could potentially provide quantitative validation for the findings obtained
through qualitative inquiry (Lo et al., 2020), it was not considered in this study due to
time and resource constraints within the scope of our research. Quantitative research,
however, could be explored in a different context with a distinct research question. For
instance, it could be utilized to validate conclusions derived from a prior qualitative
study on data governance strategies. Such a quantitative study would be amenable to
generalization based on statistical sampling and analytical methods.
Research Design
Under the pragmatic inquiry approach, an exploratory multiple case study design
has been adopted, which facilitates a thorough investigation using participant interviews
as data collection method. Importantly, it provides the flexibility to study one or multiple
bounded cases without external interference, as highlighted by Takahashi et al (2020)
study. In this study, each company’s implementation of data governance constitutes an
individual case study, with one IT manager participating in the interview for each case.
The pragmatic inquiry design has been selected as the most appropriate approach
for this study due to its emphasis on gaining a profound understanding of the data
governance strategies employed by IT managers. This design, operating within the
pragmatic framework, facilitates a comprehensive investigation through various data
collection methods such as observations, interviews, and secondary data analysis. It
grants the freedom to explore specific cases or multiple instances without external
interference, as outlined by Takahashi et al. (2020). In this study, the objective is to
comprehend data governance strategies by engaging in interviews with IT managers
from various organizations, making it inherently a multiple case study.
Alternative research designs, such as ethnographic and phenomenological
approaches, were not chosen because they do not align with the primary research focus
on IT managers’ strategies in the realm of data engineering. Ethnographic research
excels in uncovering shared patterns within cultural groups (Gherardi, 2019), a
perspective that does not apply to the specific goals of this study. Likewise, the
phenomenological approach concentrates on individual life experiences (Greening,
2019), making it unsuitable for investigating data engineering topics like data
governance.
Data saturation is a pivotal concept in qualitative research, particularly in
methodologies such as grounded theory, ethnography, content analysis, and multiple case
studies. It signifies the point at which gathering additional data no longer yields novel
insights or information pertinent to the research questions or objectives (Shaheen &
Pradhan, 2019; Fusch & Ness, 2015; Sebele-Mpofu, 2020). This indicates that I have
acquired a sufficient amount of data to comprehensively comprehend the phenomenon
under investigation. The process can be summarized as follows, as depicted in Figure 8:
1. Initial Participant Selection: Commence with a pool of 36 prospective
participants.
2. Interview and Transcription: Select one participant and schedule a live
interview. Following the interview, transcribe it meticulously.
3. Member-Checking Review: Arrange a member-checking review with the
interviewee after transcribing the interview.
4. Theme Extraction and Comparison: Analyze the transcript to identify and
extract themes. Compare these themes with the existing list of themes in the
research topic.
5. Decision-Making Point: Make a decision based on the comparison. If new
themes are identified, repeat the process by selecting another participant from
the list. If no new themes are discovered and the number of processed
candidates is less than 12, continue the process. If data saturation is achieved
(i.e., no new themes emerge, and at least 12 candidates have been processed),
terminate the data collection process.
Figure 8
Data Saturation Process
The decision to involve a minimum of 12 participants for a multiple case study
was made to ensure an adequately sized sample for comprehensive analysis. Detailed
considerations regarding data saturation are elaborated upon in the Population and
Sampling section of this study.
Population and Sampling
Population
This study focused on IT managers within Fortune 500 companies in the United
States. We aim to assemble a final sample with a minimum of 12 participants, each
representing a different company. To be eligible, IT managers were required to possess a
minimum one-year tenure with their respective companies, demonstrate expertise in data
governance, and have a deep understanding of their company’s data governance
program. Additionally, a pre-requisite for the company is that it must have a well-
established data governance program in place for at least 18 months. In the Participants
section, we outlined the approach to identify potential participants, who were primarily
sourced from professional networks like LinkedIn and Data Governance groups. We
initiated contact by sending out a solicitation, and upon their response, followed up with
a formal email containing a consent form. This form was essential, as it ensured their
voluntary participation in our data collection interviews. To build a rapport with
participants, I employed strategies such as establishing a LinkedIn connection or
engaging in professional discussions related to data governance topics. These
interactions helped create a comfortable working relationship. Subsequently, I
coordinated with each participant to arrange a live interview at a mutually convenient
time. These interviews were conducted using platforms Microsoft Teams for a seamless
experience. After the interview sessions, I transcribed the content and scheduled a
member-checking interview with each participant. This step aims was to validate the
accuracy of the transcripts against their initial responses, further enhancing the overall
data quality (Motulsky, 2021) Sampling
The study employed purposive sampling with a size of 12, as recommended for
qualitative research or until data saturation is achieved (Shaheen & Pradhan, 2019).
Purposive sampling enhances data validity and trustworthiness. Nevertheless, it is
essential to mitigate sampling bias (Johnson et al., 2020) and ensure sample diversity to
adequately address the research question and achieve data saturation. I interviewed 11
participants, each with over ten years of experience in data governance across various
industries, including retail, pharmaceutical, travel and hospitality, healthcare, and finance
(see Table 6).
Table 6
Participants and Industries
Participant DG experience in industries
Participant 1 IT, health
Participant 2 IT, pharmaceutical
Participant 3 Banking, financial services
Participant 4 IT
Participant 5 Oil and gas, tech real-estate marketplace, health insurance
Participant 6 Pharmaceuticals
Participant 7 IT
Participant 8 Pharmaceuticals
Participant 9 Manufacturing
Participant 10 Financial services and banking, manufacturing, insurance, retail
Participant 11 Travel technology, tech real-estate marketplace
I initially chose purposive sampling in my research design to select target
participants from LinkedIn, ensuring a good mix of industries. However, enlisting
volunteers proved challenging as potential participants did not commit. To address this, I
switched to snowball sampling, starting with one or two participants, interviewing them,
and obtaining additional participants through their referrals. Snowball sampling is a
nonprobability technique often used in qualitative research, particularly with hard-to-
reach or hidden populations, relying on referrals to generate new subjects. This method
was highly effective. Once a participant accepted the invitation, I had them agree to a
consent form sent via email. Subsequently, I conducted semi-structured interviews over
Microsoft Teams to gather data on their experiences implementing Data Governance
programs, following the interview protocol I designed. To enhance data reliability, I
followed the interviews with member-checking. Data collection was followed by data
analysis using the classic Data Analysis Spiral strategy in qualitative research.
I employed the Data Analysis Spiral strategy for data analysis, which involved
the following steps: (1) Managing and organizing data, (2) Reading in-depth and
memoing emerging ideas, (3) Describing and classifying codes into themes, (4)
Developing and assessing interpretations, (5) Representing and visualizing data, (6)
Providing an account of the findings.
Data Saturation
The objective of data saturation in qualitative research is to ensure the precision
and validity of the data, analogous to the statistical validity of a sample in quantitative
research. Given the pragmatic nature of our study (multiple case study), I opted for
purposive sampling, as previously discussed. It is worth noting that although an ideal
sample size of no more than 12 is recommended for this type of study, caution must be
exercised to avoid both excessively small and overly large samples. Due to the absence
of clear guidelines for sampling and generalizing findings in qualitative research, the
concept of data saturation has emerged as a widely accepted tool among researchers to
enhance the rigor, quality, and trustworthiness of their findings.
As data saturation can be assessed in various ways, Sebele-Mpofu (2020)
suggests that researchers incorporate additional criteria into their study to confirm the
attainment of data saturation. Fusch and Ness (2015) have presented indicators for
verifying data saturation, including the absence of new data, absence of new themes,
absence of new coding, and the ability to replicate the study with new participants from
the same population and time frame. In qualitative research, the quality of data is of
paramount importance compared to the quantity of data, underscoring the need for rich
and insightful information. To ensure data quality, our strategy involves targeting an
initial population of 36 companies (three times the recommended sample size of 12).
Subsequently, i conducted interviews until the optimal sample size of 12 participants or
data saturation was achieved, whichever came first. In the data collection, data saturation
occurred after 11 participants were interviewed. This approach was designed to enhance
the overall quality and reliability of the findings of this research.
Ethical Research
Responsible academic research including a doctoral study must obey to ethical
principles. Estalblished in 1976, the Belmont protocol constitutes a reference for ethical
requirements that must govern research in order to protect individuals and society from
adverse consequences (Nagai et al., 2022). The Belmont protocol rests on three core
principles: beneficence, respect of autonomy, and justice. These standards encompass
principles such as voluntary participation, the right to withdraw, informed consent, as
well as the protection of privacy and confidentiality for all participants. This study
conforms to those principles.
Firstly, Ethical research requires prioritizing benevolence, ensuring that the
study’s design and execution do not harm or bias participants. Stringent ethical standards
are essential to safeguard the well-being of individuals and organizations. The study
employed a robust data collection strategy to guarantee the confidentiality and
anonymity of participants and their affiliated companies (Saunders et al., 2015).
Interview transcripts featured coded participant names and their respective companies for
enhanced organization and confidentiality. Company identities and any confidential or
proprietary information remains undisclosed, except when such information is already in
the public domain, as in reports and industry publications.
Ensuring the ethical protection of research participants is a paramount
responsibility for any researcher. In this study, I ensured the appropriate ethical
protection of participants by following a rigorous process. First and foremost, I sought
ethical approval from the Walden Institutional Review Board (IRB) before commencing
any research involving human participants. The IRB thoroughly reviewed my research
proposal to ensure it aligns with ethical standards and granted an approval (03-19-
241059331). This is a mandatory step for doctoral student researchers before conducting
interviews or other data collection activities with participants. Furthermore, I relied on
my dissertation chair and committee members to meticulously assess my study design.
They helped me identify and address any gaps related to the ethical protection of
participants in my research. Lastly, the informed consent form used in this study
emphasized the voluntary nature of participation and the right to withdraw at any time.
This not only respects participants’ autonomy but also contributes to the ethical
protection of participants in the research. In summary, safeguarding the ethical protection
of research participants is a multi-step process that involves obtaining IRB approval,
seeking input from advisors, and emphasizing voluntary participation in the consent
form.
Secondly, the principle of autonomy was satisfied in that, for this study, each
participating IT manager was presented with a consent form before the interview with
me. In qualitative research, a consent form serves as the essential instrument that
empowers participants by granting them the freedom to willingly engage in a study or
discontinue their participation at any point without facing any form of pressure or
negative repercussions (Nagai et al., 2022). This consent form is designed to ensure the
safeguarding of the rights and well-being of not only me and participant but also Walden
University. All participants were required to sign the Participant Consent Form. The
consent form also indicated that participation to the interviews is voluntary and
participants could withdraw anytime they wished. To initiate withdrawal from the study,
participants would follow these steps. First, when during the interview, participants
could simply notify me of their decision to withdraw. In response, I would promptly halt
the interview and proceed to delete all recorded data. Otherwise, if a participant wished
to withdraw after the live interview, they would have to submit a written withdrawal
request to me. Once I received this request, I would take the necessary steps to delete the
participant’s interview raw data and records from the transcript. Furthermore, no
coercion was exerted in participants in any shape or form to participate in the study.
Participants received no form of incentive, whether financial or otherwise, for their
participation in the study. The Interview Protocol (see Appendix B) reinforced some of
these aspects.
Thirdly, to respond to the justice principle, ethical considerations were upheld
throughout the study, particularly in participant selection (Johnson et al., 2020). To
ensure ethical standards were met, the following steps were taken:
1. Privacy and Confidentiality: The study strictly adhered to company privacy
and confidentiality regulations, as well as respect individual privacy and
freedoms. Consent was obtained from each participant prior to conducting
interviews (see Appendix B). Also the interview was confidential, that is,
data collected was not shared with third party, and the identities of
participants and their companies were revealed in the analysis and
conclusions of the study. Ensuring privacy and confidentiality was imperative
throughout the entire research data lifecycle. Consequently, I explicitly
apprised interview participants and informed the eventual readers of the study
that all raw interview data will be securely stored for a duration of five years,
employing either a cloud drive or the Walden file server.
2. Avoidance of Sensitive Information: The survey carefully avoided delving
into sensitive or confidential company information. I provided a guarantee
that both company and individual privacy were to be rigorously maintained.
Participant and company names were coded and omitted from the transcripts
3. Non-Discrimination: Participant selection did not discriminate based on
gender, religious affiliation, or age, ensuring a fair and unbiased
representation of the workforce. From a legal and regulatory perspective, the
United States requires non-discrimination as mandated by both state and
federal governments (FTC.gov, n.d.).
4. Minimum Tenure Requirement: Participants were required to have a
minimum one-year tenure at the company. This ensures that participants
possessed sufficient knowledge about the DG programs implemented within
the company.
5. Relationship Building: Prior to conducting interviews, an effort was made to
establish a positive and respectful rapport with the selected participants.
Building a strong researcher-participant relationship fosters trust and open
communication (Knott et al., 2022).
By implementing these ethical guidelines, the study upheld at best the highest
ethical standards and ensured the rights and privacy of all participants were respected
throughout the research process.
Data Collection
Effective research relies on solid data as evidence to address primary research
inquiries. The data collection phase is a pivotal step in the research process. During this
phase, researchers employ suitable instruments, strategies, and techniques to gather data
while also implementing methods to organize the collected information. The credibility
of qualitative research hinges on several key factors: a well-defined research question
supported by a conceptual framework, the choice of an appropriate research
methodology, my self-awareness and consideration of their own biases, and
meticulousness in both data collection and analysis (Johnson et al., 2020). Boosting the
trustworthiness of a study involves adhering to best practices, utilizing computer
software for data collection and analysis, engaging in peer reviews, conducting audits,
implementing triangulation, and exploring alternative case analyses. The following
sections delve into a detailed discussion of instruments, collection techniques, and data
organization strategies.
Instruments
The aim of this research was to comprehensively comprehend the strategies
employed by IT managers in implementing data governance within institutions that have
already adopted this process. This exploration delved deep into these strategies, with the
overarching objective of recognizing both challenges and opportunities in this context. In
terms of data collection instruments, it was essential that they align closely with the
chosen research method and design. Accordingly, for this study, semi-structured
interviews were selected as the primary data collection instrument (please refer to the
Interview/Survey Questions section). I developed an interview protocol document (refer
to Appendix B) that served as the consistent framework for interviewing all participants.
I refered to this protocol while conducting interviews with participants. This approach
ensured uniformity in wording and constructs, effectively mitigating any potential bias or
variations in the interview administration.
Semi-structured interviews are well-suited for open-ended inquiries and are
intended to probe participants’ opinions, beliefs, and insights related to the phenomenon
under investigation (Adeoye‐Olatunde & Olenik, 2021). The interview questions were
thoughtfully crafted to align with the core research topic and incorporate elements from
both Khatri and Brown’s (2010) Unified Framework for Data Governance and Abraham,
R., Schneider, J., & Vom Brocke’s (2019) Data Governance Conceptual Framework. By
embedding dimensions from these frameworks into the interview questions, this study
aimed to shed light on IT managers’ motivations and strategies for implementing data
governance through the lens of these comprehensive frameworks.
To enhance data reliability and accuracy, this study incorporated a
memberchecking phase. During this phase, I held a live meeting with the participant to
carefully review interview transcripts to confirm their fidelity to the original interviews
(Motulsky, 2021). This approach mitigates potential researcher bias and minimizes the
risk of misinformation introduced during the transcription of audio recordings into
summarized
text.
Data Collection Technique
For this study, I conducted semi-structured interviews with experienced
individuals responsible for data governance programs in US-based Fortune 500
companies, each with over ten years of relevant expertise. Semi-structured interviews
serve as a data collection method that bridges the gap between structured and
unstructured interviews. They offer the benefit of adaptability and in-depth exploration
of topics, but are accompanied by challenges pertaining to time, training, analysis, and
potential bias. One notable advantage is their flexibility in questioning (Ruslin et al.,
2022). Interviewers can tailor the interview protocol questions based on the participant’s
responses and adjust the format to suit different participants. This flexibility fosters a
more comfortable environment for participants, encouraging richer and more
comprehensive responses that yield valuable insights into the research phenomenon.
Additionally, semi-structured interviews provide an abundance of rich data. Participants
can thoroughly explore their responses, offering a profound insight into complex
socialbehavioral research inquiries. This rich dataset can be further enhanced by
triangulating it with other data sources, such as surveys, to gain a more comprehensive
understanding of the research topic. However, despite these advantages, semi-structured
interviews do have their drawbacks. One common concern is the potential for
interviewer bias, which is inherent in most qualitative research involving human
participants. Furthermore, they can be time and resource-intensive both in preparation
and delivery. Their open-ended nature can result in response variability, posing
challenges in drawing clear and consistent conclusions. Finally, conducting a
comprehensive analysis of semi-structured interview data can be intricate, necessitating
expertise in identifying themes and extracting meaningful insights to ensure the
credibility and validity of the findings. In the following paragraph, I provide a concise
overview of my approach to conducting semi-structured interviews with specific
participants.
Authorization from the Walden University Institutional Review Board (IRB) was
a prerequisite for conducting these interviews. The interviews were conducted using the
Microsoft Teams video-conferencing platform, allowing both audio and video recording
capabilities. My initial interaction involved introducing myself and explaining the
interview’s purpose, emphasizing the importance of safeguarding participants’
information. Although I intended to pre-screen participants before the interviews, I
reconfirmed their seniority within the company and their experience in data governance
during the interview. These qualifying questions served as a foundation for exploring
core topics, including the factors driving data governance implementation, challenges,
emerging trends, and opportunities. To maintain research rigor and minimize potential
researcher bias, I formulated open-ended questions, fostering participant input without
leading or adding excessive context. To enhance data reliability, I employed clarifying
follow-up questions and take detailed notes during the interviews.
To improve data reliability and accuracy, this study included a member-checking
phase. In this phase, I conducted a live meeting with the participant to thoroughly
examine interview transcripts and verify their alignment with the original interviews
(Motulsky, 2021).
Data Organization Techniques
Efficiently organizing text data is a crucial step in the data collection process,
setting the foundation for effective data analysis. To uphold the principles of reliability,
validity, security, and traceability in collected data, it is imperative for researchers to
methodically catalog and store the information within a database (Johnson & Chauvin,
2020). This requirement aligns with the guidelines outlined in Walden University’s
Doctoral Study guidelines. My approach involved utilizing Turboscribe software, a
reliable online audio-to-text transcription software, to convert the Microsoft Teams audio
recording of the interview into a text file. Teams offered an audio-to-text transcription
feature, but I preferred Turboscribe. Subsequently, I meticulously reviewed and made
necessary edits to enhance the accuracy and clarity of the transcription. Furthernore, I
conducted a member-checking interview with participants to boost the data accuracy
(Motulsky, 2021). Finally, I ensured the safekeeping of all raw data ( the original
interview audio files and the edited transcripts) by storing them encrypted in my
Microsoft OneDrive secure cloud drive location for 5 years. This systematic approach
not only ensures the integrity of the data but also facilitates seamless access and analysis
when conducting my research.
Data Analysis Technique
Data analysis is a pivotal research stage where qualitative data is systematically
processed, transforming it into meaningful insights. Qualitative research poses unique
challenges in data analysis due to its open-ended nature. Thematic analysis is a
commonly preferred method for qualitative research. Castleberry and Nolen (2018) have
outlined a five-step methodology for conducting thematic analysis, which includes:
compiling data, disassembling data, reassembling themes,
interpreting themes, and concluding the analysis. These phases are
interdependent and collectively serve as the backbone of qualitative research, facilitating
the extraction of valuable insights and understanding from the collected data. This study
employed Castleberry and Nolen’s
(2018) methodology for conducting a thematic analysis of the collected data.
Data Compilation
The initial phase involved the compilation of qualitative data collected. This
process can be likened to piecing together a puzzle, where each data point, interview
transcript, or observation is meticulously organized and structured. This preliminary
organization serves as the cornerstone for subsequent analytical stages. During this
phase, I utilized Turboscribe audio-to-text transcription software to transcribe the raw
interview recordings into a readable text format. Following this transcription, I
thoroughly reviewed it, enabling me as a researcher to gain a deeper understanding of the
phrases and meanings extracted from the transcripts. This step remains incomplete until
the member-checking interview is conducted. The purpose of member-checking is to
enhance data quality (Motulsky, 2021), as it involves validating the transcribed responses
with the participant to ensure fidelity to the original conversation. The outcome of this
phase is a consistent and well-organized transcript, which served as the input for the next
phase of analysis.
Data Disassembly
Disassembling data involves utilizing coding techniques to discern and organize
meaningful groupings within the transcribed interview text by identifying themes,
concepts, and ideas. This coding process serves as a fundamental aspect of qualitative
research, characterized by its inductive approach, as opposed to the deductive approach
employed in quantitative research. In this qualitative endeavor, researchers permit
meaning to naturally emerge from the data. A crucial step in this process is the careful
consideration of a coding strategy before initiating the analysis (Castleberry & Nolen,
2018). These codes can often be drawn from existing research on the topic found within
the body of literature. For this particular study, I constructed codes based on the data
governance domains and dimensions outlined in both Khatri and Brown’s (2010) Unified
Framework for Data Governance and Abraham, R., Schneider, J., & Vom Brocke’s
(2019) Data Governance Conceptual Framework. This analysis can be done either
manually or with software. For this study, I used the manual approach, with in-depth
reading of transcripts, noting codes, and grouping into themes. The advantage of the
manual approach is that it does not require the cost of expensive software. The
disadvantage is that it is cumbersome
The alternative approach to facilitate this analysis, NVivo software can be used,
complemented by manual processes as needed. NVivo can assist in processing the text
file, which is derived from the the audio-to-text transcription of the interview (Allsop et
al., 2022). This software can aid in code generation and the identification of emerging
themes. An example of thematic coding with NVivo software is shown in Figure 9.
Codes are clustered together in the text data to form groups. The size of nodes represents
the frequency of each code, while colors represent code families that can be grouped into
themes. Subsequently, emerging themes cqna be organized into thematic groups.
Furthermore, it’s essential to perform manual adjustments to ensure that these themes are
in alignment with the primary research question, this study’s conceptual frameworks, and
the established academic knowledge on data governance.
Figure 9
Example of Thematic Analysis Coding With NVivo Software
Data Reassembly
Data reassembly, a pivotal phase in the research process, involves the
consolidation of coded data into cohesive themes and narratives. This step serves as the
foundation for the subsequent thematic interpretation. During this phase, I established
connections among coded segments to uncover overarching insights and derived
meaningful conclusions (Castleberry & Nolen, 2018). Furthermore, meticulous
documentation of codes, code-to-theme mappings, and theme descriptions was essential
at this stage.There are two primary methods for mapping code to themes: hierarchies and
matrices. Thematic hierarchies involve creating visual graphs by clustering codes,
generating hierarchies with varying levels of granularity. On the other hand, matrices are
constructed by organizing participant roles, themes, variables, and emerging concepts in
rows and columns.
Thematic hierarchies are the more commonly employed approach for mapping
codes to themes and establishing relationships among themes. However, it’s important to
note that reassembling these hierarchies requires analytical skills on my part to justify the
relationships and hierarchy of themes.
Eight major themes emerged from data analysis as follows, in no particular order
of importance:
1. Data Quality, Security/Privacy, and Regulatory Compliance: These are the
main drivers for Data Governance (DG) programs.
2. New Drivers: Data discoverability, data cataloging, and automation are
emerging as significant motivators for DG.
3. Third-Party Data Assets: These should be treated with the same rigor as
internal data to preserve security and privacy.
4. Data as an Asset: While most companies recognize the value of data, few
treat it as an asset unless it impacts the bottom line, either as a revenue
generator (e.g., data as a product, enabling AI innovation) or as a liability
(e.g., risk, compliance).
5. Executive Sponsorship: Vital for DG success. Without it, there is no
organizational alignment, authority, or budget support, leading to program
failure.
6. Value Justification: One of the biggest challenges is justifying the value of
DG upfront to both executive and lower management.
7. Data Competencies: DG success requires relevant competencies for all DG
roles.
8. Targeted DG Programs: DG programs should be targeted, use case-bound,
and antecedent-driven, avoiding cumbersome, large-scale initiatives
undertaken without a clear purpose.
Interpreting
Data interpretation is a crucial stage in the research process, where researchers
draw conclusions based on their analysis of codes and derived themes. To ensure a
highquality interpretation, there are five key qualities that researchers should aim to
achieve (Castleberry & Nolen, 2018):
1. Clarity: The interpretation should be presented in a clear and logical fashion,
demonstrating how it was derived from the data.
2. Reproducibility: The interpretations should be reproducible when using the
same dataset, ensuring that others can arrive at similar conclusions.
3. Fidelity to Raw Data: It is essential to maintain fidelity to the original raw
data, ensuring that interpretations are grounded in the data without distortion.
4. Knowledge Enhancement: Interpretations should contribute to a deeper
understanding of the subject matter, adding valuable insights and knowledge
to the research field.
5. External Validity and Trustworthiness: Interpretations should be robust
enough to withstand external scrutiny by other researchers, ensuring their
validity and trustworthiness.
Thematic interpretation provides the bridge between raw data and the research’s
ultimate objectives. The themes extracted in the previous step are now refined and
ranked, keeping those that best align with the research question at the forefront of this
study. Table 7 provides an illustration of the themes that have surfaced through code
analysis.
Table 7
Example of Themes Emerging From Keywords
Domain Keyword, phrase Theme (as
framework
dimension)
Topics in data
governance
Definition, architectures, objectives Domain scope
Industries Multiple industries: finance, healthcare,
retail, highly regulated industries
Domain scope
Business drivers efficiency, accuracy, transparency,
regulatory compliance, save time and
money, cleaner data, data security
Antecedents
Data principles
Strategies policies, data access policies and
control, who/what, use cases,
standards, classification (sensitive,
nonsensitive, PII), data documentation,
metadata management (ingestion,
certification) cross-functional, adjacent
to IT
Data access
Data Lifecycle
Governance
mechanisms
Organization Who is responsible, all, exec sponsor C-
level, stand-alone lead, crossfunctional,
strategic vision
Organizational
scope
Challenges Leading challenges
Organizational:
- convincing organizations that it is
worth it;
- strategy first then technology
Regulatory compliance
Costs
- high
- applications are prohibitively
expensive Technical
applications and solutions not
automated
Data quality Data
scope
Antecedents
Trends Key future trends of data governance
- regulatory compliance
- big data security and governance
commoditization of governance
capabilities: established platforms
packaging data governance within
their
own suite of capabilities
Consequences
As part of this interpretive process, I constructed a thematic map that provides a
detailed representation of the identified themes and their interrelationships. These themes
were linked to the main research topic and the dimensions of data governance within the
conceptual framework. This visual representation aids in conveying the complex
relationships between themes and their relevance to the broader research context.
For this study, triangulation was used to boost the validity of the interpretation
and findings of the study. Triangulation is a valuable technique in the qualitative data
analysis process, serving to enhance both external validity and the trustworthiness of
research results (Farquhar et al., 2020; Noble & Heale, 2019). It is widely recognized as
an effective approach to bolstering the validity of a study, whether it is qualitative,
quantitative, or a mixed-methods study (Farquhar et al., 2020). This is particularly
relevant in the context of case study research, such as the one conducted in this study,
where triangulation plays a crucial role in attaining a higher level of validity and
confidence in the findings. Triangulation in research involves the utilization of various
methods, theories, or data sources to gain a more comprehensive understanding or
explanation of the same phenomenon. The core concept underlying triangulation is
corroboration or convergence, which means that these different sources, methods, or
theories should lead to the same conclusions regarding the observed phenomenon.
Consequently, triangulation serves to enhance both internal and external validity.
There are tree types of triangulation approaches. Firstly, data source triangulation
is a commonly employed technique in case study research (Farquhar et al., 2020). It
involves gathering data from various sources, such as conducting interviews with
multiple participants, observing the phenomenon at different times, and incorporating a
combination of primary and secondary sources derived from publicly available
publications and other research materials. Secondly, researcher triangulation is another
valuable method in case study research. In this approach, two or more researchers
independently investigate the same phenomenon to arrive at similar interpretations and
conclusions, thus enhancing the reliability of the findings. Lastly, theoretical
triangulation is employed to enrich understanding. This technique involves examining
the same set of data from different theoretical perspectives, which can provide valuable
insights and a more comprehensive understanding of the subject under investigation.
Concluding
The interpretive phase involves extracting themes and creating thematic maps
from the coded data. These themes are then thoroughly examined to derive meaningful
interpretations. The ultimate step in the data analysis process is drawing conclusions,
where these interpretations are formulated to address the primary research question,
considering the conceptual frameworks underpinning the study. In this particular study,
the conclusions were tailored to provide insights into the strategies employed by IT
managers for the implementation of data governance programs. These conclusions were
validated for alignment with the data governance domains and dimensions specified in
both Khatri and Brown’s (2010) Unified Framework for Data Governance and Abraham,
R., Schneider, J., & Vom Brocke’s (2019) Data Governance Conceptual Framework.
Reliability and Validity
In qualitative research, ensuring the reliability and validity of data is crucial to
establish trustworthiness and to generate meaningful findings that contribute to the
research body and address the main research question. Daniel (2019) has introduced a
framework consisting of four key dimensions known as Trustworthiness, Auditability,
Credibility, and Transferability (TACT) to guide student researchers in achieving these
goals. In the subsequent sections, I will delve into the reliability and validity aspects of
this research, along with the strategies I employed to ensure these crucial elements were
safeguarded.
Reliability
Reliability is a critical metric evaluating the resilience and robustness of both the
research instrument and the overall study. Essentially, it examines the consistency and
stability of research findings (Cichy & Rass, 2019). In quantitative studies, assessing
reliability is relatively straightforward as it measures the repeatability of the study’s
findings. In other words, it gauges whether the same results would hold when applied to
a different population sample with similar traits. However, in the realm of qualitative
research, determining study reliability can be more complex and, at times, less
meaningful (Cichy & Rass, 2019). This challenge is especially pronounced in research
methodologies like case studies and ethnographic studies, which are tailored to specific
populations with unique characteristics. Consequently, expecting identical outcomes
when these methodologies are applied to different populations may not be realistic.
For a qualitative study, the concept of reliability takes on several dimensions.
First is internal consistency, that is the consistency of the measurement methods.
Secondly, inter-coder consistency in the interpretation of results when more than one
researcher is involved. Third, the audit-trail to maintain an organized record of data
collection, and analayis processes. Finally, triangulation by using multiple sources to
enhance reliability of the study.
Consistent Data Collection
In this study, prioritizing the reliability of our instruments was paramount, and I
rigorously upheld this during the interview process. This involved a standardization of
data collection procedures and protocols with a meticulous design of interview questions
and a comprehensive interview protocol (see Appendix B), as detailed in the Data
Collection section. Achieving a consistent data collection also involves confirmability,
that is, to achieve objectivity and neutrality of research findings. As discussed in the
Role of the Researcher section, reflexivity is a common strategy, that I have used in this
study for reducing researcher bias in data collection (Dodgson, 2019).
Member Checking
Additionally, this study emphasizes the importance of fidelity in transcribing
interview responses to guarantee the accuracy of the participant’s original statements. To
ensure this accuracy, I validated the interview transcripts with each participant in an
interview follow up with a member-checking interview (Motulsky, 2021). This was an
opportunity for participants to validate the accuracy of their contributions to ensure their
perspectives and represented without a researcher bias. This was discussed in the Data
Collection section.
Dependability and Credibility
This aspect measures the reliability of the data as it pertains to both the internal
validity and external validity. For this purpose, the study employed data triangulation
(Farquhar et al., 2020). I employed multiple data sources, including participant interview
transcripts, public domain literature and white papers and methods to cross-verify
findings, reducing the risk of bias and increasing the dependability of results. The
following six secondary sources were considered for data triangulation: Atlanta Regional
Commission (2019), Biggenden (2022), ESG (2022), Nephos Technologies (2022),
OvalEdge Team (2022).
Validity
Validity is a critical gauge of a research or study’s credibility as it measures the
accuracy and truthfulness of research findings. Validity encompasses two key
dimensions: internal and external validity (Daniel, 2019). Internal validity ensures that
the study’s outcomes consistently address the primary research questions or hypotheses.
Meanwhile, external validity ensures that the study’s results align with established
knowledge in the field, findings from analogous studies conducted by other researchers,
and industry norms.
In this study, I assessed external validity by utilizing conceptual frameworks and
critical success factors for data governance (Alhassan et al., 2019; Almeida et al., 2019;
Alsousi & Shah, 2022; Bento et al., 2022). Specifically, the study aimed to determine if
the data governance strategies and motivations of IT managers align with the decision
domains outlined by Khatri and Brown (2010) as well as Abraham et al. (2019) in their
conceptual frameworks for data governance. The focus was on ensuring that the
conclusions drawn from this study harmonize with the antecedents decision domain of
the conceptual frameworks when it comes to motivations. Additionally, the researcher
intended to align the findings on challenges and strategies with the organizational
dimensions of these conceptual frameworks. Furthermore, conclusions regarding
strategies should be consistent with the established success factors and best practices for
data governance (Alhassan et al., 2019; Almeida et al., 2019; Alsousi & Shah, 2022;
Bento et al., 2022), as documented in our literature review. It’s essential to acknowledge
that this study takes an inductive approach rather than a deductive one, implying that it
does not anticipate a perfect alignment with the existing frameworks. Instead, the
primary objective was to uncover insights into the data governance strategies employed
by IT managers and contribute new knowledge to the existing body of literature in this
field.
In this study, internal validity was assessed by evaluating how effectively the
findings address the primary research question. This involved uncovering plausible
motivations (antecedents) behind IT managers’ decisions to implement data governance
in their respective companies, along with the associated strategies. Additionally, internal
validity was gauged by confirming that data saturation (Sebele-Mpofu, 2020; Fusch &
Ness, 2015) is achieved within the retained participant sample. To enhance the accuracy
and fidelity of data collection, member-checking (Motulsky, 2021), was conducted
whereby each participant confirmed that the transcript was faithful to their original
responses to the interview with me. Finally the TACT framework (Daniel, 2019) was
leveraged to gauge credibility and transferability this study.
Confirmability
The rigor of qualitative research hinges on four essential criteria of the Lincoln
and Guba’s Framework for trustworthiness research (Devakirubai, 2020): credibility,
confirmability, transferability, and dependability. Confirmability assesses the extent to
which the conclusions of a qualitative study are firmly grounded in the collected data and
the subsequent analysis, as opposed to being influenced by the researcher’s personal
biases, values, or preconceptions. In this study, I ensured confirmability by maintaining
transparency throughout the research process. This will was achieved by meticulously
documenting every step, including data collection, analysis, and data archiving. The
appendices contain comprehensive documentation of these procedures, encompassing the
participant interview protocol, member-checking protocol, participant consent form, and
records of the code and themes extraction.
Transferability
Transferability in qualitative research assesses the applicability of study insights
and conclusions to similar contexts, populations, or settings. Unlike quantitative
research, qualitative research, such as multiple case studies, is inherently context-specific
and typically lacks generalizability (Johnson et al., 2020; Devakirubai, 2020; Dodgson,
2019). To enhance the transferability of this study, my approach was based on three key
pillars. Firstly, by anchoring it in the particular domains and dimensions of established
theoretical frameworks, such as Abraham, R., Schneider, J., & Vom Brocke’s (2019)
Data Governance Conceptual Framework and Khatri and Brown’s (2010) Unified
Framework for Data Governance, this research employs concepts that serve as a
foundation for making our findings more applicable to broader contexts. Secondly, to
enhance the study’s transferability, I employed reflexivity to acknowledge and mitigate
my personal biases as a researcher (Dodgson, 2019). Lastly, my data collection approach
involved a diverse range of participants from various companies, spanning different sizes
and cultures. This approach ensured that the findings were not confined to a single
company or location in the United States. Furthermore, meticulous documentation of the
data collection and analysis promotes transferability through transparency and
confirmability.
Transition and Summary
As the study design, Section 2 offered a comprehensive overview of the overall
research design. It encompassed the selection of the research method and specific design,
participant criteria, sampling methods, data saturation, data collection and analysis
procedures, as well as an exploration of validity and reliability. The study employed an
exploratory multiple case study approach with a pragmatic design and purposive
sampling of participants. These participants consisted of IT managers with expertise in
data governance, having a minimum tenure of one year, and employed at Fortune 500
companies in the United States with a functioning data governance program in place for
at least 18 months. Ethical considerations are of utmost importance, and I trictly adhered
to the Belmont Protocol to ensure the protection of both participants and their respective
institutions. I also addressed reflexivity by identifying personal bias and discussed
strategies on how to mitigate their adverse impact on the internal validity and
conclusions of the study.
Data collection was executed through carefully structured online interviews with
participants, following a consent form process. Additionally, member-checking
interviews were conducted as a follow-up to enhance the quality of the collected data.
Raw data collected from interviews was be kept encrypted, secure and confidential using
my Microsoft OneDrive cloud storage. Qualitative data analysis methods were applied,
including the compilation of interview data, transcription of interviews, and the
incorporation of member-checking to bolster data quality. This process culminated in the
identification of key themes that formed the basis for drawing conclusions. To maintain
the study’s reliability and validity, it adhered closely to the TACT Framework (Daniel,
2019).
Section 3 explores the presentation of study findings derived from the research
design, data collection, and analysis outlined in Section 2. It involves a comprehensive
discussion of the study’s findings, accompanied by recommendations that elucidate their
relevance to the broader field of data governance. The conclusions drawn from this
research endeavor pinpoint the strategies discerned for the implementation of data
governance programs by IT managers within American companies. This discussion also
encompasses the challenges encountered in this process. Furthermore, Section 3 delves
into the implications of these findings for the professional practice of data governance. It
also explores the potential ramifications for driving social change through improved data
governance practices. Additionally, this section identifies limitations within the study,
gaps in existing knowledge, and suggests areas for future research that can further
enhance the insights gained from this study.
Section 3: Application to Professional Practice and Implications for Change
This study focused on exploring strategies employed by IT managers to
implement successful DG programs. In this section, I present the findings from data
collection and analysis. Additionally, I discuss how these findings contribute to the
existing body of knowledge in the DG field and their implications for social change.
Finally, I suggest potential future research directions and conclude with my personal
reflections as the researcher.
Overview of Study
This qualitative pragmatic study aimed to explore strategies IT managers use to
implement DG programs, focusing on factors that facilitate or hinder these
implementations. Primary data were collected through semistructured interviews with
mid-level IT managers, including those in director-level roles, who had experience in
DG across various industries. Secondary data from organizational documents and
publicly available data from governance reports were also used for triangulation. This
study revealed several key findings and relationships among them that influence and
guide the strategies to implement DG:
1. Driving factors for DG: Data quality, security, privacy, and compliance are
the top needs driving data governance. Additionally, data discoverability,
data cataloging, and data governance automation have emerged as new
drivers.
2. Data as an asset: Companies acknowledge the importance of data, but its
value is best demonstrated when it impacts the bottom line as a revenue
generator, such as through innovation.
3. Third-party data assets: These should be governed with the same rigor as
internal data sources but with established standards for quality and privacy
protection.
4. Executive support: Executive support and sponsorship are crucial for the
success of DG programs. However, justifying the value of these programs to
executives and middle management remains a significant challenge.
5. New CSFs: The study identified two new success factors for data governance
programs:
•Ensuring data competencies for all roles driving the DG program.
•DG programs should have a clear strategy, be targeted, use case bound,
and be antecedent driven to achieve success (see Figure 10).
Figure 10
Major Themes and Relationships
Presentation of the Findings
The research question of this study was the following: What strategies do IT
managers use to implement data governance programs in practice? The aim was to
uncover insights that would contribute to the body of knowledge on DG and provide
valuable information for practitioners in the field. This section presents the findings from
my research, organized into eight identified themes.
Theme 1: Effective Data Governance Strategies Must Address the Challenges of
Data Quality, Security/Privacy, and Regulatory Compliance
One theme emerging from data analysis was that data quality, security/privacy,
and regulatory compliance are the main drivers for a successful DG implementation.
Most interviews with participants highlighted that these factors represent both essential
needs and significant challenges for organizations, directly affecting the overall
effectiveness of DG. As a result, IT managers must develop effective strategies to
address these challenges and meet these needs.
Data quality encompasses characteristics that enhance the value and trust
worthiness of data; these include accuracy, completeness, reliability, and relevance.
Companies and organizations need high-quality data to make informed business
decisions and often mission-critical ones. For regulated industries such as health care and
financial, high data quality is necessary for compliance with laws and standards. Poor
data quality can lead to of high costs and risks that adversely impact business
performance and the bottom line. Data security is concerned with protecting data from
unauthorized access, corruption, and theft throughtout the life cycle; this includes aspects
of data access, privacy, authorization, and authentication. Protection of sensitive data
such as personally identifiable information (PII), company financials, and intellectual
property falls in this category. Data breaches can result in significant financial losses due
to fines, legal fees, and the cost of remediation. Data privacy, on the other hand, is
concerned with proper handling of data (collection, storage, and sharing) ensuring that it
is used in a manner that complies with laws and allowing individuals to control how their
personal data is used and shared. Finally, businesses and organizations are required by
law to meet regulatory compliance or face financial penalties. Strict regulations exist that
govern certain industries as is the case of HIPAA for the health care industry, EU GDPR
and the California Consumer Privacy Act (CCPA) for consumer data privacy in retail
and other industries.
All 11 interview partitipants indicated either data quality, security/privacy, or
regulatory compliance as the main drivers (or antecedents) for DG programs.
Participants shared views from experiences in various industries including regulated
(health care, financial), retail, oil and gas, transportation, hospitality, and real estate. Data
quality driver was a common denominator among all participants. Participant 1 shared
data quality concerns that most businesses have with inaccurate data impacting usability
and trustworthiness. This participant also spoke of the importance of data security and
privacy to comply with GDPR and CCPA regulations; this view was shared by
Participant 5 with regard to HIPAA regulation for the health care industry. Participant 4
and Participant 7 shared a similar view that governing the data is usually data quality,
data security, and data privacy.
Furthermore, in the era of AI, companies are honing in on innovation. Participant
7 mentioned the need for high data quality and trustworthiness to feed into the large
language models used by Gen AI and other AI tools. Participant 2, Participant 6, and
Participant 8, from pharmaceuticals, also shared data quality issues as the main concern,
but added that businesses need to discover where data resides and know the content of
data assets and lineage sources to provide an intimate knowledge of the data. All this is
needed to pass the various compliance audits. Developing a data strategy to understand
where the organization wants to go and how the organization wants to use data becomes
critical. For the financial and banking industries, Participant 3 and Participant 10 insisted
on regulatory compliance as the main driver for data governance besides data quality,
data security, and privacy. The Financial Industry Regulatory Authority has a mandatory
audit of financial institutions to show the lineage of data for all transactions. Data quality
is important for accuracy and trust of the transactions. Security is paramaount to avoid
data breaches, which could allow identity theft leading to litigation and regulatory
penalties.
Interview participants suggested various strategies to enhance data governance,
focusing on data quality, security and privacy, and regulatory compliance. For data
quality, Participant 1 proposed a governance strategy centered on ranking the business
importance of data assets. This involves cross-functional engagement to identify which
data is most critical to the business. Participant 8 also emphasized improving data quality
by governing access, monitoring usage, and democratizing data across the company,
ensuring that data is made available in a centralized location. Data lineage was
highlighted by Participant 3, particularly in the context of regulatory audits within the
banking sector. Participant 3 stressed the importance of understanding where data
originates because this is foundational to any further data quality initiatives. Establishing
data lineage is crucial before addressing other aspects of data governance. For
information security, the strategy involves classifying data to determine its level of
sensitivity, as mentioned by Participant 4. This includes setting access controls and
adhering to data privacy regulations, which are especially stringent in regulated
industries such as financial services, health care, and life sciences. Participant 4 also
noted the importance of these strategies in avoiding costly fines and ensuring compliance
with regulations such as HIPAA, Federal Deposit Insurance Corporation, and Securities
Exchange Commission. Finally, the overarching strategy for regulatory compliance
involves understanding the organization’s data management goals and setting measurable
targets for data quality, security, and utilization. This approach ensures that data
governance aligns with the organization’s broader objectives and regulatory
requirements.
Methodological triangulation was achieved with organizational documents and
reports fully supporting this theme with a 100% consistency across all four sources (see
Table 8). In an effort to enhance transit data interoperability, quality, sharing, and
costeffectiveness, the Atlanta Regional Commission (2019) conducted research to
identify barriers to data governance. The study identified two primary challenges that a
robust data governance framework could resolve. First, inconsistencies in data structures,
formats, and semantics among regional transit authorities hindered data quality and
standardization, making data integration and comprehension difficult across systems.
Second, issues related to data security and privacy arose due to a lack of clear data
ownership and rights, complicating access and sharing. These challenges underscored the
need for implementing data governance to address the issues faced by the Atlanta
Regional Commission.
Table 8
Frequency of First Major Theme
Source of data collection n
Interviews 11
Documents 4
Note. Theme 1: Effective data governance strategies must address the challenges of data
quality, security/privacy, and regulatory compliance; n = frequency.
The Enterprise Strategy Group (ESG, 2022) surveyed 220 business and IT
professionals in data governance roles to understand how organizations implement DG,
focusing on the key drivers and challenges. The research identified data quality as the
leading factor in several areas: (a) how organizations define data governance (47%), (b)
the primary challenge in maximizing the return on DG efforts (45%), and (c) the top
driver for DG programs (41%). Recognizing the increasing adoption of DG by
enterprises, Nephos Technologies (2022) conducted a study to identify the needs and
challenges driving DG programs. The findings align with Theme 1, highlighting the
following key needs: (a) improved data quality, (b) enhanced data analytics, (c)
strengthened data security and privacy, (d) regulatory compliance, (e) minimized
financial risk, (f) increased efficiencies and cost savings, and (g) faster access to relevant
data. Needs A and B correspond with the data quality driver, Need C with data security
and privacy, and Need D with regulatory compliance. Needs E, F, and G are typically
achieved as positive outcomes when DG effectively addresses the top three drivers.
OvalEdge Team (2022) identified five key applications of DG across industries:
data privacy compliance, data discovery and literacy enhancement, centralized data
access management, creation of a standardized business term repository, and
collaborative analytics for new data product development. The data privacy use case
aligns with the security and privacy drivers discussed in Theme 1. The report noted that
regulations such as GDPR, CCPA, and International Association of Privacy
Professionals aim to protect consumer PII but have also imposed significant costs and
litigation risks on organizations. However, Gartner (2022) predicted that by 2024, 75%
of global user data will be protected by privacy regulations, with large organizations
budgeting $2.5 million or more to comply.
The extensive literature on DG reinforces Theme 1’s findings, identifying data
quality, data security and privacy, and regulatory compliance as the primary drivers for
data governance programs. Walsh et al. (2022) identified the data principles, data access,
and data quality components of Khatri and Brown’s (2010) DG framework as key
organizational motivations for implementing DG. For data principles, the primary
motivation was to enhance compliance with external regulations and internal business
and data processes. The goal of implementing data quality was to establish data
credibility and trustworthiness by ensuring accuracy, relevance, and accessibility for
users. The data access domain focused on security and privacy, aiming to restrict data
access to authorized users, monitor access to prevent unauthorized use, and protect data
assets, especially in cloud environments. Mahanti (2022) emphasized that improved data
quality and governance, particularly in terms of accuracy and credibility, play a crucial
role in helping companies achieve regulatory compliance and avoid accounting scandals
and financial disasters. Hikmawati et al. (2021) similarly emphasized the importance of
enhancing data quality, governance, and access through the implementation of MDM.
This Theme 1 aligns with the DG theories outlined by the conceptual frameworks
used in this study: Khatri and Brown’s (2010) unified Framework for DG and Abraham
et al.’s (2019)DG conceptual Framework. Khatri and Brown’s framework identifies five
decision domains critical for successful data governance: data principles, data quality,
metadata, data access, and data life cycle. Data principles, which include value addition,
compliance, performance, and decision rights, represent the business motivations for data
governance. Abraham et al.’s antecedents dimension aligns with these data principles.
The data quality domain focuses on the accuracy, timeliness, completeness,
trustworthiness, and credibility of data, which are essential for reliable decision making
and efficient operations. The data access domain addresses data protection and access
control, including defining access levels to meet internal and external compliance
requirements, essentially covering data security, privacy, and regulatory compliance.
From this analysis, the findings of Theme 1—specifically data quality, data
security, privacy, and regulatory compliance as key drivers of data governance programs
—align closely with the Data Principles, Data Quality, and Data Access domains in
Khatri and Brown’s (2010) framework. It is important to note that while the conceptual
framework positions these domains as justifications and guides for implementing data
governance programs, it does not necessarily rank them as top drivers. To ensure
effective data governance, a comprehensive program should address all relevant data
within an organization, focusing on the key domains of Data Quality, Data
Access, Metadata, and Data Lifecycle. A critical challenge and necessity in this process is
the ability to discover and catalog all data sets, ensuring they have a unified meaning
across the organization.
Theme 2: Data Discoverability, Data Cataloging, and Automation Are Emerging as
Significant Motivators for Successful DG Implementation
Another theme that emerged from the data analysis was data discoverability, data
cataloguing, and data governance automation. Data governance strategies must address
these needs. Data discoverability is the ease with which users can find, access, and
understand the data they need. This aspect of data management is crucial, particularly in
large organizations or systems with vast amounts of data in various formats and
sensitivity levels. In such situations, there’s a high likelihood of dark data—information
that organizations collect, process, and store without using or analyzing. This data holds
no value and can even pose a compliance risk. Effective data discoverability enables
quick access to the right data, leading to better decision-making, more efficient
operations, and improved outcomes. Achieving data discoverability requires high data
quality and robust metadata management to provide essential information about the data,
such as its type, format, source, and sensitivity. Automation tools are also necessary to
search, index, and organize data, allowing users to locate and interpret it effectively. This
process often results in the creation of a data catalog—a centralized repository that helps
users discover and understand available data within an organization, including metadata,
data lineage, and other relevant details. However, challenges such as data silos, large
data volumes, complexity, and the need to balance security and privacy—especially in
highly regulated industries like healthcare and finance—can complicate data
discoverability.
Improved discoverability not only facilitates better data management but also supports
regulatory compliance by ensuring data traceability and enabling usage audits.
Improving data discoverability involves a combination of technology, processes, and
governance practices, all aimed at making data easier to find, access, and use.
All participants emphasized the critical need for data discoverability and the
implementation of a data catalog to enhance data management, governance, and
democratization across organizations. Participants from the pharmaceutical industry
(Participants 2 and 8) highlighted the growing challenge of locating data stored in
various repositories, which is essential for operationalizing that data. They noted that
data discoverability and cataloging have become vital capabilities for companies
managing and governing their data assets. The participants stressed that, beyond just
discovering data, organizations must establish a company-wide data catalog. This catalog
would serve as a comprehensive data dictionary, providing meaning and, more
importantly, ensuring data lineage necessary for compliance and audits in regulated
industries. Participants 6 and 8 further emphasized that a data catalog, coupled with
governance automation, enhances data democratization. Participants from the financial
and banking sectors (Participants 3 and 10) underscored the importance of traceability—
tracking data lineage to show its sources, transformations, and final destinations. They
also noted that data discovery aids in classifying sensitive and PII data. Participants from
IT, manufacturing, and hospitality industries (Participants 1, 4, 7, 9, and 11) agreed that
the need for data discoverability is growing across companies, with data catalogs
becoming a priority.
They particularly stressed the importance of data dictionaries and, above all, data
lineage.
Participants 4 and 11 emphasized the need for automating data governance rules
and processes. As organizations manage increasing volumes and complexity of data
across various systems and platforms, automation becomes crucial (Nadal et al, 2022).
Manual governance methods are inadequate for scaling due to several challenges and
opportunities. First, the sheer volume and complexity of data, sourced from databases,
cloud services, and IoT devices, make it difficult to enforce consistent governance
policies manually. Secondly, regulatory compliance demands, such as GDPR and CCPA,
require organizations to maintain strict data privacy and security standards. Automation
ensures these policies are consistently applied, helping organizations stay compliant. In
highly regulated industries, like banking and finance, audits are a necessity. Automated
systems enhance tracking and reporting capabilities, making it easier for organizations to
demonstrate compliance during audits.
Most interviews with participants revealed that companies struggle with dark data
and a lack of centralized data definitions. These challenges underscore the need for data
discoverability and a data catalog to enhance data governance. In the financial sector,
data lineage is a regulatory requirement for auditing financial data. Implementing a data
catalog with lineage capabilities is a strategy that not only strengthens data governance
but also ensures regulatory compliance. Dark data, when left unused, can pose a liability
from both regulatory and security perspectives. Furthermore, the absence of a data
catalog prevents users across the company from accessing a unified understanding of
data, hindering the democratization of data and data skills. Consequently, IT managers
must develop effective strategies to implement data discoverability and a comprehensive
data catalog with lineage capabilities. Participants 1 and 3 emphasize a strategy focused
on implementing data discovery, data lineage, data classification, and a company-wide
data catalog as part of the data management lifecycle. They underscore the importance of
establishing a business glossary or data dictionary that spans from the business level to
individual applications, extending up to the enterprise level. Although essential, this
aspect of data management has often been deprioritized in their experiences. They argue
that understanding what data you have, where it resides, and its nature is crucial, as it
enables other functions like data lineage and classification to be effective. Data lineage,
in particular, is vital in regulated industries such as healthcare and finance. Data
classification, which includes categorizing information as highly classified, internal, or
public, is typically handled by information security groups. Participant 4 also supports
the implementation of data lineage and a data catalog but highlights the significant
challenges companies face in executing these strategies. The lack of automation, coupled
with the sheer volume of data and various drivers, makes data governance complex and,
at times, constrained. To overcome these challenges, Participant 4 suggests adopting the
right tools and technologies to facilitate cataloging, lineage tracking, metadata
management, and policy enforcement. Automation plays a critical role in reducing
manual effort, which, in turn, contributes to the success of data governance initiatives.
Additionally, transparent reporting on data governance metrics is crucial to
demonstrating the value of data quality efforts.
Table 9
Frequency of Second Major Theme
Source of data collection n
Interviews 11
Documents 4
Note. Theme 2: Data discoverability, data cataloging, and automation are emerging as
significant motivators for successful DG implementation; n = frequency.
Methodological triangulation was achieved for Theme 2, with all (4 out of 4)
articles and reports affirming the necessity of data discoverability and data catalogs for
companies. The Atlanta Regional Commission (2019) highlighted the lack of data
discovery—specifically, knowing who possesses what data—as a significant challenge
for agencies when sharing data with stakeholders. To address this, a data governance
framework was proposed to establish protocols and rules that enhance data sharing
efficiency by adopting data discovery services for cross-organizational data searches.
ESG (2022) identified key challenges in data management and governance, including
visibility and data quality issues. The study revealed that 42% of respondents reported at
least half of their data as “dark,” while 46% expressed a lack of confidence in the quality
of source data, hindering effective data utilization across their value chains. Additionally,
35% of respondents cited the lack of automation as a barrier to implementing data
governance. Overall challenges summed up to cataloging for data discovery (26%),
access control (24%), and visibility (23%). The OvalEdge Team (2022) identified that
the absence of data discovery tools hindered companies from identifying critical
information within large datasets, impeding the development of new data products.
Additionally, the lack of a centralized data catalog containing standardized business
terms made it difficult for non-technical users to analyze data, contributing to low data
literacy across the organization, data business terms, made it difficult for non-technical
people to analyze data, contribution to low data literacy across the company.
This Theme 2 aligns with the Metadata and Data Lifecycle dimensions of the
Data Governance theories outlined by Khatri and Brown’s (2010) Unified Framework
for Data Governance used in this study. The metadata domain involves capturing and
managing data descriptions. Metadata provides essential context and understanding of
data elements, facilitating effective data governance practices and enabling efficient data
discovery and utilization. Data catalog and lineage maps to the Master Data Management
and traceability aspects of the Data Lifecycle domain. Furthermore, the concept of data
discoverability aligns closely with Khatri and Brown’s (2010) emphasis on data quality
and stewardship. Ensuring that data is easily discoverable within an organization
enhances its quality by enabling better data usage and decision-making. Data catalogs
can be seen as a practical implementation of the Unified Framework’s principles,
particularly in relation to decision rights and data stewardship. By cataloging data assets,
organizations can better manage who has access to what data, enforce accountability, and
maintain data quality standards.
Theme 3: Successful DG Strategies Treat Third-Party Data Assets With the Same
Rigor as Internal Data Assets to Preserve Security and Privacy
The analysis revealed a key theme: from a governance standpoint, third-party
data should be managed with the same rigor as internal data assets, with additional
checks to ensure security and privacy. Businesses and organizations generate data
internally but also rely on external sources and providers. Third-party data, collected by
entities without direct relationships with the individuals or businesses to whom the data
pertains, is typically aggregated from various external sources. This data is widely used
in digital marketing, advertising, and analytics to enhance targeting, audience
segmentation, and personalization efforts. Other sectors, such as healthcare,
pharmaceuticals, and finance, also contribute to third-party data sources. Therefore,
third-party data should be included in an organization’s Data Governance programs and
processes. From a data management perspective, it’s essential to establish quality
standards and ensure privacy protection through Governance, Risk, and Compliance
(GRC) practices, with a particular focus on maintaining security and privacy.
All interview participants agree that third-party data sets should be treated as
another data source from a data governance perspective, with additional checks to ensure
contractual security and privacy. Participant 1 emphasized the importance of
understanding the standards, obligations, and applicable regulations when onboarding
third-party data, particularly regarding contractual usage rights. Many third-party data
sets, especially in regulated industries like healthcare and finance, come certified for
quality and completeness, such as SOC 1 (financial controls) and SOC 2 (availability,
security, processing integrity, confidentiality, and privacy) standards in the banking
sector. Participant 2 highlighted the need for thorough data governance quality checks to
ensure accuracy, cleanliness, and completeness. Participant 3, drawing on banking sector
experience, agreed with Participant 2 and stressed that the acquiring company is
responsible for the governance of any data it uses, whether internal or third-party.
Therefore, additional checks, including data quality and data movement checks, are
necessary. Participant 4 echoed Participant 3’s views, emphasizing that once third-party
data enters an ecosystem, the company assumes full liability for any breaches, security
issues, privacy concerns, and compliance. Participant 6 stressed the importance of
including checks to remove redundancy, while Participant 7 noted that third-party data
often contains sensitive information, such as PII, and that proper anonymization
techniques must be established. For HIPAA-regulated data, data protection and handling
standards must be strictly followed to ensure privacy and confidentiality. Participant 9
stated that “third-party governance is an integral part of data security and privacy
governance.” Participant 11 further discussed the importance of a structured process for
handling third-party data, emphasizing the need for close collaboration with other
stakeholders who have expertise in legal, privacy, and IT aspects. This process involves
three layers: legal (understanding legal implications and contractual obligations), privacy
(ensuring compliance with GDPR, CCPA, etc.), and IT (handling data engineering and
analytics based on expectations set by the legal and privacy teams).
Methodological triangulation for Theme 3 was not achieved, as none of the
articles or reports specifically addressed third-party data sets in the context of data
governance.
Table 10
Frequency of Third Major Theme
Source of data collection n
Interviews 11
Documents 0
Note. Theme 3: Successful DG strategies treat third-party data assets with the same rigor
as internal data assets to preserve security and privacy; n = frequency.
Theme 3 focused on the governance of third-party data, which requires
heightened checks on security and privacy, more so than data quality. In Theme 1, I
extensively discussed the three key drivers for data governance. Similar to Theme 1,
findings from Theme 3—particularly regarding data quality, security, and privacy—
closely align with the Data Principles, Data Quality, and Data Access domains in Khatri
and Brown’s (2010) framework. Additionally, there is a process element concerning how
organizations manage third-party data to meet contractual obligations, involving
collaboration with data owners, stewards, and legal, privacy, and IT departments. This
process aspect corresponds with the Data Policies component of the Governance
Mechanisms dimension in Abraham, Schneider, and Vom Brocke’s (2019) Conceptual
Framework for Data Governance, which includes formal guidelines for managing,
accessing, and utilizing data across the organization, with specific rules on privacy,
security, and quality.
Theme 4: Successful DG Strategies Tie the Value of Data as an Asset to Its Impact
on the Bottom Line
Most companies recognize the value of data, but traditionally, they only treat it as
an asset when it impacts the bottom line—either as a revenue generator (e.g., data
products or AI-driven innovation) or as a liability (e.g., risk and compliance). However,
with the rise of AI, companies are increasingly leveraging data to drive innovation and
enhance value. Governance strategfies in the AI era focus on innovation with two
aspects. First, companies focusing on AI innovation understand the critical need for data
quality, security, and privacy, and are increasingly recognizing the importance of data
governance (DG). Secondly, the growing significance of DG is particularly evident in
the context of big data.
The participants collectively recognize the growing importance of treating data as
a strategic asset within organizations, yet they highlight different perspectives and
challenges in achieving this. The participants agree on the importance of data as an asset,
but they emphasize the need for robust governance, cultural shifts, and a holistic
approach to fully realize its value in driving business transformation and innovation.
Participant 1 emphasized the evolving recognition of data as a valuable asset, with
companies increasingly exploring ways to monetize it. Three main approaches include
cost reduction through internal improvements, selling or licensing data, and enhancing
products through AI-driven insights. Participant 2 pointed out that while many
companies acknowledge the value of data, cultural and governance challenges often
prevent them from treating it as a true asset. Companies that directly monetize data
naturally view it as an asset, but others struggle due to fragmented data silos and
ineffective governance, often leading to missed opportunities.
Participant 3 highlighted that data governance is a relatively recent field and
underscores the necessity of understanding and ensuring data quality to genuinely
consider it an asset. The value of data as an asset is closely tied to its quality and
relevance. Participant 4 distinguished between organizations that merely claim data as an
asset and those that actively leverage it to transform their business. A truly data-driven
organization uses data to enhance efficiency, drive growth, and improve products and
services. Participant 5 noted that historically, data was not treated as an asset, but the rise
of AI has shifted this perception. The AI boom has made companies more aware of
data’s value, although many still treat data as an afterthought, addressing issues
reactively rather than strategically. Participant 6 reinforced this shift in perception,
noting that companies are now recognizing data as a valuable asset due to factors like
regulatory fines, advanced analytics, and AI advancements. Participant 11 emphasized
the need to understand data as a 360-degree asset, considering not only its production but
also its usage and consumption. Effective data governance requires a comprehensive
approach, recognizing the interconnectedness of data within an organization, rather than
viewing it in isolated parts.
Methodological triangulation was achieved with a single organizational
documents and reports fully supporting this theme. Data is a critical asset for any
business and is essential to a wide range of users across the enterprise, from executives
and line-of-business teams to IT decision-makers and staff (ESG, 2022). Many
organizations face challenges with data quality, including issues related to accessibility
and visibility, and often have large amounts of unutilized “dark” data. This underscores
the need for robust data governance solutions that offer transparency across the data
landscape, integrated with business context.
Table 11
Frequency of Fourth Major Theme
Source of data collection n
Interviews 11
Documents 2
Note. Theme 4: Successful DG strategies tie the value of data as an asset to its impact on
the bottom line; n = frequency.
The extensive literature on Data Governance reinforces Theme 4’s findings,
which position data as a strategic organizational asset. Khatri and Brown (2010) define
data as a strategic enterprise asset, emphasizing the need for governance to establish
accountability and regulate access and processes related to data (Walsh et al., 2022).
Similarly, Abraham et al. (2019) view data as a strategic asset, advocating for a
crossfunctional governance framework that specifies decision rights, formalizes policies,
and monitors compliance. Both frameworks underscore the necessity of applying
governance to maximize the value of data. Nielsen et al. (2017) reviewed 62 peer-
reviewed publications and conference proceedings, defining data governance as
company-wide processes that align decision-making rights and responsibilities with
organizational goals for effective data management. Satar (2021) highlights the critical
importance of Master Data due to its valuable organizational insights, ranking it as a top
priority for management. Yebeness and Zorilla (2019) also emphasize that effective
governance is essential for maximizing data’s value as a key organizational asset.
Theme 5: Executive Sponsorship Is Critical for Successful DG Implementation
Executive support, particularly an executive sponsor, is crucial for the success of
data governance (DG) programs. Without it, there is a lack of organizational alignment,
authority, and budget support, leading to program failure. Gaining executive support and
their buy in is one of the top challenges for DG programs. IT managers must identify an
executive sponsor and develop strategies to secure their backing. One effective strategy,
as revealed by the analysis, is to demonstrate the value of DG through use-case-driven
implementation. This can involve starting with small-scale programs or proof of concept
(POC) with modest budgets to showcase the impact and value of DG, particularly its
influence on AI innovation.
All participants agree that an executive sponsor is crucial for the success of a data
governance program. Without executive backing, even well-designed plans may falter
due to insufficient budget and resources (Participant 10). The Nephos Technologies
(2022) report also highlights the lack of executive support as a significant barrier to
adopting data governance programs. According to the report, while 98% of respondents
have ongoing governance initiatives, 33% feel they do not receive adequate executive
support. Participants discussed various strategies for gaining executive sponsor buy-in
for data governance:
1. Pilot Projects and Proof of Concepts: Starting with a pilot project or
smallscale proof of concept (POC) can effectively demonstrate the value of
data governance (DG) to executive leadership. Using key performance
indicators
(KPIs) for reporting can further illustrate the benefits and secure support
(Participants 1, 3).
2. Company-Wide Integration: Although data governance ideally should be a
company-wide initiative supported by executive leadership and embedded in
the company culture, it is often siloed with departments driving use cases
specific to their needs (Participant 2).
3. C-Suite Involvement: Securing buy-in from C-suite executives is crucial.
Data governance should not be solely IT-driven; it requires strong support
and involvement from top leadership. An executive sponsor or champion can
address challenges and advocate for data governance across the company
(Participants 4, 5).
4. Socialization and Incentivization: Having an executive champion can help
build belief in data governance through initiatives like roadshows,
newsletters, and gamification. These low-cost efforts aim to socialize and
incentivize data governance, making it more engaging for the business
(Participant 5).
5. Adoption Across Business Units: For data governance and related initiatives
(e.g., data quality, cataloging) to succeed, they need adoption across all
business units. Grassroots efforts alone are insufficient without executive
support (Participant 6).
6. Define clear Roles in Data Governance:
•Executive Sponsor : Champions the overall program
•Data Owner: Accountable for data.
•Data Steward: Manages data on a daily basis.
•Data Custodian: Handles technical aspects and system-specific
responsibilities.
•Data User: Anyone who interacts with data.
•Steering Committee: Provides oversight, similar to a legislative body.
•Data Governance Officer and Committee: Offers strategic guidance and
oversight (Participants 6, 9).
Table 12
Frequency of Fifth Major Theme
Source of data collection n
Interviews 11
Documents 2
The extensive literature on Data Governance reinforces Theme 5’s findings,
identifying the criticality of executive sponsor for the success of data governance
programs. Data governance requires the active involvement and commitment of all staff,
with robust support from both management and senior-level executives (Al-Ruithe et al.,
2019). Yebenes and Zorilla (2019) define the executive sponsor as the individual or
department responsible for driving, funding, supervising, and guiding the data
governance (DG) initiative. The executive sponsor also clarifies the initiative’s scope,
establishes milestones and goals, and ensures compliance with all relevant data laws and
regulations. Satar (2021) highlights top management support as a critical factor for the
success of Master Data Quality programs, emphasizing that top management must
recognize the importance of master data quality and endorse related management
activities. Abraham, Schneider, and Vom Brocke’s (2019) Conceptual Framework for
Data Governance categorizes the executive sponsor and other roles within the
Governance Mechanisms dimension. This framework identifies structural governance
mechanisms that define roles, responsibilities, and accountabilities for data governance.
Key roles and governance bodies include the executive sponsor, data governance leader,
data owner, data steward, data governance council, data governance office, data
producer, and data consumer. The executive sponsor plays a crucial role in providing
strategic direction, prioritizing business needs, and securing funding for data
management initiatives.
Theme 6: Effective Data Governance Strategies Must Address the Challenge of
Justifying the Value of DG Programs
The analysis highlighted a significant challenge for IT managers: justifying the
value of Data Governance (DG) programs to both executive leadership and business unit
managers. This challenge primarily stems from a lack of understanding among key
stakeholders. Executive leaders and lower management often struggle to see the
immediate ROI of DG, especially when it isn’t directly tied to compliance issues or risks
like data breaches. They’re concerned about the program’s length and potential to slow
down operations. Middle managers and end users are often even less familiar with DG.
To overcome this, IT managers must implement strategies to secure buy-in from both
executive leadership and mid-level management. Interview participants recommended
several approaches, including using agile methodologies for DG processes, focusing on
specific use cases rather than trying to address everything at once, building proofs of
concept before full implementation, and educating the entire organization on the benefits
of DG—such as improved data quality, security, and privacy. Additionally, employing
KPIs to measure and report on DG’s impact can effectively demonstrate its value. In
conclusion, reducing risks related to data security and privacy while ensuring regulatory
compliance is essential. Defining a KPI matrix to evaluate the value of data governance
(DG), along with smaller, targeted implementations, can help IT managers secure the
necessary support from both executive and middle management.
Five participants emphasized the challenge of justifying the value of DG to
organizations. Participant 4 noted that many companies treat DG as an afterthought to
their existing data strategy and transformation efforts, making it difficult to gain
executive and business buy-in. Successful DG implementation often requires extensive
collaboration with business stakeholders and IT, as well as ensuring projects stay within
budget and on schedule. A recommended strategy is to start small, perhaps with a proof
of concept, and demonstrate value through measurement and monitoring. Similarly,
Participant 6 pointed out that the challenge often lies in how DG is perceived. Many
view it as dull or overly regulatory, disconnected from tangible business outcomes. It can
be difficult to convey the value DG brings until stakeholders see concrete results. To
address this, IT managers should focus on using business language rather than DG
jargon. Instead of saying “let’s implement data governance,” they should frame it as
improving processes, saving time and money, and automating tasks. Additionally,
avoiding a textbook approach to DG and tailoring it to specific contexts and use cases
can further enhance its perceived value.
Participant 7 also noted that people often fail to understand the value of data
governance (DG) until they see quick wins. When presenting DG to executives,
discussions typically revolve around the practice’s involvement of people, processes, and
technologies. However, executives often ask for the vision and roadmap, and when told
that full implementation might take 12 months, it’s a hard sell—especially when
justifying the time, budget, and resources required. Executives usually evaluate DG’s
ROI against the company’s bottom line and are concerned about the length and potential
slowdowns associated with DG initiatives. To address this, the participant suggested
using an agile methodology rather than a waterfall approach. By delivering value
quickly, IT managers can gain executive buy-in. For instance, the first few sprints could
focus on establishing a data quality control and monitoring system. Breaking the process
down into milestones helps executives understand and appreciate the incremental value
being delivered. The participant also mentioned a selling strategy based on current
organizational pain points, such as poor data quality. For example, if poor data quality is
causing a significant financial loss, the case for DG becomes more compelling. By
framing DG as a solution to overcome these challenges, executives are more likely to see
the big picture and the long-term benefits of implementing a DG framework. Regulatory
compliance is another area where DG is easier to justify to executives. Failure to comply
with regulations directly impacts the bottom line, making it a more straightforward sell.
The participant suggested recommending regular audits of tools, technologies, processes,
and policies to prevent security breaches and ensure compliance. These strategies are key
to winning executive support for DG initiatives. Participant 10 also mentioned that
securing executive support for data governance (DG) is challenging and emphasized the
need to clearly demonstrate the program’s value to gain their backing. Similarly,
Participant 11 highlighted that the biggest hurdle is achieving buy-in, compounded by the
lack of standardized reference implementations for DG. Despite the existence of some
frameworks, there’s no universally accepted reference architecture, leading organizations
to develop their own approaches. Table 13
Frequency of Sixth Major Theme
Source of data collection n
Interviews 5
Documents 0
Note. Theme 6: Effective data governance strategies must address the challenge of
justifying the value of DG programs; n = frequency.
The challenge of justifying the value of data governance (DG) aligns with both
the Organizational Scope and Consequences dimensions of the Data Governance
Conceptual Framework by Abraham, Schneider, and Vom Brocke (2019). In terms of
Organizational Scope, the focus is on the intra-organizational level, where DG is
implemented within a single organization. This involves managing the quality, integrity,
security, privacy, and compliance of organizational data assets. Achieving this requires
collaboration with various stakeholders, business groups, and IT departments to address
their interests and secure their buy-in. However, this is often challenging due to the
diverse needs and interests across the organization. Business use cases vary, and
perceptions of ROI can differ significantly between mid-level management and
executives. The Consequences dimension addresses two key areas: intermediate
performance effects and the management of data-related risks. The performance effects
of DG refer to improvements in areas such as data quality, data access, and metadata
management, as outlined by Kathri and Brown (2019). Enhancements in data quality—
such as accuracy, availability, completeness, consistency, and timeliness—are tangible
and easier to recognize. The second aspect is the management of data-related risks,
particularly those arising from non-compliance with regulations, as well as security and
privacy breaches. Without concrete implementations and measurable outcomes, it can be
challenging to convince executives and IT managers of the positive impacts of DG
upfront. This underscores the core challenge highlighted in this theme.
Theme 7: Successful DG Implementation Requires Relevant Data Competencies for
All DG Roles
The success of data governance (DG) relies heavily on data competencies. Data
stewards and data owners need strong acumen in data management and a solid
understanding of data lifecycle components. IT managers implementing DG programs
should prioritize training participants in essential data skills, including data querying,
data classification, data analysis, and data reporting.
Two participants highlighted the importance of technical and data skills in their
responses, a point corroborated by two reports used for methodological triangulation.
Participant 7 emphasized the need for both technical skills to use data governance (DG)
tools and data skills to understand key aspects such as data quality, classification, and
security. For example, when IT integrates technology to implement the data lifecycle,
including governance, there is often a focus on the technology itself, sometimes at the
expense of the people side—change management and user adoption, which are crucial
for success. Participant 7 noted that while consultants may provide solutions for
challenges like data integration, quality, and lineage cataloging, these efforts can fail if
the people expected to own and maintain the solutions lack the necessary understanding.
Often, once the consultants leave, the organization struggles because those responsible
for the solution haven’t received proper training to maintain or enhance it. Participant 10
reinforced this point, suggesting that individuals hired to manage DG, whether on the
technical or process side, are not always experts in the field. DG is often mistakenly seen
as a project management task, but it requires deep technical knowledge. The right person
must have data skills, such as SQL proficiency and a solid understanding of data
platforms, to effectively run a DG program. Nephos Technologies (2022) further
supports this by identifying people as one of the top barriers to successful DG. Their
survey of DG professionals in the UK revealed that 52% of respondents felt they lacked
the necessary skills to deliver effective data governance. Similarly, ESG (2022) reported
that a skills shortage or gap represents 40% of the barriers to DG success.
Table 14
Frequency of Seventh Major Theme
Source of data collection n
Interviews 2
Documents 2
Note. Theme 7: Successful DG implementation requires relevant data competencies for
all DG roles; n = frequency.
Abraham et al. (2019) identified key aspects of data management that data
governance aims to address, including Data Ownership, Data Stewardship, Data Policies
and Standards, Data Quality Management, Data Security and Privacy, Data Compliance
and Risk Management, and Data Integration and Interoperability. These aspects align
with the Governance Mechanisms dimension of their framework, which includes
structural, procedural, and relational mechanisms—essentially the processes and policies
that guide data governance. Each of these aspects requires specific skills to ensure
successful implementation. Without the appropriate expertise, data governance efforts
are likely to struggle or even fail.
Theme 8: Successful DG Programs Should Have a Clear Strategy, Be Targeted, Use
Case Bound, and Be Antecedent Driven
The final theme that emerged from the analysis complements, rather than aligns
with, the conceptual frameworks discussed earlier (Alhassan et al., 2019). First, an
established data strategy is crucial. While a fully developed data lifecycle is not required
for a data governance (DG) program, DG must function within the context of existing
data infrastructure and processes. A clear data strategy is essential for guiding any DG
roadmap and implementation. Second, DG implementation should be targeted and use
case-bound. Begin with a focused, specific use case driven by a business need or
challenge, such as improving product sales. The use case should directly address this
need.
All participants agreed on the importance of establishing a clear data strategy as
the foundation for a data governance program. However, their strategies varied. Some
advocated for starting small, while others emphasized embedding data governance from
the outset to enhance compliance, privacy, security, and risk mitigation. Additional
approaches included making data governance antecedent-driven, adopting a “governance
by design” model, and following a “don’t boil the ocean” strategy.
Participant 1 emphasizes an agile, iterative approach, embedding DG into the data
lifecycle from the outset to ensure compliance, enhance privacy and security, and
mitigate risks, particularly in regulated industries. Participant 2 echoes the need for a
clear data strategy but stresses that DG should be driven by specific business needs or
challenges, rather than being IT-led, to empower better data management and
decisionmaking. Participant 3 aligns with Participant 1, advocating for embedding DG
from the ground up in data management processes to ensure a comprehensive
understanding of data sources and usage. Participant 4 highlights that many companies
treat DG as an afterthought, which is problematic without executive and business buy-in.
They stress that securing this support should be a priority. Participant 5 also supports a
“governance by design” approach, where governance controls are integrated from the
start, with a focus on one use case at a time. They emphasize the importance of
budgeting for DG, having the right tools, and engaging stakeholders early. Participant 6
and Participant 7 advocate for starting with small, specific use cases rather than large,
enterprise-wide programs, which can be costly and prone to failure. They recommend a
comprehensive data strategy as the first step. Participant 8 suggests beginning with data
quality, lineage, and master data management, aligning with FAIR principles, and
iteratively implementing DG. Participant 9 underscores that strategy is fundamental,
guiding the creation of tactical steps for implementation, including prioritizing based on
business alignment, defining roles and responsibilities, and measuring success with KPIs.
Participant 10 advises starting small to build momentum and gain executive support,
using a proof of concept to demonstrate the value of DG. Finally, Participant 11 stresses
that DG should be integrated from the start as an ongoing process, with buy-in from data
producers who are responsible for governance implementation.
Table 15
Frequency of Eighth Major Theme
Source of data collection n
Interviews 9
Documents 0
Note. Theme 8: Successful DG programs should have a clear strategy, be targeted, use
case bound, and be antecedent driven; n = frequency.
Applications to Professional Practice
Data Governance (DG) is increasingly crucial, particularly in the context of
advanced analytics, AI, security, privacy, and regulatory compliance. As a strategic
enterprise asset, data requires a cross-functional DG framework that outlines decision
rights and responsibilities for managing it effectively (Abraham et al., 2019). This study
offers valuable insights for IT managers, identifying strategies to overcome the key
challenges in DG implementation. By gathering knowledge and experiences from
seasoned DG professionals across various industries, the study highlights key drivers and
practical considerations essential for successful DG initiatives.
The study’s findings are significant because they align with the data governance
decision domains outlined in Khatri and Brown’s (2010) Unified Framework for Data
Governance and the dimensions of the Data Governance Conceptual Framework by
Abraham, Schneider, and Vom Brocke (2019). These frameworks provide guiding
principles for understanding the implementation of DG, including its challenges and
successes, through six dimensions of analysis. Methodological triangulation with several
industry reports on DG confirmed the study’s findings.
Effective strategies such as data discoverability and data cataloging with lineage
are essential for IT managers to meet regulatory compliance requirements, particularly in
regulated industries. A major challenge for DG managers is often justifying the value of
DG to executive leadership to secure their buy-in. This study suggests strategies to
address this challenge, such as starting small, focusing DG programs on specific use
cases or data domains, building pilot projects and proof of concept programs, recording
performance KPIs, and demonstrating value. Additionally, the study highlights the
importance of training and socializing data skills to empower those in DG roles to
implement and drive programs successfully. Socialization and incentivization, supported
by an executive champion, can help incorporate DG into the organization’s culture and
processes.
Some participants in the study recommended a shift in mindset towards data
engineering with DG embedded from the outset. This approach is akin to the early 2000s
when software was often designed without security in mind, leading to frequent
malicious attacks. Companies like Microsoft led the industry in changing this approach
by designing software with security at its core. The study suggests a similar shift for DG,
advocating for its integration from the beginning rather than as an afterthought in data
management. These strategies enhance the critical success factors discussed in the
study’s literature (Alhassan et al., 2019), providing IT managers with additional
knowledge to implement successful DG programs.
Implications for Social Change
Technological innovation should prioritize societal well-being, driving positive social
change through enhanced data quality, security, privacy, and the democratization of data
governance. Effective data governance improves the quality, trustworthiness, and exchange
of data (Hikmawati et al., 2021), enabling companies to harness advanced analytics and AI to
extract value from Big Data. This study identified data quality, security, privacy, and
regulatory compliance as key drivers of data governance.
By focusing on these drivers, IT managers can contribute to better data quality across
various sectors, including healthcare, finance, and retail. Stronger data governance helps
prevent data breaches that could harm institutions and, ultimately, consumers. Additionally,
high-quality data is crucial for the AI-driven innovations transforming industries and
enhancing people’s lives. Improved data governance also ensures compliance with
regulations, helping companies avoid costly fines that could be passed on to consumers.
Recommendations for Action
This study confirmed that data is a strategic organizational asset that must be
governed with the same priority and rigor as any other company asset. Data quality,
security, and privacy are critical for both companies and users to enable strategic
business decisions and drive innovation. Abraham et al. (2019) emphasized the
importance of data governance in ensuring that data assets are properly managed to
maximize value and minimize risks.
Company executives must recognize data as a strategic asset within the data
management ecosystem and allocate sufficient resources to enhance its quality, security,
privacy, and compliance. Crucially, they must move beyond treating data governance as
an afterthought and instead fund robust programs that address the challenges and needs
associated with these key drivers. IT managers should focus on implementing data
discoverability, data cataloging, and the standardization of data terms across the
organization, while also democratizing data skills.
To overcome the key challenge of securing executive buy-in for DG, IT managers
should adopt the strategies unearthed in this study. In addition to following best practices
outlined in critical success factors (CSF), IT managers can demonstrate the value of data
governance through quick wins, agile methodologies, proof-of-concept initiatives, and
pilot programs, all supported by clear KPIs. Finally, technical and data skills are
essential for the success of data governance. IT managers must seek funding and
resources to properly train those in DG roles, ensuring they can effectively execute data
governance programs.
As promised in the interview protocol and during the interviews, I will share the
results of this study with the participants, many of whom expressed eagerness to learn
about the findings. My hope is that this cumulative knowledge will assist them in their
DG implementation efforts and benefit other professionals as well. I also plan to explore
ways to disseminate the study’s findings to a broader audience, including publishing a
summary in the Data Governance Professionals Organization (DGPO) journal and other
relevant publications. Additionally, I intend to reach out to Yebeness and Zorilla (2022)
to discuss strategies for Big Data governance.
Recommendations for Further Study
This study highlighted several areas that warrant further exploration to deepen the
understanding and enrich the body of knowledge on Data Governance (DG). One key
area is DG for Big Data. Yebeness and Zorilla (2022) laid a solid foundation with a
framework and tentative reference architecture for understanding the motivations and
challenges of DG in the context of Big Data. Future research could build on their work
by focusing on the specific challenges IT professionals face in implementing DG for Big
Data and Cloud Data. Another important area for further study is the implementation of
data discoverability and data catalogs as enablers of effective data governance.
Interviews with IT managers and professionals could provide valuable primary data and
insights across industries, shedding light on the needs and challenges associated with
these aspects of DG.
Reflections
My primary motivation for choosing this research topic was to gain a deep
understanding of the theoretical foundations of Data Governance (DG), enabling me to
better grasp its challenges, opportunities, and importance in data management. As a
professional data engineer and architect, I have developed solutions for clients where DG
was claimed to exist, supported by a dedicated team. However, my assessment of data
quality and processes often revealed significant gaps. This research allowed me to
identify common challenges companies face, such as siloed organizations, DG efforts
driven by individual business teams, and a lack of executive support, budget, and buyin
—all factors that hinder the success of DG programs.
Through interviews with experienced DG managers across industries and an
extensive review of the literature, I have come to fully appreciate the critical role of DG
in managing data as a vital enterprise asset. This newfound knowledge will enable me, as
a data engineering professional, to design data platform solutions with DG integrated
from the ground up.
Summary and Study Conclusions
The primary motivation for this study was to explore strategies employed by IT
managers to effectively implement data governance (DG) programs. The findings
confirmed that data quality, security and privacy, and regulatory compliance are the main
drivers behind DG initiatives. Additionally, the study revealed a growing trend:
companies are increasingly treating data as a strategic asset, a shift from previous
practices.
Two key challenges to DG programs were identified: securing executive support
and justifying the value of DG initiatives to both leadership and mid-management. The
data collected through participant interviews provided actionable strategies for IT
managers. Chief among these is winning executive buy-in by demonstrating the value of
DG programs through pilot projects, proof of concepts, and agile methodologies.
Another successful approach involves designing DG programs with a clear, use-case-
driven strategy. Methodological triangulation was employed to validate the themes and
findings, comparing them against industry DG reports from various companies and
organizations.
This study offers valuable insights for IT managers, highlighting strategies to
overcome the main challenges in DG implementation, particularly in securing executive
support, obtaining funding, and demonstrating ROI through quick-win strategies. The
study’s findings also have broader implications for social change by emphasizing the
role of effective data governance in enhancing data quality, security, and privacy,
ultimately protecting end users from identity theft and misuse of personal data.