For this assignment, I chose to research and discuss “dirty data”. The main article I chose to read
and analyze is “A Taxonomy of Dirty Data” by Won Kim, Choi Byoung-Ju, Hong Eui-Kyeong,
Kim Soo-Kyung, and Doheon Lee. This article was written to help the audience understand
what dirty data is, where it comes from, how to pinpoint it and how to mitigate it. The article
starts with the explanation of data cycle- that includes capture, storage, update, transmission,
access, archive, restore, deletion and purge (Kim et al., 2003). Dirty data is categorized as such
when the end user received data that does not correspond with the parameters set or receives
wrong results. Understanding the causes of dirty data is also important. Data can be dirty or
wrong based on many factors such as bad data input (human or computer), incorrect data update,
data transmission errors, or bugs or corrupt material in data processing system (Kim et al., 2003).
This article digs deep into the taxonomy of dirty data and how it impacts data mining. Many
companies today use ERP systems that houses data, but is the data clean and free of duplication,
wrong information and radicals. This is when we need to understand our ERP systems, how and
where the information is coming from, and how to request information from it’s data warehouse
with similar language and specific queries.
According to our textbook, “dirty data” or bad data is a well-known problem in AIS. Our
textbook describes that “bad data” occurs “when instances of missing data, extra data, or
incorrect data occur, preventing the data from being of optimal use for decision-making”
(Gelinas et al., 2018). The textbook uses the U.S. Post Office as an example and frankly that is
still an ongoing problem. I send vendor payment on a weekly basis and there are many instances
that weeks later, the payments still have not been received. At this point, I don’t know if the
blame is due to bad data or laziness in general.
While researching “dirty data”, I came across another article from the CPA Journal “Detecting
and Resolving ‘Dirty’ Data” by Dana Hermanson, James Lawson, and Daniel Street. This article
is interesting as it discusses how businesses are more and more dependent on technology to make
operational and business decisions. Yet the question that is asked in this article is, “how can a
business solely rely on clean set of data if the data is incomplete, invalid, or simply inaccurate”
(Hermanson et al., 2022). CPAs provide business insight on data analytics depending on the
validity, completeness, and accuracy of underlying data. CPAs cannot sign off or independently
attest on accuracy or completeness of information if the data is dirty. This article discusses how
ignoring dirty data can result in higher costs in the future. Identifying dirty data and resolving
the problem is not easy feat and takes time but once fixed, it will benefit the business with
accurate and precise information based on clean data. The article discusses a 10 step process and
recommends that public, private, big and small business pursue these steps. The most critical
step is obviously identifying the root cause of the dirty data. The ten steps are: 1)understanding
the business process represented by the data, 2)analyze the source and process of the data,
3)determine which elements the data set should contain, 4)scan a sample of recent data,
5)summarize the data of each table, 6) document expectations for the data within each field,
7)document expectations for the relationships among fields, 8)test each expectation by creating
exception reports, 9)evaluate exceptions and identify root causes, and 10)act to clean the data
and address root cause. When looking at these steps, we can assume that most businesses follow
some or all of these steps, but the number one problem is understanding the business setting
which will help understand what data is expected or needed in order to get to the root cause of
the problem. Understanding how data is input into the ERP system is another key to solving the
problem. Is the entry manual or is AI involved in the process. Looking at a visual system flow
chart will help with understanding what the data represents or should represent.
“Having clean data to work with-data that’s free of duplication, broken sequences and other
discrepancies or inaccuracies, like system errors and process gaps- is so critical to success” (CPA
Canada, 2022). According to this article, dirty data is everywhere and it’s not just an IT person’s
problem, it’s a business integration problem. Data input based on business intelligence can curb
financial risks and reduce the chance for business errors. “Clean data generates powerful
analytics that are responsive to change, and that translates into more accurate solutions at every
step- to improve operations, understand sales trends or identify patterns for customer
personalization to boost profit margins” (CPA Canada, 2022). While understanding that clean
data starts at the inception of a CRM or ERP system. It all starts with the configuration of the
CRM or ERP system- correctly configuring the data can help mitigate dirty data and provide
clean data entry. While setting up a new ERP or CRM system, having a understanding of the
business model is crucial to setting up new systems. Once the new system is in place, to help
mitigate dirty data, on the job continuous training must be done. With everything being
computerized, accountants and IT professionals need to become subject matter experts and
understand their business systems inside and out. Proper controls need to be put in place and
audited several times a year for consistency and accuracy. If dirty data is found, data cleansing
must be performed, which identifies and fixes errors, duplicates, and irrelevant data from a raw
database. Data cleansing allows for accurate, defensible data that generates reliable
visualizations, models, and business decisions. According to It Focus, a company based put of
Europe, a business should cleanse it data every three to six months (if it’s a big company) and at
least once a year if it’s a small company (IT Focus, 2021). Accurate versus inaccurate data can
made a huge difference with business decision-making processes. Proverbs 30:5 says, “Every
word of God is tested; He is a shield to the who take refuse in Him” (ESV, 2011). Every
decision we make, business or personal, is based on some form of information. When that
information is bad or inaccurate, it can have huge impact on our decision-making process and
affect a lot more than just financial aspect of a business or personal finances. If the information
is flawed so are our decisions and end-results.
Cited Works:
CPA Canada. (2022, February 27). Data-driven BI is changing everything for CPAs. CPA
Canada. Retrieved February 19, 2023, from
https://www.cpacanada.ca/en/members-area/profession-news/2020/february/data-driven-
bi-changing-everything-for-cpas
English Standard Version. ESV Bible. (2011). Web
Gelinas, U. J., Dull, R. B., Wheeler, P. R., & Hill, M. C. (2018).JAccounting information
systemsJ(11th ed.) Boston, MA: Cengage Learning
Hermanson, D. R., Lawson, J. G., & Street, D. A. (2022, October 25). Detecting and resolving
'dirty' data. The CPA Journal. Retrieved February 19, 2023, from
https://www.cpajournal.com/2022/10/25/detecting-and-resolving-dirty-data/
How often should you spring clean your data? IT Focus Telemarketing. (2021, April 27).
Retrieved February 19, 2023, from https://www.itfocus-tm.com/how-often-should-you-
spring-clean-your-data/#:~:text=A%20large%20business%20will%20collect,at%20least
%20once%20a%20year.
Kim, W., Byoung-Ju Choi, Eui-Kyeong Hong, Soo-Kyung, K., & Lee, D. (2003). A Taxonomy of
Dirty Data.&Data Mining and Knowledge Discovery,J7(1), 81-99.
https://doi.org/10.1023/A:1021564703268
Powered by TCPDF (www.tcpdf.org)