1 / 64100%
Ethical Implications of Data Mining by
Government Institutions
Introduction
Data mining can be defined as the process of
extracting useful information from the large
amount of data stored in databases. With the
rapid development in computer and information
technologies, large databases are used to gather,
store and retrieve data (Vaidya & Clifton 2004).
Data mining systems exploit these data sources
for the purposes of creating new knowledge.
Data mining systems have the capacity to search
through large amount of data that would be
meaningless and convert it to useful
information. Though data mining is a knowledge
creation tool, it use for obtaining personal
information has been widely criticized and is
seen as unethical and an infringement to an
individual’s privacy rights.
Most private companies use data mining
techniques to study consumer behaviour so as to
reveal certain trends that can be exploited to
increase their sales and profits. Government
agencies, on the other hand, use information
mining techniques to improve security and
social governance.
Government security agencies combine private
and public databases with personal information
so as to identify patterns that link particular
individuals to terrorism, crime and corruption
(Seifert 2007).
Due to the power vested upon most government
agencies, they have the authority to access
information from both private and public
databases. For example, government agencies
can retrieve data stored by mobile phone
companies so as to track their client calls and
messages. The same case applies to emails and
websites.
Mining of data by government agencies have
long been criticized and is seen as a threat to
information privacy. Critics of personal data
mining insist that it infringes on the rights of an
individual and result to the loss of sensitive
information (Fule & Roddick 2004). Privacy
laws and policies should ideally be used to
protect personal information from any misuse or
transfer.
According to Olson (2007), collecting personal
information is unethical and unlawful as it
violates privacy laws. On the other hand,
government agencies insist that such mining
activities are done with the intention of
improving the national security, preventing
terrorism and improving social governance. In
this paper, a critical analysis of legal and ethical
issues surrounding data mining by government
agencies was done.
Literature review
Various studies show that even though data
mining has many advantages, it use for
surveying an individual without his consent can
be potentially dangerous. Various publications
discuss the importance of data mining to both
private and public bodies. According to
Wahlstorm et al. (2006), data mining is evolving
rapidly due to the improvement in computing
and information technologies.
Methods of data collections have become
sophisticated in the recent past with digital
systems replacing manual methods. These
digital methods include: biometric tags, radio
frequency identification tags (RFID), cell
phones, bar code readers, smart cards and
Geographical Positioning System (GPS)
location. These gadgets have enhanced the
process of data collection.
Wahlstorm et al. (2006) however cautions that,
advancement in computing and information
technology exposes a lot personal data which
can be used without the consent of the owner.
Tavani (2004) adds that, with the current
advancements in technology, more privacy and
ethical issues are arising from data usage,
storage and mining.
According to Wahlstorm et al. (2006), ethical
issues about data mining are a result of exploring
individual’s data. Though most public and
private organizations have strict policies on
personal data privacy, government agencies
have the authority to extract personal
information from private and public databases.
For example, the Austrian taxation office (ATO)
in their quest to investigate fraud among
taxpayers sought for waiver on privacy policies
so as to allow them extract personal information
about an individual’s property, employment,
earning from various databases(Parnell, 2011).
Though the main aim of this ATO excise is to
identify tax evaders, harnessing of personal
financial data is unethical and should not be
done without prior consent.
While government organizations may have a
pertinent reason for mining personal data, it is
widely seen as unethical since it infringes on an
individual’s privacy. Many government
agencies around the world perform large scale
data mining activities for the purposes of
improving security, governance and other social
services.
In the USA, government data mining operations
should follow policies and principles that show
respect to the rule of law. A draft by The
Constitution Project (2010) indicates that the US
government is using advanced data mining
techniques so as to prevent incidences such as
9/11 from occurring.
The report also indicates that this process can
encroach on civil and constitutional rights of
individuals which include privacy, freedom of
expression and protection for all.
The report highlights that innocent people have
been mistaken for terrorist and travel bans
imposed on them while other people use such
information to ruin the reputation of famous
people. For example, in the 2008 USA
presidential elections, private information was
used to ruin the reputation of these candidates
(The Constitution Project, 2010).
Mining of personal medical records of many
individuals by private and public agencies have
always resulted to controversies (Ashwinkumar
2010).
Obtaining medical information from a patient
without his consent is unethical even though
many researchers in the medical field use this
information without prior consent. Parker
(2001) argues that, government agencies and
medical researchers need to inform patients on
the information sought for before obtaining it.
Discussion
The ethical position of data mining by
government agencies for security reasons still
remains a controversial issue. There is a
contradiction between public interest and a
person’s rights and freedom. Though
governments claim that they unearth pertinent
information from personal records, mining this
information is legally wrong, unethical and also
violates the privacy of an individual (Wel &
Royakkers 2004).
Though complete privacy is not possible as
people within a society must communicate,
personal information should only be accessed by
the owner and he should decide on what to share
with other people. It is therefore not morally
right for the government to access and use this
information in attempt to improve security.
If an individual feels that part of his information
is confidential, then, it is inappropriate to access
that data as he may be prone to many
unforeseeable risks. To prevent privacy issues, it
would be prudent for the government agencies
to seek approvals from individuals before using
such data.
Policy and lawmakers around the world give an
individual the right to control the flow of his
personal information. These laws document that
an individual has the right to privacy and any
information obtained from him should not be
transmitted or used for any other purpose except
for the purpose it was obtained for (Sarre &
Prenzler 2005).
For example, a banker should not share credit
card details of his clients with any other entity
without the client consent. Thus, from a legal
perspective, data mining is illegal. It is therefore
ethically wrong for government agencies to use
personal data for secondary purposes.
Data mining methods entail searching through
numerous records from different sources. This
means that the quality and authenticity of this
data cannot be guaranteed. Though data cleaning
and analysis is done, most of the sources of
information may give a wrong perception about
an individual. Data mining is prone to
inaccuracies, errors and poor quality results.
Such low quality data may have severe
implication on an individual or society. Reliance
on such data may result to false accusations
which affect an individual’s career, family and
social life. If a person is labelled as suspect,
negative consequences such as discrimination,
injury or death during shoot out, loss of
reputation and lawsuits occur (Fule & Rodrick
2004).
Trust is another ethical issue that results from
data mining. With increased data mining
activities by government agencies and private
companies, most individuals are gradually
losing trust on these bodies and are unwilling to
share information.
In summary, data mining by government
agencies compromises on an individual privacy
which is unethical and illegal. However, the
government actions seem justifiable in the quest
to fight corruption, crime and terrorism. From
the discussion, there are three possible solutions
to the current issue. First, the government can be
given the right to infringe on personal privacy in
the hope of preventing crime and terrorism.
This would lead to abuse of ethics and privacy
rules for the sake of improving security and
social governance. Though this solution is good,
there is no quantifiable evidence that data
mining will positively identify criminals and
improve security. The second solution entails a
compromise between data mining, ethics and
privacy. That is, allow the government to
perform unethical practices on special cases.
This compromise would set the limits or
situations where the government should access
personal information. Such rules would dictate
the type and amount of data to be collected and
methods of handling it. Though this solution is
good, it does not deter the government from
collecting information through their secret
agencies provided their claim is justified. The
third solution would be to ban data mining of
personal information.
This would mean that the government must rely
on other sources of information for them to
identify and deter crime, corruption and
terrorism. This solution would force government
agencies to uphold good ethical and moral
behaviour and enforce privacy laws.
In order to deal with security and social
governance, the government would be forced to
use other methods to fight terrorism and improve
governance. This would be the most feasible
solution as it upholds good ethics and morals.
Conclusion
In conclusion, the ethical issue of data mining by
government agencies still remains very
controversial. Laws and policy making bodies
support and uphold privacy as one of the key
rights of an individual. It is therefore ethically
wrong for government agencies to infringe on
this basic right.
Though data mining can yield potentially
beneficial results to curtail crime and terrorist
activities, infringing on individual private
information can have detrimental effects and is
unethical. The government should therefore
identify other means of improving governance
and security. The government should also
enforce strict privacy laws deterring all
organizations from collecting personal
information during data mining.
References
Ashwinkumar, M 2010, ‘Ethical and Legal
Issues for Medical Data Mining’, International
Journal of Computer Applications, vol.1.no. 28,
7-10.
Fule, P & Roddick, J 2004, Detecting Privacy
And Ethical Sensitivity In Data Mining Results,
Dunedin, New Zealand.
Olson, D 2007, ‘Ethical Aspects of web Log
Data Mining’, International Journal of
Information Technology and
Management, vol.10.no1, 1-11.
Parkers, S 2001, ‘Legal Aspects of Record
Based Medical Research’, Archives of Disease
in Childhood, vol.89.no.1, 899- 901.
Parnell, S 2011, ATO Seeks Waiver to Hunt Data
on Taxpayers’ Investments. Web.
Sarre, R & Prenzler, T 2005, The Law of Private
Security in Australia, Thomson learning,
Pyrmont, NSW.
Seifert, J 2007, Data Mining and Homeland
Security: An Overview, CRS Report for
Congress, Congressional Research Service,
New York.
Tavani, H 2004, ‘Genomic research and data
mining technology: Implications for personal
privacy and informed consent’, Ethics and
Information Technology, vol.6.no.1, 15–28.
The Constitution Project 2010, Principles for
Government Data Mining: Preserving Civil
Liberties in the Information Age. Web.
Wahlstrom, K, Roddick, J, Vladimir, E. &
Denise, D 2006, On the Ethical and Legal
Implications of Data Mining School of
Computer and Information Science, University
of South Australia, Mawson Lakes, South
Australia.
Vaidya, J & Clifton, C 2004, ‘Privacy-
preserving data mining: why, how, and when?
Security & Privacy’, IEEE, vol.3.no.6,19-27.
Wel,L & Royakkers,L 2004, ‘Ethical Issues in
Web Data Mining’, Ethics and Information
Technology, vol.6.no.1,129-140.
Data Mining Technologies
Introduction
According to Mihai & Crisan (2010), in the
world today, almost every transaction that takes
place is recorded and kept in files for later
references. This has resulted to an increase in the
data produced and stored in various fields of
activity. In fact, every enterprise has
accumulated data on operations, activities and
performance.
All these data hold valuable information, e.g.,
trends and patterns, which can be utilized to
improve business decisions and optimize
success. Mihai & Crisan (2010), point out that
that traditional methods of data analysis that are
based mostly on humans dealing directly with
the data, basically do not scale to handle these
large amounts of data sets. For this reason,
various advanced technologies called data
mining techniques have been developed to
process the huge volumes of data.
According to Han & Kamber (2000), data
mining is the process of discovering
correlations, patterns, trends or relationships by
searching through a large amount of data that in
most circumstances is stored in repositories,
business databases and data warehouses. The
data mining process is employed by several
sectors to manage and deal with troubles that are
normally related to clients and which hider
efficient operations of various entities.
Important features of data mining tools
Data mining tools integrate many operations and
provide an easy-to-use way to perform the data
mining process. There are many different types
of data mining tools. The tools are diverse in
design and implementation. However, there are
several important features for data mining tools
that enable them to be accommodated by users
in the most efficient and effective manner.
The first important feature is the capability to
access various data sources. Usually, data is
obtained from a variety of sources in diverse
designs. A good tool should enable the user to
access different data sources without difficulties.
It should also posses the ability to run on huge
data sets. This is very important in
circumstances where an entity stores large
amount of data (Mihai & Crisan, 2010).
Han & Kamber (2000) argue that good data
mining tools should be user friendly. According
to them, this is significant since most of the
times the persons using the tools are not
specialists. The data-mining tools should
possess the ability to process data in a most
efficient manner since this is often considered
crucial in solving problem. Connolly, Begg &
Holowczak (2008), point out that data mining
tools should also enable good data and model
visualization.
This is will enable users to properly analyze the
data and make sound decisions. They further
state that the data mining tools should also
enable users to easily integrate different
techniques during the mining process. This is
because there is no one tool that can effectively
handle all prevailing dilemmas. Therefore,
superior data-mining tools should allow users to
incorporate a variety of procedures in order to
deal with various dilemmas.
Han & Kamber (2000), assert that since the rate
of innovation is very rapid, new techniques and
algorithms are always present in the market, it is
also important that data mining tools provide
good extensibility means. This design allows
users to monitor the tools in the most efficient
and effective approach. Good data mining tools
should also enable interoperability with other
tools as well as data exchange of model and
provide support for good data mining standards.
Data warehouse
A data warehouse is a site where information is
stored. The idea of data warehousing was
developed out of the need by different parties to
have easy access to structured store of quality
data that can be used for decision-making.
Generally, information is a very powerful asset
that can provide important benefits to any
enterprise and a competitive advantage in the
business world.
The massive amount of data possessed by firms
has made it difficult for the firms to access it and
make use of it. This is because it is in many
different formats, exists on many different
platforms, and resides in many different file and
database structures developed by different
vendors. Data warehousing offers a better
approach to managing these data (Connolly,
Begg & Holowczak, 2008).
How data mining can realise the value of data
warehouses
According to Connolly, Begg & Holowczak
(2008), data mining represents one of the most
important applications for data warehousing.
This is because most of the information that can
be used to analyze various problems is
accumulated in a data warehouse. The data
mining techniques enables users to extract all
relevant information needed for making good
decisions and which is hard to obtain in most
cases.
This kind of data enables realization of the value
of a data warehouse when used appropriately.
Han & Kamber (2000), point out that data
mining can also realize the value of data
warehouse by making use of advanced data
analysis techniques for strategic management to
interpret the information stored in the data
warehouses. It is crucial that decisions that are
taken by administrators are taken on informed
basis and not based exclusively on the talent and
knowledge of the administrator.
This application of data mining techniques
became possible by making predictions based on
the data that an enterprise has access to i.e. data
from its own databases (data warehouses). This
is very important since the data collected is
available in time for analysis when required.
According to Mihai & Crisan (2010), data
mining can also aid in designing data
warehouses for a specific application. In this
way, the value of the data warehouse can more
easily be realized because the amount of pre-
processing required before data is mined can be
determined according to the data available.
For instance, if the data is stored in relational
databases, it is easier to analyze it and most of
the data mining tools can be used without
difficulties. Data mining also transforms the
detailed level of operational data stored in the
data warehouse to a relational form that makes
the information to be more amenable to
analytical processing.
Conclusion
Data mining is an exceptionally valuable tool to
explore the essential data to create reasonable
advantage in the ever-changing environment.
Data mining is employed by several sectors to
manage and deal with troubles that are normally
related to clients and which hider efficient
operations of the various entities.
There are many different types of data mining
tools. The tools have different features and are
diverse in design and implementation. Data
warehouses are sources of new information and
are built to provide simple means to access
source of high quality data. Therefore, data
mining can easily realise value of data
warehouses by making use of the stored
information.
Large Volume Data Handling: An Efficient
Data Mining Solution Proposal
Introduction
The project proposal is specially designed to
highlight the problem of large volume data
handling and provides an efficient data mining
solution. This project proposal is specifically
designed keeping an eye on communication
service delivering problems and provides its
solution in a most approximate way. The
proposal starts with basic concepts of data
mining, related terms used in data mining,
company background and business problem, in
later sections this proposal highlights existing
problem with the system and later on proposed
solution and discussion. The proposal ends with
conclusion and references.
Data Mining
Data mining is a commonly used term in
computer field. Data mining is the process of
sorting huge amount of data and finding out the
relevant data. Usually ERP systems are used for
sorting data in large organizations. Data mining
is commonly known as knowledge discovery.
This is the process of analyzing data from
numerous perspectives and summarizing it into
useful information for further processing. There
are numerous companies following data mining
techniques in order to make their system
effective and time saving. Data mining is widely
used for the maintenance of data which helps a
lot to an organization in order to organize its
resources and capital in a proper way. In other
words, data mining is the process of finding
relationship between dozens of fields. Database
system gets affected if data mining techniques
are not properly applied in a certain domain.
Data mining is widely sued by large firms with
strong consumer focus. Data mining techniques
empowers organizations to identify the
relationship between entities and to create a
strong relationship among internal factors such
as cost, positioning, staff skills and also it gives
flexible path to create relationship among
external factors like economic indicators,
competition and customer demographics. Data
mining also helps in determining the impact on
sales caused by internal factors changes.
Terms
Data: Data is a raw material usually found
in bulk quantity. There are three types of
data operational or transactional data, non
operational data and Meta data.
Information: The pattern, relationship and
association among all this data could
provide information.
Knowledge: Information can be converted
into knowledge if related to previous facts
and figures and previous statistics of an
organization (Han & Kamber, 2000).
Data warehouses: Dramatic place for
storage of data, data warehouse gives
flexible opportunity to integrate new data
with a previous one. Data warehousing is
commonly defined as a process of proper
data management and retrieval. Data
warehousing gives a concept of storing all
data centrally in large organizations.
Knowledge discovery in databases (KDD):
KDD is a commonly used term in databases it’s
a non-trivial extraction of previously unknown
data and information from large databases.
Company Background
PCCW is the largest communication network in
Hong Kong. PCCW Limited (PCCW) is one of
the best communication companies in HKT
(PCCW, 2008). HKT Group Holdings Limited,
Hong Kong’s premier telecommunications are
the most renowned provider and a world-class
candidate in transferring information and
communication technologies. PCCW also holds
the great interest of foreign investors. The
PCCW posses a remarkable position in market
and it employs a total of 16,200 employees. Its
headquarter is located in Hong Kong and is
renowned in maintaining a presence in Europe,
the Middle East, Africa, the Americas, mainland
China and many other regions of Asia.
HKT has gained so much popularity in
telecommunication business HKT Group
Holdings Limited (HKT) was founded in 2008
with the aim of providing telecommunication
services, media and IT solutions. They are
renowned as the Hong Kong’s first quadruple-
play experience provider, PCCW/HKT
announces a wide range of media content and
services in following four domains fixed-line,
broadband Internet, TV and mobile. They have
gained a great success worldwide and posses
following credits: Hong Kong’s leading
telecoms player, genuine quadruple-play
experience, Expert in ICT solutions, Expanding
into international markets. They offer following
services:
Voice Services, Data Services, Internet Services,
Mobile Service, Equipment Solutions, ICT
Solutions, Contact Center Services, Telecom
and IPTV Solutions, Interconnect Services,
TSCM Services. PCCW is also famous in
outsourcing flagship.
Business process
Business process starts from setting up a
wireless telecommunication network using
different routers and switches. In order to
provide fault free network number of employees
and tools are used. The basic problem arises is
the management of bulk amount of data
efficiently. As, they provide telecommunication
and IT services, the main problem arises is of
data redundancy and data consistency. Both
these problems are the main hurdles in providing
valuable services. There is a huge amount of
customers data also there to be deal with.
Existing Problem
Problem question: How to deal with large
amount of customer data and services info in
order to provide speedy communication system?
PCCW is a large organization and its primary
responsibility is to provide best communication
network to all its customers. There is a high risk
of loosing customers if network gets fail due to
huge network traffic. They offer services in
telecommunication and also offer IPTV
solutions. IPTV solutions give complete
opportunity to integrate satellite systems. The
main problem occur in providing speedy
connections is data redundancy, time used in
searching a particular record and noise distortion
over large networks. It is really important that
the service provided to all customers should be
cost effective, speedy and based on fault free
network.
Goals and objectives
The main objective of using data mining
techniques is to reduce data redundancy over a
large network. It also helps company to better
utilize its resources and gives an opportunity to
allocate resources in an effective manner. Data
mining techniques provides strong facts and data
which help in decision making. They also
provide a path for better growth and allocation
of resources.
Proposed solution
There are numerous data mining techniques are
available that suits above scenario. There are
many techniques provides excellent data
handling over large networks if applied properly.
Distortion and clustering techniques is proposed
in order to solve above stated problem.
Distortion techniques are specifically designed
keeping an eye on the changing needs of explore
data, this technique also helps in data
exploration process by emphasizing on details
and preserves an overview of the complete data.
The main objective of distortion techniques is to
explore high level of detail with the combination
of lower level of data detail. For
multidimensional data sets a dynamic projection
method is widely used to change the overall
projections. In order to solve this problem
distortion technique will help a lot. PCCW team
need to implement a structure in which data
handling must be strong i.e. as they use ERP
solution for data handling so according to the
proposed solution they need to obtain the
relationship between fields and then define a
structure to link high level of data with lower
detail of data based on details or attributes of
data. When there is a link between both details,
so when any particular data is called the search
result would be according to the requirements
and lots of time will be saved. Browsing is very
difficult over large networks where a bulk
quantity of data is available. With the help of
interactive filtering and division of large data
into smaller groups along with the relationship
between fields this problem could be solve up to
high extent.
Clustering is a process widely used for
portioning data sets in meaningful classes for
further effective processing. Clustering is
commonly known as unspecified classification
of data without the combination of predefined
classes. Clustering is a techniques used for
division of large volume of data into small
identical groups. There is a numerous
perspective to classify clustering techniques in
data mining domain. Clustering plays a pivotal
role in data mining applications. Clustering has
become a significant problem in past few years’
databases, graphics, pattern recognition, neural
networks and computer graphics. Clustering
technique can solve the identified problem as
PCCW poses a wide network so if the data will
be divided in small groups, according to their
Meta data and will be stored in a central
respiratory system would be beneficial and save
time. ERP solutions provide well defined data
structure but still numerous companies are using
other software along with ERP as integration
problem is associated with an ERP solution.
PCCW needs a well defined integrated structure
of data for effective service. If clustering
technique will be applied so the data would be
stored in different groups, whenever a particular
data will be searched the crawler or pointer will
first check its Metadata and then enter in the
group. By this way data redundancy problem
can be solved as division of data based on Meta
data would not allow the same entry with same
Meta data. If in PCCW structure there would be
no data redundancy so automatically it will save
lots of time in finding a particular record.
Results and deliverables of this approach may
vary due to increasing amount of customers day
by day. The proposed solution is significant in
handling of large volume of data over large
databases.
Sample Process Model
Figure 1. in above figure Perl script is applied in
order to define the paths for data. PCCW is a
network where data travels from different
directions so it’s really necessary that data
follows the correct path so the network traffic
could be handled properly and data storage can
be made easy.
Deliverables
Details
Before
After
Time required
1-2 minutes
40-50 seconds
Project Load
Uneven & distracted
Organized & Balanced
Data Placement
Uneven & unorganized
Well utilized
Searching Time
1-2 minutes
40-50 seconds
Discussion
There are lots of advantages of using these both
techniques in PCCW environment as PCCW is a
very large network and posses bulk quantity of
data over large network. It’s harder to manage
the complexity of large data with the rate of
increasing customers. There is an issue of
mishandling of data also involves in such cases.
In order to solve this problem it’s really
necessary to detect the exact problem and then
proposed technique is applied in order to get
perfect results. Another alternate approach is
neural network can be applied in such
environment. Every algorithm and proposed
solution poses some advantages and
disadvantages. The selection of solution
depends on environment, requirements and level
of fitness in order to solve the problem.
Conclusion
PCCW is leading firm in Hong Kong offers
telecommunication services. It has a huge list of
customers and the rate of upcoming customers is
also very high. PCCW is a wide network and its
being ruling its position from last many years.
PCCW faces problems in handling large volume
of data. Proposed data mining techniques help a
lot in establishing a fault tolerant network and
also helps in proper allocation of resources and
staff. The proposal gives proper justification and
solution to the identified problem. Management
can make decision on the basis of fair and free
data obtained with the aid of proposed model.
Data Mining Tools and Data Mining Myths
Data Mining Tool
Data mining is a popular method of software
analysis used to find patterns in large datasets.
Today, various companies use this technology to
build marketing strategies, manage credit risk,
detect fraud, filter spam, or even define users’
sentiment. Today many specialists are
concerned with the security and privacy
problems that data mining evolves. There are
two significant problems highlighted: data
anonymization and validating external sources
(Bhuiyan et al., 2018). Both issues are noted
based on the different companies’ experiences.
The first problem is correlated with keeping the
identity of the person evolved in data mining
secret. The anonymization process contributes
to a more reliable systematic analysis because
the risk of data misuse is minimized. For
example, this technique can be applied in
hospitals and for insurance coding and billing
(Bhuiyan et al., 2018). Even though there are
various methods to prevent information fraud,
the problem is substantial due to the existence of
hackers who can implement de-anonymization
techniques. The second issue is the validation of
external sources, which the companies often
overlook. This process requires many budget
allocations and seems insignificant at the first
glimpse. However, it is vital the ensure high-
quality data protection. Therefore, the second
problem is also substantiated by the
irresponsibility of the stakeholders and
companies’ administrators.
Data Mining Myths
One of the major myths regarding data mining is
that it can replace domain knowledge. The
market is unexpected, and without the domain
expertise specialist, the company will have only
a few chances to advance. Data mining tools
provide the compression of data that a specialist
should interpret (Raiker, 2019). Another myth is
that only huge companies need to implement
data mining. This tool can be efficiently used by
any company disregarding of its status or size.
The small amount of money can be efficiently
used to analyze particular issues (Raiker, 2019).
It is always much more convenient to work with
the particular converted information rather than
with the whole database.
Ethnography and Data Mining in
Anthropology
Introduction
Ethnography refers to the study of specific
cultures. The study of cultures is of great
importance under normal circumstances to
enhance the understanding of the same. It is
against this background that the study of cultures
occupies a central role in society. Anthropology
involves the study of human behavior. This
involves the history and cultures of people.
However, anthropology involves a general
approach to the whole aspect of human culture,
behavior, and experiences. Ethnography on the
other hand focuses on the specific aspect of
culture. Normally ethnography involves the
selection and study of a specific culture.
This is considered vital since it brings more
accuracy and authenticity to the whole field of
anthropology. In this case, the knowledge that
could not have been achieved through
anthropology can be achieved by ethnography.
Therefore ethnography is considered as a better
way of understanding human behavior
specifically culture. Henceforth, ethnography
complements anthropology in many ways.
Ethnography works in many ways, the collection
of information and process of study involves
several methods and parameters. One common
method used by ethnography is data mining.
Data mining aids ethnography since it is through
it that the necessary information is obtained,
analyzed, and evaluated. This paper aims to take
a keen look at the concept of ethnography. To
succeed in this endeavor the paper will also
analyze data mining and its essence. The paper
will refer to several articles and sources in the
discussion of the whole concept.
Culture and Psychology
Ethnography plays a key role in the study of
psychological behaviors. These behaviors are
shaped by culture and human experience
(Atkinson & Hammersley 2007). It is against
this background that through data mining,
ethnography collects information and evaluates
bringing out the essence of the same.
Psychological behaviors are those features that
emanate from the status of the mind of students
at any given time. The psychological status of an
individual at any given time determines the
effect of whatever activity is prevailing.
Psychological behaviors might take the shape of
personality and other symptoms of mental
disorder. Under normal circumstances,
psychological behaviors involve several
conditions. Most of these conditions represent a
malfunction of the various mental faculties.
Examples of psychological behaviors include
anxiety, attitude, and motivation (Havemeyer
2007).
Various studies conducted have indicated that
anxiety is a major cause of poor performance
among students from certain cultural
backgrounds. It is a proven fact that anxiety
contributes to a high degree of low performance
by the students. However, this does not mean
that only anxiety causes poor performance since
virtually all psychological behaviors tend to
harm the performance of the students. Fears of
all kinds and aspects of mental and
psychological disorders all have a profound
effect on the aspect of performance of the
students. Yet as far as the learning process is
concerned, performance is the most important of
all the aspects (Fetterman 2009). The
performance of the students is an indicator of the
success or failure of the whole program. When a
student’s performance is affected it leads to a
kind of situation where the program is of no
significance.
Children who have psychological disorders lack
the necessary concentration and focus that is
necessary for learning (Keong 2006). Under
normal circumstances, the process of learning
requires a lot of focus and attention at the same
time. There is therefore a very clear relationship
between concentration and the learning process.
Without adequate concentration and attention,
the learning process is rendered ineffective. As a
result, the role played by factors of
psychological nature such as anxiety is great and
cannot be underestimated. Such factors don’t get
limited to the learning process alone (Larose et
al 2007). Their impact goes beyond the learning
environment. Under normal circumstances, the
students get problems in almost all the other
areas of life. For instance, the social lives of such
students are also affected by such factors of
psychological nature.
Conclusion
Ethnography plays an important role in the field
of anthropology. Ethnography complements the
process of anthropology. However, ethnography
is more effective, successful, and specific than
anthropology. This is because ethnography
focuses on a specific subject in its study. Under
normal circumstances, ethnography involves the
study of cultures. This is done in a way in which
a specific culture is selected for the study. In this
way, the process is more objective and
successful on many counts. Ethnography
involves several methodologies and aspects. For
the whole process to be successful several
parameters are needed. This is how the process
of ethnography becomes successful on many
counts. As a result, ethnography involves the
process of data mining. Data mining refers to the
process in which the data is sought for study.
Ethnography cannot operate without data
mining. Data mining is the success secret of
ethnography. Through data mining, ethnography
established the necessary information needed
for the analysis and understanding of human
culture. This works in a manner where the data
involving the selected culture is sought and
analyzed thoroughly to be used in making
conclusions. The paper has discussed fully the
concept of ethnography. Since there are several
related aspects, the paper has focused on several
parameters. This was done to navigate through
all the necessary factors of ethnography.
Data Mining and Customer Relationship
Management
Most marketing executives, along with
advertising practitioners, understand the
intrinsic value of collecting customer-related
information, but also comprehend the challenges
involved in leveraging this knowledge to
generate intelligent, proactive conduits that
could be used to add value to the customer as
well as the organization. This is where the
interplay between customer relationship
management and data mining comes in to
facilitate a process that assists organizations to
sift through stratums of ostensibly unrelated data
for significant relationships, where they can
proactively anticipate, rather than merely react
to, customer needs and expectations (Linoff &
Berry, 2008). Within the broad scope of data
mining, this paper purposes of evaluating some
underlying issues in customer relationship
management.
In data mining, the term ‘lift’ denotes a measure
of the performance of a particular model at
forecasting or grading cases and events through
the application of statistical modeling, random
choice model, or segmentation. Consequently,
the term is mostly used as a simple correlation
measure of the improvement in response
between two or more causal agents, not
mentioning that it also assists businesses and
other ventures to filter out misleading ‘strong’
associations of the form item A is dependent on
item B (Han & Kamber, 2006).
Parvatiyar & Sheth (2001) defines Customer
Relationship Management (CRM) as “…a
comprehensive strategy and process of
acquiring, retaining, and partnering with
selective customers to create superior value for
the company and the customer” (p. 5). As such,
CRM not only entails the integration of
marketing, sales, customer service, and supply
chain capabilities of the firm to attain elevated
efficiencies and effectiveness in conveying
customer value, but it obliges the organization to
know and comprehensively understand its
markets and customers in order to select the
most profitable customers as well as identify
those no longer worth targeting (Rygielski &
Wang, 2002).
The above explanation demonstrates the
convergence between CRM and data mining
the extraction of concealed extrapolative
information from large databases with a view to,
among other things, identify the most valuable
customers and predict future behaviors in order
to initiate proactive, knowledge-driven
decisions (Parvatiyar & Sheth, 2001). These are
precisely the most important benefits of CRM
ability to identify the most valuable customers,
ability to predict future behaviors, and; ability to
make proactive, knowledge-driven decisions
that enhance customer value as well as an
organizational value. Added to these benefits,
CRM is not only effective in managing
relationships between businesses and consumers
(B2C) but also in business-to-business (B2B)
environments (Rygielski & Wang, 2002).
In B2B environments, for example, it is a well-
known fact that modern business environments
are characterized by numerous transactions,
diverse custom contracts, and ever more
complicated pricing schemes. In such a scenario,
therefore, CRM can be used to facilitate the
processes when various agents of seller and
buyer organizations communicate and
collaborate. In equal measure, some CRM
initiatives such as customized catalogs, e-mail
alerts, personalized business portals, new
product information, and targeted product offers
can abridge the procurement process and
improve effectiveness and efficiencies for both
organizations, or assist in enhancing the
effectiveness of the sales pitch (Rygielski &
Wang, 2002).
The scalability of a CRM system simply means
that the system is capable of handling additional
volume and growth in the event of either
planned or spontaneous economic expansion
(Shelly et al., 2010). Predicting the future needs
and expectations of customers, along with future
shifts in behaviors, according to these authors, is
not an exact science and, as such, organizations
need to invest in scalable CRM systems that are
always informed by careful research and
planning informed by various methodologies,
including data mining.
Reference List
Han, J., & Kamber, M. (2006). Data mining:
Concepts and techniques. San Francisco, CA:
Morgan Kaufmann Publishers. Web.
Linoff, G.S., & Berry, M.J. (2011). Data mining
techniques: For marketing, sales, and customer
relationship management. Indianapolis, IN:
Wiley Publishing, Inc. Web.
Parvatiyar, A., & Sheth, J.N. (2001). Customer
relationship management: Emerging practice,
process, and discipline. Journal of Economic &
Social Research, 3(2), 1-34. Web.
Rygielski, C., Wang, J.C., & Yen, D.C. (2002).
Data mining techniques for customer
relationship management. Technology in
Society, 24(1), 483-502. Web.
Shelly, G.B., Cashman, T.J., & Rosenblatt, H.J.
(2010). System analysis and design, 8th Ed.
Boston MA: Cengage Learning. Web.
Students also viewed