Data Mining in Social Networks:
Linkedin.com
Background
Social networks have gained mass popularity in
recent years. While social networks help share
information and connect like-minded people,
they can also lead to misuse of or unintended use
of data that they share through data mining
attempts. One of the most commonly used
professional social networks is linkedin.com.
Linkedin.com was launched in 2003 and as of
October 2009, has 50 million users worldwide
with 11 million users in Europe alone (Weiner,
2009). India is the fastest-growing country with
3 million users and rising day by day (Weiner,
2009). As Jeff Weiner (2009) says, LinkedIn’s
mission is “to connect the world’s professionals
to make them more productive and successful”.
How far LinkedIn has been successful in this
attempt is interesting research by itself but more
importantly, I find it more interesting to explore
whether data mining is assisting LinkedIn in
achieving its goal of creating further
roadblocks? Getting a user to join LinkedIn is
one thing, but being able to retain that user is
another thing. So have users found LinkedIn as
useful as it claims and have they continued to
use LinkedIn or have they moved on? Survival
data mining techniques can help answer this
question. The survival data mining technique is
an interesting technique that provides “rapid
feedback about the customers and their
behaviors, while at the same time providing a
solid basis for quantifying customer value and
measuring customer loyalty” (Linoff, 2004).
Aims
The project aims to understand LinkedIn users,
improve user retention and a user lifetime to
improve profit margins using survival data
mining techniques:
1. to understand value addition such data
mining can result into.
2. to understand users, their behaviors, and
preferences.
3. to understand what strategies work best for
LinkedIn
4. to understand how LinkedIn data can be
misused, used against the user, or made
unintentional use of without permission
and how that can undermine the LinkedIn
goals.
5. to understand what steps can be taken and
what needs to be taken to reduce the
chances of data being misused thereby
improving chances of user retention.
Methodology
One of the ways to achieve the aim is to
understand how users view data mining of their
data on LinkedIn. This can be achieved by
interviews or surveys of LinkedIn users whose
expectations and concerns regarding the usage
of their data can be noted down and their
experience regarding the same can be found out.
The second methodology to be employed is to
harvest data directly from LinkedIn. This will
serve as a proof of concept as to how easy or
difficult it is to conduct data mining from
LinkedIn to automatically gather large amounts
of data and draw conclusions on usage trends.
Since there is a multitude of questions to be
answered through this research, a variety of
techniques will be used including cluster
analysis, classification and prediction, and
statistical analysis. This data will also be used to
perform survival data mining to understand how
users can be retained.
*Research Questions
The research seeks to find the answers to the
following questions:
1. How can survival data mining provide
insights into users?
2. What are the different types of LinkedIn
users?
3. How can survival data mining be used to
improve user retention?
4. How can data mining be used on LinkedIn
to provide more value-added features to its
users?
5. How can data mining reveal usage trends
useful for conducting market research?
6. Can data mining be used on LinkedIn to
help identify crime associations, terrorist
activities as can be seen as possibilities
with another social network?
7. How can data mining lead to misuse of
information provided or unintended use of
information that takes place without the
user’s permission?
8. What steps can a LinkedIn user take to
protect his privacy? What are the tradeoffs
in that case?
*Review of the literature
LinkedIn as a business-oriented social
networking site has caught the attention of many
researchers (e.g. O’Murchu et al, 2004;
Churchill et al, 2005); however, it is used more
as an example or for comparison with other
social networking sites. There has been little
research on the usefulness, privacy, and other
data mining aspects although there are blogs and
news articles on LinkedIn that are interesting. In
a related field, there has been similar research
conducted on social networking sites aimed at
having more friendly than professional contacts
including FaceBook and MySpace (e.g. Jones et
al, 2005). Data mining on social networks, in
general, has also been researched well enough
with security and privacy on social networks
receiving the most attention. For example,
Clifton et al (1996) discuss the security and
privacy implications on social networking and
provide several methods that could be used to
prevent data mining such as fuzzing the data,
eliminating unnecessary groupings, audits,
augmenting the data, etc. Seifert (2007) and Pant
et al (2009) discuss data mining on social
networks as a means to detect fraud and terrorist
activities.
Survival data mining approaches have been
researched in the medical fields and other areas
but not applied to social networks like LinkedIn.
Linoff (2004) explains how survival data mining
can be applied to a subscription-based business
model.
Expected Outcomes, Significance, or
Rationale
LinkedIn is growing as a professional
networking site and the data stored there could
have lots of potential for positive research both
for the benefit of the user as well as for the
business. However, understanding how it can be
misused is more important since it could cause
harm to the professional image of a user of this
site. The research will look at ways that data
mining could benefit the users of this site and the
business itself positively. The research will also
look at how data mining could prove detrimental
Phase
Start
Finish
Literature Review
Ongoing
15thJuly 2010
Chapter Writing – Draft
1stJuly 2010
16thSept 2009
Data Collection – Survey
16thSept 2010
10thJan 2011
Chapter Writing – Draft
11thJan 2011
15thFeb 2011
Data Collection – Data Harvesting
20thFeb 2011
15thDec 2011
Chapter Writing – Draft
16thNov 2011
15thJan 2012
Survival Data Mining Analysis
15thJan 2012
15thJun 2012
Chapter Writing – Draft
15thJun 2012
15thSept 2012
Reporting
16thSept 2012
1stNov 2012
Chapter Writing – Draft
1stNov 2012
1stJan 2013
to the users and what they can do to prevent such
incidents from happening.
*Timetable
Finalization and Editing
1stJan 2013
1stMar 2013
Thesis Submission Date
1stMar 2013
List of References
ATA, N., ÖZKÖK, E. & KARABEY, U. (2005)
“SURVIVAL DATA MINING: AN
APPLICATION TO CREDIT CARD
HOLDERS”. Journal of Engineering and
Natural Sciences, 26(1), 33-42.
Bishop, K., Draskovich, J., Hottenroth, A., Lee,
B. & Pesavento, S. (2005) Business Uses of Data
Mining and Data Warehousing. Web.
Breiger, R. L., Carley, K. M. & Pattison, P.
(2003) “Dynamic Social Network Modeling and
Analysis: workshop summary and
papers”. National Academics Press.
Churchill, E. F. & Halverson, C. A. (2005)
“Social Networks and Social
Networking”. IEEE Computer Society, 2005,
14-19
Clifton, C. & Marks, D. (1996) “Security and
Privacy Implications of Data
Mining”. Proceedings of the 1996 SIGMOD
Workshop on Data Mining and Knowledge
Discovery.
Collica, R. (2004) “Data Mining Galore: For
Business Applications”. Web.
Han, J. & Kamber, M. (2006) “Data mining:
concepts and techniques “. Elsevier Inc..
Han, J. & Kamber, M. (2006) “Data mining:
concepts and techniques “. Morgan Kaufmann.
Jensen, D. & Neville, J. (n.d.) Data Mining in
Social Networks. Web.
Jones, H. & Soltren, J. H. (2005) Facebook:
Threats to Privacy. Web.
Kleinberg, J. M. (2007) “Challenges in mining
social network data: processes, privacy, and
paradoxes”. Proceedings of the 13th ACM
SIGKDD international conference on
Knowledge discovery and data mining.
KUSIAK, A., DIXON, B. & SHAH, S. (2006)
“Predicting survival time for kidney dialysis
patients: a data mining approach”. Computers in
Biology and Medicine, 35(4), 311-327.
Linoff, G. S. (2004) “Survival Data Mining”.
Web.
Linoff, G. S. (2004) Survival Data Mining for
Customer Insight. Web.
Mohammed, Z & Kotze, D (2005) “Survival
data mining in the telecommunications
industries: usefulness and complications “. Data
Mining XI: Data Mining, Text Mining and Their
Business Applications., 505-512
Olson, D. L. & Delen, D. (2008) “Advanced
Data Mining Techniques”. Springer.
O’Murchu, I., Breslin, J. G., and Decker, S.
(2004): “Online Social and Business
Networking Communities”. Technical Report.
Pant, D. & Sharma, M. K. (2009) “Web Mining
and Social Network Analysis in Cyber war, to
warn about terrorist attacks”. Web.
Potts, W. (2006) Survival Data Mining. Web.
Seifert, J. W. (2009) Data Mining and Homeland
Security: An Overview . Web.
Wang, J. (2009) “Encyclopedia of Data
Warehousing and Mining “. Information Science
Reference.
Weiner, J. (200) LinkedIn: 50 million
professionals worldwide. Web.
Data Mining Classifiers: The Advantages and
Disadvantages
Decision Trees: C4.5 Classifier
Advantages
Classifiers produced by C4.5 are either
expressed as decision trees or rulesets. Our focus
in this discussion is on decision trees, and we
look at their advantages. The C4.5 classifier is
based on the ID32 algorithm, whose primary
aim is to find small decision trees. Based on this,
we can say that the decision trees produced are
small and simple to understand (Witten, Frank,
2000).
Another advantage of this classifier is that the
source code is readily available. Thirdly,
unavailable attribute values are accounted for in
C4.5 by assessing the gain using the records
where the particular attribute is defined. A fourth
advantage is that; it can handle both continuous
and discrete attributes. Attributes with a
continuous range, we create partitions based on
a predetermined pattern in the training set and
calculate the gain on each partition, (Quinlan,
1993). Then the partition that maximizes the
gain is picked. Lastly, tree pruning after creation
ensures tree simplicity by replacing some
branches with leaf nodes.
Disadvantages
C4.5 classifiers are basically slower in terms of
processing speed. For instance, a task that will
take C4.5 15hours to complete; C5.0 will take
only 2.5 minutes. Secondly, it is inefficient in
memory usage meaning that some tasks will not
complete on 32-bit systems (Witten, Frank,
2000). In terms of accuracy, the rule sets
produced by C4.5 classifiers basically contain
many errors. C4.5 has limited data types and
lacks the facility to label data as not applicable,
(Quinlan, 1993). The decision trees produced by
C4.5 are relatively bigger but these are catered
for in the C5.0 version. All classification errors
are treated equally but some are more serious
than others and this is another disadvantage.
Lack of a provision to quantify the importance
of cases is a disadvantage because not all cases
are equally important. Lastly, attributes need to
be winnowed before a classifier is generated and
C4.5 lacks this facility.
KStar Algorithm classifiers
Advantages
Firstly, In terms of accuracy, this algorithm is
comparable to C4.5 on voluminous UCI
benchmark datasets. KStar performs better and
with a higher speed than C4.5 on big numbers of
text classifications. Thirdly, this algorithm is
low in time complexity that means it is very fast.
Its speed can be compared to that of naive
Bayes. In addition, this algorithm can be
speeded up by combining it with other scaling-
up methods (Cormen et al 1990). Another
advantage is that it uses entropy as a measure of
distance hence providing a consistent approach
to management of symbolic attributes, unknown
values and continuous value attributes. The
results presented compare satisfactorily with
many machine learning algorithms. This
algorithm is an instance-based classifier. It
classifies an instance by comparison; an instance
is compared to pre-classified examples stored in
a database. New instances to be added to the
instance database and choice of instances from
the database to be used in classification are
determined by a concept description updater
(Gray, 1990). This helps reduce memory space
requirements and also improve tolerance to
noise in data.
Disadvantages
One of the major disadvantages of this algorithm
is the fact that it has to generate distance
measures for all the recorded attributes (Cormen
et al 1990). Another disadvantage is that it is not
cost effective. Both building and learning
processes are quite expensive. In terms of
memory space, as much as the algorithm uses an
instance updater to select and determine the
instances to be used, the storage of sample
instances is also involving in terms of memory
space (Gray, 1990). Although the algorithm can
be speeded up by combining it with other
algorithms, it is sometimes a disadvantage since
the combination process involves extra costs.
Bayesian Network classifier
Advantages
The first advantage of this classifier is its
computational efficiency. The representation of
large and complex computational problems is
decomposed into smaller and simpler self
sufficient models to enhance efficiency.
Secondly, it simplifies the incorporation of
domain knowledge into the model design by
using the structure of problem domain. Thirdly,
the natural combination of EM algorithm and the
probabilistic representation helps address the
problems with missing data. The ability of this
classifier to indicate all the possible classes a
new sample may belong and the probability, is
another advantage since in case of a
misclassification, it is easy to tell which other
class the sample may belong, (Mitchell,
1997).Another advantage is the presence of an
updater that enables the classifier to learn from
new data samples. Bayesian networks have the
advantage of being able to capture the
complexity of decision making. Lastly, the
system’s rules of classification can be
semantically determined and justified to both the
novice and the expert (Jensen, Graven 2007).
Disadvantages
The mismatch between the data likelihood and
the actual label prediction accuracy tends to
make the learning method suboptimal. Secondly,
Bayesian networks require an expert to give
domain information for the creation of the
network (Jensen, Graven 2007). Despite the
substantial amount of research carried out, the
creation of these networks is still limited to data
sets consisting of only a few variables which are
very informative. Another disadvantage of these
networks is the fact that their interpretation and
efficiency is quite limited when rulesets are
drawn from the network. In comparison to rules
derived from decision trees whose interpretation
is simple and direct, Bayesian networks are
more complex (Mitchell, 1997).
Advantages
The first advantage of this classifier is the use of
propositional rule learner which ensures that
errors are minimal. Also, the phase by phase
implementation of the algorithm ensures that the
overall results are near perfect. During the rule
set growing phase, the condition with the highest
information gain is picked and this ensures that
the rule set is perfect (Arthur, 1996). Another
advantage is that the algorithm is made shorter
and simpler by pruning any useless parts of a
rule. This makes the classifier easy to
understand and interpret. The ease of generation
of this classifier is a fourth advantage. This
classifier is highly expressive; it is comparable
to a decision tree. Even in terms of performance,
it can be ranked the same level as a decision tree
[Bishop, 1997]. JRip classifier has the ability to
classify new instances rapidly. Lastly, it is easy
for this classifier to handle missing values as
well as numeric attributes.
Disadvantages
JRip algorithm requires a large investment in
terms of time to learn the algorithm and test the
features that can be customized. The fact that the
RIPPER algorithm uses some induced rules,
which sometimes have to be replaced by expert-
derived rules for some applications is another
disadvantage. In addition, the accuracy of JRip’s
results sometimes vary accordingly. This is
because; the results produced differ depending
on the option of the rule voting method used
(Authur, 1996). Another disadvantage is that,
highly accurate results can only be achieved by
running the algorithm many times. Assignment
of salience which is an order of that has been
prescribed for firing rules may lend the expert
system’s inference engine powerless. This may
also have negative effects on the performance of
such a system which is rule-based (Bishop,
1997).
K-Nearest Neighbour Algorithm
Advantages
This algorithm uses local information. This kind
of information can produce highly adaptive
behaviour. Secondly this algorithm is robust to
noisy training data and it is simple to implement.
It is also very easy to use in parallel
implementations. Another advantage of K-
Nearest neighbour algorithm is the ease and
simplicity of learning. In addition the training
speed is very high and the results are nearly
optimal in the case of a large sample limit. This
means that the algorithm is more effective if the
set of training data is large [Witten, 2005] The
ability of this algorithm to approximate complex
concepts of the target locally and also differently
for every new instance is another advantage of
this algorithm. Lastly, the algorithm is intuitive
and easy to understand. This makes
implementation and modification easy. It also
avails a generalisation accuracy that is
favourable on many domains (Duda, Hart,
2000).
Disadvantages
One disadvantage of this algorithm is that it has
large memory space requirements. It needs to
store all the data and hence the need for large
memory space. Secondly, during instance
classification, all training occurrences have to be
visited and this makes the algorithm slow during
this procedure. A third disadvantage is that the
algorithm is easily fooled by irrelevant data
attributes; this means that the accuracy
decreases with increase in irrelevant data
attributes. Also, the fact that the accuracy
decreases with increased noise in the set of
training data is another disadvantage. Another
shortcoming of this algorithm is its
computational complexity [Witten, 2005]. This
is due to its intensive computational recall
(Duda, Hart, 2000). This algorithm is a
supervised learning and this means it runs
slowly. The algorithm is highly vulnerable to the
dimensionality curse and lastly, it is highly
biased by the value of k.
Naive Bayes classifier
Advantages
Naive Bayes classifiers have very simplified
assumptions and naive design and hence easy to
build. The structure is basically the same and
this eliminates the structure learning procedure.
Model building is highly scalable and it is
possible to parallelize scoring no matter the
algorithm. Secondly, these classifiers have
worked favourably in numerous real world
scenarios. It has outperformed many complex
classifiers in on large numbers of datasets,
despite its simplicity (Box, Tiao 1992). Thirdly,
this classifier requires a small amount of training
data to approximate parameters required for
classification which include averages and
variances of these variables. A fourth advantage
is that, it is only the class variances that need be
resolved and not the variance of the entire
dataset (Witten, Frank 2003). Both binary and
multiclass classification problems can be solved
by this algorithm. This algorithm relies on basic
probability rules making it simple in operation;
also being probabilistic the results are presented
in a form favourable for incident management
policy. Lastly, it facilitates a broader set of
model parameters to be used.
Disadvantages
The assumption that every variable is
independent of others is sometimes a problem.
The class estimates are sometimes absurd and
the threshold must be harmonized not set
analytically. This classifier has been
outperformed by newer approaches such as
boosted trees. It lacks the ability to solve more
complex problems in classification. Naive
Bayes algorithm is used in Bayesian spam
filtering and it is vulnerable to Bayesian
Poisoning (Box, Tiao 1992). Also the spam filter
is beaten by replacing text with pictures.
Another disadvantage is that if parameter
estimates are improved, the effectiveness of
such a classification will be affected. When
using the Bayesian filter, a user specific database
must be consulted on every message (Witten,
Frank, 2003). This database contains word
probabilities that are used to detect spam. Lastly,
the initialisation of the naive Bayes based filter
is a bit time consuming.
Works Cited
Bishop, Christopher. Pattern recognition and
Machine Learning. New York, NY: Springer,
Box, G., and Tiao, G. Bayesian Inference in
Statistical Analysis. New York. John Wiley &
Sons, 1992.
Duda, Richard; Hart Peter; Pattern
classification, 2nd ed. David Stork. 2000.
Frank, Eibe; Holmes, Geoffrey; and Witten,
Ian. Naive Bayes for regression. Machine
Learning, 2003.
Jensen, Finn. An Introduction to Bayesian
Networks. New York, NY: Springer-Verlag,
1996.
Jensen Finn; Thomas Graven. Bayesian
Networks and Decision Graphs. New York, NY:
Springer, 2007.
Gray, Robert. Entropy and Information
Theory. New York, NY: Springer-Verlag, 1990.
Cormen, Leiserson, and Rivest,
Ronald. Introduction to Algorithms. Cambridge.
MIT Press, 1990.
Quinlan, J. R. C4.5: Programs for Machine
Learning. Morgan Kaufmann Publishers, 1993.
Riel, Arthur. Object oriented Design
Heuristics. Addison Wesley. 1996.
Witten, Ian. Data Mining: Practical machine
learning tools and techniques. Morgan
Kaufmann, San Francisco, 2005.
Witten, Ian; Frank Eibe. Data Mining: Practical
machine learning tools and techniques with java
implementations. Morgan Kaufmann, San
Francisco, 2000.
Summary of C4.5 Algorithm: Data Mining
C4.5 algorithm is a decision tree with unlimited
number of paths within the node. This algorithm
can work only with discrete dependent attribute,
that is why it can solve only classification tasks.
C4.5 algorithm is considered to be one of the
most famous and widely used algorithms of
generating decision trees. It is necessary to
follow the next demands for working with C4.5
algorism:
1. Each record from set of data should be
associated with one of the offered classes,
it means that one of the attributes of the
class should be considered as a class mark.
It may be concluded that all the samples
should belong to the same class, otherwise
the mistakes are inevitable.
2. Each class should be discrete. Each sample
should belong to one of the classes.
3. The number of classes should be much
fewer from the number of samples in the
considered scope of data.
One should understand that C4.5 algorithm
works slowly with very large scale set of data.
Using the concept of information entropy, C4.5
builds the decision trees based on the set of data,
like ID3 algorithm. Filestem.ext is the form for
the files which are read and written within C4.5
algorithm (filestem is a file name, and ext is a
file extension which is aimed at defining the file
type). Working with the program, one is
expected to have at least two files, the first one
is with the file name and class definition and the
second one is with the date which gathers the set
of objects described by the value of the class
attributes. Considering the structure of a
decision tree based on C4.5 algorithm, it may be
either a leaf, which is predicted to identify a
class or a decision node with a number of
branches and sub trees, which show the possible
outcome of the trial (Quinlan 5).
There are two ways how this algorithm can
generate decision trees, batch mode and iterative
model. Batch mode (often called default mode)
generates a single decision tree. This tree covers
all the data available for the decision. Another
kind of this algorithm, iterative mode, is based
on the random basis. The set of data is selected
randomly. Then, a decision tree is generated
with adding some specific objects which have
been misclassified.
The actions are repeated and the decision tree is
continued until it is classified in a correct way or
it is found out that there is no any progress.
Keeping in mind that iterative model is based on
the subset selected randomly, many trials may be
used for generating decision trees based on the
same data. Keeping in mind that there can be
many different decision trees due to the multiple
trials, the presence of the filestem.unpruned is
necessary. This file is created with the purpose
to collect the decision trees in the process. If the
similar data is used for generating decision trees,
the latest variant of the tree is used. The machine
saves the best generated decision tree in the file
filestem.tree.
“Data Mining and Customer Relationship
Marketing in the Banking Industry“ by Chye
& Gerry
Summary of the Article
The article under consideration, written by Chye
and Gerry (2002), focuses on data mining as a
part of newly emerging tools for the business
assessment processes. The authors state that data
warehousing and knowledge management are
other components of new information systems
(Chye & Gerry, 2002). According to the article,
data mining aims to identify “valid, novel,
potentially useful, and understandable
correlations and patterns in data” (Chye &
Gerry, 2002, p. 1). It could be stated with
certainty that the advancement of computer
technologies, both in terms of hardware and data
mining software improvements, significantly
facilitates the accessibility and affordability of
data mining for different business companies
(Chye & Gerry, 2002). The article is subdivided
into five parts, which explore various aspects of
data mining in the context of customer
relationship management. This paper aims to
summarize each of the article’s sections
subsequently.
Customer Relationship Management
First of all, the article generally elaborates on the
notion of customer relationship management
(CRM), which is defined as “the process of
predicting customer behavior and selecting
actions to influence that behavior to benefit the
company” (Chye & Gerry, 2002, p. 2). Further,
the article states that there are several main
objectives of CRM. The first objective is to get
closer to the customer by utilizing the data from
the enterprise databases (Chye & Gerry, 2002).
Thus, a company can predict the customer’s
behavior in advance (Chye & Gerry, 2002).
Secondly, it is essential for companies to
become more customer-oriented. This goal is
reached by a greater focus on customer
profitability rather than line profitability (Chye
& Gerry, 2002). Other objectives include better
customer response and loyalty, more efficient
lead management, and increased cross-selling
possibilities.
Data Mining Methodology
According to the authors, data mining is a
relatively recent practice in the business sphere
since it emerged in 1994 (Chye & Gerry, 2002).
There are five primary stages of data mining:
sampling, exploring, modifying, modeling, and
assessing the gathered data (Chye & Gerry,
2002). Sampling is required when data is too
voluminous or it is needed to avoid problems of
generalization (Chye & Gerry, 2002).
Exploration and modification stages refer to the
processes of understanding the information and
developing meaningful insights. In the modeling
stage, the actual analysis is performed through a
set of different approaches, such as traditional
statistical methods, neural networks, decision
trees, etc. (Chye & Gerry, 2002). It should be
stated that data mining tools are considerably
variable and numerous, and thus they are
traditionally categorized into three groups
according to their purpose: “description and
visualisation, association and clustering,
classification and estimation” (Chye & Gerry,
2002, p. 4). Overall, the implementation of data
mining gives the company a competitive
advantage.
Banking Applications
There is a considerable amount of academic
literature and research dedicated to the use of
data mining in business. For example, the
authors state that, according to the professional
and trading literature, numerous companies use
data mining to be more competitive in their
industries (Chye & Gerry, 2002). Considering
the application of data mining to banking, it is
possible to mention several examples. Firstly,
data mining could be used by banks as a part of
their risk management, namely operating the
credit risks (Chye & Gerry, 2002). Another
sphere where this approach is used is customer
acquisition (Chye & Gerry, 2002). The
employment of enterprise databases makes it
possible to predict the customer’s response to
the bank’s marketing campaigns.
Application to Churn Modelling
One of the most evident examples of using the
data mining methodology is its application to
churn modeling. The authors propose a fictional
situation where it is needed to consider a
customer relation application (or churn
modeling) for a fictitious banking company,
ZBANK (Chye & Gerry, 2002). ZBANK faces
increasing competition on the market due to
customer defections (Chye & Gerry, 2002).
Through the process of implementing various
data mining tools, it is identified that the
decision tree model is capable of predicting the
number of customers who are inclined to the
voluntary churn (Chye & Gerry, 2002). Overall,
it should be stated that data mining, with the
employment of demographic and transactional
information, can accurately identify churners
and non-churners.
Data Mining Limitations
Considering the limitations of this approach, it
is possible to mention three primary aspects.
Firstly, the authors state that some products of
random fluctuations will certainly emerge
during efficient exhaustive mining of data (Chye
& Gerry, 2002). This limitation is particularly
true in the context of big data set with different
variables. Secondly, the data mining method is
well-designed for modeling, but it does not show
the same efficiency with the assessment of the
results (Chye & Gerry, 2002). Finally, the
application of data mining tools requires both
profound pieces of knowledge of the sphere to
which the method is applied and proficiency in
data mining (Chye & Gerry, 2002).
Critiques for the Paper
It should be noted that the article is of
considerably high academic quality. It provides
a profound overview of data mining
methodology in the context of customer
relationship management. The article is highly
applicable to various spheres of business, and
thus it would be useful for a wide range of
businessmen and managers. One of the primary
strengths of the article is that the authors provide
a vast amount of information about data mining,
supporting their claims with the evidence from
academic literature and real-life examples of
banks, which successfully employ the
methodology. Among the weaknesses of the
article, one can identify that it is primarily
focused on banking applications, while other
industries are not given enough insights on the
implementation of data mining.
Research
Further, regarding possible recommendations
for the development of more valuable research,
one can state that the article’s primary weakness,
which is mentioned in the previous section,
should be the main reference point for new
research. Particularly, it means that the authors
should focus more on the applications for other
business spheres. It would bring additional value
to the paper. Also, since the authors are primarily
focused on the banking industry, it would be
considered beneficial to add an application of
data mining methodology to a real-life banking
situation rather than a fictional one.
Further Additions to the Paper
Considering future work that I may add to the
paper, I should state that it will be a significant
improvement for the paper if a more profound
literature review is given. As it was already
mentioned, the authors employed a considerably
vast number of sources; however, they did not
provide specific information on the topics of
customer relationship management and risk
management. In my opinion, these two aspects
of managerial work are immensely important for
any company, especially in the banking industry
since various risks are involved.
Other Academic Sources on the Topic
Also, it is essential to mention two academic
sources, which explore the topics related to the
scope of the research under consideration. The
first source is an article, written by Khodakarami
and Chan (2014). It is chosen because it explores
the topic of customer relationship management
more profoundly, and it also gives profound
insights on the employment of data mining tools
in CRM. The second article by Chen, Deng,
Wan, Zhang, Vasilakos and Rong (2015)
investigates the peculiarities of the
implementation of data mining methodology to
the Internet of Things, which is a considerably
perspective scope of the research.
Data Mining Techniques and Applications
Select two application areas for data mining
NOT discussed in the textbook and briefly
discuss how data mining is being used to solve
a problem (or to explore an opportunity)?
Data mining involves rearranging large volumes
of data to create comprehensible information
that can be used to solve problems. There are
several ways in which data mining can be
applied in the real world (Han et al. 76). It can
be used to solve problems and explore
opportunities.
Data Mining and the Detection of
Disturbances in the Ecosystem
The use of data mining to detect disturbances in
the ecosystem can help to avert problems that
are destructive to the environment and to
society. Such calamities include floods and
droughts (Kumar and Bhardwaj 258). Remote
sensing and earth science techniques are used to
understand the radical changes in the
environment. Data is collected and archived. It
is later mined and used to detect disturbances.
Data Mining in Sports
Data mining can be used to predict sporting
activities. A case in point is the Advanced Scout
System developed by IBM (Leung and Kyle
715). The application is used by coaches to
improve the performance of players. In most
cases, fans predict games by watching. They
may also use archived data, which is mined and
statistically used to make predictions based on
the history of the game.
What is Association Rule Mining? And
explain how Market-basket analysis helps
retail business to maximize the profit from
business transactions?
Association Rule Mining
It is the retrieval of data based on the
relationship between a given set of objects. It
takes into consideration the ‘togetherness’ of
these objects and how they appear in a database.
It involves the identification of connections and
correlations between objects (Ramageri 304).
Market-Based Basket Analysis and Retail
Business
Market basket analysis and association rule
mining can be used to maximize profits and
improve transactions in the retail business. It is
used to study the behavior of customers and their
shopping trends. Marketers use the information
to design catalogs and undertake customer
behavior analysis (Han et al. 99). Consequently,
the information can be used in marketing and
advertisement to maximize profits and improve
business transactions.
Discuss k-Nearest Neighbor (KNN) learning
algorithm. What is the significance of the
value of k in k-NN?
K-Nearest Neighbor (KNN) Learning
Algorithm
The algorithm is a method that is used to classify
data obtained from sources with similar sets of
parameters. It uses a set of data based on the
known classifications of the existing database. It
makes use of separate classes to predict a new
pattern and classify the new data. The
‘neighbors’ in this case are the separate sets of
data with common characteristics (Bhatia and
Vandana 304). For instance, a bank may get a
customer who wants a loan, but the entity lacks
time to calculate the credit rating of the
applicant. The bank can use previous credit
ratings of people with similar characteristics,
such as earnings and collaterals.
The Significance of the Value of k in k-NN
The k represents the number of classes used in
the comparison. Lower values of this component
are more accurate compared to higher values.
On the other hand, increasing the random data
point raises the percentage error of
approximation (Bhatia and Vandana 304). As
such, k can be used to obtain the most accurate
approximation in data classification and
regression.
Discuss the two estimation methods of
classification-type data mining models while
considering ANN as a classifier
Supervised Learning
It is one of the estimation methods of
classification data mining models in artificial
neural networks (ANN). In this case, a set of
example pairs is provided. The objective is to
identify or ‘estimate’ a function. The function
has to lie within the permitted cluster of
functions (Nikam 15). In addition, it has to
reflect the given examples.
Unsupervised Learning
In this estimation method, the ANN works with
a given set of data. The data is usually denoted
as x. The cost function to be minimized is also
provided. The latter can be a random function
of x. It can also be the output of the network. The
output is usually denoted as f. The cost function
relies on what the network is trying to model
(Nikam 16). It is also affected by the
assumptions made.
Works Cited
Bhatia, Nitin, and Ashev Vandana. “Survey of
Nearest Neighbor Techniques.” International
Journal of Computer Science and Information
Security, vol. 8, no. 2, 2010, pp. 302-305.
Han, Jiawei, et al. Data Mining: Concepts and
Techniques. 3rd ed., Morgan Kaufmann
Publishers, 2011.
Kumar, Dharminder, and Deepak Bhardwaj.
“Rise of Data Mining: Current and Future
Application Areas.” International Journal of
Computer Science Issues, vol. 8, no. 5, 2011, pp.
256-260.
Leung, Carson, and Joseph Kyle. “Sports Data
Mining: Predicting Results for the College
Football Games.” Procedia Computer
Science, vol. 35, 2014, pp. 710-719.
Nikam, Sagar. “A Comparative Study of
Classification Techniques in Data Mining
Algorithms.” Oriental Journal of Computer
Science & Technology, vol. 8, no. 1, 2015, pp.
13-19.
Ramageri, Bharati. “Data Mining Techniques
and Applications.” Indian Journal of Computer
Science and Engineering, vol. 1, no. 4, 2011, pp.
301-305.
E-Commerce: Mining Data for Better
Business Intelligence
Abstract
Business enterprises operate in competitive
environments, which compel them to use data
mining methodologies to convert unstructured
data into information for business intelligence.
The data consisting of millions of transactions is
used to fragment existing and target markets
according to consumer behavior.
Here, the business intelligence information
gained from data mining is also used to detect
fraud, understand consumer behavior, gain
deeper insight into business operations, and to
benchmark business operations based on data
classification, prediction, and association rules.
Introduction
Many organisations struggling to remain
competitive in the market use data mining to
gain business intelligence information. Business
intelligence is important for forecast demand
management, supply chain management, cost
management, category management, and to
exploit existing business intelligence
opportunities to improve the core business
processes (Giudici 69).
For instance, Intel uses data mining to
strategically access business intelligence
information, which yields accurate information
for effective decision making to enhance value
creation opportunities at different strategic
levels of the organization (Giudici 89).
Literature review
Many organizations have established that data
mining is a new powerful tool that has been
integrated into the business practices of different
firms for analyzing large amounts of
unstructured data to discover new and emerging
trends in the fiercely competitive markets. Here,
business organizations use automatic methods to
analyze data from large databases to establish
previously unknown patterns of data.
The amounts of business data, which is
generated from millions of transactions, cannot
be processed promptly using traditional
mechanisms. However, with the advent of new
technologies, which have high processing
powers, large amounts of unstructured raw data
can be mined and processed in a short time form
large data warehouses (Delmater and Hancock
2).
Intel is one of the examples of firms, which mine
data from different data warehouses to evaluate
the company’s performance by benchmarking
her business processes against successful
business practices. The tangible benefits that
accrue from lintel’s business processes include
the elimination of unnecessary costs using
analytical reports to determine the performance
and trend of the business.
The data mining process consists of data
classification, prediction, and association rules.
To ensure that the data, which is mined for
business intelligence is used appropriately to
meet Intel’s business goals and objectives, Intel
first classifies unknown data according to
established classification rules. The process
includes identifying 90% of the portions of data
of the organisation, which has not been
classified for current and future use by ensuring
that the classification process factors the goals
or the purpose for, which the data to be used.
The classification process captures data from
different sources in different unstructured forms,
which when processes into information is used
to determine the economies of scale, switching
costs, business capital requirements, the degree
of fragmentation and concentration of the
business, the potential for globalisation,
bargaining power of suppliers and buyers,
threats of substitutes, business growth rates, the
cost advantages of the business, and the degree
of intensity of competition (Kudyba and
Hoptroff 34).
The data is derived from the millions of
transactions involving the suppliers, information
about the customers, and the day to day
transactions and operations.
Data mining for business intelligence
A case in point is Intel. Intel uses data mining to
shift through large amounts of data to establish
the relationship between data sets, anomalies in
business activities, significant business facts and
patterns, existing and emerging trends, and to
make decisions on the approach to use to make
sales.
In addition, it is used to determine product
ranges, establish and develop better marketing
strategies, and determine the strategic
approaches for creating customer royalties
(Sofaer 23).
Market segmentation is used by companies to
determine and classify customers with the same
buying behavior.
In addition, the data provides accurate
predictions of trends of loyal customers for
competitors to predict and avoid customer
churn, determine fraudulent transactions,
identify the most appropriate interactive
marketing strategies, establish the purchasing
trends of customers, and establish the
purchasing trends and behavior of the customers
(Larson 39).
Sources of data
Organisations look for different sources of data,
which is mined to address different business
needs in different environments and the typical
sources of data include website data mining,
journals, and e-commerce stores. Such data can
be precious when launching new products into
the market.
Problem statement
Many business organisations make inaccurate
decisions because they have not integrated data
mining tools to make real time and accurate
information to establish, predict, and determine
accurately the market segments, detect fraud,
and determine the marketing strategies for
effective decision making.
To address the problem, this study will analyse
the data mining processes, and its benefits,
which organisations use as a source of business
intelligence reports for accurate and effective
decision making.
Objectives
• Determine how organisations use data
mining as a tool to generate business
intelligence information
• Identify the data mining process that is
used by organisations to organise
unstructured data for decision making
• To determine areas where organizations
can use data mining for effective decision
making
Scope
The scope of the study is to cover study data
mining as a process, the areas where data mining
can be applied in business organizations, its use
in business organizations such as Intel, and the
sources of data that is used in data mining.
Research Methodology (Qualitative)
The method of inquiry used to achieve the
objectives of the study was qualitative research
method (Sofaer 45). The method allowed the use
of Intel and an example to build the study and
the literature on data mining for business
intelligence to analyze the findings.
Model
The research model consists of the process for
data mining, the use of a typical real-world
example, which use data mining to organize
unstructured data into information for business
intelligence as graphically presented below
The results of the qualitative study, which were
based on the objectives and scope of the study,
are discussed below.
Objective one
The results of the study showed that business
organisations use data mining for business
intelligence, which is crucial for determining the
current patterns of data consumer behavior.
Intel is one example of the companies, which
combine different technologies, architectures
and methodologies to mine data for business
intelligence, which is used for accurate decision
making, which is necessary for conducting cost
effective business transactions. In addition, the
study shows that Intel uses the data to
benchmark its business operations and
performance to determine the benefits that
accrue from its business operations.
Objective two
The data mining process that companies use to
accurately organise unstructured data into
business intelligence information includes
classifying the data into known and unknown
categories. The known category of data exists in
established patterns and the unknown category
of data has to be organised into specific patterns
using the rules that for classifying the data into
the known categories.
The prediction process is used to establish the
probability of the certain business trends such as
the changing buyer behavior according to the
dynamic business environment.
Gradually, the association rules are used to
determine what and why certain actions have to
be performed to achieve business goals and
objectives.
Objective three
Data mining provides business intelligence
information that can be used to make decisions
to determine the right market segments,
marketing strategies, to detect instances of
fraud, and determine the rate of customer
turnover to other business competitors (Howson
75).
Conclusions
The motivation for the study was to determine
the rationale for using data mining as a strategic
business tool for business intelligence. It was
established that data mining is a crucial tool that
competing organisations such as Intel use to
organise millions of unstructured business
transactions into information to influence
decision making.
The study showed that data mining can be used
to determine market competiveness, consumer
behavior, market segmentation, market trends,
marketing strategies, and customer movements.
The approach used automated predictions and
discovery mechanisms to organise unknown
patterns of data into information for decision
making. Here, the benefits from the study
included Information integration, enhanced
presentation capabilities, insight creation, better
organisational memory, benchmarking, and
effective decision making.
Works Cited
Delmater, Rhonda, and M. Hancock. Data
mining explained: a manager’s guide
to customer-centric business intelligence. New
York, Digital Press, 2001. Print.
Giudici, Paolo. Applied data mining: statistical
methods for business and industry. New York:
John Wiley & Sons, 2005. Print
Howson, Cindi. Successful business
intelligence. New Dell: Tata McGraw-Hill
Education, 2007.
Kudyba, Stephan, and R. Hoptroff. Data mining
and business intelligence: A guide to
productivity. New York: IGI Global, 2001. Print.
Larson, Brian. Delivering Business Intelligence
with Microsoft SQL Server 2005. New York,
McGraw-Hill, 2006. Print.
Sofaer, Shoshanna. “Qualitative research
methods.” International Journal for Quality
in Health Care 14.4 (2002): 329-336.Print.
Data Mining: A Critical Discussion Analytical
Introduction
In recent times, the relatively new discipline of
data mining has been a subject of widely
published debate in mainstream forums and
academic discourses, not only due to the fact that
it forms a critical constituent in the more general
process of Knowledge Discovery in Databases
(KDD), but also due to the increased realization
that this discipline can be applied in a number of
areas to enhance decision making processes,
efficiency, and competitiveness in contemporary
organizations (Kusiak, 2006).
The basic concept behind the emergence of data
mining, and which has contributed immensely to
its admissibility as one of the increasingly used
strategies in business establishments as well as
scientific and research undertakings, is that by
automatically sifting through large volumes of
information which may primarily appear
irrelevant, it should be possible for interested
parties to extract nuggets of useful knowledge
which can then be used to drive their agenda
forward (Adams, 2010).
Goth (2010) observes that the emergence of data
mining has been primarily informed by the rapid
growth in data warehouses as well as the
recognition that this heap of operational data can
be potentially exploited as an extension of both
business and scientific intelligence.
The present paper seeks to critically discuss the
discipline of data mining with a view to
illuminate knowledge about its origins,
concepts, applications, and the legal and ethical
issues involved in this particular field.
Definition & History of Data Mining
Although data mining as a concept has been
defined differentially in diverse mediums, this
report will adopt the simple definition given by
Payne & Trumbach (2009), that “…data mining
is the set of activities used to find new, hidden or
unexpected patterns in data” (p. 241-242).
The purpose of data mining, as observed by
these authors, is to extract information that
would not be readily established by searching
databases of raw data alone. Through data
mining, organizations are now able to combine
data from incongruent sources, both internal and
external, from across a multiplicity of platforms
with a view to assist in a variety of business
applications.
At its most elemental state, data mining utilizes
proved procedures, including modeling
techniques, statistical investigation, machine
learning, and database technology, among
others, to seek prototypes of data and fine
relationships in the sifted data with the main
objective of deducing rules and intricate
relationships that will inarguably permit the
extrapolation of future outcomes (Pain &
Trumbach, 2009; Adams, 2010).
Researchers and practitioners are in agreement
that the capability of both generating and
collecting data from a wide variety of sources
has greatly impacted the growth trajectories of
data mining as a discipline.
This capability, according to Adams (2010) and
Chen (2006), was precipitated by a number of
variables, which can be categorized into the
following:
1. increased computerization of business,
scientific, and government transactions
with the view to increase efficiency and
productivity,
2. extensive usage of electronic cameras,
scanners, publication devices, and
internationally recognized bar codes for
most business-related products,
3. advances in data gathering instruments
ranging from scanned documents and
image platforms to global positioning and
remote sensing systems,
4. the development and popularization of the
World Wide Web and the internet as
widely accepted global information
systems.
This explosive growth in stored or ephemeral
data brought us to the information age, which
was, and continues to be, characterized by an
imperative need to develop new techniques,
procedures and automated tools that can astutely
assist us in transforming and making sense of the
huge quantities of data collected via the above
stated protocols (Goth, 2010).
To dig a bit deeper into the history of data
mining, research has been able to establish that
the term ‘data mining’, which was introduced in
the 1990s, has its origins in three interrelated
family lines. It is important to note that the
convergence of these family lines to develop a
unique discipline in the context of data mining
certainly gives it its scientific foundation
(Adams, 2010).
This notwithstanding, extant research (Adams,
2010; Chez, 2006) demonstrate that the longest
of these family lines to be credited with the
gradual development of data mining as a fully-
fledged discipline is known as classical
statistics.
Researchers are in agreement that it would not
have been possible to develop the field of data
mining in the absence of statistics as the latter
provides the foundation of most technologies on
which the former is built, such as “regression
analysis, standard distribution, standard
deviation, standard variance, discriminant
analysis, and confidence intervals” (Goth, 2010,
p. 14).
All these concepts, according to this author, are
used to study data and data relationships –
central aspects in any data mining exercise.
The second longest family line that has
contributed immensely to the emergence of data
mining as a fully-fledged field is known as
artificial intelligence, or simply AI. Extant
research demonstrate that the AI discipline,
which is developed upon heuristics as opposed
to statistics, endeavors to apply human-thought-
like processing to statistical challenges while
using computer processing power as the
appropriate medium (Talia & Trunfio, 2010).
It is important to mention that since this
approach was tied to the availability of
computers and supercomputers to undertake the
heuristics, it was not practical until the early
1980s, when computers started trickling into the
market at reasonable prices (Goth, 2010).
The third family line to have influenced the field
of data mining is what is generally known as
machine learning or, better still, the
amalgamation of statistics and AI (Adams,
2010). Here, it is of importance to note that
while AI could not have been viewed as a
commercial success during the formative years,
its techniques and strategies were largely co-
opted by machine learning.
It is also important to note that machine learning,
while able to take the full benefit of the ever-
improving price/performance quotients
provided by computers in the decades of the
1980s and 1990s, found usage in more
applications because the entry price was lower
that that of AI, not mentioning that it was largely
considered as an evolved facet of AI as it was
effectively able to blend AI heuristics with
complex statistical analysis (Chen, 2006).
Review of how Data Mining is used Today
and how it could be used in the Future
Presently, there exist broad consensus that data
mining is mostly based on the machine leaning
techniques; that is, it is fundamentally perceived
as the adaptation of machine learning techniques
and concepts to a wide variety of areas, such as
business and scientific applications (Adams,
2010).
Therefore, the present-day data mining can only
be described as the amalgamation of historical
and recent developments, particularly in
statistics, artificial intelligence, and machine
learning, with a view to developing a software
program that can run on a standard computer to,
among other things, make diverse decisions
based on the data under study, use statistical
concepts and applications to establish various
relationships among the data, and also use more
advanced artificial intelligence heuristics and
algorithms to achieve its major goal (Talia &
Trunfio, 2010).
Extant research demonstrate that the major
objective of current data mining applications is
to sift through huge volumes of data to extract
nuggets of useful data, which can then be used
to establish previously-hidden trends or patterns.
Today, more than ever before, data mining is
used in the business arena to boost corporate
profits by improving customer relations and
targeting new customers (Cary et al, 2003).
According to these authors, “…AT&T Wireless
was able to increase it’s subscriber base by 20%
in less than a year when it contracted with a data-
mining company to identify customers that
would likely to be interested in AT&T’s new
flat-fee wireless service” (p. 158).
The AT&T story demonstrates that visions of
achieving good returns continue to drive
businesses toward embracing data mining
technology.
Data mining is bound to be used along the same
lines in the future to enable enterprises make
critical decisions from a knowledge-oriented
perspective. Consecutive studies have
demonstrated that most business organizations
fail to wade through the harsh economic waters
of modern times due to their perceived
inadequacy to base their most important
decisions on knowledge and evidence (Adams,
2010; Goth, 2010).
However, it is now evident that data mining can
be used to endear organizations closer to a
knowledge-based economy, which basically
translates into the use of knowledge to generate
economic benefits.
Chen (2006) observes that a knowledge-based
economy necessitates data mining processes to
become more goal-oriented with the view to
generating an enabling environment where more
tangible results can be achieved.
Consequently, data mining should be used in the
future not only to facilitate the uncovering of
concealed knowledge beneath the ocean of data
readily found in a multiplicity of mediums and
applications, but also to ensure that it makes
important contributions to the knowledge-based
economy with the express intention of coming
up with more tangible business and scientific
outcomes (Chen, 2006; Adams, 2010).
Types of Data Mining Applications
There exist a multiplicity of data mining
applications which can be used in diverse
situations and environments depending on the
major objective for usage. Some data mining
applications, according to Chen (2006), are
simple to use and may be offered for free, while
others are complex and require a sizeable
investment to operationalize.
This section will discuss some data mining
applications based on the sector of practice, and
will mainly focus attention to the banking and
finance, retail, and the healthcare sectors of the
economy.
Data Mining Applications in the Banking &
Finance Sector
Most banking institutions have over the years
employed a multiplicity of data mining
applications to model and predict credit fraud, to
assess borrower risk, to undertake trend
analysis, and to evaluate profitability, as well as
to assist in the initiation and management of
direct marketing activities (Seifert, 2004).
In equal measure, most finance and credit
companies have over the years employed a
variety of neural networks and other data mining
applications “…in stock-price forecasting, in
option trading, in bond rating, in portfolio
management, in commodity price prediction, in
mergers and acquisitions, as well as in
forecasting financial disasters” (p. 191).
Here, it can be noted that the Neural
Applications Corporation has developed an
effective application known as NETPROPHET,
which is increasingly being used by finance
companies to make stock predictions by
illustrating the real and predicted stock values
depending on the type of data that has been
keyed into the system (Groth, 1999; Chen,
2006).
The banking sector is continuously been faced
with fraud cases, and data mining applications
such as HNC Falcon has assisted the institutions
to monitor payment-card applications,
decreasing fraud cases by almost 75 percent
while increasing applications for payment card
accounts by as much as 50 percent on a yearly
basis (Groth, 1999).
The importance of banks to develop data mining
applications that could be used in cross-selling
and maintenance of customer loyalty has been
well documented in literature.
These applications, according to Groth (1999),
mainly assist banking institutions to model the
behavior of their customers in such a manner
that the resulting relationships could be used to
establish the needs and demands of their
customers, as well as make objective predictions
into the future.
RightPoint software, Security First, and
BroadVision are some of the vendors primarily
interested in integrating predictive technologies
with consumer interaction points to ensure
customer needs and demands are efficiently
dealt with (Groth, 1999; Chen, 2006), in
addition to using predictive technologies to
integrate one-to-one marketing strategies to
their clients banking sites (Adam, 2010).
According to Groth (1999), “…the RightPoint
Real-Time Marketing Suite takes data-mining
models and leverages them within real-time
interactions with customers” (p. 194). This
application is unique in that it is designed to
develop, manage and deliver one-to-one
marketing initiatives for high-end industries that
heavily depend on direct customer interaction to
undertake business (Goth, 2010).
As a general prerogative, it is important to note
that majority of the data mining applications
used in the banking and finance sector attempt
to ensure that each customer interaction seizes
the prospect of enhancing customer satisfaction,
loyalty, motivation, and profit-generation
potential (Talia & Trunfio, 2010; Zhang &
Segall, 2008).
Data mining Applications in Retail
Intense competition and slim profit margins
have obliged retailers to embrace data
warehousing strategies earlier than other sectors.
As observed by Groth (1999), “…retailers have
seen improved decision-support processes lead
directly to improved efficiency in inventory
management and financial forecasting” (p. 198).
It is a well known fact that expansive retail and
supermarket chains are in possession of huge
quantities of point-of-sale data that is not only
information-rich, but could be employed using
appropriate data mining applications to improve
the stated decision-support strategies, improve
efficiency in financial predictions and inventory
management, and analyze customer shopping
patterns (Seifert, 2004).
In the retail sector, the AREAS Property
Valuation product from HNC software, as well
as SABRE Decision Technologies, serves as
good examples on how data mining applications
can be used in the retail sector to perform
valuations, projection and forecasting, customer
purchasing behavior analysis, and customer
retention analysis, with the underlying purpose
of increasing profitability, enhancing customer
experience, and making better and more
informed business decisions (Zhang & Segall,
2008).
In evaluating customer profitability in the retail
sector, a software vendor referred to as Dovetail
Solutions has developed a data mining
application known as Value, Activity, and
LoyaltyTM (VALTM), with a view to utilize
transactional business data from the retailers to
synthesize information about customer activity
and processes, churn rate, and anticipated future
purchases (Groth, 1999; Chen, 2006).
Data Mining Applications in the Medical
Field
The vast amount of data available within the
healthcare industry, including the associated
data collected via medical research, biotechs,
and the pharmaceutical industry, have provided
a fertile ground for data mining applications to
grow. The knowledge that data mining has been
employed expansively in the medical industry is
in the public domain.
For example, we are aware that the vendor
NeuroMedical Systems ingeniously employed
neural networks to create a pap smear diagnostic
aid, while both Vysis Rochester Cancer Center
and the Oxford Transplant Center continues to
employ a data mining application known as
KnowledgeSEEKER, which utilizes a decision
tree technology, to assist in various research
undertakings (Groth, 1999; Adams, 2010; Chen,
2006).
It is important to note that these applications are
beneficial in the medical sector as they enable
health practitioners to come up with accurate
diagnosis even without subjecting patients to
physical examination (Koh & Tan, 2008).
Governments and other interested health
agencies can utilize data mining applications,
such as MapInfo, KnowledgeSEEKER, and
LEADERS, among others, to: demonstrate
average costs of health services; show efficiency
of a particular prescription over time; reveal
efficacy rates of diverse pathogens over time;
develop superior diagnosis and treatment
protocols; show patient location in order to
deliver superior health services; or assist
healthcare insurers to detect fraud (Koh & Tan,
2008; Wen-Chung et al, 2010).
Legal & Ethical Issues in Data Mining
As is the case in other disciplines, the field of
data mining is faced with a complexity of legal
and ethical issues which needs to be addressed
for the applications to succeed.
In the legal arena, it is important to evaluate how
organizations should employ data mining
applications while remaining focused on
protecting the private information of their
customers so as to avoid customer
dissatisfaction or even being subjected to legal
action by customers who may feel that the
organizations intruded into their privacy (Cary
et al, 2003; Wen-Chung et al, 2010).
In terms of ethical issues, it is a well known fact
the spread of personal information, as is the case
in many data mining applications, can lead to
elevated risks of customer identity theft (Cary et
al, 2003).
According to Payne & Trumbach (2009), data
mining processes brings into the fore a scenario
where “…the consumer loses aspects of privacy
as all of their basic demographic information,
personal interests, correspondence and activities
are stored in databases and available to be
combined together” (p. 243).
Such a scenario has obvious ethical
ramifications since this information can be used
to the disadvantage of the customers.
Another ethical query arises from the fact that
consumers lose the control over what happens to
their personal information held in large
databases, implying that such kind of
information can be used to the disadvantage of
the providers if it happens to fall into the wrong
hands (Payne & Trumbach, 2009; Cary et al
2003; McGraw, 2010).
What’s more, customers who provide personal
information to organizations face a more
ominous challenge that may entail potential
discrimination based on the personal
information they either provide or refuse to
provide to the organizations.
Another important factor to consider when
evaluating ethical concerns in data mining is that
there is no fine line distinguishing if it is indeed
necessary for an organization to use the private
information of its customers to enhance its
profitability or if such information should be
sorely used to improve customer satisfaction and
maintain consumer trust (Payne & Trumbach,
2009).
Lastly, it is well known that a number of data
mining processes may yield incorrect
conclusions, which may be costly to the
organization as well as to the customers
(McGraw, 2010).
Conclusion
This discussion has brought into the fore
important aspects of data mining, its current and
future uses, as well as perceived limitations in
terms of legal and ethical constraints.
The general consensus among academics and
practitioners is that data mining represents the
new frontier of growth, particularly in nurturing
mutually fulfilling customer relationships,
ensuring that customer’s needs and demands are
satisfactorily met, and in facilitating
organizations to forecast and predict future
growth and decision patterns (Adams, 2010;
Kusiak, 2006; Wen-Chung et al, 2010).
The task is therefore for developers to continue
investing heavily in effective and efficient data
mining applications to ensure that such tools are
able to achieve what they were originally
intended to achieve. Consequently, research and
development into these applications and tools is
of primary importance.
Reference List
Adams, N.M. (2010). Perspectives on data
mining. International Journal of Market
Research, 52(1), 11-19. Retrieved from
Business Source Premier Database
Cary, C., Wen, H.J., & Mahatanankoon, P.
(2003). Data mining: Consumer privacy, ethical
policy, and systems development
practices. Human Systems Management, 22(4),
157-168. Retrieved from Business Source
Premier Database
Chen, Z. (2006). From data mining to behavior
mining. International Journal of Information
Technology & Decision Making, 5(4), 703-711.
Retrieved from Business Source Premier
Database
Goth, G. (2010). Turning data into
knowledge. Communications of the ACM,
53(11), 13-15. Retrieved from Business Source
Premier Database
Groth, R. (1999). Data mining: Building
competitive advantage. Upper Saddle River, NJ:
Prentice Hall
Koh, H.C., & Tan, G. (2008). Data mining
applications in healthcare. Journal of
Healthcare Information Management, 19(2),
64-72
Kusiak, A. (2006). Data mining: Manufacturing
and service applications. International Journal
of Production Research, 44(18/19), 4175-4191.
Retrieved from Business Source Premier
Database
McGraw, D. (2010). Data identifiability and
privacy. American Journal of Bioethics, 10(9),
30-31. Retrieved from Academic Search
Premier Database
Payne, D., Trumbach, C.C. (2009). Data mining:
Proprietary rights, people and
proposals. Business Ethics: A European Review,
18(3), 241-252. Retrieved from Business Source
Premier Database
Seifert, J.W. (2004). Data mining: An overview.
Retrieved from
<https://fas.org/irp/crs/RL31798.pdf>
Talia, D., & Trunfio, P. (2010). How distributed
data mining tasks can thrive as knowledge
services. Communications of the ACM, 53(7),
132-137. Retrieved from Business Source
Premier Database
Wen-Chung, S., Chao-Tung, Y., & Shian-
Shyong, T. (2010). Performance-based data
distribution for data mining applications on grid
computing environments. Journal of
Supercomputing, 52(2), 171-198. Retrieved
from Academic Search Premier Database
Zhang, Q., & Segall, R.S. (2008). Web mining:
A survey of current research, techniques, and
software. International Journal of Information
Technology & Decision Making, 7(4), 683-720.
Retrieved from Business Source Premier
Database
Data Mining Role in Companies
The increasing adoption of data mining in
various sectors illustrates the potential of the
technology regarding the analysis of data by
entities that seek information crucial to their
operations. Data mining tools enable entities to
establish relationships such as associations,
classes, clusters and sequential patterns.
The analysis of data can occur using a variety of
techniques such as genetic algorithms, rule
induction, neural networks and data
visualization. Information obtained through data
mining has transformed business operations by
aiding companies in decision-making and the
prediction of crucial factors such customer’s
behavior, which helps companies to gain a
competitive advantage.
Benefits of data mining to businesses
The generation of predictive scores for
organizational elements is crucial in the analysis
of the behavior of customers. Predictive
analytics enable companies to optimize
marketing strategies by modeling trends in
customers’ responses. This information helps
businesses to allocate funds for various
campaigns based on their potential of
succeeding.
In addition, it minimizes wastage of time and
money caused by the use of manually analyzed
marketing strategies. Research shows that
manual analysis of marketing methods is a
cumbersome process that is error-prone due to
the extensive skills required. Predictive
analytics eliminate guesswork in the
identification of marketing methods by
providing reliable data on customers’
preferences and habits (Pyle, 2003).
Analyzing past and present trends about
customers’ habits provides patterns that aid in
making decisions on future undertakings of a
business. A company that implements measures
in response to future patterns of customers’
behavior is likely to gain a competitive
advantage.
Data mining enables businesses to establish
relationships among items in a transaction. The
association technique facilitates identification of
products that customers purchase frequently.
Using the association rule, businesses can
determine the way one product influences the
sale of another product.
For example, an association of Bagels and
Potato Chips provides insight on products that a
fast foods business can sell with Bagels so that
the sale of Potato Chips increases.
Market based analysis provides businesses with
important information that guides the
implementation of marketing campaigns by
enabling them to establish hypotheses for
customers’ buying patterns. Relevant marketing
strategies boost sales and promote higher profits.
Web mining enables business entities to
establish patterns of customers’ behavior from
the web. Using data mining techniques,
companies can identify customers’ interests on
the web concerning textual or multimedia data
(Soares, 2010).
Web usage mining facilitates the analysis of
target demographics. Web content and
structured mining enable companies to monitor
brands and analyze the content and structure of
competitors. Such undertakings create strategic
advantages.
Clustering enables companies to identify
distinct groups of customers and implement
strategies to retain customers that are above a
cluster and gain the confidence and loyalty of
customers that are below the cluster.
Customer-relationship management uses
clusters to segment customers based on
particular variables indentified through data
mining. Variables such as customer-retention
probability help companies in the identification
of marketing opportunities.
Reliability of data mining algorithms
The reliability of data mining algorithms
depends on the nature of data under analysis.
Some datasets contain information that has
errors or is invalid. Research shows that
algorithms have diverse responses to errors and
thus compromise the results of data analysis in
different manners. Assumptions such as noise-
free data influence the accuracy levels of data
mining algorithms.
Another factor that interferes with the accuracy
of data mining algorithms is the size of data. The
search space varies depending on the
dimensions in a domain space.
Research shows that the relationship between
the search space and dimensions in a domain
space is exponential. This relationship
introduces a phenomenon known as curse of
dimensionality, which interferes with the
reliability of data mining algorithms
(Kantardzic, 2003).
Various errors arise due to factors that affect
data-mining algorithms. Systematic errors are
likely to arise due to assumptions on clean data.
Preprocessing of data helps to minimize such
errors. Other errors include training and
pessimistic errors that arise due to invalid data
and assumptions such as noise-free data.
Data mining infringement on privacy
Data for mining purposes raises many privacy
concerns. First, data intended for profiling
customers and analyzing their behavior contains
a lot of personal information. The collection and
storage of confidential information about
individuals introduces controversies due the
possibility of illegal access to the information.
Another issue concerns the dissemination of
implicit information about an individual or a
group of customers. Thirdly, data mining
discovers valuable information that is subject to
sale. This creates loopholes for the distribution
of confidential information without control.
To address the issue of privacy protection in data
mining, concerned bodies have established
measures that promote reliable data mining
results while meeting privacy requirements.
OECD Guidelines on data mining extensively
cover the use of personal information obtained
through data mining by providing various
guidelines (Aggarwal & Yu, 2008).
First, executors of data mining should clearly
inform subjects on the process and the intended
use of the collected data. This will ensure that
people participate under their free will.
Secondly, the Forthcoming Policy regulates the
use of data mining results by stipulating the
purpose of data, its allowed use, and persons
who should access the information. The
Disclosure policy enables subjects of data
mining to determine purposes for disclosure of
knowledge by giving or denying consent on the
anticipate use of data.
Businesses that have used predictive analysis
to gain a competitive advantage
Flight 540 employed the predictive analytics
strategy to boost their customer appeal and gain
competitive advantage.
They used data from the internet, customer
spending history, comment cards and various
surveys to model their flights in such a manner
that customers could easily make purchase
decisions depending on the flight packages on
offer. Instead of planning flights on a
standardized basis, the company modified them
as per groups of customers with similar
preferences.
Netflix became a key player in the video rental
business by using predictive results on customer
preferences to model a customer-centric
business approach. The company introduced
products such as video rentals that did not
constitute a lateness fee, and the provision of
video streaming based on customers’ requests.
Apart from expanding its market share, the
company succeeded in reducing its promotional
expenses.
Tusk supermarket was able to increase its sales
considerably by using predictive analytics to
establish feasible pricing strategies using data on
customer survey feedback. The data facilitated
estimation of price sensitivity and the
determination of price ranges that would have
minimal impacts on sales. This promoted
customer retention for the supermarket, increase
market share, and stabilized generation of
revenue.
Conclusion
Data mining provides companies with a basis for
analyzing customers’ behavior and responses of
potential customers by enabling them to
establish relationships among internal and
external factors of business operation. In this
regard, companies can effectively determine the
role of factors such as price, competition and
demographics in influencing customers’
behavior.
Commercial Uses of Data Mining
Data mining is a computer-based classification
and summarization of a large set of data into
grouped information based on the correlated
data. The data mining tools present the
information in a manner that is to interpret and
predict. Data mining process entails the use of
large relational database to identify the
correlation that exists in a given data.
Data used in this process is drawn from different
sources that are related to the outcome of the
target phenomenon (Chattamvelli 18). Experts
use data mining software to predict the
occurrence of various situations and events
(Miner 35). This paper analysis how data mining
is used commercially to earn money, optimize
performance and improve service delivery.
Business organizations use data mining
technology to improve business management.
These organizations use customized data mining
applications to generate business trends and
patterns from their vast relational database. The
applications have specialized inbuilt software
algorithm that performs correlation detection.
The principal role of the applications is to sift
the data to identify correlations. Eventually, the
software presents the correlations as statistical
trends and patterns. Various business
departments use data mining for management as
follows (Chattamvelli 27).
Marketing departments use data mining to
perform market analysis. Market analysis is the
process of monitoring changes in marketing
pattern and trends. Therefore, marketers use data
mining applications to investigate the possible
causes of changes in marketing trends (Miner
53). Research asserts that common changes in
market analysisss include customer complaints,
sales decline, and loss of customer’s loyalty.
In addition, marketing departments use market
analysis to monitor new products launched in the
market. Specialized data mining applications
display customer buying trends, post purchasing
behavior, customer reviews, and demand trends
among others. Marketers use the prevailing
market trends to identify the best promotional
strategy that will outdo the competitors
(Rahman 68).
Marketing departments also use data mining
technology to facilitate direct and interactive
marketing. Direct marketing is the process of
developing a comprehensive business mailing
list. The mailing list contains the addresses of
the all the business stakeholders (Han & Kamber
56).
Business experts assert that direct marketing is
one of the most efficient ways of winning
customer and supplier’s loyalty. On the other
hand, interactive marketing is the process of
optimizing the business web accessibility. The
process ensures the business website provides
the user with satisfactory information.
Successful interactive marketing leads to
improved customer loyalty and online purchases
(Soares & Ghani 58).
Customer care departments use customer data
mining applications to improve service delivery
on customer relations. The main function of
these applications is to automate customer
handling process to minimize response time.
Some advanced data mining tools have
automated inbuilt emailing feature that
automatically respond to general customer
concerns (Miner 65). Marketing departments
also use this feature to promote direct marketing.
Furthermore, the departments use specialized
data mining applications that automate
answering of the frequently asked questions
(FAQ). FAQ tool provides customers with a
quick online guide of solving various problems
without visiting the customer care. This practice
improves service delivery leading to improved
customer loyalty (Soares & Ghani 67).
Customer care department also uses customized
data mining applications to manage business
promotions. Promotional data mining
applications automate the process of selecting
winners in uplift modeling promotions (Rahman
45). Furthermore, the departments use data
clustering applications to automate customer
segmentation analysis. Market segments are
distinct groups of customers who buy similar
goods or services.
Market segmentation process helps the business
in identification of the most profitable customers
group for promotions and marketing. In
addition, customer care departments heavily rely
on data mining software for catalogue
marketing. The departments manage huge
database that host profiles of all other
departments in the organization. Therefore, the
department uses product profiles at their
disposal to produce product catalogues for sales
promotion (Han & Kamber 77).
Commercial companies widely used data
mining analysis in the human resource (HR)
department. HR departments use customized
applications to monitoring and optimize
employee’s performance. The application
provides information on employees working
history.
This information is essential in planning for
recruitment, promotion, training, and rewarding
(Miner 84). Nevertheless, HR department use
Strategic Management application to generated
key performance indicators (KPI). KPIs analysis
provides business management with a
performance progress report. The report
provides a summary of the possibility of
achieving the organization’s goals. Therefore,
production and sales managers use KPIs to
correct performance decline hence improving
the profit margin (Rahman 73).
Entrepreneurs use data mining technology to
make money. For instance, software
development experts design and sell data mining
software to companies, businesses, and
institutions. Research reveals that hundreds of
entrepreneurs have become millionaires from
the sale of data mining software. In addition, the
entrepreneurs does not require established
companies to sale their software.
Other entrepreneurs make money from data
mining by operating data mining agency. These
professionals use the available data mining
software to open a data mining consulting
agency. Later they approach institutions,
businesses and companies to help them run their
businesses. The agency charges consultation fee
at a favorable rate depending on the amount of
work and urgency (Rahman 80).
Business people can also use data mining
technology to save or make money. If the
business owners are experienced in data mining,
they can use this knowledge to analyze market
trends. From the analysis, they can predict
market threats like decline of price or loss of
market for their goods and services.
The business can use this opportunity to reduce
production or supply of services hence avoiding
losses. On the other hand, the business people
can predict a possible rise of demand of their
goods or services. In that case, the business
increases supply of its goods and services to
meet the expected demand. Research asserts that
businesses using this approach are highly
profitable since they always avoid losses and
maximize profits (Soares & Ghani 126).
Web and Search Engine Optimization (SEO)
designers use data mining to make money
online. Using reliable data mining software, the
designers identify the most researched keyword
for their website in different locations. The
designers use this keyword to optimize their
website content.
Such websites receive many organic visitors,
which is profitable for the businesses. Research
shows that most successful online selling
websites use this approach to generate traffic.
The method is also very profitable for SEO
experts who are paid on a commission basis.
When their website ranks high in the search
engines, they are likely to get more clicks or
signups hence earning more commission
(Rahman 142).
In conclusion, commercial organizations widely
use data mining as follows. Sales department of
commercial companies use data mining to carry
out customer churning. Customer churning
involves using the buying behavior to predict
whether the company is losing its customers to
its rivals. Management uses the customer
churning results to regulate production hence
averting losses. Marketing department uses data
mining technique to facilitate market analysis,
segmentation, maximizing profits and
minimizing expenditure.
Human resource department uses data mining
technology to oversee staff monitoring,
development and termination. Finally,
entrepreneurs have made a lot of money using
data mining software. In summary, use of data
mining is very instrumental for business
organization management and investment
(Soares & Ghani 158).
Chattamvelli, Rajan. Data mining algorithms.
Oxford: Alpha Science International, 2011.
Print.
Han, Jiawei, and Micheline Kamber. Data
mining: concepts and techniques. 3rd ed.
Amsterdam: Morgan Kaufmann, 2012. Print.
Miner, Gary. Practical text mining and
statistical analysis for non-structured text data
applications. Waltham, MA: Academic Press,
2012. Print.
Rahman, Hakikur. Ethical data mining
applications for socio-economic development.
Hershey: Information Science Reference, 2013.
Print.
Soares, Carlos, and Rayid Ghani. Data mining
for business applications. Amsterdam: IOS
Press, 2010. Print.