advantages and disadvantages of using a proprietary file system or format with regard to compliance and governance

profilevars
datamining1.docx

Running Head: Data Mining in The Cloud 1

Data Mining in The Cloud 13

Data Mining in The Cloud

Student’s Name:

Institution:

Instructor:

Big data mining on the cloud

Big data mining techniques

Abstract.

Management and analysis of data is becoming a nightmare in every organization day by day. This is because there is flooding of data. This data can only be analyzed by using Information Governance and big data mining techniques. This paper aims to look at some of the big data mining techniques which can be used to analyze data in organizations with flooding of data. It will also show how information governance support big data. The paper begins with an overview of data mining, narrows down to the big data mining techniques and then finally the ways in which Information governance support big data.

Introduction

Data mining is the way toward looking at tremendous amounts of information so as to make a factually likely expectation. Data mining can be utilized, for example, to recognize when high going through clients connect with your business, to figure out which advancements succeed, or investigate the effect of the climate on your business. Information mining standards have been around for a long time related to information distribution centers, and have now taken on more noteworthy pervasiveness with the appearance of Enormous Information. Information examination and the development in both organized and unstructured information has likewise incited information mining strategies to change, since organizations are currently managing bigger informational collections with progressively fluctuated substance (Khan, Anjum, Soomro and Tahir, 2015). Also, man-made brainpower and AI are mechanizing the procedure of data mining.

Despite the methods applied, data mining involves three steps. These steps include exploration, modelling and deployment. The data must first be prepared and sorted out to is needed and what is not needed. This helps one to do away with useless data or even duplicates and ensuring that the final data that is sampled is the only one that is crucial and needed the most. Creating the statistical models with the aim of determining the one which will give the best and most accurate forecasting. This however can consume a lot of time as there are various and different models to the same data set which is applied severally to the sets of data respectively and finally analysis of data should be done. Lastly, in the last step, the model has to be tested against the old and the current data (Milani & Navimipour, 2017). This helps an individual to determine the results which he or she should expect in future.

Big data mining techniques

Data mining is a very significant and effective method when proper techniques are applied. One should be able to choose or to select the most effective technique depending on the situation or the data being analyzed. This is because the big data mining techniques are several and some may or may not perform better in every analysis of data (Benjelloun, Lahcen & Belfkih, 2015, March). Therefore, the major big data mining techniques that an individual should consider using include, classification analysis, clustering, anomaly or outlier detection, regression analysis, prediction or the induction rule, summarization, sequential patterns, decision tree learning, tracking patterns, statistical techniques, visualization, neural networks, data warehousing, association rule learning and long-term memory processing.

Classification analysis

This type of technique is used to put data into various and different classes. It is similar to clustering in that it breaks down information into portions which are referred to as classes. When using this technique, the composition of every dataset is determined and analyzed. an individual should consider the different types of data at hand so that he or she can classify it into classes according to the number of classes that he or she wants for instance an individual may decide to classify data as inbox or incoming messages and outgoing messages. This will be put in different classes.

In most cases the classification is discrete and do not necessary show or portray order. The simplest method of classifying data is through the use of binary classification. The aimed feature contains only two potential values for example a high rate and a low rate. Multiclass targets are the ones which consist of more than two values. The values in the multiclass may show low rate, medium or high rate. The classification frameworks are tested by comparing the expected values to the targeted values in the datasets. This technique can be applied in organizations for different purposes and in different areas (Sajana, Rani & Narayana, 2016). For example, it can be used in segmenting or even classifying customers, modelling businesses, marketing and also in analysis of credits.

Clustering

Clustering in data mining alludes to putting of particular and specific objects in groups depending on the features, their factors and what they possess in similar. This procedure classifies the data that best fit to the desired analysis using a special join algorithm. However, this classification permits an object not to be strictly part of the group. The quality of every group depends on the method used. Clustering can also be referred to as data segmentation because it partitions large data sets into classes based on their similarities.

This type of data mining technique can be significant in several fields. Some of these fields include in which the technique can be applied include marketing, biology, and library. There are several procedures involved in the clustering process. These procedures include partitioning whereby several clusters are made and the determined and analyzed based on the provided criteria. The second procedure is the hierarchical method whereby the sets of data are arranged in order using a certain criterion. There is also the density-based method whereby the procedure is based on density and connectivity (Yusof, Zurita-Milla, Kraak & Retsios, 2016). Finally, there is the grid-based method whereby it is based on the double resolution of the grid data framework.

Anomaly or outlier detection

Anomaly identification is the way toward discovering anomalies in a given data set. Anomalies are the information questions that stand apart among different items in the data set and don't comply with the ordinary conduct in a data set. Abnormality discovery is an information science application that joins various information science errands like order

Classification, regression, and clustering. The target variable to be predicted is whether a transaction is an outlier or not. Since clustering tasks identify outliers as a cluster, distance-based and density-based clustering techniques can be used in anomaly detection tasks.

Outliners can be of two types that is univariate and multivariate. Univariate is seen when analyzing a broad range of values in a single characteristic space while multivariate outliners is mostly seen in a n-dimensional space. This can be very difficult and challenging to human beings when it comes to that analysis of data and that is why a model is created to help an individual analyze the data. The common causes of outliners in a data set include mistakes made during entry of data, measurement mistakes, mistakes made during testing and the mistakes made during processing of data. Mistakes committed during the process of sampling can also lead to presence of outliners in the data.

Regression analysis

This is a statistical procedure used for approximating the link between variables which aids an individual to know and be able to explain the feature or the behavior of the dependent variable. It is usually used for forecasting. It helps in determining if any one of the independent variables is changed, so if you adjust one variable, the other one will also adjust. This technique of data mining is applied in several companies and organizations for business marketing and planning. It is also used for analyzing of trends and predicting what is likely to happen in future. There are several forms and methods of regression analysis. These forms include linear regression, standard multiple regression, stepwise multiple regression, hierarchical regression and setwise regression.

Prediction or Induction rule

This involves a data mining procedure of clarifying whether the if-then policies from a dataset. This decision illustrates and explains the relation between the features and the class labels in the data set. Many real-life situations are based on the use of this technique. This method makes use of the previous or old data and information to forecast what is likely to happen in future. This method assumes that if a certain action occurs frequently, then it is most likely that the same action will keep on occurring every now and then in future.

Big firms and organizations use this technique to predict what is may affect their businesses in future. This helps to prevent or curb risks which the organization may face in the coming days. For instance, if a bank wants to lend out a loan to a client, the bank may use this technique to determine if the client will be able to repay back the loan in good time by looking at the past credit records of the client. If the client has a history of delaying payments, them with the help of this rule the bank may be forced to deny the client the loan with an assumption that he or she will still delay the payments (Benjelloun, Lahcen & Belfkih, 2015, March). This method of data mining can also make the organization to come with risk mitigation mechanisms and disaster recovery plans since this technique will help it to know the risks which may face the organization or the firm in question in the coming days.

Summarization

Summarization is a procedure in which a code is used to summarize or compress data. Data summarization is important is important because the world is becoming more digital. Employees also work with large sets of data bringing up the need of the data to be summarized and compresses so that it can be easily understood. Data volume is inevitable because when data is extracted from different sources, it cannot be termed or determined the amount which will be stored or retained in the dataset. Because of this, data becomes very much complicated and a lot of time is needed to analyze it. One needs to extract information following the class and the kind of data you would like to store. This will help an individual to filter the crucial data and ensure that the stored data is the most critical (Yu, Li, Xiang & Zhang, 2018). Summarization of data can do with the aid of tables using excel whereby different formulas are applied.

Sequential patterns

This technique is concerned with identifying statistically related and relevant patterns between data examples in areas or situations where the values are presented in a sequence. It is usually assumed that the values are discrete and therefore time series mining is closely related but considered a different activity. This technique also incorporates the discovery of interesting, useful and unexpected patterns in data sets (Yang & Gidófalvi, 2018). Different and various forms of patterns can be identified in sets of data such as the items that appear frequently, sub graphs, sequential rules and also periodic patterns.

Traditionally, sequential pattern mining was applied to determine subsequences that appear frequently in a sequence database. For instance, in some situations sequential patterns can be applied to identify the sequences of objects that are frequently brought by customers. This will help the firms or organizations to be able to understand the behavior of the customers, therefore enabling them to come up sound marketing decisions.

Visualization

Data visualization in data mining involves identifying and visualizing information in a clear and simple way without any form of reading or writing. The results of the data is mostly displayed in the form of pie charts, bar graphs, statistical representation and also by use of graphical forms as well. The primary goal of this technique is to be able to bring out the information efficiently and clearly without deviating from the expected or from the topic under study. The graphs help the reader or the one analyzing the data to be able to comprehend it easily. This also saves time and ensures that the analyzer does not waste a lot of time interpreting the data. It also minimizes confusion and errors which could occur in the process of analyzing the data.

Neural Networks

Neural network is a data mining technique that involves the processing of the gathered as well as extraction of data by the recognition of the patterns which are existing in the database patterns, which is done by using neural network which is artificial. Through the techniques of the artificial neural network, there involves the structuring of the neural network in the humans, with the neurons being the conduit with regard to five senses. This artificial neural network is used as conduit for the purpose of data input despite the fact that it is usually a complicated mathematical equation for the processing of data rather than presenting it as sensory input. Basically, the neural network involves the interconnection of different groups of the artificial neurons, which involve the processing of information by the approach of connectionist for the purposes of computation. In the businesses and organizations, neural networks as a technique of data mining help in turning data which is raw to useful information (Mathan, Kumar, Panchatcharam, Manogaran & Varadharajan, 2018). Additionally, it helps organizations in looking for the patterns in data batches, giving the organization the opportunity to learn more about the customers, which in turn helps in directing the marketing strategy of the organization to attracting customers thereby increasing the sales and reduction of the costs.

Additionally, through the neural network, it becomes easy to identify cases of fraud in the organization through the character and the image recognition techniques. Most of the techniques of the artificial neural network help in the identification of various inputs and aid in the processing of the hidden and the complicated data, with relationships which are non-linear. As a result, it becomes easy for the organization to position itself for the recognition of characters, for instance handwriting, used for the purposes of detecting fraud and the assessment of security nationwide. Therefore, the artificial neural network through the image recognition, it can be applied in different platforms such as detection of cancer in healthcare organizations and facilities and the facial recognition on the platforms of social media.

Data Warehousing

Data mining is usually enhanced by the existence of data warehousing. Data warehouses can be defined as the databases storing the structured data and for the processing of this data as well as preparation for the activities of mining. Some of the tasks handled in the data warehousing include the following: data sorting, classification of data, discarding of the unusable data and the metadata setting. Data warehousing can be defined as the technique used for the collection and the management of data from different sources for the provision of insights to the business which are meaningful. Rationale for the adoption of the data warehousing as a technique for data mining is to help organizations and business in making decisions based on this data. Therefore, the data warehouse aids decision making through the provision of architecture as well as tools for the systematic organization and understanding of this data from the different databases. Businesses can apply the technique of data warehousing for the purposes of telecommunication, that is in the promotions of products, making decisions about the sales and the decisions about the distribution (Arora & Gupta, 2017). From the data warehousing, all the decisions of the business are derived and made possible.

Association Rule Learning

This is a technique used for the identification of the relations which are interesting as well as interdependencies existing between various variables in the large databases. This technique helps in finding the patterns which tend to be hidden with regard to data, which may tend to unclear or even not obvious. It is mostly used in the machine learning. Application of the association rule learning in firms and organizations is to help in the detection of the existence of anomalies, a procedure used for searching the items or even events that tend not to correspond to the pattern which is familiar (Agrawal & Agrawal, 2015). For instance, in supermarkets, this technique can be used for inferring purchasing pattern of the customers.

Ways in which information governance support big data

Information governance alludes to an approach to the management of corporate data. It consists the implementation of multi-disciplinary structures, procedures, metrics control frameworks, policies and procedures at an enterprise level. It aims at increasing or expanding the value and quality of information while trying to reduce the costs and risks of holding it. It regulates information at the data level, making sure that the management and maintenance of accurate and high quality data through the implementation of appropriate systems and process (Smallwood, 2019).  Generally, information governance consists of the frameworks which includes rules and regulations, procedures and technology by which information and data is controlled and kept safe.

Information governance is very crucial in data integration and quality, security and privacy. A successful information project requires data which is integrated and can be trusted. Therefore, information governance aids in maintaining the quality of the stored information by availing data that suits this purpose. With information governance an enterprise can be able to come up with its data quality metrics to bring to board quality issues and also come with remediation plans. A growing concern with big data projects is security and maintenance of privacy. Big data has a lot of influence on the security and privacy issues of data. Big data affects privacy and security of data legally, ethically, politically and morally. This way information governance plays a big role in handling of big data.

Conclusion

Data mining is very vital while handling data. In organizations especially where large sets of data are handled data mining is inevitable. This is because it is significant in several ways. Data mining can be applied in marketing and aid in analyzing the characteristics of consumers. It helps to identify the trends in demand and supply. This will also help the marketer to be able to identify the factors that are affecting the purchasing power of the consumers. In schools and colleges, data mining can be used to help analyze and illustrate the behaviors of the learners. The teachers will be able to know which sector or which department in the education system needs some kind of improvement.

Data mining can also be used to solve natural disasters. Information is collected on the most frequent disasters for example floods and earthquakes and the places in which they are mostly witnessed in. This will help in the creation of disaster management systems to be able to curb the disasters which may occur. Therefore, every organization should make sure that it integrates the best big-data mining techniques so that it can be able to analyze the data well and give out results that are valid and reliable.

References

Agrawal, S., & Agrawal, J. (2015). Survey on anomaly detection using data mining techniques. Procedia Computer Science, 60, 708-713.

Arora, R. K., & Gupta, M. K. (2017). e-Governance using data warehousing and data mining. International Journal of Computer Applications, 975, 8887.

Benjelloun, F. Z., Lahcen, A. A., & Belfkih, S. (2015, March). An overview of big data opportunities, applications and tools. In 2015 Intelligent Systems and Computer Vision (ISCV) (pp. 1-6). IEEE.

Khan, Z., Anjum, A., Soomro, K., & Tahir, M. A. (2015). Towards cloud based big data analytics for smart future cities. Journal of Cloud Computing, 4(1), 2.

Liu, Z., Dev, H., Dontcheva, M., & Hoffman, M. (2016). Mining, pruning and visualizing frequent patterns for temporal event sequence analysis. In Proceedings of the IEEE VIS 2016 Workshop on Temporal & Sequential Event Analysis (pp. 2-4).

Mathan, K., Kumar, P. M., Panchatcharam, P., Manogaran, G., & Varadharajan, R. (2018). A novel Gini index decision tree data mining method with neural network classifiers for prediction of heart disease. Design Automation for Embedded Systems, 22(3), 225-242.

Milani, B. A., & Navimipour, N. J. (2017). A systematic literature review of the data replication techniques in the cloud environments. Big Data Research, 10, 1-7.

Rajeswari, S., Suthendran, K., & Rajakumar, K. (2017, June). A smart agricultural model by integrating IoT, mobile and cloud-based big data analytics. In 2017 International Conference on Intelligent Computing and Control (I2C2) (pp. 1-5). IEEE.

Sajana, T., Rani, C. S., & Narayana, K. V. (2016). A survey on clustering techniques for big data mining. Indian journal of Science and Technology, 9(3), 1-12.

Sari, A. (2015). A review of anomaly detection systems in cloud networks and survey of cloud security measures in cloud storage applications. Journal of Information Security, 6(02), 142.

Smallwood, R. F. (2019). Information governance: Concepts, strategies and best practices. John Wiley & Sons.

Wright, A. P., Wright, A. T., McCoy, A. B., & Sittig, D. F. (2015). The use of sequential pattern mining to predict next prescribed medications. Journal of biomedical informatics, 53, 73-80.

Yang, C., & Gidófalvi, G. (2018). Mining and visual exploration of closed contiguous sequential patterns in trajectories. International Journal of Geographical Information Science, 32(7), 1282-1304.

Yu, C., Li, Y., Xiang, H., & Zhang, M. (2018). Data mining-assisted short-term wind speed forecasting by wavelet packet decomposition and Elman neural network. Journal of Wind Engineering and Industrial Aerodynamics, 175, 136-143.

Yusof, N., Zurita-Milla, R., Kraak, M. J., & Retsios, B. (2016). Interactive discovery of sequential patterns in time series of wind data. International journal of geographical information science, 30(8), 1486-1506.