Final Project: You will assemble your final report as follows:

profilesam6990
assignment3bigdata.docx

1

Running head: BIG DATA

6

BIG DATA

MIT 681 : MSIT capstone

Student name: sunil patel

Student number: 15T1FG68

Professor name: Mark o’connell

Introduction

With the invention of new technology data management has been viewed as one aspect where technology should be implemented. Despite the existence of traditional approaches and methodologies in data storage and management the way data is handled differs from one organization to the other. Big data analysis and storage remains a challenge to many, however much has been achieved and individuals are reaping from the changes which come with new technology techniques such as cloud computing and many other data management technologies (Curry 2016). With individuals having data sources such as the Amazon websites and the yahoo among others proper data storage and accessing tools are very important. It is very important to appreciate the fact that for one to attain high integrity levels in big data management and analysis there are several factors to put into consideration.

These factors include but are not limited to: Flexibility The method of analyzing and storing individual’s data should be able to meet the ever changing technological changes and demands. Maintainability The mechanism should be maintainable by the users Accountability and reliability Users should be able to trust the chosen system and rely on it without fear of losing their data and any other critical information.

The term big data has brought a lot of influences in the data storage, while it is used to refer to a collection of several data from different sources majority are yet to accept the fact that this has been there for ancients with traditional methods applied to access and store the relevant data. Today due to raise in technology security threats have diversified something which calls for users to be alert and practice ethical secure methods when using big data (Gandomi, & Haider 2015). Likewise users need to put into consideration the necessity of the data and ensure they observe the set policies so as to ensure that they do not go against the laid procedures and regulations while using the data. Additionally there are several big data approaches all of which have certain capabilities in managing data.

To be able to succeed in big data approach individuals ought to put into consideration the security concerns and the responses each model offer. Data access, management and storage is a very critical aspect therefore necessary frameworks should be considered before anything. It is very important to establish the programming languages used in the data storage systems; the language extensions and query language since this are key determinants of how successful the chosen approach will be. A clear programming language which is able to accommodate the technological demands remains crucial; the extension languages should also be reliable and manageable. Any storage approach which is not flexible cannot adapt changes this demands for a new data storage system whenever there is a change in technology and data type (Li, Gai, Qiu, Qiu & Zhao (2017).

Good and reliable big data architecture should be able to support all the three major components of a generic architecture which are, extraction, and transforming. The chosen design should be able to allow users to extract data for use following the necessary procedures and observing all the necessary security precautions. Also the design should be able to transform the data in a certain way such that it can be integrated in the system and finally enable users to be able to load the information (Guirguis, Pareek & Wilkes 2016). In order to establish the correct database for use then it is important for individuals to consider the following;

Online analytical processing (OLAP)

Some architectures do support an online analytical processing something which limits users who do not up constant network supply. If network connectivity is a challenge then it is very advisable to avoid this architecture.

Query language

In this one needs to establish the language they are relying on to access or manage their data. This can either be procedural, language extensions or descriptive language.

Cloud computing option

There are architectures which support cloud computing while others do not. When making a decision on which form of database to use it is good to consider this option since cloud computing offers more reliable and secure data management techniques (Coronel & Morris 2016).

Scalability

Individuals need to consider whether their choice will be able to meet the computing demand, data size and organization requirements. Fault tolerance. The system should be able to recover from any fault in an accountable and transparent matter. Programming language This establishes the programming language and design language supported by the system. The choice of language is very important since different people will come into contact with the system in different occasions, simple and understandable language is very important (Abadi, Bajda-Pawlikowski, Abouzied, & Silberschatz 2016).

Type of database.

In order to have an effective database and management system NOSQL database is the best system since it offers quality storage and retrieval services. Additionally it has high fault tolerance and does not necessarily rely on distributed file systems to function appropriately.like wise the system is able to accommodate new technology requirements and perform within the specified time (Lourenço Cabral, Carreiro Vieira, & Bernardino 2015). Since this database supports data in forms of documents it offers dynamic schemas, auto sharing capabilities and replication abilities making it a very reliable database system to be used in a business and even educational set up.

References

Abadi, Bajda-Pawlikowski, Abouzied, & Silberschatz, (2016). U.S. Patent No. 9,495,427. Washington, DC: U.S. Patent and Trademark Office. Curry (2016). The big data value chain: definitions, concepts, and theoretical approaches. In New horizons for a data-driven economy (pp. 29-37).

Cengage Learning. Gandomi & Haider (2015). Beyond the hype: Big data concepts, methods, and analytics. International journal of information management, 35(2), 137-144.

Guirguis, Pareek & Wilkes (2016). U.S. Patent No. 9,298,878. Washington, DC: U.S. Patent and Trademark Office.

Li, Gai, Qiu, Qiu & Zhao (2017). Intelligent cryptography approach for secure distributed big data storage in cloud computing. Information Sciences, 387, 103-115.

Lourenço Cabral, Carreiro Vieira, & Bernardino (2015). Choosing the right NoSQL database for the job: a quality attributes evaluation. Journal of Big Data, 2(1), 18.

Springer, Cham. Coronel & Morris (2016). Database systems: design, implementation, & management.