literature review on "Measures to Reduce Data Dump in Minimal Time"

profileEmir7
RevisedObjectiveSummary6.docx

6

Measures to Reduce Data Dump in Minimal Time

Student Name

Department, Institution

Course Title

Instructor Name

Due Date

Measures to Reduce Data Dump in Minimal Time

Note Sheet 1: From Ad-Hoc Data Analytics to DataOps

This research was a combined effort of not less than five authors, including David Issa and Anas Dakkak. This study was used to explain the whole concept of DataOps so that real-time streaming could be well understood. The authors collected the data for this study in 2020 (Munappy et al., 2020). The data for the research was collected in Stockholm, Sweden. The empirical method that was used for the collection of data was compliant with a number of questions to be answered by the respondents in interviews. The study's findings were very informative in revealing the DataOps and all its relevant concepts, as per the patient responses. The study also helped design a strategy for the evolution of data handling methods when dealing with big data. The authors then proceed to define DataOps in a way that is beneficial for all institutions dependent on big data. DataOps has the unique capability of minimizing the time for the end-to-end cycle from identification to developing insight.

. The study encompassed several faces of employees from all departments, including engineers and data analytics. The researchers took the review of the Multi-Vocational Literature as the initial step of the research methodology. The researchers then went ahead to determine the data analytics process that Ericsson employees were compliant with. The data collection results in the research were significant in guiding the interview. The time needed to deliver real-time insights is incredibly reduced, right from collecting ad-hoc, to the analytics and monitoring stage, using the capabilities of the maturity levels, also known as evolution stages. Among the crucial details revealed in this study include Orchestration, automation, and collaboration, which helped to enhance real-time streaming insights from DataOps.

Note Sheet 2: A Roadmap Towards Big Data Opportunities, Emerging Issues and Hadoop as a Solution

In the case of this study, Rida Qayyum is the single author who contributed to the whole publication from Pakistan from the University Sialkot in the year 2020. The researcher undertook the study to reveal the concepts of big data and how big data can be transferred from one storage to another easily (Qayyum, 2020). The data for the study was obtained from secondary sources, including journals, websites, and trusted sources like the IEEE. The author selected other texts that would provide information to boost understanding of big data. The research presented the concepts and gave a solution for those who handle big data. The study helped the researchers to identify several aspects that big data encompasses, including its features, types, opportunities, and emerging issues like large storage spaces and complexity in the data structures. The overall performance of hard disks is rising, whereas the rate of disk transfer is not increasing at an equal rate. Of the three components that Hadoop tends to bring along, including Hadoop Distributed File System (HDFS), kernel, and MapReduce, the first is the one responsible for processing the data in a parallel manner and distributing it throughout the file system.

The address of the researcher who accomplished this study is the Government College Women's University Sialkot, 51040, Pakistan, and the name of the researcher was Rida Qayyum, from the Department of Computer Science. The researchers behind this study about big data were aimed at revealing the different aspects that are unignorably when in that concept. The storage of large amounts of data as well as the fast speed of processing such data, was also a key concern for completing this study. The data for the study was conducted in 2020 because it was obtained from secondary literature. The study was conducted in Sialkot, 51040, Pakistan. The data was collected from a series of literature on handling big data. The resources were used to reveal the various features of big data that make it special in today’s organizations. This particular research was performed to show what characteristics, types, opportunities, and emergent issues the concept of big data presents. Included in the research was the solution of Hadoop to help in handling big proportions of data.

Note Sheet 3: DOD-ETL: distributed on-demand ETL for near real-time business intelligence

The article information revealed that it was written by several authors whose first names are as follows; Cunha, Oliveira Pereira, and Machado, and one of the researchers was reliable and affiliated to the Department of Computer Science, Universidade Federal de Minas Gerais, Belo Horizonte, Brazil. To establish and understand the framework with which Business Intelligence (BI) can facilitate a near real-time approach (Machado et al., 2019). The study was completed to review (BI) and the process of Extract Transform Load (ETL). The researchers wanted to develop an ETL solution near real-time and implement it using Demonstrated on Demand (DOD) ETL. The study proposed the DOD-ETL as a technology that can achieve near real-time ETL through a number of multiple strategies. The study finally compares DOD-ETL with other related works. Data for the study was collected in 2019 and published in Open Access. The research was conducted in a higher learning institution in Brazil known as the Universidade Federal de Minas Gerais, Belo Horizonte. The data was collected from secondary sources to solve the research problems, including integration of data sources, mastering data overheads, degradation of performance, and backing up data. The study also covered several publications, other publications which covered the frameworks of Stream Processing that help to solve real-time ETL. The experiments in the study revealed that DOD-ETL significantly increases the speed of Spark. The DOD-ETL contains the In-Memory Table Updater data dump from Message Queue but is still able to process data at very high rates than the baseline. The study also showed that DOD-ETL customizations have no negative impact on the fault tolerance and scalability of Spark Streaming. In other words, DOD-ETL techniques and strategies help reduce ETL's run time. This model outperforms a modern framework for Stream Processing.

An individual conducted the study from Universidade Federal de Minas Gerais, Belo Horizonte, Brazil, named Oliveira, alongside others like Machado, Cunha, and Pereira, and ultimately used Open Access to publish it in 2019. The research was conducted to discover how people establish near real-time BI.]. Some of the problems included the integration of data sources, backup, and mastering data overheads. The researchers also considered publications that contain information on Stream Processing. The baseline has a higher rate of data processing than the DOD-ETL, and DOD-ETL can perform this well despite the fact that it also encompasses an In-Memory Table Updater data dump originating from Message Queue. The study also showed that DOD-ETL customizations have no negative effect on either the tolerance of fault or scalability of Spark Streaming. DOD-ETL techniques and strategies help to reduce the run time of ETL. This model outperforms a modern framework for Stream Processing.

References

Machado, G. V., Cunha, Í., Pereira, A., & Oliveira, L. B. (2019). DOD-ETL: distributed on-demand ETL for near real-time business intelligence . Journal of Internet Services and Applications, 10(1), 1-15. https://link.springer.com/article/10.1186/s13174-019-0121-z

Munappy, A. R., Mattos, D. I., Bosch, J., Olsson, H. H., & Dakkak, A. (2020, June). From ad-hoc data analytics to dataops. In  Proceedings of the International Conference on Software and System Processes (pp. 165-174). https://research.chalmers.se/publication/521464/file/521464_Fulltext.pdf

Qayyum, R. (2020). A roadmap towards big data opportunities, emerging issues and Hadoop as a solution. Rida Qayyum." A Roadmap Towards Big Data Opportunities, Emerging Issues and Hadoop as a Solution", I nternational Journal of Education and Management Engineering (IJEME), 10(4), 8-17. https://j.mecs-press.net/ijeme/ijeme-v10-n4/IJEME-V10-N4-2.pdf