literature review on "Measures to Reduce Data Dump in Minimal Time"
4
Pogue (2019) presented critical insights regarding how people may “cleanse” data. When humans need to checkup on their cars, finances, or health, they usually find mechanics, accountants, or doctors. However, their digital lives are usually left unattended as many people can stay for years of disregarding their data, particularly the one stored digitally. Storage devices such as hard drives are bound to be damaged, even manufacturers are aware of that fact and indicate it in the manual’s fine print as “M.T.B.F.,” which stands for “mean time between failures.” Therefore, if an individual’s drive stores only copies of photos and other files, the implication is that all that data is on the verge of being lost. Therefore, there is a need for a continuous and automatic back up system. According to Munappy et al. (2020), only 6 percent of people have set up a backup system for their digital files. It is important to consider offsite backups like the cloud, which is immune to such natural catastrophes as fiendish burglars, floods, and fire. Pogue (2019) clearly points to the issue of data dump whereby people have not been prudent enough to backup data in safer locations that cannot be easily compromised.
Mitigation
Reduction of data dump can be seamlessly achieved when appropriate measures are taken into account, which includes protection and preservation of precious data that has already been cleaned up. Data users need to secure their data through backup, which should be accompanied by periodic and regular checks and tests. Data dump may be prevented through appropriate investments in backup solutions and creation of a backup schedule. According to a study by Machado et al. (2019), 60 percent of people utilize a backup solution to cushion themselves from potential data loss, but the data is usually compromised due to lack of updated backup solution or the backup is faulty. Apart from data security, data dump may be significantly lowered through proper data organization. This includes the formation of a neat folder structure that is properly organized and file names are logically arranged, which may help in identifying important record and documents faster and more efficiently. Organization of data prevents potential data loss through accidentally deleting files. Data users may need to create a “dump” folder aimed at storing all uncertain data. Reducing data dump can also be achieved through safe cleaning and protection of hardware. According to Qayyum, (2020), dust is one of the components that negatively influence the cooling system of a computer, thus causing the device to overheat. This may lead to crashing of the operating system and, consequently, potential data loss. It is advised that data users and device owners should utilize the necessary cleaning products for electronic devices and avoid drinking or eating near the computer.
Recommendation
In most cases, people do not always have control over the type or format of data they import from an external source of data like a web page, text file, or even a database. Therefore, it is recommended that before data analysis, the data is cleaned up. The data cleaning process may leverage such software as Microsoft Excel and Spell Checker, which helps in precisely formatting the data and clean up misspelled words in every column containing descriptions or comments. The Remove Duplicates dialog box may help in removing duplicate rows. Another recommendation is the need to create a backup copy of the original data in a different workbook. Data users should be keen towards ensuring that the data is in a tabular format of columns and rows with no blank rows within range, all rows and columns visible, and similar data in every column. This can be easily achieved through the use of an Excel table. In dealing with the problem of data dump, it is important to first execute tasks that do not need manipulation of columns like checking spelling or the use of Find and Replace dialogue box. This step should be followed by tasks that do not need the columns to be manipulated. Column manipulation includes such actions as inserting a new column next to the preceding one that requires to be cleaned. A formula that is capable of transforming the data at the top of new column also needs to be added. By end of the process, the new data result is “clean,” and it is interesting to note that the process takes very minimal time. The clean data should always be backed up, particularly in the cloud to prevent data loss, which may pose huge negative implications to the data user. When encountering data loss, some people tend to attempt data recovery through own at-home remedies, which may cause more harm than good. Therefore, it is recommended that an individual should immediately contact a data recovery professional for assistance.
References
Machado, G. V., Cunha, Í., Pereira, A., & Oliveira, L. B. (2019). DOD-ETL: distributed on-demand ETL for near real-time business intelligence. Journal of Internet Services and Applications, 10(1), 1-15. https://doi.org/10.1186/s13174-019-0121-z
Munappy, A. R., Mattos, D. I., Bosch, J., Olsson, H. H., & Dakkak, A. (2020, June). From ad-hoc data analytics to dataops. In Proceedings of the International Conference on Software and System Processes (pp. 165-174). https://doi.org/10.1145/3379177.3388909
Pogue, D. (2019). How to do a data “cleanse.” New York Times. https://www.nytimes.com/2019/02/01/smarter-living/how-to-do-a-data-cleanse.html
Qayyum, R. (2020). A roadmap towards big data opportunities, emerging issues and hadoop as a solution. Rida Qayyum." A Roadmap Towards Big Data Opportunities, Emerging Issues and Hadoop as a Solution", International Journal of Education and Management Engineering (IJEME), 10(4), 8-17. https://j.mecs-press.net/ijeme/ijeme-v10-n4/IJEME-V10-N4-2.pdf