R Anomaly detection coding and short report (Please don't bid if you cant guarantee a mark of 60).

profiledara17
Anomalydetection.211.docx

Anomaly detection algorithm 2

ANOMALY DETECTION ALGORITHM

Name

Course

Institution

Word count: 2260

Part 1

Developed systems at times malfunction. Hence, there ought to be a manner in which one can detect the various means of anomalies that would require running programs that would indicate the area that is being affected at a particular point.

Hence, data science has made it easy for engineers to detect the primary component that is affected. Therefore, critical incidents are easy to see from the case investigated (Kagosh, 2021). For this paper, the main focus would be on household electricity consumption. This is owing to the data that was obtained from the Singapore States database.

Hence, through the data, one can take a keen look and establish the main issues. This is from the various mal incidences along the power lines. Technology has made it easy due to the synchronization of the power systems in place (Kagosh, 2021). It has made things to be easy and more advanced. Comment by VICTORIOUS: What word is this Comment by VICTORIOUS: What has technology made easy

The power supply companies can note where the issues are emerging from. Thus, they would either cut out the power system from the primary source. Move to the location that has been affected with ease (Kagosh, 2021). Thus, they would be best safeguarded in a manner that would be out to change society's outlook.

Electricity poses an immediate risk to the lives of people. Hence, having systems that can detect the arising issues is key to changing the framework that one is using or operating on. Thus, it is a system that would be best suited in making systems work and, therefore, be fruitful in all their working environment (Kagosh, 2021). Thus, minimizing the hazardous issues that are prone to arise and, therefore, being productive in the mode of operation.

S1 Select and apply essential analytical techniques

Anomaly detection looks into identifying the unforeseen items or information in relation to the data that was available.

The case study identified several issues that were arising owing to the task. This was based on the nineteen various unsupervised anomalies that were detected basing on the algorithms. There were more than ten different datasets that were looked into (Kagosh, 2021). This was based on multiple applications that were set on numerous domains. Comment by VICTORIOUS: Based

The anomaly detection was based on data science to identify the various data points in the broadest sense. This was based on looking at the multiple datasets that were showing mal behavioral trends. The study looked at the unusual trends witnessed on the electrical power lines that supplied households in Singapore.

Thus, the energy sector in Singapore was observed to be one that was marred with various datasets. Hence, it was based on the anomaly detected based on the more frequent application on unlabeled data. Thus, it was known as a basis for unsupervised mal detection. Comment by VICTORIOUS: remove Comment by VICTORIOUS: please rephase this sentence

The assumptions were that;

1. Anomalies would be based on the dataset though they rarely occur.

2. The system is tamper-proof; hence the errors would be present.

Thus, throughout the system, the errors are flagged. This is through the identification of the various error points. Therefore, it leads to a form of deviation from the standard statistical properties based on the distribution. These include attributes such as mean, median, mode, and quantiles. Comment by VICTORIOUS: there is too much useage of “Thus”

Anomaly detection would be the primary step in data mining. It would be based on the various data points, events, and observations that would not be similar to the dataset that is of normal behavior. Thus, the unusual data can be based on the events that are out to check on critical incidents, for instance, the technical failures or potential threats. This would be noted by the forecast presented by the data about the graphical representation. Thus, they would be much affected based on the ways of analysis.

Hence, the use of R would be best suited in analyzing the available data. Big data would be based on analysis that would look into the data science lifecycle of modern data science. Hence, it would change the outlook of the contemporary programing language. Comment by VICTORIOUS: can’t use hence to start a paragraph

The data that was used was from Singapore moreover. It was based on the critical outlook and space provided based on the open license portal that allowed access to the data. Based on the Energy Market Authority, frequency is based on the year regime and the year outlook.

Data analysis from the forecast would then be used to come up with the best means of resolving the issues present. Thus, it would be best suited to conform to the outliers and the basis of systems that would be looked into. This would be based on the statistical elements that were based on the various methods that were identified. Comment by VICTORIOUS: please rephase the sentence

Utilizing the system would be one of the best techniques used in resolving the stalemates that are arising. This would be based on systems present due to the means of analysis and forecast of studies based on frameworks. They would be forming a baseline and also a means of corroboration. This would be based on understanding what is happening or the issues that are being transformed. Comment by VICTORIOUS: what system are you refering to. What type of systems are present for forcast analysis

S2 Conduct pre-processing

Time Series Anomaly Detection would be the first step of detecting the issues that are arising. Thus, basing on the duration of the events, one would be able to come up with the best possibilities. They would be the ones to skew the anomalies that are being brought forward (Cuadra-Sánchez and Aracil, 2015). This is based on the three steps that would be taking place.

Decompose the time series that would be an underlying variablesunderlying variable based on the trends, seasonality, and residue. Creating an upper and lower threshold value (Cuadra-Sánchez and Aracil, 2015). Identification of the data points would be based on the outside thresholds of the issues that are raised.

Visualization of the anomalies would be based on the anomalies plot function

When the algorithm runs, the average peaks are identified and indicated as above. This is based on the automated get_time_scale_template. Thus, the logical frequencies are placed at various platforms. This would be looking into the trend based on the scale of data (Cuadra-Sánchez and Aracil, 2015). It would be the first step to offer the data analysis to detect the issues.

The red points which are shown in the figures above are the areas where issues of likely anomaly are. They are prone to have anomalies that would need to be looked into. This is based on the various circumstances outlined (Mehrotra, Mohan, and Huang, 2017). They would be based on flagging the red areas and noting where the place would be having issues.

Hence, from the analysis, a scale of one day would look into the data covered in a day. Based on the legend, it would be the outlook of a data point. This would look into the frequency based on the seven days or a week (Mehrotra, Mohan, and Huang, 2017). Therefore, the analysis was prepared under a three months framework of which was observed to have a specific yield and outlook when it came to the basics of handling data.

Therefore, when we are decreasing alpha, we are increasing the bands. Hence, making the outliers fall out of place (Mehrotra, Mohan, and Huang, 2017). Therefore, the probability and presence of outliers wouldn't be possible based on the bands set out. They would be the basis of the twice big in size modes.

https://editor.analyticsvidhya.com/uploads/56330alpha1.png

Max Anoms

Maximization of the anoms would be presented by the maximum control percentage of the information obtained from the anomaly. Thus, it would be grouped on the adjusted alpha = 0.3 of a much-observed outlier (Mehrotra, Mohan, and Huang, 2017). Therefore, a comparison would be based on the max atoms that are = 0.2, twenty percent of which is allowed, and the 0.05 equivalent of five percent is permitted.

Thus, the electrical faults would occur within a certain threshold. This implies that the engineers have accepted a certain margin of error on the household power lines. Nevertheless, when they go overboard, an alarm would be signaled (Mehrotra, Mohan, and Huang, 2017). It would be the first step into coming up with solutions that would be inclined to changing the outlook.

Adjusting Local Parameters

Local parameter adjustments are evaluated by tweaking the in-function parameters. Thus, we can adjust the trends based on two weeks (Mehrotra, Mohan, and Huang, 2017). Therefore, it looks to be some of a misfit basing on the trendline.

max anams

Figure: Forecast of 20% anomalies

anams

Figure: Forecast of 5% anomalies

From the two graphs above looking at the 5% and 20%. One can observe that issues are based on a broader scale as opposed to a lesser scale (Mehrotra, Mohan, and Huang, 2017). This would imply that they would be keen on making sure that the anomalies are easily detected. Thus, they would be based on having the best spirits in place.

alpha

Figure: anomalies at alpahe0.05

Using the ‘timetk’ package

The tool is used in the prediction of the trend basing on time series.

Interactive Anomaly Visualization

Thus, the Visualization would be based on the function inclined to the timetk’s plot_anomaly_diagnostics(). Hence, it is a system that would be tweaked basing on the unusual parameters based on the fly mechanisms (Mehrotra, Mohan, and Huang, 2017).

timetk

Figure : Anomaly diagnostic trends Comment by VICTORIOUS: Label all your figure for easy referencing

The anomalies were noted to be on the 450 in the year 2005. Nevertheless, the engineers put up systems that would lead to changes in how the system was working. Thus, it enabled them to develop a clear outlook (Mehrotra, Mohan, and Huang, 2017). In the year 2020, the faults were witnessed at 900. Thus, it was a higher value as opposed to the previous values.

Interactive Anomaly Detection

To find the exact data points that are anomalies, we use the tk_anomaly_diagnostics() function.

Now, we can extract the actual data points, which are anomalies. For that, the following code can be run.

Adjusting Alpha and Max Anoms

The alpha and max_anoms are the two parameters that control the anomalies() function. H

Alpha

We can adjust alpha, which is set to 0.05 by default. By default, the bands cover the outside of the range.

Conclusion

The approach has been successfully deployed in data mining. Developed systems at times malfunction. Hence, there ought to be a manner in which one can detect the various means of anomalies that would require running programs that would indicate the area that is being affected at a particular point. Hence, through the data, one can take a keen look and establish the main issues. This is from the various mal incidences along the power lines. The power supply companies can note where the problems are emerging from. Thus, they would either cut out the power system from the primary source—society's outlook.

Electricity poses an immediate risk to the lives of people. Hence, having systems that can detect the arising issues is key to changing the framework that one is using or operating on. Thus, it is a system that would be best suited in making systems work and, therefore, be fruitful in all their working environment. Anomaly detection looks into identifying the unforeseen items or information in relation to the data that was available. More than ten different datasets were examined. This was based on multiple applications set on various domains. In the broadest sense, anomaly detection was based on data science to identify the different data points. Power lines that supplied households in Singapore.

Thus, the energy sector in Singapore was observed to be one that was marred with various datasets. Hence, it was based on the anomaly detected based on the more frequent application on unlabeled data. Thus, it was known as a basis for unsupervised mal detection. Therefore, throughout the system, the errors are flagged. This is through the identification of the various error points. This would be noted by the forecast presented by the data in relation to the graphical representation. Thus, they would be much affected based on the ways of analysis.

Utilizing the system would be one of the best techniques used in resolving the stalemates that are arising. Time Series Anomaly Detection would be the first step of detecting the issues that are occurring. Thus, basing on the duration of the events, one would be able to come up with the best possibilities. When the algorithm runs, the average peaks are identified and indicated as above. This is based on the automated get_time_scale_template. Thus, the logical frequencies are placed at various platforms. Hence, from the analysis, a scale of one day would look into the data covered in a day. Based on the legend, it would be the outlook of a data point. This would look into the frequency based on the seven days or a week. Maximization of the anoms would be presented by the maximum control percentage of the information obtained from the anomaly.

Thus, the electrical faults would occur within a certain threshold. This implies that the engineers have accepted a certain margin of error on the household power lines, from the two graphs above looking at the 5% and 20%. One can observe that issues are based detected on a broader scale as opposed to a lesser scale. Thus, the visualization would be found on the function inclined to the timetk’s plot_anomaly_diagnostics(). The anomalies were noted to be on the 450 in the year 2005. Nevertheless, the engineers put up systems that would lead to changes in how the system was working. The alpha and max_anoms are the two parameters that control the anomalies() function. We can adjust alpha, which is set to 0.05 by default. By default, the bands cover the outside of the range.

References

Cuadra-Sánchez, A. and Aracil, J. (2015). Traffic Anomaly Detection. Elsevier Science.

Dunning, T. and Friedman, E. (2014). Practical Machine Learning: A New Look at Anomaly Detection. O'Reilly Media, Inc.

Kagosh. (2021). A Case Study to Detect Anomalies in Time Series Using Anomalize Package In R. Retrieved from https://www.analyticsvidhya.com/blog/2020/12/a-case-study-to-detect-anomalies-in-time-series-using-anomalize-package-in-r/. Data accessed May 6, 2021.

Mehrotra, K., Mohan, C. and Huang, H. (2017). Anomaly Detection Principles and Algorithms. Springer.

Appendix: R Studio script

‘Anomalize’ package

Visualize the Anomalies

Extracting the Anomalous Data Points