1600 word min due 12/21. 4 Apa format scholarly sources
Measures of Dispersion and Central Tendency
In statistics, measures of dispersion describe the scattering of data. It gives a clear understanding of how a given dataset can be dispersed as well as explaining the distribution of the given data set. The various measures of dispersion include the standard deviation, variance and inter quarterly range. Standard deviation is considered the most commonly used measure of dispersion. On the other hand, central tendency measures attempt to describe the central position of a given single data set. According to (Dimitriadis et al., 2019), the various central tendency measures include mean, median and mode. This paper will examine the differences between dispersion and deviation and give a discussion of how they compare to measures of central tendency. It will also give a clear definition of standard deviation, index of dispersion and the range, as well as explaining why they are not as clear as measures of central tendency
Measures of dispersion can be used to give a satisfying description of a data set compared to how the same data set could be described using central trendy measures. For instance, when finding the mean of two different data sets, the outcome could be the same. Therefore, the difference between the two data sets may be difficult to identify to find their means. Measures of dispersion can therefore be used to find the variability between the two different data sets. Through analysis of central tendency, it is possible to identify whether a data set has a weak or a strong central tendency based on its dispersion measures.
Dispersion can also be described as the statistic of distribution that explains how a given data set's values are dispersed or spread out around the central tendency. When in a given data set, the values are broad, the values in the set are said to be scattered. On the other hand, when a given data set has small numbers, the values in the set are said to be squeezed. For example, a data set containing values such as 1.2,3,4,5 is said to be squeezed contrary to the data set containing 10, 253, 470,600,1050, which is described as scattered. It is also possible to determine the spread or distribution of a given data set using a range of descriptive statistics involving measures of dispersion.
On the contrary, it is not possible to determine the distribution of a data set using central tendency measures. The spread of a data set can always be illustrated using figures as well as using graphical methods. The graphical method of illustrating the spread of a data set is simple and easy to understand. However, both the measures of central tendency and dispersal measures give a similar mean for a similar data set. That is, they use the same method in calculating the mean of a data set.
Different data sets possess different variations of data that can only be measured using standard deviation. Standard deviation is a measure of dispersion that gives the numerical measure of the variability in a given data set (Roitman et al., 2017). Again, by finding the value of the standard deviation of a given data set, it is possible to determine how closely the values are concentrated near the mean as well as how widely the values are spread far from the mean. The standard deviation is always positive and can no longer be negative. When the values are concentrated towards the mean, the standard deviations becomes small, while on the other hand, it is described as large when the values in a data set are more spread out. The standard deviation is calculated based on the deviations. For example, if x is a number, then "x-mean" is the deviation. Lastly, the standard deviation may not be of help when the distribution is skewed. This limitation is a result of the two different types of spread present in skewed distribution; hence when the distribution is skewed, an individual employs the use of mean,
Index of dispersion refers to the ratio of the standard deviation to the mean of a given data set. The index of dispersion normally relates to the coefficient of variance and dispersion. It can only be calculated when the values in a data set can be measured on a ratio scale or when the data set values can be counted. The dispersion index can assume three types of distributions for the count data. That is the gamma, poison and geometric distributions. The dispersion of index is different for each distribution that the count data assumes.
The range is another simple measure of the dispersion of a distribution. It is commonly used for normal data or partially ordered data sets. According to (Ali et al., 2019), the range refers to the difference between the largest and the smallest value in a data set. Even though the range of data is very simple to calculate, it is very sensitive to outliers since it does not make use of all the values in a data set hence its main disadvantage. In finding the range of a data set, an individual needs to first identify the maximum and the minimum value in a data set; So, it is a more informative measure of dispersion.
In conclusion, Measures of dispersion and central trendy are extremely important techniques when dealing with data. The measures of dispersion mainly describe how data is spread, while the measures of dispersion estimate normal values of a data set. They tend to give clear descriptive information regarding values in a data set, hence improving our understanding concerning different data types of data in a data set.
Ali, Z., Bhaskar, S.B., Sudheesh, K. (2019). Descriptive Statistics: Measures of central tendency, dispersion, correlation and regression. Airway, 2(3), 120.
Dimitriadis, T. Patton, A. J., And Schmidt, P. (2019). Testing Forecast Rationality for Measures of central Tendency. ArXiv preprint arXiv: 1910.12545
Roitman, H., Erera, S., and Weiner, B. (2017, October). Robust Standard deviation estimate for query performance prediction. In Proceedings of the ACM SIGIR International Conference on Theory of Information Retrieval (pp. 245-248).