Math301 (NEED QUICK TURNAROUND)

profilencstudent1457
math301_ip4.docx

Running head: Gym Statistics 2

Silver’s Gym Driven Statistics

Deirdre Martin

MATH301-1501B-01 Data Driven Statistics

Colorado Technical University

March 17, 2015

Descriptive Statistics

Descriptive

Fat

Weight

Mean

18.94

178.92

Median

19.00

176.5

Range

45.1

244.65

Standard Deviation

7.75

29.39

Interpretation

The mean is found by dividing the sum of values by how ever many the number of observations. It is considered a good estimate for predicting subsequent data points (Anthony Atiknson, 2000). So, 18.94 Kg is the average body fat and 178.92 is the average body weight. This implies that we expect most values representing body fat and weight to be around the neighborhood of 18.94 and 178.92 respectively. Values far from the two are taken to be outliers or abnormalities.

Median refers to the middle value of a given data set if it has odd number of values, or the average of the two middle values of a set of data with even number of values. It is helpful when data is to be separated into two equal sizes.

Then, 19 and 176.5 would be the median for body fat and weight respectively. This means that the center most values of fat and weight are 19 and 176.5. If we divide the data across these values, we should get exact two equal halves.

Range refers to the difference between the largest and the smallest value in the data set. It establishes how well the central tendency describes the data. A large range is not indicative of the central tendency as a description of the data. The range for body fat would be 45.1 and weight would be 244.65. The range for fat is a bit small; indicating that the central tendency is a better description of the data. Weight values with a bit large range indicatives the central tendency is a poor measure (Belseley, 2005).

Standard deviation show how close the set of data is to the average values, but this will just give you an idea. Small standard deviations will show the data is tightly grouped and thus precise. Data sets with large standard deviations have data spread out over a wide range of values. The body fat and weight have standard deviations of 7.75 and 29.39 respectively. The body fat shows a data that is much closer to the mean and thus a better estimate while the body weight has a higher standard deviation indicating a poor estimate.

>Importance of the mean and median

In statistics, the mean and median are two ways of giving the center of a data and provide two different interpretations of the same especially when there is presence of outliers. It is used as a measure of central tendency. It is flexible as it can be used with both continuous and discrete data. It includes all the values in the data set as part of the calculation. It is also the only measure where the summation of deviations of each value from it is always zero of central tendency. However, it is easily affected by outliers. The mean is important since it provides an idea of where most values are concentrated and may be used as a platform for setting the hypothesized mean in hypothesis testing (Carlberg, 2014).

The median is the middle point or score of a given set of data that has been arranged in order of importance. It has the advantage in that it is not affected by outliers and biased data and hence offers a good measure of tendency. Its properties make it favorable for hypothesis testing especially when dealing with ordered statistics.

>The best measure for determining the where the salary fall would be the median. This is because the median is preferred when dealing with skewed data or ordinal data. In our case, salary is ordinal (highest to lowest), thus median becomes the best the measure. The mean on the other hand, is preferred when the underlying distribution is continuous and symmetrical. In our case, the salary is discrete and thus mean is not the best measure to use.

>The standard deviation and range serve as measures of dispersion. Despite the mean, the standard deviation is important because, it provides a tremendous amount of information. Its important because the distribution is out over a broad range or bunched up closely to the mean. Accuracy of prediction increases with decreasing values of standard deviation. The range represent the difference between the highest value and the lowest value. It has the same function as the standard deviation.

>Hypothesis:

Null Hypothesis: (the mean body fat in men is 20%)

Vs.

Alternative Hypothesis: (the mean body fat in men is not 20%)

We shall use a z test for the means because the sample size is large (252>30). It also represents a comparison of one sample to a specific number. We are only comparing the mean for the body fat in this particular case.

At 0.05 the z-table is -1.96

Since z<-1.96, there is sufficient evidence for rejecting the null hypothesis. Therefore we conclude that the average body fat was not 20%.

The best advice would to report to the boss that the claim that the average body fat is 20% is not justified at 5% level of significance.

References

Anthony Atiknson, M. R. (2000). Robust Diagonistic Regression Analysis. New York: Springer Science & Business Media. Belseley, D. A. (2005). Regression Diagonistics: Identifying Influential Data and Sources. Chicago: John Wiley & sons. Carlberg, C. (2014). Statistical Analysis: Microsoft Excel 2013. New York: Que Publishing.