Just need correction on statistics paper-Deliverable 1 - Descriptive Statistics
Inferential Statistics and Analytics
Deliverable 1
10/18/17
Deliverable 01 Worksheet
1. Introduce your scenario and data set.
· Provide a brief overview of the scenario you are given and describe the data set.
· Describe how you will be analyzing the data set.
· Classify the variables in your data set.
· Which variables are quantitative/qualitative?
· If it is a quantitative variable, is it discrete or continuous?
· Describe the level of measurement for each variable included in the data set (nominal, ordinal, interval, ratio).
Answer and Explanation:
Provide a brief overview of the scenario you are given above and the data set that you will be analyzing.
A client is interested in the salary distributions of jobs in the state of Minnesota that range from $40,000 to $120,000 per year. As a Business Analyst my boss asked me to research and analyzes the different salary distributions. I was given a data sheet that contained a listing of 364 job titles and the annual salary for each job title. I will be classifying the variables for option #1.
Classify the variables in your data set.
· Which variables are quantitative/ qualitative
Qualitative variables are the jobs listed while the quantitative variables are the salaries.
· Which variables are continuous/discrete
Discrete variables are both the salary amounts and jobs. No continuous variables present.
· Describe the level of measurement for each variable included in your data set
There is a total of 364 different jobs listed in the data set. The salary listed in the data set represented by the total U.S dollar annually that each job pays on average.
Jobs are nominal variables since they have no numerical values. Individuals can be categorized based on the job occupation.
Salaries are classified as interval variables because we can find the interval between the salaries of different jobs.
2. Discuss the importance of the Measures of Center.
· Name and describe each measure of center.
· Discuss the advantages and/or disadvantages of each.
Answer and Explanation:
A measure of central tendency is a summary of measure that tries to describe a whole set of data with a single value that represent the middle of its distribution. They include:
Mode: This is the most recurring (most commonly repeating) value in a data set.
Advantages: It can be found both numerically and categorical (non-numerical) data, easy to understand and simple to calculate, not affected by extreme large or small values, can be identified by inspection in ungroup data and discreet frequency distribution, useful in qualitative data, can be computed in an open-end frequency table and can also be located graphically.
Disadvantages: It may not reflect the center of the data very well, there may exist more than one mode in a data set limiting the ability to describe the data using the mode, not based on all values, not capable of further mathematical treatment and continuous data may lack a mode at all.
Mean: Is the average of the data set. It is computed by summing up all the values in a data set then dividing it by the number of observations.
Advantages: it can be used for both continuous and discrete numeric data sets, easy to understand, used with interval level data and includes all the values in a data set.
Disadvantages: cannot be calculated for categorical data since the values cannot be summed up. It is influenced by the presence of outliers and skewed distributions because it includes every value in the data set.
Median: Is the middle value in a distribution when the data set is arranged in an ascending or descending order. Thus, it divides the data set into two equal halves.
Advantages: It is less affected by the presence of outliers and skewed data sets unlike the mean thus preferable for unsymmetrical distributions, easy to compute and comprehend and can be determined for ratio, interval and ordinal measures of scale.
Disadvantages: it cannot be identified for categorical nominal data sets because it cannot be logically ordered, does not consider the precise values of each observation and hence do not use all information available in a data set, not amenable for further mathematical calculations and cannot be expressed in terms of individual medians of o poll of observation of two or more groups.
3. Discuss the importance of the Measures of Variation.
· Name and describe each measure of variation.
· Discuss the advantages and/or disadvantages of each.
Answer and Explanation:
Range: Is the difference between the highest and the lowest values in a data set.
Advantages: Easy and simple to compute, useful when evaluating the whole dataset is necessary, important in showing the spread within the data set and for comparing the spread between the dataset.
Disadvantages: not reliable when the either the smallest or the largest values is exceptionally high or small (outliers) resulting in a range that is not typical of the variability within the dataset.
Inter-quartile Range: is a measure that determines the spread of the 50% of the values in a data set are dispersed from each other. It is based on the median.
Advantages: Not affected by outliers therefore a good measure of variability and it is useful when the extreme values are not recorded accurately.
Disadvantages: Based only on two values from the dataset like the range thereby not a good representative of the spread in the whole dataset.
Standard deviation: is a measure that approximates the amount by which each value in a dataset varies from the mean.
Advantages: considers every value in a dataset, shows how much data is clustered around the mean, gives a more accurate data on how the data is distributed, it is not easily affected by outliers and can be used to detect skewness when used together with the mean.
Disadvantages: does not give full range of the data, can be tiresome to calculate, only used with data where an independent variable is plotted against its frequency, assumes a normal distribution, only used when the mean is used as a measure of central tendency (i.e. symmetric numerical data) and is an inappropriate measure of dispersion for skewed data.
Variance: is the average squared from the mean.
Advantages: treats all the deviations from the mean the same regardless of direction.
Disadvantages: gives added weights to values far from the mean through squaring them thus may give skew interpretations of the data, not easily interpreted and can be stressful to calculate.
4. Calculate the measures of center and measures of variation from the data set and list them below. Be sure to include (a) an interpretation of each measure in context of the scenario (for example, if the median is larger than the mean, what does it mean? What does the value of standard deviation tell you?) and (b) correct units of measurement. Show your calculations in your spreadsheet. You do not need to include Excel functions in your written answer below.
· Mean
· Median
· Mode
· Midrange
· Range
· Variance
· Standard deviation
Answer and Explanation:
Mean = $62,306
Explanation – the average salary for the different jobs in the state of Minnesota is $62,306
Median = $56,520
Explanation – the median salary for the different jobs in the state of Minnesota is $56,520. This implies that 50% of the salaries lies below $56,520 and the remaining 50% lies above $56,520 of the salaries in Minnesota.
Mode = $40,170, $40,590, $45,510, $46,100, $52,200, $52,330, $52,340, $55,870, $55,990, $56,600, $57,230 and $96,290
Explanation – the most paid salaries in the state of Minnesota are the modal salaries given above.
Midrange = $80,010
Explanation – the average middle salary difference for the different jobs in the state of Minnesota is $80,010. This is in the consideration of the average of the highest and the lowest paid jobs in the state.
Range = $79,680
Explanation – the difference between the highest paid and the least paid salary in the state of Minnesota is $79,680.
Variance = $366,692,391.30
Explanation = the squared variability of the salaries in the state of Minnesota is given as the variance.
Standard deviation = $19,149.21
Explanation – the deviation of the salaries of the different job in the state of Minnesota from the average salary is $19,149.21
Interpretation
Since the median salary is less than the mean salary, then it implies that majority of the jobs in the state of Minnesota are paid below average. Similarly, the difference between the highest paid and the least paid is quite large ($79,680) which is more than average salary. This implies that there is a lot disparity of salaries in the state of Minnesota as also evidenced by the standard deviation and the midrange.