Scenario (information repeated for deliverable 01, 03, and 04) A major client of your company is interested in the salary distributions of jobs in the state of Minnesota that range from $30,000 to $200,000 per year. As a Business Analyst, your boss asks
Deliverable 01 Worksheet
1. Introduce your scenario and data set.
· Provide a brief overview of the scenario you are given and describe the data set.
· Describe how you will be analyzing the data set.
· Classify the variables in your data set.
· Which variables are quantitative/qualitative?
· If it is a quantitative variable, is it discrete or continuous?
· Describe the level of measurement for each variable included in the data set (nominal, ordinal, interval, ratio).
Answer and Explanation:
Enter your step-by-step answer and explanations here.
· The client wants to know the salary distribution of jobs that pay between $30,000 and $200,000 in the state of Minnesota. I have obtained data on 364 individuals in various job titles who are within the required range from the Bureau of Statistics.
· I will analyze my data using Excel and present my results using graphical tools such as charts, histograms, scatter plots, histograms and tables to present my data. I intend to use descriptive statistics such as frequency tables, median, mean, mode, variance, skewness and frequency tables to describe the distribution of the data obtained.
· Classification of my data
· A Job title is a qualitative categorical variable because it takes names as values. It has a nominal level of measurement because it uses words.
· Salary is a quantitative continuous variable because it takes numeric values. It has an interval-ratio level of measurement.
2. Discuss the importance of the Measures of Center.
· Name and describe each measure of center.
· Discuss the advantages and disadvantages of each.
Answer and Explanation:
Enter your step-by-step answer and explanations here.
Measures of the center are used to approximate and explain the middle value or average of the given data set to understand where the center of data distribution is located. The three measures of center are mean, median and mode.
1. The mean is defined as the arithmetic average of a data set and is calculated by adding all the values in the data and dividing by the number of values. Its advantage is that it can be used as a basis for statistical analysis and tests of significance. One disadvantage of the mean is that it is very sensitive to extreme values or outliers.
2. The median is the middle value or observation of a data set arranged in numerical order. This means that half of the values are below the median and half of the values are above it. Its advantage is that it is not sensitive to outliers. It is also a better descriptive measure when dealing with a skewed data set. Its disadvantage is that it cannot be used to test statistical significance in a dataset.
3. The mode refers to the value that appears most frequently in a data set. It is easy to locate which is one of its advantages. The mode is also good in measuring datasets with nominal variables. A disadvantage of the mode is that there may be two or more mode in a data set the same time. The mode is also sensitive to the mean.
3. Discuss the importance of the Measures of Variation.
· Name and describe each measure of variation.
· Discuss the advantages and disadvantages of each.
Answer and Explanation:
Enter your step-by-step answer and explanations here.
Measuring variation helps researchers measure the spread of data hence distinguishing if it is from systematic trends or differences. There measures of variation; variance, standard deviation, and range.
1. Variance is a measure of the spread of a set of values from the average value or the mean. Its advantage is that it helps us explain the volatility of a value. Its disadvantage is that it is hard to calculate manually and it is affected extreme outliers
2. Standard deviation is a measure of how close or far away from values are spread from the mean. Its advantage is that it helps us understand how data is clustered around the mean. It is also not sensitive to outliers. Its disadvantage is that it does not give a full range of the data. It is also difficult to compute.
3. The range is a measure between the largest and smallest values of a dataset. It is easier to calculate and can be used to measure variation in cases where accuracy is not necessary. Its disadvantage is that it is only affected by two values.
4. Calculate the measures of center and measures of variation from the data set and list them below. Be sure to include (a) an interpretation of each measure in the context of the scenario (for example, if the median is larger than the mean, what does it mean? What does the value of standard deviation tell you?) and (b) correct units of measurement. Show your calculations in your spreadsheet. You do not need to include Excel functions in your written answer below.
· Mean
· Median
· Mode
· Midrange
· Range
· Variance
· Standard deviation
Answer and Explanation:
Enter your step-by-step answer and explanations here.
Measure of center
1. Mean – $71,879
2. Median – $66,525
3. Mode – $71, 420
The average salary for individuals in different jobs titles that earn salaries between $30,000 and $200,000 in Minnesota is $71,879, and the mode is $71,420.The median is $66,525 which is smaller than the mean. This means that the distribution of this range of salaries is skewed to the right. It also indicates that there are outliers in the higher end.
1. Mid-range – $66,525
2. Range – $167,760
3. Variance – 546,033,522
4. Standard deviation – 23,367
The range for the salary is $167,760 which is very large and indicates that there is a greater dispersion of the salaries between $30,000 and $200,000. The variance of the data is 546,033,522, and the standard deviation is 23,367 which indicate that the salaries for the different job titles deviate from the average salary for the group by $23,376. This is a large deviation, and this might be due to the possibility of outliers as we have seen above when interpreting the measure of center. If you eliminated the outliers from the data, the standard deviation might be smaller.