STATISTICS. Week 4

profileOKGRMXL
bus_308_equal_pay_student_worksheet.xlsx

Data

ID Sal Compa Mid Age EES SER G Raise Deg Gen1 Gr
8 23 1.000 23 32 90 9 1 5.8 1 F A PAY GRADES
10 22 0.956 23 30 80 7 1 4.7 1 F A A B C D E F
11 23 1.000 23 41 100 19 1 4.8 1 F A 23 34 41 57 69 75
14 24 1.043 23 32 90 12 1 6 1 F A 22 36 42 50 65 77
15 24 1.043 23 32 80 8 1 4.9 1 F A 23 34 47 55 58 76
23 23 1.000 23 36 65 6 1 3.3 0 F A 24 35 40 47 66 77
26 24 1.043 23 22 95 2 1 6.2 0 F A 24 27 43 49 60 76
31 24 1.043 23 29 60 4 1 3.9 1 F A 23 28 64 72
35 24 1.043 23 23 90 4 1 5.3 0 F A 24 28 56
36 23 1.000 23 27 75 3 1 4.3 0 F A 24 60
37 22 0.956 23 22 95 2 1 6.2 0 F A 24 65
42 24 1.043 23 32 100 8 1 5.7 1 F A 23 62
* 19 24 1.043 23 32 85 1 0 4.6 1 M A 22 60
25 24 1.043 23 41 70 4 0 4 0 M A 24 66
40 25 1.086 23 24 90 2 0 6.3 0 M A 24
3 34 1.096 31 30 75 5 1 3.6 1 F B 24
18 36 1.161 31 31 80 11 1 5.6 0 F B 25
20 34 1.096 31 44 70 16 1 4.8 0 F B
39 35 1.129 31 27 90 6 1 5.5 0 F B
2 27 0.870 31 52 80 7 0 3.9 0 M B
32 28 0.903 31 25 95 4 0 5.6 0 M B
34 28 0.903 31 26 80 2 0 4.9 1 M B
7 41 1.025 40 32 100 8 1 5.7 1 F C
13 42 1.050 40 30 100 2 1 4.7 0 F C
16 47 1.175 40 44 90 4 0 5.7 0 M C
27 40 1.000 40 35 80 7 0 3.9 1 M C
41 43 1.075 40 25 80 5 0 4.3 0 M C
22 57 1.187 48 48 65 6 1 3.8 1 F D
24 50 1.041 48 30 75 9 1 3.8 0 F D
45 55 1.145 48 36 95 8 1 5.2 1 F D
5 47 0.979 48 36 90 16 0 5.7 1 M D
30 49 1.020 48 45 90 18 0 4.3 0 M D
17 69 1.210 57 27 55 3 1 3 1 F E
48 65 1.140 57 34 90 11 1 5.3 1 F E
1 58 1.017 57 34 85 8 0 5.7 0 M E
4 66 1.157 57 42 100 16 0 5.5 1 M E
12 60 1.052 57 52 95 22 0 4.5 0 M E
33 64 1.122 57 35 90 9 0 5.5 1 M E
38 56 0.982 57 45 95 11 0 4.5 0 M E
44 60 1.052 57 45 90 16 0 5.2 1 M E
46 65 1.140 57 39 75 20 0 3.9 1 M E
47 62 1.087 57 37 95 5 0 5.5 1 M E
49 60 1.052 57 41 95 21 0 6.6 0 M E
50 66 1.157 57 38 80 12 0 4.6 0 M E
28 75 1.119 67 44 95 9 1 4.4 0 F F
43 77 1.149 67 42 95 20 1 5.5 0 F F
6 76 1.134 67 36 70 12 0 4.5 1 M F
9 77 1.149 67 49 100 10 0 4 1 M F
21 76 1.134 67 43 95 13 0 6.3 1 M F
29 72 1.074 67 52 95 5 0 5.4 0 M F
The column labels in the table mean:
ID – Employee sample number Sal – Salary in thousands
Age – Age in years EES – Appraisal rating (Employee evaluation score)
SER – Years of service G – Gender (0 = male, 1 = female)
Mid – salary grade midpoint Raise – percent of last raise
Grade – job/pay grade Deg (0= BS\BA 1 = MS)
Gen1 (Male or Female) Compa - salary divided by midpoint, a measure of salary that removes the impact of grade

Week 1

MICHAEL LYBARGER
Week 1. Describing the data.
1 Using the Excel Analysis ToolPak function descriptive statistics, generate and show the descriptive statistics for each appropriate variable in the sample data set.
a. For which variables in the data set does this function not work correctly for? Why?
ANSWER: The variables in the data set that the descriptive statistics function does not work for are the variables that are non-numeric: Gen 1 and Gr.
Also, for this function to work properly, the data needs to be interval data
SAL COMPA MID AGE EES
MichaellybargeR: MichaellybargeR: APPRAISAL RATING SCORE
SER
MichaellybargeR: MichaellybargeR: YEARS OF SERVICE
G
MichaellybargeR: MichaellybargeR: GENDER 0=MALE 1=FEMALE
RAISE
MichaellybargeR: MichaellybargeR: % OF LAST RAISE
DEG
MichaellybargeR: MichaellybargeR: 0=BS/BA 1=MS

MichaellybargeR: MichaellybargeR: SALARY DIVIDED BY MIDPOINT, A MEASURE OF SALARY THAT REMOVES THE IMPACT OF GRADE

MichaellybargeR: MichaellybargeR: SALARY GRADE MIDPOINT
Mean 45 Mean 1.06 Mean 41.76 Mean 35.72 Mean 85.9 Mean 8.96 Mean 0.5 Mean 4.938 Mean 0.5
Standard Error 2.72 Standard Error 0.01 Standard Error 2.30 Standard Error 1.167 Standard Error 1.61 Standard Error 0.81 Standard Error 0.07 Standard Error 0.12 Standard Error 0.07
Median 42.5 Median 1.051 Median 40 Median 35 Median 90 Median 8 Median 0.5 Median 4.9 Median 0.5
Mode 24 Mode 1.043 Mode 23 Mode 32 Mode 95 Mode 8 Mode 0 Mode 5.7 Mode 0
Standard Deviation 19.20 Standard Deviation 0.08 Standard Deviation 16.23 Standard Deviation 8.25 Standard Deviation 11.41 Standard Deviation 5.72 Standard Deviation 0.51 Standard Deviation 0.87 Standard Deviation 0.51
Sample Variance 368.69 Sample Variance 0.01 Sample Variance 263.45 Sample Variance 68.08 Sample Variance 130.30 Sample Variance 32.69 Sample Variance 0.26 Sample Variance 0.75 Sample Variance 0.26
Kurtosis -1.45 Kurtosis -0.17 Kurtosis -1.52 Kurtosis -0.77 Kurtosis -0.04 Kurtosis -0.40 Kurtosis -2.09 Kurtosis -0.79 Kurtosis -2.09
Skewness 0.24 Skewness -0.32 Skewness 0.16 Skewness 0.26 Skewness -0.82 Skewness 0.73 Skewness -0.00 Skewness -0.16 Skewness -0.00
Range 55 Range 0.34 Range 44 Range 30 Range 45 Range 21 Range 1 Range 3.6 Range 1
Minimum 22 Minimum 0.87 Minimum 23 Minimum 22 Minimum 55 Minimum 1 Minimum 0 Minimum 3 Minimum 0
Maximum 77 Maximum 1.21 Maximum 67 Maximum 52 Maximum 100 Maximum 22 Maximum 1 Maximum 6.6 Maximum 1
Sum 2250 Sum 53.124 Sum 2088 Sum 1786 Sum 4295 Sum 448 Sum 25 Sum 246.9 Sum 25
Count 50 Count 50 Count 50 Count 50 Count 50 Count 50 Count 50 Count 50 Count 50
2 Sort the data by Gen or Gen 1 (into males and females) and find the mean and standard deviation for each gender for the following variables: sal, compa, age, sr and raise
Use either the descriptive stats function or the Fx functions (average and stdev)
ANSWER: I have provided BOTH the descriptive stats function and the Fx functions below:
Note: My work is shown in the Week 1 Work tab, detailing and showing the functions used.
FEMALE - DESCRIPTIVE STATS FUNCTION
SAL COMPA AGE SER RAISE
Mean 38 Mean 1.07 Mean 32.52 Mean 7.92 Mean 4.88
Standard Deviation 18.29 Standard Deviation 0.07 Standard Deviation 6.88 Standard Deviation 4.91 D 0.92
MALE - DESCRIPTIVE STATS FUNCTION
SAL COMPA AGE SER RAISE
Mean 52 Mean 1.06 Mean 38.92 Mean 10 Mean 5.00
Standard Deviation 17.78 Standard Deviation 0.08 Standard Deviation 8.39 Standard Deviation 6.36 Standard Deviation 0.83
FEMALE - Fx FUNCTIONS
SAL COMPA AGE SER RAISE
Mean 38 Mean 1.07 Mean 32.52 Mean 7.92 Mean 4.88
Standard Deviation 18.29 Standard Deviation 0.07 Standard Deviation 6.88 Standard Deviation 4.91 Standard Deviation 0.92
MALE - Fx FUNCTIONS
SAL COMPA AGE SER RAISE
Mean 52 Mean 1.06 Mean 38.92 Mean 10 Mean 5.00
Standard Deviation 17.78 Standard Deviation 0.08 Standard Deviation 8.39 Standard Deviation 6.36 Standard Deviation 0.83
3 What is the probability for a:
a.       Randomly selected person being a male in grade E?
b.      Randomly selected male being in grade E?
c. Why are the results different?
ANSWER:
A. The probability of a randomly selected person being a male in grade E is 24%. (P=e/o). This particular event (grade E) occurrs 12 times, divided by 50 possible outcomes (total employees)
B. The probability of a randomly selected male being in grade E is 40%. ( P=e/o). This particular event (male in grade E) occurrs 10 times divided by 25 possible outcomes (total males).
C. The results are different simply because the variables are asking two different questions. One requests one out of only two outcomes (male/female), and the other requests one out of six outcomes (grades)
4 Find:
A) The z score for each male salary, based on only the male salaries.
MALE ID SAL Z-SCORE MEAN STANDEVA
19 24 -1.575 52 17.7763888346
25 24 -1.575 z=(datapoint-mean)/standarddeviation
40 25 -1.519
2 27 -1.406
32 28 -1.350
34 28 -1.350
16 47 -0.281
27 40 -0.675
41 43 -0.506
5 47 -0.281
30 49 -0.169
1 58 0.338
4 66 0.788
12 60 0.450
33 64 0.675
38 56 0.225
44 60 0.450
46 65 0.731
47 62 0.563
49 60 0.450
50 66 0.788
6 76 1.350
9 77 1.406
21 76 1.350
29 72 1.125
B) The z score for each female salary, based on only the female salaries.
FEMALE ID SAL Z-SCORE MEAN STANDEVA
8 23 -0.820 38 18.294
10 22 -0.875
11 23 -0.820
14 24 -0.765
15 24 -0.765
23 23 -0.820
26 24 -0.765
31 24 -0.765
35 24 -0.765
36 23 -0.820
37 22 -0.875
42 24 -0.765
3 34 -0.219
18 36 -0.109
20 34 -0.219
39 35 -0.164
7 41 0.164
13 42 0.219
22 57 1.039
24 50 0.656
45 55 0.929
17 69 1.695
48 65 1.476
28 75 2.023
43 77 2.132
C) The z score for each female compa, based on only the female compa values.
FEMALE ID COMPA Z-SCORE MEAN STANDEVA
8 1.000 -0.977 1.069 0.0703
10 0.956 -1.607
11 1.000 -0.982
14 1.043 -0.370
15 1.043 -0.370
23 1.000 -0.982
26 1.043 -0.370
31 1.043 -0.370
35 1.043 -0.370
36 1.000 -0.982
37 0.956 -1.607
42 1.043 -0.370
3 1.096 0.384
18 1.161 1.309
20 1.096 0.384
39 1.129 0.853
7 1.025 -0.626
13 1.050 -0.270
22 1.187 1.679
24 1.041 -0.398
45 1.145 1.081
17 1.210 2.006
48 1.140 1.010
28 1.119 0.711
43 1.149 1.138
D) The z score for each male compa, based on only the male compa values.
MALE ID COMPA Z-SCORE MEAN STANDEVA
19 1.043 -0.155 1.056 0.0838
25 1.043 -0.155
40 1.086 0.358
2 0.870 -2.220
32 0.903 -1.826
34 0.903 -1.826
16 1.175 1.420
27 1.000 -0.668
41 1.075 0.227
5 0.979 -0.919
30 1.020 -0.430
1 1.017 -0.465
4 1.157 1.205
12 1.052 -0.048
33 1.122 0.788
38 0.982 -0.883
44 1.052 -0.048
46 1.140 1.002
47 1.087 0.370
49 1.052 -0.048
50 1.157 1.205
6 1.134 0.931
9 1.149 1.110
21 1.134 0.931
29 1.074 0.215
E) What do the distributions and spread suggest about male and female salaries?
ANSWER: Since almost all the z-score values are less than the average (mean) the data tells that there are very few variances in salary between genders.
F) Why might we want to use compa to measure salaries between males and females?
ANSWER: The compa data discloses the amount(s) of increase or decrease in salaries per gender.
This data, along with the z-scores, shows that the salaries are comparable to both the mean and standard deviations.
5) Based on this sample, what conclusions can you make about the issue of male and female pay equality?
ANSWER: Based on the data provided, this control group contains both male and females that are being paid based on their levels of education, performance reviews, and age/experience.
However, the data does disclose to me that the male control group are in fact paid more when compared to the female control group with the same service times.
6) Are all of the results consistent with your conclusion? If not, why not?
ANSWER: Although the raise percentages are consistent with performance reviews, the data does show an inconsistency where men are paid more than women with equal time of service and degree.
One variable that is positive is that in all three scenarios detailed below, the females were given a higher raise % than their male counterparts.
SALARY AGE EES SERVICE RAISE DEG M/F Scenarios
76 36 70 12 4.5 1 M A female with 12 years and an EES rating of 90 makes $52,000 less than a male with 12 years and an EES rating of 70
65 39 75 20 3.9 1 M A female with 7 years and an EES rating of 80 makes $18,000 less than a male with 7 years and an EES rating of 80
64 35 90 9 5.5 1 M A female with 9 years and and EES rating of 90 makes $41,000 less than a male with 9 years and an EES rating of 90
62 37 95 5 5.5 1 M
60 45 90 16 5.2 1 M
40 35 80 7 3.9 1 M
34 30 75 5 3.6 1 F
24 32 90 12 6 1 F
24 32 80 8 4.9 1 F
24 29 60 4 3.9 1 F
23 32 90 9 5.8 1 F
22 30 80 7 4.7 1 F

Week 2

Week 2 Testing means with the t-test <Note: use right click on row numbers to insert rows to perform analysis below any question>
For questions 2 and 3 below, be sure to list the null and alternate hypothesis statements. Use .05 for your significance level in making your decisions.
For full credit, you need to also show the statistical outcomes - either the Excel test result or the calculations you performed.
1 Below are 2 one-sample t-tests comparing male and female average salaries to the overall sample mean.
Based on our sample, how do you interpret the results and what do these results suggest about the population means for male and female salaries?
Males Females
Ho: Mean salary = 45 Ho: Mean salary = 45
Ha: Mean salary =/= 45 Ha: Mean salary =/= 45
Note when performing a one sample test with ANOVA, the second variable (Ho) is listed as the same value for every corresponding value in the data set.
t-Test: Two-Sample Assuming Unequal Variances t-Test: Two-Sample Assuming Unequal Variances
Since the Ho variable has Var = 0, variances are unequal; this test defaults to 1 sample t in this situation
Male Ho Female Ho
Mean 52 45 Mean 38 45
Variance 316 0 Variance 334.6666666667 0
Observations 25 25 Observations 25 25
Hypothesized Mean Difference 0 Hypothesized Mean Difference 0
df 24 df 24
t Stat 1.9689038266 t Stat -1.9132063573
P(T<=t) one-tail 0.0303078503 P(T<=t) one-tail 0.0338621184
t Critical one-tail 1.7108820799 t Critical one-tail 1.7108820799
P(T<=t) two-tail 0.0606157006 P(T<=t) two-tail 0.0677242369
t Critical two-tail 2.0638985616 t Critical two-tail 2.0638985616
Conclusion: Do not reject Ho; mean equals 45 Conclusion: Do not reject Ho; mean equals 45
Interpretation:
2 Based on our sample results, perform a 2-sample t-test to see if the population male and female salaries could be equal to each other.
3 Based on our sample results, can the male and female compas in the population be equal to each other? (Another 2-sample t-test.)
4 What other information would you like to know to answer the question about salary equity between the genders? Why?
5 If the salary and compa mean tests in questions 3 and 4 provide different results about male and female salary equality,
which would be more appropriate to use in answering the question about salary equity? Why?
What are your conclusions about equal pay at this point?

Week 3

Week 3 Testing multiple means with ANOVA <Note: use right click on row numbers to insert rows to perform analysis below any question>
For questions 3 and 4 below, be sure to list the null and alternate hypothesis statements. Use .05 for your significance level in making your decisions.
For full credit, you need to also show the statistical outcomes - either the Excel test result or the calculations you performed.
1.      Based on the sample data, can the average(mean) salary in the population be the same for each of the grade levels? (Assume equal variance, and use the analysis toolpak function ANOVA.)
Set up the input table/range to use as follows: Put all of the salary values for each grade under the appropriate grade label.
Be sure to incllude the null and alternate hypothesis along with the statistical test and result.
A B C D E F ANOVA - Single Factor
23 27 41 47 58 76 SUMMARY H0: µA=µB=µC=µD=µE=µF
22 34 42 57 66 77 Groups Count Sum Average Variance H1: µ are not equal
23 36 47 50 60 76 A 15 353 23.533 0.6952
24 34 40 49 69 75 B 7 222 31.714 14.9048
24 28 43 55 64 72 C 5 213 42.600 7.3000
24 28 56 77 D 5 258 51.600 17.8000
23 35 60 E 12 751 62.583 14.8106
24 65 F 6 453 75.500 3.5000
24 62
24 65 ANOVA
24 60 Source of Variation SS df MS F P-value F crit
23 66 Between Groups 17686.0214285714 5 3537.2042857143 409.5941199692 1.03856156090236E-35 2.4270401198
22 Within Groups 379.9785714286 44 8.6358766234
25 Total 18066 49
24
Interpretation The average salaray in the population is not the same for each grade level.
The probability exceeds the p=.05 standard for statistical significance
2 The table and analysis below demonstrate a 2-way ANOVA with replication. Please interpret the results.
Grade
Gender A B C D E F
M 24 27 40 47 56 76
25 28 47 49 66 77
F 22 34 41 50 65 75
24 36 42 57 69 77
Ho: Average salaries are equal for all grades
Ha: Average salaries are not equal for all grades
Ho: Average salaries by gender are equal
Ha: Average salaries by gender are not equal
Ho: Interaction is not significant
Ha: Interaction is significant
Perform analysis:
Anova: Two-Factor With Replication
SUMMARY A B C D E F Total
M
Count 2 2 2 2 2 2 12
Sum 49 55 87 96 122 153 562
Average 24.5 27.5 43.5 48 61 76.5 46.8333333333
Variance 0.5 0.5 24.5 2 50 0.5 364.5151515152
Interpretation Salaries are significantly different across the pay grades. Variance within pay grade is low except C and E.
Variance across grades is high
F
Count 2 2 2 2 2 2 12
Sum 46 70 83 107 134 152 592
Average 23 35 41.5 53.5 67 76 49.3333333333
Variance 2 2 0.5 24.5 8 2 367.3333333333
Interpretation Salaries are significant different across the pay grades. Variance within pay grade is low except D
Variance across grades is high
Total
Count 4 4 4 4 4 4
Sum 95 125 170 203 256 305
Average 23.75 31.25 42.5 50.75 64 76.25
Variance 1.5833333333 19.5833333333 9.6666666667 18.9166666667 31.3333333333 0.9166666667
Interpretation Combining the genders
Variance is high in grade B and grade E
ANOVA
Source of Variation SS df MS F P-value F crit
Sample 37.5 1 37.5 3.8461538462 0.0734833371 4.7472253467 1.50E-10 0.0000000001
Columns 7841.8333333333 5 1568.3666666667 160.8581196581 0.0000000001 3.1058752391 Note: a number with an E after it (E9 or E-6, for example)
Interaction 91.5 5 18.3 1.8769230769 0.1723082608 3.1058752391 means we move the decimal point that number of places.
Within 117 12 9.75 For example, 1.2E4 becomes 12000; while 4.56E-5 becomes 0.0000456
Total 8087.8333333333 23
Do we reject or not reject each of the null hypotheses? What do your conclusions mean about the population values being tested?
Interpretation: Salaries in grade are quite different.
The salaries within grade are not statistically significant.
Neither men or women are singled out in salary, whether for good or bad reasons
We reject the null hypothesis because the population values being tested do not show a definitive gap in salary within grade and are above .05
3.    Using our sample results, can we say that the compa values in the population are equal by grade and/or gender, and are independent of each factor?
Grade Be sure to include the null and alternate hypothesis along with the statistical test and result.
Gender A B C D E F
M 1.043 0.870 1.000 0.979 0.982 1.134 H0: µA=µB=µC=µD=µE=µF
1.086 0.903 1.075 1.020 1.157 1.149 H1: µ are not equal
F 0.956 1.096 1.025 1.041 1.140 1.119
1.000 1.161 1.050 1.187 1.210 1.149
Anova: Two-Factor With Replication
SUMMARY A B C D E F Total
M
Count 2 2 2 2 2 2 12
Sum 2.129 1.773 2.175 1.999 2.139 2.283 12.498
Average 1.0645 0.8865 1.0875 0.9995 1.0695 1.1415 1.0415
Variance 0.0009245 0.0005445 0.0153125 0.0008405 0.0153125 0.0001125 0.0101348182
F
Count 2 2 2 2 2 2 12
Sum 1.999 2.257 2.075 2.228 2.35 2.268 13.177
Average 0.9995 1.1285 1.0375 1.114 1.175 1.134 1.0980833333
Variance 0.0037845 0.0021125 0.0003125 0.010658 0.00245 0.00045 0.0057559015
Total
Count 4 4 4 4 4 4
Sum 4.128 4.03 4.25 4.227 4.489 4.551
Average 1.032 1.0075 1.0625 1.05675 1.12225 1.13775
Variance 0.002978 0.020407 0.0060416667 0.0082029167 0.0096309167 0.00020625
ANOVA
Source of Variation SS df MS F P-value F crit
Sample 0.0192100417 1 0.0192100417 4.3647199159 0.0586593857 4.7472253467
Columns 0.0516077083 5 0.0103215417 2.3451608933 0.1051237912 3.1058752391
Interaction 0.0703757083 5 0.0140751417 3.1980175899 0.0459226589 3.1058752391
Within 0.0528145 12 0.0044012083
Total 0.1940079583 23
Interpretation The p-value for average compa is .058 and the p-value for the selected average is .105
Since both are greater than .05 so the null hypothesis is not rejected
H0: µA=µB=µC=µD=µE=µF
Ha: µ are not equal
4.   Pick any other variable you are interested in and do a simple 2-way ANOVA without replication. Why did you pick this variable and what do the results show?
Variable name: Be sure to include the null and alternate hypothesis along with the statistical test and result.
Gender A B C D E F TOTAL We will be using the age variable.
M 32.333 34.333 34.667 40.500 40.800 45.000 37.939 Hint: use mean values in the boxes.
F 29.833 33.000 31.000 38.000 30.500 43.000 34.222
TOTAL 31.083 33.667 32.833 39.250 35.650 44.000 36.081
Anova: Two-Factor Without Replication
SUMMARY Count Sum Average Variance
29.8333 5 175.5 35.1 28.3
31.0833 5 185.4 37.08 21.0812986112
34.3333 2 66.66665 33.333325 0.2222111113
34.6667 2 63.83335 31.916675 1.6805861112
40.5 2 77.25 38.625 0.78125
40.8 2 66.15 33.075 13.26125
45 2 87 43.5 0.5
ANOVA
Source of Variation SS df MS F P-value F crit
Rows 9.801 1 9.801 5.9003982945 0.0720530227 7.7086474222
Columns 190.8808972225 4 47.7202243056 28.7285307731 0.0033183668 6.3882329087
Error 6.6442972225 4 1.6610743056
Total 207.326194445 9
Interpretation I picked this variable primarily because it seemed the simplest to decipher the results
The results show that the ages across grades are statistically significant.
The results show that the ages between genders is statistically significant.
5.   Using the results for this week, What are your conclusions about gender equal pay for equal work at this point?
Interpretation The results that stuck out the most to me are that 21 of 25 women are in grade D and below.
Only 16% of the women are in Grade E or F.
56% of men are in Grade E or F.
All said, the salaries across grades are significantly different.
There is not enough data to represent the hypothesis that unequal pay for unequal work is present between genders.

Week 4

Week 4 Confidence Intervals and Chi Square (Chs 11 - 12) Let's look at some other factors that might influence pay. Q1 Q2 <Note: use right click on row numbers to insert rows to perform analysis below any question>
For question 3 below, be sure to list the null and alternate hypothesis statements. Use .05 for your significance level in making your decisions. Gr Deg Gen1 Sal
For full credit, you need to also show the statistical outcomes - either the Excel test result or the calculations you performed. A 0 F 34
1 One question we might have is if the distribution of graduate and undergraduate degrees independent of the grade the employee? A 0 F 41
(Note: this is the same as asking if the degrees are distributed the same way.)
Based on the analysis of our sample data (shown below), what is your answer?
Ho: The populaton correlation between grade and degree is 0. C 0 F 77
Ha: The population correlation between grade and degree is > 0
Perform analysis:
OBSERVED A B C D E F Total
COUNT - M or 0 7 5 3 2 5 3 25
COUNT - F or 1 8 2 2 3 7 3 25
total 15 7 5 5 12 6 50
EXPECTED
7.5 3.5 2.5 2.5 6 3 25 <Highlighting each cell with show how the value
7.5 3.5 2.5 2.5 6 3 25 is found: row total times column total divided by
15 7 5 5 12 6 50 grand total.>
By using either the Excel Chi Square functions or calculating the results directly as the text shows, do we
reject or not reject the null hypothesis? What does your conclusion mean?
Interpretation:
2 Using our sample data, we can construct a 95% confidence interval for the population's mean salary for each gender.
Interpret the results. How do they compare with the findings in the week 2 one sample t-test outcomes (Question 1)?
Males Mean St error Low to High
52 3.6587793957 44.4482793272 59.5517206728 Results are mean +/-2.064*standard error
Females 38 3.6227541769 30.5226353789 45.4773646211 2.064 is t value for 95% interval
<Reminder: standard error is the sample standard deviation divided by the square root of the sample size.>
Interpretation:
C 0 F 55
D 1 M 77
3 Based on our sample data, can we conclude that males and females are distributed across grades in a similar pattern within the population? D 1 M 60
4 Using our sample data, construct a 95% confidence interval for the population's mean service difference for each gender.
Do they intersect or overlap? How do these results compare to the findings in week 2, question 2?
5 How do you interpret these results in light of our question about equal pay for equal work?

Week 5

Week 5 Correlation and Regression
For each question involving a statistical test below, list the null and alternate hypothesis statements. Use .05 for your significance level in making your decisions.
For full credit, you need to also show the statistical outcomes - either the Excel test result or the calculations you performed.
1 Create a correlation table for the variables in our data set. (Use analysis ToolPak function Correlation.)
a. Interpret the results. What variables seem to be important in seeing if we pay males and females equally for equal work?
2 Below is a regression analysis for salary being predicted/explained by the other variables in our sample (Mid,
age, ees, sr, raise, and deg variables.) (Note: since salary and compa are different ways of
expressing an employee’s salary, we do not want to have both used in the same regression.)
Ho: The regression equation is not significant.
Ha: The regression equation is significant.
Ho: The regression coefficient for each variable is not significant
Ha: The regression coefficient for each variable is significant
Sal The analysis used Sal as the y (dependent variable) and
SUMMARY OUTPUT mid, age, ees, sr, g, raise, and deg as the dependent
variables (entered as a range).
Regression Statistics
Multiple R 0.9921549762
R Square 0.9843714969
Adjusted R Square 0.9817667464
Standard Error 2.5927763074
Observations 50
ANOVA
df SS MS F Significance F
Regression 7 17783.6554628284 2540.5222089755 377.9139268848 8.44042689148567E-36
Residual 42 282.3445371716 6.7224889803
Total 49 18066
Coefficients Standard Error t Stat P-value Lower 95% Upper 95% Lower 95.0% Upper 95.0%
Intercept -4.009 3.775 -1.062 0.294 -11.627 3.609 -11.627 3.609
Mid 1.220 0.030 40.674 0.000 1.159 1.280 1.159 1.280
Age 0.029 0.067 0.439 0.663 -0.105 0.164 -0.105 0.164
EES -0.096 0.047 -2.020 0.050 -0.191 -0.000 -0.191 -0.000
SR -0.074 0.084 -0.876 0.386 -0.244 0.096 -0.244 0.096
G 2.552 0.847 3.012 0.004 0.842 4.261 0.842 4.261
Raise 0.834 0.643 1.299 0.201 -0.462 2.131 -0.462 2.131
Deg 1.002 0.744 1.347 0.185 -0.500 2.504 -0.500 2.504
Interpretation: Do you reject or not reject the regression null hypothesis?
Do you reject or not reject the null hypothesis for each variable?
What is the regression equation, using only significant variables if any exist?
What does result tell us about equal pay for equal work for males and females?
3 Perform a regression analysis using compa as the dependent variable and the same independent
variables as used in question 2. Show the result, and interpret your findings by answering the same questions.
Note: be sure to include the appropriate hypothesis statements.
4 Based on all of your results to date, is gender a factor in the pay practices of this company? Why or why not?
Which is the best variable to use in analyzing pay practices - salary or compa? Why?
5 Why did the single factor tests and analysis (such as t and single factor ANOVA tests on salary equality) not provide a complete answer to our salary equality question?
What outcomes in your life or work might benefit from a multiple regression examination rather than a simpler one variable test?

Sheet1

2 way ANOVA with replication
SUMMARY A B C D E F TOTAL
MALE
COUNT 2 2 2 2 2 2 12
SUM 2.129 1.773 2.075 1.999 2.139 2.283 12.398
AVERAGE 1.0645 0.8865 1.0375 0.9995 1.0695 1.1415 1.033
VARIANCE 0.000925 0.000545 0.015313 0.000841 0.015313 0.000113 0.010135
All averages are similar other than B
SUMMARY A B C D E F TOTAL
FEMALE
COUNT 2 2 2 2 2 2 12
SUM 1.956 2.257 2.075 2.228 2.35 2.268 13.134
AVERAGE 0.978 1.1285 1.0375 1.114 1.175 1.134 1.095
VARIANCE 0.003785 0.002113 0.000313 0.010658 0.00245 0.00045 0.005756
All averages are similar. Grade D has the largest variance
SUMMARY A B C D E F
COUNT 4 4 4 4 4 4
SUM 4.085 4.03 4.15 4.227 4.489 4.551
AVERAGE 1.02125 1.0075 1.0375 1.05675 1.12225 1.13775
VARIANCE 0.0031249 0.020407 0.001042 0.008203 0.004215 0.00105
There is not much differentiation in the average data or the variance data
ANOVA
Source of Variation SS df MS F P-value F crit
Sample 0.01921 1 0.01921 4.36472 0.058659 4.747225
Columns 0.05161 5 0.010322 2.345161 0.105124 3.105875
Interaction 0.07038 5 0.014075 3.198018 0.045923 3.105875
Within 0.05282 12 0.004401
Total 0.19401 23