Statistics Analysis Report Correction

profileSylviameng
ResearchReport3-Boffa-Weng-Lacroix.docx

Cayla Boffa, Weiqian Meng, Connor Lacroix

Research Report #3

PSY 1100

For Research Report# 3, we collected data to research if the amount of hours of exercise per week differ between people in their 20s, people in their 30s and people in their 40s. The reason we choose to collect this data was to evaluate if people exercise more or less as they get older and to evaluate if there was a significant difference in how many hours per week one exercises when moving from their 20s to 30s to 40s. The three variables we collected were Age (20s, 30s, 40s) which was our independent variable and how many hours per week they exercise as our dependent variable. The way we collected our data was through asking our family and friends via email/text. Each person in the group collected data for 3 persons in each group. What we could control was the age range of the individuals. A couple things that we could not control was if the people we were polling were health fanatics or not, were a healthy weight or overweight, wanted to lose weight or had any health issues. For our data, we collected data from 9 people in their 20s, 9 people in their 30s, and 9 people in their 40s. With this study we expected to find out that people in their 40s work out more hours per week than people in their 20s or 30s as it is more difficult to lose weight as you get older.

The table below shows the raw data collected for people in their 20s, 30s and 40 and the hours per week they exercise. For our study, we collected data from 9 people in their 20s, 9 people in their 30s, and 9 people in their 40s. “Group 1” represents people in their 20s. “Group 2” represents people in their 30s and “Group 3” represents people in their 30s. The numbers provided represents hours per week spent exercising.

GROUP 1

GROUP 2

GROUP 3

5

5

4

7

5

10

2

5

6

10

5

12

7

6

4

4

6

6

8

6

7

7

6

4

5

6

7

The three assumptions associated with an independent samples t-test are normality, homogeneity and independence of observations. According to chapter 14 notes, normality states that the dependent variable should be normally distributed in the population from which we draw our samples. The current data was only obtained from 9 and 9 and 9 subjects which is a small portion of the population. The research was specific to the population of people in their 20s, 30s and 40s. Below are the histograms, first is the one represented by people in their 20s, the second is represented by people in their 30s and third one represents people in their 40s.

From the three histograms you can see that Group 1, Group 2 and Group 3 are unimodal. Given that the sample size is small and the histograms above are not symmetric we would assume violation of normality.

The second assumption is homogeneity of population variance which is the situation in which two or more population variances have equal variance. The rule is a 4 times difference of the variances when comparing one set to the other. In our study the standard deviation for Group 1-20s was 2.368; the standard deviation for Group 2-30s was .527 and the standard deviation for Group 3-40s was 2.78. There is a 4 times difference in variance, therefore, we would assume violation of homogeneity of population variance is not violated.

Descriptive Statistics

Dependent Variable: Hours of exercise per week

Age

Mean

Std. Deviation

N

20's

6.1111

2.36878

9

30's

5.5556

.52705

9

40's

6.6667

2.78388

9

Total

6.1111

2.10006

27

The third and MOST important assumption is independence of observation, which assumes that samples are independent of one another. Each of us provided data for 3 individuals for each group. Each person was individually sent the question of how many hours a week they exercise. There was no discussion amongst the people regarding their answers. Therefore, the question and answers were independent of one another. For this reason, the assumption of independence of observation has not been violated.

To compute the F value using ANOVA, we have 3 groups, with 9 subjects in each group. Therefore, we calculate df for group as k-1 (3-1=2). Df for Error is n-k (27-3)= 24 and df total equals n-1(27-1=26).

GROUP 1

GROUP 2

GROUP 3

5

5

4

7

5

10

2

5

6

10

5

12

7

6

4

4

6

6

8

6

7

7

6

4

5

6

7

∑x= 55 ∑x= 50 ∑x= 60

∑x2= 381 ∑x2= 280 ∑x2= 462

To compute the F value using ANOVA, we have 3 groups, with 9 subjects in each group. Therefore, we calculate df for group as k-1 (3-1=2). Df for Error is n-k (27-3)= 24 and df total equals n-1(27-1=26).

The F critical value for F .05 (2, 24) = +3.40. This value is taken from the table.

From there we calculate SSTotal, SSGroup and SSError.

Tests of Between-Subjects Effects

Dependent Variable: Hours of exercise per week

Source

Type III Sum of Squares

df

Mean Square

F

Sig.

Partial Eta Squared

Corrected Model

5.556a

2

2.778

.611

.551

.048

Intercept

1008.333

1

1008.333

221.792

.000

.902

Age

5.556

2

2.778

.611

.551

.048

Error

109.111

24

4.546

Total

1123.000

27

Corrected Total

114.667

26

a. R Squared = .048 (Adjusted R Squared = -.031)

Boffa 5