SPSS assignment

profileDanyah
lecture__11_chi-square_analysis_sp_15.pptx

Chi-Square Analysis

Chi-Square Statistics

The Chi-square test uses frequency data to generate a statistical results.

Chi-Square analysis is used to:

assess the relationship between two qualitative variables.

test whether a frequency fits a specific pattern (expected frequencies).

Where O is the observed frequency and E is the expected frequency

Types of Chi-Square Tests

Chi-square for goodness-of- fit:

This test is used to test how closely an observed distribution matches an expected distribution.

The null hypothesis: Observed distribution fits an expected distribution

Contingency table analysis –

Chi-square for independence:

This test is used to assess the association between two qualitative variables

The null hypothesis: There is no association between the two variables

Chi-square test for homogeneity:

This test is used to test the claim that different populations have the same proportions of some characteristics.

The null hypothesis: The distribution of the categorical (qualitative) variable is the same across the population

Types of Chi-Square Tests

McNemar’s test :

This test is use to assess the relationship between two paired (related) qualitative variables.

It is similar to paired t-test for paired quantitative variables

The null hypothesis: the proportion of subjects with the specific characteristic (or event) is the same before and after the exposure or the intervention

1. Suppose there are n observations.

2. Each observation falls into a cell (or class).

3. Observed frequencies in each cell: O1, O2, O3, … , Ok.

Sum of the observed frequencies is n.

4. Expected, or theoretical, frequencies: E1, E2, E3, . . . , Ek.

Sum of the expected frequencies is n

Goodness- of- Fit Test

5

Example: A group of researchers from the nursing school has suggested that normal births do not take place randomly throughout the day. The researchers observed delivery frequency of 28 deliveries within 24 hours and obtained the following results:

Observed
12- 6 am 12
6 – 12 am 5
12 – 6 pm 3
6 – 12 pm 8

6

Answer:

H0: Observed number of births do not vary overtime

Observed Expected Test Statistic
12- 6 am 12 7 3.57
6 – 12 am 5 7 0.57
12 – 6 pm 3 7 2.29
6 – 12 pm 8 7 0.14
Total 28 28 6.57
Expected value=(12+5+3+8)/4=7
Degree of freedom= n – 1= 4-1=3 Critical value at alpha of 0.05 = 7.815 * n is the number of groups

7

8

Results:

Decision: Fail to reject H0.

Conclusion: At the 0.05 level of significance, there is no evidence to suggest delivery varies overtime

Test statistic χ 2 = 6.57

Critical value χ2 = 7.815

9

Contingency Table Analysis

Test of Independence

Data is sorted into cells, and the observed frequency in each cell is reported in a cross tabulation.

Cross tabulation involves two qualitative variables

Typical question: Are the two variables independent or dependent?

Are the socioeconomic status and levels of physical activity independent?

Is there any association between exercise status at baseline and gender?

The null hypothesis: There is no association (relationship) between the two variables

10

H0: Smoking status at baseline is independent of gender

Example: Does smoking status at baseline depend on gender?

BASELINE SMOKING STATUS * GENDER
Gender Total
Male Female
Smoking Status Smoker 18 19 37
Non-smoker 149 240 389
Total 167 259 426

Observed Values

11

H0: Smoking status at baseline is independent of gender

Example: Does smoking status at baseline depend on gender?

BASELINE SMOKING STATUS * GENDER
Gender Total
Male Female
Smoking Status Smoker (37*167)/426=14.5 (37*259)/426=22.5 37
Non-smoker (389*167)/426=152.5 (389*259)/426=236.5 389
Total 167 259 426

Expected Values

12

H0: Smoking status at baseline is independent of gender

Example: Does smoking status at baseline depend on gender?

BASELINE SMOKING STATUS * GENDER
Gender Total
Male Female
Smoking Status Smoker Observed 18 19 37
Expected 14.5 22.5
Non-smoker Observed 149 240 389
Expected 152.5 236.5
Total 167 259 426

Observed and Expected Values

13

BASELINE SMOKING STATUS * GENDER
Gender Total
Male Female
Smoking Status Smoker Observed 18 19 37
Expected 14.5 22.5
Non-smoker Observed 149 240 389
Expected 152.5 236.5
Total 167 259 426

Chi-square analysis

Test Statistics

=====1.52

14

Chi-square test statistic = 1.52

3.84

Degrees of freedom: (rows - 1)(columns - 1) = 1

Critical value from table (α = 0.05) = 3.84

y

x

FTR

15

Results

Decision: Fail to reject H0

Conclusion: at the 0.05 level of significance, there is no association between smoking status at baseline and gender

16

Contingency Table Analysis

Test of Homogeneity

Data is sorted into cells, and the observed frequency in each cell is reported in a cross tabulation.

Cross tabulation involves two variables

This test is used to test the claim that different populations have the same proportions of some characteristics (to test the equality of proportions in different populations).

The null hypothesis: The distribution of the categorical (qualitative) variable is the same across the populations.

It is computed exactly the same as the chi-square test for independence.

17

Example: Is the proportion of the students who drive their own or their parents’ cars the same at all three schools ?

Drive their own or their parents’ cars * School
School
A B C Total
Drive their own or their parents’ cars Yes 18 22 16 56
No 32 28 34 94
Total 50 50 50 150

Observed Values

18

Expected Values

Drive their own or their parents’ cars * School
School
A B C Total
Drive their own or their parents’ cars? Yes (56*50)/150=18.67 (56*50)/150=18.67 (56*50)/150=18.67 56
No (94*50)/150=31.33 (94*50)/150=31.33 (94*50)/150=31.33 94
Total 50 50 50 150

Example: Is the proportion of the students who drive their own or their parents’ cars the same at all three schools ?

19

Drive their own or their parents’ cars * School
School Total
A B C
Drive their own or their parents’ cars? Yes Observed 18 22 16 56
Expected 18.67 18.67 18.67
No Observed 32 28 34 94
Expected 31.33 31.33 31.33
Total 50 50 50 150

Observed and Expected Values

20

Chi-square analysis

Test Statistics

Drive their own or their parents’ cars * School
School Total
A B C
Drive their own or their parents’ cars Yes Observed 18 22 16 56
Expected 18.67 18.67 18.67
No Observed 32 28 34 94
Expected 31.33 31.33 31.33
Total 50 50 50 150

=======1.60

21

Chi-square test statistic = 1.596

5.991

Degrees of freedom: (rows - 1)(columns - 1) = 2

Critical value from table (α = 0.05) = 5.991

y

x

FTR

22

Results

Decision: Fail to reject H0

Conclusion: at the 0.05 level of significance, the proportion of the students who drive their own or their parents’ cars is the same at all three schools

23

General rules to follow in using Chi-square:

No more than 20% of the cells should have an expected frequency less than 5.

No expected frequency should be less than 1.

If the expected frequencies are too small then FISHER’S EXACT TEST should be used in place of the Chi-square.

Observations MUST be considered to be independent. This means that it cannot be used on a ‘before and after’ problem.

If independence cannot be assumed then another non-parametric test called the MCNEMAR test must be used.

24

Example: Does baseline exercise level depend on martial status?

H0: Exercise level at baseline is independent of marital status

Assumption not met

We will use Fisher’s Exact test to answer this question because the assumption of Pearson Chi-Square was not met (The minimum expected count is less than 1)

Example: Does the number of smokers change after 6 weeks compared to baseline?

We will use McNemar test to answer this question because it is before and after situation and the variables are categorical.

Example: Does the number of smokers change after 6 weeks compared to baseline?

We will use McNemar test to answer this question because it is before and after situation and the variables are categorical.

Results

Decision: Fail to reject H0

Conclusion: at the 0.05 level of significance, the number of smokers does not change after 6 weeks compared to baseline

28

2

2

()

i

i

i

OE

E

c

-

n

O

O

O

O

k

=

+

+

+

+

L

3

2

1

n

E

E

E

E

k

=

+

+

+

+

L

3

2

1

k Categories

1st2nd3rd

kth

Total

Observed FrequencyO

1

O

2

O

3

O

k

n

Expected Frequency

E

1

E

2

E

3

E

k

n

L

Sheet1

k Categories
1st 2nd 3rd kth Total
Observed Frequency O1 O2 O3 Ok n
Expected Frequency E1 E2 E3 Ek n
&A
Page &P

Sheet2

&A
Page &P

Sheet3

&A
Page &P

Sheet4

&A
Page &P

Sheet5

&A
Page &P

Sheet6

&A
Page &P

Sheet7

&A
Page &P

Sheet8

&A
Page &P

Sheet9

&A
Page &P

Sheet10

&A
Page &P

Sheet11

&A
Page &P

Sheet12

&A
Page &P

Sheet13

&A
Page &P

Sheet14

&A
Page &P

Sheet15

&A
Page &P

Sheet16

&A
Page &P

2

2

()

6.57

i

i

i

OE

E

c

-

=å=

(

)

(

)

RowtotalColumntotal

ExpectedValues

Grandtotal

=

2

2

()

1.52

i

i

i

OE

E

c

-

==

å

c

0

123

1

:

:At

Hppp

Hleastoneproportionisdifferentfromtheoth

ers

==

2

2

()

1.60

i

i

i

OE

E

c

-

==

å