Advance Biostats SPSS Assignment: Multiple Logistic Regression in Action
PART3
Step-by-Step Guide to Assignment 7.3
Multivariable Logistic Regression
Problem 3. Multivariable Logistic Regression
a. Run a multivarible binary logistic regression model using SPSS and Hypertension as the dependent variable, Chol_Cat, Age_Cat, Obese, and Sex as the Covariates. Include the output in your submission.
Step 1. Open the SPSS data set used in Problems 7.1 and 7.2. Go to Analyze ( Regression ( Binary Logistic
Step 2. Click Reset to remove the entries made in Problem 2.
Step 3. Place Hypertension in the Dependent box. Place Chole_Cat, Age_Cat, Obese, and Sex in the covariates box. Click on Categorical.
Step 4. Make sure Contrast is set to Indicator. Change Reference Category to First. To do this, highlight Age_Cat by clicking on it. Change Reference Category from Last to First and click Change. The word (first) will appear next to Age_Cat.
Step 5. Repeat the steps 3 and 4 for each of the remaining categorical variables. Click Continue.
Step 6. Click on Options
Step 7. Check the boxes the same as below. Click Continue.
Step 8. Save your output file. (The Output file must be submitted with the Application.)
SPSS Output:
b. Identify the Odds Ratio and the significance of the Odds Ratio for each of the covariates. How has the relationship between Chole_Cat and Hypertension changed with the addition of the other variables (compare to the output from # 2)?
ORs for hypertension with Chole_Cat:
|
Variables in the Equation |
|||||||||
|
|
B |
S.E. |
Wald |
df |
Sig. |
Exp(B) |
95% C.I.for EXP(B) |
||
|
|
|
|
|
|
|
|
Lower |
Upper |
|
|
Step 1a |
Chole_Cat |
|
|
10.989 |
2 |
.004 |
|
|
|
|
|
Chole_Cat(1) |
1.294 |
.537 |
5.816 |
1 |
.016 |
3.648 |
1.274 |
10.443 |
|
|
Chole_Cat(2) |
2.667 |
.828 |
10.369 |
1 |
.001 |
14.400 |
2.840 |
73.018 |
|
|
Constant |
-1.569 |
.492 |
10.182 |
1 |
.001 |
.208 |
|
|
|
a. Variable(s) entered on step 1: Chole_Cat. |
ORs for Hypertension with Chole Cat controlling for other variables
|
Variables in the Equation |
|||||||||
|
|
B |
S.E. |
Wald |
df |
Sig. |
Exp(B) |
95% C.I.for EXP(B) |
||
|
|
|
|
|
|
|
|
Lower |
Upper |
|
|
Step 1a |
Age_Cat(1) |
.221 |
.394 |
.313 |
1 |
.576 |
1.247 |
.576 |
2.700 |
|
|
Chole_Cat |
|
|
9.317 |
2 |
.009 |
|
|
|
|
|
Chole_Cat(1) |
1.202 |
.554 |
4.708 |
1 |
.030 |
3.328 |
1.123 |
9.859 |
|
|
Chole_Cat(2) |
2.546 |
.849 |
8.993 |
1 |
.003 |
12.762 |
2.416 |
67.408 |
|
|
sex(1) |
-.217 |
.386 |
.314 |
1 |
.575 |
.805 |
.378 |
1.717 |
|
|
Constant |
-1.502 |
.528 |
8.091 |
1 |
.004 |
.223 |
|
|
|
a. Variable(s) entered on step 1: Age_Cat, Chole_Cat, sex. |
(Exp(B)(exponentialtion of the B coefficients) gives the odds ratio)
In your response, be sure to discuss the OR for Age_Cat and Sex and how they contribute to the model of the relationship between Cholesterol and BP. Include a discussion of the meaning of the odds ratios, the change in the model, and the significance of each of the variables in the model.
c. Test the assumption that the model fits the data using using the Hosmer-Lemeshow Goodness of Fit test. Interpret the Chi Square statistic given in the output of this test and state what it means in terms of the assumptions needed to use logistic regression with this data.
In the SPSS Output in the Hosmer-Lemeshow (H-L) Test table and the Contingency Table for H-L Test:
|
Hosmer and Lemeshow Test |
|||
|
Step |
Chi-square |
df |
Sig. |
|
1 |
3.135 |
5 |
.679 |
|
Contingency Table for Hosmer and Lemeshow Test |
||||||
|
|
Hypertension = No |
Hypertension = Yes |
Total |
|||
|
|
Observed |
Expected |
Observed |
Expected |
|
|
|
Step 1 |
1 |
11 |
11.872 |
3 |
2.128 |
14 |
|
|
2 |
13 |
12.128 |
2 |
2.872 |
15 |
|
|
3 |
12 |
9.396 |
3 |
5.604 |
15 |
|
|
4 |
14 |
15.511 |
13 |
11.489 |
27 |
|
|
5 |
12 |
12.617 |
10 |
9.383 |
22 |
|
|
6 |
12 |
12.477 |
12 |
11.523 |
24 |
|
|
7 |
3 |
3.000 |
9 |
9.000 |
12 |
The H-L test compares the observed cases to the number predicted by the logistic regression model (expected). If the H-L goodness of fit test statistic is greater than 0.05, we fail to reject the null hypothesis implying the at the model’s estimates fit the data at an acceptable level. Well-fitting models show non-significance in the H-L goodness of fit test. An outcome of non-significance indicates the model prediction does not differ significantly from the observed cases.
State what the H-L test results mean for this model.
d. Use the save function to create the following new variables: Predicted Probabilities, Deviance Residuals, and Cook’s Distance. Evaluate the model using these variables and the following Scatter Plots.
Step 1. Go to Analyze ( Regression ( Binary Logistic. Click Save.
Step 3. In the Save window, check Probabilities, Cook’s, and Deviance. Click Continue.
Step 4. In the Logistic Regression window, click OK.
Step 5. SPSS will re-run the changes and the Output window will appear. Open Variable View. You should see 3 new variables: PRE_1 (Predicted probabilities), COO_1 (Cook’s distance), and DEV_1 (Deviance residuals).
· Create a Scatter Plot of the Deviance and the variable ID: Are there any outliers? What does this mean when evaluating your model?
Step 1. Select Graphs (Legacy Dialogs ( Scatter/Dots.
Step 2. Click on Simple Scatter then click Define.
Step 3. Click on Deviance value (DEV_) and move it to the Y Axis box. Click on ID and move it to the X Axis box. Click OK.
SPSS Output:
Be sure to discuss any outliers in this scatter plot. Explain what these large residuals could mean for your model.
· Create a Scatter Plot of Cook’s Distance and the variable ID: Are there any influential cases? What does this mean when evaluating your model?
Repeat the above steps 1-2. For step 3, transfer Deviance value (DEV_1) in the Y Axis box back to the variables storage box using the arrow.
Step 4. Place Analog of Cook’s Influence (COO_1) in the Y Axis box. Click OK.
SPSS Output:
Discuss the presence of outliers in this scatter plot and what this may mean for your model.
· Create a Scatter Plot of Deviance and the Predicted Probabilities. Is there anything in the scatterplot that could cause some concern in terms of you model?
Step 1: Return to Variable view and repeat steps for the above exercises to get to the Scatterplot menu. Click Reset.
Step 2: Move Deviance value (DEV_1) to the Y Axis box. Move Predicted probability (PRE_1) to the X Axis box. Click OK.
SPSS Output:
Discuss the presence of outliers in this scatter plot and what this may mean for your model.