Uregent 1
Regression Diagnostics and Model Evaluation
Regression Diagnostics and Model Evaluation Program Transcript
[MUSIC PLAYING]
MATT JONES: We've gone over estimating bivariate and multiple regression models, but one thing we haven't talked about up to this point are some of the assumptions of multiple regression models. It's very important to adhere to these assumptions to have proper interpretation of our models. These assumptions include linearity, independence of error, homoscedasticity, multicollinearity, undue influence, and normal distribution of errors. Let's go back to SPSS to see how we can test these assumptions and evaluate our models.
Let's go ahead and estimate a multiple regression model using respondent's socioeconomic status index is the dependent variable, respondent's highest education as an independent variable, and occupational prestige score as an independent variable. But this time, let's request some additional information to perform some diagnostics around our model.
Go to analyze, regression, and linear, since we are still using an ordinary least squares method. We'll scroll down and enter my dependent variable first, respondent socioeconomic index. My independent variables of occupational prestige and highest year of school completed. I want to go over to statistics and request some additional information. I will request collinearity diagnostics, and Durbin-Watson of the residuals.
Click continue. And we'll also click on plots and request a plot, my predicted against my residuals. I enter those and click continue. Lastly, I want to click on Save and request Cook's-- this is also called Cook's distance-- and my standardized residuals. Click continue. And once I click OK, the model will be estimated, along with some tests for diagnostics.
Looking at my first piece of output, the model summary, I want to pay attention to the Durbin-Watson statistic. The Durbin-Watson statistic has values from zero to 4.0. The Durbin-Watson statistic provides us with some information about independence of errors. A value of 2.0 for Durbin-Watson indicates there's absolutely no correlation between the residuals.
As a very general rule, the values below 1.0 and values above 3.0 are considered quite dangerous and one might expect that the model suffers from rather serious serial correlation. Moving down to the ANOVA, we see that our overall model is significant. We have our coefficients output where we can interpret our individual predictors, but as part of our diagnostics, we want to pay attention to VIF. This is the variance inflation factor.
©2016 Laureate Education, Inc. 1
Regression Diagnostics and Model Evaluation
As a general rule, values close to 10 and definitely above 10 indicate serious multicollinearity in the model. That means the independent variables have a high level of correlation between each other. We see here that the value of 1.4, for both of our predictor variables, are well below that 10.0 general rule. Therefore, we can assume that we've met the assumption.
We requested a Cook's distance, which tells us something about undue influence that is specific outliers one or the variables that might be causing undue influence on the model. They might have a significant impact. We can go to our Cook's distance and look at the descriptives on our residual statistics.
Again, as a general rule, Cook's distance values of 1.0 or greater are considered problematic and further diagnostics should be performed to evaluate for possible undue influence on the model. We see here that our Cook's distance values range from a minimum of 0.0 to 0.025, well below the general rule of 1.0. We can assume that we have no undue influence in this model.
After examining the Cook's distance, we can examine a histogram of the distribution of our errors. The assumption on multiple regression is the normal distribution of errors. As you can see from our histogram, our distribution is fairly normal. Therefore, we can conclude that we have met this assumption. Or, at the least, we do not have a significant deviation from normality. I should note that many modern statisticians see this assumption as of little importance to estimating regression models as it has little impact on the model.
Next, let's look at the scatter plot which provides us with information about homoscedasticity, or whether our residuals at each level the predictor are equal in variance. As we can see here, there is no discernible pattern with the spread of scatter. If our model suffered from heteroscedasticity, we would see a grouping of scatter at one end that funnels out into a discernible pattern, often looking like a trumpet.
If we double click our scatterplot, we obtain the chart editor. And I'm just going to go add a fit line to provide a reference line. Once I add this reference line, you can see that there's no discernible pattern to the scatter. There are a couple of areas in the middle where I notice that the scatter funnels out slightly but it certainly doesn't look like a funnel or a cone shaped pattern.
Scatterplot also provides us with some information about linearity as well. Remember that an assumption of regression is that the variables have some linear relationship to each other. Scatter, again, tells me that there is this linear relationship. If there was not a linear relationship, I might see a nonlinear relationship, in which the scatter perform a u shaped pattern.
Regression diagnostics is a difficult topic, especially once you start to understand all the different ways in which your model can go wrong. There are entire
©2016 Laureate Education, Inc. 2
Regression Diagnostics and Model Evaluation
textbooks written, and classes on regression diagnostics, and so I would encourage you to look at other resources and to view this video again. Lastly, don't forget to reach out to your instructor if you have questions.
[MUSIC PLAYING]
©2016 Laureate Education, Inc. 3