Psychology week five assignment
all right well hello everyone hopefully some folks will join in the upcoming minutes but i will go ahead and get started in the meantime so welcome to week five you all you are officially on the downhill slide in the course and so last Last week, you learned about correlation and regression, and this week, we're going to step up to multiple regression. So why do we need multiple regression? Well, we are interested in studying complex psychological phenomena, and those rarely have just one single cause. So typically, the behavior that we're interested in studying emerges from multiple influences. Additionally, our cognitive processes involve many separate components. And on top of that, social factors interact with our individual differences. So being able to account for those other factors in the same analysis is very valuable. So say, for example, you're interested in overall academic performance measured by GPA. Well, lots of things can impact that. So, study habits, sure. Test anxiety, sure. The amount of sleep you have, of course. No one variable is going to explain all of your academic performance. And there are other factors as well. So, multiple regression is really useful in parsing apart the effect of different variables on an outcome that you're interested in. so to layer on your learning from last week remember those kind of inferential uh statistical core concepts so remember that we are we never have access to the whole population so we are always making inferences about those populations based on our samples uh through the scientific method we are testing hypotheses about the relationships when we get the results, we work to understand whether or not we have statistical significance and the practical significance of that, looking at confidence intervals, effect sizes, that sort of thing. So, this week, we are expanding to handle multiple predictors simultaneously. So, same concept, but we're just adding in additional independent variables. So, we can do that a couple of different ways. We can do it lots of different ways. We'll cover a few. So, you can handle multiple predictors overall, either simultaneously or step-by -step. So, multiple regression is simultaneous and hierarchical is step-by-step. so say we had a research question what predicts academic performance well as we said earlier with this no one variable is going to explain this relationship so we need to move beyond simple correlations to understand how multiple predictors work together additionally we want to understand the unique contributions of each factor. We want to be able to piece that apart to investigate that complex interplay of variables. You know, do two of three variables have a greater impact on academic performance than another combo of two or three? And additionally, it allows us to control for confounding factors that we know are related but are not part of our primary question. So when we think about this research question, we first ask, what are the things that influence student GPA? So academic performance. Well, as we already said earlier, so things like study habits, test anxiety levels, sleep quality, these are things we know for sure affect student GPA. So then we have to ask ourselves which of those factors matters most and how do they work together? So multiple and hierarchical regression could allow us to study this. So the key questions in multiple regression are these four. So in this specific research scenario so one how well do these factors collectively so all together study habits test anxiety and sleep quality let's look at them all together first how well did they collectively predict gpa so academic performance and then what's unique about each factor so what's unique about study habits, about test anxiety, about sleep quality. Another important question to ask is, should we consider some factors before others? So the literature is going to tell us this. If we know that study habits account for most of academic performance, then we might consider an order, a hierarchical regression where we test the thing we expect to have the biggest effect first, and then we layer on. So should we consider some factors before others? You have to know your literature to know that. And then four, how much do these influences overlap with one another? So how much variance do they share? So when we look at the results of regression, we're going to be, you know, looking at beta values, but the most direct way to interpret is through R squared. That gives us the percentage of variance explained. So anytime you run a regression and a multiple regression in this instance, say that we did that and our model explained 45% of the variance in GPA. So when we combine study habits, test anxiety, and sleep quality, that explains almost half of the variance in GPA. So what does that mean altogether? This would be the R squared value for the overall model. So that would mean that 45% of the variance in grades is explained by the combination of study habits, test anxiety, and sleep quality, but 55 % of the variance would be due to other factors. So then we would need to consider the practical significance of our findings. So with multiple and hierarchical regression, when I am talking about parsing apart the influence of each individual variable, we want to look at the unique contribution. So we do that with semi-partial correlations, which you'll do for your assignment. So say that we found that test anxiety specifically uniquely explained 16% of the variance in GPA. So what that would tell us, what that one semi -partial correlation, just looking at anxiety out of this whole model, that would tell us anxiety's contribution after controlling for all the other factors, specifically meaning what's the effect of test anxiety independent of study time and sleep. And then we can actually discern the actual practical implications for academic interventions. Okay, so when you get the results from your regressions, students often get confused about the standardized versus unstandardized betas. So anytime you see that B, we call that a beta if it is this um squiggly looking uh symbol b that is the standardized value and we typically prefer this because it's comparable well i'm sorry that's wrong to say we compare we prefer this we use them both so the the standardized beta say that value was 0.40. This is useful because it's actually comparable across predictors and shows the relative importance of each individual predictor to the overall percentage of variance explained. So this is a scale-free comparison, and it's very useful. But now the unstandardized beta weights, So you can see this is the same value represented in standardized versus unstandardized. So the unstandardized beta is negative 0.25 points GPA per anxiety unit. So this has more of a practical meaning. The unstandardized values are useful for practical interpretation and predictions because it retains the original units. But that standardized beta is more useful for comparisons. Okay, so now sometimes the order of your multiple regression is theory-driven. So if we dug into the literature, we would find most likely these three factors in this order. So known factors first. We know that the amount of time you spend studying is directly related to your academic performance. There's a direct positive correlation there. As the number of hours you study increases, your grade tends to increase. So it would make sense to put this in a hierarchical regression where we would first account for study time. And then we could add in psychological variables like anxiety after that, and then environmental variables like sleep. So think of hierarchical regression like telling a story through building understanding. So you start with the obvious predictors, and then you add complexity. And at each step, when you add a new variable, the hierarchical regression is testing the improvement at each step. So say we wanted to predict GPA and we were going to do a hierarchical regression. So the foundation here, this pinkish reddish bar at the bottom would be the foundation. We would expect study habits to have it the biggest impact. So basic academic behavior would be step one. At step two, we would add test anxiety to measure those emotional or cognitive factors. And then at step three, we would add sleep quality for that physiological or environmental factor. So if you, when you do this in your software of choice, you are going to get your output. And so what you want to look, oh, excuse me, jumped at, what you want to look for, you first jump to the overall model fit. The R squared value will be a decimal value and you move the decimal two places to the right. And that gives you the percentage of variance explained by that unique contribution um so we we're going to look at that r squared value we're going to look at significance tests so we're going to look for the p value is it less than 0.05 we're going to look at the beta values for those individual contributions at each step along with the p values we're looking for that change in between steps so each new variable that we add is it improving the model? Are we getting closer to a more accurate prediction? And then all along, we have to make sure that we're meeting those assumptions of regression and demonstrate that we've checked those assumptions and are analyzing accordingly. So multiple regression helps you understand complex behavior, account for multiple influences, ultimately make better predictions that can inform better interventions. Okay, so we talk a lot about the assumptions in all of the statistical tests that we conduct. And so for linear regression, there are these six major assumptions that need to be met. So, the first is linearity, meaning the relationships between predictors and independent variables should be linear or approach a straight line. So, this also applies to the collective relationship of all predictors with the dependent variable. Two, there should be independence of observations, meaning each data point should be independent of the others. no repeated measures or clustered data without special handling of that. Number three is homoscedasticity, which is super fun to say. That refers to equal variance of residuals, not raw values, but residuals, the error across those predicted values. So, the spread of those residuals should be consistent to meet that assumption of heteroskenasticity. The residuals should approximate a normal distribution. So, the residuals or errors should be normally distributed. Again, not the raw variables, but they're residuals. We want to avoid multiculinarity. So, no perfect multiculinarity. meaning that the predictors, when you have multiple predictors, you don't want them to be too highly correlated with one another. So, the variable inflation factors or the VIFs ideally need to be under 10. You might get some editors or reviewers argue that they need to be under 5. And then your tolerance values need to be above 0.1. Finally, you need to make sure you don't have any significant outliers. So look for influential cases that could distort the results, use Cook's distance, leverage values, et cetera, to handle those outliers if they are true outliers. All right. So you're going to have to do this in Jamovi. So I just want to will give you an idea of your options. You can also do it in SPSS, but those of you using Jamovii, I just want to tell you what you're looking at when you click on assumption checks. So the autocorrelation test, that tests independence of observations using the Durbin Watson statistic. This is typically important if the data might be time dependent. The culinarity statistics that test for that multi-culinarity provides those VIF and tolerance values. So, this is absolutely essential when your predictors might be highly correlated. The normality test that lets you test if your residuals are normally distributed. It's the statistical test for normality. And this is well complemented by a visual QQ plot. Speaking of that, the QQ plot of residuals is a visual check for normality. It shows you your plots expected versus observed residual values to see if you're meeting that normality assumption. So a straight line indicates a normal distribution in that QQ plot. Residuals plots, that's going to test homoscedasticity, show you the pattern of results, and help you identify potential problems. And then finally, this data summary, Cook's distance, that helps you identify potential outliers and any potentially problematic cases that might be important for your data screening needs. all right so let's touch on your assignment now so you'll be using the same data file this week that you used last week so it contains those same variables so this is how they're labeled but what those actually mean is number of visits to health professionals number of physical health symptoms number of mental health symptoms and a stressful life events score okay Okay. So for part one, you're going to explore the data. So create visualizations showing the relationships between doctor visits and each predictor. So you're going to need to think about the things that you've been doing all term. So what types of visualization do you need to create for which types of variables? So think through that for me. Examine those potential univariate and bivariate outliers. So I want you to document any concerning patterns that you see and tell me how you would handle them, but do not actually remove any data. Keep all data points in the analysis, even if you do identify outliers. And then I want you to calculate and interpret the correlations between variables. So part two, the standard multiple regression. If you go to analysis and regression and then linear regression, you will get a screen that pops up that looks just like this. You will set visits to the doctor as your dependent variable, so time DRS, and then move all the other variables to the covariance box, request model fit measures and coefficient statistics. And then for part and partial correlations, I'm going to give you some more specific detail about that. You'll actually need separate analyses for each predictor for the partial correlation bit. So kind of put a pin in that just for a second. But in your output from the standard multiple regression, I will be looking for the R value, the R squared, and the adjusted R squared, the ANOVA results. So that's the omnibus test to see if the overall model fit is significant. And then the beta coefficients with their significance tests, the part and partial correlations, which I'll dig into, And then at least one visualization supporting this analysis. Let me dig into the hierarchical regression. So after the standard multiple regression, I want you to do a hierarchical regression. So what you need to do here, it's in the same linear regression options. in Jamovi, but before you get to the analysis, you need to determine a theoretically justified order for entering your predictors, document your reasoning for this order with at least one journal article to support that assertion, and then enter the variables in sequence, examining the changes at each step. All right, so required analysis here, the R-squared change at each step, the significance of each change, the final model coefficients, and I want you to then compare the results of a hierarchical regression with those earlier results from the multiple regression. So just as a note, to do this hierarchically, you will need to do it in the model builder. So this is block one. We're just looking at physical health. I mean, you'll put them in the theoretical order that makes sense, but I just put them in the order here. So you just add one new block at a time. Each variable is its own block in hierarchical regression. okay so for the write-up your results section make sure you talk about your data screening so describe the distributions discuss any outliers or patterns and then give a summary of those bivariate relationships from the correlation analysis part two the standard regression results, give the overall model evaluation, talk about the individual contributions of each predictor, interpret the effect sizes, and then talk about the practical significance from those findings. Then you're going to do that again only with a hierarchical regression analysis. So, this needs to include a justification for the order of your variables that you input those blocks. I want to see the changes at each step and then that final model interpretation. And then finally, for the visual presentation, make sure you include any relevant plots or figures from your data screening or what have you. make sure you properly format any tables that you include. And for any figures, tables, anything, make sure that involves clear labeling and titles. So don't use the variable names from the code book, like time DRS. The chart shouldn't say time DRS. It's visits to the doctor. So, think about trying to present your results to a reader who didn't have access to the data set. So, they don't know what time DRS is. So, make sure that you're using common language, okay? All right. Then, as promised, I want to dive in a little bit to the partial and semi-partial correlations. So, if you're using SPSS, it's going to provide those for you with a single checkbox. Jamovi does require a little bit more of a detailed approach, but it can actually deepen your understanding of these concepts. And I think there are a couple of ways to go about this, but this is how I did it. So, first, let's understand the concepts. a partial correlation shows you the relationship between two variables after controlling for other variables. A semi-partial or part correlation shows the unique contribution of a predictor to the dependent variable. Okay, so without that, that same control. All right, so how do you do this in JMOBI? So for each predictor, you're actually going to need to run a separate analysis. So you'll go to analysis, click on regression, and then partial correlation will be an option. Under correlation type, once you get there, you'll choose semi-partial. So you don't need to check report significance, and you'll get this pop-up window. So for each analysis in the variables box, so start off by putting visits to the doctor and physical health in the variables box, and then put the other two predictors in the control box, and you will run that analysis. And then you need to, so you're going to need to run three analyses total to capture all of those relationships. So these screenshots that I have here, I'm showing you, this is what you'll, these are all three of the analyses that you'll need to run. So the first time DRS and physical health are in the variables box and mental health and stress are in the control. Then you will switch to times at the doctor with mental health, with stress and physical health control. Finally, you will do times at the doctor with stress and then control for physical and mental health. So that will get you all of the semi-partial part and partial correlations that you need. Okay, so how do you interpret those results? So the coefficient, this coefficient value in the output is your semi-partial correlation. You square that value to get the unique variance explained. And then you can compare those values to understand each predictor's unique contribution. So, an example interpretation of this specific relationship says the semi -partial correlation between physical health and doctor visits controlling for mental health and stress was 0.33, indicating that physical health uniquely explains 10.89% of the variance in doctor visits. Excuse me. All right. So I think that should give you enough to get you going for your assignment. And then for the week five discussion, I want you all to read over the chapter on effect sizes. So don't get bogged down in formulas or the methods of how to calculate effect sizes or anything like that. Just understand it conceptually. And then read Graviter and Walnow 8.5, Concerns About Hypothesis testing, measuring effect size, that Cohen article, The Earth is Round, if you haven't run that already, and then describe the importance of effect sizes in the field of psychology. All right. And finally, sadly, you also have a quiz this week. So this class is a lot of work, but you're all doing great, and I believe in you. Stop sharing and see. Nope. No one made it in today, but that's all right. If anyone has questions, please reach out to me. I am more than happy to answer your questions.