Psychology week one assignment

profilebmarlurer8
en-Rec-Sep220253_04PM-QuantitativeResearchII.mp4.txt

and share. All righty, you all. So, welcome to RSM 801, your Quantitative Research II course. I am Dr. Chastity Ratliff, and we've got a lot to cover today, so I'm not going to go into a lot of details about my background, but this is just a visual depiction of my education journey from Southeast Missouri State University and doing studies on animal behavior learning to my later graduate school education, where I focused on psychology and law and societal threat, prejudice, those types of things while I was working on my PhD. So I've been doing research for a lot of years. I have a few, I believe I have four peer-reviewed journal publications from the research that I've done, and I've been teaching at the undergraduate and graduate level for about nine years altogether now. Okay, so I like to start off my stats and methods courses with a little just pause to try to get you all to really start thinking differently and try to shift your approach to learning in these kinds of classes. Because I get a lot of students who do really well in their theory classes that come into these stats and methods courses and struggle when they lose points or don't hit all the marks. So I like to try to take a few minutes to explain the difference and what you'll need to do to be successful. So your theory courses test your ability to understand and analyze concepts, but methods and statistics courses require you to actually apply those concepts to solve novel problems. So the feedback that I give you all, and I spend a lot of time giving you feedback, and I hope that you use it, it focuses on precision and accuracy rather than those broader theoretical interpretations. So my feedback is really going to focus on how you are applying those things within the bounds of the scientific method and the research process. So this is also a new type of academic writing. Results sections require very precise standardized formatting. And a lot of students find that tedious or annoying that you have to put this thing in italicies and use so many decimal points. And we think those kinds of things don't matter and you might not pay that much attention to it. But it really does matter for the reader who is trained in these types of fields to be able to know how to read and understand what you're trying to communicate. So statistical writing emphasizes clarity and accuracy over elaboration. Forget all that stuff about trying to fluff up your papers to hit the page requirement. We want clear, accurate result sections instead of fluff. So APA style guidelines must be followed exactly. And the sooner that you just get in the habit of writing in that way, it just becomes second nature and it improves the overall quality and readability of your writing. So this type of writing overall is more technical and concise than what you will find in theoretical courses. So essentially, learning statistics, you all, is like learning a whole new language. You're learning a new way to communicate your research findings. And for this reason, regular practice is essential. You cannot cram for statistical understanding or substitute your own understanding with that of a type of AI or something like that. So mistakes are so normal and they are actually a very valuable part of the learning process. us. And I want you to make mistakes early on and listen to the feedback that I give you and incorporate that because those are such valuable learning experiences. So, like all of these kinds of classes, the skills that you build each week are going to build on each other throughout the whole term. As I've mentioned, the feedback I give you is going to be detailed and specific. I tend to focus on the issue that needs to be addressed because I have a lot of grading and I want to zero in on the things that need work. So please don't take that personally. My focus is on helping you master the technical skills. So my feedback to you is going to address whether you actually completed the assignment or the task as required. So on execution, rather than on how much you tried or if you understood it, but you didn't demonstrate that or communicate that, I have to address that in my feedback. But trust me, I'm spending all this time giving you all feedback and guidance now to make you stronger down the line when you face your comprehensive exams and your ultimate dissertation. So, some things that you can do to be successful is do not underestimate the amount of time that you need to focus on this class. It's a lot of work, and if you struggle with this kind of material, then I would say you need to plan for about 15 to 20 hours per week of dedicated work. I know that feels overwhelming, but if you're a student who struggles, I'm especially talking to you. So start your assignments early so that you have time to ask me questions and make revisions. Make sure you're in the class four or five times per week at minimum, and it will do you well. It will serve you well to schedule regular practice sessions with your statistical software of choice. Now, you have to be an active learner on your end. all of these doctoral classes are only eight weeks. So, but you are getting all of the material and learning that you would in a 16 week course. We're just cramming it into these eight weeks to let you all work on your PhDs from wherever you are and still work. So you have to be active on your own. We only get one live hour together per week, but I am going to spend that time working. I'm going to give you a little lecture each week and then try to walk you through examples step by step. I encourage you to create your own practice problems. I've had students in this course in the past, the most successful students, got together with their classmates and formed study groups. They were immensely helpful, and I could always tell the students who were working in the study group. one of them would ask questions. They would work together to fill in the pieces that they missed. So cannot encourage that enough. Rely on each other. I encourage you to maintain an error log to track and learn from your mistakes so that you can correct that in later work. So all building on that, use my feedback effectively. Like I said, I take a lot of time to give it to you, not to beat you up, but to improve your work. So please review all of the feedback carefully before starting new assignments. Keep a running list of common errors to check for in your future work so that you're not making the same mistakes over again. And always please ask for clarification if assignments or feedback are unclear. So on that, what you can expect from me throughout this term, lots of support you all. I am, I'm here to support you. I want to see you succeed. I'll be holding weekly live sessions to review the concepts and go over the assignments. There are clear rubrics for all of the assignments. I suggest you look at those before submitting your work. I am always available for individual meetings. All you have to do is send me an email with three days and times that you're available. Just save that time in the beginning if you want to meet with me or talk to me just say hey i'd love to have a meeting i'm available monday from you know whatever to whatever tuesday and wednesday send me your availability and then i'll get back to you uh within 24 uh work hours monday through friday about a meeting we'll make it happen and i'm also absolutely willing to give you additional resources and examples if you find yourself stuck. I just needed to reach out to me. So when you reach out, you can expect responses to emails within 24 hours on weekdays. I will make regular course announcements and try to give you very clear assignment instructions and explicit grading criteria through the rubrics. My teaching approach is supportive. I try to give you all step-by-step instruction for new concepts with with examples. These assignments are scaffolded, meaning they build upon one another progressively, and overall the focus is on practical application. Now, what do I expect from you all? I really expect you to uphold academic standards, meaning that your work should demonstrate graduate level attention to detail, precise adherence to APA format, clear documentation of all of your statistical procedures, and professional academic writing. So on that note, I do not want to see AI generated work. You can use AI in early stages of research for brainstorming and things like that. And at the end to kind of polish your work, but the bulk of the work needs to be yours. I want to see the messy struggle because I've seen students come through this program enough now to see those who are over -relying on AI get into upper level classes, into comps, and really struggle because they didn't take the time to really learn the earlier material and it starts to show up. And so that's unacceptable. And I will be on the lookout for that and really do not want to see you all using AI generated work in this class. So I want you to authentically engage with your classmates in discussions. I need you to be proactive with me and communicate about any challenges you might be facing. I expect timely submission of assignments, and I do not accept late work without reaching out to me. So if before the due date comes and you realize you're not going to be able to make it due to some circumstance, shoot me an email. If you shoot me an email before it's due and let me know whatever's happening, give me just a little cliff note summary of whatever is happening, I most likely won't take any points off. But if you don't email me until after the fact, I'll go from there. But if you don't email me at all and you just submit a late assignment, I'm just not going to grade it. So that's kind of the deal we have here. I will absolutely be flexible and understanding and kind about whatever you have going on, but you have to communicate that with me, okay? All right, so that timely submission, actively engaging with my feedback, and then this class is challenging for a lot of students because you have to develop some technical skills for the analyses. So, you've got to get some proficiency going with the statistical software. You need to be able to learn to follow those statistical procedures exactly, to be able to clearly present your output, and to accurately interpret your results. All right, so on that note, as I've mentioned, you are free to either use Jamovi or SPSS. If you choose to use SPSS, you do need the grad pack. The standard does not have everything you need to do all of the analysis. So you need the more expensive package. But if you choose to use Jamovi, it is free. That's what I will be doing all of my demonstrations in, and the add-ons are free in it as well. So, it's capable of doing everything that we will do in this class. It is what I will be using, but if you're more comfortable with SPSS, you're free to use that, and there are some additional resources in the course for SPSS users. Okay, so now I'm going to jump into week one. Week one stresses students out in this class, you all, just as a heads up, because we jump right in. We're essentially jumping right in at the point where if you had your data, if you had your results, you would dive in and start trying to understand the quality of your data okay so that's where we're starting this week we're building on foundations from quant one the focus here is on advanced data handling and analysis so we do jump right in but this is of critical importance for everything that comes after particularly your dissertation research and the emphasis throughout is on real world challenges in conducting research. All right. So when we talk about data quality, the first two things we need to look at are outliers and missing data. So let's start with outliers. What is an outlier and how do you determine if you have them? We'll see these little dots out here that don't follow the rest of the pattern. This is a univariate outlier, the multivariate outlier. So univariate varies in terms of a single variable and a multivariate outlier is one which involves more than one variable. So what you're looking at here are scatter plots. These are what we use to one of the ways to determine if we have outliers looking visually. Another way is to use box plots. We can also use Z -scores, which we will do. So after you've gone through these techniques and identified your outliers, then you have to investigate potential causes for those outliers. So, was it an error in data collection or was it a valid but unusual data point? So, how do you handle outliers if you find them? Well, first, you check for error. So, any typos or data entry mistakes could be corrected or removed. So, for example, if someone said they worked 150 -hour work week, that clearly isn't valid. We would likely remove that. We have to consider the theoretical relevance. So, is the score, is that outlier unusual but meaningful for your population? So, for example, a very high burnout score might represent an actual at-risk teacher or employee. It can be an outlier, but still valid and meaningful. so to handle these you want to use transparent common rules so you want to keep the outlier if it is within the plausible range of your measure you want to transform it through a log transformation for example if your distribution is skewed another technique is to Windsorize or to cap the extreme values if the outlier is legitimate but overly influential. It's skewing your results. But remove it only if you can clearly justify why that case doesn't belong to your study population. In general, we try to keep outliers and you must if it represents a valid score in your sample. Finally, you need to document your choice of how you identified the outlier. So, a Z -score greater than or less than 3.3 box plots, scatter plots, or another way. State what you did and why in your results and methods section. Okay, so we've got outliers handled. Now, let's talk about missing data. um the general steps for analysis with missing data are to one identify the patterns and reasons for the missing data and recode correctly if necessary you need to understand the distribution of missing data and finally decide on the best method of analysis this. So going to step one, understand your data. Why is it missing? Is it due to attrition? So some social or natural processes that cause people to drop out of your study. So maybe you were studying students and some of them graduated or dropped out. Or if you're studying an elderly population, maybe some of them passed on. What are the reasons for the missing data? It could be due to a skip pattern. So, for example, certain questions only ask respondents who indicated they were married. So, sometimes only certain participants are asked to answer a question. So, that's a legitimate skip pattern. It could be intentional. So, missing as part of the actual data collection process where only one condition sees one set of questions and another sees a different set. It could be random data collection issues, or it could be respondent refusal or just non-response. In other words, they didn't answer the question. Step two, so we've dug into it. We've kind of figured out what we think the reason is for why the data is missing. And then we have to look at that probability distribution of missingness. So considering that probability is asking ourselves in looking at the data, are certain groups more likely to have missing values? So for example, are respondents in service occupations less likely to report their income with TIPS. Another way to consider it is are certain responses more likely to be missing? So, for example, respondents with high income are less likely to report their income. So, we need to see if that's the case. And then, so certain analysis methods, those assumptions that we check before we do our analyses, this is part of it because certain analysis methods assume a certain probability distribution. So if your data doesn't meet that assumption, you need to do something about it. So, missing data mechanisms, we can just break this down into three categories. So, data that's missing completely at random, MCAR data. So, this is when the missing data doesn't depend on either observed or unobserved values. So, for example, lab equipment that was recording the time, randomly malfunctioned, causing some missing measurements. That is missing completely at random. Missing at random is when the missing data does deserve on observed variables, but not unobserved values. So, for example, service workers are less likely to report their income and that is related to their occupation type so that's missing at random but it's not completely at random because it is tied to part of what we're studying and then there is missing not at random so this is missing data that depends on the unobserved values themselves So, for example, those high-income individuals that are less likely to report their income, they will just not answer. Okay. So, when we're exploring missing data mechanisms, you can never be 100% sure about the probability of the missing data because we don't actually know the missing values. So there are things we could do, like we could test for missing completely at random using t -tests, but that's not actually totally accurate. So many missing data methods assume that it's missing completely at random or missing at random, but in reality, our data are often missing not at random. So, we also need methods specifically for that, including a selection model like Heckman or pattern mixture models. So, we don't have to go too far into the weeds here, just giving you an overview. Because this is what we really need to dig into. How do you deal with your missing data? Well, you have to use what you know about why the data is missing and anything you know about the distribution of missing data. Are there any meaningful patterns? And then you can decide on the best analysis strategy to get you the least biased estimates. So your options are deletion methods like list -wise or pair-wise deletion. There are single imputation methods, so mean or mode substitution, the dummy variable method, or single regression. And then there are model -based methods, so maximum likelihood modeling and multiple imputation modeling. Let me dig into those deletion methods. So list -wise deletion is when you're looking at one complete case and doing, I'm sorry, so list-wise deletion, you would take out all of the responses. And pair-wise deletion, we're looking at pairs of responses there. So a little more concrete example here. So with list -wise deletion, we're only going to analyze cases with the data available on each variable. That's list-wise. So looking here, we would not analyze at all these first two measures because they don't have data on each variable. So we would essentially ignore all of that data. So the advantages of that are that it's simple and it gives us easy comparability across analyses. But big disadvantages, you're losing data. So it's reducing statistical power, you're lowering your sample size, you're not using all available information. And with this kind of list-wise deletion method, if the data is not missing completely at random, then your estimates may end up biased. So as a note, list-wise deletion often produces unbiased regression slope estimates, as long as missingness is not a function of that outcome variable. So here's an illustration of pair-wise deletion. so this is saying go ahead and analyze um all of them in which the variables of interest are present so that would be saying in this example if we were only looking at eighth grade or let's say we wanted to look at eighth and twelfth grade test scores but we wanted to to eliminate gender. So then we could analyze the third, fourth, and fifth cases. We could actually go ahead and analyze the sixth case because gender is missing, but that's not one of the variables we're looking at. So that would make sense for pairwise deletion. Hopefully that makes sense. You all, if not, feel free to interrupt me at any time if you have questions. Dr. Ratliff, I do have one question. Yeah, please. So is it either, is it one or the other, or can you use a combination of both with your data? You'll want to choose one way or the other. Okay. Yeah. So you will, you're going to spend a lot of time when you get your hands on data, you're going to want to dig in and try to figure out why there are any outliers or missing data and then decide how you're going to go about handling it so list wise is saying like they either completed everything or we're not taking them and pairwise is saying as long as they completed everything for this analysis we're going to keep it okay thank you yeah so advantages of this it does keep as many cases as possible and uses all information possible within each analysis But the disadvantage here is we can't make direct comparisons across analyses and the data set because your sample is slightly different each time. So, say one time we wanted to look at gender and eighth grade math test scores here. That would give us a certain number of pairwise options to work with. And then if we wanted to look at gender and 12th grade math score. So, that's the disadvantage because we end up with a slightly different sample each time by using pairwise deletion. Okay. So, some other methods would be the single imputation methods. So, mean mode substitution, dummy variable controls, and conditional mean substitution. So, I'll go through these with you. mean mode substitution replaces missing values with the sample mean or mode for that variable and then you would run the analysis as if all the cases were complete so advantage of mean mode substitution is that you can use a complete case analysis method but the disadvantage here is that it reduces variability. When you have missing data and you just plug in the mean or the mode, that reduces variability, which is this flattening line that you're seeing here. Reduces variability and weakens the covariance and correlation estimates in the data because it's ignoring that real relationship between variables and assuming the mean. Now, dummy variable adjustment is another option. This is where we create an indicator for the missing value. So, we would say one, we would enter a one if the value is missing for the observation, and a zero if the value is observed for that observation, for that question. Then we would impute the missing values to a constant, such as the mean, and include that missing indicator in a regression analysis. So doing this, using the dummy variable adjustment, does use all available information about that missing observation. But disadvantage is that it absolutely results in biased estimates and is not theoretically driven. So the dummy variable adjustment is really the best when the value is missing because of a legitimate skip. So when it was legitimately skipped, then it makes sense to use the dummy variable adjustment. Otherwise, you run the risk of bias results. All right, now there's regression imputation. So what this does is replace missing values with a predicted score from a regression equation. So this is a little more precise and complex. Advantage here is that it's using information from the actual observed data. But the disadvantage here is that regression imputation tends to overestimate the model fit and the correlation estimates, and it weakens variance. So, I'm sure you're picking up on it. With each of these methods, you kind of have to weigh your pro and con options. Okay, so let's move on to the model-based methods, which are more preferred when possible. So, maximum likelihood and multiple imputation. So, model-based methods like MLE, like maximum likelihood estimation, identify the set of parameter values that produces the highest log likelihood. So, we don't have to go too deep down the math theory rabbit hole here. So, just try to understand the broader concepts. But the maximum likelihood estimate is the value that is most likely to have resulted in the observed data. So conceptually, this process is actually the same with or without missing data. So advantages of maximum likelihood estimates are that they use full information, so both complete cases and incomplete cases, to calculate the log likelihood. So this gives excellent unbiased parameter estimates when we have missing completely at random or missing at random data. Disadvantages, there have to be some, are that standard errors tend to be biased downward in these models. However, those can be adjusted by using an observed information matrix. X. Okay, so essentially just trying to give you a little illustration here with this image. So here they assumed the general Gaussian bell curve shape, but had to infer the parameters which determine the location of the curve along the X axis, as well as the quote fatness of the curve the data distribution could like could look like any of these curves it's all being estimated and so maximum likelihood estimation essentially would tell us which curve has the highest likelihood of fitting our data okay all right and then there's multiple imputation which if models like that might make your head spin a little it's okay um so imputation is when data is filled with imputed values using a specified regression model so then that step is repeated however many times m times resulting in a separate data set each time then analyses are performed within each data set the results from each of those are pooled into one estimate for the values for the missing data. So, advantages of this are that variability is more accurate with multiple imputations for each missing value. So, this type of modeling considers variability due to sampling and variability due to imputation. So kind of gold standard here. You get the highest quality here, but the disadvantage is cumbersome coding. When you're running that many analyses, you have to impute each of the data sets, analyze each of the data sets, pool them, and then make final estimates. So there's a lot of cumbersome coding, which leaves room for human error when specifying the models. Okay, moving right along. So that's your little mini lecture of the material for this week. And now I want to walk you through your tasks, the assignments that are due from you. So, you have a discussion post that's due for week one. Your prompts are what is the most important thing to determine when understanding missing data? What are the advantages and disadvantages of common missing data methods? And that Osborne and Overby paper that's linked here in your materials discusses outliers. So what are the main causes of outliers? And then provide me with an example of a univariate outlier. So a single variable in psychology and those readings that are attached to your discussion assignment that should give you everything you need there. And then the scary part, you all, it's not so bad, I promise. but you're going to dig right into some data. You have a result writing assignment. So you'll need to download the actual assignment document as well as the data file. So the .sav is for SPSS and the .onv is in Jamovi. So I really want to stress to you all how important it is that you spend most of your time writing up the results section of the assignment. This is the most important piece the format is very important so take the time to polish it make sure your text tables and figures all strictly follow apa format myself and your other professors create these assignments to give you practice in doing the things you'll do later so in running analyses and statistical software as well as writing up the results so we're always hoping that these assignments provide you with the practice you need for capstone dissertation and other research. And I highly recommend keeping a journal or a diary of how you went through these assignments so that you can come back to it later. I promise it'll be helpful. All right. So I am going to, I have screenshots here and I'm walking you through everything that is required of you. I've just blocked out the actual results, okay? So, for your assignment, for the continuous variables that you're going to be working with, so age and household size, you will open Jamovi and load your data set. You will click analyses, exploration, and then descriptives, you will get a box that looks like this and you will move over your variable. So you just click it and hit the arrow over. So you'll move over age and household size. Under this little statistics dropdown, you'll check mean and median for distribution. You will check standard deviation minimum maximum distribution we want skewness and kurtosis with their standard errors the sample size and missing counts so um as a note i want you to examine if the data is approximately normal by looking at the skewness and standard error ratio so a value greater than two indicates a significant skew. So, this is a model of what to follow for the first part. Now, for your categorical variables, so marital status, education, and work. In that same descriptives window, you'll add your categorical variables to that variables box, and then you'll check this little frequency tables box here, this little checkbox, you'll check that. And that will provide counts, percentages, valid percentages that include missing data, as well as cumulative percentages. So when you have continuous variables like we had before these continuous variables, you repeat or you report measures of central tendency, like the mean, median, standard deviation, all of those kinds of things. But for categorical variables, it doesn't, there's no mean of a categorical variable. So those are just categories. And so we have to look at those frequencies, the descriptives that we look at for categorical variables are the counts and percentages. All right. Now this is where this is so useful. This is very helpful for you all in the future. Often what you need to do is transform some of your scale items. So on a depression scale, for example, some of the items will be reverse scored. So So if most of the questions say things like, I feel very sad most days, I do not feel motivated to get up most days, and that's the way most of the scale goes, then some of the items say, I wake up feeling ready to confront my day. That would be a reverse scored item. So we need to fix that in our data set before we move forward. Does that make sense, you all? I hope so. Stop me if not. Okay, so here I'm just trying to walk you through the steps that you'll need to do in Jamoge. So you'll click on the data tab at the top. You will click this compute button and then for each item to be reversed. So you all, there are one, two, three, four items that you need to reverse score. So you need to name the new variable. So in this box here where it says CES4REV. So I've named the CES004 as CES4REV for reverse coded. Then under this formula box, you will enter the formula three minus, because it's based on the actual scale item. So you'll enter the formula three minus the original item. So for here, I said CES for reverse equals three minus the original variable name, CES underscore 0004. Okay. And then doing that, entering that gives me that reverse scored variable. And you'll need to do that for each of those other variables. So each of those four. And I just walked you exactly through the first one. So then you will repeat that process with the other three variables from your assignment. Okay. Once you've done that, you need to check that it worked, that it's accurate. So look at the data view to verify the transformations. Original zeros should have become threes. Original ones should have become twos. Twos became ones and threes became zeros. Okay. So when you get there, you can spot check a few rows to ensure that it worked correctly. You want to make sure that they actually all did get reverse scored. okay so now after you've reverse scored those items it's time to actually create your depression scale so when you have multiple item scales with like 15 questions after you adjust the reverse scored items then you're going to create the total depression scale variable so you will click compute from the data tab you will name the new variable cesd tot for total and then note here that your formula you're going to be adding together items but you need to make sure you include the reverse scored items and not the original items so i've given you the exact formula here with the reverse items and then the others. Feel free to rewatch and pause the video later and type that formula in exactly. That should work for you as long as you've reverse scored everything accurately. All right. So then at this point, you're going to, after you've created the scale, you want to document the number of missing responses for each case, calculate the total score. So, if the participant answered at least 80% of the items, so 16 out of 20 of the depression scale items, you want to make note of any cases that do need to be excluded due to too many missing responses, and then document your decision process in your write-up. Do it professionally as if you were reporting to a journal. All right. So, part three, you're going to assess the quality of your data. You will run descriptive statistics for the scale scores. So, to do that, you'll go analysis, exploration, descriptives will be a pop down. For the total depression for CESD tote and media use, I would like for you for both of those to look at their central tendencies, so mean and medians, the dispersion, so standard deviation, and the interquartile range, the shape of the distribution, so look at skewness, kurtosis, and the range, so the minimum and maximum for those continuous variables. then you're going to do your outlier detection so for both of those same variables you're going to click the compute button in the data tab and then you want to name your new variables systematically so we're creating z scores to check for outliers so in this formula builder box here You will click Functions and then choose Statistical and select Scale. Your formula should look like this. So, Scale with the variable name in parentheses. So, here I'm doing Scale, C-E-S-D-T-O-T. So, note that the Scale function automatically standardizes your variable by subtracting the mean and dividing by the standard deviation. thus creating z-scores. So you'll need to, I've given you the exact formula for the depression scale to create a z-score variable there. You will need to also do the same for the media variable, the media use variable. You need to create a z-score variable. And then for each of them, you want to verify that it z-scored properly. So the mean should be approximately zero. The standard deviation should be approximately one. To check that, you'll run descriptives. So, you'll do mean and standard deviation on those new Z-score variables. Okay. All right. So, once you've done that, in the data tab, you're going to click filters. and then for outliers beyond 3.3, you'll enter this formula. So negative 3 .3 is less than the Z score of CESD total. So we're looking for values in between. So we don't want anything less than or greater than this 3.3 so note here that you use z on the original variable not on the computed z score variable okay so don't check for outliers in the z score variable that's make sure you're actually using the original variable and then rows that are filtered out are the outliers and they'll be grayed out in DataView. So, it makes it very nice. It's a real quick, easy thing. So, then after you've done that, you want to create visualizations of your outliers. So, go to Analysis, Exploration, Descriptives, add your original variables, and under Plots, select box plot, QQ plot, and histograms. In your write-up, you're going to report the number of outliers identified, their specific Z scores, whether they appear to be valid data points or potential errors, and then your justified decision for handling the cases, whether you would keep, transform, or remove them. For your missing data analysis, you want to document the number of missing cases, calculate the percentage missing, look for patterns. Are certain variables more likely to be missing? Are there patterns across related variables? And then develop your handling strategy. So case -wise deletion, mean substitution, multiple imputation, what have you, and then document and justify your choice. For your write-up, you're going to include sample characteristics, so the sample size and demographics, descriptive statistics for age and household size, frequencies for the categorical variables. You're going to talk about your scale preparation, so your reverse scoring process, how you handled the missing data, the final descriptive statistics for that depression scale, and then you're going to talk about the characteristics of the distribution, how you identified and handled outliers, and any missing data patterns and the solutions. So, there are also some additional notes about APA -style reporting. All right, and then we're running out of time here, but we do get participation questions every week for listening to the live lecture. And so for week one, your participation question is this, what are you A, most excited about for this term, and B, most terrified about for this term? So just share your honest feelings with me at the start of the term and get those easy participation points. All right. Any questions? This is you all like poor grad students doing okay. Okay. I am going to come back out here and stop sharing. So good to see a lot more people joined us. So unfortunately, because of the week that we're having we will normally have live sessions on mondays and i'll have lots more time after but i will have i only have a few minutes until i have another live starting for my other class so if you have questions please unmute yourself and ask everyone it is hi dad direct this is lisa hi how are you i'm good how are you i had trouble getting um in so i joined a couple minutes later when you started your lecture but i did have a question okay when um i just got done completing obviously the first you know course the quantitative research with dr tudiz and dr tudiz informed us that we should be prepared to utilize the same textbook for this course now i did see that there is in their textbook of the statistics um and behavioral sciences by graviteer and i order my textbooks online through the kaiser university online bookstore and they unfortunately do not have um the textbook so that i believe this class it's the bordens and well i might be mixed up i'm teaching two different classes so will you make will you email that to me lisa just shoot me an email and actually everyone if you have i will i will absolutely answer any questions that you have i know everything feels a little overwhelming in week one but email me all of your questions and we'll take the time to answer you and then see you all next monday at live where we won't have this time crunch at the end okay thank you i can do that thank you thank you all i'm really excited for this semester i see a lot of familiar names so all right you all i'm glad to be with you and sorry to rush out this week one please reach out with all of your questions i want i want all of your emails tomorrow morning when I start working again. Okay. All right. Have a good night. You too. Thank you.