Week 5 Assignment
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 1
Research Design and Program Evaluation NCE Module Program Transcript DR. ELIZABETH VENTURA: Hi, everyone. My name is Dr. Elizabeth Ventura, and I am core faculty at the CMHC program. I'm super excited to be a part of your journey studying for the National Counselor Exam. Bringing this workshop series to each of you was a team effort because we understand the inherent anxiety that comes with taking the exam. Having been in your shoes and prepping for the same exam, we know that the more prepared you are and feel, the better the experience. The National Counselor Exam is a 200 item multiple choice test that is comprised of 160 actual score test items, 40 items that NBCC uses as test questions for future exams. For those of you that are not aware, the exam focuses on the following content areas-- human growth and development, theories and counseling, social and cultural foundations and counseling, helping relationships, group counseling theories and process, career counseling and lifestyle development, assessment and counseling, research and program evaluation, and professional orientation to counseling. Additionally, the sections are not represented equally on the exam. The following illustrates an example of how the questions are broken down-- 36 questions are the helping relationships, 29 are professional orientation and ethics, 20 questions are career and lifestyle development, 20 appraisal, 16 research and program evaluation, 16 group work, 12 human growth and development, and 11 for social and cultural foundations. Many students that have previously taken the exam reported that multiple study guides and resources are keys to successful outcomes. This workshop is just one piece of your study plan. Please, use this workshop to guide you in recognizing areas that you feel confident in and those that need a bit more attention. Please, know that we're here to help guide you through this process, and we hope you find the information valuable. So let's talk a bit about how this presentation is going to work. Because there's so much information to cover, I will not be participating in the chat throughout the presentation, but please feel free to follow along and use that chat to ask questions or post comments that other key faculty in the room will do their best to address. I've specifically carved out 30 minutes at the end to answer any concerns you have and to catch up on any outstanding questions that overflow from the chat also to follow up on the quiz. Hoping that now we have most of the housekeeping out of the way, let's get started. So we're going to be talking about counseling research. And let me start out by saying I do understand the inherent anxiety that exists for many of you when we
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 2
talk about research. I know none of us are math majors. And so when we hear statistics and research, immediately there is this level of panic that I think sets in. Hopefully, what this presentation is going to do, it's going to help you feel a little bit more at ease around the questions that are going to be asked on the exam and eliminate some of the information that you think you have to know but may not for the exam specifically. So just to get us started, in counseling research, it's important to understand that these questions are central to the inquiry process but they're not our starting point. This is a little bit different than scientific research where you start out with a research question and then you move away from that. There are essential steps to counseling research that are unique. For example, you look at your own bias to begin and what your beliefs tell you about the population you are working with- - whether that be a vulnerable population, like the elderly or individuals maybe in prison, extending far out into just observational research of the general population. You're going to find a place among the research paradigms. Why is the researcher wanting to conduct even relevant? And this is where your literature review, where your literature research comes in very handy. Asking research questions-- this is a very large process for many people because you're going to ask so many questions that it feels overwhelming. But it's iterative, and what I mean by that is you start out very broad and you work your way down until you get a finely tuned research question then asks exactly what it is you're wanting to know. Choosing a method based on what you want to find out. The method is very important and we're going to talk about some of those because it needs to align with who you are as a researcher. So these are examples of some of the research paradigms in counseling research and the methods and the tools that we use to be effective. For example, there is the initial paradigm, which is one that many of you may have learned early in college or in high school. And this is what looks most like scientific research. So this is a positivist approach. And what this means is you're looking at very factual information. And from what's actually observed and factual, you are then going to derive research in a quantitative manner. So you're going to look at a hypothesis. You're going to run experiments. You're going to have subjects that are going to be-- the sample size is something that's going to be completely blind to the researcher. We then move into more of a constructivist approach, which looks that qualitative methods. We're going to go over both of these throughout the workshops so you'll feel more comfortable at the end knowing what this means. But from a qualitative perspective, you're not doing, quote unquote, "numbers" so much as
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 3
you would be in quantitative. You're using more interviews, observations, subjective experiences, interview styling to gather your data. Transformative-- the transformative paradigm is also a qualitative design, but you can blend in some quantitative methods as well. So this looks a little bit more mixed methods. Transformative, the basis for the paradigm is that it's actually transforming something that we already currently know. So it's transformational in the results. It will change the way in which we look at educational research or counseling research. It shifts paradigms for people within a frame of mind. So how society may have been working under one dynamic, transformative research will shift it in a completely different direction. Pragmatic is qualitative and/or quantitative. And it's used with tools from both a positivist and constructivist paradigm. So you can see the way that these are set up-- starting off individually and then as the paradigms go on the list, transformative and pragmatic incorporate elements of all. So we start out with literature reviews, and why do we even need to do these? What's the point of a literature review before even having to start research? Just to give you an understanding as to how canceling research and research in general is done, using the literature review helps you distinguish what's been done in the past and what you want to do so that you feel unique and that you are topic's relevant. You have to provide a new perspective. Otherwise, you're basically just redoing what's somebody's already done. Identifying relationships between ideals and practice, establishing context of topic or problem, and rationalizing the significance of the problem-- literature review helps you to do all of this. It really helps you understand the subject matter, the vocabulary that's used. It helps you provide a structure for how you're going to set up your design. It also helps you relate ideas and theory to actual application in the field. So no real credible research can ever be done without a literature review. And typically, it has to be exhaustive, relevant to what you're going to be actually investigating. So what do our research questions do for us specifically? They help us organize our research and provide context. It gives us boundaries to work within. It helps us provide some focus and direction for how we're going to conduct the actual research. It helps us to find the theoretical underpinning or the theoretical lens that we're going to conduct the research through. It helps us understand our methodology and how we're going to collect the data. So how do we develop these research questions, and what do I want to find out? Possibilities as to how or what you're going to study, delineating the parts of the larger question. Like I had said before, it's an iterative process. You start out with
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 4
this narrowing of topics till you get to a very small piece that you want to focus on and that will drive your research. You do it again and again and again and even again until you get down to where you absolutely need to be where you know your research is going to be unique. . In research planning for counseling research, this is going to be no different than it was for scientific research. You asked the questions of what, how, why, and who. The reason we do this is because the what question is the question. What is it that we're going to be studying? What are we trying to find out? How are we going to go about doing this? How are we going to collect our data? Why are these even relevant? Why is this research worth doing? Who has said something about this before? And I even like to add, who is going to want to read this? So you have to make your research something unique enough that other people are going to feel that they would want to use this in their study as part of their literature review. Or it's going to some way, shape, or form influence our profession. So now that we kind of have the basic understanding of why this is even important to understand and conduct research, let's move into research ethics because this is an area that's really transformed within our profession and in the profession of psychology. Human subject research is very much restricted now more so than ever before. And what I mean by that is there used to be very loose, at times no guidelines that existed for protecting human subjects. And through the unfortunate studies or processes of having people being exploited through research, there has now been safeguards in place and boundaries for protecting human subjects. There is a difference between practicing on somebody versus an actual research of testing to determine whether or not an intervention, a drug, a test is going to be effective. And the difference really involves the way in which you go about implementing the research. So balancing the need of the field, the good of all, and the rights of the few. Protecting human subjects. A certification is now required by all IRB, or Institutional Review Board, submissions. So let's say a little bit about what the IRB is. The Institutional Review Board was developed in order to screen research studies that are being proposed. There is not one single IRB. Usually, various institutions are linked to institutional review boards so that they can have faculty or scientific research evaluated to determine if it's going to be safe for subjects. ACA Code of Ethics has several sections and subsections that speak to doing research. And I would familiarize myself with those if I were you because that kind of overlaps two sections on the NCE exam, both ethics and research.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 5
There are specific protections for vulnerable populations. So these would be considered children under the age of 18, pregnant women, prisoners, mentally disabled. When you submit a research proposal to go before the IRB, what you are doing is you're asking for there to be an expedited review, a full board review. There are various types of reviews that you can submit for. A full board review is an exhaustive process. It takes much longer than an expedited review. Because in those cases, you're often working with these vulnerable populations. So any time you would submit for a research study to be approved, and you would be working with one of these vulnerable populations, it would be presented as a full board review. So the Belmont Report-- the Belmont Report was published in 1979. And the Belmont Report outlines the boundaries between research and practice. The respect for people and ethical principles that many of you may remember from your ethics course were delineated in the Belmont Report-- so beneficence, justice, no malfeasance, autonomy, and fidelity-- beneficence, to do good; no malfeasance, to do no harm. These are the core principles that were outlined in order to help us understand how research needs to be guided. The overall outcomes of the Belmont Report is when is the informed consent process, privacy, confidentiality, and anonymity. So now what we look at is when even clients come into our practice and we discuss privacy rights, give them their informed consent, discuss with them that the confidentiality of being in session. That is a product of aspects of the Belmont Report. So I love this cartoon, and I wanted to put this in here for you. How do clients and research participants make informed decisions, and which therapy do you prefer? And you see this individual standing there clearly looking distressed, having no idea what the scientist is asking. And part of this speaks to the prevalence of past research problems that have existed and why there is now a need for informed consent. Informed consent gives the subject or the client all of the information upfront and allows them to make a decision as best they can with always the opportunity to terminate at any point without fear of retribution. There have been hallmark cases in our counseling profession slash psychology that speak to ethical issues in research. Many of you, I'm sure, remember the Tuskegee Syphilis experiment, the Milgram Obedience study, and the Zimbardo Stanford Prison Experiment. I picked these three because they clearly are the most popular of the studies that exist for outlining ethical issues in research. And students that take the NCE exam will come back and tell us often at least one of these was mentioned. So let me just take a minute and briefly remind you what these were about. So the Tuskegee Syphilis experiment was taking African-
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 6
American males, 600 I believe, in Alabama, some of which already had contracted syphilis. Close to 300 males did not already have the condition. At that point, this was run by the federal government. And these individuals were under the guise that they were going to be getting medical treatment, that they were going to be given burial fees, that they were going to be given meals, they were going to even care at some point. And part way through the actual experiment, funding was lost. None of the individuals enrolled in the study ever received treatment and were never told. And so this was a tragedy given what had happened, the lack of awareness that these individuals that had, the lack of intervention even though intervention was possible. The Milgram Obedience study-- this was a very interesting study, which I'm sure many of you remember because this involved males again in obedience. In this study, the males were given the opportunity to, quote unquote, "inflict pain" on an innocent person who was, in fact, innocent but was given the scenario by the researcher that there were as various other aspects that had occurred in this individual's life that made them not so innocent or, otherwise, deserving of being shocked. In this study, what they found was that despite the actions of these men feeling conflicted over inflicting pain on the individual that they couldn't necessarily see, but they could hear, they would do it anyway. And so this study was very much indicative of obedience, following figures that seem to have authority. And even though the actions were conflicting with their conscience, they did it anyway. The last example is the Zimbardo Stanford Prison Experiment. And in this experiment, the study was conducted where it was a Stanford University and college students were taken and divided up in between prison guards and inmates. And the idea was that Philip Zimbardo, who was the psychologist overseeing the experiment, was going to see how strongly these individuals took on their role. And, again, the federal government and the US Navy and military were very interested in the results of the study because of the implications of being a prison guard or being an inmate. And what had happened is that they observed the individuals took on these roles far past what they ever expected. And prison guards were even at some point subjecting the inmates to psychological torture. And the inmates were even taking on the persona of being an inmate and were becoming more submissive and were even taking directive to psychologically exploit or torture other inmates based on the directive from the prison guard. So let's move into research and evaluation and the connection between the two. There are differences between research and evaluation and often these terms are interchangeable for a lot of people but there are distinct differences. Research and evaluation are closely related, but they differ in many different
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 7
ways. The purpose for which they're conducted, the concern regarding external validity or the ability for it to be generalized from one sample to the population, their relationships to theory, and the applicability of the findings. What is evaluation? Evaluation usually concerns the validity or the application of a particular program. So what factors make the program good or bad? You'd have an evaluation or a program assessment that may occur where an evaluation of even an individual looks at factors that make it productive or thrive versus changes that may need to occur. Based on these evaluations, you may have a consultant come in and work with programs on how to better manage some of the deficiencies that the evaluation brought forward. Here are some steps that are commonly linked to providing an evaluation. So identifying the purpose and what you're even going to begin to look at. Determine who's going to make the decisions once the data is collected. When working in an agency, this can look like the program director or it can be contracted with another agency to work together. It also can look like a consultant, as I said before, an expert coming in. What is the criteria that's going to be used to making the decisions? What are the sources of process and outcome data? What's going to be the methods for data collection? How are you going to collect the data, analyzing the data, and interpreting the data? And how are you going to use the data to make a decision about the program in question? And then circling back to the top part, who were you going to bring in to actually implement the process? So research and evaluation continued, we look at steps for counselor accountability-- needs assessment for the client served, establish program objectives based on client needs, design the interventions to achieve the objectives, develop an outcome measure to determine success. What does this mean? What this means is that when you do a needs assessment, a needs assessment looks at maybe a population or even a specific individual to determine what are we missing. What do we need to do with this population or this client that we may not already be doing? So when I was just finishing graduate school, I worked at a program for pregnant addicted women. And we found that many women were not coming to their appointments, and they were canceling pretty regularly. And we worried that there were barriers that we weren't aware of so we did a needs assessment. And what we found out is that in this population with pregnant addicted women, they had little to no options for childcare. And so based on the needs assessment, when they didn't have child care for their other children, they were skipping their appointments. So what we did was we incorporated a daycare into our program that allowed the women to come to
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 8
treatment, bring their kids, and then not have to worry about missing out for a lack of personal childcare. This increased our retention rate with clients up to 90%. So this was a huge point of awareness for us as clinicians to know we're missing something. And then actually being able to implement a process that helped to retain the client that the clinic was really successful. Now, I want everybody to just take a moment here because we're going to get into quantitative methods in counseling research. And I am well aware that this is what brings a majority of anxiety on for many people. So I'm going to do my best to describe some of these concepts to you in ways that you're able to really understand without using a lot of technical terminology that we can get lost in. If at any point, when we go through this, you feel like you're very confused, you're not sure what's going on, you want to revisit some of the terminology that we're going to be discussing-- just jot it down or make a notation of it in the chat so that we can get back to it later. When we talk about the types of data, discreet data is a limited number of choices. So these are types of data that are going to be produced. Nominal data is a type of discrete data and limited number of choices. When you give a yes or no item choice, dead or alive, disease free or not, male or female-- these are yes and no, two choices, very simplistic. Just like we have said earlier on the paradigms of research, as we go through these types of data sets, you're going to see that they get a little bit more complex. So categorical, for example, is just another type of nominal data that is usually unordered, meaning race, age, group. Thees, again, are more than two choices. It's getting a little bit more complex, but it still remains unordered. When we go into ordinal data, these are more than two choices, but then they're ordered. So now we go from unordered to adding order-- stages of cancer, a Likert scale for responsive on a survey, and I gave you an example below. Getting more complex looking at continuous data-- continuous data is theoretically infinite possible values, including fractional values. So 5 and 1/2 feet, 6 and 1/4 inch, 34 and 1/2, 120.3-- these are examples of continuous data. Continuous data can be interval. Interval between measures has meaning. This is a very important feature of understanding interval data. Interval data is measured along a scale in which each position is equidistant from another. So this allows for the distance between two pairs to be equivalent in some way. This is often used in psychological experiments that measure attributes along an arbitrary scale between two extremes. Interval data can't be multiplied or divided. So let's go with an example. My level of happiness is rated from 1 to 10. Interval data is measured along a scale in which each position equidistant from one another. That fits, my level of
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 9
happiness is rated from 1 to 10. This allows for the difference between two pairs to be equivalent in some way. It can't be multiplied or divided. So that's just one example. When you take the NCE exam, these questions will likely come up. The best way to help you with this is to remember my example. So when you get the example on the exam, you can compare it to the one you know already and see if it's in line and it has similar features. When we talk about ratio data, ratio of the measure has meaning so weight and height. Numbers can be compared as multiples of one another. One person can be twice as tall as another. The number 0 has meaning. This is the only category of data where the number 0 has meaning. The difference between a person of 35 and a person of 38 is the same as the difference between people who are 12 and 15. A person can also have an age of 0. So, again, as we go through these examples, review these and think about how the example they're going to give you on the exam may or may not fit with the one you already know. That's a really good way for you to understand better how to answer the question quickly. Interval and ratio data measure quantities are quantitative because they're measured on the scale. The other two are not. So why are types of data important? The type of data defines the summary of what we're going to do. The mean, the standard deviation for continuous data and their proportions for the discrete data. We also need to know what test to run in order to determine our analysis. And so because of that, if we use the t-test or an ANOVA or a chi-square, that's going to depend on the type of data. So everything is connected. In order to make one decision, we need to know how the other factors are going to affect it. So without understanding your literature review, you're research questions, how you're going to drive your research, knowing what type of data you're going to have to use, the population you're going to pick-- you're not going to be able to run the appropriate testing. So one decision goes into another. When we look at how to categorize data sets-- hopefully, many of you may remember this from your classes-- histograms, frequency distribution, box and whisker plots. These are all very popular pictorial descriptions of data. There's also numeric descriptions and we know those from being a mean, median, standard deviation, interquartile range. These are going to be explained in just a couple of minutes so you have a better understanding of what these look like. This is just a pictorial version of a histogram for you to see. Instead of reading words of what it is, this is what it looks like. I know that in past exam students would say that they just were asked to label pictures and what type of data sets that they were useful for. So histograms are great for continuous data.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 10
They don't segment the data off into groups like a frequency distribution does. So I have this up here to show you a contrast between this and a histogram. This chart will show you how it does segment the data into the groups. And you can see, because at the bottom you're able to determine in the key, that's listed on the graph, where the graph falls in terms of segments. The histogram doesn't do that. It's all kind of clumped together. Box and whisker plots-- this is just another example, but this one shows how it's broken up into percentiles, where the median falls, where the mean would fall, and then the percentiles being both at the bottom of the whisker plots versus the top. You can see that on the first one to your left Group A, the median is the average or in the middle. On the other, it's skewed. And so more data points fell at the top than they did at the bottom. And this is just an example of box and whisker plots, which is useful for comparing data graphically instead of with numbers or in a narrative. So what are some of these numeric descriptions descriptive statistics-- measures of central tendency of data, or the mean, median, and mode. These probably I'm sure go way back into the early years of math that we're going to revisit just for a moment. There's also measures of variability of data-- the standard deviation, standard error, interquartile, which we'll discuss in a minute. So what is a mean? The mean is the sum of all the values in a sample divided by the number of values. I'm sure many of you know that. It is the most commonly used measure of central tendency. It's best applied in normal distributed data. And you can't use it in categorical data. So I hope that makes sense because remember in categorical data, we're using names. So you can't have a mean in that instance. It just wouldn't even make sense. Sample median in this case is used to indicate the average in a skewed population. It's usually always reported with the mean. And if the mean and the median are the same, the sample is always normally distributed. And the middle value is how you know the median. So if there is an odd number of values, it's the middle one. If an even number of values, it's the average of the two middle values. So if you would have any list of numbers, you would then take the middle two if there was an even number, add them together, divide by that, and then that would give them the median for that data set. The mode-- the mode is very rarely reported as a value in studies, but it is listed as the most commonly occurring value, more frequently used to describe a distribution of data. So if you have data that's distributed and there is a value that is reoccurring within a data set, you may then notice an outlier because that data number is going to be very off base from where the others are falling on a curve. So what does that all mean? So what that means is when you have a survey, and you have a data set and people have responded, you may get a large
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 11
majority of information that falls within each other. So 10 people have an A, seven people have a B, three people have a C-- they're falling within a similar average. You may have one or two people that have or one person that has an F, that failed the test. That f is considered an outlier. But the mode, or like the most commonly occurring value, would be in A or, in this case, let's say a 95%. So that's how you would understand the mode, the median, and the mean. Like I said before, if we're using nominal data, like I said A, B, C, you're not going to be able to determine mean. If these transferred into numbers, so a 95 95, 96, 97, they would be then able to be valued. Standard deviation and error-- standard deviation is often abbreviated as such or as simply SD. It provides an indication of how far the individual response to a question varies from the mean. So how far away does your answer deviate from what the mean is? So let me give you an example. If the average IQ is 100, the standard deviation may be 15. So somebody who scores a 130 as an IQ is going to be two standard deviations from the mean. If a standard deviation is 15, on standard deviation was 115, two standard deviations is 130. Standard deviation tells the researcher how spread out responses are or are they concentrated around the mean or scattered far and wide. Standard error is an indication of the reliability of the mean. So a small standard error value is an indication that the sample mean is a more accurate reflection of the actual population mean. So what does this mean? A small error is an indication that the sample mean that you are used for your study is an accurate reflection of the actual population. The larger sample size that you have is normally results in a smaller standard deviation because-- I'm sorry, standard error. Because what you're saying is you have so many data points and there's so many people that you were able to sample, the more people you're able to sample, the more people are going to be representative of the larger population. So your standard error is going to be smaller. If you had 10 people and you tried to then take those 10 answers and generalize it to the entire population, clearly your standard error would be much larger. What is the significance of standard error? The significance is it's the basis of confidence intervals. And we'll talk about what confidence intervals really are here in a minute. But confidence intervals-- a 95% confidence interval, for example, is defined by taking the sample mean, plus or minus this figure, times a standard error. You are not going to need to know this equation, but I put it in here for you to understand why or how we can come up with these confidence intervals and what it is that they mean.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 12
What a confidence interval actually means is you can be 95% confident that your data reflects what it reflects without error. You can be 99% confident. You can be 90% confident. It controls the error that exists within the study. So the larger the study or the sample size, the smaller the confidence intervals. And the greater the position of the estimate. So let's use this in terms of an example. When you have an educational research and you would decide that you're going to have a 95% confidence interval. What that means is that you're going to allow for five or so percent that is going to be in that entire data set, 5% you're allowing for there to be error. You certainly would hope that in scientific research-- for example, I have here which might-- a study conducted by the FDA used most frequently. I would certainly not want there to be a 5% chance of error if I'm taking a new medication that's life sustaining. I would really like there to be a 99% chance that it's going to work. I would rather leave that 1% chance than a 5% chance. So typically in scientific research, you're going to have much more stringent restrictions on the research making it so that the confidence interval is going to be at 99%. You're going to be 99% confident. So these are always fun to remember and super confusing at times, because we get into type I and type II errors. Claiming there is a difference between two samples when there is no difference is a type I error. This is definitely where you do not want to be. You as a researcher are saying that your study showed a difference between two things or that there was an effect in some way. There was significance when they're in fact was none. So let's think of an example that might be helpful. Let's say that there was a research study that claimed there was definitely a difference between using CBT on obsessive compulsive disorder and using person centered therapy. And so this research was published, clinics changed the entire way that they did things. They went and said OK, there's a difference between these two. Let's say person centered-- the research showed, let's say, person centered was the most effective way to handle obsessive compulsive disorder. And these clinics that were using CBT changed the way that they were doing things and went and said, nope, person centered is the way to go, but your research was wrong. You claimed that there was a difference between the two samples when there wasn't one. So what does that do? That then allows for research to be published and, in fact, isn't credible. When you look at a type II error, a type II error is that you're really not giving yourself credit as the researcher. You're claiming there's no difference between two samples when, in fact, there is. So you worked really hard to get the research and you actually do have evidence of something being different, yet you say that there was no difference.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 13
This is a chart that simply just outlines for you the null and the alternative or the alternative is often called the researcher's hypothesis. This is sometimes for people a much simpler way to remember it. People who like pictures. For other people, it's easier just to remember it by using examples. So how do we determine sample size? Power analysis is the term that you may see on the exam because power analysis is also linked in understanding how to collect in what is the sample size. When designing a study, you need to determine how large the study is going to be and how large it needs to be in order to get your confidence intervals and your data set and your generalizability to the population. Powers the ability of a study to avoid a type II error. And remember the type II error is claiming there is no difference between two samples when there is one. Sample size calculation yields the number of studies subjects needed. So it tells us how many we need in order to have a desirable outcome or avoiding errors. There are certain kinds of statistical tests that we need to run once we do-- the steps prior, where we're collecting the data, we're understanding the subject size, the data sets, looking at the research questions that we have. We have pooled all of this together, now, what do we do once we have the data? Once we have the data, we have to run our statistical tests or our analysis. Parametric tests are used for continuous data that are normally distributed. Non- parametric tests are run for data sets that are continuous not normally distributed, categorical, or ordinal in nature. We're going to give some examples here in a minute. t values, you're able to look those up in a table. So this is a very simple way to understand the significance. This is one of these questions where you are not going to have to know the t values, you're not going to need the table for the exam. But you do have to know what the definition is. So this is one of those questions where understanding what a t value is important, but you're not going to have to calculate those. So parity tests-- parity tests uses that change before and after intervention in a single individual. So it looks at pre and post for an individual-- reduces the degree of variability between the groups-- given the same number of patients has greater power to detect a difference between groups. This is going to look much different than an analysis of variance, where this is used to determine if two or more samples are from the same population. If two samples-- it's the same as the t-test. You need three or more samples in order to do an analysis of variance. An ANOVA allows one to determine whether the differences between the samples are due to random sampling error or whether there are systematic
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 14
treatment effects that cause the mean in one group to differ from the mean in another. So in an ANOVA, you're looking at three or more samples and being able to differentiate why they may be different from one another. When we work with ANOVAs, all populations involved follow a normal distribution. Populations have the same variance or standard deviation, and the samples are randomly selected and independent of one another. These are the three assumptions that we need to follow in order to conduct an ANOVA. Non-parametric tests-- I just gave you a couple examples because these are the most popular. Testing proportions for categorical data-- a chi-square is used to compare data that you observe from data you would expect. Testing ordinal variables-- one of the most common is a Mann-Whitney U test, which is basically the opposite of a t-test. It compares population means that come from the same population. It also used to test whether two population means are equal or not equal. So the Mann-Whitney is a non-parametric test that's the opposite of what we've already learned if a t-test. And a chi-square uses categorical data, and it takes what we've observed and how that's different than what we would expect. If there's a substantial difference between observed and expected, then it's likely that the null hypothesis would be rejected. These are often graphs as a two by two table, which it's good to familiarize yourself with what that may look like. But, typically, whenever you would be asked the question that would involve nominal data, you're going to know pretty quickly that it's going to usually be a chi-square. Most non-parametric tests are based on ranks or other non-value related methods. So the use of these non-parametric tests is really through interpretation to understand is the p-value significant . And remember that p-value of 0.05 or 0.005 then ties back into our confidence intervals. Many of you might be more familiar with correlation. A lot of things people commonly say is correlation doesn't mean causation. So even though something is correlated, it doesn't mean there a cause and effect relationship. I have some examples to show you what correlation actually looks like when it's graphed. This is an example of a perfect correlation. This is showing an example in letter E of a correlation coefficient of 0, meaning it's completely loose. There's really no formidable consistency that you can make of the data points. A correlation coefficient of 0.3 we know from before would be a low connection. The strength of the association of the data points being linked together is low. You don't get a moderate connection until you're at like 0.4 to 0.6 and 0.7.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 15
So these are looking at like more moderate strength for the correlation of the data points. And just to reiterate, understanding that these values mean certain strengths for how the data points are going to be correlated is going to be important for you to know. I remember on my own NCE exam and don't even ask how I understand-- how I remember this. But on my own NCE exam, I remember that there was a graph. And it showed data points just like I showed you. And it asked me where the correlation coefficient would be. And the correlation coefficient is a small r. So my choices would have been r equals 0.2, 0.6, 0.7, or 1. And based on what the image of the data points was then, I had the answer what I felt was that accurate representation. So that's why I really put these graphs on here so you could see for yourself what these were representative of. So as we move into understanding independent, dependent, and controlled variables. Independent variable is what you change. Dependent variable is what you observe. And the controlled variable is what you keep the same. Let's just get rid of all of the fancy definitions. And if you can remember this, this will really help you in the moment. OK, we're done with quantitative research. So everybody can take a moment and breathe because we're done with that. We're putting it to the side. Hopefully, you see now that it wasn't maybe as overwhelming as you anticipated. You've jotted down a few issues or questions that you have to follow up on. And you kind of know for yourself that maybe you need to go back and review those data sets. Maybe you want to take a minute and really understand what a t-test is or how a chi-square may look. Now that you know what some of the really important components are to understanding research for the exam, you can kind of take a minute and step back and say these are the three areas I need to remember. But, wow, I really feel good. I know I did great understanding the other aspects. So I don't necessarily need to spend as much time as I thought I did on that. Moving into qualitative research-- qualitative research is normally a lot easier to understand from the standpoint of what it looks like, how it's conducted. It speaks more to us as counselors because the majority of counseling research sometimes falls within a qualitative paradigm. But qualitative research is any kind of research that produces findings that are not arrived at by means of statistical procedures. When I was in my doctoral program, many of my cohort members, my classmates, would say, are you going to do a quantitative or qualitative? And it was almost the look of craziness, and we would look at each other and say, you're going to do quantitative? You're actually going to do that? Because for
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 16
many of us, the questions that we want to ask in our field aren't able to be measured. in terms of using numbers. So we look at qualitative research to be a good blend and to bring in another dynamic for us to understand concepts and subjective experiences of other people that can't be quantified. This is a relatively recent development in using qualitative research in educational environments. It has a very rich history in anthropology and sociology. Ethnography is describing primitive cultures, dating back to the late 1800s, when we would try to understand cultures inductively, observing behavioral patterns how and why people live the way they live in order to help other people make sense of their own culture. From 1900 to World War II, there was this model of the lone ethnographer. And what this meant was that people were spending extended periods of time doing observations among natives and in a distant land, utilized observation, interviewing, artifact gathering in order to determine how cultures were living and how to compare one society or one social group norm to another. Similar studies were conducted in Chicago looking at how ordinary people live their life and what this looks like, making the city a social laboratory. We then go into kind of the post-World War II phase through the mid 1970s when the methods of qualitative research became more formalized. Scholars became more self-conscious about the research approach. And there was more of a push to finding some balance between a positivist approach and searching for that validity and reliability and being able to generalize the research with constructivist models of doing the research. So in the 1970s, there was really more of a crystallization of these two together in order to help researchers find this balance between the research being credible and having some semblance of a subjective experience. This is where we start to use the qualitative research model within education, looking at it from a phenomenological approach-- critical theory, feminism began to be recognized as all being credible avenues to conducting research. There were some boundaries between the social sciences and the humanities that were becoming blurred. Interpretive methods, such as being able to experience a subjective experience and to be able to use this for adaptive-- to be able to use this for qualitative analysis was becoming more and more difficult to do. So it became obvious that there needed to be a way to legitimize qualitative work, and to move it in the direction of educational research to where you would be able to conduct this type of research that would then drive quantitative hypotheses. So what do I mean by that? Qualitative research would be conducted. And out of that, hypotheses would be given, which would then be the basis for more quantitative research that can be done later.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 17
What are some of the characteristics of qualitative research? It needs to be in a natural setting. Participate perspectives are of utmost importance. Researcher is the data gatherer or the instrument for collecting the data-- extended first hand engagement with the participants. There's a wholeness and complexity to understanding a problem. It's a subjective experience, an emergent design, inductive data analysis, and reflexivity. What does that even mean? So we're going to go through each step by step. Within a natural setting, what we're doing is we're trying to understand how individuals make sense of their everyday lives. So when research settings are controlled and manipulated, you're not getting the actual authentic experience. There's some artificial context that is overlaid that allows people to act differently because it's a slice of life. It's almost like when you do an initial assessment. You see a person and you do initial assessment. And you feel like you have all the information. And you make conclusions based on that and a diagnoses. Well, the beauty of being able to and the hope that you go back to and revamp and refine the diagnosis is because as things emerge and as time evolves, there is so much more that you understand that places it in context. And so when we look at artificial situations where somebody is only given a few moments or are able to only act narrowly, it seems as if it's not representative of what we're actually trying to investigate. So we look at a participant perspective. Individuals act on the world based not on some supposed objective reality, but on perceptions that they actually experience. So we need to gain insight into their own lived experience, their own perspective. And we do that through the methods of qualitative research, like interviewing and observation. The researcher is the data gatherer. So in qualitative data, what we do is we use field notes from observing participants. We transcribe interviews with informants. We may use other data points, such as artifacts from research sites, records related to the social phenomenon under investigation, existing data sets that may have been conducted prior to. Data take on no significance until they are processed using human intelligence-- so using ourselves. What does this look like? So in my dissertation when I was completing graduate school, I did a qualitative research study on the lived experience of students at the master's level going through internship and practicum and being able to examine what it was like for them to sit across from somebody a client who was giving them a trauma story. And did they feel prepared, what was their lived experience of having gone through that? Do we need to do more as a program to give people more preparedness in the area of trauma? So I was getting all of this information as the study was being conducted, but it was each individual story. And I really could do nothing with it until I was done
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 18
collecting the data and starting to synthesize what I had learned. So what I needed to make sure that happened was spend enough time with the participants in those contexts to feel confident that I got when I needed to get. So this is technically called saturation of the data. And what that means is when I get to a point where I've interviewed enough people, that the stories are kind of linking the same, then I know that I've reached saturation of the data. And I can start to actually analyze the data. One of the problems that can exist with qualitative work is that the researcher spends too little time in research settings and does not reach saturation of the data and, therefore, terminates the research too soon. So when you look at centrality of meaning in terms of a qualitative approach, we're describing the meanings individuals use to understand social circumstances rather than trying to identify the social facts that comprise the positivist social theory. Human beings act toward things on the basis of the meaning that they have in their own life. The meaning of such things as derived from or arises out of the social interaction that one has with one's fellows. These meanings are handled in, and sometimes modified through, an interpretive process used by individuals in dealing with the things they encounter. So when we look at it through this symbolic interactionists theory, we see that human beings act on meaning that they have for themselves. So that's why it's important to look at the meaning in their life and not just meaning as it is a societal fact or a social theory. Wholeness and complexity-- individuals are unique. They're dynamic and complex. And there are various factors that are unique for each individual, even who's experienced similar circumstances. So if I would interview 10 people that would have been a part of 9/11, even though they may have all been in the same office building or elevator, very unique experiences are so different that I need to understand the wholeness and complexity that exist for each individual. The subjectivity of this is important because it requires the researchers to move from descriptive towards interpretation. I need to be able to describe what I know and then interpret that. I can't pretend to be objective because this is something that I'm inserting myself in. So I'm concentrating on how to apply my own boundaries in a way to be reflexive through the process and how to understand my motive and assumptions and even those of my participants that would be in the research. So one of the things that I did part of my dissertation process was I was certain that I had supervision. I kept a journal. I was able to work with another researcher, the chair of my dissertation, my mentor, to be able to help do what would be kind of a blind coding. So she was able to take a look at some of the
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 19
data sets that I had in terms of the respondent information and give me her take on it that was completely independent of mine. What is an emergent design? Studies change as they're implemented. So research questions and methods and other elements of designs are altered as the study unfolds. If I would be going along in a process and realize that this isn't working, then I would be able to change how something was being implemented and that would be totally acceptable in a qualitative project. It would not be that way in a quantitative project. So when we look at inductive data analysis, when we move from specifics to generalizations or generalizations to specifics, you're not putting together a puzzle. Whose picture you already know. You are constructing a picture that takes shape as you collect and examine the parts. So in qualitative research, the researchers do not begin with a hypothesis. We don't start out with something that we're trying to prove or disprove. Through our work, we come up with hypotheses that are a result of what we found. So we're not on a mission when we start as some quantitative researchers are. We're not attempting to prove or disprove anything. We're looking for information to come us. We mentioned this before and I gave you some examples of being reflexive is really being able to keep track of the influence on his setting-- bracket bias, monitor emotional responses. When I was conducting my research as a grad student, I wasn't able to have anybody in my study, that was a current student of mine while I was teaching there. Because I felt that it was best to keep that part of my study separate from my work as an educator. And so that was one way that I was able to bracket bias. The process of personally and academically reflecting on lived experiences in ways that reveal deep connections between the writer and his or her subject is the way that you have somewhat control over your study. Like I had mentioned earlier, these are some examples of ways that researchers can be reflexive. And I would encourage you just to kind of remember these because they're commonly listed and good examples-- go to examples, I would say as well, that if you can remember these, if a question is asked, you're going to be able to understand pretty quickly what it's asking for. So these are just some types of qualitative research that I have listed. Hopefully, some of these at least look a little bit familiar for you. Focus groups-- in my study, I did both a focus group and I did individual interviews. So there was a little bit of redundancy in mine but that was one of the ways that I worked to make it more credible. You can look at narratives studies, phenomenological studies, a case study, grounded theory studies. Some of these are ones maybe familiar to you. Others you may not be so sure what they mean. Certainly, I would not spend a tremendous amount of time knowing the
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 20
intricate details of all of them. I would certainly be familiar with what type of qualitative research-- these types of qualitative research and maybe being able to come up with an example. So I've felt a little bit about my dissertation to you in hopes to give you some context for qualitative research, but just to say a few more points on that. In order to conduct successful qualitative research, it has to be through a theoretical lens. So just as quantitative research comes up with a hypotheses, it's equally critical for us to use a theoretical lens for qualitative research. So phenomenological is kind of the way in which I approach my study. So from a phenomenological perspective, I wanted to see the world through another person's eyes. I wanted to study their subjective experience. I wanted to know their lived experience of what it felt like for them to be in the session when something was disclosed and trauma and they felt unprepared. So I would get informants that would tell me, I felt like time stood still. I felt like the room was closing in on me. I started to feel my palms sweat and my stomach got nauseous. I didn't know what to do. I felt like getting up and running. I wanted to run to the bathroom and cry. I felt so overwhelmed. I felt like I was going to re- traumatize the client. I'm not sure how I would ever begin to measure that quantitatively. And so that's the best example I can give you to show why it's important to have both elements of research available to us. So the next thing we're going to cover is assessment. And assessment starts back very early in terms of being able to look at career counseling and its linkage to that. And even in modern day counseling when we have clients come in for an initial assessment. We have them go over a battery of informational tests and maybe even a Beck Depression Inventory, personality inventory-- all of these things fall under assessments. So this section really just highlights for you some of the more important aspects of assessment to know. It's certainly not an exhaustive review, but it's definitely picking out some of the more key players that are involved in assessment and some tests that would be beneficial for you to know. So Francis Galton developed the first intelligence theory. And he was an avid researcher who believed intelligence was normally distributed just like height and weight and that it was mostly genetic. We move into Alfred Binet who was the first modern intelligence test in 1995, which many of you may know as the Binet- Simon scale and later collaborated with Stern to develop what we now know was the IQ. This test was originally used with children with intellectual disabilities and was normed on children aged 3 to 11. The Stanford-Binet IQ was developed a little bit later as a refinement of the Binet-Simon. And the skills include average-- I'm
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 21
sorry, below average, average, above average and are still used today. It's a standardized measure because the scoring and administration are uniform. So what that means is within assessment, one of the ways to control for error is to make it so that administration procedures are specific and precise across every time it's being administered. Guidelines exist to control from various areas of the testing process. So we know instructions are given consistently. And this differs greatly from non-standardized tests. Many of you may remember if you took ACT or GRE-- place your pencils down, please close the booklet. Those types of things are instructions that are always given at the same time for standardized test. These are just some examples of popular intelligence tests I thought would be helpful for you to know. It's sometimes hard to weed through all of the information, so I just picked out a couple that if you can remember a few points about these. We talked about the Stanford-Binet, but the Wessler is an intelligence scale specifically for children. It targets children through late adolescence from 6 to 16 years old. The Wessler IQ test, or the WAIS-III, is targeted for adults and each has a different refinement than the other. There's also a difference between aptitude and achievement tests. Aptitude are tests that are used to predict an individual's likelihood to pass or perform in school. Achievement tests or those that measure what the student has already learned. So the SAT is the most popular aptitude test, and that's used to predict an individual's likelihood to perform well in college as is the GRE for graduate school. There's generally no need to study or prepare for an aptitude test, although, I know very well that there are many programs and booklets and study guide and study aids that do help. But, basically, the aptitude test in general is used to measure kind of what you already are aware of. These yield higher results with increased preparation by the test taker if we're referring to achievement tests. One example of an achievement test is the ACT, which many people take alternatively to the SAT. Because of the format of the test, individuals taking the ACT sometimes can perform differently or better than those same individuals taking SAT. So interest inventories are other popular tests that are used often in vocational counseling. They measure interests for employment and placement and even finding a mate. So when you think of online dating sites, the questionnaires that they're giving you are interest inventories. And the Strong Interest Inventory is the most popular interest inventory used, and it's really helpful for vocational or career counselors. Personality inventories are fun, and they're originally developed during World War I to screen recruits with mental illness. As these personality inventories have
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 22
evolved, the MMPI or the Minnesota Multiphasic Personality Inventory is clearly the most widely used and popular. It's also exhaustive. It's incredibly long, and it takes a while for people to take. And it can feel overwhelming. If you Google personality inventories online, you'll see-- if you like pink, you'll like this kind of person. If you like red, you have this type of personality. Those are obviously fun and exciting to take and you see them all time on social media, but the reality is that personality inventories really need to be normed. And they absolutely are something that employers use and military can use to determine goodness of fit for specific employment jobs. When you have personality inventories what you have to worry sometimes as a researcher about is somebody faking the answer. So the questions on the assessments will be normed for that. So it will be able to look at response bias or faking good or faking bad, depending on the way the individual answers the question. For example, the question may say, "I never get angry with my children." I have a lot of parents that I see that are going through issues of custody in my practice. And they may come to me after the MMPI and say, there was a question on there that said, "I never get angry with my children." And I said that I do get angry, and I think that that's probably not going to go well for me, but I just couldn't lie. And I'll say to them, what person do you know would answer that that was true? That they have never once gotten angry at their children? That's not a normal response. That's not a normal way to answer that question. Every person at some point or another have gotten angry with their children. And so if individual circles true, that they never get angry with their children, that's a red flag for researchers. That appears as if you're faking good. Projective tests are fun. And they are a series of assessments that were developed early in the 1900s for measuring psychopathology and personality disorders. So the roots of these tests were in free association, which we know from Freudian and Jungian time. And it was developed in 1921, the specific test that we're referring to, the Rorschach. And it uses inkblots to be presented for interpretation. And I know many of you have seen these before. So, hopefully, this is something you're familiar with, and you don't feel like you need to spend a lot of time on it. But what we do need to know is what are these projective tests and what are they useful for. So projective tests are getting the subjective experience from an individual. So, essentially, we use the Thematic Apperception Test. This test uses pictorial descriptions of scenarios that are fictitious in nature. And you show these pictures to individuals who will then tell you a narrative that links to this story. You'll have some clients that may look at that picture and say
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 23
his wife just passed away. You'll have some that may say he just murdered his wife. You'll have some say she's very sick and he's overwhelmed, and he doesn't know how he's going to handle raising the children without her. And from the information that you get with the client, what you're then able to do is say I'm wondering how their narrative fits with what is going on in their own life, in their past. So the responses in the stories made up to these ambiguous pictures reveal their underlying motives or concerns or the way in which they see the world. It's projective in nature, because it brings forth the unconcious. What's critical to know about projective tests is that you can't use them alone for diagnostic purposes. Because there is so much subjectiveness that involves projective measures, you have to use this in combination with other assessments. So using the House-Tree-Person, the Thematic Apperception Test, the Rorschach are viable and useful and definitely speak to parts of the self that wouldn't normally be exposed and doing other types of tests. At the same time, it can't be used alone. When we look at how we understand test purposes and the audience that it should be used for, the way in which tests can be normed, data sets-- we look to the Mental Measurement Yearbook. The Mental Measurement Yearbook is essentially like the DSM for diagnosis, but the Mental Measurement Yearbook is that for testing and assessments. So it basically provides an overview of the quality of assessment measures for each assessment published. Sections are notated according to the audience it's intended for, how to administer it, the price to purchase the test, how long it takes, the delivery method, who the publisher is and the date it was published. Many factors go into test construction. The test measures a certain set of behaviors or constructs and can be designed in a variety of ways. Item discrimination refers to the extent of which the items are going to elicit a response that will differentiate test takers. You can have a subjective format or an objective format. So for a subjective format, you're going to see short answers or essays. Objective, you might see true/false, multiple choice, et cetera. Free choice format is a short answer or test format in which more subjective is going to be in the scoring. However, the length of the answers are not as in- depth as those of an essay. So a short answer might be one or two questions, where an essay might be a paragraph. Other examples are going to look like fill in the blank, or it may even be a response item. For example, in terms of my dissertation, I had to defend it. I had to orally respond to certain questions and defend my position on my research so that's an example. Many people get a little bit confused when we talk about validity and reliability. So really the best way that I can help you understand this is to simply know the definition. Once you know in the definition, I think it's going to be easier for you to
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 24
apply it to a situation. And I'm going to give you an example that incorporates both. So what is validity? Validity of a test is really whether it measures what it's supposed to measure. So does it do what it's supposed to do? So the Beck Depression Inventory has really strong validity. It measures what it's supposed to measure. It measures depression. The content validity is the extent to which the items on the test actually do test the construct. So if depression is being assessed, do the questions cover the symptoms of depression? The Beck does that as well. Criterion validity-- type of validity that shows how effective an instrument is at predicting an individual's performance or assessing current status. Construct validity-- how well does an instrument measure a theoretical idea or concept? Going back to the Beck, it's going to have very strong construct validity. Experimental design validity-- it involves an experiment to show the instrument measures a certain construct. For example, have a therapist give Depression Inventory before and after a therapy to determine whether therapy was effective. Factor analysis is a statistical technique to analyze relationships between instruments items. Are subscales on the Depression Inventory related to each other and the overall concept of depression? So does it all tie together? Does it all-- is it all encompassing? You can have convergent validity. Validity that looks at whether assessment is related to what it should be. So is Depression Inventory positively related to the Beck Depression Inventory? You can have discriminant validity. Depression Inventory scores are not related to scores from an achievement test. So they're not one and the same. They shouldn't be. Or I guess maybe they could be, but I think we're getting way off track if we go that road. But in thinking for the NCE exam, discriminant validity-- a good example to remember is Depression Inventory scores are not related to scores from an achievement test. You're discriminating between the two. Faith validity-- does an instrument look credible? Can you tell from the content asked what the test is purporting to measure? So if somebody hands you the Beck Depression Inventory, the face validity on that is obviously very strong. The title of it is the Beck Depression Inventory. And the questions use the terms like, how unhappy are you? Have you ever been so sad that you've had thoughts of harming yourself? You probably get an idea from those questions that it's not about your skills at being a good baker. So there's a big difference between questions on an assessment that have very strong face validity versus those that don't.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 25
Concurrent validity-- validity that looks at relationship between test results and other currently obtainable measures. And lastly, predictive validity-- validity that looks at relationships between test results now and a measure collected in the future. So if we're going to compare Depression Inventory with hospital admission two years after the assessment and see growth and see difference and see how it can be compared. When we look at reliability, reliability is the consistency of scores by the same person over multiple administrations of the same test. So this can be explained in many ways, but I'm going to give you an example. Jane's been struggling with an eating disorder and weighs herself every day. Throughout the course of a day, she will weigh herself five times. She questions the reliability of her scale, even though every time she gets on it, it tells her the same number with only 0.5 to 1 ounce difference. From this information, we can assume that her scale is reliable. Test/retest reliability is when you look at scores on two different administrations of the same test. Please, remember that in this case, there are certain factors that affect this. So if I were to give first graders a test on August 15 and give them the exact same test on October 15, that may be a decent frame for a test/retest for reliability. But if I attempt to give them the test on August 15 and not again until the end of the school year, so much time has passed that it will affect my results. Alternate form reliability-- reliability that compare scores from two equivalent forms of the same test. Also called parallel form reliability. So, for example, if we're going to deter students from cheating, Professor A creates two forms of a midterm that are equivalent in difficulty. By comparing the scores from the two tests, the form of reliability can be obtained. So we have internal consistency and split-half reliability. Internal consistency measures consistency of responses from one test item to the next during the same testing session. And split-half reliability is the internal consistency that correlates one half of a test against the other. These are some ways to further understand or even calculate reliability. I want to jump down to reliability coefficient because this is the most commonly discussed correlation coefficient. The reliability coefficient in this example is when you look at a test reliability, the closer to 1 the value is the better. So this is also called a Pearson r and that's just the formula and the procedure in which you go about understanding it and the developer of the formula. So in the Pearson r value, it ranges from negative 1 to positive 1. So, again, when you look at the reliability coefficient, remember that a test reliability is measured the closer to 1, the better.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 26
So if it has 0.7, it's going to be pretty good compared to when you get-- it's almost like-- when we remember correlation coefficients and you remember the scatter plots, this is going to be a similar range that you might. So there are certain errors in testing. And some of these may seem like common sense to many of you, but I just want to review some of them just to jog your memory. Certain factors affect the scores of tests that are outside the psychometrics themselves. So we can look at the test taker, the condition of the environment, the quality of the administrator, and reasons behind it was given in the first place. So like I had given you the example earlier of the MMPI with clients that I have, parents that I have that are going through custody cases. They may be so overwhelmed with anxiety around how their answers are going to affect the welfare of their children that their answers are skewed. You may have an adolescent who is going to take the SAT and broke up with a boyfriend last night or girlfriend. And so in that case, the conditions of which they're in, the test taker themselves, are all going to affect the overall results. So let's take this a step further, can test scores be reliable but not valid? Yes, if we remember our example of Jane, validity means that a test-- or scale in her case-- measures what it purports to measure. If Jane steps on the scale and every time it reads one pound, it is reliable. It's producing the same result each time she steps on the scale, but it clearly is not valid because it's not accurately measuring her weight. Can test scores be valid but not reliable? No-- if a test score reflects that the assessment measured what it was intended to measure, then each time it will produce similar results. When we think about test construction, we look at item analysis and difficulty and the ability for the items to discriminate from one to another. So when we think about the NCE exam, the NCE includes 40 test items on the exam. Based on how individuals respond, they're going to determine the quality of those questions-- whether or not they should be included on future tests. The percent of test takers who answer an item correctly, the value is represented as a p value. The higher the value, the easier the item. You want to kind of range in the middle here, so it's not super difficult and not super easy. And item discrimination, we talked about it earlier. But just to remember, it's the degree to which a test item differentiates test takers. If the Beck Inventory intends to detect depressive symptoms then depressed individuals will give different answers than those who don't. Criterion referenced assessment-- assessments that compare a person scored to pre-determined standards. So the NCE is a great example of this.
Research Design and Program Evaluation NCE Module
© 2018 Laureate Education, Inc. 27
The NCE uses criterion referenced assessments to determine passing rates. Pass rates need to be within each criterion within the overall test. And then it's on each subset as they're listed that there. You can have scoring developmentally, age equivalent, or grade equivalent. So developmental scores are going to describe an individual's location on a developmental continuum. And then you can compare others the same age. These types of tests are helpful for pediatricians, for example, when you're looking at milestones. Age equivalent scores are going to do something very similar, comparing individual scores with average scores of those the exact same age. It's usually reported in years and months to be even more specific. And these are helpful for school psychologists when they're trying to identify students with disabilities or do IEPs. And then grade equivalent scores are comparing individuals with average scores of those in the same grade. So for test interpretation, you do clearly define what the test is intended to measure. There's a limitation of the scores related to error that sometimes you can't control. There can be misinterpretations. Use appropriate understandable language whenever you're giving it, making sure that for standardized tests the administration is standardized. And that there's cultural sensitivity. When interpreting test results, it's unethical to simply list values without explanation. Assigning labels to an individual as a result of the test in the absence of anything else or further explanation is also unethical-- providing broad based generalizations or using technical language that is outside the scope of the individual's ability to understand. So now that we have concluded the majority of the material for the research section of the presentation, what we're going to do now is I'm going to give you a few moments to take the NCE practice quiz that goes along with the material presented. I just want you to take a couple of minutes and go over-- I believe there's six questions that this asked you. And you can jot down the answers at the end. We'll go through and list the answers that are for each question. And then if you have questions, you can list those in the chat and we can address those.