U4A1 - Your Research Question....Please follow all instructions and read the attachments. DUE SUNDAY BY 9PM CST. My Field is Public Service Leadership
WHAT THEY DO NOT TELL YOU ABOUT STATISTICS
Missing Data It is not unusual for datasets to have missing values for a given variable. Various methods for addressing missing data have been developed. Three of the most common include:
1. Doing nothing 2. Deleting cases with missing values 3. Substituting the mean of the missing variables
Each of these methods has pros and cons and each has some risk of introducing error into your analysis. Deciding how to deal with the missing data will depend on the analysis being performed and your understanding of the interrelationships between variables.
Relationships in ANOVA
Description
Consider this analogy for a moment: You go to the doctor and the nurse weighs you and takes your temperature, your pulse, your blood pressure, and your rate of respiration and writes these vital statistics down in your chart. She asks you what brings you into the office. Next, the doctor comes in and looks at your chart. If she were to say "Hey, great... your temperature is 98.5, you can go home," you would want to start looking for another doctor. You expect the doctor to analyze the various findings in relation to one another.
In the same way, understanding the relationships within the ANOVA (ANalysis Of VAriance) test will help you head off error in your analysis. All the numbers mean something and you have to look at all the numbers in a holistic manner and consider how they relate to one another.
One relationship to be aware of is the P value. The P value will help you determine whether to accept or reject the null hypothesis. If the P value is less than .05, then you need to reject the null hypothesis - all the means being analyzed are equal and there is no difference. If the P value is greater than .05, then the difference between the means is significant and you need to reject the null hypothesis.
Another important relationship that you will examine in ANOVAs is the f statistic. Think of f as the answer divided by error. Obviously, you want to minimize error and maximize answer.
And finally, remember that sample size does not affect the ANOVA analysis. ANOVA is not sensitive to differences in sample size - a sample of 500 cases can be compared to a sample of 50 cases.
Illustration
ANOVA
Source of Variation SS df MS F *4 P-value F crit
*Notes from table 1. Treatments 2. Error 3. You always want the "between groups" to be MUCH higher than the "within groups," as to account
for as much variance as possible (i.e. decrease or control for error) in your analysis. 4. A low F suggests in advance a high probability for a high p-value. 5. Low df suggests small sample size, so probability of finding significance is smaller too. Recall as
little "n" approaches big "N" the central limit.
Is More Always Better? We would expect that the more data we have, the more reliable to conclusions we draw from it. However, it is not always true. Imagine that you are trying to gather data about the emergency room visits. Your survey design calls for you to gather data from one of 3 hospitals in your community by asking people to fill out a short survey and return it to the nurse's station. You do this for 4 weeks from 12:00 PM until 2:00 AM and collect 500 surveys. Now imagine that you conducted the same survey, but this time, you survey at all three hospitals for one week. In this survey design, the survey is given out 24 hours a day and at the end of the week, you have 200 surveys.
You may find that your data for the first survey is not as representative of the population as a whole as is the data gathered by the second survey.
The point here? That study design needs to address more than just the number of cases obtained.
Quality of Data High quality data is data that represent the reality of the population, community, or pool that it represents. that contain errors. Dirty data can contain various mistakes such as errors in spelling or punctuation, having been put into the wrong field, incomplete or outdated data or even data that has somehow been duplicated. Data can be compromised in a number of ways:
• Human error during collection or transcription. • Errors that happen during the transmission of data sent from one computer to another. • Software bugs or viruses. • Hardware malfunctions, such as disk crashes. • Physical damage due to natural events such as fires or floods.
They can't be using the same data...can they? Sometimes when we listen to people taking about issues, it seems as if they are getting their information from two completely unrelated sources. How the data is analyzed, what variables are considered, and how the results are presented can all affect what the data appears to be showing.
ANOVA
Between Groups *1 62.8 2 31.4 *3 2.837349 0.0979 3.8853
Within Groups *2 132.8 12 11.07 *3
Total 195.6 14 *5
"The U.S. unemployment rate averaged 4.7% from 2001-2007, under George W Bush. This compares with a 5.2% average rate during President Clinton's term of office." - September 3rd Wall Street Journal ("Bush Has a Good Economic Record")
In 1993, the unemployment rate was 7.3%. When Clinton left office in 2001, it was 4.2%. In 2001, when Bush took office, the rate was 4.2%. When that article was written, it had grown to 5.7%.
CREDITS Subject Matter Expert: Interactive Design:
Instructional Design:
Project Manager:
Dr. Nick Coppola
Christina Adams
Felicity Pearson
Kyle Huppert
L i c e n s e d u n d e r a C r e a t i v e C o m m o n s A t t r i b u t i o n 3 . 0 L i c e n s e .