stats assignment
EDB 601 --- Descriptive Statistics
Here we are at Descriptive Stats. All you need is a small hand-held calculator, one that has a square root function. Descriptive Statistics finds the Mode (score that appears the most times), the Mean (the average of the scores), the Median (the score that sits right in the middle of all the scores), the Variance (the amount of spread among your scores), and the Standard Deviation (the square root of the VARIANCE).
None of these is difficult to find as long as you know how to add, subtract, multiply, divide, and hit the square root button on your calculator.
So to get you started, I’m giving you 4 worksheets, along with the answer keys. Obviously I want you to work the practice sheets first and then compare your answers. There really isn’t anything magic in this. You should be able to figure it out after doing a couple of the worksheets. But remember, I’m only a call or email away if you get stuck.
So let’s get started with TABLE 1. This distribution table already has the data arranged in the far left column, starting with the smallest number and going down to the largest. The data is designated as X and in this case it is weight in kg.
The second column is Frequency or f and indicates how many times the number to its left appeared. Rather than having a distribution table with 100 lines (since there are 100 participants in this study), it’s easier to group the participants that weigh the same. So you see that there are 5 men who weigh 74 kg. and 34 who weigh 77 kg.
The third column is Xf or the number in the first column multiplied by the number in the second column. So in this column on the first line you would put the product of 74 x 5 because 5 men weighed 74 kg. You do this for all the lines, multiplying the X times the f to get the Xf.
The fourth column is the cumulative frequency. This column adds up the numbers in the f column. So on the first line, the cumulative frequency is 5. On the next line you take that 5 and add it to the number on the second f line which is 7. You get 12, so 12 goes on the second line in that column. Then you take 12 and add it to the number on the third f line which is 26. You get 38 and that number is the cumulative frequency for line three. You do this all the way down. If you added correctly, you should end up with 100, the number of men in the study.
Now that you’ve filled in the distribution chart, find the MODE, the number in the X column that appears the most times. Look at the f column to find the largest number. The MODE will be the weight to its left.
To find the MEAN add up all the numbers in the Xf column. Since you’ve already multiplied the X times the f, adding up the numbers in this third column will include all the 100 participants’ weights. To find the MEAN, you just take the total of column three and divide it by the total number of participants (100).
Finding the MEDIAN takes you to the last column, the cumulative frequency column. The MEDIAN is the score that sits in the middle of your Xs. That means that if you have 100 scores, as you do here, the middle number is the one that has as many numbers above it as below it. When you are dealing with an uneven number of scores, you just divide the total by 2 and round up one. For example, if you had only 99 scores, you divide 99 by 2 and get 49 ½. You round up to 50. There are 49 numbers below 50 and another 49 above it.
But you have 100 scores to deal with. If you divide 100 by 2, you get 50. But 50 can’t be the exact middle because there are 49 numbers below it and another 50 above it and the MEDIAN has to be in the exact middle. So here we split the difference. The MEDIAN is between 50 and 51 or 50 ½. But you don’t have anyone who is participant 50 ½. True, so very often the MEDIAN is a number that is not one of your actual scores. If this doesn’t make sense, don’t feel too bad. Many people find this really difficult to grasp.
Let’s look at your TABLE 1. Since we have 100 weights, the MEDIAN weight is the weight the lays between participant 50 and participant 51. If these two participants have the same weight, that weight is the MEDIAN and you don’t have to make any other correction. If Participant 50 and Participant 51 have different weights, then the MEDIAN weight is going to be the number that falls exactly between the two. Look in the last column and find where 50 and 51 would fall. Remember that the numbers in the cumulative frequency column includes a lot of numbers. For example, in the cumulative frequency column line 1, you have 5. That means that participants 1, 2, 3, 4, and 5 are included in that number. As you go to the second line, they are included with participants 6, 7, 8, 9, 10, 11, and 12. Now try to find where 50 and 51 fall. What is their weight?
Try to complete this before looking at the answer sheet. But once you’ve done this, go to TABLE 2. You will see that this is the same data, with an additional 2 columns. Copy your data from the TABLE 1 distribution table. Now we’ll figure out how to find the VARIANCE and the STANDARD DEVIATION.
__
The fifth column is (X – X)2 or the X minus the MEAN and then the difference squared. You got the MEAN (that’s the X with the bar over it) by adding up all the weights and dividing that number by the number of weights, or averaging the weights. Now I want you to take that number, and subtract it from each of the Xs or weights. But because some of the weights may be smaller than the MEAN weight, you may end up with a negative number. We don’t like to use negative numbers, so we get rid of them by squaring or taking that difference and multiplying it by itself. So if the difference between the weight and the MEAN is -3, we multiply -3 x -3 and get +9, since multiplying two negative numbers results in a positive number. Now whether the difference is a positive or negative number, you have to square it. Do this for each X.
The sixth column just takes what you got in the 5th column and multiplies it by the frequency or the number of times that weight shows up on the distribution table. So whatever you got on line 1, you multiply by 5 (the frequency) and put that number on line 1 in the sixth column. Do that all the way down, for each weight. When you’ve done that, add up all the numbers in that last column. That number is then divided by the total number of weights (in this case 100) and is your VARIANCE. But because that number has been squared, to find the STANDARD DEVIATION you just hit the square root button on your calculator and that square root will be your SD. Simple, huh?
This is really merely a lot of attention to detail. There isn’t any difficult math. Just take it step by step. And again, try to complete the table by yourself and then look at the answer key to see how you did.
Once you’ve completed these two worksheets, there is an assignment for you to try. It requires you to take the scores and arrange them in order. You cannot find the MEDIAN if you don’t put the scores in numerical order from smallest to largest (or largest to smallest). Some scores may appear more than once. Rather than write them down twice, you use that second frequency column to show the number of times they appear. While most of your numbers in this column will be 1, you may have 2 or more.
And finally, I ask you for the actual scores within one or two SDs. How do you do this? Easily. Take the MEAN and add the SD and then take the MEAN and subtract the SD. So for example, most IQ tests have a MEAN of 100 with a SD of 15. The scores that would fall within 1 SD would be 100 + 15 and 100 – 15 or from 85 to 115. For two SDs you’d add 30 and subtract 30 from the MEAN of 100 and get scores from 70 to 130. I asked for actual scores, so find the smallest of your scores that would fit within the 1 + SD, even if it’s larger than the smallest possible number. For example, using our IQ scores, suppose I had scores of 68, 72, 87, 101, and 116. Within one SD would be scores 87 and 101, even though the whole band would run from 85 to 115. Two SDs would include scores 72 through 116, but not 68 because it’s below the 70 that is the low end of the 2SDs. Does that make sense?
And finally, I ask you to describe the shape of the distribution. In a normal distribution, or that old bell-shaped curve we all know and love, most of the scores are arranged around the middle with only a few at each of the extreme ends. The MEAN and the MEDIAN will be the same number. If, however, the MEAN and MEDIAN are different, then the shape is described as being negatively or positively skewed. Negatively skewed means that the extreme scores (the scores that are most unlike the majority of scores) are low and so they pull down the MEAN so that it is less than the MEDIAN. Positively skewed means that the extreme scores are higher than the majority of the scores and so they pull the MEAN up over the MEDIAN.
When doing these sheets, if you end up with a MEAN that has a fraction, round up or down. You don’t have to deal with lots of decimals that way and then neither do I.