Statistics Homework
WEEK 2 HW 2 (Based on Lane C3 and Illowsky C2.5 – 2.8)
|
TAKE THE TIME TO READ THE “INTRODUCTIONS TO STATISTICS” SECTIONS IN BOTH TEXTS. They will give you a better perspective on what we are covering and why. As always, email or “message” me with any questions. I check these at least once per day, but I am NOT online 24/7. AND, if you can help a classmate out, please do so without waiting for me. If an unanswered question makes you unable to complete an assignment on time, let me know. |
|
THIS IS A VERY IMPORTANT WEEK. ALL THE BASICS OF STATISTICS ARE INTRODUCED HERE. MAKE SURE YOU ARE CLEAR ON ALL OF THIS. HERE ARE THE TERMS FOR THIS WEEK:
|
· MEAN · MEDIAN · MODE · VARIANCE · STANDARD DEVIATION |
· DISCRETE/CONTINUOUS DATA · FREQUENCY · RELATIVE FREQUENCY · CUMULATIVE RELATIVE FREQUENCY · HISTOGRAM |
· PERCENTILES · QUARTILES (Q1, Q2, Q3) · IQR (= Q3 – Q1) · BOX PLOT (5 number summary) · UNUSUAL/RARE DATA VALUES |
Last week we talked about how to collect samples from a population of interest. Random sampling is intended to insure that we collect data representative of the entire population. The key is RANDOM with EVERY selection having the same probability of being collected as every other so that no sources of error or bias are introduced. Of course in some less ethical sampling programs, bias is intended and counted on to sway unknowing decision makers (e.g. political ads).
So, how do we spot BIAS if we don’t know how the data were collected? Often, we must read opposing studies that may have their own bias in them; then, WE can balance that bias (critical thinking) and come up with our own conclusions. However, we don’t really look at these raw data, rather we review STATISTICS calculated FROM those data. Unfortunately, these same statistics can be calculated from biased data just as easily as from truly representative data. The math is the same. The validity, meaning usefulness to us, isn’t (unless we are the ones fudging the analyses).
OK, just what statistics are we talking about? Just three: MEANS, VARIANCES, AND STANDARD DEVIATIONS. With properly collected samples the means and variances of the sample data are good predictors (unbiased) of population means and variances. The standard deviations are not as valid (biased).
Just how well do our sample data statistics reflect the true population’s statistics? This is where PROBABILITY comes in. Do we want to be (or can we be) 90% certain? 95% or even 99% ? Yes. As you might imagine, sample size has a great effect on our confidence level.
That’s all there is to statistics. We collect data from samples of the population, calculate means, variances and/or standard deviations and then try to make probability-based statements about the population in regard to the parameter being analyzed: health, wealth, effectiveness, life span, etc. NOW, onto the homework problems. Each weekly HW assignment will have 10 problems and they are based on the material in the Lane and Illowsky text chapters assigned.
So, let’s start the HOMEWORK with MEANS and WHAT they can tell us (and what they can’t).
PROBLEM #1. The VALUE (usefulness) of the MEAN
Let’s say we randomly select ten (10) comparable model pick-up trucks manufactured by three (3) different vehicle companies. The data are the number of years before that vehicle required a major repair.
CALCULATE THE MEANS FOR THESE THREE SETS OF TRUCKS AND HOW THAT MEAN HELPED OR DIDN’T HELP YOU DECIDE ON WHICH BRAND TO BUY.
|
Vehicle |
COMPANY "A" |
COMPANY "B" |
COMPANY "C" |
|
1 |
9 |
8.5 |
7 |
|
2 |
5 |
5.5 |
7 |
|
3 |
8 |
7 |
7 |
|
4 |
6 |
9 |
7 |
|
5 |
12 |
5 |
7 |
|
6 |
2 |
7 |
7 |
|
7 |
13 |
8 |
7 |
|
8 |
1 |
6 |
7 |
|
9 |
10 |
9.5 |
7 |
|
10 |
4 |
4.5 |
7 |
|
MEAN |
|
|
|
PROBLEM #2. CONFIDENCE IN THE MEAN: CALCULATION OF VARIANCES AND STANDARD DEVIATIONS
in Week 1 you tried various descriptive ways of showing data distributions (e.g., bar charts). Let’s look at a basic “SCATTER PLOT” of the truck repair data from the above three vehicle manufacturers:
Do you SEE the value of data displays (DESCRIPTIVE STATISTICS) verses just a basic statistic like the mean? BUT, are there statistics that QUANTIFY the spread of data around the mean? YES ! These statistics are the VARIANCE, which simply adds up the square of all the distances (squaring makes the negative distances positive as in -3 X -3 = +9), and its SQUARE ROOT, THE STANDARD DEVIATION (the average distance of the data points to the mean, but this average is calculated a little differently from the mean as you will see).
(a) SO, LET’S CALCULATE THE VARIANCE AND STANDARD DEVIATION FOR OUR THREE SETS OF DATA
For the mean of out ten data points we added them up and divided by 10 (the number of data points in a sample data set is referred to as “n”). TO GET THE VARIANCE, HOWEVER, WE DIVIDE BY (n – 1). WHY?
This may be a little confusing, but since we already know the mean of our ten numbers, only 9 of the squares can vary as they like, but the 10th must be fixed. For example, if I give you 5 numbers: 3, 6, 7, 9, ___, but tell you the total is 30, then the missing number MUST be ____? AND, if I told you that the mean of these 5 numbers was 6, you would get the same value of the fifth number (6 * 5 = 30). Long story short, FOR SAMPLES we divide the sum of the squared distances by (n – 1) which would be (10 – 1) = 9 in these 10-number data sets. (Later, you will see that for entire POPULATIONS, we would divide by simply N (the small “n” is for samples). This (n – 1) subtraction in later chapters gets called the “Degrees of Freedom”.
(b) SO, do the variance and STANDARD DEVIATION along with the MEAN help you make a truck buying decision? How & Why ?
|
KNOW THIS: For many REAL data sets (but not all-as there are conditions that apply) approximately 68% of the data points are within plus or minus one standard deviation from the mean, about 94% are within + two SD’s of the mean and 99% are within + three SD’s of the mean. REMEMBER THIS. For example if the mean is 10 and the SD is + 3, then in theory we could expect 68% of our data points to be between 7 and 13. |
PROBLEM #3: There are other statistics and displays used for various purposes that we can calculate or generate for data sets: medians, modes, quartiles, IQR’s and outliers. Additional displays include the “5-Number Summary” and the BOX PLOT. These are still all DESCRIPTIVE STATISTICS, which simply apply to SAMPLE data sets. In later weeks when we use our SAMPLE data to predict things about the POPULATION we will be using INFERENTIAL STATISTICS, which involves probability, BUT the means and variance are the heart of these topics as well.
NOW, CALCULATE THE MEDIAN, MODE AND RANGE FOR OUR THREE DATA SETS AND EXPLAIN IF ANY OF THESE STATISTICS AFFECT YOUR INTREPRETATION OF THE DATA AND WHICH VEHICLE YOU MIGHT BUY? WHY OR WHY NOT?
|
Vehicle |
COMPANY "A" |
COMPANY "B" |
COMPANY "C" |
|
1 |
9 |
8.5 |
7 |
|
2 |
5 |
5.5 |
7 |
|
3 |
8 |
7 |
7 |
|
4 |
6 |
9 |
7 |
|
5 |
12 |
5 |
7 |
|
6 |
2 |
7 |
7 |
|
7 |
13 |
8 |
7 |
|
8 |
1 |
6 |
7 |
|
9 |
10 |
9.5 |
7 |
|
10 |
4 |
4.5 |
7 |
|
MEAN |
|
|
|
|
MEDIAN |
|
|
|
|
MODE |
|
|
|
|
RANGE |
|
|
|
MOVING ON TO ILLOWSKY CHAPTER 1
PROBLEM #4:
This problem is a review of your knowledge of expressing fractions. Let’s say we have 40 students enrolled in a Stat 200 class: 1/5 live outside the U.S., 25% have a slow internet connection, and 12 are under age 30. Compute the following:
(a) How many and what % of students live outside the U.S. ?
(b) How many and what fraction (e.g., ½, 1/6, ?) and what decimal fraction do NOT have a slow internet connection?
(c) How many and what percent are OVER age 30
(d) IF you were to write all the ages on separate pieces of paper, put them in a hat and blindly draw one out, what are the chances (probability) that age would be under 30 years? (This is where we go next with our studies)
PROBLEM #5: NOW, Let’s start to make the transition from DESCRIPTIVE to INFERENTIAL statistics. This is the introduction to the PROBABILITY (what are the odds) part of Statistical Analyses.
HERE IS A DATA SET WITH 30 VALUES (REFERRED TO AS “DISCRETE” SINCE ALL ARE WHOLE NUMBERS. IF THEY WERE FRACTIONAL THEY WOULD BE CONSIDERED “CONTINUOUS”. THIS MAKES A DIFFERENCE LATER ON:
80, 71, 81, 99, 1, 54, 55, 16, 20, 27, 61, 62, 79, 68, 35, 37, 38, 41, 45, 49, 50, 21, 27, 50, 51, 55, 55, 60, 61, 70
SO, WHAT DO WE DO NOW?
(a) The FIRST thing that should occur to you is that these values are all mixed up and have repeats. RANK ORDER them from LOWEST to HIGHEST (1 to 99 in this case). Do it and make sure you did not leave out any data points. (I put in the header numbers as you will need them later on for QUARTILES).
|
1 |
2 |
3 |
4 |
5 |
6 |
7 |
8 |
9 |
10 |
11 |
12 |
13 |
14 |
15 |
16 |
17 |
18 |
19 |
20 |
21 |
22 |
23 |
24 |
25 |
26 |
27 |
28 |
29 |
30 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
(b) This data set has only 30 values. Imagine if you were dealing with a data set that had 10,000 or more data points. SECOND thing is to sort these data into logical groups like the number of DATA POINTS between 1 and 10, 11 and 20, etc., up to 91 to 100. The important thing here is that each range must be the same, and recognize too that some range(s) may have zero data in them. DO IT.
|
RANGE |
# OF DATA POINTS IN EACH RANGE (FREQUENCY) |
|
1-10 |
|
|
11-20 |
|
|
21-30 |
|
|
31-40 |
|
|
41-50 |
|
|
51-60 |
|
|
61-70 |
|
|
71-80 |
|
|
81-90 |
|
|
91-100 |
|
LET’S INTRODUCE THREE MORE STATISTICS TERMS: FREQUENCY, RELATIVE FREQUENCY, AND CUMULATIVE RELATIVE FREQUENCY (likely Quiz and Final Exam problems)
As included in the Table header above, FREQUENCY is simply the actual number of data points in any given range. If you had 5 data points in the 11-20 range, that FREQUENCY would be 5.
RELATIVE FREQUENCY transitions us into probability in that it is the number of data points in a given range divided by the total number of data points. Continuing our example of 5 data points in the 11-20 range, the would be a RELATIVE FREQUENCY of 5 / 30 = 0.167 or 16.7% Meaning that about 16.7% of our data is in this 11-21 range.
Look at this another way: let’s say that each of our data values was written on a ping pong ball and put in a bucket. We reach in blindly and pull out one ball. What would you say is the likelihood (“probability” in statistical terms) that the data value on this ball is between 11 and 21? You had better say 16.7% or about 17 chances out of 100. Not great odds, but better than Vegas.
CUMULATIVE REALTIVE FREQUENCY (CFR) has us account for ALL possibilities. AFTER you have calculated all the Relative Frequencies simply ADD THEM UP as you go down the column. for example if 3 data points were in the 1-10 range, that would be a relative frequency of 3/30 = 0.10 or 10%, That number would go in the top box of the Cumulative Frequency column. We earlier determined that the Relative Frequency of data points in the 11-20 range is 16.7%, SO in that Cumulative frequency box we add the 0.10 to the 0.167 to get the Cumulative Relative Frequency 0.267, which means that about 26.7% of our data are in the first two ranges accounting for data values from 1 to 20.
Keep adding up the Relative Frequencies as you go down the column with the bottom box ALWAYS ending up at 1.00, since by that point all 30 (100% or 1.00) of the data points are accounted for. THE CUMULATIVE REALTIVE FRQUENCY IN ANY BOX IS THE PERCENT OF DATA POINTS AT OR BELOW THE MAXIMUM VALUE OF THAT BOX’S RANGE. For example, If the box’s range is 31 to 40 then the CRF would be the percentage of data AT OR BELOW 40. KEEP READING THIS PART OVER UNTIL IT MAKES SENSE AS IT IS THE HEART OF STATISTICS.
(c) FILL IN THIS TABLE:
|
RANGE |
FREQUENCY (FROM ABOVE) |
RELATIVE FREQUENCY |
CUMULATIVE REL FREQ |
|
1-10 |
|
|
|
|
11-20 |
|
|
|
|
21-30 |
|
|
|
|
31-40 |
|
|
|
|
41-50 |
|
|
|
|
51-60 |
|
|
|
|
61-70 |
|
|
|
|
71-80 |
|
|
|
|
81-90 |
|
|
|
|
91-100 |
|
|
|
|
REVIEW: The frequency is a whole number. The relative frequency (and cumulative rel freq) when first set up are fractions (e.g., 6/30) then calculated decimal fractions like 6/30 = 0.20, which can be converted to a percentage 0.20 x 100% = 20%. Get used to these versions. |
PROBLEM #6 Interpreting Frequency Distributions (USE THE CUMULATIVEREALTIVE FREQUENCY COLUMN IN THE TABLE ABOVE)
a. What is the percentage of this data between 21 and 60?
b. What percentage of these data are 71 or GREATER?
c. What is the relative frequency of these data at or under 50?
d. What is the cumulative relative frequency for these data less than 100? (kind of a trick question)
PROBLEM #7 DATA PLOTS (HISTOGRAMS)
We need to DISPLAY our data as a special bar chart called a HISTOGRAM. A histogram gives us a picture of the shape of our data distribution, which is critical to know, as well as a way to calculate percentages (probabilities) of data values in any range or set of ranges.
This is the ONLY time you will construct one of these so do it by hand unless you are facile with Excel and can get it to widen the bars to reduce the gap between them to zero. (Points deducted if bars don’t touch).
The ranges are along the horizontal x-axis ( for simplicity just label them range 1, 2, 3, etc.) and the number of data points in each range is up the vertical y-axis. Use the FREQUENCY column in the earlier Table.
(a) DRAW THIS HISTOGRAM FOR FREQUENCIES
(b) THEN, DRAW A HISTOGRAM FOR RELATIVE FREQUENCIES (HOW DO THESE TWO COMPARE?)
(c) LOOK at the plots – Do they look like a BELL CURVE (the Normal Distribution) OR are they SKEWED (positive or negative)?
(d) Sketch a plot that has a POSITIVE and one that has a NEGATIVE skew (recognizing skewness has been a typical final exam question, so know what they look like)
PROBLEM #8 PERCENTILES (e.g., QUARTILES) (Related to LANE C-1 - Formula on pages 29-31)
We will use the FORMULA approach (Lane’s third definition) to calculate PERCENTILES (SEE BLUE BOX). QUARTILES are specific percentiles: 25th, 50th, 75th. Q1 is the 25th percentile and means that 25% of the data set is below that data value. Q2 is the 50th percentile also called the MEDIAN and means that 50% of the data are below that data value. Logically, the 75% percentile means that 75% of the data are BELOW that data value. Usually, we would say something like you need to be in top 25th percentile (top 25% of your high school graduating class) to get into a college. BUT, we determine this in reverse meaning that you need to be HIGHER than the 75th percentile value. KEEP THIS WAY OF THINKING IN MIND. IF 75% are below then 25% MUST be above that 75th percentile value since all data values are accounted for.
|
CALCULATING RANK = percentile/100 * (n+1) AND if the Rank ends up with a fraction (FR in Lane), e.g. 6.4, multiply that fraction (0.4) times the difference between the actual data points with the ranks above and below 6.4 which in this case are the data points with ranks 6 and 7. Add the result to the lower data point value. |
USE THE EARLIER RANK ORDER TABLE OF THE 30 data points arranged from LOWEST to HIGHEST. Don’t confuse RANK with percentile. The RANK is just the position of a data point in the ranked list of all data points. Once you calculate the rank, just count down that many data points to the data value.
(a) Look at the list and “EYEBALL” (don’t calculate) what you think is the 25th percentile (this is the FIRST QUARTILE or Q1), then the MEDIAN (this is the 50th percentile or Q2), and finally the 75th percentile (the THIRD QUARTLE or Q3). (You can also try using Lane’s first and second definitions of “percentile”. Excel gives slightly different results in some cases, but doesn’t say what formula it is using)
(b) NOW, using the FORMULA, calculate these three QUARTILES: Q1, Q2 and Q3. How close did you come?
(c) Calculate the Inter-Quartile Range (IQR) which is simply Q3 – Q1.
(d) FINALLY, complete a 5-Number Summary and BOX PLOT for YOUR data. LABEL OR DESCRIBE WHAT EACH OF THE PARTS OR LINES IN THE BOX PLOT TELL US ABOUT YOUR DATA SET.
PROBLEM #9. REVIEW: If you are in the 75th PERCENTILE this means that 75% of the scores (data points) are below yours. To actually enroll to be on-campus at many large State Universities (like Penn State) you need to be in the top 10% of your high school graduating class. What percentile would this be (tricky)?
We have been calculating RANK PERCENTILES, but what if we want to know just what percentile a given data point actually is? Here is that formula: SPECIFIC DATA POINT’S Percentile = [ (x + 0.5 y) / n ] x 100%
Where the “x” is the number of data points BELOW the given point added to 0.5 times the number of duplicate scores at that given data point (y), all divided by the total number of data points (n). Multiplying that decimal fraction times 100% gives us the percentile for our specific data point. Using our 30 data points determine the following:
(a) What percentile are the scores of 80 and 54 ?
(b) If you need to be in 80th percentile, what score is that? (uses the RANK formula)
PROBLEM #10. The last concept we want to cover is “rare events” or “unusual values”. How do we identify them? This does NOT mean that if a data point is “unusual” we can simply delete it. Any deletions MUST be justifiable. Imagine if a pharmaceutical company deleted the one death in 100 patients taking a new medication because it was “unusual”. Want to take that pill?
An UNUSUAL data point can be characterized two ways (for us) DO BOTH FOR OUR 30 DATA POINTS:
(a) Any data point that is more than TWO (2) standard deviations ABOVE OR BELOW the mean.
(b) Any data point that is more than Q3 + (1.5 * IQR) ABOVE OR BELOW the mean.
(c) ARE ANY DATA POINTS “UNUSUAL”?
|
Next week we move into PROBABILITY which is the basis for INFERENTIAL STATISTICS. This allows us to test hypotheses about a POPULATION, based on samples taken from that population. This is what statistics is all about. It starts with PROBABILITY, meaning that life is a “crap shoot”. Failing to realize that, we have the famous last words: “What could go wrong?” |