Quantitative Assignment- (Statistics Assignment)
Quantitative Methods/Choosing a Sample.pptx
Choosing a Sample
Leedy, P., and Ormrod, J., Practical Research. (8th ed.)
Fink, A. 1995. From the Survey Toolkit published by Sage.
Choosing a Sample to Survey
Population – the group to be covered by your research plan
Sample – a subset of your population
Generalize results – only if the sample is representative of the population
Probability sampling
Non-probability sampling
2
Probability sampling – Random Sampling
Each member of the population has an equal chance of being selected.
3
Probability sampling – Stratified Random Sampling
Take equal samples from each group (layers, strata).
4
Probability sampling – Proportional Stratified Sampling
Take equal proportions of samples from each group (layers, strata).
5
Probability sampling – Cluster Sampling
Take equal proportions of samples from certain regions only.
6
Non-probability Sampling – Convenience Sampling
No attempt to have a representative sample
Examples:
Survey people in your neighborhood.
Customer satisfaction cards in a restaurant.
Survey all companies who have had projects done by NWMOC.
Survey all Human Resources Managers at Stout Career Fair.
7
Sample size?
Entire population, if N<100
20-50% of population, if 100 < N < 2000
About 400, if N > 2000
Affects the time and cost of the study, the precision of statistical results
Be sure to consider the response rate
8
Sampling Bias
Bias – an influence, condition, or set of conditions which distort the data
Sampling bias – is the sample random?
Examples:
Political polls by phone interview
A mail survey of alumni satisfaction, with 30% response rate
9
10
Dilbert
Click to edit Master text styles
Second level
Third level
Fourth level
Fifth level
Quantitative Methods/Confidence Intervals.pptx
Confidence Intervals
Some adapted from http:// stattrek.com/estimation/confidence-interval.aspx and http://www.stat.yale.edu/Courses/1997-98/101/confint.htm
1
Confidence Interval
Statisticians use a confidence interval to describe the amount of uncertainty associated with a sample estimate of a population parameter.
It gives an estimated range of values which is likely to include an unknown population parameter,
the estimated range is calculated from a set of sample data.
Confidence Interval
Gives the probability that the interval produced by the sample method includes the true value of the parameter
You must assume a normal distribution
Confidence Interval Selection
Common choices for the confidence level are 0.90, 0.95, and 0.99. These levels correspond to percentages of the area of the normal density curve. For example, a 95% confidence interval covers 95% of the normal curve --
Normal Distribution
Confidence Intervals
Suppose that a 90% confidence interval states that the population mean is greater than 100 and less than 200. How would you interpret this statement?
It does not mean there is a 90% chance that the mean of the ENTIRE population falls between 100 and 200. The population mean is a constant, not a random variable. It does not change
Confidence Intervals
The confidence level describes the uncertainty associated with a sampling method. Suppose we used the same sampling method to select different samples and to compute a different interval estimate for each sample. Some interval estimates would include the true population parameter and some would not
Confidence Interval
A confidence interval based on a sample does not predict that the true value of the parameter has a particular probability of being in the confidence interval given the data actually obtained.
To construct a confidence interval you need to know the Z value for your interval
Confidence Interval
You must determine the percentage for your confidence Interval – generally assume 95%
For a 99% confidence interval
Z= 2.576
For a 95% confidence interval
Z= 1.960
For a 90% confidence interval
Z=1.645
Confidence interval on mean
10
Confidence interval on mean: Example
You sample number of orders shipped per day
n = 64 days, M = 120 orders per day, s = σ = 25 orders per day
Confidence interval is:
120–(1.96(25))/√64 < μ < 120+(1.96(25))/ √64
120-(49/ 8) < μ < 120+(49/8)
120-6.1 < μ < 120+6.1
Or, 113.9 < μ < 126.13
So, the number of orders per day is between 113.9 and 126.1
11
Confidence interval on mean: Graphically
Example
For the daily order data, are the number of orders significantly different than X=115 per day? Different than X=100 per day?
You cannot disprove that the average number of daily orders is 115
You can disprove that the average number of daily orders is 100
Graph of Values
113.9
126.1
120 – sample mean
115
100
Outside the confidence interval –
Reject null hypothesis here and
Accept the alternative hypothesis
Within the confidence interval –
Retain null hypothesis here
Confidence interval on mean: Small n
Notice that the width of the confidence interval increases as n decreases:
For n = 200, 116.5 < μ < 123.5
For n = 64, 113.9 < μ < 126.1
For n = 30, 111 < μ < 129
A larger sample is better
A smaller sample gives you less information
Interpretation of results
Relate the statistical results to the original research problem
Compare results to existing literature
Is there practical significance to the results?
What are limitations of the study?
Examples
In Hawaii, surfing is popular
Most important factor for good day of surf is size of swell (wave height), the average sizes of swells throughout a given timeframe, and the consistency (ride length & wave direction) of a swell
CI – measure probability of how high and what wave direction waves will travel at a given timeframe
Example
A business might estimate a machine will use 10 lbs of plastic for each unit
No machine will always exactly use precisely 10 lbs per unit, a CI is created to give a range
The company might predict that there is a 95% chance that the machine uses, on average, between 9.85 and 10.5 lbs of plastic per unit
0
10
20
30
40
50
60
90100110120130140150
Mean of 120
Mean of 115
Mean of 100
Quantitative Methods/Correlation Introduction.pptx
Correlation
1
Correlation is a way to measure how associated or related two variables are.
The researcher looks at things that already exist and determines if and in what way those things are related to each other.
The purpose of doing correlations is to allow us to make a prediction about one variable based on what we know about another variable.
Correlation
For example, there is a correlation between income and education. We find that people with higher income have more years of education. (You can also phrase it that people with more years of education have higher income.) When we know there is a correlation between two variables, we can make a prediction. If we know a group’s income, we can predict their years of education.
Correlation
A key thing to remember when working with correlations is never to assume a correlation means that a change in one variable causes a change in another.
Sales of personal computers and athletic shoes have both risen strongly in the last several years and there is a high correlation between them, but you cannot assume that buying computers causes people to buy athletic shoes (or vice versa).
However, you use dependent and independent variables – the practice problem will require this
Correlations
We can make predictions about things when we know about correlations – assuming there is foundation in theory
If two variables are correlated, we can predict one based on the other. For example, SAT scores and college achievement are positively correlated.
So when college admission officials want to predict who is likely to succeed at their schools, they choose students with high SAT scores.
Correlation
The problem with the correlation method is the assumption that because variables are significantly correlated, one or more variables cause a change in another variable
Take a minute and say to yourself: Correlation is not Causation!
Correlation
It is a measure of the relation between two or more variables. The measurement scales used should be interval or ratio scales
Correlation coefficients can range from -1.00 to +1.00.
The value of -1.00 represents a perfect negative correlation
While a value of +1.00 represents a perfect positive correlation.
A value of 0.00 represents a lack of correlation
What is correlation
Pearson correlation coefficient, r
Strength:
The closer to +1 or -1, the stronger the correlation
The closer the data follows a line of best fit
Correlation coefficient
8
Examples
Positive correlation, r > 0
Negative correlation, r < 0
Strong correlation, r → 1
Weak correlation, r → 0
9
We know that education and income are positively correlated.
We do not know if one caused the other.
It might be that having more education causes a person to earn a higher income.
It might be that having a higher income allows a person to go to school more.
It might also be some third variable.
Correlation
There is a relationship between years of education and salary
The problem is you don’t know if the correlation is significant
Correlation
To determine if the correlation is significant you use regression: Line of best fit equation:
Y-intercept, b, where line crosses Y axis
Slope, m, change in Y over change in X
Line equation, Y = mX + b
Line of best fit (regression)
Data follows a line of best fit
Data does not follow line of best fit
12
Tools, Data Analysis
Select Correlation, OK
Input data range: Highlight ALL the variables
Labels in first row
Output range: give top left corner
Correlation on Excel
Tools, Data Analysis
Select Regression, OK
Input Y range: dependent variable label and data in a column
Input X range: independent variable label and data in a column
Labels in first row
Confidence level, 95%
Output range: give top left corner
Regression analysis on Excel
14
Multiple R = r, correlation coefficient
You have to do a regression to find out about the correlation significance – we don’t cover the other uses of multiple regression in INMGT 700
CRITICAL – Unless otherwise specified – you ALWAYS use a .05 p value for significance in correlation.
If your result is .05 or SMALLER, the correlation is significant – if it is larger than .05 (.051) then it is NOT significant
Excel regression output:
15
Is correlation significant?
In Excel output, you look at the p value OF THE INDEPENDENT VARIABLE in the output – if it is less than .05 then the correlation is significant, if it is over .05 it is NOT significant
Excel regression output:
Quantitative Methods/Correlation Regression - Example and Practice.xlsx
Example
| Data from a Study on SAT scores, High School grades in Math, Science and English. | |||||
| Their University GPA is after 3 semesters. These are all Computer Science Majors at University | |||||
| (Note: this university has a different GPA system - not a 4 point scale) | |||||
| Observ. | SAT-M | SAT-V | HSM | HSS | GPA |
| 1 | 640 | 530 | 8 | 6 | 4.35 |
| 2 | 670 | 600 | 9 | 10 | 4.08 |
| 3 | 600 | 400 | 8 | 8 | 5.21 |
| 4 | 570 | 480 | 7 | 7 | 4.34 |
| 5 | 510 | 530 | 6 | 8 | 3.4 |
| 6 | 750 | 610 | 10 | 9 | 3.43 |
| 7 | 650 | 460 | 8 | 9 | 4.48 |
| 8 | 720 | 630 | 10 | 10 | 5.73 |
| 9 | 760 | 500 | 10 | 10 | 5.8 |
| 10 | 640 | 670 | 9 | 6 | 4 |
| 11 | 640 | 490 | 10 | 9 | 5.16 |
| 12 | 520 | 360 | 9 | 8 | 4.73 |
| 13 | 700 | 520 | 7 | 8 | 3.07 |
| 14 | 490 | 550 | 6 | 8 | 3.82 |
| 15 | 640 | 520 | 10 | 10 | 5.12 |
| 16 | 550 | 290 | 9 | 7 | 4.25 |
| 17 | 600 | 520 | 10 | 10 | 4.93 |
| 18 | 710 | 530 | 10 | 9 | 4.83 |
| 19 | 750 | 670 | 9 | 10 | 5.1 |
| 20 | 620 | 480 | 9 | 9 | 4.87 |
| 21 | 630 | 440 | 10 | 10 | 5.61 |
| 22 | 770 | 720 | 10 | 7 | 4.75 |
| 23 | 610 | 560 | 10 | 10 | 5.26 |
| 24 | 640 | 570 | 10 | 10 | 5.67 |
| 25 | 650 | 480 | 10 | 10 | 5.3 |
| 26 | 660 | 630 | 10 | 10 | 5.62 |
| 27 | 570 | 480 | 7 | 8 | 4.55 |
| 28 | 690 | 550 | 9 | 7 | 5.25 |
| 29 | 670 | 500 | 7 | 7 | 4.21 |
| 30 | 660 | 460 | 10 | 9 | 4.5 |
| 31 | 600 | 630 | 8 | 8 | 5.03 |
| 32 | 447 | 320 | 9 | 10 | 3.92 |
| 33 | 580 | 470 | 6 | 8 | 4.7 |
| 34 | 630 | 630 | 9 | 7 | 4.96 |
| 35 | 600 | 560 | 10 | 10 | 4.76 |
| 36 | 550 | 560 | 9 | 10 | 5.4 |
| 37 | 630 | 500 | 8 | 8 | 4.48 |
| 38 | 750 | 760 | 10 | 10 | 5.86 |
| 39 | 491 | 391 | 9 | 8 | 4.62 |
| 40 | 550 | 500 | 7 | 8 | 5.72 |
Practice
| Observ. | HSM | SAT-V | GPA |
| 1 | 8 | 530 | 4.35 |
| 2 | 9 | 600 | 4.08 |
| 3 | 8 | 400 | 5.21 |
| 4 | 7 | 480 | 4.34 |
| 5 | 6 | 530 | 3.4 |
| 6 | 10 | 610 | 3.43 |
| 7 | 8 | 460 | 4.48 |
| 8 | 10 | 630 | 5.73 |
| 9 | 10 | 500 | 5.8 |
| 10 | 9 | 670 | 4 |
| 11 | 10 | 490 | 5.16 |
| 12 | 9 | 360 | 4.73 |
| 13 | 7 | 520 | 3.07 |
| 14 | 6 | 550 | 3.82 |
| 15 | 10 | 520 | 5.12 |
| 16 | 9 | 290 | 4.25 |
| 17 | 10 | 520 | 4.93 |
| 18 | 10 | 530 | 4.83 |
| 19 | 9 | 670 | 5.1 |
| 20 | 9 | 480 | 4.87 |
| 21 | 10 | 440 | 5.61 |
| 22 | 10 | 720 | 4.75 |
| 23 | 10 | 560 | 5.26 |
| 24 | 10 | 570 | 5.67 |
| 25 | 10 | 480 | 5.3 |
| 26 | 10 | 630 | 5.62 |
| 27 | 7 | 480 | 4.55 |
| 28 | 9 | 550 | 5.25 |
| 29 | 7 | 500 | 4.21 |
| 30 | 10 | 460 | 4.5 |
| 31 | 8 | 630 | 5.03 |
| 32 | 9 | 320 | 3.92 |
| 33 | 6 | 470 | 4.7 |
| 34 | 9 | 630 | 4.96 |
| 35 | 10 | 560 | 4.76 |
| 36 | 9 | 560 | 5.4 |
| 37 | 8 | 500 | 4.48 |
| 38 | 10 | 760 | 5.86 |
| 39 | 9 | 391 | 4.62 |
| 40 | 7 | 500 | 5.72 |
Sheet3
Quantitative Methods/Descriptive Statistics.pptx
Descriptive Statistics
1
Describe what one data set looks like:
Measures of central tendency
Measures of the dispersion of the data
Measures of the shape of the plotted data
Descriptive statistics
2
Scores of central tendency:
For data values X1, X2, ….Xn
Mean = Average or center of balance (for normal data)
M or μ = ∑ X / n
Median = Middle position value (for skewed data)
For ranked data, position (n+1)/2 or average of n/2 and (n+1)/2
Mode = Number which occurs most frequently
3 4 5 5 6 9 15 17 125
What is the mean? The median? The mode?
Mean = 21, median = 6, mode = 5
Central Tendency: the central point around which the data revolve
3
Variance = how close the data are to the mean.
Large variance = widely scattered
Small variance = close to mean
Who cares?
Example – Team decision-making – when evaluating criteria - the larger the variance the less agreement
Measures of Variability
4
Range = Highest score minus lowest score
Standard deviation (of a sample) = index of a distribution’s spread
σ or s = √ ∑ (X – M)2/n-1
Variance is standard deviation squared
σ 2 or s 2 = ∑ (X – M)2/n-1
For previous example, what are R and s 2?
R = 122, s = 39.3, s 2 = 1545.25
Measures of Variability:
5
| Scores | X-m | (x-m)2 | Selected Excel Output | |
| 3 | -18 | 324 | Mean | 21 |
| 4 | -17 | 289 | Std Error | 13.103 |
| 5 | -16 | 256 | Median | 6 |
| 5 | -16 | 256 | Mode | 5 |
| 6 | -15 | 225 | Std Deviation | 39.31 |
| 9 | -12 | 144 | Sample Variance | 1545.25 |
| 15 | -6 | 36 | Count | 9 |
| 17 | -4 | 16 | Range | 122 |
| 125 | 104 | 10816 | ||
| (∑(x-m)2)/n-1 | 1545.25 = s2 | |||
| √(∑(x-m)2)/n-1 | 39.310 = s |
Calculating Variance & Standard Deviation
Assuming the distribution is normal –
With the standard deviation you can compute the percentile rank of any number
It is used for inferential statistical tests to find significance
Used to estimate the population when it is not feasible to test entire population
Why is it important?
1 Standard Deviation contains ~ 68% of the data
2 Standard Deviations contain ~ 95% of the data
3 Standard Deviations contain ~ 99.7% of the data
Standard Deviation
Shape of the data plot: Distributions
Normal distribution
Uniform distribution
Exponential distribution
9
Skewness
Positive (right) skew
Negative (left) skew
10
Negatively Skewed Test Data
Histogram
Frequency 2 4 6 8 10 12 14 16 18 20 More 0 0 0 3 4 10 12 12 12 8 0Bins
Frequency
Tools, Data Analysis
Select Histogram, OK
Input range: include label and data in a column
Labels, in first row of data
Output range: give top left corner
Chart output, others as desired
Need to enter “bins” – a bin is a range” whereby you can group your individual data points. The bin size will vary depending on your range. You need to have a column in your data set with the top value of the range.
Histograms on Excel
12
Tools, Data Analysis
Select Descriptive Statistics, OK
Input range: include label and data in a column
Labels in first row
Output range: give top left corner where you want the output to load
Summary statistics, others as desired
You need to do a separate data analysis if you want a histogram
Descriptive statistics on Excel
13
Low
Mid
High
Low
Mid
High
Low
Mid
High
Low
Mid
High
Low
Mid
High
Quantitative Methods/Hypotheses.pptx
Variables and Hypotheses
Adapted from Practical Research
And other material
1
Variable –
any characteristic on which the elements of a sample or population differ from each other
Height, weight, sex, national origin, age, grade, number of sick days, etc.
Variables
Defines the principle focus of research interest
It is presumably affected by one or more independent variables
The independent variables are presumed to determine the value of the dependent variable
In a research study on the relationship between number of mosquitoes and number of mosquito bites, the number of mosquitoes bites is the dependent variable
Dependent Variable
Antecedent conditions that are presumed to affect a dependent variable
They are manipulated by the researcher or are observed/measured by the researcher
In a research study on the relationship between mosquitoes and mosquito bites, the number of mosquitoes per acre of ground would be an independent variable
The number of mosquitoes will influence the number of bites
Independent Variable
5
Hypotheses
Definition: A starting point for further investigation from known facts.
Hypotheses are never proved or disproved; they are either supported or not supported by data.
When data does not support a particular hypothesis, the researcher rejects the hypothesis.
6
Hypotheses, Continued
The hypotheses provide specific statements about what the investigator has tested and reported.
The hypotheses are derived directly from the statement of the problem.
7
Criteria for Good Hypotheses, continued
There should be a basis for the formulation of the hypothesis– it should be:
derived from theory,
from the findings of the related empirical research of others, or
from logical argument based on expert opinion and/or personal experience.
The Literature
8
Criteria for Good Hypotheses, continued
Hypotheses should be testable.
The researcher should be certain that any stated hypothesis can be tested by some objective means.
Hypotheses should be as concise and clear as possible.
“The simplest way is the best way.”
9
Criteria for Good Hypotheses
Hypotheses should be stated in declarative sentence form and should state an expected relationship or difference between two or more variables.
The direction or nature of the relationship(s) should be specified in the hypotheses. If a relationship (or difference) can be hypothesized to exist, then the nature of that relationship (or difference) can be hypothesized.
Criteria for Good Hypotheses
For example
The greater the number of mosquitoes per acre the greater the number of mosquito bites
10
11
Criteria for Good Hypotheses, continued
The independent and dependent variables should be identified and should be operationalized in measurable terms in a hypothesis.
For example, rather than saying “achievement of students,” the variable would be operationalized as grade point average or score on the XYZ achievement test.
Students that score higher on the XYZ achievement test will have higher GPAs
12
Hypothesis Testing
Hypothesis testing will determine whether particular experimental results are due to the manipulation of the independent variable or to chance fluctuations in the population.
Example: Did the safety training program (independent variable) improve the safety record of the department (dependent variable)
Hypothesis testing
Testing the credibility of a specific statistical hypothesis.
A statistical hypothesis is a mathematical expression (it is what the statistics test for) which should be related to your research hypothesis. It should also be represented in words
Use your statistical results to interpret your research hypothesis.
Null and alternate hypotheses
Null hypothesis, H0
Tests equivalence
What you hope to disprove
“The difference in test scores before and after training = 0”
Alternate hypothesis, H1
Tests inequality
What you hope to prove
“The difference in test scores before and after training ≠ 0”
Null & Alternative
Null: The safety training program (independent variable) did not significantly improve the safety record of the department (dependent variable)
Alternative: The safety training program (independent variable) significantly improved the safety record of the department (dependent variable)
15
16
Null Hypothesis
A statement that no difference exists between the populations being compared:
Null hypothesis: There is no significant difference in college graduation rates between students who were referred to the counseling service and students who utilized the counseling service at their own initiative.
Alternative Hypotheses
A statement that there is a statistically significant difference between the populations being compared:
Alternative hypothesis: There is a statistically significant difference in college graduation rates between students who were referred to the counseling service and students who utilized the counseling service at their own initiative.
17
Hypothesis test results
| Conclusion: training not shown beneficial | Conclusion: training was beneficial | |
| Reality: training was not beneficial | Correct conclusion | Incorrect conclusion, Type I error, α |
| Reality: training was beneficial | Incorrect conclusion, Type II error, β | Correct conclusion |
Type I and Type II errors
Type I error, α, (significance level)
You are making a claim which is false
This is more significant
Controlled - usually set to 0.05
Type II error, β
You fail to make a claim which is true
Uncontrolled – depends on α, sample size, and reliability of your data
Two ways that test hypotheses
Confidence interval:
Shows the possible population means that could generate your sample data.
Example: Does the confidence interval on the difference in test scores contain 0?
P-value:
Is the probability that the data supports the null hypothesis.
Example is the P-value small enough that the difference in test scores is not 0?
Quantitative Methods/Statistical Experiments.pptx
Experiments
Experimental Design
Experimental design attempts to prove cause-and-effect relationships
The methodology must be planned carefully to insure proper statistical results
You can statistically prove correlation or cause-and-effect
You cannot prove lack of correlation or cause-and-effect. You may just lack the evidence.
Control
Independent variable
Confounding variables
Dependent variable
Methods to control confounding
Keep some things constant
Include a control group
Randomly assign subjects to groups
Use matched pairs (repeated measures)
Expose participants to all conditions
Statistical control
Experimental Case Study (1)
Descriptive statistics only
X1 X2 X3 X4
:
Stats
One Group Pretest-Posttest (2)
Demonstrates change only
Not necessarily cause-and-effect
XA1 XA2 XA3 XA4
:
Stats
XB1 XB2 XB3 XB4
:
D1 D2 D3 D4
:
Control Group Designs (3, 6)
Is there cause-and-effect?
It depends on the group assignments
X1 X2 X3 X4
:
Stats
Y1 Y2 Y3 Y4
:
Stats
Is there a difference?
Pretest-Posttest Control Group (4,7,8)
Demonstrates cause-and-effect
XA1 XA2 XA3 XA4
:
Stats
XB1 XB2 XB3 XB4
:
DX1 DX2 DX3 DX4
:
YA1 YA2 YA3 YA4
:
Stats
YB1 YB2 YB3 YB4
:
DY1 DY2 DY3 DY4
:
Is there a difference?
What statistics do we need?
Statistical analysis of:
A single data set – use the tool of Descriptive Statistics
Difference between two data sets – use the tool of Confidence Interval
Paired differences between two data sets - Hypothesis tests (correlation/t-test/z-test) Inferential statistics
0
50
100
1st Qtr2nd Qtr3rd Qtr4th Qtr
Sales
Quantitative Methods/T-TEST and Z-TEST.pptx
Statistical Estimation – T-Tests & Z-Tests
1
One or Two-Tailed Tests
In practice, you should use a one‐tailed test only when you have good reason to expect that the difference will be in a particular direction. A two‐tailed test is more conservative than a one‐tailed test because a two‐tailed test takes a more extreme test statistic to reject the null hypothesis.
Two-Tailed Test
You suspect that a particular class's performance on a proficiency test is not representative of those people who have taken the test. The national mean score on the test is 74.
A test statistic in either tail of the distribution (positive or negative) will lead to the rejection of the null hypothesis of no difference
One-Tailed Test
The Acme Drug Company develops a new drug, designed to prevent colds. The drug is said to be more effective for women than for men. The test is a simple random sample of 100 women and 200 men from a population of 100,000 volunteers.
The alternative hypothesis would only be accepted if the women’s score was significantly higher than the men’s score at a .05 level
P Value
To determine significance with these 2 tests we use the p parameter
What Does P-Value Mean? The level of marginal significance within a statistical hypothesis test, representing the probability of the occurrence of a given event.
P Value
The p-value used to provide the smallest level of significance at which the null hypothesis would be rejected.
The smaller the p-value, the stronger the evidence is in favor of the alternative hypothesis
For social science research a p value of .05 is used
P Value
That means – of the p value is .05
We say the results of a t-test or z-test with a p value of .05 or less is significant
There is a 5% chance that the difference in the variables occurs by chance
What is a Z-Test
A statistical test used to determine whether two population means are different when the variances are known and the sample size is large.
The test statistic is assumed to have a normal distribution and parameters such as standard deviation/variance should be known in order for an accurate z-test to be performed.
Hypothesis tests between two data sets
For n>=30 you use a Z-Test
Both sample sizes MUST be 30 or more
Calculate sample means, M1 and M2, and sample standard deviations σ1 and σ2 and variance
In Excel, use Data Analysis, Descriptive Statistics & Z-test: 2-Sample for Means
Enter the information required
If P(Z<=z) two-tail < 0.05, there is a significant difference
Remember use .05 UNLESS specifically told differently
What is a t-test
The t-test assesses whether the means of two groups are statistically different from each other.
This analysis is appropriate whenever you want to compare the means of two groups
Why you use a T-Test
Often, you haven't the time or money to measure every single item in a “collection of stuff”. Sometimes, it's just not practical, either. Let's say you want to see how much force it takes to break new laptop computer. If you break them all, you won't have any left to sell. Not a good idea. Or a particularly smart business plan.
Why use T-Test
That's why you measure a smaller sample. But the standard deviation of a small sample of data doesn't necessarily tell you anything useful about how wildly the larger group's values vary around their average. And that distribution's important.
T-Test
Because sometimes the average of a small sample comes in where you want it to, but the sample's values are so widely spread around that you can't be sure the larger group's average will come in about the same place as the sample's
T-Tests
The t-score factors in:
the average of the values in your sample
the supposed average of the larger population your sample is drawn from
the standard deviation of your sample's values
the number of values in your sample.
14
t-test
In the figure below – the difference in means is IDENTICAL – but is it significant?
t-test
This leads us to a very important conclusion: when we are looking at the differences between scores for two groups, we have to judge the difference between their means relative to the spread or variability of their scores. The t-test does just this.
Hypothesis tests between two data sets: small n
For n<30, assume equal variances
Use a t-test
This is the formula for a t-test
If you did by hand – you take the t value and look it up on a t-table – we will use Excel to do it for us
In Excel, use Data Analysis, T-test: 2-Sample Assuming Equal Variances
If P(T<=t) two-tail < 0.05, there is a significant difference
Hypothesis tests on paired differences between two data sets
Calculate differences between each pair of observations
Calculate the sample mean, M, and standard deviation, s
In Excel, use Data Analysis, T-test: Paired 2-Sample for Means
If P(T<=t) two-tail < 0.05, there is a significant difference
Interpretation of results
Relate the statistical results to the original research problem
Compare results to existing literature
Is there practical significance to the results?
What are limitations of the study?
Paired Sample Choice for t-test
For t-tests you can have a 2 sample or a paired sample which is
E.g., a pre-score and a post-score from the same person for the same thing
A response from the same person about two different things
Quantitative Methods/Using Excel for Statistics.pptx
Using Excel for Statistics
1
You must be sure that you have the Excel ToolPak loaded into Excel – it is what we will use for our statistics section.
It is not automatically loaded when first installing Excel.
For this portion of the course you must use a PC – not a Mac. The Mac analysis will not do all the statistics required for the class
Excel – Analysis ToolPak
2
In Excel 2010 you
Click on File
Click on options
Click on Add-In
Highlight Analysis ToolPak
Click on Go
New screen will pop-up – be sure the Data analysis is highlighted and click OK
It will add Data Analysis to the Data Ribbon
Excel – Analysis ToolPak
3
In 2007 you click on the “Office” button
Select Excel Options
Select Add-In
Highlight Analysis ToolPak
Click on Go
New screen will pop-up – be sure the Data analysis is highlighted and click OK
It will add Data Analysis to the Data Ribbon
Excel – Analysis ToolPak
4
In Excel 2003
Click onTools
Select Add-ins
Select Analysis ToolPak
It should load it for you and you will see it added to the Tool menu
(I’m going by what others told me – my memory on 2003 is fading)
Excel – Analysis ToolPak
5