Hypothesis Testing and Examples of Hypothesis Testing in Research

profileclw0dq8
Week3Lecture.docx

Week 3 Lecture

This week we move from describing data and distributions – and the insights gained from these activities – to making formal inferences about populations based upon the sample results.  This involves the hypothesis testing procedure and statistical tests.

We sample in both research and everyday life to get the information we need to make decisions and take actions.  At the same time, we know that the results we get from a random, representative sample are generally “reasonably close” to, but not exactly equal to the population parameters we are interested in knowing.  So, how do we interpret our results?  If we are trying to see if things have change (for example, reducing the production line’s reject rate), how does a “reasonably close” estimate help us make our decision?  How do we account for the error in the sample statistic’s estimate?

The answer lies in the hypothesis testing process.  This approach considers both the sample mean and the variation in the population (as expressed in the sample’s variation) to give us an estimate of how likely the population result differs from a claim.  So, let’s see how this works.

Taking our example of a production line’s error rate; let’s assume the existing reject rate is 10%.  Now, we know that we will not see this rate every time we sample the outcomes – just like tossing a pair of dice – we will sometimes see a rate slightly higher and sometimes a rate slightly lower.  This is sometimes called normal, or systematic, variation; it is just part of manufacturing process as nothing is perfect.  So, if the previous rejection rate was 10%, and the sample shows us a rejection rate of 8%; does this mean we have improved the process?  Because of the systematic variation that exists, we can rarely use the sample outcome directly to answer this question; could this have just been one of the better production runs and no real change in the long-term average actually has occurred?

We can use the hypothesis testing procedure to tell us if this difference is “real” – that is, a change has actually changed – or if our outcome is merely due to chance alone (often called a sampling error).  The basic approach is to compare the difference in the means (sample minus the known average) divided by the variation in the sample outcome.  This result is compared to a probability distribution to determine the likelihood of getting a result as large or larger purely by sampling error (chance) alone.

This process can be used to compare single sample results against a standard or known value, two sample results (such as different production line error rates) against each, other, multiple samples, distributions (we graphically examined this last week, but in the upcoming weeks we will see a statistical technique to quantify our visual judgements), etc.

This week we will look at one sample situations.  We will examine multiple sample situations next week.

Hypothesis Testing Procedure

The hypothesis testing procedure is a standardized decision-making process that ensures we make our decisions on whether things are different or not based on the data, and not some other factors.  Many times, our results are more conservative than individual managerial judgements – that is, a statistical decision will call fewer things significantly different than the judgement calls of managers.  This is, at times, frustrating for managers who want to show that things have changed.  It is nicer if we are hoping that things – such as error rates – have not gotten worse.

While a lot of statistical texts have slightly different versions of the hypothesis testing procedure (fewer or more steps), they are essentially the same – and are a spinoff of the scientific method.  For this class, we will use the following six steps:

1. State the null and alternate hypothesis

2. Select a level of significance

3. Identify the statistical test to use

4. State the decision rule

5. Perform the analysis

6. Interpret the result.

Step 1

A hypothesis is a claim about an outcome.  It comes in two forms.  The first is the null hypothesis – sometimes called the testable hypothesis, as it is the claim we perform all of our statistical tests on.  It is termed the “Null” hypothesis, shown as Ho, as it basically says “no difference exists.  Even if we want to test for a difference, such as males and females having a different average compa-ratio; in statistics, we test to see if they do not.

Why?  It is easier to show that something differs from a fixed point than it is to show that the difference is meaningful – I mean how can we focus on “different?”  What does “different” mean?  So, we go with testing no difference.  The key rule about developing a null hypothesis is that it always contains an equal claim, this could be equal (=), equal to or less than (<=), or equal to or more than (=>).

Examples.  Here are some one sample examples of research questions and related null hypothesis statements:

Ex 1:   Question: Is the rejection rate mean = 10%?

Ho: Rejection rate mean = 10%.

Ex 2:   Q: is the rejection rate greater than 10%?

            Ho: Rejection rate <= (less than or equal to) 10%.

Ex. 3:  Q: Is the rejection rate less than 10%?  In this case, the null again is the opposite of what the question asks:

            Ho: Rejection rate => 10%.

Since we can only test a null that has an = sign in it, each of the questions we want answered need to be translated into a null hypothesis (testable claim) that contains an equal.  If our question asks for an equality (example 1), this is simple; the null is exactly what we are asking.  Other example of questions having an equality element are:

· Is the reject rate at least (=>) 10%? The associated null would be Ho: Rejection rate mean => 10%.

· Is the reject rate no greater than 10% (or 10% or less)? The associated null would be Ho: Rejection rate mean <= 10%.

A null hypothesis is always coupled with an alternate hypothesis.  The alternate is the opposite claim as the null.  The alternate hypothesis is shown as Ha.  Between the two claims, all possible outcomes must be covered.  So, for our three examples, the complete step 1 (state the null and alternate hypothesis statements) would look like:         

Question: Is the rejection rate mean = 10%?

Ho: Rejection rate mean = 10%.

Ha: Rejection rate mean =/= (not equal to) 10%.

Ex 2:   Q: is the rejection rate greater than 10%?

            Ho: Rejection rate <= (less than or equal to) 10%.

            Ha: Female compa-ratio mean > 10%.

Ex. 3:  Q: is the rejection rate less than 10%?

            Ho: Rejection rate => (equal to or greater than) 10%.

            Ha: Female compa-ratio mean < 10%. 

(Again, note that in the last two examples, the alternate hypothesis is the question being asked, but the null is what we always use as the test hypothesis.)

Guidelines.  When developing the null and alternate hypothesis,

1. Look at the question being asked.

2. If the wording implies an equality could exist (equal to, at least, no more than, etc.), we have a null hypothesis and we write it exactly as the question asks.

3. If the wording does not suggest an equality (less than, more than, etc.), it refers to the alternate hypothesis. Write the alternate first.

4. Then, for whichever hypothesis statement you wrote, develop the other to contain all the other possible outcomes. An = null should have a =/= alternate, an => null should have a < alternate; a <= null should have a > alternate, and vice versa.

5. The order the variables are listed in each hypothesis must be the same – if we list males first in the null, we need to list males first in the alternate. This minimizes confusion in interpreting results.

Note: the hypothesis statements are claims about the population parameters/values based on the sample results.  So, if we have; for example; a sample mean that is obviously different from the standard we are comparing it against, we could still accept the null hypothesis claim and say that the mean equals the standard – we are talking about the population mean value, and not the sample outcome.

If you look at the examples, you can notice two distinct kinds of null hypothesis statements.  One has only an equal sign in it, while the other contains an equal sign and an inequality sign (<= or =>).  These two types correspond to two different research questions and test results.  If we are only interested in whether something is equal or not, such as if the male average salary equals the female average salary, we do not really care which is greater – just if they could be the same in the population or not.  This is considered a two-tail test, as either of two conditions would cause us to reject the null.  This would be the case if our new rejection rate was > (greater than) or < (less than) the previous rate. 

The other condition we might be interested in, and we need a reason to select this approach, occurs when we want to specifically know if the mean exceeds or is less than the claim or standard.  In this situation, we care about the direction of the difference.  For example, only if the reject rate male mean is either greater than or less than the previous rejection rate mean.  If we tested to see if the rejection rate was less than 10%, and we found it was more than 10%, we would not reject the null hypothesis even though the rate had changed.  This is a critical distinction in the directional (one-tail) test.

Step 2

The level of significance is another concept that is critical in statistics, and is often not used in typical business decisions.  One senior manager told the author that their role was to ensure that the “boss’ decisions were right 50% +1 of the time rather than 50% -1.”  This suggests that the level of confidence that the right decisions are being made is around 50%.  In statistics, this would be completely unacceptable.

A typically statistical test has a level of confidence that the right decision is being made is about 95%, with a typical range from 90 to 99%.  This is done with our chosen level of significance.  For this class, we will always use the most common level of 5%, or more technically alpha = 0.05.  This means we will live with a 5% chance of saying a difference is significant when we really have only a chance sampling error.

Remember, no decision that does not involve all the possible information that can be collected will ever have no possibility of being wrong.  So, saying we are 95% sure we made the right call is great.  Marketing studies often will use an alpha of .10, meaning that are 90% sure when they say the marketing campaign worked.  Medical studies will often use an alpha of 0.01 or even 0.001, meaning they are 99% or even 99.9% sure that the difference is real and not a chance sampling error.

Step 3

Choosing the statistical test and test statistic depends upon the data we have and the question we are asking.  For this week, we will look at two tests, one that tests the variance against a standard claim and one that tests a mean against a related standard or claim. 

In the quality improvement world, one of the strategies for improving performance of a process is to first look at and reduce the variation in the data – after all, if the data has a lot of variation, we cannot really trust the mean to be very reflective of the entire data set.  After we have answered the question about variance equality, we can then focus on the question of the mean. 

Step 4

One of the rules in researching questions is that our decision rule – how we are going to make our decision once the analysis is done – should be stated upfront and before we even get to the data.  This helps ensure that our decision is data driven rather than being made by emotional factors to get the outcome we want rather than the outcome that fits the data.

The decision rule for our class is very simple, and will always be the same:

Reject the null hypothesis if the p-value is less than our alpha of .05.  (Note: this would be the same as saying that if the p-value is not less than 0.05, we would fail to reject the null hypothesis.)

We introduced the p-value in our week 1 discussion of probability, it is the probability of our outcome being as large or larger than we have by pure chance alone.  The further from the actual population mean a sample mean is, the less chance we have of getting a value that differs that that much or more purely by chance alone; the closer to the actual mean, the greater our chance would be of getting that difference or more purely by sampling error.

Our decision rule ties our criteria for significance of the outcome – the step 2 choice of alpha – with the results that the statistical tests will provide (and, the Excel tests will give us the p-values for us to use in making the decisions).

Step 5

Once we know how we will analyze and interpret the results, it is time to get our sample data, and set it up for input into an Excel statistical function.  Some examples of how this data input works will be discussed in the third lecture for this week.

What is constant about this step is the need to:

1. Look at the appropriate p-value (and it is indicated in the test outputs, as we will see below).

2. Compare the p-value with our value for alpha (0.05).

3. Make a decision – if the test p-value is less than (<) 0.05, we will reject the null hypothesis. If the test p-value is more than or equal (=>) 0.05, we will fail to reject the null hypothesis.

Rejecting the null hypothesis means that we feel the alternate hypothesis is the more accurate statement about the populations we are testing.  This is the same for all of our statistical tests.

Step 6

Once we have made our decision to reject or fail to reject the null hypothesis, we need to close the loop, and go back and answer our original question.  We need to take the statistical result or rejecting or failing to reject the null and turn it into an “English” answer to the question.  Doing so depends on how the original question lead to the hypothesis statements. 

If the question is worded the same as the null hypothesis (example 1 above), then rejecting the null means our question answer is “no.”  Not rejecting the null translates to an answer of yes to the question.  If question is worded the opposite of what the null hypothesis states (examples 2 and 3 above), rejecting the null claim means our answer to the question is yes, if we do not reject the null the answer to our question is no.

One-Sample Statistical Tests

There are three generally performed one-sample statistical tests, one variance equality, one for proportion equality, and one for means.  Additionally, we will look at a test that focuses on distributions/counts.  We will look at each of these, and then provide examples of how they fit into the hypothesis testing procedure.

Examples

Let’s set up a situation to illustrate the use of the hypothesis testing procedure with several 1-sample tests.  A few assumptions are needed.  Let us take as a given that the population defect rate is 10%, and the variance (of the reject counts) is 0.0009 (both are long term averages) while the mean number of defects equals 1.0, for a proportion of 10%.  A process improvement project was conducted with the aims of making the process more consistent (less variable from one sample to the next) and reducing the rejection rate. 

Below is a screen shot of the results of 10 samples taken from the new production process.  The statistics for this outcome are:

· Proportion of defects = 8/100 = .08 = 8%

· Variance, based on proportion, = .08*.92/100 = 0.000736

· Mean number of defects, based on 10 samples of 10 each = 8/10 = 0.8

· Variance of defect count, based on 10 samples of 10 each = 0.4

Side note: for this example, a defect rate, the proportion values are a more accurate indication of the population parameters than the results based on the 10 individual sample outcomes, as these are more likely to show a wider variance in defects per sample than the larger group.  For purposes of showing the one-sample test of means in the same context, we will use the 10 group sample outcomes as well.

Variance

The one sample variance test compares a population variance (estimated by a sample result) to a known or desired value, which could be a long-term average or quality standard.  The one-sample variance test uses the Chi Square statistic with this formula (not available in Excel):

Chi Square = (n-1) * sample variance estimate/standard variance value,

With, n = sample size and the degrees of freedom (df) = (n-1).

We can perform either a one- or two-tailed test with this statistic, depending on whether we are interested in knowing if variances are the same (or different) or if one variance is larger (or smaller) than the other. The outcome of this formula is interpreted as discussed in last week’s section on probability distributions.

Proportions

A proportion is a simple ratio of the number of success divided by the number of opportunities (like the definition of the binominal probability).  Equality (or differences) between a proportion and a standard value is tested using the Z-Test, based on normal curve probabilities.  The formula (not available in Excel) is:

Z = (proportion – standard)/sqrt(p*q/n),

where p is the success proportion, q = 1-p, and n is the sample size

We can perform either a one- or two-tailed test with this statistic, depending on whether we are interested in knowing if the proportions are the same (or different) or if one proportion is larger (or smaller) than the other. The outcome of this formula is interpreted as discussed in last week’s section on probability distributions.

Means

The one-sample test on the equality of a mean when compared against is done with the t-test and t-distribution probabilities.  The formula (not available in Excel) is:

t = (sample mean – standard)/sqrt(variance/n),

where n is the sample size

We can perform either a one- or two-tailed test with this statistic, depending on whether we are interested in knowing if the means are the same (or different) or if one mean is larger (or smaller) than the other. The outcome of this formula is interpreted as discussed in last week’s section on probability distributions.

Distributions

Testing distributions/shapes involve making decisions on whether or not the observed sample could reasonably have come from a known (or assumed) population distribution.  This is accomplished using the Chi Square Goodness of Fit test, which is available within Excel’s list of statistical procedures, in the Fx, Formulas, or Data Analysis groups.   For consistency, the formula is:

Chi Square = (Observed count – Expected Count)^2/Expected.

This test can also be performed using the Chi.Test function

This test is not a directional test, and the result shows if the observed distribution could have reasonably come from the described population distribution.  The outcome of this formula is interpreted as discussed in last week’s section on probability distributions.  The following example looks at how employees are distributed across 6 grades (from a low of grade A to the highest grade of F).  The comparison is with an approximately pyramid distribution.

 

Observed

Sum

Grade

A

B

C

D

E

F

Employees

15

7

5

5

12

6

50

Expected

Grade

A

B

C

D

E

F

Employees

16

11

9

6

5

3

50

Cell Chi square values each cell value = (observed - expected)^2/expected

0.06

1.45

1.78

0.17

9.80

3.00

16.26

 

The associated hypothesis test is shown below.