six sigma case study?
Lecture 2
Review of
basic statistics
IMPROVEMENT ROADMAP
Uses of Probability Distributions
- Establish baseline data characteristics.
Project Uses
- Identify and isolate sources of variation.
- Use the concept of shift & drift to establish project expectations.
- Demonstrate before and after results are not random chance.
Breakthrough
Strategy
Phase 4:
Control
Characterization
Phase 1:
Measurement
Phase 2:
Analysis
Optimization
Phase 3:
Improvement
Measurements are critical...
- If we can’t accurately measure something, we really don’t know much about it.
- If we don’t know much about it, we can’t control it.
- If we can’t control it, we are at the mercy of chance.
Types of Data
- Data where the metric is composed of a classification in one of two (or more) categories is called Attribute data. This data is usually presented as a “count” or “percent”.
- Good/Bad
- Yes/No
- Hit/Miss etc.
- Data where the metric consists of a number which indicates a precise value is called Variable data.
- Time
- Miles/Hr
- Use Minitab to have students record the results and have the students display using Graph..Histogram
- Note how “rough” the graph looks
- Redo using Basic Statistics …. Descriptive Statistics and display using the Graphical Summary. Walk through the normal curve transform:
- Mean (Arithmetic Average)
- Standard Deviation
- Skew (How off center the data is skewed -=left)
- Kurtosis (How flat or peaked the data is -=flat)
- Show the Box Plot:
- Quartile (25% of the Data Points)
- Median (50% of the Data Points on Each Side)
- Show the 95% Confidence Interval and Explain how it relates to the data.
Probability and Statistics
- Probability and Statistics influence our lives daily
- Statistics is the universal language for science
- Statistics is the art of collecting, classifying,
presenting, interpreting and analyzing numerical
data, as well as making conclusions about the
system from which the data was obtained.
Population Vs. Sample (Certainty Vs. Uncertainty)
- A sample is just a subset of all possible values
population
sample
- Since the sample does not contain all the possible values, there is some uncertainty about the population.
- Hence any statistics, such as mean and standard deviation, are just estimates of the true population parameters.
Descriptive Statistics
Descriptive Statistics is the branch of statistics which
most people are familiar. It characterizes and summarizes
the most prominent features of a given set of data (means,
medians, standard deviations, percentiles, graphs, tables
and charts.
Descriptive Statistics describe the elements of
a population as a whole or to describe data that represent
just a sample of elements from the entire population
Inferential Statistics
Inferential Statistics is the branch of statistics that deals with
drawing conclusions about a population based on information
obtained from a sample drawn from that population.
While descriptive statistics has been taught for centuries,
inferential statistics is a relatively new phenomenon having
its roots in the 20th century.
We “infer” something about a population when only information
from a sample is known.
Probability is the link between
Descriptive and Inferential Statistics
USES OF PROBABILITY DISTRIBUTIONS
Primarily these distributions are used to test for significant differences in data sets.
To be classified as significant, the actual measured value must exceed a critical value. The critical value is tabular value determined by the probability distribution and the risk of error. This risk of error is called a risk and indicates the probability of this value occurring naturally. So, an a risk of .05 (5%) means that this critical value will be exceeded by a random occurrence less than 5% of the time.
Critical Value
Critical Value
Common Occurrence
Rare Occurrence
Rare Occurrence
WHAT IS THE MEAN?
The mean is the most common measure of central tendency for a population. The mean is simply the average value of the data.
n=12
Mean
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
x
i
=
-
å
2
mean
x
x
n
i
=
=
=
-
=
-
å
2
12
17
.
WHAT IS THE MEDIAN?
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
If we rank order (descending or ascending) the data set for this distribution we could represent central tendency by the order of the data points.
If we find the value half way (50%) through the data points, we have another way of representing central tendency. This is called the median value.
Median Value
Median
50% of data points
WHAT IS THE MODE?
If we rank order (descending or ascending) the data set for this distribution we find several ways we can represent central tendency.
We find that a single value occurs more often than any other. Since we know that there is a higher chance of this occurrence in the middle of the distribution, we can use this feature as an indicator of central tendency. This is called the mode.
Mode
Mode
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
MEASURES OF CENTRAL TENDENCY, SUMMARY
X
X
n
i
=
=
-
=
å
2
12
17
.
X
MEAN ( )
(Otherwise known as the average)
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
n/2=6
n/2=6
}
Mode = 0
Mode = 0
MEDIAN
(50 percentile data point)
Here the median value falls between two zero values and therefore is zero. If the values were say 2 and 3 instead, the median would be 2.5.
MODE
(Most common value in the data set)
The mode in this case is 0 with 5 occurrences within this data.
Median
n=12
SO WHAT’S THE REAL DIFFERENCE?
MEAN
The mean is the most consistently accurate measure of central tendency, but is more difficult to calculate than the other measures.
MEDIAN AND MODE
The median and mode are both very easy to determine. That’s the good news….The bad news is that both are more susceptible to bias than the mean.
SO WHAT’S THE BOTTOM LINE?
MEAN
Use on all occasions unless a circumstance prohibits its use.
MEDIAN AND MODE
Only use if you cannot use mean.
Example of tossing a coin 200 time
(probability of getting heads)
1
3
0
1
2
0
1
1
0
1
0
0
9
0
8
0
7
0
6
0
0
5
0
0
4
0
0
3
0
0
2
0
0
1
0
0
0
Number of occurrences
What are some of the ways that we can easily indicate the dispersion (spread) characteristic of the population?
Three measures have historically been used; the range, the standard deviation and the variance.
WHAT IS THE RANGE?
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
Range
Range
x
x
MAX
MIN
=
-
=
-
-
=
4
5
9
(
)
The range is a very common metric which is easily determined from any ordered sample. To calculate the range simply subtract the minimum value in the sample from the maximum value.
Range
Max
Min
WHAT IS THE VARIANCE & STANDARD DEVIATION?
The variance (s2) is a very robust metric which requires a fair amount of work to determine. The standard deviation(s) is the square root of the variance and is the most commonly used measure of dispersion for larger sample sizes.
DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
X
X
i
-
-5-(-.17)=-4.83
-3-(-.17)=-2.83
-1-(-.17)=-.83
-1-(-.17)=-.83
0-(-.17)=.17
0-(-.17)=.17
0-(-.17)=.17
0-(-.17)=.17
0-(-.17)=.17
1-(-.17)=1.17
3-(-.17)=3.17
4-(-.17)=4.17
(-4.83)2=23.32
(-2.83)2=8.01
(-.83)2=.69
(-.83)2=.69
(.17)2=.03
(.17)2=.03
(.17)2=.03
(.17)2=.03
(.17)2=.03
(1.17)2=1.37
(3.17)2=10.05
(4.17)2=17.39
61.67
(
)
s
X
X
n
i
2
2
1
61
67
12
1
5
6
=
-
-
=
-
=
å
.
.
X
X
n
i
=
=
-
=
å
2
12
-.17
MEASURES OF DISPERSION
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
ORDERED DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
Min=-5
R
X
X
=
-
=
-
-
=
max
min
(
)
4
6
10
Max=4
DATA SET
-5
-3
-1
-1
0
0
0
0
0
1
3
-6
-5
-4
-3
-2
-1
0
1
2
3
4
5
6
4
X
X
n
i
=
=
-
=
å
2
12
-.17
(
)
s
X
X
n
i
2
2
1
61
67
12
1
5
6
=
-
-
=
-
=
å
.
.
X
X
i
-
-5-(-.17)=-4.83
-3-(-.17)=-2.83
-1-(-.17)=-.83
-1-(-.17)=-.83
0-(-.17)=.17
0-(-.17)=.17
0-(-.17)=.17
0-(-.17)=.17
0-(-.17)=.17
1-(-.17)=1.17
3-(-.17)=3.17
4-(-.17)=4.17
(-4.83)2=23.32
(-2.83)2=8.01
(-.83)2=.69
(-.83)2=.69
(.17)2=.03
(.17)2=.03
(.17)2=.03
(.17)2=.03
(.17)2=.03
(1.17)2=1.37
(3.17)2=10.05
(4.17)2=17.39
61.67
s
s
=
=
=
2
5
6
2
37
.
.
RANGE (R)
(The maximum data value minus the minimum)
VARIANCE (s2)
(Squared deviations around the center point)
STANDARD DEVIATION (s)
(Absolute deviation around the center point)
SO WHAT’S THE REAL DIFFERENCE?
VARIANCE/ STANDARD DEVIATION
The standard deviation is the most consistently accurate measure of central tendency for a single population. The variance has the added benefit of being additive over multiple populations. Both are difficult and time consuming to calculate.
RANGE
The range is very easy to determine. That’s the good news….The bad news is that it is very susceptible to bias.
SO WHAT’S THE BOTTOM LINE?
VARIANCE/ STANDARD DEVIATION
Best used when you have enough samples (>10).
RANGE
Good for small samples (10 or less).
SO WHAT IS THIS SHIFT & DRIFT STUFF...
The project is progressing well and you wrap it up. 6 months later you are surprised to find that the population has taken a shift.
-12
USL
LSL
-10
-8
-6
-4
-2
0
2
4
6
8
10
12
SO WHAT HAPPENED?
Time
All of our work was focused in a narrow time frame. Over time, other long term influences come and go which move the population and change some of its characteristics. This is called shift and drift.
Historically, this shift and drift primarily impacts the position of the mean and shifts it 1.5 s from it’s original position.
Original Study
VARIATION FAMILIES
Variation is present upon repeat measurements within the same sample.
Variation is present upon measurements of different samples collected within a short time frame.
Variation is present upon measurements collected with a significant amount of time between samples.
Sources of Variation
Within Individual Sample
Piece to Piece
Time to Time
SO WHAT DOES IT MEAN?
To compensate for these long term variations, we must consider two sets of metrics. Short term metrics are those which typically are associated with our work. Long term metrics take the short term metric data and degrade it by an average of 1.5s.
IMPACT OF 1.5s SHIFT AND DRIFT
Z
PPM
ST
C
pk
PPM
LT
(+1.5
s
)
0.0
500,000
0.0
933,193
0.1
460,172
0.0
919,243
0.2
420,740
0.1
903,199
0.3
382,089
0.1
884,930
0.4
344,578
0.1
864,334
0.5
308,538
0.2
841,345
0.6
274,253
0.2
815,940
0.7
241,964
0.2
788,145
0.8
211,855
0.3
758,036
0.9
184,060
0.3
725,747
1.0
158,655
0.3
691,462
1.1
135,666
0.4
655,422
1.2
115,070
0.4
617,911
1.3
96,801
0.4
579,260
1.4
80,757
0.5
539,828
1.5
66,807
0.5
500,000
1.6
54,799
0.5
460,172
1.7
44,565
0.6
420,740
Here, you can see that the impact of this concept is potentially very significant. In the short term, we have driven the defect rate down to 54,800 ppm and can expect to see occasional long term ppm to be as bad as 460,000 ppm.
IMPROVEMENT ROADMAP
Uses of Probability Distributions
Breakthrough
Strategy
Phase 4:
Control
Characterization
Phase 1:
Measurement
Phase 2:
Analysis
Optimization
Phase 3:
Improvement
- Baselining Processes
- Verifying Improvements
Common Uses
Data points vary, but as the data accumulates, it forms a distribution which occurs naturally.
Location
Spread
Shape
Distributions can vary in:
PROBABILITY DISTRIBUTIONS,
WHERE DO THEY COME FROM?
Sheet: Sheet1
Sheet: Sheet2
Sheet: Sheet3
Sheet: Sheet4
Sheet: Sheet5
Sheet: Sheet6
Sheet: Sheet7
Sheet: Sheet8
Sheet: Sheet9
Sheet: Sheet10
Sheet: Sheet11
Sheet: Sheet12
Sheet: Sheet13
Sheet: Sheet14
Sheet: Sheet15
Sheet: Sheet16
X
COMMON PROBABILITY DISTRIBUTIONS
-4
-3
-2
-1
0
1
2
3
4
0
1
2
3
4
5
6
7
-4
-3
-2
-1
0
1
2
3
4
0
1
2
3
4
5
6
7
0
1
2
3
4
0
1
2
3
4
5
6
7
Original Population
Subgroup Average
Subgroup Variance (s2)
Continuous Distribution
Normal Distribution
c2 Distribution
Population and Sample Symbology
s
2
s
2
x
P
P
Cp
Value
Population
Sample
Mean
m
Variance
Standard Deviation
s
s
Process Capability
Cp
Binomial Mean
Z TRANSFORM
-1s
+1s
68.26%
2 tail = 32%
1 tail = 16%
+/- 1s = 68%
-2s
+2s
95.46%
2 tail = 4.6%
1 tail = 2.3%
+/- 2s = 95%
-3s
+3s
99.73%
2 tail = 0.3%
1 tail = .15%
+/- 3s = 99.7%
Common Test Values
Z(1.6) = 5.5% (1 tail a=.05)
Z(2.0) = 2.5% (2 tail a=.05)
The Focus of Six Sigma…..
Y = f(x)
All critical characteristics (Y) are driven by factors (x) which are “downstream” from the results….
Attempting to manage results (Y) only causes increased costs due to rework, test and inspection…
Understanding and controlling the causative factors (x) is the real key to high quality at low cost...
Probability distributions identify sources of causative factors (x). These can be identified and verified by testing which shows their significant effects against the backdrop of random noise.
HOW DO POPULATIONS INTERACT?
ADDING TWO POPULATIONS
mnew
snew
Population means interact in a simple intuitive manner.
m1
m2
Means Add
m1 + m2 = mnew
Population dispersions interact in an additive manner
s1
s2
Variations Add
s12 + s22 = snew2
HOW DO POPULATIONS INTERACT?
SUBTRACTING TWO POPULATIONS
mnew
snew
Population means interact in a simple intuitive manner.
m1
m2
Means Subtract
m1 - m2 = mnew
Population dispersions interact in an additive manner
s1
s2
Variations Add
s12 + s22 = snew2
TRANSACTIONAL EXAMPLE
- Orders are coming in with the following characteristics:
- Shipments are going out with the following characteristics:
- Assuming nothing changes, what percent of the time will shipments exceed orders?
X = $53,000/week
s = $8,000
X = $60,000/week
s = $5,000
TRANSACTIONAL EXAMPLE
X
X
X
shipments
orders
shipments
orders
-
=
-
=
-
=
$60
,
$53
,
$7
,
000
000
000
(
)
(
)
s
s
s
shipments
orders
shipments
orders
-
=
+
=
+
=
2
2
2
2
5000
8000
$9434
$7000
$0
Shipments > orders
X = $53,000 in orders/week
s = $8,000
X = $60,000 shipped/week
s = $5,000
Orders
Shipments
To solve this problem, we must create a new distribution to model the situation posed in the problem. Since we are looking for shipments to exceed orders, the resulting distribution is created as follows:
The new distribution looks like this with a mean of $7000 and a standard deviation of $9434. This distribution represents the occurrences of shipments exceeding orders. To answer the original question (shipments>orders) we look for $0 on this new distribution. Any occurrence to the right of this point will represent shipments > orders. So, we need to calculate the percent of the curve that exists to the right of $0.
TRANSACTIONAL EXAMPLE, CONTINUED
X
X
X
shipments
orders
shipments
orders
-
=
-
=
-
=
$60
,
$53
,
$7
,
000
000
000
(
)
(
)
s
s
s
shipments
orders
shipments
orders
-
=
+
=
+
=
2
2
2
2
5000
8000
$9434
To calculate the percent of the curve to the right of $0 we need to convert the difference between the $0 point and $7000 into sigma intervals. Since we know every $9434 interval from the mean is one sigma, we can calculate this position as follows:
$0
m
0
74
-
=
-
=
X
s
s
$0
$7000
$9434
.
Look up .74s in the normal table and you will find .77. Therefore, the answer to the original question is that 77% of the time, shipments will exceed orders.
$7000
Shipments > orders
CORRELATION ANALYSIS
Correlation Analysis is necessary to:
show a relationship between two variables. This also sets the stage for potential cause and effect.
IMPROVEMENT ROADMAP
Uses of Correlation Analysis
- Determine and quantify the relationship between factors (x) and output characteristics (Y)..
Common Uses
Breakthrough
Strategy
Phase 4:
Control
Characterization
Phase 1:
Measurement
Phase 2:
Analysis
Optimization
Phase 3:
Improvement
KEYS TO SUCCESS
Always plot the data
Remember: Correlation does not always imply cause & effect
Use correlation as a follow up to the Fishbone Diagram
Keep it simple and do not let the tool take on a life of its own
WHAT IS CORRELATION?
Input or x variable (independent)
Output or y variable (dependent)
Correlation
Y= f(x)
As the input variable changes, there is an influence or bias on the output variable.
- A measurable relationship between two variable data characteristics.
Not necessarily Cause & Effect (Y=f(x))
- Correlation requires paired data sets (ie (Y1,x1), (Y2,x2), etc)
- The input variable is called the independent variable (x or KPIV) since it is independent of any other constraints
- The output variable is called the dependent variable (Y or KPOV) since it is (theoretically) dependent on the value of x.
- The coefficient of linear correlation “r” is the measure of the strength of the relationship.
- The square of “r” is the percent of the response (Y) which is related to the input (x).
WHAT IS CORRELATION?
TYPES OF CORRELATION
Strong
Weak
None
Positive
Negative
Y=f(x)
Y=f(x)
Y=f(x)
x
x
x
CALCULATING “r”
Coefficient of Linear Correlation
- Calculate sample covariance ( )
- Calculate sx and sy for each data set
- Use the calculated values to compute rCALC.
- Add a + for positive correlation and - for a negative correlation.
(
)
(
)
s
x
x
y
y
n
xy
i
i
=
-
-
-
å
1
r
s
s
CALC
s
xy
x
y
=
s
xy
While this is the most precise method to calculate Pearson’s r, there is an easier way to come up with a fairly close approximation...
APPROXIMATING “r”
Coefficient of Linear Correlation
W
L
Y=f(x)
x
r
W
L
»
±
-
æ
è
ç
ö
ø
÷
1
r
»
-
-
æ
è
ç
ö
ø
÷
=
-
1
6
7
12
6
47
.
.
.
+ = positive slope
- = negative slope
W
L
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
6.7
12.6
- Plot the data on orthogonal axis
- Draw an Oval around the data
- Measure the length and width of the Oval
- Calculate the coefficient of linear correlation (r) based on the formulas below
HOW DO I KNOW WHEN I HAVE CORRELATION?
- The answer should strike a familiar cord at this point… We have confidence (95%) that we have correlation when |rCALC|> rCRIT.
- Since sample size is a key determinate of rCRIT we need to use a table to determine the correct rCRIT given the number of ordered pairs which comprise the complete data set.
- So, in the preceding example we had 60 ordered pairs of data and we computed a rCALC of -.47. Using the table at the left we determine that the rCRIT value for 60 is .26.
- Comparing |rCALC|> rCRIT we get .47 > .26. Therefore the calculated value exceeds the minimum critical value required for significance.
- Conclusion: We are 95% confident that the observed correlation is significant.
Ordered
Pairs
r
CRIT
5
.88
6
.81
7
.75
8
.71
9
.67
10
.63
15
.51
20
.44
25
.40
30
.36
50
.28
80
.22
100
.20
CENTRAL LIMIT THEOREM
- For this module you will need 12 dice and flip charts.
The Central Limit Theorem is:
- the key theoretical link between the normal distribution and sampling distributions.
- the means by which almost any sampling distribution, no matter how irregular, can be approximated by a normal distribution if the sample size is large enough.
IMPROVEMENT ROADMAP
Uses of the Central Limit Theorem
Common Uses
Breakthrough
Strategy
Phase 4:
Control
Characterization
Phase 1:
Measurement
Phase 2:
Analysis
Optimization
Phase 3:
Improvement
- The Central Limit Theorem underlies all statistic techniques which rely on normality as a fundamental assumption
WHAT IS THE CENTRAL LIMIT THEOREM?
Central Limit Theorem
For almost all populations, the sampling distribution of the mean can be approximated closely by a normal distribution, provided the sample size is sufficiently large.
Normal
What this means is that no matter what kind of distribution we sample, if the sample size is big enough, the distribution for the mean is approximately normal.
This is the key link that allows us to use much of the inferential statistics we have been working with so far.
This is the reason that only a few probability distributions (Z, t and c2) have such broad application.
If a random event happens a great many times, the average results are likely to be predictable.
Jacob Bernoulli
HOW DOES THIS WORK?
As you average a larger and larger number of samples, you can see how the original sampled population is transformed..
Parent Population
n=2
n=5
n=10
ANOTHER PRACTICAL ASPECT
s
n
x
x
=
s
1
n
This formula is for the standard error of the mean.
What that means in layman's terms is that this formula is the prime driver of the error term of the mean. Reducing this error term has a direct impact on improving the precision of our estimate of the mean.
The practical aspect of all this is that if you want to improve the precision of any test, increase the sample size.
So, if you want to reduce measurement error (for example) to determine a better estimate of a true value, increase the sample size. The resulting error will be reduced by a factor of . The same goes for any significance testing. Increasing the sample size will reduce the error in a similar manner.
Team experiment
- Break into 3 teams
Team one will be using 2 dice
Team two will be using 4 dice
Team three will be using 6 dice
- Each team will conduct 100 throws of their dice and record the average of each throw.
- Plot a histogram of the resulting data.
- Each team presents the results in a 10 min report out.
(
)
X
X
i
-
2
X