Analysis
ANA:3:11/13 - 1 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Quantitative Analysis: Tools
Learning Objectives
1. Understand how to calculate three measures of central tendency and interpret the effect of extreme values on each measure. (p. 6)
2. Understand how to calculate three measures of dispersion. (p. 11)
3. Describe the role of the standard deviation in a normal distribution and explain the differences between a normal and skewed distribution. (p. 14)
4. Draw a simple histogram showing individual losses versus annual totals. (p. 19)
5. Forecast future losses by describing and calculating confidence intervals and simple linear regression. (p. 22)
ANA:3:11/13 - 2 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
I. Fundamentals of Statistics
A. Why should risk managers study statistics?
To make the best possible risk control and risk financing decisions
ANA:3:11/13 - 3 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #3
Frequency History for Unbelievable Company
Year No. of Losses Total Losses X1 250 $281,250
X2 250 $281,250
X3 250 $281,250
X4 250 $281,250
X5 250 $281,250
X6 ? ?
Severity per loss for X6 has been separately predicted to be an average of $1,125 per loss. (Number of Losses) x (Severity of Loss) = (Total Losses)
250 x $1,125 = $281,250 1. Is this prediction reasonable?
2. Is it reasonable to expect X6 will follow exactly the same as the prior five years?
3. Does this example reflect the real world?
4. What can be done to make better predictions?
ANA:3:11/13 - 4 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #4
Frequency History of Smooth-On and Jumping Jack
Smooth-On Year Jumping Jack 240 X1 120
260 X2 383
230 X3 247
270 X4 301
250 X5 199
? X6 ?
Severity per loss for X6 has been separately predicted to be an average of $1,125 per loss for both facilities.
1. What will predicted losses be in X6?
2. What is the range of possible losses that might be predicted to occur?
3. Can we assign probability or determine a degree of certainty for losses not exceeding some number?
ANA:3:11/13 - 5 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #5
Smooth-On Year Jumping Jack 240 X1 120
260 X2 383
230 X3 247
270 X4 301
250 X5 199
1,250 Total 1,250
Mean is 5
1250 = 250
Frequency predicted:
250 X6 250
Total $ loss predicted for X6:
expected frequency expected severity
250 $1,125 = $281,250
1. Is this prediction reasonable?
2. Can we take the different patterns of experience into consideration?
3. Does this example reflect what really happens in the real world?
ANA:3:11/13 - 6 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Learning Objective #1: Understand how to calculate three measures of central tendency and interpret the effect of extreme values on each measure.
II. Measures of Central Tendency
A. Measures of central tendency (based on normal distribution)
1. Mean – sum of all observations divided by the number of observations (also known as the average or arithmetic mean). This measure is highly susceptible to extreme values or outlying observations.
2. Median – midpoint of the observations ranked in order of value; half the observations lie below and half above; the middle value (also known as the 50th percentile). If an even number of observations is at the midpoint, the median is the average of the middle two. This measure reduces the impact of extreme observations.
3. Mode – observation that occurs most often in the sample; the highest frequency. There may be none, one, or more than one mode. The population mode is the observation that has the highest probability of occurring.
Note: Sample – a subset of a larger group having the same characteristics of the group
Population – the entire group of observations
ANA:3:11/13 - 7 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #6
1. Calculate the three measures of central tendency for the following seven numbers:
1, 4, 2, 1, 1, 7, 5
2. Recalculate the three measures of central tendency for the following eight numbers:
1, 4, 2, 1, 1, 7, 5, 100
A B
Mean
Median
Mode
Note: Most business and economic data, including loss data, is right-skewed since there is a zero lower bound and no effective upper bound on the data, e.g., revenues, costs, individual income, or losses.
ANA:3:11/13 - 8 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #7 Calculate the three measures of central tendency using the following information:
Total Return on the S&P 500
1989 31.23%
1988 16.34
1987 5.67
1986 18.54
1985 31.06
1984 5.97
1983 22.31
1982 20.37
1981 (4.85)
1980 31.48
Mean
Median
Mode
ANA:3:11/13 - 9 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
B. Measures of central tendency in a sample versus the population
1. The previous examples illustrate calculations for a
sample.
2. When applied to a population, the mean is referred to as an “expected value”. The expected value is calculated by weighting values with the likelihood (probability) of occurring.
C. Probability concepts
1. Probabilities vary from 0 to 1.00
2. Probabilities must total 1.00 for all possible outcomes
3. Details are provided in Appendix A
ANA:3:11/13 - 10 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #8
Expected Loss (Probability Weighted Mean) What is the effect on the expected loss to an asset related to a fire exposure with and without loss control procedures in place (automatic sprinkler system)?
1. Without loss control
Severity Probability (Prob x Loss)
$ 0 0.79 $ 0 10,000 0.12 1,200 50,000 0.06 3,000 100,000 0.03 3,000 1.00 $ 7,200 Expected loss = $7,200
2. With loss control (impacts only severity) Severity Probability (Prob x Loss)
$ 0 0.79 $ 0 5,000 0.12 600 7,500 0.06 450 50,000 0.03 1,500 1.00 $ 2,550 Expected loss = $2,550
ANA:3:11/13 - 11 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Learning Objective #2: Understand how to calculate three measures of dispersion.
III. Measures of Dispersion
A. Measures of dispersion – range, variance and standard deviation
1. Range – the difference between the largest and smallest values
Note: Range is a statistic consisting of a single value; some people erroneously refer to the range as being 100 to 900, but that is not the range. The values may range from 100 to 900 (where range is used as a verb); however, the statistical range is 800. If the largest possible property loss for an organization is $500,000, losses can vary from $0 to $500,000, but the range is $500,000
Pitfall: cannot distinguish between the two below-listed scenarios
50/50 99/1
0
25
50
75
0 100
F re q u e n c y
0 15 30 45 60 75 90
105 120
0 100
F re q u e n c y
ANA:3:11/13 - 12 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
2. Variance – the modified average of the squared deviation of each value from the arithmetic mean of those values. The “modified average” refers to dividing the sum of the squared deviations by n-1 (n being the number of data items) rather than by n; this is a statistical modification used when dealing with samples rather than populations
3. Standard deviation – the square root of the variance. The standard deviation of a sample group of numbers is calculated by using the following steps:
a. Calculate the mean of the numbers
b. Calculate the difference between each individual observation and the mean or “deviations” from the mean
c. Square each of the deviations
d. Add the squared deviations together
e. Divide the sum of the squared deviations by n-1, which is the number of data items less one resulting with the variance
f. Take the square root of the variance resulting in the standard deviation
ANA:3:11/13 - 13 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Calculation of the Mean, Variance and
Standard Deviation (Using the array of values on page 7)
Actual
Outcome Value
Expected Value (Mean)
Deviation: Actual Outcome
less Mean Squared Deviation
1 3 -2 4 4 3 1 1 2 3 -1 1 1 3 -2 4 1 3 -2 4 7 3 4 16 5 3 2 4
21 34
n = number of observations = 7 n-1 or 7-1 = 6 mean = average = 21/7 = 3 variance = 34/6 = 5.67 The square root of the variance, or the Standard deviation = 67.5 = 2.38
ANA:3:11/13 - 14 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Learning Objective #3: Describe the role of the standard deviation in a normal distribution and explain the differences between a normal and skewed distribution.
IV. Confidence Intervals Using Normal Distributions
A. Normal distributions
1. A normal distribution, often called a bell curve, has the same value at the high point of the graph
x = mean = median = mode
Frequency
Variable Values
ANA:3:11/13 - 15 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
2. The Empirical Rule states that nearly all values lie within 3 standard deviations of the mean for a normal distribution to forecast final outcomes.
a. 68.2% of values lie within 1 standard
deviation of the mean
b. 95.4% of values lie within 2 standard deviations of the mean
c. 99.7% of values lie within 3 standard deviations of the mean
The significance of a standard deviation in a normal distribution is that statistical measures of central tendency can be readily used to forecast losses.
1 S 68.2% 2 S 95.4% 99.7% -3S -2S -1S +1S +2 S +3S
M E A N
ANA:3:11/13 - 16 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
3. Central Limit Theorem
a. The Central Limit Theorem states that with an appropriately large sample, commonly values 30, that sample’s average can be treated as if it were drawn from a normal distribution
b. A normal distribution or a smooth, bell-
shaped curve makes statistical analysis simpler because it is necessary only to determine measures of central tendency and dispersion to fully describe the distribution
c. Insurance companies rely on the Law of
Large Numbers to ensure a normal distribution when setting rates and premiums
d. Most organizations face a small sample size
that may require adjustments to the data for skewness and kurtosis (distribution of data from the mean)
Note: Appendix A illustrates the power of this theorem for the total when tossing a pair of six-sided dice
ANA:3:11/13 - 17 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
e. The Central Limit Theorem can be illustrated by comparing an individual loss profile with a total loss profile using a bar chart such as the following:
In $1,000’s, the losses in a particular year for an organization are as follows:
10, 10, 10, 20, 40, 50, 70 = 210
Note: Each loss has a 1/7 likelihood of occurring (1/7 = 0.143). There are three losses of $10 out of the 7 losses. The probability of a $10 loss is 3/7 or 0.429.
0
0.15
0.3
0.45
10 20 30 40 50 60 70
Loss Amount
P ro
b a
b il
it y
ANA:3:11/13 - 18 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
B. Skewed distributions have differing values at the high point of the graph. Skewness is the measure of the degree of asymmetry of a frequency distribution.
1. In a positive or right-skewed distribution, the mean is to the right of the median; the average is greater than the median value, and the “tail” is on the right.
2. Right-skewed distributions are common in risk
management. For example, individual severity (loss size) distributions are skewed to the right, as most losses are small relative to the possible maximum loss. In this case, the values missed by using standard deviations to “capture” the losses would be the infrequent but severe losses.
3. In a negative or left-skewed distribution, the mean
is to the left of the median; the average is less than the median value, and the “tail” is on the left.
4. The Empirical Rule does not hold for right-
skewed or left-skewed distributions. It is used only with normal distributions.
Right-skewed distribution
Frequency Variable Values
Median
Mode Mean
ANA:3:11/13 - 19 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Learning Objective #4: Draw a simple histogram showing individual losses versus annual totals.
V. Simple Histogram
A. Purpose of a histogram – used to display frequencies for broader ranges of losses; shows what proportion of losses fall into each range
B. Steps for drawing a histogram
Data: In $1,000’s, the losses are: 40, 10, 50, 70, 10, 10, and 20 for a total of 210
1. Sort the data in size order
10, 10, 10, 20, 40, 50, 70
2. Select the number and size (range) of bins (groups of data). Generally, it is best to use bins of equal size. Trial and error may be required.
3. Sort the data into the selected bins
0 to 25 = 4 occurrences (10, 10, 10, 20) 26 to 50 = 2 occurrences (40, 50) 51 to 75 = 1 occurrence (70)
ANA:3:11/13 - 20 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
4. Calculate percentages
0 to 25 = 4/7 = 0.571 = 57.1%
26 to 50 = 2/7 = 0.285 = 28.5%
51 to 75 = 1/7 = 0.143 = 14.3%
5. Draw the graph
Bin ranges are on the X-axis
Percentages (probabilities) are on the Y-axis
0
0.1
0.2
0.3
0.4
0.5
0.6
0 to 25 26 to 50 51 to 75
Loss Ranges ($1,000s)
P ro
b a
b il
it y
ANA:3:11/13 - 21 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
VI. Total Loss Profiles
These profiles illustrate probable total ultimate losses. Ultimate losses can be defined as “fully developed losses” during a given time period, usually a year.
Such profiles can be shown for a single organization over time or for a number of different organizations at a point in time. In either case, as with a sample, the number of total losses increases, and profiles commonly approach a normal- like distribution.
The following graph illustrates annual total losses for a number of organizations in a particular year. Losses illustrated on page 17 total $210,000, and are included in the “2 to 3” range bar.
0.10
0.25
0.30
0.20
0.14
0.01
0
0.05
0.1
0.15
0.2
0.25
0.3
0.35
0.4
0 to 1 1 to 2 2 to 3 3 to 4 4 to 5 5 to 6
Total Ultimate Losses ($100,000)
P ro
b a b
il it
y
Total loss profiles, as well as frequency distributions, tend to be more “normal”.
ANA:3:11/13 - 22 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Learning Objective #5: Forecast future losses by describing and calculating confidence intervals and simple linear regression.
VII. Confidence Interval
Confidence interval – a set of numbers believed to include an unknown population parameter. Associated with the interval is a measure of the confidence that the interval does contain the parameter. In normal distributions, the following intervals contain the stated percentage of the values:
+ 1 standard deviation 68.2% of the time
+ 2 standard deviations 95.4% of the time
+ 3 standard deviations 99.7% of the time
1 S 68.2% 2 S 95.4% 99.7% -3S -2S -1S +1S +2 S +3S
M E A N
ANA:3:11/13 - 23 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #9 Calculating confidence intervals to forecast expected losses for year X6 Smooth-On Year Jumping Jack 240 X1 120 260 X2 383 230 X3 247 270 X4 301 250 X5 199
1,250 Total 1,250 1,250/5 = 250 Mean 1,250/5 = 250 250 Frequency predicted for X6 250
15.81 (16) Standard deviation (rounded) 99.75 (100)
31.62 (32) + 2 standard deviations 199.50 (200)
95% confidence interval 218 to 282 (mean + 2 standard deviations) 50 to 450 (frequency x average severity ($1,125) = total expected losses)
Confidence Intervals of Total Expected Losses
$245,250 to $317,250 $56,250 to $506,250
ANA:3:11/13 - 24 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Appendix B includes calculations for the standard deviations and an additional example to illustrate confidence intervals.
The examples use sample sizes of only five years for
illustration. Samples are best used to predict an average year. Predicting individual years from samples, especially samples this small, is risky at best. For simplicity, these illustrations assume losses are normally distributed and they ignore the complications of sample size.
Given the mean and standard deviation for both organizations, determine the 95% confidence intervals, the range of outcomes that should occur 95 times in 100. For simplicity, focus on the upper range only.
The mean is 250 losses for each of the organizations. Average
predicted losses for each are $1,125 x 250 = $281,250. For Smooth-On, the standard deviation is calculated to be 15.81.
Two times the standard deviation is 31.62, rounded to 32. Add 32 to the mean and the upper end of the 95% confidence level is 282 losses. This result times an average severity of $1,125 is $317,250.
For Jumping Jack, the standard deviation is 99.75. Two times
the standard deviation is 199.5, rounded to 200, added to the mean of 250 is 450 losses. 450 times $1,125 is $506,250.
Interpret the results, and compare the two organizations.
ANA:3:11/13 - 25 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
VIII. Regression Analysis for Obvious Trends or Patterns
A. Regression – statistical technique of modeling the relationship between variables by fitting the “best” line to a scatter of dots
1. Regression analysis will produce a better
projection than a confidence interval when a time trend or other relationship is strong
2. Simple linear regression
a. Models the relationship between one or more
variables when one is a dependent variable (Y) and the others are independent (X) based on the premise that changes in X cause changes in Y
b. A linear relationship is fit to the data such that the sum of the squared deviations is minimized
c. Obtained via software
ANA:3:11/13 - 26 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Scatter Plot and Graph
Dependent Variable y • • • • m • • • • c • • • c •
x
Independent Variable
y = mx + c where m = slope of the regression c = intercept on y axis
Notes:
More than one variable can be used (multiple regressions)
Additional details are provided in Appendix C
ANA:3:11/13 - 27 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
B. Correlation – measure of the strength of a linear relationship between two variables
1. Correlation coefficient r
-1 < r < +1
2. When r = 0, there is no correlation
3. When r is positive, the variables move together; the closer to +1, a perfect, positive correlation, the stronger the relationship
4. When r is negative, the variables move in opposite directions; the closer to –1, a perfect, negative correlation, the stronger the inverse relationship
C. Causality – relationship between one variable and another variable in which the second variable is a direct consequence of the first. However, correlation between two variables does not necessarily imply causality. Additionally, statistical correlation does not imply a meaningful direct relationship, e.g., hemlines or super bowl winners and the level of the stock market.
ANA:3:11/13 - 28 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
D. Coefficient of determination – r2 is a descriptive measure of the strength of the regression relationship, or how well the regression line fits the data. r2
measures the percentage of the variation in the y variable explained by the regression
1. r2 ranges from 0 to 1.00
2. When r2 = 1.00, 100% of the variation in y is explained by the x variable (a perfect fit)
3. When r2 = 0, the regression line explains nothing
4. The higher the r2, the better the fit and the higher the degree of confidence in the regression
5. The error term, e, represents the unexplained variance
6. Rough guidelines for r2 (which statisticians probably will not appreciate):
a. r2 > 0.90 very good predictor
b. r2 = 0.80-0.89 good predictor
c. r2 = 0.60-0.79 fair predictor
d. r2 < 0.60 poor predictor
Note: Lower-case r2 indicates only a single independent variable (simple regression) while upper-case R2 indicates more than one independent variable (multiple regression). Excel always uses R2.
ANA:3:11/13 - 29 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Skills Application Scenario #10
A new organization, Up-Up-and-Away, is added to the two organizations from the previous Scenario #9 on page 23. Up-Up- and-Away experienced the same trend of loss frequencies as did Jumping Jack. However, there is one major difference; the trend for Up-Up-and-Away indicates the frequency is worsening over time, not jumping around like Jumping Jack.
Year Smooth-On Jumping Jack Up-Up-and-Away X1 240 120 120 X2 260 383 199 X3 230 247 247 X4 270 301 301 X5 250 199 383 Total 1,250 1,250 1,250 1. Forecast X6’s frequency of loss using a 95% confidence
interval
Mean 250 250 250
1 standard deviation 16 100 100
2 standard deviations 32 200 200
Intervals (± 2SD) 218 to 282 50 to 450 50 to 450
ANA:3:11/13 - 30 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Notes:
• The numbers for Smooth-On and Jumping Jack are found in Skills Application Scenario #9.
• The calculations for Smooth-On and Jumping Jack are the same because their frequencies are the same; they are listed in a different order.
2. Forecast X6’s frequency using regression analysis, if appropriate.
a. The graphs on page 32 were created using Microsoft® Excel's Chart Wizard. An "XY Scatter" was chosen as the chart type. "Series Options" were selected to display the regression equation and r2.
b. Smooth-On’s r2 is .09 and Jumping Jack’s r2 is .015. Both are too low to indicate a trend. The confidence intervals calculated in Skills Application Scenario #9 should be used instead of regression analysis.
ANA:3:11/13 - 31 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
c. Up-Up-and-Away’s r2 is unbelievably high at .991. Therefore, it is an ideal candidate for forecasting X6 using regression analysis.
1) The regression equation is shown on the graph to be:
y = mx + c
y = 62.8x + 61.6
2) The calculations for X6's projected frequency are:
# losses = 62.8(# of yrs) + 61.6
= 62.8 (6) + 61.6
= 376.8 + 61.6
= 438.4 rounded to 438
ANA:3:11/13 - 32 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Trend Analysis of the Frequency of Losses
y = 7.6x + 227.2 R² = 0.0145
0
50
100
150
200
250
300
350
400
450
0 1 2 3 4 5 6
F re q u e n cy o f L o s s
Year
Jumping Jack
y = 62.8x + 61.6 R2 = 0.9909
0
50
100
150
200
250
300
350
400
450
0 1 2 3 4 5 6
F re q u e n cy o f L o s s
Year
Up-Up-and-Away
y = 3x + 241 R2 = 0.09
225
230
235
240
245
250
255
260
265
270
275
0 1 2 3 4 5 6
F re q u e n c y o f L o s s
Year
Smooth-On
ANA:3:11/13 - 33 -
© 2013 Certified Risk Managers International. All Rights Reserved
SM
Review of Learning Objectives
1. Understand how to calculate three measures of central tendency and interpret the effect of extreme values on each measure. (p. 6)
2. Understand how to calculate three measures of dispersion. (p. 11)
3. Describe the role of the standard deviation in a normal distribution and explain the differences between a normal and skewed distribution. (p. 14)
4. Draw a simple histogram showing individual losses versus annual totals. (p. 19)
5. Forecast future losses by describing and calculating confidence intervals and simple linear regression. (p. 22)
**CardStock**
Certified Risk Managers International a proud member of The National Alliance for Insurance Education & Research
www.TheNationalAlliance.com
Quantitative Analysis: Tools Appendix
© 2010. The National Alliance for Insurance Education & Research. All Rights Reserved. This outline or any part thereof may not be reproduced in any form or by any means or stored in any information retrieval system without the express written consent of the author. This publication includes copyrighted material of Insurance Services Office, Inc. with its permission.
ANA:3:04/10 Appendix Page 1 © 2010 Certified Risk Managers International. All Rights Reserved.
Appendix A: Probability Concepts
Probability is the chance of something occurring, ranging from 0 (an impossibility) to 1.0 (a certainty). Flip a coin, and the probability of heads is one in two, or 1/2, or 0.5, or 50%.
If we flip a coin, the probability of “heads” is .50; if we flip two coins, the probability of heads on each is .25, etc.
outcomes probability outcomes probability
HH .25 HH .25
HT .25 TH or HT .50
TH .25 TT .25
TT .25 1.00
1.00 Let us graph these outcomes and their probabilities:
0
0.25
0.5
0.75
0 1 2
Possibilities (# of heads)
P ro
ba bi
lit y
ANA:3:04/10 Appendix Page 2 © 2010 Certified Risk Managers International. All Rights Reserved.
If we do the same thing with two dice: there are 36 possible outcomes from one throw of two dice. As you can see, the possibility of a seven is 6/36 or 1/6 or 0.167.
Numbers Thrown Dice Total # of times Occurring Probability
1-1 2 1 1/36 1-2, 2-1 3 2 2/36 1-3, 2-2, 3-1 4 3 3/36 1-4, 2-3, 3-2, 4-1 5 4 4/36 1-5, 2-4, 3-3, 4-2, 5-1 6 5 5/36 1-6, 2-5, 3-4, 4-3, 5-2, 6-1 7 6 6/36 2-6, 3-5, 4-4, 5-3, 6-2 8 5 5/36 3-6, 4-5, 5-4, 6-3 9 4 4/36 4-6, 5-5, 6-4 10 3 3/36 5-6, 6-5 11 2 2/36 6-6 12 1 1/36 36 36/36 Graphing the results:
0.000 0.028 0.056 0.084 0.112 0.140 0.168 0.196
2 4 6 8 10 12
Total (sum) of two dice
P ro
ba bi
lit y
ANA:3:04/10 Appendix Page 3 © 2010 Certified Risk Managers International. All Rights Reserved.
Note that for both examples, the sum of the probabilities of all possible outcomes is 1.0 (e.g. - The chance of either heads OR tails is 0.5 + 0.5 = 1.0, i.e. 100%; all dice probabilities total 36/36.)
Joint probability - The probability of both A and B occurring, i.e. the joint probability of two independent events, is equal to the product of the two. P(A and B, both) = P(A) × P(B). So, the chance of two heads (HH) is 1/2 times 1/2 = 1/4 or 0.25. The events (each coin dropping is an event) are independent, that is, the occurrence of one has no effect on the other. Also note 1/4 + 1/4 + 1/4 + 1/4 = 1.0 (all possibilities). Notice HT and TH will look the same when they land, but each is a separate alternative. The probability of either one of these occurring in one toss is 1/4 + 1/4 = 1/2 or .50, a 50% chance. Example: A shipping line has two “independent” ships. The probability of a total loss to Ship A by exploding is .02, and of Ship B grounding is .04, during a given period of time. What is the probability of these specific losses (A explodes and B runs aground) happening to BOTH ships in that period? This is a joint probability with independent events: p(A and B) = A × B = .02 × .04 = .0008 or 8 in 10,000 chances. What is the probability A will NOT explode? That B will NOT run aground?
p(not A) = 1.0 - p(A); p(not A) = 1.0 - .02 = .98; p(not B) = 1.0 - .04 = .96
What is the probability NEITHER ship will suffer their respective losses (A won’t explode, and B won’t ground) during the same period? p(not A) × p(not B) = .98 × .96 = .9408
Note p(A) × p(not B) = .02 × .96 = .0192 and p(B) × p(not A) = .98 × .04 = .0392;
Now, .0392 + .0192 + .9408 + .0008 = 1.0000; so all possibilities are accounted for.
ANA:3:04/10 Appendix Page 4 © 2010 Certified Risk Managers International. All Rights Reserved.
Alternative probability - The probability any ONE of 2 or more events will occur (but not both or more than one). Calculations vary depending on whether or not the events are mutually exclusive (can occur during the same period) or not.
Example: A third Ship C may sink (.03), run aground (.04) or explode (.02).
If the events are mutually exclusive, as in a single coin toss (H or T can occur, not both), the events are additive, as above:
p(sink or ground, but not both) = .03 + .04 = .07
If the events are not mutually exclusive, as with most real-life situations, calculation of alternative probability must subtract out the JOINT probability to arrive at “either or both” and avoid double counting (Venn Diagrams help illustrate):
p(sink or explode or both) = p(sink) + p(explode) – p(both) = .03 + .02 - (.03 × .02) =.05 – .0006 = .0494
Sink Explode (or both) (or both) Sink Both Explode 0.03 0.02 .0006 Each of these events has a little overlap that each probability includes (i.e., the chance of both events occurring). To calculate alternative probability excluding the possibility of both, we subtract the joint probability again:
p(sink or explode but not both) = .0494 - .0006 = .0488
ANA:3:04/10 Appendix Page 5 © 2010 Certified Risk Managers International. All Rights Reserved.
Another way of looking at it is that p(sinking but not exploding) = 0.03 – 0.0006 = 0.0294 and p(exploding but not sinking) = 0.02 – 0.0006 = 0.0194, and the alternate probability of these two events, the p(sink but not explode) or p(explode but not sink) = 0.0194 + 0.0294 = 0.0488. Only Only Sink Explode 0.0294 0.0194 Conditional Probability - When one event may follow from another event, we say the second event has a conditional probability: Example: The probability of Ship C (having already grounded) then exploding is .15; i.e., p(explode given grounding) = p(e|g) = .15 Joint Probability - Dependent Events – The probability of events occurring in a given order is the probability of the first event times the probability of the conditional probability of the event:
p(ground then explode) = p(g) × p(e|g) = .04 × .15 = .006
ANA:3:04/10 Appendix Page 6 © 2010 Certified Risk Managers International. All Rights Reserved.
Appendix B: Calculating Standard Deviations & Confidence Intervals Standard Deviations for Example 6
Loss Frequency for Smooth-On, Inc. Xi – X = deviation deviation
2 (# losses)
240 – 250 = (10) 100 260 – 250 = 10 100 230 – 250 = (20) 400 270 – 250 = 20 400 250 – 250 = 0 0
1000 ÷4 Variance =250
S = 250 = 15.811…
Loss Frequency for Jumping Jack, Ltd. Xi – X = deviation deviation
2 (# losses)
120 – 250 = (130) 16,900 383 – 250 = 133 17,689 247 – 250 = ( 3) 9 301 – 250 = 51 2,601 199 – 250 = ( 51) 2,601
39,800 ÷4 Variance =9,950
S = 9,950 = 99.749…
ANA:3:04/10 Appendix Page 7 © 2010 Certified Risk Managers International. All Rights Reserved.
Example: Mean, Standard Deviation, and 95% Confidence Interval
Standard Deviation Calculation
ABC Inc.'s Casualty Losses Over Time ($millions) Year 1 $ 4.4 Year 2 4.2 Year 3 5.1 Year 4 4.6 Year 5 4.8 23.1 / 5 = $4.62 million (avg annual loss)
)1(
)( 1
2
−
− =
∑ =
n
XX S
n
i i
Xi – X = deviation deviation
2 (outcome)
4.40 – 4.62 = (0.22) 0.0484 4.20 – 4.62 = (0.42) 0.1764 5.10 – 4.62 = 0.48) 0.2304 4.60 – 4.62 = (0.02) 0.0004 4.80 – 4.62 = 0.18) 0.0324
0.4880 ÷4 Variance = 0.122
S = 0.122 = $ 0.349… million
ANA:3:04/10 Appendix Page 8 © 2010 Certified Risk Managers International. All Rights Reserved.
Confidence Interval Calculation
NOTE The example uses a sample size of only five years for illustration. Samples are best used to predict an average year. Predicting individual years from samples, especially samples this small, is risky at best. For simplicity, these illustrations assume losses are normally distributed and they ignore the complications of sample size.
However We can say statistically that the loss that this firm should
sustain in any given year will be: 68.2 % of the time between $4.27 and $4.97 million ($4.62 ± 0.35) 95.4 % of the time between $3.92 and $5.32 million ($4.62 ± 0.70) 99.7 % of the time between $3.57 and $5.67 million ($4.62 ± 1.05) ± 1 S 68.2% ± 2 S 95.4% 2 S 1 S M 1 S 2 S
ANA:3:04/10 Appendix Page 9 © 2010 Certified Risk Managers International. All Rights Reserved.
Appendix C: Regression issues
A. Assumptions 1. The value(s) of the independent variable x are not correlated with the
random error term. 2. The error terms are normally distributed with mean 0 and constant
variances. Constant variance means that y is no easier/harder to predict for high or low x values. The errors are uncorrelated with each other in successive observations.
B. Heteroscedasticity: the width of the scatter plot of the residuals increases or
decreases as the x variable increases. This violates the assumption of constant variance in (2) above and a more complex generalized least squares method must be utilized.
C. Serial Correlation: there is a pattern between the error term and x. This
violates assumption (2) above and might result from:
1. Missing variables. There should be no trend in the residuals (error terms) when plotted against time or any other possible independent variable not already in the regression equation. If a trend is found to exist, those variable(s) should be included in the model along with x resulting in a multiple regression model.
2. Nonlinear relationships. If the relationship between x and y is curved
or nonlinear, forcing a straight line to fit the data will result in a poor fit. The residuals are not random and independent and show curvature. This can be corrected by adding the variable x 2 to the model. This entails utilizing a multiple regression technique.
ANA:3:04/10 Appendix Page 10 © 2010 Certified Risk Managers International. All Rights Reserved.
y • • • • • • • • • • • Residuals • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • •• • • • • • • • • • x x D. Regression model as a predictor Predictions should be limited to the
region of the data used in the estimation process. Extrapolation outside the estimation range is risky, as the estimated relationship may not hold outside this range, e.g. oil prices over time.
E. Multiple regression analysis involves the use of several independent
variables in the regression equation. More realistic and thorough relationships can be modeled in this way.
F. Notes
1. Adding independent variables will raise the r2 (capital R with multiple
regression) if the variable adds explanatory information not already provided by the existing variables. Raising the r2 to .99 or even 1.0 would be extremely rare, however. Very high r2's do not imply that the model is correct or that it can be used to predict the dependent variable with any degree of accuracy.
2. The F test is utilized to determine if there is a regression relationship
between the dependent variable y and any of the proposed explanatory x variables.
ANA:3:04/10 Appendix Page 11 © 2010 Certified Risk Managers International. All Rights Reserved.
Total Return on the S&P 500 1989 31.23% 1988 16.34% 1987 5.67% 1986 18.54% 1985 31.06% 1984 5.97% 1983 22.31% 1982 20.37% 1981 -4.85% 1980 31.48%
for the Mean (AVERAGE)
Result
17.81%
for the MEDIAN
Result
19.46%
for the MODE
Result
#N/A
(indicating no mode)
Using Excel's "Paste Function" (fx ) for Section 3, Example 4's Mean, Median & Mode
Appendix D: Using Excel
ANA:3:04/10 Appendix Page 12 © 2010 Certified Risk Managers International. All Rights Reserved.
Total Return on the S&P 500
1989 31.23 1988 16.34 1987 5.67 1986 18.54 1985 31.06 1984 5.97 1983 22.31 1982 20.37 1981 (4.85) 1980 31.48
Column1
Mean 17.812 Standard Error 3.90593 Median 19.455 Mode #N/A Standard Deviation 12.35163 Sample Variance 152.5629 Kurtosis -0.55354 Skewness -0.566778 Range 36.33 Minimum -4.85 Maximum 31.48 Sum 178.12 Count 10
Notes: (1) Use the "Tools" pull-down menus. If you do not see "Data Analysis" among the options, then select "Add-Ins..." and click on the box to the left of "Analysis ToolPak." You will probably need the original Excel CD and it might take a few seconds while Excel loads the ToolPak.
(2) Not all of these descriptive statistics are covered in this session. The ones discussed have been highlighted in the results.
Results
Using Excel's "Tools" -- "Data Analysis" for Section 3, Example 4's Descriptive Statistics
ANA:3:04/10 Appendix Page 13 © 2010 Certified Risk Managers International. All Rights Reserved.
Year Up-Up-and-Away 1 120 The "Chart Wizard" icon looks like a 3-dimensional bar chart at the 2 199 top of the Excel screen. It provides the simplest way to plot Y against 3 247 X and display a regression trendline. We display a straight line 4 301 relationship, but you will see that there are nonlinear alternatives. 5 383
While in Step 3, also remove the check next to "Major To add the trend line, (1) right-click on any dot and select Gridlines" under "Gridlines" and "Show Legend" under "Add Trendline..." (2) Under "Options", click on "Legend." Then click on "Finish." This will give you "Display equation on chart" and also on the graph shown here. "Display R-squared value on chart". (3) Click OK.
Using Excel's "Chart Wizard" to Graph a Regression for Section 3, Example 7
Up-Up-and-Away
y = 62.8x + 61.6 R2 = 0.9909
0 50
100 150 200 250 300 350 400 450
0 1 2 3 4 5 6
Year
Fr eq
ue nc
y of
L os
s
ANA:3:04/10 Appendix Page 14 © 2010 Certified Risk Managers International. All Rights Reserved.
Year Up-Up-and-Away Use the "Tools" pull-down menus. If you do not see "Data 1 120 Analysis" among the options, then select "Add-Ins…" and click 2 199 on the box to the left of "Analysis ToolPak." You will probably 3 247 need the original Excel CD and it might take a few seconds 4 301 while Excel loads the ToolPak. 5 383
SUMMARY OUTPUT The "Regression" option of the "Data Analysis" tool prints
Regression Statistics many statistics beyond the scope of this session. The ones Multiple R 0.99545 discussed are highlighted in the results to the left and below. R Square 0.99091 Adjusted R Square 0.98789 An advantage of the "Regression" option over the "Chart Standard Error 10.97877 Wizard" is that it allows you to include multiple X-variables Observations 5 as causal influences on Y. The X variables must be in
consecutive columns. ANOVA
df SS MS F Significance F Regression 1 39438.4 39438.4 327.19912 0.00037 Residual 3 361.6 120.5 Total 4 39800.0
Coefficients Standard Error t Stat P-value Lower 95% Upper 95% Intercept 61.60000 11.51463 5.34972 0.01278 24.95528 98.24472 Year 62.80000 3.47179 18.08865 0.00037 51.75120 73.84880
Using Excel's "Tools" -- "Data Analysis" to Perform a Regression for Section 3, Example 7