Correlation and Regression & Examples of Correlation and Regression

profileclw0dq8
Week5Lecture.docx

Week 5 Lecture

Our investigation changes focus a bit this week.  We started the class by finding ways to describe and summarize data sets – finding measures of the center and dispersion of the data with means, medians, standard deviations, ranges, etc.  We moved on to looking at distributions, the patterns the data values demonstrated within the samples.  As interesting as these were, in Weeks 3 and 4 we learned how to ask questions about differences and how important different sample outcomes were.  We found that all differences were not important, and that for many relatively small sample result differences we could safely ignore them for decision making purposes – they were due to simple sampling errors.  We found that this idea of sampling error could extend into work and individual performance outcomes observed over time; and that over-reacting to such differences did not make much sense.

This week’s focus changes from detecting and evaluating differences to looking at relationships.  As students often comment, finding significant differences in measures does not explain why these differences exist.  This week’s tools, correlation and regression, move us towards understanding why outcomes exist and what impacts them.  Key underlying questions for this week are: If some measure changes does another measure change as well?  And, if so, can we use this information to make predictions and/or understand what underlies this common movement? 

Our tools in doing this involve correlation, the measurement of how closely two variables move together; and regression, an equation showing the impact of inputs on a final output.  A regression is like a recipe for a cake or other food dish; take a bit of this and some of that, put them together, and we get our result.

Linear Correlation

When two things seem to move in a somewhat predictable way, we say they are correlated.  This correlation could be direct or positive, both move in the same direction, or it could be inverse or negative, where when one increases the other decreases.  Some positive correlations we understand include:

· All other things being equal, the more we eat the more we weigh.

· Kids, up to the early teens, grow taller the older they get.

· If we consistently speed, we will get more speeding tickets than those who obey the speed limit.

· The more efforts we put into studying, the better grades we get.

· The Law of Demand and price in economics: the more demand existing for something, the more we can charge for it.

On the other hand, a very common example of a in direct correlation is, the Law of Supply and price in economics; the more supply exists for something, typically the less we can charge for it. These are all examples of correlations

Correlations exist in many forms, and, for the most part we will focus on linear correlations – those that can be graphed in a straight line.  A perfect linear correlation means that if we graphed the values, they would fall exactly on a straight line, either increase from bottom left to top right (direct or positive) or from top left to bottom right (indirect or negative).  As the absolute value of the correlation shrinks towards 0, the plotted data points move away from a straight line. Curvilinear correlations also exist, but will not be covered in this course. 

A somewhat specialized correlation was the Chi Square contingency table test (for multi-row, multi-column tables) we looked at last week, if we find the distributions differ, then we say that the variables are related/correlated.  This correlation would run from 0 (no correlation) thru positive values (the larger the value the stronger the relationship).

Probably the most commonly used correlation is the Pearson Correlation Coefficient, symbolized by r.  It measures the strength of the association – the extent to which measures change together – between interval or ratio level measures.  Excel’s Fx Correl, and the Data Analysis Correlation both produce Pearson Correlations.

Another type of correlation is the Spearman’s rank order correlation, symbolized by rs.  This correlation, called rho, is interpreted the same way as the Pearson’s Correlation.  It is used on ordinal or any ranked data variables.  Calculation of Spearman’s is not done automatically in Excel, and is a bit tedious.  One of several calculation formulas for Spearman’s rank order correlation is

rs  = 1 – 6*(sum of rank differences)^2/(n*(n-1)).

To calculate a Spearman’s correlation:

· Rank each variable data list from low to high separately, ignore any duplicate values for now, give each a unique rank

· If several values are the same in a list, average their ranks (see examples in the screen shot below)

· Find the difference between the related ranks

· Square this value

· Sum the squares of the rank differences

· Multiply this sum by 6

· Divide this result by n*(n-1), where n is the number of pairs

· Subtract this result from 1.

The screen shot below shows an example of Spearman’s rank order correlation for three grades and their related salaries for the sample looking at equal pay for equal work.

Both correlations, Pearson’s and Spearman’s, are interpreted the same way.  They range from -1.0 (perfect indirect or negative) thru 0 to +1.0 (perfect direct or positive) correlation.  The stronger the absolute value (ignoring the sign), the stronger the correlation and the more the data points would form a straight line when plotted on a graph.

While the Fx Correl and the Data Analysis Correlation both produce a Pearson correlation coefficient, they output for each is different.  Correl shows simply a single correlation value for the two variables, while Correlation produces a table that includes the variable names and can include multiple variables at the same time.  The following screen shot shows these different results.  The data used is a more complete sample of information that might be used to look at the question of equal pay for equal work between males and females.  The exact values are not critical for this example.

Strength

The strength of a correlation is shown by the value (regardless of the sign).  For example, a correlation of +.78 is just as strong as a correlation of -.78; the only difference is the direction of the change.  If we graphed a +.78 correlation the data points would run from the lower left to the upper right and somewhat cluster around a line we could draw thru the middle of the data points.  A graph of a -.78 correlation would have the data points starting in the upper left and run down to the lower right.  They would also cluster around a line.

Correlations below an absolute value of around .70 are generally not considered to be very strong.  The reason for this is due to the coefficient of determination.  This equals the square of the correlation and shows the amount of shared variation between the two variables.  The more the shared variation, the more one variable can be used to predict the other.  If we square .70 we get .49, or about 50% of the variation being shared.  Anything less is too weak of a relationship to be of much help.

Students often feel that a correlation shows a “cause-and-effect” relationship; that is, changes in one thing “cause” changes in the other variable.  In some cases, this is true – height and weight for pre-teens, weight and food consumption, etc. are all examples of cause-and- effect relationships.  And, in research, we cannot say that one thing causes or explains another without having a strong correlation present.

However, just as our favorite detectives find what they think is a cause for someone to have committed the crime, only to find that the motive did not actually cause that person to commit the crime; a correlation does not prove cause-and-effect.  An example of this is the example the author heard in a statistics class of a perfect +1.00 correlation found between the barrels of rum imported into the New England region of the United States between the years of 1790 and 1820 and the number of churches built each year.  If this correlation showed a cause-and-effect relationship, what does it mean?  Does rum drinking (the assumed result of importing rum) cause churches to be built?  Does the building of churches cause the population to drink more rum?

As tempting as each of these explanations is, neither is reasonable – there is no theory or justification to assume either is true.  This is a spurious correlation – one caused by some other, often unknown, factor.  In this case, the culprit is population growth.  During these years – many years before Carrie Nation’s crusade against Demon Rum – rum was the common drink for everyone.  It was even served on the naval ships of most nations.  And, as the population grew, so did the need for more rum.  At the same time, churches in the region could only hold so many bodies (this was before mega-churches that held multiple services each Sunday); so, as the population got too large to fit into the existing churches, new ones were needed.

At times, when a correlation makes no sense we can find an underlying variable fairly easily with some thought.  At other times, it is harder to figure out, and some experimentation is needed.  The site   http://www.tylervigen.com/spurious-correlations (Links to an external site.)Links to an external site. is an interesting website devoted to spurious correlations, take a look and see if you can explain them. 😊 

Testing for Significance

The next question is, of course, which of these correlations are significant?  Testing a correlation for statistical significance uses the same 6-step hypothesis testing procedure that we have used for the F, T, and ANOVA tests.  The null states that the correlation equal 0 (or may be directional for testing a positive or negative correlation) with the alternate stating the opposite (not equal (=/=, positive (>), or negative (<) claim).  Technically, to answer this question we would need to perform the hypothesis testing procedure (all 6 steps) on each of the correlations as they are independent results.  What a pain!

While realizing that we always use the hypothesis testing procedure as our guide, we can shorten the process of determining if a correlation is significant or not.  Both Pearson’s and Spearman’s correlation are evaluated for statistical significance with the t-test for correlation:

t = r * sqrt(n-2)/sqrt(1-r^2), df = n-2 for either r or rs; n equals the number of data point  pairs used in the correlation. 

Using this formula, we can find the associated t value for any and all correlations in a correlation table or individually calculated.  We can then use fx’s T.DIST.2T(t value, df), where df = number of pairs – 2, to determine the p-value for each correlation. 

Still somewhat of a pain.  So, if we have a table of values and want to easily determine which are statistically significant, we can modify the t formula to give us the related r value for any value of t. 

The formula to use in finding the minimum correlation value that is statistically significant is r = sqrt(t^2/(t^2 + n-2))

We would find the appropriate t value by using the t.inv.2T(alpha, df) with alpha = 0.05 and df = n-2, where n equals the number of data pairs used in each correlation.  For example, with 50 subjects, each correlation would have 50 pairs of data points, and the 0.05 t value for a two tail significance test would be T.INV.2T(0.05, 48) = 2.011.

Putting 2.011 and 48 (n-2) into our formula gives us a r value of 0.278; therefore, in a correlation table based on 50 pairs, any correlation greater or equal to 0.278 would be statistically significant.

Technical Point.  If you are interested in how we obtained the formula for determining the minimum r value, the approach is shown below.  If you are not interested in the math, you can safely skip this paragraph.

t = r* sqrt(n-2)/sqrt(1-r2)

Multiplying gives us t *sqrt (1- r2) = r* sqrt(n-2)

Squaring gives us: t2 * (1- r2) = r2* (n-2)

Multiplying out gives us: t2– t2* r2 = n r2-2* r2

Adding gives us: t2= n* r2-2*r2+ t2 *r2

 Factoring gives us t2= r2 *(n -2+ t2)

Dividing gives us t2 / (n -2+ t2) = r2

Taking the square root gives us r = sqrt (t2 / (n -2+ t2)

Once we have determined the r value related to the significance level we desire, we can then identify any value greater than this to know which are the significant values in a correlation table.

Multiple Correlation

As interesting as linear correlation is, multiple correlation is even more so.  It correlates several independent (input) variables with a single dependent (output) variable.  For example, it would show the shared variation (multiple R squared) for compa-ratio with the other variables in the data set.

While we can generate this value by itself, in general we obtain it as part of a multiple regression equation.  So, let’s move on to regression – linear and multiple.

Regression

Even if the correlation is spurious, we can often use the data in making predictions until we understand what the correlation is really showing us.  This is what regression is all about.  Earlier correlations between age, height, and even weight were mentioned.  In pediatrician offices, doctors will often have charts showing typical weights and heights for children of different ages.  These are the results of regressions, equations showing relationships.  For example (and these values are made up for this example), a child’s height might be his/her initial height at birth plus 4 inches per year.  If the average height of a newborn child is about 19 inches, then the linear regression would be:

Height = 19 inches plus 4 inches * age in years, or in math symbols:

Y = a + b*x, where y stands for height, a is the intercept or initial value at age 0 (immediate birth), b is the rate of growth per year, and x is the age in years.

In both cases, we would read and interpret it the same way: the expected height of a child is 19 inches plus 4 inches times its age.  For a 12-year old, this would be 19 + 4*12 = 19 + 48 = 67 inches or 5 feet 7 inches (assuming the made-up numbers are accurate).  One of the important issues about regression is that the equations generally are only valid for a restricted range of values – obviously, children stop growing at some point, so equating height to age stops being useful at some point around the early teen age years.  This limitation is true for virtually all regressions or trends, at some point they are no longer valid and new relationships are needed.

That was an example of a linear regression having one output and a single, independent variable as an input.  A multiple regression equation is quite similar but has several independent input variables.  It is similar to a recipe for a cake:

Cake = 1*cake mix + 2* eggs + 1½ * cup milk + ½ * teaspoon vanilla + 2 tablespoons* butter.

A regression equation, either linear or multiple, shows us how “much” each factor is used in or influences the outcome.  The math format of the multiple regression equation is similar to that of the linear regression, it just includes more variables:

Y = a + b1*x1 + b2*X2 + b3*X3 + …; where a is the intercept value when all the inputs are 0, the various b’s are the coefficients that are multiplied by each variable value, and the x’s are the values of each input. 

Excel

Creating regression outputs in Excel uses the Data Analysis Regression function, for both the linear and multiple regression situations.  The only difference in data entry involves the number of variable columns entered in the Input X range data entry box.  Here is a screen shot of a complete Regression set-up for a regression equation for a variable called compa-ratio.  (Note: while we have not mentioned this variable before, it is a common HR compensation measure that is the result of an employee’s salary divided by the grade midpoint.  It is often used to show how salaries are dispersed within a salary range.)

Data range entry for the Y (or outcome) and the X (or input) variables are done separately by either typing in the ranges or using dragging the cursor over the data range after clicking on the up arrow at the right end of the data entry boxes.  The same is done with the data entry box after clicking the circle for Output range.

There are a number of options to consider.  First, of course, is the need to click the labels box if your data ranges include labels.  A second option is the Constant is Zero equation.  This would force the regression equation to pass thru the X = 0 and Y = 0 origin, even if this is not the best fit.  Use this with caution, even though it might make no sense to have Y = 0 when all the X variables are 0, using this option may not give us the equation that best fits the data.

The residuals box provides a way to see how well each of the plotted data points fits with the predicted results.  This will often allow us to see outliers – cases that do not fit with the rest of the data set.  Outliers are sometimes indications of data entry errors or, in the case of salary, they may be paid using a different approach.  One such example would be a commission salesperson being included with employees that are paid on a straight salary, the basis of pay is so different these two should not be analyzed in the same study.  Other options here allow for the results to be turned into Z-scores (Standardized Residuals), plotted on a graph, or have linear plots made for the output and each separate input.  Normal Probability Plots are rather complicated to discuss, and it is left to the student to explore this if desired.

After completing the set-up box, click on OK to produce the result.

Here is a screen shot of a multiple regression analysis for the question of what factors influence the outcome variable compa-ratio.  Note:  we will split the discussion of the output into two screen shots.

The first 4 steps of the hypothesis procedure should be familiar by now.  The null for a multiple regression equation is always Ho: The regression equation is not significant versus the alternate of Ha: The regression equation is significant.  In step 3, we say the F statistic from the ANOVA-Reg (short for ANOVA – Regression Table).

The first table in the output provides some summary statistics.  Two are important for us – the multiple correlation, shown as R, which equals 0.655, a moderate value; and, the R square or the multiple coefficient of determination showing that about 43% of the variation in compa-ratio values can be explained by the shared variation in the variables used in the analysis.

The second table shows the results of the actual statistical test of the regression.  Similar to the ANOVA tables we looked at last week, it has two rows that are used to generate our F statistic (4.51) and the p-value (Significance F) of 0.0008.

Since the p-value is less than our alpha, we reject the null in step 5, and in step 6 say that some of the compa-ratio outcomes can be explained by the selected variables.  We used the phrase “some of” since the equation only explains 43% of the variance, less than half.

Once we reject the null hypothesis, our attention changes to the actual equation, the variables and their corresponding coefficients.  The third table provides all the details we need to reach our conclusions.

As with the correlations in question 1, we will use the hypothesis testing process, but will write it only once and use the p-values to make decisions on each of the possible equation variables.  The Multiple Regression equation is similar to the linear regression example given above except it has more independent terms: Y = a + b1*X1 + b2*X2 + B3*X3 + ….  The b’s stand for the coefficients that are multiplied by the value of each variable (represented by the X’s).

In first column (L in the screen shot) are the possible regression elements starting with the intercept, which is always a part of the equation.

The next column (M) and the fifth column are the really important columns.  Column P, labeled p-value, tells us which variables are statistically significant.  Just as with our previous tests, if the p-value is less than (<) our chosen alpha, we reject the related null hypothesis and accept the alternate that the coefficient’s value is different than 0, and the related variable should be included in the final equation.

For our example, we find that only 3 variables are statistically significant; the midpoint, the performance rating (missing its yellow highlighting), and the gender.  The results for all variables are shown in both the ANOVA table and reproduced in the step 5 table in rows 58 – 61.  With these 3 variables and the intercept, the statistically significant regression equation is:

Compa-ratio = 0.9545 + 0.0034*midpoint -0.0024*performance rating + 0.0562*gender.

So, what does this equation mean?  How do we interpret it?  The intercept (0.9545) is somewhat of a place holder – it centers the line in the middle of the data points, but has little other meaning for us.  The three variables, however tell us a lot.  Changes in each of them impact the compa-ratio outcome independently of the others – it is as if we can consider the other factors being held constant as we examine each factor’s impact.  So, all other things the same, each dollar increase in midpoint increases the compa-ratio value by 0.0034.  This relates to what we found last week that compa-ratio is not independent of grade.  At the same time, and possibly surprisingly, every increase in an employee’s performance rating causes the compa-rating to decrease by .0024!  Finally, the equation says that gender is an important factor.  This factor alone means that the company is violating the equal pay act.  But, what might be surprising is that for a change from male (coded 0) to female (coded 1) the compa-ratio goes up by 0.0562!  Females get a higher compa-ratio (percent of midpoint) when all other things are equal than males do, since the female gender results in adding 0.0562*1 to the compa-ratio while the male gender has 0.0562 * 0 (or 0) added to their compa-ratio.