U6D1-64 - Identify, Indicate and Describe Application of Correlation from your professional life or career specialization..see details attached

CrdrAble34
Chapter10-BivariateCorrelation.docx

Chapter 10 - Bivariate Correlation

CORRELATIONS MAY be computed by making use of the SPSS command Correlate. Correlations are designated by the lowercase letter r, and range in value from −1 to +1. A correlation is often called a bivariate correlation to designate a simple correlation between two variables, as opposed to relationships among more than two variables, as frequently observed in multiple regression analyses or structural equation modeling. A correlation is also frequently called the Pearson product-moment correlation or the Pearson r. Karl S. Pearson is credited with the formula from which these correlations are computed. Although the Pearson r is predicated on the assumption that the two variables involved are approximately normally distributed, the formula often performs well even when assumptions of normality are violated or when one of the variables is discrete. Ideally, when variables are not normally distributed, the Spearman correlation (a value based on the rank order of values) is more appropriate. Both Pearson and Spearman correlations are available using the Correlate command. There are other formulas from which correlations are derived that reflect characteristics of different types of data, but a discussion of these goes beyond the scope of this book. See the IBM SPSS Statistics Guide to Data Analysis for additional information.

10.1 What is a Correlation?

Perfect positive (r = 1) correlation: A correlation of +1 designates a perfect, positive correlation. Perfect indicates that one variable is precisely predictable from the other variable. Positive means that as one variable increases in value, the other variable also increases in value (or conversely, as one variable decreases, the other variable also decreases).

Perfect correlations are essentially never found in the social sciences and exist only in mathematical formulas and direct physical or numerical relations. An example would be the relationship between the number of hours worked and the amount of pay received. As one number increases, so does the other. Given one of the values, it is possible to precisely determine the other value.

Positive (0 < r < 1) correlation: A positive (but not perfect) correlation indicates that as the value of one variable increases, the value of the other variable also tends to increase. The closer the correlation value is to 1, the stronger is that tendency, and the closer the correlation value is to 0, the weaker is that tendency.

An example of a strong positive correlation is the relation between height and weight in adult humans (r = .83). Tall people are usually heavier than short people. An example of a weak positive correlation is the relation between a measure of empathic tendency and amount of help given to a needy person (r = .12). Persons with higher empathic tendency scores give more help than persons who score lower, but the relationship is weak.

No (r = 0) correlation: A correlation of 0 indicates no relation between the two variables. For example, we would not expect IQ and height in inches to be correlated.

Negative (−1 < r < 0) correlation: A negative (but not perfect) correlation indicates a relation in which as one variable increases the other variable has a tendency to decrease. The closer the correlation value is to −1, the stronger is that tendency. The closer the correlation value is to 0, the weaker.

An example of a strong negative correlation is the relation between anxiety and emotional stability (r = −.73). Persons who score higher in anxiety tend to score lower in emotional stability. Persons who score lower in anxiety tend to score higher in emotional stability. A weak negative correlation is demonstrated in the relation between a person’s anger toward a friend suffering a problem and the quality of help given to that friend (r = −.13). If a person’s anger is less the quality of help given is more, but the relationship is weak.

Perfect negative (r = −1) correlation: Once again, perfect correlations (positive or negative) exist only in mathematical formulas and direct physical or numerical relations. An example of a perfect negative correlation is based on the formula distance = rate × time. When driving from point A to point B, if you drive twice as fast, you will take half as long.

10.2 Additional Considerations

10.2.1 Linear versus Curvilinear

It is important to understand that the Correlate command measures only linear relationships. There are many relations that are not linear. For instance, nervousness before a major exam: Too much or too little nervousness generally hurts performance while a moderate amount of nervousness typically aids performance. The relation on a scatter plot would look like an inverted U, but computing a Pearson correlation would yield no relation or a weak relation. The chapters on simple regression and multiple regression analysis ( Chapters 15 and 16 ) will consider curvilinear relationships in some detail. It is often a good idea to create a scatter plot of your data before computing correlations, to see if the relationship between two variables is linear. If it is linear, the scatter plot will more or less resemble a straight line. While a scatter plot can aid in detecting linear or curvilinear relationships, it is often true that significant correlations may exist even though they cannot be detected by visual analysis alone.

10.2.2 Significance and Effect Size

There are three questions that you, as a researcher, must answer about the r statistic. The first question is, “Are you fairly certain that the strength of relationship tested by this statistic is real instead of random?” As with most other statistical procedures, a significance or probability (or p value) is computed to determine the likelihood that a particular correlation could occur by chance. A significance less than .05 (p < .05) means that there is less than a 5% probability this relationship occurred by chance. SPSS has two different significance measures, one-tailed significance and two-tailed significance. To determine which to use, the rule of thumb generally followed is to use two-tailed significance when you compute a table of correlations in which you have little idea as to the direction of the correlations. If, however, you have prior expectations about the direction of correlations (positive or negative), then the statistic for one-tailed significance is generally used.

The second question that you need to ask about the r statistic is, “How big is it?” For all inferential statistics you have to think about the effect size to answer this question. For most statistics, there is an additional effect size measure. With correlation, you are lucky: r itself is both the test statistic and the measure of effect size. When r is close to 0, the size of the relationship is small, and when r is close to 1 or −1, the size of the relationship is large. The third question that you need to ask about the r statistic is, “Is it important?” To answer that question, you need to consider whether you are fairly certain that the effect is not due to chance, how big the effect is, and how much you care (for theoretical or practical reasons). A small r may not be important if one of your variables is length of beard but may be very important if one of your variables measures the cure for cancer.

10.2.3 Causality

Correlation does not necessarily indicate causation. Sometimes causation is clear. If height and weight are correlated, it is clear that additional height causes additional weight. Gaining weight is not known to increase one’s height. Also the relationship between gender and empathy shows that women tend to be more empathic than men. If a man becomes more empathic this is unlikely to change his gender. Once again, the direction of causality is clear: gender influences empathy, not the other way around.

There are other settings where direction of causality is likely but open to question. For instance self-efficacy (the belief of one’s ability to help) is strongly correlated with actual helping. It would generally be thought that belief of ability will influence how much one helps, but one could argue that one who helps more may increase their self-efficacy as a result of their actions. The former answer seems more likely but both may be partially valid.

Thirdly, sometimes it is difficult to have any idea of which causes which. Emotional stability and anxiety are strongly related (more emotionally stable people are less anxious). Does greater emotional stability result in less anxiety, or does greater anxiety result in lower emotional stability? The answer, of course, is yes. They both influence each other.

Finally there is the third variable issue. It is reliably shown that ice cream sales and homicides in New York City are positively correlated. Does eating ice cream cause one to become homicidal? Does committing murders give one a craving for ice cream? The answer is neither. Both ice cream sales and murders are correlated with heat. When the weather is hot more murders occur and more ice cream is sold. The same issue is at play with the reliable finding that across many cities the number of churches is positively correlated with the number of bars. No, it’s not that church-going drives one to drink, nor is it that heavy drinking gives one an urge to attend church. There is again the third variable: population. Larger cities have more bars and churches while smaller cities have fewer of both.

10.2.4 Partial Correlation

We mention this issue because partial correlation is included as an option within the context of the Correlate command. We mention it only briefly here because it is covered in some detail in Chapter 14 in the discussion about covariance. Please refer to that chapter for a more detailed description of partial correlation. Partial correlation is the process of finding the correlation between two variables after the influence of other variables has been controlled for. If, for instance, we computed a correlation between GPA and total points earned in a class, we could include year as a covariate. We would anticipate that fourth-year students would generally do better than first-year students. By computing the partial correlation, that “partials out” the influence of year, we mathematically eliminate the influence of years of schooling on the correlation between total points and GPA. With the partial correlation option, you may include more than one variable as a covariate if there is reason to do so.

The file we use to illustrate the Correlate command is our example introduced in the first chapter. The file is called grades.sav and has an N = 105. This analysis computes correlations between five variables in the file: gender, previous GPA (gpa), the first and fifth quizzes (quiz1, quiz5), and the final exam (final).

10.3 Step by Step

10.3.1 Describing Subpopulation Differences

To access the initial SPSS screen from the Windows display, perform the following sequence of steps:

Mac users: To access the initial SPSS screen, successively click the following icons:

After clicking the SPSS program icon, Screen 1 appears on the monitor.

Step 2

Create and name a data file or edit (if necessary) an already existing file (see Chapter 3 ).

Screens 1 and 2 (displayed on the inside front cover) allow you to access the data file used in conducting the analysis of interest. The following sequence accesses the grades.sav file for further analyses:

Whether first entering SPSS or returning from earlier operations the standard menu of commands across the top is required. As long as it is visible you may perform any analyses. It is not necessary for the data window to be visible.

After completion of Step 3 a screen with the desired menu bar appears. When you click a command (from the menu bar), a series of options will appear (usually) below the selected command. With each new set of options, click the desired item. The sequence to access correlations begins at any screen with the menu of commands visible:

Screen 10.1 The Bivariate Correlations Window

After clicking Bivariate, a new window opens ( Screen 10.1 , below) that specifies a number of options available with the correlation procedure. First, the box to the left lists all the numeric variables in the file (note the absence of firstname, lastname, and grade all nonnumeric). Moving variables from the list to the Variables box is similar to the procedures used in previous chapters. Click the desired variable in the list, click , and that variable is pasted into the Variables box. This process is repeated for each desired variable. Also, if there are a number of consecutive variables in the variables list, you may click the first one and then press shift key and click the last one to select them all. Then a single click of will paste all highlighted variables into the active box.

In the next box labeled Correlation Coefficients, the Pearson r is selected by default. If your data are not normally distributed, then select Spearman. You may select both options and see how the values compare.

Under Test of Significance, Two-tailed is selected by default. Click on One-tailed if you have clear knowledge of the direction (positive or negative) of your correlations.

Flag significant correlations is selected by default and places an asterisk (*) or double asterisk (**) next to correlations that attain a particular level of significance (usually .05 and .01). Whether or not significant values are flagged, the correlation, the significance accurate to three decimals, and the number of subjects involved in each correlation will be included.

For analyses demonstrated in this chapter we will stick with the Pearson correlation, the Two-tailed test of significance, and also keep Flag significant correlations. If you wish other options, simply click the desired procedure to select or deselect before clicking the final OK. For sequences that follow, the starting point is always Screen 10.1 . If necessary, perform whichever of Steps 1–4 (pages 141–143) are required to arrive at that screen.

To produce a correlation matrix of gender, gpa, quiz1, quiz5, and final, perform the following sequence of steps:

(

Additional procedures are available if you click the Options button in the upper right corner of Screen 10.1 . This window ( Screen 10.2 , below) allows you to select additional statistics to be printed and to deal with missing values in two different ways. Means and standard deviations may be included by clicking the appropriate option, as may Cross-product deviations and covariances.

Screen 10.2 The Bivariate Correlations: Options Window

To Exclude cases pairwise means that for a particular correlation in the matrix, if a subject has one or two missing values for that comparison, then that subject’s influence will not be included in that particular correlation. Thus correlations within a matrix may have different numbers of subjects determining each correlation. To Exclude cases listwise means that if a subject has any missing values, all data from that subject will be eliminated from any analyses. Missing values is a thorny problem in data analysis and should be dealt with before you get to the analysis stage. See Chapter 4 for a more complete discussion of this issue.

The following procedure, in addition to producing a correlation matrix similar to that created in sequence Step 5, will print means and standard deviations in one table. Cross-product deviations and covariances will be included in the correlation matrix:

What we have illustrated thus far is the creation of a correlation matrix in which there are equal number of rows and columns. Often a researcher wishes to create correlations between one set of variables and another set of variables. For instance she may have created a 12 × 12 correlation matrix but wishes to compute correlations between 2 new variables and the original 12. The windows format does not allow this option and it is necessary to create a “command file,” something familiar to users of the PC or mainframe versions of SPSS. If you attempt anything more complex than the sequence shown below, you will probably need to consult a book that explains SPSS syntax: See the Reference section, page 374.

Screen 10.3 The SPSS Syntax Editor Window

[with the command file from Step 5b (in the following page) included]

In previous editions of this book we had you type in your own command lines, but due to some changes in SPSS syntax structure, we find it simpler (and safer) to begin with the windows procedure and then switch to syntax. Recall that the goal is to create a 2 × 5 matrix of correlations between gender and total (the rows) and with year, gpa, quiz1, quiz5, and final (in the columns). Begin by creating what looks like a 7 × 7 matrix with variables in the order shown above. Then click the Paste button and a syntax screen ( Screen 10.3 , previous page) will appear with text very similar to what you see in the window of Screen 10.3 . The only difference is that we have typed in the word “with” between total and year, allowing a single space before and after the word “with.” All you need to do then is click the button (the arrow (↗) on Screen 10.3 identifies its location) and the desired output appears.

You may begin this sequence from any screen that shows the standard menu of commands across the top of the screen. Below we create a two by five matrix of Pearson correlations comparing gender and total with year, gpa, quiz1, quiz5, and final.

When you use this format, simply replace the variables shown here by the variables you desire. You may have as many variables as you wish both before and after the “with.”

Upon completion of Step 5, 5a, or 5b, the output screen will appear (Screen 1, inside back cover). All results from the just-completed analysis are included in the Output Navigator. Even when viewing output, the standard menu of commands is still listed across the top of the window. Further analyses may be conducted without returning to the data screen.

10.4 Printing Results

Results of the analysis (or analyses) that have just been conducted require a window that displays the standard commands (File Edit Data Transform Analyze …) across the top. A typical print procedure is shown below beginning with the standard output screen (Screen 1, inside back cover).

To print results, from the Output screen perform the following sequence of steps:

To exit you may begin from any screen that shows the File command at the top.

Note: After clicking Exit, there will frequently be small windows that appear asking if you wish to save or change anything. Simply click each appropriate response.

10.5 Output

10.5.1 Correlations

This output is from sequence Step 5 (page 144) with two-tailed significance selected and significant correlations flagged.

Notice first of all the structure of the output. The upper portion of each cell identifies the correlations between variables accurate to three decimals. The middle portion indicates the significance of each corresponding correlation. The lower portion records the number of subjects involved in each correlation. Only if there are missing values is it possible that the number of subjects involved in one correlation may differ from the number of subjects involved in others. The notes below the table identify the meaning of the asterisks and indicate whether the significance levels are one-tailed or two-tailed.

Correlations

The diagonal of 1.000s (some versions delete the “.000”) shows that a variable is perfectly correlated with itself. Since the computation of correlations is identical regardless of which variable comes first, the half of the table above the diagonal of 1.000s has identical values to the half of the table below the diagonal. Note the moderately strong positive relationship between final and quiz5 means we can be fairly certain is not due to chance (r = .472, p < .001). As described in the introduction of this chapter, these values indicate a positive relationship between the score on the fifth quiz and the score on the final. Those who scored higher on the fifth quiz tended to score higher on the final as well.

Exercises

Answers to selected exercises are downloadable at www.spss-step-by-step.net.

1. Using the grades.sav file create a correlation matrix of the following variables: id, ethnic, gender, year, section, gpa, quiz1, quiz2, quiz3, quiz4, quiz5, final, total; select one-tailed significance; flag significant correlations. Print out results on a single page.

· Draw a single line through the columns and rows where the correlations are meaningless.

· Draw a double line through cells where correlations exhibit linear dependency.

· Circle the one “largest” (greatest absolute value) NEGATIVE correlation (the p value will be less than .05) and explain what it means.

· Box the three largest POSITIVE correlations (each p value will be less than .05) and explain what they mean.

· Create a scatterplot of gpa by total and include the regression line (see Chapter 5, page 98 for instructions).

2. Using the divorce.sav file create a correlation matrix of the following variables: sex, age, sep, mar, status, ethnic, school, income, avoicop, iq, close, locus, asq, socsupp, spiritua, trauma, lsatisy; select one-tailed significance; flag significant correlations. Print results on a single page. Note: Use Data Files descriptions (page 364) for meaning of variables.

· Draw a single line through the columns and rows where the correlations are meaningless.

· Draw a double line through the correlations where there is linear dependency

· Circle the three “largest” (greatest absolute value) NEGATIVE correlations (each p value will be less than .05) and explain what they mean.

· Box the three largest POSITIVE correlations (each p value will be less than .05) and explain what they mean.

· Create a scatterplot of close by lsatisy and include the regression line (see Chapter 5, page 98 for instructions).

· Create a scatterplot of avoicop by trauma and include the regression line.

George, Darren. IBM SPSS Statistics 23 Step by Step, 14th Edition. Routledge, 20160322. VitalBook file.