rewriting task
Probability Theory, Chi-Squared Test, and Pearson Correlation
This next chapter in Sebastian M. Rasinger’s Quantitative Research in Linguistics: An Introduction is chock full of information. We are looking now at thorough data analyses that go beyond simple description; that is, we are using probability theory and two ways of measuring the relationship between two or more different variables, including the chi-squared test and the Pearson correlation. Because of the sheer amount of new information from this week’s reading, I’ve divided my reading report into five sections.
probability theory
Where do we begin? With definitions, of course! Probability theory, as its very name suggests, is concerned with how probable or likely a certain event is to occur (Rasinger 149). An event, then, is a collection of results or outcomes of an experiment we are conducting. We can calculate the probability of an event by dividing the number of time a particular outcome occurs by the number of different events we could obtain. Rasinger lists three basic properties of probabilities:
· The probability of an event always lies between 0 and 1 is inclusive;
· The probability of an event that cannot occur is 0;
· and the probability of an event that must occur is 1–based on this it is easy to calculate the probability of an event not occurring. This is called the complementation rule.
We can begin by approximating the probability of a simple event. To do so, we look t its relative frequency; that is, how often a particular outcome occurs when the event is repeated multiple times (Rasinger 151). Yet, when is an event ever simple? Not often. Especially not in linguistics. For this, we look at compound events, which are events in which more than one simple event occurs together (Rasinger 152). There are two types of compound events: (1) the union of events, which is the probability of selecting either A or B; and (2) the intersection of events, or probability of selecting both A and B. We calculate the union of events by means of the addition rule: we add the probabilities of the two simple events, and subtract the probability of the intersected event from the result in order to avoid double counting (Rasinger 153). Just like in any other statistical analysis, when using probabilities, especially in situations where we have to approximate probabilities via the relative frequencies, we need to make sure we have a sufficiently large sample, so as to avoid any outlying values skewing your results. What else is new?
the chi-squared test
The chi-squared test is for use with data based on a nominal or categorical scale. It is essentially based on the comparison of the observed values with the appropriate set of expected values. In other words, it compares whether our expectations and our actual data correspond (Rasinger 157).
First things first: we have to calculate a contingency table with expected values based on our actual values. How do we do this? Well, multiplying the row and column totals with each other, and dividing the result by the total number in the sample is how we calculate expected values. Simple enough. What now? Recall that we are interested in the differences. The closer the expected values are to zero, the less likely our variables are to be independent. What does this all mean? Something interesting, indeed. Values close to zero are a very strong indicator that there is a relationship between categorical data.
pearson correlation
But what if our data is based on a ratio scale? We can’t very well use the chi-squared test for that, now can we? What we can use, however, is the Pearson correlation. The Pearson correlation is based on the variances of the two variables to determine the strength of the relationship between two ratio-scale variables. -1 < r < 1, whereby 1=perfect positive, -1=perfect negative correlation, and 0=no correlation at all (Rasinger 167).
Before delving into the principles about the correlation of variables, let’s refresh ourselves with the scatter plot. When looking at a scatter plot, we can immediately see if two variables are related. the stronger the correlation between two variables is, the more the individual dots on the graph form a straight line. Our hope is that the variables are related in such a way that we can connect all individual points to a straight line. If this is, indeed, the case, there is a perfect correlation between two variables. Usually this isn’t the case. Rasinger outlines five basic principles about the correlation of variables:
· The Pearson correlation coefficient r can have any value between -1 and 1;
· r=1 indicates a perfect positive correlation, that is, the two variables increase in a linear fashion, that is, all the data points lie in a straight line just as if they were plotted with the help of a ruler;
· r=-1 indicates a perfect negative correlation, that is, while one variable increases, the other decreases, again, linearly;
· r=0 indicates that the two variables do not correlate at all; that is, there is no relationship between them;
· and the Pearson correlation only works for normally distributed data.
Our data might be distributed in such a way that the coefficient turns out to have a particular value simply and only because the arithmetic make it to, yet the result may be useless. as with the chi-square test, we need to determine the significance level first. Even before calculating the significance level, however, we have to decide whether our test is something called a 1-tailed test or a 2-tailed test. On the one hand, in a 1-tailed test we have a pretty good idea in advance in which direction the relationship between the variables is going. In a 2-tailed test, on the other hand, we do not predict this direction. Now we can calculate the significance level. For a given significance level, r must be equal or exceed the critical value of this level.
Rasinger pauses here to emphasize the following: while the Pearson correlation provides us with a useful tool for investigating the relationship between variables, we should be careful in accepting the results (166). He gives the following rationale:
· The correlation coefficient as a measure of association only tells us that there is a relationship between the two variables, and it also indicates the strength of the relationship.
· It does not tell us anything about causality. just because there is a strong correlation does not mean that one caused the other.
· It does not tell us which variable is the independent and which one is the dependent variable.
· We must consider the influence of latent variables.
· We must consider the phenomenon whereby independent variables strongly correlate with each other, which is known as collinearity.
So, let’s consider a partial correlation. The partial correlation gives us a more accurate picture of the true correlation between two variables because it controls for a third variable (Rasinger 166). If one of our coefficients does not turn out to be significant, we won’t need to use a partial correlation.
One last point to make about the Pearson correlation: It is not only useful for the analysis of our data, but it is also useful to test whether our measurement is reliable (Rasinger 184). Let’s think this through. A reliable measure should give us similar results if applied at two different points in time. If our method delivers such difference results in the retest, it means that our data is not stable. Note, however, that using Pearson’s r to evaluate a method’s reliability only works if other variables are kept constant between the test and the re-test.
causality and significance
Does this mean anything? Well, it might. If we square the Pearson correlation coefficient, the r-squared can tell us how much the independent variable accounts for the outcome of the dependent variable. That is to say, causality can be approximated via r-squared (Rasinger 171). While we can say that there is a relatively strong influence, we cannot say there one causes the other.
When we present measures such as the Pearson’s r, we must also give the significance value. Otherwise, our result is meaningless. Significance in statistics is the probability of our results being due to chance or due to something else that might mean something important. It shows us the likelihood that our result is reliable. Statistical significance is denoted with a p, with p fluctuating between 0 and 1 inclusive, translating into 0 to 100 percent. So, the smaller the p, the less likely our result is to be due to chance (Rasinger 172). In this way, we are covering all our bases. We are admitting the possibility of being wrong.
regression analysis models
In this last section of the chapter, we go back to our scatter plots. Let’s think about those again. Specifically, how we can calculate an imaginary straight line into our data on our scatter plot. This line is called a regression line (or trend line). The regression line an idealized representation of our data, and it is calculated in such a way that it lies in the middle of all data points, no matter how spread apart our data happens to be (Rasinger 174). Here are four properties that define a regression line:
· If the data correlates positively, the slope will be from bottom-left to top-right;
· If the correlation is negative, its slope will be from top-left to bottom-right;
· The stronger the correlation is, the steeper the slope; the weaker, the flatter it is;
· and if there is no correlation, there will be no slope and the trendline horizontal.
Now we can make forecasts! When we do forecasts, we have to assume that our sample data is representative for our population. In order to make predictions, we could simply use the scatter plot with its regression line, extend the line and simply read the values of the x and y-axis. In this way, and with little effort and a small data set, we can make interesting forecasts about data that is not even there (Rasinger 178). But it is important that we interpret our statistical output in the light of our theoretical and methodological framework before we can come to any conclusions. After all, a statistical result is only as good as what you can make of it.
With our simple regression model, we can forecast models containing one dependent and one independent variable. But, what is we have not one but many independent variables? Using a simple regression won’t cut it. Cue the multiple regression analysis. This tool will allow us to build a model that contains several independent variables (Rasinger 181). We may even be able to tell how much each of them influences our dependent variable!