5-6 Forestry paper with graphs, due 4.15

profileNancy 00001
stats_primer.docx

A brief primer on statistical analysis

For this lab, you will be performing statistical analyses, with the help of some Excel tools. If you have already taken FRST231 or are presently taking it, this should be a refresher for you. Otherwise, you may not have ever performed statistical analysis, and will come across some confusing terms and numbers you may not know how to interpret. Hopefully this will clear up some confusion.

Why do we need to use statistics?

In this lab, we are seeking to explain genetic variation in tree growth by comparing measurements across taxonomic varieties, and analysing the relationship between our measurements and the climates of our trees’ provenances. When collecting data, there will always be some degree of unaccounted-for variation in our measurements. This variation may be caused by measurement error, natural variation caused by genetic differences, environmental variation, or more likely some combination of all these. As a result, real relationships are “messy”. Below are two figures showing a hypothetical relationship between provenance temperature and heights measured in a common garden. While not perfect, figure 1 clearly shows a strong relationship between these variables. But what about figure 2? It’s not as clear. Statistics allow us to ask “How likely is it that the pattern we’re seeing is the result of a real underlying relationship, rather than just random chance?” We can quantify this likelihood and assess the probability that our results reflect real relationships.

C:\Users\jcdegner.stu\AppData\Local\Microsoft\Windows\INetCache\Content.Word\stats_fig1.png In this lab we will be performing two different statistical analyses: linear regression, and t-tests. Linear regression is a method to determine whether a relationship exists between two numeric variables, e.g., height and provenance mean annual temperature, as shown in figures 1 and 2. A t-test is used to determine whether a relationship exists between one numeric variable and one categorical variable, e.g., height and taxonomic variety, as shown in figures 3 and 4.

C:\Users\jcdegner.stu\Documents\Jon Degner\MS-PhD\Teaching\FRST 210\stats_fig2.png

Figure 2: The effect of provenance mean annual precipitation on root-to-shoot ratios in 5-year old Pseudotsuga menziesii trees grown in a common garden in Campbell River, BC (n=5-7 per provenance).

Figure 1: The effect of provenance mean annual temperature on height in 7-year old Pinus contorta trees grown in a common garden in Prince George, BC (n=10 per provenance).

p-values

The cornerstone of most statistical analyses is the p-value. The p-value is a statistic that reflects the probability that an observed trend could happen by random chance, and it varies from 0 to 1. If p is close to 0, there is a very low chance that the pattern we see is due to chance (i.e., it is likely to reflect a real relationship). If p is close to 1, there is no difference between our data and data generated at random, so we should not interpret there to be a relationship here. Figure 1 shows a trend with a p-value of 0.0001, meaning that if we randomly generated 10,000 datasets, we could expect one of them to show a trend this strong. Figure 2 shows a trend with a p-value of 0.1, meaning that random data could create a pattern at least as strong as this 10% of the time.

We set an arbitrary threshold on p-values to assign statistical significance to a trend, meaning we are willing to accept some probability that our results are due to chance, but that number is generally low. In most scientific fields, we set a significance threshold of 0.05, meaning we are willing to accept a 5% chance that our results are due to random chance. A p-value less than 0.05 means we can consider that result "statistically significant" and therefore a reflection of a real pattern in the data. By this criterion, figure 1 shows a statistically significant relationship, while figure 2 does not. P-values are used in both linear regression and t-tests. If a linear regression has a p-value less than 0.05, we say that that the variables used in that analysis have a significant relationship. If a t-test has a p-value less than 0.05, we say that the two groups we are comparing show significant differences.

Linear regression

Linear regression is a commonly-used method across many fields of biology. In this method, we are comparing data with two dimensions: an independent variable (generally meaning one that we assign or that is known ahead of time), and a dependent variable (generally one that we measure, also known as a response variable). In this lab, our independent variables will be some aspects of climate from various provenances of Garry oak. Our dependent variables will be tree heights and diameters. Linear regressions are always presented as scatter plots, where the independent variable goes on the x-axis (the horizontal axis of the scatter plot), and the dependent variable goes on the y-axis (the vertical axis of the scatter plot). In this way, we can interpret the way that our dependent variable changes in relation to our independent variable i.e., in figure 1 if the provenance mean annual temperature changes by 1°C, how much does tree height change? This is the relationship quantified by linear regression.

Linear regression uses a set of equations to draw a line of fit through our data. This is shown by the grey diagonal lines in figures 1 and 2. A line of fit is a straight line that best describes the relationship in the data, and can be used to quantitatively describe that relationship using its slope. For example, the linear regression of figure 1 has a slope of 1.17m/°C, meaning for every 1°C that provenance mean annual temperature increases, we can expect height to increase by 1.17m. Another important aspect of linear regression is called the coefficient of determination, which has the mathematical symbol R2. This is a measurement of how strongly the two variables in our regression relate to one another. R2 ranges from 0 to 1, with 0 meaning there is no correlation between our data, and 1 meaning that the data is perfectly correlated i.e. all of the points in our scatter plot would fall exactly along the line of best fit. The relationship in figure 1 has an R2 of 0.62, which can be interpreted as “provenance mean annual temperature explains 62% of the variation in mean provenance heights”. Both slope and coefficient of determination are quantitative descriptions of significant relationships. Therefore, they should not be reported or interpreted for regressions with p-values greater than 0.05.

t-tests

Often, we wish to know whether two groups differ for some variable of interest. In this lab, we will be determining whether two taxonomic varieties of Garry oak differ in their average height or number of stems. A simple statistical method for this is called the t-test. A t-test determines whether the means of two groups of data are significantly different. This is determined not only by the means of those two groups, but the amount of variation in each group. This variation is referred to as error in statistics, although that term is misleading. As mentioned earlier, variation in our data is expected even if we measure things perfectly. We will quantify this variation in our groups using a parameter called standard deviation. This number is the average amount that each measurement in a group differs from the mean. For example, in figure 3, the mean height for Pinus contorta var. latifolia individuals is 9.9m, with a standard deviation of 1.1m. This means that, on average, individuals range from 8.8-11m within this group (the upper and lower end of the error bars in this figure).

C:\Users\jcdegner.stu\Documents\Jon Degner\MS-PhD\Teaching\FRST 210\stats_fig4.png C:\Users\jcdegner.stu\Documents\Jon Degner\MS-PhD\Teaching\FRST 210\stats_fig3.png

Figure 4: Belowground biomass of 5-year old Pseudotsuga menziesii trees grown in fertilized (n=40) and unfertilized (n=35) plots in a common garden in Campbell River, BC. Error bars represent standard deviation.

Figure 3: Mean heights of 7-year old Pinus contorta var. latifolia (n=100) and P. contorta var. contorta (n=40), grown in a common garden in Prince George, BC. Error bars represent standard deviation.

In figures 3 and 4, the groups both appear to differ. Pinus contorta var. latifolia individuals appear taller than var. contorta (fig. 3), and fertilized Pseudotsuga menziesii appear to grow more roots in fertilized plots (fig. 4). However, the error bars in figure 4 are larger than those in figure 3, meaning the data in figure 3 has more variation. A t-test can tell us whether these differences are statistically significant. The outputs of a t-test are a t-value and a p-value. The t-value is difficult to interpret without a much deeper dive into statistics, but a low value (near 0) represents less difference between groups than a larger value. However, the p-value is the same as discussed earlier. If p < 0.05, then the groups are significantly different, if p > 0.05, the groups are not significantly different. In figure 3, p = 0.001. These groups are significantly different. In figure 4, p = 0.09. These groups do not differ significantly. If our groups differ significantly, we can discuss the relative difference between them e.g., in figure 3, var. contorta individuals have an average height of 8.5m, 15% shorter than var. latifolia individuals. If the groups do not differ significantly, we infer that there is no difference between the mean values of our two groups and should not attempt to quantify this.