Critique of two research articles
INFERENTIAL STATISTICAL ANALYSIS Dr Eve Corner
PhD, MRes, BSc (hons)
PLAN…
• Sample size calculations
• One and two tailed hypotheses
• Parametric and non-parametric tests
• Comparing groups
• Correlations
SAMPLE SIZE CALCULATIONS How many is enough?
Get enough people to get a representative sample and to find an effect
Don’t have time or money to recruit every person in the population
Point one:
Sample needs to be sufficiently
large to represent the population in
which it was derived/represents.
Running speed in seconds (100 metres)
9.23 10.5 10.5
8.98 10.8 10.8
10.56 11.1 11.9
17.5 15.3 25.5
18.12 13.6 26
12.3 14.8 17.5
16.5 12.1 20.1
14.3 11.7 8.7
9.2 11.3 11.2
Mean 13.7 seconds (4.54 SD)
Running speed in seconds (100 metres)
9.23 10.5 10.5
8.98 10.8 10.8
10.56 11.1 11.9
17.5 15.3 25.5
18.12 13.6 26
12.3 14.8 17.5
16.5 12.1 20.1
14.3 11.7 8.7
9.2 11.3 11.2 Mean: 17.32 (SD 6.75)
Running speed in seconds (100 metres)
9.23 10.5 10.5
8.98 10.8 10.8
10.56 11.1 11.9
17.5 15.3 25.5
18.12 13.6 26
12.3 14.8 17.5
16.5 12.1 20.1
14.3 11.7 8.7
9.2 11.3 11.2 Mean: 15.0 SD (6.44)
• Population mean (n=24): 13.7 seconds (SD 4.53)
• Sample means (n=3): 17.32 seconds (SD 6.75)
• Sample mean (n=6): 15.0 seconds (SD 6.44)
• Sample mean (n=9): 13.9 seconds (SD 5.35)
SAMPLE SIZE
The larger the sample sizes the less variability between the sample means i.e. the smaller the standard error
What if we had two samples,
and we want to know the
difference between the
groups?
Get enough people to get a representative sample and to find an effect
Don’t have time or money to recruit every person in the population
Is there a difference in 100m sprint time between those who exit
first and the remaining candidates?
Increased sample size = decreased variability
= greater power to detect an effect if one exists
SAMPLE SIZE CALCULATIONS
• We can calculate the sample size required to find a true difference (or effect) using a sample size calculation
• Trying to minimise the chance of a type I or type II error
• Type I error: Finding a difference (effect) when there isn’t actually one (false positive)
• Type II error: Not finding a difference (effect) when there is actually one (false negative)
TYPE 1 ERROR (false positive)
TYPE 2 ERROR (false negative)
SAMPLE SIZE
• The ability of a statistical test to detect a true difference (effect) is known as it’s power.
• The power of a test depends on the size of the sample and size of the difference (effect)
• The bigger the sample size the better chance of it finding a true difference (effect) if it’s there
• If the sample size is big it can detect a small difference (effect)
• If the sample size is small a test could still be powerful if the difference (effect size) is very big
EFFECT SIZE AND SAMPLE SIZE
100 metre sprint time
Case B
Case A
CALCULATING SAMPLE SIZE: T-TEST.
NB The method used to calculate the sample size depends on the statistical test that will be used to analyse the primary outcome
If using a t-test, to calculate a sample size required to find a difference between groups you need:
1. To know the size of the difference (effect) and an estimate of the variability within the population
2. To decide what risk are you willing to take that you’ll get a type I or type II error
SAMPLE SIZE CALCULATION FOR GROUP COMPARISONS
Info needed Description/Reason
α level (i.e. p < ?) Set the cut-off for the probability of a type I error.
Conventionally set at α = 0.05 (p < 0.05). This tells
you the probability of getting a false positive is 5%.
Probability of sending an innocent man to jail.
Power Set the cut-off for the probability of avoiding a type
II error. Often set at 80% or 90% (this is the
probability that you’ll detect a true effect if there is
one).
Probability of a guilty man walking free.
Predicted difference
(effect size)
How large do you expect it to be?
Predicted sample
variability
What’s the variability within the population?
Predicted difference (effect
size)
How large do you expect it to be?
How do you know this??
Clinical judgement of what’s important Changes observed in previous studies
Minimal Detectable change or Minimal
Clinically Important Difference
Predicted sample variance What’s the variation within the population?
i.e. what’s the standard deviation
How do you know this??
Standard deviation from previous studies
Pilot studies
What if I the information required isn’t in the literature?
• Pilot study
IMPORTANT POINT!
• 80% power means 80% chance of finding a true difference if it exists.
• BUT…. 20% chance of not finding it….
• AND if you under recruit, your power drops.
You may let a guilty man walk…
i.e. it may be a false negative i.e. no difference found, when a difference exists.
HOW DO THESE EFFECT MY SAMPLE SIZE?
Info needed Change Impact on sample size
α = 0.05 e.g. α = 0.01
Power = 80% e.g. power = 90%
Effect size
e.g.
difference =
2 seconds
e.g. Difference = 5 secs
Standard
deviation
e.g. 3 s
e.g. SD = 6 seconds
ONE OR TWO TAILED TESTS Which do I use?
ONE OR TWO TAILED TESTS
Trying to detect a difference in either
direction
Trying to detect a difference in
one direction
One-tailed test
Two-tailed test
ONE-TAILED OR TWO-TAILED?
• Do a two-tailed test if you think the difference between the groups could be positive or negative
e.g. Do 6MWT, do an exercise intervention, repeat 6MWT
The 6MWT results in the second test could be better or worse than in the first test…..therefore need a two-tailed test
NB: Unless it’s physically impossible to have change in either direction (i.e. positive or negative) use a two-tailed test
(If you choose a one-tailed test, you’re more likely to find a positive effect, even if it’s not there, i.e. type I error)
STATISTICAL TESTS For this module: compare groups & examining associations
COMPARING GROUPS
Within group Between groups
2 conditions 3 or more conditions
2 different groups 3 or more different groups
Parametric Paired t-test One-way
repeated measures Analysis of Variance (ANOVA)
Independent t-test One-way Analysis
of Variance (ANOVA)
Non- parametric
Wilcoxon test (ordinal data or ratio/interval data with a skewed
distribution)
McNemar’s test (nominal data)
Friedman test (ordinal data or ratio/interval data with a skewed
distribution)
Mann-Whitney U test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
Kruskal Wallis test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
ASSESSING CORRELATIONS
2 variables
Parametric test Pearson test
Non-parametric test Spearman test (ordinal
data or ratio/interval
data with a skewed
distribution)
Cohen’s Kappa
coefficient (nominal
data)
5 STEPS TO STATS
Step 1: What type of data is it and is it normally distributed?
Step 2: State the null and alternative hypothesis
Step 3: Set a significance level
Step 4: Choose a test statistic and conduct the test
Step 5: Interpret the result
PARAMETRIC OR NON-PARAMETRIC Which do I use?
PARAMETRIC VERSUS NON- PARAMETRIC
• A parametric statistical test makes assumptions about the parameters
(defining properties) of the population distribution(s) from which the
data are drawn:
• That data is normally distributed.
• That data is ratio or interval.
• Parametric tests are more sensitive or “powerful” than non-parametric tests
• Therefore more likely to detect a true effect if there is one
Parametric
Non-
parametric
ALTERNATIVES…
Within group Between groups
2 conditions 3 or more conditions
2 different groups 3 or more different groups
Parametric Paired t-test One-way
repeated measures Analysis of Variance (ANOVA)
Independent t-test One-way Analysis
of Variance (ANOVA)
Non- parametric
Wilcoxon test (ordinal data or ratio/interval data with a skewed
distribution)
McNemar’s test (nominal data)
Friedman test (ordinal data or ratio/interval data with a skewed
distribution)
Mann-Whitney U test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
Kruskal Wallis test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
EXAMPLES • 1. Comparing 6MWT results (metres) between two groups (positive skew)?
• Ratio but skew
• Two groups
• = Mann Whitney
• 2. Comparing Oxford scale grading of strength between two time points?
• Ordinal data
• One group
• = Wilcoxon test
• 3. Difference in number of men versus women categorised as overweight?
• Nominal
• Two groups
• = Chi-squared test
COMPARING GROUPS Two groups of people, one time point
COMPARING GROUPS
Within group Between groups
2 conditions 3 or more conditions
2 different groups 3 or more different groups
Parametric Paired t-test One-way
repeated measures Analysis of Variance (ANOVA)
Independent t-test One-way Analysis
of Variance (ANOVA)
Non- parametric
Wilcoxon test (ordinal data or ratio/interval data with a skewed
distribution)
McNemar’s test (nominal data)
Friedman test (ordinal data or ratio/interval data with a skewed
distribution)
Mann-Whitney U test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
Kruskal Wallis test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
Is there a difference in 100m sprint time between those who exit
first and the remaining candidates?
• Step one: What type of data?
• Time in seconds
• RATIO
• Is it normally distributed?
Categorical data
Nominal Ordinal
Continuous data
Interval Ratio
• Step two: State the null and alternative hypothesis:
• Null (H0): There is no difference between running speed and half way survival in adults in the Hunger Games.
• Alternative (H1): There is a difference between running speed and half way survival in adults in the Hunger Games. (two tailed)
• Step 3: Set a significance level?
• Probability of type 1 error
i.e. how willing are you to accept that you may find a difference when no difference exists?
• < 5% ?
• <1% ?
• α<.05
“Send an innocent man to jail”
STEP 4: CHOOSE A STATISTICAL TEST- PARAMETRIC
• To test for differences between groups we generally calculate a statistic known as “t”
• Paired t-test comparing one group that are measured on two occasions (e.g. pre and post)
• Independent t-test comparing two separate groups (e.g. intervention versus control)
• Which would you use?
Two groups = independent
T-TEST
Difference between means for group 1 and group 2
Standard error t =
Standard error: SD of the original distribution divided by the square route of N
Mean first 12: 15 seconds
Mean remaining 12: 9 seconds
Difference = 6 seconds
t = 0.238, p = 0.813
STEP 5: INTERPRETING THE RESULTS
p = 0.813
p is not < 0.05 therefore do not reject the null hypothesis
There is no real difference
COMPARING GROUPS One group of people, two time points
COMPARING GROUPS
Within group Between groups
2 conditions 3 or more conditions
2 different groups 3 or more different groups
Parametric Paired t-test One-way
repeated measures Analysis of Variance (ANOVA)
Independent t-test One-way Analysis
of Variance (ANOVA)
Non- parametric
Wilcoxon test (ordinal data or ratio/interval data with a skewed
distribution)
McNemar’s test (nominal data)
Friedman test (ordinal data or ratio/interval data with a skewed
distribution)
Mann-Whitney U test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
Kruskal Wallis test (ordinal data or ratio/interval data with a skewed
distribution)
Chi-squared test (nominal data)
Does the percentage of Love Island contestants in a relationship change as a result of the ‘intervention’
from 2012-2019?
TIME POINT ONE: Entry To Love Island
Single Single
Single Single Single
Single Single
In a
relationship
Single Single Single
1/11 in a relationship = 9.1%
TIME POINT TWO: Exit To Love Island
Single
6/11 in a relationship = 54.5%
COMPARING PERCENTAGES
Year Pre Post Difference
2018 9.1% 54.5% 45.40%
2017 0% 45.5% 45.50%
2016 0% 63.6% 63.60%
2015 9.1 % 54.5% 45.40%
2014 0% 54.5% 54.50%
2013 9.1 % 72.7% 63.60%
2012 0% 45.5% 45.50% Mean (SD) 3.9 % (0.05) 55.83% (0.10) 51.93% (7.99%)
• Step one: What type of data?
• Percentages
• Ratio
• Is it normally distributed?
• Step two: State the null and alternative hypothesis:
• Null (H0): There is no difference in proportion of Love Island Contestants who are in a relationship before and after they appear on the show between 2012 and 2018
• Alternative (H1): There is a difference in the in proportion of Love Island Contestants who are in a relationship before and after they appear on the show between 2012 and 2018 (two tailed)
• Alternative (H1): There is a higher proportion of Love Island Contestants who are in a relationship before and after they appear on the show between 2012 and 2018 (one tailed)
• Step 3: Set a significance level?
• Probability of type 1 error
i.e. how willing are you to accept that you may find a difference when no difference exists?
• < 5% ?
• <1% ?
• α<.05
“Send an innocent man to jail”
STEP 4: CHOOSE A STATISTICAL TEST- PARAMETRIC
• To test for differences between groups we generally calculate a statistic known as “t”
• Paired t-test comparing one group that are measured on two occasions (e.g. pre and post)
• Independent t-test comparing two separate groups (e.g. intervention versus control)
• Which would you use?
One group = Paired
Mean difference is 51.93 %
Standard deviation is 7.99 %
Year Pre Post Difference
2018 9.1% 54.5% 45.40%
2017 0% 45.5% 45.50%
2016 0% 63.6% 63.60%
2015 9.1 % 54.5% 45.40%
2014 0% 54.5% 54.50%
2013 9.1 % 72.7% 63.60%
2012 0% 45.5% 45.50%
Mean (SD) 3.9 % (0.05) 55.83% (0.10) 51.93% (7.99%)
STEP 4: CHOOSE A STATISTICAL TEST
• To test our null hypothesis the test-statistic we use is:
𝑡 = ҧ𝑑 − 0 𝑠 𝑛
Is the mean difference different to zero
Standard error (variability of sample means) t =
p = <.05
STEP 5: INTERPRET THE RESULTS
Ask yourself:
1. Is the difference between groups statistically significant?
2. What is the effect size? What is the mean difference between the groups relative to the group means?
In other words…
1. Is p <0.05?
2. If p <0.05 is the mean difference of 51.93% between groups big?
What’s meaningful change?
CORRELATIONS
ASSESSING CORRELATIONS
2 variables
Parametric test Pearson test
Non-parametric test Spearman test (ordinal
data or ratio/interval
data with a skewed
distribution)
Cohen’s Kappa
coefficient (nominal
data)
CORRELATIONS
• You may want to look at how two variables are associated with another
• You have two measures (or variables) within one group of people
• If the aim is to investigate the relationship between two variables need to calculate a correlation coefficient
• -1.0 (perfect negative association between two variables)
• +1.0 (perfect positive association between two variables)
• 0 (no association)
• Written as r • e.g. r= 0.4, r = -0.7
• Research question:
• Is there an association between 100m speed and position on the leaderboard?
• Step one: What type of data and is it normally distributed?
• Time in seconds
• RATIO
• Ranking
• ORDINAL
STEP TWO: STATE THE NULL AND ALTERNATIVE HYPOTHESIS:
• Null (H0): There is no association between running speed and ranking on the leaderboard in adults in the Hunger Games.
• Alternative (H1): There is an association between running speed and ranking on the leaderboard in adults in the Hunger Games. (two tailed)
• Alternative (H1): There is an association between faster running speed and better ranking on the leaderboard in adults in the Hunger Games. (one tailed)
• Step 3: Set a significance level?
• Probability of type 1 error
i.e. how willing are you to accept that you may find an association between running speed and ranking when no association exists?
• α<.05
“Send an innocent man to jail”
STEP 4: CHOOSE THE TEST PARAMETRIC OR NON- PARAMETRIC?
• Is the data ratio or interval?
• If so, is it normally distributed?
• Is the data ratio or interval?
• If so, is it normally distributed?
Parametric test is Pearson’s test a.k.a. Pearson’s
correlation coefficient (r)
Non-parametric equivalent is Spearman
test to calculate Spearman’s rho (r or rs)
0
5
10
15
20
25
30
35
0.00 10.00 20.00 30.00
L e
a d
e rb
o a
rd p
o si
ti o
n
Running speed
0
5
10
15
20
25
30
0.00 10.00 20.00 30.00
L e
a d
e rb
o a
rd p
so it o
n
Running speed
STEP 5: INTERPRET THE RESULTS
Perfect positive r= 1, p<.05 Perfect negative r= -1, p <.o5)
The faster you run, the more likely you are to win
or
The faster you run, the more likely you are to lose
0
5
10
15
20
25
30
0.00 5.00 10.00 15.00 20.00 25.00 30.00
L e
a d
e rb
o a
rd p
so it o
n
Running speed
CORRELATION BETWEEN RUNNING SPEED AND POSITION ON LEADER
BOARD
No association between running speed and leader
board position
r= ?
p= ?
CONFOUNDING
SUMMARY
• To answer the aim of your study you usually need to use conduct a statistical test
Step 1: Identify type of data and distribution (if interval/ratio)
Step 2: Develop a null hypothesis and alternative hypothesis
Step 3: Decide on a significance level
Step 4: Decide on a statistical test (parametric or non-parametric) and conduct the statistical test
Step 5: Interpret the results; both effect size (mean difference or r) and p value
QUIZ: WHAT TEST SHOULD I USE?
1. Investigating association between weight and pain in people with OA. Both are ratio data and data is normally distributed.
2. Investigating the difference in QoL between people with stroke and people without stroke. QoL is ordinal data.
3. Investigating the change in participation after a stretching programme in people with MS. Participation is ratio data and is not normally distributed.
4. Investigating the association between height and running speed in athletes. Both are ratio data and data is not normally distributed.
5. Investigating the difference in the prevalence of diabetes between people with and without obesity. Data is nominal.
USEFUL RESOURCES
• Martin Bland (2000) An Introduction to Medical Statistics. 3rd edition. Oxford: Oxford University Press.
• https://www.youtube.com/user/how2stats
• https://www.youtube.com/user/ProfAndyField
• Andy Field (2013) Discovering statistics using IBM SPSS Statistics: and sex and drugs and rock ‘n’ roll. 4th edition. Sage.