Critique of two research articles
Populations and Samples
Validity, Reliability, Accuracy
What do they actually mean?
Neil O’Connell PH2600 Research Methods 2018
Learning Outcomes By the end of this lecture you should be able to:
Define the terms population and sample
Discuss some of the issues related to participant recruitment
Describe the various types of sampling methods
Discuss the importance of sample size and power
Define the terms Validity and Reliability and Accuracy
Discuss the issues of validity and reliability with regards to outcome measures
Discuss the issues of internal and external validity with regards to studies
TARGET POPULATION
SAMPLE
ACCESSIBLE POPULATION
Populations and samples
Population: Every person in the group of interest (target population)
Accessible population: Everyone in that group who you have access to
Sample: A subgroup of the population who you actually recruit.
You will only ever have a sample. The goal is to make it as representative of the population as possible
Inclusion/ Exclusion criteria
Sampling Methods
What people do I have access to? What’s my time period for recruitment?
Sampling Method
Can introduce known and unknown biases through the methods used to recruit sample
Selection bias: The way you select people may mean your sample isn’t representative of your population
Example: Sampling obese people from people attending a weight-loss clinic – obese people attending a weight-loss clinic may be severely obese and/or they may be motivated to lose weight
Volunteer bias: People who volunteer to participate in research may be different to those who refuse
Sampling Method Definition Advantages Disadvantages
Simple random sampling
Each member of the population has an equal chance of being selected. (e.g.use a random numbers table)
Should result in a truly representative sample - bias minimised
Pragmatically difficult, particularly in large populations
Systematic sampling Select every nth person in the accessible population
More efficient than random sampling
The order of the list can introduce a systematic bias
Stratified sampling Sampling ensures a proportion of representation across specific characteristics
Ensures balance on characteristics that are known to be important
Getting info difficult, time consuming and sometimes arbitrary
Probability Sampling Methods
Stratified Sampling
Consider: 270 children with CP on a database - Want to randomly select participants
Random sampling could result in an unequal number of boys and girls
Randomly select participants from each strata e.g. 10 boys and 10 girls
Non probability sampling methods
Sampling Methods
Description Advantages Disadvantages
Convenience (incidental) sampling
Take who you can get!
Easy and efficient
Easily affected by bias- sample may not be representative
Snowball sampling*
Get those you have to ask around!
Purposive sampling*
Handpick those who meet your needs
* More common in qualitative research
SIZE MATTERS
How big a sample do you need?
Bigger generally= better
Larger samples more representative
Larger samples give more precision
More sensitive to detect differences between groups/ conditions
For a survey you need enough people to be sure that your sample is representative.
For an experiment you need enough to be sensitive to detect a change
The larger the true effect the easier it is to detect and so the smaller the required sample
TYPE 1 ERROR TYPE 2 ERROR
FALSE POSITIVE:
DETECTING AN EFFECT THAT ISN’T ACTUALLY THERE
(FAILURE TO ACCEPT THE NULL HYPOTHESIS)
FALSE NEGATIVE:
FAILING TO DETECT AND EFFECT THAT IS ACTUALLY THERE
(ERRONEOUSLY ACCEPTING THE NULL HYPOTHESIS)
The two errors
Sample Size Calculation
Finger in the air method common and profoundly dodgy
Better to perform a formal sample size calculation
But for this you need to establish some background data
The method varies depending on the study design
This will be discussed in detail in the lecture on inferential statistics
Measurement
Selecting an outcome measure
All outcome measures should be accurate and feasible
What does accurate mean? Valid Reliable
How feasible is it for me to use?
Validity & Reliability of specific measures
We need to know whether our outcome measures are valid and reliable
We can only be sure that we are accurately quantifying something if both have been established
Validity
‘Measures what it intends to measure’
………..a ruler cannot measure the weight of an object!
(Hicks, 2009)
FACE VALIDITY
CRITERION VALIDITY
CONSTRUCT VALIDITY
CONTENT VALIDITY
Construct validity- What is a construct?
An artificial framework that is not directly observable….
Abstract ideas that explain observable behaviours
E.G: depression, wellness, IQ…
Construct validity
Do scores on the test accurately reflect the “construct” being measured?
Is the test a consistent reflection of the underlying theory of the “construct”
Construct Validity
Can be established by reviewing the literature of the theory that pertains to the topic being researched.
Can also be established by examining the convergent and divergent validity of a measure.
Construct validity
Convergent Validity Divergent
Validity
compares the target test with other measures believed to
measure the same construct.
The results should correlate highly if the same construct is
reflected in both tests.
compares the target test with other measures believed to
measure different characteristics or traits.
A low correlation is expected in this case.
Convergent and Divergent Validity
M cG
il l
Q oL
Q ue
st io
nn ai
re
Ro la
nd M
or ri
s di
sa bi
li ty
qu
es ti
on na
ir e
SF-36 SF-36
Content Validity
Does the test measure a range of behaviours that it purports to measure based on the theoretical concepts on which the test is founded?
Typically refers to questionnaires
Content Validity A wide-range of observable, quantifiable
behaviours should be contained in the measure.
The content should logically follow on from the literature reviewed for the construct validity stage as well as personal and professional experience.
Can be established by using expert or user review, or by generating and testing hypotheses.
Content Validity
Example: Assessing ability of people with COPD to do ADLs using a questionnaire
Conduct a literature review to establish the ADLs that are important to people with COPD
Develop the questionnaire based on the literature review
Ask people with COPD to review the questionnaire and decide if it covers to topics that are important to them and if there’s anything missing
Criterion Validity Criterion validity is established by comparing the new
measure with an accepted gold standard of measurement a.k.a. a criterion measure
kcal per minute kcal per minute
Indirect calorimetry
Type Definition
Face Validity Does the test look as though it is measuring what it is supposed to?
Construct Validity
Does the test measure the theory relating to the topic under investigation?
Content Validity Does the test measure a full range of behaviours expected to emerge from the theory?
Criterion validity
Does the test show good levels of agreement with an accepted gold standard?
Adapted from Hicks 2009 p.269
Reliability
The consistency or repeatability of a measure
NOTHING is 100% reliable
INTRA RATER RELIABILITY
• Degree of consistency of a set of measures taken by the same
person
INTER RATER
RELIABILITY
• Degree of consistency in the measures taken by 2 or more people
INSTRUMENT RELIABILITY
INTER/ INTRA
• The degree of consistency associated
with the specific measurement tool
Threats to Reliability What threats to
reliability might arise during goniometry?
Intra-rater Inter-rater Instrument
….intrasubject?
TRUTH
Unreliable and not valid
Valid but not reliable
Reliable but not valid Valid and Reliable
Test Accuracy
How accurate is my test? Classification Accuracy: how well a test correctly
identifies or excludes a condition
e.g. does the Lachman test correctly identify people who have and don’t have an ACL rupture
Sensitivity: how well a test correctly identifies a condition
e.g. does the Lachman test correctly identify people who have an ACL rupture
Specificity: how well a test correctly excludes a condition
e.g. does the Lachman test correctly identify people who don’t have an ACL rupture
For example……
ACL rupture
No ACL rupture
total
Positive Test 5 1 6
Negative Test
5 9 14
total 10 10 20
A physio does the Lachman test on 20 people to test for ACL rupture
Arthroscopy indicated that 10 people had an ACL rupture and 10 people didn’t
For example……
ACL rupture
No ACL rupture
Positive Test 5 1 6
Negative Test
5 9 14
10 10 20
A physio does the Lachman test on 20 people to test for ACL rupture
Arthroscopy indicated that 10 people had an ACL rupture and 10 people didn’t
ACL rupture
No ACL rupture
Positive Test 5 1
Negative Test 5 9
10 10
True positives False positives
False negatives True negatives
ACL rupture
No ACL rupture
Positive Test 5 1
Negative Test 5 9
10 10
True positives
Sensitivity = Number of people correctly identified as having an ACL rupture/all people with an ACL rupture
i.e. Sensitivity = True positives/(true positives + false negatives)
Sensitivity = 5/10 = 50%
ACL rupture No ACL rupture
Positive Test 5 1
Negative Test 5 9
10 10
True negatives
Specificity = People correctly identified as not having an ACL rupture correctly identified/All people without an ACL rupture
i.e. Specificity = True negatives/(false positives + true negatives)
Specificity = 9/10 = 90%
ACL rupture No ACL rupture
Positive Test 5 1
Negative Test 5 9
10 10
True positives
Classification accuracy = People with and without ACL rupture correctly identified/All people who were tested
i.e. Classification Accuracy = True positives + true negatives/(positives + negatives)
Accuracy = 5+9/20 = 70%
True negatives
Validity at the “study level”
Just to confuse you, the term “validity” is also often used in the critical appraisal of whole research studies.
In this case it means something a little different
When we ask “Is this study valid?” we are asking 2 overall questions…..
CAN I TRUST THESE RESULTS?
ARE THE METHODS ROBUST?
WHAT IS THE RISK OF BIAS?
ARE THE RESULTS GENERALIZABLE ?
INTERNAL VALIDITY EXTERNAL VALIDITY
Bias
Bias – the presence of systematic error in a study
Brings you further away from the truth
Should I believe the results?
Internal validity
Refers to how well controlled a study is. Internal validity is high when:
◦ The risk of bias is low ◦ The risk of confounding is low ◦ The methods are tightly controlled
Threats to internal validity
Experimenter effects Participant effects History/ Maturation Regression to the mean Unreliable/ invalid measurements Selection/ Assignment Confounding Attrition/ dropout
External Validity
Refers to how far the results can be generalized to wider populations
External validity is high when: ◦ Sampling is broad and representative ◦ The study conditions mimic real world
conditions
Threats to external validity
Non random sampling (recruitment bias)
Restrictive inclusion/ exclusion criteria
Tightly controlled experimental conditions
Study conducted in a highly unique environment
Small study size
How big is the effect?
Useful/ relevant outcomes?
Appropriate/ achievable
intervention?
Similar Patient Group?
Do I even have enough information?
A Tension
Internal validity
Tight control of all variables
Strict sampling Rigorous measurement
Experimental manipulation
External validity
Broad and inclusive Looser control – variability
allowed Reflects real world practice
and measures
Summary
How you select your sample is a critical methodological issue
Acceptable validity and reliability are vital properties for a good outcome measure
Understanding what is meant by both and threats to both is important
At the study level validity (internal/ external) refers to how rigorous a study is and how generalizable its results are to a broader population
Recommended reading
On selection and assignment of participants: Chapter 9 Carter RE, Lubinsky J, Domholdt (2011) Rehabilitation Research (4th edition). USA: Elsevier
On validity and reliability of a measure: Pages 237-244 Carter RE, Lubinsky J, Domholdt (2011) Rehabilitation Research (4th edition). USA: Elsevier Pages 267 and 268 Hicks, C. (2009)
On study validity: Chapter 8 Carter RE, Lubinsky J, Domholdt (2011) Rehabilitation Research (4th edition). USA: Elsevier