Critique of two research articles

profileMichelle_Michy
Lecture4samplingvalidityandreliability.pdf

Populations and Samples

Validity, Reliability, Accuracy

What do they actually mean?

Neil O’Connell PH2600 Research Methods 2018

Learning Outcomes By the end of this lecture you should be able to:

 Define the terms population and sample

 Discuss some of the issues related to participant recruitment

 Describe the various types of sampling methods

 Discuss the importance of sample size and power

 Define the terms Validity and Reliability and Accuracy

 Discuss the issues of validity and reliability with regards to outcome measures

 Discuss the issues of internal and external validity with regards to studies

TARGET POPULATION

SAMPLE

ACCESSIBLE POPULATION

Populations and samples

Population: Every person in the group of interest (target population)

Accessible population: Everyone in that group who you have access to

Sample: A subgroup of the population who you actually recruit.

You will only ever have a sample. The goal is to make it as representative of the population as possible

Inclusion/ Exclusion criteria

Sampling Methods

What people do I have access to? What’s my time period for recruitment?

Sampling Method

 Can introduce known and unknown biases through the methods used to recruit sample

 Selection bias: The way you select people may mean your sample isn’t representative of your population

 Example: Sampling obese people from people attending a weight-loss clinic – obese people attending a weight-loss clinic may be severely obese and/or they may be motivated to lose weight

 Volunteer bias: People who volunteer to participate in research may be different to those who refuse

Sampling Method Definition Advantages Disadvantages

Simple random sampling

Each member of the population has an equal chance of being selected. (e.g.use a random numbers table)

Should result in a truly representative sample - bias minimised

Pragmatically difficult, particularly in large populations

Systematic sampling Select every nth person in the accessible population

More efficient than random sampling

The order of the list can introduce a systematic bias

Stratified sampling Sampling ensures a proportion of representation across specific characteristics

Ensures balance on characteristics that are known to be important

Getting info difficult, time consuming and sometimes arbitrary

Probability Sampling Methods

Stratified Sampling

Consider: 270 children with CP on a database - Want to randomly select participants

Random sampling could result in an unequal number of boys and girls

Randomly select participants from each strata e.g. 10 boys and 10 girls

Non probability sampling methods

Sampling Methods

Description Advantages Disadvantages

Convenience (incidental) sampling

Take who you can get!

Easy and efficient

Easily affected by bias- sample may not be representative

Snowball sampling*

Get those you have to ask around!

Purposive sampling*

Handpick those who meet your needs

* More common in qualitative research

SIZE MATTERS

 How big a sample do you need?

 Bigger generally= better

 Larger samples more representative

 Larger samples give more precision

 More sensitive to detect differences between groups/ conditions

 For a survey you need enough people to be sure that your sample is representative.

 For an experiment you need enough to be sensitive to detect a change

 The larger the true effect the easier it is to detect and so the smaller the required sample

TYPE 1 ERROR TYPE 2 ERROR

FALSE POSITIVE:

DETECTING AN EFFECT THAT ISN’T ACTUALLY THERE

(FAILURE TO ACCEPT THE NULL HYPOTHESIS)

FALSE NEGATIVE:

FAILING TO DETECT AND EFFECT THAT IS ACTUALLY THERE

(ERRONEOUSLY ACCEPTING THE NULL HYPOTHESIS)

The two errors

Sample Size Calculation

 Finger in the air method common and profoundly dodgy

 Better to perform a formal sample size calculation

 But for this you need to establish some background data

 The method varies depending on the study design

 This will be discussed in detail in the lecture on inferential statistics

Measurement

Selecting an outcome measure

All outcome measures should be accurate and feasible

What does accurate mean?  Valid  Reliable

How feasible is it for me to use?

Validity & Reliability of specific measures

 We need to know whether our outcome measures are valid and reliable

 We can only be sure that we are accurately quantifying something if both have been established

Validity

‘Measures what it intends to measure’

………..a ruler cannot measure the weight of an object!

(Hicks, 2009)

FACE VALIDITY

CRITERION VALIDITY

CONSTRUCT VALIDITY

CONTENT VALIDITY

Construct validity- What is a construct?

 An artificial framework that is not directly observable….

 Abstract ideas that explain observable behaviours

 E.G: depression, wellness, IQ…

Construct validity

 Do scores on the test accurately reflect the “construct” being measured?

 Is the test a consistent reflection of the underlying theory of the “construct”

Construct Validity

 Can be established by reviewing the literature of the theory that pertains to the topic being researched.

 Can also be established by examining the convergent and divergent validity of a measure.

Construct validity

Convergent Validity Divergent

Validity

compares the target test with other measures believed to

measure the same construct.

The results should correlate highly if the same construct is

reflected in both tests.

compares the target test with other measures believed to

measure different characteristics or traits.

A low correlation is expected in this case.

Convergent and Divergent Validity

M cG

il l

Q oL

Q ue

st io

nn ai

re

Ro la

nd M

or ri

s di

sa bi

li ty

qu

es ti

on na

ir e

SF-36 SF-36

Content Validity

 Does the test measure a range of behaviours that it purports to measure based on the theoretical concepts on which the test is founded?

 Typically refers to questionnaires

Content Validity  A wide-range of observable, quantifiable

behaviours should be contained in the measure.

 The content should logically follow on from the literature reviewed for the construct validity stage as well as personal and professional experience.

 Can be established by using expert or user review, or by generating and testing hypotheses.

Content Validity

 Example: Assessing ability of people with COPD to do ADLs using a questionnaire

 Conduct a literature review to establish the ADLs that are important to people with COPD

 Develop the questionnaire based on the literature review

 Ask people with COPD to review the questionnaire and decide if it covers to topics that are important to them and if there’s anything missing

Criterion Validity  Criterion validity is established by comparing the new

measure with an accepted gold standard of measurement a.k.a. a criterion measure

kcal per minute kcal per minute

Indirect calorimetry

Type Definition

Face Validity Does the test look as though it is measuring what it is supposed to?

Construct Validity

Does the test measure the theory relating to the topic under investigation?

Content Validity Does the test measure a full range of behaviours expected to emerge from the theory?

Criterion validity

Does the test show good levels of agreement with an accepted gold standard?

Adapted from Hicks 2009 p.269

Reliability

 The consistency or repeatability of a measure

 NOTHING is 100% reliable

INTRA RATER RELIABILITY

• Degree of consistency of a set of measures taken by the same

person

INTER RATER

RELIABILITY

• Degree of consistency in the measures taken by 2 or more people

INSTRUMENT RELIABILITY

INTER/ INTRA

• The degree of consistency associated

with the specific measurement tool

Threats to Reliability  What threats to

reliability might arise during goniometry?

 Intra-rater  Inter-rater  Instrument

 ….intrasubject?

TRUTH

Unreliable and not valid

Valid but not reliable

Reliable but not valid Valid and Reliable

Test Accuracy

How accurate is my test?  Classification Accuracy: how well a test correctly

identifies or excludes a condition

e.g. does the Lachman test correctly identify people who have and don’t have an ACL rupture

 Sensitivity: how well a test correctly identifies a condition

e.g. does the Lachman test correctly identify people who have an ACL rupture

 Specificity: how well a test correctly excludes a condition

e.g. does the Lachman test correctly identify people who don’t have an ACL rupture

For example……

ACL rupture

No ACL rupture

total

Positive Test 5 1 6

Negative Test

5 9 14

total 10 10 20

A physio does the Lachman test on 20 people to test for ACL rupture

Arthroscopy indicated that 10 people had an ACL rupture and 10 people didn’t

For example……

ACL rupture

No ACL rupture

Positive Test 5 1 6

Negative Test

5 9 14

10 10 20

A physio does the Lachman test on 20 people to test for ACL rupture

Arthroscopy indicated that 10 people had an ACL rupture and 10 people didn’t

ACL rupture

No ACL rupture

Positive Test 5 1

Negative Test 5 9

10 10

True positives False positives

False negatives True negatives

ACL rupture

No ACL rupture

Positive Test 5 1

Negative Test 5 9

10 10

True positives

Sensitivity = Number of people correctly identified as having an ACL rupture/all people with an ACL rupture

i.e. Sensitivity = True positives/(true positives + false negatives)

Sensitivity = 5/10 = 50%

ACL rupture No ACL rupture

Positive Test 5 1

Negative Test 5 9

10 10

True negatives

Specificity = People correctly identified as not having an ACL rupture correctly identified/All people without an ACL rupture

i.e. Specificity = True negatives/(false positives + true negatives)

Specificity = 9/10 = 90%

ACL rupture No ACL rupture

Positive Test 5 1

Negative Test 5 9

10 10

True positives

Classification accuracy = People with and without ACL rupture correctly identified/All people who were tested

i.e. Classification Accuracy = True positives + true negatives/(positives + negatives)

Accuracy = 5+9/20 = 70%

True negatives

Validity at the “study level”

 Just to confuse you, the term “validity” is also often used in the critical appraisal of whole research studies.

 In this case it means something a little different

 When we ask “Is this study valid?” we are asking 2 overall questions…..

CAN I TRUST THESE RESULTS?

ARE THE METHODS ROBUST?

WHAT IS THE RISK OF BIAS?

ARE THE RESULTS GENERALIZABLE ?

INTERNAL VALIDITY EXTERNAL VALIDITY

Bias

 Bias – the presence of systematic error in a study

 Brings you further away from the truth

Should I believe the results?

Internal validity

 Refers to how well controlled a study is. Internal validity is high when:

◦ The risk of bias is low ◦ The risk of confounding is low ◦ The methods are tightly controlled

Threats to internal validity

 Experimenter effects  Participant effects  History/ Maturation  Regression to the mean  Unreliable/ invalid measurements  Selection/ Assignment  Confounding  Attrition/ dropout

External Validity

 Refers to how far the results can be generalized to wider populations

 External validity is high when: ◦ Sampling is broad and representative ◦ The study conditions mimic real world

conditions

Threats to external validity

 Non random sampling (recruitment bias)

 Restrictive inclusion/ exclusion criteria

 Tightly controlled experimental conditions

 Study conducted in a highly unique environment

 Small study size

How big is the effect?

Useful/ relevant outcomes?

Appropriate/ achievable

intervention?

Similar Patient Group?

Do I even have enough information?

A Tension

Internal validity

Tight control of all variables

Strict sampling Rigorous measurement

Experimental manipulation

External validity

Broad and inclusive Looser control – variability

allowed Reflects real world practice

and measures

Summary

 How you select your sample is a critical methodological issue

 Acceptable validity and reliability are vital properties for a good outcome measure

 Understanding what is meant by both and threats to both is important

 At the study level validity (internal/ external) refers to how rigorous a study is and how generalizable its results are to a broader population

Recommended reading

 On selection and assignment of participants: Chapter 9 Carter RE, Lubinsky J, Domholdt (2011) Rehabilitation Research (4th edition). USA: Elsevier

 On validity and reliability of a measure: Pages 237-244 Carter RE, Lubinsky J, Domholdt (2011) Rehabilitation Research (4th edition). USA: Elsevier Pages 267 and 268 Hicks, C. (2009)

 On study validity: Chapter 8 Carter RE, Lubinsky J, Domholdt (2011) Rehabilitation Research (4th edition). USA: Elsevier