Health Care Law and Legislation, Statistics Policies
Chapter 9
Multivariable Methods
Learning Objectives (1 of 2)
- Define and provide examples of dependent and independent variables in a study of a public health problem
- Explain the principle of statistical adjustment to a lay audience
- Organize data for regression analysis
Learning Objectives (2 of 2)
- Define and provide an example of confounding
- Define and provide an example of effect modification
- Interpret coefficients in multiple linear and multiple logistic regression analysis
Definitions
- Confounding—the distortion of the effect of a risk factor on an outcome
- Effect modification—a different relationship between the risk factor and an outcome depending on the level of another variable
Confounding
- A confounder is related to the risk factor and also to the outcome.
- Assessing confounding
- Formal tests of hypothesis
- Clinically meaningful associations
- We wish to assess the association between obesity and incident cardiovascular disease.
Example 9.1.
Confounding (1 of 2)
Example 9.1.
Confounding (2 of 2)
- Is age a confounder?
- A clinical trial is run to assess the efficacy of a new drug to increase HDL cholesterol.
H0: m1 = m2 versus H1:m1 ≠ m2
Z = –1.13 is not statistically significant.
Example 9.2.
Effect Modification (1 of 2)
- Is there effect modification by gender?
Example 9.2.
Effect Modification (2 of 2)
Effect Modification
Cochran-Mantel-Haenszel Method
- Technique to estimate association between risk factor and outcome accounting for confounding
- Data are organized into stratum and associations are estimated in each stratum and combined.
Correlation and Simple Linear
Regression Analysis
- Two continuous variables
Y = dependent, outcome variable
X = independent, predictor variable
- Relationship between age and SBP, number of hours of exercise and percent body fat, caffeine consumption and blood sugar level.
Correlation and Simple
Linear Regression
- Correlation—nature and strength of linear association between variables
- Regression—equation that best describes relationship between variables
Scatter Diagram
Chart1
| 10 |
| 14 |
| 15 |
| 17 |
| 19 |
| 21 |
| 25 |
| 28 |
| 30 |
| 35 |
| 39 |
Sheet1
| X | Y |
| 10 | 5 |
| 14 | 8 |
| 15 | 3 |
| 17 | 4 |
| 19 | 9 |
| 21 | 12 |
| 25 | 16 |
| 28 | 10 |
| 30 | 17 |
| 35 | 18 |
| 39 | 21 |
Sheet2
Sheet3
Correlation Coefficient
- Population correlation r
- Sample correlation r, –1 ≤ r ≤ +1
- Sign indicates nature of relationship (positive or direct, negative, or inverse).
- Magnitude indicates strength.
Direct Relationship Between
X and Y, r = 0.6
Chart1
| 10 |
| 14 |
| 15 |
| 17 |
| 19 |
| 21 |
| 25 |
| 28 |
| 30 |
| 35 |
| 39 |
Sheet1
| X | Y |
| 10 | 5 |
| 14 | 8 |
| 15 | 3 |
| 17 | 4 |
| 19 | 9 |
| 21 | 12 |
| 25 | 16 |
| 28 | 10 |
| 30 | 17 |
| 35 | 18 |
| 39 | 21 |
Sheet2
Sheet3
Inverse Relationship Between
X and Y, r = –0.6
Chart1
| 40 |
| 37 |
| 38 |
| 32 |
| 28 |
| 25 |
| 17 |
| 19 |
| 12 |
| 15 |
| 10 |
Sheet1
| X | Y |
| 40 | 5 |
| 37 | 8 |
| 38 | 3 |
| 32 | 4 |
| 28 | 9 |
| 25 | 12 |
| 17 | 16 |
| 19 | 10 |
| 12 | 17 |
| 15 | 18 |
| 10 | 21 |
Sheet2
Sheet3
Sample Correlation Coefficient
Example (1 of 6)
- Suppose we are interested in the relationship between body mass index (BMI; computed as the ratio of weight in kilograms to height in meters squared) and systolic blood pressure in males 50 years of age.
Example (2 of 6)
- A random sample of 10 males 50 years of age is selected and their weights, heights, and systolic blood pressures are measured.
- Their weights and heights are transformed into body mass index scores (see next slide).
- In this analysis, the independent (or predictor) variable is BMI and the dependent (or response) variable is systolic blood pressure.
Example (3 of 6)
- Data
X = BMI Y = SBP
18.4 120
20.1 110
22.4 120
25.9 135
26.5 140
28.9 115
30.1 150
32.9 165
33.0 160
34.7 180
X = BMI (X – ) (X – )2
18.4 –8.89 79.032
20.1 –7.19 51.696
22.4 –4.89 23.912
25.9 –1.39 1.932
26.5 –0.79 0.624
28.9 1.61 2.592
30.1 2.81 7.896
32.9 5.61 31.472
33.0 5.71 32.604
34.7 7.41 54.908
272.9 286.669
= 27.29
Example (4 of 6)
Y = SBP (Y – ) (Y – )2
120 –19.5 380.25
110 –29.5 870.25
120 –19.5 380.25
135 –4.5 20.25
140 0.5 0.25
115 –24.5 600.25
150 10.5 110.25
165 25.5 650.25
160 20.5 420.25
180 40.5 1640.2
1395 5072.50
= 139.5
Example (5 of 6)
(X – ) (Y – ) (X – )(Y – )
–8.89 –19.5 173.355
–7.19 –29.5 212.105
–4.89 –19.5 95.355
–1.39 –4.5 6.255
–0.79 0.50 –0.395
1.61 –24.5 –39.455
2.81 10.5 29.505
5.61 25.5 143.055
5.71 20.5 117.055
7.41 40.5 300.105
1036.95
cov (X,Y) = 1036.95/9
= 115.22
Example (6 of 6)
Sample Correlation Coefficient
Example (1 of 2)
- Suppose in the same study we also measure the number of hours of vigorous exercise per week.
- Is there a relationship between the number of hours of exercise and SBP in males 50 years of age?
Example (2 of 2)
- Data
X = Hours of Exercise Y = SBP
4 120
10 110
2 120
3 135
3 140
5 115
1 115
2 165
2 160
0 180
Sample Correlation Coefficient
Simple Linear Regression
Y = Dependent, outcome variable
X = Independent, predictor variable
= b0 + b1 x
b0 is the Y-intercept, b1 is the slope
Simple Linear Regression Assumptions
- Linear relationship between X and Y
- Independence of errors
- Homoscedasticity (constant variance) of the errors
- Normality of errors
Least Squares Estimates of
Regression Parameters
Regression Analysis: BMI and SBP
Using Regression Equation
- What is expected SBP for a male with BMI = 20?
- Compare two males whose BMIs differ by 2 units. How do SBPs compare?
Person with higher BMI will have SBP that is 2(3.61) = 7.22 units higher.
Regression Analysis: Exercise and SBP
Example 9.6.
Linear Regression Analysis
- Clinical trial to assess the efficacy of a new drug to increase HDL cholesterol
where 1 = new drug and 0 = placebo
Multiple Linear Regression
Y = continuous outcome variable
X1, X2, …, Xp = set of independent or predictor variables
Multiple Regression Analysis (1 of 2)
- Model is conditional, parameter estimates are conditioned on other variables in model.
- Perform overall test of regression.
- If significant, examine individual predictors.
- Relative importance of predictors by p-values (or standardized coefficients)
Multiple Regression Analysis (2 of 2)
- Predictors can be continuous, indicator variables (0/1), or a set of dummy variables.
- Dummy variables (for categorical predictors)
- Race: white, black, Hispanic
Black (1 if black, 0 otherwise)
Hispanic (1 if Hispanic, 0 otherwise)
Example 9.7.
Multiple Linear Regression Analysis
Outcome = infant birth weight, grams
Independent Regression
Variable Coefficient t p-value
Intercept –3850.92 –11.56 0.0001
Male gender 174.79 6.06 0.0001
Gestational age, weeks 179.89 22.35 0.0001
Mother’s age, years 1.38 0.47 0.6361
Black race –138.46 –1.93 0.0535
Hispanic race –13.07 –0.37 0.7103
Other race –68.67 –1.05 0.2918
Simple Logistic Regression Analysis
- Outcome is dichotomous (1 = event, 0 = non-event) and p = P(event).
- Outcome is modeled as log odds.
Multiple Logistic Regression Analysis
- Outcome is dichotomous (1 = event, 0 = non-event) and p = P(event).
- Outcome is modeled as log odds.
Example 9.8.
Logistic Regression Analysis
- Study to assess the relationship between obesity and incident CVD
Estimation of Regression Coefficients
- Model parameters are estimated using maximum likelihood techniques.
b1 is the log odds ratio.
exp(b1) is the odds ratio estimate from a logistic regression model.
Interpretation of Regression Coefficients in Logistic Regression (1 of 2)
- With a dichotomous predictor X, b1 is a log odds ratio for success for group1 versus group2.
- With a continuous predictor X, b1 is a log odds ratio for success per unit change in X.
b1 = 0 No association between Y and X
b1 > 0 Probability of success increases as X increases
b1 < 0 Probability of success decreases as X increases
Interpretation of Regression Coefficients in Logistic Regression (2 of 2)
Multiple Logistic Regression Model for Hypertension (Y/N)
Predictor b p OR (95% CI for OR)
Intercept –5.407 0.0001
Age 0.052 0.0001 1.053 (1.044 – 1.062)
Male –0.250 0.0007 0.779 (0.674 – 0.900)
BMI 0.158 0.0001 1.171 (1.146 – 1.198)
1.78
0.086
0.153
60/700
46/300
RR
CVD
=
=
=
1.44
0.13
0.18
RR
and
1.43
0.07
0.10
RR
50
Age
|
CVD
50
Age
|
CVD
=
=
=
=
+
<
0
5
10
15
20
25
051015202530354045
X
Y
0
5
10
15
20
25
051015202530354045
X
Y
1
-
n
)
X
-
(X
Σ
=
s
2
2
x
1
-
n
)
Y
-
(Y
Σ
=
s
2
2
y
2
y
2
x
s
s
Y)
cov(X,
=
r
1
-
n
)
Y
-
(Y
)
X
-
(X
Σ
=
Y)
cov(X,
X
X
1
-
n
)
X
-
(X
Σ
=
s
2
2
x
852
.
31
9
286.669
=
s
2
x
=
Y
1
-
n
)
Y
-
(Y
Σ
=
s
2
2
y
611
.
563
9
5072.5
=
s
2
y
=
1
-
n
)
Y
)(Y
X
-
(X
Σ
=
Y)
cov(X,
-
X
0.86
.611)
31.852(563
115.22
s
s
Y)
cov(X,
=
r
2
y
2
x
=
=
0.75
11)
7.73(563.6
49.33
s
s
Y)
cov(X,
=
r
2
y
2
x
-
=
-
=
y
ˆ
x
y
1
s
s
r
=
b
X
b
-
Y
=
b
1
0
61
.
3
852
.
31
611
.
563
86
.
0
s
s
r
=
b
x
y
1
=
=
98
.
40
)
29
.
27
)(
61
.
3
(
5
.
139
X
b
-
Y
=
b
1
0
=
-
=
x
3.61
40.98
y
ˆ
+
=
113.81
(20)
3.61
40.98
y
ˆ
=
+
=
38
.
6
73
.
7
611
.
563
75
.
0
s
s
r
=
b
x
y
1
-
=
-
=
0b = Y − 1b X =139.5− (−6.38)(3.2) =159.9
0
b
= Y-
1
b
X=139.5-(-6.38)(3.2)=159.9
ŷ = 159.9 – 6.38 x
ˆ
y = 159.9 – 6.38 x
Treatment
0.95
9.21
y
ˆ
+
=
Men: ŷ = 39.06 + 6.19 Treatment Women: ŷ = 39.24 – 0.36 Treatment
Men:
ˆ
y=39.06 + 6.19 Treatment
Women:
ˆ
y=39.24 – 0.36 Treatment
x
b
+
.
.
.
+
x
b
+
x
b
+
b
=
y
ˆ
p
p
2
2
1
1
0
X
b
b
X
b
b
1
0
1
0
e
1
e
p
ˆ
+
+
+
=
x
b
b
p
1
p
ln
logit(p)
log(odds)
1
0
+
=
÷
÷
ø
ö
ç
ç
è
æ
-
=
=
ln p̂ 1 – p̂ ⎛
⎝ ⎜
⎞
⎠ ⎟= b0 + b1x1 + b2x2 + ... + bpxp
ln
ˆ
p
1 –
ˆ
p
æ
è
ç
ö
ø
÷
=b
0
+b
1
x
1
+ b
2
x
2
+ ... + b
p
x
p
ln p̂ 1− p̂ ⎛
⎝ ⎜
⎞
⎠ ⎟= −2.592 + 0.415 Obesity + 0.655 Age Group
ÔR = exp(0.415) = 1.52
ln
ˆ
p
1-
ˆ
p
æ
è
ç
ö
ø
÷
=-2.592+0.415 Obesity + 0.655 Age Group
ˆ
OR = exp(0.415) = 1.52
ln p̂ 1 – p̂ ⎛
⎝ ⎜
⎞
⎠ ⎟= −2.367+ 0.658 Obesity
ÔR = exp(0.658) = 1.93
ln
ˆ
p
1 –
ˆ
p
æ
è
ç
ö
ø
÷
=-2.367+0.658 Obesity
ˆ
OR = exp(0.658) = 1.93