Statistics in Health Care Management
Chapter 9
Multivariable Methods
Objectives
• Define and provide examples of dependent and
independent variables in a study of a public
health problem
• Explain the principle of statistical adjustment
to a lay audience
• Organize data for regression analysis
Objectives
• Define and provide an example of confounding
• Define and provide an example of effect
modification
• Interpret coefficients in multiple linear and
multiple logistic regression analysis
Definitions
• Confounding – the distortion of the effect of a
risk factor on an outcome
• Effect Modification – a different relationship
between the risk factor and an outcome
depending on the level of another variable
Confounding
• A confounder is related to the risk factor and
also to the outcome
• Assessing confounding
– Formal tests of hypothesis
– Clinically meaningful associations
Example 9.1.
Confounding
We wish to assess the association between obesity and
incident cardiovascular disease.
Incident CVD
No CVD
Total
Obese 46 254 300
Not Obese
60 640 700
Total 106 894 1000
1.78 0.086
0.153
60/700
46/300 RR
CVD
Example 9.1.
Confounding
Is age a confounder?
Age
< 50
CVD No CVD
Total Age 50+
CVD No CVD
Total
Obese 10 90 100 Obese 36 164 200
Not Obese
35 465 500 Not Obese
25 175 200
Total 45 555 600 Total 65 335 400
1.44 0.13
0.18 RR and 1.43
0.07
0.10 RR
50 Age|CVD50Age|CVD
Example 9.2.
Effect Modification
A clinical trial is run to assess the efficacy of a new drug
to increase HDL cholesterol.
N Mean Std Dev
New drug 50 40.16 4.46
Placebo 50 39.21 3.91
H0: m1m2 versus H1:m1≠m2
Z=-1.13 is not statistically significant
Example 9.2.
Effect Modification
Is there effect modification by gender?
Women N Mean Std Dev
New drug 40 38.88 3.97
Placebo 41 39.24 4.21
Men N Mean Std Dev
New drug 10 45.25 1.89
Placebo 9 39.06 2.22
Effect Modification
34
36
38
40
42
44
46
Women Men
M e
a n
H D
L
Gender
Placebo
New Drug
Cochran-Mantel-Haenszel Method
• Technique to estimate association between risk
factor and outcome accounting for
confounding
• Data are organized into stratum and
associations are estimated in each stratum and
combined
Correlation and Simple Linear Regression
Analysis
• Two continuous variables
– Y= dependent, outcome variable
– X=independent, predictor variable
Relationship between age and SBP, number of
hours of exercise and percent body fat, caffeine
consumption and blood sugar level.
Correlation and Simple Linear Regression
• Correlation – nature and strength of linear
association between variables
• Regression – equation that best describes
relationship between variables
Scatter Diagram
0
5
10
15
20
25
0 5 10 15 20 25 30 35 40 45
X
Y
Correlation Coefficient
• Population correlation r
• Sample correlation r, -1 < r < +1
• Sign indicates nature of relationship (positive
or direct, negative or inverse)
• Magnitude indicates strength
Direct Relationship Between X and Y, r = 0.6
0
5
10
15
20
25
0 5 10 15 20 25 30 35 40 45
X
Y
Inverse Relationship Between X and Y, r = -
0.6
0
5
10
15
20
25
0 5 10 15 20 25 30 35 40 45
X
Y
Sample Correlation Coefficient
1 -n
)X - (X Σ = s
2 2
x 1 -n
)Y - (Y Σ =s
2 2
y
2
y
2
x ss
Y)cov(X, =r
1 -n
)Y - (Y )X - (X Σ = Y)cov(X,
Example
Suppose we are interested in the relationship
between body mass index (computed as the
ratio of weight in kilograms to height in meters
squared) and systolic blood pressure in males
50 years of age.
Example
A random sample of 10 males 50 years of age is selected
and their weights, heights and systolic blood pressures are
measured. Their weights and heights are transformed into
body mass index scores and are given below. In this
analysis, the independent (or predictor) variable is body
mass index and the dependent (or response) variable is
systolic blood pressure.
Example
• Data X = BMI Y = SBP
18.4 120
20.1 110
22.4 120
25.9 135
26.5 140
28.9 115
30.1 150
32.9 165
33.0 160
34.7 180
X = BMI (X- ) (X- )2
18.4 -8.89 79.0322
20.1 -7.19 51.696
22.4 -4.89 23.912
25.9 -1.39 1.932
26.5 -0.79 0.624
28.9 1.61 2.592
30.1 2.81 7.896
32.9 5.61 31.472
33.0 5.71 32.604
34.7 7.41 54.908
272.9 286.669
X
X
1-n
)X - (X Σ = s
2
2
x
X
= 27.29
852.31 9
286.669 = s
2
x
Y = SBP(Y- ) (Y- )2
120 -19.5 380.25
110 -29.5 870.25
120 -19.5 380.25
135 -4.5 20.25
140 0.5 0.25
115 -24.5 600.25
150 10.5 110.25
165 25.5 650.25
160 20.5 420.25
180 40.5 1640.2
1395 5072.50
Y
1-n
)Y - (Y Σ =s
2
2
y
= 139.5Y
Y
611.563 9
5072.5 =s
2
y
(X- ) (Y- ) (X- )(Y- )
-8.89 -19.5 173.355
-7.19 -29.5 212.105
-4.89 -19.5 95.355
-1.39 -4.5 6.255
-0.79 0.50 -0.395
1.61 -24.5 -39.455
2.81 10.5 29.505
5.61 25.5 143.055
5.71 20.5 117.055
7.41 40.5 300.105
1036.95
1-n
)Y)(YX - (X Σ = Y)cov(X,
Y
cov (X,Y) = 1036.95/9
= 115.22
X YX
Sample Correlation Coefficient
0.86 .611)31.852(563
115.22
ss
Y)cov(X, =r
2
y
2
x
Example
Suppose in the same study we also measure
the number of hours of vigorous exercise per
week. Is there a relationship between the
number of hours of exercise and SBP in males
50 years of age?
Example
• Data X = # Hrs Exercise Y = SBP
4 120
10 110
2 120
3 135
3 140
5 115
1 115
2 165
2 160
0 180
Sample Correlation Coefficient
0.75 11)7.73(563.6
49.33
ss
Y)cov(X, =r
2
y
2
x
Simple Linear Regression
Y = Dependent, Outcome variable
X = Independent, Predictor variable
= b0 + b1 x
b0 is the Y-intercept, b1 is the slope
ŷ
Simple Linear Regression
Assumptions
• Linear relationship between X and Y
• Independence of errors
• Homoscedasticity (constant variance) of the errors
• Normality of errors
Least Squares Estimates of Regression
Parameters
x
y
1 s
s r = b
X b - Y = b 10
Regression Analysis: BMI and SBP
61.3 852.31
611.563 86.0
s
s r = b
x
y
1
98.40)29.27)(61.3(5.139X b - Y = b 10
x3.61 40.98 ŷ
Using Regression Equation
• What is expected SBP for a male with BMI=20?
• Compare 2 males whose BMIs differ by 2 units – how do SBPs compare?
Person with higher BMI will have SBP that is 2(3.61) = 7.22 units higher
113.81 (20) 3.61 40.98 ŷ
Regression Analysis: Exercise and SBP
38.6 73.7
611.563 75.0
s
s r = b
x
y
1
9.159)2.3)(38.6(5.139X b - Y = b 10
x6.38 - 159.9 ŷ
Example 9.6.
Linear Regression Analysis
Clinical trial to assess the efficacy of a new drug to
increase HDL cholesterol:
Treatment 0.95 9.21ŷ
Treatment 0.36- 24.39ŷ :WOMEN
Treatment 6.19 06.39ŷ :MEN
where 1=new drug and 0=placebo
Multiple Linear Regression
Y = continuous outcome variable
X1, X2, …, Xp = set of independent or predictor
variables
x b + . . .+ x b + x b + b = ŷ pp22110
Multiple Regression Analysis
• Model is conditional, parameter estimates are
conditioned on other variables in model
• Perform overall test of regression
– If significant, examine individual predictors
– Relative importance of predictors by p-values (or
standardized coefficients)
Multiple Regression Analysis
• Predictors can be continuous, indicator
variables (0/1) or a set of dummy variables
• Dummy variables (for categorical predictors)
– Race: white, black, Hispanic
• Black (1 if black, 0 otherwise)
• Hispanic (1 if Hispanic, 0 otherwise)
Example 9.7.
Multiple Linear Regression Analysis
Outcome = infant birth weight, grams
Independent Regression
Variable Coefficient t p-value
Intercept -3850.92 -11.56 0.0001
Male gender 174.79 6.06 0.0001
Gestational age, weeks 179.89 22.35 0.0001
Mother’s age, years 1.38 0.47 0.6361
Black race -138.46 -1.93 0.0535
Hispanic race -13.07 -0.37 0.7103
Other race -68.67 -1.05 0.2918
Simple Logistic Regression
Analysis
• Outcome is dichotomous (1=event, 0=non-event) and
p=P(event)
• Outcome is modeled as log odds
Xbb
Xbb
10
10
e1
e p̂
xbb p1
p lnlogit(p)log(odds)
10
Multiple Logistic Regression
Analysis
• Outcome is dichotomous (1=event, 0=non-event) and
p=P(event)
• Outcome is modeled as log odds
pp22110 xb ... xb xbb
p̂-1
p̂ ln
Example 9.8.
Logistic Regression Analysis
Study to assess the relationship between obesity and
incident CVD.
1.52 exp(0.415) RÔ
Group Age 0.655 Obesity 0.4152.592 p̂-1
p̂ ln
1.93 exp(0.658) RÔ
Obesity 0.6582.367 p̂-1
p̂ ln
Estimation of Regression
Coefficients
Model parameters are estimated using maximum
likelihood techniques
b1 is the log odds ratio
exp(b1) is the odds ratio estimate from a logistic
regression model
Interpretation of Regression Coefficients in
Logistic Regression
With a dichotomous predictor X, b1 is a log odds ratio for success for group1 versus group2
With a continuous predictor X, b1 is a log odds ratio for success per unit change in X
Interpretation of Regression Coefficients in
Logistic Regression
b1 = 0 No association between Y and X
b1 > 0 Probability of success increases as
X increases
b1 < 0 Probability of success decreases as
X increases
Multiple Logistic Regression Model for
Hypertension (Y/N)
Predictor b p OR (95% CI for OR)
Intercept -5.407 0.0001
Age 0.052 0.0001 1.053 (1.044-1.062)
Male -0.250 0.0007 0.779 (0.674-0.900)
BMI 0.158 0.0001 1.171 (1.146-1.198)