Statistics in Health Care Management

profilegregueira82
chapter9.pdf

Chapter 9

Multivariable Methods

Objectives

• Define and provide examples of dependent and

independent variables in a study of a public

health problem

• Explain the principle of statistical adjustment

to a lay audience

• Organize data for regression analysis

Objectives

• Define and provide an example of confounding

• Define and provide an example of effect

modification

• Interpret coefficients in multiple linear and

multiple logistic regression analysis

Definitions

• Confounding – the distortion of the effect of a

risk factor on an outcome

• Effect Modification – a different relationship

between the risk factor and an outcome

depending on the level of another variable

Confounding

• A confounder is related to the risk factor and

also to the outcome

• Assessing confounding

– Formal tests of hypothesis

– Clinically meaningful associations

Example 9.1.

Confounding

We wish to assess the association between obesity and

incident cardiovascular disease.

Incident CVD

No CVD

Total

Obese 46 254 300

Not Obese

60 640 700

Total 106 894 1000

1.78 0.086

0.153

60/700

46/300 RR

CVD 

Example 9.1.

Confounding

Is age a confounder?

Age

< 50

CVD No CVD

Total Age 50+

CVD No CVD

Total

Obese 10 90 100 Obese 36 164 200

Not Obese

35 465 500 Not Obese

25 175 200

Total 45 555 600 Total 65 335 400

1.44 0.13

0.18 RR and 1.43

0.07

0.10 RR

50 Age|CVD50Age|CVD 



Example 9.2.

Effect Modification

A clinical trial is run to assess the efficacy of a new drug

to increase HDL cholesterol.

N Mean Std Dev

New drug 50 40.16 4.46

Placebo 50 39.21 3.91

H0: m1m2 versus H1:m1≠m2

Z=-1.13 is not statistically significant

Example 9.2.

Effect Modification

Is there effect modification by gender?

Women N Mean Std Dev

New drug 40 38.88 3.97

Placebo 41 39.24 4.21

Men N Mean Std Dev

New drug 10 45.25 1.89

Placebo 9 39.06 2.22

Effect Modification

34

36

38

40

42

44

46

Women Men

M e

a n

H D

L

Gender

Placebo

New Drug

Cochran-Mantel-Haenszel Method

• Technique to estimate association between risk

factor and outcome accounting for

confounding

• Data are organized into stratum and

associations are estimated in each stratum and

combined

Correlation and Simple Linear Regression

Analysis

• Two continuous variables

– Y= dependent, outcome variable

– X=independent, predictor variable

Relationship between age and SBP, number of

hours of exercise and percent body fat, caffeine

consumption and blood sugar level.

Correlation and Simple Linear Regression

• Correlation – nature and strength of linear

association between variables

• Regression – equation that best describes

relationship between variables

Scatter Diagram

0

5

10

15

20

25

0 5 10 15 20 25 30 35 40 45

X

Y

Correlation Coefficient

• Population correlation r

• Sample correlation r, -1 < r < +1

• Sign indicates nature of relationship (positive

or direct, negative or inverse)

• Magnitude indicates strength

Direct Relationship Between X and Y, r = 0.6

0

5

10

15

20

25

0 5 10 15 20 25 30 35 40 45

X

Y

Inverse Relationship Between X and Y, r = -

0.6

0

5

10

15

20

25

0 5 10 15 20 25 30 35 40 45

X

Y

Sample Correlation Coefficient

1 -n

)X - (X Σ = s

2 2

x 1 -n

)Y - (Y Σ =s

2 2

y

2

y

2

x ss

Y)cov(X, =r

1 -n

)Y - (Y )X - (X Σ = Y)cov(X,

Example

Suppose we are interested in the relationship

between body mass index (computed as the

ratio of weight in kilograms to height in meters

squared) and systolic blood pressure in males

50 years of age.

Example

A random sample of 10 males 50 years of age is selected

and their weights, heights and systolic blood pressures are

measured. Their weights and heights are transformed into

body mass index scores and are given below. In this

analysis, the independent (or predictor) variable is body

mass index and the dependent (or response) variable is

systolic blood pressure.

Example

• Data X = BMI Y = SBP

18.4 120

20.1 110

22.4 120

25.9 135

26.5 140

28.9 115

30.1 150

32.9 165

33.0 160

34.7 180

X = BMI (X- ) (X- )2

18.4 -8.89 79.0322

20.1 -7.19 51.696

22.4 -4.89 23.912

25.9 -1.39 1.932

26.5 -0.79 0.624

28.9 1.61 2.592

30.1 2.81 7.896

32.9 5.61 31.472

33.0 5.71 32.604

34.7 7.41 54.908

272.9 286.669

X

X

1-n

)X - (X Σ = s

2

2

x

X

= 27.29

852.31 9

286.669 = s

2

x 

Y = SBP(Y- ) (Y- )2

120 -19.5 380.25

110 -29.5 870.25

120 -19.5 380.25

135 -4.5 20.25

140 0.5 0.25

115 -24.5 600.25

150 10.5 110.25

165 25.5 650.25

160 20.5 420.25

180 40.5 1640.2

1395 5072.50

Y

1-n

)Y - (Y Σ =s

2

2

y

= 139.5Y

Y

611.563 9

5072.5 =s

2

y 

(X- ) (Y- ) (X- )(Y- )

-8.89 -19.5 173.355

-7.19 -29.5 212.105

-4.89 -19.5 95.355

-1.39 -4.5 6.255

-0.79 0.50 -0.395

1.61 -24.5 -39.455

2.81 10.5 29.505

5.61 25.5 143.055

5.71 20.5 117.055

7.41 40.5 300.105

1036.95

1-n

)Y)(YX - (X Σ = Y)cov(X,

Y

cov (X,Y) = 1036.95/9

= 115.22

X YX

Sample Correlation Coefficient

0.86 .611)31.852(563

115.22

ss

Y)cov(X, =r

2

y

2

x



Example

Suppose in the same study we also measure

the number of hours of vigorous exercise per

week. Is there a relationship between the

number of hours of exercise and SBP in males

50 years of age?

Example

• Data X = # Hrs Exercise Y = SBP

4 120

10 110

2 120

3 135

3 140

5 115

1 115

2 165

2 160

0 180

Sample Correlation Coefficient

0.75 11)7.73(563.6

49.33

ss

Y)cov(X, =r

2

y

2

x

 

Simple Linear Regression

Y = Dependent, Outcome variable

X = Independent, Predictor variable

= b0 + b1 x

b0 is the Y-intercept, b1 is the slope

Simple Linear Regression

Assumptions

• Linear relationship between X and Y

• Independence of errors

• Homoscedasticity (constant variance) of the errors

• Normality of errors

Least Squares Estimates of Regression

Parameters

x

y

1 s

s r = b

X b - Y = b 10

Regression Analysis: BMI and SBP

61.3 852.31

611.563 86.0

s

s r = b

x

y

1 

98.40)29.27)(61.3(5.139X b - Y = b 10 

x3.61 40.98 ŷ 

Using Regression Equation

• What is expected SBP for a male with BMI=20?

• Compare 2 males whose BMIs differ by 2 units – how do SBPs compare?

Person with higher BMI will have SBP that is 2(3.61) = 7.22 units higher

113.81 (20) 3.61 40.98 ŷ 

Regression Analysis: Exercise and SBP

38.6 73.7

611.563 75.0

s

s r = b

x

y

1 

9.159)2.3)(38.6(5.139X b - Y = b 10 

x6.38 - 159.9 ŷ 

Example 9.6.

Linear Regression Analysis

Clinical trial to assess the efficacy of a new drug to

increase HDL cholesterol:

Treatment 0.95 9.21ŷ 

Treatment 0.36- 24.39ŷ :WOMEN

Treatment 6.19 06.39ŷ :MEN



where 1=new drug and 0=placebo

Multiple Linear Regression

Y = continuous outcome variable

X1, X2, …, Xp = set of independent or predictor

variables

x b + . . .+ x b + x b + b = ŷ pp22110

Multiple Regression Analysis

• Model is conditional, parameter estimates are

conditioned on other variables in model

• Perform overall test of regression

– If significant, examine individual predictors

– Relative importance of predictors by p-values (or

standardized coefficients)

Multiple Regression Analysis

• Predictors can be continuous, indicator

variables (0/1) or a set of dummy variables

• Dummy variables (for categorical predictors)

– Race: white, black, Hispanic

• Black (1 if black, 0 otherwise)

• Hispanic (1 if Hispanic, 0 otherwise)

Example 9.7.

Multiple Linear Regression Analysis

Outcome = infant birth weight, grams

Independent Regression

Variable Coefficient t p-value

Intercept -3850.92 -11.56 0.0001

Male gender 174.79 6.06 0.0001

Gestational age, weeks 179.89 22.35 0.0001

Mother’s age, years 1.38 0.47 0.6361

Black race -138.46 -1.93 0.0535

Hispanic race -13.07 -0.37 0.7103

Other race -68.67 -1.05 0.2918

Simple Logistic Regression

Analysis

• Outcome is dichotomous (1=event, 0=non-event) and

p=P(event)

• Outcome is modeled as log odds

Xbb

Xbb

10

10

e1

e p̂

 

xbb p1

p lnlogit(p)log(odds)

10 

  

 

Multiple Logistic Regression

Analysis

• Outcome is dichotomous (1=event, 0=non-event) and

p=P(event)

• Outcome is modeled as log odds

pp22110 xb ... xb xbb

p̂-1

p̂ ln 

  

Example 9.8.

Logistic Regression Analysis

Study to assess the relationship between obesity and

incident CVD.

1.52 exp(0.415) RÔ

Group Age 0.655 Obesity 0.4152.592 p̂-1

p̂ ln



 

  

1.93 exp(0.658) RÔ

Obesity 0.6582.367 p̂-1

p̂ ln



 

  

Estimation of Regression

Coefficients

Model parameters are estimated using maximum

likelihood techniques

b1 is the log odds ratio

exp(b1) is the odds ratio estimate from a logistic

regression model

Interpretation of Regression Coefficients in

Logistic Regression

With a dichotomous predictor X, b1 is a log odds ratio for success for group1 versus group2

With a continuous predictor X, b1 is a log odds ratio for success per unit change in X

Interpretation of Regression Coefficients in

Logistic Regression

b1 = 0  No association between Y and X

b1 > 0  Probability of success increases as

X increases

b1 < 0  Probability of success decreases as

X increases

Multiple Logistic Regression Model for

Hypertension (Y/N)

Predictor b p OR (95% CI for OR)

Intercept -5.407 0.0001

Age 0.052 0.0001 1.053 (1.044-1.062)

Male -0.250 0.0007 0.779 (0.674-0.900)

BMI 0.158 0.0001 1.171 (1.146-1.198)