Provide reflection of this week that contains 1-2 paragraph, which address at least one or two of the following topics

profiledream86
Biostatistics_CH09_MultivariableMethods-2.pptx

Chapter 9

Multivariable Methods

Learning Objectives (1 of 2)

Define and provide examples of dependent and independent variables in a study of a public health problem

Explain the principle of statistical adjustment to a lay audience

Organize data for regression analysis

Learning Objectives (2 of 2)

Define and provide an example of confounding

Define and provide an example of effect modification

Interpret coefficients in multiple linear and multiple logistic regression analysis

Definitions

Confounding—the distortion of the effect of a risk factor on an outcome

Effect modification—a different relationship between the risk factor and an outcome depending on the level of another variable

Confounding

A confounder is related to the risk factor and also to the outcome.

Assessing confounding

Formal tests of hypothesis

Clinically meaningful associations

We wish to assess the association between obesity and incident cardiovascular disease.

Example 9.1. Confounding (1 of 2)

Example 9.1. Confounding (2 of 2)

Is age a confounder?

A clinical trial is run to assess the efficacy of a new drug to increase HDL cholesterol.

H0: m1 = m2 versus H1:m1 ≠ m2

Z = –1.13 is not statistically significant.

Example 9.2. Effect Modification (1 of 2)

Is there effect modification by gender?

Example 9.2. Effect Modification (2 of 2)

Effect Modification

Cochran-Mantel-Haenszel Method

Technique to estimate association between risk factor and outcome accounting for confounding

Data are organized into stratum and associations are estimated in each stratum and combined.

Correlation and Simple Linear Regression Analysis

Two continuous variables

Y = dependent, outcome variable

X = independent, predictor variable

Relationship between age and SBP, number of hours of exercise and percent body fat, caffeine consumption and blood sugar level.

Correlation and Simple Linear Regression

Correlation—nature and strength of linear association between variables

Regression—equation that best describes relationship between variables

Scatter Diagram

Y 10 14 15 17 19 21 25 28 30 35 39 5 8 3 4 9 12 16 10 17 18 21

X

Y

Correlation Coefficient

Population correlation r

Sample correlation r, –1 ≤ r ≤ +1

Sign indicates nature of relationship (positive or direct, negative, or inverse).

Magnitude indicates strength.

Direct Relationship Between X and Y, r = 0.6

Y 10 14 15 17 19 21 25 28 30 35 39 5 8 3 4 9 12 16 10 17 18 21

X

Y

Inverse Relationship Between X and Y, r = –0.6

Y 40 37 38 32 28 25 17 19 12 15 10 5 8 3 4 9 12 16 10 17 18 21

X

Y

Sample Correlation Coefficient

Example (1 of 6)

Suppose we are interested in the relationship between body mass index (BMI; computed as the ratio of weight in kilograms to height in meters squared) and systolic blood pressure in males 50 years of age.

Example (2 of 6)

A random sample of 10 males 50 years of age is selected and their weights, heights, and systolic blood pressures are measured.

Their weights and heights are transformed into body mass index scores (see next slide).

In this analysis, the independent (or predictor) variable is BMI and the dependent (or response) variable is systolic blood pressure.

Example (3 of 6)

Data

X = BMI Y = SBP

18.4 120

20.1 110

22.4 120

25.9 135

26.5 140

28.9 115

30.1 150

32.9 165

33.0 160

34.7 180

X = BMI (X – ) (X – )2

18.4 –8.89 79.032

20.1 –7.19 51.696

22.4 –4.89 23.912

25.9 –1.39 1.932

26.5 –0.79 0.624

28.9 1.61 2.592

30.1 2.81 7.896

32.9 5.61 31.472

33.0 5.71 32.604

34.7 7.41 54.908

272.9 286.669

= 27.29

Example (4 of 6)

Y = SBP (Y – ) (Y – )2

120 –19.5 380.25

110 –29.5 870.25

120 –19.5 380.25

135 –4.5 20.25

140 0.5 0.25

115 –24.5 600.25

150 10.5 110.25

165 25.5 650.25

160 20.5 420.25

180 40.5 1640.2

1395 5072.50

= 139.5

Example (5 of 6)

(X – ) (Y – ) (X – )(Y – )

–8.89 –19.5 173.355

–7.19 –29.5 212.105

–4.89 –19.5 95.355

–1.39 –4.5 6.255

–0.79 0.50 –0.395

1.61 –24.5 –39.455

2.81 10.5 29.505

5.61 25.5 143.055

5.71 20.5 117.055

7.41 40.5 300.105

1036.95

cov (X,Y) = 1036.95/9

= 115.22

Example (6 of 6)

Sample Correlation Coefficient

Example (1 of 2)

Suppose in the same study we also measure the number of hours of vigorous exercise per week.

Is there a relationship between the number of hours of exercise and SBP in males 50 years of age?

Example (2 of 2)

Data

X = Hours of Exercise Y = SBP

4 120

10 110

2 120

3 135

3 140

5 115

1 115

2 165

2 160

0 180

Sample Correlation Coefficient

Simple Linear Regression

Y = Dependent, outcome variable

X = Independent, predictor variable

= b0 + b1 x

b0 is the Y-intercept, b1 is the slope

Simple Linear Regression Assumptions

Linear relationship between X and Y

Independence of errors

Homoscedasticity (constant variance) of the errors

Normality of errors

Least Squares Estimates of Regression Parameters

Regression Analysis: BMI and SBP

Using Regression Equation

What is expected SBP for a male with BMI = 20?

Compare two males whose BMIs differ by 2 units. How do SBPs compare?

Person with higher BMI will have SBP that is 2(3.61) = 7.22 units higher.

Regression Analysis: Exercise and SBP

Example 9.6. Linear Regression Analysis

Clinical trial to assess the efficacy of a new drug to increase HDL cholesterol

where 1 = new drug and 0 = placebo

Multiple Linear Regression

Y = continuous outcome variable

X1, X2, …, Xp = set of independent or predictor variables

Multiple Regression Analysis (1 of 2)

Model is conditional, parameter estimates are conditioned on other variables in model.

Perform overall test of regression.

If significant, examine individual predictors.

Relative importance of predictors by p-values (or standardized coefficients)

Multiple Regression Analysis (2 of 2)

Predictors can be continuous, indicator variables (0/1), or a set of dummy variables.

Dummy variables (for categorical predictors)

Race: white, black, Hispanic

Black (1 if black, 0 otherwise)

Hispanic (1 if Hispanic, 0 otherwise)

Example 9.7. Multiple Linear Regression Analysis

Outcome = infant birth weight, grams

Independent Regression

Variable Coefficient t p-value

Intercept –3850.92 –11.56 0.0001

Male gender 174.79 6.06 0.0001

Gestational age, weeks 179.89 22.35 0.0001

Mother’s age, years 1.38 0.47 0.6361

Black race –138.46 –1.93 0.0535

Hispanic race –13.07 –0.37 0.7103

Other race –68.67 –1.05 0.2918

Simple Logistic Regression Analysis

Outcome is dichotomous (1 = event, 0 = non-event) and p = P(event).

Outcome is modeled as log odds.

Multiple Logistic Regression Analysis

Outcome is dichotomous (1 = event, 0 = non-event) and p = P(event).

Outcome is modeled as log odds.

Example 9.8. Logistic Regression Analysis

Study to assess the relationship between obesity and incident CVD

Estimation of Regression Coefficients

Model parameters are estimated using maximum likelihood techniques.

b1 is the log odds ratio.

exp(b1) is the odds ratio estimate from a logistic regression model.

Interpretation of Regression Coefficients in Logistic Regression (1 of 2)

With a dichotomous predictor X, b1 is a log odds ratio for success for group1 versus group2.

With a continuous predictor X, b1 is a log odds ratio for success per unit change in X.

b1 = 0  No association between Y and X

b1 > 0  Probability of success increases as X increases

b1 < 0  Probability of success decreases as X increases

Interpretation of Regression Coefficients in Logistic Regression (2 of 2)

Multiple Logistic Regression Model for Hypertension (Y/N)

Predictor b p OR (95% CI for OR)

Intercept –5.407 0.0001

Age 0.052 0.0001 1.053 (1.044 – 1.062)

Male –0.250 0.0007 0.779 (0.674 – 0.900)

BMI 0.158 0.0001 1.171 (1.146 – 1.198)