final_math_411_winter_2016_post.docx

7

MATH 411, Winter 2016, Take-Home Final Exam

Instructions

Answers must be submitted in typed hard copy form by 5 PM on Tuesday, March 15, 2016.

Exams may be submitted to any of the course instructors or left for us at the Math Department office.

On March 15 ONLY, exams may be submitted in Korman 207 (Math lounge) across the hall from the Math Department office, between 4 and 5 PM.

Do not submit answers with handwritten text, although symbols, equations, and figures may be produced

by hand if desired. Submit only your answers. Do not submit copies of any of the exam questions.

The first hard copy submission of the exam will be the ONLY basis for determining the grade.

No additional printed submissions, electronic submissions, or e-mailed corrections will be accepted.

Further guidelines

1. There are three problems, each worth about one-third of a total of 100 for the exam.

2. Answers should be concise and display knowledge of any statistical tests used, as well as adequate

justification of conclusions. Clarity of presentation is important, and points will be deducted if the answers

can’t be understood, or contain material not relevant to the specific problem.

3. Do not use a text font smaller than 12 points. Limit the text part of the answers to all three questions to no

more than three pages, although less text should be sufficient to get full credit.

4. The exam is open book, open notes, and open reference. Any books or web sites are fair sources of

information. You may NOT use live experts, especially professors, graduate students, or consulting services,

as resources.

5. You may discuss your work, ideas, and analyses with others in the course. However, you may not copy

others’ work, or provide your work to others.

You must hand in your own best answer in your own words, and you must be able to explain it yourself.

If essentially identical exams with essentially identical text, graphs, etc. are handed in,

or if answers are copied from any other source, the exam scores will be 0.

1. A data set obtained to study lung function includes the following variables.

pemax: maximum expiration pressure (cm H2O)

height: height in centimeters

weight: weight in kilograms

bmp: body mass (weight/height2) as percent of median for age

The goal was to model pemax in terms of the three other variables. Here is the scatterplot matrix of the data,

including histograms of each variable and pairwise least squares fits.

The first model examined included all three predictors. Ch, Cw, and Cb are symbols for the partial slopes.

Model 1: pemax = intercept + Ch*height + Cw*weight + Cb*bmp

Estimate Std. Error t value Pr(>|t|)

(Intercept) 245.3936 119.8927 2.047 0.0534

height -0.8264 0.7808 -1.058 0.3019

weight 2.7717 1.1377 2.436 0.0238

bmp -1.4876 0.7375 -2.017 0.0566

Residual standard error: 25.24 on 21 degrees of freedom

R-squared: 0.5015, Adjusted R-squared: 0.4302

F-statistic: 7.041 on 3 and 21 DF, p-value: 0.00187

a. Explain what can be concluded from the F test. Include a statement of the null hypothesis..

b. What can be concluded from the t tests and why?

Three additional models were also examined.

Model 2: pemax = intercept + Cw*weight + Cb*bmp

Estimate Std. Error t value Pr(>|t|)

(Intercept) 124.8297 37.4786 3.331 0.00303

weight 1.6403 0.3900 4.206 0.00037

bmp -1.0054 0.5814 -1.729 0.09780

Residual standard error: 25.31 on 22 degrees of freedom

R-squared: 0.4749, Adjusted R-squared: 0.4271

F-statistic: 9.947 on 2 and 22 DF, p-value: 0.0008374

Model 3: pemax = intercept + Cw*weight

Estimate Std. Error t value Pr(>|t|)

(Intercept) 63.5456 12.7016 5.003 4.63e-05

weight 1.1867 0.3009 3.944 0.000646

Residual standard error: 26.38 on 23 degrees of freedom

R-squared: 0.4035, Adjusted R-squared: 0.3776

F-statistic: 15.56 on 1 and 23 DF, p-value: 0.0006457

Model 4: pemax = intercept + Cb*bmp

Estimate Std. Error t value Pr(>|t|)

(Intercept) 59.0802 44.7444 1.320 0.20

bmp 0.6392 0.5652 1.131 0.27

Residual standard error: 33.24 on 23 degrees of freedom

R-squared: 0.05268, Adjusted R-squared: 0.01149

F-statistic: 1.279 on 1 and 23 DF, p-value: 0.2698

c. Based on the regression summaries of Models 1 and 2, which one appears to be the better model? Explain,

including brief discussion of the p-values, Residual standard error, and Adjusted R-squared, also

explaining how the latter differs from R-squared.

d. For Model 2, explain the meanings of the partial slopes for weight and for bmp, in context, in terms of

pemax, weight, and bmp, including correct units.

e. Are Models 3 and 4 consistent with Model 2? How, in particular, can it make sense that Model 4 has a

positive slope for bmp, while Model 2 has a negatuve partial slope for bmp, as well as a much smaller

p-value for that partial slope? The plot on the next page may help explain the answer.

2. Three diets for hamsters (labeled "I","II","III") were tested for differences in weight gain (measured as grams

of increase) after a specified period of time. Six inbred laboratory lines (labeled "A","B","C", "D", "E", "F")

were used to represent the responses of different genotypes to the various diets. The lines were treated as

blocks and all three diets were assigned randomly within each block. The data consisted of a total of 18

observations. Both one-way ANOVA and two-way ANOVA models were fit to the data. Shown below are an

interaction plot, and the ANOVA tables for 2 different models of the data.

Model 1: gain = Intercept + Cl*line + Cd*diet

Df Sum Sq Mean Sq F value Pr(>F)

line 5 71.17 14.23 7.491 0.00365

diet 2 36.33 18.17 9.561 0.00477

Residuals 10 19.00 1.90

Model 2: gain = Intercept + Cd*diet

Df Sum Sq Mean Sq F value Pr(>F)

diet 2 36.33 18.167 3.022 0.0789

Residuals 15 90.17 6.011

a. What are your conclusions and the reasons for those conclusions about the different diets, based on:

i. Model 1?

ii. Model 2?

b. Which model and conclusion is a better one? Justify your answer by discussing the effects of blocking

for these data. Refer in particular to the residual sums of squares and residual mean squares, and their

effects on the p-values in the two models.

c. i. What does the interaction plot suggest about whether or not line and diet interact, and why is this so?

ii. What is the implication of your answer to part i. for the validity of the randomized complete block

design model?

d. What kind of design model corresponds to the one-way ANOVA? Explain this design.

e. i. State the null hypothesis about the factor diet that is tested in the Model 1 ANOVA table.

ii. If you reject this null hypothesis, what is the next step in the analysis of the effect of diet on gain?

3. A data set obtained to study the protective effect of helmets used in a contact sport includes the following

variables.

Model: one of 4 helmet models (brands): A, B, C, or D

Side: helmet side of impact: Front or Back

Severity: a continuous measure of severity of impact transmitted through the helmet

Severity was measured 10 separate times for each combination of helmet model and side of impact. The

severity of impact was modeled as a function of the two categorical variables, model and side.

Two versions of an interaction plot are shown below. On the left the solid dots are the means of the 10

replications of each of the 8 combinations of helmet model and helmet side of impact. On the right, all the

data are shown, where the F and B symbols are values for Front and Back impacts, respectively. Many of the

individual data points are plotted by symbols that overlap.

Tables of mean Severites, by Model, by Side, and by Model and Side, follow:

A B C D

1070.30 1069.85 1116.65 1359.65

Back Front

1217.35 1090.88

Back Front

A 974.5 1166.1

B 1022.1 1117.6

C 1376.3 857.0

D 1496.5 1222.8

Two models of the data were examined. (Side:Model = interaction term between Side and Model)

Model 1: Severity = intercept + Cs*Side + Cm*Model

Df Sum Sq Mean Sq F value Pr(>F)

Side 1 319918 319918 6.936 0.0103

Model 3 1155478 385159 8.351 7.33e-05

Residuals 75 3459110 46121

Residual standard error: 214.8 on 75 degrees of freedom

R-squared: 0.299, Adjusted R-squared: 0.2616

Model 2: Severity = intercept + Cs*Side + Cm*Model + Csm*Side:Model

Df Sum Sq Mean Sq F value Pr(>F)

Side 1 319918 319918 12.61 0.000682

Model 3 1155478 385159 15.18 9.46e-08

Side:Model 3 1632156 544052 21.44 4.99e-10

Residuals 72 1826954 25374

Residual standard error: 159.3 on 72 degrees of freedom

R-squared: 0.6298, Adjusted R-squared: 0.5938

Residual plots for the two models are shown below.

a. What conclusions do you reach from the F tests for Model 2? Explain.

b. Compare the average percent accuracies of the two models. Refer, in part, to the residual standard errors.

c. Interpret the difference in the values of R-square for the two models.

d. Compare the residual plots for the two models, and explain which ones are better and why.

e. Summarize your conclusions from Model 2. Include statements about the main effects and interactions.

Refer to the interaction plot and the tables of means. Be specific about helmet models and helmet sides.

10

12

14

16

18

line

gain

ABCDEF

Diet

I

II

III

800

1000

1200

1400

1600

Helmet Model

Severity

ABCD

H$Side

Back

Front

800

1000

1200

1400

1600

Helmet Model

Severity

ABCD

H$Side

Back

Front

F

F

F

F

F

F

F

F

F

F

B

B

B

B

BB

B

B

B

B

F

F

F

F

F

F

F

F

F

F

B

B

B

B

B

B

B

B

B

B

F

F

F

F

F

F

F

F

F

F

B

B

B

B

B

B

B

B

B

B

F

F

F

F

F

F

F

F

F

F

B

B

B

B

B

B

B

B

B

B

10001100120013001400

-400

-200

0

200

400

600

Model 1

Fitted values

Residuals

Residuals vs Fitted

5

11

53

-2-1012

-2

-1

0

1

2

Model 1

Theoretical Quantiles

Standardized residuals

Normal Q-Q

5

11

53

900110013001500

-400

-200

0

200

400

Model 2

Fitted values

Residuals

Residuals vs Fitted

60

5

70

-2-1012

-2

-1

0

1

2

3

Model 2

Theoretical Quantiles

Standardized residuals

Normal Q-Q

60

5

70

x

Frequency

pemax

11013015017065758595

60

80

100

120

140

160

180

200

110

120

130

140

150

160

170

180

x

Frequency

height

x

Frequency

weight

20

30

40

50

60

70

60100140180

65

70

75

80

85

90

95

204060

x

Frequency

bmp

203040506070

60

80

100

120

140

160

180

200

weight

pemax

68

65

64

67

93

68

89

69

67

68

89

90

93

93

6670

70

92

69

7286

86

97

71

95

plotted numbers = bmp values