Academic Giant Statistic research method exam
Question 1
A study looked at the effects of anti-tobacco ads on smoking using data from a very large survey. It found that the presence of ads was associated with a lower prevalence of smoking by 0.2 percentage points (e.g., reducing smoking from 30.0% of the population to 29.8% of the population). The relationship was statistically significant.
Answer the following questions:
a. What do you think of the practical significance of this result? Explain. Make sure to briefly define practical significance.
b. Why do you think this result (i.e., such a small difference in percentages) is still statistically significant?
c. What kind of statistical test could have been used to determine statistical significance? (Hint: what kind of variable (qualitative or categorical) is the dependent variable? What kind of variable is the independent variable?)
You are interested in sampling fellow college students on how well they enjoy the restaurant selections on main campus. Main campus has approximately 16,000 students. You want your results to fall within a 3% margin of error (0.03 precision). Hint: N = (1/margin of error)^2
Answer the following questions:
a. How many individuals will you need to sample to meet your desired level of precision (0.03)?
b. Assume that only 30% of individuals you approach will actually agree to complete the survey. Now, how many will you need to sample to maintain that same level of precision (0.03)?
Question 3
A regression analysis considered predictors of getting a mammogram among women. The dependent variable was set up such that getting a mammogram in the past year = 1 and = 0 if otherwise.
|
|
Adjusted Odds Ratio |
Lower Bound |
Upper Bound |
|
Family history of breast cancer |
1.1 |
1.0 |
1.3 |
|
Type of insurance |
|
|
|
|
Private |
1.2 |
.6 |
2.3 |
|
No insurance |
.3 |
.1 |
.8 |
|
Medicare |
2.8 |
1.2 |
6.4 |
|
Pseudo R-squared |
.10 |
The first column presents the different variables used in the study. The second column presents the point estimate of the odds ratio, while the third and fourth columns respectively report the lower and upper bounds of each point estimate (95% confidence interval assumed).
Complete the following tasks:
a. Describe how well the model fits the data.
b. Women who had Medicaid insurance were also included in this study. Why was this category not included as a dummy variable?
c. How would you interpret the number 0.3 (no insurance) in this table?
d. With 95% confidence, what is the true odds ratio for getting a mammogram among women with Medicare insurance?
e. What are the odds of getting a mammogram among women with a family history of breast cancer?
Question 4
An on-line statistics course was developed at a major research university known for the quantitative skills of its students. The on-line course was open source and available to anyone on the internet. The developers evaluated the course using a randomized experiment at that university. Only students who volunteered to be in the study could be randomized and used as study subjects. Those subjects were randomized to either the on-line course or to the face-to-face course. Results showed that the on-line course was more effective in teaching statistics than the face-to-face course.
Answer the following:
a. What advantages (if any) does this study have over other possible studies?
b. What kinds of human artifact problems could the experiment have and what could be the consequences of such artifacts?
c. Discuss the generalizability of the study, particularly from the perspective of a community college considering using the on-line course.
Question 5
Researchers estimated the effect of a law requiring fast food restaurants in New York City to post the amount of calories in each meal by comparing the average amount of calories purchased by customers in NYC and a nearby city before and after the law.
Answer the following:
a. Discuss the strengths and weaknesses of this study.
b. How generalizable is it? Why?
c. Does it have good internal validity? Why or why not?
d. Was the treatment exogenous to the outcome? Why or why not?
Question 6
Describe two types of studies you might use to evaluate the effects of training on employee productivity. Compare and contrast the quality of the causal evidence of the two types of studies.
Question 7
A survey asked people to rate their feelings toward paying taxes on a ‘feeling thermometer’ ranging from 0 to 100, with 0 being ‘very cold’ or very unfavorable and 100 being ‘very warm’ or very favorable. In this model, age (in years) and income (in thousands of dollars) are quantitative variables, and homeowner is a dummy variable that assumes the value of ‘1’ if an individual owns a home and ‘0’ if otherwise.
|
Variable |
Beta |
|
Constant |
90 |
|
Income (thousands of dollars) |
-0.5 |
|
Homeowner |
-4 |
|
Age |
-0.5 |
Using the table, answer the following questions:
a. Write out the regression equation for this model. Clearly label your equation using appropriate units.
b. Interpret the value of the constant (even if it doesn’t make any logical sense). What is it telling you?
c. What would the expected rating be for someone who is 40 years old, a homeowner, and earns $60,000 per year?
d. What kind of variable is a ‘Homeowner’? Is it quantitative or dichotomous (categorical)?
e. What is the independent effect of being a homeowner on feelings toward taxes? In other words, interpret the beta coefficient for the variable ‘Homeowner.’
Question 8
A study found that poverty is related to drug abuse. Discuss how you would decide whether there is a causal relationship between these two variables.
Specifically, consider the following:
a. Identify and explain a plausible causal mechanism that could explain this relationship. What is your independent variable? What is your dependent variable? What is your intervening variable (or variables)?
b. Are there potential common causes or alternative explanations? Identify at least three of these factors and explain how (and in which direction) they are likely to bias the relationship between poverty and drug use.
c. Sketch out the relationships you’ve discussed in Parts A and B above using a path diagram. Clearly label your diagram (i.e., use arrows, signs, etc.).
d. What kind of study you would use to assess this relationship? Would you adopt an observational or experimental design? Why?
e. How would you measure your independent and dependent variables?
Question 9
The table on the next page is taken from “Maternal Employment and Teenage Childbearing” by Leonard Lopoo in the Journal of Policy Analysis and Management 24(1) 2004. We will only consider the OLS regression results. The unit of analysis is the mother. The variable “daughter has a birth at age 17 or 18” is a dummy variable with the meaning implied. “Mother’s education” refers to the years of education that a mother has. “Mother’s annual work hours” is the hours worked during the year.
Use the table to answer the following questions:
a. What is the dependent variable in this study?
b. Is whether Daughter has a birth at 17 or 18 statistically significant? If so, at which level (1%, 5%, or 10%)?
c. Holding constant all other independent variables, for every additional year of education a mother has, by how much are her annual work hours predicted to increase? Is this relationship statistically significant? If so, at which level (1%, 5%, or 10%)?
d. African American and Hispanic are dummy variables to proxy race/ethnicity. Are the effects of these variables statistically significant? What is the comparison group against which the effects of being African American or Hispanic compared?