Need one who is good at Economic data analysis and using Stata(A Software) (Undergraduate)

profileSharonpapapa
Thirddeliverableforfinalproject.pdf

Eco311 Project, Final Deliverable, Due Thursday 12/7 at 5 p.m.

Note. There will be a 20 point penalty for each day (or part thereof) that the assignment is late.

This deliverable should include individual results for your secondary topic and a discussion of your conclusions. For your secondary topic, provide the following steps in your analysis. Your grade on the final project will be based on the content of your analysis, but also whether you are able to generate a document that is professional in its appearance and content. Your intended audience is someone who would have the knowledge that is expected of someone who has mastered the content in Economics 311.

1. A title page that includes a descriptive title, the author’s name, a date, and a subtitle indicating that it is a deliverable to Prof. William Even for Eco311 in Fall 2017.

2. Provide an introductory section that with the following: a. A review of the main findings of your first two deliverables. This shouldn’t be a discussion

of the details of your data and variables, but rather a simple discussion of the key findings from your regression. For example, “In our earlier deliverables, we learned that there are several important determinants of whether a person over age 55 is employed. First, …..).

b. A discussion of what is new in this deliverable. For example, “In this deliverable, I will extend our earlier analysis to examine differences in employment rates between men and women and attempt to understand why women have lower employment rates than men. I will also investigate whether the aging has a differential effect on employment rates. I find that ….”.

3. Background. This should include a discussion of the main hypotheses you are testing, the data you

will use, and a brief summary of your major findings.

4. Provide a table of summary statistics for the dependent and control variables for the two groups created by your secondary variable (i.e. by sex, race, location, marital status, or year). If your secondary variable is year, be sure to convert all variables measured in dollars into current dollars using the CPI. Included in the summary statistics should be a t-statistic that tests the null hypothesis that the means are equal for the two groups and asterisks indicating whether the difference in means is significantly different from zero at the .10 (*), .05(**) or .01(***). You may either use the stata command ttest, or a regression of the relevant dependent variable on the group dummy. For example,

ttest incss, by(female)

or

reg incss female

The means, test statistic, and p-value for the null hypothesis should be included in a single professional table. Be sure your variable names are self-explanatory, that you have an appropriate title, and that your footnotes clearly define the sample for your analysis.

NOTE: You should try to get your table to fit on a single page. This is much simpler if you start the table at the top of a new page. Your table should not be split across pages unless it is impossible to fit it all on a single page.

Table 1. Summary Statistics by Sex for Analysis of Determinants of Number of Children. (Sample size =446,480)a

Variable Mean for Single People

Mean for Married People

t-statistics for equality (p-value in parentheses)

Number of childrenb 0.57 1.40 233.31 (0.000)

Age 28.33 30.17 155.98 (0.000)

Etc… for all the other controls

a Sample is drawn from 2016 American Community Survey and restricted to people aged 21-35. Excludes people living in group quarters. b Number of children represents number of own children living in same household.

5. Provide a regression analysis that allows you to test whether the between group difference in the

dependent variable is “explained” by differences in the control variables. After performing your regression analysis, discuss at least two sets of key variables (e.g. a group of education dummies would count as one set of key variables) in your regression and indicate whether they help explain why there is a gap in the dependent variable between the two groups. A sample table is provided below. The first column gives the raw difference in the dependent variable across groups (married in this case). Notice that this exactly matches the difference in the dependent variable provided in table 1. The second specification is from a regression of number of children on all of your control variables. The third specification is included for the final question. Keep in mind that this table has only age and its square as a control variable. Your table should have all or most of your control variables. If all of the control variables aren’t listed, provide a list of other controls that were included in a footnote to your table.

Table 2. Regression Analysis of Determinants of Number of Children.a Specification 1 Specification 2 Specification 3 Married 0.829*** 0.701*** 0.462**

(233.3) (197.6) (2.480) Age

-0.0733*** 0.00143

(-11.51) (0.148) Age2

0.00253*** 0.000871***

(22.95) (5.087) Education (omitted group has less than a high school degree)

High school Degree --- --- Some College --- --- College Degree

---

--- Married*Age --- -0.0219*

--- (-1.668) Married*Age2 --- 0.00101*** --- (4.430) Constant 0.574*** 0.573*** -0.180

(204.2) (6.345) (-1.335)

Observations 446,880 446,880 446,880 R-squared 0.109 0.161 0.164 F-test/p-valuec --- a Sample is drawn from 2014 American Community Survey and restricted to people aged 21-35. Excludes people living in group quarters. b t-statistics are in parentheses and are calculated using robust standard errors. *** indicates p-value below .01; ** below 0.05, and * below 0.1. c F-test and associated p-value are for null hypothesis that …..

You should discuss your results and refer to the relevant table and/or specification. An example of such writing is below: Based on a White/Breusch-Pagan test using the residuals from specification (2), it was determined that the regression model had heteroscedasticity. As a consequence, all of the t-statistics in table 2 are based on robust standard errors. Based on a comparison of the coefficients on the married dummy variable in specifications 1 and 2 of table 2, we can see that the control variables we added account for married people having .128 more children than single people). One explanation for this is that married people are, on average, 1.84 years older than single people (see table 1). Moreover, over the 21-35 year old age range in the sample, age has a positive marginal effect on the number of children that is increasing in the number of children

based on estimates of the quadratic in age in the regression.1 Consequently, an important reason that married people have more children than single people is that they are, on average, older. You should discuss at least 2 “important” variables (or sets of variables) and indicate whether they help explain why the dependent variable differs across your two groups.

6. Provide a test of the null hypothesis that the effect of at least one key variable differs across your two groups. Describe the results of your test and explain the implications for how the variable has differential effects on the dependent variable for the two groups you are examining.

To provide a test that a variable has a differential impact across groups, use interaction terms. For example, if you want to test that age has a differential effect across married and single people, create interactions between married and age, age2. I have included these in specification 3 of table 2. Be sure to discuss the statistical and economic significance of the interaction terms. If you have multiple interaction terms, perform a test for the joint significance of them all and include this in your table (as illustrated in table 2). For example,

In specification (3) of table 2, interactions between married and age and its square are added to the regression. Given the quadratic in age, a simple comparison of the marginal effect of age on children for married and single people is not simple. The marginal effect for single people is given by

𝑑𝑑𝑑𝑑𝑑𝑑ℎ𝑖𝑖𝑖𝑖𝑑𝑑 𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑

= .00143 + .000871 ∗ 𝑑𝑑𝑑𝑑𝑑𝑑

The marginal effect for married people is

𝑑𝑑𝑑𝑑𝑑𝑑ℎ𝑖𝑖𝑖𝑖𝑑𝑑 𝑑𝑑𝑑𝑑𝑑𝑑𝑑𝑑

= −.02047 + .001881 ∗ 𝑑𝑑𝑑𝑑𝑑𝑑

A comparison of these marginal effects reveals that the marginal effect of age is greater for married than single people for ages 22-35. The marginal effect of age is slightly larger for singles than married at age 21.

An f-test of the null hypothesis that the coefficients on the interaction terms is zero is provided at the bottom of specification (3) in table 2. The results indicate that the null ……

1 In fact, the quadratic in age implies that the marginal effect of age is negative until age 14.5 and is positive for all ages beyond 14.5