Economic Assignment Report
Part 4. Specification Analysis
Items in round brackets are optional depending on the question. Economics 381 deals with items in square
brackets. You can leave these sections until you get to that course.
(1) Introduction
(a) Research question
(i) What is the main connection you are working on? We are interested in estimating part
of the money value of a local pollution externality. We will investigate the relationship
residential housing prices with exposure to airborn pollution.
(A) What are the variables? The two main variables are house prices and the level of
airborn automobile pollution.
(B) What is the direction of causality? We expect changes in pollution levels to cause
changes in house prices all else equal.
(ii) What parameters are you trying to measure? We want to measure the relative sensitivity
of house prices to changes pollution levels; so we’ll try to measure the elasticity of prices
with respect to pollution levels.
(A) In what units will the parameters be measured? The units are the percent change
in house prices per one percent increase in airborne pollution.
(b) (Policy analysis)
(i) How does answering the question contribute to policy making? In the face of scarce
resources, local pollution regulation must compete with demand for other public goods
and services. To decide the merit of any proposal to limit local automobile pollution we
require some measure of the expected benefit of the regulation.
(ii) How important is the policy involved?
(2) [Literature Survey]
The idea for this study comes from a 1978 Journal of Envirnonmental Economics and Management
paper “Hedonic housing prices and the demand for clean air” [1].
(1) The Model
(a) What other connections besides those identified in 1(a)i will you include in the analysis? House
prices depend on a great many other factors. Those that we should think about including are
the features of a house that are likely to correlate with exposure to automobile pollution. These 1
2
are likely to be associated with proximity to major road ways and areas of concentration of auto
traffic. As an example, it’s been established by urban economists and geographers that house
prices tend to be higher the closer the neighborhood is to employment nodes [cite needed]. At
the same time, we have to allow for the possibility that neighborhood automobile pollution levels
may be higher near employment nodes because of higher traffic densities. Without accounting
for these relationships we may produce attenuated estimates of the elasticity of house prices
with respect to pollution levels.
(b) The Path Diagram: Figure 0.1 shows the path diagram for the base model.
(i) What are the causal connections? The main causal connections run from pollution to
price, distance to employment to pollution, distance from employment to price, and crime
to price.
(ii) What are connections that represent covariance? The covariances are between crime and
distance to employment, and crime and pollution.
(iii) What are the signs of the relationships represented by all the paths?
(A) We expect the causal connection between pollution and house prices to be negative:
the general idea is that all negative externalities affecting a property will tend to
depress its market value relative to properties not so exposed.
(B) We hypothesize a negative impact of distance to employment on pollution levels; we
expect proximity to employment nodes entails higher levels of exposure to automobile
traffic, and hence higher levels of airborne pollution.
(C) We expect a positive (direct) impact of distance to employment nodes on house
prices; distance is a proxy for travel costs to work, one of the costs of home occupancy.
In competitive markets, we expect house prices to be higher, ceteris paribus, the
closer we are to employment nodes.
(D) We expect a negative impact of crime rates on house prices; crime is an example of a
negative externality. Differences in criminal activity are reflected in insurance costs
of home occupancy. As well, differences in crime rates are associated with different
level of apprehension about personal safety.
(E) Distance to employment and crime rates are likely to covary in the presence of a
common cause, opportunities for criminal activity. We epxect more opportunities
lead to higher crime rates. As well, we expect that opportunities for crimes like
3
break and enter, mugging, arson, fraud, and the like to be concentrated in and
around employment nodes. Consequently, we expect the covariance between crime
and distance to employment nodes to be negative.
(F) Finally, because we expect distance to employment nodes to be a cause of pollution
levels, and pollution and crime covary, we expect pollution and crime to covary
positively.
(c) What mathematical specification will you use as your base model (i.e. before any respecification
based on analysis of the first regression results)? Based on experience and the literature, we
choose the following linear in parameters specification for the population regression function.
The discussion above of the signs of relationships implies β1 < 0, β2 < 0, and β3 < 0.
ln(price)i = β0 + β1ln(pollutioni) + β2ln(distancei) + β3crimei + ui
price
u
pollution
Distance to
employment crime
Figure 0.1. Path Diagram for Base Model
(d)
(2) The Data
(a) What is the source of the data? The data for this study is associated with the Harrison
and Rubenfeld paper [1]. The units of analysis are 506 census tracts in the Boston Standard
Metropolitan Statistical Area. The authors employed several sources of information to construct
the dataset.
(i) The house price measure was extracted from the 1970 Census.
4
(ii) Information about pollution is based on estimates from a simulation model (Transporta-
tion and Air Shed Simulation - TASSIM).
(iii) Crime rates were obtained from the Federal Bureau of Investigation.
(iv) Distance to employment nodes were retrieved from a 1973 Harvard doctoral disserta-
tion[need citation].
(b) [How are the variables in the model operationalized?]
(i) House prices are measured by the median house price (nominal 1970 $) reported in the
1970 Census for each of the 506 Census Tracts.
(ii) The TASSIM model generates surface level concentrations of nitogen oxides in parts per
million (ppm) conditional on the emissions characteristics of the 1970 automobile fleet in
the Boston SMSA. These estimates are then checked (calibrated) against actual surface
level pollutant data from 19 monitoring stations.
(iii) It’s not clear in the original paper whether they used the Uniform Crime Report produced
by the FBI and the DOJ. Recent work on the UCR suggests that local information on
crime is highly unreliable.
(iv) The distance to employment nodes is measured in its logarithm. So, regression results
using the variable need to be interpreted with this transformation in mind.
(c) What are the univariate properties of the data? [see /Users/Terry/Desktop/Base_model_results.smcl]
(i) Central tendency
(ii) Spread
(iii) Shape
(iv) Pattern and exceptions
(d) What are the multivariate properties of the data?
(i) What are the partial correlations for the paths in your model under 1b? Figure 0.2 shows
the partial correlations for the sample of 506 Census Tracts. The null hypothesis that
the partial correlation between the logarithm of the concentration of nitrogen oxides and
crime rates is zero cannot be rejected (p-value = .4148). Otherwise, the null hypotheses
can be rejected at the 5 percent α-level or less. The signs of the correlations aggree
with the discussion in 1(b)iii. It’ apparent that we must account for the contributions
5
of distance and crime rates if we are to obtain a ceteris paribus estimate of the effect of
pollution on prices.
log(price)
u
Log(nox)
log(distance) crime
-.3557
-.4100
-.1846 -.7733
.0361
-.1304
Figure 0.2. Partial Correlations
(ii) (Are there any issues you need to deal with as a result of the analysis for outliers in
2(c)iv?) The presence of extreme values for the crime rate variable should be addressed.
We set up an indicator variable to identify those Census Tracts with crime rates greater
than 14.466 per capita. We can modify the population regression function by adding this
variable to the model as in
log (pricei) = β0 + β1ln(pollutioni) + β2ln(distancei) + β3crimei + β4extremei + ui
The parameter β4 allows us to determine if there is an extra discount to housing prices in areas
experiencing unusually high crime rates.
(3) The results
(a) What are the results of the estimation of your base model? Table 1 shows the OLS estimates
of the base model, a model that includes the indicator variable that indentifies Census Tracts
with extreme values on the crime rate variable, and for comparison a model that drops the 30
Census Tracts so identified.
(i) Estimates
6
Base Model Model 2 Model 3 ln(nox) −1.113
(0.111) −1.162 (0.111)
−0.911 (0.116)
ln(distance) −0.048 (0.010)
−0.051 (0.010)
−0.047 (0.010)
crime −0.018 (0.003)
−0.011 (0.004)
−0.031 (0.006)
extreme – −0.328 (0.100)
–
constant 12.075 12.160 11.763 R2 .4003 .4155 .2922 N 506 506 476
Table 1. OLS Estimates
(A) For each estimate, the results of the hypothesis test, and an interpretation. Robust
standard errors are reported in brackets below estimates. We can reject the null
hypothesis that the elasticity of house prices with respect to concentration of nitrogen
oxides is zero (p=value < .001) for all three models. More importantly, we fail to
reject the null hypothesis β1 = −1 for all three models. That is we cannot rule out
the possibility that the population elasticity is a one percent decline in price for a one
percent increase in nitogen oxide levels. The elasticity of housing price with respect
to distance to employment nodes is about -.05 percent per one percent increase in
distance. We reject the null hypothesis that these elasticities are zero in all three
models (p=value < .001). An increase of one crime per capita is associated with
a 1.1 percent decline in price in the Base Model (p=value < .001), a decline of 1.8
percent in Model 2 (p-value < .001), and a decline of 3.1 percent in Model 3 (p-value
= .005). The addition of the indicator to identify Census Tracts with extreme crime
rates improves the Base Model; the estimate indicates that there is an additonal
discount of 32.8 percent for Census Tracts with exceptionally high crime rates.
(B) Does the result match the prediction of the direction of the relationship in 2(c)iv?
The signs of the estimates agree with the predicted relationships discussed in 1(b)iii.
(C) How well does the sample prediction function explain the dependent variable? The
Base Model and Model 2 explain about 40 percent of the variance of the logarithm
of median prices. Inclusion of the indicator in Model 2 improves the overall fit (F =
10.77, p-value = .0011). Model 3 explains about 30 percent of the variance of the
7
logarithm of prices for those Census Tracts which do not exhibit extreme crime rate
values.
(ii) Assessment of the specification
(A) Do the residuals appear to match the assumption of independence from any of the
included variables? We plotted the predicted log(price) values for Model 2 [use Stata
cmnd predict mod2_fit, xb to obtain fitted values] against the residuals for Model
2. There is no apparent relationship between the predictions and the residuals that
would lead us to consider respecifying the regression function to include nonlinear
components. Plots of the reisuals against individual independent variables confirm
this conclusion.
(B) Can we leave out any of the independent variables based on tests of exclusion re-
strictions? This is not an issue since we reject all of the individual null hypotheses
for the four parameters of the model.
(C) If you’ve decided to, what modifications did you make to the base model (e.g adding
or removing variables, transforming the dependent or independent variables, etc.)
The dataset contains other variables that are probably associated with the median
house proces in the Census Tracts. These are listed below. The risk of leaving any
of these out is that they are associated with automobile pollution and housing prices
and hence sources of omitted variable bias.
• The average number of rooms (rooms) in house in the Census Tract.
• The average student teacher ratio (stratio) in schools in the Tract.
• The percent of the Tract population in the lower socio-economic status category
(lowstat).
• A measure access to Boston’s ring road system (radial).
(D) Repeat 3(a)i and 3(a)ii.
The OLS estimated sample regression function, when we augment Model 2 with these
additional variables, is
l̂pricei = 11.522 − .562 (.095)
lnoxi − .051 (.008)
disti − .010 (.003)
crimei − .042 (.091)
extremei
+ .109 (.026)
roomsi − .041 (.004)
stratioi − .028 (.004)
lowstati + .004 (.002)
radiali
8
The R-square is .7568. What is most noticeable is the drop in the elasticity of house prices with
respect to automobile pollution. While we continue to resject the null hypothesis that house
price and exposure to airborne pollution are not associated [H0 : β1 = 0] the estimate of the
elasticity is now .562 percent decrease in house price per one percent increase in nitorgen oxide
concentration. [This is a good example of omitted variable bias.]
(4) Conclusions
(a) What answer(s) to the main research question(s) do the results of your work provide? Based
on the sample we have, we estimate the elasticity of housing prices with respect to increasing
nitrogen oxide pollution to be -.562 percent per one percent increase in pollution.
(b) (What are the policy implications of these answers?) If our results are valid, they provide us a
way to monetize the reduction in welfare arising from higher levels of air pollution. The ceteris
paribus decline in house prices as pollution increases measures the reduction, arising from the
negative externality, in the flow of welfare generating services that a house provides. So if we
can predict, ceteris paribus, that a one perncet increase in airborne pollution reduces the price
of a house by $2,000, this becomes the basis for assignng annual costs to the homeowner of the
pollution externality.
(c) [What limitations are your answers subject to?]
References
[1] David Harrison Jr. and Daniel L Rubinfeld. Hedonic housing prices and the demand for clean air. Journal of
Environmental Economics and Management, 5(1):81–102, 3 1978.