Background: My datasets turned out to be a bit smaller than I initially
thought they would be. The American Community Survey (ACS) has
two kinds of datasets--5-year averages, and 1-year data. When I was
initially looking at the sets, I was using the 5-year datasets. These are
not appropriate, though, because each year's data represents another
point in a running average. My data would be too interrelated, I think,
to provide a meaningful result. Going to the 1-year datasets, there are
far fewer counties per year (just over 800), and there are
only
10 years
of data available, and the datasets have about 8300 observations. The
ACS did not capture health insurance coverage status or Medicaid
enrollment by county before 2010, and has not published its usual
tables for 2020 or 2021 because of COVID. I would like to figure out a
way to get 10,000 records (or at least get closer to 10,000).
I think I have a couple of options, but would appreciate others'
thoughts on the pros/cons of the approaches:
1. ACS has state-level experimental health insurance and
demographic information for 2020 that I could use to
extrapolate county-level data, albeit very roughly. That would
get me another 800-ish records.
2. Because I'm interested in how things changed with Medicaid
expansion and the implementation of the Affordable Care Act, I
could oversample the 2010 data to account for an imbalanced
target--for practical purposes, the ACA took effect in 2011, so
the dataset only has one pre-change year.
3. I could do both--that would get me over 10K records.
4. Something else?
What do you think? Do you need more info to make a suggestion (A
fuller description of one of my datasets, on health insurance coverage
by race, follows below)? Thanks for your ideas!
Resha
Health Insurance by Race (American Community 1-Year Surveys,
2010-2019, via Census Bureau)
• 8306 observations of 65 variables
• 6 race categories X 3 age categories X two health insurance
status:
o American Indian (AmerInd), Asian, Black, Hispanic/Latino
(His/Lat), Native Hawaiian/Pacific Islander (NHPI), White
o Youth, Adult, Senior
o Insured, Not Insured
• County-level records: 10 years of data, 843 distinct counties (not
all counties have 10 years of records)
• Quality: High quality, reliable data, drawn from Census Bureau’s
annual American Community Statistics (estimated, not actual
counts).
• Data Challenges: disparity in racial distribution, high frequency
of “null” for smaller racial groups. Nulls are comprehensive-
either all 65 variables in the observation are null, or none are. For
2018:
o NHPI: total pop: 612,346; 8041 out of 8306 observations
are “null”
o American Indians: total pop: 2.7M; 6847 out of 8306
observations are “null”
o Asian: total pop: 18.3M; 4676 out of 8306 observations are
“null”
o Black: total pop: 40.4M; 2596 out of 8306 observations are
“null”
o Hispanic/Latino: total pop: 59.0M; 1860 out of 8306
observations are “null”
o White: total pop: 194.8M; 103 out of 8306 observations
are “null”
• Narrow analysis to Black, Hispanic/Latino, White, Other
o Other total pop:21.7M; 4226 out of 8306—about half—are
zero (null)
▪ Even aggregated, this group possibly has too many
missing values to be useful.