I already have used visualizations in my data prep process, and the
more in-depth data analysis will require additional visualizations,
moving from the largely univariate analysis-focused visualizations
that informed data cleansing and other preparation decisions, to
using visualizations to perform multivariate analysis. Scatterplots will
be used to look for visual clues about the relationships between the
variables of interest (the sets of variables that describe: Medicaid
expansion status, health insurance coverage, and Medicaid
enrollment) and the outcome variables (health outcomes, education
outcomes, and employment/income outcomes. R offers a relatively
simple function for creating a scatterplot matrix, called pairs, that can
operate on the entire dataframe to generate pairwise scatterplots.
These visualizations may offer additional insight into either variables
that can be removed (to cut down on dimensionality) or to at least
prioritize the relationships to test in the models.
The initial set of predictive models that will be used are association
analysis (using the Apriori algorithm in R), the conditional inference
tree algorithm (also in R), and Ensemble models in SAS Enterprise
Miner. In all cases, the models will be run on a randomly generated
subset of the data (training data) and then validated on a separate
subset (test/validation data set).
The Apriori method will require discretization of nearly all of the
variables, which are generally continuous, numeric variables. These
variables will be divided into groups using the “discretize” function in
R, primarily with equal interval or cluster methods, given the
skewness of the data, the size of the dataset, and the difficulty of
identifying breakpoints manually. The Apriori method for generating
association rules is applicable to this project because it does not
comprehensively model every combination of every level of the
discretized variables to identify strong associations. The method
identifies only evaluates those relationships that meet the user-
determined minimum support level to then output a confidence level,
that again must meet a user-determined threshold. The strength of
the correlation is captured in the lift value: greater than or less than 1
means that there is some dependency (positive or negative,
respectively) between the two items, while a lift value of 1 means the
two events are independent. Using lift helps mitigate the potentially
misleading/overstating of relationships that are simply frequent, but
not statistically correlated. These models should offer further insight
(building on the scatterplots) to identify relationships that not only
appear visually, but those which may be hidden by relatively fewer
cases or those which may be overstated with a simple one-to-one
mapping.
Supervised classification models using conditional inference decision
trees should refine class memberships that may show up less clearly
in Apriori results. The conditional inference approach operates in a
strict if/then structure that resolves the limitations of Apriori models,
which allow for overlapping class memberships that can muddy
insights and create ambiguity in predictions of future case outcomes.
This model will not require discretization of the variables, because
conditional inference trees can take both discrete and continuous
variables. This type of approach has the added benefit of being
agnostic to skewed or normal distributions, which is relevant given
the skewness of several of the posited outcomes of interest, and of
participation in Medicaid expansion itself. This approach is also
valuable because it does not require the explanatory variables to be
independent. For example, this project does not evaluate whether the
decision to expand Medicaid or not is predicated on a particular
income level or set of health outcomes—in other words, the causality
of the relationship may be the reverse of what is being tested. This
issue makes conditional inference particularly useful. Because
conditional inference trees can handle collinear models and select the
best predictor, they also offer the benefit of being able to take all or
most of the variables identified the previous section has having some
degree of collinearity, rather than requiring the researcher to make
the choice or test multiple combinations.
f f f f f f f f f f Finally, Ensemble Models generated in SAS Enterprise Miner
will also be used. The decision trees and boosted trees in the random
forest, gradient boosting, and bagging models are, like conditional
inference trees, typically robust to asymmetric or unbalanced
datasets. SAS Enterprise Miner offers the advantage over R of being
able to relatively quickly generate model results, which allows the
user to try different mechanisms to tweak (and improve) the results.
For example, the models can be tuned with adjusted cutoff
thresholds or cost adjustments to improve predictive strength. The
power of the SAS Enterprise Miner application is that it enables the
identification of optimal thresholds that, in turn, increase the
robustness of the model’s predictive strength across multiple
evaluative dimensions: recall, specificity, and precision (in addition to
general accuracy). The application also enables comparisons of more
and less complex models, to ultimately arrive at a predictive model
that is as simple as possible with as robust an outcome as possible.
The benefits of this tool are directly relevant to the project at hand
because of the size and complexity of the dataset, as described
previously; the apparent collinearity among the variables of interest;
and the need to be able to tell as clear an analytic story as possible to
inform health care policy.
This last item is particularly important and particularly well-served by
the ability available through SAS EM to iterate through dozens of
versions of models if necessary, to arrive at the most understandable,
explainable, and actionable implications. the ultimate goal of this
project is to determine what, if any, relationship exists between the
policy decision and implementation of Medicaid expansion and
positive societal results, in order to inform future policy choices on
health care.