I already have used visualizations in my data prep process, and
the more in-depth data analysis will require additional
visualizations, moving from the largely univariate analysis-focused
visualizations that informed data cleansing and other preparation
decisions, to using visualizations to perform multivariate analysis.
Scatterplots will be used to look for visual clues about the
relationships between the variables of interest (the sets of
variables that describe: Medicaid expansion status, health
insurance coverage, and Medicaid enrollment) and the outcome
variables (health outcomes, education outcomes, and
employment/income outcomes. R offers a relatively simple
function for creating a scatterplot matrix, called pairs, that can
operate on the entire dataframe to generate pairwise scatterplots.
These visualizations may offer additional insight into either
variables that can be removed (to cut down on dimensionality) or
to at least prioritize the relationships to test in the models.
The initial set of predictive models that will be used are
association analysis (using the Apriori algorithm in R), the
conditional inference tree algorithm (also in R), and Ensemble
models in SAS Enterprise Miner. In all cases, the models will be
run on a randomly generated subset of the data (training data)
and then validated on a separate subset (test/validation data set).
The Apriori method will require discretization of nearly all of the
variables, which are generally continuous, numeric variables. These
variables will be divided into groups using the “discretize” function
in R, primarily with equal interval or cluster methods, given the
skewness of the data, the size of the dataset, and the difficulty of
identifying breakpoints manually. The Apriori method for
generating association rules is applicable to this project because it
does not comprehensively model every combination of every level
of the discretized variables to identify strong associations. The
method identifies only evaluates those relationships that meet the
user-determined minimum support level to then output a
confidence level, that again must meet a user-determined
threshold. The strength of the correlation is captured in the lift
value: greater than or less than 1 means that there is some
dependency (positive or negative, respectively) between the two
items, while a lift value of 1 means the two events are
independent. Using lift helps mitigate the potentially
misleading/overstating of relationships that are simply frequent,
but not statistically correlated. These models should offer further
insight (building on the scatterplots) to identify relationships that
not only appear visually, but those which may be hidden by
relatively fewer cases or those which may be overstated with a
simple one-to-one mapping.
Supervised classification models using conditional inference
decision trees should refine class memberships that may show up
less clearly in Apriori results. The conditional inference approach
operates in a strict if/then structure that resolves the limitations
of Apriori models, which allow for overlapping class memberships
that can muddy insights and create ambiguity in predictions of
future case outcomes. This model will not require discretization of
the variables, because conditional inference trees can take both
discrete and continuous variables. This type of approach has the
added benefit of being agnostic to skewed or normal distributions,
which is relevant given the skewness of several of the posited
outcomes of interest, and of participation in Medicaid expansion
itself. This approach is also valuable because it does not require
the explanatory variables to be independent. For example, this
project does not evaluate whether the decision to expand
Medicaid or not is predicated on a particular income level or set
of health outcomes—in other words, the causality of the
relationship may be the reverse of what is being tested. This issue
makes conditional inference particularly useful. Because
conditional inference trees can handle collinear models and select
the best predictor, they also offer the benefit of being able to
take all or most of the variables identified the previous section
has having some degree of collinearity, rather than requiring the
researcher to make the choice or test multiple combinations.
e e e e e e e e e e Finally, Ensemble Models generated in SAS Enterprise
Miner will also be used. The decision trees and boosted trees in
the random forest, gradient boosting, and bagging models are, like
conditional inference trees, typically robust to asymmetric or
unbalanced datasets. SAS Enterprise Miner offers the advantage
over R of being able to relatively quickly generate model results,
which allows the user to try different mechanisms to tweak (and
improve) the results. For example, the models can be tuned with
adjusted cutoff thresholds or cost adjustments to improve
predictive strength. The power of the SAS Enterprise Miner
application is that it enables the identification of optimal
thresholds that, in turn, increase the robustness of the model’s
predictive strength across multiple evaluative dimensions: recall,
specificity, and precision (in addition to general accuracy). The
application also enables comparisons of more and less complex
models, to ultimately arrive at a predictive model that is as simple
as possible with as robust an outcome as possible. The benefits of
this tool are directly relevant to the project at hand because of
the size and complexity of the dataset, as described previously;
the apparent collinearity among the variables of interest; and the
need to be able to tell as clear an analytic story as possible to
inform health care policy.
This last item is particularly important and particularly well-served
by the ability available through SAS EM to iterate through dozens
of versions of models if necessary, to arrive at the most
understandable, explainable, and actionable implications. the
ultimate goal of this project is to determine what, if any,
relationship exists between the policy decision and
implementation of Medicaid expansion and positive societal
results, in order to inform future policy choices on health care.