I already have used visualizations in my data prep process, and the
more in-depth data analysis will require additional visualizations,
moving from the largely univariate analysis-focused visualizations
that informed data cleansing and other preparation decisions, to
using visualizations to perform multivariate analysis. Scatterplots
will be used to look for visual clues about the relationships between
the variables of interest (the sets of variables that describe:
Medicaid expansion status, health insurance coverage, and Medicaid
enrollment) and the outcome variables (health outcomes, education
outcomes, and employment/income outcomes. R offers a relatively
simple function for creating a scatterplot matrix, called pairs, that
can operate on the entire dataframe to generate pairwise
scatterplots. These visualizations may offer additional insight into
either variables that can be removed (to cut down on
dimensionality) or to at least prioritize the relationships to test in
the models.
The initial set of predictive models that will be used are association
analysis (using the Apriori algorithm in R), the conditional inference
tree algorithm (also in R), and Ensemble models in SAS Enterprise
Miner. In all cases, the models will be run on a randomly generated
subset of the data (training data) and then validated on a separate
subset (test/validation data set).
The Apriori method will require discretization of nearly all of the
variables, which are generally continuous, numeric variables. These
variables will be divided into groups using the “discretize” function
in R, primarily with equal interval or cluster methods, given the
skewness of the data, the size of the dataset, and the difficulty of
identifying breakpoints manually. The Apriori method for generating
association rules is applicable to this project because it does not
comprehensively model every combination of every level of the
discretized variables to identify strong associations. The method
identifies only evaluates those relationships that meet the user-
determined minimum support level to then output a confidence
level, that again must meet a user-determined threshold. The
strength of the correlation is captured in the lift value: greater than
or less than 1 means that there is some dependency (positive or
negative, respectively) between the two items, while a lift value of 1
means the two events are independent. Using lift helps mitigate the
potentially misleading/overstating of relationships that are simply
frequent, but not statistically correlated. These models should offer
further insight (building on the scatterplots) to identify relationships
that not only appear visually, but those which may be hidden by
relatively fewer cases or those which may be overstated with a
simple one-to-one mapping.
Supervised classification models using conditional inference decision
trees should refine class memberships that may show up less clearly
in Apriori results. The conditional inference approach operates in a
strict if/then structure that resolves the limitations of Apriori
models, which allow for overlapping class memberships that can
muddy insights and create ambiguity in predictions of future case
outcomes. This model will not require discretization of the variables,
because conditional inference trees can take both discrete and
continuous variables. This type of approach has the added benefit of
being agnostic to skewed or normal distributions, which is relevant
given the skewness of several of the posited outcomes of interest,
and of participation in Medicaid expansion itself. This approach is
also valuable because it does not require the explanatory variables
to be independent. For example, this project does not evaluate
whether the decision to expand Medicaid or not is predicated on a
particular income level or set of health outcomes—in other words,
the causality of the relationship may be the reverse of what is being
tested. This issue makes conditional inference particularly useful.
Because conditional inference trees can handle collinear models and
select the best predictor, they also offer the benefit of being able to
take all or most of the variables identified the previous section has
having some degree of collinearity, rather than requiring the
researcher to make the choice or test multiple combinations.
Finally, Ensemble Models generated in SAS Enterprise Miner
will also be used. The decision trees and boosted trees in the
random forest, gradient boosting, and bagging models are, like
conditional inference trees, typically robust to asymmetric or
unbalanced datasets. SAS Enterprise Miner offers the advantage
over R of being able to relatively quickly generate model results,
which allows the user to try different mechanisms to tweak (and
improve) the results. For example, the models can be tuned with
adjusted cutoff thresholds or cost adjustments to improve predictive
strength. The power of the SAS Enterprise Miner application is that
it enables the identification of optimal thresholds that, in turn,
increase the robustness of the model’s predictive strength across
multiple evaluative dimensions: recall, specificity, and precision (in
addition to general accuracy). The application also enables
comparisons of more and less complex models, to ultimately arrive
at a predictive model that is as simple as possible with as robust an
outcome as possible. The benefits of this tool are directly relevant
to the project at hand because of the size and complexity of the
dataset, as described previously; the apparent collinearity among
the variables of interest; and the need to be able to tell as clear an
analytic story as possible to inform health care policy.
This last item is particularly important and particularly well-served
by the ability available through SAS EM to iterate through dozens of
versions of models if necessary, to arrive at the most
understandable, explainable, and actionable implications. the
ultimate goal of this project is to determine what, if any, relationship
exists between the policy decision and implementation of Medicaid
expansion and positive societal results, in order to inform future
policy choices on health care.