Data preparation and cleaning for my datasets entailed merging
the reformatting the datasets and combining them. To prepare the
CHR&R dataset for data merge, the most relevant variables were
selected for analysis and all other remaining variables were
dropped. Because the data is organized by county and the state
level data is needed, the county level data was aggregated by
state. There was not much that needed to be done to prepare the
SAMHSA dataset for merge, since it was already properly
formatted. The two datasets were merged using R which
facilitates joining two data frames. In this case, a left join was
used to join the SAMHSA (on the left) and CHR&R (on the right)
datasets based on state. This means that the final output table
contains all the rows and columns of data in the SAMHSA dataset
and the corresponding matching column data in the CHR&R
dataset. Data cleaning including handling missing data and
removing outliers, all of which was done using R.
The analytical tools used in my project are Tableau, R and Python.
Tableau is a great user-friendly tool for creating visualizations of
geographic data in the form of data maps. The drawback to using Tableau
is that it does not offer users many customization options. Overall,
however, Tableau’s click-and-drag functionality make it a good choice for
creating quick and informative maps. Mapped visualizations are important
in this exploration because the mental health landscape will be compared
across states. Initially, it will be important to explore various health and
environmental parameters by state to gain a better understanding of the
data.
R is also great tool choice for visualizations and is used in this project for
data preparation, cleaning, and initial data exploration. R was specifically
built for statistical analysis and is particularly well known for generating
some of the most attractive data visualizations (IBM, 2021). In R it is easy
to generate statistical summaries in one line of code by using simple
commands like str() and summary(). It is also easy to create visualizations
for univariate and multivariate analysis of each field in the dataset. It is
also capable of handling large datasets.
e e e e e e e e Python is the tool of choice for the predictive model development
since it is a general-purpose tool widely used for big data analysis and
machine learning. As of 2022, it is considered the top programming
languages by popularity (Luna, 2022). Although R is also useful for
modeling large datasets, Python is more intuitive to use since the code is
more readable. Since Python is open source, it will also be easy to find
solutions to any potential issues through online forums, analytics libraries,
and other online documentation. Python will make it is easy to model the
data and generate easy to interpret summaries of the results. Several
machine learning methods will be employed in this project and the
performance of each will be compared. Since the target variable is
categorical, neural networks, SVM, random forest, and ensemble models
will be used which can all be implemented using Python.
Sources:
IBM Cloud Team & IBM Cloud (2021). Python vs. R: what’s the
difference? IBM. https://www.ibm.com/cloud/blog/python-vs-r
Luna, J. (2022). Top programming languages for data scientists in 2022.
Datacamp. https://www.datacamp.com/blog/top-programming-languages-
for-data-scientists-in-2022