1 / 1100%
After combining the data from the two data sources and performing
data cleaning, the final dataset has just over 4 million records and 29
variables which can be used to predict the binary outcome variable
depression. The outcome variable is split such that 70% of the patients
do not have depression and 30% have depression, which makes the
dataset slightly imbalanced but not enough so that it requires
balancing.
Because the outcome variable is binary, predicting depression in this
project is a classification problem, so all the machine learning models
explored are classifiers. The models explored include logistic
regression, k nearest neighbors, decision trees, and random forest.
For building the models the data is divided into training and testing
sets using a 70/30 split. With 70% in the training set and 30% in the
test set. Each model is run on the training set and then assessed using
the test set. The model parameters are tuned to optimize model
performance. The performance of each model is then compared using
several measures including AUC, precision, recall, and specificity.
Several measures are used for model evaluation to get a more detailed
understanding of the performance and avoid making inaccurate
conclusions.
Trying to run the models with millions of records proved to be time
consuming and so the dataset was randomly sampled using the
stratified random sampling technique to create a smaller subset to
work with that maintains that has the same class proportions as the
original dataset. Currently I am still tweaking the model inputs to get
better results as model performance has been poor so far.
Students also viewed