1 / 1100%
Now that I have selected datasets for the project, I plan to follow the
standard steps in analyzing data:
1. Data Cleaning is the first and arguably most important step in an
analytics project since any errors or inconsistencies in the data
can have drastic impacts on results. This is typically the step
where duplicates, irrelevant variables, missing values,
inconsistent data formatting, outliers, and data transformations
are addressed. A quick exploration of one of my datasets showed
that the data for each county is stored in rows, rather than in
columns (i.e., each county has multiple rows). Because my
second data set has the data for each county stored in columns
(i.e., each county has one row), I will have to transform the first
dataset before I can merge.
2. Exploratory Data Analysis is important for gaining a deeper
understanding of the dataset through statistical various
measures and visualizations. This is the portion of the analysis
that is the most fun since there is ample opportunity for
creativity in exploring and generating visualizations.
3. Variable Creation or feature engineering is where features are
chosen or created for input into a model. There are multiple
options for this step including feature extraction, feature
selection, feature construction or feature learning (Bharadwa,
2021). The results of EDA will help in determining which
approach to take.
4. Data Modeling is where a model will be fit to the data to address
the problem statement. I have not yet determined a specific
modeling approach, but will likely compare the performance of
multiple model types and choose the best.
Sources:
Bharadwa, A. (2021). 7 Steps to a successful data science project.
Towards Data Science. https://towardsdatascience.com/7-steps-to-
a-successful-data-science-project-b452a9b57149
Students also viewed