1 / 2100%
A common mistake made in data science projects is rushing into
data collection and analysis, which precludes spending sufficient
time to plan and scope the amount of work involved, understanding
requirements, or even framing the business problem properly.
The following data analytic lifecycle will be applied to in achieving
the scope of this project
Phase 1—Discovery: In Phase 1, the team learns the business
domain, including relevant history such as whether the organization
or business unit has attempted similar projects in the past from
which they can learn. The team assesses the resources available to
support the project in terms of people, technology, time, and data.
Important activities in this phase include framing the business
problem as an analytics challenge that can be addressed in
subsequent phases and formulating initial hypotheses (IHs) to test
and begin learning the data.
Phase 2—Data preparation: Phase 2 requires the presence of an
analytic sandbox, in which the team can work with data and
perform analytics for the duration of the project. The team needs to
execute extract, load, and transform (ELT) or extract, transform and
load (ETL) to get data into the sandbox. The ELT and ETL are
sometimes abbreviated as ETLT. Data should be transformed in the
ETLT process so the team can work with it and analyze it. In this
phase, the team also needs to familiarize itself with the data
thoroughly and take steps to condition the data (Section 2.3.4).
Phase 3—Model planning: Phase 3 is model planning, where the
team determines the methods, techniques, and workflow it intends
to follow for the subsequent model building phase. The team
explores the data to learn about the relationships between variables
and subsequently selects key variables and the most suitable
models.
Phase 4—Model building: In Phase 4, the team develops datasets for
testing, training, and production purposes. In addition, in this phase
the team builds and executes models based on the work done in the
model planning phase. The team also considers whether its existing
tools will suffice for running the models, or if it will need a more
robust environment for executing models and workflows (for
example, fast hardware and parallel processing, if applicable).
Phase 5—Communicate results: In Phase 5, the team, in
collaboration with major stakeholders, determines if the results of
the project are a success or a failure based on the criteria
developed in Phase 1. The team should identify key findings,
quantify the business value, and develop a narrative to summarize
and convey findings to stakeholders.
Phase 6—Operationalize: In Phase 6, the team delivers final reports,
briefings, code, and technical documents. In addition, the team may
run a pilot project to implement the models in a production
environment.
Analysis plan: In this stage resources like technology, tools. data,
models. systems and people, from analytical software packages,
such as Python, Tableau, R, SAS etc. will be h used on file extracts
and datasets for testing purposes. This assesses the validity of the
model and its results. For instance, determine if the model accounts
for most of the data and has robust predictive power. At this point,
refine the models to optimize the results, such as by modifying
variable inputs or reducing correlated variables where appropriate.
which will be confirmed or denied once the models are executed.
When immersed in the details of constructing models and
transforming data, many small decisions are often made about the
data and the approach for the modeling.
https://learning.oreilly.com/library/view/data-science-
and/9781118876138/10_chapter-02.html#c02-04
Students also viewed