1 / 2100%
A common mistake made in data science projects is rushing into
data collection and analysis, which precludes spending sufficient
time to plan and scope the amount of work involved,
understanding requirements, or even framing the business problem
properly.
The following data analytic lifecycle will be applied to e e in
achieving the scope of this project
Phase 1—Discovery: In Phase 1, the team learns the business
domain, including relevant history such as whether the
organization or business unit has attempted similar projects in the
past from which they can learn. The team assesses the resources
available to support the project in terms of people, technology,
time, and data. Important activities in this phase include framing
the business problem as an analytics challenge that can be
addressed in subsequent phases and formulating initial hypotheses
(IHs) to test and begin learning the data.
Phase 2—Data preparation: Phase 2 requires the presence of an
analytic sandbox, in which the team can work with data and
perform analytics for the duration of the project. The team needs
to execute extract, load, and transform (ELT) or extract, transform
and load (ETL) to get data into the sandbox. The ELT and ETL are
sometimes abbreviated as ETLT. Data should be transformed in
the ETLT process so the team can work with it and analyze it. In
this phase, the team also needs to familiarize itself with the data
thoroughly and take steps to condition the data (Section 2.3.4).
Phase 3—Model planning: Phase 3 is model planning, where the
team determines the methods, techniques, and workflow it
intends to follow for the subsequent model building phase. The
team explores the data to learn about the relationships between
variables and subsequently selects key variables and the most
suitable models.
Phase 4—Model building: In Phase 4, the team develops datasets
for testing, training, and production purposes. In addition, in this
phase the team builds and executes models based on the work
done in the model planning phase. The team also considers
whether its existing tools will suffice for running the models, or if
it will need a more robust environment for executing models and
workflows (for example, fast hardware and parallel processing, if
applicable).
Phase 5—Communicate results: In Phase 5, the team, in
collaboration with major stakeholders, determines if the results of
the project are a success or a failure based on the criteria
developed in Phase 1. The team should identify key findings,
quantify the business value, and develop a narrative to summarize
and convey findings to stakeholders.
Phase 6—Operationalize: In Phase 6, the team delivers final
reports, briefings, code, and technical documents. In addition, the
team may run a pilot project to implement the models in a
production environment.
Analysis plan: In this stage resources like technology, tools. data,
models. systems and people, from analytical software packages,
such as Python, Tableau, R, SAS etc. will be e e used on file extracts
and datasets for testing purposes. This assesses the validity of
the model and its results. For instance, determine if the model
accounts for most of the data and has robust predictive power. At
this point, refine the models to optimize the results, such as by
modifying variable inputs or reducing correlated variables where
appropriate. which will be confirmed or denied once the models
are executed. When immersed in the details of constructing
models and transforming data, many small decisions are often
made about the data and the approach for the modeling.
https://learning.oreilly.com/library/view/data-science-
and/9781118876138/10_chapter-02.html#c02-04
Students also viewed