1 / 2100%
A common mistake made in data science projects is rushing into data
collection and analysis, which precludes spending sufficient time to
plan and scope the amount of work involved, understanding
requirements, or even framing the business problem properly.
The following data analytic lifecycle will be applied to in achieving the
scope of this project
Phase 1—Discovery: In Phase 1, the team learns the business domain,
including relevant history such as whether the organization or
business unit has attempted similar projects in the past from which
they can learn. The team assesses the resources available to support
the project in terms of people, technology, time, and data. Important
activities in this phase include framing the business problem as an
analytics challenge that can be addressed in subsequent phases and
formulating initial hypotheses (IHs) to test and begin learning the data.
Phase 2—Data preparation: Phase 2 requires the presence of an
analytic sandbox, in which the team can work with data and perform
analytics for the duration of the project. The team needs to execute
extract, load, and transform (ELT) or extract, transform and load (ETL)
to get data into the sandbox. The ELT and ETL are sometimes
abbreviated as ETLT. Data should be transformed in the ETLT process
so the team can work with it and analyze it. In this phase, the team also
needs to familiarize itself with the data thoroughly and take steps to
condition the data (Section 2.3.4).
Phase 3—Model planning: Phase 3 is model planning, where the team
determines the methods, techniques, and workflow it intends to
follow for the subsequent model building phase. The team explores
the data to learn about the relationships between variables and
subsequently selects key variables and the most suitable models.
Phase 4—Model building: In Phase 4, the team develops datasets for
testing, training, and production purposes. In addition, in this phase
the team builds and executes models based on the work done in the
model planning phase. The team also considers whether its existing
tools will suffice for running the models, or if it will need a more robust
environment for executing models and workflows (for example, fast
hardware and parallel processing, if applicable).
Phase 5—Communicate results: In Phase 5, the team, in collaboration
with major stakeholders, determines if the results of the project are a
success or a failure based on the criteria developed in Phase 1. The
team should identify key findings, quantify the business value, and
develop a narrative to summarize and convey findings to stakeholders.
Phase 6—Operationalize: In Phase 6, the team delivers final reports,
briefings, code, and technical documents. In addition, the team may
run a pilot project to implement the models in a production
environment.
Analysis plan: In this stage resources like technology, tools. data,
models. systems and people, from analytical software packages, such
as Python, Tableau, R, SAS etc. will be used on file extracts and
datasets for testing purposes. This assesses the validity of the model
and its results. For instance, determine if the model accounts for most
of the data and has robust predictive power. At this point, refine the
models to optimize the results, such as by modifying variable inputs or
reducing correlated variables where appropriate. which will be
confirmed or denied once the models are executed. When immersed in
the details of constructing models and transforming data, many small
decisions are often made about the data and the approach for the
modeling.
https://learning.oreilly.com/library/view/data-science-
and/9781118876138/10_chapter-02.html#c02-04
Students also viewed