1 / 28100%
Module 5
Exploring Data Visually
A. Exploratory Data Analysis
As a first step in an analytical study, it is critical to examine and explore the data.
This exploratory data analysis (EDA) makes heavy use of descriptive statistics and
visualization to gain an initial understanding of the data. The objectives of EDA include
(1) detection of errors, missing values, and any other unusual observations; (2)
characterization of the distribution of the values for the individual variables; and (3)
identification of patterns and relationships between variables. Visual display is an
essential principle of EDA, as it allows the analyst to translate the information contained
in the rows and columns of data into charts, providing “first looks at the data” that
achieve the EDA objectives.
To understand the challenges in exploring data, we consider its structural
dimensions. Tall data occurs when the number of records (rows) is large. Wide data
occurs when the number of variables (columns) is large. As data grow taller or wider, the
possibility of data errors and missing values increases. In addition, wide data becomes
increasingly arduous to explore because there are a large number of possible
combinations of variables to examine.
Let us consider an example involving Espléndido Jugo y Batido, Inc. (EJB), a
company that manufactures bottled juices and smoothies. EJB produces its products in
five fruit flavors (apple, grape, orange, pear, and tomato) and four vegetable flavors (beet,
carrot, celery, and cucumber), and it ships these products from distribution centers (DCs)
in Idaho, Mississippi, Nebraska, New Mexico, North Dakota, Rhode Island, and West
Virginia. EJB management has retrieved data on each order it has received over the past
three years and stored it in the file EJB.
Within the intricacies of dataset representation, it is pivotal to comprehend the
fundamental structure where each record encapsulates the quantity of a singular product.
This product, delineated by a unique combination of category and flavor, is integral to the
datasets core. Notably, it’s crucial to recognize that a single order can span across
multiple records within the dataset. In this nuanced system, the granularity of the data
extends beyond a mere order level, emphasizing the individual products that constitute
each order. This granular approach facilitates a detailed examination of the products
Students also viewed