STATISTICS FOR DECISION MAKING
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 1/45
3.1 Survey sampling
Descriptive and inferential statistics
Two types of statistical analysis exist to describe survey data: descriptive statistics and inferential statistics.
Descriptive statistics focuses on summarizing survey data about a sample drawn from a population. Summary statistics include measures of central tendency such as mean, median, and mode; and dispersion such as range and standard deviation. Descriptive statistics cannot make conclusions based on the data. Rather, descriptive statistics is a way to present data in a meaningful way.
Inferential statistics focuses on using information from the sample to make conclusions about the population from which the sample was drawn. The two primary methods of inferential statistics are con�dence intervals, which specify the range within which a parameter falls with a given probability, and hypothesis testing, which allows differences between population parameters to be compared.
Surveys
Surveys are conducted to allow statisticians to make generalizations about a population.
A population is any collection of objects, people, or things about which statistical inference are made. A parameter of a population is a numerical characteristic of a population, such as mean, median, or standard deviation.
A sampling unit is an individual in the population on which a measurement can be taken.
The sampling frame is the subset of the population from which a sample is drawn.
The sample is composed of the sampling units that provide data to be collected.
A statistic is a numerical characteristic of a sample, rather than the population.
The following animation shows the relationship between the population, sampling unit, sampling frame, and sample.
PARTICIPATION ACTIVITY 3.1.1: Sampling a population.
Animation captions:
1. A sample is a representative subset of a population that is used to measure a parameter of the population.
2. A sampling unit is an individual in the population from which a parameter can be measured. 3. The sampling frame is the subset of the population from which samples can be drawn. 4. The sample is the subset of the sampling frame from which measurements are actually taken.
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 2/45
The following animation shows the relationship between a parameter and a statistic.
PARTICIPATION ACTIVITY 3.1.2: Parameters and statistics.
Example 3.1.1: The Liraglutide Effect and Action in Diabetes: Evaluation of Cardiovascular Outcome Results (LEADER) clinical trial.
The LEADER clinical trial was initiated in 2010 at 410 hospitals in 32 countries to evaluate the effect of liraglutide, a drug for treatment of type 2 diabetes, on the frequency of cardiovascular diseases such as heart attack, stroke, and heart failure . The populations under study were type 2 diabetes patients with excessively high blood sugar taking either liraglutide or a placebo (an inactive drug). The parameters measured include blood sugar level, kidney function measurement, frequency of adverse effects and complications, and mortality rate. The overall goal of the trial was to determine whether liraglutide treatment was effective for treating type 2 diabetes without increasing the danger of cardiovascular complications.
Identify the sampling unit, sampling frame, and surveys conducted.
Solution
The sampling unit was a type 2 diabetes patient at any of the hospitals at which the trial was conducted. The sampling frame was the subset of type 2 diabetes patients who ful�lled the criteria for inclusion into the clinical trial, such as age, cardiovascular disease status, other drugs taken, and whether informed consent to participate in the study was given. The surveys included the medical tests that were performed to measure blood sugar level and other health information, as well as observations of the frequency and severity of adverse effects, complications, and deaths.
Animation captions:
1. A preschool's population has aged between and . Sally, aged , is one such child, as is Joey, aged . The dot plot shows the number of children of each age.
2. Sample A consists of children, including Sally but not Joey, selected from the population of .
3. Sample B consists of other children selected from the population, this time including Joey but not Sally.
4. The population mean is a parameter. The mean of each sample is a statistic.
15 1 5 5 4
5 15
5
1
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 3/45
PARTICIPATION ACTIVITY 3.1.3: Distinguishing populations and samples.
1) An analyst obtains the salaries of all federal appeals court judges in
2014 and computes the mean salary to be . Are the judges a population or a sample?
2) An analyst obtains the salaries of all federal appeals court judges in
2014 and computes the mean salary to be . Is a parameter or a statistic?
3) An analyst surveys registered nurses across the U.S. and computes their mean earnings to be /hr. The analyst reports that U.S. nurses have mean earnings of /hr . Are the
nurses a population or a sample?
4) An analyst surveys registered nurses across the U.S. and computes their mean earnings to be /hr. The analyst reports that U.S. nurses have mean earnings of /hr. Is /hr a parameter or a statistic?
5) An analyst is asked to determine how many miles each employee commutes at a -person company. Should the analyst collect data from the population or from a sample?
167
$211, 200 167
Population
Sample
167
$211, 200 $211, 200
Parameter
Statistic
1, 000
$27
$27 2 3
1, 000
Population
Sample
1, 000
$27
$27 $27
Parameter
Statistic
20
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 4/45
6) An analyst is asked to determine how many miles each employee commutes at Microsoft, which has over employees . Should the analyst collect data from the population or on a sample?
Bias
In statistics, a bias is a difference between the parameter predicted from a survey from the true value of the parameter in the population. Two broad categories of statistical bias include selection bias and response bias.
Selection bias exists when the sampling units selected from a population are not representative of the entire population, and are instead biased toward certain subsets of the population. A population should be surveyed in such a way to minimize sampling bias. Several types of selection bias follow.
Undercoverage occurs when certain members of a population are inadequately represented in a sample.
Nonresponse bias occurs when a sample is biased toward members of a population that participate in a survey.
Voluntary response bias occurs when a sample is biased toward members that self-select for participation in a survey.
Response bias can result if the responses of survey participants are affected by how a question is asked or the behaviors or attitudes of the participant. Several types of response bias follow.
Acquiescence bias occurs when respondents tend to agree with a statement in a survey.
Extreme responding occurs when respondents tend to select the most extreme options available.
Social desirability bias occurs when respondents tend to answer questions in a way that is socially accepted by others. In other words, a social desirability bias exists when respondents over-report "good" behaviors or under-report "bad" behaviors.
Example 3.1.2: Types of bias.
For each situation below, determine the most likely type of selection bias.
Population
Sample
100, 000 4
Population
Sample
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 5/45
a. A website survey b. A survey about the frequency of alcohol consumption c. In-person survey conducted at a mall in an a�uent neighborhood d. A survey conducted over landline phones
Solution
a. A comments section of a website soliciting survey responses on a controversial issue in which most of the participants express extreme viewpoints for or against may exhibit voluntary response bias because members of the population who are indisposed to the issue are not adequately represented.
b. Social desirability bias may exist in a survey about the frequency of alcohol consumption among subsets of the population with differing social attitudes or prohibitions toward alcohol.
c. Undercoverage. An in-person survey conducted at a shopping center in an a�uent neighborhood may inadequately represent members of the population without the economic or transportation means to travel to or shop at the mall.
d. Nonresponse bias. A survey conducted over landline phones exhibits nonresponse bias because members of the population who exclusively use a cell phone are not adequately represented.
PARTICIPATION ACTIVITY 3.1.4: Types of bias.
Match each description to the correct type of bias.
A survey asking how often respondents send text messages while driving routinely underestimates the actual frequency of texting while driving.
Students conducting a survey on a campus-wide issue only conducted the survey in front of the main
Acquiescence bias Nonresponse bias Undercoverage Extreme responding
Social desirability bias
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 6/45
engineering building, and the concerns of humanities students on the other side of the campus were not adequately addressed. An internet survey sent to a remote rural area with poor internet penetration had only a response rate.
The majority of reviews on a restaurant's social media site are either one-star or �ve-star reviews.
A question on a political ballot contains the question "Is Freedom Important?", which is a relatively non- controversial statement with which
of voters agreed.
Sampling methods
Different sampling methods can help mitigate certain types of statistical bias.
In simple random sampling, a sample is constructed by random selection from the population. Mathematically, simple random sampling is a sampling method in which all possible samples consisting of units selected from a population of units are equally likely.
In systematic sampling, every th unit from a population of units is selected to be in a sample.
In strati�ed sampling, the population is �rst divided into groups, or strata, depending on some characteristic. Next, samples within each stratum are randomly selected in a proportional manner.
In cluster sampling, the population is �rst divided into groups, or clusters, depending on some characteristic. Next, the sample is constructed by randomly selecting one or more clusters.
In convenience sampling, units are drawn from a subset of the population that is readily available.
Example 3.1.3: Sampling methods.
Determine which sampling method is used in each situation.
a. A community college contains students in the School of Arts and Sciences, students in the School of Engineering, and students in the School of
3%
98%
Reset
n N
k N
3000 1000 1000
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 7/45
Performing Arts. From each school, of students are randomly selected for participation in the survey, for a total of students from the School of Arts and Sciences, students from the School of Engineering, and students from the School of Performing Arts.
b. Participants for the survey are recruited by �agging down students crossing the main quad of the community college until the necessary number of students have been recruited.
c. The student body of the community college consists of 1st year students, 2nd year students, 3rd year students, 4th year students, and transfer students. 3rd year students were randomly selected for participation in the survey.
d. students are randomly selected from the student body of for participation in the survey.
e. Participants for the survey are recruited by selecting every th name from a list of students at the community college.
Solution
a. Strati�ed sampling b. Convenience sampling c. Cluster sampling d. Simple random sampling e. Systematic sampling
PARTICIPATION ACTIVITY 3.1.5: Sampling methods.
Match each description to the correct sampling method.
A major polling company in the United States randomly selects Alaska, Illinois, Texas, Florida, and Pennsylvania as the states from which to select households for a survey.
10% 300
100 100
500 5000
10
Simple random sampling Strati�ed sampling Cluster sampling
Systematic sampling
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 8/45
To select households for the survey, the company consults the property tax rolls in a speci�c area, and selects every household with an odd street number for the survey. To select households for the survey, the company consults the property tax rolls in a speci�c area, and selects households at random for the survey.
To select households for the survey, the company consults the property tax rolls in a speci�c area, and subdivides households into several property value brackets. Households are then randomly and proportionally selected from each bracket.
References
(*1) Marso, Steven P., et al. "Liraglutide and Cardiovascular Outcomes in Type 2 Diabetes." The New England Journal of Medicine, 375:311-322, 28 July 2016, DOI: 10.1056/NEJMoa1603827
(*2) "Providers and Service Use Indicators NURSES AND PHYSICIAN ASSISTANTS." Henry J Kaiser Family Foundation, kff.org/other/state-indicator/total-registered-nurses/
(*3) "Registered Nurse (RN) Salary." Payscale.com, 2016, www.payscale.com/research/US/Job=Registered_Nurse_(RN)/Hourly_Rate, 2016 data
(*4) "Facts About Microsoft." Microsoft, 2015, http://news.microsoft.com/facts-about- microsoft/#sm.0001d1d65h4x8e63v142jztf9potu
Reset
3.2 Measures of center
The mean
Large amounts of data can be overwhelming. A single number can summarize information about a dataset, such as the central tendency or the dispersion of the dataset. A common data summary is the arithmetic mean or mean, which is the sum of the data values in a dataset divided by the number of values in the dataset.
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 9/45
Mathematically, the mean of a set of data values is denoted and is de�ned as follows.
A similar quantity called the weighted mean is often calculated in addition to the mean. The weighted mean is a measure of center where some values are counted more than once. Weights are often expressed either as positive integers or percentages. Mathematically, the weighted mean of a set
with corresponding weights is de�ned as follows.
The following animation shows the relationship of the mean to the values in a dataset.
PARTICIPATION ACTIVITY 3.2.1: The mean.
Example 3.2.1: Finding weighted means.
Find the weighted mean of the following dataset.
Data Weight
Solution
n , , … ,x1 x2 xn x̄
= =x̄ 1
n ∑ i=1
n
xi + + ⋯ +x1 x2 xn
n
, , … ,x1 x2 xn , , … ,a1 a2 an
= =x̄ ∑ni=1 ai xi
∑ni=1 ai
+ + … +a1 x1 a2 x2 anxn + + … +a1 a2 an
Animation content:
undefined
Animation captions:
1. The mean summarizes data, computed as the sum divided by the number of values. 2. Graphically, the mean is a value that balances the data values.
6 2
8 2
17 1
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 10/45
Data with weights that are greater than are counted more than once, which means that the dataset above is the same as . Using the formula, the weighted mean is
PARTICIPATION ACTIVITY 3.2.2: The mean.
1) What is the mean of the dataset ?
2) What is the mean of the dataset ?
3) What is the mean of the dataset ? Type as: #.#
4) What is the mean of the dataset ?
5) What is the weighted mean of the dataset if the weights are , , and respectively?
1 6, 6, 8, 8, 17
= = = = 9x̄̄̄ ∑ni=3 ai xi
∑ni=1 ai
2 ⋅ 6 + 2 ⋅ 8 + 1 ⋅ 17
2 + 2 + 1
45
5
12, 1, 2
Check Show answer
2, 6, 4
Check Show answer
2, 3, 4, 1
Check Show answer
−3, 15
Check Show answer
20, 4, 1 1 2 2
Check Show answer
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 11/45
6) What is the weighted mean of the dataset if the weights are , , and respectively?
PARTICIPATION ACTIVITY 3.2.3: Spreadsheets: Mean and weighted mean.
Video 3.2.1: Mean rent in 6 cities in the United States.
FDA zyBook: Mean rent in 6 citiesFDA zyBook: Mean rent in 6 cities
3, 0, 37 2 4 2
Check Show answer
Animation content:
undefined
Animation captions:
1. Two column are �lled with data that requires analysis. The columns contain random data with corresponding weights.
2. The mean of entries A2 through A5 is found using the AVERAGE function. The cursor is dragged to an empty cell, D1, where the formula is entered.
3. The data from cells A2 to A5 are selected. Pressing enter displays the output in cell D1. 4. The SUMPRODUCT and SUM functions are to �nd the weighted mean. 5. For the SUMPRODUCT function, the data for the weights in cells B2 to B5 are also selected.
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 12/45
The median
The median is the middle value in a sorted dataset. To �nd the median of a dataset, the dataset must �rst be sorted in ascending or descending order. The method of �nding the median depends on whether the number of data values is even or odd.
If is odd, the median is the middle value of the sorted dataset. Speci�cally, the median is the
th value.
If is even, the median is the mean of the middle two values of the sorted dataset. Speci�cally,
the median is the mean of the th and th values.
Example 3.2.2: The median.
Find the median of each of the following datasets.
a. b. c.
Solution
a. The dataset contains values and is already sorted in ascending order. Thus, the
median is the th data value, which is .
b. The dataset contains values and is already sorted in descending order. Thus, the
median is the mean of the th data value and the th data value,
which is .
c. The dataset must �rst be arranged in ascending or descending order to �nd the median. In ascending order, the dataset is . Thus, the median is the
rd data value, which is .
PARTICIPATION ACTIVITY 3.2.4: The median.
1) What is the median of the dataset ?
n
n
( )n + 1 2
n
( )n 2
( + 1)n 2
10, 20, 20, 30, 60, 60, 80 99, 80, 60, 60, 30, 20, 20, 10 100, 3, 6, 9, 2
7
= 4 7 + 1
2 30
8
= 4 8
2 + 1 = 5
8
2
= = 45 60 + 30
2
90
2
2, 3, 6, 9, 100
= 3 5 + 1
2 6
1, 4, 5, 9, 11
6
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 13/45
2) What is the median of the dataset ?
3) What is the median of the dataset ?
4) What is the median of the dataset ?
PARTICIPATION ACTIVITY 3.2.5: Spreadsheets: Median.
Outliers
The dataset in a previous example illustrates an advantage of the median over the mean. The data value is much larger than the other data values. Such a value is an outlier, or a data value that is either much greater than or much less than the rest of the data and not representative of the rest of the data being considered. Compared to the median of , the mean is
5
4, 3, 2, 6, 7
2
4
2, 3, 5, 18
3
4
5
7
−1, −5, −3, 6, 7
−3
−5
−1
Animation content:
undefined
Animation captions:
1. Two sets of data �ll the columns A and B. 2. The median of entries A2 through A4 is found using the MEDIAN function. The cursor is
dragged to an empty cell, D1, where the formula is entered. 3. The data from cells A2 to A4 are selected. Pressing enter displays the output in cell D1. 4. The median for Set B is calculated similarly.
100, 3, 6, 9, 2 100
6
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 14/45
, which is much larger than due to in�uence by the outlier of
.
As a practical example, the net worth (in USD) of particular individuals in Medina, Washington in 2015 was , , , , and . The outlier of
is due to Bill Gates, a co-founder of Microsoft, living in Medina. The mean net worth of is thus a poor data summary because the mean suggests that all people are wealthy multimillionaires.
The following animation shows the relationship of the median to the mean and the values in a dataset, including outliers.
PARTICIPATION ACTIVITY 3.2.6: The median.
Example 3.2.3: Pensions in San Diego.
In 2011, the mean pension among the 10 highest earning city employees was (the median was ).
In recent years, various scandals have been reported relating to exorbitant pensions that city employees approved for themselves, sometimes resulting in city staff later being found guilty of crimes and imprisoned. The following table summarizes data for a major city.
Pensions are paid for the pensioners' remaining life, often 30 years or more. The following table lists the pension amounts and the person's last job position, which the above mean and median summarize.
Last job position Pension amount
Assistant City Attorney
Investment O�cer
Fire Battalion Chief
Assistant Police Chief
= = 24 2 + 3 + 6 + 9 + 100
5
120
5 6
100
5 $300, 000 $400, 000 $250, 000 $80, 000, 000, 000 $600, 000
$80, 000, 000, 000 $16, 000, 000, 000 5
Animation captions:
1. To �nd the median, the data must �rst be sorted. 2. The median is the middle value among the sorted values. 3. The value of an outlier affects the mean but not the median.
$239, 940 $231, 922
$307, 758
$255, 509
$244, 435
$242, 947
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 15/45
City Librarian
Fire Chief
Fire Battalion Chief
Deputy City Attorney
Fire Battalion Chief
Assistant Water Department Director
Example 3.2.4: Misleading lawyer ads.
Let us represent you: Our mean settlement amount is .
Let us represent you: Our median settlement amount is .
Law �rms sometimes report mean settlement amounts rather than the more appropriate median. The idea is to lure clients into believing they may win a larger settlement than is actually likely, since the mean is in�uenced by a few big wins. The following provides two summaries of data for a law �rm whose win data is shown below.
Examining the table below, one can see that the above mean is misleading, due to being unduly in�uenced by the top 2-3 settlements. The data summary using the median more appropriately re�ects what a client might expect to win.
Group settlement for motor vehicle accident
Federal verdict against Veteran's Administration for not detecting neurological condition and performing surgery before woman became paralyzed
Medical malpractice involving undiagnosed kidney failure prior to baby delivery
Nursing home neglect leading to bedsores
Nursing home neglect
Nursing home neglect leading to a fall
Nursing home neglect leading to death
$234, 091
$229, 753
$228, 392
$224, 863
$217, 649
$214, 007
$1, 529, 000
$675, 000
$7, 500, 000
$5, 700, 000
$4, 750, 000
$3, 000, 000
$1, 400, 000
$1, 000, 000
$950, 000
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 16/45
Motor vehicle accident resulting in death
Nursing home neglect leading to bedsores
Medical malpractice involving death due to overdose
Nursing home neglect leading to infection
Product liability case involving leg injury from an ATV
Client struck by a limousine
Nursing home neglect resulting in death
Medical malpractice involving pressure sores
Medical malpractice involving pressure sores in hospital
Motor vehicle collision
Client struck in the arm by a stray bullet while driving
Client suffered neck injury from a motor vehicle accident
Client injured knee falling through a restaurant trap door
Example 3.2.5: Mean and median age of marriage.
The median age at �rst marriage in a certain country increased from for men and for women in 2000 to for men and for women in 2010. In this case, the data is likely skewed to the right. People do not usually marry much earlier than , but can marry as old as or later. The outliers here are not a few high numbers, but rather a skewing of data that increases the mean. Thus, the median is more often reported compared to the mean.
PARTICIPATION ACTIVITY 3.2.7: Mean and median.
1) The pension data above can be summarized nearly equally well using either the mean or the median.
$830, 000
$750, 000
$700, 000
$650, 000
$600, 000
$550, 000
$500, 000
$400, 000
$375, 000
$300, 000
$230, 000
$215, 000
$194, 000
26.8 25.1 28.2 26.1
18 50
True
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 17/45
2) The settlement data above can be summarized nearly equally well using either the mean or the median.
3) In the marriage example above, the data presenter likely chose median because the presenter is trying to lure people into believing marriage age is younger than justi�ed by the data.
4) Law �rms are not the only companies that report the mean when the median would be more appropriate.
5) Home prices for a given city or state are commonly summarized in news articles using the median price.
The mode
The mode is the most frequently-occurring value in a dataset and is another measure of center. A dataset may have multiple modes if multiple values have the same maximum frequency. A dataset with only unique values does not have a mode.
Example 3.2.6: The mode.
Find the mode or modes of each dataset.
a. b. c.
False
True
False
True
False
True
False
True
False
1, 2, 2, 2, 3, 3, 4 1, 2, 2, 3, 3, 4 1, 2, 3, 4
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 18/45
Solution
a. Since is the value with the highest frequency , the mode is . b. Since both and have the highest frequency , the dataset has two modes,
and . c. Since every value in the dataset is unique, the dataset does not have a mode.
PARTICIPATION ACTIVITY 3.2.8: The mode.
1) What is the mode of the dataset ?
2) Which of the following statements is true about the dataset , , , ?
3) After a baseball tournament, the number of runs scored by each player is . What is the mode?
PARTICIPATION ACTIVITY 3.2.9: Spreadsheets: Mode.
2 (3) 2 2 3 (2) 2
3
1, 4, 4, 5, 5, 9, 9, 9
1
5
9
2 5 6 7
No mode exists
All values in the dataset are modes.
0, 0, 0, 0, 0, 0, 0, 3, 3, 6, 10
0
7
Animation content:
undefined
Animation captions:
1. Two sets of data �ll the columns A and B.
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 19/45
Spreadsheet functions
Table 3.2.1: Spreadsheet-Functions: Mean, median, and mode.
Function Syntax Description
AVERAGE AVERAGE(number1, [number2], ...)
Returns the arithmetic mean of given numbers. The input can be a cell range.
SUMPRODUCT SUMPRODUCT(array1, [array2], [array3], ...)
Multiplies corresponding components of arrays.
SUMPRODUCT, in addition to SUM, is needed to calculate weighted means.
MEDIAN MEDIAN(number1, [number2], ...)
Returns the median of given numbers. The input can be a cell range.
MODE MODE(number1, [number2],...)
Returns the mode of given numbers. The input can be a cell range.
MODE.MULT MODE.MULT((number1, [number2],...)
Returns an array of modes in the selected cells. CTRL+SHIFT should be pressed to return the output of a function array.
Challenge activities
CHALLENGE ACTIVITY 3.2.1: Mean, median, mode.
2. The mode of the data in cells B2 through B5 is found using the MODE function. The cursor is dragged to an empty cell, D1, where the formula is entered.
3. The data from cells A2 to A4 are selected. Pressing enter displays the output in cell D1. 4. To �nd the modes for a multimodal dataset, the MODE.MULT function is used. Since the
output is an array, an array formula is used, by selecting more than one cell. 5. The CTRL+SHIFT are pressed after selecting the cells containing the dataset to return the
modes. 6. The output in cell D4 is #N/A because the dataset contains two modes, not three.
Start
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 20/45
CHALLENGE ACTIVITY 3.2.2: Spreadsheets: mean, median, mode.
Click this link to download the spreadsheet for use in this activity.
2 3 4 5
Enter the mean of the following dataset: 13, 11, 12, 6
Ex: 3.2=x̄̄̄
Check Next
Start
2 3 4
The �rst sheet of the spreadsheet linked above contains the scores of 50 students on 4 differ What is the mean exam score of Student 47 for all four exams?
Ex: 50.8
Check Next
1
1
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 21/45
3.3 Measures of variability
Variance and standard deviation
Variability is the difference between values in a dataset and the center of the dataset. A measure of center alone does not indicate the extent of variability. Ex: the data set , , , and and the data set
, , , and both have a mean of and a median of . However, the values in the �rst dataset have a larger variability. Two common measures of variability are variance and standard deviation. Variance, is the average of the square difference from the mean. Standard deviation is the square root of the variance. By de�nition, the variance is the square of the standard deviation.
The formula for variance and standard deviation depends on whether the dataset contains the whole population or or a subset of the population. The sample standard deviation is denoted by , while the population standard deviation is denoted by . The formulas are given below.
In the formulas above, is the number of data values, are the data values, is the sample mean, and is the population mean. The numerator of the fraction in both variance formulas is often referred to as the sum of the square differences. A large standard deviation indicates a more spread-out data set. A small standard deviation indicates a more tightly clustered data set.
The following �gure shows histograms of data with a low standard deviation and a high standard deviation.
Figure 3.3.1: Data sets with low and high standard deviations.
1 2 8 9 4 5 5 6 5 5
s σ
= (Population variance)σ2 ∑ni=1 ( − μ)ai
2
n
σ = (Population standard deviation) ∑ni=1 ( − μ)ai
2
n
− −−−−−−−−−−−
√
= (Sample variance)s2 ∑ni=1 ( − )ai x̄
2
n − 1
s = (Sample standard deviation) ∑ni=1 ( − )ai x̄
2
n − 1
− −−−−−−−−−−−
√
n , , … ,a1 a2 an n x̄ μ
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 22/45
Example 3.3.1: Finding the variance and standard deviation.
Find the variance and standard deviation of the dataset: , , ,
a. assuming that the dataset is a subset of the population b. assuming that the dataset represents the whole population
Solution
a. Since the dataset is a subset of measurements from a population, the sample variance and sample standard deviation are obtained.
First, the sample mean is found.
Using a table, the sum of the square differences can be obtained.
From the table above, the sum of the square differences is
Thus, the sample variance and sample standard deviation are
4 5 6 13
= = = 7x̄̄̄ 4 + 5 + 6 + 13
4
28
4
ai x̄̄̄ −ai x̄̄̄ ( −ai x̄̄̄)2
4 7 −3 9
5 7 −2 4
6 7 −1 1
13 7 6 36
= 9 + 4 + 1 + 36 = 50∑ i=1
4
( − )ai x̄ 2
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 23/45
b. Since the dataset represents the entire population, the population variance and standard deviation are obtained. The population mean is and the sum of the square differences is , using the same set of calculations as shown above.
Thus, the population variance and population standard deviation are
Analysis
Although the population mean and sample mean are the same in this example, and are generally different. In most cases, is unknown or di�cult to calculate. The sample mean can be used to estimate the population mean, but the sample mean is strongly susceptible to the presence of extreme values.
The sample variance and standard deviation are always greater than the population variance and standard deviation, because the value depends strongly on the elements of the subset taken during sampling. Thus, samples display greater variability than the entire population.
PARTICIPATION ACTIVITY 3.3.1: Variance and standard deviation.
Consider the dataset taken from a subset of a population.
1) What is the sample mean?
2) What is the sum of the squares of the differences between each data value
= = ≈ 16.667s2 ∑4i=1 ( − )ai x̄
2
4 − 1
50
3
s = ≈ 4.082 50
3
−−− √
μ = 7 = 50∑4i=1 ( − μ)ai
2
= = = 12.5σ2 ∑4i=1 ( − μ)ai
2
4
50
4
σ = ≈ 3.536 ∑4i=1 ( − μ)ai
2
4
− −−−−−−−−−−−
√
x̄ μ μ
1, 2, 4, 5
Check Show answer
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 24/45
and the sample mean?
3) What is the sample variance? Type as: #.###
4) What is the sample standard deviation? Type as: #.###
5) Suppose the data represents measures from the entire population. What is the population variance? Type as: #.#
6) Suppose the data represents measures from the entire population. What is the population standard deviation? Type as: #.###
Example 3.3.2: Course student evaluation data summary.
Below is a data summary of student evaluations for a particular course at a major university. The summary includes and compares three sets of data: the course (and professor), all courses within that course's department, and all courses at the university. For each set, the summary provides the mean (Mean), median (Med), and standard deviation (SD). Items to note:
Check Show answer
Check Show answer
Check Show answer
Check Show answer
Check Show answer
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 25/45
The median conveys little useful information for this data (e.g., being 5.0 for nearly all questions under Course); the mean is clearly superior. Per the standard deviation, this course professor's scores for question 13 (the main question of interest for professors) have less variation (0.5) than for the department (0.8) or university (0.9), indicating students were more consistent in rating this professor (highly). The counts per rating category are provided, to provide further insight into how the ratings were distributed. "Percentiles" (shown as "% tile") are also indicated showing how the course ranked compared to department or university courses. Ex: For question 19, this course was rated higher than 83% of all courses at the university (in other words, in the top 17% of all courses).
Mean absolute deviation
Mean absolute deviation (MAD) is the mean of the absolute difference between each value and the mean of the values. The MAD uses the absolute value instead of the square root of a sum of squares to avoid negative distances. The formula for the MAD is given below.
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 26/45
PARTICIPATION ACTIVITY 3.3.2: Mean absolute deviation (MAD).
PARTICIPATION ACTIVITY 3.3.3: Computing mean absolute deviation.
Use the dataset to answer the following.
1) What is the mean?
2) What is the sum of the absolute differences of each data value and the mean?
3) What is the mean absolute deviation of the data values? Type as: #.#
Spreadsheet functions
(Mean absolute deviation)
| − |∑ i=1
n
ai x̄
n
Animation captions:
1. A measure of center alone does not indicate the extent of the variability of the data values. 2. The mean absolute deviation is the mean of the distances of the data values from the mean
of the data values. 3. Data has the same mean but a smaller mean absolute deviation if the data points have a
smaller variability.
1, 2, 4, 5
Check Show answer
Check Show answer
Check Show answer
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 27/45
Table 3.3.1: Spreadsheet-Functions: Measures of variability.
Function Syntax Description
VAR.S VAR.S(number1, [number2], ...)
Returns the sample variance of the given numbers. The input can be a cell range.
VAR.P VAR.P([number1, [number2], ...)
Returns the population variance of the given numbers. The input can be a cell range.
STDEV.S STDEV.S(number1, [number2], ...)
Returns the sample standard deviation of the given numbers. The input can be a cell range.
STDEV.P STDEV.P([number1, [number2], ...)
Returns the population standard deviation of the given numbers. The input can be a cell range.
AVEDEV AVEDEV([number1, [number2], ...)
Returns the mean absolute deviation of the given numbers. The input can be a cell range.
Challenge activities
CHALLENGE ACTIVITY 3.3.1: Measures of variability.
Start
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 28/45
2 3 4
Which dataset has the smallest standard deviation?
Check Next
3.4 Box plots
Minimum, maximum, and range
In addition to standard deviation and variance, the minimum, maximum, and range of a dataset can describe the spread of the dataset.
The maximum of a dataset is the largest value in the dataset. The minimum of a dataset is the smallest value in the dataset. The range of a dataset is the difference between the maximum and minimum of the dataset.
Example 3.4.1: Minimum, maximum, and range.
Find the minimum, maximum, and range of the dataset .
Solution
The largest value in the dataset is . Thus, the maximum is . The smallest value in the dataset is . Thus, the minimum is . The range is the difference between the maximum and the minimum, or .
PARTICIPATION ACTIVITY 3.4.1: Minimum, maximum, and range.
Use the dataset to answer the following.
1) What is the maximum?
−5, 3, 0, −1, 4, 7
7 7 −5 −5
7 − (−5) = 12
−3, 5, 8, 1, −6, 4
1
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 29/45
2) What is the minimum?
3) What is the range?
Percentiles
The th percentile of a dataset is the data value such that percent of the data falls at or below that value. Three percentiles are particularly important.
The �rst quartile is the th percentile. One-quarter of the data fall at or below . The �rst quartile is the median of the lower half of the data. The third quartile is the th percentile. Three-quarters of the data fall at or below . The third quartile is the median of the upper half of the data. Because half of the data fall at or below the median, the median is also the th percentile of a dataset.
Collectively, the minimum and maximum values, , median, and form a set of descriptive statistics called the �ve-number summary.
Example 3.4.2: Creating a �ve-number summary.
The number of receptions made by players on a certain American football team are given by the dataset . Create a �ve-number summary of this data.
Solution
First, the data should be sorted in ascending or descending order. In ascending order, the dataset is
Check Show answer
Check Show answer
Check Show answer
n n
(Q1) 25 Q1
(Q3) 75 Q3
50
Q1 Q3
3, 37, 23, 61, 36, 65, 6, 24, 1, 19, 72, 1, 13, 40, 1
1, 1, 1, 3, 6, 13, 19, 23, 24, 36, 37, 40, 61, 65, 72
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 30/45
Thus, the minimum is and the maximum is . The dataset contains values, so the
median is the th value of . Thus, the lower half of the data is
and the upper half of the data (the median belongs to both the upper and lower halves). The median of the lower half is
and the median of the upper half is .
The �ve-number summary is
Minimum
(�rst quartile)
Median
(third quartile)
Maximum
PARTICIPATION ACTIVITY 3.4.2: Five-number summary.
Complete the �ve-number summary for the dataset .
1) Minimum
2)
3) median
4)
1 72 15
= 8 15 + 1
2 23
1, 1, 1, 3, 6, 13, 19, 23 23, 24, 36, 37, 40, 61, 65, 72
= = 4.5 3 + 6
2
9
2 = = 38.5
37 + 40
2
77
2
1
Q1 4.5
23
Q3 38.5
72
0, −6, 10, 5, 8, 2, −12, 11, −2
Check Show answer
Q1
Check Show answer
Check Show answer
Q3
Check Show answer
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 31/45
5) Maximum
Introduction to box plots
A box plot is a data visualization that uses a box and several lines to depict the distribution of data in a dataset. A box spans the middle of the data, with as the lower boundary of the box and as the upper boundary of the box. The median is shown as a line inside the box. Two lines, known as whiskers, extend from the lower boundary of the box to the minimum and from the upper boundary of the box to the maximum. The whiskers represent the lower and upper of the data.
The following animation shows the creation of a box plot using data from a previous example.
PARTICIPATION ACTIVITY 3.4.3: Creating a box plot.
PARTICIPATION ACTIVITY 3.4.4: Characteristics of a box plot.
Match the value to the corresponding term, based on the following box plot:
Check Show answer
50% Q1 Q3
25%
Animation content:
undefined
Animation captions:
1. To create a box plot, the dataset must be sorted in ascending or descending order. 2. An axis with the minimum and maximum data points as the endpoints shows the range of the
data. 3. The median is the middle number in the ordered data and is represented as a line inside the
box. 4. is the median of the lower half of the data and forms the lower boundary of the box.
5. is the median of the upper half of the data and forms the upper boundary of the box.
6. The box is formed and the whiskers are extended from the lower boundary of the box to the minimum and from the upper boundary of the box to the maximum.
Q1 Q1 = (3 + 6) ÷ 2 = 4.5 Q3 Q3 = (37 + 40) ÷ 2 = 38.5
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 32/45
The minimum value appearing in the data set
The maximum value appearing in the data set
The median value of the data set
The number of data points in the data set
The data points within the box
A box plot helps visualize a data set's distribution, giving more information than just the mean or median. The box plot below shows the distribution of the percentages of the total population of the United States for each of the states. The box plot shows the median is , whereas the mean is . The box plot shows that the upper of states range from to of the total population, which signi�cantly affects political representation, resource distribution, and other important factors.
10 2.5 Unknown 1 7 50% 2
Q1
Q3
Reset
50 1 1.37% 1.96% 25% 2.2% 12.15%
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 33/45
The skew is the difference between the mean and the median. A positive skew means that the distribution is skewed to the right, while a negative skew means that the distribution is skewed to the left. In the box plot below, the skew is , which means that the distribution is skewed to the right.
Figure 3.4.1: Box plot showing U.S. population distribution by state.
PARTICIPATION ACTIVITY 3.4.5: Interpreting a box plot.
The following box plot shows the distribution of the per capita real gross domestic product (GDP) in 2017 of each U.S. state . The per capita GDP for the entire U.S. is .
1.96% − 1.37% = 0.59%
Source U.S. Census Bureau
2 $51,749
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 34/45
1) The real GDP of is a good representation per capita GDP of the United States.
2) The national mean income of is a good representation of each state's individual mean income.
Detecting outliers
One way to detect outliers using a box plot is to determine how far each data element is from either or . The interquartile range (IQR) of a dataset is the difference between and
, or the length of the box in a box plot. A data value greater than or less than is considered an outlier. Often, an outlier is not included in either whisker and is instead represented in the plot as a marker such as an open circle or a triangle.
Example 3.4.3: The interquartile range.
For the dataset , and . What is the IQR, and does the dataset contain any outliers?
Solution
Source: U.S. Bureau of Economic Analysis
$51,749
True
False
$51,749
True
False
Q1 Q3 Q3 Q1 (Q3 − Q1) Q3 + 1.5(IQR)
Q1 − 1.5(IQR)
3, 37, 23, 61, 36, 65, 6, 24, 1, 19, 72, 1, 13, 40, 1 Q1 = 4.5 Q3 = 38.5
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 35/45
The IQR is . An outlier is either greater than or less than . Since no data values are greater than
or less than , no outliers exist.
PARTICIPATION ACTIVITY 3.4.6: Detecting outliers.
mtcars is a historical dataset from a 1974 issue of Motor Trend comparing the performance of cars. The �rst few rows of the data are given below as well as the �ve number summary.
Unnamed: 0 mpg cyl disp hp drat wt qsec vs am gear carb 0 Mazda RX4 21.0 6 160.0 110 3.90 2.620 16.46 0 1 4 4 1 Mazda RX4 Wag 21.0 6 160.0 110 3.90 2.875 17.02 0 1 4 4 2 Datsun 710 22.8 4 108.0 93 3.85 2.320 18.61 1 1 4 1 3 Hornet 4 Drive 21.4 6 258.0 110 3.08 3.215 19.44 1 0 3 1 4 Hornet Sportabout 18.7 8 360.0 175 3.15 3.440 17.02 0 0 3 2
min 1.513000 25% 2.581250 50% 3.325000 75% 3.610000 max 5.42400
1) What is the interquartile range for the weights data?
2) What is the upper bound for the whiskers?
3) What is the lower bound for the whiskers?
4) Is a car with a weight of (
38.5 − 4.5 = 34 Q3 + 1.5(IQR) = 38.5 + 1.5(34) = 89.5 Q1 − 1.5(IQR) = 4.5 − 1.5(34) = −46.5 89.5
−46.5
32
3.911
−0.108
1.029
1.544
5.154
5.424
1.038
2.581
1.513
5.250 5, 250
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 36/45
lbs) an outlier?
5) Is a car with a weight of ( lbs) an outlier?
References
(*1) United States Census Bureau. "State Population Totals and Components of Change: 2010-2017." census.gov. 8 May 2018. Web. 4 Jun. 2018.
(*2) United States Bureau of Economic Analysis. "Per capita real GDP by state (chained 2009 dollars)" bea.gov. 4 May 2018. Web. 4 Jun. 2018.
Yes
No
1.513 1, 513
Yes
No
3.5 Histograms
Histograms with evenly-sized bins
A frequency distribution is a table that displays how often an outcome occurs for a sample. To construct a frequency distribution, the data set is divided into mutually exclusive classes. A class is either a value of a categorical variable or an interval of a continuous variable. The frequency of a class is the number of events or values that fall under each class. Ex: An informal poll among a group of friends tallies how many people have gaming applications on their phone. The results of the poll can be summarized in the frequency distribution below.
Table 3.5.1: Frequency distribution showing the number of people in a group having gaming apps on their phone.
Gaming apps Tally Frequency
||||
|||| ||
||||
x
x
(x)
0 4
1 7
2 5
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 37/45
|||
||
The most common graphical representation of a frequency distribution is a histogram. A histogram depicts data values by splitting a continuous variable into a number of class intervals, each known as a bin. The simplest and most common type of histogram has bins of equal size. When bin sizes are equal, bins have rectangular bars with heights representing the frequency, which is the number of values in a bin.
The -axis contains a continuous number line with ticks that represent bin boundaries. A bin includes values equal to or greater than the lower boundary, but less than the upper boundary (lower value
upper). Gaps between rectangles are removed to show that the data is continuous.
PARTICIPATION ACTIVITY 3.5.1: Histogram of number of tickets per miles per hour over speed limit.
PARTICIPATION ACTIVITY 3.5.2: Histogram fundamentals.
Consider the following histogram showing speeding tickets issued by O�cer Brown.
1) speeding tickets were issued
3 3
4 2
x ≤
<
Animation captions:
1. Data of MPH over the speed limit for tickets can be represented as a histogram. 2. The -axis of the histogram is continuous with evenly-spaced bin intervals. In this case, each
interval is MPH, but other intervals are possible. 3. The -axis represents the frequency, or the number of tickets. 4. Each bin frequency is represented by a rectangular column.
15 x
5 y
3
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 38/45
between mph.
2) The most speeding tickets were given in the mph range.
3) The number of speeding tickets issued between mph is unknown.
4) A ticket issued that is mph over the speed limit should be placed in the
mph bin.
Histogram bin size
A key goal of a histogram is to estimate the probability density function of the continuous variable on the -axis. In short, the goal is to �t a smooth curve over the most rectangles, while minimizing the white space under the curve.
When creating a histogram, multiple bin sizes should be attempted to determine the best distribution of the data. A good rule of thumb is to start with a bin size so that the number of bins is roughly equal to the square root of the number of values. Ex: For O�cer Smith's tickets seen in the animation above, a good number of bins to start with are bins. Since the tickets are as much as mph over the limit, a good initial bin size would be mph bins mph bins.
The �gure below shows red distribution curves and histograms for O�cer Smith's tickets for various bin sizes. Histogram-1 and Histogram-2 contain many gaps between bins, leaving too much white space under the curve. Thus, bin sizes 1 and 2 are not good options. Histogram-15 leaves little white space under the curve, but offers little insight about O�cer Smith's ticketing trends. Finally, when compared to Histogram-10, Histogram-5 shows less white space under the curve, and thus, is the best option.
Figure 3.5.1: Histogram bin-size comparison for speeding tickets issued by O�cer Smith.
5 − 10
True
False
15 − 20
True
False
0 − 5
True
False
5
0 − 5
True
False
x
15 = 3.9 ≈ 415
−−√ 28 28 /4 = 7
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 39/45
Several basic distribution patterns should be looked for when selecting a bin size, because some situations are known to follow a certain distribution. The standard bell curve is a unimodal distribution pattern. A unimodal distribution occurs when there is one (uni) prevalent peak (mode) in the histogram. Ex: In Histogram-5, the mph bin has the highest frequency of all bins, and thus, is the single mode.
Other common distribution patterns are listed below.
Bimodal: Contains two prevalent modes Multimodal: Contains multiple prevalent modes Skewed left: Contains a mode on the right with a tail of low-frequency bins on the left
10 − 15
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 40/45
Skewed right: Contains a mode on the left with a tail of low-frequency bins on the right
PARTICIPATION ACTIVITY 3.5.3: Common histogram distributions.
The following histograms show common distributions to look for when selecting bin size:
Unimodal
Bimodal
Multimodal
Skewed left
Skewed right
Histograms with unevenly-sized bins
Histogram bins are not always equally sized. Ex: Consider the table below containing data for 2014 motor vehicle crash deaths. Most age data is represented in year intervals (green). However, some age intervals are larger or smaller than years (red). Thus, a histogram representing the crash data cannot have equally-sized bins.
Table 3.5.2: Insurance Institute for Highway
Histogram b Histogram a Histogram e Histogram d Histogram c
Reset
5 5
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 41/45
Safety (IIHS) motor vehicle crash deaths per people data, showing unequal
interval sizes.
Age Deaths
Green: -year bins Red: Non- -year bins
100, 000
0 − 12 872
13 − 15 380
16 − 19 2, 243
20 − 24 4, 047
25 − 29 3, 250
30 − 34 2, 567
35 − 39 2, 155
40 − 44 2, 067
45 − 49 2, 196
50 − 54 2, 712
55 − 59 2, 414
60 − 64 1, 976
65 − 69 1, 517
70 − 74 1, 228
75 − 79 1, 107
80 − 84 872
85+ 985
5 5
Source: IIHS fatality facts, 20141
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 42/45
The �gure below shows an incorrect histogram for the IIHS crash data with rectangles of equal width, despite different bin sizes. The �rst red bin represents ages and the second red bin represents ages, while the blue bins represent ages.
By comparing bin heights, the �gure's histogram correctly indicates that year-olds are more likely to die in a crash than year-olds. Comparing the two bins is reasonable because both bins have the same bin size: years. However, using the same rectangle width to represent bins of different sizes can lead to incorrect conclusions about the likelihood of each bin. Ex: The histogram visually indicates, incorrectly, that a year-old is more likely to die in a crash than a year-old (see correct histogram further below).
Figure 3.5.2: Incorrect histogram for IIHS crash data, showing 2014 motor vehicle crash deaths per people.
To compare likelihoods of two unequally-sized bins, a unit area histogram must be created. A unit area histogram has rectangle heights equal to the bin frequency divided by the bin size. The following table shows rectangle heights for a unit area histogram being computed for the IIHS crash data.
Table 3.5.3: IIHS motor vehicle crash deaths per people data, showing rectangle
heights (deaths per age) being computed for a unit area histogram.
Age Deaths Bin size Deaths per age
13 3 5
20 − 24 25 − 29
5
10 15
100, 000
100, 000
0 − 12 872 13 872/13 = 67
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 43/45
The �gure below shows the unit area histogram for the IIHS crash data. Notice that the y-axis has changed to re�ect the Deaths per age metric and that some bins are different sizes. With different bin sizes, bin frequency is determined by rectangle area, instead of rectangle height. Ex: The unit area histogram shows deaths per age for year-olds: deaths per age years in bin
deaths for year-olds.
The incorrect histogram above visually suggested that a year-old is more likely to die in a crash than a year-old. Even though the year-old bin frequency is higher than the year-old bin frequency , the unit area histogram shows that a child between the ages of years-old is less likely (shorter rectangle) to die in a crash than a year-old child.
Figure 3.5.3: Correct unit area histogram for IIHS crash data, showing 2014 motor vehicle crash deaths per people.
13 − 15 380 3 127
16 − 19 2, 243 4 561
20 − 24 4, 047 5 809
25 − 29 3, 250 5 650
30 − 34 2, 567 5 513
35 − 39 2, 155 5 431
40 − 44 2, 067 5 413
45 − 49 2, 196 5 439
50 − 54 2, 712 5 542
55 − 59 2, 414 5 483
60 − 64 1, 976 5 395
65 − 69 1, 517 5 303
70 − 74 1, 228 5 246
75 − 79 1, 107 5 221
80 − 85 872 5 174
85+ 985 15 66
650 25 − 29 (650 )×(5 )= 3, 250 25 − 29
10 15 0 − 12 (872) 13 − 15
(380) 0 − 12 13 − 15
100, 000
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 44/45
PARTICIPATION ACTIVITY 3.5.4: Unit area histogram.
Consider the following unit area histogram showing tips earned for a waitress:
1) The waitress is _______ likely to earn a tip than a tip.$5 − $20 $25 − $30
less
equally
more
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
2/3/2021 QNT/275T: Statistics for Decision Making home
https://learn.zybooks.com/zybook/QNT_275T_54402574/chapter/3/print 45/45
2) The waitress is _______ likely to earn a tip than a tip.
3) The waitress earned a _______ tip more often than any other tip.
4) The waitress is most likely to earn a _______ tip.
References
(*1) "General Statistics" Insurance Institute for Highway Safety Highway Loss Data Institute 2014, http://www.iihs.org/iihs/topics/t/general-statistics/fatalityfacts/overview-of-fatality-facts.
$10 $27
less
equally
more
to $5 $20
to $20 $25
to $35 $40
to $5 $10
to $20 $25
to $35 $40
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574
©zyBooks 02/03/21 13:10 922949 Julio Romero
QNT_275T_54402574