STATISTICS REPRESENTATION
1 1. MISLEADING STATISTICAL VISUALIZATIONS
Problem 1.
A bar graph was created to represent the number of cars sold by two different car dealers, A
and B, over the past year. The heights of the bars indicate the number of cars sold, but the scales
on the y-axis are different. Dealer A’s bar is twice as tall as Dealer B’s bar, even though Dealer B
actually sold 10 more cars than Dealer A. Dealer A’s bar is labeled as 100 cars sold.
a) If Dealer A actually sold 50 cars, how many cars did Dealer B sell?
b) If Dealer B actually sold 85 cars, how tall should its bar be in the graph?
c) Discuss why this bar graph is misleading and suggest a better way to represent the data.
Solution 1.
a) Let’s denote the number of cars sold by Dealer B as B. From the information given, we know
that Dealer A’s bar is twice as tall as Dealer B’s, and Dealer A’s bar is labeled as 100 cars sold.
So, we have: 100 = 2B
Solving for B, we get: B=100
2= 50
Thus, Dealer B actually sold 50 cars.
b) Let’s denote the height of the bar representing the number of cars sold by Dealer B as H.
We are given that Dealer B actually sold 85 cars. We can set up a proportion to find the height of
the bar:
85
50 =H
100
Solving for H, we get: H=85
50 ×100 = 170
So, the bar representing Dealer B should be 170 units tall in the graph.
c) This bar graph is misleading because it exaggerates the difference in sales between Dealer
A and Dealer B. Although Dealer A’s bar is taller, Dealer B actually sold more cars. A better way to
represent the data would be to use a consistent scale on the y-axis for both dealers, ensuring that
the height of the bars accurately reflects the number of cars sold by each dealer. Alternatively, a
side-by-side bar graph could be used to clearly compare the sales of both dealers without distorting
the data.
2 2. DATA DUMPING AND SKEWING ANALYSIS
Problem 2.
Suppose a dataset contains the following values: 5, 8, 10, 12, 15, 18, 20, 22, 25. Determine
the measures of central tendency and dispersion for this dataset.
Solution 2.
a) To find the measures of central tendency:
•Mean (¯x) = 5+8+10+12+15+18+20+22+25
9=135
9= 15
•Median = The middle value of the dataset, since the dataset is in ascending order, the median
is the 5th value which is 15
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.
•Mode = The value/s that appear/s the most frequently. In this case, there is no mode as all
values appear only once.
b) To find the measures of dispersion:
•Range = 25 −5= 20
•Variance = P(xi−¯x)2
n=(5−15)2+(8−15)2+...+(25−15)2
9=(−10)2+(−7)2+(−5)2+(−3)2+02+32+52+72+102
9
=100+49+25+9+0+9+25+49+100
9=366
9= 40.67 (rounded to 2 decimal places)
•Standard Deviation = √V ariance =√40.67 = 6.38 (rounded to 2 decimal places)
Therefore, the measures of central tendency for the dataset are mean = 15, median = 15,
mode = none. And the measures of dispersion for the dataset are range = 20, variance 40.67, and
standard deviation 6.38.
3 3. INACCURATE SAMPLING TECHNIQUES
Problem 3.
A researcher wants to estimate the average height of students in a university. She decides
to use convenience sampling by taking the heights of the first 30 students she encounters in the
student center.
Assume the true average height of all students in the university is 65 inches with a standard
deviation of 3 inches.
a) Calculate the bias in the researcher’s estimate of the average height based on this conve-
nience sample.
b) Determine the standard error of the estimate based on this convenience sample.
Solution 3.
a) The bias in the estimate of the sample average height can be calculated as the difference
between the expected value of the sample average and the true population average height.
The expected value of the sample average height is equal to the true population average height,
which is 65 inches.
Therefore, the bias in the estimate is 0 inches.
b) The standard error of the estimate is given by the formula:
SE =σ
√n
where σ= 3 is the population standard deviation and n= 30 is the sample size.
Substitute the values,
SE =3
√30 =3
√30 ≈3
5.48 ≈0.55
Thus, the standard error of the estimate based on this convenience sample is approximately
0.55 inches.
4 4. MISINTERPRETATION OF CONFIDENCE INTERVALS
Problem 4. In a study investigating the average height of students at a university, a 95%
confidence interval for the mean height was reported as (165.2 cm, 170.8 cm).
a) Misinterpret this confidence interval.
b) What does the interpretation of this confidence interval actually mean?
c) If another study reported a 90% confidence interval for the mean height as (166.0 cm, 170.0
cm), can we be more confident in the findings of the first study?
Solution 4.
a) The misinterpretation of this confidence interval would be to say that there is a 95% probability
that the true mean height of all students at the university lies between 165.2 cm and 170.8 cm. This
is incorrect because the true mean either falls within this interval or it does not; it is not a matter of
probability.
b) The correct interpretation of a 95% confidence interval is that if we were to take multiple
samples and construct confidence intervals using the same method, we would expect about 95%
of those intervals to contain the true population mean height of the students at the university.
c) The width of the confidence interval is a reflection of the precision of the estimate. Since the
95% confidence interval from the first study is narrower than the 90% confidence interval from the
second study, we can say that the first study provides a more precise estimate of the mean height
of the students. However, the level of confidence is not a measure of the accuracy of the estimate,
so we cannot necessarily say that the first study’s findings are more reliable based solely on the
confidence level.
4.1 5. BIASED SURVEY QUESTIONS
Problem 5. A company conducted a survey to determine customer satisfaction with their prod-
uct, but the survey question was biased. The question asked was "How much do you love our
amazing product?" with answer choices ranging from "A lot" to "Not at all."
According to the company’s records, 80
a) What is the potential bias in this survey question? b) How could the survey question be
rephrased to reduce bias? c) If the survey results are used to make decisions about product im-
provements, what potential issues could arise?
Solution 5. a) The potential bias in this survey question is that it is leading and suggestive. By
using language like "amazing product" and providing extreme options like "A lot" and "Not at all,"
respondents may feel pressured to give positive feedback even if they do not truly feel that way.
b) To reduce bias, the survey question could be rephrased to be more neutral and open-ended.
For example, the question could be changed to "Please share your thoughts on our product. What
do you like about it, and what areas do you think could be improved?"
c) If the biased survey results are used to make decisions about product improvements, poten-
tial issues could arise because the data collected may not accurately reflect customer satisfaction.
Improvements made based on skewed or misleading survey results could lead to wasted resources
and possibly worsen customer satisfaction if the changes do not address real concerns or prefer-
ences. It is crucial to gather unbiased and meaningful feedback to make informed decisions about
product improvements.
5 6. MISLABELING VARIABLES IN ANALYSIS
Problem 6. A researcher is studying the relationship between study hours and exam scores of
students. She mistakenly labels the study hours variable as "exam scores" and the exam scores
variable as "study hours" in her analysis. As a result, she calculates a correlation coefficient of
-0.85 between the two variables. The true correlation between study hours and exam scores is
actually 0.75.
Solution 6. Let’s denote the true study hours variable as Xand the true exam scores variable
as Y.
a) The researcher labeled the variables incorrectly in her analysis. If she actually calculated
a correlation coefficient of -0.85 between the mislabeled variables, what is the true correlation
coefficient between study hours and exam scores?
b) Can we trust the regression analysis results when the variables are mislabeled?
c) What steps should the researcher take to correct the mislabeling issue?
Solution 6. a) To determine the true correlation coefficient between study hours and exam
scores, we need to consider the relationship between the mislabeled variables and the true vari-
ables. Let’s denote the mislabeled study hours variable as Y′and the mislabeled exam scores
variable as X′.
Given that the researcher calculated a correlation coefficient of -0.85 between X′and Y′, we
have r=−0.85.
Since the variables were mistakenly swapped, the true correlation should be the negative of
this value: rX Y =−rX′Y′= 0.85.
Therefore, the true correlation coefficient between study hours and exam scores is 0.85.
b) We cannot trust the regression analysis results when the variables are mislabeled. Using the
mislabeled variables would lead to incorrect coefficient estimates, standard errors, and hypothesis
tests.
c) To correct the mislabeling issue, the researcher should swap the variables back to their
correct labels. In this case, she should use the true study hours variable as the independent
variable and the true exam scores variable as the dependent variable in her analysis. This will
ensure that the regression analysis produces valid results based on the correct variables.
5.1 7. INCONSISTENCIES IN DATA REPORTING
Problem 7. In a survey about favorite ice cream flavors, 150 people were asked to choose their
top three flavors. The data collected showed the following inconsistencies:
•20 people chose only vanilla as their favorite flavor.
•30 people chose only chocolate as their favorite flavor.
•25 people chose only strawberry as their favorite flavor.
•15 people chose both vanilla and chocolate as their favorite flavors.
•10 people chose both chocolate and strawberry as their favorite flavors.
•8 people chose both vanilla and strawberry as their favorite flavors.
•7 people chose all three flavors as their favorite.
a) How many people selected at least one of the three flavors as their favorite?
b) How many people did not choose vanilla as one of their top three flavors?
c) What percentage of people selected exactly two of the three flavors as their favorite?
Solution 7.
a) To find the number of people who selected at least one of the three flavors, we can sum the
number of people who chose each flavor individually and subtract the overlap between them:
20 + 30 + 25 −15 −10 −8 + 7 = 49
Therefore, 49 people selected at least one of the three flavors as their favorite.
b) To find the number of people who did not choose vanilla as one of their top three flavors, we
can subtract the number of people who chose only vanilla from the total number of people:
150 −20 −15 −8 + 7 = 114
So, 114 people did not choose vanilla as one of their top three flavors.
c) To find the percentage of people who selected exactly two of the three flavors, we need to
add the number of people who chose exactly two flavors, which is the sum of people who chose
each combination of two flavors without overlap:
15 + 10 + 8 = 33
Therefore, the percentage of people who selected exactly two of the three flavors as their fa-
vorite is:
33
150 ×100% = 22%
6 8. IGNORING OUTLIERS IN DATA SETS
Problem 8. Consider a data set with values: {12, 15, 18, 20, 22, 25, 30, 100}.
a) Compute the mean (average) of this data set.
b) Compute the median of the data set.
c) Compute the median after excluding the outlier value from the data set.
Solution 8.
a) To find the mean of the data set, we sum all the values and divide by the total number of
values:
Mean = 12+15+18+20+22+25+30+100
8=242
8= 30.25
Therefore, the mean of the data set is 30.25.
b) To find the median of the data set, we first need to arrange the values in ascending order:
{12, 15, 18, 20, 22, 25, 30, 100}
Since there are 8 values, the median is the average of the two middle values: (20 + 22) / 2 =
21.
Therefore, the median of the data set is 21.
c) Excluding the outlier value (100) from the data set, the adjusted data set is: {12, 15, 18, 20,
22, 25, 30}.
Arranging these values in ascending order:
{12, 15, 18, 20, 22, 25, 30}
Since there are 7 values, the median is the 4th value, which is 20.
Therefore, the median of the data set after excluding the outlier is 20.
7 9. MISUNDERSTANDING CORRELATION VS. CAUSATION
Problem 9. In a study analyzing the relationship between ice cream sales and drowning inci-
dents, the following data was collected for the past 12 months:
Month Ice Cream Sales (in gallons)
1 100
2 110
3 120
4 130
5 140
6 150
7 160
8 170
9 180
10 190
11 200
12 210
Month Drowning Incidents
1 30
2 40
3 50
4 60
5 70
6 80
7 90
8 100
9 110
10 120
11 130
12 140
a) Calculate the correlation coefficient between ice cream sales and drowning incidents.
b) Based on the correlation coefficient, what can you conclude about the relationship between
ice cream sales and drowning incidents?
Solution 9. a) To calculate the correlation coefficient between ice cream sales and drowning
incidents, we can use the formula:
r=nPXY −PXPY
p(nPX2−(PX)2)(nPY2−(PY)2)
Where: - nis the number of data points (12 in this case), - PXY is the sum of the product of
ice cream sales and drowning incidents, - PXis the sum of ice cream sales, - PYis the sum of
drowning incidents, - PX2is the sum of the squares of ice cream sales, - PY2is the sum of the
squares of drowning incidents.
Calculating the necessary values:
XX= 100 + 110 + . . . + 210 = 1680
XY= 30 + 40 + . . . + 140 = 840
XX2= 1002+ 1102+. . . + 2102= 39900
XY2= 302+ 402+. . . + 1402= 8190
XXY = (100 ×30) + (110 ×40) + . . . + (210 ×140) = 33120
Plugging these values into the correlation coefficient formula:
r=12(33120) −(1680)(840)
p(12 ×39900 −16802)(12 ×8190 −8402)
r=331200 −1411200
p(478800 −282240)(97800 −70560)
r=−1080000
p(196560)(27240)
r≈ −0.843
b) The correlation coefficient obtained is approximately -0.843. This negative value indicates
a strong negative correlation between ice cream sales and drowning incidents. However, it is
important to note that correlation does not imply causation. In this case, it is highly unlikely that ice
cream sales directly cause an increase in drowning incidents. Other factors might be influencing
both variables simultaneously.
8 10. CONFUSION BETWEEN MEAN, MEDIAN, AND MODE
Problem 10. The ages of 8 students in a statistics class are as follows: 20, 22, 23, 25, 25, 26,
27, 29.
a) Find the mean, median, and mode of the ages.
Solution 10. a) To find the mean, we add up all the ages and divide by the total number of
students:
Total sum of ages = 20 + 22 + 23 + 25 + 25 + 26 + 27 + 29
= 197.
Number of students (n) = 8
Mean = Total sum of ages
n=197
8= 24.625.
To find the median, we first arrange the ages in ascending order: 20, 22, 23, 25, 25, 26, 27,
29. Since there are 8 ages, the median is the average of the middle two ages, which in this case
are the 4th and 5th ages: 25 and 25.
Median = 25+25
2= 25.
To find the mode, we look for the age that appears most frequently. In this case, the mode is
25 as it appears twice, more than any other age.
Therefore, the mean is 24.625, the median is 25, and the mode is 25.
9 11. FLAWED HYPOTHESIS TESTING
Problem 11. A researcher is investigating the effectiveness of a new drug in lowering blood
pressure. The null hypothesis (H0) is that the drug has no effect, while the alternative hypothesis
(H1) is that the drug does lower blood pressure. The researcher conducts a hypothesis test at a 5
a) State the possible errors that could have occurred in the hypothesis test.
b) Explain which of these errors is more severe in this context.
c) Discuss the implications of the hypothesis test outcome in relation to the effectiveness of the
new drug.
Solution 11.
a) Possible errors that could have occurred in the hypothesis test are:
- Type I error: Rejecting the null hypothesis when it is actually true. - Type II error: Failing to
reject the null hypothesis when it is actually false.
b) In this context, a Type II error (failing to reject the null hypothesis when it is actually false) is
more severe. This is because if the new drug does actually lower blood pressure but the researcher
fails to detect this effect, patients who could benefit from the drug may not receive the treatment
they need. This could have serious health consequences.
c) The outcome of failing to reject the null hypothesis implies that the data did not provide
enough evidence to conclude that the new drug effectively lowers blood pressure. This does not
mean that the drug has no effect, but rather that the evidence was not strong enough to support
the alternative hypothesis. Further studies may be needed to explore the effectiveness of the drug
more thoroughly.
10 12. OVERRELIANCE ON P-VALUES
Problem 12. A researcher conducted an experiment to test the effectiveness of a new drug for
lowering blood pressure. The researcher obtained a p-value of 0.03 for the hypothesis test. The
significance level was set at 0.05.
a) Interpret this p-value in the context of the hypothesis test.
b) What conclusion can be drawn based on the p-value and significance level?
c) Discuss the limitations of relying solely on p-values for drawing conclusions in hypothesis
testing.
Solution 12.
a) The p-value of 0.03 means that if the null hypothesis (usually stating that there is no effect or
difference) is true, there is a 3% chance of obtaining the observed results or more extreme results
by random chance alone.
b) Since the p-value of 0.03 is less than the significance level of 0.05, we reject the null hypoth-
esis. This suggests that there is evidence to support the alternative hypothesis, which in this case
could be that the new drug is effective in lowering blood pressure.
c) Relying solely on p-values for drawing conclusions in hypothesis testing has limitations. P-
values do not provide information about the effect size or the practical significance of the results.
Additionally, p-values can be influenced by sample size, leading to statistically significant results
that may not be practically significant. It is important to consider other factors such as effect size,
confidence intervals, and the context of the research when interpreting the results of hypothesis
tests.
11 13. MISAPPLICATION OF NORMAL DISTRIBUTION
Problem 13. The average weight of a large population of adult horses is known to be 1200 pounds
with a standard deviation of 150 pounds. A horse farm claims that their horses are on average
heavier than the population average. To test this claim, a random sample of 25 horses from the
farm was selected and their weights were recorded. The sample mean weight was found to be
1230 pounds. Assuming that the weights of horses at the farm are normally distributed, conduct a
hypothesis test at a significance level of 0.05 to determine if there is enough evidence to support
the farm’s claim.
Solution 13.
a) We need to set up the null and alternative hypotheses:
Null Hypothesis (H0): The average weight of the farm horses is the same as the population
average, H0:µ= 1200 pounds.
Alternative Hypothesis (H1): The average weight of the farm horses is heavier than the popu-
lation average, H1:µ > 1200 pounds.
b) Calculate the test statistic:
The test statistic for a one-sample Z-test for the mean is given by:
Z=¯
X−µ0
σ
√n
where ¯
Xis the sample mean, µ0is the population mean, σis the population standard deviation,
and nis the sample size.
Plugging in the values, we get:
Z=1230 −1200
150
√25
=30
30 = 1
c) Determine the critical value and make a decision:
At a significance level of 0.05 for a one-tailed test, the critical value for Z is approximately 1.645
(from the Z-table).
Since our calculated Z-value of 1 is less than the critical value of 1.645, we fail to reject the null
hypothesis. There is not enough evidence to support the claim that the farm horses are heavier on
average than the population average.
12 14. OVERSIMPLIFICATION OF DATA ANALYSIS
Problem 14.
A company conducted a survey to determine how satisfied customers were with their service
on a scale of 1 to 5 (1 being very dissatisfied and 5 being very satisfied). The data collected is as
follows:
Satisfaction Rating Number of Customers
1 15
2 25
3 40
4 30
5 20
a) Calculate the mean satisfaction rating.
b) Calculate the median satisfaction rating.
c) Calculate the mode of the satisfaction ratings.
Solution 14.
a) To calculate the mean satisfaction rating, we use the formula:
Mean =PRating ×Frequency
PFrequency
Calculating the mean:
Mean =(1 ×15) + (2 ×25) + (3 ×40) + (4 ×30) + (5 ×20)
15 + 25 + 40 + 30 + 20
Mean =15 + 50 + 120 + 120 + 100
130
Mean =405
130 = 3.11
Therefore, the mean satisfaction rating is 3.11.
b) To calculate the median satisfaction rating, we first arrange the data in ascending order: 1,
1, ..., 2, 2, ..., 3, 3, 3, ..., 4, 4, 4, 4, ..., 5, 5.
Since there are 130 customers in total, the median will be the average of the 65th and 66th
observations. These observations correspond to a satisfaction rating of 3. Therefore, the median
satisfaction rating is 3.
c) The mode of the satisfaction ratings is the value that appears most frequently in the data. In
this case, the mode is 3 since it appears 40 times, which is more frequent than any other rating.
13 15. MISREPRESENTATION OF STANDARD DEVIATION
Problem 15. A sample of 10 students is taken and their heights (in cm) are recorded as follows:
152, 158, 162, 156, 150, 155, 160, 148, 164, 154. Calculate the standard deviation of the heights.
Solution 15.
a) First, calculate the mean height:
¯x=1
n
n
X
i=1
xi
¯x=1
10(152 + 158 + 162 + 156 + 150 + 155 + 160 + 148 + 164 + 154)
¯x=1
10 ×1609 = 160.9cm
b) Next, calculate the squared deviations from the mean for each value:
xixi−¯x(xi−¯x)2
152 -8.9 79.21
158 -2.9 8.41
162 1.1 1.21
156 -4.9 24.01
150 -10.9 118.81
155 -5.9 34.81
160 -0.9 0.81
148 -12.9 166.41
164 3.1 9.61
154 -6.9 47.61
c) Now, calculate the variance:
s2=1
n−1
n
X
i=1
(xi−¯x)2
s2=1
9(79.21+8.41+1.21+24.01+118.81+34.81+0.81+166.41+9.61+47.61) = 490.9
9≈54.54 cm2
d) Finally, calculate the standard deviation:
s=√s2=√54.54 ≈7.39 cm
Therefore, the standard deviation of the heights of the 10 students is approximately 7.39 cm.
14 16. LACK OF COMMUNICATION IN DATA INTERPRETATION
Problem 16. Consider a dataset of exam scores for 10 students: 78, 85, 92, 64, 71, 89, 80,
83, 76, 95.
a) Calculate the mean and median of the dataset.
b) Calculate the range and interquartile range of the dataset.
c) Determine if there are any outliers in the dataset using the 1.5 * IQR rule.
Solution 16.
a) To calculate the mean of the dataset:
Mean =78 + 85 + 92 + 64 + 71 + 89 + 80 + 83 + 76 + 95
10 =797
10 = 79.7
To calculate the median, first arrange the dataset in ascending order: 64, 71, 76, 78, 80, 83,
85, 89, 92, 95.
Since there are 10 values, the median is the average of the middle two values, which are 80
and 83.
Therefore, Median =80+83
2= 81.5.
b) To calculate the range, subtract the minimum value from the maximum value:
Range = 95 - 64 = 31.
To calculate the interquartile range (IQR), first find the first quartile (Q1) and third quartile (Q3).
Q1 is the median of the lower half of the dataset: 71.
Q3 is the median of the upper half of the dataset: 89.
IQR=Q3-Q1=89-71=18.
c) First, calculate 1.5 times the IQR:
1.5×18 = 27.
Any value below Q1 - 27 or above Q3 + 27 is considered an outlier. In this dataset, the values
outside this range are 64 and 95. Therefore, both 64 and 95 are considered outliers.
I. The following data represents the number of hours spent studying for an exam by a group of
students:
9, 12, 14, 11, 10, 8, 7, 13
a) Calculate the mean number of hours spent studying. b) Calculate the median number of
hours spent studying.
II. A survey was conducted to determine the number of siblings each student in a class has.
The data collected is as follows:
2, 1, 3, 4, 1, 2, 3, 0, 1, 4
a) Calculate the range of the number of siblings. b) Calculate the mode of the number of
siblings.
III. The following data represents the ages of participants in a marathon race:
25, 29, 32, 27, 28, 31, 38, 40
a) Calculate the standard deviation of the ages. b) Determine the value of the variance.
IV. The weights (in kg) of a group of students are as follows:
56, 62, 58, 54, 60, 65, 70
a) Calculate the mean weight of the students. b) Determine the standard deviation of the
weights.
V. The number of goals scored by a football team in their last 10 matches are as follows:
2, 3, 1, 4, 0, 2, 3, 1, 1, 2
a) Calculate the median number of goals scored. b) Determine the mode of the number of goals
scored.
VI. The following data represents the heights (in cm) of a group of individuals:
160, 165, 170, 155, 175, 180, 168, 172
a) Calculate the mean height of the group. b) Determine the range of the heights.
15 17. ERROR IN PROBABILITY CALCULATIONS
Solution I. a) To calculate the mean number of hours spent studying, we add up all the hours and
divide by the total number of students: Mean = (9 + 12 + 14 + 11 + 10 + 8 + 7 + 13) / 8 Mean = 84
/ 8 Mean = 10.5 hours
b) To calculate the median number of hours spent studying, we first arrange the data in ascend-
ing order: 7, 8, 9, 10, 11, 12, 13, 14 Since there are 8 data points, the median is the average of
the 4th and 5th values: Median = (10 + 11) / 2 Median = 10.5 hours
Therefore, the mean number of hours spent studying is 10.5 hours, and the median number of
hours spent studying is also 10.5 hours.
Continue with similar detailed solutions for the other problems.
16 18. IMPROPER USE OF STATISTICAL SOFTWARE
Problem 18. A researcher is analyzing data on the heights (in inches) of 10 students in a class.
The data is as follows:
63, 67, 70, 66, 61, 72, 68, 62, 69, 65
a) Calculate the mean height of the students.
b) Calculate the median height of the students.
c) Calculate the standard deviation of the heights.
Solution 18.
a) To calculate the mean height, we sum up all the heights and divide by the total number of
students:
Mean = 63 + 67 + 70 + 66 + 61 + 72 + 68 + 62 + 69 + 65
10 =703
10 = 70.3inches.
Therefore, the mean height of the students is 70.3 inches.
b) To calculate the median height, we first rearrange the heights in ascending order:
61, 62, 63, 65, 66, 67, 68, 69, 70, 72.
Since we have an even number of data points, the median is the average of the two middle
values:
Median = 66 + 67
2= 66.5inches.
Therefore, the median height of the students is 66.5 inches.
c) To calculate the standard deviation, we first calculate the variance:
Variance = P(Xi−¯
X)2
n−1.
Where P(Xi−¯
X)2= (63 −70.3)2+ (67 −70.3)2+... + (65 −70.3)2.
After calculating the variance, we take the square root to find the standard deviation:
Standard Deviation = √V ariance.
After calculating, we find that the standard deviation is approximately 3.51 inches.
17 19. INADEQUATE EXPLANATION OF DATA SOURCES
Problem 19. A researcher is studying the relationship between hours of study and exam scores.
She collects data from a group of 20 students. The dataset is as follows:
Hours of Study (x) Exam Score (y)
3 65
5 70
6 72
8 75
a) Calculate the mean hours of study and mean exam score.
b) Determine the median hours of study and median exam score.
c) Find the standard deviation of hours of study and exam scores.
Solution 19. a) To calculate the mean hours of study and mean exam score, we use the
formula:
Mean =Pvalues
number of values
For hours of study:
Mean hours of study =3+5+6+8
4=22
4= 5.5
For exam scores:
Mean exam score =65 + 70 + 72 + 75
4=282
4= 70.5
So, the mean hours of study is 5.5 and the mean exam score is 70.5.
b) To determine the median hours of study and median exam score, we arrange the data in
ascending order: Hours of study: 3, 5, 6, 8 Exam scores: 65, 70, 72, 75
Median for hours of study: Median = (5 + 6) / 2 = 5.5 Median for exam scores: Median = (70 +
72) / 2 = 71
Therefore, the median hours of study is 5.5 and the median exam score is 71.
c) To find the standard deviation, we use the formula:
Standard Deviation =rP(Xi−¯
X)2
n
For hours of study:
X(Xi−¯
X)2= (3 −5.5)2+ (5 −5.5)2+ (6 −5.5)2+ (8 −5.5)2
= 2.52+ 0.52+ 0.52+ 2.52= 6.25 + 0.25 + 0.25 + 6.25 = 13
Standard deviation of hours of study =r13
4≈1.8028
For exam scores:
X(Yi−¯
Y)2= (65 −70.5)2+ (70 −70.5)2+ (72 −70.5)2+ (75 −70.5)2
= 5.52+ 0.52+ 1.52+ 4.52= 30.25 + 0.25 + 2.25 + 20.25 = 53
Standard deviation of exam scores =r53
4≈3.863
Therefore, the standard deviation of hours of study is approximately 1.8028 and the standard
deviation of exam scores is approximately 3.863.
18 20. LIMITED CONSIDERATION OF SAMPLING BIAS
Problem 20. A company is conducting a survey to estimate the average income of residents in
a particular city. They randomly select 100 residents from a database of 1,000 residents. However,
the database is known to have a bias towards higher-income residents, with 60% of the residents
having an income above $50,000. The sample average income is found to be $55,000.
a) Calculate the sampling bias in the survey.
b) Given the sample average income, suggest a method to adjust the estimate for the true
average income of all residents in the city.
Solution 20.
a) The sampling bias in the survey can be calculated by comparing the proportion of higher-
income residents in the sample to the proportion in the population.
In the sample: Proportion of residents with income above $50,000 = 1 - 0.6 = 0.4
In the population: Proportion of residents with income above $50,000 = 0.6
Sampling bias = Population proportion - Sample proportion
Sampling bias = 0.6 - 0.4 = 0.2 or 20%
Therefore, the survey has a sampling bias of 20%.
b) One way to adjust the estimate for the true average income is by using weighting.
Since the sample is biased towards higher-income residents, we can assign weights to each
respondent based on their income category. This means giving more weight to responses from
lower-income residents and less weight to responses from higher-income residents.
Let’s say we assign a weight of 0.8 to responses from lower-income residents and a weight of
1.2 to responses from higher-income residents.
Adjusted average income = (0.8 * Avg. income of lower-income residents + 1.2 * Avg. income
of higher-income residents) / (0.8 * Number of lower-income residents + 1.2 * Number of higher-
income residents)
This adjusted average income would give a more accurate estimate by accounting for the sam-
pling bias towards higher-income residents.