Healthcare Statistics
Chapter 4
Summarizing Data Collected in the Sample
Learning Objectives (1 of 3)
- Distinguish between dichotomous, ordinal, categorical, and continuous variables
- Identify appropriate numerical and graphical summaries for each variable type
- Compute a mean, median, standard deviation, quartiles and range for a continuous variable
Learning Objectives (2 of 3)
- Construct a frequency distribution table for dichotomous, categorical, and ordinal variables
- Provide an example of when the mean is a better measure of location than the median
- Interpret the standard deviation of a continuous variable
Learning Objectives (3 of 3)
- Generate and interpret a box plot for a continuous variable
- Produce and interpret side-by-side box plots
- Differentiate between a histogram and a bar chart
Variable Types
- Dichotomous variables have two possible responses (e.g., yes/no).
- Ordinal and categorical variables have more than two responses, and responses are ordered and unordered, respectively.
- Continuous (or measurement) variables assume in theory any values between a theoretical minimum and maximum.
Biostatistics
- Two areas of applied biostatistics
- Descriptive statistics—summarize a sample selected from a population
- Inferential statistics—make inferences about population parameters based on sample statistics.
Vocabulary
- Data elements/data points
- Subjects/units of measurement
- Population versus sample
Sample vs. Population
- Any summary measure computed on a sample is a statistic.
- Any summary measure computed on a population is a parameter.
n = Sample Size
N = Population Size
Example 4.1.
Dichotomous Variable
Frequency Distribution Table
Relative Frequency Bar Chart for Dichotomous Variable
Sample: n = 50
Population: Patients at health center
Variable: Marital status
Categorical Outcome (1 of 2)
| Marital Status | Number of Patients |
| Married | 24 |
| Separated | 5 |
| Divorced | 8 |
| Widowed | 2 |
| Never married | 11 |
| Total | 50 |
Categorical Outcome (2 of 2)
Frequency Distribution Table
| Marital Status | Number of Patients (f) | Relative Frequency (f/n) |
| Married | 24 | 0.48 |
| Separated | 5 | 0.10 |
| Divorced | 8 | 0.16 |
| Widowed | 2 | 0.04 |
| Never married | 11 | 0.22 |
| Total | 50 | 1.00 |
Frequency Bar Chart
Sample: n =50
Population: Patients at health center
Variable: Self-reported current health status
Ordinal Outcome (1 of 2)
| Health Status | Number of Patients |
| Excellent | 19 |
| Very good | 12 |
| Good | 9 |
| Fair | 6 |
| Poor | 4 |
| Total | 50 |
Ordinal Outcome (2 of 2)
Frequency Distribution Table
| Heath Status | Freq. | Rel. Freq. | Cumulative Freq. | Cumulative Rel. Freq. |
| Excellent | 19 | 38% | 19 | 38% |
| Very good | 12 | 24% | 31 | 62% |
| Good | 9 | 18% | 40 | 80% |
| Fair | 6 | 12% | 46 | 92% |
| Poor | 4 | 8% | 50 | 100% |
| 50 | 100% |
Relative Frequency Histogram
Example 4.2.
Ordinal Variable
Frequency Distribution Table
Relative Frequency Histogram
for Ordinal Variable
- Assume, in theory, any value between a theoretical minimum and maximum
- Quantitative, measurement variables
Continuous Variable (1 of 9)
- Population: Patients 50 years of age with coronary artery disease
- Sample: n = 7 patients
- Outcome: Systolic blood pressure (mmHg)
Continuous Variable (2 of 9)
Sample data
X 100 110 114 121 130
130 160
Continuous Variable (3 of 9)
X 100 110 114 121 130
130 160
865
Continuous Variable (4 of 9)
Consider a second sample from the same population.
We record SBP on each subject in the second sample:
120 121 122 124 125 126 127
n = 7
= 865 / 7 = 123.6.
What is different between the two samples?
Continuous Variable (5 of 9)
*
- Dispersion
Continuous Variable (6 of 9)
| X | (X – ) |
| 100 | –23.6 |
| 110 | –13.6 |
| 114 | –9.6 |
| 121 | –2.6 |
| 130 | 6.4 |
| 130 | 6.4 |
| 160 | 36.4 |
| 865 | 0 |
- Dispersion
Mean absolute
deviation (MAD):
Continuous Variable (7 of 9)
| X | (X – ) |
| 100 | –23.6 |
| 110 | –13.6 |
| 114 | –9.6 |
| 121 | –2.6 |
| 130 | 6.4 |
| 130 | 6.4 |
| 160 | 36.4 |
| 865 | 0 |
- Sample variance
X (X – ) (X – )2
100 –23.6 556.96
110 –13.6 184.96
114 –9.6 92.16
121 –2.6 6.76
130 6.4 40.96
130 6.4 40.96
160 36.4 1324.96
865 0 2247.72
Continuous Variable (8 of 9)
Continuous Variable (9 of 9)
- Sample standard deviation
- Standard summary
n = 7, X = 123.6, s = 19.4
Median
Median
100 110 114 121 130 130 160
- Median—holds 50% of values above and 50% of values below
- Order data
For n odd—median is middle value
For n even—median is mean of two middle values
Quartiles
- Q1 = first quartile holds approximately 25% of the scores at or below it.
- Q3 = third quartile holds approximately 25% of the scores at or above it.
- Q2 = ??
Continuous Variable
Median
Order data
100 110 114 121 130 130 160
Q1
Q3
Box and Whisker Plot
100 110 120 130 140 150 160
Min Q1 Median Q3 Max
Comparing Samples with
Box and Whisker Plots
100 110 120 130 140 150 160
Summarizing Location and Variability
- When there are no outliers, the sample mean and standard deviation summarize location and variability.
- When there are outliers, the median and interquartile range (IQR) summarize location and variability, where IQR = Q3 – Q1.
Sample: n = 51 participants in a study of cardiovascular risk factors.
Variable: age (years)
60 62 63 64 64 65 65 65 65 65 65
66 66 66 66 66 67 67 67 68 68 68
70 70 70 71 71 72 72 73 73 73 73
73 73 75 75 75 76 76 77 77 77 77
79 82 83 85 85 87
Example (1 of 2)
Example (2 of 2)
Sample mean:
Sample variance:
Sample standard deviation:
Standard summary: n = 51, X = 71.3, s = 6.4
Outliers
IQR = Interquartile Range = Q3 – Q1
= Range of middle half of the data
- Outliers are values that either:
- Exceed Q3 + 1.5 IQR
- Fall below Q1 – 1.5 IQR
- Or, are outside ± 3s
Check for Outliers in Example
- Q1 = 66, Q3 = 76, IQR = 10
- Lower = 66 – 1.5(10) = 51
- Upper = 76 + 1.5(10) = 91
- ± 3s = 52.1 to 90.5
Presenting Data (1 of 2)
- Suppose we collapse ages into five mutually exclusive and exhaustive categories
Age Class Number of Individuals (freq.) 60–64 5
65–69 17
70–74 12
75–79 12
80–84 2
85–89 3
Presenting Data (2 of 2)
Cumulative
Age Class Freq. Rel. Freq. Freq. Rel. Freq.
60-64 5 0.10 5 0.10
65-69 17 0.33 22 0.43
70-74 12 0.24 34 0.67
75-79 12 0.24 46 0.91
80-84 2 0.04 48 0.95
85-89 3 0.06 51 1.00
Total 51 1.00
Frequency Histogram
Example 4.3.
Summarizing Continuous Variables
- Diastolic blood pressures in n = 10 randomly selected participants attending the seventh examination of the Framingham Offspring Study
76 64 62 81 70
72 81 63 67 77
Summarizing Location
- What is a typical diastolic blood pressure?
Sample mean:
= Sum of diastolic blood pressures/n
= 713/10 = 71.3
Notation
- Let X represent the outcome of interest (e.g., X = diastolic blood pressure)
Summarizing Variability
- Sample range:
= maximum – minimum = 81 – 62 = 19
- Sample variance:
Sample Variance (1 of 2)
DBP Deviation from Mean
76 (76 – 71.3) = 4.7
64 (64 – 71.3) = –7.3
62 (62 – 71.3) = –9.3
81 9.7
70 –1.3
72 0.7
81 9.7
63 –8.3
67 –4.3
77 5.7
S X = 71.3 S Deviations from Mean = 0
Sample Variance (2 of 2)
DBP Deviation from Mean Squared Deviations
76 (76 – 71.3) = 4.7 22.09
64 (64 – 71.3) = –7.3 53.29
62 (62 – 71.3) = –9.3 86.49
81 9.7 94.09
70 –1.3 1.69
72 0.7 0.49
81 9.7 94.09
63 –8.3 68.89
67 –4.3 18.49
77 5.7 32.49
S X = 71.3 S Deviations = 0 S Deviations2 = 472.10
Sample Variance and
Sample Standard Deviation
Median
- Median holds 50% of values above and 50% of values below
- Order data
- For n odd—median is middle value
- For n even—median is mean of two middle values
Median = 71
62 63 64 64 70 | 72 76 77 81 81
Quartiles
- Q1 = first quartile holds 25% of values below it
- Q3 = third quartile holds 25% of values above it
Median = 71
62 63 64 64 70 | 72 76 77 81 81
Q1 Q3
Determining Outliers
- Outliers—values below Q1 – 1.5(Q3 – Q1) or above Q3 + 1.5(Q3 – Q1)
- In Example 4.3: lower limit = 64 – 1.5(77 – 64) = 44.5 and upper limit = 77 + 1.5(77 – 64) = 96.5
- Outliers?
- Mean or median?
- s or IQR?
Box Plot for Continuous Variable
- Dichotomous and categorical
- Frequencies and relative frequencies
- Bar charts (freq. or relative freq.)
- Ordinal
- Frequencies, relative frequencies, cumulative frequencies, and cumulative relative frequencies
- Histograms (freq. or relative freq.)
Numerical and Graphical
Summaries (1 of 2)
Numerical and Graphical
Summaries (2 of 2)
- Continuous
- Mean, standard deviation, minimum, maximum, range, median, quartiles, interquartile range
- Box plot
0
5
10
15
20
25
30
35
40
PoorFairGoodVery GoodExcellent
Health Status
%
6
.
123
7
865
n
X
X
=
=
=
å
n
X
X
mean
Sample
å
=
=
X
n
|
X
-
X
|
Σ
=
MAD
X
X
1
n
)
X
Σ(X
s
2
2
-
-
=
374.6
6
2247.72
s
2
=
=
s
=
s
2
4
.
19
6
.
374
=
=
s
71.3
=
51
3637
=
n
X
Σ
=
X
41.4
=
50
/51
)
(3637
-
261,439
=
1
-
n
/n
)
X
(
Σ
-
X
Σ
=
s
2
2
2
2
6.4
=
41.4
=
s
0
2
4
6
8
10
12
14
16
18
60-
64
65-
69
70-
74
75-
79
80-
84
85-
89
Age Class
Frequency
1
n
)
x
(x
s
2
2
-
-
=
å
46
.
52
9
10
.
472
1
n
)
x
(x
s
2
2
=
=
-
-
=
å
2
.
7
46
.
52
1
n
)
x
(x
s
2
=
=
-
-
=
å
60
65
70
75
80
dbp
60
65
70
75
80
dbp