Statistic in Health Care Management Week 2

profilegregueira82
chapter4.pdf

Chapter 4

Summarizing Data Collected in

the Sample

Learning Objectives

• Distinguish between dichotomous, ordinal,

categorical, and dichotomous variables

• Identify appropriate numerical and graphical

summaries for each variable type

• Compute a mean, median, standard deviation,

quartiles and range for a continuous variable

Learning Objectives

• Construct a frequency distribution table for

dichotomous, categorical and ordinal variables

• Provide an example of when the mean is a

better measure of location than the median

• Interpret the standard deviation of a continuous

variable

Learning Objectives

• Generate and interpret a box plot for a

continuous variable

• Produce and interpret side-by-side box plots

• Differentiate between a histogram and a bar

chart

Variable Types

• Dichotomous variables have 2 possible responses

(e.g., Yes/No)

• Ordinal and categorical variables have more than two

responses and responses are ordered and unordered,

respectively

• Continuous (or measurement) variables assume in

theory any values between a theoretical minimum and

maximum

Biostatistics

Two Areas of Applied Biostatistics:

Descriptive Statistics

– Summarize a sample selected from a population

Inferential Statistics

– Make inferences about population parameters based on sample statistics.

Vocabulary

• Data elements/data points

• Subjects/units of measurement

• Population Vs. Sample

Sample vs Population

• Any summary measure computed on a sample

is a statistic

• Any summary measure computed on a

population is a parameter

n = sample size

N = population size

Example 4.1.

Dichotomous Variable

Frequency Distribution Table

Hypertension Treatment

Frequency Relative Frequency (%)

No 2313 65.5%

Yes 1219 34.5%

3532 100.0%

Relative Frequency Bar Chart for

Dichotomous Variable

Categorical Outcome

Sample: n=50

Population: Patients at health center

Variable: Marital status

Marital Status Number of Patients

Married 24

Separated 5

Divorced 8

Widowed 2

Never Married 11

Total 50

Categorical Outcome

Frequency Distribution Table

Marital Status Number of

Patients (f)

Relative Frequency (f/n)

Married 24 0.48

Separated 5 0.10

Divorced 8 0.16

Widowed 2 0.04

Never Married 11 0.22

Total 50 1.00

Frequency Bar Chart

Ordinal Outcome

Sample: n=50

Population: Patients at health center

Variable: Self-reported current health status

Health Status Number of Patients

Excellent 19

Very Good 12

Good 9

Fair 6

Poor 4

Total 50

Ordinal Outcome

Frequency Distribution Table

Heath Status Freq. Rel. Freq. Cumulative Freq

Cumulative Rel. Freq.

Excellent 19 38% 19 38%

Very Good 12 24% 31 62%

Good 9 18% 40 80%

Fair 6 12% 46 92%

Poor 4 8% 50 100%

50 100%

Relative Frequency Histogram

0

5

10

15

20

25

30

35

40

Poor Fair Good Very Good Excellent

Health Status

%

Example 4.2.

Ordinal Variable

Frequency Distribution Table

Blood Pressure Categories

Frequency Relative Frequency (%)

Normal 1206 34.1%

Pre-hypertension 1452 41.1%

Stage I hypertension 653 18.5%

Stage II hypertension 222 6.3%

Total 3533 100.0%

Relative Frequency Histogram for Ordinal

Variable

Continuous Variables

• Assume, in theory, any value between a

theoretical minimum and maximum

• Quantitative, measurement variables

Continuous Variable

• Population: Patients 50 years of age with

coronary artery disease

• Sample: n = 7 patients

• Outcome: Systolic blood pressure (mmHg)

Continuous Variable

Sample data

X

100

110

114

121

130

130

160

Continuous Variable

6.123 7

865

n

X X 

X

100

110

114

121

130

130

160

865

n

X X mean Sample

 

Continuous Variable

Consider a second sample from the same population.

We record SBP on each subject in the second sample:

120 121 122 124 125 126 127

n = 7

= 865 / 7 = 123.6.

What is different between the 2 samples? X

Continuous Variable

• Dispersion

X (X- )

100 -23.6

110 -13.6

114 -9.6

121 -2.6

130 6.4

130 6.4

160 36.4

865 0

X

Continuous Variable

• Dispersion

X (X- )

100 -23.6

110 -13.6

114 -9.6

121 -2.6

130 6.4

130 6.4

160 36.4

865 0

X Mean Absolute Deviation (MAD):

n

| X - X| Σ = MAD

Continuous Variable

X X

1n

)XΣ(X s

2 2

 

374.6 6

2247.72 s

2 

Sample Variance:

X (X- ) (X- )2

100 -23.6 556.96

110 -13.6 184.96

114 -9.6 92.16

121 -2.6 6.76

130 6.4 40.96

130 6.4 40.96

160 36.4 1324.96

865 0 2247.72

Continuous Variable

• Sample Standard Deviation:

s = s 2

4.196.374 s

Standard Summary: n=7, X = 123.6, s=19.4

Median

Median

100 110 114 121 130 130 160

Median holds 50% of values above and 50% of values below

Order data For n odd – median is middle value For n even – median is mean of 2

middle values

Quartiles

Q1 = first quartile holds approximately 25% of the

scores at or below it and

Q3 = third quartile holds approx. 25% of the

scores at or above it

Q2 = ??

Continuous Variable

Median

Order data

100 110 114 121 130 130 160

Q1 Q3

Box and Whisker Plot

100 110 120 130 140 150 160

Min Q1 Median Q3 Max

Comparing Samples with

Box and Whisker Plots

100 110 120 130 140 150 160

Summarizing Location and Variability

• When there are no outliers, the sample mean

and standard deviation summarize location

and variability

• When there are outliers, the median and

interquartile range (IQR) summarize location

and variability, where IQR = Q3-Q1

Example

Sample: n=51 participants in a study of

cardiovascular risk factors.

Variable: age (years)

60 62 63 64 64 65 65 65 65 65 65

66 66 66 66 66 67 67 67 68 68 68

70 70 70 71 71 72 72 73 73 73 73

73 73 75 75 75 76 76 77 77 77 77

77 79 82 83 85 85 87

Example

Sample mean: 71.3 =

51

3637 =

n

XΣ = X

Sample variance:

41.4 = 50

/51)(3637 - 261,439 =

1 -n

/n)X(Σ - XΣ = s

222

2

Sample standard deviation:

6.4 = 41.4 = s

Standard Summary: n=51, X = 71.3, s=6.4

Outliers

IQR = Interquartile Range = Q3 - Q1

= range of middle half of the data

Outliers are values which either:

exceed Q3 + 1.5 IQR, or

fall below Q1 - 1.5 IQR

Or outliers are outside + 3s X

Check for Outliers in Example

• Q1=66, Q3=76, IQR=10

– Lower=66-1.5(10)=51

– Upper=76+1.5(10)=91

• + 3s = 52.1 to 90.5X

Presenting Data

• Suppose we collapse ages into 5 mutually exclusive and

exhaustive categories:

Age Class Number of Individuals (freq.)

60-64 5

65-69 17

70-74 12

75-79 12

80-84 2

85-89 3

Presenting Data

Cumulative

Age Class Freq Rel Freq Freq Rel Freq

60-64 5 0.10 5 0.10

65-69 17 0.33 22 0.43

70-74 12 0.24 34 0.67

75-79 12 0.24 46 0.91

80-84 2 0.04 48 0.95

85-89 3 0.06 51 1.00

Total 51 1.00

Frequency Histogram

0 2 4 6 8

10 12 14 16 18

60-

64

65-

69

70-

74

75-

79

80-

84

85-

89

Age Class

F r e q

u e n

c y

Example 4.3.

Summarizing Continuous Variables

Diastolic blood pressures in n=10 randomly

selected participants attending the seventh

examination of the Framingham Offspring

Study

76 64 62 81 70

72 81 63 67 77

Summarizing Location

• What is a typical diastolic blood pressure?

Sample Mean

= Sum of diastolic blood pressures/n

= 713/10 = 71.3

Notation

• Let X represent the outcome of interest (e.g.,

X=diastolic blood pressure)

n

X X mean Sample

 

Summarizing Variability

• Sample range

= maximum–minimum=81–62 = 19

• Sample variance

1n

)x(x s

2

2

 

Sample Variance

DBP Deviation from Mean

76 (76 - 71.3) = 4.7

64 (64 - 71.3) = -7.3

62 (62 - 71.3) = -9.3

81 9.7

70 -1.3

72 0.7

81 9.7

63 -8.3

67 -4.3

77 5.7

S X = 71.3 S Deviations from Mean = 0

Sample Variance

DBP Deviation from Mean Squared Deviations

76 (76 - 71.3) = 4.7 22.09

64 (64 - 71.3) = -7.3 53.29

62 (62 - 71.3) = -9.3 86.49

81 9.7 94.09

70 -1.3 1.69

72 0.7 0.49

81 9.7 94.09

63 -8.3 68.89

67 -4.3 18.49

77 5.7 32.49

S X = 71.3 S Deviations = 0 S Deviations2 = 472.10

Sample Variance and Sample Standard

Deviation

46.52 9

10.472

1n

)x(x s

2

2 

 

2.746.52 1n

)x(x s

2

 

 

Median

• Median holds 50% of values above and 50%

of values below

– Order data

– For n odd – median is middle value

– For n even – median is mean of 2 middle values

Median = 71

62 63 64 64 70 | 72 76 77 81 81

Quartiles

• Q1 = first quartile = holds 25% of values

below it

• Q3 = third quartile = holds 25% of values

above it

Median = 71

62 63 64 64 70 | 72 76 77 81 81

Q1 Q3

Determining Outliers

• Outliers are values

below Q1-1.5(Q3-Q1) or

above Q3+1.5(Q3-Q1)

• In Example 4.3,

lower limit = 64-1.5(77-64) = 44.5

and upper limit=77+1.5(77-64) = 96.5

• Outliers?

• Mean or Median? s or IQR?

Box Plot for Continuous Variable

60

65

70

75

80

d b p

Numerical and Graphical Summaries

• Dichotomous and categorical

– Frequencies and relative frequencies

– Bar charts (freq. or relative freq.)

• Ordinal

– Frequencies, relative frequencies, cumulative frequencies and cumulative relative frequencies

– Histograms (freq. or relative freq.)

Numerical and Graphical Summaries

• Continuous

– Mean, standard deviation, minimum, maximum, range, median, quartiles, interquartile range

– Box plot