250 words/ two scholarly sources due 12/2!
AMWA Journal / V30 N1 / 2015 / amwa.org 31
by Thomas M. schindler, PhD/ Head Medical Writing Europe, Boehringer Ingelheim Pharma GmbH & Co. KG, Biberach, Germany
Meaning It! a Refresher on Mean, Median, and Mode
Y ou are likely to have used them a hundred times in
your writing. The terms mean, median, and mode are
in almost every text about clinical studies or medi-
cal investigations. Writers may be so used to these terms that
they sometimes do not fully appreciate their very specific
meanings (Table). This short piece is designed to serve as a
little reminder of the underlying concepts and the limita-
tions of these terms.
Medical writers are frequently called upon to describe
data and draw conclusions from the analysis of data. Very
often, 2 or more sets of data are compared. The most direct
way to compare groups is by listing their values side by side.
This way, all data are presented, and the reader can compare
them one by one. However, such a procedure is only practi-
cal for very small groups, and even then the comparison of
individual values is not very informative.
What we really want is to summarize data of groups into
single measures and then use these measures to compare the
groups. To be meaningful, these measures need to be repre-
sentative of the group, ie, they need to embody the center or
focal point of the data. When we plot the data, we can visual-
ize their distribution. We can then see whether the data clus-
ter around a particular value. We can also see to which extent
the data vary around the central point.
Thus, the measure of central tendency is the single value
that best represents all of the data. There are 3 major mea-
sures of central tendency: the mean, the median, and the
mode. When we use one of these measures, we abstract
from the individual observations to obtain a single measure
for an entire set of data. The measures of central tendency
are a kind of shorthand for the distribution of the data. For
an appropriate description of numerical data, we also need
a measure of variation or dispersion around the central
tendency.
The Mode Let’s start with the one that is least used in biomedical texts,
the mode. The mode is actually a very handy method to
describe the central tendency when used for the appropri-
ate kind of data. The mode is calculated by counting how
often each value in our data occurs. Once we have identi-
fied the most frequently occurring value, we have identified
the mode. The mode is best used to describe categorical (ie,
non-numerical) data. Examples for categorical data are dis-
ease stages, such as the New York Heart Association stages
of heart failure (I to IV ), blood groups, or sex. When several
values appear equally often in our data, we have several
modes, and the initial purpose to have one single value char-
acterizing our data is betrayed (eg, bi-modal distributions
with 2 peaks). Likewise, the mode would be unhelpful if each
value in our data appears only once. This is usually the case
with continuous numerical data such as height or clinical
laboratory values. In these cases, we need to choose another
measure of central tendency, eg, the median or the mean.
Because the mode represents the most frequently occur-
ring value, it does not need to be accompanied by a measure
of dispersion. If we use the mode, we have decided that the
most frequent value in our data is the best representation of
the entire data set, and therefore as a measure of dispersion
it is not really helpful. However, we should always provide
the number of potential categories (for NYHA it would be 4)
should the number of categories not be obvious.
The Median To find the median, we first need to order our data, from low
to high or vice versa. We then look for the one data point
right in the middle, with 50% of the values above and 50%
below, and this is our median. The median divides our list in
half. In a data set with an odd number of data, the median
YoUR sTATs REFREshER!
32 AMWA Journal / V30 N1 / 2015 / amwa.org
YOuR sTaTs REFREsHER!
value is the true middle value. If we have an even number of
data, the median value is the mean of the 2 middle values. The
median is best used to describe continuous numerical data,
such as age, height, or heart rate. The median is the best rep-
resentation of the central tendency when the data are not nor-
mally distributed, ie, they do not form a nice symmetrical and
bell-shaped curve when plotted. Instead the resulting curve
might be leaning to either side—that is, be skewed to the left
or the right (Figure). The median is very robust because it is
not affected by extreme values (outliers), and it is indepen-
dent of the shape of the distribution. There are several appro-
priate measures of dispersion that could be reported with the
median. Most simply, the minimum and maximum value could
be reported to provide the reader with the overall range of the
data. As the median is the 50% point of a distribution, a good
indication of the spread is the interquartile range. This is the
range of the values between the 25th and 75th percentile of a
distribution (ie, a subtraction of the value at 25% of the data
from the value at 75% of the data).
The Mean The mean (also called the arithmetic mean or the average) is
probably the most widely used measure of central tendency.
It has the advantage of taking all data into account, both the
number of observations and all their values. The mean is the
result of adding all the values in the data and dividing the
total by the number of data points. The mean is well suited to
describe continuous numerical data such as height, weight, or
blood pressure. Since the mean takes every value into account,
it is sensitive to extreme values or outliers. If extreme values
are entered into the calculation, the resulting mean will not
accurately represent the central tendency. Furthermore, the
mean is only then an appropriate representation of the cen-
tral tendency when the data are normally distributed, ie, when
the shape of the plotted values resembles the outline of a bell
(Figure). The shape needs to be symmetrical with the mean in
the middle. We need to realize that we might not have a single
representative of the mean value in our observations. If we
measured 150 participants and their mean height was 168.4
cm, we may have not any individual who is exactly that tall.
How well the mean represents the central tendency can be
evaluated when we also know the measure of dispersion.
Figure. a) Data that have a normal distribution. In this case mean, median, and mode all have the same value. b) Data are negatively skewed. The data cluster toward the right side of the x-axis; mean, median, and mode are different. c) Data are positively skewed. The data cluster towards the left side of the x-axis; again mean, median, and mode are different.
Table. Measures of Central Tendency and Their Uses
Measure of central tendency calculation when applied
appropriate measure of dispersion
Mode Most frequently occurring value in the data
With categorical data, eg, sex, blood groups, disease stages
Not required
Median Order data from low to high or vice versa, pick the value that divides the data in 50% above and 50% below
Continuous numerical data that is non-normally distributed (“skewed”), eg, height, weight, lab values
Range (minimum– maximum) and/or
Interquartile range (25th-75th percentile)
Mean Add all data together and divide the total by the number of data points
Numerical continuous normally distributed data (“bell-shaped”), eg, height, weight, lab values
Standard deviation
Low High Scores
Fr e q
u e n
cy
Mean Median Mode
(a)
{
Low High Scores
Fr e q
u e n
cy
Mean Median Mode
(b)
Low High Scores
Fr e q
u e n
cy
Mean Median Mode
(c)
AMWA Journal / V30 N1 / 2015 / amwa.org 33
YOuR sTaTs REFREsHER!
For the mean, an appropriate measure of dispersion is the
standard deviation. The standard deviation tells us the spread
of the data around the mean. The smaller the standard devia-
tion, the steeper the slopes of the bell-shape, and the smaller
the average differences of the values from the mean. The
reverse is also true: the greater the standard deviation, the flat-
ter the slopes of the bell-shape and the greater the average dif-
ference of the values from the mean. The size of the standard
deviation depends on our sample size. The more data we have,
the smaller the standard deviation.
If the mean and median and the mode have the same
value, we have a symmetrical distribution of the data. We can
crudely estimate the skewedness of data when we subtract the
median from the mean; the larger the difference, the greater
the skewedness of the data. When the mean is substantially
greater than the median, the data are right skewed (Figure).
In summary, there are 3 measures of central tendency: the
mode, the median, and the mean. Each one has a special pur-
pose and should only be used within its limitations to avoid
unintentional misrepresentation of data.
REsOuRcEs
Lang TA, Secic M. How to Report Statistics in Medicine.
American College of Physicians, 2nd edition 2006.
Gonzales VA, Ottenbacher KJ. Measures of central tendency
in rehabilitation research: what do they mean? Am J Phys Med
Rehabil. 2001;80:141-146.
Swingler MV, Bishop P, Swingler K. SUMS: A flexible approach
to the teaching and learning of statistics. Measurements
of Central Tendency, Statistics Tutorial, SUMS Statistical
Understanding Made Simple, University of Stirling, www.gla.
ac.uk/sums.
Measures of central tendency. Statistical language,
Australian Bureau of Statistics, www.abs.gov.au/websitedbs/
a3121120.nsf/home/statistical+language+-+measures+of+
central+tendency. Updated July 3, 2013.
Measures of shape. Statistical language, Australian Bureau
of Statistics www.abs.gov.au/websitedbs/a3121120.nsf/
home/statistical+language+-+measures+of+shape. Updated
July 3, 2013.
Know the central Tendency of Your Data No matter the text you are working on or the audi-
ence you are writing or editing for, understanding
the central tendency will enable you to appropriately
describe or explain medical research.
■ If you have a study in which an intervention was
tested in a large group with a mean age of 43
years and a standard deviation of 4 years, the
results from the study will have little bearing
on how the treatment would fare in teens or
the elderly.
■ If you have just a few data points in a study of a
new miracle weight loss treatment, and the mean
weight loss was 7.2 pounds, consider whether a
few individuals lost a lot of weight while others
stayed unchanged or even gained weight. The
median with the minimum and maximum (or with
interquartile range) will more accurately represent
the data. Your description of the results needs to
acknowledge the small sample size and the wide
variation in results.
■ When you have a fairly large sample and don’t
know the distribution of the data, have your stat-
istician (or your computer) calculate all 3 measures
of central tendency (mean, median, mode). If they
are very similar you can assume a normal distri-
bution. In this case you can use the mean and the
standard deviation for reporting; if the 3 measures
are not very similar, the median is more appropri-
ate for continuous values.
Send your ideas for articles for Your Stats Refresher! to [email protected].➲
Copyright of AMWA Journal: American Medical Writers Association Journal is the property of American Medical Writers Association (AMWA) and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use.