250 words/ two scholarly sources due 12/2!

profileWahonda7
SS1.pdf

AMWA Journal / V30 N1 / 2015 / amwa.org 31

by Thomas M. schindler, PhD/ Head Medical Writing Europe, Boehringer Ingelheim Pharma GmbH & Co. KG, Biberach, Germany

Meaning It! a Refresher on Mean, Median, and Mode

Y ou are likely to have used them a hundred times in

your writing. The terms mean, median, and mode are

in almost every text about clinical studies or medi-

cal investigations. Writers may be so used to these terms that

they sometimes do not fully appreciate their very specific

meanings (Table). This short piece is designed to serve as a

little reminder of the underlying concepts and the limita-

tions of these terms.

Medical writers are frequently called upon to describe

data and draw conclusions from the analysis of data. Very

often, 2 or more sets of data are compared. The most direct

way to compare groups is by listing their values side by side.

This way, all data are presented, and the reader can compare

them one by one. However, such a procedure is only practi-

cal for very small groups, and even then the comparison of

individual values is not very informative.

What we really want is to summarize data of groups into

single measures and then use these measures to compare the

groups. To be meaningful, these measures need to be repre-

sentative of the group, ie, they need to embody the center or

focal point of the data. When we plot the data, we can visual-

ize their distribution. We can then see whether the data clus-

ter around a particular value. We can also see to which extent

the data vary around the central point.

Thus, the measure of central tendency is the single value

that best represents all of the data. There are 3 major mea-

sures of central tendency: the mean, the median, and the

mode. When we use one of these measures, we abstract

from the individual observations to obtain a single measure

for an entire set of data. The measures of central tendency

are a kind of shorthand for the distribution of the data. For

an appropriate description of numerical data, we also need

a measure of variation or dispersion around the central

tendency.

The Mode Let’s start with the one that is least used in biomedical texts,

the mode. The mode is actually a very handy method to

describe the central tendency when used for the appropri-

ate kind of data. The mode is calculated by counting how

often each value in our data occurs. Once we have identi-

fied the most frequently occurring value, we have identified

the mode. The mode is best used to describe categorical (ie,

non-numerical) data. Examples for categorical data are dis-

ease stages, such as the New York Heart Association stages

of heart failure (I to IV ), blood groups, or sex. When several

values appear equally often in our data, we have several

modes, and the initial purpose to have one single value char-

acterizing our data is betrayed (eg, bi-modal distributions

with 2 peaks). Likewise, the mode would be unhelpful if each

value in our data appears only once. This is usually the case

with continuous numerical data such as height or clinical

laboratory values. In these cases, we need to choose another

measure of central tendency, eg, the median or the mean.

Because the mode represents the most frequently occur-

ring value, it does not need to be accompanied by a measure

of dispersion. If we use the mode, we have decided that the

most frequent value in our data is the best representation of

the entire data set, and therefore as a measure of dispersion

it is not really helpful. However, we should always provide

the number of potential categories (for NYHA it would be 4)

should the number of categories not be obvious.

The Median To find the median, we first need to order our data, from low

to high or vice versa. We then look for the one data point

right in the middle, with 50% of the values above and 50%

below, and this is our median. The median divides our list in

half. In a data set with an odd number of data, the median

YoUR sTATs REFREshER!

32 AMWA Journal / V30 N1 / 2015 / amwa.org

YOuR sTaTs REFREsHER!

value is the true middle value. If we have an even number of

data, the median value is the mean of the 2 middle values. The

median is best used to describe continuous numerical data,

such as age, height, or heart rate. The median is the best rep-

resentation of the central tendency when the data are not nor-

mally distributed, ie, they do not form a nice symmetrical and

bell-shaped curve when plotted. Instead the resulting curve

might be leaning to either side—that is, be skewed to the left

or the right (Figure). The median is very robust because it is

not affected by extreme values (outliers), and it is indepen-

dent of the shape of the distribution. There are several appro-

priate measures of dispersion that could be reported with the

median. Most simply, the minimum and maximum value could

be reported to provide the reader with the overall range of the

data. As the median is the 50% point of a distribution, a good

indication of the spread is the interquartile range. This is the

range of the values between the 25th and 75th percentile of a

distribution (ie, a subtraction of the value at 25% of the data

from the value at 75% of the data).

The Mean The mean (also called the arithmetic mean or the average) is

probably the most widely used measure of central tendency.

It has the advantage of taking all data into account, both the

number of observations and all their values. The mean is the

result of adding all the values in the data and dividing the

total by the number of data points. The mean is well suited to

describe continuous numerical data such as height, weight, or

blood pressure. Since the mean takes every value into account,

it is sensitive to extreme values or outliers. If extreme values

are entered into the calculation, the resulting mean will not

accurately represent the central tendency. Furthermore, the

mean is only then an appropriate representation of the cen-

tral tendency when the data are normally distributed, ie, when

the shape of the plotted values resembles the outline of a bell

(Figure). The shape needs to be symmetrical with the mean in

the middle. We need to realize that we might not have a single

representative of the mean value in our observations. If we

measured 150 participants and their mean height was 168.4

cm, we may have not any individual who is exactly that tall.

How well the mean represents the central tendency can be

evaluated when we also know the measure of dispersion.

Figure. a) Data that have a normal distribution. In this case mean, median, and mode all have the same value. b) Data are negatively skewed. The data cluster toward the right side of the x-axis; mean, median, and mode are different. c) Data are positively skewed. The data cluster towards the left side of the x-axis; again mean, median, and mode are different.

Table. Measures of Central Tendency and Their Uses

Measure of central tendency calculation when applied

appropriate measure of dispersion

Mode Most frequently occurring value in the data

With categorical data, eg, sex, blood groups, disease stages

Not required

Median Order data from low to high or vice versa, pick the value that divides the data in 50% above and 50% below

Continuous numerical data that is non-normally distributed (“skewed”), eg, height, weight, lab values

Range (minimum– maximum) and/or

Interquartile range (25th-75th percentile)

Mean Add all data together and divide the total by the number of data points

Numerical continuous normally distributed data (“bell-shaped”), eg, height, weight, lab values

Standard deviation

Low High Scores

Fr e q

u e n

cy

Mean Median Mode

(a)

{

Low High Scores

Fr e q

u e n

cy

Mean Median Mode

(b)

Low High Scores

Fr e q

u e n

cy

Mean Median Mode

(c)

AMWA Journal / V30 N1 / 2015 / amwa.org 33

YOuR sTaTs REFREsHER!

For the mean, an appropriate measure of dispersion is the

standard deviation. The standard deviation tells us the spread

of the data around the mean. The smaller the standard devia-

tion, the steeper the slopes of the bell-shape, and the smaller

the average differences of the values from the mean. The

reverse is also true: the greater the standard deviation, the flat-

ter the slopes of the bell-shape and the greater the average dif-

ference of the values from the mean. The size of the standard

deviation depends on our sample size. The more data we have,

the smaller the standard deviation.

If the mean and median and the mode have the same

value, we have a symmetrical distribution of the data. We can

crudely estimate the skewedness of data when we subtract the

median from the mean; the larger the difference, the greater

the skewedness of the data. When the mean is substantially

greater than the median, the data are right skewed (Figure).

In summary, there are 3 measures of central tendency: the

mode, the median, and the mean. Each one has a special pur-

pose and should only be used within its limitations to avoid

unintentional misrepresentation of data.

REsOuRcEs

Lang TA, Secic M. How to Report Statistics in Medicine.

American College of Physicians, 2nd edition 2006.

Gonzales VA, Ottenbacher KJ. Measures of central tendency

in rehabilitation research: what do they mean? Am J Phys Med

Rehabil. 2001;80:141-146.

Swingler MV, Bishop P, Swingler K. SUMS: A flexible approach

to the teaching and learning of statistics. Measurements

of Central Tendency, Statistics Tutorial, SUMS Statistical

Understanding Made Simple, University of Stirling, www.gla.

ac.uk/sums.

Measures of central tendency. Statistical language,

Australian Bureau of Statistics, www.abs.gov.au/websitedbs/

a3121120.nsf/home/statistical+language+-+measures+of+

central+tendency. Updated July 3, 2013.

Measures of shape. Statistical language, Australian Bureau

of Statistics www.abs.gov.au/websitedbs/a3121120.nsf/

home/statistical+language+-+measures+of+shape. Updated

July 3, 2013.

Know the central Tendency of Your Data No matter the text you are working on or the audi-

ence you are writing or editing for, understanding

the central tendency will enable you to appropriately

describe or explain medical research.

■ If you have a study in which an intervention was

tested in a large group with a mean age of 43

years and a standard deviation of 4 years, the

results from the study will have little bearing

on how the treatment would fare in teens or

the elderly.

■ If you have just a few data points in a study of a

new miracle weight loss treatment, and the mean

weight loss was 7.2 pounds, consider whether a

few individuals lost a lot of weight while others

stayed unchanged or even gained weight. The

median with the minimum and maximum (or with

interquartile range) will more accurately represent

the data. Your description of the results needs to

acknowledge the small sample size and the wide

variation in results.

■ When you have a fairly large sample and don’t

know the distribution of the data, have your stat-

istician (or your computer) calculate all 3 measures

of central tendency (mean, median, mode). If they

are very similar you can assume a normal distri-

bution. In this case you can use the mean and the

standard deviation for reporting; if the 3 measures

are not very similar, the median is more appropri-

ate for continuous values.

Send your ideas for articles for Your Stats Refresher! to [email protected].➲

Copyright of AMWA Journal: American Medical Writers Association Journal is the property of American Medical Writers Association (AMWA) and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use.