mod 2 due in 48 hours
attached
3 years ago
15
ModuleII.docx
Unit3.pdf
Unit4.pdf
ModuleII.docx
Module II : Measures of Central Tendency
What are Measures of Central Tendency?
List and explain 3 measures of central tendency. Use criminal justice examples to earn high marks.
note: No Plagiarism. You will lose points significantly if you do plagiarize.
Unit3.pdf
Unit 3
Measures of Central Tendency
Learning Objectives
Explain the purposes of measures of central tendency and interpret the information they
convey.
Calculate, explain, and compare and contrast the mode, median, and mean.
Explain the mathematical characteristics of the mean.
Select an appropriate measure of central tendency according to level of measurement and
skew.
Use SPSS to produce means, medians, and modes.
Unit Outline
Using Statistics
Introduction
The Mode
The Median
The Mean
Three Characteristics of the Mean
Choosing a Measure of Central Tendency
Using Statistics
Measures of central tendency are used to find the typical case or average score on a single
variable. They can, for example:
Identify the most commonly purchased car in the United States.
Compare public opinion on the Affordable Health Care act over time.
Measure the median income in Detroit, Michigan.
Track changes in age at first birth over time.
Introduction
There are three measures of central tendency. Each one is a way of describing a typical
case or average score in a distribution. They include:
The mode
The median
The mean
The mode of a distribution is the value that occurs most frequently.
The mode is most useful when working with nominal level variables.
The mode is the only measure of central tendency appropriate for nominal
variables.
Limitations of the mode:
Some distributions have no mode at all. This occurs when no value occurs more
than once, or when all values occur at the same frequency.
Some distributions have multiple modes. This occurs when more than one value
(but not all of the values) occur at the same frequency.
The mode for an ordinal or interval-ratio variable may not be central to the
distribution as a whole.
The mode can be found in these two distributions by identifying the value with the
highest frequency. In Example A, the mode is Protestant. In Example B. the mode is 93.
The median (Md) is always at the exact center of a distribution.
Half of the scores in a distribution are higher than the median, and half of the scores are
lower than the median.
Calculating the median
The median can be calculated for ordinal and interval-ratio level data.
Before determining the median, all of the scores must be arranged in order, from low to
high.
When there are an odd number of cases (N):
The median is the exact middle case. Find the case number of the median
by using the formula below.
Example:
The Median is 7.
When there are an even number of cases (N):
The median is the value between the two middle-most
cases. Find the case number of the median by adding the
two middle cases and dividing by two.
Find the first middle case by dividing N by 2. The first
middle case below is 7. The second middle case is the
next case. In the example below, the second middle case
is 5. Add 7 plus 5 and divide by two.
Example:
2
1N
The median is 6.
Limitations of the median: Nominal variables do not have a median because their categories
cannot be ranked from low to high.
Although the median is the centermost score, it represents only one
point in the data. It may not be very representative of the other scores.
The mean is the arithmetic average of all scores in a distribution.
The mean is the most commonly used measure of central tendency, but
it can only be used for interval-ratio level data.
-A sample mean is denoted with the symbol:
-A population mean is denoted with the symbol:
Calculating the mean:
Where:
x is the sample mean
is the Summation of all the scores
x
N
x x
i
ix
N is the total number of scores
Example
Calculate the mean of these grades on homework assignments:
85, 92, 78, 86, 94, 80
85 + 92 + 78 + 86 + 94 + 80 = 515 = 85.83 = 14.31
6 6
Limitations of the mean: The mean is most appropriate for interval-ratio variables.
However, because of some useful properties of the mean, it is
sometimes calculated on ordinal variables.
The mean can be deceiving when data are skewed
Characteristics of the Mean
-The mean balances all the scores. -The mean is like a fulcrum that “balances” all scores in a distribution.
-The mean is the central point around which all scores “cancel out” each other. One low score,
cancels out a high score
The mean minimizes the variation of the scores. This is also called the “least squares” principle of the mean.
The mean is the point in a distribution around which the variation (or
differences) in scores is minimized.
The mean is closer to all of the scores in a distribution than any other
measure of central tendency.
The mean can be misleading if the distribution is skewed.
N
x x
i
0xxi
minimum 2
xxi
A “skew” occurs when an otherwise normal distribution has a few extremely high
or a few extremely low scores.
Positively skewed distributions have a few high scores that “pull” the
mean higher.
Negatively skewed distributions have a few low scores that “pull” the
mean lower.
Skewed data can best be better described with the median.
Below are two examples of skewed distributions.
In a normal distribution (unskewed, symmetrical), the mode, median and mean
will all be equal.
Choosing a Measure of Central Tendency
Two main criteria are used when choosing a measure of central tendency:
The level of measurement
For nominal variables, only the mode is possible.
For ordinal variables, the mode and median are possible,
but the median is usually preferred (except in some
inferential statistics, when the mean is used).
For interval-ratio variables, all three measures are
possible, but the mean is usually preferred (except in
cases of skewness, when the median is used).
Unit4.pdf
Unit 4
Measures of Dispersion
Learning Objectives
Explain the purpose of measures of dispersion, and the information they convey.
Compute and explain the range (R), the inter-quartile range (Q), the standard deviation
(s), and the variance (s 2
).
Select an appropriate measure of dispersion and correctly calculate and interpret the
statistic.
Describe and explain the mathematical characteristics of the standard deviation.
Analyze a box-plot.
Use SPSS to produce the standard deviation and range.
Unit Outline
Using Statistics
Introduction
The Range and Interquartile Range
The Standard Deviation and Variance
Interpreting the Standard Deviation
Visualizing Dispersion: Box-plots
Using Statistics
Measures of dispersion are used to describe the variability or diversity in a set of scores.
They can be used to describe the:
Variability in hospital patients’ body weight
Diversity in life styles across different social settings
Differences in the level of need for health care subsidies across social class levels
Variations in income inequality across nations over time
Introduction
Examine these two distributions of ambulance response times. Service A has low
dispersion because it has more similar values. Service B has high dispersion because it
has more dissimilar values.
The Range
The range (R) is the distance between the highest and lowest scores in a
distribution.
The range is easy to calculate and serves as a quick measure of variability in a
distribution.
The greater the value, the more dispersion in the distribution.
Computing the Range
Arrange the scores from low to high.
The range is equal to the difference between the lowest value and the highest value.
Limitation of the range:
The range is based on only two scores in the distribution, the two most
extreme (the highest, and the lowest).
The range provides no information on the variation of the scores
between the highest and the lowest values.
The Interquartile Range
The interquartile range (Q) is the distance between the third quartile (Q 3 ) and first
quartile (Q 1 ) in a distribution.
The interquartile range provides more information than the range.
The interquartile range focuses only on the middle 50% of values in a distribution.
Computing the Interquartile
Arrange the scores from low to high.
valuelowestvaluehighestR
13 QQQ
Find the third quartile (Q 3 ), or the point at which 75% (three-quarters) of a
distribution falls at or below.
Multiply N by 0.75. This yields the case number of the third quartile.
(If you get a fractional value, round to the nearest case number.)
Find the first quartile (Q 1 ), or the point at which 25% (one-quarter) of a distribution
falls at or below.
Multiply N by 0.25. This yields the case number of the first quartile.
(If you get a fractional value, round to the nearest case number.)
Limitations of the interquartile range:
Like the range, the interquartile range is based on only two scores.
However, the interquartile range focuses on the middle, or most
typical, 50% of the values
The interquartile range ignores the bottom 25% and top 25% of the
distribution.
Computing the Range and Interquartile Range
Example: Calculate the range and interquartile range for the following
distribution.
The Standard Deviation and Variance
The ideal measure of dispersion
Uses all the scores in a distribution.
Describes the average or typical deviation of the scores.
Increases in value as the scores become more diverse.
A measure of dispersion focused on deviations (the differences between each individual score
and the mean) meets these three criteria. In the formula below, the deviation from the mean is
calculated for one value.
-The standard deviation (s) describes the dispersion of a distribution.
-The variance (s 2
) is primarily used in inferential statistics
Note: For both formulas, use N in the denominator for populations and use N-1 in
the denominator for samples.
To compute the standard deviation (and variance) set up a table with three columns: a
score (x) column, a deviation column, and a deviation-squared column.
Sum up the values in the deviation-squared column – this gives you the numerator. N (for
populations) or N-1 (for samples) is your denominator. Divide these values and then take the
square-root.
Example: Calculating the dispersion in students’ ages
xxDeviation i
N
xx s
i
2
N
xx s
i
2
2
Interpreting the Standard Deviation
The standard deviation can be interpreted as:
An index of variability that increases in value as the distribution
becomes more variable.
The minimum possible value is zero – where there is no
variation or when all cases are the same.
A method of comparing one distribution with another.
A tool for determining the area under the normal curve.
Interpreting Box Plots
Boxplots provide a helpful way to visualize and analyze dispersion.
Boxplots display the median, the range, and the Interquartile range.
In this example, each box, and its accompanying lines and “whiskers” (the t-shaped
endings) shows the birth rates per 1,000 population in nations with low, lower middle,
upper middle, and high income.
Summary of Measures of Dispersion
Measures of dispersion summarize information about the heterogeneity, or variety,
in a distribution of scores. While measures of central tendency locate the central
points of the distribution, measures of dispersion indicate the amount of diversity in
the distribution.
The range (R) is the distance from the highest to the lowest score in the distribution.
The interquartile range (Q) is the distance from the third to the first quartile (the
“range” of the middle 50% of the scores). These two ranges can be used with
variables measured at either the ordinal or interval-ratio level.
The standard deviation (s) is the most important measure of dispersion because of
its key role in many more advanced statistical applications. The standard deviation
has a minimum value of zero (indicating no variation in the distribution) and
increases in value as the variability of the distribution increases. It is used most
appropriately with variables measured at the interval-ratio level.
Boxplots provide a visual way of analyzing dispersion by graphically illustrating
the range, the interquartile range, outliers, and extreme outliers.
Basic Terms
Boxplot
Deviation
Dispersion
Interquartile range (Q)
Measures of dispersion
Range (R)
Standard deviation
Variance