1 / 6100%
Homework Notes: Week 1
I. The Investigative Process of Statistics
a. Formulate questions
b. Collect data needed to answer questions
c. Describe the data
d. Draw conclusions using appropriate methods
II. The Terminology of Statistics
a. Population and Sample
i. Statistics is the study of procedures for collecting, describing,
and drawing conclusions from information.
1. A statistic is a number that describes a sample.
ii. A population is the entire collection of individuals about which
information is sought.
1. A parameter is a number that describes a population.
iii. A sample is a subset of a population containing the individuals
that are actually observed.
III. Constructing a Simple Random Sample
a. A simple random sample of size n is a sample chosen by a method in
which each collection of n population items is equally likely to make up
the sample.
i. A simple random sample is analogous to a lottery.
IV. Sample of Convenience
a. In some cases, it is difficult or impossible to draw a sample in a truly
random way. In these cases, often the best one can do is to sample
items by some convenient method.
i. A sample of convenience is a sample that is not drawn by a well-
defined random method.
b. The problem with samples of convenience is that they may differ
systematically in some way from the population.
V. Stratified, Cluster, Systematic, and Voluntary Response Sampling
a. In stratified random sampling, the population is divided up into groups,
called strata, then a simple random sample is drawn from each
stratum.
i. Stratified sampling is useful when the strata differ from one
another, but the individuals within a stratum tend to be alike.
b. In cluster sampling, items are drawn from the population in groups, or
clusters.
i. Cluster sampling is useful when the population is too large and
spread out for simple random sampling to be feasible.
c. In systematic sampling, items are ordered and every kth item is chosen
to be included in the sample.
i. Systematic sampling is sometimes used to sample products as
they come off an assembly line, in order to check that they meet
quality standards.
d. Voluntary response samples are often used by the media to try to
engage the audience. For example, a radio announcer will invite people
to call the station to say what they think.
i. Voluntary response samples are never reliable for the following
reasons:
1. People who volunteer an opinion tend to have stronger
opinions than is typical of the population.
2. People with negative opinions are often more likely to
volunteer their response.
VI. Qualitative and Quantitative Data
a. Data can be divided into two types:
i. Qualitative Data-Classify individuals into categories.
ii. Quantitative Data-Tell how much or how many of something
there is.
VII. Ordinal and Nominal
a. Qualitative data can be further divided into ordinal and nominal
variables.
i. Ordinal variables-Have a natural ordering.
ii. Nominal variables-Do not have a natural ordering.
VIII. Discrete and Continuous
a. Quantitative data can be further divided into discrete and continuous
variables.
i. Discrete variables-Possible values can be listed.
ii. Continuous variables-Can take on any value in some interval.
IX. Ratio and Interval Data
a. Quantitative variables can be categorized as having a ratio or an
interval level of measurement.
i. Ratio Level of Measurement:
1. Zero represents the absence of quantity.
2. Ratios are meaningful.
ii. Interval Level of Measurement:
1. Zero does not represent the absence of quantity.
2. Ratios are not meaningful.
X. Frequency and Relative Frequency Distributions for Qualitative Data
a. Frequency Distribution
i. The frequency of a category is the number of times it occurs in
the data set.
ii. A frequency distribution is a table that presents the frequency
for each category.
b. Relative Frequency
i. A frequency distribution displays how many observations are in
each category. Sometimes, we are interested in the proportion of
observations in each category.
ii. The proportion of observations in a category is called the
relative frequency of the category.
iii. The relative frequency of a category is the frequency of the
category divided by the sum of all frequencies.
XI. Bar Graphs and Pie Charts for Qualitative Data
a. Bar Graphs
i. A bar graph is a graphical representation of a frequency or
relative frequency distribution.
1. Consists of rectangles of equal width, with one rectangle
for each category.
2. The height of the rectangles represents the frequencies or
relative frequencies of the categories.
ii. Pareto Chart-a bar graph in which the categories are presented
in order of frequency or relative frequency.
1. Useful when it is important to see clearly which are the
most frequently occurring categories.
iii. Horizontal Bars-Sometimes more convenient when the
categories have long names.
iv. Side-by-side Bar Graphs-two bar graphs that have the same
categories constructed on the same axis, putting bars that
correspond to the same category next to each other.
b. Pie Charts
i. A pie chart is an alternative to the bar graph for displaying
relative frequency information.
1. A circle which is divided into sectors, one for each
category.
2. The relative sizes of the sectors match the relative
frequencies of the categories.
XII. Frequency Distributions and Histograms for Quantitative Data
a. Frequency Distribution
i. To summarize quantitative data, we use a frequency distribution
just like those for qualitative data. However, since these data
have no natural occurrences, we divide the data into classes.
ii. Classes are intervals of equal width that cover all values that are
observed in the data set.
1. The lower class limit of a class is the smallest value that
can appear in that class. Ex: 0-4, 5-9
2. The upper class limit of a class is the largest value that
can appear in that class. Ex: 0-4, 5-9
3. The class width is the difference between consecutive
lower class limits. Ex: 5-0=5 (0-4=5, 5-9=5)
a. Guidelines for Choosing Classes
i. Every observation must fall into one of the
classes.
ii. The classes must not overlap.
iii. The classes must be of equal width.
iv. There must be no gaps between classes.
Even if there are no observations in a class,
it must be included in the frequency
distribution.
iii. Constructing a Frequency Distribution
1. Step 1: Choose a class width.
2. Step 2: Choose a lower limit for the first class. This should
be a convenient number that is slightly less than the
minimum data value.
3. Step 3: Compute the lower limit for the second class, by
adding the class width to the lower limit for the first class.
4. Step 4: Compute the lower limits for each of the
remaining classes, by adding the class width to the lower
limit of the preceding class. Stop when the largest data
value is included in a class.
5. Step 5: Count the number of observations in each class,
and construct the frequency distribution.
b. Relative Frequency Distribution
i. Given a frequency distribution, a relative frequency distribution
can be constructed by computing the relative frequency for each
class.
1. Relative frequency=Frequency/sum of all frequencies
c. Histogram
i. Once we have a frequency distribution or a relative frequency
distribution, we can put the information in graphical form by
constructing a histogram.
1. A histogram is constructed by drawing a rectangle for
each class. The heights of the rectangles are equal to the
frequencies or the relative frequencies, and the widths are
equal to the class width.
d. Choosing the Number of Classes
i. In general, it is good to have more classes rather than fewer, but
it is also good to have reasonably large frequencies in some of
the classes.
1. There are two principles that can guide the choice:
a. Too few classes produce a histogram lacking in
detail.
b. Too many classes produce a histogram with too
much detail, so that the main features of the data
are obscured.
XIII. Shapes of Distributions
a. A histogram gives a visual impression of the “shape” of a data set.
Statisticians have developed terminology to describe some of the
commonly observed shapes.
i. Skewed Histograms
1. A histogram with a long right-hand tail is said to be
skewed to the right, or positively skewed.
2. A histogram with a long left-hand tail is said to be skewed
to the left, or negatively skewed.
ii. Symmetric Histograms
1. A histogram is symmetric if its right half is a mirror image
of its left half.
2. A symmetric histogram with a peak in the middle is said
to be bell-shaped.
3. A symmetric histogram in which all classes have
approximately equal frequencies is said to be uniformly
distributed.
iii. Unimodal and Bimodal Histograms
1. A peak, or high point, of a histogram is referred to as a
mode. A histogram is unimodal if it has only one mode,
and bimodal if it has two clearly distinct modes.
XIV. Frequency Polygons and Ogives
a. Class Midpoints
i. Some graphs used for representing frequency or relative
frequency distributions require class midpoints. The midpoint of
a class is the average of its lower class limit and the lower class
limit of the next class.
1. Class midpoint = lower limit for a class + lower limit of
next class / 2
b. Frequency Polygon
i. Although histograms are the most commonly used graphical
display for representing distribution, there are others.
1. Frequency Polygon-constructed by plotting a point for
each class.
a. The x coordinate of the point is the class midpoint
and the y coordinate is the frequency.
b. All points are connected with straight lines.
c. Ogive
i. Ogives plot values known as cumulative frequencies.
ii. The cumulative frequency of a class is the sum of the
frequencies of that class and all previous classes.
iii. An ogive is constructed by plotting a point for each class.
1. The x coordinate of the point is the upper class limit and
the y coordinate is the cumulative frequency.
2. All points are connected with straight lines.
XV. Stem-and-Leaf Plots
a. Stem-and-leaf plots are a simple way to display small data sets.
i. The right most digit is the leaf, and the remaining digits form the
stem.
1. Ex: 14.8 Stem is 14 Leaf is 8
2. Ex: 2,730 Stem is 273 Leaf is 9
ii. When two data sets have values similar enough so that the
same stems can be used, their shapes can be compared with a
back-to-back stem-and-leaf plot.
b. A dotplot is a graph that can be used to give a rough impression of the
shape of a data set. It is useful when the data set is not too large, and
when there are some repeated values.
i. For each value in the data set, a vertical column of dots is drawn
with the number of dots in the column equal to the number of
times the value appears in the data set.
XVI. Construct a Time-Series Plot
a. A time-series plot may be used when the data consist of values of a
variable measured at different points in time.
i. The horizontal axis represents time.
ii. The vertical axis represents the value of the variable measured.
iii. The horizontal axis is plotted with dots and connected with
straight lines.
Powered by TCPDF (www.tcpdf.org)
Students also viewed