lOMoARcPSD|21919833
Statistics Notes Vocab Formula
Introduction to Probability and Statistics (Liberty
University)
lOMoARcPSD|21919833
Vocabulary Chart
Nominal Data Categorical data with qualities that cannot be ordered or
ranked
Ordinal Data Categorical data with qualities that can be ordered or
ranked.
Qualitative
(Categorical) Data
Data that describes. It can't be measured or used for
arithmetic.
Quantitative
(Numerical) Data
Data that is numerical. It can be measured and it can be
used for arithmetic. .
Raw Data Unorganized, unprocessed and not summarized.. Typically,
this is data that is not already available
Data Information used in a study to answer a statistical question
Bias Data The systematic favoring of certain outcomes in a study.
There are many ways to introduce bias into a study
Available Data Data collected by some other entity - a government
organization or private company.
Sample/Sampling A subset of the population. There are many ways to select a
sample.
Representative
Sample
A sample that accurately reflects the population.
Population The entire set of individuals from which to sample
census Using the entire population to obtain data
Random sample A sample that has been selected in a manner where every
member of the population has some predetermined chance
of being selected for the sample
Random selection The method of obtaining a random sample
probablitilty
sampling plan
The way to collect a random sample the guarantees a
certain likelihood for each member of the population to be
selected
lOMoARcPSD|21919833
Random number
generator
A method of collecting a sample that utilizes technology to
select random numbers corresponding to individuals in the
population
Random number
table
A method of collecting a sample to select random numbers
corresponding to individuals in the population/ each
individual is assigned a number which are then selected
from table
Simple random
sample
A method of selection that guarantees that every sample of
a certain size has an equal chance of being the selected
sample
Systematic random
sample
A sampling method where every K th individual is selected
from the sample (e.d every 2nd, 4th or 10th individual)
Cluster Smaller subgroups of the population not necessarily similar
in any way besides all being together in one place making
the individuals easier to sample together
Stratum/strata The homogenous groups in a Stratfied random sample. All
individuals in each stratum have something in common and
we would like to see how that affects the outcome of the
sample
Cluster sample A sampling method where the population is separated into
groups typically geographically and random selection of
clusters is made. Each individual in the cluster becomes
part of the sample
Subjects/Participants The people or things being examined in an observational
study.
Retrospective Study A study that observes what happened to the subjects in the
past, in an effort to understand how they became the way
they are in the present.
Prospective Study A study that begins by selecting participants, then tracks
them and keeps data on the subjects as they go into the
future.
Observational Study A type of study where researchers can observe the
participants, but not affect the behavior or outcomes in any
way.
Replication Repeating the experiment on multiple subjects/experimental
units. This principle of experimental design that states that
a larger experiment with more subjects/experimental units
will allow us to more clearly see differences between the
treatments.
Randomization The principle of experimental design that requires that the
lOMoARcPSD|21919833
subjects/experimental units be assigned to groups using
some random process. This ensures that the two groups are
roughly equal prior to assigning treatments.
Blinding The practice of making sure that certain individuals do not
know which subjects are receiving which treatment.
Convenience Sample A sample that is easily obtained. It is often not
representative of the population.
Random Digit
Dialing
A method of contacting people on the phone. Random
numbers are dialed, so this allows researchers to sample
people with unlisted phone numbers.
Double-Blind
Experiment
An experiment where neither the subjects, nor anyone in
contact with them, has any knowledge of which subjects are
receiving which treatment.
Deliberate Bias The purposeful misrepresentation of data for the purpose of
advancing an agenda.
Representative
Sample
A sample that accurately reflects the population.
Single-Blind
Experiment
An experiment where either the subjects have no
knowledge of which subjects are receiving which
treatment, or people in contact with the subjects have no
knowledge of
which subjects are receiving which treatment, but not both
Self-Selected
(Voluntary Response)
Sample
A sample that the participants choose to be a part of.
Selection Bias A bias that occurs when certain groups are systematically
left out of the sample. This is a systematic error.
Measurement Bias A mistake in the measurements taken in the study. This is a
systematic error.
Continuous
Probability
Distribution
A probability distribution where probabilities are related by
a mathematical function, and the outcomes can take any
value within a given range.
Law of
Averages/Gambler's
Fallacy/Gambler's
Ruin
A misapplication of the Law of Large Numbers, where
people try to apply long-run probabilities to short-run
events. The false "Law of Averages" is not a mathematical
phenomenon, but rather a psychological trick people play
on themselves to convince themselves that favorable
outcomes are about to occur, using past behavior to
influence their reasoning.
Selection Bias A bias that results from systematically excluding certain
lOMoARcPSD|21919833
subsets of the population from the sample. It is not
necessarily intentional.
Scatterplot A graphical display that allows us to see
the relationship between two
quantitative variables.
Multiple Data
Sets
Plotting more than one data set on a scatterplot
requires that we use different colors or symbols
for the different data sets so we can see the
relationships separately
Unintentional Bias Bias that is not purposeful. It exists because of errors in the
design of the study
Explanato
ry
Variable
The variable whose increase or decrease we
believe helps explain a tendency to increase
or decrease in some other variable.
Response Variable The variable that tends to increase or decrease
due to an increase or decrease in the
explanatory variable.
Best-Fit
Line/Trend
Line/Regression
Line
A line that closely approximates the response
values for given explanatory values when the
form of the scatter plot is linear.
Least-squares Line The regression line where the sum of the
squares of the residuals are the smallest.
Poisson
Distributi
on
A distribution used for rare events. It can find
the probability of exactly a certain number of
successes within a given timeframe, assuming
that events occur independently.
lOMoARcPSD|21919833
General Notes:
This is also called the Monte Carlo method, named after Monte Carlo
Casino because it uses physical objects like dice.
Four top steps
○ Collect
○ Analyze
○ Interpret
○ Present
1. Collect. Collect the information from a variety of sources
2. Analyze. Analyze the information that you've collected
3. Interpret. Interpret what that analysis means
4. Present. Present it in a way that anyone can
understand 5.
Statistical study
A way to collect information from individuals
descriptive statistics, you are going to analyze what's going on at a particular point and use
statistics to describe the information that you've obtained.
inferential statistics, you are going to use statistics that you've obtained and make a
generalization about the population at large.
Discrete vs. Continuous Data:
● Discrete:
○ can only take on certain values within a range
○ Data that can only take so many different values
● Continuous:
○ Can take any value within a range
Correlation:
lOMoARcPSD|21919833
The correlation coefficient is always between -1 and +1.
Correlation
The strength and direction of a linear association between two quantitative
variables.
Correlation coefficient (r)
The numerical value between -1 and +1 that measures the correlation
between two quantitative variables.
The correlation is measured using a numerical value known as the
correlation coefficient. The correlation coefficient is a variable called "r"
and is unit-less. It is expressed as a number between negative 1 and
positive 1 and indicates the strength of the linear association.
Numbers that are close to negative 1 or positive 1 are associated with
a strong association between the two variables--a 1 indicating a
strong positive association, and a negative 1 indicating a strong
negative association. Numbers near zero represent almost no linear
relationship.
=CORREL(B2:B6,C2:C6)