marketing assignment (including math)
8/28/2017
38
Survey Research Chapter 7
Len Hostetter Fall 2017
SURVEY TERMINOLOGY
• Respondents – People who answer the questions.
• Survey – Data collection methodology with representative sample of people.
• Sample Survey – Surveying respondents who are a representative sample of the
target population. • Response Rate
– The number of questionnaires returned divided by the number of eligible people asked to participate in the survey.
• Factors that Bias the Response Rate – Persons predisposed to complete a survey (or not). – Respondents usually better educated & more likely a homeowner. – Person filling out survey is not the intended subject.
113
SURVEY RESEARCH
• Survey Objectives – Collect information to describe what is happening (e.g. people’s
beliefs, attitudes, values, likes/dislikes). • Survey research is Descriptive Research
– Identifies and describes characteristics of target markets and/or purchasing patterns.
– Measures consumer attitudes. – “Paints a picture.” – May be both quantitative and qualitative.
114
8/28/2017
39
SURVEY RESEARCH
• Advantages – Quick – Inexpensive – Efficient – Accurate – Straightforward statistical tools to analyze data – Flexible
• Disadvantages – Results are no better than the quality of the sample and
answers obtained – Errors lead to misleading results
115
SURVEY RESEARCH
• Sampling Error – Results from the sample not being representative of the population.
• Systematic Error – Results from a shortcoming in the research design (administrative
error) or research execution (respondent error). • Respondent Error – related to research execution • Administrative Error – related to a shortcoming in the research design
• Sample Bias – Occurs when the results of a sample deviate from the true value of
the population parameter.
116
SURVEY RESEARCH – RESPONDENT ERROR
• Results from some respondent action or inaction such as non- response or response bias. – Non-respondents are people who are not contacted or who refuse to
cooperate in the research. – Self-selection bias occurs because people who feel strongly about a
subject are more likely to respond to survey questions versus those who are indifferent.
• Response Bias occurs when respondents unconsciously or deliberately answer questions in a way that misrepresents the truth. – Unconscious misrepresentation.
• Misunderstanding the question
• Unable to recall details
• Unprepared response to an unexpected question
• Inability to translate feelings into words
• After-event underreporting
117
8/28/2017
40
SURVEY RESEARCH – RESPONDENT ERROR
• Deliberate Falsification: people deliberately give false answers – Misrepresent answers to appear intelligent – Conceal personal information – Avoid embarrassment
• Types of Response Bias – Acquiescence Bias
• A tendency to agree with all or most questions. – Extremity Bias
• The tendency to use extremes when responding to questions. – Interviewer Bias
• Interviewer presence influences respondents’ answers. – Social Desirability Bias
• Bias in responses caused by respondents’ desire, either conscious or unconscious, to appear more socially acceptable or prestigious
118
SURVEY RESEARCH – ADMINISTRATIVE ERROR
• Caused by improper execution of the research. – Data-processing error: incorrect data entry, incorrect data
tabulation or other procedural errors during data analysis. – Sample selection error: improper sample design or execution
of the sampling procedure. – Interviewer error: mistakes made by interviewers failing to ask
or record survey responses correctly. – Interviewer cheating: interviewer fills in fake answers or
falsifies questions.
119
SURVEY RESEARCH METHODOLOGY
• Interactive Surveys – Two-way interaction between the interviewer and the respondent. – Can be either personal or electronic.
• Non-interactive Surveys – Do not facilitate two-way communication whereby respondents give
answers to static questions. • Personal Interview
– Direct communication whereby an interviewer asks respondents questions face-to-face.
– Versatile, flexible and interactive – Advantages:
• Opportunity for feedback, Probe complex answers, props, high participation rate
– Disadvantages • Interviewer influence, lack of anonymity, costs
120
8/28/2017
41
SURVEY RESEARCH METHODOLOGY
• Email – Include a questionnaire in the body of an e-mail. – Distribute questionnaire as an attachment. – Include a hyperlink within the body of an e-mail.
• Internet – Self-administered questionnaire posted on a website. – Respondents provide answers to questions displayed online by
highlighting a phrase, clicking an icon, or keying in an answer.
• Text Surveys via Mobile Phone – Newest survey approach.
– Likely see more applications in the near future.
– Area codes not necessarily geographic.
– Can be cumbersome.
121
SURVEY RESEARCH RESPONSE RATE
• Increasing Response Rate – Cover letter – Incentives – Interesting questions – Follow-ups – Advance notification – Survey sponsorship – Keying mail questionnaires with codes
122
GLOBAL CONSIDERATIONS
123
• Variations in willingness to participate. – Sensitivity to interview subject matter – Cultural considerations – Beliefs about appropriate business conduct
• Challenges with identifying and contacting respondents
8/28/2017
42
PERSONAL INTERVIEW
124
• Done Poorly https://www.youtube.com/watch?v=U4UKwd0KExc
• Done Well https://www.youtube.com/watch?v=eNMTJTnrTQQ
APPROPRIATE SURVEY APPROACH
125
• Questions to be answered: – Is the assistance of an interviewer necessary? – Are respondents interested in the issues being investigated? – Will cooperation be easily attained? – How quickly is the information needed? – Will the study require a long and complex questionnaire? – How large is the budget?
PRE-TESTING THE SURVEY
126
• A screening procedure involving a trial run with a group of respondents to correct problems in the survey design.
• Pre-test methodologies – Screen the questionnaire with other research professionals. – Have the client or the research manager review the finalized
questionnaire. – Collect data from a small number of respondents.
8/28/2017
43
SURVEY RESEARCH ETHICS
127
• Respondents’ right to privacy. • Use of deception. • Respondents’ right to be informed. • Need for confidentiality. • Need for honesty in collecting data. • Need for objectivity in reporting data.
Sampling Design Chapter 12
Len Hostetter Fall 2017
A SAMPLE
• A sample is a more accurate way to obtain information about a population than a census – A census is when all elements of the population are used to obtain
information – Obtaining a sample is more efficient – cheaper and faster than
obtaining a census – Destruction of test units – For a given budget, a sample is more accurate than a census – A small, representative and unbiased sample is preferred to a
larger, biased sample
129
8/28/2017
44
SAMPLING
• Houston’s Green Bank wants to learn about households with annual income above $500,000, with at least one child, and who own a house in the West University neighborhood.
• The bank wants to learn their preferences for internet banking. The bank’s management wants to find out why such households do not use internet banking.
130
SAMPLING
• A sampling plan is described in terms of a population, population element, sampling unit, sampling frame, sample and extent. – A population is all the elements about which information is sought. – A population Element is the unit about which information is sought.
• Here it is a household residing in West University, with an annual income above $500k and at least one child.
– A sampling unit is the unit of a sample which can provide information about the population element.
• Here it is the head of the household making the banking decisions. – A sampling frame is a list of sampling units
• Many forms: email address list, phone numbers, homeowners association meeting attendees, etc. Multiple frames may be available.
– The sample is the participants from the sampling frame. – Extent is the time and place where the sampling takes place.
• Feb – Mar 2017 data gathering in the West University neighborhood.
131
SAMPLING
• Oftentimes the population element and sampling unit are the same entity, but not always. – B2B Research (element and unit are different)
• Population element: Mid-size firms in the U.K. with annual sales of $5 million
• Sampling unit: CEOs of mid-size firms in the U.K. with annual sales of $5 million
– B2C Research (element and unit are the same) • Population element: Professionals in Greater Houston who shop at
HEB • Sampling unit: Professionals in Greater Houston who shop at HEB
132
8/28/2017
45
SAMPLING
• A researcher must select a sampling frame based on the degree of over-adjustment and under-adjustment in the frame. – This is referred to as over-coverage or under-coverage error.
• Over-adjustment (over-coverage) occurs when the sampling frame contains units that should not be part of the population. – This problem is correctible through pre-screening of respondents.
• Under-adjustment (under-coverage) occurs when the sampling frame systematically excludes units that are part of the population. – This problem cannot be easily corrected.
Note: Systematic (non-sampling) errors typically result from the nature of a study’s design and the correctness of execution, and are not a result of sampling errors
133
SAMPLING
• Critical to obtaining a good sample is understanding the over- and under-adjustment in it. – A sampling frame attendees at the West University Neighborhood
Association meeting has under-coverage because some homeowners may not attend.
– The over-coverage may be due to some people who attend the meeting, but rent the house and do not own the house.
• The over-coverage error can easily be fixed by simply including a screening question in the survey: “Do you currently own a house and reside in the City of West University?” Those answering no can be excluded from the survey.
134
SAMPLING
• Lexus is planning a survey to determine which shades of blue and red will be preferred by potential buyers. There are two sampling frames: – A list of telephone numbers with AHHI > $100K. – An email list of people residing in affluent areas with AHHI > $200K.
• The business school dean needs to quickly measure the opinions of currently enrolled full-time MBA students. – Available sampling frames included email addresses, students sitting
in a core course during the second semester, student mail boxes and those attending the weekly “coffee with the dean.”
135
8/28/2017
46
NON-RESPONSE ERROR
• Once the researcher defines a population and selects a sampling frame, the next step is deciding how to obtain a representative sample of sampling units from the sampling frame. – The term representative does not mean a large sample. – It is possible and desirable to have a small and representative
sample rather than a large and non-representative sample. • How can you obtain a large and non-representative
sample? – Using the example of the business school dean who wants to
survey full-time MBAs. The school has 250 full-time MBAs at any given point in time. Recall, the email list (obtained from the IT office) is the sampling frame. Two strategies were available (see next slide):
136
NON-RESPONSE ERROR
• #1: An email blast with a survey link was sent to the entire email list on Friday evening. By Monday morning, 125 surveys had been received, for a response rate of 50%. Data from these 125 respondents was analyzed.
• #2: From the 250 names on the sample list, we randomly chose 100 names. A survey was sent to the sample of 100 randomly chosen units on Friday evening. By Monday morning, 65 surveys were returned. A reminder with a request to participate was sent to the remaining 35 students on Monday, which yielded another 25 surveys, for a total of 90 surveys, or a 90% response rate. Data from these 90 students was analyzed.
137
NON-RESPONSE ERROR
• Despite a larger sample size, why is the first sample less representative? – Though the survey was sent out to the entire sampling frame,
a sub-set responded. It is very likely that the responders were systematically different than non-responders. For example:
• There may be students who do not check their emails on weekends. • There may be students who were out of town for the weekend. • There may be students who are busy during the weekend.
• In summary, among those receiving the survey, the non- responders are likely to be systematically different than responders. Thus, despite a larger sample size, we systematically excluded certain sub-groups of students. – This will bias the results.
138
8/28/2017
47
NON-RESPONSE ERROR
• What happens in the second sampling strategy? – Here we randomly choose a small initial sample of 100
students from the sampling frame, and then try to maximize the response rate.
– Why is this a better strategy? • Statistically, a carefully chosen random sample is highly representative of the
sampling frame. • From the initial random sample, we have tried to maximize the response rate
and achieved a response rate of 90% [90 completed surveys divided by 100 surveys initially sent].
• Because of the high response rate, we can be certain that the final sample of 90 is fairly similar to the initial random sample of 100
• Because the initial sample was randomly chosen, it should be very similar to the sampling frame.
• Thus, to select a representative sample, choose a small initial random sample and then try to maximize the response rate.
139
NON-RESPONSE ERROR
• Tips on Sample Representativeness – The sampling frame should not be confused to be the sample. A
sampling frame is the list of sampling units from which a sample (i.e. sample units) must be chosen.
– Sending the survey to all units in the sampling frame does not guarantee that responders are similar to the rest of the units in the frame. This is simply obtaining a large, convenient sample.
– From the sampling frame, select an initial random sample, preferably not too large. From this initial random sample, maximize response rate so that the initial random sample and the final sample are very similar.
– Always try to spend resources: 1) on obtaining a sample list (or sampling frame) that does not suffer from under-adjustment and 2) to select a small initial random sample and maximize* response rate.
*Note: W e can maximize response rate through: i) respondent incentives (e.g., small cash amount, chance to win a prize) and ii) reminders and requests to fill out the survey.
140
UNDER COVERAGE
• In most cities, over 50% of the population have unlisted phone numbers, and frequently these tend to be higher income professionals.
– Therefore, as a sampling frame, the phone book is inherently non- representative.
• To find out if MBA students believe Valhalla is a good bar on the Rice University campus, a survey was distributed to students at Valhalla on Thursday and Friday evenings.
– 150 students filled out the survey. – Results showed 78% of the students find it to be a “good/great” bar. – How different would the answer be of the under-covered group, and how
would the answer be different? • The 78% rating of “good/great” based on the sample is an over-estimate
because we only interviewed people at the bar. • These people are there because they like the bar. • The under-coverage of MBA students who do not like the bar, and hence do not
go there, is a serious flaw. • As a result, we have over-estimated the liking of the bar. In a representative
sample, the estimated percentage rating of “good/great” should be less.
141
8/28/2017
48
UNDER COVERAGE
142
SUMMARY OF NON-RESPONSE ERROR
143
• Practical limitations may prevent researchers from obtaining a representative sampling frame.
• Researchers assess the under/over-coverage to understand the extent to which they may have over- or under-estimated the quantity of interest (e.g. the 78% of students who believe the Valhalla is a “good/great” bar).
• Issues to consider when evaluating a sample – Identify groups systematically under/over represented in the final
sample. – Does their exclusion/inclusion increase or decrease the estimate
(average, percentage agreeing, etc.) of the variable of interest? • Estimate by how much
– When drawing conclusions, clearly state if the sample estimate is relatively higher or lower because of coverage issues and/or non- random initial sample.
SAMPLING ERROR
144
• Known as “Margin of Error” as reported in surveys. • Reflects the uncertainty with which the conclusions can be applied to
the sampling frame. • To the extent the sampling frame is representative of the population,
the conclusions and the uncertainty can be applied to draw inferences about the population.
• Using the Valhalla as an example for interpreting the sampling error and drawing conclusions about a population. – Recall a survey of 150 students at the bar on Thursday/Friday
showed 78% believe it to be a good/great bar. – The margin of error, based on a sample size of 150, is 6.8%, at the
95% level of confidence. – Based on this, the 95% confidence interval is 71.2% to 84.8%. – This implies we can be 95% confident that “in the population,
between 71.2% to 84.8% of the people believe the bar to be good.”
8/28/2017
49
SAMPLING ERROR
145
SAMPLING ERROR
146
• From a statistician’s perspective: – The 95% Confidence Interval (C.I.) implies if we repeat the
same process an infinite number of times, then 95% of the time the percentage will fall between 71.2% and 84.8%.
– Thus, among the infinite number of samples, 95% of the samples will have a proportion of people who believe the bar to be good in this range.
– Of course, this presumes the sample is randomly chosen and therefore representative of the population.
SAMPLING ERROR
147
• Is the sample representative of the sampling frame? – The population is the entire MBA student body
• Sampling frame is the people at the bar on Thursday and Friday evenings. • Assume the survey was given to all 250 people at the bar, and 150 of them
completed the survey on Thursday and Friday evening. • How representative are the 150 respondents of the sampling frame?
– Some MBA students may have been too drunk and did not fill out the survey – Some non-MBA students may have completed the survey, if not screened. – Perhaps the survey was distributed inside the bar, and not to those waiting outside.
– Because the sample was not randomly selected from the sampling frame, the sample of 150 respondents may not be representative of the sampling frame of 250 people present at the bar.
– Irrespective of the margin of error being large or small (6.8% versus say another number), we are concerned about the validity of applying the estimate of 78% “being a good bar” rating to the sampling frame, i.e., 250 people at the bar. Why (see next slide)?
8/28/2017
50
SAMPLING ERROR
148
• To the extent that students who really like the bar, but were waiting outside, 78% may be an under-estimation. – Perhaps the correct estimate was 88% +/- 6.8%
• Similarly, students really drunk in the bar were super-enthusiastic and were generous in rating the bar. – Perhaps the correct estimate was 65% +/- 6.8%
• The non-representative sample affects the actual estimate of “being a good bar” (78% or 88% or 65%). – The margin of error, which is purely a function of sample size, is
unaffected, staying at 6.8% • To ensure a good estimate, draw a small random sample from the
sampling frame and maximize the response rate from that sample. – If we selected a small random sample of 100 people, and ensured
they all filled out the survey, the result is a more representative sample.
SAMPLING ERROR
149
• Is the sample representative of the population of interest? – If a sample is representative of the sampling frame, we can’t
assume the sample is also representative of the population. – A sampling frame of Thursday/Friday attendees at the
Valhalla is not a representative sampling frame as it under- represents MBA students who like to go to different bars, were in class on Thursday and Friday, who do not drink or smoke, and those who leave campus because they commute.
– By systematically excluding these MBA’s from the sampling frame, they will never be part of the final sample.
• If the sampling frame is unrepresentative, then the final sample is unrepresentative, even if the sample is randomly selected from the sampling frame.
• The estimate will be incorrect, irrespective of the margin of error associated with it.
SAMPLING ERROR
150
• What does the sampling error (i.e., margin of error) tell us? – It does not tell us anything about the quality of the estimate. – It only tells us about the uncertainty associated with the estimate
based on the size of the sample from which the estimate is derived: • The sampling error (+/- 6.8%) does not tell us anything about the quality of our
estimate (i.e. 78%) • To ascertain the quality of the estimate, we need to (i) examine if the final
sample is representative of the sampling frame, and (ii) the extent to which the sampling frame is representative of the population
• Sampling involves a trade-off between quality and quantity. – Determining and obtaining a representative sampling frame,
obtaining an initial random sample from the frame and maximizing the response rate is expensive and time consuming.
• However, it yields a superior and representative sample.
8/28/2017
51
MEASUREMENT ERROR
151
MEASUREMENT ERROR
• These issues are related to measurement error and is unrelated to the sample size.
• Measurement error is associated with the design, interpretation and analysis of the survey.
• Developing a good survey, analyzing the data, and interpreting the results typically involves hiring experts with substantial research experience and is resource intensive.
• A poorly designed survey can yield bad data and the entire research study can be derailed.
• Measurement error can lead to erroneous conclusions despite low margin of error.
152
TOTAL SURVEY ERROR
• Total Survey Error = Sampling frame (or coverage error) + Non- response error + Sampling error (margin of error) + Measurement error – Sampling error may be able to quantified. – The other errors can only be assessed by understanding the soft
issues of the survey methodology. • Sampling frame error: Was the list of potential respondents representative of
the desired population? What types of population members may be systematically included or excluded from the sample? How will this exclusion or under-representation affect the final estimate of key variables?
• Non-response error: Did the researcher simply blast the survey to everyone on the sample list or did they choose a random sample and then administer the survey? What is the response rate (higher is better)? How similar are responders and non-responders to each other? How similar or different are the results of the survey to previous studies?
• Sampling error: Based on sample size, what is the margin of error? • Measurement error: Ask for a copy of the survey to understand how different
constructs and variables are measured. Did the study use reasonable scales? How were the scales coded and analyzed?
153
8/28/2017
52
OBTAINING RESPONDENTS FOR A SAMPLE
• Sampling Services – Firms (List Brokers) specializing in providing lists or databases of
specific populations. • Online Panels
– Lists of respondents who have agreed to participate in marketing research.
– Contains millions of potential respondents. – The more specific the profile requested, the more expensive the
panel.
154
PROBABILITY AND NON-PROBABILITY SAMPLING
• Probability sampling — every population element has a known, non-zero probability of selection – Simple random sample is the best-known probability sample
• Non-probability sampling — probability of any member of the population being chosen is unknown – Judgment Sampling relies upon the researcher’s judgment and is
quite arbitrary – Convenience Sampling obtains people or units that are
conveniently available – Quota Sampling ensures various sub-groups of a population are
represented – Used to obtain a large number of completed questionnaires quickly
and economically, or when obtaining a sample through other means is impractical
155
Big Data Basics
Chapter 13 Len Hostetter Fall 2017
8/28/2017
53
DESCRIPTIVE STATISTICS AND BASIC INFERENCES
• Raw data are simply numbers and words – with little meaning • Basic statistical tools for summarizing information from data include:
– Frequency distributions – Proportions – Measures of central tendency and dispersion
• Metrics provide a means of comparison • The combination of summary metrics from basic statistics and a valid
sample are necessary for effective marketing/business decisions • Inferential statistics allow inferences about a whole population from a
sample • Two applications of statistics:
– To describe characteristics of the population or sample and – To generalize from a sample to a population
157
SAMPLE STATISTICS AND POPULATION PARAMETERS
• The purpose of inferential statistics is to make a judgment about the population
• The sample is a subset of the total number of elements in the population
• Sample statistics are measures computed from sample data
• Population parameters are measured characteristics of a specific population
• We generally use Greek lowercase letters to denote population parameters (e.g., P or V) and English letters to denote sample statistics (e.g., X or S)
158
FREQUENCY DISTRIBUTIONS
• Constructing a frequency table or frequency distribution is a common means of summarizing a set of data – The frequency of a value is the number of times a particular value
of a variable occurs – A distribution of relative frequency, or a percentage distribution, is
developed by dividing the frequency of each value by the total number of observations, and multiplying the result by 100
– Probability is the long-run relative frequency with which an event will occur
159
8/28/2017
54
FREQUENCY DISTRIBUTIONS
• Frequency Distribution of Bank Deposits
160
FREQUENCY DISTRIBUTIONS
• Frequency Distribution of Bank Deposits
161
PROPORTIONS
• A proportion indicates the percentage of population elements that successfully meet some standard on the particular characteristic – May be expressed as a percentage, a fraction, or a decimal
number
162
8/28/2017
55
FREQUENCY DISTRIBUTIONS
• Frequency Distribution of Bank Deposits
163
TOP-BOX/BOTTOM-BOX SCORES
• A top box score refers to the portion of respondents who choose the most favorable response toward a company – the portion that would highly recommend a business to others or the portion expressing the highest likelihood of doing business again – The logic is that respondents who choose the most extreme
response are unique compared to the others • Managers should examine the bottom-box score – the portion of
respondents who choose the least favorable response to some question about customer opinion – More diagnostic of customer problems – Often signals a need for some managerial reaction
164
CENTRAL TENDENCY METRICS – THE MEAN
• The arithmetic average • A common measure of central tendency • The sum of all the observations divided by the number of
observations – A sample mean, 𝑋 (read as “X bar”), can be calculated when there
is not enough data to calculate the population mean, μ • Can sometimes be misleading, particularly when extreme values
or outliers are present
165
8/28/2017
56
CENTRAL TENDENCY METRICS – THE MEAN
166
Number of Sales Calls Per Day by Salesperson
CENTRAL TENDENCY METRICS – THE MEDIAN
• The midpoint of the distribution, or the 50th percentile • The value below which half the values in the sample fall • A better measure of central tendency in the presence of extreme
values or outliers
167
CENTRAL TENDENCY METRICS – THE MODE
• The measure of central tendency that identifies the value that occurs most often
• Determined by listing each possible value and noting the number of times each value occurs
• Used for data that is less than interval, with one large peak
168
8/28/2017
57
DISPERSION METRICS
• Accurate analysis of data requires knowing the tendency of observations to depart from the central tendency
• Another way to summarize the data is to calculate the dispersion of the data, or how the observations vary from the mean
169
DISPERSION METRICS
• Sales Levels for Two Products with Identical Average Sales
170
THE RANGE
• The simplest measure of dispersion – the distance between the smallest and largest values of a frequency distribution – Does not take into account all the observations – Indicates the extreme values of the distribution – In a skinny distribution, values are a short distance from the mean;
in a fat distribution values are spread out • The interquartile range encompasses the middle 50 percent of the
observations, i.e., the range between the bottom quartile and the top quartile
171
8/28/2017
58
THE RANGE
• Low Dispersion versus High Dispersion
172
DEVIATION SCORES
• A method of calculating how far any observation is from the mean is to calculate individual deviation scores – A deviation of any observation from the mean can be calculated by
subtracting the mean from that observation
173
WHY USE THE STANDARD DEVIATION?
• It is the most valuable index of spread, or dispersion • Other measures of dispersion that may be used:
– Average deviation – determined by calculating the deviation score of each observation value (i.e., its difference from the mean) and summing these scores; then dividing by the sample size (n)
– Variance – useful for describing the sample variability; will equal to zero if and only if each and every observation in the distribution is the same as the mean
174
8/28/2017
59
WHY USE THE STANDARD DEVIATION?
• Calculating a Standard Deviation: Number of Sales Calls per Day for Eight Salespeople
175
DISTINGUISH BETWEEN POPULATION, SAMPLE AND SAMPLE DISTRIBUTION
• The normal distribution – One of the most common probability distributions in statistics – Also known as the normal curve – Bell-shaped – Almost all (99 percent) of its values are within ±3 standard
deviations from its mean
176
DISTINGUISH BETWEEN POPULATION, SAMPLE AND SAMPLE DISTRIBUTION
• Normal Distribution: Distribution of Intelligence Quotient (IQ) Scores
177
8/28/2017
60
THE STANDARDIZED NORMAL DISTRIBUTION
• A specific normal curve with several characteristics: – It is symmetrical about its mean – The mean identifies its highest point (the mode) and vertical line
about which this curve is symmetrical – The normal curve has an infinite number of cases (it is a
continuous distribution), and the area under the curve has a probability density equal to 1.0
– The standardized normal distribution has a mean of 0 and a standard deviation of 1
178
THE STANDARDIZED NORMAL DISTRIBUTION
• Standardized Normal Distribution
179
THE STANDARDIZED NORMAL DISTRIBUTION AND Z SCORES
• The standardized normal distribution is valuable because we can transform any normal variable, X, into the standardized value, Z – This has implications for the marketing researcher – A typical standardized normal table allows us to evaluate the
probability of the occurrence of certain events without any difficulty • Computing Z Scores
180
8/28/2017
61
EXAMPLE – STANDARDIZED VALUES
• Suppose that a toy manufacturer has experienced mean sales, μ, of 9,000 units and a standard deviation, σ, of 500 units during the month of September – The production manager wishes to know if wholesalers will demand
between 7,500 and 9,625 units during the month of September this year
– Because there are no tables in the back of our textbook showing the distribution for a mean of 9,000 and a standard deviation of 500, we must transform our distribution of toy sales, X, into the standardized form with our simple formula
181
EXAMPLE – STANDARDIZED VALUES
182
EXAMPLE – STANDARDIZED VALUES
• Standardized Normal Table: Area under Half of the Normal Curve
183
8/28/2017
62
EXAMPLE – STANDARDIZED VALUES
• Standardized Distribution Curve
184
POPULATION DISTRIBUTION AND SAMPLE DISTRIBUTION
• Population distribution – a frequency distribution of the population elements – Mean and standard deviation represented by the Greek letters μ
and σ • Sample distribution – a frequency distribution of a sample is
called the – The sample mean is designated with 𝑋 and the sample standard
deviation is S
185
SAMPLING DISTRIBUTION
• Illustrates the functional relation between the possible values of some characteristic of n cases drawn at random and the probability associated with each value over all possible samples of size n
• The sampling distribution’s mean is called the expected value of the statistic – The expected value of the mean of the sampling distribution is
equal to μ – The standard deviation of the sampling distribution is called the
standard error of the mean (𝑆 𝑋 ) and is approximately equal to σ/ 𝑛 – As sample size increases, the spread of the sample mean around μ
decreases
186
8/28/2017
63
CENTRAL LIMIT THEOREM
• As the sample size, n, increases, the distribution of the mean, 𝑋, of a random sample approaches a normal distribution, with a mean μ and a standard deviation, σ/ 𝑛
• The central-limit theorem works regardless of the shape of the original population distribution
• Theoretical knowledge about distributions can help solve practical research problems – Estimating parameters – Determining sample size
187
CENTRAL LIMIT THEOREM
• Population Distribution: Hypothetical Toy Expenditures
188
CENTRAL LIMIT THEOREM
• Calculation of Population Mean
189
8/28/2017
64
ESTIMATION OF PARAMETERS AND CONFIDENCE INTERVALS
• Point estimates – Our goal in utilizing statistics is to make an estimate about
population parameters – The population mean μ, and the standard deviation σ, are
constants, but usually unknown – Point estimate: an estimate of the population mean in the form of a
single value, usually the sample mean – One would be extremely lucky if the sample estimate were exactly
the same as the population value
190
CONFIDENCE INTERVALS
• A confidence interval estimate is based on the knowledge that μ = 𝑋 ± a small sampling error
• After calculating an interval estimate, we can determine how probable it is that the population mean will fall within this range of statistical values
• The confidence level is a percentage or decimal that indicates the long-run probability that the results will be correct – Traditionally, researchers have utilized the 95 percent confidence
level
191
SAMPLE SIZE
• Random error and sample size – Random sampling error varies with samples of different sizes – Increasing the sample size decreases the width of the confidence
interval at a given confidence level • When the standard deviation of the population is unknown, a
confidence interval is calculated by using the following formula: – Confidence interval = 𝑋 ± 𝑍 𝑆𝑛
192
8/28/2017
65
Basic Data Analysis Chapter 14
Len Hostetter Fall 2017
INTRODUCTION
• Researchers infer whether or not a condition exists in a population based on what is observed in a sample.
• Researchers use statistics to search for a pattern within the data.
194
CODING QUALITATIVE RESPONSES
1- RECREATIONAL ACTIVITIES
2 – RESTAURANT SELECTION
3 – LUXURIOUS ROOMS
195
8/28/2017
66
CODING QUALITATIVE RESPONSES
• Coding is the process for assigning meaning to a response – Codes represent the meaning in data – Assign a measurement symbol to categorize responses (i.e. a
number, letter or word) – Mistakes in coding can change the conclusions
14–196
1 2 3
196
CODING QUALITATIVE RESPONSES
• Dummy coding assigns a 0 to one response category and a 1 to the other response category
• Effects coding assigns a +1 to one category and a -1 to the other category
14–197
197
CODING QUALITATIVE VARIABLES
• Class coding assigns numbers to categories in an arbitrary way • Dummy coding, Effects Coding and Class Coding provide the
ability to statistically analyze qualitative responses for categorical coding
198
14–198
8/28/2017
67
DESCRIPTIVE ANALYSIS
• Transforming data to describe characteristics such as central tendency, distribution and variability – Descriptive statistics include mean (average), median, mode,
variance, range and standard deviation – The upside is that by using selected statistics, a researcher can
summarize responses from a large numbers of respondents – Descriptive statistics from a sample are used to make inferences
about characteristics of the entire population
199
14–199
DESCRIPTIVE ANALYSIS
200
14–200
TABULATION
• Arranging data in a table or summary format – Shows frequency that each result occurs
• Counting the different ways respondents answered a question and arranging them in a tabular form yields a frequency table – The actual number of responses to each category is a variable’s
frequency distribution
201
14–201
8/28/2017
68
CROSS TABULATION
• A tool for comparing two or more variables simultaneously • Creating a table with multiple columns and rows from raw
data as raw data is divided into sub-groups – Shows how a dependent variable changes among sub-groups – Most popular method of data analysis in market research – Uncovers relationships among variables that are otherwise
unclear
202
14–202
CROSS TABULATION
203
14–203
Consumer Responses to Different Ways of Obtaining Music in the United States (in millions)
CONTINGENCY TABLE
• A table in a matrix format displaying the frequency distribution of the variables – Two-way contingency tables involving two variables are used most
often – Beyond three variables, a contingency table is difficult to analyze
and explain – The row and column totals are often called marginals, because they
appear in the table’s margin
204
8/28/2017
69
CONTINGENCY TABLE
205
MODERATOR VARIABLE
• Oftentimes, a third variable, or a moderator variable is introduced after analysis of the relationship between two initial variables – Introduced into the analysis to improve understanding
• Understand the conditions under which the relationship between the first two variables is strongest and weakest
– Analysis involves basic cross-tabulation within various sub-groups of the sample
– Changes the relationship between the original independent and dependent variables
206
MODERATOR VARIABLE
207
8/28/2017
70
HOW MANY CROSS TABULATIONS TO COMPLETE?
• Surveys asking dozens of questions and having hundreds of categorical variables are stored in a data warehouse – Software (e.g. CHAID: chi-square automatic interaction detection)
assists researchers in cross-tabulating every combination of categorical variables
– Very common in exploratory research
208
DATA TRANSFORMATION (DATA CONVERSION)
• The process of changing data from its original form to a more suitable format for performing a data analysis – Involves recoding raw responses into modified or new variables – Combining adjacent categories of a variable is a common form of
data transformation to reduce the number of categories
209
14–209
INDEX NUMBER
• A data transformation allowing the tracking of a variable’s value over time and comparing with other variable(s) – Calibrated to a base period or base number – Base year identified with a time related number – Require ratio measurement scales
210
NOTE: The basis is 1968 (base year) and US Consumption in 1968 of 4 liters/person/year
8/28/2017
71
DISPLAYING DATA
211
• Tables, graphs and charts simplify and clarify data – Range from a computer printout to elaborate graphs – Software programs and tools quickly produce these – Facilitates summarization and communication – Bar charts (histograms), pie charts, curve/line diagrams and scatter
plots are among the most widely used tools
HYPOTHESIS TESTING USING BASIC STATISTICS
212
• Empirical testing involves inferential statistics – An inference made about some population based that can be
verified based on observations of a sample representing that population, or experience rather than logic or theory
– Types of analysis depend on # of variables involved • Univariate statistical analysis tests hypotheses with 1 variable • Bivariate statistical analysis tests hypotheses with 2 variables • Multivariate statistical analysis tests hypotheses and models
involving 3 or more variables
HYPOTHESIS TESTING PROCEDURE
213
• Hypothesis is tested by comparing an educated guess with empirical (i.e. observable) reality. The process: ─ A specific and theoretically sound hypothesis is derived from the
research objectives ─ A sample is obtained from the population and the relevant
variables are measured ─ The measured value obtained in the sample is compared to the
value either stated or implied in the hypothesis ─ The hypothesis is supported or not supported by the measured value
8/28/2017
72
HYPOTHESIS TESTING PROCEDURE
214
• Example: Univariate Hypothesis – On average, females represent < 30% of
college football television viewership • Sample is completed and an average of 28% of
television viewership is female, therefore the hypothesis is supported
• If sample is completed and an average of 45% of television viewership is female, therefore the hypothesis is not supported
SIGNIFICANCE LEVELS
215
• We can’t make a statement about a sample with absolute certainty…there’s always a chance for error
• The probability of committing a Type I Error is called the Significance Level – A Type I Error occurs when a condition that is true in the
population is rejected based on statistical observation in the sample
• Example: An observed sample average leads to the conclusion that the mean is < 30% when in fact the true population mean is 31%
• The probability of committing a Type II Error is called Beta – A Type II Error occurs when the sample data suggests that a
relationship does not exist when in fact, a relationship does exist in the population
• Example: An observed sample average leads to the conclusion that the mean is > 30% when in fact the true population mean is < 30%
P - VALUES
216
• When performing a hypothesis test in statistics, a p-value helps determine the significance of the results.
• Hypothesis tests, test the validity of a claim that is made about a population (this claim that’s “on trial” is called the null hypothesis).
• The alternative hypothesis is the one you’d believe if the null hypothesis is not true.
– The evidence in the trial is the data and supporting statistics – Hypothesis tests use a p-value to weigh the strength of the evidence (what
the data says about the population) – The p-value is a number between 0 and 1 and interpreted in the following
way: • A small p-value (typically ≤ 0.05) indicates strong evidence against the null hypothesis, so you reject the
null hypothesis • A large p-value (> 0.05) indicates weak evidence against the null hypothesis, so you accept the null
hypothesis • p-values very close to the cutoff (0.05) are considered to be marginal (could go either way)
8/28/2017
73
P - VALUES
217
• Maruca’s Pizza on the Seaside Heights, NJ boardwalk claims that on average, their delivery times are 30 minutes or less.
─ The null hypothesis (Ho), is that the average delivery time of 30 minutes or less. You don’t believe this.
─ Your alternative hypothesis (Ha) is that the average delivery time is greater than 30 minutes.
─ You randomly sample some delivery times, run the data through the hypothesis test, and the p-value = 0.001, which is much less than 0.05.
─ Typically the null hypothesis is rejected when the p-value is less than 0.05. ─ You conclude that their average delivery time > 30 minutes. The pizzeria’s
claim is wrong. You believe the alternative hypothesis.
UNIVARIATE T - TEST
218
• Used to determine if the difference between two means is significantly different, or due to random chance – The t-distribution, like the standardized normal curve, is a
symmetrical, bell-shaped distribution with a mean of 0 and a standard deviation of 1.0
– When sample size (n) is < 30, use a t-distribution – The shape of the t-distribution is influenced by its degrees of
freedom (df)
UNIVARIATE T - TEST
219
• The number of degrees of freedom is the number of values in the final calculation of a statistic that are free to vary – We calculate degrees of freedom to understand the number of
independent ways a dynamic system can move, without violating any constraint imposed on it
– In the case of a univariate t-test, the degrees of freedom are equal to the sample size (n) minus one
– Today, with computerized software packages, the number of degrees of freedom is provided automatically for most tests