Statistic test for marketing class
Section 11
Testing Best Practices
Rhonda Knehans Drake
Assistant Professor, New York University
Statistical Measurements
Copyright © 2010
*
Introduction – Why We Test
- The ability to easily test new marketing concepts, products or lists and read results is what sets direct marketers apart from other marketers.
- By conducting various A/B split and multivariate tests, we can determine:
- Best direct mail package
- Best creative or offer
- The best landing pages
- The best banner ad creative and offers
- The most appropriate emails to current customers
- Which ads yield the highest registration rates
- Which e-mail lists are the best
- What is the best subject line for increasing our open rates
- Which PPC ads yield the best prospects.
Copyright 2010
Introduction – Why We Test
The basic steps of the scientific method of test design and analysis.
Devise Hypothesis
Design Experiments
Do Experiments
Analyze Results
Make Rollout
Recommendations
Copyright 2010
Sampling Methods
A sample is a subset of customers or members or prospects and a random selection from the universe of interest. For example, ACME Direct, a direct marketer of books, music, videos and magazines is interested in testing a new product offering. The universe of interest is the Core Customer Segment on their database; therefore, the sample will be comprised of names selected from this universe. A random and representative sample of names from this Core Segment will be test promoted as shown in the figure below.
Copyright 2010
ACME Database
(10,000,000 names)
Core Customer
Segment
(3,000,000 names)
Secondary
Customer Segment
(2,500,000 names)
Tertiary Customer
Segment
(4,500,000 names)
Sampling Methods (Cont.)
Representative Samples - To be meaningful, the sample must be representative of the entire population of concern. A representative sample is a sample truly reflecting the population of interest from which the direct marketer draws inferences. For a sample to be representative, no members of the population of interest are purposely excluded from the sample.
To determine the effectiveness of a new format test sent to a specific segment of customers residing on the database, for example, the direct marketer cannot restrict the sample to only those names living in New York. If this new format is sent only to New Yorkers, the expected results will not be representative of the entire population, but only that of New Yorkers. Some direct marketers overlook this very important concept and assume they can apply test results from one population to another. This may work in some cases but not always. Be careful!
Copyright 2010
Sampling Methods (Cont.)
Typically, the only names that should be eliminated from testing are names eliminated in roll-outs such as:
- DMA do-not-promotes
- Frauds
- Credit risk accounts
When testing new promotions, some integrated marketers also consider eliminating:
- Names recently promoted for other marketing tests
- States or cities such as Washington D. C. known to have strict promotional
restrictions. This is especially important if the new test promotion has not yet
been fully reviewed by legal counsel.
Copyright 2010
Sampling Methods (Cont.)
Random Samples - A sample not taken randomly yields biased and misleading results. A random sample is one in which every member of the sample is equally likely to be chosen, ensuring a composition similar to that of the population. Pulling names one after another from the beginning of a geographically sequenced customer database will result in a geographically biased sample. In this case, depending on the size of the sample being drawn in comparison to the size of the population in total, some regions may not be represented in the sample.
To ensure random samples, many direct marketers utilize what is called “nth selects.” For example, ACME Direct maintains a database of 10,000,000 names. To test a new format to a random sample of 10,000 names from their entire database, the direct marketer will begin by selecting one name on the database, choosing every 1,000th (10,000,000/10,000) name thereafter. The result is a random sample of 10,000 names representative of the entire population or database.
Copyright 2010
Sampling Methods (Cont.)
Blocking – Sometimes, when sampling there may be important variables you will want to control for to ensure a clean read of your tests versus the control.
A good example is in the telemarketing world. For example, if you are about to conduct a new outbound telemarketing scripting test, you may be well advised to block based on the years of experience the telemarketing reps have. Doing this will ensure that you do not end up in a situation where the control test has reps with more years experience than the test. Or you might consider blocking on time of day the calls are made.
By blocking on years of experience you will ensure that the same proportion of experienced reps are assigned to both the test and control groups. By blocking on the time of day the calls are made will ensure that the control is not loaded up in the AM and the test in the PM.
Copyright 2010
| Rep Experience | ||
| Script | Less Than 1 Year | Over 1 Year |
| Control Script | 25,000 | 25,000 |
| Test Script | 25,000 | 25,000 |
You have two options for assessing tests:
Hypothesis Testing
Confidence Intervals
Analyzing Test Results
Copyright 2010
Confidence Intervals allows a marketing manager to determine:
- a range in which the response rate or average sales is likely to fall in roll-out based on the sample results, or
- a range in which the difference in response rates or average sales between your test and control package truly lies based on the sample test results.
Hypothesis testing, on the other hand, is much less revealing and only gives you an answer to the question “Did the test beat the control?”
As such, it is much less revealing. Whereas the confidence interval will not only answer this question but also tell you by how much the test has beaten the control.
Analyzing Test Results
*
Copyright 2010
Consider the following example to help explain what one misses by only considering a hypothesis test for purposes of test assessment.
Based on the hypothesis test alone, one might consider the test a failure.
However, by viewing the results of the range (via the confidence interval) in which the test is likely to fall, we see that the risk of rolling-out is minimal compared to the large upside potential.
Analyzing Test Results
*
Copyright 2010
Sheet1
| Confidence Interval | |||||
| Test Panel | Names | RR | Sign @ 95% | LB | UB |
| Control Format | 75,000 | 1.05% | -- | -- | -- |
| Test Format | 75,000 | 1.15% | No | 1.04% | 1.36% |
Sheet2
Sheet3
Analyzing Test Results
There are two types of confidence intervals that can be created.
A confidence interval around a single test result
You will use this formula when interest revolves around assessing the results of a single vertical list or new customer segment.
A confidence interval around the difference between two test results
You will use this formula when interest revolves around assessing the difference between the control and new format, offer or creative tests.
*
Copyright 2010
Understanding Sampling Error
You conduct a new email format test to a 100,000 random sampling of names drawn from your core universe of concern and receive a response rate of 1.19%.
Can you run to the bank with the 1.19% response rate?
*
Copyright 2010
Understanding Sampling Error
Absolutely Not!
Because you did not test the whole universe available, but only a sample, the response rate obtained is therefore only an estimate of what might be.
In fact, each time you conduct such a test you will get a different response rate.
*
Copyright 2010
Let’s say you conducted not 1, but 10 tests of the same format to 10 unique samples of names randomly drawn from the same universe.
The results may look as follows:
Understanding Sampling Error
*
Copyright 2010
Test # % Resp Test # % Resp
1 1.23 6 1.17
2 0.99 7 1.22
3 1.06 8 0.97
4 1.15 9 1.08
5 1.19 10 0.95
Your one test panel result of 1.19%
Every time you test a list, you will get a different response rate.
Some tests will yield results above the true response rate of the entire list and some below the response rate of the entire list.
So, how can we possibly make any decisions with any level of confidence about what we can expect to receive in roll-out given the fact that there is the potential for so much variation in our testing results?
The “Central Limit Theorem” is the answer!
Understanding Sampling Error
Copyright 2010
Let’s assume that instead of repeating the experiment 10 times we did it 1,000 times.
And, assume you tallied the response rates received and created a histogram (bar chart) of the response rates using the Chart Wizard feature in ExcelTM.
The Central Limit Theorem
Copyright 2010
The tally of response rates might look as follows:
The Central Limit Theorem
Copyright 2010
Sheet1
| Test Response Rate | Frequency |
| 0.60% - 0.70% | 17 |
| 0.70% - 0.80% | 73 |
| 0.80% - 0.90% | 125 |
| 0.90% - 1.00% | 170 |
| 1.00% - 1.10% | 226 |
| 1.10% - 1.20% | 174 |
| 1.20% - 1.30% | 122 |
| 1.30% - 1.40% | 74 |
| 1.40% - 1.50% | 19 |
| TOTAL | 1,000 |
Sheet2
Sheet3
And, if we produce the histogram from this data using the Chart Wizard feature of Excel, it will look as follows:
The Central Limit Theorem
*
Copyright 2010
Chart6
| 0.60% - 0.70% |
| 0.70% - 0.80% |
| 0.80% - 0.90% |
| 0.90% - 1.00% |
| 1.00% - 1.10% |
| 1.10% - 1.20% |
| 1.20% - 1.30% |
| 1.30% - 1.40% |
| 1.40% - 1.50% |
Sheet1
| Test Response Rate | Frequency |
| 0.60% - 0.70% | 20 |
| 0.70% - 0.80% | 63 |
| 0.80% - 0.90% | 113 |
| 0.90% - 1.00% | 176 |
| 1.00% - 1.10% | 255 |
| 1.10% - 1.20% | 180 |
| 1.20% - 1.30% | 110 |
| 1.30% - 1.40% | 64 |
| 1.40% - 1.50% | 19 |
| TOTAL | 1,000 |
Sheet1
Sheet2
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
Sheet3
Voila! The symmetric bell shaped or “normal” curve is revealed.
The Central Limit Theorem
*
Copyright 2010
Chart6
| 0.60% - 0.70% |
| 0.70% - 0.80% |
| 0.80% - 0.90% |
| 0.90% - 1.00% |
| 1.00% - 1.10% |
| 1.10% - 1.20% |
| 1.20% - 1.30% |
| 1.30% - 1.40% |
| 1.40% - 1.50% |
Sheet1
| Test Response Rate | Frequency |
| 0.60% - 0.70% | 20 |
| 0.70% - 0.80% | 63 |
| 0.80% - 0.90% | 113 |
| 0.90% - 1.00% | 176 |
| 1.00% - 1.10% | 255 |
| 1.10% - 1.20% | 180 |
| 1.20% - 1.30% | 110 |
| 1.30% - 1.40% | 64 |
| 1.40% - 1.50% | 19 |
| TOTAL | 1,000 |
Sheet1
Sheet2
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
| 0 |
Sheet3
By way of The Central Limit Theorem, the following has been proven*:
- The distributional shape of sample averages (or response rates as in our case) will closely resemble a bell shaped and symmetric normal distribution.
________________________
* These statements are only true for large sample sizes (to be revealed later) and become more true as the sample sizes approach infinity.
The Central Limit Theorem
*
Copyright 2010
For any set of data that is distributed normally (having a symmetric bell shaped curve), the following can be said:
- 99.7% of the observations in the data set will lie within 3 standard deviations of the mean,
- 99% within 2.575 standard deviations of the mean,
- 95% within 1.96 standard deviations of the mean,
- 90% within 1.645 standard deviations of the mean, and
- 68% within 1 standard deviation of the mean.
The Central Limit Theorem
*
Copyright 2010
And, we have just derived the confidence interval formula based on The Central Limit Theorem.
The Central Limit Theorem
*
Copyright 2010
Confidence Intervals/
A Single Test Response Rate
or CTR
We often assess the results of a new vertical list or new customer segment or a new decile on their own merit compared to some break-even or BAU response rate. In other words, for these type of tests we are not making a comparison to a control. When this is the case we will use this formula.
Constructing this confidence interval will allow you, the marketer, to assess the likely range in which the true response rate will lie based on your test.
To calculate a confidence interval around a single test response rate, the following information is required.
- The sample response rate p obtained from the test
- The sample size n of the test
- The desired confidence level c
Where n * p and n * (1 - p) must both be greater than or equal to 5 to assume normality of the data. And, when this condition is not meet we have other means of assessing our tests which we will review shortly.
*
Copyright 2010
To determine the confidence level c, you must answer the following question:
“How confident do I want/need to be that the interval I construct around my test response rate will contain the true response rate I can expect to achieve in the large mailing (roll-out)?”
Do you need to be 90% confident? 95%? 99%?
Confidence Intervals/
A Single Test Response Rate
*
Copyright 2010
The value chosen for c will guarantee, with the same level of probability, that the interval constructed around the test response rate will contain the true population response rate you can expect to receive in roll-out.
…Assuming that you roll-out to the same population that you tested and if nothing else has changed since the time of the test.
It is recommended all confidence intervals be constructed with a 90% or better confidence level. Employing lower levels will yield more risk than you should be willing to take.
We will come back to this later in this section.
Confidence Intervals/
A Single Test Response Rate
or CTR
Copyright 2010
In simplistic terms, a confidence interval around your test response rate is constructed by adding and subtracting from it a multiple of the “sampling error” associated with the test. The “multiple” will depend on the desired confidence level chosen.
The formula for the lower bound of the confidence interval is:
p - (z) *
And, the upper bound:
p + (z) *
The “standard deviation” associated with the test response rate p.
The “multiplier” which equals 1.645, 1.96 and 2.575 for a 90%, 95% and 99% confidence level respectively.
Confidence Intervals/
A Single Test Response Rate
or CTR
Copyright 2010
( p )( 1 - p )/n
( p )( 1 - p )/n
Consider the following example:
Assume the Marketing Manager in charge of external acquisition at American Express is testing a new vertical list for the Blue Card. The size of the test was 65,000 names. The test results yielded a gross response rate of 1.02%.
Before getting too excited, the Marketing Manager decides to assess the potential for the new vertical by constructing a confidence interval. In particular, she needs to calculate an interval around this sample response rate and desires to be 95% confident. The true response rate she can expect in a large mailing roll-out will fall within the bounds.
Confidence Intervals/
A Single Test Response Rate
or CTR
Copyright 2010
Luckily, you do not need to use the complicated formula previously given to calculate confidence intervals.
With the help of “The Plan-alyzer,” a software package created by Drake Direct, you can easily create confidence intervals around a single test response rate
Copyright 2010
5.psd
The Senior Marketing Manager will construct the 95% confidence interval by inputting the response rate, sample size and confidence level in “The Plan-alyzer” as shown below.
Confidence Intervals/
A Single Test Response Rate
or CTR
To download the plan-alyzer go to: www.optimaldm.com/optimaldm/software.cfm
*
Copyright 2010
The Senior Marketing Manager can be 95% confident, should she decide to roll-out with this new vertical, the response rate will not:
- fall below 0.94%, or
- exceed 1.10%.
It is recommended that she use the lower bound to determine the worst case scenario in terms of profitability by running a cost benefit analysis.
Based on her findings, she will decide whether to:
1. Drop the vertical from further consideration
2. Retest to a larger sample
3. Roll-out to the entire vertical tested
Confidence Intervals/
A Single Test Response Rate
or CTR
*
Copyright 2010
Setting the Confidence Level
At what percent should you set the confidence level of your interval?
Confidence Intervals/
A Single Test Response Rate
or CTR
*
Copyright 2010
Confidence Intervals/
A Single Test Response Rate
or CTR
In general, the rule is to set the confidence level at 95%. Your stake in the ground. Then…
- If the results are favorable at 95% you will also see if they are favorable at 99% to make you feel even better about your decision.
- If the results are not favorable at 95% you will then see if they are favorable at 90%. Not that you want to lower your confidence by that much but at least this will give you some additional information upon which to base your decision.
- See the following slide.
*
Copyright 2010
Confidence Intervals/
A Single Test Response Rate
or CTR
A no brainer. Let’s roll!
That’s okay…at a minimum let’s consider a partial to full roll out.
Okay, so we have something here. Let’s either retest or go for a partial rollout.
Not good. We should scrape this from further consideration.
*
Copyright 2010
Evaluate your test response rate at 95% confidence level
Significant?
Yes
No
Is it significant at the 99% confidence level?
Not that we want to go that low but is it significant at the 90% confidence level?
Yes
No
Yes
No
Confidence Intervals/
A Single Test Response Rate
or CTR
So….if the minimum requirement for your outside list prospecting is a 1% response rate or higher what should we do as the next step?
*
Copyright 2010
Confidence Intervals/
The Difference Between
Two Test Response Rates
When interest revolves around determining if one test has beaten the champion offer or creative, you will be interested in examining the difference between response, click or conversion rates for a test versus a control – new price versus control, new OE versus control, new terms versus control, new image versus control image, new subject line versus control subject line, etc.
This is the formula you will most use.
For example, you conduct a new creative test against the control creative based on samples of size 100,000 each. You receive a 0.63% response rate for the new creative and a 0.58% response rate for the control format.
Can you go to the bank assuming the new creative has in fact beaten the control?
*
Copyright 2010
Confidence Intervals/
The Difference Between
Two Test Response Rates
Absolutely Not!
Because both results are based on samples, each will have a certain amount of associated “sampling error.” The true difference is not .05% but something more or less than this percent.
A confidence interval constructed around the difference between the two test response rates will allow you to determine the range in which the true difference actually lies by taking into account the amount of error associated with both tests.
*
Copyright 2010
To calculate a confidence interval around the difference between two test response rates, the following information is required.
- The sample response rates p1 and p2 for both tests
- The sample sizes n1 and n2 for both tests
- The desired confidence level c
Where n* p and n* (1 - p) for both samples must be greater than or equal to 5 to assume normality of the data. And, when this condition is not meet we have other means of assessing our tests which we will review shortly.
Confidence Intervals/
The Difference Between
Two Test Response Rates
Copyright 2010
In simplistic terms, you construct the confidence interval around the difference in your test response rates by adding and subtracting from it a multiple of the “sampling error” associated with the difference in test response rates. The “multiple,” as before, will depend on the desired confidence level chosen.
The formula for the lower bound of the confidence interval is:
( p1 - p2 ) - (z) *
And, the upper bound:
( p1 - p2 ) + (z) *
The “standard deviation” associated with the difference in test response rates p1 and p2
The “multiplier” which equals 1.645, 1.96 and 2.575 for a 90%, 95% and 99% confidence level respectively.
Confidence Intervals/
The Difference Between Two
Test Response Rates
*
Copyright 2010
(( p1)(1 - p1)/n1) + (( p2)(1 - p2)/n2)
(( p1)(1 - p1)/n1) + (( p2)(1 - p2)/n2)
Consider the following example:
In the March 2008 campaign, the Marketing Manager for Sports Illustrated tested a larger and more expensive banner ad. The results of the new banner test and the current control banner ad are shown below.
The Marketing Manager determined the new and larger banner ad must yield four-tenths of an additional click per 1,000 impressions versus the control in order to cover its higher costs*. In other words, the new banner must beat the control by at least 0.04% (.0004).
In order to properly assess the value of the new banner ad, the Marketing Manager will construct a 95% confidence interval (our stake in the ground) around the observed difference in test CTR's.
Confidence Intervals/
The Difference Between
Two Test Response Rates
* Note for this test we are assuming no difference in the conversion rate for the test versus the control. Are you okay with this fact?
*
Copyright 2010
Regarding how the additional 2 orders per 1,000 names mailed was determined, we will discuss break-even calculations a bit later today.
|
|
Number of Customers Mailed |
Number of Customers who clicked |
CTR |
|
Control Banner |
100,000 |
348 |
0.348% |
|
Larger Banner |
100,000 |
430 |
0.430% |
Put away your Bayer Aspirin!
We can also use “The Plan-alyzer”
to help determine the confidence interval surrounding the difference between
two test response rates.
*
Copyright 2010
Confidence Intervals/
The Difference Between
Two Test Response Rates
The Senior Marketing Manager will construct the 95% confidence interval by inputting the control and test data in “The Plan-alyzer” as shown below.
*
Copyright 2010
With 95% confidence, the new larger banner ad is guaranteed to outperform the current ad by anywhere from .027% (.00027) to .137% (.00137)
Note, the lower bound of the difference in the CTR rates (.00027) is less than the minimum required difference (.0004), so…
what should we do?
Confidence Intervals/
The Difference Between
Two Test Response Rates
*
Copyright 2010
Unique Aspects of Analyzing tests further down the funnel
Suppose you interested in assessing a banner ad test in terms of upfront clicks and back end conversions.
- Be aware the sample for assessing the CTR for a banner test versus a control banner will be impressions.
- And,…the sample size for assessing the conversion rate for a banner test versus a control banner will be the number of clicks, NOT IMPRESSIONS!
So, plan accordingly. If you did not properly plan for such an analysis by testing a large enough sample you will not be able to substantiate any difference in conversion rates or other measures of customer value.
*
Copyright 2010
Confidence Interval Facts
Regardless of the amount of time spent planning a test, the results of a confidence interval are only valid if nothing outside of your control occurs from the time of the test to roll-out which may impact your test universe in some way.
Various scenarios which could negatively impact the likelihood that the true response rate obtained in roll-out will not fall within the bounds of your confidence interval include:
- the timing of the test vs. roll-out (e.g., test was conducted in May with a
roll-out the following January)
- a major competitor emerged on the market after the test
- increased competition from the time of the test
- bad press regarding integrated marketing practices since the time of the test and before roll-out
- a major natural disaster in a specific region of the country since the time of the test and before roll-out that might affect traditional mail delivery or internet connectivity
- changes in the promotion or offer (including legal verbiage changes or changes in the application) prior roll-out
*
Copyright 2010
- Remember when we started….
- Based on the hypothesis test alone, one might consider the test a failure when in reality the risk of rolling-out with it is minimal compared to the large upside potential.
Coming Back Full Circle
A Word on Hypothesis Testing
*
Copyright 2010
Sheet1
| Confidence Interval | |||||
| Test Panel | Names | RR | Sign @ 95% | LB | UB |
| Control Format | 75,000 | 1.05% | -- | -- | -- |
| Test Format | 75,000 | 1.15% | No | 1.04% | 1.36% |
Sheet2
Sheet3
Analyzing Averages --
Dollars Spent, Web Page Views
In addition to analyzing the response rates of our tests, we may also be interested in analyzing averages such as dollars spent, balance transfers, page views, etc.
You build confidence intervals around average much like that of response rates. However, keep in mind that the formulas are a bit different and the sample data required to create such intervals is also different.
Plan-alyzer Screen shots are below;
.
Copyright 2010
Planning Marketing Tests/
Sample Size Determination
- Prior to conducting your marketing tests, the most important part of the planning process is to determine the appropriate sample sizes for your test panels.
- Without adequate sample sizes, the time and money spent creating and developing your marketing tests will be wasted.
- This is true for any type of test discussed so far.
- Without proper sample sizes your test results will have so much error (a large standard deviation) that the results will be unreliable.
- As such, odds will be the test results will not even come close to resembling what you can expect in roll-out.
Copyright 2010
Two formulas will be reviewed:
- Sample size determination when concerned with the accuracy of a single
test response, click or conversion rate which is not to be compared to a
control such as a new vertical list or customer segment or new decile.
- Sample size determination when concerned with accurately measuring the
difference between two test panels such as a new banner image versus
the control.
Planning Marketing Tests/
Sample Size Determination
*
Copyright 2010
Sample Size Estimation/
A Single Test Response Rate
or CTR
To calculate the sample size required when concerned with the reliability of a sample response, click or conversion rate, the following information is required:
- The maximum allowable error variance E
- An estimate of the expected response rate p
- The desired confidence level c
*
Copyright 2010
To determine the value of E, you need to answer the following question:
After your test results are final and you build a confidence interval around the sample response rate, what is the maximum width you are willing to accept for the confidence interval?
The answer to this question is the value of E.
For example, you need the test result to be within plus or minus 1/10th of 1 percentage point or 0.10% of the true response , click or conversion rate you can expect in roll-out. In this case E equals .0010.
Sample Size Estimation/
A Single Test Response Rate
or CTR
*
Copyright 2010
You will also need an estimate of the expected test response rate p. Base it on past experience. You do not have to be exact but try to estimate as best you can.
Sample Size Estimation/
A Single Test Response Rate
or CTR
*
Copyright 2010
The formula for deriving the sample size n is:
n = ((z2)( p)(1- p))/E2
Where p is the estimated response, click or conversion rate
E is the maximum error you are willing to accept
z equals 1.645, 1.96 and 2.575 for a 90%, 95% and 99% confidence level
NOTE: The sample size formula is derived from the formula for a confidence interval for a single test response rate previously discussed.
Sample Size Estimation/
A Single Test Response Rate
or CTR
*
Copyright 2010
Consider the following example:
The Marketing Manager in charge of prospecting at Chefs Catalog is prepared to test a new source of names from a list broker to determine the list’s viability as a new name generator. The Marketing Manager believes the response rate for this new vertical list will be no greater than 1.00% (the maximum response rate conceivable by the Marketing Manager for this list). If she requires the test response rate to be within 0.10% (or .0010) of the actual response rate to expect in roll-out with 90% confidence, how many names should she test?
Sample Size Estimation/
A Single Test Response Rate
or CTR
*
Copyright 2010
Luckily, you do not need
to use the complicated formula
previously given to calculate the
sample size for a single test.
With the help of “The Plan-alyzer”
you can easily determine the appropriate sample size for your test.
Copyright 2010
The Marketing Manager will calculate the required sample size by (1) inputting the estimated response rate for the test, (2) inputting 10% as the error rate and (3) selecting the confidence level as shown below.
Sample Size Estimation /
A Single Test Response Rate
Copyright 2010
Therefore, if the Marketing Manager tests 26,785 names she will be guaranteed, with 90% confidence, the response rate obtained from the test will be within 0.10% of the true response rate she can expect in roll-out.
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
Question: If the Marketing Manager decides to increase the confidence level of her test response rate from 90% to 95%, do you think the number of names she must test will go up or down?
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
In order to have more confidence in the results of her list test, all else being equal, she will need to sample 38,030 names.
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
Question: If the Marketing Manager decides to keep the confidence level at 90% but needs more precision in her test estimate and therefore reduces the allowable error from .10% to .05%, do you think the number of names she must test will go up or down?
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
In order to have less error variance in the results of her list test, all else being equal, she will need to sample 107,140 names.
Sample Size Estimation/
A Single Test Response Rate
Copyright 2010
Determining the Tolerable Error Variance
To determine the tolerable error variance, the Marketing Manager must take into consideration how close the expected response rate for the new list is to the minimum response rate required for roll-out. The closer the expected response rate of the new list is to this cutoff level, the more important it will be to have a test estimate with little error variance.
Assume the Marketing Manager in our example will roll-out with a new vertical list as long as it exceeds 0.90% (the break-even response rate level). If the Marketing Manager truly expects the list test to yield at least a 1.00% response rate, then for decision making purposes she only needs to test enough names to ensure that her resulting confidence interval is no wider than ± 0.10% (derived by taking 1.00% minus 0.90%).
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
Determining the Tolerable Error Variance (Cont.)
All that is needed for decision making purposes (rent the names or do not rent the names) is to be certain that the response rate of the new vertical list in roll-out will not be below 0.90%.
Testing more names to yield less variance would be a waste of testing dollars. It is that simple!
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
Setting the Confidence Level
Keep in mind that we put our stake in the ground at 95%.
Sample Size Estimation/
A Single Test Response Rate
or CTR
Copyright 2010
Sample Size Estimation/The
Difference Between Two Test
Response Rates
- When interest revolves around measuring the difference in response rates between two test panels in order to be able to accurately determine if one test has beaten another in terms of response, click or conversion rates, you will employ the formula provided in this section to determine the required sample sizes.
- This is one of the most commonly used formulas when designing new format, creative concept, pricing, offer, image or subject line tests while needing to compare the results against your control.
Copyright 2010
To calculate the sample sizes required when concerned with the reliability of the difference in two test response , click or conversion rates, the following information is required:
- An estimate of the expected rate p1 for one of the two samples
- The minimum difference d to detect as significant
- The desired confidence level c
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
You will base the estimate for one of the test panels p1 on past experience. In all likelihood, one of the two test panels for comparison will be the control package for which you should have historical response rate figures. You do not need to be exact but you should try to be as close as possible.
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
Also required is the minimum difference d to detect as significant. This value represents the difference between the two test response, click or conversion rates (p1 and p2 ) you will want to detect as significant. It should represent a “marketing critical” difference for decision making purposes.
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
For example, your current offer is known to yield a .85% response rate. The addition of a premium to entice response is being tested for which one additional response per 1,000 names promoted is needed to break-even.
Therefore, the minimum difference you are interested in being able to detect between the control and test is 0.10% or .001.
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
The formula for deriving the sample sizes n1 and n2 is:
n1= n2= [( z2)] [( p1)(1 - p1) + ( p2)(1 - p2)] / [( d2)]
Where p1 is the estimated rate for the control test panel
p2 is the estimated rate for the second test panel
d is the minimum difference to be detected and must equal the
difference between p1 and p2
z equals 1.645, 1.96 and 2.575 for a 90%, 95% and 99% confidence
level
NOTE: The sample size formula is derived from the formula for a confidence interval for a difference between two test response rates previously discussed.
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
Consider the following example:
The Marketing Director is testing the addition of a new “Benefits Buck Slip” to the current direct mail format for Net Flix which will explain all the great things one gets if they accept the invitation.
- Assume the control format is known to yield a response rate of 0.85%.
- Also assume the Marketing Director determined that 1 additional
response per 1,000 names mailed are required to cover the additional
costs of the “Buck Slip”.
- Therefore, the Marketing Director will require the test to yield a response
rate 0.10% higher than the control or .95% (which equates to a 11.8% lift
in response).
If the Marketing Director does not observe this difference with statistical significance, he knows he will not be able to use the new “Buck Slip”.
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
So, what sample sizes should the control and test panels be to ensure the Marketing Director will be in a good position to make a confident decision should the observed difference between the test and control response rates yield the required 0.10% absolute difference?
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
Put away that drink!
We can also use “The Plan-alyzer” to help us determine the required sample sizes when concerned with the difference between two test response rates.
Copyright 2010
To be able to read a 0.10% absolute difference in response rates as statistically significant with 95% confidence (should that be the difference observed), the Marketing Director will be required to test 68,522 names per panel.
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Page *
Copyright 2010
Show them how much a 99% confidence level would impact the sample size required.
Consider another example:
- The Marketing Director at TD Ameritrade is testing a new background color for their current banner ad.
- Assume the control banner ad is known to yield a CTR of 0.25%.
- If the Marketing Director wants to be able to detect a 10% lift with significance how many impressions should he order test and control?
- Do you think a 10% lift is a likely event for this type of test?
Sample Size Estimation/The
Difference Between Two Test
Response Rates
Copyright 2010
The Plan-alyzer also lets you determine sample sizes when concerned with accurately measuring differences in average spend, average page views or average time on site – test versus control.
Screen shots are provided below.
Sample Size Estimation When
Interested in Assessing Averages
Copyright 2010
Show them how much a 99% confidence level would impact the sample size required.
Regardless of the amount of time spent planning a test, the results of what you might achieve in roll-out are only valid if nothing outside of your control occurs from the time of the test to roll-out which may impact your test universe in some way.
- the timing of the test vs. roll-out (e.g., test was conducted in May with a
roll-out the following January)
- a major competitor emerged on the market after the test
- increased competition from the time of the test
- bad press regarding integrated marketing practices since the time of the test and before roll-out
- a major natural disaster in a specific region of the country since the time of the test and before roll-out that might affect traditional mail delivery or internet connectivity
- changes in the promotion or offer (including legal verbiage changes or changes in the application) prior roll-out
Sample Size Estimation Facts
Copyright 2010
Planning Marketing Tests/
The Five Rules of Test Design
When preparing to test new formats, copy alternatives, prices or offers several rules must be followed to ensure your results will be readable, reliable and projectable.
Rule 1: For mailers, combine the Control and Test Mailing at the Lettershop for Shipping
If (a) your mail quantities cannot cost justify the co-mingling of the test and champion offers or (b) your test promotion pieces cannot be co-mingled with the champion offer due to it being too different in size as defined by the post office for bundling purposes then …
the promotional pieces associated with the test mailing will be delivered to the post office in one batch and the promotional pieces associated with the champion offer in a separate batch. As such the test pieces will receive different handling and delivery than the champion mailing which will be larger in size.
When this is the case the test mailing (of smaller quantity) will typically be sent to the SCF closest to the letter shop while the bulk mailing (of larger quantity) will be trucked to either the delivery SCF’s (3 digit zip codes) or the appropriate BMC’s or sometimes to the local branches (5 digit zip codes).
Copyright 2010
Rule 2: Back Test Package Changes
As previously mentioned, when changing to a new promotional format one should always back test (re-test) the old promotional format to validate the lift in response you had forecasted. Without back tests you will not be able to determine if the cause for a campaign that is under forecast is due to the list selection, the promotion, or a general downturn in the business or any combination of these three.
Consider the following example….
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
Rule 2: Reverse Test Package Changes (Cont.)
Last year you conducted a new format email test and received a .25% response rate versus a .225% for the control format (a +11% lift). You decided to roll-out with the new format to all eligible names (1,450,000 in size). You forecasted a .25% for this email campaign. The results of the roll-out are as follows:
Your boss is not happy with the results of the campaign and has questioned your decision to change to the new format. Do you have appropriate information to defend your decision?
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
|
|
Number of Customers Blasted |
Number of Customers who Responded |
Response Rate |
|
New Format |
1,450,000 |
2,828 |
.195% |
Rule 2: Reverse Test Package Changes (Cont.)
Now, suppose I told you we conducted a reverse test of the old format within your production plan with results as shown below. Can you now defend your decision to your boss? What happened?
A best practice might be to only back test when the results of the test were not significant at 99% but only 95% or lower.
And, in general, one back test is enough for confirmation.
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
|
|
Number of Customers Blasted |
Number of Customers who Responded |
Response Rate |
|
New Format |
1,350,000 |
2,633 |
.195% |
|
Reverse Test of Old Format |
100,000 |
176 |
.176% |
Rule 3: Test Each Change in a Package Separately
When applicable, test various changes to the control mailer, banner, email separately. Otherwise, you may be misled by the testing results and be left in a situation where no action can be taken.
For example, suppose you are planning to test a new creative format and new copy approach in your next mail campaign. Do not test the two changes combined into one test. Test each change separately. You will gain much more information from two separate test results than if they were combined into one test panel.
The same holds true if testing price and incentive tests.
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
Rule 3: Test Each Change in a Package Separately (Cont.)
In this example, it may be that the copy change increases response while the new format decreases response. This will not be apparent if tested together.
Consider the following test series results where we only tested both changes simultaneously:
Question: What would the decision have been had we not tested each element separately but only combined in the above test series?
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
|
Test Panel |
Response Rate |
Lift versus Control |
|
Control Package |
1.10% |
-- |
|
As Control with New Format and Copy Change |
1.09% |
-1% |
|
As Control with New Format |
0.99% |
-10% |
|
As Control with Copy Change |
1.19% |
+8% |
Rule 3: Test Each Change in a Package Separately (Cont.)
Had we tested each separately, notice how we would have taken a different action based on the results.
We would have never determined that the copy change was a winning element here had we not tested each separately.
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
|
Test Panel |
Response Rate |
Lift versus Control |
|
Control Package |
1.10% |
-- |
|
As Control with New Format and Copy Change |
1.09% |
-1% |
|
As Control with New Format |
0.99% |
-10% |
|
As Control with Copy Change |
1.19% |
+8% |
Rule 4: Test for Only Meaningful Element Interactions
Generally it is unnecessary to test every possible element combination in your test plan. For example, the Marketing Manager may be interested in testing the following changes to the control email:
- Price increase
- Addition of a premium
- Color change to the background
- New creative layout
Testing every possible combination of price, premium, color and creative yields a total of 16 test panels. Testing all 16 is called a “full factorial test design.”
When should a direct marketer consider a full factorial test design?
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
Rule 4: Test for Only Meaningful Element Interactions (Cont.)
The only reason a marketer would test a full factorial test design is if it was truly believed interactions will occur between all four elements with respect to response.
In this example, the only possible interaction to be concerned with would be one between price and premium. In other words, if testing a higher price, perhaps the minus in response (due solely to pricing) would be less for the package with a premium versus the package without the premium.
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
Rule 4: Test for Only Meaningful Element Interactions (Cont.)
Assuming you are only interested in assessing a possible interaction between price and premium, the test series would appear as shown below:
Question: Based on this test series, how would you determine if the addition of a premium to the control package offset any or all of the negative effect that a price increase might have on response?
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
|
Test Panel |
Description |
|
1 |
Control |
|
2 |
Price test - as control with $2 price increase |
|
3 |
Premium test - as control but with premium for order added |
|
4 |
Creative layout test - as control but with a new creative |
|
5 |
Background color change test - as control but new background color |
|
6 |
Price and premium test - as test panel #3 with $2 price increase |
Rule 5: Define the Universe for Testing Carefully
Careful consideration must be given to the names selected for testing.
If you want to compare test results to one another for purposes of determining winning formats for example, you will need consistency in your test universe definitions.
You cannot take the results of a test conducted to one group of customers and assume you will receive the same response rate if used on a different group of customers.
Planning Marketing Tests/
The Five Rules of Test Design
Copyright 2010
Introduction to Full Factorial Test Design
You work for eHarmony.com. You are testing some new banners. In particular, you are testing three rates (RC, R1, and R2) and three incentives (IC, I1 and I2), a full factorial test design will consist of 3 x 3 = 9 test panels:
Panel 1 RC, IC (Your control rate and incentive)
Panel 2 RC, I1
Panel 3 RC, I2
Panel 4 R1, IC
Panel 5 R1, I1
Panel 6 R1, I2
Panel 7 R2, IC
Panel 8 R2, I1
Panel 9 R2, I2
A full factorial test design considers all possible combinations of the elements or factors being tested.
Copyright 2010
Such a test design will allow you to examine the interaction between rate and incentive.
In other words, such a design will allow you to assess if the effect of rate on response is dependent on the incentive being extended.
However, it will sometimes be the case that some of the factors being tested are independent of one another. By this we mean that the drop in response to a rate increase, for example, is the same regardless of the incentive being extended. If this is the case then a full factorial test design is not necessary.
Copyright 2010
Introduction to Full Factorial Test Design
The question arises: How do we know whether the factors are independent or not and therefore whether a full factorial test design is required?
- Base your decision on prior experience and business knowledge.
- If you are unsure then you will test some if not all of the interaction panels.
Copyright 2010
Introduction to Full Factorial Test Design
Why not play it safe and test a full factorial test design every time?
- Because your test plan may be unnecessarily complicated which may result in possible execution errors.
- You may lack the resources (people or dollars) to create all the various package combinations.
- You may lack enough product codes for full factorial testing.
- Time is critical and you must test and get to market.
- Unless you have an unlimited testing budget to ensure an adequate sample size per panel, you may put yourself in a situation where you are assessing response rates for some of the interaction panels where the sampling error is quite high (see example on the next slide).
Copyright 2010
Introduction to Full Factorial Test Design
American Express
Full Factorial Test Design Example
Your current champion package for the American Express Gold Card is comprised of the following elements:
- Initial rate of 0% for 12 months
- 6.99% fixed go-to rate
- 1.99% LOB transfer rate
You have decided to test the following:
- Two additional terms for the initial rate (0% for 6 months and 0% for 24 months),
- Two additional go-to rates (9.99% fixed and 4.99% fixed),
- And no balance transfer option.
Copyright 2010
American Express
Full Factorial Test Design Example
How do we design this test to gain the most information possible thereby ensuring we can build the best package for the next major campaign?
Copyright 2010
American Express
Full Factorial Test Design Example
You will start with a full factorial test design.
In this example we have three initial rate terms, three go-to rates and two balance transfer options.
Therefore, a full factorial test design will yield 18 = 3 x 3 x 2 test panels as shown on the next slide.
Copyright 2010
American Express
Full Factorial Test Design Example
The resulting full factorial test design is as shown in the figure below:
With such a test design you can determine if initial term, the go-to rate and the balance transfer option are independent of one another or if interactions exist.
However, keep in mind, as previously taught, a full factorial test design is typically not required for most situations.
Copyright 2010
Sheet1
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Missing Test Panel | 0% for 15 months | BT 3.99 LOB | 7.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
Sheet2
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 13 | Test 16 |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | Test 11 | Test 14 | Test 17 |
| Go-To Rate 4.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 | Test 15 | Test 18 |
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | ||
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | |||
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 16 | |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 |
Sheet3
American Express
Full Factorial Test Design Example
The rule you will be asked to follow next is to take the full factorial design and eliminate all cells only keeping those that are testing one change at a time versus the control as shown below:
Keep in mind that at a minimum this must be your test design. This is called your “main effects” design.
And, if you truly believe no interactions will exist among these three factors, then we are done and the above “main effects” design is your final design.
Copyright 2010
Sheet1
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Missing Test Panel | 0% for 15 months | BT 3.99 LOB | 7.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
Sheet2
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 13 | Test 16 |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | Test 11 | Test 14 | Test 17 |
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 | Test 15 | Test 18 |
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | ||
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 4.99% Fixed | Test 3 | Test 6 | Test 9 | |||
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 16 | |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 |
Sheet3
American Express
Full Factorial Test Design Example
Next you will determine which cells to add back in.
In this example, because we can assume possible interactions among the three factors, we will need to add back in some of the interaction test cells.
Given this, the question then becomes “which interaction panels should we add back in?”
Copyright 2010
American Express
Full Factorial Test Design Example
With the “main effects” design shown below again, please realize that you cannot take the plus observed when increasing the initial 0% rate from 12 months to 24 months if you are also going to be taking away the BT option. The plus observed when comparing Test #1 versus Test #7 may be less when the BT option is dropped. Be careful!
So, to determine which cells to add back in, you must determine the other questions besides those on Slide 19 that you will need answered to ensure you will be in a good position to build the best package possible based on your testing results.
Copyright 2010
Sheet1
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Missing Test Panel | 0% for 15 months | BT 3.99 LOB | 7.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
Sheet2
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 13 | Test 16 |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | Test 11 | Test 14 | Test 17 |
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 | Test 15 | Test 18 |
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | ||
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 4.99% Fixed | Test 3 | Test 6 | Test 9 | |||
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 16 | |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 |
Sheet3
American Express
Full Factorial Test Design Example
Questions that you should be concerned about answering are:
Does the drop in response observed when decreasing the initial 0% rate from 12 months to 6 months lessen if the go-to rate is lowered (Test 6)
Does the drop in response observed when dropping the BT option lessen if the initial 0% rate is changed from 12 months to 24 months (Test 16)
Does the drop in response observed when increasing the go-to rate from 6.99% fixed to 9.99% fixed lessen if the initial 0% rate is changed from 12 months to 24 months (Test 8)
Does the plus in response observed when taking the go-to rate from 6.99% fixed to 4.99% fixed lessen if the BT option is dropped (Test 12)
Does the plus in response observed when taking the go-to rate from 6.99% fixed to 4.99% fixed lessen if the initial 0% rate is changed from 12 months to 6 months (Test 6)
Copyright 2010
American Express
Full Factorial Test Design Example
Therefore, our final test design for the Gold Card is as shown below:
Copyright 2010
Sheet1
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 3.99 LOB | 7.99% fixed |
| Test 1 | 0% for 6 months | BT 3.99 LOB | 7.99% fixed |
| Test 2 | 0% for 12 months | BT 3.99 LOB | P + 3.99% fixed |
| Missing Test Panel | 0% for 15 months | BT 3.99 LOB | 7.99% fixed |
| Test 3 | 0% for 15 months | BT 3.99 LOB | P + 3.99% fixed |
| Test 4 | 0% for 15 months | BT 3.99 LOB | P + 5.99% fixed |
| Test 5 | 0% for 15 months | BT 3.99 LOB | P + 2.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 6 | 0% for 15 months | None | P + 2.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
| Test Panel | Initial Rate | Balance Transfer Terms | Go-to Rate |
| Control | 0% for 12 months | BT 0 for 12 | 7.99% fixed |
| Test 1 | 0% for 12 months | BT 0 for 12 | P + 3.99% fixed |
| Missing Test Panel | 0% for 12 months | None | 7.99% fixed |
| Test 2 | 0% for 12 months | None | P + 3.99% fixed |
Sheet2
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 13 | Test 16 |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | Test 11 | Test 14 | Test 17 |
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 | Test 15 | Test 18 |
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | ||
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 12.99% Fixed | Test 3 | Test 6 | Test 9 | |||
| Balance Transfer 1.99 LOB | No Balance Transfer Option | |||||
| Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | Initial Rate 0% for 12 Months | Initial Rate 0% for 6 Months | Initial Rate 0% for 24 Months | |
| Go-To Rate 6.99% Fixed | Test 1 (Control) | Test 4 | Test 7 | Test 10 | Test 16 | |
| Go-To Rate 9.99% Fixed | Test 2 | Test 5 | Test 8 | |||
| Go-To Rate 4.99% Fixed | Test 3 | Test 6 | Test 9 | Test 12 |
Sheet3
Netflix Full Factorial In-Class Exercise
Let’s now try an in-class exercise.
Assume you work for Netflix and are preparing to test some new banner ad variations.
Your current banner is comprised of the following elements:
Yellow background
Young girl with blonde hair image
Headline = first two months free
You have decided to test the following:
Light blue background
Two new images of a young girl (same girl with brown hair, same girl with black hair)
Headline = first month free
All else the same.
Copyright 2010
Netflix Full Factorial In-Class Exercise
How many test cells does a full factorial test design contain?
What does the full factorial test design look like?
Copyright 2010
Netflix Full Factorial In-Class Exercise
What does the “main effects” design look like assuming no interactions among the three factors being tested?
What questions will this “main effects” design answer?
Copyright 2010
Netflix Full Factorial In-Class Exercise
Should we be concerned with any interactions and if so which factors of the three being tested do we believe will interact with one another?
How might we design our final test plan?
Copyright 2010
Netflix Full Factorial In-Class Exercise
Should we be concerned with any interactions and if so which factors of the three being tested do we believe will interact with one another?
How might we design our final test plan?
Copyright 2010
Confidence Interval
Test PanelNamesRRSign @ 95%LBUB
Control Format75,0001.05%------
Test Format75,0001.15%No1.04%1.36%
Test Response RateFrequency
0.60% - 0.70%17
0.70% - 0.80%73
0.80% - 0.90%125
0.90% - 1.00%170
1.00% - 1.10%226
1.10% - 1.20%174
1.20% - 1.30%122
1.30% - 1.40%74
1.40% - 1.50%19
TOTAL1,000
Histogram of Response Rates
0
50
100
150
200
250
300
0.60% - 0.70%0.70% - 0.80%0.80% - 0.90%0.90% - 1.00%1.00% - 1.10%1.10% - 1.20%1.20% - 1.30%1.30% - 1.40%1.40% - 1.50%
Histogram of Response Rates
0
50
100
150
200
250
300
0.60% - 0.70%0.70% - 0.80%0.80% - 0.90%0.90% - 1.00%1.00% - 1.10%1.10% - 1.20%1.20% - 1.30%1.30% - 1.40%1.40% - 1.50%
mean
mean-1sd
For symmetric and bell-
shaped distributions,
the mean falls exactly in
the middle.
For symmetric and bell-shaped
distributions, 50% of the observations fall to
the right side of the mean and the remaining
50% fall to the left side of the mean.
mean-3sdmean-2sdmean+3sd
mean+1sdmean+2sd
Number
of Customers
Mailed
Number of
Customers who
clicked
CTR
Control Banner 100,000 348 0.348%
Larger Banner 100,000 430 0.430%
Confidence Interval
Test PanelNamesRRSign @ 95%LBUB
Control Format75,0001.05%------
Test Format75,0001.15%No1.04%1.36%
Number
of Customers
Blasted
Number of
Customers who
Responded
Response Rate
New Format 1,450,000 2,828 .195%
Number
of Customers
Blasted
Number of
Customers who
Responded
Response Rate
New Format 1,350,000 2,633 .195%
Reverse Test of
Old Format
100,000 176 .176%
Test Panel
Response Rate
L
ift
versus
Control
Control Package
1.10%
--
As Control with New Format and Copy Change
1.09%
-
1%
As Control with New Format
0.99%
-
10%
As Control with Copy Change
1.19%
+8%
Test Panel Description
1 Control
2 Price test - as control with $2 price increase
3 Premium test - as control but with premium for order added
4 Creative layout test - as control but with a new creative
5 Background color change test - as control but new background color
6 Price and premium test - as test panel #3 with $2 price increase
Balance Transfer 1.99 LOB
No Balance Transfer Option
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Go-To Rate 6.99% FixedTest 1 (Control)Test 4Test 7Test 10Test 13Test 16
Go-To Rate 9.99% FixedTest 2Test 5Test 8Test 11Test 14Test 17
Go-To Rate 4.99% FixedTest 3Test 6Test 9Test 12Test 15Test 18
Balance Transfer 1.99 LOB
No Balance Transfer Option
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Go-To Rate 6.99% FixedTest 1 (Control)Test 4Test 7Test 10
Go-To Rate 9.99% FixedTest 2Test 5Test 8
Go-To Rate 4.99% FixedTest 3Test 6Test 9
Balance Transfer 1.99 LOB
No Balance Transfer Option
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Go-To Rate 6.99% FixedTest 1 (Control)Test 4Test 7Test 10
Go-To Rate 9.99% FixedTest 2Test 5Test 8
Go-To Rate 4.99% FixedTest 3Test 6Test 9
Balance Transfer 1.99 LOB
No Balance Transfer Option
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Initial Rate
0% for 12 Months
Initial Rate
0% for 6 Months
Initial Rate
0% for 24 Months
Go-To Rate 6.99% FixedTest 1 (Control)Test 4Test 7Test 10Test 16
Go-To Rate 9.99% FixedTest 2Test 5Test 8
Go-To Rate 4.99% FixedTest 3Test 6Test 9Test 12