ATT is a large telecommunications company and they have really good data about phone calls globally.
Class 2: Consumer Data, Target Marketing, Segmentation, Customer Value
What data is available about consumers and where is it?
How to evaluate and use a target marketing model
Marketing segmentation – what it is and how to use it
How to build a customer value measure
1
Homework 1: Prospect Solicitation
2
I’m a senior manager at a large bank. I oversee about 10 products that we provide to our customers. I have about 10 million credit card accounts, but only about 100,000 of these also have an installment loan with us.
I’d like to do a solicitation campaign to try to get people who are currently noncustomers (prospects) to sign up for an installment loan.
2
Consumer Data Generally Available
3
Data in your company - data about your customers: Mostly behavioral data
How/when customers came to you
How they behave on existing accounts (behavioral data)
This data is fairly accurate
Data outside of your company - data about your customers and prospects: Mostly demographic, psychographic data
Age, income, gender, occupation, education, marital status, race, presence of children, household size, spending categories, interests…
This data is approximate, estimated, inferred, modeled…
Comes from many sources: warranty registrations, magazine subscriptions, census bureau, property records, other public records, surveys, professional licenses, online cookies, surveys, social sites
Some Available U.S. Data
4
Consumer demo/psychographics
Estimates at consumer level (about 250 million adults)
Age, gender, income, education, occupation, marital status…
Personal interests: travel, shopping, health, electronics,…
Census bureau
Zip code level summaries. ~40,000 zip codes (ZCTAs)
Population, land area, % races, distributions of age, gender, children, income, home value…
Public records
Home prices, court records, criminal records, sex offenders…
Zillow
Home values, home properties, past sales…
Financial markets
Histories of stocks, bonds, derivatives…
Social network data
PII (name, date of birth…), friends, interests, images…
Pretty good discussion of and list of data selling companies:
https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=&cad=rja&uact=8&ved=2ahUKEwj7g-erpKPsAhXlCjQIHcA0Bi8QFjAAegQIAxAC&url=https%3A%2F%2Fwww.fastcompany.com%2F90310803%2Fhere-are-the-data-brokers-quietly-buying-and-selling-your-personal-information&usg=AOvVaw0UuqxSI8S7MYrnv0xmtY_M
or just search for consumer data for purchase fastcompany.com
Show Infutor
5
Business Problem: Decide Whom to Solicit
Businesses make many offers to many consumers to buy their products. Making these offers costs the business money.
Would like to narrow down the list of people to make offers to in order to lower costs. Don’t make offers to people who are unlikely to accept them.
Solution – A target marketing score that predicts the likelihood a consumer will accept a product offer.
Target marketing scores are used to rank-order prospects for marketing offers.
Sort by the score and make offers only to the top scoring people, who have the highest likelihood or accepting.
Saves substantial marketing costs.
6
Cumulative Lift over random
Response score bin
How to Use a Target Marketing Model
Build a binary classification model to predict who will respond to an offer (respond/don’t respond)
Score a large population of prospects for the offer
Rank order the prospect list by the response score
Offer the product to the top of the list
Penetrate the list as deeply as the financials support: Maximize expected profit
7
Expected profit = Expected revenue – Contact cost
Average revenue x probability of response
Cost
Expected revenue
Expected profit
List penetration %
Offer to top ~25% of the ranked prospects
Show Target Marketing Notebook
8
Citibank Zodiac Story
Asked to build a target marketing model for credit insurance
Asked to build models for 4 additional products
Citi’s existing “Zodiac” system: rotate through 12 products, one each month.
They said they have a problem that they can only offer 12 products because of this Zodiac system.
Assignment: come up with a new system that can offer unlimited number of products.
9
Build a separate model independently for each product (column)
Each model predicts the expected profit if we offer that product to that customer
10
Each number is the expected profitability for that customer for that product offer
Many millions
Many 10’s
Model for Product A
Model for Product B
Model for Product E
Model for Product D
Model for Product C
Model for Product G
Model for Product F
Separate, distinct, independent models
How We Framed/Solved the Problem
Next, Choose Which Offer For Each Customer
First, cross off any disallowed offers (red x’s)
Then run a profit optimization algorithm to select one offer for each customer (green circle) while fulfilling constraints. This requires a complex constrained optimization algorithm.
11
Many millions
Many 10’s
Must offer at least 50,000
Must offer between 100,000 and 400,000
Cannot offer more than 300,000
Divide entities (customers, products, households…) into distinct groups based on some criteria in order to make different actions for the different segments.
We typically use a small to moderate number of segments because we want to take customized actions with each segment. Typically a few to a dozen or so segments.
Example customer segmentation:
Each customer is in one particular segment.
We might have different communications, offers and/or strategies between segments.
How do we decide how to divide?
A business expert uses his judgement and experience, or
We use a machine learning algorithm, typically clustering or a decision tree
Segmentation for Marketing Purposes
12
Income
Age
Golden Eagles
Grey Havens
Steady Success
Blue Neckware
Rising Stars
Algorithms For Marketing Segmentation
13
Decision trees (e.g., CART) (typically supervised): Carves up independent variable space x into boxes based on similarity of a dependent variable y. Metric of closeness is typically similarity of y.
K means (typically unsupervised): Finds natural groupings of points (entities) in independent variable space x. Metric of closeness is typically Euclidean distance in x space, thus need to make sure all dimensions are properly scaled.
Hierarchical Clustering Algorithms (HCA): Iteratively divides/aggregates points into groups based on nearness. Could be top-down or bottom up. Needs a metric of closeness, could be anything.
After we finish the segmentation process we have divided our things (customers, products…) into separate groups
Now we want to describe the unique and differentiated characteristics of the different segments.
Make a table and calculate the averages of important characteristics
Can divide by averages to easily see relative differences
14
How to Describe the Marketing Segments
Younger
Lower income
Lower value
Older
Higher income
High value
Demographic Profiles of PA Population
15
Examine the Differences of Characteristics Across the Segments
Demographic characteristics
12 segments
Yellow is above average
Green is below average
Average
Example slide from a large credit card company’s segmentation project
15
Segment Descriptions and Names
Segment 1: “Diamonds in the Rough”
Age (18+, avg: 49), Income (< $35K, avg: $25K), BCI (< $3500, avg: $1700)
Low income across all age groups. Very high response rates.
Single and younger, fresh faces and working class who could be new to the credit world. Regardless, this group is credit hungry and possesses the youngest average revolving age on trades across all the segments. These “Diamonds in the Rough” strive to establish their place and could have the most potential to bloom into valuable customers. They are not likely to be home owners, so probably are not bogged down yet by mortgage payments. This group is the least likely to have opened a bank card recently. Because of this, these fresh faces are highly responsive but somewhat riskier.
Segment 2: “Golden Years”
Age (49+, avg: 61), Income ($35K-$100K, avg: $64K), BCI (< $3500, avg: $2100)
Middle income and older with very low revolving balances.
These are older couples who may be close to retirement. They have more time on their hands to entertain themselves and their grandchildren. Their income will not increase with their growing age, yet these individuals tend to respond favorably to new credit offers. They are not likely to have opened a bank card recently so they tend to be more aware of and possibly seek out “good deals”.
* Averages Based on Demographic Profiles for PA and SC
16
Example slide from a large credit card company’s segmentation project
16
Customer Value Measures (CVM)
17
What is it: Measure the value of each customer to my business
Why: Use for diagnosis, analysis, understanding of who brings what value, differentiated treatments, prospect targeting
Generally, Customer Value Measures are a measurement formula, combining known data values, with perhaps some models
How: Basic form:
CVM = profit from each customer = revenue – cost
Sometimes add in “expert factors”, driven primarily by marketing people rather than financial
One Simple Way to Make a Customer Value Measure
18
Start with a base value:
Base value = profit from that customer = revenue – cost
Use whatever data is available to measure revenue and cost at the customer level. Do the best you can.
We generally don’t include fixed costs, that are the same for each customer (facilities, corporate overhead…). Only include revenue and cost explicitly for that customer.
Interview the people requesting the measure. Are there any other possibly important factors?
Add in expert factors as multiplicative enhancements:
CVM = base value x f1 x f2 x f3…
Can use smoothing formula for the expert factors:
This formula gives up to a 50% boost depending on the value of “field”.
Function to Smoothly Transition Between Values
19
Frequently we want a value to smoothly transition from one number to another. Can use a logistic formula:
“Smoothed” value
Two transition parameters:
nmid is the value of n where the smoothed value is halfway between Ylow and Yhigh
c is a measure of how quickly it transitions
Smoothing counter n
Could be integer or continuous
Yhigh
Ylow
nmid
2c
Different parameter choices make different shapes
nmid = 30
c = .1
nmid = 50
c = 10
nmid = 20
c = 3
Can make a step function with a low value of c
The value smoothly transitions between Ylow and Yhigh as the counter n increases
Homework 2: ATT New Business
20
ATT is a large telecommunications company and they have really good data about phone calls globally.
They would like to build new business services around the nice data assets that they have, similar to a data broker. One beautifully rich data set they have is Call Detail Records (CDRs). They would like to build a data broker service around this internal data. (google: call detail records).
There are privacy restrictions around data that they can sell, but they can derive new data fields from their rich proprietary data. If averaged at a large enough geographical region (zip9, zip5, zip3?) these derived fields can be sold.
They see a company called Claritas and others that make a lot of money selling highly descriptive consumer segments (google: Claritas PRIZM segments).
An ATT executive want to build such a data broker business, selling such demographic-like segments built using the CDRs. Let’s start with the U.S. only. Design a process to build PRIZM-like segments using CDR records.
20
Product AProduct BProduct CProduct DProduct EProduct FProduct G
Customer 16456415533562
Customer 22323435314246
Customer 315425464268384
Customer 4635763776441
Customer 536523129742325
Customer 64383975253524
Customer 72729151542246
Customer 841363734367373
Customer 952415915195226
Customer 1013274764244135
Customer 112294235422584
Sheet1
| Product A | Product B | Product C | Product D | Product E | Product F | Product G | … | |
| Customer 1 | 64 | 5 | 64 | 15 | 53 | 35 | 62 | |
| Customer 2 | 23 | 23 | 43 | 53 | 14 | 24 | 6 | |
| Customer 3 | 15 | 42 | 54 | 64 | 26 | 83 | 84 | |
| Customer 4 | 6 | 35 | 76 | 37 | 7 | 64 | 41 | |
| Customer 5 | 36 | 52 | 31 | 29 | 74 | 23 | 25 | |
| Customer 6 | 43 | 83 | 9 | 75 | 25 | 35 | 24 | |
| Customer 7 | 27 | 29 | 15 | 15 | 42 | 24 | 6 | |
| Customer 8 | 41 | 36 | 37 | 34 | 36 | 73 | 73 | |
| Customer 9 | 52 | 41 | 59 | 15 | 19 | 52 | 26 | |
| Customer 10 | 13 | 27 | 47 | 64 | 24 | 41 | 35 | |
| Customer 11 | 2 | 29 | 42 | 35 | 42 | 25 | 84 | |
| … |
Sheet2
Sheet3
Income
Age
Golden Eagles
Grey Havens
Steady Success
Blue Neckware
Rising Stars
Income
Age
Golden Eagles
Grey Havens
Steady
Success
Blue Neckware
Rising Stars
Segment 1 Segment 2 Segment 3 Segment 4 … Average Age 62 37 25 46 … 42.5
Income 120,000 96,000 56,000 74,000 … 86,500 Monthly Spend 62 45 18 24 … 37.25 Customer Value 7.4 4.2 1.6 3.2 … 4.1
… … … … … … …
Segment 1Segment 2Segment 3Segment 4…Average
Age62372546…42.5
Income120,00096,00056,00074,000…86,500
Monthly Spend62451824…37.25
Customer Value7.44.21.63.2…4.1
… ………………
Segment 1 Segment 2 Segment 3 Segment 4 … Average Age 1.46 0.87 0.59 1.08 … 1
Income 1.39 1.11 0.65 0.86 … 1 Monthly Spend 1.66 1.21 0.48 0.64 … 1 Customer Value 1.80 1.02 0.39 0.78 … 1
… … … … … … …
Segment 1Segment 2Segment 3Segment 4…Average
Age1.460.870.591.08…1
Income1.391.110.650.86…1
Monthly Spend1.661.210.480.64…1
Customer Value1.801.020.390.78…1
… ………………
DIRGDYSRDSKLDLPFSGASACACWOPMGTACBHWTALL
Age
48614053535252535253495152
Household Income
$25,562$64,791$61,800$124,525$105,942$55,889$63,461$128,793$86,589$107,386$26,487$55,768$86,380
Gender (Male)
49.9%48.6%49.0%47.0%47.5%48.6%48.1%46.8%47.3%47.2%48.4%49.2%47.8%
Married
32.7%57.0%50.1%67.1%70.7%60.2%61.4%77.3%0.0%100.0%38.6%55.2%65.3%
# of children
0.160.370.310.440.520.370.430.610.330.670.220.380.48
1 Child
5.21%9.05%7.20%8.89%10.14%10.30%10.44%10.83%4.99%13.76%6.62%9.68%10.16%
2 Children
1.49%2.37%2.14%2.15%2.76%3.34%3.46%3.07%1.19%4.82%2.09%3.07%3.21%
3+ Children
2.77%7.82%6.52%10.21%12.11%6.58%8.48%14.70%8.53%14.42%3.62%7.56%10.52%
Missing
90.53%80.77%84.14%78.75%74.99%79.80%77.62%71.40%85.28%67.00%87.66%79.70%76.10%
Home value
<$150K
77.6%67.3%66.0%34.1%37.7%72.8%49.8%22.3%24.7%20.2%65.2%64.8%34.6%
$150-$250K
13.9%25.4%25.6%39.3%40.5%21.5%35.9%41.8%33.3%35.4%20.8%26.0%34.6%
$250K+
8.6%7.3%8.5%26.6%21.8%5.7%14.3%35.9%42.0%44.4%14.1%9.2%30.8%
Wealth
Highest 20%
5.2%9.0%10.6%26.9%27.7%12.8%23.5%44.9%47.7%52.8%12.0%14.0%34.9%
High 20%
11.5%19.9%21.5%28.9%31.1%23.1%28.9%31.0%25.4%24.7%19.3%25.2%25.4%
Middle 20%
19.7%24.8%25.1%22.9%23.1%27.0%24.2%15.8%14.8%13.2%24.1%27.0%18.8%
Low 20%
28.9%25.5%23.9%14.4%13.3%24.2%16.6%6.6%8.7%7.0%26.8%22.6%13.4%
Lowest 20%
34.7%20.8%18.9%7.0%4.8%12.9%6.9%1.7%3.4%2.3%17.8%11.2%7.4%
Education
College graduate
19.6%21.4%20.8%28.6%31.3%24.1%31.8%40.5%42.2%46.9%25.8%26.0%37.2%
Some college
27.0%30.5%29.4%31.8%29.4%26.9%26.6%24.8%22.3%20.0%28.0%30.2%24.5%
High school graduate
48.5%43.3%44.8%34.0%34.4%44.4%36.8%29.1%29.3%27.3%41.5%39.7%33.0%
LOR
61091111109118116910
Dwelling (single)
61.3%80.2%74.6%82.9%89.3%86.9%88.3%93.0%84.1%96.4%73.2%83.2%87.9%
Occupation
Professional
23.8%27.8%27.7%34.0%35.3%29.9%35.2%38.7%36.9%40.6%29.8%32.2%36.6%
Admin/management
8.9%11.2%10.9%14.2%14.4%12.0%14.1%16.7%14.6%18.3%11.1%12.4%15.6%
Sales/service
8.0%6.0%6.1%6.0%6.2%6.0%6.6%6.5%7.9%7.1%7.9%6.5%6.8%
Clerical/white collar
12.4%12.1%12.1%9.9%9.9%11.2%9.5%8.6%8.7%6.6%11.0%11.1%8.6%
Blue Collar
22.9%24.6%23.6%16.4%16.4%23.3%17.2%11.5%9.9%9.5%18.2%20.8%14.0%
Others
24.2%18.2%19.6%19.5%17.8%17.7%17.5%18.0%22.0%17.8%22.0%17.0%18.3%
Affluence
High
10.6%20.4%26.3%42.4%43.6%21.7%35.5%60.8%60.3%61.9%18.2%30.1%45.6%
Medium
45.1%47.1%45.9%43.3%44.0%49.7%46.1%32.7%30.2%29.6%49.0%46.7%38.2%
Low
44.3%32.6%27.8%14.4%12.3%28.6%18.4%6.5%9.5%8.5%32.8%23.2%16.2%
%Pop 4.95%5.48%4.56%1.70%11.01%3.27%11.70%7.03%11.52%31.02%3.26%4.51%100.00%
%Response 0.59%0.42%0.40%0.30%0.23%0.27%0.24%0.18%0.21%0.14%0.41%0.34%0.25%
%Book 0.33%0.27%0.24%0.21%0.18%0.21%0.18%0.14%0.16%0.11%0.27%0.24%0.17%
f1(field) = 1 + 1.5 � 1
1 + e�(field�fieldmid)/c
< l
a t
e x
i t
s
h a
1 _
b a
s e
6 4
= "
o J
V 5
6 x
8 1
r d
v d
b E
w Q
y U
E f
V R
f Y
8 d
A =
" >
A A
A C
S 3
i c
b V
D L
S 8
M w
H E
7 n
e 7
6 q
H r
0 E
N 0
E R
Z y
O K
e h
B E
L x
4 V
n A
r b
L G
n 2
6 x
Z M
2 p
K k
w i
j 9
/ 7
x 4
8 e
Y /
4 c
W D
I h
5 M
t x
3 m
4 w
e B
L 9
8 j
j y
9 I
B N
f G
8 1
6 c
0 t
j 4
x O
T U
9 E
x 5
d m
5 +
Y d
F d
W r
7 W
c a
o Y
1 F
k s
Y n
U b
U A
2 C
R 1
A 3
3 A
i 4
T R
R Q
G Q
i 4
C e
7 P
C v
3 m
A Z
T m
c X
R l
e g
m 0
J O
1 E
P O
S M
G k
v 5
b l
C t
h j
7 Z
y J
p K
4 p
C D
a O
e b
+ B
g T
v I
U z
U t
v H
2 5
g 0
Y 5
s f
M I
U J
8 r
t s
e 9
R v
P S
M 7
v 4
8 l
L 8
7 Z
w S
z P
q 1
X f
r X
g 1
r z
/ 4
L y
B D
U E
H D
u f
D d
5 2
Y 7
Z q
m E
y D
B B
t W
4 Q
L z
G t
j C
r D
m Y
C 8
3 E
w 1
J J
T d
0 w
4 0
L I
y o
B N
3 K
+ l
3 k
e N
0 y
b R
z G
y q
7 I
4 D
4 7
m s
i o
1 L
o n
A +
u U
1 H
T 1
b 6
0 g
/ 9
M a
q Q
k P
W x
m P
k t
R A
x A
Y X
h a
n A
J s
Z F
s b
j N
F T
A j
e h
Z Q
p r
h 9
K 2
Z d
q i
g z
t v
6 y
L Y
H 8
/ v
J f
c L
1 b
I 3
u 1
o 8
v d
y s
n p
s I
5 p
t I
r W
0 A
Y i
6 A
C d
o H
N 0
g e
q I
o U
f 0
i t
7 R
h /
P k
v D
m f
z t
f A
W n
K G
m R
X 0
Y 0
o T
3 +
E H
r j
Q =
< /
l a
t e
x i
t >
Value = Ylow + Yhigh � Ylow
1 + e�(n�nmid)/c <latexit sha1_base64="jZkk48A5rkKp6gru5QJXUklwLsg=">AAACO3icbVC7TiMxFPWwPMMrQEljkSCBUMJMKNgGKWIbyoBIACVh5HFuEgs/RraHVTSa/9qGn6DbZpstQIiWHidMwetIls49517b90QxZ8b6/l9v6sf0zOzc/EJhcWl5ZbW4tt4yKtEUmlRxpS8jYoAzCU3LLIfLWAMREYeL6ObX2L+4BW2Ykud2FENXkIFkfUaJdVJYPCuX044WuEV4Ahk+wlfhpObqd7aX5sWQDYZZ5Z3TUe5OHOzBdVrZkRUZpoL1st19mmXlclgs+VV/AvyVBDkpoRyNsHjf6SmaCJCWcmJMO/Bj202JtoxyyAqdxEBM6A0ZQNtRSQSYbjrZPcPbTunhvtLuSIsn6vuJlAhjRiJynYLYofnsjcXvvHZi+z+7KZNxYkHSt4f6CcdW4XGQuMc0UMtHjhCqmfsrpkOiCbUu7oILIfi88lfSqlWDg2rttFaqH+dxzKNNtIV2UIAOUR2doAZqIor+oH/oAT16d95/78l7fmud8vKZDfQB3ssrGe+s7A==</latexit>
0
0.5
1
1.5
2
2.5
0 10 20 30 40 50 60 70 80 90 100
0
0.5
1
1.5
2
2.5
0102030405060708090100
0
0.5
1
1.5
2
2.5
0 20 40 60 80 100
0
0.5
1
1.5
2
2.5
0 20406080100
0
0.5
1
1.5
2
2.5
0 20 40 60 80 100
0
0.5
1
1.5
2
2.5
0 20406080100
0
0.5
1
1.5
2
2.5
0 20 40 60 80 100
0
0.5
1
1.5
2
2.5
0 20406080100