ATT is a large telecommunications company and they have really good data about phone calls globally.

profileMichelle_Michy
20201016222929dso_510_example_homework3.pdf

Coupon Optimization

Solution Overview:

Find what to print on coupon based on the association rules.

Solution Details:

1. Generate frequent itemsets from a list of items using Apriori principle: Support > minsup • What I want to do in Step 1 is to find all frequent itemsets with support > minsup (as

shown in the graph below).

Here, I hope to clarify a few concepts I will use to build up the solution: itemsets is the combination of items sold in Home Depot support is the fraction of transaction containing certain itemset.

Support of X = +,-./-0123./ 03.1-2.2.4 5

+31-6 .789:, 3; 1,-./-0123./ minsup is a threshold decided based on professional judgement and used to filter the itemsets. Apriori principle is a methodology that helps us find frequent itemsets in a more efficient way (compared to create millions of combinations of all products and then compare their support to the minsup). It tells us that all subsets of a frequent itemsets must also be frequent. That is, if we remove an item from a certain item, the support of the itemset will either go up or remain the same. In other words, if we realized the support of {light, toilet paper} is below the minsup, and an itemset with any item added to it, for example, {light, toilet paper, pillow} will never go beyond minsup too. The reason I want to filter the itemsets is that, if some itemset have a very low support, then I don’t have enough information on the association between its items, hence we can barely draw conclusion from such rule. Let’s get started with generating frequent itemsets.

• Create itemsets having only one item sold in Home Depot. The itemsets will look similar to this: {light bulb}, {toilet paper}, {pillow}, {quilt}.

• Seek frequent itemsets. Choose itemsets with Support > minsup, let’s say we filtered out {quilt}.

• Generate all itemsets of length 2 and then trim the ones with support < minsup. Itemsets of length 2 will looks like: {light bulb, toilet paper}, {toilet paper, pillow}, {light bulb, toilet paper}

• Generate all itemsets of length N+1 and then trim the ones with support < minsup.

• Stop until when all itemsets of length K are below the minsup. Now we have the Maximal frequent itemset, where no item can be added so that the itemsets still remains above the minsup threshold.

2. Generate all possible rules from the frequent itemsets: lift > minconf • What I try to do in Step 2 is to identify rules with lift > minconf

Like what I did in Step 1, I’d like to define the concepts I will use in the following discussion: Support of X to Y: The fraction of having both itemset X and itemset Y in the one transaction

Confidence of X to Y: The conditional probability of occurrence of Y (the consequent) in the cart given that the cart already have itemsets X (the antecedent)

Minconf: It is the threshold we built to select rules that fall above a minimum confidence level Anti-monotone property: It’s a trait of confidence of rules that helps us filter in a more efficient way. It tells us: Confidence of (A, B, C→ D) ≥ (B, C → A, D) ≥ (C → A, B, D). Because the numerator is the same for all three, and the denominator keeps increasing (transaction containing D is > transaction containing A and D > transaction containing A, B and D). Thus, the more items we have in X, the higher the confidence. That is to say, we can trim the rules in the way we did for the frequent itemsets (as shown in the graph below)

Let’s start with selecting association rules

• Start with one of the frequent itemsets. Form rules with only one consequent (Y). For example, for a frequent itemset {A, B, C, D}, the rule will look like this {A, B, C à D}, {B, C, D à A} etc.

• Trim the rules with confidence < Minconf

• Form new rules using a combination of consequences (Y) from the remaining ones, prune the ones < Minsconf

• Repeat until there is only one item left in X (the antecedent) 3. Seek the subset of rules that give the highest lift

• What I want to do in Step 3 is to identify subset of rules that give the highest lift. The reason why I hope to calculate lift is that, if we only measure rules using confidence, there’s a big problem. If the consequence is frequent (for example toilet paper), the confidence is always going to be high even if the association between X and Y is weak. By calculating the lift of rules, we take into consider the increase in occurrence of Y. Lift of X to Y: It measures the lift that {X} provides to our confidence for having Y on the cart. It is the probability of having {Y} in the cart given {X} is there, over the probability of having {Y} in the cart

4. Use the selected rules with the highest lift to determine the coupon

After the previous three steps, we got a rule, for example, saying: {toilet paper, quilt, pillow} has the highest lift to {sheet}, then we will print out coupon related to sheet on the receipt that has toilet paper, quilt and pillow in the purchase list.

References [1] Anisha Garg (2018, Sep 17). Complete guide to Association Rules (2/2) Retrieved December 4, 2019, from https://towardsdatascience.com/complete-guide-to-association-rules-2-2- c92072b56c84 [2] Anisha Garg (2018, Sep 3). Complete guide to Association Rules (1/2) Retrieved December 4, 2019, from https://towardsdatascience.com/association-rules-2-aa9a77241654