ATT is a large telecommunications company and they have really good data about phone calls globally.
Problem Statement Home Depot wants to use historical transaction data to generate targeting coupons to increase sales. There are three main things to consider: - Coupons about what the customers bought - What might the customers buy next - From what other categories Data Structure The data will be organized by the transaction level which could be used for models and forecasts. Home Depot will have a great accuracy on goods sold on each transaction and the data will most likely to look like the following:
Transaction ID Product ID Product Detail Categor Quantit Unit Price Discount % Total
1 G43 Banana Grocery 3 0.96 0 2.89
1 S356 Skateboard Sports 1 74.28 10% 66.85
2 B125 Beer Beverage 12 1.63 20% 15.65
The same transaction ID could have different products and totals because each transaction could have multiple items. Association Rule Mining The main purpose of association rule mining is to discover strong measures of relationships in databases. This is a very common practice where POS system data is used to discover regularities between products. The most classic example is the discovery of a strong association between beer and diapers. The transaction data will be aggregated to a matrix only involving item names and its associated transaction ID. Each row will contain 1 if the item is present in the transaction, or 0 if it is absent. The data will look something like this: (for simplicity sake assume only these 5 items could be discounted by a coupon)
Transaction ID pl wood saw rebar ruler screws
1 1 1 0 0 0
2 0 0 1 0 0
3 0 0 0 1 1
4 1 1 1 0 0
5 0 1 0 0 0
Then an association rule for this data set is rebar, saw => plywood , meaning that customers buying rebar and saw also bought plywood. For Home Depot s matrix dataset, association rules need to be supported by hundreds or thousands (threshold determined by exports) of examples to make it a rule. Rules for Home Depot s case means that customers who buy a particular itemset X of items are most likely interested in the Y item(s) as well. Additionally, an itemset is defined as a combination of two or more items in the dataset. For itemsets, there are mathematical formulas to calculate the: ¾ S o How often the itemset appears in the dataset ¾ Confidence How often a rule is found to be true
¾ Lif The test for independence: if 1 then independent and rule is useless. <1 means negative association and >1 means the rule is dependent therefore has strong implication of the rule
¾ Con ic ion ratio of the frequency in an itemset X occurs without Y ¾ R le Po e Fac o intensity of positive relationship of items in the rule With the aforementioned concepts, we can use expert advice to set a constraint for the minimum support and minimum confidence to filter for only the top association rules. Out of those rules, we can check the lift (the larger the better), the conviction (the smaller the better) and rule power factor (the larger the better). By rank ordering the top-ranking association rules by support or confidence, we could employ coupon strategies to boost revenue for Home Depot. For coupon distribution, the top ranking association rule could be for example screws, ruler => plywood with high support and high confidence, we could offer screw coupons for customers who buy ruler or ruler coupons to those who buy screws as well as offer combination of screws and ruler buyers plywood coupons. If the accuracy of the rules needs to be tested, we can split the data into training and testing. Associate rules will be mined from the training set and tested on the test set to see the accuracy. Lastly, it is advised to give out coupons first in smaller scales as an experiment using the associate rules and test if the coupon strategy indeed increased revenue. Supplementar Solutions If we are unable to match a strong association rule with a purchase, we could either: 1) Give a coupon based on seasonality or time series trends 2) Give no coupons if no good trends match or if the customer is unlikely to return (could be proven
if a local store asks for zip code and customer gives a zip code outside a set radius) For the first option based on seasonality, we can explore historic data and see what items are most bought in a particular time of the month or year. Based on this, we can give out coupons on items that is going to up trend in sales in advance. It is important to note that this solution might need to only employed on items that have a high margin. Additionally, it is best to conduct an experiment first to confirm the actual boost in revenue and compare it with past data when the coupon was not being given out (some items will sell regardless if discounts are present, so therefore we need to ensure that these items are either excluded or more quantities of them need to be sold to accommodate the loss on profit from discounts.) Additionally, any trendy items that are up trending in sale also could be in included in this supplementary solution. We need to perform time series models to forecast the demand of these items. This strategy also requires an experiment stage where boost of sales is confirmed in a smaller scale before full deployment.