July 2022 | Przegląd Piekarski i Cukierniczy (Baking and Confectionery Review)
The application of association rules of the Apriori algorithm
Today, practically every object and every service can be bought over the internet. The scale of internet sales has considerably exceeded the volume of traditional sales. The retail revolution has not bypassed the confectionery and baking trade either.
What are association rules?
At the beginning, internet shops conducted a fairly passive sales policy. Goods were displayed on websites, there were promotions, sales and basic prices — in a word, everything we see in traditional shops was there. With time the internet filled up with shops which began to compete intensively with one another.
Competing in internet sales does not consist solely of a price race or a position in displayed search results. On the internet the most important thing is to gain the buyer’s interest, because this increases popularity. A relationship with the buyer and their interest can also be built by cleverly suggesting to them what to buy and what is the best choice for them. Usually buyers on the internet enter into a deeper relationship with the seller (who, by the way, is not a person but a machine) than if they were buying in an ordinary shop. An internet buyer compares, reads descriptions, characteristics and opinions, looks for bargains, and is very susceptible to any suggestions and prompts on the part of virtual assistants.
W-MOSZCZYNSKI ppic 7-22When we buy a product on the internet, a suggestion appears that we should also buy other goods. These suggestions are most often astonishingly accurate. We often discover that we had forgotten something which the shop’s virtual assistant has just reminded us of. Behind this seemingly simple mechanism stand very sophisticated algorithms, which I will now briefly describe.
Grouping customers
Grouping customers is the most commonly used method serving to recommend products to customers. The algorithm collects data from the search history conducted by customers. Then the data is filtered. On the basis of the choices made, the system creates groups of customers behaving similarly.
In the past, in classic marketing, we dealt with customer segmentation, grouping them according to age, education, or, for example, social status. In this case the choice of customers for clusters was based solely on the characteristic of customer behaviour.
In the case of internet shopping, customer behaviour is recorded in the registers of websites. There one can see what the customer did, how much time they devoted to reading opinions, where they looked, and what products they were interested in. On the basis of an enormous quantity of data, a relatively simple grouping algorithm creates several or a dozen or so groups called clusters. People in clusters usually have very similar personality traits and ways of behaving, and, what is most important, they look for similar things.
The machine adapts to the behaviour of each group, applying a huge range of marketing techniques and tricks. To some it suggests opinions, to others it proposes additional products.
The mechanism groups similar customers, looking for common patterns of behaviour, similar preferences and temperaments. At the same time the algorithm filters purchases made in clusters in order to find similar products, in order to recommend them to potential customers.
The Apriori algorithm for grouping pairs of products
Searching for association rules is a common technique for preparing an offer for the customer. Let us use the example of a chain of confectionery cafés.
A chain of cafés decided to offer customers a second cake at 80% of the price. Thanks to this the confectioner’s shop hoped to reduce the quantity of returns by encouraging customers to make larger purchases.
It was not wanted to offer every cake to every customer. It was wanted to do this seemingly occasionally. This trick was meant to make the customer feel singled out. The problem, however, was that it was not very clear what „exceptional” cake to offer the „exceptional” customer. Of course one can propose whichever is currently on promotion. However, customers of the confectioner’s shop usually have their own tastes, and such a proposal might not reach them. Besides, a cake on promotion should be sold at at least a 50% discount, and they wanted to sell the second cake at 80% of the basic price. Besides, the point was to reduce the size of the returns, and cakes from the promotion will not solve this problem.
To find out what customers bought most often, it was enough to review all the sales transactions from the last 3 years. There were 37 thousand of them. Fortunately the chain of confectioners’ shops had a computer with a spreadsheet. The bookkeeper downloaded the sales register from the program and opened it in Excel.
By means of simple tables, the data was transformed into the summary contained in Table 2. For the summary to be legible, the names of the goods were replaced by their abbreviations (e.g. the abbreviation CH denotes coffee, and MK denotes poppy-seed cake).
Table 1. A summary of the confectioner’s shop’s sales
| Receipt date | Goods | Sale price | Transaction number |
|---|---|---|---|
| 2019-01-02 | WZka | 2.70 | 1 |
| 2019-01-02 | Galaretka | 2.90 | 2 |
| 2019-01-02 | WZka | 2.70 | 3 |
| 2019-01-02 | WZka | 2.70 | 4 |
| 2019-01-02 | WZka | 2.70 | 5 |
| … | … | … | … |
| 2021-06-30 | WZka | 2.70 | 1599624 |
| 2021-06-30 | Makowiec | 3.10 | 1599625 |
| 2021-06-30 | WZka | 2.70 | 1599626 |
| 2021-06-30 | Szarlotka | 2.40 | 1599627 |
| 2021-06-30 | Szarlotka | 2.40 | 1599628 |
Table 2 presents a purchase matrix, where the columns denote the sales of 14 defined goods, and the rows — the individual receipts. When a customer bought coffee (denoted by the abbreviation CH) and cheesecake (denoted as SP), ones appear in the table in the columns CH and SP. For the remaining products in that transaction, zeros appear, because the customer did not buy those products.
Table 2. The purchase matrix (columns: CH – coffee, DA – …, HV – fruit sorbet, HY – …, IN – …, IT – jelly cake, MK – poppy-seed cake, NA – …, PD – …, SP – cheesecake, ST – cheese strudel, SZ – apple pie, TM – …, WE – …)
| Receipt no. | CH | DA | HV | HY | IN | IT | MK | NA | PD | SP | ST | SZ | TM | WE |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 1 |
| 1 | 1 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
| 3 | 1 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| 4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 1 |
| … | … | … | … | … | … | … | … | … | … | … | … | … | … | … |
| 35467 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 |
| 35468 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 35469 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| 35470 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 1 | 0 | 1 | 0 | 0 |
| 35471 | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 1 |
The practical application of the Apriori algorithm
In our example, the owner of the confectioner’s shop wanted to sell the next cake at 80% of its value. However, they wanted to sell at a reduced price specific cakes, which would be proposed to the buyer during the purchase. The goal was to reduce the size of the daily returns and to accustom customers to larger purchases at the confectioner’s shop. For the action to succeed, it is necessary to offer customers those cakes which they will want to buy. So one has to find pairs of products which sell well together. The shop assistant at the confectioner’s shop, after ringing up the first cake on the till, should receive a prompt from the system as to what cake to propose next.
In order to launch the Apriori algorithm, we enter the free Jupyter Notebook environment and import the appropriate libraries.
from mlxtend.frequent_patterns import apriori
from mlxtend.frequent_patterns import association_rules
Then we enter the command for creating the table of association rules. As can be seen, the whole task took us four lines of Python code.
rules = association_rules(frequent_itemsets, metric=”lift”, min_threshold=0.01)
rules
The Apriori algorithm found 47 association rules; 12 of them are in Table 3.
The second column in the first row means that the algorithm chose coffee (CH) and noticed that apple pie (denoted by the abbreviation SZ) and cheese strudel (denoted as ST) are also often bought with coffee. So with coffee, customers often bought two further specific cakes as well.
The antecedents column denotes the preceding purchase, whereas the consequents column denotes the subsequent purchase. It is worth noting that on the receipts there is no order of purchases made. The algorithm carries out a simulation of the probabilities of the order of purchases.
Association rules
Below we will discuss some of the association rules which appear in Table 3.
Table 3. The association rules table
| antecedents | consequents | antecedent support | consequent support | support | confidence | lift | leverage | conviction | |
|---|---|---|---|---|---|---|---|---|---|
| 39 | (CH) | (SZ, ST) | 0.313402 | 0.035972 | 0.020439 | 0.065215 | 1.812948 | 0.009165 | 1.031284 |
| 47 | (SZ) | (CH, WE) | 0.304522 | 0.048658 | 0.024583 | 0.080726 | 1.659041 | 0.009765 | 1.034884 |
| 23 | (SZ) | (PD) | 0.304522 | 0.055706 | 0.026218 | 0.086095 | 1.545530 | 0.009254 | 1.033252 |
| 10 | (CH) | (TM) | 0.313402 | 0.059286 | 0.027825 | 0.088783 | 1.497531 | 0.009244 | 1.032371 |
| 33 | (SZ) | (TM) | 0.304522 | 0.059286 | 0.020100 | 0.066006 | 1.113350 | 0.002046 | 1.007195 |
| 45 | (CH) | (WE, SZ) | 0.313402 | 0.065291 | 0.024583 | 0.078438 | 1.201368 | 0.004120 | 1.014267 |
| 40 | (SZ) | (CH, ST) | 0.304522 | 0.065601 | 0.020439 | 0.067117 | 1.023112 | 0.000462 | 1.001625 |
| 14 | (IT) | (HV) | 0.099853 | 0.066278 | 0.026556 | 0.265951 | 4.012688 | 0.019938 | 1.272017 |
| 21 | (ST) | (MK) | 0.140167 | 0.068251 | 0.021087 | 0.150442 | 2.204253 | 0.011521 | 1.096746 |
| 2 | (CH) | (MK) | 0.313402 | 0.068251 | 0.028699 | 0.091571 | 1.341687 | 0.007309 | 1.025671 |
| 15 | (HV) | (IT) | 0.066278 | 0.099853 | 0.026556 | 0.400681 | 4.012688 | 0.019938 | 1.501948 |
| 46 | (WE) | (CH, SZ) | 0.149216 | 0.107606 | 0.024583 | 0.164746 | 1.531010 | 0.008526 | 1.068410 |
Antecedent support — measures the probability of the purchase of the chosen product. In the first row is coffee (denoted as CH). It appears in 31% of all purchases made in the cafés.
Consequent support — measures the probability of the purchase of the chosen product or pair of products. In the first row, in the consequents column, is apple pie and strudel. Receipts on which the customer bought these two cakes at the same time make up 3.5%. Customers rarely buy these two cakes at the same time.
Support — says how popular the set of products in the antecedents and consequents columns is. The simultaneous purchase of coffee, apple pie and strudel constitutes barely 2% of all transactions. For this association rule, the order in which the products were bought does not matter.
Confidence — this is an association rule which says what the probability is that if someone buys one item, they will also buy the other. In the first row, when someone buys coffee, they also buy apple pie and strudel. The probability of this phenomenon amounts to 6.5%. Please note another pair from Table 3. If a customer buys jelly cake (denoted as IT), then according to the confidence rule, there is a 27% chance they will also buy fruit sorbet (denoted as HV). In turn, when a customer orders fruit sorbet first, there is a 40% chance they will also order jelly cake. Both products are strongly connected with each other, even though they are rarely bought by customers (as can be seen in the antecedent support column).
Lift — measures the relationship between the probability of a given rule occurring and its expected probability if the items were independent of one another. Lift categorically determines the possibility of the purchase being analysed. When lift is greater than zero, the second item will probably be bought.
How does one make use of association rules?
The detected association rules define various features. High support means that the rule occurs very often, while the accompanying high confidence rule indicates that it has great predictive power (which means that when the purchase of the first product occurs, the next product from the association rule usually also appears). Unfortunately, there are no defined thresholds or reference values which could tell us whether the associative bond is relatively strong or not. The assessment of the bond depends on many things, including the specifics of the object of study. This assessment will be based to a considerable degree on the researcher’s subjective feelings.
A classic example of assessment is the relationship between caviar and champagne. Both products are bought relatively rarely individually, but are often bought together. It is therefore expected that the association rule will show a high level of lift and a low level of support. We had a similar situation with the products jelly cake (IT) and fruit sorbet (HV).
Summary
Association rules provide objective information about the mutual relations between products. Nevertheless the final word lies with the researcher, who has to say to what extent the rule is useful. Subjective feelings and conclusions based on the researcher’s experience are an important element of assessing patterns and similarities.
Wojciech Moszczyński
Wojciech Moszczyński — graduate of the Department of Econometrics and Statistics of Nicolaus Copernicus University in Toruń; specialist in econometrics, finance, data science, and management accounting. He specializes in the optimization of production and logistics processes. He conducts research in the area of the development and application of artificial intelligence. For years he has been engaged in the popularization of machine learning and data science in business environments.

Dodaj komentarz