Data Science in the bakery
This is another paper on Data Science in Bakeries and Candy Factories, and in previous publications, I've shown that bakeries can be the source of a huge amount of data, and I've shown how that data can be used effectively.
Building econometric models based on the available information generated every day by the company can be a source of great savings and optimization.
Based on the information automatically collected from the bakery operating systems, these models can accurately predict future events with precision for each index, time and location.
Today's the time to discuss a little embarrassing problem of malpractice by candy vendors and bakers.
W-MOSZCZYNSKI-2019-3-03Data science in the field of anomaly detection
One of the most important areas of Data Science is tracking misconduct in millions of banking, telecommunications and settlement transactions, and recently there has been a huge increase in so-called big data analytics to detect and explain anomalies in the scale of giant transaction processes.
Today, every single transaction made with a credit card is analyzed for its characteristics, and if the customer has never paid a large sum of money in the middle of the night, the system will consider such an event to be an anomaly and will ask for additional confirmation via SMS.
Banks lose huge sums of money every year from ATM card theft and debit card number theft.
Any deviation from typical customer behaviour, such as unusual and long connections to a country in Africa, will activate security systems, the connection may be blocked.
Malversations
The scale of transactions in the candy retail network is incomparably smaller than the scale of international capital settlements, banks or retail networks, but if we compare the percentage of malpractice in total sales, we would get similar proportions.
Conscious fraud is usually marginal in size, but it does result in significant losses in the company's operations.
Misappropriation can occur in every area of a bakery’s operations, beginning with flour deliveries, continuing through production and ending with the sale of finished products.
The larger the scale of the activity, the more tempting the malpractices will be. At the same time, the larger the size of the activity and the more diverse the products, the more difficult it is to detect these phenomena.
A simple way to detect malpractice at the candy retail store
There are two types of sales fraud: one is cheating the customer by charging the wrong price or charging a non-existent sale; the other is cheating the employer by selling outside the system.
Both the first and second kinds of malfeasance strike the owners of the candy store because a customer once cheated is unlikely to return to our candy store in the future.
Both types of malpractice can be detected by daily inventorying of candy magazines and daily returns, and that doesn't require data science analysis, but determination and systematicity, and that's how we can find out if the sale doesn't match the inventory.
What if there are five people working in a retail outlet, and some of them are working on a rotating basis, every day in a different store? Or should you punish the crew after the first discovery of wrongdoing?
Once the misconduct is discovered, those with a dirty conscience will stop their fraud for a time, which will make it even more difficult to explain the procedure in the future.
Mass statistical analysis of the data
In order to accurately understand processes and detect their anomalies, it is necessary to use mass data analysis. If a company has several dozen retail outlets, the data scale can reach tens of thousands of transactions per month. The more data, the better they build the phenomenon statistics.
The data is organized into certain patterns defined in statistics, called probability distributions. Each person in sales has an assigned system login to the transaction. Hundreds of transactions made by individual employees will build the statistical characteristics of the transactions made and indicate their uniqueness.
For example, someone regularly sells very few cookies, while the rest of the staff sells a lot of cookies during that time, and the checkout registry doesn't agree more often when a specific person comes to the change, and someone has an exceptionally high number of complaints about cookies, even though there seems to be no explanation for that.
Data must be collected from different sources
The more data there is, the better it describes the reality. But you don't have to rely solely on the information from the sales system. You have to collect data from different, even very exotic, places. Gathering data from different sources can be very effective.
If the candy store's Wi-Fi server recorded dozens of calls and there was no turnover at the time, you might ask what the customers were doing at the candy store?
You can compare sales with camera recordings, you can analyze the ratio of complaints to sales volume or the number of unauthorized openings at the POS, you can analyze the mass of flour imported with the average mass of the last few weeks, you can analyze the duration of transactions or the number of transactions on particularly expensive products.
It is possible to compare the volume of sales of a particular product simultaneously from different stores.
This can easily be done in Excel, and we recommend using statistical tools on a large scale, and data science can be very helpful again.
You need to know your processes
One of the biggest mistakes managers make is to rely on final results without examining the processes in detail.
They know their processes best, they know the flaws in programs, they know how to bypass security and how to falsify documentation, and relying solely on system data without empirical verification can lead to wrong conclusions.
A few years ago, I was working on a sales system in a large cafe in the center of Warsaw. The system had several drawbacks, one of which was cancelling the bill when an additional person was added to the order of a group of people.
The waiters were taking the money into their pockets, and it took us weeks to solve this puzzle, and the inventories showed a lack of raw materials, and the data analysis and the timing, and the waiters and the products they ordered, and they didn't create statistical accuracy because there was no error in the system.
Employees are the first to learn about the loopholes in the security system, and they accidentally discover how to open a tax safe without authorization, how to withdraw an order or reset the system, so besides statistical process control, you have to keep up with the bakery processes yourself.
Data Science, through statistical process control, provides enormous opportunities for monitoring the business, and malpractice is a recurring phenomenon, and once someone cheats on a system, they're going to replicate that behavior in the future.
Data Science data analysis involves combining large amounts of data from different sources, matching them together, and looking for repeated errors and anomalies, so that any anomaly will be explained in the long run.
Even if the transaction is not in the sales system, the event leaves a lot of traces in various other, random records, and usually the mere awareness that someone is analyzing the data is enough to discourage employees from malpractice and fraud.
Wojciech Moszczyński — graduate of the Department of Econometrics and Statistics of Nicolaus Copernicus University in Toruń; specialist in econometrics, finance, data science, and management accounting. He specializes in the optimization of production and logistics processes. He conducts research in the area of the development and application of artificial intelligence. For years he has been engaged in the popularization of machine learning and data science in business environments.

Dodaj komentarz