December 2021 | Przegląd Piekarski i Cukierniczy (Baking and Confectionery Review)
Data scientist in the bakery
For some time I have been publishing articles promoting methods of optimisation in everyday business activity. These methods are usually not too sophisticated or difficult. Nevertheless, not everyone shares the view that they are needed and can contribute to a significant improvement in the efficiency of production processes, reduce costs, minimise quality errors or maximise profit. Ten years ago I had the opportunity to apply a set of simple linear regression models, which turned out to be extremely necessary in the process of planning the production of a large bakery. What is interesting is that I received the commission to create forecasting models from the owner of the bakery, a person who knew little about the existence of linear regression models. They were not interested in statistics or econometrics, but in the process of production and sale of bakery products.
Let us imagine that we want to build a model forecasting the number of cars crossing a bridge in some large city. It is known that the intensity of traffic depends on the hour. We will record different traffic at rush hour than at night or at dawn.
W-MOSZCZYNSKI ppic 12-21The next factor is the day of the week. Sunday traffic cannot be compared with Thursday traffic. The next factor is the weather. When it rains, the city’s inhabitants will not be inclined to commute to work by bicycle or on foot.
We therefore have three variables for our model:
x₁ — hour of measurement
x₂ — day of the week of measurement
x₃ — quantity of precipitation at the moment the measurement was taken
On the basis of these three variables, our model is to determine the probable level of traffic intensity on the bridge. This time I will not go into the technical aspects of building a linear regression model. I will only say that the model builds itself. Our task is only to enter into the first column of the spreadsheet the full hours (variable x₁). In the second column one should enter what day of the week it was then, and in the next, what the level of precipitation was in mm at that time. The most important is the last column, in which one should record the number of cars that crossed the bridge at a defined hour and day of the week. This last column contains the dependent variables, also called the outcome values.
The spreadsheet will do the rest of the work for us. Most spreadsheets, including Excel, are equipped with a tool that independently creates a linear regression model. A tool having at its disposal several hundred thousand hours across the days of the week in which traffic intensity on the bridge was recorded. On the basis of this data, a special mathematical mechanism is created, which, trained on history, is able to predict what the level of traffic on the bridge will be in the future, at a defined hour, on a defined day of the week, and on the basis of the forecast level of precipitation.
This is precisely how, in practice, the oldest, commonly used linear regression model is created.
A case study
This is a case study which really took place and which constituted a serious business problem. A certain bakery had a chain of a dozen or so bread shops. They were located in the immediate vicinity of railway stations in a number of suburban localities whose inhabitants commuted to work every day by public transport.
The bread shops worked from 6 a.m. until the evening. In the morning, customers bought bread on their way to work, and in the afternoon, when returning from work. Baked goods were delivered to the shops twice a day. The first delivery in the morning, and the second before the afternoon rush hour.
Both underproduction and overproduction of bread for the second, afternoon delivery caused the bakery to incur a daily loss on its activity. When there was too little bread, customers left the shop disappointed and dissatisfied, and the bakery did not earn as much as it could have earned. It is worth emphasising that evening purchases of bread were also key for the inhabitants of the locality. Returning from work, they bought fresh bread for the next day. When they did not find it in the bakery’s shops, they looked for it in other shops.
What decided the demand for bread?
The bakery’s employees noticed that demand for bread in the afternoon hours depended on the day of the week. Increased demand occurred on Mondays, Wednesdays and Saturdays. They also noticed that the level of demand was indirectly influenced by the weather, the season and the month in which the sale was conducted. The employees responsible for production planning took these variables into account when planning production. However, they were not always able to correctly interpret so many factors at once, which created configurations of mutual dependencies. The process of planning by „fortune-telling” was very labour-intensive and frustrating. The bakery had a dozen or so shops, and a forecast had to be made for several dozen products in each of these shops daily, and it had to be done relatively quickly.
In an act of desperation, the owner came up with the idea of producing for the second delivery exactly as much bread as had been sold precisely on that day of the week a year ago. This idea drove the production manager to the brink of a nervous breakdown. Every morning he had to go into the history of the sales system and find the correct day from a year before. Then compare the sales levels for each of the thirty-odd products in each of the dozen or so shops. IT specialists came to the rescue in time, combining all the data from several years into one spreadsheet. From that time on, the production manager could arrange a production plan even several weeks in advance. Unfortunately this was not an effective method. Sales from many years ago diverged worryingly from sales on specific days. It was therefore decided to base the plan on multi-year averages for individual days.
Reinventing the wheel?
In order to precisely create a linear regression model, a spreadsheet is enough, in which the describing variables and the outcome variable will be located.
The bakery’s employees, in the course of their attempts to predict the future, had organised the history of sales volumes. Every row denoted a single transaction made at a defined time, on a defined day of the week, and concluded with a defined quantity of bread sold. As we remember, the bakery’s employees indicated that the weather had an indirect influence on sales. Now it was enough to add historical weather information to the data in order to obtain a complete data set for creating a linear regression model. The database was therefore supplemented with weather data, such as temperature, pressure, and the size and type of precipitation.
Data from morning sales was removed from the database, which was then divided into 30 separate sets corresponding to the 30 products mentioned. As a result, 30 separate models were created. Linear regression models have the form of a simple equation, which can easily be written as a function in a spreadsheet.
Each of the 30 products sold on the second shift had its own forecasting model based on thousands of transactions from previous years. These models were written as spreadsheet functions. It was enough to enter the day of the week, the month and a few pieces of information from the weather forecast in order to find out what future sales would be at the chosen location.
This simple spreadsheet of linear regression models forecast sales with an accuracy of 65–82%, which was enough to obtain a profit.
The situation described really took place, which shows that scientific methods are not detached from reality. Having at one’s disposal similar sales databases and modern technology consisting of contemporary forecasting models and neural networks, one can detect unusual interdependencies, coincidences and trends. On the basis of the size and diversity of individual purchases, one can group customers, identify their preferences or sensitivity to promotions or price increases. Sales data constitutes a treasury of knowledge which one day in the future may be effectively used.
Wojciech Moszczyński
(in collaboration with Ewa Moszczyńska)
Wojciech Moszczyński — graduate of the Department of Econometrics and Statistics of Nicolaus Copernicus University in Toruń; specialist in econometrics, finance, data science, and management accounting. He specializes in the optimization of production and logistics processes. He conducts research in the area of the development and application of artificial intelligence. For years he has been engaged in the popularization of machine learning and data science in business environments.

Dodaj komentarz