May 2026 | Przegląd Piekarski i Cukierniczy (Baking and Confectionery Review)
The baker has always faced the same question: how much and what to bake. It sounds simple, but in practice it is one of the most difficult problems in the entire industry. Bread, rolls, kaiser rolls, croissants, challahs, sliced bread, dark bread, light bread, with grain, without grain – each of these products sells differently. Differently on Monday, differently on Friday, differently before a holiday, differently after a holiday. Differently in the morning, differently in the afternoon. Differently in a small neighbourhood shop, differently in an outlet on a busy street. And the decision has to be made before the customer walks into the shop. Often it has to be made at night, when the work is only just beginning.
This is precisely where the heart of the problem lies. Bread cannot be made up without cost and without time. If a bakery produces too much, part of the goods will stay on the shelf. Then it has to be marked down, returned, disposed of or thrown away. Each of these paths means a loss. The loss does not concern the bread alone. The loss also concerns the flour, the electricity, the gas, people’s work, the oven’s working time, the work of the vehicles, the space on the shelf and the whole effort put into production. And if the bakery produces too little, the consequences are different, but not lighter at all. The customer will come for bread and will not get it. The next day they may go somewhere else. After several such situations, the opinion about the shop and about the brand begins to deteriorate. Then the loss no longer lies in a single unsold item or in a single shortage. Then the loss concerns trust.
In a simple picture it looks like this. When a baker is to make a hundred loaves and sells only seventy, thirty items become a burden. When they make seventy and the customers want a hundred, thirty customers may leave empty-handed. On one side there is surplus. On the other side, shortage. Both hurt. Both cost. That is why the age-old problem of bakers is not only about how well to bake. It is also about how well to predict the market’s demand a day in advance.
W-MOSZCZYNSKI ppic 5-26Experience then and now
For many years this problem was solved mainly by experience. A good baker and a good shop owner „felt” roughly how much had to be made. They knew their district, they knew the people, they knew the rhythm of the week. Such knowledge still has great value. It cannot be disregarded. However, with a larger number of points of sale and with a wide assortment, intuition alone is increasingly not enough. When a company has one shop, many things can still be assessed by sight and memory. When there are several or a dozen or so shops, chaos begins. In one outlet more rolls sell in the morning. In the second, more bread in the afternoon. In the third, heavy traffic is on Friday. In the fourth, before closing, only the cheaper products still sell. A person may know some of these phenomena, but is not able to accurately calculate everything in their head every day.
Today yet another important problem has appeared. It is no longer enough to know how much to send to a given shop for the whole day. Increasingly, one has to know how much to give for the first shift and how much for the second. In other words: it is no longer only about the question „how much bread is to reach the shop”, but also about the question „at what time is it to get there”. This changes the scale of difficulty. Because if too much is sent in the morning, part of the goods may lie there too long. If too little is sent in the morning, the shelves will be empty before noon. And if the second delivery is badly calculated, the problem will return in the afternoon. As a result, the bakery must predict not only the size of demand, but also its distribution over time.
This is not exclusively a logistics problem. It is above all a production problem. Production is not planned at one o’clock in the afternoon, when it is already clear that something is missing. Production has to be planned earlier. Often, already on the day before the sale, one has to know how many items of a given product will be needed in the morning in shop A, how many in the afternoon in shop B, how much light bread, how much dark, how many small rolls, how much family bread. This means that a good sales decision must be made before the bakery’s night work begins. First, demand has to be predicted. Then that forecast has to be translated into a baking plan. Then that plan has to be carried out. Only at the end does the customer see the finished goods on the shelf.
This can be described with a simple example. The shop is to be open from six in the morning until seven in the evening. In the morning people come for fresh bread for work and for school. In the afternoon other customers come, often for small purchases for the evening. If someone counts the whole day as one number, e.g. „a hundred and twenty items will be sold”, they still do not know whether eighty of them are to be ready in the morning and forty in the afternoon, or the other way round. And that makes a great difference. Because it determines how much has to be baked at night, how much to leave for later, how to load the vehicle and how to distribute the goods between the shops.
Is forecasting necessary?
At this point it is clearly visible that the lack of a good forecast is not a minor inconvenience. It is a real cost. When a company operates on intuition, it is easy to make wrong decisions which repeat themselves every day. One day with a surplus is a single loss. Ten days with a surplus is already a constant leak of money. One day with shortages means a few dissatisfied customers. Many days with shortages means the loss of part of the market. Added to this is the unnecessary work of people. Someone makes the dough, someone operates the oven, someone loads the vehicle, someone arranges the goods, and then it turns out that part of that work produced no result, because the goods found no buyer. Raw materials are wasted as well.
On the other hand, well-calculated production brings profit not because something extraordinary is happening. It brings profit because the chaos disappears. The goods go where they have the greatest chance of being sold. The shops do not stand there with an empty counter. There is no need to patch up shortages in a hurry.
That is why the question „how much bread to bake” is not a simple question. It is one of the main management questions in a bakery. It combines sales, production, transport, people’s work and the relationship with the customer. Whoever answers it badly pays every day, sometimes without seeing the full scale of the losses. Whoever answers it well gains not only a better financial result, but also greater order in the whole company. That is precisely why demand forecasting is neither a fashion nor an ornament. It is a tool for limiting losses and for steering the bakery better.
This article deals with precisely this problem. It is not about theory for theory’s sake. It is about a practical question: how to predict how much and what kind of bread will be needed in a given shop, at a given time, on a given day, so as to limit losses and not lose the customer. Because in a bakery a forecast error does not end in a table. A forecast error ends on the shelf, in the bin or in the customer’s opinion.
Should we forecast the mean, or perhaps the conditional mean?
The first reflex is often simple. Since a given shop usually sells 40 loaves of bread a day, one may assume that tomorrow it will also sell 40. This way of thinking is easy and convenient. It gives a quick result. The problem is that the life of a shop does not run along a ruler. Sales are not the same every day. Sometimes they are higher, sometimes lower. The mean alone smooths the world too strongly. It shows the middle, but it hides the differences which matter a great deal in a bakery.
That is why someone may go a step further and say: let us not take one mean for everything, but let us calculate a conditional mean. That is, a mean for defined conditions. For example, separately for Monday, separately for Friday. Separately for the morning, separately for the afternoon. Separately for dry days, separately for rainy ones. This is already a better approach, because it takes into account the fact that not every day is the same.
The conditions can be very simple. Day of the week. Time of day. Weather. The time before a holiday. The time after a holiday. School open or school holidays. All of this influences customer traffic. So when someone speaks of a conditional mean, in practice they are speaking of a mean calculated for a specific situation. Not for the whole year at once, but for a slice of reality.
The example is simple. On Thursdays a shop may sell on average 50 loaves. But when those Thursdays are split into two types, the picture may change. On Thursdays without rain, sales may amount to an average of 58 items. On Thursdays with rain, only 41. One ordinary mean gives the number 50, but this number fits well neither the dry day nor the wet one. It is in the middle. And the shop does not operate in the middle. The shop operates in specific weather. All right, but does it work similarly in winter and in summer? And what about spring?
The same is visible during the day. If someone calculates one mean for the whole twenty-four hours, they may conclude that a given outlet sells 120 rolls. But nothing follows from that yet for the work plan. It may be that 80 sell in the morning and 40 in the afternoon. And it may be the other way round. One daily mean does not say when the customer will come. And for a bakery this is a crucial matter, because bread has to be baked earlier and it has to be established on which shift which baking will take place. Perhaps it was better in the old days, when there was one shop and 3 assortments?
The conditional mean is therefore a step forward, but it still has its limits. The biggest problem is that the world does not consist of simple, separate drawers. Factors combine with one another. The day of the week works together with the weather. The weather works together with the time of day. The time of day works together with the type of shop. And this leads to non-linear relationships.
The wretched non-linear relationships
A non-linear relationship means that a change in one factor does not always produce the same change in the result. This is not an arrangement of the type: one step to the left, minus five items; one step to the right, plus five items. In real trade the influence is sometimes uneven, abrupt or dependent on other conditions.
The first simple example: rain. Light rain may hardly change sales in a neighbourhood shop. People will go out for bread anyway. But heavy rain in the afternoon may already clearly reduce traffic. That is, the influence of the weather does not grow evenly. Light precipitation and heavy precipitation do not act in the same way.
The second example: the time of day. In the morning, a shortage of a few loaves may mean a large loss, because traffic is then fast and the customer wants to buy immediately. In the afternoon, the same shortage of a few items may be of lesser importance, because traffic is weaker or customers are buying other goods. The same number of shortages therefore does not produce the same effect at every time of day. This too is non-linearity.
In practice this means one thing: even a good conditional mean is sometimes too poor. It works well when the world is simple and calm. However, when a shop lives by the rhythm of the day, the week, the weather and local customs, the mean alone begins to lose. It is too rigid. It sees too little. It cannot combine many influences at once well.
That is why at a certain point it is worth going further. Not to reject experience. Not to reject simple statistics, but to treat them as a starting point and not as a final goal. The mean and the conditional mean can give a first approximation. They can show the base. However, where accuracy, smaller losses and a better baking plan matter, models are needed which see more than a simple average day. That is precisely why modern sales forecasts do not end with the mean. They begin with it, but they go further.
Linear regression as the first step towards a better forecast
Calculating the conditional mean alone sounds reasonable only until someone tries to actually do it. If a bakery has 30 products and 15 shops, the situation quickly becomes difficult. One would have to calculate means separately for bread, separately for rolls, separately for croissants. On top of that, separately for each shop. Then separately again for Monday, Tuesday, Wednesday. Then for the morning and for the afternoon. Then again for days with rain and without rain. At a certain point this turns into a thicket of numbers from which it is hard to draw a simple decision. This is hard, back-breaking work. Exhausting not only for the bakery owner; it is a „way of the cross” even for an experienced analyst. And what is worse, the effect is still sometimes too weak for today’s needs.
That is why it is worth going a step further. Such a step is linear regression. Here the level of labour intensity drops significantly. There is no need to build a complicated system straight away. At a basic level this can be calculated even in Excel. And that is its great advantage. It gives a simpler, faster and more orderly way of thinking about sales.
Put most simply, linear regression is a formula which combines several important pieces of information and, on that basis, calculates the predicted sales. Into this formula one inserts variables, that is, such data as have an influence on demand. For example, temperature, precipitation, wind strength, day of the week, shop number or time of day. Then the formula calculates how many items of a given product should be sold.
This can be imagined very simply. The mean says: „usually 40 items are sold”. Linear regression says something more: „if tomorrow is Thursday, if it is cold, if it rains, and the sales concern the shop by the housing estate and the morning hours, then the predicted sales amount to such and such”. This is no longer looking at one average day. This is calculating a specific situation.
It is precisely here that the advantage of such an approach can be seen. We know many things in advance. We know what day of the week tomorrow will be. We can also check the weather forecast. We can know the expected temperature, precipitation and wind strength. We also know which shop the goods are to reach. We also know whether we are talking about sales before noon or in the afternoon. Since this information is available in advance, it can be used to predict sales the day before baking. Yes, it is precisely thanks to this that we will be able to plan the appropriate level of production one or two days in advance.
This is also important because shops are not the same. A shop opposite a petrol station works differently from a shop on a housing estate. In one outlet the traffic may be fast and in the morning. In another, higher sales may be spread throughout the day. In one place customers drop in on their way. In another they buy calmly, in the course of their everyday shopping. That is why the shop number, and more broadly the shop itself as a place of sale, is also an important variable. This is not a detail. It is one of the basic pieces of information.
It is similar with the time of day. Sales before noon and in the afternoon often look different. In the morning customers buy fresh bread for the start of the day. In the afternoon traffic may be smaller or different. Whoever calculates only daily sales still does not know how much stock should be in the shop in the morning and how much in the afternoon. That is why it is worth adding one more variable to the formula: the time of day. It is a very simple move, and it can change a great deal.
What is this regression?
It is simply a way of calculating, a formula, like a recipe for baking a poppy-seed cake: 0.5 kg of poppy seed, 0.02 kg of yeast and 10 kg of type 500 wheat flour. Regression works in exactly the same way. Wind influences sales by lowering them by 4%, sun raises sales by 12%, on Thursdays sales are higher by 20%. Regression is simply a recipe for calculating the volume of sales. In practice such a formula works in such a way that every variable contributes its part of the information. Temperature says something about customer behaviour. Precipitation says something about the inclination to leave home. Wind strength may also have an influence, especially in worse weather. The day of the week shows the rhythm of the shop’s life. The shop number shows the difference between locations. The time of day shows when traffic is greater and when it is smaller. Linear regression gathers this into one and gives a final number.
It is like a recipe. Several ingredients go into the bowl: the day of the week, the weather, the shop, the time of day. Mixed together, they give a forecast. There is no longer any need to calculate hundreds of means by hand for every possible combination. One formula does it in an orderly way.
This is precisely what makes linear regression a good starting point. It is simpler than the laborious calculation of conditional means. It is also more practical, because it allows many things to be taken into account at once. And at the same time it is not off-putting. It can be implemented. It can be calculated. It can be explained. The bakery owner does not have to know all the technical details straight away. It is enough that they understand a simple thing: if important information about tomorrow can be gathered in advance, then tomorrow’s sales can also be predicted better.
Linear regression will not solve every problem straight away. But it gives the first real step forward. Instead of guessing, one begins to calculate. Instead of building dozens of little tables by hand, one builds a single coherent formula. And that is precisely why, for many bakeries, this may be a sensible beginning of the road to better production, smaller losses and calmer planning.
And who is supposed to do it? This regression…
At this point many bakery owners have the same thought: this sounds sensible, but who is supposed to do it. Numbers, tables, formulas, calculations again. Not everyone has to like it. Not everyone has to be able to do it. And there is nothing wrong with that. The baker is to know the oven, the dough, the proving, the rhythm of production and sales. For calculations there are people who have been trained in it.
It is worth saying this openly. Building a simple linear regression is not a task for a professor of mathematics. It is one of the basic activities in the work of people who deal with financial analysis, reporting, controlling or forecasts. For a person with a degree in economics, finance or statistics, or even for a junior who has worked in Excel and has had contact with simple data analysis, it is an ordinary matter. Not calculating by hand on a sheet of paper, but normal work in a spreadsheet.
This is important news, because many entrepreneurs immediately assume that once the subject of a model comes up, an expensive artificial intelligence specialist has to be hired. No. Linear regression is the simplest forecasting model. That is precisely why it is a good start. It is simple enough to be implemented quickly, and useful enough to give a real improvement.
The data, too, is usually already there. And that is the second piece of good news. The data does not have to be collected from scratch. It is in the points of sale, that is, in the POS systems. Every sale leaves a trace. This data accumulates in the sales register. The accounting department knows very well what that is. In practice it is digital data, in the form of an Excel or txt spreadsheet ready for export from the sales system. It shows what was sold, where, when and in what quantity.
Added to this is weather data. It is available immediately from the internet. One can download the forecast of temperature, precipitation and wind strength. That is already enough for a simple model. The next thing is the day of the week. This information is given by the date itself. Excel can determine it with a simple function. Whether tomorrow is Thursday, Friday or Saturday is known in advance. There is no obstacle here at all.
It is also worth adding the shop as separate information. This is very important. An outlet by a petrol station works differently, an outlet on a housing estate differently, a shop by a school differently, and one by a busy street differently again. It is also good to split sales into the time before noon and the afternoon, because traffic varies. As a result, the model receives several simple pieces of data: the date, the day of the week, the weather, the shop and the time of day. That already gives a sensible basis for a forecast. Importantly, the more data the better. If we have data from 5 years, so much the better; for Excel it makes no difference whether the database consists of 200 lines or 200 thousand.
The regression analysis is done by a low-qualified specialist, once a year, and it takes them 2 days
Here one thing has to be emphasised. It is not about the bakery owner sitting over Excel themselves and learning statistics at night. It is about them knowing that this can be commissioned from someone. Such a model can be prepared by an office worker, an analyst, an economist, a person who handles reports or someone from outside on a short commission. It does not have to be a full-time post. The building of the model can be commissioned once, and then one only has to check from time to time whether it still works well.
And that is exactly how it is worth looking at it. Not as a great technological project, but as an ordinary service. Just as an oven, a car or a cash register is serviced, in the same way a forecasting model can be checked from time to time. The model itself does not have to be created anew every week. If it is made correctly and based on decent data, it can work stably for many months. After that it only requires review, correction and refreshing.
In short: data is usually not lacking, the tools are widespread, and the very skill of building a simple linear regression is widely known among economists and analysts. For the baking industry this is good news. Because it means that there is no need to be afraid of mathematics. One only has to know whom to commission the work from and what to require from that person.
This is basic knowledge. For an analyst as basic as, for a baker, the knowledge of how yeast works and when dough rises well and when badly. That is why the bakery owner should not ask: will I manage to calculate it myself. A better question is: who is to build it for me efficiently and how quickly can I start using it.
The next step is simple. First, data is taken from the sales register. Then the weather, the day of the week, the shop and the time of day are added. And then a person who knows Excel and the basics of forecasting builds the first model. Without great theory. Without fear. Simply in order to limit losses and plan the work better.
When Excel is no longer enough
At this point it is worth saying a simple thing. The bakery owner does not have to understand all this mechanics. They do not have to know the names of the models, they do not have to know the formulas, they do not have to be able to calculate it. That is not their role. Their role is to know that it can be done better than in Excel and that it can be commissioned. A well-made system can work in the background, download the day of the week by itself, download the weather forecast by itself, calculate the predicted sales by itself, and show the result on a private company page, an ordinary website accessible only to the owner and authorised persons. In the morning, or the day before, one can simply go there and see: how much sales are predicted in a given shop, for a given product, before noon and in the afternoon of the following day, in two days’ time.
Who is to do it?
This is no longer the level of Excel. This is the level of a person who can work in Python and build forecasting models. Such a specialist does not have to sit in the bakery every day. They can prepare the solution once, launch it, and then only check and refresh it from time to time. If the data is decent and the market around the shop does not change abruptly, such a model can work stably for many months. It usually does not break down from one day to the next. More often the problem appears when the business environment changes: a supermarket opens next door, a new bakery starts up, a competitor closes down or the traffic in the area changes.
The time to build such an application on a website is about one week. One can allow oneself 2 weeks. That is already the maximum time for building such a solution. The price? Let us assume that an expert will take 200 PLN per hour. That is a very high price, but there are not too many such specialists, everyone fights for them, so they value themselves highly. Let us count 80 hours at 200 PLN, which gives 16 thousand PLN + VAT. Now – how much will we save on this? Is it worth it? This system is to run without servicing for at least 12 months, in practice 24 months. There is nothing here to break down. Even if it starts forecasting more weakly, it will always forecast better than regression. On the scale of 24 months this comes out at 25 PLN a day. If the model has saved more than 24-30 PLN a day, it means that the investment has paid off. We are no longer counting the fact that the toil of calculating on our own falls away, along with the stress of responsibility for losses, whose hot breath we feel on our backs every day. We are moving with the times; after all, it is a competitive advantage on the digital battlefield.
Better than regression
It is also worth saying honestly that linear regression is a good beginning, but not the end of the road. It is a simple and useful model. However, there are methods better suited to this type of task. In a bakery one does not predict classes of the type „little”, „medium”, „much”, but a specific number of items. Models which count non-negative quantities fit such data well.
Now there will be a message for the potential contractor of the „digital fortune-teller”.
In scikit-learn there is, among others, PoissonRegressor for this, that is, the Poisson model, and a step higher stands HistGradientBoostingRegressor with the setting loss=”poisson”. The latter copes better with non-linearity, and the scikit-learn documentation also states that for larger sets it is much faster than classic gradient boosting. In an example comparison in scikit-learn, both PoissonRegressor and HistGradientBoostingRegressor gave a better agreement of forecasts with observations than the simpler linear model of the Ridge type (a variant of linear regression).
So if a junior wants to develop, it is precisely here that the next step begins. First they build a simple linear regression. Then they learn count models, which they will still find in Excel. And then they move on to tree-based models, which better capture the relationships hidden in the data. For the baker the most important conclusion is simple: it is possible to have a forecast better than the one from Excel, but for that a person who knows these tools is already needed.
Whom to give it to?
Here, however, an important matter appears. The bakery owner should not hand over the subject blindly. They do not have to understand everything, but they should know what to require. Firstly, the model is to predict the number of items. Secondly, it is to be assessed not by an abstract indicator, but by the error in items. They are not interested in the fact that the indicator looks nice in a report. They are interested in by how much on average the model was wrong in the number of loaves, rolls or croissants. That is precisely why, to start with, the MAE measure, that is, the mean absolute error, is much better. In scikit-learn, MAE is described as a measure of regression error whose best value is 0. For business this sounds clear: „we are wrong on average by 4 items”. R², in turn, may be useful for an analyst, but for the owner it is sometimes not very readable, and the documentation states outright that it can even be negative.
Thirdly, such a model has to be checked the way time works. One must not mix yesterday with tomorrow in the learning. The scikit-learn documentation recommends TimeSeriesSplit for data ordered in time, because ordinary splitting methods could lead to learning on the future and evaluating on the past. This is very important. In a bakery the model is to imitate real life: first it knows the old data, and only then is it to predict a new day. If the specialist has not understood: I am talking about the split into training and test variables in the supervised learning method.
For the specialist, therefore, the guidance is short. If they are to make a model better than Excel, they should start from a target in the form of the number of items sold, treat the data as a regression problem with a time order, not do random mixing of train and test, measure the error in items, and consider PoissonRegressor or – more often – HistGradientBoostingRegressor(loss=”poisson”) as sensible candidates. PoissonRegressor is a generalised linear model with a Poisson distribution, and HistGradientBoostingRegressor explicitly permits the poisson loss.
The digital baker
For the bakery owner something else is important. Such a model is not to be a black box closed forever. It should show every day not only the forecast, but also its own error. Most simply: the model gave a forecast for tomorrow, the bakery planned the baking according to it, and a day later the sales and the returns are already visible. Then one can check by how much the model was wrong. Such data on the size of returns can be entered into Excel by the production manager. If for many weeks it is wrong by little, it is working well. If it suddenly starts being clearly more wrong, that is a sign that it has to be corrected or refreshed. In this way the owner does not have to know mathematics. It is enough that they look every day at one simple thing: what the forecast was and what the error was. One can also, as an option, ask the contractor for the model to keep learning by itself every month on new data. This is a normal functionality of such systems. Then its durability and the high quality of its forecasts will be counted in years.
That is precisely what gives the greatest peace of mind. There is no need to take the model at its word. It can be controlled. It can be watched over. It can be required to report its own misses. Thanks to this, the model does not live beside the company, but works for the company. The model is to keep learning by itself. It is to be as trouble-free as an app in a phone or an oven for baking bread.
Conclusion, a happy ending
And this is a good ending to the whole subject. First there was intuition. Then the mean. Then the conditional mean. Then linear regression in Excel. And at the end a more mature solution appears: a model made by a specialist, working automatically, showing the forecast on a private page and regularly reporting its own error. Without great philosophy. Without fashionable slogans. A maintenance-free model which every month improves its abilities, because it keeps learning automatically. Simply as a tool for a better baking plan, smaller losses and calmer running of the bakery. Of course, first there will appear the stress that we have to change something, then the effort of finding someone who will do it. A little stress, a little embarrassment and nerves. We are moving forward, it is worth it, it is really worth it.
Wojciech Moszczyński
Wojciech Moszczyński — graduate of the Department of Econometrics and Statistics of Nicolaus Copernicus University in Toruń; specialist in econometrics, finance, data science, and management accounting. He specializes in the optimization of production and logistics processes. He conducts research in the area of the development and application of artificial intelligence. For years he has been engaged in the popularization of machine learning and data science in business environments.

Dodaj komentarz