Data Science in Failure Diagnostics

Data Science in the bakery

This is the fourth article to describe the possibilities of using the latest trends in management and control in bakery production.

The enormous amount of data and the increasing complexity of processes are driving the use of increasingly advanced diagnostic and analytical technologies, and today we're going to talk about statistical process controls to prevent major failures.

W-MOSZCZYNSKI-2019-4-04

Accidents in the bakery

The bakery has always been, and will always be, a community-based facility, strongly connected to the people around it. Bread is a symbol of the home, family, the local environment. The bakery owner has a huge responsibility for the continuity and quality of production. It's hard to even imagine that all of a sudden there will be a shortage of bread in all the local shops.

The baking process has become very dependent on the reliability of the machines. Today, the manufacturing processes have become very complex and technologically conditioned. What would be the losses if the cold, freezing temperatures suddenly thaw? What will happen to the bread if the oven loses half its capacity?

Sometimes we find a breakdown when the bread we sell is uncooked or sticky, and any drastic decline in the quality of the product means a loss of customer confidence.

Any failure of the slime pumps or a lack of raw materials in the silos caused by bad readings is a production halt. A serious failure means a few days' stoppage, and consequently a reduction or even a halt in supply.

The bakery could lose revenue and, worse, the trust of the local environment.

History of diagnosis based on normal distribution

The basic method for preventing failures worldwide, across all industries and across all economies, is the so-called Statistical Process Control (SPC), which is based on the theory of the Normal Density of Probability of Events.

For a very long time, this theory has not been used in research diagnostics.

It was first discovered in 1733 by the French mathematician Abraham de Moivre, and its discovery was forgotten for many years before being reborn in two independent discoveries: the Pierre de Laplace and Carl Gauss.

In 1915, the English biologist and evolutionary statistician Sir Ronald Fisher introduced the methodology of normal-order research tests, and the real father of statistical process control was his doctoral student Walter A. Shewhart.

It was a kind of diagnostic revolution, where they were both exposed to hateful attacks from conservative scientific authorities for the rest of their lives.

The normal distribution

As I mentioned earlier, the basis for predicting failure is the normal distribution, which resembles the shape of the bell that has the most commonly encountered value in a population, and the area under the curve is defined as the area of density where the probability of events occurring on the curve.

To shed some light on this enigmatic description, I'll use a simple example.

If we select a group of 100 men, we'll find that about 70 of them have a height range of 171 to 185 cm. The probability of hitting someone who is 205 cm or 146 cm is close to zero. The greatest probability of being selected is when the height is close to the arithmetic mean of the selected group.

So this is how you can picture the mechanics of normal decomposition. Natural phenomena with a stable nature tend to concentrate around the mean.

Diagnostic systems in a bakery

Today, virtually every more complex device has its own diagnostic system, the most common of which consists of counting the cycles or hours of work to determine the date of the next technical breakdown.

It used to be a good practice for brigadiers to record all the quantitative data in their shift logs, recording process temperatures, production volumes, raw material consumption, and any observations like strange vibrations or sounds, and this kind of diagnostic method worked particularly well to explain the causes of the failure.

Unfortunately, it was difficult to predict future production problems this way, because the values recorded were not processed, not modeled and processed, and integrated production management systems have been common for almost 20 years.

Such systems combine readings from multiple machines and devices, process them and form process maps and are usually present where the production lines are a composite system supplied in full by one manufacturer.

Unfortunately, bakeries and confectioneries are most often equipped with devices from different manufacturers, so they do not have integrated production control systems. Bakeries have their own measurement system, freezers and bakeries. In this case, the only system that integrates the whole is the raw material dosing system, based on a prescription protocol.

To integrate the entire production for the purposes of the Data Science diagnostic system, measuring devices should be installed, which should communicate with the server that collects all the data in the form of time series.

The task of Data Science is to create a self-learning research system that will be able to detect anomalies and changes in statistical characteristics in advance, predicting major failures or minor defects.

The scope of such a data set is almost unlimited, depending on the level of knowledge of the data science professionals employed in the project, the diagnostic system is fully autonomous and requires virtually no maintenance, and the only requirement is a systematic review of reports and warnings.

In recent years, there has been a very intense development in the study of process processes. Today, it is not only the behaviour of individual machines that is studied, but above all the interaction of many related processes that are analyzed.

His method of experimental design is currently the most important method of predicting future processes.

Statistical control of processes

Why is the normal distribution so important in the statistical control process? Because it analyzes deviations from the mean. Measurement systems are constantly gathering information about temperature, vibration and humidity. They put it together into endless time series, data sets from the Data Science environment.

This data generates process statistics, where deviations play the most important role, wherever the measurement level exceeds the so-called upper or lower boundary of variability, there is talk of a special cause of variability, and each such phenomenon is interpreted as an anomaly warning of a future, serious failure.

The normal distribution describes a perennial principle of nature that all autonomous processes are balanced. Interestingly, this 18th-century principle did not become widespread in diagnostic analysis until the mid-1980s.

Mankind has come up with the seemingly simplest solution for nearly 250 years.

Wojciech Moszczyński — graduate of the Department of Econometrics and Statistics of Nicolaus Copernicus University in Toruń; specialist in econometrics, finance, data science, and management accounting. He specializes in the optimization of production and logistics processes. He conducts research in the area of the development and application of artificial intelligence. For years he has been engaged in the popularization of machine learning and data science in business environments.

Bądź pierwszy, który skomentuje ten wpis!

Dodaj komentarz

Twój adres email nie zostanie opublikowany.


*