February 2020 | Przegląd Zbożowo-Młynarski (Grain and Milling Review)
Data science is a new field of study, which came into being not quite a few years ago. At present it is regarded as the basic and leading discipline in the area of building forecasting models.
In the following articles of this series I will discuss the individual methods of forecasting and classification applied in determining future grain prices and predicting trends on raw material markets.
What is data science?
Data science is a field of study combining many disciplines, such as mathematics, statistics, econometrics and computer science. The most important stimulus for its emergence was the dynamic development of data processing systems and the enormous growth in the quantity of available data which accompanied it. Thanks to the development of computer science, theoretical methods of forecasting and modelling worked out in the past have „come alive” and have become possible to use in practice.
W-MOSZCZYNSKI pzm 2-20In the past, without computers, applying complex models of neural networks or decision trees involved long-lasting, laborious calculations. The labour-intensiveness and time of the calculations, together with the absence of easy-to-use data, constituted an effective barrier to the spread of innovative analytical methods in the area of business.
The term data science has appeared in various contexts only in recent years. The breakthrough moment in the emergence of this field was the article by the American data researcher Dhanurjay Patil, published in 2012 in the journal „Harvard Business Review”. In the article „Data Scientist: The Sexiest Job of the 21st Century” he stated that the data specialist is a „new breed”, and that „the shortage of data scientists is becoming a serious constraint in some sectors”. He went on to describe the scope of the new science and its significance in the economy of the future. After 2017 the concept of data science became commonly known throughout the world. It has therefore existed in public space for barely three years.
Why did data science appear?
Data science is very strongly connected with the technology of collecting and processing data. Its emergence is connected not only with the dynamic development of computer science, but above all with the dramatic growth in the quantity of data. In the past, economic processes generated only modest, basic data. Together with the growth in the complexity of industrial processes and the transformation of the economy there came an avalanche-like growth in available information. Every machine, every vehicle and every sales system today produces enormous quantities of data in the form of textual time series. Part of the data is of fundamental importance in modelling, in optimising processes and thereby in achieving competitive advantage. A business need therefore appeared for the effective use of data.
In the new situation, the application of traditional analytical tools turned out to be impractical. Spreadsheets, commonly used by analysts, turned out to be powerless in the face of enormous quantities of data and the complexity of processes. In turn, the tools used hitherto in the area of statistical analyses turned out to be inflexible and incapable of quickly adapting new analytical solutions.
Programming environments in data science
Data science analyses can be conducted in spreadsheets. Unfortunately, commonly used spreadsheets of the Excel type are able to process textual databases effectively up to a size of 0.2 GB. Unfortunately, the majority of forecasting models operate on considerably larger textual databases, in the range from 0.1 to as much as 30 GB.
The large range of data therefore forced a move away from the traditional spreadsheet towards a programming environment.
The basic format of data science databases is text files, including the most popular standards: csv and txt. The reason is very simple: the majority of devices and systems in the surrounding economy generate data in the form of flat textual tables. Data science is a research science; it does not specialise in database systems and rarely uses relational databases.
The main programming languages in data science are R and Python.
The R language was developed by John Chambers at Bell Laboratories in 2000. It is for the most part written in the C and Fortran languages. It is commonly held that because of Fortran it is slower than the competing Python, which was written entirely in the C language. Python was written in the early 1990s by the American scholar Guido van Rossum. The name of the language comes from the comedy series broadcast in the 1970s by the BBC, „Monty Python’s Flying Circus”.
The R language is used more often in academic centres, whereas Python can be encountered more often in analytical centres. At present a tendency to move away from the R language in favour of the Python language is noticeable.
Programming languages are used in special editors, among which the most popular are the editors Jupyter Notebook and PyCharm.
It is commonly held that the best operating system for working with the R and Python languages is Ubuntu. Of course these languages can work on other operating systems such as MS Windows or Mac OS.
When creating advanced mathematical models which heavily burden computing resources, the Apache virtual environments are used, among others: Hadoop, Spark, Kafka, Cassandra, Hive and others. These systems are most often applied in large research centres and analytical centres.
Both the R and Python languages, like their editors Jupyter and PyCharm and the Ubuntu operating system, are available free of charge on the Internet.
Libraries in data science
Although data science is very strongly embedded in programming languages, analytical work consists not in programming in these languages, but in using specialist libraries.
Libraries are a kind of sub-language specialised for particular tasks. Every library has its own set of commands and a specific syntax. Most often the syntax of the libraries is similar to the syntax of their base languages, Python or R. It is commonly held that Python has considerably richer libraries than the R language.
The most important libraries applied in Python in the area of neural networks are: TensorFlow and, recently gaining in popularity, PyTorch. For graphical analyses the libraries Seaborn, Matplotlib and Plotly are used. Forecasting models are handled by the libraries Sklearn, Statmodels and Numpy. Undoubtedly a more important library is also Pandas. The name comes from the abbreviation „panel data”. This library serves to fetch and prepare data for use in other libraries.
Libraries, like other data science working tools, are free of charge.
What changes does data science bring?
The new field of study which is data science introduces a series of revolutionary changes which have for ever changed the face of conducting research and analyses. From the point of view of analysts there are three most important changes:
1. The growth in the importance of programming languages and their libraries
This is a very large change, because until now work of this type was carried out on expensive forecasting programs of the type: Minitab, Statistica or SAS.
In contrast to the analytical systems mentioned, libraries undergo frequent modifications in which the latest achievements of science are implemented. Libraries have the ability to cooperate with one another, as a result of which they are more flexible in operation, in contrast to the programs used hitherto.
2. The growth in the availability of knowledge, together with its currency and its rapid transmission
At present knowledge, in the form of practical advice, tutorials and ready-made scripts, is available practically instantly on the Internet. Every week new mathematical and programming solutions appear which often fundamentally change the way research is conducted and models are created. The phenomenon has appeared of the premature obsolescence of books, even before they are printed.
3. Work in the analytical cloud
The growth in the intensity of data processing processes has led to the emergence of a new kind of service consisting in remote work with the use of leased computing power of external computers. This is a completely new phenomenon, possessing enormous development potential.
The 21st century is called the age of information. Data science is an area of activity which fits ideally into the present times.
Wojciech Moszczyński
Wojciech Moszczyński — graduate of the Department of Econometrics and Statistics of Nicolaus Copernicus University in Toruń; specialist in econometrics, finance, data science, and management accounting. He specializes in the optimization of production and logistics processes. He conducts research in the area of the development and application of artificial intelligence. For years he has been engaged in the popularization of machine learning and data science in business environments.

Dodaj komentarz