Researchers at HSE in St Petersburg Develop Superior Machine Learning Model for Determining Text Topics

They also revealed poor performance of neural networks on such tasks
Topic models are machine learning algorithms designed to analyse large text collections based on their topics. Scientists at HSE Campus in St Petersburg compared five topic models to determine which ones performed better. Two models, including GLDAW developed by the Laboratory for Social and Cognitive Informatics at HSE Campus in St Petersburg, made the lowest number of errors. The paper has been published in PeerJ Computer Science.
Determining the topic of a publication is usually not difficult for the human brain. For example, any editor can easily tag this article with science, artificial intelligence, and machine learning. However, the process of sorting information can be time-consuming for a person, which becomes critical when dealing with a large volume of data. A modern computer can perform this task much faster, but it requires solving a challenging problem: identifying the meaning of documents based on their content and categorising them accordingly.
This is achieved through topic modelling, a branch of machine learning that aims to categorise texts by topic. Topic modelling is used to facilitate information retrieval, analyse mass media, identify community topics in social networks, detect trends in scientific publications, and address various other tasks. For example, analysing financial news can accurately predict trading volumes on the stock exchange, which are significantly influenced by politicians' statements and economic events.
Here's how working with topic models typically unfolds: the algorithm takes a collection of text documents as input. At the output, each document is assessed for its degree of belonging to specific topics. These assessments are based on the frequency of word usage and the relationships between words and sentences. Thus, words such as ‘scientists,’ ‘laboratory,’ ‘analysis,’ ‘investigated,’ and ‘algorithms’ found in this text categorise it under the topic of ‘science.’
However, many words can appear in texts covering various topics. For example, the word ‘work’ is often used in texts about industrial production or the labour market. However, when used in the phrase ‘scientific work,’ it categorises the text as pertaining to ‘science.’ Such relationships, expressed mathematically through probability matrices, form the core of these algorithms.
Topic models can be enhanced by creating embeddings—fixed-length vectors that describe a specific entity based on various parameters. These embeddings serve as additional information acquired through training the model on millions of texts.
Any phrase or text, such as this news item, can be represented as a sequence of numbers—a vector or a vector space. In machine learning, these numerical representations are referred to as embeddings. The idea is that measuring spaces and detecting similarities becomes easier, allowing comparisons between two or more texts. If the similarities between the embeddings describing the texts are significant, then they likely belong to the same category or cluster—a specific topic.
Scientists at the HSE Laboratory for Social and Cognitive Informatics in St Petersburg examined five topic models—ETM, GLDAW, GSM, WTM-GMM and W-LDA, which are based on different mathematical principles:
- ETM is a model proposed by the prominent mathematician David M. Blei, who is one of the founders of the field of topic modelling in machine learning. His model is based on latent Dirichlet allocation and employs variational inference to calculate probability distributions, combined with embeddings.
- Two models—GSM and WTM-GMM—are neural topic models.
- W-LDA is based on Gibbs sampling and incorporates embeddings, but also uses latent Dirichlet allocation, similar to the Blei model.
- GLDAW relies on a broader collection of embeddings to determine the association of words with topics.
For any topic model to perform effectively, it is crucial to determine the optimal number of categories or clusters into which the information should be divided. This is an additional challenge when tuning algorithms.
Sergey Koltsov, primary author of the paper, Leading Research Fellow, Laboratory of Social and Cognitive Informatics
Typically, a person does not know in advance how many topics are present in the information flow, so the task of determining the number of topics must be delegated to the machine. To accomplish this, we proposed measuring a certain amount of information as the inverse of chaos. If there is a lot of chaos, then there is little information, and vice versa. This allows for estimating the number of clusters, or in our case, topics associated with the dataset. We applied these principles in the GLDAW model.
The researchers investigated the models for stability (number of errors), coherence (establishing connections), and Renyi entropy (measuring the degree of chaos). The algorithms' performance was tested on three datasets: materials from a Russian-language news resource Lenta.ru and two English-language datasets - 20 Newsgroups and WoS. This choice was made because all texts in these sources were initially assigned tags, allowing for evaluation of the algorithms' performance in identifying the topics.
The experiment showed that ETM outperformed other models in terms of coherence on the Lenta.ru and 20 Newsgroups datasets, while GLDAW ranked first for the WoS dataset. Additionally, GLDAW exhibited the highest stability among the tested models, effectively determined the optimal number of topics, and performed well on shorter texts typical of social networks.
Sergey Koltsov, primary author of the paper, Leading Research Fellow, Laboratory of Social and Cognitive Informatics
We improved the GLDAW algorithm by incorporating a large collection of external embeddings derived from millions of documents. This enhancement enabled more accurate determination of semantic coherence between words and, consequently, more precise grouping of texts.
GSM, WTM-GMM and W-LDA demonstrated lower performance than ETM and GLDAW across all three measures. This finding surprised the researchers, as neural network models are generally considered superior to other types of models in many aspects of machine learning. The scientists have yet to determine the reasons for their poor performance in topic modelling.
See also:
Scientists Create Open Dataset for Studying Concentration
A team of Russian researchers, including scientists from HSE University–St Petersburg, has developed the first open multimodal dataset containing recordings of brain activity, heart function, and video observations to help researchers understand what happens in the human brain during deep concentration. In the future, the dataset could accelerate the development of neural interfaces, rehabilitation technologies, and AI systems. The article has been published in Scientific Data.
Scientists Propose Method for More Efficient Resource Use in Machine Learning
An international group of researchers, including mathematicians from the AI and Digital Science Institute at the HSE Faculty of Computer Science, has provided a theoretical justification for a simple and computationally efficient method of estimating uncertainty in Stochastic Gradient Descent (SGD). The paper has been published on the scientific preprint server arXiv.org and presented at AISTATS 2026.
Team Success: Aligning Means with Objectives
In corporations, sports, and academia, people often face challenges they cannot handle alone. In such cases, selecting the right team is crucial. Tatiana Mayskaya, Associate Professor at the HSE Faculty of Economic Sciences and the International College of Economics and Finance, together with colleagues from foreign universities, examined team characteristics and found that less diverse teams are better suited to objectives where a high average performance is important, whereas more diverse teams are preferable when avoiding failure is critical. The paper has been published in Economic Theory.
Economists Propose More Effective Approach to Reducing Smoking
Economists at HSE University have examined how smokers respond to changes in cigarette prices. When tobacco prices increase, cigarette consumption does not always decline. In fact, spending on tobacco may even rise: according to the researchers, a 1% decrease in cigarette affordability leads to a 0.28% increase in per capita tobacco expenditure. The findings suggest that to reduce smoking rates, tobacco prices must rise faster than household incomes. The study has been published in Voprosy Statistiki.
Biologists Discover Unique Properties of MiR-93-5p MicroRNA in Prostate Cancer
Researchers at the International Laboratory of Microphysiological Systems of the HSE Faculty of Biology and Biotechnology investigated how different isoforms of the same microRNA influence gene function in prostate adenocarcinoma. The study found that in some cases, microRNAs can reinforce each other’s effects by targeting and suppressing the same genes. This finding offers a fresh perspective on the molecular mechanisms underlying tumour development and on the search for disease biomarkers. The results have been published in PeerJ.
HSE Researchers Provide the World’s First Legal Definition of a Digital Ecosystem
Digital ecosystems have evolved from a technological innovation into a fundamental institution of the modern economy over the past few years. According to HSE University’s latest estimates, they account for 8.5% of Russia’s GDP. Previously, no jurisdiction had a statutory definition of what constitutes a digital ecosystem. HSE University researchers have addressed this gap by proposing the first legal concept of a digital ecosystem. Their article, ‘The Digital Ecosystem as a Novel Economic Phenomenon and Legal Concept,’ has been published in the BRICS Law Journal.
HSE Economists Use Search Queries to Forecast Birth Rates
Researchers from the HSE Faculty of Economic Sciences have shown that the accuracy of birth rate forecasts for Russia can be improved by almost 50% by incorporating the dynamics of online search queries related to pregnancy and childbirth into forecasting models. In the best-performing models, the forecasting error fell from 4.6% to 3.2%. The findings have been published in Populations and Economics.
When Looking at Their Own Faces, Men Forget Everything
In an experiment involving 15 healthy men, scientists at HSE University investigated how different phases of the cardiac cycle influence the excitability of the motor cortex when participants viewed either their own photograph or the faces of strangers. The researchers found that when participants looked at their own image, the brain’s response to signals from the heart was weaker, meaning that the influence of cardiac activity on the motor cortex decreased. This finding came contrary to expectations, as self-focused attention was thought to enhance the brain's sensitivity to internal bodily signals. The study has been published in Frontiers in Signal Processing.
HSE Researchers Discover Who Eats Out in Russia—And Why
Around one-third of Russians (31.3%) rarely eat out or buy ready-made meals. The core group of active consumers—those who eat out or purchase prepared food almost every day or several times a week—accounts for only about 9% of the population. These are the findings of a study conducted by the HSE Institute for Social Policy. According to the researchers eating out is no longer a marker of high social status in Russia.
Scientists Model How Interactions Between Societies Can Trigger Chaotic Behaviour
Scientists at HSE MIEM have proposed a mathematical model explaining how interactions between societies can influence their stability. Based on the classical theory of evolutionary games, the study reveals an unexpected effect: even a weak informational influence of one society on another can cause one society to remain stable while the other exhibits chaotic behaviour among its individual members. The study has been published in the International Journal of Bifurcation and Chaos.


