bg
Medicine and healthcare
07:26, 29 August 2026
views
7

Neural Network for Peptides: Russian Scientists Speed Search for New Drugs

In Russia, scientists have developed a compact neural network that helps assess the properties of peptides – short chains of amino acids. The model requires far fewer computing resources than large protein neural networks and could accelerate the search for promising molecules for new drugs.

The search for a new drug begins long before a candidate reaches the laboratory. Researchers must sift through vast numbers of molecules to determine which ones work – suppressing bacteria, reducing inflammation, affecting metabolism or exhibiting other useful properties. Today, some of this work can be moved from the laboratory to a computer. Scientists at the ITMO University Artificial Intelligence in Chemistry Center are developing just such an approach. They have created a neural-network model called dcBiLSTM-AE that analyzes peptides and predicts their biological and physicochemical properties.

The new model contains about 90,000 parameters, thousands of times fewer than large protein language models. As a result, it does not require enormous computing resources or massive datasets to operate.

Why Peptides Matter

Peptides are short chains of amino acids. They are involved in a wide range of biological processes, and some are already used in medicine or being explored as the basis for new drugs. Some peptides, for example, have antimicrobial or anti-inflammatory activity. Others could be of interest for treating diseases associated with metabolic disorders. Still others are used as part of targeted drug approaches.

But a molecule's useful properties cannot be determined from its name alone. What matters is which amino acids it contains and the order in which they are arranged. Scientists therefore have to test many variants. The traditional approach is to synthesize a molecule and send it for laboratory testing, which takes time, materials and researchers' labor. If a computer model can eliminate some candidates in advance, the laboratory receives a shorter list of the most promising options. That is precisely what the ITMO University system is designed to do.

A New Solution to Old Problems

The Russian researchers developed dcBiLSTM-AE, a compact autoencoder whose internal peptide representations are inherently linked to physicochemical characteristics. The model does not require additional training to explain its results. It can not only generate a prediction but also reveal which specific peptide property influenced that prediction.

The model was tested on eight tasks, including antimicrobial, anti-inflammatory, antidiabetic, antioxidant and hemolytic activity, resistance to biofouling, and water solubility. The eighth task involved predicting the minimum peptide concentration required to inhibit the growth of Escherichia coli. Across all eight tasks, the new model delivered results comparable to those of large language models and ranked second among six comparable models.

Room for Further Development

The researchers now plan to extend the neural network to longer chains and rare amino acids outside the standard set of 20 naturally occurring amino acids, as well as adapt the model to peptides with unusual structures – for example, those closed into rings or containing side branches. Such structures can enhance a molecule's bioactivity or help it remain in the bloodstream longer, extending a drug's effectiveness.

Prospects for Medical Research

For pharmaceutical companies, the technology could become part of the early stages of drug development. Computer models could be used for initial screening, with the most promising compounds then moving on to experimental testing. If the algorithm can eliminate unsuccessful candidates more quickly, researchers can focus on molecules that genuinely merit further study. This could potentially reduce the number of unnecessary experiments and conserve resources.

Another important factor is access to computing resources. Working with enormous models requires substantial computing power. A compact algorithm can run at much lower computational cost, potentially making such tools accessible to a broader range of researchers.

For Russian science, this is a particularly promising direction because, rather than competing to build ever-larger neural networks, researchers can develop compact, specialized models that solve specific problems while remaining interpretable and accessible to scientists. The ultimate goal of this work is both straightforward and concrete – to find molecules that could become new drugs more quickly.

Our model contains about 90,000 parameters, whereas popular protein language models operate with hundreds of millions or billions. To pretrain dcBiLSTM-AE, we used about 155,000 peptides, significantly fewer than are used to train large protein language models. The compact architecture and the use of the physicochemical characteristics of amino acids make it possible to obtain informative peptide representations with a substantially smaller model and training dataset
quote

like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next