bg
Science and new technologies
07:51, 29 September 2026
views
1

Neural Network Taught to Doubt Itself: MEPhI Researchers Are Rethinking the Rules for Trusting AI

Researchers at the National Research Nuclear University MEPhI (MEPhI) have developed a new artificial intelligence architecture – a stochastic transformer that can not only generate a prediction but also assess how confident it is in its own decision.

Most modern neural networks have one notable weakness: they deliver answers with the same apparent confidence whether they are certain that the information is reliable or simply making an educated guess. Researchers at MEPhI have proposed an architecture that changes this logic. The stochastic transformer not only generates a prediction but also assesses how much its own conclusion can be trusted. If the model is uncertain, it signals that to a human. At the same time, the technology protects AI from so-called adversarial attacks, in which an attacker makes subtle changes to input data to cause the algorithm to produce a wrong result.

Right to Be Wrong

A conventional neural network operates with fixed parameters. A stochastic transformer works differently: its calculations include random values. The model runs a task through multiple variants and checks whether they produce consistent results. If the answers converge, the decision is considered reliable. If they diverge, the system registers a high level of uncertainty.

This approach addresses two problems at once. First, it reduces the kind of blind confidence that AI should not have. Second, it helps protect the model against manipulation. Adversarial attacks often exploit neural networks’ blind spots: for example, an imperceptible change to a pixel in an image can cause a model to mistake a stop sign for a speed-limit sign. If an algorithm can assess its own uncertainty, it is more likely to detect that something is wrong.

Five Years of the Reliability Race

The idea of trustworthy AI did not emerge overnight. In 2021, the research community began openly discussing how vulnerable neural networks were to even minor data distortions. It also became clear that accuracy alone was no longer enough – AI systems also needed resilience.

In 2022, Russian researchers presented self-learning models designed to protect industrial facilities. The models could adapt to new attack scenarios rather than simply respond to known patterns. It was an important step: the system was no longer static.

By 2023, the focus had shifted to explainability. Researchers at Innopolis University showed how to make models less dependent on incidental features in data. If an algorithm understands what it is actually looking at, it becomes harder to fool.

In 2024, a new problem emerged: attacks on large language models using specially crafted prompts. Prompt injection could make an AI system violate its own instructions. It became clear that protection was needed not only for classifiers but also for generative systems.

In 2025, the concept of resilience moved beyond cybersecurity. A neural network controlling noise reduction in the LIGO gravitational-wave detector learned to adapt to physical interference from the equipment. Reliability was no longer exclusively an IT problem.

Finally, in 2026, MEPhI introduced Mamba Shield, an architecture designed to resist data poisoning. The model was designed to remain operational even when its training dataset was deliberately corrupted. This is no longer just an individual algorithm but a broader philosophy for building AI systems.

When the Cost of an Error Is Too High

A stochastic transformer could be useful in situations where the cost of an error is too high. Consider medicine, where an algorithm that confidently makes a diagnosis without acknowledging uncertainty can put a patient at risk. A model that can say, “I need more data,” or “a doctor should review this,” is a very different proposition. The same applies to banking, where AI evaluates transactions for signs of fraud. In these settings, false confidence can be more dangerous than acknowledging uncertainty.

The researchers plan to adapt the approach for large language models using techniques such as LoRA, which would make it possible to embed the protection into large neural networks without fully retraining them. That matters because retraining a model with hundreds of billions of parameters is a resource-intensive task.

An Industry Shift From Speed to Confidence

The challenge of making AI reliable knows no borders. European and US research labs are working on similar questions: how to make models resilient to attacks, how to teach them to explain their own decisions and how to prevent cascading failures. The Russian development fits into this broader effort.

It is still too early to talk about directly exporting the architecture itself – the work remains at the research stage. But the direction has commercial potential. Demand for protected AI solutions is growing: corporations are reluctant to deploy algorithms they cannot trust in critical scenarios. What can be said with confidence is that the Russian development addresses a highly relevant problem with international significance.

The work by MEPhI researchers is part of a broader shift in the industry. If the race over the past decade was to make AI smarter, more accurate and faster, a different requirement is now coming to the fore: AI must be reliable, able to express uncertainty and recognize the limits of its own capabilities. Systems with these characteristics could ultimately earn a place in energy, medicine, transportation and government.

We have created not just another neural network but a theoretically grounded framework. Our work shows that stochasticity – randomness – in a model’s parameters does not get in the way. It actually helps: it makes AI both more secure and more reliable
quote
like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next