bg
Science and new technologies
13:29, 10 August 2026
views
5

Moscow State University Researchers Are Improving AI Assistants’ Memory

Researchers at Moscow State University’s Faculty of Computational Mathematics and Cybernetics have proposed a new approach to how artificial intelligence systems can remember and use information during long-term interactions with users.

LLMs handle conversational tasks well, but they have one major limitation: no built-in long-term memory. A model relies only on the current message context, and once that context is trimmed, it forgets facts, mixes up details and loses consistency in its persona. Longer contexts also drive up the cost of running the model.

AI does not have a single, universal form of memory. Instead, it relies on a cascade of memory layers, each with its own characteristics, lifespan and operating rules. How those layers work is one of the main factors determining just how “smart” and consistent a conversation with an AI system can feel.

Anyone who has carried on a long conversation with a modern chatbot has probably encountered the same problem: after a few dozen messages, it “forgets” the beginning of the conversation, mixes up details, repeats questions or simply loses the thread. The reason is straightforward – a language model operates within a limited context window, and anything that falls outside that window effectively ceases to exist for the model. Researchers at Moscow State University’s Faculty of Computational Mathematics and Cybernetics have proposed an approach that could change the underlying logic of how people interact with AI.

Don’t Store Everything, Choose What Matters

At the Lomonosovskie chteniya (Lomonosov Readings) conference, the researchers presented a long-term memory architecture for AI assistants. The central idea is that instead of mechanically accumulating an entire conversation history, a separate subsystem decides for itself what deserves to be retained. It determines which information is genuinely important, what should be deleted, what can be condensed, tagged or reformulated. Before the assistant generates a response, the system retrieves relevant records, assembles them into a compact context and passes only that context to the language model. The result is less information noise, less duplication and greater consistency across long conversations.

For now, this is a research project rather than a market-ready product. But the problem it addresses is one of the key technological limitations facing today’s AI agents.

Not the First Word, but an Important Step

Russia’s market is undergoing a noticeable shift: companies are moving beyond experimental chatbots toward full-fledged AI assistants embedded in corporate workflows. According to CNews, agentic systems became one of the main areas of development for Russia’s AI sector in 2026. Existing products include Robin AI, BPMSoft AI, SimpleOne GenAI, Elma Cortex and Sherpa AI. Sooner or later, all of them run into the same question: how can an assistant remember its history with an employee, customer or process without sending gigabytes of past conversations to the model with every request?

For everyday users, that raises the prospect of an assistant that does not need to have its context explained from scratch every time. For businesses, it could mean a corporate assistant capable of working with months or even years of accumulated data without prohibitive computing costs.

The concept of structured long-term memory for language models is not new. In 2023, the international MemGPT project, later renamed Letta, proposed separating a model’s limited context from external memory that the agent manages itself. In 2024, researchers introduced HippoRAG, an approach inspired by how the human hippocampus works. That same year, ChatGPT gained its Memory feature for remembering user preferences, and a year later Google expanded Gemini’s memory capabilities. By 2026, analysis of previous conversations had also become available in Gemini’s free version: long-term context is becoming a standard feature of mainstream AI services.

In parallel, an entire class of specialized systems is growing – Mem0, Letta and Graphiti – with the emphasis shifting from simply storing information to selecting it, updating it and deleting outdated records. The MSU project conceptually belongs to this same line of development.

Looking Ahead

It is too early to export the research architecture itself: the published information describes an approach, not a commercial platform. Memory architectures, however, are tied to neither a particular language nor an industry, making them potentially applicable to any LLM-based system. If the method passes real-world testing and demonstrates an advantage in comparative benchmarks, it could become a sought-after component of Russian enterprise assistants, particularly in settings where interactions extend over months.

Memory in LLMs is always built from several layers. One part works quickly and directly during your conversation, another stores a user’s personal interests, while a third holds corporate knowledge or documents. This architecture is what gives modern AI its flexibility and convenience, but it also introduces important limitations. Anything that does not make it into the active memory layer becomes inaccessible to the model, even if it is still being stored somewhere on a server behind the scenes. The main point of new developments is not to try to ‘stuff’ more data into the model. It is to teach the machine to forget what it does not need
quote
like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next