Yandex Translate Adds Karelian
Yandex Translate now supports the Karelian language. The launch of the new language model was timed to coincide with Russia’s Indigenous Peoples’ Languages Day.

To train the machine-translation algorithms, engineers assembled a specialized digital corpus containing more than 86,000 sentences. These bilingual texts include an original and an accurate Russian translation. The neural network analyzes these pairs and learns to recognize recurring language patterns.
The team also uploaded about 100,000 untranslated sentences to the system. This larger body of text helps expand the model’s context. The algorithm studies Karelian grammar, syntax and word frequency, allowing it to “understand” sentence structure and correctly inflect words even when there are no direct equivalents in Russian.
Training neural networks for languages with limited amounts of data requires advanced mathematical models. Engineers use transfer-learning techniques: the system draws on baseline knowledge acquired from processing widely used languages and adapts it to the characteristics of the Finno-Ugric language group. This significantly speeds up model training and improves the final accuracy of translations.

Sources of Data and the Role of Experts
Building a high-quality text corpus required linguists and IT professionals to work as a team. The data came from literary works, online encyclopedias and archival materials. Authors and correspondents of the OmaMua newspaper made a significant contribution to the database, because publications in national languages contain contemporary vocabulary and current expressions.
Scientific and educational institutions across the republic also became actively involved. Developers used data from VepKar, an open corpus of the Veps and Karelian languages. The academic database contains annotated texts with morphological and syntactic analysis. Another source was LiPaS, a multimedia dictionary of the Karelian language. It contains not only text pairs but also examples showing how words are used in different contexts.
The Government of Karelia and Yandex signed an agreement to implement the project in 2026 at the St. Petersburg International Economic Forum. Karelian-language teachers, local journalists and volunteers joined the initiative. Native speakers even manually reviewed machine-generated hypotheses and corrected neural-network errors, which made a major contribution to the quality of the model’s training.

Technologies for Low-Resource Languages
Integrating regional languages into global digital services presents a difficult engineering challenge, because standard translation models require millions of sentence pairs to reach an acceptable level of quality. Karelian is considered a low-resource language, and the amount of available digitized text is objectively limited by historical and demographic factors.
Developers compensate for the lack of data through synthetic text expansion. The algorithms generate different variations of existing sentences by changing grammatical cases and tenses. The neural network also identifies hidden relationships between words, allowing the system to correctly translate complex grammatical forms found in Finno-Ugric languages.
The service supports translation of individual words, phrases and entire paragraphs and is available on both PCs and mobile devices.

Preserving Linguistic Heritage and Plans for the Future
The addition of Karelian to Yandex Translate is more than a new feature in a popular service. Digitalization is becoming a key tool for preserving the country’s linguistic diversity. When a language is integrated into widely used IT services, it gains a new environment in which it can remain in active use and develop naturally.
The service now supports more than 20 languages spoken by the peoples of Russia. Integrating each new language requires months of painstaking work by professionals. Yandex continues to expand the country’s language map: by the end of 2026, the company plans to add support for Veps, marking another step toward digitizing the linguistic heritage of northwestern Russia.
The project offers a successful model of public-private partnership for cultural preservation. Regional authorities provide access to academic databases and bring in experts, while the IT company supplies computing resources and advanced machine-learning algorithms. This collaboration among government, academia and business can help preserve Indigenous peoples’ cultures in the digital age.









































