bg
Culture, sports and media
20:24, 05 October 2026
views
1

Bashkortostan Is Building Datasets to Teach AI Its Native Language

The Republic of Bashkortostan has launched a large-scale effort to build open datasets for training AI systems to understand the Bashkir language.

The region is the first in Russia to take a systematic approach to preserving a national language by digitizing archives, literature and audio recordings. The resulting datasets are expected to serve as a foundation for voice assistants, translation tools and educational platforms.

Collecting Big Data for AI

The project is coordinated by an interagency working group under the republic's government and the ANO for the Preservation and Development of the Bashkir Language. Government agencies have already signed agreements with key content holders in the region. The collaboration includes the Bashkortostan Television and Radio Broadcasting Company, Kitap Publishing House, the Bashkortostan Republican Special Library for the Blind, the Bashkir Institute for Education Development and Bashkir State Pedagogical University.

Developers and linguists face the substantial task of cleaning and structuring the material, since much of the archival data cannot be used to train AI systems in its current form. So far, the teams have processed about 75,000 pages of text and roughly 130 hours of audio recordings. The database is set to include not only contemporary books but also newspaper archives, magazine articles, television and radio broadcasts, and specialized educational materials.

The first result of the project is an open collection of narrated audiobooks. It contains 476 audio clips from 14 works, including the folk epics Akbuzat and Aldar i Zukhra (Aldar and Zukhra). Each audio clip is paired with synchronized text in Bashkir and Russian, a crucial resource for training machine translation and speech recognition algorithms.

Digitizing Dialects and Historical Sources

Today's AI tools cannot handle the recognition of Bashkir text in old archives on their own. When digitizing Soviet-era newspapers, algorithms struggle with similar-looking characters, complex word breaks, a mixture of Russian and Bashkir typography, and outdated page layouts. Developers use frequency analysis and spelling dictionaries to check the recognized text. In one pilot project, linguists assembled a corpus containing 840,000 word forms.

Newspapers and magazines do more than increase the size of the text corpus. They contain news stories, interviews, official terminology, conversational language, satire and examples of wordplay. An AI model needs to understand not only the literal meaning of expressions but also how people write and communicate in different contexts.

The teams are also accounting for regional variations in speech. Different parts of Bashkortostan have their own dialects, ranging from the northwestern dialect to the speech of people in the Kugarchinsky District. Historical materials, folklore, place names and recordings of spoken language are all being brought into digital form. The more diverse the dataset, the more accurately a language model can understand Bashkir.

Bringing Bashkir Into Digital Services

High-quality datasets could pave the way for a full-fledged digital ecosystem in Bashkir. The language is already integrated into major commercial services. Yandex Translate, for example, has supported Bashkir since 2015. The company's services now support 24 languages spoken by peoples of Russia, with advanced speech synthesis and recognition technologies available for seven of them, including Bashkir.

Machine translation, however, is only one potential application. Models trained on the new datasets could power smart keyboards with predictive text, automatic video captioning systems and educational platforms. Voice assistants could respond to users speaking Bashkir while recognizing its distinctive phonetics and intonation, and algorithms could pick commands out of speech even in the presence of background noise.

The open datasets could also provide a foundation for independent development. Startups and IT companies will gain access to cleaned and labeled data they can use to build their own specialized applications.

Technological Sovereignty and Cultural Preservation

All materials go through an approval process with copyright holders, while documentation for state-owned works is prepared to make them available for open use. This makes the datasets legally usable and protects developers from potential lawsuits.

Preserving a language in the 21st century requires giving it a place in the digital world. If AI systems cannot understand Bashkir, speakers have to switch to more widely used languages when interacting with search engines and smart devices. Building high-quality language models can make their native language more viable in the global information environment.

Successful implementation of the project would demonstrate the effectiveness of public-private partnerships in language technology. In the future, similar interagency working groups and centers of expertise could replicate these methods to digitize the cultural heritage of other Indigenous peoples in Russia.

Today, developing a language means not only preserving the texts, books, audio recordings and other materials that already exist, but also preparing them for use in the digital world. Language models are trained on data, so the capabilities a language gains in technology depend directly on which texts, speech recordings and other materials we collect and how well we prepare them
quote
like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next