Bashkortostan Is Building Datasets to Teach AI Its Native Language
The Republic of Bashkortostan has launched a large-scale effort to build open datasets for training AI systems to understand the Bashkir language.

The region is the first in Russia to take a systematic approach to preserving a national language by digitizing archives, literature and audio recordings. The resulting datasets are expected to serve as a foundation for voice assistants, translation tools and educational platforms.
Collecting Big Data for AI
The project is coordinated by an interagency working group under the republic's government and the ANO for the Preservation and Development of the Bashkir Language. Government agencies have already signed agreements with key content holders in the region. The collaboration includes the Bashkortostan Television and Radio Broadcasting Company, Kitap Publishing House, the Bashkortostan Republican Special Library for the Blind, the Bashkir Institute for Education Development and Bashkir State Pedagogical University.
Developers and linguists face the substantial task of cleaning and structuring the material, since much of the archival data cannot be used to train AI systems in its current form. So far, the teams have processed about 75,000 pages of text and roughly 130 hours of audio recordings. The database is set to include not only contemporary books but also newspaper archives, magazine articles, television and radio broadcasts, and specialized educational materials.
The first result of the project is an open collection of narrated audiobooks. It contains 476 audio clips from 14 works, including the folk epics Akbuzat and Aldar i Zukhra (Aldar and Zukhra). Each audio clip is paired with synchronized text in Bashkir and Russian, a crucial resource for training machine translation and speech recognition algorithms.

Digitizing Dialects and Historical Sources
Today's AI tools cannot handle the recognition of Bashkir text in old archives on their own. When digitizing Soviet-era newspapers, algorithms struggle with similar-looking characters, complex word breaks, a mixture of Russian and Bashkir typography, and outdated page layouts. Developers use frequency analysis and spelling dictionaries to check the recognized text. In one pilot project, linguists assembled a corpus containing 840,000 word forms.
Newspapers and magazines do more than increase the size of the text corpus. They contain news stories, interviews, official terminology, conversational language, satire and examples of wordplay. An AI model needs to understand not only the literal meaning of expressions but also how people write and communicate in different contexts.
The teams are also accounting for regional variations in speech. Different parts of Bashkortostan have their own dialects, ranging from the northwestern dialect to the speech of people in the Kugarchinsky District. Historical materials, folklore, place names and recordings of spoken language are all being brought into digital form. The more diverse the dataset, the more accurately a language model can understand Bashkir.

Bringing Bashkir Into Digital Services
High-quality datasets could pave the way for a full-fledged digital ecosystem in Bashkir. The language is already integrated into major commercial services. Yandex Translate, for example, has supported Bashkir since 2015. The company's services now support 24 languages spoken by peoples of Russia, with advanced speech synthesis and recognition technologies available for seven of them, including Bashkir.
Machine translation, however, is only one potential application. Models trained on the new datasets could power smart keyboards with predictive text, automatic video captioning systems and educational platforms. Voice assistants could respond to users speaking Bashkir while recognizing its distinctive phonetics and intonation, and algorithms could pick commands out of speech even in the presence of background noise.
The open datasets could also provide a foundation for independent development. Startups and IT companies will gain access to cleaned and labeled data they can use to build their own specialized applications.

Technological Sovereignty and Cultural Preservation
All materials go through an approval process with copyright holders, while documentation for state-owned works is prepared to make them available for open use. This makes the datasets legally usable and protects developers from potential lawsuits.
Preserving a language in the 21st century requires giving it a place in the digital world. If AI systems cannot understand Bashkir, speakers have to switch to more widely used languages when interacting with search engines and smart devices. Building high-quality language models can make their native language more viable in the global information environment.
Successful implementation of the project would demonstrate the effectiveness of public-private partnerships in language technology. In the future, similar interagency working groups and centers of expertise could replicate these methods to digitize the cultural heritage of other Indigenous peoples in Russia.









































