Built with AI4Bharat, the new speech recognition system is designed to handle India’s diverse languages, accents and scripts, with separate models aimed at higher accuracy and broader multilingual coverage.
BodhanAI, an initiative associated with IIT Madras, has introduced Indic-Transcribe, a speech recognition system developed in collaboration with AI4Bharat and designed specifically for India’s multilingual environment. The model supports 26 Indian languages in addition to English and is aimed at improving the accuracy of speech-to-text applications across diverse regional use cases.
The 1.2-billion-parameter model has been developed to account for differences in languages, scripts and accents, which can make speech recognition particularly challenging in India. BodhanAI plans to make Indic-Transcribe accessible through its website from September 5.
Professor Mitesh Khapra announced the model on X, describing it as a system built to understand the way people across India actually speak, reflecting the country’s wide linguistic diversity.
Indic-Transcribe is being offered in two variants, Core and Flex. According to BodhanAI, the Core version prioritises recognition accuracy, while Flex is designed to provide broader coverage, including speech involving multiple scripts and Romanised text.

Focus on India’s diverse speech patterns
Speech recognition has become an increasingly important area for artificial intelligence companies as voice-based applications expand. India presents a particularly complex environment because people frequently switch between languages, use different scripts and speak with region-specific accents.
BodhanAI says Indic-Transcribe has been developed with these characteristics in mind. Its supported languages include Hindi, Bengali, Odia, Bhojpuri, Punjabi, Nepali, Haryanvi and Maithili, among others.
The model is also designed to handle situations that go beyond conventional single-language transcription. These include mixed-language speech, such as cricket commentary containing words from different languages, as well as railway announcements involving multiple scripts or languages.
BodhanAI says the system can also process Sanskrit shlokas and transcribe songs, potentially extending its applications beyond conventional business and conversational speech recognition.
The model's coverage also extends to languages spoken by smaller populations. Bodo, used in Assam, and Santali, spoken in parts of Jharkhand and other regions, are among the languages included in the system.
Education pilot covers 12 languages
BodhanAI has also begun pilots involving Indic-Transcribe across 12 languages. The initiative has accumulated around 12,000 hours of speech data from students in Classes 1 to 10, with the organisation stating that parental consent was obtained for the programme.
The benchmark results shared by BodhanAI indicate a Word Error Rate (WER) of 8.9 for Indic-Transcribe Core, compared with 10.1 reported for Sarvam’s Saaras V3. The Flex variant recorded a WER of 11.1, while BodhanAI cited 18.9 for Google Gemini 3 Pro in its comparison.
BodhanAI is the national Centre of Excellence for Artificial Intelligence in Education and operates through the IIT Madras Bodhan AI Foundation. Supported by the Ministry of Education, the organisation is focused on developing and promoting AI capabilities for India’s education ecosystem.
With Indic-Transcribe, the initiative is extending that focus into speech technology, targeting applications where language diversity, regional accents and mixed-language communication remain significant challenges for conventional speech recognition systems.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.




