The new models support speech, translation and OCR across multiple Indian languages, offering open weights and hosted APIs on sovereign infrastructure to help developers, educators and institutions build locally relevant artificial intelligence applications.
Bodhan AI, an AI-focused Centre of Excellence in Education incubated at the Indian Institute of Technology Madras (IIT Madras), has introduced four foundational artificial intelligence models aimed at strengthening India’s multilingual technology ecosystem. The models have been developed in collaboration with AI4Bharat and are being made available as Digital Public Goods.
The newly launched portfolio addresses four key areas: speech recognition, speech generation, machine translation and optical character recognition (OCR). By making the models accessible through open weights and hosted application programming interfaces (APIs), Bodhan AI aims to allow developers, educational institutions and technology organisations to customise the models and create applications suited to Indian users.
Models target India’s multilingual needs
The speech recognition model supports 27 Indian languages, while the OCR model can process content across 23 languages. Bodhan-Translate supports 22 languages, and the text-to-speech model covers 23 languages, expanding the scope for voice- and vision-based applications in education and other sectors.
The models have been trained and optimised using NVIDIA Nemotron open models and technologies, including the NVIDIA NeMo framework. Bodhan AI has also post-trained NVIDIA Nemotron 3.5 ASR to improve its ability to handle Indian languages, including regional dialects and accents.
Building sovereign AI infrastructure
The initiative is part of the broader Bharat EduAI Stack, which is being developed as sovereign digital public infrastructure for education. The objective is to strengthen AI capabilities for India’s diverse linguistic environment and support the development of locally relevant educational technologies.
IIT Madras Director V. Kamakoti said India’s AI journey requires technology that “understands India”, describing the new voice and vision models as an important step towards building sovereign digital public infrastructure for AI in education.
Bodhan AI and NVIDIA are also expected to work together on datasets, training approaches and evaluation methods for future foundational models targeting Indian languages. AI4Bharat will contribute its experience in open-source datasets, tools and models to further advance Indian-language AI development.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.




