IISc Launches SraVaani: Speech AI Model for 65 Indian Languages

TECHNOLOGY
Whalesbook Logo
AuthorAnanya Iyer|Published at:
IISc Launches SraVaani: Speech AI Model for 65 Indian Languages

The Indian Institute of Science has released SraVaani, an open-source speech AI model covering 65 languages to support 25 crore people. Developed with ARTPARK and Google, the tool aims to expand inclusive AI access for Indian regional dialects.

The Indian Institute of Science (IISc) has launched SraVaani, a new speech recognition AI model designed to bridge the language gap in artificial intelligence. Developed by the SPIRE Lab at IISc in partnership with ARTPARK and with support from Google, the model is built to understand 65 different Indian languages and dialects.

This release aims to address the needs of over 250 million people whose primary languages have historically been underserved by global AI models. While many existing speech technologies focus on widely used scheduled languages, SraVaani incorporates 20 scheduled and 45 additional regional languages. By making this model open-source on the platform Hugging Face, IISc allows developers and startups to use it freely under an MIT license, which could lower the costs and technical barriers for companies building localized Indian applications.

Technical Foundation and Performance

The model is built upon the expansive Vaani dataset, a large-scale project that collected over 31,000 hours of spontaneous speech from 156,000 individuals across 165 districts in India. The focus was on capturing natural, real-world speech rather than rehearsed audio. This training data helped the model perform significantly better on less common languages than existing systems. For example, in tests involving the Garo language, SraVaani achieved a 9.5% word error rate, showing a major improvement compared to other systems which often struggle with high error rates in regional dialects.

Impact on the Indian AI Ecosystem

For the Indian technology sector, the release of SraVaani is a step toward building "Sovereign AI," where technology is built specifically for local needs rather than relying on global models that may not understand regional nuances. By removing the need for manual language tagging and providing automated language identification, SraVaani helps developers integrate voice features into apps without needing extensive technical expertise. This is particularly valuable for sectors like banking, e-governance, and e-commerce, where reaching users in their native language is a key growth strategy.

Challenges to Monitor

While the model shows promise, technology experts note that maintaining high accuracy across 65 diverse languages remains a complex challenge. Real-world conditions, such as background noise in busy environments or varying accents, can still impact the performance of speech recognition tools. As companies and developers begin to test SraVaani in commercial applications, the consistency of the model across different demographics and audio environments will be a key factor for its widespread adoption. Investors and industry participants may watch how this open-source tool influences the cost of developing vernacular-focused digital services in the coming months.

Disclaimer: This article is published for informational purposes only. This is not a buy sell recommendation.