Indian artificial intelligence startup Sarvam AI has launched Saaras V4, the latest version of its automatic speech recognition (ASR) model, with a focus on multilingual speech, Indian languages and real-world audio conditions.
The new model is designed to process speech across Indian languages and English while supporting different forms of transcription, including verbatim text, transliteration and translation. Sarvam AI has also built capabilities for handling code-mixed conversations, dialect variations and noisy audio into the model.

Credits: Sarvam AI
Saaras V4 Offers Five Speech Output Modes
Saaras V4 combines an audio encoder with a 3-billion-parameter hybrid state-space language model developed by Sarvam AI. The architecture is designed to help the model process speech that may not follow clean, standardised audio patterns.
According to Sarvam AI, the model can convert the same audio into five different output formats: verbatim transcription, normalised text, code-mixed text, transliteration and translation.
The company says these capabilities are integrated directly into the model instead of relying on separate post-processing systems.
For English speech, Saaras V4 was evaluated across seven benchmarks covering areas such as Indian English, international accents, meetings, financial conversations and media. Sarvam AI said the model recorded the lowest average Word Error Rate (WER) across the datasets it evaluated.
For Indian languages, the model was tested on the Vistaar benchmark across 10 languages. The evaluation used both standard WER and LLM-WER. While WER measures transcription errors, LLM-WER also considers the semantic impact of an error, helping distinguish changes that affect meaning from differences in formatting or spelling conventions.
Model Targets Indian Languages and Noisy Audio
Indian language support is a central part of Saaras V4’s design.
The model can automatically identify the language being spoken and generate the transcript in its corresponding native script. Sarvam AI reported a 5.22% language identification error rate across 22 Indian languages, while the rate stood at 2.9% across the 10 most widely spoken languages included in its evaluation.
Saaras V4 also supports keyterm prompting, allowing developers to provide names, product terms, acronyms and other words that could otherwise be difficult for a speech recognition model to identify correctly.
On the IndicContextEval benchmark, Sarvam AI reported a WER of 16.03% in the L5 keyword-prompting setting.
The model’s ability to handle code-mixed speech could also be important for applications involving conversations where speakers switch between Indian languages and English. Such speech patterns are common in informal conversations, customer interactions and other real-world settings.

Credits: Swarajya
Saaras V4 Brings Streaming and Developer Access
Sarvam AI has also focused on making Saaras V4 suitable for applications that require near-real-time speech recognition.
The company says the model supports streaming with a time to first token of less than 150 milliseconds. It can also process long-form audio, expanding its potential use beyond short voice commands.
Developers can access Saaras V4 through Sarvam AI’s API, with support for Python and Node.js SDKs. The model can also be integrated with platforms including Vercel AI SDK, LiveKit Agents and Pipecat Agents.
The API provides access to the model’s different speech processing capabilities through modes including transcribe, translate, translit, verbatim and codemix.
With Saaras V4, Sarvam AI is positioning speech recognition as a broader multilingual infrastructure layer rather than simply a transcription tool. Its focus on Indian languages, code-mixed speech, contextual terms and different output formats reflects the challenges involved in deploying speech AI outside controlled audio environments.
The launch also expands Sarvam AI’s focus on building AI systems tailored to India’s linguistic diversity. By combining language identification, contextual prompting and multiple transcription formats, Saaras V4 is aimed at use cases such as customer support, voice interfaces, meetings and other applications where accurate speech processing across languages is important.



