Sarvam AI launches Saaras V4, a multilingual speech‑recognition model covering 22 Indian languages and English
Sarvam AI announced Saaras V4 on 24 August 2026, a new generation speech‑recognition system that supports 22 Indian languages plus English, offers five built‑in transcription modes and claims state‑of‑the‑art word‑error rates across all supported languages.

On 24 August 2026 Sarvam AI unveiled Saaras V4, the latest iteration of its speech‑recognition platform. The announcement positioned the model as a single‑system solution for the 22 officially recognised Indian languages together with English, a coverage scope that has been rare among commercial speech engines.
The technical description supplied by Sarmam AI indicates that Saaras V4 combines a dedicated audio encoder with a hybrid state‑space large language model (LLM) decoder. The decoder contains three billion parameters and was trained entirely in‑house, meaning that Sarvam AI retained full control over the data pipeline, model architecture and optimisation procedures.
Integrated transcription modes remove post‑processing steps
A notable feature of Saaras V4 is the inclusion of five transcription modes directly within the engine: transcribe, verbatim, codemix, translit and translate. By embedding these capabilities, the model eliminates the need for separate post‑processing stages that traditionally introduce alignment errors, formatting mismatches or language‑identification mistakes.
The transcribe mode delivers clean, punctuated text suitable for general applications. Verbatim retains every filler, hesitation and non‑lexical sound, useful for legal or medical records. Codemix processes speech that blends multiple languages within a single utterance, a common pattern in Indian multilingual contexts. Translit converts spoken content into a target script, while translate produces an English rendering of the original utterance, enabling cross‑language accessibility without external translation pipelines.
Performance claims and external validation limits
Sarvam AI states that Saaras V4 achieves the lowest word‑error rate (WER) reported on all 22 Indian languages and on seven English benchmark sets that include global accents and Indian English variants. These claims are presented as internal benchmarks, and the company has not yet released independent replication studies or third‑party audit results.
The absence of external validation means that organisations must treat the reported WER figures as provisional. While internal testing can demonstrate strong performance under controlled conditions, real‑world deployments often encounter acoustic environments, speaker demographics and dialectal variations that differ from the training corpus. Until peer‑reviewed evaluations appear, the exact margin of improvement over competing systems remains uncertain.
Robustness to noise, code‑mixing and dialects
According to Sarvam AI’s internal tests, Saaras V4 maintains robustness in noisy recordings, handles code‑mixed speech and adapts to dialectal shifts. The company reports a language‑identification error rate of 2.9 % for the ten most spoken Indian languages and 5.22 % across the full set of 22 languages. These figures suggest that the model can reliably detect the language of an utterance even when speakers alternate between languages within a single sentence.
The robustness claims are tied to internal evaluation protocols that simulate background noise and mixed‑language inputs. No public datasets or reproducible scripts have been released, which limits the ability of external researchers to confirm the model’s tolerance to adverse acoustic conditions.
Streaming latency and deployment considerations
Saaras V4 is advertised as supporting low‑latency streaming. The documentation specifies a time‑to‑first‑token of less than 150 milliseconds, and the engine can process multi‑minute recordings at a speed of one second of audio per second of compute. These metrics are relevant for real‑time transcription services, call‑center analytics and live captioning.
However, the self‑hosting guide currently only covers version 3 of the platform. The lack of on‑premises documentation for version 4 creates a barrier for organisations that require local deployment due to data‑privacy regulations or network constraints. Customers may need to rely on Sarvam AI’s managed cloud offering until the v4 deployment guide is published.
- Three‑billion‑parameter hybrid state‑space LLM decoder
- Five built‑in transcription modes (transcribe, verbatim, codemix, translit, translate)
- Language‑identification error rate: 2.9 % (top 10 languages), 5.22 % (all 22)
- Streaming latency: first token <150 ms, processing speed 1 × real‑time
Pricing is transparent and tiered. Sarvam AI lists a cost of 30 ₹ per hour for real‑time or batch transcription, while the addition of speaker diarisation raises the rate to 45 ₹ per hour. These rates apply to usage of the hosted service; the model weights themselves are not released as open‑source artefacts, preventing direct modification or redistribution.
The closed‑source nature of the model weights means that organisations cannot audit the underlying architecture for bias, security vulnerabilities or compliance with local regulations. While the pricing structure is simple, the lack of model openness may influence procurement decisions for entities that prioritise transparency.
From a practical standpoint, enterprises looking to adopt Saaras V4 must weigh the advantages of integrated multilingual capabilities against the uncertainties surrounding external validation and on‑premises deployment. The model’s ability to handle code‑mixing and dialectal variation aligns with the linguistic reality of many Indian markets, potentially reducing the need for multiple specialised engines.
Nevertheless, organisations should plan for a validation phase that includes pilot testing on their own audio corpora, measurement of latency in their network environment, and assessment of language‑identification accuracy for the specific dialects they encounter. Such a trial can reveal any gaps between the internal benchmarks reported by Sarvam AI and the performance observed in production.
In summary, Saaras V4 represents a significant technical step for Sarvam AI, offering a unified solution for a broad set of Indian languages and English with multiple transcription options and low‑latency streaming. The company’s internal claims of state‑of‑the‑art accuracy and robustness are promising, but the absence of independent verification and limited deployment documentation introduce measurable risk. Organisations that adopt the service should conduct thorough internal testing, consider the implications of a closed‑source model, and monitor forthcoming updates to the self‑hosting guide before committing to large‑scale, mission‑critical deployments.
Sources
- Introducing Saaras V4 - Sarvam AISarvam AI · September 25, 2026
- Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - MarkTechPostMarkTechPost · September 26, 2026



