Overview
Speech to Text REST (Saaras V3) - Process short audio files with immediate response. Best for quick transcriptions and testing with a maximum duration of 30 seconds.
Speech to Text Websocket (Saaras V3) - Transform audio into text in real-time with our WebSocket-based streaming API. Built for applications requiring immediate speech processing with minimal delay.
Highlights
- Indic-first ASR across 23 languages - 22 Indian languages + English, with automatic language detection and code-mixed audio support.
- Five output modes in one model - transcribe, translate (to English), verbatim, transliterate (Roman script), and code-mix, selectable per request.
- Production-grade for real workloads - optimized for 8 kHz telephony audio, with utterance-level timestamps and intelligent proper-noun/entity preservation; accepts common audio + telephony codecs. Trained on 1M+ hours of real Indian speech.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.g6.xlarge Inference (Batch) Recommended | Model inference on the ml.g6.xlarge instance type, batch mode | $4.00 |
ml.g6e.xlarge Inference (Real-Time) Recommended | Model inference on the ml.g6e.xlarge instance type, real-time mode | $4.00 |
Vendor refund policy
Contact support@sarvam.ai fot details!
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Make Sarvam's STT powered by in-house saaras:v3.1 available on AWS Sagemaker
Additional details
Inputs
- Summary
Audio to be transcribed — a call recording, voice note, meeting clip, or broadcast segment. Provide a single audio file in any supported format. Optionally specify the language code (or leave unset for auto-detection) and select an output mode.
- Limitations for input type
- Audio length is capped at 30s for realtime API. For longer audios, websocket api can be used
- Input MIME type
- multipart/form-data, application/json
Resources
Vendor resources
Support
Vendor support
Saaras (STT) model docs:
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.