Sarvam Document Digitization provides enterprise-grade document processing powered by Sarvam Vision, our state-of-the-art multimodal model. Transform any document into structured, searchable, and machine-readable data with world-class accuracy.
Sarvam Vision is a 3B-parameter, state-space Vision Language Model (VLM) purpose-built for high-accuracy Document Intelligence. Part of Sarvam's sovereign model series, it treats document understanding as knowledge extraction rather than simple text capture - extracting text, converting complex tables, and preserving layout, reading order, and hierarchy from PDFs and scanned images.
Sarvam Vision delivers world-class accuracy across 23 languages (22 official Indian languages + English), with native support for Indian scripts where most global models fall short. It achieves best-in-class scores on global benchmarks (olmOCR-Bench, OmniDocBench V1.5) for English, and leading accuracy on the Sarvam Indic OCR Bench for Indian languages - outperforming frontier models on Indic document tasks.
Ideal for digitizing scanned archives, Indic OCR at scale, and table-heavy documents such as scientific literature, financial reports, government bulletins, historical manuscripts, textbooks, and newspapers.
Highlights
Indic-first document intelligence - native, high-accuracy OCR across 22 official Indian languages plus English (23 total).
Knowledge extraction, not just text - parses complex tables, charts, and multi-column layouts while preserving reading order and document structure; outputs clean HTML or Markdown.
Efficient & benchmark-leading - a compact 3B state-space VLM that leads on the Sarvam Indic OCR Bench and scores competitively on global benchmarks (olmOCR-Bench, OmniDocBench V1.5).
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour for the compute instance that runs Sarvam Vision document inference. Pricing splits into two processing modes. Batch mode runs on ml.g6 instances for asynchronous, bulk document jobs. Real-time mode runs on both ml.g6 and ml.g6e instances for immediate responses. Within each mode, you choose an instance size from xlarge up to 48xlarge. Larger sizes add compute capacity and cost more per hour. You select the mode and size that fit your workload, then pay only for the hours the instance runs.
Top-of-mind questions for buyers
What does one HostHrs unit cover, and when do charges start and stop?
One HostHrs unit is one hour that a chosen instance runs the Sarvam Vision model. Charges accrue for each hour the instance is active. When you stop or terminate the instance, software charges stop. Underlying AWS infrastructure fees may still apply for stored resources.
How does batch mode differ from real-time mode for billing purposes?
Batch mode runs on ml.g6 instances and processes documents asynchronously in bulk, suited to large jobs. Real-time mode runs on ml.g6 and ml.g6e instances for immediate responses. Both bill per instance-hour. Real-time mode adds ml.g6e instance options that batch mode does not offer.
What input limits apply when running Sarvam Vision, and do they affect how many hours I pay for?
Each PDF is capped at 10 pages and files are limited to 200 MB. Split larger PDFs before processing. These caps set per-job scope but do not change the hourly billing. You still pay for each hour the instance runs, regardless of document count.
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
Version release notes
Launching Sarvam's vision model supported on L4 and L40s GPU series with GPU Cluster support
Additional details
Inputs
Outputs
Usage instructions
Sample notebooks
Inputs
Summary
Summary
Documents to be digitized - PDFs or scanned page images. Provide a single PDF, individual PNG/JPG page images, or a flat ZIP archive of page images (up to 10). Optionally specify the target language code and output format.
Input MIME type
application/pdf
Limitations for input type
Upload up to 10 pages for sync endpoint. For higher number of pages use async endpoint
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Sarvam AI's APIs are a modular set of AI tools built for developers creating scalable AI applications tailored to Indian languages. They include APIs for speech-to-text, text-to-speech, translation, transliteration, language identification, analytics, and document parsing. Designed for accuracy with code-mixed and accented speech, these APIs enable seamless integration into web, mobile, and voice systems. They are production-ready with low latency, flexible pricing, and extensive documentation, giving developers the building blocks they need to power real-world AI solutions across diverse platforms.
Sarvam Arya unifies your enterprise data from CRMs, reports, PDFs, tables, and data lakes into a single, secure, searchable system. It enables teams to retrieve information quickly and accurately by connecting scattered data sources with AI-driven precision search. Once the data foundation is in place, businesses can deploy AI agents to automate repetitive tasks, streamline workflows, and reduce manual effort. Sarvam Arya helps enterprises eliminate bottlenecks, make informed decisions, and scale automation.
Sarvam Samvaad, our enterprise-grade Conversational AI Platform, is built to help enterprises build, test, and launch AI agents tailored for India. Exceptionally fluent in 11 Indian languages with authentic accents, it powers seamless interactions across telephone, WhatsApp, web, and apps, ensuring users can engage effortlessly, no matter the channel. Designed for production-grade performance, Sarvam Samvaad deploys AI agents that work around the clock, follow every instruction, and handle complex phrases, alphanumerics, and proper nouns with high precision. The platform makes it easy to move from pilot to production with rapid time-to-value, supporting every step from crafting and customizing the agent to integrating knowledge bases and connecting with top telephony providers. Once live, advanced analytics help data teams monitor every conversation performance and extract deeper insights from call logs.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.