Overview
Nemotron 3.5 ASR - Enterprise Multilingual Speech-to-Text
Nemotron 3.5 ASR is a production-ready, GPU-accelerated speech-to-text appliance powered by NVIDIA, designed to help organizations add accurate voice transcription to contact centers, AI assistants, meeting platforms, media workflows, and more. Because it runs entirely within your AWS VPC as an AMI, all audio data stays on your infrastructure - no data leaves the instance, providing a strong data-sovereignty posture for privacy-sensitive workloads.
Key Capabilities
- Real-time streaming and batch transcription from a single deployment
- 40 language-locales with optional automatic language detection
- Built-in punctuation and capitalization for clean, readable transcripts
- Fine-tuning support to improve recognition for specialized vocabularies and industry terminology
- Configurable performance settings to balance latency and accuracy per workload
Use Case Example: Contact Center Post-Call Analytics
A financial services contact center processing thousands of calls daily can deploy Nemotron 3.5 ASR to transcribe recordings in batch, feeding transcripts into downstream analytics and compliance systems. The workflow is straightforward: audio files are ingested from S3, transcribed via the ASR API running on a GPU instance, and the resulting text is stored back in S3 or streamed to a CRM or quality-management platform. Because processing happens within the VPC, sensitive customer conversations never leave the organization's cloud boundary - a critical requirement for regulated industries.
Deployment Details
- Supported AWS instance types: GPU-accelerated instances (e.g., g4dn.xlarge, g4dn.2xlarge, g4dn.4xlarge, g4dn.8xlarge) with sufficient GPU memory for model inference
- Operating system: Ubuntu-based AMI 26.04 with NVIDIA GPU drivers pre-installed
- Networking: Configure security groups to allow API traffic on the designated service port, port 8000
- Launch steps: Start the AMI on a supported GPU instance, allow 5 minutes for the service to fully launch, and begin sending audio via the streaming or batch API
Data Handling and Security
- All audio processing occurs locally on the EC2 instance within your VPC
- No audio data or transcription results are transmitted to external servers
- Supports encryption at rest via AWS EBS encryption
- No telemetry or usage data is sent outside your environment
Why Nemotron 3.5 ASR
Unlike per-minute cloud API services, Nemotron 3.5 ASR runs on your own GPU instances, giving you predictable costs at scale and full control over data residency. A single multilingual model eliminates the complexity of managing separate per-language deployments. High GPU utilization means more concurrent transcription sessions per instance, reducing infrastructure spend compared to alternatives that require dedicated resources per language or workload type.
Whether you are building conversational AI, automating customer interactions, creating searchable media archives, or transcribing business meetings, Nemotron 3.5 ASR provides enterprise-grade speech recognition optimized for challenging real-world audio conditions.
Highlights
- Single Multilingual Model: Transcribe speech across 40 language-locales using one multilingual model, with optional automatic language detection.
- Punctuation and Capitalization: Automatically produces transcripts with proper punctuation and capitalization for polished, readable output.
- Cost Efficient and Performant: Optimized to maximize GPU performance, allowing more concurrent transcription sessions per instance and helping organizations reduce the cost of large-scale speech recognition deployments.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Free trial
Dimension | Cost/hour |
|---|---|
g4dn.xlarge Recommended | $0.09 |
g4dn.2xlarge | $0.09 |
g4dn.4xlarge | $0.09 |
g4dn.8xlarge | $0.09 |
Vendor refund policy
Refunds may be considered on a case-by-case basis.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
64-bit (x86) Amazon Machine Image (AMI)
Amazon Machine Image (AMI)
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Production-ready release.
Additional details
Usage instructions
Once deployed, ensure port 8000 is allowed in the EC2 instance's security group. Then, replace the public DNS endpoint and sample audio file in the following curl command and run it to test the service:
curl http://ec2-123-123-123-123.compute-1.amazonaws.com:8000/v1/audio/transcriptions -F file=@TestAudio.mp3 -F model=nemotron-3.5-asr-streaming-0.6b
The exposed API endpoint adheres to the OpenAI spec and can be used as such.
Support
Vendor support
Getting Help
For deployment support, configuration assistance, performance tuning, and troubleshooting, contact the support team at: support@salientengineering.com
What Support Covers
- Initial deployment and AMI launch guidance
- Performance optimization and GPU utilization tuning
- Troubleshooting transcription errors or service health issues
Getting Started
After launching the AMI on a supported GPU instance, wait 5 minutes then verify the service is running by checking the health endpoint. The ASR API accepts audio input via streaming or batch modes.
Requesting Refunds
For billing questions or refund requests related to AWS Marketplace charges, contact support@salientengineering.com with your AWS account ID and subscription details.
Please include your AWS instance ID, region, and a description of the issue when contacting support to help expedite resolution.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products
