NVIDIA Nemotron are open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents
Nemotron-3-Super-120B-A12B-FP8 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template.
The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Distinct from the Nano model, the Super model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using NVFP4 quantization to maximize compute efficiency. The model has 12B active parameters and 120B parameters in total.
The supported languages include: English, French, German, Italian, Japanese, Spanish, and Chinese
This model is ready for commercial use.
Highlights
Architecture Type: Mamba2-Transformer Hybrid Latent Mixture of Experts (LatentMoE) with Multi-Token Prediction (MTP)
Network Architecture: Nemotron Hybrid LatentMoE
Number of model parameters: 120B Total / 12B Active
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the host hour based on the AWS instance type you run the model on. Pricing is usage-based, so you are charged for the time each instance runs. One option uses the ml.g5.48xlarge instance in batch mode, which processes inference in groups rather than live. The other four options run real-time inference on the ml.p4de.24xlarge, ml.p5.48xlarge, ml.p5e.48xlarge, and ml.p5en.48xlarge instances, which return results on demand. Choose an instance type to match your workload and hardware needs.
Top-of-mind questions for buyers
What is the difference between batch mode and real-time inference for billing?
Both meter by host hour on the instance you run. Batch mode processes inference requests in groups, so you run the ml.g5.48xlarge instance to handle grouped jobs. Real-time mode returns results on demand and uses the ml.p4de, ml.p5, ml.p5e, or ml.p5en instances. You pay for each hour the instance runs.
Am I charged when the model is deployed but not actively processing requests?
Charges apply per host hour while the instance runs, whether or not it is actively processing inference. The meter tracks running time, not request volume. To stop software charges, you must shut down the endpoint. Underlying AWS infrastructure fees may apply separately based on instance state.
What does one host hour cover on these instance types?
One host hour is one hour of running a single instance of the chosen type. It covers running the Nemotron 3 Super 120B model on that hardware. Each instance type differs in GPU count and memory, so choose one that matches your workload and throughput needs.
docs.nvidia.com
Helpful?
Vendor refund policy
No Refunds.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
NVIDIA Nemotron-3-Super-120B-A12B accepts JSON requests via the /invocations API. The request contains a list of messages and optional generation controls. Reasoning behavior is controlled via the chat_template_kwargs parameter, supporting reasoning on, reasoning off, and low-effort reasoning modes. Both streaming (invoke_endpoint_with_response_stream) and non-streaming (invoke_endpoint) are supported.
Limitations for input type
Other Properties Related to Input: Maximum context length up to 1M tokens. Supported languages include: English, French, German, Italian, Japanese, Spanish, and Chinese.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Run sensitive AI workloads on shared GPU infrastructure without data exposure. Stained Glass Output Protection protects tokens as they are generated by the LLM. Pairing with Stained Glass Transform (not included in this marketplace product) gives round-trip LLM protection.
Llama 3.3 Nemotron Super 49B V1.5 is a significantly upgraded version of Llama 3.3 Nemotron Super 49B V1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and agentic tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.
Nemotron-3-Nano-30B-A3B is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response.
Llama 3.1 Nemotron Ultra 253B V1 is a LLM which is a derivative of Meta Llama-3.1-405B-Instruct. It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.