Cerebras Inference Cloud on AWS Marketplace brings lightning-fast, on-demand performance to the latest open-source LLMs, including Llama, Qwen, OpenAI GPT-OSS, and more.
Built for real-time interactivity, multi-step reasoning, and complex agentic workflows, Cerebras delivers the speed, scale, and simplicity needed to go from API key to production in under 30 seconds.
Powered by a fast AI accelerator, the Wafer-Scale Engine (WSE), and the CS-3 system, the Cerebras Inference Cloud offers ultra-low latency and high-throughput inferencing via a drop-in, OpenAI-compatible API.
Up to 70X faster than GPUs: With throughput exceeding 2,500 tokens per second, Cerebras eliminates lag, delivering near-instant responses, even from large models.
Full reasoning in under 1 second: No more multi-step delays, Cerebras executes full reasoning chains and delivers final answers in real time.
Instant API access to top open-source models: Skip GPU setup and launch models like Llama, Qwen, OpenAI GPT-OSS in seconds, just bring your prompt.
Access real-time vendor security and compliance information through their Trust Center powered by Drata or Vanta. Review certifications and security standards before purchase.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay through a single usage-based dimension: Cerebras Consumption Units. Your bill reflects actual usage, so cost scales with how much inference you run. There are no fixed tiers or seat licenses here. Consumption Units act as the metering measure that tracks your workload volume across the models you access. As you send more requests, you consume more Units. When usage is low, you consume fewer Units. This structure fits both testing small workloads and scaling to real-time production applications, since spending rises and falls directly with what you use.
Top-of-mind questions for buyers
What is a Cerebras Consumption Unit tied to when I run inference?
A Consumption Unit measures the volume of inference work you run. Usage grows with the number of tokens your requests process across the models you access. Different models consume Units at different rates based on their token throughput and pricing, so heavier or larger workloads draw down more Units.
How does my cost change when my inference volume goes up or down?
Charges track actual usage, so cost rises as you send more requests and falls when activity drops. There is no fixed minimum tied to the Consumption Units dimension. Idle periods with no inference calls consume no Units, so you pay only for the work you run.
Does the type of model I choose affect how many Units I use?
Yes. Each model processes tokens at its own rate and price. Running a model that costs more per million tokens draws down Units faster for the same request volume. Choosing a lower-cost model for a given workload consumes fewer Units per request.
www.cerebras.ai
Helpful?
Vendor refund policy
Payment obligations are non-cancelable once incurred, and Fees paid are non-refundable.
Request a private offer to receive a custom quote.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
Learn more about our supported models, rate limits, pricing, and more in our documentation (docs.cerebras.ai/cloud). For 24x7 technical support, contact support@cerebras.net or +1 (650) 933-4980.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Cerebra offers customized AI-powered Digital Assistants for specific use cases that help operations, production and quality teams manage day-to-day industrial operations and improve their decision making while managing large-scale, complex operations
This product has charges associated with it for support and maintenance. Nengo is an open-source Python library for building and simulating large-scale neural networks and brain-inspired cognitive models.
Some of the most important datasets for image classification research, including
CIFAR 10 and 100, Caltech 101, MNIST, Food-101, Oxford-102-Flowers, Oxford-IIIT-Pets,
and Stanford-Cars. This is part of the fast.ai datasets collection hosted by
AWS for convenience of fast.ai students. See documentation link for citation and
license details for each dataset.