Valar is an inference platform purpose-built for agentic and other high-volume AI workloads. We run leading open-weight models on a throughput-optimized stack to deliver 5-20x lower cost.
Valar is the inference platform for agentic and async workloads. We run leading open-weight models on an optimized stack across GPUs and alternative accelerators, delivering 5-20x lower cost than frontier APIs with better reliability and throughput.
That cost advantage lets teams run longer and more capable background agents - larger models, deeper reasoning loops, broader search, more retries, and stronger validation on every task. The result is smarter long-running AI workflows, higher throughput, better reliability, and dramatically better intelligence per dollar.
Highlights
Run popular open-weight models - Kimi K, GLM, DeepSeek, MiniMax M3, Qwen3.5, gpt-oss and more, behind one OpenAI-compatible API, optimized for the economics of long-running agent and async work rather than the speed of a single call.
Build agentic and compound systems on the primitives you already use: tool calls, structured (JSON-schema) outputs, background/async mode, webhooks, and idempotency across the Responses and Chat Completions APIs.
Pick a completion window per request and trade wall-clock time for price: Standard runs about 50% below competitor serverless rates, Flex about 65%, with automatic prefix caching on every request.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor, and additional usage. You pay upfront or in installments according to your contract terms with the vendor. This entitles you to a specified quantity of use for the contract duration. Usage-based pricing is in effect for overages or additional usage not covered in the contract. These charges are applied on top of the contract price. If you choose not to renew or replace your contract before the contract end date, access to your entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
This contract listing bills through two usage-based dimensions measured in units. Additional Usage covers consumption beyond your committed amount, so you pay for extra inference work as it happens. Shared Deployment supports the Enterprise setup, where you run workloads on shared infrastructure managed through a single API. The two dimensions work together: your Enterprise deployment defines how work runs, while Additional Usage scales cost with the volume of inference you process. Both meter units of work rather than fixed seats or servers, so charges rise as your job volume grows.
Top-of-mind questions for buyers
What counts as one unit for billing under these dimensions?
A unit measures inference work processed by the service. The service optimizes cost per completed task rather than per request, so units track the volume of model calls and jobs you run. Charges rise with the amount of inference work you process, not with fixed seats or servers.
What happens to my cost when I exceed my committed amount?
Work beyond your committed amount bills through the Additional Usage dimension. You pay for the extra inference as it happens, so consumption above your commitment continues without interruption. Cost scales with the added volume of units you process past your commitment.
How do the Shared Deployment and Additional Usage charges combine on my bill?
Shared Deployment supports the Enterprise setup, where you run workloads on shared infrastructure through a single API. Additional Usage meters inference beyond your committed amount. Both bill in units and appear together. Additional Usage grows your bill as job volume rises past your commitment.
valarhq.ai
Helpful?
Vendor refund policy
All fees are non-cancellable and non-refundable except as required by law.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
CompactifAI API empowers organizations with ultra-efficient and scalable AI models that slash compute and energy costs, accelerate deployment, and fuel innovation, all without compromising performance or reliability!
Trieve Vector Inference is an in-VPC solution for fast, unmetered embedding vector inference.
Get fastest-in-class embeddings using any private, custom, or open-source models from dedicated embedding servers hosted in your own cloud.
GigaOps GPU Monitor on DeepSeek-R1 with WebUI. Includes GPU utilization tracking, model health checks, automatic restart on failure, and memory alerts. Run AI inference workloads with built-in observability.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.