Jamba 1.5 Large is the first of its kind hybrid Mamba-Transformer architecture at a production grade level offering unmatched efficiency. With an unprecedented context window length (256K), it offers superior quality output for tasks needing large input context & low latency, at a competitive price point for its class.
AI21 Labs Jamba 1.5 Large is a foundation model built from groundbreaking hybrid architecture, leveraging both the novel Mamba architecture and traditional Transformer architecture to achieve leading quality at the best price.
By drawing on its SSM-Transformer hybrid architecture, as well as its impressive 256K context window, Jamba 1.5 Large efficiently solves a variety of text generation and comprehension use cases for the enterprise. Its 94B active parameters and 398B total parameters lead to superior accuracy in responses.
Jamba 1.5 Large is ideal for enterprise workflows with tasks that are data-heavy and require a model to be able to ingest a large amount of information in order to produce an accurate and thorough response, such as summarizing lengthy documents or enabling question answering across an extensive organizational knowledge base. Jamba 1.5 Large is a model designed for superior quality responses, high throughput and attractive price compared to other models in its size class.
Use cases:
Multi-document analysis: Summarize or compare across multiple documents at once to identify key points and insights. For example, summarizing 8 years worth of 10-K financial reports of a public company.
Multi-document question answering: Query multiple documents, records, or policies in a database. For example, powering a customer-facing chatbot and 15 thorough Q&A exchanges referencing 50 pages of support articles per question.
Organizational search assistant: Improve the retrieval stage of a RAG system for organizational data, resulting in higher quality answers. For example, an internal agent assistant can hold up to 800 pages of organizational data in its context window to provide accurate and thorough responses.
Highlights
At least 2.5X faster than competitors in its size category.
Supports 256K context window, the largest openly available.
Deliver competitive quality performance, ranking amongst the top models in its size class.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay based on how much you use this model, not a fixed subscription. Two separate dimensions drive your cost. The first charges per 1 million input tokens, which covers the text you send to the model. The second charges per 1 million output tokens, which covers the text the model generates back. Your total depends on both volumes combined. Because pricing scales with tokens, longer prompts and longer responses raise costs, while shorter interactions cost less. This structure lets you match spending to actual usage rather than committing to a set term or quantity.
Top-of-mind questions for buyers
What counts as a token for input and output billing?
A token is a chunk of text, roughly a few characters or part of a word. Input tokens cover the text you send to the model, including prompts and any documents. Output tokens cover the text the model generates back. You are charged per 1 million of each.
Which dimension usually drives most of my cost — input or output tokens?
It depends on your workload. Both charges bill independently and appear together on the same invoice. Long-document tasks with brief replies lean toward input token cost. Tasks that generate lengthy responses lean toward output token cost. The model handles a 256K context window, so large inputs are common.
Do I pay when the model is idle or not processing requests?
No. Charges apply only when you send requests. There is no fixed subscription or commitment. You pay for the tokens processed, so idle time carries no software cost. Your bill scales directly with how much you use the model.
www.ai21.com
Helpful?
Vendor refund policy
No Refunds
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
We'd love to hear any thoughts, feedback, or questions directly. Or just email us to tell us about the awesome app you're building with our models: support-amazon@ai21.com
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
OpenAI GPT-5.6 Sol brings OpenAI flagship GPT-5.6 model for advanced reasoning, coding, scientific research, cybersecurity, and agentic workflows to Amazon Bedrock.
OpenAI GPT-5.5 brings OpenAIs most capable model for complex professional work to coding, analysis, software operation,and long-running agentic tasks through OpenAI APIs and Amazon Bedrock.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.