OpenAI GPT-6 Luna is the most efficient model from OpenAI for focused, high-volume tasks. It brings GPT-6 intelligence to repeatable work with a clear goal, giving developers a practical way to serve many users or process large volumes of documents and requests.
Luna is well suited for summarizing documents, extracting key information, and powering assistants that answer focused questions using business data returned by application tools. It improves on GPT-5.6 Luna in factual reliability, communication, and selected coding and workflow evaluations. Developers can tune reasoning effort to balance the quality, speed, and cost their workload requires.
Through Amazon Bedrock, organizations can use GPT-6 Luna in the AWS environment where their applications and workflows already run, alongside the AWS security, governance, procurement, and billing processes they already use.
Highlights
Scale focused, repeatable tasks such as document summarization, information extraction, and question answering grounded in business data.
Bring improved factual reliability and clearer communication to everyday applications, with adjustable reasoning effort to balance quality, speed, and cost.
Access GPT-6 Luna through Amazon Bedrock to bring the most efficient model from OpenAI to high-volume applications in your existing AWS environment.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay only for what you use, measured per token. Charges split by token role: input, output, cache reads, and 30-minute cache writes. Each role is priced separately, so writing to cache costs differently than reading from it. Three service modes affect rate: Standard, Priority, and Flex. Long-context dimensions apply when your requests use extended context windows. A parallel set of global dimensions covers global routing. Not every combination exists—Priority applies to short-context roles only, while Flex and Standard extend across long-context roles. Total cost depends on token volume, role, mode, and context length.
Top-of-mind questions for buyers
What counts as one billed unit for these token dimensions?
Each unit is one token, a small piece of text the model processes. Input tokens cover text you send. Output tokens cover text the model returns. Cache read and cache write tokens meter reuse of stored context. You are charged per token consumed in each role.
How do the different token charges combine into one bill?
Each token role bills independently and adds together. A single request can generate input, output, cache read, and cache write charges at the same time. Output tokens usually cost more per token than input. Long-context requests shift charges to the long-context dimensions. Your total sums every role, mode, and context type used.
What is the difference between the Standard, Priority, and Flex service modes?
The three modes set the per-token rate for the same work. Faster processing carries a higher rate than standard handling; lower-cost handling trades speed for price. Priority applies to short-context roles only. Flex and Standard extend across long-context roles. You select the mode that fits each request's speed needs.
openai.com
Helpful?
Vendor refund policy
No refunds, all sales are final.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
OpenAI GPT-5.6 Terra brings a balanced GPT-5.6 model for everyday work, software engineering, knowledge workflows, and scalable AI applications to Amazon Bedrock.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.