Overview

Product video
Backboard Capsules convert AI from a per-token dependency on outside vendors into an owned asset running on your own AWS GPU infrastructure. A Capsule fuses a leading open model with your company's context - documents, code, tickets, and history - into one endpoint that you own and run inside your AWS account.
HOW A CAPSULE IS MADE
- Prepare. Your data is sorted and cleaned by Backboard's data engine.
- Mint data into the model.
- Compress. BBQuant compression makes the model up to 70%+ smaller with 98-101% quality recovery so large models fit on one GPU and decode up to 2.7x faster.
- Deploy. Run it in your AWS account, inside your VPC, or on an air-gapped network. You call it like any model.
WHY TEAMS CHOOSE CAPSULES OVER RAG AND CONVENTIONAL POST-TRAINING - RAG rents your context; a Capsule owns it. Retrieval stuffs documents into every prompt and still guesses. A Capsule pays for context once, at mint.
- Built for work that can't guess. Regulated, high-stakes, and air-gapped teams need a model that abstains instead of inventing an answer.
- Nothing leaves. Data, weights, and inference stay inside your environment. No calls to an outside API.
- No token fees after build. Inference runs on compute you already pay for, including committed AWS spend.
- New models are a tailwind, not a migration project. When a better open base model ships, your Capsule can be re-minted on it.
PROOF POINTS - Up to 70% compression with 98-101% quality recovery versus full precision (example: 152 GB models compressed to 40-44 GB).
- Up to 2.7x faster decode; 68-74% memory savings; one GPU does the work of two or three.
- 99.95% on LoCoMo, the leading agent-memory benchmark, on a quantized 35B Capsule.
- In one customer engagement on the same data set, a Capsule was built with roughly 21x less compute than a conventional post-train, on 1 GPU instead of 8, in about a week instead of a month (single engagement; illustrative of that use case, not a general benchmark).
WHO IT'S FOR
Enterprises and public sector organizations with sovereignty, IP, or security requirements that rule out sending data to outside model APIs - financial services, energy, mining, healthcare, defence-adjacent, and government - plus ISVs embedding an owned model into their own product, and any team running self-hosted models that want to increase speed and accuracy. DEPLOYMENT AND AWS INTEGRATION Capsules deploy as containers on customer-controlled AWS GPU infrastructure (Amazon ECS or Amazon EKS), keeping data, weights, and inference inside your account. Compression cuts the GPU instance count needed per workload, so you can afford to run more owned models on AWS. Enterprise engagements transact through AWS Marketplace, letting you apply committed AWS spend. Backboard is the full post-pre-training stack: memory, routing, action, compression, and Capsules. A Capsule can run standalone or alongside the Backboard AI Platform (benchmark-leading memory and agentic RAG) and Backboard R-CLI (secure enterprise AI coding).
Highlights
- 1. Own your model outright. Your documents, code, and history are minted into a leading open model, then deployed as one endpoint in your AWS account, VPC, or air-gapped network. Data, weights, and inference never leave your environment, and there are no per-token fees after the build.
- 2. Near-lossless compression built in. Capsules compress up to 70% with 98-101% quality recovery and decode up to 2.7x faster, so large models fit on a single GPU and one GPU does the work of two or three, making it feasible and cost effective to self-host.
- 3. New base models are a tailwind, not a migration. When a better open model ships, your Capsule can be re-minted, keeping your owned model current without rebuilding your stack.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/unit |
|---|---|---|
Capsule Tokens - Tier XS (per 1M tokens) | 1M input + output tokens on an XS Capsule. Tier set by Capsule size and configuration. | $0.02 |
Capsule Tokens - Tier S (per 1M tokens) | 1M input + output tokens on an S Capsule. Tier set by Capsule size and configuration. | $0.25 |
Capsule Tokens - Tier M (per 1M tokens) | 1M input + output tokens on an M Capsule. Tier set by Capsule size and configuration. | $1.00 |
Capsule Tokens - Tier L (per 1M tokens) | 1M input + output tokens on an L Capsule. Tier set by Capsule size and configuration. | $5.00 |
Capsule Tokens - Tier XL (per 1M tokens) | 1M input + output tokens on an XL Capsule. Tier set by Capsule size and configuration. | $20.00 |
Vendor refund policy
Capsule builds are custom engagements; fees for completed minting, compression, and deployment work are non-refundable. If a deployed Capsule fails to meet the acceptance criteria agreed in your order or private offer, contact us within 30 days and we will remediate or refund the affected fees. Refund requests: support@backboard.io . Infrastructure charges from AWS are governed by AWS's own policies.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Software as a Service (SaaS)
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
Resources
Vendor resources
Support
Vendor support
Backboard provides direct support for all Capsule deployments.
Email: support@backboard.io
Capsule engagements include deployment assistance into your AWS account (Amazon ECS/EKS), guidance on GPU instance selection and sizing, and support through minting, compression, and re-minting cycles. Enterprise customers receive a dedicated contact and agreed response times as part of their engagement. For air-gapped deployments, support procedures are agreed during onboarding to match your network constraints.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.