Connect OpenAI, Anthropic, Cohere and Ollama apps to Amazon Bedrock: point them at your own gateway. 100+ models, plus retrieval, batches, images and voice. Runs in your AWS account.
Most AI tools speak only the OpenAI, Anthropic, or Cohere APIs and cannot reach Amazon Bedrock or the AWS AI services. stdapi.ai is an AI gateway you run in your own AWS account that makes them reachable from the tools you already use - Open WebUI, n8n, Claude Code, LangChain and hundreds more. Pointing an application at it is a client-side change, and the model it names comes from the whole catalog, not one vendor's list.
WHY CHOOSE STDAPI.AI
80+ Endpoints Across Four Protocols - Chat, Responses, conversations, embeddings, reranking, vector stores, batches, images, video, audio (speech, transcription, translation, realtime), files, and moderation, on the OpenAI, Anthropic, Cohere and Ollama protocols at once. Standard SDKs connect on the base URL alone.
100+ AI Models, Discovered for You - Anthropic Claude, OpenAI GPT and xAI Grok (via Amazon Bedrock Mantle), Moonshot Kimi, DeepSeek, Amazon Nova, Meta Llama, Alibaba Qwen, Mistral, and Cohere, in a typical multi-region catalog. Bedrock, Bedrock Mantle, Polly, Transcribe and Comprehend surface as models on one endpoint, found automatically as AWS adds them; Amazon Translate backs translation. A retired model ID redirects to its replacement.
Retrieval the Model Runs Itself - Vector stores index your files on Amazon S3 Vectors in your account and search them by meaning, or a store can be an Amazon Bedrock knowledge base you already run. On the Responses API the model runs the searches a turn needs and cites the files it used.
Runs in Your AWS Account - AWS-only by design, not a generic multi-provider proxy. Inference runs on the AWS services and regions you enable, and Amazon Bedrock does not share prompts with model providers or use them for training. Reasoning modes, prompt caching, guardrails, service tiers, inference profiles, prompt routers and batch pricing are reachable through standard API parameters. AWS compliance certifications apply to the services and regions you choose, not to stdapi.ai.
POPULAR USE CASES
Private ChatGPT Alternative and Team Bots - Deploy Open WebUI, LobeHub or LibreChat with multi-modal chat, document RAG and web search, or Q&A bots on Slack, Teams and Telegram.
Voice Assistants - A spoken conversation runs over one WebSocket, with interruption and a live transcript, in sessions up to 8 minutes. Transcription streams phrase by phrase; LiveKit Agents or Pipecat add browser and telephony front ends.
Workflow Automation - Connect AI to hundreds of services via n8n (Make and Zapier via generic HTTP modules) for support, content creation and data processing, with large jobs at Bedrock's batch price.
AI Coding Assistants - Use Claude Code, Cline, OpenCode, OpenAI Codex CLI or JetBrains AI Assistant in your IDE or terminal, backed by Claude, Moonshot Kimi or Qwen Coder.
Autonomous Agents & MCP - 80+ API operations are exposed as Model Context Protocol tools over Streamable HTTP and SSE. OpenClaw, Claude Code, LangGraph and other MCP clients connect with no HTTP client code and discover every tool.
KEY BENEFITS
Regional Quota and Retry - Bedrock quota is per region, so every region you enable adds its own. Eligible throttling and availability failures retry in another enabled region; a streamed call retries only before the stream opens.
Production Infrastructure - Two Terraform or OpenTofu commands deploy ECS Fargate with an HTTPS load balancer, auto-scaling, KMS encryption and private subnets, following the AWS Well-Architected Framework; WAF and CloudWatch alarms are optional.
Authentication Per Caller - Amazon Cognito user pool tokens let each user reach the API with their own credential, instead of or alongside the API key, with no AWS call on the request path; unauthorized responses publish where to authenticate, so an agent finds the authorization server itself. API keys are held in AWS Systems Manager, with CORS and SSRF protection.
Per-User Cost Attribution - Model calls can run under a short-lived role session tagged with the end user, so AWS reports each user's spend in Cost Explorer and the Cost and Usage Report - from the invoice, not an estimate. Off by default.
No Markup, No Lock-in - Amazon Bedrock is billed to you directly by AWS at AWS rates, with no per-seat fee or minimum commitment; private offers cover custom terms. The APIs are the standard ones and the gateway also ships as a free AGPL-3.0 Community edition, so leaving is the base-URL and model-name change that brought you in.
EVIDENCE
AWS Qualified Software. 6,000+ automated tests at 95%+ branch coverage run against real AWS services and the real OpenAI, Anthropic, Cohere and Ollama endpoints. 20 third-party clients are driven end to end against a live gateway. Gateway overhead is under 1 ms.
GET STARTED
Start with the 14-day free trial, then deploy the Terraform module: terraform init and terraform apply. Guides for Open WebUI, n8n, RAG, voice and coding assistants at https://stdapi.ai/?utm_source=aws-marketplace.
Highlights
Serves hundreds of OpenAI, Anthropic, Cohere and Ollama compatible applications - Open WebUI, n8n, Cline, Claude Code, LangChain, and agent frameworks such as OpenClaw, LangGraph, and CrewAI. 80+ endpoints across four protocols cover chat, conversations, vector search, batches, embeddings, images, speech, transcription, realtime voice, files, and moderation, and 80+ operations are exposed as MCP tools. 100+ models including Claude, GPT, Kimi, DeepSeek, Nova, and Qwen are addressed by name.
The gateway runs in your own AWS account, so no third party sits between your users and your models. Inference stays on the AWS services and regions you enable; AWS compliance certifications apply to those services and regions and are not inherited by stdapi.ai. Amazon Cognito can authenticate each caller with their own token, and model calls can be tagged per end user so AWS reports their spend in Cost Explorer. Terraform deploys ECS Fargate with auto-scaling, HTTPS, VPC and KMS encryption.
0% markup on model usage: Amazon Bedrock is billed to you directly by AWS at AWS rates, on the invoice you already receive, so the gateway license is the only charge from us. No per-seat fees, no minimum commitment, and private offers cover custom terms for organizations that need them. A free AGPL-3.0 open-source Community edition of the same gateway is also available, so leaving is the base-URL and model-name change that brought you in.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 14 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
stdapi.ai - OpenAI, Anthropic, Cohere & Ollama AI Gateway for Bedrock
You pay one usage-based rate for each hour a container runs. Billing is metered per container-hour and flows onto your existing AWS invoice. The deployment defaults to one container per Availability Zone, and it scales on CPU load up to five times that count. So your total cost tracks how many containers run and for how long. There is no separate charge here for model usage; you pay AWS Bedrock rates directly for the underlying models. This gateway charge is the only dimension on Marketplace.
Top-of-mind questions for buyers
What counts as one container-hour for billing?
One container-hour is one running container for one hour. The deployment defaults to one container per Availability Zone. Under CPU load, it scales up to five times that count. Your bill sums the hours each running container accrues, so more containers running longer means more container-hours charged.
Does my container-hour charge include the cost of the AI models I call?
No. The container-hour rate covers only the gateway you run. You pay AWS Bedrock rates directly for model usage, with no markup added. Model charges appear on your AWS invoice from AWS, separate from this gateway charge. So two cost streams apply: gateway hours plus underlying model usage.
What happens to my cost when traffic rises and the gateway scales up?
Scaling adds containers, and each running container accrues its own container-hours. Under CPU load, the count can grow to five times the baseline of one per Availability Zone. When load drops and containers stop, those stopped containers no longer accrue charges. Your cost tracks how many containers run and for how long.
stdapi.ai
Helpful?
Vendor refund policy
stdapi.ai offers refunds on a case-by-case basis. We encourage you to try our free tier first to evaluate the product before purchasing.
To request a refund, contact support@stdapi.ai with your AWS Marketplace order ID and reason. Our team will review your request promptly.
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Containers are lightweight, portable execution environments that wrap server application software in a filesystem that includes everything it needs to run. Container applications run on supported container runtimes and orchestration services, such as Amazon Elastic Container Service (Amazon ECS) or Amazon Elastic Kubernetes Service (Amazon EKS). Both eliminate the need for you to install and operate your own container orchestration software by managing and scheduling containers on a scalable cluster of virtual machines.
Ready-to-deploy Terraform examples: https://github.com/stdapi-ai/samples
(single-region production, multi-region EU/GDPR and US, and application stacks.)
OVERVIEW
Hardened container image for production use. Typical deployment is ECS Fargate via the official Terraform module, which creates VPC, ALB with HTTPS, auto-scaling, S3, CloudWatch, IAM and KMS; WAF and CloudWatch alarms are optional, off by default. Manual deployment on ECS is also supported.
Billing is $0.10 per container-hour through AWS Marketplace, with a 14-day free trial and 0% markup on Amazon Bedrock usage, which AWS bills to you directly. The module defaults to one task per Availability Zone, so plan for about $216 per month in license cost in a 3-AZ region. For custom terms, duration, or committed usage, request a private offer at https://stdapi.ai/contact/?utm_source=aws-marketplace.
The Terraform module is published at https://registry.terraform.io/modules/stdapi-ai/stdapi-ai/aws/latest and produces a complete, AWS Well-Architected deployment from a few input variables. Variables cover domain/HTTPS, auto-scaling, allowed Bedrock regions, authentication, WAF, monitoring and existing-VPC integration.
Chat, images, audio, conversations, realtime voice and streamed transcription need no resource beyond the deployment itself and the IAM permissions above. Five optional features stay off until you create the resource they use: batch inference (an S3 bucket and the IAM role Amazon Bedrock assumes for it), vector stores (an Amazon S3 vector bucket), knowledge base vector stores (an allowlist of the knowledge bases this deployment may address), Amazon Cognito authentication (a user pool and its app clients), and per-user cost attribution (a role for the end-user sessions). Long text-to-speech uses the regional bucket the gateway already has, except on generative voices.
API REFERENCE
OpenAI-compatible endpoints (/v1/, including conversations, vector stores, batches and the realtime WebSocket), Anthropic-compatible endpoints (/anthropic/v1/, including message batches), Cohere-compatible endpoints (/cohere/*), and native /search_models for agents:
https://stdapi.ai/api_overview/?utm_source=aws-marketplace
Container runs as non-root with minimal attack surface. Vulnerability scans and prompt patching. Region allow-lists restrict inference to the AWS regions you select; AWS compliance certifications apply to the AWS services and regions you choose and are not inherited by stdapi.ai. CloudWatch audit logs. Security details at https://stdapi.ai/operations_authentication_security/?utm_source=aws-marketplace.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Expert deployment of stdapi.ai, the OpenAI, Anthropic, Cohere and Ollama AI gateway for Amazon Bedrock, into your own AWS account. Production infrastructure configured and secured by specialists.
Eden AI is a unified AI API gateway that helps developers and enterprises access, compare, and orchestrate multiple AI models and AI providers through a single secure API. The platform centralizes access to leading LLMs, generative AI models, computer vision, OCR, speech-to-text, text-to-speech, translation, embeddings, moderation, document parsing, and other AI APIs. Eden AI helps teams simplify AI integrations, reduce vendor lock-in, monitor usage, control AI costs, manage provider fallback, and improve reliability in production. Built for developers, SaaS companies, and enterprise AI teams, Eden AI provides a scalable way to build AI-powered applications, automate workflows, and deploy AI features faster with unified billing, usage analytics, API keys, and provider management.
One-stop API access to Z.ai's full suite of GLM models, covering language, vision, image generation, and video capabilities. A single API key unlocks all AI capabilities for intelligent dialogue, content creation, visual understanding, and multimedia generation.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.