This is a repackaged open source software product wherein additional charges apply for hardening, security configuration, and setup support.
LocalAI is an open-source, drop-in OpenAI-compatible API server delivered as a single statically linked Go binary - chat, embeddings, images, and audio served from your own VM with no GPU and no cloud egress. This Lynxroute build is hardened and ready out of the box: per-instance LOCALAI_API_KEY at first boot, host Nginx with TLS in front of the API, the LocalAI process bound to loopback, UFW firewall pre-configured, and a CIS Level 1 hardened Ubuntu 24.04 LTS base.
MIT license - fully auditable, no vendor lock-in.
This is a repackaged open source software product wherein additional charges apply for hardening, security configuration, and setup support.
WHAT IS LOCALAI
LocalAI is an open-source, OpenAI-compatible inference server delivered as a single statically linked Go binary. It exposes the same REST surface as OpenAI - /v1/chat/completions, /v1/embeddings, /v1/audio/transcriptions, /v1/audio/speech, /v1/images/generations, /v1/models, plus a built-in /models gallery API - so any client written for OpenAI (openai-python, openai-node, LangChain OpenAI provider, LlamaIndex, IDE plugins) talks to it unchanged with one base URL swap. Inference runs entirely on CPU using llama.cpp and Whisper.cpp backends; GGUF model files load from local disk under /var/lib/localai/models and persist across restarts - no external database, no Redis, no cloud egress to OpenAI or Anthropic. Bring your own GGUF model (LLaMA family, Mistral, Qwen, Phi, Gemma) or install one in two clicks from the bundled gallery (mudler/LocalAI gallery index). MIT license, no vendor lock-in.
WHAT THIS AMI ADDS
Security hardening:
Per-instance LOCALAI_API_KEY (24-char random) generated at first boot, enforced by LocalAI for every /v1/, /models, /embeddings, /audio, /images call - never baked into the AMI
The same API key authenticates OpenAI-SDK REST clients (Authorization: Bearer) and the Web UI ("Login with API Token") - one credential, one rotation surface
LocalAI bound to 127.0.0.1:8080 - the API process never listens on a public interface
Host Nginx fronts port 443 with a self-signed certificate; HTTP redirects to HTTPS; security headers (X-Content-Type-Options, X-Frame-Options, Referrer-Policy) applied
UFW firewall pre-configured - only TCP 22, 80, 443 are exposed
fail2ban, AppArmor
CVE scan - every image is scanned for vulnerabilities before release
OS hardening (CIS Level 1):
CIS Ubuntu 24.04 LTS Level 1 benchmark applied via ansible-lockdown
CIS Conformance Report at /etc/lynxroute/cis-report.html
CIS Tailored Profile at /usr/share/doc/lynxroute/CIS_TAILORED_PROFILE.md
Highlights
LocalAI security baked in: per-instance LOCALAI_API_KEY at first boot, LocalAI bound to 127.0.0.1, host Nginx with TLS on :443, all auth enforced by LocalAI for /v1/, /models, /embeddings, /audio, /images.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 5 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
LocalAI - Hardened Self-Hosted OpenAI-Compatible API Server
You pay by the hour based on the EC2 instance size you launch. Five instance options are available. The t3.medium and t3.large run on burstable general-purpose hardware for lighter workloads. The m6i.large, m6i.xlarge, and m6i.2xlarge run on general-purpose hardware, scaling up in CPU and memory as you move to the next size. The hourly rate rises with instance capacity, so you match cost to the compute you need. All options ship as the same hardened, self-hosted API server image. AWS bills you hourly through the Marketplace.
Top-of-mind questions for buyers
Am I charged for the software when an instance is stopped or powered off?
Billing meters running instance-hours. A stopped instance stops accruing the hourly software charge. You may still pay underlying AWS fees, such as storage for the attached volume, while the instance is stopped. Only running hours count toward the software rate.
What do I get with each instance option, and how does capacity differ between them?
Each option maps to one running EC2 instance of the named size. The t3.medium and t3.large use burstable CPUs for lighter loads. The m6i.large, m6i.xlarge, and m6i.2xlarge add CPU and memory at each step. All run the same hardened API server image.
Is this pay-as-you-go, or do I commit to a term upfront?
This is usage-based billing with no upfront commitment. You pay only for the hours each instance runs. Start or stop instances any time, and charges follow actual running time. This suits variable workloads where you want cost tied directly to usage.
lynxroute.com
Helpful?
Vendor refund policy
We do not offer refunds for this product. AWS infrastructure charges (EC2, EBS, data transfer) are billed separately by AWS and are not refundable by us.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
LocalAI v4.9.0
Fix: scheduled log rotation no longer restarts the LocalAI service. Rotation previously signalled the server in a way LocalAI treats as a shutdown request, which dropped every loaded model and every in-flight request once a day. Log files are now rotated in place and the server is not signalled.
LocalAI 4.9.0 on Ubuntu 24.04 LTS - minor update (from 4.8.2). MIT unchanged
Security: authentication is deny-by-default - every HTTP route requires credentials unless it appears in an explicit public registry, closing a bypass in which unprefixed aliases such as /moderations, /models, /backends and /mcp/chat/completions fell outside the old protected-prefix list
Chat context compression, opt-in per model: older complete turns are compressed before inference, preserving system prompts, the newest messages and whole tool-call units
One page per resource - /app/models owns Explore and Installed, /app/backends owns Catalog and Installed; existing /app/manage bookmarks still work
Reversible request-scoped PII pseudonyms, global HTTP admission control with live backend traces, and parallel Hugging Face downloads
Certbot pre-installed - enable a trusted HTTPS certificate with one command: sudo certbot --nginx -d yourdomain.com
Rebuilt on the latest CIS Level 1 hardened Ubuntu 24.04 LTS base
Additional details
Usage instructions
Launch instance (t3.large recommended; t3.medium minimum for the smallest quantized models)
Open Security Group - allow TCP 443 from your IP only
Open https://<PUBLIC_IP>/ in your browser - accept the self-signed certificate warning
Click "Login with API Token", paste the API key from the credentials file
Web UI: Models tab -> pick a model from the gallery -> Install. Or drop a GGUF file into /var/lib/localai/models and run sudo systemctl restart local-ai
REST API (OpenAI-compatible) with curl:
curl -k https://<PUBLIC_IP>/v1/chat/completions
-H "Authorization: Bearer <api-key>"
-H "Content-Type: application/json"
-d '{"model":"<your-model>","messages":[{"role":"user","content":"hi"}]}'
The same API key authenticates Web UI sessions and OpenAI-SDK REST clients (Authorization: Bearer).
Credentials are saved to /root/localai-credentials.txt at first boot.
Models are NOT pre-loaded - install on demand from the gallery or upload your own GGUF files.
Replace the self-signed TLS certificate with a CA-signed certificate for production use:
sudo certbot --nginx -d YOUR_DOMAIN
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
This is a repackaged open source software product wherein additional charges apply for hardening, security configuration, and setup support.
vLLM is a high-throughput OpenAI-compatible inference server for open-source LLMs (Python + PyTorch CPU build), bundled with Open WebUI as a browser chat front end. This Lynxroute build is secured and ready out of the box: a random 32-byte API key generated at first boot, Bearer-token auth enforced on every /v1 request, vLLM and Open WebUI bound to loopback behind Nginx TLS, a default tiny model (facebook/opt-125m) preloaded so the API and chat work immediately, UFW firewall pre-configured, and a CIS Level 1 hardened Ubuntu 24.04 LTS base.
vLLM is Apache-2.0 licensed; Open WebUI uses a source-available license (see Long Description for details).
Preconfigured LocalAI virtual machine for running AI models on your own infrastructure. Unified platform for AI Chat, image and voice generation, LLMs and embeddings with OpenAI API compatibility.
Available in both CPU and GPU configurations.
AISIX Cloud combines a managed control plane with a data plane in your AWS account/VPC. Use one API for Amazon Bedrock, OpenAI, Anthropic, Google Vertex AI, Azure OpenAI, and compatible endpoints; enforce model access, budgets, rate limits, guardrails, and observability without routing prompts through the AISIX control plane.
REQUIRES PRIVATE OFFER
To purchase LiteLLM Enterprise Self-Hosted, please reach out to sales@berri.ai for a Private Offer.
LiteLLM is an OpenAI compatible Proxy Server (LLM Gateway) to call 2,000+ LLM APIs using the OpenAI format Bedrock, Huggingface, VertexAI, TogetherAI, Azure OpenAI, OpenAI, etc. Get started with Opensource LiteLLM here: https://github.com/BerriAI/litellm (40,000+ Github Stars)
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.