Instead of relying on external APIs, you can deploy and manage state-of-the-art LLMs directly within your own infrastructure, giving you a competitive edge and peace of mind.
This solution is ideal for businesses that handle sensitive data, require deep customization, or need predictable performance and cost management.
Key Benefits:
Unparalleled Data Security & Privacy: With our self-hosted solution, your data never leaves your environment. All prompts, responses, and proprietary information remain securely within your private cloud or on-premise infrastructure. You have complete control over who can access your data and how it's used, eliminating the risk of third-party data breaches or having your confidential information used to train external models.
Ultimate Flexibility & Customization: Break free from vendor lock-in and API limitations. Our solution allows you to:
Fine-tune models on your specific, proprietary datasets to create a truly unique and powerful AI assistant tailored to your business needs.
Choose from a wide range of open-source models to find the perfect balance between performance and resource requirements.
Integrate seamlessly with your existing enterprise systems, databases, and workflows.
Optimize performance for your specific workloads, controlling latency, throughput, and scalability without being subject to external rate limits or policy changes.
Product Features:
Easy Deployment: Our pre-configured templates, documentation, and support simplify the deployment process, allowing you to get up and running in minutes on popular cloud platforms or on-premise.
Model-Agnostic: Support for a wide variety of leading open-source LLMs, giving you the freedom to choose the best model for your application.
Scalable Architecture: Designed to scale with your business, whether you need to run a single model for a small team or a cluster of models for a large enterprise.
Enterprise-Grade Security: Implement your own security protocols and access controls, including end-to-end encryption, to ensure the highest level of protection for your AI applications.
Stop compromising on security and flexibility. Empower your team to innovate with the full power of AI on your own terms.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour for the AWS EC2 instance type you run this LLM server on. Pricing scales with the compute resources each instance provides. The general-purpose t3.nano and t3.xlarge options suit lighter workloads. The c4.xlarge and c4.2xlarge options add more CPU capacity. The g4ad.xlarge, g4ad.4xlarge, g5.xlarge, g5.4xlarge, and p3.2xlarge options include GPU acceleration for heavier model workloads. Larger instances within each family cost more per hour. You choose the instance type at launch based on your performance and budget needs, then pay only for the hours you run it.
Top-of-mind questions for buyers
What does the hourly rate cover, and am I charged when the instance is stopped?
The rate covers software running on the chosen AWS EC2 instance, billed per hour it runs. Stopped instances stop accruing software charges. You still pay any underlying AWS fees, such as storage for attached volumes, while the instance is stopped.
What do the GPU instance options give me over the CPU-only options?
The g4ad, g5, and p3 options include GPU acceleration, which speeds up running larger language models. The t3 and c4 options rely on CPU only, suiting lighter workloads. You choose based on model size and response speed needs, then pay the hourly rate for that instance type.
If I need more capacity, do I switch instance types or does billing scale automatically?
Billing does not scale automatically. You select one instance type at launch and pay its hourly rate. To get more capacity, you launch a different, larger instance type. There is no automatic tier upgrade; the change is a manual choice you make when configuring the deployment.
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Build for AWS Marketplace
Additional details
Usage instructions
Log in can be executed with the user ubuntu with the selected key file and using port 22. The backend and frontend are executed as services so they will restart automatically between reboots without interventions. For more insights please refer to the full guide at: https://quasiscience.com/resources/private-llm-server-docs
Access our documentation or contact our developers directly with the following links:
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
This product has charges associated with it for the provision and deployment of the application and AMI support. Unlock the full potential of DeepSeek AI by self hosting it on your own infrastructure. This powerful large language model offers advanced reasoning, coding, and mathematical capabilities, making it an ideal solution for AI enthusiasts, developers, and enterprises looking for privacy, speed, and scalability without reliance on third-party providers.
Odysseus is a self-hosted AI workspace that brings together LLM chat, web search, autonomous agents, and MCP tool integrations in a single interface. Deploy on your own AWS infrastructure and keep full control of your data. Launch in minutes, no configuration required.
Generate Enterprise by Iterate is a private, secure agent builder and AI assistant that enables large organizations to analyze, summarize, search, and query documents, spreadsheets, and databases using natural language. Generate Enterprise deploys entirely on secure servers with the ability of running local AI models or connect models running in private cloud environments, ensuring complete data privacy and security compliance. The platform delivers richly formatted answers and insights while maintaining enterprise-grade security with granular access controls by teams, security levels, or individuals. Generate Enterprise combines the power of advanced language models with the security requirements of large enterprises, enabling organizations to unlock AI-driven document intelligence without compromising sensitive data or regulatory compliance.
This product has charges associated with it for seller support. Wekan is an open-source Kanban board application that provides a visual and collaborative way for teams to manage their projects and tasks. With an easy-to-use interface, users can create, move, and track tasks as cards on boards, making it a versatile tool for project management and task tracking.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.