Overview
Ollama AI on Ubuntu 24.04 LTS provides a ready-to-use local large language model deployment environment. Built on Ubuntu 24.04.4 LTS with Long Term Support until 2034, this AMI includes Ollama runtime and pre-installed Qwen2.5 7B model optimized for GPU inference.
This AMI is designed for enterprises and developers who need to run AI models privately without sending data to third-party cloud services. All inference happens locally on your AWS GPU instance, ensuring data privacy and compliance requirements.
Key components included:
- Ubuntu 24.04.4 LTS (Long Term Support until 2034)
- NVIDIA Driver 580.x (Tesla T4, V100, A100 support)
- CUDA Toolkit 12.5 (GPU computation framework)
- Ollama 0.5.x (LLM inference runtime, MIT licensed)
- Open WebUI (Web interface for chat, optional)
- Qwen2.5 7B Instruct (Pre-installed, Apache 2.0 licensed)
The system supports multiple deployment scenarios:
- CLI usage: Run models directly via 'ollama run' command
- API server: Ollama serves OpenAI-compatible API on port 11434
- Web UI: Install Open WebUI for browser-based interaction
Recommended use cases:
- Enterprise internal AI assistant
- Privacy-sensitive data processing
- Development and testing of AI applications
- Cost-effective AI inference (pay for instance time only)
Highlights
- Ready to use - Pre-installed Qwen2.5 7B model. Run 'ollama run qwen2.5' immediately after launch without any setup.
- Privacy-first - All AI inference happens locally on your instance. No data leaves your infrastructure. Ideal for sensitive workloads.
- Flexible deployment - Use CLI, OpenAI-compatible API, or optional Web UI. Supports multiple models beyond the pre-installed one.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Cost/hour |
|---|---|
g4dn.xlarge Recommended | $350.00 |
g4dn.2xlarge | $350.00 |
g4dn.4xlarge | $350.00 |
g4dn.8xlarge | $350.00 |
g4dn.12xlarge | $350.00 |
g4dn.16xlarge | $350.00 |
g4dn.metal | $350.00 |
p4d.24xlarge | $350.00 |
p3.2xlarge | $350.00 |
p3.8xlarge | $350.00 |
Vendor refund policy
NO Fefund
Custom pricing options
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
64-bit (x86) Amazon Machine Image (AMI)
Amazon Machine Image (AMI)
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Local LLM deployment AMI built on Ubuntu 24.04 LTS with Ollama runtime and pre-installed Qwen2.5 7B model. Run large language models privately on your own GPU infrastructure without relying on cloud APIs.
Included Software:
- Ubuntu 24.04.4 LTS (Base OS)
- NVIDIA Driver 580.x (GPU driver)
- CUDA Toolkit 12.5 (GPU computation)
- Ollama 0.5.x (LLM runtime, MIT licensed)
- Open WebUI (Web interface, optional, MIT licensed)
- Qwen2.5 7B Instruct (Pre-installed model, Apache 2.0 licensed)
Features:
- One-command model serving via CLI
- Open WebUI for browser-based chat (optional installation)
- GPU-accelerated inference on NVIDIA Tesla T4, V100, A100
- Privacy-first: all data stays on your instance
- API compatible with OpenAI
Security:
- SSH key pair authentication only
- Automatic security updates enabled
- Run Ollama as local user
Optimizations:
- Pre-installed Qwen2.5 7B model ready to use
- CUDA environment variables pre-configured
- Web UI auto-start option
Additional details
Usage instructions
SSH to the instance and login as 'ubuntu' using the key pair specified at launch. For more details on connecting to a Linux instance, see: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AccessingInstancesLinux.html
Once connected, view the built-in README for complete instructions:
- Run: cat ~/ami-readme.md
The README includes:
- How to run the pre-installed Qwen2.5 model
- How to install additional models
- How to enable Open WebUI for browser access
- How to use the OpenAI-compatible API
- GPU verification commands
Quick start:
- Run model: ollama run qwen2.5
- Check GPU: nvidia-smi
- API endpoint: http://localhost:11434
For Open WebUI installation, refer to the README.
Support
Vendor support
If you encounter problems in the process of using the system, please feel free to contact us by email: support@thinkclouds.ai . Thank you!
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.