This product has charges associated with it for seller support. Deploy advanced retrieval-augmented workflows with local or remote LLMs like OpenAI, Deepseek, Qwen, Mistral, complete control, high flexibility, and enterprise-grade security
This is a repackaged open source software product wherein additional charges apply for support by TechLatest.net.
Important: For step by step guide on how to setup this vm , please refer to our Getting Started guide
Deploy a ready-to-use virtual machine powered by Ragflow and Ollama, fully loaded with lead-ing open-source language models and optimized for high-performance GPU inference.
Ragflow is an open-source framework purpose-built for Retrieval-Augmented Generation (RAG) pipelines for deep document understanding. It lets you easily build, manage, and deploy AI systems that combine LLM reasoning with your proprietary or domain-specific data.
Offering features such as:
Deep Document Understanding: Intelligent layout analysis, template-based chunking (for PDFs, tables, resumes, legal docs, etc.), visual chunking, and explainable citations that reduce hallucinations and support traceability Quality-In, Quality-Out: High-fidelity input leads to accurate, grounded outputs, even with large contexts or complex formats
Broad Multimodal Support: Works across diverse sources, including Word, PPT, Excel, images, scanned docs, web pages, structured data
Seamless Pipeline Orchestration: Provides both Workflow and Agentic Workflow, a unified canvas for low-code and prompt-driven logic, simplifying complex orchestration
Deep Research Multi Agent Engine: Built-in template enabling dynamic, iterative ex-ploration of user queries across internal and external sources, using a robust agent hi-erarchy and prompt-engineered decision flows:
2 - Ollama: Local LLM Inference
Ollama allows you to run large language models locally with ease. It is designed for perfor-mance, portability, and low latency, making it perfect for developers and enterprises alike.
Preinstalled and ready to go with GPU acceleration, Ollama on this VM includes the following models:
Deepseek-R1: family of open reasoning models
Qwen 3.5: High-performing general-purpose model
Mistral: Compact and efficient model for reasoning tasks
Gemma 3: Open, lightweight LLM by Google
nomic-embed-text:latest: High-quality semantic embeddings + text use cases
LLaMA 3.3 - optimized for dialogue/chat use cases
mxbai-embed-large Accurate embeddings at scale
phi4 Compact, powerful AI model
3 - NVIDIA GPU Support
Fully configured GPU-ready environment
Harness the power of GPU-accelerated inference to drastically reduce latency and in-crease throughput for LLM tasks
Works seamlessly with Ollama and Ragflow for high-speed GenAI workflows
Use Cases
Deep Research Agents: Autonomously break down research tasks into sub-tasks, re-trieve across multiple sources, and synthesize executive-level reports.
Document Q&A & Knowledge Assistants: Tap into structured data across formats with accurate citation and transparency.
AI Copilots & Knowledge Workers: Leverage visual and text inputs to power multi-modal assistants.
Secure, Scalable RAG Applications: Everything runs within your own cloud environ-ment with full workflow control and observability.
Low-Latency LLM APIs: Direct deployment of Ollama LLMs for high-performance AI endpoints.
Why Choose This VM?
Full Data Control & Security: Everything runs in your isolated cloud environment giv-ing you Full control over your environment and data, Ideal for sensitive workloads, in-ternal documents, and enterprise-grade compliance.
Flexible Model Support: Use your own embeddings, documents, and vector DBs with Ragflow. Comes with preinstalled LLMs (Deepseek-R1, Qwen 2.5, Mistral, Gemma, Llama, LLaVA) and allows you to easily add your own models via Ollama or any Other LLM provider, giving you complete control over what models you use and how you run them.
All-in-One: Everything you need for GenAI development in a single VM
Instant Setup: No need to install anything , spin up and start working
Multimodal Ready: Includes LLaVA for image+text inference
Disclaimer: Other trademarks and trade names may be used in this document to refer to either the entities claiming the marks and/or names or their products and are the property of their respective owners. We disclaim proprietary interest in the marks and names of others.
Highlights
Agentic RAG with Ragflow, Local Ollama LLMs, GPU-Acceleration & Secure AI Workflow Platform
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour for the EC2 instance type you run this software on. Every dimension maps to a specific AWS instance, and pricing scales with the compute, memory, and storage that instance provides. Smaller instances like large and xlarge sizes cost less per hour, while metal and high-core sizes cost more. You can pick CPU-only instances or NVIDIA GPU instances such as g4dn for faster inference. The minimum supported configuration is 4 vCPUs and 16 GB memory. Choose the size that matches your workload, then adjust anytime by switching instance types.
Top-of-mind questions for buyers
What does one hourly unit cover, and what do I get for the default instance type?
Each unit is one running EC2 instance billed per hour. The software runs as a ready-to-use virtual machine. The default is t2.xlarge with 4 vCPUs and 16 GB memory. That 4 vCPU, 16 GB setup is also the minimum supported configuration for running this software.
Am I charged the hourly software fee when the instance is stopped?
The hourly software charge applies only while the instance runs. A stopped instance stops accruing software charges. You may still pay underlying AWS storage fees for the disk that keeps your data while stopped. To pause billing fully, terminate the instance instead of stopping it.
Do I pay separately for the language models and GPU features, or are they included in the hourly rate?
The hourly rate covers the single instance you run. The virtual machine comes pre-installed with the RAG framework, local inference tooling, and several open-source models. No separate charge applies for these. Choosing a GPU-capable instance such as g4dn changes only the hourly instance rate, not a separate software fee.
www.techlatest.net+1
Helpful?
Vendor refund policy
Will be charged for usage, can be canceled anytime and usage fee is non refundable
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Open putty, paste the IP address and browse your private key you downloaded while deploying the VM, by going to SSH- >Auth->Credentials , click on Open.
Login as ubuntu user.
Update the password of ubuntu user using below command :
sudo passwd ubuntu
Once ubuntu user password is set, access the GUI environment using RDP on Windows machine or Remmina on Linux machine.
Copy the Public IP of the VM and paste it in the RDP. Login with ubuntu user and its password.
To access the RAGFlow web interface , open your browser and copy paste the public IP of the VM as https://public_ip_of_vm
Create your first admin account on the registration page by clicking Sign up button.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
DBLab Standard Edition (DBLab SE) delivers fast, cost-effective database cloning, enabling agile teams to create instant, full-size production clones for development, testing, and staging, all managed like code, integrated into CI/CD pipelines, and optimized with transparent pricing.
Cazena's Instant Data Lake on AWS, data storage, security, operations and 24x7 support. Easily load, store and analyze data with SQL, R, Python, Scala; Connect apps & notebooks or run Spark, Impala, Kafka & Kudu. Start quickly with Cazena Instant Data Lake with Cloudera on AWS.
AllSecure is an on-demand CyberSecurity SaaS platform providing AWS customers with Penetration Testing as a Service and Cloud Security Posture Management. Instantly run comprehensive security assessments to ensure the highest level of protection for your AWS instances. Keywords: cybersecurity, SaaS, on-demand assessments, penetration testing, cloud security posture management, AWS security
PodWatcher automatically identifies non-running objects within your cluster, instantly surfacing the relevant error codes and actionable troubleshooting commands. By providing rapid diagnostics and cross-cloud compatibility, PodWatcher reduces Mean Time to Recovery (MTTR) and simplifies management across diverse EKS and self-managed Kubernetes environments. This product is a buyer-deployed container image, giving you full control within your own VPC.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.