A self-hosted, production-ready OpenJev 27B decision server deployed into your AWS environment with a single click. Because everything runs entirely within your private cloud, your data stays secure, isolated, and fully under your control. Best of all, unlimited usage.
This is a self-hosted deployment of the OpenJev 27B decision server. It runs as a single GPU-powered EC2 instance allowing you to keep your data private and evaluate typed decisions with no per-token charges. Access is via HTTP on port 3000. Once the instance is powered on, the server requires up to 5 minutes to load the model before it is ready to serve requests. Highlights of the OpenJev model include:
Typed decisions about text, web pages, and screenshots. One forward pass per question, with up to 52 options. The answer is a choice, a yes/no probability, or a score, each with a probability for every option. There is no free-form text to parse and no chain of thought.
Labels live in the request. No labelled data and no training run. A new task, label set, or domain is a change to the JSON, not a new model.
Built for agents as well as text. It reads DOM, JSON, and one screenshot per request, picks a target element, picks the next operation, and says when the task is finished. Treat DONE as the model's opinion and confirm completion in your own loop.
Prompts up to 16,384 tokens. On this image, one NVIDIA L40S serves the FP8 weights with up to 160 sequences in flight. The model card's 256-sequence and latency figures were measured on one 80 GB H100.
On a held-out set of 10,000 text questions from 34 public sources, OpenJev scored 84.0% (8,403 of 10,000). Shuffling the option order changed the answer in 2.3% of cases.
The model card lists English, German, French, Hindi, Chinese, and Japanese. Held-out XNLI inference in German, French, Hindi, and Chinese scored 82.5% on 240 questions.
Weights are released under CC BY-NC 4.0: free for research and other non-commercial use, with attribution. Commercial use of the weights requires permission from the OpenJev project. The helper and serve files are Apache 2.0. OpenJev is an independent project and is not affiliated with TypeSafe.
Example API call (replace PUBLIC_IP with the instance address):
curl -s http://PUBLIC_IP:3000/v1/systemone -H 'Content-Type: application/json' -d '{"model":"openjev","state":"Customer message: I was charged twice for my order last week and nobody has replied.","questions":{"route":{"type":"choice","instructions":"Which team should handle this?","criteria":{"billing":null,"shipping":null,"technical":null}},"angry":{"type":"noul","instructions":"Is the customer angry?"}}}'
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 5 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
You pay by the hour for OpenJev 27B running on a single g6e.4xlarge GPU instance. This is usage-based pricing with one dimension, so your cost scales directly with how many hours you keep the instance running. There are no tiers or upfront commitments. You control spend by starting and stopping the instance as needed. The model runs one forward pass per question to return typed decisions, so throughput depends on your workload rather than on separate per-decision charges.
Top-of-mind questions for buyers
What hardware do I get on the g6e.4xlarge instance for this hourly rate?
You get one g6e.4xlarge GPU instance running OpenJev 27B. The model's measured serving recipe fits on a single GPU using FP8 quantization. The bfloat16 checkpoint is about 54 GB and the FP8 checkpoint about 29 GB, so it runs on one GPU of that class.
Am I charged when the instance is stopped or paused?
Hourly software charges accrue only while the instance runs. Stop the instance and the hourly charge stops. You can start and stop it around your workload. Note that stopped instances may still incur underlying AWS storage fees for attached volumes, separate from the software rate.
Does my cost change based on how many questions or decisions the model handles per hour?
No. You pay per instance-hour, not per decision. The model runs one forward pass per question and handles up to 52 options in a single pass. Higher throughput per hour lowers your effective cost per decision, since the hourly rate stays the same regardless of question volume.
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Configured for production environments, please allow up to 5 minutes once the instance is powered on for OpenJev to load the model. The decision API is exposed on port 3000.
Health: curl http://<PUBLIC_IP>:3000/v1/version
Test via HTTP with: curl -s http://<PUBLIC_IP>:3000/v1/systemone -H 'Content-Type: application/json' -d '{"model":"openjev","state":"Customer message: I was charged twice for my order last week and nobody has replied.","questions":{"angry":{"type":"noul","instructions":"Is the customer angry?"}}}'
Additional details
Usage instructions
Deploy the EC2 instance, configure the Security Group to only allow inbound port 22 and 3000 from your trusted IP address(es)
After the instance is powered on, allow up to 5 minutes for OpenJev to load the model. GET /v1/version returns 200 when ready.
Access the OpenJev 27B decision API via HTTP on port 3000. POST /v1/systemone accepts {model, state, questions}. vLLM stays on localhost port 8000 and is not exposed.
Test via HTTP with: curl -s http://<PUBLIC_IP>:3000/v1/systemone -H 'Content-Type: application/json' -d '{"model":"openjev","state":"Customer message: I was charged twice for my order last week and nobody has replied.","questions":{"angry":{"type":"noul","instructions":"Is the customer angry?"}}}'
Our team is happy to assist with deployment and configuration issues.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
A self-hosted, production-ready Bonsai 2 27B model deployed into your AWS environment with a single click. Because everything runs entirely within your private cloud, your data stays secure, isolated, and fully under your control. Best of all, unlimited tokens.
OpenSesame offers a marketplace of 60,000+ courses from top publishers, but its value goes far beyond content. With AI-powered tools to identify skill gaps and a proprietary talent growth engine, learners get personalized learning paths tailored to their role and career growth.
Need custom content? OpenSesame's intuitive course creation tool makes it easy to build high-quality, bespoke courses, no instructional design experience required. Plus, courses can be translated into 70+ languages for global teams.
L&D teams trust OpenSesame to scale learning effortlessly and prove the impact of their programs, ensuring employees gain the right skills to succeed.
Cloud Optimized OpenCV delivers a high-performance build of OpenCV, enabling faster computation of core computer vision operations such as resize, adaptive gaussian, contour detection functions. This optimized edition is designed for accelerated computer vision workloads on AWS Graviton and ARM-based environments, helping developers achieve improved efficiency for AI, ML, and image processing applications.
Initial release of MedGemma 27B Multimodal (google/medgemma-27b-it) on AWS Marketplace. This version delivers Google's most capable open medical AI model, supporting both medical image comprehension and clinical text reasoning including FHIR-based EHR data, across radiology, dermatology, histopathology, and ophthalmology modalities.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.