Send a ticket, log line, document or JSON object plus typed questions and get back a choice, a yes/no probability or a score level, never generated text. Runs inside your AWS account on a SageMaker endpoint, so no text is sent to a third-party API.
Runs inside your own AWS account: your data is processed on a SageMaker
endpoint you control, inside your VPC, and is never sent to a third-party
API.
Ask typed questions about unstructured input and get answers your code can
act on, instead of free text to parse. Send one "state" (text, or any JSON
object or array) and up to 64 questions. Every answer is one of the options
you supplied:
choice: one of your named options, with a probability for each
noul: the probability that a yes/no question is true
score: a level on a rubric you define, with the most likely level and the
expected level
Typical uses are routing support tickets, flagging churn, fraud or abuse
signals, classifying log lines by severity, and scoring reviews. Every
answer carries a confidence, so you can act automatically on confident
answers and send the rest to a person. Questions about one state are
evaluated together, so adding a question is cheap.
The request and response shapes follow the public Jev System One API, so
existing Jev clients work unchanged. Real-time endpoints and SageMaker batch
transform (JSON Lines, one request per line) are supported. Use a GPU
instance (ml.g4dn.xlarge or ml.g5.xlarge) for production speed; CPU
instances (ml.m5.xlarge) suit low volume. Batch transform is validated on
ml.m5.xlarge and ml.g4dn.xlarge.
English-centric. Other languages are not evaluated.
The context window is 4096 tokens. Your questions and options are kept
intact first, and a long state is truncated from the end.
One evaluation runs at a time per instance. Add instances for more
throughput.
Score questions are the least accurate type. Act on "level" and check
"confidence".
Model and training data
Strands Decider 2B from AWS (Apache-2.0), a LoRA fine-tune of Qwen3.5-2B,
served as published. The authors' training data is described in the model
card: https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19.
The container runs in network isolation, so no data leaves your account.
Measured performance
Accuracy of 94 percent on our validation set, the same on GPU and CPU.
One question, median and 95th percentile: ml.g4dn.xlarge 451 ms and
959 ms; ml.g5.xlarge 389 ms and 994 ms; ml.m5.xlarge 2.2 seconds.
Twelve questions in one request: ml.g5.xlarge 0.6 seconds (about 18
questions per second); ml.g4dn.xlarge 1.6 seconds (about 7 per second);
ml.m5.xlarge 14.1 seconds.
Highlights
Typed answers, not generated text: pick one of your options, a yes/no probability or a rubric level, each with calibrated confidence
Runs inside your own AWS account and VPC - your text is never sent to a third-party API
94 percent accuracy on our validation set; about 0.4 s per question on a GPU instance
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour for each instance that runs the model, measured in host hours. Pricing splits across three instance families: ml.m5 (CPU), ml.g4dn, and ml.g5 (GPU). Within each family, you choose a size, either xlarge or 2xlarge. Each option also has a mode: real-time inference serves live requests, while batch inference processes grouped jobs. You pick the instance and mode that match your workload, and billing scales with the hours those instances run. The model deploys in your own AWS account.
Top-of-mind questions for buyers
What does one host hour mean for billing on this model?
A host hour is one hour that a single instance runs your model. Billing counts each running instance by the hour. If you run two instances for one hour, that counts as two host hours. Charges accrue only while an instance is active.
How does real-time inference billing differ from batch inference?
Both bill per host hour on the chosen instance. Real-time mode keeps the instance running to serve live requests, so it bills for continuous uptime. Batch mode runs the instance only while processing grouped jobs, so hours accrue during those job runs. Pick the mode that matches your request pattern.
How do the CPU and GPU instance families affect what I pay?
You choose one instance family: ml.m5 is CPU-based, while ml.g4dn and ml.g5 are GPU-based. Within each, you pick xlarge or 2xlarge sizing. Each instance bills separately by the hour. GPU families suit heavier compute; CPU suits lighter workloads. Your cost follows the family, size, and hours run.
www.sigmodata.com
Helpful?
Vendor refund policy
No refunds offered but you may cancel at any time
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
Version release notes
Initial release.
Strands Decider 2B typed decisions: choice, noul and score
JSON requests with up to 64 questions per state
Real-time endpoints and SageMaker batch transform (JSON Lines)
GPU recommended (ml.g4dn.xlarge, ml.g5.xlarge); runs on GPU with default settings
Additional details
Inputs
Outputs
Usage instructions
Sample notebooks
Inputs
Summary
A JSON object with "state" (text, or any JSON object or array) and "questions" (a map of question name to a typed question). For batch transform, one such object per line (JSON Lines).
Input MIME type
application/json, application/jsonlines
Real-time inference sample input data
{"state": "Ticket #4411: I was charged twice for my March invoice and support has not replied in 3 days. Fix this or I am cancelling.", "questions": {"route": {"type": "choice", "instructions": "Which team should handle this ticket?", "criteria": {"billing": "payments, invoices, refunds", "technical": "bugs, outages, errors", "sales": "new purchases, upgrades"}}, "churn": {"type": "noul", "instructions": "Is the customer threatening to cancel?"}, "urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["low", "medium", "high", "critical"]}}}
Batch transform sample input data
{"state": "Ticket #4411: I was charged twice for my March invoice and support has not replied in 3 days. Fix this or I am cancelling.", "questions": {"route": {"type": "choice", "instructions": "Which team should handle this ticket?", "criteria": {"billing": "payments, invoices, refunds", "technical": "bugs, outages, errors", "sales": "new purchases, upgrades"}}, "churn": {"type": "noul", "instructions": "Is the customer threatening to cancel?"}, "urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["low", "medium", "high", "critical"]}}}
Input data descriptions
The following table describes supported input data fields for real-time inference and batch transform.
Field name
Description
Constraints
Required
state
The input to decide about, as text or any JSON object or array
-
Yes
questions
Map of question name to a typed question (choice, noul or score), up to 64 per request
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
FICO® Decision Modeler enables business users to author, test, and optimize decision logic that delivers their organizations higher profitability, lower development costs, and the flexibility to rapidly iterate and adapt to current and future business climates. (Private Offer Only)
FICO® Platform brings data, analytics, and decisioning together into a single decision intelligence platform, connecting previously siloed experiences into an operating framework that drives value creation across the customer lifecycle.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.