Listing Thumbnail

    Agentarium: AI Assistant Testing & Evaluation on AWS

     Info
    Sold by: ElementX 
    Automated testing harness for enterprise AI assistants. Run multi-turn scenarios with different personas, capture full transcripts and citations, and automatically score assistant performance against your evaluation criteria.

    Overview

    Open image

    What You're Solving

    You've built (or deployed) an AI assistant to handle real work: customer support, employee onboarding, policy guidance. Now you need to know: is it actually working? Will it handle edge cases? Does it give consistent answers? And can you prove it to stakeholders?

    Manual testing doesn't scale. Point-and-click QA consoles make it hard to run realistic, repeatable scenarios, and export workflows are clunky. You end up with spreadsheets and screenshots instead of evidence.

    What Agentarium Does

    Agentarium is an instrumented lab where AI agents run experiments on other AI agents: an AI testing agent does the science, and a test harness records everything. You define scenarios with personas, context, and success criteria; the testing agent drives the conversations and scores the results; the harness captures the evidence.

    In each test run:

    • The testing agent converses with your assistant through its real API, the same interface your users reach
    • Multi-turn scenarios play out like real interactions: follow-ups, misunderstandings, changes of direction
    • Every exchange is recorded automatically: transcripts with raw payloads, tool-call traces, latency, and citations
    • The testing agent files a pass / partial / fail scorecard on each conversation and flags anything needing a human look
    • Flagged conversations land in a review queue; run notebooks collate findings into exportable reports

    Why It Matters

    Methodology stays flexible. Criteria, personas, and rubrics live in the testing agent's prompts, not in code. Change how you judge success without a code change; update rubrics mid-quarter and rerun scenarios on demand, directly comparable against the baseline.

    Full traceability. Every run is attributed to who triggered it. Tests hit draft environments by default; targeting a live assistant is always an explicit opt-in. Rate limits protect your assistant and keep costs predictable. You own all the data; export and analyse it offline any time.

    Realistic testing. Your assistant doesn't sit in isolation. Agentarium tests it the way real users will, so you see failures on the journey before your users do.

    Real example: a university assistant answered "I want to bring my dog to campus" correctly, with citations. The follow-up "Does my flatmate need to register her assistance dog?" returned "no relevant information", even though the same content had retrieved one turn earlier. The testing agent flagged it on the spot, and the recorded tool trace showed exactly what the retrieval step returned: a retrieval-robustness gap caught before production.

    Agentarium works today with IBM watsonx Orchestrate; its connector layer is designed so other assistant platforms plug in.

    What You Get in the Service

    Silver tier: Setup & Baseline Testing

    • Deploy Agentarium on your AWS infrastructure (ECS, RDS, S3: all standard, all yours)
    • Training on test methodology and scenario design
    • Initial baseline test runs (5 to 10 scenarios across key user journeys)
    • Persona templates and scoring rubric starter kit
    • 30 days of email support, handover documentation, and a user guide

    Gold tier: Full Testing Programme

    • Everything in Silver, plus:
    • Higher-throughput configuration for larger test programmes (rate limits stay in place to protect your assistant)
    • Custom persona workshops and bespoke rubrics aligned to your business outcomes
    • Unlimited test runs for 6 months; you pay only your own AWS and model costs, kept predictable by rate limits
    • Monthly trend reports, comparative analysis, and baseline coaching (interpreting fail-to-pass shifts, spotting regressions early)
    • 90 days ongoing support; we can join your QA workflows or run tests on your schedule

    Integration with AWS

    Agentarium runs in your AWS account: ECS for the harness and web UI, RDS PostgreSQL for transcript storage, S3 for exports and archives, CloudWatch for monitoring and audit logs, and IAM roles scoped to exactly what it needs. It makes outbound calls to your assistant's API; nothing needs inbound exposure beyond your team's access to the dashboard. Transcripts, scores, and reports stay in your account. (Test conversations flow to your assistant's platform and the testing agent's model provider over the same paths your real traffic already uses.)

    Before You Start

    We'll need:

    • AWS account access (we'll work through IAM roles with you)
    • Network access to your assistant endpoints, or API keys / service credentials
    • Model access for the testing agent; Amazon Bedrock works well and keeps everything in your AWS account
    • A 30 to 60 minute workshop to map out your personas and success criteria
    • Ongoing access to someone who can clarify what "correct" looks like for your domain

    Highlights

    • Automated scoring, your rubric: an AI testing agent evaluates every run against criteria you define, so no more manual spreadsheet QA
    • Full evidence capture: every turn, citation, and tool-call trace recorded automatically and exportable for audit and offline analysis
    • Methodology without code changes: refine evaluation criteria by updating prompts, then rerun scenarios on demand against your baseline

    Details

    Sold by

    Delivery method

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Pricing

    Custom pricing options

    Pricing is based on your specific requirements and eligibility. To get a custom quote for your needs, request a private offer.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Resources

    Vendor resources

    Support

    Vendor support

    Support & SLA

    Response time: 24 hours for non-emergency questions; 4 hours for blockers

    Email: You'll be assigned a point of contact; we provide a shared Slack channel for Gold tier

    Monthly check-ins: We review test trends, flag patterns, suggest rubric adjustments

    Training: Initial onboarding plus refreshers if your team rotates

    Software associated with this service