Overview
Overview dashboard
The overview. A week of runs, a live count of what needs review, and suite health. Each experiment's recent run shows as a colour strip, so a regression is visible before you click anything.
Overview dashboard
Every Assistant, catalogued
Organised the way testing really happens
What You're Solving
You've built (or deployed) an AI assistant to handle real work: customer support, employee onboarding, policy guidance. Now you need to know: is it actually working? Will it handle edge cases? Does it give consistent answers? And can you prove it to stakeholders?
Manual testing doesn't scale. Point-and-click QA consoles make it hard to run realistic, repeatable scenarios, and export workflows are clunky. You end up with spreadsheets and screenshots instead of evidence.
What Agentarium Does
Agentarium is an instrumented lab where AI agents run experiments on other AI agents: an AI testing agent does the science, and a test harness records everything. You define scenarios with personas, context, and success criteria; the testing agent drives the conversations and scores the results; the harness captures the evidence.
In each test run:
- The testing agent converses with your assistant through its real API, the same interface your users reach
- Multi-turn scenarios play out like real interactions: follow-ups, misunderstandings, changes of direction
- Every exchange is recorded automatically: transcripts with raw payloads, tool-call traces, latency, and citations
- The testing agent files a pass / partial / fail scorecard on each conversation and flags anything needing a human look
- Flagged conversations land in a review queue; run notebooks collate findings into exportable reports
Why It Matters
Methodology stays flexible. Criteria, personas, and rubrics live in the testing agent's prompts, not in code. Change how you judge success without a code change; update rubrics mid-quarter and rerun scenarios on demand, directly comparable against the baseline.
Full traceability. Every run is attributed to who triggered it. Tests hit draft environments by default; targeting a live assistant is always an explicit opt-in. Rate limits protect your assistant and keep costs predictable. You own all the data; export and analyse it offline any time.
Realistic testing. Your assistant doesn't sit in isolation. Agentarium tests it the way real users will, so you see failures on the journey before your users do.
Real example: a university assistant answered "I want to bring my dog to campus" correctly, with citations. The follow-up "Does my flatmate need to register her assistance dog?" returned "no relevant information", even though the same content had retrieved one turn earlier. The testing agent flagged it on the spot, and the recorded tool trace showed exactly what the retrieval step returned: a retrieval-robustness gap caught before production.
Agentarium works today with IBM watsonx Orchestrate; its connector layer is designed so other assistant platforms plug in.
What You Get in the Service
Silver tier: Setup & Baseline Testing
- Deploy Agentarium on your AWS infrastructure (ECS, RDS, S3: all standard, all yours)
- Training on test methodology and scenario design
- Initial baseline test runs (5 to 10 scenarios across key user journeys)
- Persona templates and scoring rubric starter kit
- 30 days of email support, handover documentation, and a user guide
Gold tier: Full Testing Programme
- Everything in Silver, plus:
- Higher-throughput configuration for larger test programmes (rate limits stay in place to protect your assistant)
- Custom persona workshops and bespoke rubrics aligned to your business outcomes
- Unlimited test runs for 6 months; you pay only your own AWS and model costs, kept predictable by rate limits
- Monthly trend reports, comparative analysis, and baseline coaching (interpreting fail-to-pass shifts, spotting regressions early)
- 90 days ongoing support; we can join your QA workflows or run tests on your schedule
Integration with AWS
Agentarium runs in your AWS account: ECS for the harness and web UI, RDS PostgreSQL for transcript storage, S3 for exports and archives, CloudWatch for monitoring and audit logs, and IAM roles scoped to exactly what it needs. It makes outbound calls to your assistant's API; nothing needs inbound exposure beyond your team's access to the dashboard. Transcripts, scores, and reports stay in your account. (Test conversations flow to your assistant's platform and the testing agent's model provider over the same paths your real traffic already uses.)
Before You Start
We'll need:
- AWS account access (we'll work through IAM roles with you)
- Network access to your assistant endpoints, or API keys / service credentials
- Model access for the testing agent; Amazon Bedrock works well and keeps everything in your AWS account
- A 30 to 60 minute workshop to map out your personas and success criteria
- Ongoing access to someone who can clarify what "correct" looks like for your domain
Highlights
- Automated scoring, your rubric: an AI testing agent evaluates every run against criteria you define, so no more manual spreadsheet QA
- Full evidence capture: every turn, citation, and tool-call trace recorded automatically and exportable for audit and offline analysis
- Methodology without code changes: refine evaluation criteria by updating prompts, then rerun scenarios on demand against your baseline
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Pricing
Custom pricing options
How can we make this page better?
Legal
Content disclaimer
Resources
Vendor resources
Support
Vendor support
Support & SLA
Response time: 24 hours for non-emergency questions; 4 hours for blockers
Email: You'll be assigned a point of contact; we provide a shared Slack channel for Gold tier
Monthly check-ins: We review test trends, flag patterns, suggest rubric adjustments
Training: Initial onboarding plus refreshers if your team rotates
Software associated with this service
