AI BENCHMARK PROFILE
VAmoS Bench
VAmoS Bench evaluates complete voice-agent systems in a stateful customer-support task, with 100 scenarios, a simulated caller, real SQL tools, and binary assertions assessed against full interaction traces.
- Released
- 2026-07-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It measures end-to-end call containment and correct backend mutations, going beyond component metrics to capture task-level correctness in realistic scenarios.
Motivation
Production voice agents span cascaded, speech-to-speech, and hybrid architectures.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.