Benchmark Radar
AI BENCHMARK PROFILE

VAmoS Bench

General AIMultimodal PerceptionVeris AI

VAmoS Bench evaluates complete voice-agent systems in a stateful customer-support task, with 100 scenarios, a simulated caller, real SQL tools, and binary assertions assessed against full interaction traces.

Released
2026-07-29
Readiness
Runnable
Primary field
General AI

Why it matters

It measures end-to-end call containment and correct backend mutations, going beyond component metrics to capture task-level correctness in realistic scenarios.

Motivation

Production voice agents span cascaded, speech-to-speech, and hybrid architectures.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.