MetaRoute-Bench
MetaRoute-Bench is an open framework for comparing meta-decision policies in agentic workflow routing. Contains 180 synthetic task profiles, eight routing policies, and 30 seeds, evaluating success, cost, and latency through 43,200 traces. Metrics include success rate and paired confidence intervals.
- Released
- 2026-07-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Meta-decisions in agentic systems are often embedded and evaluated only via aggregate task accuracy. MetaRoute-Bench provides a shared execution model to compare routing policies, enabling analysis of tradeoffs between success, cost, and latency.
Motivation
Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.