Benchmark Radar
AI BENCHMARK PROFILE

MetaRoute-Bench

General AIKnowledge & Reasoning

MetaRoute-Bench is an open framework for comparing meta-decision policies in agentic workflow routing. Contains 180 synthetic task profiles, eight routing policies, and 30 seeds, evaluating success, cost, and latency through 43,200 traces. Metrics include success rate and paired confidence intervals.

Released
2026-07-31
Readiness
Paper only
Primary field
General AI

Why it matters

Meta-decisions in agentic systems are often embedded and evaluated only via aggregate task accuracy. MetaRoute-Bench provides a shared execution model to compare routing policies, enabling analysis of tradeoffs between success, cost, and latency.

Motivation

Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.