Benchmark Radar
AI BENCHMARK PROFILE

OrchBench

General AIKnowledge & Reasoning

OrchBench evaluates multi-agent orchestration plans in isolation using deterministic simulation. It constructs DAGs from real-world tasks and scores plans on result quality, makespan, and token cost without executing worker agents.

Released
2026-07-28
Readiness
Paper only
Primary field
General AI

Why it matters

OrchBench provides a fast, token-efficient evaluation of orchestration plans, decoupling planning quality from worker capabilities and environmental noise. Its simulated scores correlate strongly with real executions, enabling cost-effective comparison and diagnosis of multi-agent planners.

Motivation

Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.