OrchBench
OrchBench evaluates multi-agent orchestration plans in isolation using deterministic simulation. It constructs DAGs from real-world tasks and scores plans on result quality, makespan, and token cost without executing worker agents.
- Released
- 2026-07-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
OrchBench provides a fast, token-efficient evaluation of orchestration plans, decoupling planning quality from worker capabilities and environmental noise. Its simulated scores correlate strongly with real executions, enabling cost-effective comparison and diagnosis of multi-agent planners.
Motivation
Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.