PeakBench
PeakBench is a benchmark of executable multi-tool workflows for evaluating resource-aware tool invocation in LLM agents. It includes execution-grounded dependency annotations and measured resource profiles, with a two-part evaluation framework distinguishing logical planning from physical scheduling. Code is available on GitHub.
- Released
- 2026-08-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent benchmarks overlook parallelization and resource-constrained scheduling, creating practical failure modes. PeakBench provides a testbed to diagnose resource-aware agent behavior, helping improve safe and efficient tool execution.
Motivation
LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.