Benchmark Radar
AI BENCHMARK PROFILE

PeakBench

General AIKnowledge & Reasoning

PeakBench is a benchmark of executable multi-tool workflows for evaluating resource-aware tool invocation in LLM agents. It includes execution-grounded dependency annotations and measured resource profiles, with a two-part evaluation framework distinguishing logical planning from physical scheduling. Code is available on GitHub.

Released
2026-08-25
Readiness
Runnable
Primary field
General AI

Why it matters

Existing agent benchmarks overlook parallelization and resource-constrained scheduling, creating practical failure modes. PeakBench provides a testbed to diagnose resource-aware agent behavior, helping improve safe and efficient tool execution.

Motivation

LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.