AI BENCHMARK PROFILE
Agent Planning Benchmark
Agent Planning Benchmark (APB) is a diagnostic benchmark with 4,209 multimodal cases across 22 domains, evaluating planning capabilities in five settings including tool noise and unsolvable tasks.
- Released
- 2026-06-03
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a planning-specific evaluation that isolates failures from execution, enabling targeted improvement of agent planning and refusal behaviors.
Motivation
Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.