Benchmark Radar
AI BENCHMARK PROFILE

Agent Planning Benchmark

General AISafety & TrustworthinessAPB Team

Agent Planning Benchmark (APB) is a diagnostic benchmark with 4,209 multimodal cases across 22 domains, evaluating planning capabilities in five settings including tool noise and unsolvable tasks.

Released
2026-06-03
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a planning-specific evaluation that isolates failures from execution, enabling targeted improvement of agent planning and refusal behaviors.

Motivation

Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.