Benchmark Radar
AI BENCHMARK PROFILE

PlanBench-XL

General AIAgentsPlanBench-XL Team

PlanBench-XL evaluates LLM tool-use agents on long-horizon planning in large-scale tool ecosystems. It includes 327 retail tasks across 1,665 tools, requiring iterative tool retrieval and use. Optional blocker mechanisms inject missing, failing, or distracting tools to test adaptive planning. Scoring is based on final answer accuracy and auxiliary metrics.

Released
2026-06-21
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks often assume full tool visibility, which underrepresents real-world agent deployment. PlanBench-XL fills this gap by testing planning under retrieval-limited visibility and tool failures, providing insight into robustness and adaptability of LLM agents in complex tool environments.

Motivation

LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.