PlanBench-XL
PlanBench-XL evaluates LLM tool-use agents on long-horizon planning in large-scale tool ecosystems. It includes 327 retail tasks across 1,665 tools, requiring iterative tool retrieval and use. Optional blocker mechanisms inject missing, failing, or distracting tools to test adaptive planning. Scoring is based on final answer accuracy and auxiliary metrics.
- Released
- 2026-06-21
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks often assume full tool visibility, which underrepresents real-world agent deployment. PlanBench-XL fills this gap by testing planning under retrieval-limited visibility and tool failures, providing insight into robustness and adaptability of LLM agents in complex tool environments.
Motivation
LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.