Benchmark Radar
AI BENCHMARK PROFILE

AdaPlanBench

General AIAgentsJiayuJeff/AdaPlanBench

Evaluates adaptive planning of LLM agents under progressively disclosed world and user constraints across 307 household tasks.

Released
2026-06-04
Readiness
Runnable
Primary field
General AI

Why it matters

Fills the gap in evaluating re-planning under dual constraints with interactive feedback, offering a testbed for reliable adaptation in LLM agents.

Motivation

Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are progressively disclosed through interaction.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.