AI BENCHMARK PROFILE
AdaPlanBench
Evaluates adaptive planning of LLM agents under progressively disclosed world and user constraints across 307 household tasks.
- Released
- 2026-06-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fills the gap in evaluating re-planning under dual constraints with interactive feedback, offering a testbed for reliable adaptation in LLM agents.
Motivation
Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are progressively disclosed through interaction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.