NL-PDDL-Bench
NL-PDDL-Bench evaluates natural-language-to-PDDL specification generation. It consists of multi-domain instances from IPC domains with planner-verified executability, difficulty scaled by object count, and a suite for parseability, solvability, and plan-level consistency.
- Released
- 2026-06-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
This benchmark addresses the lack of standardized evaluation for LLM-generated planning specifications, with an emphasis on executability and verifiability. It provides a reproducible basis for assessing model reliability in safety-sensitive planning applications, where incorrect formalization can lead to unsafe outcomes.
Motivation
Planning often requires symbolic specifications that are both executable and verifiable.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.