Benchmark Radar
AI BENCHMARK PROFILE

NL-PDDL-Bench

General AIKnowledge & Reasoningibasicplan

NL-PDDL-Bench evaluates natural-language-to-PDDL specification generation. It consists of multi-domain instances from IPC domains with planner-verified executability, difficulty scaled by object count, and a suite for parseability, solvability, and plan-level consistency.

Released
2026-06-29
Readiness
Runnable
Primary field
General AI

Why it matters

This benchmark addresses the lack of standardized evaluation for LLM-generated planning specifications, with an emphasis on executability and verifiability. It provides a reproducible basis for assessing model reliability in safety-sensitive planning applications, where incorrect formalization can lead to unsafe outcomes.

Motivation

Planning often requires symbolic specifications that are both executable and verifiable.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.