AI BENCHMARK PROFILE
Causal-Plan-Bench
Causal-Plan-Bench evaluates embodied planning across four causal dimensions (executability, effects, composition, robustness) using 1,200 instances from 12 tasks, with MCQ and rubric-based scoring.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Targets the gap where benchmarks reward linguistic prediction over physical causal reasoning, offering a diagnostic protocol to assess genuine physical agency and support training for grounded planning.
Motivation
Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.