Benchmark Radar
AI BENCHMARK PROFILE

Causal-Plan-Bench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceTHUSI-Lab

Causal-Plan-Bench evaluates embodied planning across four causal dimensions (executability, effects, composition, robustness) using 1,200 instances from 12 tasks, with MCQ and rubric-based scoring.

Released
2026-06-01
Readiness
Runnable
Primary field
Robotics & Autonomous Systems

Why it matters

Targets the gap where benchmarks reward linguistic prediction over physical causal reasoning, offering a diagnostic protocol to assess genuine physical agency and support training for grounded planning.

Motivation

Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.