RoboProcessBench
RoboProcessBench evaluates vision-language models on process-aware understanding in robotic manipulation, with 12 diagnostic question families covering static monitoring and dynamic reasoning over execution traces. The benchmark includes 58k QA pairs across 260 tasks.
- Released
- 2026-06-11
- Readiness
- Inspectable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing evaluations largely ignore fine-grained process understanding, which is crucial for VLMs used as critics or failure detectors. This benchmark provides a structured way to measure progress in this capability and supports post-training via a dedicated SFT split.
Motivation
Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.