Benchmark Radar
AI BENCHMARK PROFILE

RoboProcessBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceRoboProcessBench Project

RoboProcessBench evaluates vision-language models on process-aware understanding in robotic manipulation, with 12 diagnostic question families covering static monitoring and dynamic reasoning over execution traces. The benchmark includes 58k QA pairs across 260 tasks.

Released
2026-06-11
Readiness
Inspectable
Primary field
Robotics & Autonomous Systems

Why it matters

Existing evaluations largely ignore fine-grained process understanding, which is crucial for VLMs used as critics or failure detectors. This benchmark provides a structured way to measure progress in this capability and supports post-training via a dedicated SFT split.

Motivation

Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.