RoboSemanticBench
RoboSemanticBench (RSB) is an embodied benchmark evaluating whether vision-language-action models use instruction semantics to select and grasp the correct physical target among candidate blocks in response to math or general-knowledge questions. It includes six suites with four- or ten-choice variants, procedural arithmetic, GSM8K-style, and MMLU-style questions, with diagnostic metrics separating task success from grasp success.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
RSB addresses the evaluation gap of measuring whether VLA models actually ground instruction semantics in action prediction, separating low-level manipulation from semantic selection. It provides a controlled, repeatable protocol with held-out questions, useful for diagnosing model capabilities and guiding improvements in embodied semantic understanding.
Motivation
Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.