AI BENCHMARK PROFILE
SurgVLA-Bench
Evaluates vision-language-action models in laparoscopic surgical robotics across 8 tasks (atomic, conditional, composite) on the SurRoL simulator, using action accuracy and semantic consistency metrics.
- Released
- 2026-06-28
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Fills the lack of standardized surgical VLA benchmarks, enabling comparison of autoregressive vs. flow-matching models and identifying physical bottlenecks like limited field of view.
Motivation
Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.