MedClawBench
MedClawBench evaluates long-horizon temporal reasoning in surgical videos through 1,123 doctor-grounded questions over self-built long neurosurgery recordings and a public lecture-video test split, with fixed evaluation dimensions for comparison.
- Released
- 2026-08-14
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing VLM benchmarks fail to capture temporal dependencies in long surgical videos. MedClawBench provides a reproducible, doctor-grounded dataset to assess video reasoning capabilities beyond short clips, aiding progress in surgical AI systems.
Motivation
Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.