Benchmark Radar
AI BENCHMARK PROFILE

MedClawBench

General AIMultimodal PerceptionMedClaw team

MedClawBench evaluates long-horizon temporal reasoning in surgical videos through 1,123 doctor-grounded questions over self-built long neurosurgery recordings and a public lecture-video test split, with fixed evaluation dimensions for comparison.

Released
2026-08-14
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing VLM benchmarks fail to capture temporal dependencies in long surgical videos. MedClawBench provides a reproducible, doctor-grounded dataset to assess video reasoning capabilities beyond short clips, aiding progress in surgical AI systems.

Motivation

Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.