Benchmark Radar
AI BENCHMARK PROFILE

Endo-C6

Health & Life SciencesSafety & Trustworthiness

Evaluates temporal vision-language models on surgical endoscopy video understanding under six realistic corruptions, using public videos and standardized prompts.

Released
2026-08-14
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Provides a standardized robustness benchmark for clinical vision-language systems, exposing worst-case performance degradation under clinically relevant artifacts.

Motivation

Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.