CaST-Bench
CaST-Bench evaluates causal chain-grounded spatio-temporal reasoning in video question answering. It contains 2,066 questions over 1,015 videos with causal chains annotated as temporal segments and bounding-box tracks. The evaluation suite includes metrics for answer correctness and visual evidence grounding.
- Released
- 2026-05-22
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing video QA benchmarks lack fine-grained causal grounding. CaST-Bench provides a protocol to assess whether VLMs can identify and localize causal evidence chains, supporting progress in transparent and reliable video understanding.
Motivation
Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.