Sci-VBench
Sci-VBench evaluates text-to-video generation across 1,253 expert-annotated examples in 60 scientific subjects. It requires temporally rich videos demonstrating scientific reasoning and knowledge-grounded synthesis. Scoring covers four dimensions: prompt grounding, scientific correctness, spatiotemporal consistency, and low-level perceptual fidelity, using a rubric-based protocol with MLLM judges.
- Released
- 2026-08-10
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing video generation benchmarks focus on surface realism, leaving scientific and causal correctness unmeasured. Sci-VBench provides a public protocol and open dataset to compare models on knowledge-intensive generation, revealing gaps between visual quality and reliable scientific dynamics.
Motivation
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.