Benchmark Radar
AI BENCHMARK PROFILE

Sci-VBench

General AIMultimodal Perception

Sci-VBench evaluates text-to-video generation across 1,253 expert-annotated examples in 60 scientific subjects. It requires temporally rich videos demonstrating scientific reasoning and knowledge-grounded synthesis. Scoring covers four dimensions: prompt grounding, scientific correctness, spatiotemporal consistency, and low-level perceptual fidelity, using a rubric-based protocol with MLLM judges.

Released
2026-08-10
Readiness
Runnable
Primary field
General AI

Why it matters

Existing video generation benchmarks focus on surface realism, leaving scientific and causal correctness unmeasured. Sci-VBench provides a public protocol and open dataset to compare models on knowledge-intensive generation, revealing gaps between visual quality and reliable scientific dynamics.

Motivation

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.