VideoVIBE
Evaluates video-grounded diagnostic understanding of one-shot website generation. Approximately 1.7K Video QA instances from 6,338 verified failures across semantic-logical, visual-motion, structural-temporal, and functional categories. Scoring is based on model accuracy in multi-agent system V2Lens compared against baseline Video MLLMs.
- Released
- 2026-08-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks often score isolated artifacts or final outcomes, lacking diagnostic insight. This benchmark provides a repeatable protocol to assess failure modes in generated webpages, enabling targeted model improvements for interactive website generation.
Motivation
Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.