Benchmark Radar
AI BENCHMARK PROFILE

VideoVIBE

General AIMultimodal Perception

Evaluates video-grounded diagnostic understanding of one-shot website generation. Approximately 1.7K Video QA instances from 6,338 verified failures across semantic-logical, visual-motion, structural-temporal, and functional categories. Scoring is based on model accuracy in multi-agent system V2Lens compared against baseline Video MLLMs.

Released
2026-08-10
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks often score isolated artifacts or final outcomes, lacking diagnostic insight. This benchmark provides a repeatable protocol to assess failure modes in generated webpages, enabling targeted model improvements for interactive website generation.

Motivation

Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.