Benchmark Radar
AI BENCHMARK PROFILE

CoVEBench

General AIMultimodal PerceptionNJU-LINK Team

CoVEBench evaluates compositional video editing with 416 source videos, 626 multi-point instructions, and 9,990 checklist items, using MLLM judges and objective metrics.

Released
2026-06-07
Readiness
Runnable
Primary field
General AI

Why it matters

Realistic video editing requires handling multiple coupled edits; CoVEBench provides a diagnostic testbed to reveal failures in complex instruction compliance and preservation.

Motivation

While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.