AI BENCHMARK PROFILE
CoVEBench
CoVEBench evaluates compositional video editing with 416 source videos, 626 multi-point instructions, and 9,990 checklist items, using MLLM judges and objective metrics.
- Released
- 2026-06-07
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Realistic video editing requires handling multiple coupled edits; CoVEBench provides a diagnostic testbed to reveal failures in complex instruction compliance and preservation.
Motivation
While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.