AI BENCHMARK PROFILE
CLIP-CC-Bench
CLIP-CC-Bench evaluates paragraph-level video description quality using 5-hour movie content and expert-written references. It scores 17 VLMs via coarse- and fine-grained semantic matching with five LLM-based embedding judges.
- Released
- 2026-08-05
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It fills the gap in long-form video description benchmarks, providing a reliable framework with public scripts and data for reproducible evaluation.
Motivation
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.