Benchmark Radar
AI BENCHMARK PROFILE

CLIP-CC-Bench

General AIMultimodal PerceptionMultimodal Intelligence Lab

CLIP-CC-Bench evaluates paragraph-level video description quality using 5-hour movie content and expert-written references. It scores 17 VLMs via coarse- and fine-grained semantic matching with five LLM-based embedding judges.

Released
2026-08-05
Readiness
Runnable
Primary field
General AI

Why it matters

It fills the gap in long-form video description benchmarks, providing a reliable framework with public scripts and data for reproducible evaluation.

Motivation

Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.