Benchmark Radar
AI BENCHMARK PROFILE

CapRiCorn-1K

General AIMultimodal PerceptionCapRiCorn-1K Team

CapRiCorn-1K evaluates video captioning quality and subject referential consistency across long videos (15s-10min) with audiovisual and visual-only settings. It uses LLM judge to compute accuracy, coverage, and referential consistency metrics based on manual annotations.

Released
2026-06-20
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on short videos and overall caption quality, lacking evaluation of subject referential consistency over long horizons. CapRiCorn-1K's metrics correlate with downstream understanding and generation performance, offering practical value for selecting captioning models.

Motivation

Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.