C3-Bench
C3-Bench evaluates context-aware change captioning with 4,996 human-labeled image pairs across 51 real-world contexts, using an LLM-as-Judge framework scoring correctness, specificity, fluency, relevance, and a reversibility metric.
- Released
- 2026-06-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Change captioning performance varies with domain; C3-Bench exposes systematic failures in conventional models and LMMs, providing a comprehensive benchmark to drive generalization and trustworthiness.
Motivation
While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluation frameworks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.