Benchmark Radar
AI BENCHMARK PROFILE

C3-Bench

General AIMultimodal Perception

C3-Bench evaluates context-aware change captioning with 4,996 human-labeled image pairs across 51 real-world contexts, using an LLM-as-Judge framework scoring correctness, specificity, fluency, relevance, and a reversibility metric.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Change captioning performance varies with domain; C3-Bench exposes systematic failures in conventional models and LMMs, providing a comprehensive benchmark to drive generalization and trustworthiness.

Motivation

While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluation frameworks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.