Benchmark Radar
AI BENCHMARK PROFILE

DiCoBench

General AIMultimodal PerceptionPKU-ICST-MIPL

DiCoBench evaluates multimodal LLMs on multi-image fine-grained perception using high-resolution (up to 2K) image pairs. It includes 765 multiple-choice questions across two tracks: differential and commonality visual cues, covering 8 perception tasks, with exact-match scoring.

Released
2026-06-25
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks rely on explicit textual cues or low resolution, whereas DiCoBench targets autonomous discovery of subtle visual cues in high-resolution pairs. It provides a challenging testbed with a large human-model performance gap, aiding progress in complex multi-image understanding.

Motivation

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grained perception capabilities.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.