SIGNPOST-Bench
SIGNPOST-Bench evaluates text-vision conflict resolution in multimodal large language models via a counterfactual benchmark of image variants (Original, Blank, Similar, Random, Adversarial) for visual geolocation. It includes 5,111 groups and 25,555 variants, with metrics for localization error and target-directed shifts.
- Released
- 2026-08-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks rarely reveal how MLLMs arbitrate conflicting text and visual evidence. SIGNPOST-Bench provides a controlled framework to measure robustness to conflicting scene text, showing that localization performance degrades substantially under adversarial text edits.
Motivation
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.