Benchmark Radar
AI BENCHMARK PROFILE

SIGNPOST-Bench

General AIMultimodal Perceptioninorganicwriter

SIGNPOST-Bench evaluates text-vision conflict resolution in multimodal large language models via a counterfactual benchmark of image variants (Original, Blank, Similar, Random, Adversarial) for visual geolocation. It includes 5,111 groups and 25,555 variants, with metrics for localization error and target-directed shifts.

Released
2026-08-04
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks rarely reveal how MLLMs arbitrate conflicting text and visual evidence. SIGNPOST-Bench provides a controlled framework to measure robustness to conflicting scene text, showing that localization performance degrades substantially under adversarial text edits.

Motivation

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.