Benchmark Radar
AI BENCHMARK PROFILE

MM-Snowball

CybersecurityMultimodal Perception

MM-Snowball is a benchmark for diagnosing hallucination snowballing in multimodal multi-turn dialogue, with fine-grained analysis. It includes data and code via a project page.

Released
2026-05-30
Readiness
Inspectable
Primary field
Cybersecurity

Why it matters

Addresses the lack of benchmarks for error propagation in long-horizon interactions, and shows existing mitigation methods are ineffective, motivating new approaches.

Motivation

Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by hallucination snowballing: a phenomenon where initial errors amplify across conversational turns, leading to a collapse in coherence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.