Benchmark Radar
AI BENCHMARK PROFILE

SVHalluc

General AIMultimodal PerceptionChenshuang Zhang et al.

SVHalluc evaluates speech-vision hallucination in audio-visual large language models, focusing on semantic and temporal alignment between speech content and visual signals.

Released
2026-05-31
Readiness
Inspectable
Primary field
General AI

Why it matters

Addresses a critical gap in evaluating audio-visual LLMs, as prior benchmarks ignored speech-induced hallucinations. Provides a systematic method to assess model grounding, aiding in model development and selection for multimodal applications.

Motivation

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.