AI BENCHMARK PROFILE
SVHalluc
SVHalluc evaluates speech-vision hallucination in audio-visual large language models, focusing on semantic and temporal alignment between speech content and visual signals.
- Released
- 2026-05-31
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Addresses a critical gap in evaluating audio-visual LLMs, as prior benchmarks ignored speech-induced hallucinations. Provides a systematic method to assess model grounding, aiding in model development and selection for multimodal applications.
Motivation
Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.