Benchmark Radar
AI BENCHMARK PROFILE

MusTBENCH

General AIKnowledge & Reasoning

MusTBENCH evaluates temporal grounding in Large Audio-Language Models through five temporally grounded question-answering tasks, validated by music experts, to test alignment with audio regions.

Released
2026-05-28
Readiness
Paper only
Primary field
General AI

Why it matters

Temporal grounding is critical for music understanding, where events are localized. MusTBENCH establishes this as a missing capability and offers a challenging benchmark for improvement.

Motivation

Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.