PitchBench
PitchBench evaluates pitch hearing in audio-language models across 28 experiments spanning absolute and relative pitch perception in sequences and chords, varying acoustic conditions and response formats.
- Released
- 2026-05-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Pitch perception is foundational for musical reasoning, yet existing benchmarks probe it indirectly. PitchBench provides a systematic, controlled evaluation to identify limitations in current models and support the development of pitch-aware audio-language systems.
Motivation
Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation systems, and music production.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.