Benchmark Radar
AI BENCHMARK PROFILE

PitchBench

General AIMultimodal Perception

PitchBench evaluates pitch hearing in audio-language models across 28 experiments spanning absolute and relative pitch perception in sequences and chords, varying acoustic conditions and response formats.

Released
2026-05-25
Readiness
Paper only
Primary field
General AI

Why it matters

Pitch perception is foundational for musical reasoning, yet existing benchmarks probe it indirectly. PitchBench provides a systematic, controlled evaluation to identify limitations in current models and support the development of pitch-aware audio-language systems.

Motivation

Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation systems, and music production.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.