AI BENCHMARK PROFILE
ParaPairAudioBench
ParaPairAudioBench evaluates LALMs as judges for paralinguistic speech across five dimensions with 5,175 audio pairs.
- Released
- 2026-06-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Targets fine-grained paralinguistic distinctions that prior benchmarks overlook, enabling calibration-aware assessment of judge reliability.
Motivation
Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.