Benchmark Radar
AI BENCHMARK PROFILE

ParaPairAudioBench

General AIMultimodal Perception

ParaPairAudioBench evaluates LALMs as judges for paralinguistic speech across five dimensions with 5,175 audio pairs.

Released
2026-06-23
Readiness
Paper only
Primary field
General AI

Why it matters

Targets fine-grained paralinguistic distinctions that prior benchmarks overlook, enabling calibration-aware assessment of judge reliability.

Motivation

Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.