Benchmark Radar
AI BENCHMARK PROFILE

PolySpeech-100

General AIMultimodal PerceptionPolySpeech-100 Benchmark Team

PolySpeech-100 evaluates speech understanding in speech-large language models across 110 linguistic variants, including 19 Chinese dialects and over 80 low-resource languages, using tasks that assess semantic reasoning beyond transcription.

Released
2026-05-31
Readiness
Runnable
Primary field
General AI

Why it matters

Existing speech benchmarks are biased toward high-resource languages and focus on low-level recognition, limiting assessment of reasoning abilities and dialect robustness. PolySpeech-100 provides a broader coverage and a scoring protocol for comparing model performance on diverse speech understanding tasks.

Motivation

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.