PolySpeech-100
PolySpeech-100 evaluates speech understanding in speech-large language models across 110 linguistic variants, including 19 Chinese dialects and over 80 low-resource languages, using tasks that assess semantic reasoning beyond transcription.
- Released
- 2026-05-31
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing speech benchmarks are biased toward high-resource languages and focus on low-level recognition, limiting assessment of reasoning abilities and dialect robustness. PolySpeech-100 provides a broader coverage and a scoring protocol for comparing model performance on diverse speech understanding tasks.
Motivation
While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.