VoxSumm
VoxSumm evaluates joint summarization and translation of long-form spoken news. It comprises 10,045 BBC article-summary pairs across 24 languages and about 703 hours of speech, with scoring based on summarization and translation quality.
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks treat long-document summarization and speech translation separately, leaving a gap for cross-lingual summarization of spoken content. VoxSumm provides a publicly inspectable resource for developing systems that compress and translate long-form speech, aiding evaluation of multilingual instruction-following and cross-lingual generation.
Motivation
As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.