AI BENCHMARK PROFILE
M3-DuplexBench
M3-DuplexBench evaluates full-duplex spoken dialogue models in multi-turn, multilingual (English and Japanese), multidomain settings, with multiple dialogue context settings and turn-taking analysis.
- Released
- 2026-07-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the lack of fair multi-turn comparisons in full-duplex dialogue systems, enabling analysis across languages, domains, and context settings.
Motivation
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.