Benchmark Radar
AI BENCHMARK PROFILE

M3-DuplexBench

General AIKnowledge & Reasoning

M3-DuplexBench evaluates full-duplex spoken dialogue models in multi-turn, multilingual (English and Japanese), multidomain settings, with multiple dialogue context settings and turn-taking analysis.

Released
2026-07-31
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the lack of fair multi-turn comparisons in full-duplex dialogue systems, enabling analysis across languages, domains, and context settings.

Motivation

Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.