AI BENCHMARK PROFILE
SPEARBench
SPEARBench evaluates naturalness in streaming speech-to-speech language models via question-answer interactions, measuring latency, interruptions, speech quality, ASR robustness, language consistency, emotional naturalness, and interpersonal stance.
- Released
- 2026-07-06
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Standard speech metrics miss conversational naturalness. This benchmark provides a multidimensional protocol to assess human-like behavior in spoken interactions, critical for user acceptance.
Motivation
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.