Benchmark Radar
AI BENCHMARK PROFILE

SPEARBench

General AIMultimodal Perception

SPEARBench evaluates naturalness in streaming speech-to-speech language models via question-answer interactions, measuring latency, interruptions, speech quality, ASR robustness, language consistency, emotional naturalness, and interpersonal stance.

Released
2026-07-06
Readiness
Inspectable
Primary field
General AI

Why it matters

Standard speech metrics miss conversational naturalness. This benchmark provides a multidimensional protocol to assess human-like behavior in spoken interactions, critical for user acceptance.

Motivation

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.