Benchmark Radar
AI BENCHMARK PROFILE

CallBench

General AIKnowledge & Reasoning

CallBench is a Chinese benchmark for evaluating dual-goal coordination in phone call assistants, with 50,000 multi-turn dialogues across six scenarios and a preset-aware turn-level scoring protocol.

Released
2026-06-22
Readiness
Paper only
Primary field
General AI

Why it matters

Existing dialogue benchmarks focus on single explicit goals, while real phone assistants must balance the owner's preset and the caller's dynamic goal. CallBench provides a reusable evaluation to measure turn-level decisions under proxy constraints.

Motivation

Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.