Benchmark Radar
AI BENCHMARK PROFILE

T1-Bench

General AIKnowledge & Reasoning

T1-Bench evaluates agentic systems through realistic customer-facing, multi-domain tasks. Scenarios involve multi-turn user-assistant interactions across 25 domains with varying difficulty, measuring structured reasoning, tool utilization, and conversational quality. Automatic evaluation is complemented by human judgments.

Released
2026-06-09
Readiness
Paper only
Primary field
General AI

Why it matters

Existing agent benchmarks lack realism and domain diversity, limiting assessment of sustained multi-step reasoning. T1-Bench addresses this gap by providing interleaved, complex scenarios across multiple domains, enabling more accurate evaluation of agents' practical capabilities in real-world customer service settings.

Motivation

Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.