Benchmark Radar
AI BENCHMARK PROFILE

E-Bench

General AIKnowledge & Reasoning

E-Bench evaluates multi-step tool-use agents in synthetic state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting. It requires agents to discover hidden information and compose multiple tool calls before changing state, with deterministic grading by database-state diffs.

Released
2026-07-26
Readiness
Paper only
Primary field
General AI

Why it matters

E-Bench addresses the gap in evaluating complex tool-use agents that interact with stateful environments over multiple steps, providing a scalable and controllable alternative to existing benchmarks that often focus on isolated API calls or short trajectories.

Motivation

Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.