E-Bench
E-Bench evaluates multi-step tool-use agents in synthetic state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting. It requires agents to discover hidden information and compose multiple tool calls before changing state, with deterministic grading by database-state diffs.
- Released
- 2026-07-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
E-Bench addresses the gap in evaluating complex tool-use agents that interact with stateful environments over multiple steps, providing a scalable and controllable alternative to existing benchmarks that often focus on isolated API calls or short trajectories.
Motivation
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.