AI BENCHMARK PROFILE
TELBench
Evaluates span-level error localization in deep-research agent trajectories. TELBench comprises 1,000 instances with annotations of harmful error spans among normal exploration, failed searches, tentative hypotheses, and harmless noise, scored by span-level localization and first-error accuracy.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Final-answer evaluation does not reveal which trajectory steps make answers unreliable. This benchmark enables process-level reliability assessment and comparison of error localization methods for deep-research agents.
Motivation
Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.