AGORA
Agora evaluates agentic document reasoning across eight domain collections of 9,664 authentic workplace documents. It includes 362 questions requiring location of sparse evidence and reconciliation of terminology, units, and time conventions. The benchmark is designed to exceed model context windows, necessitating deliberate exploration.
- Released
- 2026-06-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Agora addresses the gap in evaluating archive-grounded reasoning where agents must navigate large, messy document collections. It provides a challenging and realistic testbed for assessing agentic document search and synthesis capabilities, offering practical insights for deployment in document-intensive domains.
Motivation
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.