Benchmark Radar
AI BENCHMARK PROFILE

AGORA

General AIKnowledge & ReasoningAGORA Benchmark Team

Agora evaluates agentic document reasoning across eight domain collections of 9,664 authentic workplace documents. It includes 362 questions requiring location of sparse evidence and reconciliation of terminology, units, and time conventions. The benchmark is designed to exceed model context windows, necessitating deliberate exploration.

Released
2026-06-23
Readiness
Paper only
Primary field
General AI

Why it matters

Agora addresses the gap in evaluating archive-grounded reasoning where agents must navigate large, messy document collections. It provides a challenging and realistic testbed for assessing agentic document search and synthesis capabilities, offering practical insights for deployment in document-intensive domains.

Motivation

Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.