QO-Bench
QO-Bench is a diagnostic benchmark for query-operator question answering over typed event tuples. It covers 22,984 news articles, 614 corporate events, and 18 query templates, with 785 questions. Gold answers are deterministically computed and scored by recall via exact match.
- Released
- 2026-06-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
RAG systems may retrieve relevant passages but fail to preserve typed values needed for query operators. QO-Bench exposes this gap and allows operator-level diagnosis, guiding development of retrieval systems that preserve query semantics.
Motivation
Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.