LitTraceQA
LitTraceQA evaluates literature-grounded question answering over scientific papers, requiring systems to return paper IDs, evidence locations, and answers in multiple formats. The public split includes 55 examples with gold annotations.
- Released
- 2026-08-07
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark fills a gap in evaluating verifiable scientific QA, separating retrieval, grounding, and answer accuracy, providing a reusable testbed for systems that produce evidence-backed responses.
Motivation
Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent generation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.