LakeQA
LakeQA evaluates search-centric question answering over a 9.5 TB data lake of Wikipedia and government data. Tasks require multi-hop reasoning across heterogeneous structured and unstructured sources, with expert-annotated answers.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing QA benchmarks provide explicit evidence or trivial retrieval, missing the challenge of locating and composing evidence in large-scale data lakes. LakeQA fills this gap, supporting development and assessment of agents that can search and reason over massive heterogeneous data.
Motivation
Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.