Benchmark Radar
AI BENCHMARK PROFILE

LakeQA

General AIKnowledge & Reasoning

LakeQA evaluates search-centric question answering over a 9.5 TB data lake of Wikipedia and government data. Tasks require multi-hop reasoning across heterogeneous structured and unstructured sources, with expert-annotated answers.

Released
2026-06-09
Readiness
Paper only
Primary field
General AI

Why it matters

Existing QA benchmarks provide explicit evidence or trivial retrieval, missing the challenge of locating and composing evidence in large-scale data lakes. LakeQA fills this gap, supporting development and assessment of agents that can search and reason over massive heterogeneous data.

Motivation

Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.