AI BENCHMARK PROFILE
LakeQuest
LakeQuest evaluates end-to-end question answering over data lakes with 9,846 QA pairs across three domains (AI/ML metadata, retail banking, biomedical drug info), with exact modality-aware evidence pointers, measuring retrieval and cross-modal synthesis.
- Released
- 2026-07-14
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
This benchmark fills the gap in evaluating QA systems on heterogeneous, weakly structured data lakes, exposing failure modes where high-quality retrieval does not guarantee correct reasoning, which is crucial for agentic QA development.
Motivation
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.