Benchmark Radar
AI BENCHMARK PROFILE

LakeQuest

Health & Life SciencesFinance & EconomicsKnowledge & Reasoning

LakeQuest evaluates end-to-end question answering over data lakes with 9,846 QA pairs across three domains (AI/ML metadata, retail banking, biomedical drug info), with exact modality-aware evidence pointers, measuring retrieval and cross-modal synthesis.

Released
2026-07-14
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

This benchmark fills the gap in evaluating QA systems on heterogeneous, weakly structured data lakes, exposing failure modes where high-quality retrieval does not guarantee correct reasoning, which is crucial for agentic QA development.

Motivation

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.