LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
LakeQuest is a human-validated benchmark of 9,846 question-answer pairs that evaluates end-to-end retrieval and synthesis across realistic, heterogeneous data lakes in AI/ML, retail banking, and biomedical domains. Baseline results show that even strong retrieval systems often fail at multi-step reasoning, policy grounding, and cross-modal evidence composition, revealing key limitations of current RAG and agentic QA methods.