From the 1 of 4 linked papers with an AI index.
4 papers
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Michael Solodko, Steven Gong, Guangwei Yu +3
LakeQuest is a human‑validated benchmark of 9,846 question‑answer pairs for evaluating end‑to‑end retrieval and synthesis over heterogeneous data lakes across AI/ML metadata, retai…
SynQuE: Estimating Synthetic Dataset Quality Without Annotations
Arthur Chen, Victor Zhong
We introduce and formalize the Synthetic Dataset Quality Estimation (SynQuE) problem: ranking synthetic datasets by their expected real-world task performance using only limited un…
AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
Zijian Chen, Xueguang Ma, Shengyao Zhuang +3
Deep Research agents are rapidly emerging as primary consumers of modern retrieval systems. Unlike human users who issue and refine queries without documenting their intermediate t…
Test-Time Adaptation for LLM Agents via Environment Interaction
Arthur Chen, Zuxin Liu, Jianguo Zhang +6
Large language model (LLM)-based agents struggle to generalize to novel and complex environments, such as unseen websites or new sets of functions, due to a fundamental mismatch be…