From the 1 of 4 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Michael Solodko, Steven Gong, Guangwei Yu +3
LakeQuest is a human‑validated benchmark of 9,846 question‑answer pairs for evaluating end‑to‑end retrieval and synthesis over heterogeneous data lakes across AI/ML metadata, retai…
cs.CL2025
AppSelectBench: Application-Level Tool Selection Benchmark
Tianyi Chen, Michael Solodko, Sen Wang +14
Computer Using Agents (CUAs) are increasingly equipped with external tools, enabling them to perform complex and realistic tasks. For CUAs to operate effectively, application selec…