2 papers
cs.CL2026
Benchmarking Hybrid Deep Research Across Database Querying and Web Search
Ruofan Wu, Peiran Xu, Xiaolong Li +9
While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely conf…
cs.AI2026
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
Fan Shu, Yite Wang, Ruofan Wu +4
The fast-growing demands in using Large Language Models (LLMs) to tackle complex multi-step data science tasks create an emergent need for accurate benchmarking. There are two majo…