12 papers
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
Yitong Qiao, Lei Liu, Yue Shen +4
Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often ope…
TICoder: A Repository-Level Code Generation Framework with Test-Driven Planning and Implementation-Aware Reuse
Siyu Nan, Yaling Luo, Jian Wang +2
Repository-level code generation with Large Language Models (LLMs) remains challenging, primarily due to complex dependencies and limited context windows. Recent approaches adopt r…
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
Zequn Xie, Junjie Wang, Dan Yang +4
Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cost. Driven by accuracy-focused…
WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning
Junjie Wang, Zequn Xie, Dan Yang +9
Deep Research systems based on web agents have shown strong potential in solving complex information-seeking tasks, yet their search efficiency remains underexplored. We observe th…
Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models
Renwei Meng, Bowen Zhang, Jian Wang +4
Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, o…
DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrieval
Duolin Sun, Meixiu Long, Dan Yang +9
Retrieval-augmented generation has achieved strong performance on knowledge-intensive tasks where query-document relevance can be identified through direct lexical or semantic matc…