11 papers
WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning
Junjie Wang, Zequn Xie, Dan Yang +9
Deep Research systems based on web agents have shown strong potential in solving complex information-seeking tasks, yet their search efficiency remains underexplored. We observe th…
DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrieval
Duolin Sun, Meixiu Long, Dan Yang +9
Retrieval-augmented generation has achieved strong performance on knowledge-intensive tasks where query-document relevance can be identified through direct lexical or semantic matc…
LiveClin: A Live Clinical Benchmark without Leakage
Xidong Wang, Shuqi Guo, Yue Shen +6
The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address the…
ClinAlign: Scaling Healthcare Alignment from Clinician Preference
Shiwei Lyu, Xidong Wang, Lei Liu +6
Although large language models (LLMs) demonstrate expert-level medical knowledge, aligning their open-ended outputs with fine-grained clinician preferences remains challenging. Exi…
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
Peiji Li, Linyang Li, Handa Sun +15
Large language models have demonstrated strong reasoning capabilities in complex tasks through tool integration, which is typically framed as a Markov Decision Process and optimize…
Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs
Yitong Qiao, Licheng Pan, Yu Mi +4
Hallucinations in Large Language Models (LLMs), i.e., the tendency to generate plausible but non-factual content, pose a significant challenge for their reliable deployment in high…