25 papers
ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering
Taojie Zhu, Yuan Xia, Tao Sun +8
Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical question…
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
Yitong Qiao, Lei Liu, Yue Shen +4
Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often ope…
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
Zequn Xie, Junjie Wang, Dan Yang +4
Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cost. Driven by accuracy-focused…
Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information
Renjie Gu, Jiaxu Li, Yihao Wang +8
We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified, yet still continue reasoni…
LATTE: Forecasting Peer Anchored Preference Trajectories for Personalized LLM Generation
Jinze Li, Xiaoyan Yang, Shuo Yang +5
Personalized generation with frozen large language models requires a conditioning signal that is both compact and current. Existing personalization methods typically retrieve or su…
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
Yihao Wang, Haoran Xu, Renjie Gu +10
The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, ex…