11 papers
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
Jinge Wu, Hongjian Zhou, Mingde Zeng +8
Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness an…
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
Jiayuan Zhu, Jiazhen Pan, Yuyuan Liu +2
The severe shortage of medical doctors limits access to timely and reliable healthcare, leaving millions underserved. Large language models (LLMs) offer a potential solution but st…
SWE Context Bench: A Benchmark for Context Learning in Coding
Jiayuan Zhu, Junde Wu, Minhao Hu +9
Large language models are increasingly used as coding agents for software engineering tasks. Current benchmarks mainly evaluate whether the agent can correctly solve the request or…
Git Context Controller: Manage the Context of LLM-based Agents like Git
Junde Wu, Minhao Hu, Jiayuan Zhu +4
Large language model (LLM) agents have demonstrated strong capabilities in long-horizon tasks by interleaving reasoning with tool use. However, as these agents scale to complex wor…
RiskAgent: Synergizing Language Models with Validated Tools for Evidence-Based Risk Prediction
Fenglin Liu, Jinge Wu, Hongjian Zhou +9
Large Language Models (LLMs) achieve competitive results compared to human experts in medical examinations. However, it remains a challenge to apply LLMs to complex clinical decisi…
Towards Collective Intelligence: Uncertainty-aware SAM Adaptation for Ambiguous Medical Image Segmentation
Mingzhou Jiang, Jiaying Zhou, Junde Wu +3
Collective intelligence from multiple medical experts consistently surpasses individual expertise in clinical diagnosis, particularly for ambiguous medical image segmentation tasks…