4 papers
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
Jinge Wu, Hongjian Zhou, Mingde Zeng +8
Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness an…
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Hongjian Zhou, Xinyu Zou, Jinge Wu +19
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
Ming Li, Pei Chen, Zhenhao Zhang +10
Large Language Models demonstrate strong capabilities in single-turn instruction following but suffer from Lost-in-Conversation (LiC), a degradation in performance as information i…
Provable Robust Saliency-based Explanations
Chao Chen, Chenghua Guo, Rufeng Chen +5
To foster trust in machine learning models, explanations must be faithful and stable for consistent insights. Existing relevant works rely on the distance for stability as…