13 papers
ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents
Yuxing Lu, Yushuhong Lin, Wenqi Shi +4
Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to sequential, irreversible dec…
Tree-of-Evidence: Efficient "System 2" Search for Faithful Multimodal Grounding
Micky C. Nnamdi, Benoit L. Marteau, Yishan Zhong +2
Large Multimodal Models (LMMs) achieve state-of-the-art performance in high-stakes domains like healthcare, yet their reasoning remains opaque. Current interpretability methods, su…
RobustMedSAM: Degradation-Resilient Medical Image Segmentation via Robust Foundation Model Adaptation
Jieru Li, Matthew Chen, Micky C. Nnamdi +3
Medical image segmentation models built on Segment Anything Model (SAM) achieve strong performance on clean benchmarks, yet their reliability often degrades under realistic image c…
KindSleep: Knowledge-Informed Diagnosis of Obstructive Sleep Apnea from Oximetry
Micky C Nnamdi, Wenqi Shi, Cheng Wan +4
Obstructive sleep apnea (OSA) is a sleep disorder that affects nearly one billion people globally and significantly elevates cardiovascular risk. Traditional diagnosis through poly…
LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?
J. Ben Tamo, Daniel Carlander-Reuterfelt, Jonathan Rubin +3
Despite multilingual pretraining, large language models often struggle with non-English tasks, particularly in language control, the ability to respond in the intended language. We…
EvidenceRL: Reinforcing Evidence Consistency for Trustworthy Language Models
J. Ben Tamo, Yuxing Lu, Benoit L. Marteau +2
Large Language Models (LLMs) are fluent but prone to hallucinations, producing answers that appear plausible yet are unsupported by available evidence. This failure is especially p…