4 papers
Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
Quang Bui, Shlok Jaiswal, Samuel Paik-Heintz +14
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is rou…
Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
Jiawei He, Boya Zhang, Hossein Rouhizadeh +6
Large language models (LLMs) in biomedicine face a fundamental conflict between static parameter knowledge and the dynamic nature of clinical evidence. Retrieval-Augmented Generati…
HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
Boya Zhang, Alban Bornet, Rui Yang +2
How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to…
An Evaluation Benchmark for Adverse Drug Event Prediction from Clinical Trial Results
Anthony Yazdani, Alban Bornet, Philipp Khlebnikov +4
Adverse drug events (ADEs) are a major safety issue in clinical trials. Thus, predicting ADEs is key to developing safer medications and enhancing patient outcomes. To support this…