7 papers · 1 filter
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Rui Yang, Weihao Xuan, Yi Lin +23
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive discl…
Toward Global Large Language Models in Medicine
Rui Yang, Huitao Li, Weihao Xuan +47
Despite continuous advances in medical technology, the global distribution of health care resources remains uneven. The development of large language models (LLMs) has transformed…
HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
Boya Zhang, Alban Bornet, Rui Yang +2
How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to…
Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
Rui Yang, Matthew Yu Heng Wong, Huitao Li +13
The rapid growth of medical knowledge and increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) have demonstrated value; however…
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao Xuan, Rui Yang, Heli Qi +29
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingui…
The Evolving Landscape of Generative Large Language Models and Traditional Natural Language Processing in Medicine
Rui Yang, Huitao Li, Matthew Yu Heng Wong +12
Natural language processing (NLP) has been traditionally applied to medicine, and generative large language models (LLMs) have become prominent recently. However, the differences b…