6 papers · 1 filter
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Rui Yang, Weihao Xuan, Yi Lin +23
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive discl…
HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
Boya Zhang, Alban Bornet, Rui Yang +2
How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to…
Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
Rui Yang, Matthew Yu Heng Wong, Huitao Li +13
The rapid growth of medical knowledge and increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) have demonstrated value; however…
Enabling Inclusive Systematic Reviews: Incorporating Preprint Articles with Large Language Model-Driven Evaluations
Rui Yang, Jiayi Tong, Haoyuan Wang +8
Background. Systematic reviews in comparative effectiveness research require timely evidence synthesis. Preprints accelerate knowledge dissemination but vary in quality, posing cha…
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao Xuan, Rui Yang, Heli Qi +29
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingui…
The Evolving Landscape of Generative Large Language Models and Traditional Natural Language Processing in Medicine
Rui Yang, Huitao Li, Matthew Yu Heng Wong +12
Natural language processing (NLP) has been traditionally applied to medicine, and generative large language models (LLMs) have become prominent recently. However, the differences b…