12 papers · 1 filter
Simulated patient systems powered by large language model-based AI agents offer potential for transforming medical education
Huizi Yu, Jiayan Zhou, Lingyao Li +22
Background: Simulated patient systems are important in medical education and research, providing safe, integrative training environments and supporting clinical decision making. Ad…
KScope: A Framework for Characterizing the Knowledge Status of Language Models
Yuxin Xiao, Shan Chen, Jack Gallifant +3
Characterizing a large language model's (LLM's) knowledge of a given question is challenging. As a result, prior work has primarily examined LLM behavior under knowledge conflicts,…
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
Shan Chen, Pedro Moreira, Yuxin Xiao +6
Large language models (LLMs) are increasingly envisioned as decision-support tools in clinical practice, yet safe clinical reasoning demands integrating heterogeneous knowledge bas…
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
David Restrepo, Chenwei Wu, Zhengxu Tang +14
Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising…
The use of large language models to enhance cancer clinical trial educational materials
Mingye Gao, Aman Varshney, Shan Chen +15
Cancer clinical trials often face challenges in recruitment and engagement due to a lack of participant-facing informational and educational resources. This study investigated the…
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
João Matos, Shan Chen, Siena Placino +13
Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fai…