1 citations · 1 across the 11 of their papers we have counts for
11 papers
HealMed: Multilingual Evaluation of Large Language Models in Medicine
Yingjian Chen, Fan Gao, Sherry T. Tong +42
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn…
Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
Fan Gao, Sherry T. Tong, Jiwoong Sohn +11
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local langua…
Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs
Pasin Buakhaw, Kun Kerdthaisong, Phuree Phenhiran +4
The emergence of large language models (LLMs) has opened new opportunities for creating dynamic non-player characters (NPCs) in gaming environments, enabling both functional task e…
Beyond One World: Benchmarking Super Heros in Role-Playing Across Multiversal Contexts
Perapard Ngokpol, Kun Kerdthaisong, Pasin Buakhaw +4
Large language models (LLMs) are increasingly used as role-playing agents, yet their capacity to faithfully and consistently portray version-specific characters -- for example, sup…
MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal Prostate MRI Segmentation
Yovin Yahathugoda, Davide Prezzi, Patricia A. Gutierrez +4
Active Surveillance (AS) is a treatment option for managing low and intermediate-risk prostate cancer (PCa), aiming to avoid overtreatment while monitoring disease progression thro…
On the Robustness of Answer Formats in Medical Reasoning Models
Pittawat Taveekitworachai, Natpatchara Pongjirapat, Krittaphas Chaisutyakorn +3
Medical reasoning models (MRMs) achieve superior performance on medical benchmarks compared to medical LLMs; however, high accuracy alone is insufficient for practical deployment.…