1 paper
Xiao Li, Mouxiao Bian, Zhaodi Wu +8
Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critica…