1 paper
Zixian Su, Hongkai Zhang, Fan Gao +12
Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on…