1 paper
Chaoyi Wu, Pengcheng Qiu, Jinxin Liu +5
In this study, we present MedS-Bench, a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in clinical contexts. Unlike existing benchmark…