1 paper
Zihan Guan, Qiao Jin, Guangzhi Xiong +6
The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, while in reality, medical kno…