Zheqing Li, Yiying Yang, Jiping Lang +16
Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks primarily depend on exam-style or…