1 paper
Yuechun Yu, Han Ying, Haoan Jin +5
The reliable evaluation of large language models (LLMs) in medical applications remains an open challenge, particularly in capturing the complexity of multi-turn doctor-patient int…