1 paper · 1 filter
Yunqi Zhu, Wensheng Zhang, Xuebing Yang
Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the correctness of final outputs while overlooking the quality of intermed…