1 paper
Xiuyuan Chen, Tao Sun, Dexin Su +37
Current benchmarks for AI clinician systems, often based on multiple-choice exams or manual rubrics, fail to capture the depth, robustness, and safety required for real-world clini…