2 papers
cs.CL2026
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
Jiahui Niu, Huizi Yu, Wenkong Wang +11
Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static, single-turn, or narrowly outcome-based,…
cs.CL2026
Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
Xinxin Lin, Guangxin Dai, Yi Zhong +20
Large language models (LLMs) hold transformative potential for medical decision support yet their application in psychiatry remains constrained by hallucinations and superficial re…