3 papers
cs.CL2025
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
Shuyu Liu, Ruoxi Wang, Ling Zhang +7
The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinic…
cs.CL2025
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
Pengcheng Qiu, Chaoyi Wu, Shuyu Liu +7
Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in pro…
cs.CL2025
PediaBench: A Comprehensive Chinese Pediatric Dataset for Benchmarking Large Language Models
Qian Zhang, Panfeng Chen, Jiali Li +6
The emergence of Large Language Models (LLMs) in the medical domain has stressed a compelling need for standard datasets to evaluate their question-answering (QA) performance. Alth…