3 papers
cs.CL2025
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
Pengcheng Qiu, Chaoyi Wu, Shuyu Liu +7
Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in pro…
cs.CL2025
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
Shuyu Liu, Ruoxi Wang, Ling Zhang +7
The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinic…
cs.CL2024
PediaBench: A Comprehensive Chinese Pediatric Dataset for Benchmarking Large Language Models
Qian Zhang, Panfeng Chen, Jiali Li +6
The emergence of Large Language Models (LLMs) in the medical domain has stressed a compelling need for standard datasets to evaluate their question-answering (QA) performance. Alth…