papers
Publications (2)
cs.CL2025
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
Pengcheng Qiu, Chaoyi Wu, Shuyu Liu +7
Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in pro…
cs.CL2024
Towards Evaluating and Building Versatile Large Language Models for Medicine
Chaoyi Wu, Pengcheng Qiu, Jinxin Liu +5
In this study, we present MedS-Bench, a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in clinical contexts. Unlike existing benchmark…