Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
Yuxuan Zhou, Xien Liu, Chenwei Yan +8
Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored.…
cs.CL2024
Reliable and diverse evaluation of LLM medical knowledge mastery
Yuxuan Zhou, Xien Liu, Chen Ning +2
Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing…
cs.CL2024
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
Yuxuan Zhou, Xien Liu, Chen Ning +1
Large language models (LLMs) have excelled across domains, also delivering notable performance on the medical evaluation benchmarks, such as MedQA. However, there still exists a si…