3 papers
cs.CL2026
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
Ziqian Zhang, Xingjian Hu, Yue Huang +8
Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving ad…
cs.CL2025
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
Guangfu Guo, Kai Zhang, Bryan Hoo +4
Medical question-answering (QA) is a critical task for evaluating how effectively large language models (LLMs) encode clinical knowledge and assessing their potential applications…
cs.CL2025
Fact or Guesswork? Evaluating Large Language Models' Medical Knowledge with Structured One-Hop Judgments
Jiaxi Li, Yiwei Wang, Kai Zhang +5
Large language models (LLMs) have been widely adopted in various downstream task domains. However, their abilities to directly recall and apply factual medical knowledge remains un…