3 citations · 5 across the 3 of their papers we have counts for
4 papers
CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
Qingqing Zhu, Qiao Jin, Tejas S. Mathai +10
Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publi…
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine
Yifan Yang, Qiao Jin, Robert Leaman +15
The remarkable capabilities of Large Language Models (LLMs) make them increasingly compelling for adoption in real-world healthcare applications. However, the risks associated with…
Entry-level guide to the use of large language models for medical research
Qiao Jin, Nicholas Wan, Robert Leaman +20
Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4, and DeepSeek-R1, represent a transformative class of AI tools capable of revolutionizing variou…
MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
Nikhil Khandekar, Qiao Jin, Guangzhi Xiong +14
As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answeri…