5 papers
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
Guangzhi Xiong, Qiao Jin, Sanchit Sinha +2
Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clini…
Large Language Models Lack Temporal Awareness of Medical Knowledge
Zihan Guan, Qiao Jin, Guangzhi Xiong +6
The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, while in reality, medical kno…
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
Rezarta Islamaj, Robert Leaman, Joey Chan +13
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabil…
CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
Qingqing Zhu, Qiao Jin, Tejas S. Mathai +10
Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publi…
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
Xiangru Tang, Qiao Jin, Kunlun Zhu +10
AI scientists powered by large language models have demonstrated substantial promise in autonomously conducting experiments and facilitating scientific discoveries across various d…