24 papers
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Jiazhen Pan, Bailiang Jian, Paul Hager +19
The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…
Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation
Yunhang Qian, Jiaquan Yu, Jiawei Liu +3
Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on lan…
The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography
Kaiyuan Yang, Fabio Musio, Yihui Ma +112
The paper introduces the TopCoW Challenge, a benchmark for automatically segmenting the Circle of Willis in CT and MR angiography using deep learning, and provides a new annotated…
Sparse Representation Learning for Vessels
Chinmay Prabhakar, Bastian Wittmann, Paul Büschl +3
Analyzing human vasculature and vessel-like, tubular structures, such as airways, is crucial for disease diagnosis and treatment. Current methods often rely on small sub-regions or…
Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve
Weixiang Shen, Bailiang Jian, Jun Li +6
Tool-augmented large language model (LLM) agents can orchestrate specialist classifiers, segmentation models, and visual question-answering modules to interpret chest X-rays. Howev…
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
Haozhen Gong, Xiaozhong Ji, Yuansen Liu +9
MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Comple…