From the 1 of 8 linked papers with an AI index.
2 citations · 2 across the 1 of their papers we have counts for
8 papers
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Jiazhen Pan, Bailiang Jian, Paul Hager +19
The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
Weixiang Shen, Chengzhi Shen, Yanzhu Hu +12
Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition. Real clinical workflows impose…
Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases
Jun Li, Mingxuan Liu, Jiazhen Pan +4
Clinical abnormality grounding for rare diseases is often hindered by data scarcity, making supervised fine-tuning impractical and single-pass inference highly unstable. We propose…
Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve
Weixiang Shen, Bailiang Jian, Jun Li +6
Tool-augmented large language model (LLM) agents can orchestrate specialist classifiers, segmentation models, and visual question-answering modules to interpret chest X-rays. Howev…
Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
Che Liu, Yinda Chen, Haoyuan Shi +21
The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundati…
Towards Cardiac MRI Foundation Models: Comprehensive Visual-Tabular Representations for Whole-Heart Assessment and Beyond
Yundi Zhang, Paul Hager, Che Liu +4
Cardiac magnetic resonance imaging is the gold standard for non-invasive cardiac assessment, offering rich spatio-temporal views of the cardiac anatomy and physiology. Patient-leve…