2 papers
cs.AI2026
Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering
Xupeng Chen, Binbin Shi, Chenqian Le +5
Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich visual content (tables, figures, str…
cs.AI2026
Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
Xupeng Chen, Binbin Shi, Chenqian Le +5
Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a pr…