2 papers
cs.CV2026
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Haowen Gu, Gensheng Pei, Junzhu Mao +3
Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features o…
cs.CV2026
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
Haowen Gu, Gensheng Pei, Zeren Sun +4
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational d…