Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
arXiv:2607.27122
The paper introduces a multi‑task fine‑tuning approach for small vision‑language models to improve visual question answering on GI endoscopy images, adding grounding and description tasks using existing polyp masks and weak Grad‑CAM supervision to boost accuracy and alignment.
Abstract
Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must answer free-form clinical questions about an image. While recent vision-language models (VLMs) achieve promising answer accuracy on this task, clinical adoption also requires the model's internal representations to reflect the visual evidence behind its answers. We propose a simple multi-task fine-tuning recipe that constructs auxiliary grounding and description tasks from an existing VQA dataset with minimal additional annotation: expert-annotated polyp masks are reused directly, while a GI-domain pretrained classifier with Grad-CAM localization provides weak supervision for finding categories that lack ground-truth masks. Three small VLM backbones are fine-tuned with low-rank adaptation under matched VQA-only and multi-task recipes on Kvasir-VQA-x1, and we show consistent accuracy gains together with improved implicit alignment between answer tokens and the relevant image region, evaluated on both in-distribution and out-of-distribution data.
Accepted at EMA4MICCAI 2026 (Workshop on Efficient Medical AI, MICCAI 2026)