From the 1 of 3 linked papers with an AI index.
3 papers
Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh +1
The paper introduces a multi‑task fine‑tuning approach for small vision‑language models to improve visual question answering on GI endoscopy images, adding grounding and descriptio…
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
Itbaan Safwan, Muhammad Annas Shaikh, Muhammad Haaris +2
We present a multi-task framework for the MediaEval Medico 2025 challenge, leveraging a LoRA-tuned Florence-2 model for simultaneous visual question answering (VQA), explanation ge…
Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
Muhammad Annas Shaikh, Hamza Zaman, Arbaz Asif
This paper presents a comprehensive evaluation of nine convolutional neural network architectures for binary classification of horses and motorcycles in the VOC 2008 dataset. We ad…