10 papers
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
Darakshan Rashid, Raza Imam, Ufaq Khan +9
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition arti…
An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering
Mai A. Shaaban, Tausifa Jan Saleem, Alaa Mohamed +3
The paper systematically evaluates continual learning methods for medical visual question answering across a range of heterogeneous clinical tasks, examining catastrophic forgettin…
PolyAlign: Conditional Human-Distribution Alignment
L. D. M. S. Sai Teja, Ufaq Khan, Sathira Silva +2
Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assistant behavior. While effective fo…
MMClima: A Framework for Multimodal Climate Science Data and Evaluation
Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad +2
Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, most…
PathWISE: Multi-Agent Cancer Pathway Triaging Ontology Learning from Clinical Flowcharts
Sofiat Abioye, Ufaq Khan, Shazad Ashraf +6
Clinical pathways are disseminated as visual flowcharts where spatial topology, arrow direction, colour coding, and font weight encode critical triage logic that remains inaccessib…
RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing
Sofiat Abioye, Ufaq Khan, Shazad Ashraf +4
Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual review and transcription. The orig…