10 papers
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
Darakshan Rashid, Raza Imam, Ufaq Khan +9
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition arti…
Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs
Raza Imam, Darakshan Rashid, Yutong Xie +3
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where c…
T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining
Tayeba Qazi, Ayush Maheshwari, Prerana Mukherjee +1
Thermal imaging offers a powerful alternative to visible-spectrum vision under challenging conditions such as low illumination and adverse weather, yet foundational vision-language…
Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement
Jyotirmoy Nath, Neeraj Kumar, Brejesh Lall
Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, most existing methods treat each task's prompt as a monolithic, instance-b…
Continual Segmentation under Joint Nonstationarity
Prashant Pandey, Himanshu Kumar, Devineni Sri Venkatraya Chowdary +1
Evolving data streams induce joint nonstationarity in continual semantic segmentation, where semantic classes, input distributions, and supervision availability change simultaneous…
Unified Multi-Dataset Training for TBPS
Nilanjana Chatterjee, Sidharatha Garg, A V Subramanyam +1
Text-Based Person Search (TBPS) has seen significant progress with vision-language models (VLMs), yet it remains constrained by limited training data and the fact that VLMs are not…