activity
20242026
collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Darakshan Rashid, Raza Imam, Ufaq Khan +9

Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition arti…

cs.CV2026

Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs

Raza Imam, Darakshan Rashid, Yutong Xie +3

Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where c…

cs.CV2026

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining

Tayeba Qazi, Ayush Maheshwari, Prerana Mukherjee +1

Thermal imaging offers a powerful alternative to visible-spectrum vision under challenging conditions such as low illumination and adverse weather, yet foundational vision-language…

cs.CV2026

Continual Segmentation under Joint Nonstationarity

Prashant Pandey, Himanshu Kumar, Devineni Sri Venkatraya Chowdary +1

Evolving data streams induce joint nonstationarity in continual semantic segmentation, where semantic classes, input distributions, and supervision availability change simultaneous…

cs.CV2026

Unified Multi-Dataset Training for TBPS

Nilanjana Chatterjee, Sidharatha Garg, A V Subramanyam +1

Text-Based Person Search (TBPS) has seen significant progress with vision-language models (VLMs), yet it remains constrained by limited training data and the fact that VLMs are not…

cs.CV2025

LIB-KD: Teaching Inductive Bias for Efficient Vision Transformer Distillation and Compression

Gousia Habib, Tausifa Jan Saleem, Ishfaq Ahmad Malik +1

With the rapid development of computer vision, Vision Transformers (ViTs) offer the tantalising prospect of unified information processing across visual and textual domains due to…