9 papers
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
Mainak Singha, Tanisha Gupta, Ankit Jha +3
Pretrained biomedical vision-language models (VLMs) such as BioMedCLIP perform well on average but often degrade on challenging modalities where inter-class margins are small and a…
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
Aditya Chaudhary, Sneha Barman, Mainak Singha +3
In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (H…
SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation
Prachet Dev Singh, Shyamsundar Paramasivam, Sneha Barman +4
Hyperspectral image (HSI) classification presents unique challenges due to its high spectral dimensionality and limited labeled data. Traditional deep learning models often suffer…
Reconstruction Guided Few-shot Network For Remote Sensing Image Classification
Mohit Jaiswal, Naman Jain, Shivani Pathak +4
Few-shot remote sensing image classification is challenging due to limited labeled samples and high variability in land-cover types. We propose a reconstruction-guided few-shot net…
Two-Stage Vision Transformer for Image Restoration: Colorization Pretraining + Residual Upsampling
Aditya Chaudhary, Prachet Dev Singh, Ankit Jha
In computer vision, Single Image Super-Resolution (SISR) is still a difficult problem. We present ViT-SR, a new technique to improve the performance of a Vision Transformer (ViT) e…
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
Mainak Singha, Subhankar Roy, Sarthak Mehrotra +4
In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network…