10 papers
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
Mainak Singha, Tanisha Gupta, Ankit Jha +3
Pretrained biomedical vision-language models (VLMs) such as BioMedCLIP perform well on average but often degrade on challenging modalities where inter-class margins are small and a…
bi-modal textual prompt learning for vision-language models in remote sensing
Pankhi Kashyap, Mainak Singha, Biplab Banerjee
Prompt learning (PL) has emerged as an effective strategy to adapt vision-language models (VLMs), such as CLIP, for downstream tasks under limited supervision. While PL has demonst…
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
Aditya Chaudhary, Sneha Barman, Mainak Singha +3
In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (H…
SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation
Prachet Dev Singh, Shyamsundar Paramasivam, Sneha Barman +4
Hyperspectral image (HSI) classification presents unique challenges due to its high spectral dimensionality and limited labeled data. Traditional deep learning models often suffer…
Reconstruction Guided Few-shot Network For Remote Sensing Image Classification
Mohit Jaiswal, Naman Jain, Shivani Pathak +4
Few-shot remote sensing image classification is challenging due to limited labeled samples and high variability in land-cover types. We propose a reconstruction-guided few-shot net…
Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features
Yuzhen Hu, Biplab Banerjee, Saurabh Prasad
Hyperspectral imaging (HSI) enables detailed land cover classification, yet low spatial resolution and sparse annotations pose significant challenges. We present a label-efficient…