1 citations · 1 across the 1 of their papers we have counts for
8 papers
HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models
Aryan Das, Tanishq Rachamalla, Pravendra Singh +5
We introduce HyperCap, the first large-scale hyperspectral captioning dataset designed to enhance model performance and effectiveness in remote sensing applications. Unlike traditi…
Efficient Text-Guided Convolutional Adapter for the Diffusion Model
Aryan Das, Koushik Biswas, Swalpa Kumar Roy +2
We introduce the Nexus Adapters, novel text-guided efficient adapters to the diffusion-based framework for the Structure Preserving Conditional Generation (SPCG). Recently, structu…
Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
Aryan Das, Tanishq Rachamalla, Koushik Biswas +2
We introduce a novel uncertainty-aware multimodal segmentation framework that leverages both radiological images and associated clinical text for precise medical diagnosis. We prop…
RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
Xi Xiao, Yunbei Zhang, Janet Wang +9
Accurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understandi…
SceneMixer: Exploring Convolutional Mixing Networks for Remote Sensing Scene Classification
Mohammed Q. Alkhatib, Ali Jamali, Swalpa Kumar Roy
Remote sensing scene classification plays a key role in Earth observation by enabling the automatic identification of land use and land cover (LULC) patterns from aerial and satell…
Are Vision xLSTM Embedded UNet More Reliable in Medical 3D Image Segmentation?
Pallabi Dutta, Soham Bose, Swalpa Kumar Roy +1
The development of efficient segmentation strategies for medical images has evolved from its initial dependence on Convolutional Neural Networks (CNNs) to the current investigation…