6 papers · 1 filter
UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion
Aryan Das, Koushik Biswas, Moloud Abdar +1
We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike prior methods that train separa…
FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization
Mohammed Asad Karim, Vinay Kumar Verma
In-context localization (ICL) seeks to localize a target object specified by a small set of support examples in a query image, operating on the fly without training or parameter up…
Efficient Text-Guided Convolutional Adapter for the Diffusion Model
Aryan Das, Koushik Biswas, Swalpa Kumar Roy +2
We introduce the Nexus Adapters, novel text-guided efficient adapters to the diffusion-based framework for the Structure Preserving Conditional Generation (SPCG). Recently, structu…
Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
Aryan Das, Tanishq Rachamalla, Koushik Biswas +2
We introduce a novel uncertainty-aware multimodal segmentation framework that leverages both radiological images and associated clinical text for precise medical diagnosis. We prop…
HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models
Aryan Das, Tanishq Rachamalla, Pravendra Singh +5
We introduce HyperCap, the first large-scale hyperspectral captioning dataset designed to enhance model performance and effectiveness in remote sensing applications. Unlike traditi…
Reliable or Deceptive? Investigating Gated Features for Smooth Visual Explanations in CNNs
Soham Mitra, Atri Sukul, Swalpa Kumar Roy +2
Deep learning models have achieved remarkable success across diverse domains. However, the intricate nature of these models often impedes a clear understanding of their decision-ma…