7 papers
GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence
Maram Hasan, Aman Verma, Savitra Roy +5
Remote-sensing vision-language models (RS-VLMs) have advanced Earth-observation analysis toward visual interpretation and instruction-following, yet fall short of operational geo-i…
ArcGate: Adaptive Arctangent Gated Activation
Avik Bhattacharya, Siddhant Dnyanesh Gole, Subhasis Chaudhuri +2
Activation functions are central to deep networks, influencing non-linearity, feature learning, convergence, and robustness. This paper proposes the Adaptive Arctangent Gated Activ…
CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
Mainak Singha, Sarthak Mehrotra, Paolo Casari +3
Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these models remain fragile under dom…
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing
Maram Hasan, Md Aminur Hossain, Savitra Roy +6
Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limite…
FOCUS: Bridging Fine-Grained Recognition and Open-World Discovery across Domains
Vaibhav Rathore, Divyam Gupta, Moloud Abdar +2
We introduce the first unified framework for *Fine-Grained Domain-Generalized Generalized Category Discovery* (FG-DG-GCD), bringing open-world recognition closer to real-world depl…
Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
Sanchar Palit, Subhasis Chaudhuri, Biplab Banerjee
Diffusion models have recently achieved significant success in various image manipulation tasks, including image super-resolution and perceptual quality enhancement. Pretrained tex…