3 papers
cs.CV2026
DnA: Denoising Attention for Visual Tasks
Ron Campos, Subhajit Maity, Xin Li +2
The softmax activation in multihead attention (MHA) is the de facto standard for attention-based models in visual perception tasks. However, standard softmax can produce noisy atte…
cs.CV2026
D2-V2X: Depth-Driven Cooperative V2X Reasoning for Autonomous Driving
Kevin Richard, Alphin Varghese, Colin Pham +2
Single-vehicle Vision-Language Models (VLMs) are fundamentally constrained by sensor occlusions. While Vehicle-to-Everything (V2X) systems mitigate this, current benchmarks lack th…
cs.CV2026
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
Ali K. Rahimian, Manish K. Govind, Subhajit Maity +4
Vision Transformers and their variants have achieved remarkable success in diverse visual perception tasks. Despite their effectiveness, they suffer from two significant limitation…