4 papers
FreqKD: Frequency-Decoupled Cross-Modal Knowledge Distillation for Infrared Object Detection
Keval Thaker, Venkatraman Narayanan, Abdalmalek Aburaddaha +1
Transfer learning from large-scale RGB foundation models to infrared (IR) imagery through knowledge distillation (KD) remains challenging due to fundamental differences in image fo…
SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining
Abdalmalek Aburaddaha, Venkatraman Narayanan, Keval Thaker +1
Mining hard, safety-critical scenes from driving logs is bottlenecked by the absence of difficulty labels, and no single proxy, collision risk, trajectory ambiguity, or semantic ra…
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
Rahul Ahuja, Mudit Jain, Bala Murali Manoghar Sai Sudhakar +4
Vision foundation models (VFMs) and Bird's Eye View (BEV) representation have advanced visual perception substantially, yet their internal spatial representations assume the rectil…
MambaFusion: Adaptive State-Space Fusion for Multimodal 3D Object Detection
Venkatraman Narayanan, Bala Sai, Rahul Ahuja +3
Reliable 3D object detection is fundamental to autonomous driving, and multimodal fusion algorithms using cameras and LiDAR remain a persistent challenge. Cameras provide dense vis…