2 papers
cs.CV2026
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
Bonan Ding, Umair Nawaz, Ufaq Khan +5
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…
cs.CV2025
SSLFusion: Scale & Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
Bonan Ding, Jin Xie, Jing Nie +1
Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial i…