7 papers
Mamba: CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification
Mingxiang Cao, Weiying Xie, Xin Zhang +4
Multi-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fu…
E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
Jiaqing Zhang, Mingxiang Cao, Weiying Xie +5
Multimodal image fusion and object detection are crucial for autonomous driving. While current methods have advanced the fusion of texture details and semantic information, their c…
DiffCLIP: Few-shot Language-driven Multimodal Classifier
Jiaqing Zhang, Mingxiang Cao, Xue Yang +2
Visual language models like Contrastive Language-Image Pretraining (CLIP) have shown impressive performance in analyzing natural images with language information. However, these mo…
SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection
Shuhan Dong, Yunsong Li, Weiying Xie +4
Multimodal object detection leverages diverse modal information to enhance the accuracy and robustness of detectors. By learning long-term dependencies, Transformer can effectively…
Multi-scale direction-aware SAR object detection network via global information fusion
Mingxiang Cao, Weiying Xie, Jie Lei +3
Deep learning has driven significant progress in object detection using Synthetic Aperture Radar (SAR) imagery. Existing methods, while achieving promising results, often struggle…
FoRA: Low-Rank Adaptation Model beyond Multimodal Siamese Network
Weiying Xie, Yusi Zhang, Tianlin Hui +3
Multimodal object detection offers a promising prospect to facilitate robust detection in various visual conditions. However, existing two-stream backbone networks are challenged b…