3 papers
cs.CV2025
ContrastAlign: Toward Robust BEV Feature Alignment via Contrastive Learning for Multi-Modal 3D Object Detection
Ziying Song, Hongyu Pan, Feiyang Jia +8
In the field of 3D object detection tasks, fusing heterogeneous features from LiDAR and camera sensors into a unified Bird's Eye View (BEV) representation is a widely adopted parad…
cs.CV2025
VoxelNextFusion: A Simple, Unified and Effective Voxel Fusion Framework for Multi-Modal 3D Object Detection
Ziying Song, Guoxin Zhang, Jun Xie +4
LiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existin…
cs.CV2025
FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection
Guoxin Zhang, Ziying Song, Lin Liu +1
Multimodal 3D object detection has garnered considerable interest in autonomous driving. However, multimodal detectors suffer from dimension mismatches that derive from fusing 3D p…