Unifying Local and Global Multimodal Features for Place Recognition in Aliased and Low-Texture Environments
arXiv:2403.13395 · doi:10.1109/ICRA57147.2024.10611563
Abstract
Perceptual aliasing and weak textures pose significant challenges to the task of place recognition, hindering the performance of Simultaneous Localization and Mapping (SLAM) systems. This paper presents a novel model, called UMF (standing for Unifying Local and Global Multimodal Features) that 1) leverages multi-modality by cross-attention blocks between vision and LiDAR features, and 2) includes a re-ranking stage that re-orders based on local feature matching the top-k candidates retrieved using a global representation. Our experiments, particularly on sequences captured on a planetary-analogous environment, show that UMF outperforms significantly previous baselines in those challenging aliased environments. Since our work aims to enhance the reliability of SLAM in all situations, we also explore its performance on the widely used RobotCar dataset, for broader applicability. Code and models are available at https://github.com/DLR-RM/UMF
Accepted submission to International Conference on Robotics and Automation (ICRA), 2024
References in corpus (7)
- Bootstrap your own latent: A new approach to self-supervised Learning
- OverlapTransformer: An Efficient and Rotation-Invariant Transformer Network for LiDAR-Based Place Recognition
- Visual Place Recognition: A Tutorial
- Masked Autoencoder for Self-Supervised Pre-training on Lidar Point Clouds
- Challenges of SLAM in extremely unstructured environments: the DLR Planetary Stereo, Solid-State LiDAR, Inertial Dataset
- Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling
- 6D Camera Relocalization in Visually Ambiguous Extreme Environments