A Hybrid Mamba for Audio-Visual Navigation
arXiv:2607.13110
The paper introduces Samba, a hybrid Mamba-based model for audio‑visual navigation that replaces GRUs with a Mamba State Encoder and adds an Audio Mamba Encoder to better capture global time‑frequency patterns, achieving higher success rates on Matterport3D and Replica while reducing computational cost.
Abstract
Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences. This paper proposes Samba(A Hybrid Mamba for Audio-Visual Navigation). It uses the adaptive selection-enabled Mamba State Encoder (M-SE) to replace conventional GRUs for temporal aggregation, and constructs an Audio Mamba Encoder (AME) to remedy the limitations of convolutional operators in capturing global time-frequency dependencies in spectrograms. Experiments demonstrate that Samba exhibits exceptional generalization performance when facing unheard sound sources and unseen scenes. On the Matterport3D dataset, it improves the navigation success rate (SR) by 11.3\% compared with existing state-of-the-art models, and the performance gain is even more pronounced on the Replica dataset, which features finer scene structures. Such modernized architectural reconstruction unlocks stronger embodied representation capabilities at a lower computational cost, thereby providing a highly robust technical pathway for paradigm evolution in the field of audio-visual navigation.
Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems and Man and Cybernetics 2026 (IEEE SMC 2026)