From the 1 of 6 linked papers with an AI index.
6 papers
A Hybrid Mamba for Audio-Visual Navigation
Yi Wang, Yinfeng Yu
The paper introduces Samba, a hybrid Mamba-based model for audio‑visual navigation that replaces GRUs with a Mamba State Encoder and adds an Audio Mamba Encoder to better capture g…
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
Jingzhi Huang, Junkai Huang, Wenxuan Song +4
Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories t…
Deconstructing Spatial Complexity: Hierarchical Decomposition for LLM Spatial Reasoning
Yi Wang, Haojie Lu, Zhaofan Zhang +2
LLMs have shown remarkable proficiency in general language understanding and reasoning. However, they consistently underperform in spatial reasoning that severely limits their appl…
AERR-Nav: Adaptive Exploration-Recovery-Reminiscing Strategy for Zero-Shot Object Navigation
Jingzhi Huang, Junkai Huang, Haoyang Yang +2
Zero-Shot Object Navigation (ZSON) in unknown multi-floor environments presents a significant challenge. Recent methods, mostly based on semantic value greedy waypoint selection, s…
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
Yi Wang, Yinfeng Yu, Bin Ren
Audio-visual embodied navigation aims to enable an agent to autonomously localize and reach a sound source in unseen 3D environments by leveraging auditory cues. The key challenge…
Audio-Guided Visual Perception for Audio-Visual Navigation
Yi Wang, Yinfeng Yu, Fuchun Sun +2
Audio-Visual Embodied Navigation aims to enable agents to autonomously navigate to sound sources in unknown 3D environments using auditory cues. While current AVN methods excel on…