3 papers
cs.CV2026
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
Yi Wang, Yinfeng Yu, Bin Ren
Audio-visual embodied navigation aims to enable an agent to autonomously localize and reach a sound source in unseen 3D environments by leveraging auditory cues. The key challenge…
cs.CV2025
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
Xu Zheng, Zihao Dongfang, Lutao Jiang +17
Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend…
cs.AI2025
MLLMs are Deeply Affected by Modality Bias
Xu Zheng, Chenfei Liao, Yuqian Fu +15
Recent advances in Multimodal Large Language Models (MLLMs) have shown promising results in integrating diverse modalities such as texts and images. MLLMs are heavily influenced by…