3 papers
cs.RO2026
Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations
Han Bao, Bingyi Xia, Hanjing Ye +5
Robot crowd navigation requires the ability to infer human intentions while accounting for the structural constraints of the environment. Currently, deep reinforcement learning (DR…
cs.CV2026
Tango: Taming Visual Signals for Efficient Video Large Language Models
Shukang Yin, Sirui Zhao, Hanchao Wang +4
Token pruning has emerged as a mainstream approach for developing efficient Video Large Language Models (Video LLMs). This work revisits and advances the two predominant token-prun…
cs.RO2025
SONAR: Semantic-Object Navigation with Aggregated Reasoning through a Cross-Modal Inference Paradigm
Yao Wang, Zhirui Sun, Wenzheng Chi +3
Understanding human instructions and accomplishing Vision-Language Navigation tasks in unknown environments is essential for robots. However, existing modular approaches heavily re…