4 papers
Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation
Bingqian Lin, Yi Zhu, Xiaodan Liang +2
Vision-Language Navigation (VLN) is a challenging task which requires an agent to align complex visual observations to language instructions to reach the goal position. Most existi…
Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive Learning
Bingqian Lin, Yanxin Long, Yi Zhu +4
Vision-and-language navigation (VLN) asks an agent to follow a given language instruction to navigate through a real 3D environment. Despite significant advances, conventional VLN…
DNA Family: Boosting Weight-Sharing NAS with Block-Wise Supervisions
Guangrun Wang, Changlin Li, Liuchun Yuan +5
Neural Architecture Search (NAS), aiming at automatically designing neural architectures by machines, has been considered a key step toward automatic machine learning. One notable…
AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis
Tao Tang, Guangrun Wang, Yixing Lao +5
Neural implicit fields have been a de facto standard in novel view synthesis. Recently, there exist some methods exploring fusing multiple modalities within a single field, aiming…