5 papers
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
Xudong Wang, Gan Li, Zhiyu Liu +3
Deploying vision-and-language navigation (VLN) agents requires adaptation across diverse scenes and environments, but fine-tuning on a specific scenario often causes catastrophic f…
Lifelong Embodied Navigation Learning
Xudong Wang, Jiahua Dong, Baichen Liu +3
Embodied navigation agents powered by large language models have shown strong performance on individual tasks but struggle to continually acquire new navigation skills, which suffe…
Lifelong Language-Conditioned Robotic Manipulation Learning
Xudong Wang, Zebin Han, Zhiyu Liu +5
Traditional language-conditioned manipulation agent sequential adaptation to new manipulation skills leads to catastrophic forgetting of old skills, limiting dynamic scene practica…
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning
Zebin Han, Xudong Wang, Baichen Liu +5
Sequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task navigation guided by complex, long-ho…
CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
Baichen Liu, Qi Lyu, Xudong Wang +3
Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal…