collaborators

6 papers

cs.CV2026

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Qi Lyu, Baicheng Liu, Xudong Wang +3

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all…

cs.CV2026

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

Qi Lyu, Jiahua Dong, Baichen Liu +7

Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial…

cs.CL2026

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

Haoyi Hu, Qirong Lyu, Xianghan Kong +7

While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after explicit user prompts. This par…

cs.RO2026

Lifelong Embodied Navigation Learning

Xudong Wang, Jiahua Dong, Baichen Liu +3

Embodied navigation agents powered by large language models have shown strong performance on individual tasks but struggle to continually acquire new navigation skills, which suffe…

cs.RO2026

SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning

Zebin Han, Xudong Wang, Baichen Liu +5

Sequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task navigation guided by complex, long-ho…

cs.CV2025

CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation

Baichen Liu, Qi Lyu, Xudong Wang +3

Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal…