collaborators

9 papers

cs.RO2026

Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

Zhenduo Shang, Xiyao Liu, Bohan Li +4

Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing met…

cs.CV2026

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Qi Lyu, Baicheng Liu, Xudong Wang +3

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all…

cs.CV2026

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

Qi Lyu, Jiahua Dong, Baichen Liu +7

Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial…

cs.LG2026

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

Sihan Wang, Xiyao Liu, Lianqing Liu +1

On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well…

cs.CV2026

All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation

Xudong Wang, Gan Li, Zhiyu Liu +3

Deploying vision-and-language navigation (VLN) agents requires adaptation across diverse scenes and environments, but fine-tuning on a specific scenario often causes catastrophic f…

cs.RO2026

Lifelong Embodied Navigation Learning

Xudong Wang, Jiahua Dong, Baichen Liu +3

Embodied navigation agents powered by large language models have shown strong performance on individual tasks but struggle to continually acquire new navigation skills, which suffe…