5 papers
An AI4AI Framework for Visual Token Pruning
Zhen Liu, Wenli Huang, Wei Song +3
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and…
The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation
Zhen Liu, Yuhan Liu, Jinjun Wang +3
In vision-and-language navigation (VLN), self-improvement from policy-induced experience, using only standard VLN action supervision, critically depends on balancing behavioral div…
Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation
Zhen Liu, Yuhan Liu, Jinjun Wang +3
Vision-and-Language Navigation requires agents to follow natural-language instructions in visually changing environments. A central challenge is the dynamic entanglement between la…
Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection
Zhen Liu, Xinyu Ning, Zhe Hu +8
Recent vision-language-action (VLA) systems have demonstrated strong capabilities in embodied manipulation. However, most existing VLA policies rely on limited observation windows…
HSC-VLA: Hierarchical Scene-Clearing for Robust Bimanual Manipulation in Dense Clutter
Zhen Liu, Xinyu Ning, Zhe Hu +3
Modern Vision--Language--Action models often suffer from critical instruction-following failures in high-density manipulation environments, where task-irrelevant visual clutter dil…