28 papers
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
Qu Tang, Benhui Zhuang, Bo Yuan +3
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
Jiayue Jin, Jingwei Zhang, Chen Wang +2
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work,…
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Yanghong Mei +7
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon task…
TimeThink: Reasoning with Time for Video LLMs
Handong Li, Longteng Guo, Zikang Liu +8
Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…
SurveilNav: Collaborative Object Goal Navigation with Robot and Surveillance System
Ming-Ming Yu, Qunbo Wang, Rongtao Xu +5
With the growing deployment of surveillance systems in factories, offices, and homes, integrating them with robots offers a promising direction for collaborative and efficient task…
NavWM: A Unified Navigation World Model for Foresight-Driven Planning
Yanghong Mei, Longteng Guo, Ming-Ming Yu +3
Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer a promising alternative, exis…