9 papers
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Yanghong Mei +7
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon task…
NavWM: A Unified Navigation World Model for Foresight-Driven Planning
Yanghong Mei, Longteng Guo, Ming-Ming Yu +3
Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer a promising alternative, exis…
VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Junyou Zhu +6
Vision-language-action (VLA) models have shown strong promise for robotic manipulation, but their reliability at test time remains limited by one-shot action prediction, where even…
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
Zilin Zhu, Longteng Guo, Yanghong Mei +5
Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize…
Thinking in Streaming Video
Zikang Liu, Longteng Guo, Handong Li +7
Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video re…
UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
Yanghong Mei, Yirong Yang, Longteng Guo +5
Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial…