9 papers
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
Bing Hu, Zaijing Li, Rui Shao +4
Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varyi…
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
Zaijing Li, Bing Hu, Rui Shao +5
Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception…
HiconAgent: History Context-aware Policy Optimization for GUI Agents
Xurui Zhou, Gongwei Chen, Yuquan Xie +6
Graphical User Interface (GUI) agents require effective use of historical context to perform sequential navigation tasks. While incorporating past actions and observations can impr…
Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimization
Zaijing Li, Yuquan Xie, Rui Shao +5
Developing generalist agents capable of solving open-ended tasks in visually rich, dynamic environments remains a core pursuit of embodied AI. While Minecraft has emerged as a comp…
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
Haoyu Zhang, Meng Liu, Zaijing Li +4
Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied in…
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
Haoxiang Shi, Xiang Deng, Zaijing Li +3
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural language instructions through free-form 3D spaces. Existing VLN-CE approaches typic…