5 papers
ARM: Advantage Reward Modeling for Long-Horizon Manipulation
Yiming Mao, Zixi Yu, Weixin Mao +5
Long-horizon robotic manipulation remains challenging for reinforcement learning (RL) because sparse rewards provide limited guidance for credit assignment. Practical policy improv…
Long-Term Memory for VLA-based Agents in Open-World Task Execution
Xu Huang, Weixin Mao, Yinhao Li +2
Vision-Language-Action (VLA) models have demonstrated significant potential for embodied decision-making; however, their application in complex chemical laboratory automation remai…
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
Haosheng Li, Weixin Mao, Zihan Lan +6
Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. H…
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
Wenhao Li, Chengwei Ma, Weixin Mao
In robot learning, Vision Transformers (ViTs) are standard for visual perception, yet most methods discard valuable information by using only the final layer's features. We argue t…
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
Guangyan Chen, Meiling Wang, Qi Shao +10
Developing robust and general-purpose manipulation policies represents a fundamental objective in robotics research. While Vision-Language-Action (VLA) models have demonstrated pro…