7 papers · 1 filter
How Should Vision-Language-Action Models Use Proprioceptive State?
Yiren Zhao, Ziyang Chen, Ziyang Rao +5
Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected…
Source-Lifted Flow Matching for Intervenable Multimodal Imitation
He Zhang, Ying Sun, Pengteng Li +6
Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sa…
PHASER: Phase-Aware and Semantic Experience Replay for Vision-Language-Action Models
Ziyang Chen, Shaoguang Wang, Weiyu Guo +5
Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation. However, deploying these models in open-ended environments requir…
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
Pengteng Li, Weiyu Guo, He Zhang +4
We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objec…
A Brain-inspired Embodied Intelligence for Fluid and Fast Reflexive Robotics Control
Weiyu Guo, He Zhang, Pengteng Li +7
Recent advances in embodied intelligence have leveraged massive scaling of data and model parameters to master natural-language command following and multi-task control. In contras…
Equivariant Volumetric Grasping
Pinhao Song, Yutong Hu, Pengteng Li +1
We propose a new volumetric grasp model that is equivariant to rotations around the vertical axis, leading to a significant improvement in sampling efficiency. Our model employs a…