5 papers
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
Yicheng Liu, Shiduo Zhang, Zibin Dong +12
Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often…
On The Eigenvalue Rigidity of the Jacobi Unitary Ensemble
Dan Dai, Chenhao Lu
In this paper, we prove an optimal global rigidity estimate for the eigenvalues of the Jacobi unitary ensemble. Our approach begins by constructing a random measure defined through…
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
Tianyuan Yuan, Yicheng Liu, Chenhao Lu +3
Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requir…
Galaxea Open-World Dataset and G0 Dual-System VLA Model
Tao Jiang, Tianyuan Yuan, Yicheng Liu +7
We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gath…
Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control
Chenhao Lu, Xuxin Cheng, Jialong Li +6
Humanoid robots require both robust lower-body locomotion and precise upper-body manipulation. While recent Reinforcement Learning (RL) approaches provide whole-body loco-manipulat…