18 papers
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Peiyan Li, Yuze Zhu, Yixiang Chen +10
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existi…
SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation
Nie Lin, Takehiko Ohkawa, Sijin Chen +10
Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulatio…
World Value Models for Robotic Manipulation
Zhihao Wang, Jianxiong Li, Yu Cui +4
Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. Mathematically, accurate value estimation demands deep temporal…
TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models
Wenbo Zhang, Jianxiong Li, Shuai Yang +4
Vision-Language-Action (VLA) models trained on large-scale data have made remarkable progress, but they remain vulnerable to distribution shifts at deployment time. Recent VLA mode…
Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention
Zhuohang Li, Liqun Huang, Wei Xu +5
Vision-Language-Action (VLA) models are prone to compounding errors in dexterous manipulation, where high-dimensional action spaces and contact-rich dynamics amplify small policy d…
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
Lei Lv, Yunfei Li, Yu Luo +2
Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning…