4 papers · 1 filter
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
Chenyu Su, Zhaolong Shen, Yuan Qian +8
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcem…
VINE: Taming Generative Control Policies for Reinforcement Learning
Rushuai Yang, Zhuo Han, Houlin Li +10
Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of…
GPU-Parallel Multi-Task Reinforcement Learning with Demonstration Guided Policy Optimization
Rui Zhang, Qiwei Wu, Zhengyu Zhang +5
Large scale GPU-parallel reinforcement learning has changed what can be trained in robot simulation, yet most systems still optimize one specialist policy per task. We propose a co…
Q-VGM: Q-Guided Value-Gradient Matching for Offline-to-Online RL of Flow-Matching VLA
Ziqian Wang, Rui Zhang, Yitian Liu +3
We propose Q-Guided Value-Gradient Matching (Q-VGM), an offline-to-online reinforcement learning (RL) method for fine-tuning flow-matching vision-language-action (VLA) policies wit…