3 papers
cs.LG2026
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
Xixiang He, Qiyao Sun, Ao Cheng +5
Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improvin…
cs.RO2026
From a Single Demonstration to a General Policy for Contact-Rich Manipulation
Xing Li, Oliver Brock
We present a Learning from Demonstration (LfD) framework that achieves one-shot generalization in multi-stage, contact-rich manipulation tasks. Central to our approach is the utili…
cs.RO2026
AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation
Kai Yang, Zedong Chu, Yingnan Guo +6
While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter sizes typically necessitate c…