1 paper · 1 filter
Chenhao Zhang, Hanyu Zhao, Hang Cheng +2
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world e…