1 paper · 1 filter
Siyao Xiao, Yuhong Zhang, Zhifang Liu +9
Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, this paradigm compromises the inherent generalization capabilities of Vis…