4 papers
Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models
Huoren Yang, Jianchao Zhao, Hu Yusong +7
Vision-Language-Action (VLA) models have advanced rapidly with stronger backbones, broader pre-training, and larger demonstration datasets, yet their action heads remain largely ho…
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
Jianchao Zhao, Huoren Yang, Yusong Hu +6
Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades under local deployment conditio…
KAC: Kolmogorov-Arnold Classifier for Continual Learning
Yusong Hu, Zichen Liang, Fei Yang +3
Continual learning requires models to train continuously across consecutive tasks without forgetting. Most existing methods utilize linear classifiers, which struggle to maintain a…
Multi-Token Enhancing for Vision Representation Learning
Zhong-Yu Li, Yu-Song Hu, Bo-Wen Yin +1
Vision representation learning, especially self-supervised learning, is pivotal for various vision applications. Ensemble learning has also succeeded in enhancing the performance a…