3 papers
cs.RO2026
GVLA: Geometric inductive bias for Vision-Language-Action Models
Yue Peng, Yongzhe Zhao, Artur Habuda +5
Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their…
cs.AI2025
Exploring Task-Level Optimal Prompts for Visual In-Context Learning
Yan Zhu, Huan Ma, Changqing Zhang
With the development of Vision Foundation Models (VFMs) in recent years, Visual In-Context Learning (VICL) has become a better choice compared to modifying models in most scenarios…
cs.CV2025
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
Huan Ma, Yan Zhu, Changqing Zhang +5
Vision-language foundation models have exhibited remarkable success across a multitude of downstream tasks due to their scalability on extensive image-text paired data. However, th…