7 papers
Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation
Siyu Xu, Zijian Wang, Yunke Wang +3
Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they…
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
Jiaqi Liu, Tao Huang, Chang Xu
Recent advances in autoregressive (AR) models have demonstrated their potential to rival diffusion models in image synthesis. However, for complex spatially-conditioned generation,…
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
Xiaohuan Pei, Yuxing Chen, Siyu Xu +3
Robotic manipulation with Vision-Language-Action models requires efficient inference over long-horizon multi-modal context, where attention to dense visual tokens dominates computa…
Rethinking Causal Mask Attention for Vision-Language Inference
Xiaohuan Pei, Tao Huang, YanXiang Ma +1
Causal attention has become a foundational mechanism in autoregressive vision-language models (VLMs), unifying textual and visual inputs under a single generative framework. Howeve…
Learning Mask Invariant Mutual Information for Masked Image Modeling
Tao Huang, Yanxiang Ma, Shan You +1
Masked autoencoders (MAEs) represent a prominent self-supervised learning paradigm in computer vision. Despite their empirical success, the underlying mechanisms of MAEs remain ins…
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
Siyu Xu, Yunke Wang, Chenghao Xia +3
Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions…