collaborators

7 papers

cs.RO2025

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

Siyu Xu, Zijian Wang, Yunke Wang +3

Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they…

cs.CV2025

Efficient Conditional Generation on Scale-based Visual Autoregressive Models

Jiaqi Liu, Tao Huang, Chang Xu

Recent advances in autoregressive (AR) models have demonstrated their potential to rival diffusion models in image synthesis. However, for complex spatially-conditioned generation,…

cs.RO2025

Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation

Xiaohuan Pei, Yuxing Chen, Siyu Xu +3

Robotic manipulation with Vision-Language-Action models requires efficient inference over long-horizon multi-modal context, where attention to dense visual tokens dominates computa…

cs.CV2025

Rethinking Causal Mask Attention for Vision-Language Inference

Xiaohuan Pei, Tao Huang, YanXiang Ma +1

Causal attention has become a foundational mechanism in autoregressive vision-language models (VLMs), unifying textual and visual inputs under a single generative framework. Howeve…

cs.CV2025

Learning Mask Invariant Mutual Information for Masked Image Modeling

Tao Huang, Yanxiang Ma, Shan You +1

Masked autoencoders (MAEs) represent a prominent self-supervised learning paradigm in computer vision. Despite their empirical success, the underlying mechanisms of MAEs remain ins…

cs.RO2025

VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching

Siyu Xu, Yunke Wang, Chenghao Xia +3

Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions…