3 papers
cs.AI2026
Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models
Zhiming Liu, Yujie Wei, Lei Feng +5
Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downst…
cs.RO2026
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
Yujie Wei, Jiahan Fan, Jiyu Guo +5
Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a c…
cs.RO2026
Inject Once Survive Later: Backdooring Vision-Language-Action Models to Persist Through Downstream Fine-tuning
Jianyi Zhou, Yujie Wei, Ruichen Zhen +5
Vision-Language-Action (VLA) models have become foundational to modern embodied AI systems. By integrating visual perception, language understanding, and action planning, they enab…