4 papers · 1 filter
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
Zihao Zheng, Zhihao Mao, Xingyue Zhou +9
Vision-and-Language Navigation (VLN) increasingly relies on large vision-language models, but their inference cost conflicts with real-time deployment. Token caching is a promising…
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
Zihao Zheng, Zhihao Mao, Sicheng Tian +8
Vision-Language-Action (VLA) Models have become the mainstream solution for robot control, but suffer from slow inference speeds. Speculative Decoding (SD) is a promising accelerat…
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
Zihao Zheng, Zhihao Mao, Maoliang Li +6
Vision-Language-Action (VLA) models build a token-domain robot control paradigm, yet suffer from low speed. Speculative Decoding (SD) is an optimization strategy that can boost inf…
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
Xiaoquan Sun, Zetian Xu, Chen Cao +9
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be impr…