Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
Yitong Chen, Shiduo Zhang, Jingjing Gong +1
Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus a…
cs.CV2025
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
Yicheng Liu, Shiduo Zhang, Zibin Dong +12
Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often…