2 papers
cs.CV2026
Action Images: End-to-End Policy Learning via Multiview Video Generation
Haoyu Zhen, Zixian Gao, Qiao Sun +7
World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, exist…
cs.RO2025
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
Hongyin Zhang, Shiyuan Zhang, Junxi Jin +4
Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these mod…