3 papers
cs.RO2026
Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision
Haoyang Li, Guanlin Li, Youhe Feng +9
Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally across robot platforms. We obser…
cs.CV2026
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
Chen Zhao, Zhuoran Wang, Haoyang Li +6
Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generat…
cs.CV2025
Few-Shot-Based Modular Image-to-Video Adapter for Diffusion Models
Zhenhao Li, Shaohan Yi, Zheng Liu +7
Diffusion models (DMs) have recently achieved impressive photorealism in image and video generation. However, their application to image animation remains limited, even when traine…