4 papers
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
Hokyun Im, Euijin Jeong, Andrey Kolobov +2
Vision-language-action models (VLAs) trained on large-scale robotic datasets have demonstrated strong performance on manipulation tasks, including bimanual tasks. However, because…
LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
Xiaofan Wang, Xingyu Gao, Jianlong Fu +5
The capability of performing long-horizon, language-guided robotic manipulation tasks critically relies on leveraging historical information and generating coherent action sequence…
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
Zuolei Li, Xingyu Gao, Xiaofan Wang +1
Learning transferable latent actions from large-scale object manipulation videos can significantly enhance generalization in downstream robotics tasks, as such representations are…
PromptFix: You Prompt and We Fix the Photo
Yongsheng Yu, Ziyun Zeng, Hang Hua +2
Diffusion models equipped with language models demonstrate excellent controllability in image generation tasks, allowing image processing to adhere to human instructions. However,…