1 citations · 2 across the 11 of their papers we have counts for
5 papers · 1 filter
Veo-Act: Enhancing VLA Policies with Frontier Video Models
Zhongru Zhang, Chenghan Yang, Qingzhou Lu +4
Video generation models can produce coherent vi- sual sequences depicting object motion and interactions. We in- vestigate how frontier video generation models can complement visio…
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
Yucheng Hu, Jianke Zhang, Yuanfei Luo +9
Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recen…
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
Jianke Zhang, Yucheng Hu, Yanjiang Guo +5
Building generalist robot policies that can handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage knowledge from large-scale pretraining,…
Improving Vision-Language-Action Model with Online Reinforcement Learning
Yanjiang Guo, Jianke Zhang, Xiaoyu Chen +4
Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulti…
Prediction with Action: Visual Policy Learning via Joint Denoising Process
Yanjiang Guo, Yucheng Hu, Jianke Zhang +4
Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical…