activity
20242026
most citedImproving Vision-Language-Action Model with Online Reinforcement Learning

1 citations · 2 across the 11 of their papers we have counts for

collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

Veo-Act: Enhancing VLA Policies with Frontier Video Models

Zhongru Zhang, Chenghan Yang, Qingzhou Lu +4

Video generation models can produce coherent vi- sual sequences depicting object motion and interactions. We in- vestigate how frontier video generation models can complement visio…

cs.RO2026

BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation

Yucheng Hu, Jianke Zhang, Yuanfei Luo +9

Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recen…

cs.RO2025

UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning

Jianke Zhang, Yucheng Hu, Yanjiang Guo +5

Building generalist robot policies that can handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage knowledge from large-scale pretraining,…

cs.RO2025★ 1 cited

Improving Vision-Language-Action Model with Online Reinforcement Learning

Yanjiang Guo, Jianke Zhang, Xiaoyu Chen +4

Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulti…

cs.RO2024★ 1 cited

Prediction with Action: Visual Policy Learning via Joint Denoising Process

Yanjiang Guo, Yucheng Hu, Jianke Zhang +4

Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical…