most citedImproving Vision-Language-Action Model with Online Reinforcement Learning

1 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.RO2025

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

Xiaoyu Chen, Hangxing Wei, Pushi Zhang +9

Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenar…

cs.RO20251 cited

Improving Vision-Language-Action Model with Online Reinforcement Learning

Yanjiang Guo, Jianke Zhang, Xiaoyu Chen +4

Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulti…

cs.CV2025

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Jianke Zhang, Yanjiang Guo, Yucheng Hu +3

Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-…

cs.CV2024

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Yucheng Hu, Yanjiang Guo, Pengchao Wang +6

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image…

cs.RO20241 cited

Prediction with Action: Visual Policy Learning via Joint Denoising Process

Yanjiang Guo, Yucheng Hu, Jianke Zhang +4

Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical…