most citedSeeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

1 citations · 1 across the 2 of their papers we have counts for

collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

Zhiyuan Feng, Qixiu Li, Huizhi Liang +12

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large…

cs.RO2025

Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories

Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang +6

Scaling vision-language-action (VLA) model pre-training requires large volumes of diverse, high-quality manipulation trajectories. Most current data is obtained via human teleopera…

cs.RO2025

How Do VLAs Effectively Inherit from VLMs?

Chuheng Zhang, Rushuai Yang, Xiaoyu Chen +4

Vision-language-action (VLA) models hold the promise to attain generalizable embodied control. To achieve this, a pervasive paradigm is to leverage the rich vision-semantic priors…

cs.RO2025

Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training

Rushuai Yang, Hangxing Wei, Ran Zhang +8

Vision-language-action (VLA) models have shown strong generalization across tasks and embodiments; however, their reliance on large-scale human demonstrations limits their scalabil…

cs.RO2025

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

Xiaoyu Chen, Hangxing Wei, Pushi Zhang +9

Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenar…