activity
20242026
most citedVASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

9 citations · 27 across the 17 of their papers we have counts for

collaborators
Showing cs.ROShow all

12 papers · 1 filter

cs.RO2026

Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution

Weichen Xu, Zhenhua Liu, Lin Luo +8

Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task-agnostic peri…

cs.RO2026

FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

Ruicheng Li, Qixiu Li, Ruichun Ma +8

Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by condi…

cs.RO2026

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

Zili Lin, Wenyao Zhang, Yuyang Zhang +9

Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the…

cs.RO2026

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

Rongxu Cui, Zongzheng Zhang, Jingrui Pang +11

Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address th…

cs.RO2026

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

Xingyuming Liu, Ruichun Ma, Heyu Guo +7

Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for robotics rely solely on RGB observation…

cs.RO2026

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

Zhiyuan Feng, Qixiu Li, Huizhi Liang +12

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large…