collaborators
Showing cs.ROShow all

12 papers · 1 filter

cs.RO2026

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

Yupeng Zheng, Xiang Li, Songen Gu +11

Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction obje…

cs.RO2026

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

Jiahao Liu, Zhongpu Xia, Shuai Tian +13

Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot…

cs.RO2026

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

Shuai Tian, Yupeng Zheng, Yuhang Zheng +7

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…

cs.RO2026

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models

Boyu Li, Chaoyi Xu, Haoqi Yuan +5

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse…

cs.RO2026

Posterior Optimization with Clipped Objective for Bridging Efficiency and Stability in Generative Policy Learning

Yuhui Chen, Haoran Li, Zhennan Jiang +4

Expressive generative models have advanced robotic manipulation by capturing complex, multi-modal action distributions over temporally extended trajectories. However, fine-tuning t…

cs.RO2026

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation

Jiahao Liu, Cui Wenbo, Zhongpu Xia +3

Mobile manipulation is a fundamental capability for general-purpose robotic agents, requiring both coordinated control of the mobile base and manipulator and robust perception unde…