works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.RO2026

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Yixiang Chen, Jiabing Yang, Yuan Xu +10

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture p…

cs.RO2026

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Peiyan Li, Yuze Zhu, Yixiang Chen +10

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existi…

cs.RO2026

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Yixiang Chen, Peiyan Li, Yuan Xu +13

The paper introduces FlowWAM, a dual‑stream diffusion model that uses optical flow as a unified video‑native representation of actions, enabling both action prediction and world mo…

cs.CV2026

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

Jiabing Yang, Yixiang Chen, Yuan Xu +14

Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demonstrating remarkable potential for…

cs.RO2026

SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models

Ziheng He, Yixiang Chen, Ning Yang +11

Embodied world models have emerged as a promising paradigm in robotics by predicting how robot actions affect the surrounding scene. However, the rollout inference remains computat…

cs.CV2026

Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving

Minhao Xiong, Zichen Wen, Zhuangcheng Gu +9

Vision-Language Models (VLMs) have emerged as a promising paradigm in autonomous driving (AD), providing a unified framework for perception and decision-making. However, their real…