collaborators

9 papers

cs.RO2026

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Yixiang Chen, Jiabing Yang, Yuan Xu +10

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture p…

cs.RO2026

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld Team, Angen Ye, Angyuan Ma +26

The paper introduces GigaWorld-Policy-0.5, a robot control model that learns from future visual dynamics during training but generates actions only at inference, achieving faster (…

cs.RO2026

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Yixiang Chen, Peiyan Li, Yuan Xu +13

The paper introduces FlowWAM, a dual‑stream diffusion model that uses optical flow as a unified video‑native representation of actions, enabling both action prediction and world mo…

cs.RO2026

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

Yuan Xu, Yixiang Chen, Kai Wang +5

Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all t…

cs.RO2026

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

Ning Yang, Yan Huang, Kaiwen Peng +9

Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policies that directly map observat…

cs.CV2026

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

Jiabing Yang, Yixiang Chen, Yuan Xu +14

Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demonstrating remarkable potential for…