collaborators

5 papers

cs.RO2026

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

Zhenhao Shen, Jiaqi Liang, Jasper Lu +11

Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. Howe…

cs.CV2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Haodong Yan, Junfeng Li, Junjie He +12

Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs…

cs.RO2026

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

Qize Yu, Jiadi You, Yuran Wang +10

Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the…

cs.RO2026

FLUX: Accelerating Cross-Embodiment Generative Navigation Policies via Rectified Flow and Static-to-Dynamic Learning

Zeying Gong, Yangyi Zhong, Yiyi Ding +6

Autonomous navigation requires a broad spectrum of skills, from static goal-reaching to dynamic social traversal, yet evaluation remains fragmented across disparate protocols. We i…

cs.CV2025

SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes

Dicong Qiu, Jiadi You, Zeying Gong +3

We present the Semantics-aware Dataset and Benchmark Generation Pipeline for Open-vocabulary Object Navigation in Dynamic Scenes (SD-OVON). It utilizes pretraining multimodal found…