collaborators

6 papers

cs.RO2026

MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

Yuxin Jiang, Chang Yu, Yunuo Chen +4

Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observation window, which renders long-h…

cs.CV2026

StreamingEffect: Real-Time Human-Centric Video Effect Generation

Yiren Song, Cheng Liu, Yuxin Jiang +1

Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet remains difficult due to th…

cs.CV2026

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

Yuchao Gu, Guian Fang, Yuxin Jiang +4

Few-step video generation has been significantly advanced by consistency distillation. However, the performance of consistency-distilled models often degrades as more sampling step…

cs.RO2026

Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance

Wenhao Wang, Jianheng Song, Chiming Liu +13

While Vision-Language-Action (VLA) models show strong generalizability in various tasks, real-world deployment of robotic policy still requires large-scale, high-quality human expe…

cs.RO2025

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Yue Liao, Pengfei Zhou, Siyuan Huang +11

We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-g…

cs.RO2025

EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Yuxin Jiang, Shengcong Chen, Siyuan Huang +8

Robotic imitation learning has advanced from solving static tasks to addressing dynamic interaction scenarios, but testing and evaluation remain costly and challenging due to the n…