2 papers
cs.CV2026
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
Yiren Song, Yihan Wang, Xiyao Deng +2
Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generatio…
cs.CV2026
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
Yiren Song, Xiyao Deng, Pei Yang +2
Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for e…