4 papers
Generalized Dynamics Generation towards Scannable Physical World Model
Yichen Li, Zhiyi Li, Brandon Feng +2
Digital twin worlds with realistic interactive dynamics presents a new opportunity to develop generalist embodied agents in scannable environments with complex physical behaviors.…
MultiModal Action Conditioned Video Generation
Yichen Li, Antonio Torralba
Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgen…
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
Linli Yao, Yicheng Li, Yuancheng Wei +11
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process c…
Learning Generalizable Language-Conditioned Cloth Manipulation from Long Demonstrations
Hanyi Zhao, Jinxuan Zhu, Zihao Yan +3
Multi-step cloth manipulation is a challenging problem for robots due to the high-dimensional state spaces and the dynamics of cloth. Despite recent significant advances in end-to-…