17 papers
URoPE: Universal Relative Position Embedding across Geometric Spaces
Yichen Xie, Depu Meng, Chensheng Peng +4
Relative position embedding has become a standard mechanism for encoding positional information in Transformers. However, existing formulations are typically limited to a fixed geo…
DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors
Pengcheng Wang, Kaiwen Hong, Chensheng Peng +4
Unlike chatbots, physical AI must act while the world keeps evolving. Therefore, the inter-chunk pause of synchronous executors are fatal for dynamic tasks regardless of how fast t…
Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
Jiuming Liu, Chaojun Ni, Mengmeng Liu +7
With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, benefiting various downstream do…
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
Tianshuo Xu, Yichen Xie, Depu Meng +5
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity p…
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
Yixiao Wang, Mingxiao Huo, Zhixuan Liang +8
Pretrained vision foundation models (VFMs) advance robotic learning via rich visual representations, yet individual VFMs typically excel only in specific domains, limiting generali…
R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation
William Ljungbergh, Bernardo Taveira, Wenzhao Zheng +8
Validating autonomous driving (AD) systems requires diverse and safety-critical testing, making photorealistic virtual environments essential. Traditional simulation platforms, whi…