4 papers
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
Feng Jiang, Yang Chen, Kyle Xu +8
Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for…
Unified Multimodal Models as Auto-Encoders
Zhiyuan Yan, Kaiqing Lin, Zongjian Li +10
Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection,…
Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized Worlds
Fan Wang, Pengtao Shao, Yiming Zhang +6
In-Context Reinforcement Learning (ICRL) enables agents to learn automatically and on-the-fly from their interactive experiences. However, a major challenge in scaling up ICRL is t…
BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
Yumeng Zhang, Shi Gong, Kaixin Xiong +7
World models have attracted increasing attention in autonomous driving for their ability to forecast potential future scenarios. In this paper, we propose BEVWorld, a novel framewo…