From the 1 of 10 linked papers with an AI index.
10 papers
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Xinghang Li, Jun Guo, Qiwei Li +21
The paper introduces Xiaomi-Robotics-U0, a 38‑billion‑parameter multimodal autoregressive model that extends foundation image and video generation to embodied robotics, enabling co…
OmniGen2: Towards Instruction-Aligned Multimodal Generation
Chenyuan Wu, Pengfei Zheng, Ruiran Yan +19
In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, imag…
Scaling World Model for Hierarchical Manipulation Policies
Qian Long, Yueze Wang, Jiaxi Song +13
Vision-Language-Action (VLA) models are promising for generalist robot manipulation but remain brittle in out-of-distribution (OOD) settings, especially with limited real-robot dat…
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
Huaying Yuan, Jian Ni, Zheng Liu +7
Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of…
Emu3.5: Native Multimodal Models are World Learners
Yufeng Cui, Honghao Chen, Haoge Deng +20
We introduce Emu3.5, a large-scale multimodal world model that natively predicts the next state across vision and language. Emu3.5 is pre-trained end-to-end with a unified next-tok…
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Haiwen Diao, Xiaotong Li, Yufeng Cui +6
Existing encoder-free vision-language models (VLMs) are rapidly narrowing the performance gap with their encoder-based counterparts, highlighting the promising potential for unifie…