34 papers · 1 filter
Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation
Zixuan Duan, Xunzhi Xiang, Yabo Chen +6
Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to p…
Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models
Ke Hao, Yuanzhi Liang, Tingxi Chen +5
Unified multimodal models integrate visual understanding and generation within a single network, yet the two capabilities are commonly optimized as separate tasks. We introduce Gen…
Search-to-World: Evaluation of 3D World Delivery from User Request through Web Search
Zixiao Gu, Yabo Chen, Xunzhi Xiang +5
Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to transform retrieved web content into a usable 3D world has not been s…
TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
Xin Zhang, Yabo Chen, Zixuan Duan +4
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, an…
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
Bojia Zi, Xiaoyan Yang, Yu Zhou +7
Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, t…
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
Rui Li, Yuanzhi Liang, Ke Hao +4
Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. Howeve…