collaborators
Showing cs.CVShow all

34 papers · 1 filter

cs.CV2026

Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation

Zixuan Duan, Xunzhi Xiang, Yabo Chen +6

Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to p…

cs.CV2026

Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models

Ke Hao, Yuanzhi Liang, Tingxi Chen +5

Unified multimodal models integrate visual understanding and generation within a single network, yet the two capabilities are commonly optimized as separate tasks. We introduce Gen…

cs.CV2026

Search-to-World: Evaluation of 3D World Delivery from User Request through Web Search

Zixiao Gu, Yabo Chen, Xunzhi Xiang +5

Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to transform retrieved web content into a usable 3D world has not been s…

cs.CV2026

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

Xin Zhang, Yabo Chen, Zixuan Duan +4

Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, an…

cs.CV2026

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

Bojia Zi, Xiaoyan Yang, Yu Zhou +7

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, t…

cs.CV2026

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

Rui Li, Yuanzhi Liang, Ke Hao +4

Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. Howeve…