24 papers
Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
Henglin Liu, Fangyuan Kong, Jing Wang +7
The paper introduces concentrated Implicit Preference Optimization (cIPO), a post‑training method for text‑to‑video diffusion models that derives preference signals from reconstruc…
UniVideo: Unified Understanding, Generation, and Editing for Videos
Cong Wei, Quande Liu, Zixuan Ye +5
Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVide…
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
Minghong Cai, Qiulin Wang, Zongli Ye +7
Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpainting, or interpolation, treating…
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
Zeyu Liu, Zanlin Ni, Yang Yue +5
Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely deco…
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
Liang Hou, Cong Liu, Mingwu Zheng +4
Resolution generalization in image generation tasks enables the production of higher-resolution images with lower training resolution overhead. However, a key obstacle for diffusio…
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
Shengqiong Wu, Weicai Ye, Jiahao Wang +8
To address the bottleneck of accurate user intent interpretation within the current video generation community, we present Any2Caption, a novel framework for controllable video gen…