13 papers
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
Zixuan Ye, Quande Liu, Cong Wei +5
Recently, the introduction of Chain-of-Thought (CoT) has largely improved the generation ability of unified models. However, it is observed that the current thinking process during…
In-Context Audio Control of Video Diffusion Transformers
Wenze Liu, Weicai Ye, Minghong Cai +3
Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, thes…
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
Qinghe Wang, Xiaoyu Shi, Baolu Li +7
Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which require flexible shot arrangement, coherent narrative, and…
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
Shengqiong Wu, Weicai Ye, Yuanxing Zhang +7
Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionall…
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
Shian Du, Menghan Xia, Chang Liu +4
Cascaded video super-resolution has emerged as a promising technique for decoupling the computational burden associated with generating high-resolution videos using large foundatio…
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
Xuanhua He, Quande Liu, Zixuan Ye +7
Fine-grained and efficient controllability on video diffusion transformers has raised increasing desires for the applicability. Recently, In-context Conditioning emerged as a power…