10 papers
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
Minghong Cai, Qiulin Wang, Zongli Ye +7
Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpainting, or interpolation, treating…
VINO: A Unified Visual Generator with Interleaved OmniModal Context
Junyi Chen, Tong He, Zhoujie Fu +3
We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independen…
In-Context Audio Control of Video Diffusion Transformers
Wenze Liu, Weicai Ye, Minghong Cai +3
Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, thes…
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
Shengqiong Wu, Weicai Ye, Jiahao Wang +8
To address the bottleneck of accurate user intent interpretation within the current video generation community, we present Any2Caption, a novel framework for controllable video gen…
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
Shengqiong Wu, Weicai Ye, Yuanxing Zhang +7
Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionall…
Native 3D Editing with Full Attention
Weiwei Cai, Shuangkang Fang, Weicai Ye +7
Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimiza…