3 citations · 3 across the 7 of their papers we have counts for
9 papers
VINO: A Unified Visual Generator with Interleaved OmniModal Context
Junyi Chen, Tong He, Zhoujie Fu +3
We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independen…
In-Context Audio Control of Video Diffusion Transformers
Wenze Liu, Weicai Ye, Minghong Cai +3
Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, thes…
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
Shengqiong Wu, Weicai Ye, Yuanxing Zhang +7
Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionall…
Native 3D Editing with Full Attention
Weiwei Cai, Shuangkang Fang, Weicai Ye +7
Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimiza…
SketchVideo: Sketch-based Video Generation and Editing
Feng-Lin Liu, Hongbo Fu, Xintao Wang +4
Video generation and editing conditioned on text prompts or images have undergone significant advancements. However, challenges remain in accurately controlling global layout and g…
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
Xuan Ju, Weicai Ye, Quande Liu +6
Current video generative foundation models primarily focus on text-to-video tasks, providing limited control for fine-grained video content creation. Although adapter-based approac…