11 citations · 12 across the 4 of their papers we have counts for
4 papers · 1 filter
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
Yuanhang Li, Yiren Song, Junzhe Bai +4
We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon charact…
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
Yuxin Song, Wenkai Dong, Shizun Wang +8
Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks…
CLOP: Video-and-Language Pre-Training with Knowledge Regularizations
Guohao Li, Hu Yang, Feng He +4
Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner,…
UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance
Wei Li, Xue Xu, Xinyan Xiao +8
Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusio…