8 papers
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
Binhe Yu, Zhen Wang, Kexin Li +6
Multi-subject customization aims to synthesize multiple user-specified subjects into a coherent image. To address issues such as subjects missing or conflicts, recent works incorpo…
FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
Yilei Jiang, Zhen Wang, Yanghao Wang +4
With the surge of pre-trained text-to-image flow matching models, text-based image editing performance has gained remarkable improvement, especially for \underline{simple editing}…
CoMo: Compositional Motion Customization for Text-to-Video Generation
Youcan Xu, Zhen Wang, Jiaxin Shi +6
While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods f…
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
Zheng Chong, Yanwei Lei, Shiyue Zhang +7
Despite its great potential, virtual try-on technology is hindered from real-world application by two major challenges: the inability of current methods to support multi-reference…
Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
Hongxu Chen, Zhen Wang, Taoran Mei +4
Concept Erasure, which aims to prevent pretrained text-to-image models from generating content associated with semantic-harmful concepts (i.e., target concepts), is getting increas…
Diffusion Restoration Adapter for Real-World Image Restoration
Hanbang Liang, Zhen Wang, Weihui Deng
Diffusion models have demonstrated their powerful image generation capabilities, effectively fitting highly complex image distributions. These models can serve as strong priors for…