5 papers
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
Xueji Fang, Liyuan Ma, Jianhao Zeng +3
Diffusion transformer (DiT) has been widely adopted in the generative diffusion field, advancing the denoising of query tokens through attention and Feed-Forward (\text{FFN}) layer…
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
Jianhao Zeng, Yancheng Bai, Ruidong Chen +7
Video virtual try-on technology provides a cost-effective solution for creating marketing videos in fashion e-commerce. However, its practical adoption is hindered by two critical…
Semantic Context Matters: Improving Conditioning for Autoregressive Models
Dongyang Jin, Ryan Xu, Jianhao Zeng +4
Recently, autoregressive (AR) models have shown strong potential in image generation, offering better scalability and easier integration with unified multi-modal systems compared t…
Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
Ruidong Chen, Yancheng Bai, Xuanpu Zhang +6
Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and…
Group Relative Attention Guidance for Image Editing
Xuanpu Zhang, Xuesong Niu, Ruidong Chen +6
Recently, image editing based on Diffusion-in-Transformer models has undergone rapid development. However, existing editing methods often lack effective control over the degree of…