collaborators

5 papers

cs.CV2026

FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation

Xueji Fang, Liyuan Ma, Jianhao Zeng +3

Diffusion transformer (DiT) has been widely adopted in the generative diffusion field, advancing the denoising of query tokens through attention and Feed-Forward (\text{FFN}) layer…

cs.CV2026

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on

Jianhao Zeng, Yancheng Bai, Ruidong Chen +7

Video virtual try-on technology provides a cost-effective solution for creating marketing videos in fashion e-commerce. However, its practical adoption is hindered by two critical…

cs.CV2026

Semantic Context Matters: Improving Conditioning for Autoregressive Models

Dongyang Jin, Ryan Xu, Jianhao Zeng +4

Recently, autoregressive (AR) models have shown strong potential in image generation, offering better scalability and easier integration with unified multi-modal systems compared t…

cs.CV2026

Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

Ruidong Chen, Yancheng Bai, Xuanpu Zhang +6

Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and…

cs.CV2025

Group Relative Attention Guidance for Image Editing

Xuanpu Zhang, Xuesong Niu, Ruidong Chen +6

Recently, image editing based on Diffusion-in-Transformer models has undergone rapid development. However, existing editing methods often lack effective control over the degree of…