collaborators

6 papers

cs.CV2025

Step by Step Network

Dongchen Han, Tianzhu Ye, Zhuofan Xia +4

Scaling up network depth is a fundamental pursuit in neural architecture design, as theory suggests that deeper models offer exponentially greater capability. Benefiting from the r…

cs.CV2025

IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance

Jiayi Guo, Chuanhao Yan, Xingqian Xu +4

Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-qu…

cs.CV2024

COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing

Jiangshan Wang, Yue Ma, Jiayi Guo +3

Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite e…

cs.CV2024

ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis

Zanlin Ni, Yulin Wang, Renping Zhou +5

Recently, token-based generation have demonstrated their effectiveness in image synthesis. As a representative example, non-autoregressive Transformers (NATs) can generate decent-q…

cs.CV2024

AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation

Zanlin Ni, Yulin Wang, Renping Zhou +6

Recent studies have demonstrated the effectiveness of token-based methods for visual content generation. As a representative work, non-autoregressive Transformers (NATs) are able t…

cs.CV2024

Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators

Yifan Pu, Zhuofan Xia, Jiayi Guo +9

This paper identifies significant redundancy in the query-key interactions within self-attention mechanisms of diffusion transformer models, particularly during the early stages of…