13 papers
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang +4
Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers…
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
Yanjun Guo, Zhengqiang Zhang, Pengfei Wang +3
Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory m…
VOSR: A Vision-Only Generative Model for Image Super-Resolution
Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang +4
Most of the recent generative image super-resolution (SR) methods rely on adapting large text-to-image (T2I) diffusion models pretrained on web-scale text-image data. While effecti…
GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution
Qiaosi Yi, Shuai Li, Rongyuan Wu +3
Recently, reinforcement learning (RL) has been employed for improving generative image super-resolution (ISR) performance. However, the current efforts are focused on multi-step ge…
MV2UV: Generating High-quality UV Texture Maps with Multiview Prompts
Zheng Zhang, Qinchuan Zhang, Yuteng Ye +5
Generating high-quality textures for 3D assets is a challenging task. Existing multiview texture generation methods suffer from the multiview inconsistency and missing textures on…
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
Chaodong Xiao, Zhengqiang Zhang, Lei Zhang
Transformers have achieved widespread and remarkable success, while the computational complexity of their attention modules remains a major bottleneck for vision tasks. Existing me…