6 papers
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
Bingtian Qiao, Yue Shi, Yingjie Zhou +3
Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inheri…
Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts
Zhen Sun, Yongjian Guo, Haoran Sun +6
While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deployment remains challenged by…
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
Yucheng Guo, Yongjian Guo, Zhong Guan +6
In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to t…
FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
Fuhan Cai, Yong Guo, Jie Li +3
Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, the…
Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic Teacher
Yong Guo, Shulian Zhang, Haolin Pan +3
Knowledge distillation aims to transfer knowledge from a large teacher model to a compact student counterpart, often coming with a significant performance gap between them. We find…
Enhanced Long-Tailed Recognition with Contrastive CutMix Augmentation
Haolin Pan, Yong Guo, Mianjie Yu +1
Real-world data often follows a long-tailed distribution, where a few head classes occupy most of the data and a large number of tail classes only contain very limited samples. In…