4 papers
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
Bingtian Qiao, Yue Shi, Yingjie Zhou +3
Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inheri…
Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts
Zhen Sun, Yongjian Guo, Haoran Sun +6
While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deployment remains challenged by…
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
Yucheng Guo, Yongjian Guo, Zhong Guan +6
In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to t…
FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
Fuhan Cai, Yong Guo, Jie Li +3
Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, the…