10 papers · 1 filter
AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
Haoran Qin, Zhengan Yan, Shikang Zheng +9
Diffusion Transformers (DiTs) achieve high-quality generation but are costly due to iterative sampling. Dynamic-resolution sampling reduces early-stage cost by denoising at low res…
LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
Jinshan Liu, Haoran Qin, Xiaobing Tu +9
Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical d…
Dynamic Video Generation: Shaping Video Generation Across Time and Space
Shikang Zheng, Jingkai Huang, Jiacheng Liu +5
Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens…
SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking
Zhengan Yan, Shikang Zheng, Haoran Qin +9
Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dyna…
From Sketch to Fresco: Efficient Diffusion Transformer with Progressive Resolution
Shikang Zheng, Guantao Chen, Lixuan He +4
Diffusion Transformers achieve impressive generative quality but remain computationally expensive due to iterative sampling. Recently, dynamic resolution sampling has emerged as a…
Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Efficient Diffusion Transformers
Guantao Chen, Shikang Zheng, Yuqi Lin +1
Diffusion Transformer (DiT) models have achieved unprecedented quality in image and video generation, yet their iterative sampling process remains computationally prohibitive. To a…