13 papers · 1 filter
Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache
Zhirong Shen, Rui Huang, Chang Zou +10
Diffusion Transformers have become the dominant paradigm in generative AI, but their high computational costs severely hinder real-time applications. Prediction-based feature cachi…
AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
Haoran Qin, Zhengan Yan, Shikang Zheng +9
Diffusion Transformers (DiTs) achieve high-quality generation but are costly due to iterative sampling. Dynamic-resolution sampling reduces early-stage cost by denoising at low res…
LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
Jinshan Liu, Haoran Qin, Xiaobing Tu +9
Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical d…
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
Yuhang Han, Wenzheng Yang, Yujie Chen +4
Vision-language-model-based graphical user interface (GUI) agents have shown broad automation capabilities, yet deployment is bottlenecked by a key-value (KV) cache that grows line…
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
Feihong Yan, Shaoyu Liu, Haixuan Wang +4
Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents direct generation at higher r…
SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking
Zhengan Yan, Shikang Zheng, Haoran Qin +9
Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dyna…