6 papers
HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling
Ziran Qin, Yuchen Jiang, Mingbao Lin +4
Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality generation with substantially fewer decoding steps. However, existing VAR models su…
Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
Rui Huang, Shitong Shao, Zikai Zhou +6
Diffusion models have achieved remarkable performance on a wide range of generative tasks, yet training them from scratch is notoriously resource-intensive, typically requiring mil…
PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation
Jiangshan Wang, Kang Zhao, Jiayi Guo +5
High computational costs and slow inference hinder the practical application of video generation models. While prior works accelerate the generation process through feature caching…
Elastic Diffusion Transformer
Jiangshan Wang, Zeqiang Lai, Jiarui Chen +5
Diffusion Transformers (DiT) have demonstrated remarkable generative capabilities but remain highly computationally expensive. Previous acceleration methods, such as pruning and di…
Efficient Autoregressive Video Diffusion with Dummy Head
Hang Guo, Zhaoyang Jia, Jiahao Li +5
The autoregressive video diffusion model has recently gained considerable research interest due to its causal modeling and iterative denoising. In this work, we identify that the m…
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
Hang Guo, Yawei Li, Luca Benini
Recent advances in Large Language Model (LLM) compression, such as quantization and pruning, have achieved notable success. However, as these techniques gradually approach their re…