collaborators

7 papers

cs.LG2025

DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization

Yuantian Shao, Yuanteng Chen, Peisong Wang +5

Quantization plays a crucial role in accelerating the inference of large-scale models, and rotational matrices have been shown to effectively improve quantization performance by sm…

cs.CV2025

Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation

Qirui Li, Guangcong Zheng, Qi Zhao +4

The computational demands of self-attention mechanisms pose a critical challenge for transformer-based video generation, particularly in synthesizing ultra-long sequences. Current…

cs.LG2025

Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models

Tai An, Ruwu Cai, Yanzhe Zhang +6

In the era of large language models (LLMs), N:M sparsity has emerged as a structured compression technique critical for accelerating inference. While prior work has primarily focus…

cs.CV2025

RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy

Aiyue Chen, Bin Dong, Jingru Li +4

Video generation using diffusion models is highly computationally intensive, with 3D attention in Diffusion Transformer (DiT) models accounting for over 80\% of the total computati…

cs.CV2025

Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers

Haosong Liu, Yuge Cheng, Wenxuan Miao +8

Video diffusion transformers (vDiTs) have made tremendous progress in text-to-video generation, but their high compute demands pose a major challenge for practical deployment. Whil…

cs.LG2025

Dynamic Low-Rank Sparse Adaptation for Large Language Models

Weizhong Huang, Yuxin Zhang, Xiawu Zheng +4

Despite the efficacy of network sparsity in alleviating the deployment strain of Large Language Models (LLMs), it endures significant performance degradation. Applying Low-Rank Ada…