2 papers
cs.CL2026
Correlation-Aware Structured Pruning for Large Language Models
Sicheng Xu, Hao Shi, Wei Zhang +6
Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods…
cs.LG2026
PRQuant: Permutation Residual Quantization for Low-Overhead Inference
Peiran Wang, Anqi Wang, Jiaying Zhao +10
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, bu…