4 papers
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
Pengxiang Zhao, Xiaoming Yuan
Large Language Models (LLMs) face significant deployment challenges due to their substantial resource requirements. While low-bit quantized weights can reduce memory usage and impr…
SPAP: Structured Pruning via Alternating Optimization and Penalty Methods
Hanyu Hu, Xiaoming Yuan
The deployment of large language models (LLMs) is often constrained by their substantial computational and memory demands. While structured pruning presents a viable approach by el…
FASP: Fast and Accurate Structured Pruning of Large Language Models
Hanyu Hu, Pengxiang Zhao, Ping Li +3
The rapid increase in the size of large language models (LLMs) has significantly escalated their computational and memory demands, posing challenges for efficient deployment, espec…
A Convex-optimization-based Layer-wise Post-training Pruner for Large Language Models
Pengxiang Zhao, Hanyu Hu, Ping Li +3
Pruning is a critical strategy for compressing trained large language models (LLMs), aiming at substantial memory conservation and computational acceleration without compromising p…