Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
Chao Han, Haozhe Hu, Xiaoyu Shen
Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance d…
cs.LG2026
UniRank: Unified Rank Allocation for Low-Rank LLM Compression
Chao Han, Haozhe Hu, Yongjie Du +5
Low-rank decomposition is a promising compression paradigm for large language models (LLMs), yet its effectiveness hinges on rank budget allocation across weight matrices: uniform…