2 papers
cs.LG2026
Locality-Aware Redundancy Pruning for LLM Depth Compression
Vincent-Daniel Yun, Youngrae Kim, Woosang Lim +3
Large language models are known to contain representational redundancy across network depth, making depth pruning an effective approach for improving inference efficiency. Existing…
cs.LG2026
Neural Weight Compression for Language Models
Jegwang Ryu, Minkyu Kim, Seungjun Shin +3
Efficient compression of language model weights is increasingly critical as model scale and deployment grow. Yet, most existing methods rely on handcrafted transforms and heuristic…