Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
Yixuan Wang, Haoyu Qiao, Lujun Li +2
Large Language Models (LLMs) confront significant memory challenges due to the escalating KV cache with increasing sequence length. As a crucial technique, existing cross-layer KV…
cs.LG2025
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
Lujun Li, Zhu Qiyuan, Jiacheng Wang +4
Mixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merg…