8 citations · 18 across the 43 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
Yixuan Wang, Haoyu Qiao, Lujun Li +2
Large Language Models (LLMs) confront significant memory challenges due to the escalating KV cache with increasing sequence length. As a crucial technique, existing cross-layer KV…
cs.LG2024
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
Yuzhuang Xu, Shiyu Ji, Qingfu Zhu +1
Powerful large language models (LLMs) are increasingly expected to be deployed with lower computational costs, enabling their capabilities on resource-constrained devices. Post-tra…