2 papers
cs.DC2025
KV Cache Compression for Inference Efficiency in LLMs: A Review
Yanyu Liu, Jingying Fu, Sixiang Liu +4
Withtherapid advancement of large language models (LLMs), the context length for inference has been continuously increasing, leading to an exponential growth in the demand for Key-…
cs.DC2025
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
Jie Kong, Junxiang Zhang, Jiheng Xu +7
In the field of deep learning, traditional attention mechanisms face significant challenges related to high computational complexity and large memory consumption when processing lo…