3 papers
cs.CL2025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
Haoqi Yang, Yao Yao, Zuchao Li +3
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks. However, their extensive memory requirements, particularly…
cs.CL2024
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
Luohe Shi, Yao Yao, Zuchao Li +2
Large language models (LLMs) have rapidly advanced and demonstrated impressive capabilities. In-Context Learning (ICL) and Parameter-Efficient Fine-Tuning (PEFT) are currently two…
cs.CL2024
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Luohe Shi, Hongyi Zhang, Yao Yao +2
Large Language Models (LLMs), epitomized by ChatGPT's release in late 2022, have revolutionized various industries with their advanced language comprehension. However, their effici…