2 papers
cs.CL2025
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
Dingyu Yao, Chenxu Yang, Zhengyang Tong +4
The Key-Value (KV) cache introduces substantial memory overhead during large language model (LLM) inference. Although existing vector quantization (VQ) methods reduce KV cache usag…
cs.CL2025
A Factuality and Diversity Reconciled Decoding Method for Knowledge-Grounded Dialogue Generation
Chenxu Yang, Zheng Lin, Chong Tian +6
Grounding external knowledge can enhance the factuality of responses in dialogue generation. However, excessive emphasis on it might result in the lack of engaging and diverse expr…