1 paper
Fei Li, Song Liu, Weiguo Wu +2
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quant…