1 paper
Luca Moschella, Laura Manduchi, Ozan Sener
The growing size of Large Language Models (LLMs) makes efficient inference challenging, primarily due to the memory demands of the autoregressive Key-Value (KV) cache. Existing evi…