5 citations · 15 across the 9 of their papers we have counts for
1 paper · 1 filter
Luca Moschella, Laura Manduchi, Ozan Sener
The growing size of Large Language Models (LLMs) makes efficient inference challenging, primarily due to the memory demands of the autoregressive Key-Value (KV) cache. Existing evi…