1 paper
Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya +3
Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this footprint by evi…