1 paper
Ahmed Burak Gulhan, Krishna Teja Chitty-Venkata, Murali Emani +2
In Large Language Model (LLM) inference, Key-Value (KV) caches (KV-caches) are essential for reducing time complexity. However, they result in a linear increase in GPU memory as th…