1 paper
Tan T. Nguyen, Quan V. Dang
As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually b…