1 paper · 1 filter
Wei Luo, Yi Huang, Songchen Ma +3
The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when they process long contexts. Curren…