1 paper · 1 filter
Lianjun Liu, Hongli An, Weiqi Yan +4
The growing computational and memory demands of the Key-Value (KV) cache significantly limit the ability of Large Language Models (LLMs). While KV merging has emerged as a promisin…