1 paper
Lianjun Liu, Hongli An, Weiqi Yan +4
The growing computational and memory demands of the Key-Value (KV) cache significantly limit the ability of Large Language Models (LLMs). While KV merging has emerged as a promisin…