1 paper
Manlai Liang, JiaMing Zhang, Xiong Li +1
The increasing size of the Key-Value (KV) cache during the Large Language Models long-context inference is the main obstacle for its balance between the deployment cost and task ac…