1 paper
Gang Lin, Dongfang Li, Zhuoen Chen +4
The proliferation of long-context large language models (LLMs) exposes a key bottleneck: the rapidly expanding key-value cache during decoding, which imposes heavy memory and laten…