1 paper · 1 filter
Chenxia Tang, Jianchun Liu, Hongli Xu +1
Large language models (LLMs) inference relies heavily on KV-caches to accelerate autoregressive decoding, but the resulting memory footprint grows rapidly with sequence length, pos…