1 paper
Oteo Mamo, Hyunjin Yi, Joydhriti Choudhury +2
As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. Existing input-le…