1 paper
Ngoc Bui, Shubham Sharma, Simran Lamba +2
Memory and computation remain core bottlenecks in long-horizon LLM inference due to the quadratic cost of self-attention and the ever-growing key-value (KV) cache. Existing strateg…