1 paper
Junkai Zhang, Hang Guo, Luca Benini +1
Large language models (LLMs) have shown strong performance across diverse tasks, but their inference with long input contexts is bottlenecked by memory size and bandwidth. The Key-…