1 paper
Tong Yuan, Chengxi Liao, Zeyi Wen
Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of tokens, making decoding latency a…