2 papers
cs.CL2025
Squeezed Attention: Accelerating Long Context Length LLM Inference
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +6
Emerging Large Language Model (LLM) applications require long input context in order to perform complex tasks like document analysis and code generation. For these long context len…
cs.LG2025
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4
LLMs are seeing growing use for applications which require large context windows, and with these large context windows KV cache activations surface as the dominant contributor to m…