Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4
LLMs are seeing growing use for applications which require large context windows, and with these large context windows KV cache activations surface as the dominant contributor to m…
cs.LG2024
Learned Best-Effort LLM Serving
Siddharth Jha, Coleman Hooper, Xiaoxuan Liu +2
Many applications must provide low-latency LLM service to users or risk unacceptable user experience. However, over-provisioning resources to serve fluctuating request patterns is…