2 papers
cs.AR2025
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
Minsu Kim, Seongmin Hong, RyeoWook Ko +5
Modern Large Language Model serving system batches multiple requests to achieve high throughput, while batching attention operations is challenging, rendering memory bandwidth a cr…
cs.AR2024
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
Seungjae Moon, Jung-Hoon Kim, Junsoo Kim +14
The explosive arrival of OpenAI's ChatGPT has fueled the globalization of large language model (LLM), which consists of billions of pretrained parameters that embodies the aspects…