Showing cs.ARShow all
3 papers · 1 filter
cs.AR2025
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
Minsu Kim, Seongmin Hong, RyeoWook Ko +5
Modern Large Language Model serving system batches multiple requests to achieve high throughput, while batching attention operations is challenging, rendering memory bandwidth a cr…
cs.AR2025
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
Junsoo Kim, Hunjong Lee, Geonwoo Ko +4
The growing adoption of Large Language Models (LLMs) across various domains has driven the demand for efficient and scalable AI-serving solutions. Deploying LLMs requires optimizat…
cs.AR2024
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
Seungjae Moon, Jung-Hoon Kim, Junsoo Kim +14
The explosive arrival of OpenAI's ChatGPT has fueled the globalization of large language model (LLM), which consists of billions of pretrained parameters that embodies the aspects…