2 papers
cs.AR2025
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
Minsu Kim, Seongmin Hong, RyeoWook Ko +5
Modern Large Language Model serving system batches multiple requests to achieve high throughput, while batching attention operations is challenging, rendering memory bandwidth a cr…
cs.AR2025
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
Junsoo Kim, Hunjong Lee, Geonwoo Ko +4
The growing adoption of Large Language Models (LLMs) across various domains has driven the demand for efficient and scalable AI-serving solutions. Deploying LLMs requires optimizat…