4 papers
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
Bodon Jeong, Hongsu Byun, Youngjae Kim +4
The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (…
AFLL: Real-time Load Stabilization for MMO Game Servers Based on Circular Causality Learning
Shinsuk Kang, Youngjae Kim
Massively Multiplayer Online (MMO) game servers must handle thousands of simultaneous players while maintaining sub-100ms response times. When server load exceeds capacity, traditi…
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
Kihyun Kim, Jinwoo Kim, Hyunsun Chung +3
LLM inference is essential for applications like text summarization, translation, and data analysis, but the high cost of GPU instances from Cloud Service Providers (CSPs) like AWS…
Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
Hyungwoo Lee, Kihyun Kim, Jinwoo Kim +5
Recent large language models (LLMs) face increasing inference latency as input context length and model size continue to grow. In particular, the retrieval-augmented generation (RA…