2 papers
cs.DC2026
ShuntServe: Cost-Efficient LLM Serving on Heterogeneous Spot GPU Clusters
Seungwoo Jeong, Moohyun Song, Juhyun Park +1
As large language model (LLM) services become widely adopted, the cost of GPU resources for serving these models in cloud environments has emerged as a critical concern. Spot insta…
cs.DC2026
Latency Prediction for LLM Inference on NPU Systems
Juhyun Park, Seungwoo Jeong, Jingyu Lee +1
Deploying Large Language Models (LLMs) requires exploring a large configuration space spanning parallelization strategies, batching techniques, and scheduling policies. Exhaustive…