Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
Ferran Agullo, Joan Oliveras, Chen Wang +5
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of a…
cs.DC2024
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
Beomyeol Jeon, Chen Wang, Diana Arroyo +2
This paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes…