3 papers
cs.DC2026
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
Ferran Agullo, Joan Oliveras, Chen Wang +5
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of a…
cs.PF2025
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
Ferran Agullo, Joan Oliveras, Chen Wang +5
With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds…
cs.CL2025
Towards Pareto Optimal Throughput in Small Language Model Serving
Pol G. Recasens, Yue Zhu, Chen Wang +5
Large language models (LLMs) have revolutionized the state-of-the-art of many different natural language processing tasks. Although serving LLMs is computationally and memory deman…