1 paper · 1 filter
Ferran Agullo, Joan Oliveras, Chen Wang +5
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of a…