3 papers
cs.LG2026
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
Shaoang Li, Jian Li
Edge deployment of large language models (LLMs) increasingly relies on libraries of lightweight LoRA adapters, yet GPU/DRAM can keep only a small resident subset at a time. Serving…
cs.LG2026
Near-Optimal Online Deployment and Routing for Streaming LLMs
Shaoang Li, Jian Li
The rapid pace at which new large language models (LLMs) appear, and older ones become obsolete, forces providers to manage a streaming inventory under a strict concurrency cap and…
cs.LG2025
Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits
Shaoang Li, Jian Li
Non-stationary multi-armed bandits enable agents to adapt to changing environments by incorporating mechanisms to detect and respond to shifts in reward distributions, making them…