4 papers
Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
Qianli Liu, Kaibin Guo, Zicong Hong +5
Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, w…
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
Jian Lin, Jiazhi Mi, Zicong Hong +5
Supporting long-context LLMs is challenging due to the substantial memory demands of the key-value (KV) cache. Existing offloading systems store the full cache in host memory and s…
Efficient Multi-user Offloading of Personalized Diffusion Models: A DRL-Convex Hybrid Solution
Wanting Yang, Zehui Xiong, Song Guo +3
With the impressive generative capabilities of diffusion models, personalized content synthesis has emerged as the most highly anticipated. However, the large model sizes and itera…
Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
Yao Zhang, Yuyi Mao, Hui Wang +5
Visual Simultaneous Localization and Mapping (vSLAM) is a prevailing technology for many emerging robotic applications. Achieving real-time SLAM on mobile robotic systems with limi…