4 papers · 1 filter
Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN
Tianhang Ding, Jianchun Liu, Hongli Xu
AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target…
Adaptive and Fine-grained Module-wise Expert Pruning for Efficient LoRA-MoE Fine-Tuning
Weihang Li, Jianchun Liu, Hongli Xu
LoRA-MoE has emerged as an effective paradigm for parameter-efficient fine-tuning, combining the low training cost of LoRA with the increased adaptation capacity of Mixture-of-Expe…
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
Jiaming Yan, Jianchun Liu, Hongli Xu +1
Mixture-of-Experts (MoE) has emerged as a promising architecture for modern large language models (LLMs). However, massive parameters impose heavy GPU memory (i.e., VRAM) demands,…
Mitigating Catastrophic Forgetting with Adaptive Transformer Block Expansion in Federated Fine-Tuning
Yujia Huo, Jianchun Liu, Hongli Xu +3
Federated fine-tuning (FedFT) of large language models (LLMs) has emerged as a promising solution for adapting models to distributed data environments while ensuring data privacy.…