12 papers
ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services
Yang Xu, Zihuai Xu, Hongli Xu +3
Large Language Models (LLMs) are increasingly deployed as continuously evolving services, where frequent base-model updates may invalidate previously deployed task-specific Low-Ran…
DySTop
Yizhou Shi, Qianpiao Ma, Yan Xu +4
Federated Learning (FL) has emerged as a potential distributed learning paradigm that enables model training on edge devices (i.e., workers) while preserving data privacy. However,…
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
Ying Zhu, Yang Xu, Hongli Xu +3
Training large language models (LLMs) requires massive computational resources, often necessitating the aggregation of geographically distributed data centers (\ie, cross-region tr…
Resource-Efficient Federated Fine-Tuning Large Language Models for Heterogeneous Data
Jun Liu, Yunming Liao, Hongli Xu +1
Fine-tuning large language models (LLMs) via federated learning, i.e., FedLLM, has been proposed to adapt LLMs for various downstream applications in a privacy-preserving way. To r…
A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language Models
Zuan Xie, Yang Xu, Hongli Xu +2
Recent advancements in large language models (LLMs) have catalyzed a substantial surge in demand for LLM services. While traditional cloud-based LLM services satisfy high-accuracy…
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
Zihuai Xu, Yang Xu, Hongli Xu +3
Considering the hardware-friendly characteristics and broad applicability, structured pruning has emerged as an efficient solution to reduce the resource demands of large language…