3 papers
cs.LG2025
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
Tuo Zhang, Ning Li, Xin Yuan +4
With the breakthrough progress of large language models (LLMs) in natural language processing and multimodal tasks, efficiently deploying them on resource-constrained edge devices…
cs.LG2025
FedPaI: Achieving Extreme Sparsity in Federated Learning via Pruning at Initialization
Haonan Wang, Zeli Liu, Kajimusugura Hoshino +3
Federated Learning (FL) enables distributed training on edge devices but faces significant challenges due to resource constraints in edge environments, impacting both communication…
cs.NI2025
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
Ning Li, Song Guo, Tuo Zhang +5
The powerfulness of LLMs indicates that deploying various LLMs with different scales and architectures on end, edge, and cloud to satisfy different requirements and adaptive hetero…