3 papers
cs.DC2026
From Cloud to Crowd: Democratizing LLM Service with Decentralized Edge Collaboration for RAG
Jiaxing Li, Hengzhi Wang, Feng Wang +5
The rapid advancement of large language models (LLMs) has increased demand for scalable and cost-effective deployment, especially for mobile and edge devices. Cloud-hosted LLMs are…
cs.CL2025
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
Chi Xu, Gefei Zhang, Yantong Zhu +4
N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…
cs.DC2025
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
Jiaxing Li, Chi Xu, Lianchen Jia +3
Large language models (LLMs) have demonstrated impressive capabilities in language tasks, but they require high computing power and rely on static knowledge. To overcome these limi…