Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
Xiumei Deng, Zehui Xiong, Binbin Chen +3
Large language models (LLMs) are proliferating rapidly at the edge, delivering intelligent capabilities across diverse application scenarios. However, their practical deployment in…
cs.DC2025
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi +5
The parameter size of modern large language models (LLMs) can be scaled up via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computat…