Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
-Agent: An Executable and Evolving Agent for Resource Management of Edge Generative Inference
Rui Bao, Yaping Sun, Zhiyong Chen +4
Edge deployments of generative inference increasingly face two practical realities: per-device per-model performance is often unknown at deployment time, and it is non-stationary d…
cs.LG2024
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
Nan Xue, Yaping Sun, Zhiyong Chen +6
Large Language Models (LLMs) have achieved significant success in various natural language processing tasks, but the role of wireless networks in supporting LLMs has not been thoro…