5 papers
CR^2: Cost-Aware Risk-Controlled Routing for Wireless Device-Edge LLM Inference
Nan Xue, Shengkang Chen, Zhiyong Chen +4
As large language models (LLMs) move from centralized clouds to mobile edge environments, efficient serving must balance latency, energy consumption, and accuracy under constrained…
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
Zixuan Liu, Zhiyong Chen, Nan Xue +4
While distributed device-edge speculative decoding enhances resource utilization across heterogeneous nodes, its performance is often bottlenecked by conventional token-level verif…
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
Rui Bao, Nan Xue, Yaping Sun +1
The integration of wireless communications and Large Language Models (LLMs) is poised to unlock ubiquitous intelligent services, yet deploying them in wireless edge-device collabor…
CSGO: Generalized Optimization for Cold Start in Wireless Collaborative Edge LLM Systems
Xuran Liu, Nan Xue, Rui Bao +5
While deploying large language models on edge devices promises low-latency and privacy-preserving AI services, it is hindered by limited device resources. Although pipeline paralle…
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
Nan Xue, Yaping Sun, Zhiyong Chen +6
Large Language Models (LLMs) have achieved significant success in various natural language processing tasks, but the role of wireless networks in supporting LLMs has not been thoro…