7 papers
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
Guanqiao Qu, Shuo Chen, Qian Chen +2
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: A…
TrimCaching: Parameter-sharing Edge Caching for AI Model Downloading
Guanqiao Qu, Zheng Lin, Qian Chen +4
Next-generation mobile networks are expected to facilitate fast AI model downloading to end users. By caching models on edge servers, mobile networks can deliver models to end user…
SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge
Guanqiao Qu, Tao Li, Qian Chen +2
To support on-device inference, the next-generation mobile networks are expected to support real-time model downloading services to mobile users. However, powerful AI models typica…
PartialLoading: User Scheduling and Bandwidth Allocation for Parameter-sharing Edge Inference
Guanqiao Qu, Qian Chen, Xianhao Chen +2
By provisioning inference offloading services, edge inference drives the rapid growth of AI applications at network edge. However, how to reduce the inference latency remains a sig…
AdaptSFL: Adaptive Split Federated Learning in Resource-constrained Edge Networks
Zheng Lin, Guanqiao Qu, Wei Wei +2
The increasing complexity of deep neural networks poses significant barriers to democratizing them to resource-limited edge devices. To address this challenge, split federated lear…
Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities
Zheng Lin, Guanqiao Qu, Qiyuan Chen +3
Large language models (LLMs), which have shown remarkable capabilities, are revolutionizing AI development and potentially shaping our future. However, given their multimodality, t…