2 citations · 2 across the 3 of their papers we have counts for
9 papers
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
Guanqiao Qu, Shuo Chen, Qian Chen +2
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: A…
SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge
Guanqiao Qu, Tao Li, Qian Chen +2
To support on-device inference, the next-generation mobile networks are expected to support real-time model downloading services to mobile users. However, powerful AI models typica…
PartialLoading: User Scheduling and Bandwidth Allocation for Parameter-sharing Edge Inference
Guanqiao Qu, Qian Chen, Xianhao Chen +2
By provisioning inference offloading services, edge inference drives the rapid growth of AI applications at network edge. However, how to reduce the inference latency remains a sig…
Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
Guanqiao Qu, Qiyuan Chen, Wei Wei +3
On-device large language models (LLMs), referring to running LLMs on edge devices, have raised considerable interest since they are more cost-effective, latency-efficient, and priv…
TrimCaching: Parameter-sharing AI Model Caching in Wireless Edge Networks
Guanqiao Qu, Zheng Lin, Fangming Liu +2
Next-generation mobile networks are expected to facilitate fast AI model downloading to end users. By caching models on edge servers, mobile networks can deliver models to end user…
TrimCaching: Parameter-sharing Edge Caching for AI Model Downloading
Guanqiao Qu, Zheng Lin, Qian Chen +4
Next-generation mobile networks are expected to facilitate fast AI model downloading to end users. By caching models on edge servers, mobile networks can deliver models to end user…