4 papers
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
Ce Zheng, Xinghan Wang, Jiahong Ning +3
Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent…
Wireless large AI model: shaping the AI-empowered future of 6G and beyond
Fenghao Zhu, Xinquan Wang, Siming Jiang +22
The emergence of sixth-generation and beyond communication systems is expected to fundamentally transform digital experiences through introducing unparalleled levels of intelligenc…
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
Jiahong Ning, Ce Zheng, Tingting Yang
Large language models (LLMs) have transformed natural language processing but face critical deployment challenges in device-edge systems due to resource limitations and communicati…
EdgePrompt: A Distributed Key-Value Inference Framework for LLMs in 6G Networks
Jiahong Ning, Pengyan Zhu, Ce Zheng +3
As sixth-generation (6G) networks advance, large language models (LLMs) are increasingly integrated into 6G infrastructure to enhance network management and intelligence. However,…