Showing eess.SPShow all
3 papers · 1 filter
eess.SP2026
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
Ce Zheng, Xinghan Wang, Jiahong Ning +3
Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent…
eess.SP2025
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
Jiahong Ning, Ce Zheng, Tingting Yang
Large language models (LLMs) have transformed natural language processing but face critical deployment challenges in device-edge systems due to resource limitations and communicati…
eess.SP2025
EdgePrompt: A Distributed Key-Value Inference Framework for LLMs in 6G Networks
Jiahong Ning, Pengyan Zhu, Ce Zheng +3
As sixth-generation (6G) networks advance, large language models (LLMs) are increasingly integrated into 6G infrastructure to enhance network management and intelligence. However,…