2 papers
cs.DC2026
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
Peirong Zheng, Wenchao Xu, Haozhao Wang +2
The deployment of large language models' (LLMs) inference at the edge can facilitate prompt service responsiveness while protecting user privacy. However, it is critically challeng…
cs.DC2024
Deploying Foundation Model Powered Agent Services: A Survey
Wenchao Xu, Jinyu Chen, Peirong Zheng +8
Foundation model (FM) powered agent services are regarded as a promising solution to develop intelligent and personalized applications for advancing toward Artificial General Intel…