Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Pruning-Aware Multi-Cluster Co-Inference for Large AI Models in AI-RANs
Xiaowen Cao, Zhonghao Lyu, Shicheng Chu +6
The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distri…
cs.DC2026
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
Yuchen Li, Rui Kong, Zhonghao Lyu +11
Deploying large language models (LLMs) in mobile and edge computing environments is constrained by limited on-device resources, scarce wireless bandwidth, and frequent model evolut…