1 citations · 1 across the 7 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Federated Inference for Heterogeneous LLM Communication and Collaboration
Zihan Chen, Zeshen Li, Howard H. Yang +2
Given the limited performance and efficiency of on-device Large Language Models (LLMs), the collaborations between multiple LLMs enable desirable performance enhancements, in which…
cs.DC2025
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
Mulei Ma, Xinyi Xu, Minrui Xu +3
LLMs are increasingly executed in edge where limited GPU memory and heterogeneous computation jointly constrain deployment which motivates model partitioning and request scheduling…