3 papers
cs.DC2026
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai +28
In-Network Collective (INC) acceleration holds immense potential for optimizing AI training and inference; however, its cross-layer nature has historically hindered investment and…
cs.DC2026
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
Yitao Yuan, Chenqi Zhao, Bohan Zhao +3
Efficiently harnessing GPU compute is critical to improving user experience and reducing operational costs in large language model (LLM) services. However, current inference engine…
cs.DC2026
Fine-grained MoE Load Balancing with Linear Programming
Chenqi Zhao, Wenfei Wu, Linhai Song +2
Mixture-of-Experts (MoE) has emerged as a promising approach to scale up deep learning models due to its significant reduction in computational resources. However, the dynamic natu…