2 papers
cs.DC2026
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
Zhexiang Zhang, Ye Wang, Yumiao Zhao +10
Serving large Mixture-of-Experts (MoE) models is challenging because of their large memory footprints, heterogeneous resource demands, and highly dynamic inference workloads. Most…
cs.NI2024
Identification of Path Congestion Status for Network Performance Tomography using Deep Spatial-Temporal Learning
Chengze Du, Zhiwei Yu, Xiangyu Wang
Network tomography plays a crucial role in assessing the operational status of internal links within networks through end-to-end path-level measurements, independently of cooperati…