2 papers
cs.LG2026
LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers
Enda Yu, Dezun Dong, Zhaoning Zhang +6
Mixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer laten…
cs.DC2025
Demystifying ARM SME to Optimize General Matrix Multiplications
Chencheng Deng, Weiling Yang, Jianbin Fang +1
General Matrix Multiplication (GEMM) is a critical kernel in high-performance computing and deep learning. While modern architectures like ARM's Scalable Matrix Extension (SME) int…