2 papers
cs.DC2025
Demystifying ARM SME to Optimize General Matrix Multiplications
Chencheng Deng, Weiling Yang, Jianbin Fang +1
General Matrix Multiplication (GEMM) is a critical kernel in high-performance computing and deep learning. While modern architectures like ARM's Scalable Matrix Extension (SME) int…
cs.LG2025
LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers
Enda Yu, Dezun Dong, Zhaoning Zhang +6
Mixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer laten…