2 papers
cs.PL2026
Model2Kernel: Model-Aware Symbolic Execution For Safe CUDA Kernels
Mengting He, Shihao Xia, Haomin Jia +2
The widespread adoption of large language models (LLMs) has made GPU-accelerated inference a critical part of modern computing infrastructure. Production inference systems rely on…
cs.DC2026
Fine-grained MoE Load Balancing with Linear Programming
Chenqi Zhao, Wenfei Wu, Linhai Song +2
Mixture-of-Experts (MoE) has emerged as a promising approach to scale up deep learning models due to its significant reduction in computational resources. However, the dynamic natu…