2 papers
cs.CL2026
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification
Ji Liu, Puyuan Yang, Rongzhang Zheng +18
High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizatio…
cs.AR2026
Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng +2
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kern…