3 papers
cs.LG2026
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
Gang Liao, Hongsen Qin, Ying Wang +36
Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key system challenges - model architecture div…
cs.AR2026
The xPU-athalon: Quantifying the Competition of AI Acceleration
Alicia Golden, Carole-Jean Wu, Gu-Yeon Wei +1
The push for greater efficiency in AI computation has given rise to an array of accelerator architectures that increasingly challenge the GPU's long-standing dominance. In this wor…
cs.DC2026
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
Alicia Golden, Michael Kuchnik, Samuel Hsia +4
Large model training beyond tens of thousands of GPUs is an uncharted territory. At such scales, disruptions to the training process are not a matter of if, but a matter of when --…