2 papers
cs.CL2025
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
Quentin Anthony, Yury Tokpanov, Skyler Szot +18
We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance…
cs.LG2024
IKUN: Initialization to Keep snn training and generalization great with sUrrogate-stable variaNce
Da Chang, Deliang Wang, Xiao Yang
Weight initialization significantly impacts the convergence and performance of neural networks. While traditional methods like Xavier and Kaiming initialization are widely used, th…