Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
Quentin Anthony, Yury Tokpanov, Skyler Szot +18
We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance…
cs.CL2024
Critical Data Size of Language Models from a Grokking Perspective
Xuekai Zhu, Yao Fu, Bowen Zhou +1
We explore the critical data size in language models, a threshold that marks a fundamental shift from quick memorization to slow generalization. We formalize the phase transition u…