1 citations · 1 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
AXLearn: Modular, Hardware-Agnostic Large Model Training
Mark Lee, Chang Lan, Tom Gunter +34
AXLearn is a production system which facilitates scalable and high-performance training of large deep learning models. Compared to other state-of-art deep learning systems, AXLearn…
cs.LG2024★ 1 cited
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
Dongseong Hwang
This paper establishes a mathematical foundation for the Adam optimizer, elucidating its connection to natural gradient descent through Riemannian and information geometry. We prov…