5 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 5 cited
Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington +1
The success of the Adam optimizer on a wide array of architectures has made it the default in settings where stochastic gradient descent (SGD) performs poorly. However, our theoret…
cs.LG2021
Fast Sparse Decision Tree Optimization via Reference Ensembles
Hayden McTavish, Chudi Zhong, Reto Achermann +4
Sparse decision tree optimization has been one of the most fundamental problems in AI since its inception and is a challenge at the core of interpretable machine learning. Sparse d…