7 citations · 7 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
Gathering and Exploiting Higher-Order Information when Training Large Structured Models
Pierre Wolinski
When training large models, such as neural networks, the full derivatives of order 2 and beyond are usually inaccessible, due to their computational cost. Therefore, among the seco…
cs.LG2018
Learning with Random Learning Rates
Léonard Blier, Pierre Wolinski, Yann Ollivier
Hyperparameter tuning is a bothersome step in the training of deep learning models. One of the most sensitive hyperparameters is the learning rate of the gradient descent. We prese…