Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Flatland: The Adventures of Gradient Descent with Large Step Sizes
Leonardo Galli, Curtis Fox, Wiebke Bartolomaeus +2
The training of neural networks often entails objective functions that are not globally -smooth. For these functions, it is both theoretically and practically difficult to reply…
cs.LG2024
Faster Convergence for Transformer Fine-tuning with Line Search Methods
Philip Kenneweg, Leonardo Galli, Tristan Kenneweg +1
Recent works have shown that line search methods greatly increase performance of traditional stochastic gradient descent methods on a variety of datasets and architectures [1], [2]…