2 papers
cs.LG2025
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
Atli Kosson, Jeremy Welborn, Yang Liu +2
Transferring the optimal learning rate from small to large neural networks can enable efficient training at scales where hyperparameter tuning is otherwise prohibitively expensive.…
cs.LG2019
Learning Index Selection with Structured Action Spaces
Jeremy Welborn, Michael Schaarschmidt, Eiko Yoneki
Configuration spaces for computer systems can be challenging for traditional and automatic tuning strategies. Injecting task-specific knowledge into the tuner for a task may allow…