26 citations · 76 across the 18 of their papers we have counts for
4 papers · 1 filter
Using Large Language Models for Hyperparameter Optimization
Michael R. Zhang, Nishkrit Desai, Juhan Bae +2
This paper explores the use of foundational large language models (LLMs) in hyperparameter optimization (HPO). Hyperparameters are critical in determining the effectiveness of mach…
Studying Large Language Model Generalization with Influence Functions
Roger Grosse, Juhan Bae, Cem Anil +14
When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which tr…
Benchmarking Neural Network Training Algorithms
George E. Dahl, Frank Schneider, Zachary Nado +22
Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workl…
Efficient Parametric Approximations of Neural Network Function Space Distance
Nikita Dhawan, Sicong Huang, Juhan Bae +1
It is often useful to compactly summarize important properties of model parameters and training data so that they can be used later without storing and/or iterating over the entire…