Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
u-P: The Unit-Scaled Maximal Update Parametrization
Charlie Blake, Constantin Eichenberg, Josef Dean +7
The Maximal Update Parametrization (P) aims to make the optimal hyperparameters (HPs) of a model independent of its size, allowing them to be swept using a cheap proxy model ra…
cs.LG2025
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
Björn Deiseroth, Mayukh Deb, Samuel Weinbach +3
Generative transformer models have become increasingly complex, with large numbers of parameters and the ability to process multiple input modalities. Current methods for explainin…
cs.LG2024
Mechanistic Design and Scaling of Hybrid Architectures
Michael Poli, Armin W Thomas, Eric Nguyen +9
The development of deep learning architectures is a resource-demanding process, due to a vast design space, long prototyping times, and high compute costs associated with at-scale…