3 citations · 3 across the 3 of their papers we have counts for
3 papers
Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, Johannes Von Oswald
The effect of regularizers such as weight decay when training deep neural networks is not well understood. We study the influence of weight decay as well as -regularization whe…
Learning Randomized Algorithms with Transformers
Johannes von Oswald, Seijin Kobayashi, Yassir Akram +1
Randomization is a powerful tool that endows algorithms with remarkable properties. For instance, randomized algorithms excel in adversarial settings, often surpassing the worst-ca…
Random initialisations performing above chance and how to find them
Frederik Benzing, Simon Schug, Robert Meier +5
Neural networks trained with stochastic gradient descent (SGD) starting from different random initialisations typically find functionally very similar solutions, raising the questi…