15 citations · 18 across the 4 of their papers we have counts for
4 papers
Can In-Context Learning Support Intrinsic Curiosity?
Eric Elmoznino, Sangnie Bhardwaj, Johannes von Oswald +5
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the pro…
Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, Johannes Von Oswald
The effect of regularizers such as weight decay when training deep neural networks is not well understood. We study the influence of weight decay as well as -regularization whe…
Random initialisations performing above chance and how to find them
Frederik Benzing, Simon Schug, Robert Meier +5
Neural networks trained with stochastic gradient descent (SGD) starting from different random initialisations typically find functionally very similar solutions, raising the questi…
The least-control principle for local learning at equilibrium
Alexander Meulemans, Nicolas Zucchet, Seijin Kobayashi +2
Equilibrium systems are a powerful way to express neural computations. As special cases, they include models of great current interest in both neuroscience and machine learning, su…