3 papers
cs.LG2025
Scaling can lead to compositional generalization
Florian Redhardt, Yassir Akram, Simon Schug
Can neural networks systematically capture discrete, compositional task structure despite their continuous, distributed nature? The impressive capabilities of large-scale neural ne…
cs.LG2024
Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, Johannes Von Oswald
The effect of regularizers such as weight decay when training deep neural networks is not well understood. We study the influence of weight decay as well as -regularization whe…
cs.LG2024
Learning Randomized Algorithms with Transformers
Johannes von Oswald, Seijin Kobayashi, Yassir Akram +1
Randomization is a powerful tool that endows algorithms with remarkable properties. For instance, randomized algorithms excel in adversarial settings, often surpassing the worst-ca…