3 papers
cs.LG2025
Scaling can lead to compositional generalization
Florian Redhardt, Yassir Akram, Simon Schug
Can neural networks systematically capture discrete, compositional task structure despite their continuous, distributed nature? The impressive capabilities of large-scale neural ne…
cs.LG2025
Attention as a Hypernetwork
Simon Schug, Seijin Kobayashi, Yassir Akram +2
Transformers can under some circumstances generalize to novel problem instances whose constituent parts might have been encountered during training, but whose compositions have not…
cs.LG2024
Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, Johannes Von Oswald
The effect of regularizers such as weight decay when training deep neural networks is not well understood. We study the influence of weight decay as well as -regularization whe…