4 papers · 1 filter
Sparsely gated tiny linear experts
Simon Schug
Sparsity allows scaling model parameters without proportionally increasing computational cost. While mixture of experts (MoE) models are made increasingly sparse, individual expert…
Scaling can lead to compositional generalization
Florian Redhardt, Yassir Akram, Simon Schug
Can neural networks systematically capture discrete, compositional task structure despite their continuous, distributed nature? The impressive capabilities of large-scale neural ne…
Attention as a Hypernetwork
Simon Schug, Seijin Kobayashi, Yassir Akram +2
Transformers can under some circumstances generalize to novel problem instances whose constituent parts might have been encountered during training, but whose compositions have not…
When can transformers compositionally generalize in-context?
Seijin Kobayashi, Simon Schug, Yassir Akram +5
Many tasks can be composed from a few independent components. This gives rise to a combinatorial explosion of possible tasks, only some of which might be encountered during trainin…