5 papers
Feature learning is decoupled from generalization in high capacity neural networks
Niclas Alexander Göring, Charles London, Abdurrahman Hadi Erturk +3
Neural networks outperform kernel methods, sometimes by orders of magnitude, e.g. on staircase functions. This advantage stems from the ability of neural networks to learn features…
Characterising the Inductive Biases of Neural Networks on Boolean Data
Chris Mingard, Lukas Seier, Niclas Göring +3
Deep neural networks are renowned for their ability to generalise well across diverse tasks, even when heavily overparameterized. Existing works offer only partial explanations (fo…
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
Charles London, Varun Kanade
Pause tokens, simple filler symbols such as "...", consistently improve Transformer performance on both language and mathematical tasks, yet their theoretical effect remains unexpl…
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
Yoonsoo Nam, Seok Hyeong Lee, Clementine C J Domine +5
In physics, complex systems are often simplified into minimal, solvable models that retain only the core principles. In machine learning, layerwise linear models (e.g., linear neur…
Exploiting the equivalence between quantum neural networks and perceptrons
Chris Mingard, Jessica Pointing, Charles London +2
Quantum machine learning models based on parametrized quantum circuits, also called quantum neural networks (QNNs), are considered to be among the most promising candidates for app…