3 citations · 5 across the 14 of their papers we have counts for
6 papers · 1 filter
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
Tiberiu Musat, Tiago Pimentel, Nicolas Zucchet +1
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics h…
Fast and Geometrically Grounded Lorentz Neural Networks
Robert van der Klis, Ricardo Chávez Torres, Max van Spengler +3
Hyperbolic space is quickly gaining traction as a promising geometry for hierarchical and robust representation learning. A core open challenge is the development of a mathematical…
Heads collapse, features stay: Why Replay needs big buffers
Giulia Lanzillotta, Damiano Meier, Thomas Hofmann
A persistent paradox in continual learning (CL) is that neural networks often retain linearly separable representations of past tasks even when their output predictions fail. We fo…
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Denis Sutter, Julian Minder, Thomas Hofmann +1
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracte…
Causal Estimation of Memorisation Profiles
Pietro Lesci, Clara Meister, Thomas Hofmann +2
Understanding memorisation in language models has practical and societal implications, e.g., studying models' training dynamics or preventing copyright infringements. Prior work de…
The Languini Kitchen: Enabling Language Modelling Research at Different Scales of Compute
Aleksandar Stanić, Dylan Ashley, Oleg Serikov +5
The Languini Kitchen serves as both a research collective and codebase designed to empower researchers with limited computational resources to contribute meaningfully to the field…