6 papers
Interpreting Reinforcement Learning Agents with Susceptibilities
Chris Elliott, Einar Urdshals, David Quarel +1
Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We gener…
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
Chris Elliott, Einar Urdshals, David Quarel +2
Singular learning theory characterizes Bayesian learning as an evolving tradeoff between accuracy and complexity, with transitions between qualitatively different solutions as samp…
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
Ãloïse Benito-Rodriguez, Einar Urdshals, Jasmina Nasufi +1
Understanding Large Language Models (LLMs) is key to ensure their safe and beneficial deployment. This task is complicated by the difficulty of interpretability of LLM structures,…
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
Einar Urdshals, Edmund Lau, Jesse Hoogland +2
We study neural network compressibility by using singular learning theory to extend the minimum description length (MDL) principle to singular models like neural networks. Through…
Electronic structure of liquid xenon in the context of light dark matter direct detection
Riccardo Catena, Luca Marin, Marek Matas +2
We present a description of the electronic structure of xenon within the density-functional theory formalism with the goal of accurately modeling dark-matter-induced ionisation in…
Structure Development in List-Sorting Transformers
Einar Urdshals, Jasmina Urdshals
We study how a one-layer attention-only transformer develops relevant structures while learning to sort lists of numbers. At the end of training, the model organizes its attention…