Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Interpreting Reinforcement Learning Agents with Susceptibilities
Chris Elliott, Einar Urdshals, David Quarel +1
Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We gener…
cs.LG2026
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
Chris Elliott, Einar Urdshals, David Quarel +2
Singular learning theory characterizes Bayesian learning as an evolving tradeoff between accuracy and complexity, with transitions between qualitatively different solutions as samp…
cs.LG2025
Structure Development in List-Sorting Transformers
Einar Urdshals, Jasmina Urdshals
We study how a one-layer attention-only transformer develops relevant structures while learning to sort lists of numbers. At the end of training, the model organizes its attention…