Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
Chris Elliott, Einar Urdshals, David Quarel +2
Singular learning theory characterizes Bayesian learning as an evolving tradeoff between accuracy and complexity, with transitions between qualitatively different solutions as samp…
cs.LG2025
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
Sean P. Fillingham, Andrew Gordon, Peter Lai +3
Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs…