3 papers
cs.CL2024
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
Charles O'Neill, Thang Bui
This paper introduces an efficient and robust method for discovering interpretable circuits in large language models using discrete sparse autoencoders. Our approach addresses key…
cs.LG2024
Measuring Sharpness in Grokking
Jack Miller, Patrick Gleeson, Charles O'Neill +2
Neural networks sometimes exhibit grokking, a phenomenon where perfect or near-perfect performance is achieved on a validation set well after the same performance has been obtained…
cs.LG2023
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
Jack Miller, Charles O'Neill, Thang Bui
In some settings neural networks exhibit a phenomenon known as \textit{grokking}, where they achieve perfect or near-perfect accuracy on the validation set long after the same perf…