4 papers
The Geometric Structure of Models Learning Sparse Data
Thomas Walker, T. Mitchell Roddenberry, Ahmed Imtiaz Humayun +2
The manifold hypothesis (MH) is often used to explain how machine learning can overcome the curse of dimensionality. However, the MH is only applicable in regimes where the trainin…
The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts
Thomas Walker, Ahmed Imtiaz Humayun, Randall Balestriero +1
The Linear Representation Hypothesis (LRH) identifies features of a trained deep network (DN) as linear directions in the activation spaces, i.e., output spaces of intermediate lay…
GrokAlign: Geometric Characterisation and Acceleration of Grokking
Thomas Walker, Ahmed Imtiaz Humayun, Randall Balestriero +1
A key challenge for the machine learning community is to understand and accelerate the training dynamics of deep networks that lead to delayed generalisation and emergent robustnes…
Concept Boundary Vectors
Thomas Walker
Machine learning models are trained with relatively simple objectives, such as next token prediction. However, on deployment, they appear to capture a more fundamental representati…