22 citations · 56 across the 25 of their papers we have counts for
30 papers
The Geometric Structure of Models Learning Sparse Data
Thomas Walker, T. Mitchell Roddenberry, Ahmed Imtiaz Humayun +2
The manifold hypothesis (MH) is often used to explain how machine learning can overcome the curse of dimensionality. However, the MH is only applicable in regimes where the trainin…
The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts
Thomas Walker, Ahmed Imtiaz Humayun, Randall Balestriero +1
The Linear Representation Hypothesis (LRH) identifies features of a trained deep network (DN) as linear directions in the activation spaces, i.e., output spaces of intermediate lay…
Is your algorithm unlearning or untraining?
Eleni Triantafillou, Ahmed Imtiaz Humayun, Monica Ribero +3
As models are getting larger and are trained on increasing amounts of data, there has been an explosion of interest into how we can ``delete'' specific data points or behaviours fr…
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
Md. Rezuwan Hassan, Azmol Hossain, Kanij Fatema +13
The Bengali language, spoken extensively across South Asia and among diasporic communities, exhibits considerable dialectal diversity shaped by geography, culture, and history. Pho…
GrokAlign: Geometric Characterisation and Acceleration of Grokking
Thomas Walker, Ahmed Imtiaz Humayun, Randall Balestriero +1
A key challenge for the machine learning community is to understand and accelerate the training dynamics of deep networks that lead to delayed generalisation and emergent robustnes…
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
Tian Jin, Ahmed Imtiaz Humayun, Utku Evci +4
Pruning eliminates unnecessary parameters in neural networks; it offers a promising solution to the growing computational demands of large language models (LLMs). While many focus…