6 papers
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
Yizhou Xu, Pierfrancesco Beneventano, Isaac Chuang +1
A large body of theory and empirical work hypothesizes a connection between the flatness of a neural network's loss landscape during training and its performance. However, there ha…
Topological Invariance and Breakdown in Learning
Yongyi Yang, Tomaso Poggio, Isaac Chuang +1
We prove that for a broad class of permutation-equivariant learning rules (including SGD, Adam, and others), the training process induces a bi-Lipschitz mapping between neurons and…
Proof of a perfect platonic representation hypothesis
Liu Ziyin, Isaac Chuang
In this note, we elaborate on and explain in detail the proof given by Ziyin et al. (2025) of the ``perfect" Platonic Representation Hypothesis (PRH) for the embedded deep linear n…
Heterosynaptic Circuits Are Universal Gradient Machines
Liu Ziyin, Isaac Chuang, Tomaso Poggio
We propose a design principle for the learning circuits of the biological brain. The principle states that almost any dendritic weights updated via heterosynaptic plasticity can im…
Neural Thermodynamics: Entropic Forces in Deep and Universal Representation Learning
Liu Ziyin, Yizhou Xu, Isaac Chuang
With the rapid discovery of emergent phenomena in deep learning and large language models, understanding their cause has become an urgent need. Here, we propose a rigorous entropic…
Parameter Symmetry Potentially Unifies Deep Learning Theory
Liu Ziyin, Yizhou Xu, Tomaso Poggio +1
The dynamics of learning in modern large AI systems is hierarchical, often characterized by abrupt, qualitative shifts akin to phase transitions observed in physical systems. While…