5 papers
The Spectral Dynamics and Noise Geometry of Muon
Pierfrancesco Beneventano, Mahmoud Abdelmoneum, Tomaso Poggio
Muon replaces a matrix gradient by its polar factor . This keeps the singular directions selected by the gradient, but makes the update spectrum flat. We stu…
Learning Sparse Compositional Functions with Norm-Constrained Neural Networks
Shuo Huang, Lorenzo Fiorito, Lorenzo Rosasco +1
The ability of deep neural networks to learn hierarchical features is widely regarded as a key mechanism underlying their success in high-dimensional learning. Existing theory part…
A universal compression theory for lottery ticket hypothesis and neural scaling laws
Hong-Yi Wang, Di Luo, Tomaso Poggio +2
When training large-scale models, the performance typically scales with the number of parameters and the dataset size according to a slow power law. A fundamental theoretical and p…
Hierarchical Reasoning Models: Perspectives and Misconceptions
Renee Ge, Qianli Liao, Tomaso Poggio
Transformers have demonstrated remarkable performance in natural language processing and related domains, as they largely focus on sequential, autoregressive next-token prediction…
Self-Assembly of a Biologically Plausible Learning Circuit
Qianli Liao, Liu Ziyin, Yulu Gan +3
Over the last four decades, the amazing success of deep learning has been driven by the use of Stochastic Gradient Descent (SGD) as the main optimization technique. The default imp…