16 papers
Distribution-Aware Algorithm Design with LLM Agents
Saharsh Koganti, Priyadarsi Mishra, Pierfrancesco Beneventano +1
Many optimization problems arise repeatedly from a fixed but unknown distribution. Even when the worst-case problem is hard, this distribution may carry reusable structure, such as…
Edge of Stability Selectively Shapes Learning Across the Data Distribution
Shauna Kwag, Anakha Ganesh, Tomaso Poggio +1
Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning a…
The Spectral Dynamics and Noise Geometry of Muon
Pierfrancesco Beneventano, Mahmoud Abdelmoneum, Tomaso Poggio
Muon replaces a matrix gradient by its polar factor . This keeps the singular directions selected by the gradient, but makes the update spectrum flat. We stu…
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
Mohua Das, Pierfrancesco Beneventano, Shibshankar Dey +2
Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survive…
Retrieval Is Not Enough: Why Organizational AI Needs Epistemic Infrastructure
Federico Bottino, Carlo Ferrero, Nicholas Dosio +1
Organizational knowledge used by AI agents typically lacks epistemic structure: retrieval systems surface semantically relevant content without distinguishing binding decisions fro…
Does Weight Decay Enhance Training Stability?
Marius Saether, Amir Kolic, Tomaso Poggio +1
In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, diverging from its classical role as a static regularization penalty. We investigate a…