1 citations · 2 across the 25 of their papers we have counts for
17 papers · 1 filter
Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws
Liu Ziyin, Yizhou Xu, Tomaso Poggio +1
Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training…
Edge of Stability Selectively Shapes Learning Across the Data Distribution
Shauna Kwag, Anakha Ganesh, Tomaso Poggio +1
Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning a…
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
Mohua Das, Pierfrancesco Beneventano, Shibshankar Dey +2
Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survive…
Does Weight Decay Enhance Training Stability?
Marius Saether, Amir Kolic, Tomaso Poggio +1
In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, diverging from its classical role as a static regularization penalty. We investigate a…
Too Sharp, Too Sure: When Calibration Follows Curvature
Alessandro Morosini, Matea Gjika, Tomaso Poggio +1
Modern neural networks can achieve high accuracy while remaining poorly calibrated, producing confidence estimates that do not match empirical correctness. Yet calibration is often…
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
Arseniy Andreyev, Advikar Ananthkumar, Marc Walden +2
Recent work suggests that (stochastic) gradient descent self-organizes near an instability boundary, shaping both optimization and the solutions found. Momentum and mini-batch grad…