activity
20242026
most citedFormation of Representations in Neural Networks

1 citations · 2 across the 25 of their papers we have counts for

collaborators
Showing cs.LGShow all

17 papers · 1 filter

cs.LG2026

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

Liu Ziyin, Yizhou Xu, Tomaso Poggio +1

Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training…

cs.LG2026

Edge of Stability Selectively Shapes Learning Across the Data Distribution

Shauna Kwag, Anakha Ganesh, Tomaso Poggio +1

Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning a…

cs.LG2026

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

Mohua Das, Pierfrancesco Beneventano, Shibshankar Dey +2

Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survive…

cs.LG2026

Does Weight Decay Enhance Training Stability?

Marius Saether, Amir Kolic, Tomaso Poggio +1

In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, diverging from its classical role as a static regularization penalty. We investigate a…

cs.LG2026

Too Sharp, Too Sure: When Calibration Follows Curvature

Alessandro Morosini, Matea Gjika, Tomaso Poggio +1

Modern neural networks can achieve high accuracy while remaining poorly calibrated, producing confidence estimates that do not match empirical correctness. Yet calibration is often…

cs.LG2026

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability

Arseniy Andreyev, Advikar Ananthkumar, Marc Walden +2

Recent work suggests that (stochastic) gradient descent self-organizes near an instability boundary, shaping both optimization and the solutions found. Momentum and mini-batch grad…