3 papers
cs.LG2026
A Theory of Saddle Escape in Deep Nonlinear Networks
Divit Rawal, Michael R. DeWeese
In deep networks with small initialization, training exhibits long plateaus separated by sharp feature-acquisition transitions. Whereas shallow nonlinear networks and deep linear n…
stat.ML2026
Majority-of-Three is Optimal
Divit Rawal, Nikita Zhivotovskiy
We give a short proof that the majority vote of three independent consistent classifiers is an optimal learner in the realizable PAC setting. This proves optimality for the simples…
stat.ML2026
Minimax Rates for Hyperbolic Hierarchical Learning
Divit Rawal, Sriram Vishwanath
We prove an exponential separation in sample complexity between Euclidean and hyperbolic representations for learning on hierarchical data under standard Lipschitz regularization.…