2 citations · 2 across the 3 of their papers we have counts for
4 papers
Second-order regression models exhibit progressive sharpening to the edge of stability
Atish Agarwala, Fabian Pedregosa, Jeffrey Pennington
Recent studies of gradient descent with large step sizes have shown that there is often a regime with an initial increase in the largest eigenvalue of the loss Hessian (progressive…
One Network Fits All? Modular versus Monolithic Task Formulations in Neural Networks
Atish Agarwala, Abhimanyu Das, Brendan Juba +4
Can deep learning solve multiple tasks simultaneously, even when they are unrelated and very different? We investigate how the representations of the underlying tasks affect the ab…
Temperature check: theory and practice for training models with softmax-cross-entropy losses
Atish Agarwala, Jeffrey Pennington, Yann Dauphin +1
The softmax function combined with a cross-entropy loss is a principled approach to modeling probability distributions that has become ubiquitous in deep learning. The softmax func…
Learning the gravitational force law and other analytic functions
Atish Agarwala, Abhimanyu Das, Rina Panigrahy +1
Large neural network models have been successful in learning functions of importance in many branches of science, including physics, chemistry and biology. Recent theoretical work…