6 papers · 1 filter
Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning
Thomas Chen
We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the inp…
A Theoretical Framework for Self-Play Theorem Proving Algorithms
Thomas Chen, Zhiyuan Li
Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promising empirical results in the context of formal theorem proving using Large La…
Architecture independent generalization bounds for overparametrized deep ReLU networks
Anandatheertha Bapu, Thomas Chen, Chun-Kai Kevin Chien +2
We prove that overparametrized neural networks are able to generalize with a test error that is independent of the level of overparametrization, and independent of the Vapnik-Cherv…
Gradient flow in parameter space is equivalent to linear interpolation in output space
Thomas Chen, PatrÃcia Muñoz Ewald
We prove that the standard gradient flow in parameter space that underlies many training algorithms in deep learning can be continuously deformed into an adapted gradient flow whic…
Non-Asymptotic Length Generalization
Thomas Chen, Tengyu Ma, Zhiyuan Li
Length generalization is the ability of a learning algorithm to learn a hypothesis which generalizes to longer inputs than the inputs in the training set. In this paper, we provide…
Zero loss guarantees and explicit minimizers for generic overparametrized Deep Learning networks
Thomas Chen, Andrew G. Moore
We determine sufficient conditions for overparametrized deep learning (DL) networks to guarantee the attainability of zero loss in the context of supervised learning, for the $\mat…