8 papers
Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning
Thomas Chen
We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the inp…
A Theoretical Framework for Self-Play Theorem Proving Algorithms
Thomas Chen, Zhiyuan Li
Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promising empirical results in the context of formal theorem proving using Large La…
Architecture independent generalization bounds for overparametrized deep ReLU networks
Anandatheertha Bapu, Thomas Chen, Chun-Kai Kevin Chien +2
We prove that overparametrized neural networks are able to generalize with a test error that is independent of the level of overparametrization, and independent of the Vapnik-Cherv…
Gradient flow in parameter space is equivalent to linear interpolation in output space
Thomas Chen, PatrÃcia Muñoz Ewald
We prove that the standard gradient flow in parameter space that underlies many training algorithms in deep learning can be continuously deformed into an adapted gradient flow whic…
Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees
Thomas Y. Chen
We establish the first information-theoretic limits for multimodal retrieval. Casting ranking as lossy source coding, we derive a single-letter rate-distortion function for…
Non-Asymptotic Length Generalization
Thomas Chen, Tengyu Ma, Zhiyuan Li
Length generalization is the ability of a learning algorithm to learn a hypothesis which generalizes to longer inputs than the inputs in the training set. In this paper, we provide…