collaborators

8 papers

cs.LG2026

Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning

Thomas Chen

We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the inp…

cs.LG2026

A Theoretical Framework for Self-Play Theorem Proving Algorithms

Thomas Chen, Zhiyuan Li

Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promising empirical results in the context of formal theorem proving using Large La…

cs.LG2026

Architecture independent generalization bounds for overparametrized deep ReLU networks

Anandatheertha Bapu, Thomas Chen, Chun-Kai Kevin Chien +2

We prove that overparametrized neural networks are able to generalize with a test error that is independent of the level of overparametrization, and independent of the Vapnik-Cherv…

cs.LG2026

Gradient flow in parameter space is equivalent to linear interpolation in output space

Thomas Chen, Patrícia Muñoz Ewald

We prove that the standard gradient flow in parameter space that underlies many training algorithms in deep learning can be continuously deformed into an adapted gradient flow whic…

cs.IT2025

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees

Thomas Y. Chen

We establish the first information-theoretic limits for multimodal retrieval. Casting ranking as lossy source coding, we derive a single-letter rate-distortion function for…

cs.LG2025

Non-Asymptotic Length Generalization

Thomas Chen, Tengyu Ma, Zhiyuan Li

Length generalization is the ability of a learning algorithm to learn a hypothesis which generalizes to longer inputs than the inputs in the training set. In this paper, we provide…