6 papers
ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space
Gabe Guo, Thanawat Sornwanee, Lutong Hao +3
Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundament…
A Theory of Generalization in Deep Learning
Elon Litman, Gabe Guo
We present a non-asymptotic theory of generalization in deep learning where the empirical neural tangent kernel partitions the output space. In directions corresponding to signal,…
The Origin of Edge of Stability
Elon Litman
Full-batch gradient descent on neural networks drives the largest Hessian eigenvalue to the threshold , where is the learning rate. This phenomenon, the Edge of Stabilit…
You Need Better Attention Priors
Elon Litman, Gabe Guo
We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized…
Scaled-Dot-Product Attention as One-Sided Entropic Optimal Transport
Elon Litman
The scaled-dot-product attention (SDPA) mechanism is a core component of modern deep learning, but its mathematical form is often motivated by heuristics. This work provides a firs…
Finite-Nudge Equilibrium Propagation in Thermal Ensembles
Elon Litman
We liberate Equilibrium Propagation (EP) from the limit of infinitesimal perturbations by establishing a finite-nudge foundation for local credit assignment. By modeling network st…