collaborators

6 papers

cs.LG2026

ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

Gabe Guo, Thanawat Sornwanee, Lutong Hao +3

Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundament…

cs.LG2026

A Theory of Generalization in Deep Learning

Elon Litman, Gabe Guo

We present a non-asymptotic theory of generalization in deep learning where the empirical neural tangent kernel partitions the output space. In directions corresponding to signal,…

cs.LG2026

The Origin of Edge of Stability

Elon Litman

Full-batch gradient descent on neural networks drives the largest Hessian eigenvalue to the threshold , where is the learning rate. This phenomenon, the Edge of Stabilit…

cs.LG2026

You Need Better Attention Priors

Elon Litman, Gabe Guo

We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized…

cs.LG2025

Scaled-Dot-Product Attention as One-Sided Entropic Optimal Transport

Elon Litman

The scaled-dot-product attention (SDPA) mechanism is a core component of modern deep learning, but its mathematical form is often motivated by heuristics. This work provides a firs…

cs.LG2025

Finite-Nudge Equilibrium Propagation in Thermal Ensembles

Elon Litman

We liberate Equilibrium Propagation (EP) from the limit of infinitesimal perturbations by establishing a finite-nudge foundation for local credit assignment. By modeling network st…