3 papers
cs.LG2026
Convergent Stochastic Training of Attention and Understanding LoRA
Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi +2
Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, i…
cs.LG2026
Generalization Bounds for Physics-Informed Neural Networks for the Incompressible Navier-Stokes Equations
Sebastien Andre-Sloan, Dibyakanti Kumar, Alejandro F Frangi +1
This work establishes rigorous first-of-its-kind upper bounds on the generalization error for the method of approximating solutions to the (d+1)-dimensional incompressible Navier-S…
cs.LG2025
Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
Dibyakanti Kumar, Samyak Jha, Anirbit Mukherjee
In this work, we will establish that the Langevin Monte-Carlo algorithm can learn depth-2 neural nets of any size and for any data and we give non-asymptotic convergence rates for…