Publications (59)
Nodal Sets of Random Eigenfunctions for the Isotropic Harmonic Oscillator
Boris Hanin, Steve Zelditch, Peng Zhou
We consider Gaussian random eigenfunctions (Hermite functions) of fixed energy level of the isotropic semi-classical Harmonic Oscillator on . We calculate the expected d…
Random Neural Networks in the Infinite Width Limit as Gaussian Processes
Boris Hanin
This article gives a new proof that fully connected neural networks with random weights and biases converge to Gaussian processes in the regime where the input dimension, output di…
The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport
Peter Halmos, Boris Hanin
The paper establishes an exact link between score‑based diffusion model sampling and adiabatic transport of ground states of specially constructed Schrödinger operators, providing…
Pairing of Zeros and Critical Points for Random Polynomials
Boris Hanin
Let p_N be a random degree N polynomial in one complex variable whose zeros are chosen independently from a fixed probability measure mu on the Riemann sphere S^2. This article pro…
How Data Augmentation affects Optimization for Linear Regression
Boris Hanin, Yi Sun
Though data augmentation has rapidly emerged as a key tool for optimization in modern machine learning, a clear picture of how augmentation schedules affect optimization and intera…
Scaling Limit for the Kernel of the Spectral Projector and Remainder Estimates in the Pointwise Weyl Law
Yaiza Canzani, Boris Hanin
Let (M, g) be a compact smooth Riemannian manifold. We obtain new off-diagonal estimates as λ tend to infinity for the remainder in the pointwise Weyl Law for the kernel of the sp…
Top Singular Value in Sum-Products of Random Matrices
Kevin Han Huang, Boris Hanin
We study the top singular value for a sum of independent random matrices, each of which is a product of i.i.d. Gaussian matrices. Our main conceptu…
Level Spacings and Nodal Sets at Infinity for Radial Perturbations of the Harmonic Oscillator
Thomas Beck, Boris Hanin
We study properties of the nodal sets of high frequency eigenfunctions and quasimodes for radial perturbations of the Harmonic Oscillator. In particular, we consider nodal sets on…
Les Houches Lectures on Deep Learning at Large & Infinite Width
Yasaman Bahri, Boris Hanin, Antonin Brossollet +4
These lectures, presented at the 2022 Les Houches Summer School on Statistical Physics and Machine Learning, focus on the infinite-width limit and large-width regime of deep neural…
Interface Asymptotics of Wigner-Weyl Distributions for the Harmonic Oscillator
Boris Hanin, Steve Zelditch
We prove several types of scaling results for Wigner distributions of spectral projections of the isotropic Harmonic oscillator on . In prior work, we studied Wigner d…
Bayesian Inference with Shaped Deep Non-linear MLPs
Boris Hanin, Tianze Jiang
A central aim of deep learning theory is to characterize how neural networks make predictions in the regime of simultaneously large model and training set size. Since the limits of…
Optimizing Model Selection for Compound AI Systems
Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4
Compound AI systems that combine multiple LLM calls, such as self-refine and multi-agent-debate, achieve strong performance on many AI tasks. We address a core question in optimizi…
Bayesian Interpolation with Deep Linear Networks
Boris Hanin, Alexander Zlokapa
Characterizing how neural network depth, width, and dataset size jointly impact model quality is a central problem in deep learning theory. We give here a complete solution in the…
Products of Many Large Random Matrices and Gradients in Deep Neural Networks
Boris Hanin, Mihai Nica
We study products of random matrices in the regime where the number of terms and the size of the matrices simultaneously tend to infinity. Our main theorem is that the logarithm of…
Non-asymptotic Results for Singular Values of Gaussian Matrix Products
Boris Hanin, Grigoris Paouris
This article concerns the non-asymptotic analysis of the singular values (and Lyapunov exponents) of Gaussian matrix products in the regime where the number of term in the pro…
Quantitative CLTs in Deep Neural Networks
Stefano Favaro, Boris Hanin, Domenico Marinucci +2
We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant . Und…
Maximal Initial Learning Rates in Deep ReLU Networks
Gaurav Iyer, Boris Hanin, David Rolnick
Training a neural network requires choosing a suitable learning rate, which involves a trade-off between speed and effectiveness of convergence. While there has been considerable t…
Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained Analysis
Wuyang Chen, Wei Huang, Xinyu Gong +2
Advanced deep neural networks (DNNs), designed by either human or AutoML algorithms, are growing increasingly complex. Diverse operations are connected by complicated connectivity…
Which Neural Net Architectures Give Rise To Exploding and Vanishing Gradients?
Boris Hanin
We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical…
Pairing of Zeros and Critical Points for Random Meromorphic Functions on Riemann Surfaces
Boris Hanin
We prove that zeros and critical points of a random polynomial of degree in one complex variable appear in pairs. More precisely, if is conditioned to have $p_N(ξ)…
Learning Rate Transfer in Normalized Transformers
Boris Shigida, Boris Hanin, Andrey Gromov
The Normalized Transformer, or nGPT (arXiv:2410.01131) achieves impressive training speedups and does not require weight decay or learning rate warmup. However, despite having hype…
The Principles of Deep Learning Theory
Daniel A. Roberts, Sho Yaida, Boris Hanin
This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks,…
High Frequency Eigenfunction Immersions and Supremum Norms of Random Waves
Yaiza Canzani, Boris Hanin
A compact Riemannian manifold may be immersed into Euclidean space by using high frequency Laplace eigenfunctions. We study the geometry of the manifold viewed as a metric space en…
Hyperparameter Transfer in Graph Neural Networks
Gage DeZoort, Boris Hanin
The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and weight decay. Hyperparameter transfer aim…
How to Start Training: The Effect of Initialization and Architecture
Boris Hanin, David Rolnick
We identify and study two common failure modes for early training in deep ReLU nets. For each we give a rigorous proof of when it occurs and how to avoid it, for fully connected an…
Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit
Blake Bordelon, Lorenzo Noci, Mufan Bill Li +2
The cost of hyperparameter tuning in deep learning has been rising with model sizes, prompting practitioners to find new tuning methods using a proxy of smaller networks. One such…
The Future of Artificial Intelligence and the Mathematical and Physical Sciences (AI+MPS)
Andrew Ferguson, Marisa LaFleur, Lars Ruthotto +97
This community paper developed out of the NSF Workshop on the Future of Artificial Intelligence (AI) and the Mathematical and Physics Sciences (MPS), which was held in March 2025 w…
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4
Many recent state-of-the-art results in language tasks were achieved using compound systems that perform multiple Language Model (LM) calls and aggregate their responses. However,…
Scaling of Harmonic Oscillator Eigenfunctions and Their Nodal Sets Around the Caustic
Boris Hanin, Steve Zelditch, Peng Zhou
We study the scaling asymptotics of the eigenspace projection kernels of the isotropic Harmonic Oscillator of eigenvalue $E = \hbar(N +…
Hyperparameter Transfer with Mixture-of-Expert Layers
Tianze Jiang, Blake Bordelon, Cengiz Pehlevan +1
Mixture-of-Experts (MoE) layers have emerged as an important tool in scaling up modern neural networks by decoupling total trainable parameters from activated parameters in the for…
Bayesian Inference with Deep Weakly Nonlinear Networks
Boris Hanin, Alexander Zlokapa
We show at a physics level of rigor that Bayesian inference with a fully connected neural network and a shaped nonlinearity of the form is (perturbatively) so…
Correlations and Pairing Between Zeros and Critical Points of Gaussian Random Polynomials
Boris Hanin
We study the asymptotics of correlations and nearest neighbor spacings between zeros and holomorphic critical points of , a degree N Hermitian Gaussian random polynomial in th…
Deep ReLU Networks Have Surprisingly Few Activation Patterns
Boris Hanin, David Rolnick
The success of deep networks has been attributed in part to their expressivity: per parameter, deep networks can approximate a richer class of functions than shallow networks. In R…
Ridgeless Interpolation with Shallow ReLU Networks in is Nearest Neighbor Curvature Extrapolation and Provably Generalizes on Lipschitz Functions
Boris Hanin
We prove a precise geometric description of all one layer ReLU networks with a single linear unit and input/output dimensions equal to one that interpolate a given datase…
Implicit Bias of the JKO Scheme
Peter Halmos, Boris Hanin
Wasserstein gradient flow provides a general framework for minimizing an energy functional over the space of probability measures on a Riemannian manifold . Its canonica…
Deep Neural Nets as Hamiltonians
Mike Winer, Boris Hanin
Neural networks are complex functions of both their inputs and parameters. Much prior work in deep learning theory analyzes the distribution of network outputs at a fixed a set of…
Depth Dependence of P Learning Rates in ReLU MLPs
Samy Jelassi, Boris Hanin, Ziwei Ji +3
In this short note we consider random fully connected ReLU networks of width and depth equipped with a mean-field weight initialization. Our purpose is to study the depende…
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design
Jared Quincy Davis, Boris Hanin, Lingjiao Chen +3
As practitioners seek to surpass the current reliability and quality frontier of monolithic models, Compound AI Systems consisting of many language model inference calls are increa…
Deep ReLU Networks Preserve Expected Length
Boris Hanin, Ryan Jeong, David Rolnick
Assessing the complexity of functions computed by a neural network helps us understand how the network will learn and generalize. One natural measure of complexity is how the netwo…
Don't be lazy: CompleteP enables compute-efficient deep transformers
Nolan Dey, Bin Claire Zhang, Lorenzo Noci +6
We study compute efficiency of LLM training when using different parameterizations, i.e., rules for adjusting model and optimizer hyperparameters (HPs) as model size changes. Some…
Random Fully Connected Neural Networks as Perturbatively Solvable Hierarchies
Boris Hanin
This article considers fully connected neural networks with Gaussian random weights and biases as well as hidden layers, each of width proportional to a large parameter . Fo…
Nodal Sets of Smooth Functions with Finite Vanishing Order and p-Sweepouts
Thomas Beck, Spencer T. Becker-Kahn, Boris Hanin
We show that on a compact Riemmanian manifold , nodal sets of linear combinations of any smooth functions form an admissible sweepout provided these linear combina…
Approximating Continuous Functions by ReLU Nets of Minimal Width
Boris Hanin, Mark Sellke
This article concerns the expressive power of depth in deep feed-forward neural nets with ReLU activations. Specifically, we answer the following question: for a fixed $d_{in}\geq…
Local Universality for Zeros and Critical Points of Monochromatic Random Waves
Yaiza Canzani, Boris Hanin
This paper concerns the asymptotic behavior of zeros and critical points for monochromatic random waves of frequency on a compact, smooth, Riemannian manifold …
Neural Network Approximation
Ronald DeVore, Boris Hanin, Guergana Petrova
Neural Networks (NNs) are the method of choice for building learning algorithms. Their popularity stems from their empirical success on several challenging learning problems. Howev…
Scaling asymptotics of spectral Wigner functions
Boris Hanin, Steve Zelditch
We prove that smooth Wigner-Weyl spectral sums at an energy level exhibit Airy scaling asymptotics across the classical energy surface . This was proved earlier by the au…
Global Universality of Singular Values in Products of Many Large Random Matrices
Boris Hanin, Tianze Jiang
We study the singular values (and Lyapunov exponents) for products of independent random matrices with i.i.d. entries. Such matrix products have been extensively an…
Finite Depth and Width Corrections to the Neural Tangent Kernel
Boris Hanin, Mihai Nica
We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation…
Mean of the -norm for -normalized random waves on compact aperiodic Riemannian manifolds
Yaiza Canzani, Boris Hanin
This article concerns upper bounds for -norms of random approximate eigenfunctions of the Laplace operator on a compact aperiodic Riemannian manifold We study $f…
The lemniscate tree of a random polynomial
Michael Epstein, Boris Hanin, Erik Lundberg
To each generic complex polynomial there is associated a labeled binary tree (here referred to as a "lemniscate tree") that encodes the topological type of the graph of $|p(…
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
Alan Zhu, Parth Asawa, Jared Quincy Davis +5
As the demand for high-quality data in model training grows, researchers and developers are increasingly generating synthetic data to tune and train LLMs. However, current data gen…
Interface Asymptotics of Eigenspace Wigner distributions for the Harmonic Oscillator
Boris Hanin, Steve Zelditch
Eigenspaces of the quantum isotropic Harmonic Oscillator on have extremally high multiplicites and th…
Principled Architecture-aware Scaling of Hyperparameters
Wuyang Chen, Junru Wu, Zhangyang Wang +1
Training a high-quality deep neural network requires choosing suitable hyperparameters, which is a non-trivial and expensive process. Current works try to automatically optimize or…
Hyperparameter Transfer for Dense Associative Memories
Roi Holtzman, Dmitry Krotov, Boris Hanin
Dense Associative Memory (DenseAM) is a promising family of AI architectures that is represented by a neural network performing temporal dynamics on an energy landscape. While hype…
Universal Function Approximation by Deep Neural Nets with Bounded Width and ReLU Activations
Boris Hanin
This article concerns the expressive power of depth in neural nets with ReLU activations and bounded width. We are particularly interested in the following questions: what is the m…
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
Noam Razin, Sadhika Malladi, Adithya Bhaskar +3
Direct Preference Optimization (DPO) and its variants are increasingly used for aligning language models with human preferences. Although these methods are designed to teach a mode…
C-infinity Scaling Asymptotics for the Spectral Function of the Laplacian
Yaiza Canzani, Boris Hanin
This article concerns new off-diagonal estimates on the remainder and its derivatives in the pointwise Weyl law on a compact n-dimensional Riemannian manifold. As an application, w…
Complexity of Linear Regions in Deep Networks
Boris Hanin, David Rolnick
It is well-known that the expressivity of a neural network depends on its architecture, with deeper networks expressing more complex functions. In the case of networks that compute…
Principles for Initialization and Architecture Selection in Graph Neural Networks with ReLU Activations
Gage DeZoort, Boris Hanin
This article derives and validates three principles for initialization and architecture selection in finite width graph neural networks (GNNs) with ReLU activations. First, we theo…