papers

Publications (59)

math.PR2013

Nodal Sets of Random Eigenfunctions for the Isotropic Harmonic Oscillator

Boris Hanin, Steve Zelditch, Peng Zhou

We consider Gaussian random eigenfunctions (Hermite functions) of fixed energy level of the isotropic semi-classical Harmonic Oscillator on . We calculate the expected d…

math.PR2021

Random Neural Networks in the Infinite Width Limit as Gaussian Processes

Boris Hanin

This article gives a new proof that fully connected neural networks with random weights and biases converge to Gaussian processes in the regime where the input dimension, output di…

math-ph2026

The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport

Peter Halmos, Boris Hanin

The paper establishes an exact link between score‑based diffusion model sampling and adiabatic transport of ground states of specially constructed Schrödinger operators, providing…

#score-based diffusion models#adiabatic transport#schrödinger operators#density reconstruction
math.PR2016

Pairing of Zeros and Critical Points for Random Polynomials

Boris Hanin

Let p_N be a random degree N polynomial in one complex variable whose zeros are chosen independently from a fixed probability measure mu on the Riemann sphere S^2. This article pro…

cs.LG2021

How Data Augmentation affects Optimization for Linear Regression

Boris Hanin, Yi Sun

Though data augmentation has rapidly emerged as a key tool for optimization in modern machine learning, a clear picture of how augmentation schedules affect optimization and intera…

math.SP2015

Scaling Limit for the Kernel of the Spectral Projector and Remainder Estimates in the Pointwise Weyl Law

Yaiza Canzani, Boris Hanin

Let (M, g) be a compact smooth Riemannian manifold. We obtain new off-diagonal estimates as λ tend to infinity for the remainder in the pointwise Weyl Law for the kernel of the sp…

math.PR2026

Top Singular Value in Sum-Products of Random Matrices

Kevin Han Huang, Boris Hanin

We study the top singular value for a sum of independent random matrices, each of which is a product of i.i.d. Gaussian matrices. Our main conceptu…

math-ph2017

Level Spacings and Nodal Sets at Infinity for Radial Perturbations of the Harmonic Oscillator

Thomas Beck, Boris Hanin

We study properties of the nodal sets of high frequency eigenfunctions and quasimodes for radial perturbations of the Harmonic Oscillator. In particular, we consider nodal sets on…

stat.ML2024

Les Houches Lectures on Deep Learning at Large & Infinite Width

Yasaman Bahri, Boris Hanin, Antonin Brossollet +4

These lectures, presented at the 2022 Les Houches Summer School on Statistical Physics and Machine Learning, focus on the infinite-width limit and large-width regime of deep neural…

math-ph2019

Interface Asymptotics of Wigner-Weyl Distributions for the Harmonic Oscillator

Boris Hanin, Steve Zelditch

We prove several types of scaling results for Wigner distributions of spectral projections of the isotropic Harmonic oscillator on . In prior work, we studied Wigner d…

math.ST2026

Bayesian Inference with Shaped Deep Non-linear MLPs

Boris Hanin, Tianze Jiang

A central aim of deep learning theory is to characterize how neural networks make predictions in the regime of simultaneously large model and training set size. Since the limits of…

cs.AI2025

Optimizing Model Selection for Compound AI Systems

Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4

Compound AI systems that combine multiple LLM calls, such as self-refine and multi-agent-debate, achieve strong performance on many AI tasks. We address a core question in optimizi…

stat.ML2023

Bayesian Interpolation with Deep Linear Networks

Boris Hanin, Alexander Zlokapa

Characterizing how neural network depth, width, and dataset size jointly impact model quality is a central problem in deep learning theory. We give here a complete solution in the…

math.PR2018

Products of Many Large Random Matrices and Gradients in Deep Neural Networks

Boris Hanin, Mihai Nica

We study products of random matrices in the regime where the number of terms and the size of the matrices simultaneously tend to infinity. Our main theorem is that the logarithm of…

math.PR2021

Non-asymptotic Results for Singular Values of Gaussian Matrix Products

Boris Hanin, Grigoris Paouris

This article concerns the non-asymptotic analysis of the singular values (and Lyapunov exponents) of Gaussian matrix products in the regime where the number of term in the pro…

cs.LG2024

Quantitative CLTs in Deep Neural Networks

Stefano Favaro, Boris Hanin, Domenico Marinucci +2

We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant . Und…

stat.ML2023

Maximal Initial Learning Rates in Deep ReLU Networks

Gaurav Iyer, Boris Hanin, David Rolnick

Training a neural network requires choosing a suitable learning rate, which involves a trade-off between speed and effectiveness of convergence. While there has been considerable t…

cs.LG2022

Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained Analysis

Wuyang Chen, Wei Huang, Xinyu Gong +2

Advanced deep neural networks (DNNs), designed by either human or AutoML algorithms, are growing increasingly complex. Diverse operations are connected by complicated connectivity…

stat.ML2018

Which Neural Net Architectures Give Rise To Exploding and Vanishing Gradients?

Boris Hanin

We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical…

math.CV2013

Pairing of Zeros and Critical Points for Random Meromorphic Functions on Riemann Surfaces

Boris Hanin

We prove that zeros and critical points of a random polynomial of degree in one complex variable appear in pairs. More precisely, if is conditioned to have $p_N(ξ)…

cs.LG2026

Learning Rate Transfer in Normalized Transformers

Boris Shigida, Boris Hanin, Andrey Gromov

The Normalized Transformer, or nGPT (arXiv:2410.01131) achieves impressive training speedups and does not require weight decay or learning rate warmup. However, despite having hype…

cs.LG2021

The Principles of Deep Learning Theory

Daniel A. Roberts, Sho Yaida, Boris Hanin

This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks,…

math.SP2014

High Frequency Eigenfunction Immersions and Supremum Norms of Random Waves

Yaiza Canzani, Boris Hanin

A compact Riemannian manifold may be immersed into Euclidean space by using high frequency Laplace eigenfunctions. We study the geometry of the manifold viewed as a metric space en…

cs.LG2026

Hyperparameter Transfer in Graph Neural Networks

Gage DeZoort, Boris Hanin

The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and weight decay. Hyperparameter transfer aim…

stat.ML2018

How to Start Training: The Effect of Initialization and Architecture

Boris Hanin, David Rolnick

We identify and study two common failure modes for early training in deep ReLU nets. For each we give a rigorous proof of when it occurs and how to avoid it, for fully connected an…

stat.ML2023

Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Blake Bordelon, Lorenzo Noci, Mufan Bill Li +2

The cost of hyperparameter tuning in deep learning has been rising with model sizes, prompting practitioners to find new tuning methods using a proxy of smaller networks. One such…

cs.AI2026

The Future of Artificial Intelligence and the Mathematical and Physical Sciences (AI+MPS)

Andrew Ferguson, Marisa LaFleur, Lars Ruthotto +97

This community paper developed out of the NSF Workshop on the Future of Artificial Intelligence (AI) and the Mathematical and Physics Sciences (MPS), which was held in March 2025 w…

cs.LG2024

Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems

Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4

Many recent state-of-the-art results in language tasks were achieved using compound systems that perform multiple Language Model (LM) calls and aggregate their responses. However,…

math-ph2016

Scaling of Harmonic Oscillator Eigenfunctions and Their Nodal Sets Around the Caustic

Boris Hanin, Steve Zelditch, Peng Zhou

We study the scaling asymptotics of the eigenspace projection kernels of the isotropic Harmonic Oscillator of eigenvalue $E = \hbar(N +…

cs.LG2026

Hyperparameter Transfer with Mixture-of-Expert Layers

Tianze Jiang, Blake Bordelon, Cengiz Pehlevan +1

Mixture-of-Experts (MoE) layers have emerged as an important tool in scaling up modern neural networks by decoupling total trainable parameters from activated parameters in the for…

stat.ML2024

Bayesian Inference with Deep Weakly Nonlinear Networks

Boris Hanin, Alexander Zlokapa

We show at a physics level of rigor that Bayesian inference with a fully connected neural network and a shaped nonlinearity of the form is (perturbatively) so…

math.PR2012

Correlations and Pairing Between Zeros and Critical Points of Gaussian Random Polynomials

Boris Hanin

We study the asymptotics of correlations and nearest neighbor spacings between zeros and holomorphic critical points of , a degree N Hermitian Gaussian random polynomial in th…

stat.ML2019

Deep ReLU Networks Have Surprisingly Few Activation Patterns

Boris Hanin, David Rolnick

The success of deep networks has been attributed in part to their expressivity: per parameter, deep networks can approximate a richer class of functions than shallow networks. In R…

stat.ML2021

Ridgeless Interpolation with Shallow ReLU Networks in is Nearest Neighbor Curvature Extrapolation and Provably Generalizes on Lipschitz Functions

Boris Hanin

We prove a precise geometric description of all one layer ReLU networks with a single linear unit and input/output dimensions equal to one that interpolate a given datase…

stat.ML2026

Implicit Bias of the JKO Scheme

Peter Halmos, Boris Hanin

Wasserstein gradient flow provides a general framework for minimizing an energy functional over the space of probability measures on a Riemannian manifold . Its canonica…

cond-mat.dis-nn2025

Deep Neural Nets as Hamiltonians

Mike Winer, Boris Hanin

Neural networks are complex functions of both their inputs and parameters. Much prior work in deep learning theory analyzes the distribution of network outputs at a fixed a set of…

cs.LG2023

Depth Dependence of P Learning Rates in ReLU MLPs

Samy Jelassi, Boris Hanin, Ziwei Ji +3

In this short note we consider random fully connected ReLU networks of width and depth equipped with a mean-field weight initialization. Our purpose is to study the depende…

cs.AI2024

Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design

Jared Quincy Davis, Boris Hanin, Lingjiao Chen +3

As practitioners seek to surpass the current reliability and quality frontier of monolithic models, Compound AI Systems consisting of many language model inference calls are increa…

stat.ML2021

Deep ReLU Networks Preserve Expected Length

Boris Hanin, Ryan Jeong, David Rolnick

Assessing the complexity of functions computed by a neural network helps us understand how the network will learn and generalize. One natural measure of complexity is how the netwo…

cs.LG2026

Don't be lazy: CompleteP enables compute-efficient deep transformers

Nolan Dey, Bin Claire Zhang, Lorenzo Noci +6

We study compute efficiency of LLM training when using different parameterizations, i.e., rules for adjusting model and optimizer hyperparameters (HPs) as model size changes. Some…

math.PR2023

Random Fully Connected Neural Networks as Perturbatively Solvable Hierarchies

Boris Hanin

This article considers fully connected neural networks with Gaussian random weights and biases as well as hidden layers, each of width proportional to a large parameter . Fo…

math.AP2016

Nodal Sets of Smooth Functions with Finite Vanishing Order and p-Sweepouts

Thomas Beck, Spencer T. Becker-Kahn, Boris Hanin

We show that on a compact Riemmanian manifold , nodal sets of linear combinations of any smooth functions form an admissible sweepout provided these linear combina…

stat.ML2018

Approximating Continuous Functions by ReLU Nets of Minimal Width

Boris Hanin, Mark Sellke

This article concerns the expressive power of depth in deep feed-forward neural nets with ReLU activations. Specifically, we answer the following question: for a fixed $d_{in}\geq…

math.PR2020

Local Universality for Zeros and Critical Points of Monochromatic Random Waves

Yaiza Canzani, Boris Hanin

This paper concerns the asymptotic behavior of zeros and critical points for monochromatic random waves of frequency on a compact, smooth, Riemannian manifold

math.NA2020

Neural Network Approximation

Ronald DeVore, Boris Hanin, Guergana Petrova

Neural Networks (NNs) are the method of choice for building learning algorithms. Their popularity stems from their empirical success on several challenging learning problems. Howev…

math-ph2022

Scaling asymptotics of spectral Wigner functions

Boris Hanin, Steve Zelditch

We prove that smooth Wigner-Weyl spectral sums at an energy level exhibit Airy scaling asymptotics across the classical energy surface . This was proved earlier by the au…

math.PR2025

Global Universality of Singular Values in Products of Many Large Random Matrices

Boris Hanin, Tianze Jiang

We study the singular values (and Lyapunov exponents) for products of independent random matrices with i.i.d. entries. Such matrix products have been extensively an…

cs.LG2019

Finite Depth and Width Corrections to the Neural Tangent Kernel

Boris Hanin, Mihai Nica

We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation…

math-ph2014

Mean of the -norm for -normalized random waves on compact aperiodic Riemannian manifolds

Yaiza Canzani, Boris Hanin

This article concerns upper bounds for -norms of random approximate eigenfunctions of the Laplace operator on a compact aperiodic Riemannian manifold We study $f…

math.PR2018

The lemniscate tree of a random polynomial

Michael Epstein, Boris Hanin, Erik Lundberg

To each generic complex polynomial there is associated a labeled binary tree (here referred to as a "lemniscate tree") that encodes the topological type of the graph of $|p(…

cs.CL2025

BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation

Alan Zhu, Parth Asawa, Jared Quincy Davis +5

As the demand for high-quality data in model training grows, researchers and developers are increasingly generating synthetic data to tune and train LLMs. However, current data gen…

math-ph2019

Interface Asymptotics of Eigenspace Wigner distributions for the Harmonic Oscillator

Boris Hanin, Steve Zelditch

Eigenspaces of the quantum isotropic Harmonic Oscillator on have extremally high multiplicites and th…

cs.LG2024

Principled Architecture-aware Scaling of Hyperparameters

Wuyang Chen, Junru Wu, Zhangyang Wang +1

Training a high-quality deep neural network requires choosing suitable hyperparameters, which is a non-trivial and expensive process. Current works try to automatically optimize or…

cs.LG2026

Hyperparameter Transfer for Dense Associative Memories

Roi Holtzman, Dmitry Krotov, Boris Hanin

Dense Associative Memory (DenseAM) is a promising family of AI architectures that is represented by a neural network performing temporal dynamics on an energy landscape. While hype…

stat.ML2017

Universal Function Approximation by Deep Neural Nets with Bounded Width and ReLU Activations

Boris Hanin

This article concerns the expressive power of depth in neural nets with ReLU activations and bounded width. We are particularly interested in the following questions: what is the m…

cs.LG2025

Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Noam Razin, Sadhika Malladi, Adithya Bhaskar +3

Direct Preference Optimization (DPO) and its variants are increasingly used for aligning language models with human preferences. Although these methods are designed to teach a mode…

math.AP2016

C-infinity Scaling Asymptotics for the Spectral Function of the Laplacian

Yaiza Canzani, Boris Hanin

This article concerns new off-diagonal estimates on the remainder and its derivatives in the pointwise Weyl law on a compact n-dimensional Riemannian manifold. As an application, w…

stat.ML2019

Complexity of Linear Regions in Deep Networks

Boris Hanin, David Rolnick

It is well-known that the expressivity of a neural network depends on its architecture, with deeper networks expressing more complex functions. In the case of networks that compute…

stat.ML2023

Principles for Initialization and Architecture Selection in Graph Neural Networks with ReLU Activations

Gage DeZoort, Boris Hanin

This article derives and validates three principles for initialization and architecture selection in finite width graph neural networks (GNNs) with ReLU activations. First, we theo…