3 papers
math.ST2026
Bayesian Inference with Shaped Deep Non-linear MLPs
Boris Hanin, Tianze Jiang
A central aim of deep learning theory is to characterize how neural networks make predictions in the regime of simultaneously large model and training set size. Since the limits of…
cs.LG2026
Hyperparameter Transfer with Mixture-of-Expert Layers
Tianze Jiang, Blake Bordelon, Cengiz Pehlevan +1
Mixture-of-Experts (MoE) layers have emerged as an important tool in scaling up modern neural networks by decoupling total trainable parameters from activated parameters in the for…
math.PR2025
Global Universality of Singular Values in Products of Many Large Random Matrices
Boris Hanin, Tianze Jiang
We study the singular values (and Lyapunov exponents) for products of independent random matrices with i.i.d. entries. Such matrix products have been extensively an…