4 papers
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
Maxim Bolshim, Alexander Kugaevskikh
The loss and the norm of its gradient separate the healthy and the pathological regimes of neural-network training only weakly, whilst the curvature of the empirical risk differs q…
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
Maxim Bolshim, Alexander Kugaevskikh
Modern automatic differentiation frameworks (JAX, PyTorch) return the Hessian of the loss function as a monolithic tensor, without exposing the internal structure of inter-layer in…
Local properties of neural networks through the lens of layer-wise Hessians
Maxim Bolshim, Alexander Kugaevskikh
We introduce a methodology for analyzing neural networks through the lens of layer-wise Hessian matrices. The local Hessian of each functional block (layer) is defined as the matri…
Wasserstein Regression as a Variational Approximation of Probabilistic Trajectories through the Bernstein Basis
Maksim Maslov, Alexander Kugaevskikh, Matthew Ivanov
This paper considers the problem of regression over distributions, which is becoming increasingly important in machine learning. Existing approaches often ignore the geometry of th…