3 papers
cs.LG2026
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
Maxim Bolshim, Alexander Kugaevskikh
The loss and the norm of its gradient separate the healthy and the pathological regimes of neural-network training only weakly, whilst the curvature of the empirical risk differs q…
cs.LG2026
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
Maxim Bolshim, Alexander Kugaevskikh
Modern automatic differentiation frameworks (JAX, PyTorch) return the Hessian of the loss function as a monolithic tensor, without exposing the internal structure of inter-layer in…
cs.LG2025
Local properties of neural networks through the lens of layer-wise Hessians
Maxim Bolshim, Alexander Kugaevskikh
We introduce a methodology for analyzing neural networks through the lens of layer-wise Hessian matrices. The local Hessian of each functional block (layer) is defined as the matri…