6 papers
Riemannian Optimization over Symmetric Positive Definite Matrices with the Alpha-Procrustes Geometry
Derun Zhou, Keisuke Yano, Mahito Sugiyama
In Riemannian optimization, it is well known that the condition number of the Riemannian Hessian at an optimum strongly influences the asymptotic convergence behavior of optimizati…
A Complete Decomposition of KL Error using Refined Information and Mode Interaction Selection
James Enouen, Mahito Sugiyama
The log-linear model has received a significant amount of theoretical attention in previous decades and remains the fundamental tool used for learning probability distributions ove…
Same Graph, Different Likelihoods: Calibration of Autoregressive Graph Generators via Permutation-Equivalent Encodings
Laurits Fredsgaard, Aaron Thomas, Michael Riis Andersen +2
Autoregressive graph generators define likelihoods via a sequential construction process, but these likelihoods are only meaningful if they are consistent across all linearizations…
Quadratic polarity and polar Fenchel-Young divergences from the canonical Legendre polarity
Frank Nielsen, Basile Plus-Gourdon, Mahito Sugiyama
Polarity is a fundamental reciprocal duality of -dimensional projective geometry which associates to points polar hyperplanes, and more generally -dimensional convex bodies t…
New Evidence of the Two-Phase Learning Dynamics of Neural Networks
Zhanpeng Zhou, Yongyi Yang, Mahito Sugiyama +1
Understanding how deep neural networks learn remains a fundamental challenge in modern machine learning. A growing body of evidence suggests that training dynamics undergo a distin…
On the Cone Effect in the Learning Dynamics
Zhanpeng Zhou, Yongyi Yang, Jie Ren +2
Understanding the learning dynamics of neural networks is a central topic in the deep learning community. In this paper, we take an empirical perspective to study the learning dyna…