Tangent-Space Gradient Optimization of Tensor Network for Machine Learning
arXiv:2001.04029 · doi:10.1103/PhysRevE.102.012152
Abstract
The gradient-based optimization method for deep machine learning models suffers from gradient vanishing and exploding problems, particularly when the computational graph becomes deep. In this work, we propose the tangent-space gradient optimization (TSGO) for the probabilistic models to keep the gradients from vanishing or exploding. The central idea is to guarantee the orthogonality between the variational parameters and the gradients. The optimization is then implemented by rotating parameter vector towards the direction of gradient. We explain and testify TSGO in tensor network (TN) machine learning, where the TN describes the joint probability distribution as a normalized state in Hilbert space. We show that the gradient can be restricted in the tangent space of hyper-sphere. Instead of additional adaptive methods to control the learning rate in deep learning, the learning rate of TSGO is naturally determined by the angle as . Our numerical results reveal better convergence of TSGO in comparison to the off-the-shelf Adam.
5 pages, 4 figures
References in corpus (5)
- The density-matrix renormalization group in the age of matrix product states
- Matrix Product States, Projected Entangled Pair States, and variational renormalization group methods for quantum spin systems
- Classical simulation of infinite-size quantum lattice systems in two spatial dimensions
- Scaling of entanglement support for Matrix Product States
- Tree Tensor Networks for Generative Modeling
Cited by in corpus (9)
- The Presence and Absence of Barren Plateaus in Tensor-network Based Machine Learning
- Generative machine learning with tensor networks: benchmarks on near-term quantum computers
- Tensor networks for interpretable and efficient quantum-inspired machine learning
- Residual Matrix Product State for Machine Learning
- Tensor network to learn the wavefunction of data
- Quantum-Classical Machine learning by Hybrid Tensor Networks
- Non-parametric Semi-Supervised Learning in Many-body Hilbert Space with Rescaled Logarithmic Fidelity
- Bayesian Tensor Network with Polynomial Complexity for Probabilistic Machine Learning
- Using matrix-product states for time-series machine learning