Training Algorithm Matters for the Performance of Neural Network Potential: A Case Study of Adam and the Kalman Filter Optimizers
arXiv:2109.03769 · doi:10.1063/5.0070931
Abstract
One hidden yet important issue for developing neural network potentials (NNPs) is the choice of training algorithm. Here we compare the performance of two popular training algorithms, the adaptive moment estimation algorithm (Adam) and the Extended Kalman Filter algorithm (EKF), using the Behler-Parrinello neural network (BPNN) and two publicly accessible datasets of liquid water [Proc. Natl. Acad. Sci. U.S.A. 2016, 113, 8368-8373 and Proc. Natl. Acad. Sci. U.S.A. 2019, 116, 1110-1115]. This is achieved by implementing EKF in TensorFlow. It is found that NNPs trained with EKF are more transferable and less sensitive to the value of the learning rate, as compared to Adam. In both cases, error metrics of the validation set do not always serve as a good indicator for the actual performance of NNPs. Instead, we show that their performance correlates well with a Fisher information based similarity measure.
References in corpus (8)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- On the difficulty of training Recurrent Neural Networks
- How van der Waals interactions determine the unique properties of water
- PiNN: A Python Library for Building Atomic Neural Networks of Molecules and Materials
- Temperature effects on the ionic conductivity in concentrated alkaline electrolyte solutions
- Generalization in Deep Networks: The Role of Distance from Initialization
- Pushing the limit of molecular dynamics with ab initio accuracy to 100 million atoms with machine learning