On the optimization of hyperparameters in Gaussian process regression with the help of low-order high-dimensional model representation
arXiv:2112.01374 · doi:10.1007/s10910-022-01407-x
Abstract
When the data are sparse, optimization of hyperparameters of the kernel in Gaussian process regression by the commonly used maximum likelihood estimation (MLE) criterion often leads to overfitting. We show that choosing hyperparameters (in this case, kernel length parameter and regularization parameter) based on a criterion of the completeness of the basis in the corresponding linear regression problem is superior to MLE. We show that this is facilitated by the use of high-dimensional model representation (HDMR) whereby a low-order HDMR representation can provide reliable reference functions and large synthetic test data sets needed for basis parameter optimization even when the original data are few.
16 pages, 2 figures, 2 tables
References in corpus (2)
- Data-driven kinetic energy density fitting for orbital-free DFT: linear vs Gaussian process regression
- Easy representation of multivariate functions with low-dimensional terms via Gaussian process regression kernel design: applications to machine learning of potential energy surfaces and kinetic energy densities from sparse data
Cited by in corpus (5)
- Neural network with optimal neuron activation functions based on additive Gaussian process regression
- Rectangularization of Gaussian process regression for optimization of hyperparameters
- The loss of the property of locality of the kernel in high-dimensional Gaussian process regression on the example of the fitting of molecular potential energy surfaces
- Orders-of-coupling representation with a single neural network with optimal neuron activation functions and without nonlinear parameter optimization
- A kinetic-based regularization method for data science applications