The Renyi Gaussian Process: Towards Improved Generalization
arXiv:1910.06990 · doi:10.1080/24725854.2023.2219468
Abstract
We introduce an alternative closed form lower bound on the Gaussian process () likelihood based on the Rényi -divergence. This new lower bound can be viewed as a convex combination of the Nyström approximation and the exact . The key advantage of this bound, is its capability to control and tune the enforced regularization on the model and thus is a generalization of the traditional variational regression. From a theoretical perspective, we provide the convergence rate and risk bound for inference using our proposed approach. Experiments on real data show that the proposed algorithm may be able to deliver improvement over several inference methods.
References in corpus (19)
- Practical Bayesian Optimization of Machine Learning Algorithms
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration
- Deep Gaussian Processes
- BayesOpt: A Bayesian Optimization Library for Nonlinear Optimization, Experimental Design and Bandits
- Generalization in Deep Learning
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- When Gaussian Process Meets Big Data: A Review of Scalable GPs
- Gaussian Process Behaviour in Wide Deep Neural Networks
- Rates of Convergence for Sparse Variational Gaussian Process Regression
- Product Kernel Interpolation for Scalable Gaussian Processes
- Approximate Inference for Fully Bayesian Gaussian Process Regression
- Variational Inference of Joint Models using Multivariate Gaussian Convolution Processes
- Alpha-Beta Divergence For Variational Inference
- Robust Experimental Designs for Model Calibration
- Space-Filling Designs for Robustness Experiments
- Preconditioning via Diagonal Scaling
- Scalable Gaussian Processes with Billions of Inducing Inputs via Tensor Train Decomposition