Permutation and Grouping Methods for Sharpening Gaussian Process Approximations
arXiv:1609.05372 · doi:10.1080/00401706.2018.1437476
Abstract
Vecchia's approximate likelihood for Gaussian process parameters depends on how the observations are ordered, which can be viewed as a deficiency because the exact likelihood is permutation-invariant. This article takes the alternative standpoint that the ordering of the observations can be tuned to sharpen the approximations. Advantageously chosen orderings can drastically improve the approximations, and in fact, completely random orderings often produce far more accurate approximations than default coordinate-based orderings do. In addition to the permutation results, automatic methods for grouping calculations of components of the approximation are introduced, having the result of simultaneously improving the quality of the approximation and reducing its computational burden. In common settings, reordering combined with grouping reduces Kullback-Leibler divergence from the target model by a factor of 80 and computation time by a factor of 2 compared to ungrouped approximations with default ordering. The claims are supported by theory and numerical results with comparisons to other approximations, including tapered covariances and stochastic partial differential equation approximations. Computational details are provided, including efficiently finding the orderings and ordered nearest neighbors, and profiling out linear mean parameters and using the approximations for prediction and conditional simulation. An application to space-time satellite data is presented.
Cited by in corpus (27)
- A general framework for Vecchia approximations of Gaussian processes
- A class of multi-resolution approximations for large spatial datasets
- Vecchia approximations of Gaussian-process predictions
- Highly Scalable Bayesian Geostatistical Modeling via Meshed Gaussian Processes on Partitioned Domains
- Practical Bayesian Modeling and Inference for Massive Spatial Datasets On Modest Computing Environments
- Vecchia-Laplace approximations of generalized Gaussian processes for big non-Gaussian spatial data
- Spectral Density Estimation for Random Fields via Periodic Embeddings
- Fine-scale spatiotemporal air pollution analysis using mobile monitors on Google Street View vehicles
- Modeling Massive Spatial Datasets Using a Conjugate Bayesian Linear Regression Framework
- High-dimensional Multivariate Geostatistics: A Bayesian Matrix-Normal Approach
- Modeling Extremal Streamflow using Deep Learning Approximations and a Flexible Spatial Process
- Correlation-based sparse inverse Cholesky factorization for fast Gaussian-process inference
- Estimating Atmospheric Motion Winds from Satellite Image Data using Space-time Drift Models
- Spatial Multivariate Trees for Big Data Bayesian Regression
- Bayesian nonparametric generative modeling of large multivariate non-Gaussian spatial fields
- Conjugate Nearest Neighbor Gaussian Process Models for Efficient Statistical Interpolation of Large Spatial Data
- Iterative Methods for Vecchia-Laplace Approximations for Latent Gaussian Process Models
- The Renyi Gaussian Process: Towards Improved Generalization
- Fast Bayesian inference of Block Nearest Neighbor Gaussian process for large data
- Gridding and Parameter Expansion for Scalable Latent Gaussian Models of Spatial Multivariate Data
- Gaussian Process Inference Using Mini-batch Stochastic Gradient Descent: Convergence Guarantees and Empirical Benefits
- Using spatial extreme-value theory with machine learning to model and understand spatially compounding weather extremes
- Inverse sampling intensity weighting for preferential sampling adjustment
- Generative multi-scale modeling via spatial autoregressive transport maps
- Bayesian nonstationary and nonparametric covariance estimation for large spatial data
- What is the best predictor that you can compute in five minutes using a given Bayesian hierarchical model?
- Nearest-Neighbor Neural Networks for Geostatistics