paper

Fast and accurate conditioning for large-scale Gaussian process prediction problems

arXiv:2605.02574

Abstract

Gaussian Process (GP) models provide a flexible framework for prediction and uncertainty quantification. For most covariance functions, however, exact GP prediction with points scales as , making it prohibitively expensive for large datasets or large numbers of prediction points. While nearest neighbor-based prediction can work well in certain settings, non-pathological circumstances (like measurement noise, for example) can severely restrict its efficiency. This work presents a complementary approach where one conditions on carefully designed linear combinations of data, which is particularly effective in the setting of jointly predicting many values in large connected regions of the data domain. For kernel functions that are smooth away from the origin and simple prediction domains, this method can be exponentially convergent in the number of linear combinations used for conditioning. The procedure costs work, where is the cost of solving a linear system with the data covariance matrix, and so in many cases can be computed in linear or near-linear cost by exploiting rank structure in well-behaved covariance matrices. At the cost of additional precomputation work, this approach can also provide predictions at arbitrary points of a designated region in online work, making it particularly attractive for problems where prediction points are not known in advance. After establishing favorable theoretical properties, we provide several example applications to problems in prediction and matrix approximation.

Fast and accurate conditioning for large-scale Gaussian process prediction problems · wovepaper