Learning Scalable Deep Kernels with Recurrent Structure
arXiv:1610.08936
Abstract
Many applications in speech, robotics, finance, and biology deal with sequential data, where ordering matters and recurrent structures are common. However, this structure cannot be easily captured by standard kernel functions. To model such structure, we propose expressive closed-form kernel functions for Gaussian processes. The resulting model, GP-LSTM, fully encapsulates the inductive biases of long short-term memory (LSTM) recurrent networks, while retaining the non-parametric probabilistic advantages of Gaussian processes. We learn the properties of the proposed kernels by optimizing the Gaussian process marginal likelihood using a new provably convergent semi-stochastic gradient procedure and exploit the structure of these kernels for scalable training and prediction. This approach provides a practical representation for Bayesian LSTMs. We demonstrate state-of-the-art performance on several benchmarks, and thoroughly investigate a consequential autonomous driving application, where the predictive uncertainties provided by GP-LSTM are uniquely valuable.
37 pages, 7 figures, 5 tables. Updated to the final version that appears in JMLR, 18(82):1-37, 2017
References in corpus (2)
Cited by in corpus (10)
- On Calibration of Modern Neural Networks
- Gaussian Process Regression for In-situ Capacity Estimation of Lithium-ion Batteries
- On Exact Computation with an Infinitely Wide Neural Net
- Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance
- Recurrent Attentive Neural Process for Sequential Data
- Multimodal Word Distributions
- Identification of Gaussian Process State Space Models
- Calibrated simplex-mapping classification
- Stepwise Model Selection for Sequence Prediction via Deep Kernel Learning
- Forecasting solar radiation during dust storms using deep learning