Coresets for Scalable Bayesian Logistic Regression
arXiv:1605.06423
Abstract
The use of Bayesian methods in large-scale data settings is attractive because of the rich hierarchical models, uncertainty quantification, and prior specification they provide. Standard Bayesian inference algorithms are computationally expensive, however, making their direct application to large datasets difficult or infeasible. Recent work on scaling Bayesian inference has focused on modifying the underlying algorithms to, for example, use only a random data subsample at each iteration. We leverage the insight that data is often redundant to instead obtain a weighted subset of the data (called a coreset) that is much smaller than the original dataset. We can then use this small coreset in any number of existing posterior inference algorithms without modification. In this paper, we develop an efficient coreset construction algorithm for Bayesian logistic regression models. We provide theoretical guarantees on the size and approximation quality of the coreset -- both for fixed, known datasets, and in expectation for a wide class of data generative models. Crucially, the proposed approach also permits efficient construction of the coreset in both streaming and parallel settings, with minimal additional effort. We demonstrate the efficacy of our approach on a number of synthetic and real-world datasets, and find that, in practice, the size of the coreset is independent of the original dataset size. Furthermore, constructing the coreset takes a negligible amount of time compared to that required to run MCMC on it.
In Proceedings of Advances in Neural Information Processing Systems (NIPS 2016)
References in corpus (1)
Cited by in corpus (36)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Selection via Proxy: Efficient Data Selection for Deep Learning
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Bayesian Batch Active Learning as Sparse Subset Approximation
- A New Coreset Framework for Clustering
- Real-Time EEG Classification via Coresets for BCI Applications
- Deep Learning on a Data Diet: Finding Important Examples Early in Training
- Decentralized Stochastic Gradient Langevin Dynamics and Hamiltonian Monte Carlo
- On Coresets for Logistic Regression
- Practical bounds on the error of Bayesian posterior approximations: A nonasymptotic approach
- Coresets for Minimum Enclosing Balls over Sliding Windows
- Approximations of Geometrically Ergodic Reversible Markov Chains
- Small quantum computers and large classical data sets
- Core-set Sampling for Efficient Neural Architecture Search
- PASS-GLM: polynomial approximate sufficient statistics for scalable Bayesian GLM inference
- Quantifying sources of uncertainty in drug discovery predictions with probabilistic models
- Embarrassingly Parallel Inference for Gaussian Processes
- Generic Coreset for Scalable Learning of Monotonic Kernels: Logistic Regression, Sigmoid and more
- Distributed Weight Consolidation: A Brain Segmentation Case Study
- Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds
- No Free Lunch for Approximate MCMC
- Wasserstein Measure Coresets
- Importance Sampling via Local Sensitivity
- Improved Synthetic Training for Reading Comprehension
- Online Sampling from Log-Concave Distributions
- A Randomized Algorithm to Reduce the Support of Discrete Measures
- Secure Search on the Cloud via Coresets and Sketches
- Consensus Monte Carlo for Random Subsets using Shared Anchors
- A Tale Of Two Long Tails
- Targeted stochastic gradient Markov chain Monte Carlo for hidden Markov models with rare latent states
- Iterative Teaching by Label Synthesis
- Efficient posterior sampling for high-dimensional imbalanced logistic regression
- Bayesian Coresets: Revisiting the Nonconvex Optimization Perspective
- One Backward from Ten Forward, Subsampling for Large-Scale Deep Learning
- BooVAE: Boosting Approach for Continual Learning of VAE
- Deep or Simple Models for Semantic Tagging? It Depends on your Data [Experiments]