On statistics, computation and scalability
arXiv:1309.7804 · doi:10.3150/12-BEJSP17
Abstract
How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of statistics. With the goal of identifying "time-data tradeoffs," we investigate some of the statistical consequences of computational perspectives on scability, in particular divide-and-conquer methodology and hierarchies of convex relaxations.
Published in at http://dx.doi.org/10.3150/12-BEJSP17 the Bernoulli (http://isi.cbs.nl/bernoulli/) by the International Statistical Institute/Bernoulli Society (http://isi.cbs.nl/BS/bshome.htm)
References in corpus (3)
Cited by in corpus (14)
- Ergodicity of Approximate MCMC Chains with Applications to Large Data Sets
- Sketch and Validate for Big Data Clustering
- Distributed inference for quantile regression processes
- Expanding the scope of statistical computing: Training statisticians to be software engineers
- Optimal Sampling Designs for Multi-dimensional Streaming Time Series with Application to Power Grid Sensor Data
- Joint integrative analysis of multiple data sources with correlated vector outcomes
- The Statistical Performance of Collaborative Inference
- Online Asynchronous Distributed Regression
- Statistical and Computational Tradeoff in Genetic Algorithm-Based Estimation
- A robust fusion-extraction procedure with summary statistics in the presence of biased sources
- A subsampled double bootstrap for massive data
- Partitioned Cross-Validation for Divide-and-Conquer Density Estimation
- A MOM-based ensemble method for robustness, subsampling and hyperparameter tuning
- Eigenvector-based sparse canonical correlation analysis: Fast computation for estimation of multiple canonical vectors