Challenges of Big Data Analysis
arXiv:1308.1479 · doi:10.1093/nsr/nwt032
Abstract
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data introduce unique computational and statistical challenges, including scalability and storage bottleneck, noise accumulation, spurious correlation, incidental endogeneity, and measurement errors. These challenges are distinguished and require new computational and statistical paradigm. This article give overviews on the salient features of Big Data and how these features impact on paradigm change on statistical and computational methods as well as computing architectures. We also provide various new perspectives on the Big Data analysis and computation. In particular, we emphasis on the viability of the sparsest solution in high-confidence set and point out that exogeneous assumptions in most statistical methods for Big Data can not be validated due to incidental endogeneity. They can lead to wrong statistical inferences and consequently wrong scientific conclusions.
References in corpus (14)
- Nearly unbiased variable selection under minimax concave penalty
- Pathwise coordinate optimization
- Covariance regularization by thresholding
- One-step sparse estimates in nonconcave penalized likelihood models
- Coordinate descent algorithms for nonconvex penalized regression, with applications to biological feature selection
- Coordinate descent algorithms for lasso penalized regression
- High-dimensional classification using features annealed independence rules
- Parallel Coordinate Descent for L1-Regularized Loss Minimization
- L1-Penalization for Mixture Regression Models
- Regularized rank-based estimation of high-dimensional nonparanormal graphical models
- Oracle inequalities for the lasso in the Cox model
- Factor modeling for high-dimensional time series: Inference for the number of factors
- Large Vector Auto Regressions
- Control of the False Discovery Rate Under Arbitrary Covariance Dependence
Cited by in corpus (77)
- Artificial Intelligence and Big Data in Entrepreneurship: A New Era Has Begun
- Deep Learning for Time Series Forecasting: A Survey
- Computational Intelligence Challenges and Applications on Large-Scale Astronomical Time Series Databases
- Distributed Estimation and Inference with Statistical Guarantees
- Data learning from big data
- Distributed Simultaneous Inference in Generalized Linear Models via Confidence Distribution
- Are Discoveries Spurious? Distributions of Maximum Spurious Correlations and Their Applications
- Optimal Shrinkage Estimator for High-Dimensional Mean Vector
- A Survey of Bayesian Statistical Approaches for Big Data
- A systematic data characteristic understanding framework towards physical-sensor big data challenges
- Factor Augmented Sparse Throughput Deep ReLU Neural Networks for High Dimensional Regression
- A big-data spatial, temporal and network analysis of bovine tuberculosis between wildlife (badgers) and cattle
- Non-Parametric Causality Detection: An Application to Social Media and Financial Data
- Beyond Volume: The Impact of Complex Healthcare Data on the Machine Learning Pipeline
- Large deviation theory for diluted Wishart random matrices
- Variational Bayesian Weighted Complex Network Reconstruction
- Spherical Cap Packing Asymptotics and Rank-Extreme Detection
- Quantifying the structure of strong gravitational lens potentials with uncertainty-aware deep neural networks
- Making a Case for Social Media Corpus for Detecting Depression
- Big Learning with Bayesian Methods
- Supervised clustering of high dimensional data using regularized mixture modeling
- DEFM: Delay E mbedding based Forecast Machine for Time Series Forecasting by Spatiotemporal Information Transformation
- Tensor Methods for Additive Index Models under Discordance and Heterogeneity
- Information Retrieval and Recommendation System for Astronomical Observatories
- An Overview on the Estimation of Large Covariance and Precision Matrices
- A Likelihood Ratio Framework for High Dimensional Semiparametric Regression
- Certifiably Optimal Sparse Inverse Covariance Estimation
- Curse of Heterogeneity: Computational Barriers in Sparse Mixture Models and Phase Retrieval
- ExSIS: Extended Sure Independence Screening for Ultrahigh-dimensional Linear Models
- Joint integrative analysis of multiple data sources with correlated vector outcomes
- Sorting Big Data by Revealed Preference with Application to College Ranking
- Cloud Computing - Architecture and Applications
- DotHash: Estimating Set Similarity Metrics for Link Prediction and Document Deduplication
- Applying the Delta method in metric analytics: A practical guide with novel ideas
- Optimal subsampling for quantile regression in big data
- Likelihood Ratio Test in Multivariate Linear Regression: from Low to High Dimension
- A Block Minorization--Maximization Algorithm for Heteroscedastic Regression
- Asymptotically Independent U-Statistics in High-Dimensional Testing
- Sparse transition matrix estimation for high-dimensional and locally stationary vector autoregressive models
- Detection and inference of changes in high-dimensional linear regression with non-sparse structures
- Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
- GOLFS: Feature Selection via Combining Both Global and Local Information for High Dimensional Clustering
- Optimization meets Big Data: A survey
- Strong Sure Screening of Ultra-high Dimensional Data with Interaction Effects
- Optimal estimation of functionals of high-dimensional mean and covariance matrix
- Topological Detection of Phenomenological Bifurcations with Unreliable Kernel Densities
- Distributed nonparametric regression imputation for missing response problems with large-scale data
- Distributed rank-1 dictionary learning: Towards fast and scalable solutions for fMRI big data analytics
- An Efficient Matrix Multiplication with Enhanced Privacy Protection in Cloud Computing and Its Applications
- A robust fusion-extraction procedure with summary statistics in the presence of biased sources
- Adaptive Huber Regression on Markov-dependent Data
- Top eigenpair statistics of diluted Wishart matrices
- Goodness-of-Fit Tests for Large Datasets
- Big data need physical ideas and methods
- A Survey on Nonconvex Regularization Based Sparse and Low-Rank Recovery in Signal Processing, Statistics, and Machine Learning
- A constrained L1 minimization approach for estimating multiple Sparse Gaussian or Nonparanormal Graphical Models
- GPIC - GPU Power Iteration Cluster
- Sparse recovery via nonconvex regularized -estimators over -balls
- Bayesian variable selection in linear regression models with instrumental variables
- Data-based prediction and causality inference of nonlinear dynamics
- Deep Neural Networks Guided Ensemble Learning for Point Estimation
- Selective Correlation Based Knowledge Distillation for Ground Reaction Force Estimation
- Asymptotics of empirical eigenvalues for large separable covariance matrices
- Adaptive monotonicity testing in sublinear time
- Swift Two-sample Test on High-dimensional Neural Spiking Data
- Better Solution Principle: A Facet of Concordance between Optimization and Statistics
- Positive Definite Estimation of Large Covariance Matrix Using Generalized Nonconvex Penalties
- Marginal and Interactive Feature Screening of Ultra-high Dimensional Feature Spaces with Multivariate Response
- Cross-validation and Peeling Strategies for Survival Bump Hunting using Recursive Peeling Methods
- A unified algorithm framework for quality control of sensor data for behavioural clinimetric testing
- Efficient Test-based Variable Selection for High-dimensional Linear Models
- Shamap: Shape-based Manifold Learning
- Bayes Calculations from Quantile Implied Likelihood
- Pre-processing with Orthogonal Decompositions for High-dimensional Explanatory Variables
- A Survey of Semantics-Aware Performance Optimization for Data-Intensive Computing
- Sequential Estimation under Multiple Resources: a Bandit Point of View
- Toward Compact Data from Big Data