Single-Pass PCA of Large High-Dimensional Data
arXiv:1704.07669
Abstract
Principal component analysis (PCA) is a fundamental dimension reduction tool in statistics and machine learning. For large and high-dimensional data, computing the PCA (i.e., the singular vectors corresponding to a number of dominant singular values of the data matrix) becomes a challenging task. In this work, a single-pass randomized algorithm is proposed to compute PCA with only one pass over the data. It is suitable for processing extremely large and high-dimensional data stored in slow memory (hard disk) or the data generated in a streaming fashion. Experiments with synthetic and real data validate the algorithm's accuracy, which has orders of magnitude smaller error than an existing single-pass algorithm. For a set of high-dimensional data stored as a 150 GB file, the proposed algorithm is able to compute the first 50 principal components in just 24 minutes on a typical 24-core computer, with less than 1 GB memory cost.
IJCAI 2017, 16 pages, 6 figures
References in corpus (5)
- Convolutional Neural Networks for Sentence Classification
- Improved Distributed Principal Component Analysis
- A Practical Guide to Randomized Matrix Computations with MATLAB Implementations
- RSVDPACK: An implementation of randomized algorithms for computing the singular value, interpolative, and CUR decompositions of matrices on multi-core and GPU architectures
- Single Pass PCA of Matrix Products
Cited by in corpus (6)
- Faster Tensor Train Decomposition for Sparse Data
- Computing low-rank approximations of large-scale matrices with the Tensor Network randomized SVD
- Pass-Efficient Randomized LU Algorithms for Computing Low-Rank Matrix Approximation
- Pass-efficient methods for compression of high-dimensional turbulent flow data
- Deterministic matrix sketches for low-rank compression of high-dimensional simulation data
- Random projections in gravitational-wave searches from compact binaries II: efficient reconstruction of the detection statistic