Parallel Tensor Compression for Large-Scale Scientific Data
arXiv:1510.06689 · doi:10.1109/IPDPS.2016.67
Abstract
As parallel computing trends towards the exascale, scientific data produced by high-fidelity simulations are growing increasingly massive. For instance, a simulation on a three-dimensional spatial grid with 512 points per dimension that tracks 64 variables per grid point for 128 time steps yields 8~TB of data, assuming double precision. By viewing the data as a dense five-way tensor, we can compute a Tucker decomposition to find inherent low-dimensional multilinear structure, achieving compression ratios of up to 5000 on real-world data sets with negligible loss in accuracy. So that we can operate on such massive data, we present the first-ever distributed-memory parallel implementation for the Tucker decomposition, whose key computations correspond to parallel linear algebra operations, albeit with nonstandard data layouts. Our approach specifies a data distribution for tensors that avoids any tensor data redistribution, either locally or in parallel. We provide accompanying analysis of the computation and communication costs of the algorithms. To demonstrate the compression and accuracy of the method, we apply our approach to real-world data sets from combustion science simulations. We also provide detailed performance results, including parallel performance in both weak and strong scaling experiments.
Cited by in corpus (21)
- Low-Rank Tensor Networks for Dimensionality Reduction and Large-Scale Optimization Problems: Perspectives and Challenges PART 1
- Z-checker: A Framework for Assessing Lossy Compression of Scientific Data
- A Unified Optimization Approach for Sparse Tensor Operations on GPUs
- A Distributed and Incremental SVD Algorithm for Agglomerative Data Analysis on Large Networks
- Unsupervised Machine Learning Based on Non-Negative Tensor Factorization for Analyzing Reactive-Mixing
- Improving Performance of Iterative Methods by Lossy Checkponting
- cuTT: A High-Performance Tensor Transpose Library for CUDA Compatible GPUs
- Low-Rank Tucker Approximation of a Tensor From Streaming Data
- Multiresolution Tensor Learning for Efficient and Interpretable Spatial Analysis
- Compression challenges in large scale PDE solvers
- Projection-based model reduction of dynamical systems using space-time subspace and machine learning
- Shared Memory Parallelization of MTTKRP for Dense Tensors
- Tensor clustering with algebraic constraints gives interpretable groups of crosstalk mechanisms in breast cancer
- a-Tucker: Input-Adaptive and Matricization-Free Tucker Decomposition for Dense Tensors on CPUs and GPUs
- Parallel algorithms for computing the tensor-train decomposition
- Pass-efficient methods for compression of high-dimensional turbulent flow data
- A rank-adaptive higher-order orthogonal iteration algorithm for truncated Tucker decomposition
- PASTA: A Parallel Sparse Tensor Algorithm Benchmark Suite
- Efficient Alternating Least Squares Algorithms for Low Multilinear Rank Approximation of Tensors
- Riemannian preconditioned coordinate descent for low multi-linear rank approximation
- Tensor-structured algorithm for reduced-order scaling large-scale Kohn-Sham density functional theory calculations