Generating coupled cluster code for modern distributed memory tensor software
arXiv:2409.06759 · doi:10.1021/acs.jctc.5c00219
Abstract
Using GPU-based HPC platforms efficiently for coupled cluster computations is a challenge due to heterogeneous hardware structures. The constant need to adapt software to these structures and the required man-hours makes a systematization of high-performance code development desirable, even more so for higher-order coupled cluster. This is generally achieved by introducing a high-level representation of the problem, which is then translated to low-level instructions for the hardware using a compiler/translator component. Designing such software comes with another challenge: Allowing efficient implementation by capturing key symmetries of tensors, while retaining the abstraction from the hardware. We review ways to address these two challenges while presenting design decisions which led us to the development of a general-order coupled cluster code generator. The systematically produced code shows excellent weak scaling behavior running on up to 1200 GPUs using the distributed memory tensor library ExaTENSOR. We present an open-source modular tensor framework "tenpi" for coupled cluster code development with diagrammatic derivation, visualization module, symbolic algebra, intermediate optimization and support for multiple tensor backends. Tenpi brings higher-order CC functionality to the massively parallel ExaCorr module of the DIRAC code for relativistic molecular calculations.
quantum chemistry, 13 pages, 11 figures
References in corpus (18)
- Tensor Decomposition for Signal Processing and Machine Learning
- The DIRAC code for relativistic molecular calculations
- On Tensors, Sparsity, and Nonnegative Factorizations
- BAGEL: Brilliantly Advanced General Electronic-structure Library
- On-the-fly CASPT2 surface hopping dynamics
- Neural tensor contractions and the expressive power of deep neural quantum states
- MADNESS: A Multiresolution, Adaptive Numerical Environment for Scientific Simulation
- High-Performance Tensor Contraction without Transposition
- 4-component relativistic Hamiltonian with effective QED potentials for molecular calculations
- Tensor Contractions with Extended BLAS Kernels on CPU and GPU
- Implementation of relativistic coupled cluster theory for massively parallel GPU-accelerated computing architectures
- Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
- DISTAL: The Distributed Tensor Algebra Compiler
- Automatic derivation of fermionic many-body theories based on general Fermi vacua
- TAMM: Tensor Algebra for Many-body Methods
- Assessing MP2 frozen natural orbitals in relativistic correlated electronic structure calculations
- Equation Generator for Equation-of-Motion Coupled Cluster Assisted by Computer Algebra System
- A Power Series Approximation in Symmetry Projected Coupled Cluster Theory