5 papers
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
Jakub Homola, Ondřej Meca, Lubomír Říha +1
Schur complement matrices emerge in many domain decomposition methods that can solve complex engineering problems using supercomputers. Today, as most of the high-performance clust…
Heterogeneous Memory Pool Tuning
Filip Vaverka, Ondrej Vysocky, Lubomir Riha
We present a lightweight tool for the analysis and tuning of application data placement in systems with heterogeneous memory pools. The tool allows non-intrusively identifying, ana…
Methodology for GPU Frequency Switching Latency Measurement
Daniel Velicka, Ondrej Vysocky, Lubomir Riha
The development of exascale and post-exascale HPC and AI systems integrates thousands of CPUs and specialized accelerators, making energy optimization critical as power costs rival…
Assembly of FETI dual operator using CUDA
Jakub Homola, Radim Vavřík, Ondřej Meca +2
FETI is a numerical method used to solve engineering problems. It builds on the ideas of domain decomposition, which makes it highly scalable and capable of efficiently utilizing w…
Toward an End-to-End Auto-tuning Framework in HPC PowerStack
Xingfu Wu, Aniruddha Marathe, Siddhartha Jana +7
Efficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software…