activity
20192025
most citedHPC Storage Service Autotuning Using Variational-Autoencoder-Guided Asynchronous Bayesian Optimization

15 citations · 26 across the 7 of their papers we have counts for

collaborators
Showing cs.DCShow all

6 papers · 1 filter

cs.DC20241 cited

Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey

Hammad Ather, Jean Luca Bez, Chen Wang +3

Driven by artificial intelligence, data science, and high-resolution simulations, I/O workloads and hardware on high-performance computing (HPC) systems have become increasingly co…

cs.DC202215 cited

HPC Storage Service Autotuning Using Variational-Autoencoder-Guided Asynchronous Bayesian Optimization

Matthieu Dorier, Romain Egele, Prasanna Balaprakash +5

Distributed data storage services tailored to specific applications have grown popular in the high-performance computing (HPC) community as a way to address I/O and storage challen…

cs.DC2022

SKaMPI-OpenSHMEM: Measuring OpenSHMEM Communication Routines

Camille Coti, Allen D. Malony

Benchmarking is an important challenge in HPC, in particular, to be able to tune the basic blocks of the software environment used by applications. The communication library and di…

cs.DC2021

Measuring OpenSHMEM Communication Routines with SKaMPI-OpenSHMEM User's manual

Camille Coti, Allen D Malony

This document presents the OpenSHMEM extension for the Special Karlsruhe MPI benchmark and the measurement algorithms used to measure the routines.

cs.DC2020

On-the-fly Optimization of Parallel Computation of Symbolic Symplectic Invariants

Joseph Ben Geloun, Camille Coti, Allen D. Malony

Group invariants are used in high energy physics to define quantum field theory interactions. In this paper, we are presenting the parallel algebraic computation of special invaria…

cs.DC20198 cited

Checkpoint/restart approaches for a thread-based MPI runtime

Julien Adam, Maxime Kermarquer, Jean-Baptiste Besnard +6

Fault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occu…