5 papers
HCCL: Collective Communication for Meta Training and Inference Accelerators
Wesley Bland, Tiago Antunes, Lars Paul Huse +63
We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. MTIA…
Effects of lower floating-point precision on scale-resolving numerical simulations of turbulence
Martin Karp, Ronith Stanly, Timofey Mukha +11
Modern computing clusters offer specialized hardware for reduced-precision arithmetic that can speed up the time to solution significantly. This is possible due to a decrease in da…
Generating wall-bounded turbulent inflows at high Reynolds numbers
Ronith Stanly, Timofey Mukha, Martin Karp +2
One of the main challenges in simulating high Reynolds number () turbulent boundary layers (TBLs) is the long streamwise distance required for large-scale outer-layer structure…
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
MÃ¥ns I. Andersson, Martin Karp, Niclas Jansson +1
With the emergence of new high-performance computing (HPC) accelerators, such as Nvidia and AMD GPUs, efficiently targeting diverse hardware architectures has become a major challe…
Robustness and uncertainty of direct numerical simulation under the influence of rounding and noise
Martin Karp, Niclas Jansson, Saleh Rezaeiravesh +2
Numerical precision in large-scale scientific computations has become an emerging topic due to recent developments in computer hardware. Lower floating point precision offers the p…