2 papers
cs.DC2026
Toward a Unified GPU-Aware OpenSHMEM Specification
Naveen Ravi, Nathan Wichmann, Md. Wasi-ur- Rahman +19
Leadership-class HPC systems are now accelerator-centric, with GPUs providing most floating-point throughput and memory bandwidth. As next-generation systems increasingly integrate…
cs.DC2024
Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications
Sanjif Shanmugavelu, Mathieu Taillefumier, Christopher Culver +3
Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumu…