3 papers
cs.DC2026
Exceeding the Numerical and Performance Characteristics of IEEE-754 SGEMM with BFloat16 Tensor Cores on GPUs for Scientific Computing
Harun Bayraktar, Cole Brower, John Gunnels +9
Largely due to their increased native capacity for numerical intensity and power efficiency, reduced-precision floating-point computing resources, primarily used in artificial inte…
cs.DC2025
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
Angelika Schwarz, Anton Anders, Cole Brower +8
The rapid growth of artificial intelligence (AI) has made low-precision formats such as FP16, FP8, and, most recently, block-scaled FP4 the primary focus of modern GPUs, where Tens…
math.NA2024
Hardware Trends Impacting Floating-Point Computations In Scientific Applications
Jack Dongarra, John Gunnels, Harun Bayraktar +2
The evolution of floating-point computation has been shaped by algorithmic advancements, architectural innovations, and the increasing computational demands of modern technologies,…