2 papers
cs.DC2026
Exceeding the Numerical and Performance Characteristics of IEEE-754 SGEMM with BFloat16 Tensor Cores on GPUs for Scientific Computing
Harun Bayraktar, Cole Brower, John Gunnels +9
Largely due to their increased native capacity for numerical intensity and power efficiency, reduced-precision floating-point computing resources, primarily used in artificial inte…
cs.MS2024
Cascading GEMM: High Precision from Low Precision
Devangi N. Parikh, Robert A. van de Geijn, Greg M. Henry
This paper lays out insights and opportunities for implementing higher-precision matrix-matrix multiplication (GEMM) from (in terms of) lower-precision high-performance GEMM. The d…