2 papers
cs.DC2026
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
Feng Pan, Hanfeng Gu, Paul Springer +1
Exact tensor network contraction underpins quantum circuit simulation, quantum error correction, combinatorial optimization, and many-body dynamics. The dominant parallelization st…
cs.DC2026
Exceeding the Numerical and Performance Characteristics of IEEE-754 SGEMM with BFloat16 Tensor Cores on GPUs for Scientific Computing
Harun Bayraktar, Cole Brower, John Gunnels +9
Largely due to their increased native capacity for numerical intensity and power efficiency, reduced-precision floating-point computing resources, primarily used in artificial inte…