1 paper
Hatem Ltaief, Rabab Alomairy, Qinglei Cao +7
We exploit the widening margin in tensor-core performance between [FP64/FP32/FP16/INT8,FP64/FP32/FP16/FP8/INT8] on NVIDIA [Ampere,Hopper] GPUs to boost the performance of output ac…