1 paper
Callum McLean, Luke Y. Prince, Alexandre Payot +2
Matrix multiplication performance has long been the major bottleneck to scaling deep learning workloads, which has stimulated the design of new accelerators that use increasingly l…