5 citations · 5 across the 1 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2023
Proteus: Simulating the Performance of Distributed DNN Training
Jiangfei Duan, Xiuhong Li, Ping Xu +4
DNN models are becoming increasingly larger to achieve unprecedented accuracy, and the accompanying increased computation and memory requirements necessitate the employment of mass…
cs.DC2021★ 5 cited
Characterizing and Demystifying the Implicit Convolution Algorithm on Commercial Matrix-Multiplication Accelerators
Yangjie Zhou, Mengtian Yang, Cong Guo +5
Many of today's deep neural network accelerators, e.g., Google's TPU and NVIDIA's tensor core, are built around accelerating the general matrix multiplication (i.e., GEMM). However…