1 paper
Lequn Chen, Weixin Deng, Anirudh Canumalla +4
Having large batch sizes is one of the most critical aspects of increasing the accelerator efficiency and the performance of DNN model inference. However, existing model serving sy…