27 citations · 27 across the 1 of their papers we have counts for
1 paper
Paras Jain, Xiangxi Mo, Ajay Jain +5
Serving deep neural networks in latency critical interactive settings often requires GPU acceleration. However, the small batch sizes typical in online inference results in poor GP…