1 paper · 1 filter
Ryan Quach, Yidi Wang, Ali Jahanshahi +2
As AI inference becomes mainstream, research has begun to focus on improving the energy consumption of inference servers. Inference kernels commonly underutilize a GPU's compute re…