1 paper · 1 filter
Weihang Shen, Yinqiu Chen, Rong Chen +1
The limited HBM capacity has become the primary bottleneck for hosting an increasing number of larger-scale GPU tasks. While demand paging extends capacity via host DRAM, it incurs…