2 citations · 5 across the 14 of their papers we have counts for
1 paper · 2 filters
Weihang Shen, Yinqiu Chen, Rong Chen +1
The limited HBM capacity has become the primary bottleneck for hosting an increasing number of larger-scale GPU tasks. While demand paging extends capacity via host DRAM, it incurs…