1 citations · 1 across the 1 of their papers we have counts for
3 papers
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
Zhimin Ding, Jiawen Yao, Brianna Barrow +7
An obvious way to alleviate memory difficulties in GPU-based AI computing is via CPU offload, where data are moved between GPU and CPU RAM, so inexpensive CPU RAM is used to increa…
DeePref: Deep Reinforcement Learning For Video Prefetching In Content Delivery Networks
Nawras Alkassab, Chin-Tser Huang, Tania Lorido Botran
Content Delivery Networks carry the majority of Internet traffic, and the increasing demand for video content as a major IP traffic across the Internet highlights the importance of…
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
Saeid Ghafouri, Kamran Razavi, Mehran Salmani +5
Efficiently optimizing multi-model inference pipelines for fast, accurate, and cost-effective inference is a crucial challenge in machine learning production systems, given their t…