2 papers
cs.DC2026
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +3
Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to…
eess.SY2025
ECLIP: Energy-efficient and Practical Co-Location of ML Inference on Spatially Partitioned GPUs
Ryan Quach, Yidi Wang, Ali Jahanshahi +2
As AI inference becomes mainstream, research has begun to focus on improving the energy consumption of inference servers. Inference kernels commonly underutilize a GPU's compute re…