2 papers
cs.DC2025
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +3
Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to…
cs.AR2024
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
Anirudha Agrawal, Shaizeen Aga, Suchita Pati +1
Concurrent computation and communication (C3) is a pervasive paradigm in ML and other domains, making its performance optimization crucial. In this paper, we carefully characterize…