3 citations · 3 across the 4 of their papers we have counts for
3 papers · 1 filter
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
Marco Kurzynski, Shaizeen Aga, Di Wu
Training large language models (LLMs) efficiently requires a deep understanding of how modern GPU systems behave under real-world distributed training workloads. While prior work h…
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +3
Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to…
SeqPoint: Identifying Representative Iterations of Sequence-based Neural Networks
Suchita Pati, Shaizeen Aga, Matthew D. Sinclair +1
The ubiquity of deep neural networks (DNNs) continues to rise, making them a crucial application class for hardware optimizations. However, detailed profiling and characterization…