3 citations · 3 across the 4 of their papers we have counts for
11 papers · 1 filter
CompPow: A Case for Component-level GPU Power Management
Shaizeen Aga, Mohamed Assem Ibrahim
The ever increasing demand for ML-driven intelligence in a wide spectrum of domains has led to ubiquity of GPUs. At the same time, GPUs are notorious for their power consumption ne…
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
Anirudha Agrawal, Shaizeen Aga, Suchita Pati +1
Concurrent computation and communication (C3) is a pervasive paradigm in ML and other domains, making its performance optimization crucial. In this paper, we carefully characterize…
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
Varsha Singhania, Shaizeen Aga, Mohamed Assem Ibrahim
Ubiquity of AI makes optimizing GPU power a priority as large GPU-based clusters are often employed to train and serve AI models. An important first step in optimizing GPU power co…
Global Optimizations & Lightweight Dynamic Logic for Concurrency
Suchita Pati, Shaizeen Aga, Nuwan Jayasena +1
Modern accelerators like GPUs are increasingly executing independent operations concurrently to improve the device's compute utilization. However, effectively harnessing it on GPUs…
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
Mohamed Assem Ibrahim, Mahzabeen Islam, Shaizeen Aga
With unprecedented demand for generative AI (GenAI) inference, acceleration of primitives that dominate GenAI such as general matrix-vector multiplication (GEMV) is receiving consi…
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +2
Large Language Models increasingly rely on distributed techniques for their training and inference. These techniques require communication across devices which can reduce scaling e…