activity
20202026
most citedDemystifying BERT: Implications for Accelerator Design

3 citations · 3 across the 4 of their papers we have counts for

collaborators
Showing cs.ARShow all

11 papers · 1 filter

cs.AR2026

CompPow: A Case for Component-level GPU Power Management

Shaizeen Aga, Mohamed Assem Ibrahim

The ever increasing demand for ML-driven intelligence in a wide spectrum of domains has led to ubiquity of GPUs. At the same time, GPUs are notorious for their power consumption ne…

cs.AR2024

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

Anirudha Agrawal, Shaizeen Aga, Suchita Pati +1

Concurrent computation and communication (C3) is a pervasive paradigm in ML and other domains, making its performance optimization crucial. In this paper, we carefully characterize…

cs.AR2024

FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights

Varsha Singhania, Shaizeen Aga, Mohamed Assem Ibrahim

Ubiquity of AI makes optimizing GPU power a priority as large GPU-based clusters are often employed to train and serve AI models. An important first step in optimizing GPU power co…

cs.AR2024

Global Optimizations & Lightweight Dynamic Logic for Concurrency

Suchita Pati, Shaizeen Aga, Nuwan Jayasena +1

Modern accelerators like GPUs are increasingly executing independent operations concurrently to improve the device's compute utilization. However, effectively harnessing it on GPUs…

cs.AR2024

Balanced Data Placement for GEMV Acceleration with Processing-In-Memory

Mohamed Assem Ibrahim, Mahzabeen Islam, Shaizeen Aga

With unprecedented demand for generative AI (GenAI) inference, acceleration of primitives that dominate GenAI such as general matrix-vector multiplication (GEMV) is receiving consi…

cs.AR20241 cited

T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives

Suchita Pati, Shaizeen Aga, Mahzabeen Islam +2

Large Language Models increasingly rely on distributed techniques for their training and inference. These techniques require communication across devices which can reduce scaling e…