1 citations · 2 across the 6 of their papers we have counts for
6 papers · 1 filter
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
Mohamed Assem Ibrahim, Mahzabeen Islam, Shaizeen Aga
With unprecedented demand for generative AI (GenAI) inference, acceleration of primitives that dominate GenAI such as general matrix-vector multiplication (GEMV) is receiving consi…
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +2
Large Language Models increasingly rely on distributed techniques for their training and inference. These techniques require communication across devices which can reduce scaling e…
Just-in-time Quantization with Processing-In-Memory for Efficient ML Training
Mohamed Assem Ibrahim, Shaizeen Aga, Ada Li +2
Data format innovations have been critical for machine learning (ML) scaling, which in turn fuels ground-breaking ML capabilities. However, even in the presence of low-precision fo…
Collaborative Acceleration for FFT on Commercial Processing-In-Memory Architectures
Mohamed Assem Ibrahim, Shaizeen Aga
This paper evaluates the efficacy of recent commercial processing-in-memory (PIM) solutions to accelerate fast Fourier transform (FFT), an important primitive across several domain…
Egalitarian ORAM: Wear-Leveling for ORAM
Yi Zheng, Aasheesh Kolli, Shaizeen Aga
While non-volatile memories (NVMs) provide several desirable characteristics like better density and comparable energy efficiency than DRAM, DRAM-like performance, and disk-like du…
Computation vs. Communication Scaling for Future Transformers on Future Hardware
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +2
Scaling neural network models has delivered dramatic quality gains across ML problems. However, this scaling has increased the reliance on efficient distributed training techniques…