3 citations · 3 across the 3 of their papers we have counts for
4 papers
Global Optimizations & Lightweight Dynamic Logic for Concurrency
Suchita Pati, Shaizeen Aga, Nuwan Jayasena +1
Modern accelerators like GPUs are increasingly executing independent operations concurrently to improve the device's compute utilization. However, effectively harnessing it on GPUs…
Demystifying BERT: Implications for Accelerator Design
Suchita Pati, Shaizeen Aga, Nuwan Jayasena +1
Transfer learning in natural language processing (NLP), as realized using models like BERT (Bi-directional Encoder Representation from Transformer), has significantly improved lang…
SeqPoint: Identifying Representative Iterations of Sequence-based Neural Networks
Suchita Pati, Shaizeen Aga, Matthew D. Sinclair +1
The ubiquity of deep neural networks (DNNs) continues to rise, making them a crucial application class for hardware optimizations. However, detailed profiling and characterization…
Analyzing Machine Learning Workloads Using a Detailed GPU Simulator
Jonathan Lew, Deval Shah, Suchita Pati +8
Most deep neural networks deployed today are trained using GPUs via high-level frameworks such as TensorFlow and PyTorch. This paper describes changes we made to the GPGPU-Sim simu…