150 citations · 1.3k across the 73 of their papers we have counts for
16 papers · 1 filter
SMASH: Co-designing Software Compression and Hardware-Accelerated Indexing for Efficient Sparse Matrix Operations
Konstantinos Kanellopoulos, Nandita Vijaykumar, Christina Giannoula +6
Important workloads, such as machine learning and graph analytics applications, heavily involve sparse linear algebra operations. These operations use sparse matrix compression as…
EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
Skanda Koppula, Lois Orosa, Abdullah Giray Yağlıkçı +4
The effectiveness of deep neural networks (DNN) in vision, speech, and language processing has prompted a tremendous demand for energy-efficient high-performance DNN inference syst…
DSPatch: Dual Spatial Pattern Prefetcher
Rahul Bera, Anant V. Nori, Onur Mutlu +1
High main memory latency continues to limit performance of modern high-performance out-of-order cores. While DRAM latency has remained nearly the same over many generations, DRAM b…
Refresh Triggered Computation: Improving the Energy Efficiency of Convolutional Neural Network Accelerators
Syed M. A. H. Jafri, Hasan Hassan, Ahmed Hemani +1
To employ a Convolutional Neural Network (CNN) in an energy-constrained embedded system, it is critical for the CNN implementation to be highly energy efficient. Many recent studie…
The Non-IID Data Quagmire of Decentralized Machine Learning
Kevin Hsieh, Amar Phanishayee, Onur Mutlu +1
Many large-scale machine learning (ML) applications need to perform decentralized learning over datasets generated at different devices and locations. Such datasets pose a signific…
Enabling and Exploiting Partition-Level Parallelism (PALP) in Phase Change Memories
Shihao Song, Anup Das, Onur Mutlu +1
Phase-change memory (PCM) devices have multiple banks to serve memory requests in parallel. Unfortunately, if two requests go to the same bank, they have to be served one after ano…