pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
arXiv:2104.07699 · doi:10.1109/MICRO56248.2022.00067
Abstract
Data movement between the main memory and the processor is a key contributor to execution time and energy consumption in memory-intensive applications. This data movement bottleneck can be alleviated using Processing-in-Memory (PiM). One category of PiM is Processing-using-Memory (PuM), in which computation takes place inside the memory array by exploiting intrinsic analog properties of the memory device. PuM yields high performance and energy efficiency, but existing PuM techniques support a limited range of operations. As a result, current PuM architectures cannot efficiently perform some complex operations (e.g., multiplication, division, exponentiation) without large increases in chip area and design complexity. To overcome these limitations of existing PuM architectures, we introduce pLUTo (processing-using-memory with lookup table (LUT) operations), a DRAM-based PuM architecture that leverages the high storage density of DRAM to enable the massively parallel storing and querying of lookup tables (LUTs). The key idea of pLUTo is to replace complex operations with low-cost, bulk memory reads (i.e., LUT queries) instead of relying on complex extra logic. We evaluate pLUTo across 11 real-world workloads that showcase the limitations of prior PuM approaches and show that our solution outperforms optimized CPU and GPU baselines by an average of 713 and 1.2, respectively, while simultaneously reducing energy consumption by an average of 1855 and 39.5. Across these workloads, pLUTo outperforms state-of-the-art PiM architectures by an average of 18.3. We also show that different versions of pLUTo provide different levels of flexibility and performance at different additional DRAM area overheads (between 10.2% and 23.1%). pLUTo's source code is openly and fully available at https://github.com/CMU-SAFARI/pLUTo.
References in corpus (15)
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- PULP-NN: Accelerating Quantized Neural Networks on Parallel Ultra-Low-Power RISC-V Processors
- GRIM-Filter: Fast Seed Location Filtering in DNA Read Mapping Using Processing-in-Memory Technologies
- A Resistive CAM Processing-in-Storage Architecture for DNA Sequence Alignment
- In-DRAM Bulk Bitwise Execution Engine
- Buddy-RAM: Improving the Performance and Efficiency of Bulk Bitwise Operations Using DRAM
- Benchmarking a New Paradigm: An Experimental Analysis of a Real Processing-in-Memory Architecture
- Memristive Devices for Computation-In-Memory
- NOM: Network-On-Memory for Inter-Bank Data Transfer in Highly-Banked Memories
- Performance Implications of NoCs on 3D-Stacked Memories: Insights from the Hybrid Memory Cube
- Mitigating Edge Machine Learning Inference Bottlenecks: An Empirical Study on Accelerating Google Edge Models
- Polynesia: Enabling Effective Hybrid Transactional/Analytical Databases with Specialized Hardware/Software Co-Design
- LazyPIM: Efficient Support for Cache Coherence in Processing-in-Memory Architectures
- GenStore: A High-Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis
- RowClone: Accelerating Data Movement and Initialization Using DRAM
Cited by in corpus (4)
- MARS: Processing-In-Memory Acceleration of Raw Signal Genome Analysis Inside the Storage Subsystem
- Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
- Single-Cell Universal Logic-in-Memory Using 2T-nC FeRAM: An Area and Energy-Efficient Approach for Bulk Bitwise Computation
- Darwin: A DRAM-based Multi-level Processing-in-Memory Architecture for Data Analytics