A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference
arXiv:2212.02872 · doi:10.1038/s41928-023-01010-1
Abstract
The need to repeatedly shuttle around synaptic weight values from memory to processing units has been a key source of energy inefficiency associated with hardware implementation of artificial neural networks. Analog in-memory computing (AIMC) with spatially instantiated synaptic weights holds high promise to overcome this challenge, by performing matrix-vector multiplications (MVMs) directly within the network weights stored on a chip to execute an inference workload. However, to achieve end-to-end improvements in latency and energy consumption, AIMC must be combined with on-chip digital operations and communication to move towards configurations in which a full inference workload is realized entirely on-chip. Moreover, it is highly desirable to achieve high MVM and inference accuracy without application-wise re-tuning of the chip. Here, we present a multi-core AIMC chip designed and fabricated in 14-nm complementary metal-oxide-semiconductor (CMOS) technology with backend-integrated phase-change memory (PCM). The fully-integrated chip features 64 256x256 AIMC cores interconnected via an on-chip communication network. It also implements the digital activation functions and processing involved in ResNet convolutional neural networks and long short-term memory (LSTM) networks. We demonstrate near software-equivalent inference accuracy with ResNet and LSTM networks while implementing all the computations associated with the weight layers and the activation functions on-chip. The chip can achieve a maximal throughput of 63.1 TOPS at an energy efficiency of 9.76 TOPS/W for 8-bit input/output matrix-vector multiplications.
References in corpus (2)
Cited by in corpus (20)
- A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
- Using the IBM Analog In-Memory Hardware Acceleration Kit for Neural Network Training and Inference
- Fast offset corrected in-memory training
- Synaptic-Like Plasticity in 2D Nanofluidic Memristor from Competitive Bicationic Transport
- All-in-One Analog AI Hardware: On-Chip Training and Inference with Conductive-Metal-Oxide/HfOx ReRAM Devices
- Analytical Modelling of the Transport in Analog Filamentary Conductive-Metal-Oxide/HfOx ReRAM Devices
- How to keep pushing ML accelerator performance? Know your rooflines!
- The Inherent Adversarial Robustness of Analog In-Memory Computing
- Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-Training
- CiMBA: Accelerating Genome Sequencing through On-Device Basecalling via Compute-in-Memory
- In-Materia Speech Recognition
- The Ouroboros of Memristors: Neural Networks Facilitating Memristor Programming
- A Precision-Optimized Fixed-Point Near-Memory Digital Processing Unit for Analog In-Memory Computing
- Update Disturbance-Resilient Analog ReRAM Crossbar Arrays for In-Memory Deep Learning Accelerators
- D-SELD: Dataset-Scalable Exemplar LCA-Decoder
- A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
- A Distributed Emulation Environment for In-Memory Computing Systems
- Predicting sampling advantage of stochastic Ising Machines for Quantum Simulations
- Efficient transformer adaptation for analog in-memory computing via low-rank adapters
- ThRIve: Thermally Robust CNN Inference via Low-Rank Adaptation in Heterogeneous PIM Architectures