4 papers
Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
Yung-Chin Chen, Chung Peng Lee, Ze-Wei Liou +1
Massive activation spikes in Large Language Models (LLMs) severely degrade quantization by stretching dynamic ranges. While prior hypotheses characterize these as high-level scalar…
ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits
Wenlun Zhang, Shimpei Ando, Yung-Chin Chen +1
SRAM-based Analog Compute-in-Memory (ACiM) demonstrates promising energy efficiency for deep neural network (DNN) processing. Nevertheless, efforts to optimize efficiency frequentl…
A Review of SRAM-based Compute-in-Memory Circuits
Kentaro Yoshioka, Shimpei Ando, Satomi Miyagi +2
This paper presents a tutorial and review of SRAM-based Compute-in-Memory (CIM) circuits, with a focus on both Digital CIM (DCIM) and Analog CIM (ACIM) implementations. We explore…
PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
Wenlun Zhang, Shimpei Ando, Yung-Chin Chen +3
Approximate computing emerges as a promising approach to enhance the efficiency of compute-in-memory (CiM) systems in deep neural network processing. However, traditional approxima…