6 papers
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
Zi-Wei Lin, Tian-Sheuan Chang
We present VitaLLM, a mixed precision accelerator that enables ternary weight large language models to run efficiently on edge devices. The design combines two compute cores, a mul…
A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator with In-Situ Regulation for Energy-Efficient Spiking Neural Networks
Shih-Hang Kao, Yang-Chan Hung, I-Wen Wang +7
This paper presents a PVT-resilient, subthreshold SRAM-based computing-in-memory (CIM) macro tailored for energy-efficient spiking neural networks (SNNs). The macro integrates in-s…
Computing-In-Memory Aware Model Adaption For Edge Devices
Ming-Han Lin, Tian-Sheuan Chang
Computing-in-Memory (CIM) macros have gained popularity for deep learning acceleration due to their highly parallel computation and low power consumption. However, limited macro si…
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
Bo-Yu Chen, Tian-Sheuan Chang
This paper introduces the first low-power hardware accelerator for Spiking Transformers, an emerging alternative to traditional artificial neural networks. By modifying the base Sp…
An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
Tseng-Jen Li, Tian-Sheuan Chang
Transformer-based models have become the \textit{de facto} backbone across many fields, such as computer vision and natural language processing. However, as these models scale in s…
A Low-Power Sparse Deep Learning Accelerator with Optimized Data Reuse
Kai-Chieh Hsu, Tian-Sheuan Chang
Sparse deep learning has reduced computation significantly, but its irregular non-zero data distribution complicates the data flow and hinders data reuse, increasing on-chip SRAM a…