6 papers
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
Robin Geens, Joran Heldens, Joren Dumoulin +1
Ternary weight quantization (e.g., BitNet b1.58) offers a promising path to mitigate the memory bandwidth bottleneck in Large Language Model (LLM) inference. However, conventional…
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
Xiaoling Yi, Ryan Antonio, Yunhao Deng +4
Achieving high compute utilization across a wide range of AI workloads is crucial for the efficiency of versatile DNN accelerators. This paper presents the Voltra chip and its util…
The Configuration Wall: Characterization and Elimination of Accelerator Configuration Overhead
Josse Van Delm, Anton Lydike, Joren Dumoulin +6
Contemporary compute platforms increasingly offload compute kernels from CPU to integrated hardware accelerators to reach maximum performance per Watt. Unfortunately, the time the…
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
Stef Cuyckens, Xiaoling Yi, Robin Geens +4
Emerging continual learning applications necessitate next-generation neural processing unit (NPU) platforms to support both training and inference operations. The promising Microsc…
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
Ryan Albert Antonio, Joren Dumoulin, Xiaoling Yi +4
Heterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movem…
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
Xiaoling Yi, Ryan Antonio, Joren Dumoulin +4
Deep neural networks (DNNs) face significant challenges when deployed on resource-constrained extreme edge devices due to their computational and data-intensive nature. While stand…