Showing cs.ARShow all
3 papers · 1 filter
cs.AR2024
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
Hyungkyu Ham, Wonhyuk Yang, Yunseon Shin +5
As DNNs are widely adopted in various application domains while demanding increasingly higher compute and memory requirements, designing efficient and performant NPUs (Neural Proce…
cs.AR2024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
Guseul Heo, Sangyeop Lee, Jaehong Cho +6
Modern transformer-based Large Language Models (LLMs) are constructed with a series of decoder blocks. Each block comprises three key components: (1) QKV generation, (2) multi-head…
cs.AR2019
Mixed-Signal Charge-Domain Acceleration of Deep Neural networks through Interleaved Bit-Partitioned Arithmetic
Soroush Ghodrati, Hardik Sharma, Sean Kinzer +5
Low-power potential of mixed-signal design makes it an alluring option to accelerate Deep Neural Networks (DNNs). However, mixed-signal circuitry suffers from limited range for inf…