7 papers · 1 filter
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
Saptarshi Mitra, Rachid Karami, Haocheng Xu +2
Emerging applications such as AR are driving demands for machine intelligence capable of processing continuous and/or long-context inputs on local devices. However, currently domin…
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
Faraz Tahmasebi, Michael Pelluer, Hyoukjun Kwon
The computation and memory costs of large language models kept increasing over last decade, which reached over the scale of 1T parameters. To address the challenges from the large…
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
Faraz Tahmasebi, Yian Wang, Benji Y. H. Huang +1
Recent research has shown that large language models (LLMs) can utilize low-precision floating point (FP) quantization to deliver high efficiency while maintaining original model a…
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
Rachid Karami, Sheng-Chun Kao, Hyoukjun Kwon
Among ML operators today, GEneralMatrix Multiplication (GEMM)-based operators are known to be key operators that build the main backbone of ML models. As their computational overhe…
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
Mohanad Odema, Luke Chen, Hyoukjun Kwon +1
We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems…
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
Haocheng Xu, Faraz Tahmasebi, Ye Qiao +3
Recent innovations in Transformer-based large language models have significantly advanced the field of general-purpose neural language understanding and generation. With billions o…