11 citations · 37 across the 6 of their papers we have counts for
4 papers · 1 filter
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
Bingbing Li, Geng Yuan, Zigeng Wang +6
Resistive Random Access Memory (ReRAM) has emerged as a promising platform for deep neural networks (DNNs) due to its support for parallel in-situ matrix-vector multiplication. How…
Accelerating Framework of Transformer by Hardware Design and Model Compression Co-Optimization
Panjie Qi, Edwin Hsing-Mean Sha, Qingfeng Zhuge +5
State-of-the-art Transformer-based models, with gigantic parameters, are difficult to be accommodated on resource constrained embedded devices. Moreover, with the development of te…
Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices
Yuhong Song, Weiwen Jiang, Bingbing Li +6
A pruning-based AutoML framework for run-time reconfigurability, namely RT3, is proposed in this work. This enables Transformer-based large Natural Language Processing (NLP) models…
Enabling Retrain-free Deep Neural Network Pruning using Surrogate Lagrangian Relaxation
Deniz Gurevin, Shanglin Zhou, Lynn Pepin +4
Network pruning is a widely used technique to reduce computation cost and model size for deep neural networks. However, the typical three-stage pipeline, i.e., training, pruning an…