7 citations · 20 across the 3 of their papers we have counts for
4 papers
Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices
Yuhong Song, Weiwen Jiang, Bingbing Li +6
A pruning-based AutoML framework for run-time reconfigurability, namely RT3, is proposed in this work. This enables Transformer-based large Natural Language Processing (NLP) models…
Enabling Retrain-free Deep Neural Network Pruning using Surrogate Lagrangian Relaxation
Deniz Gurevin, Shanglin Zhou, Lynn Pepin +4
Network pruning is a widely used technique to reduce computation cost and model size for deep neural networks. However, the typical three-stage pipeline, i.e., training, pruning an…
Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning
Bingbing Li, Zhenglun Kong, Tianyun Zhang +4
Pre-trained large-scale language models have increasingly demonstrated high accuracy on many natural language processing (NLP) tasks. However, the limited weight storage and comput…
FTRANS: Energy-Efficient Acceleration of Transformers using FPGA
Bingbing Li, Santosh Pandey, Haowen Fang +7
In natural language processing (NLP), the "Transformer" architecture was proposed as the first transduction model replying entirely on self-attention mechanisms without using seque…