13 citations · 15 across the 3 of their papers we have counts for
3 papers
cs.CL2022
Transkimmer: Transformer Learns to Layer-wise Skim
Yue Guan, Zhengyi Li, Jingwen Leng +2
Transformer architecture has become the de-facto model for many machine learning tasks from natural language processing and computer vision. As such, improving its computational ef…
cs.CL2020★ 2 cited
How Far Does BERT Look At:Distance-based Clustering and Analysis of BERTs Attention
Yue Guan, Jingwen Leng, Chao Li +2
Recent research on the multi-head attention mechanism, especially that in pre-trained models such as BERT, has shown us heuristics and clues in analyzing various aspects of the mec…
cs.DC2020★ 13 cited
Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity
Cong Guo, Bo Yang Hsueh, Jingwen Leng +7
Network pruning can reduce the high computation cost of deep neural network (DNN) models. However, to maintain their accuracies, sparse models often carry randomly-distributed weig…