activity
20162023
most citedLanguage model compression with weighted low-rank factorization

20 citations · 52 across the 7 of their papers we have counts for

collaborators

5 papers

eess.SP2023

To Wake-up or Not to Wake-up: Reducing Keyword False Alarm by Successive Refinement

Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching-Hua Lee +3

Keyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happ…

cs.AI20232 cited

GOHSP: A Unified Framework of Graph and Optimization-based Heterogeneous Structured Pruning for Vision Transformer

Miao Yin, Burak Uzkent, Yilin Shen +2

The recently proposed Vision transformers (ViTs) have shown very impressive empirical performance in various computer vision tasks, and they are viewed as an important type of foun…

cs.LG202220 cited

Language model compression with weighted low-rank factorization

Yen-Chang Hsu, Ting Hua, Sungen Chang +3

Factorizing a large matrix into small matrices is a popular strategy for model compression. Singular value decomposition (SVD) plays a vital role in this compression strategy, appr…

cs.CL202116 cited

Automatic Mixed-Precision Quantization Search of BERT

Changsheng Zhao, Ting Hua, Yilin Shen +2

Pre-trained language models such as BERT have shown remarkable effectiveness in various natural language processing tasks. However, these models usually contain millions of paramet…

cs.DB20167 cited

DPHMM: Customizable Data Release with Differential Privacy via Hidden Markov Model

Yonghui Xiao, Yilin Shen, Jinfei Liu +3

Hidden Markov model (HMM) has been well studied and extensively used. In this paper, we present DPHMM ({Differentially Private Hidden Markov Model}), an HMM embedded with a private…