490 citations · 789 across the 24 of their papers we have counts for
45 papers
MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation
Simiao Zuo, Qingru Zhang, Chen Liang +3
Pre-trained language models have demonstrated superior performance in various natural language processing tasks. However, these models usually contain hundreds of millions of param…
CAMERO: Consistency Regularized Ensemble of Perturbed Language Models with Weight Sharing
Chen Liang, Pengcheng He, Yelong Shen +2
Model ensemble is a popular approach to produce a low-variance and well-generalized model. However, it induces large memory and inference costs, which are often not affordable for…
No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models
Chen Liang, Haoming Jiang, Simiao Zuo +5
Recent research has shown the existence of significant redundancy in large Transformer models. One can prune the redundant parameters without significantly sacrificing the generali…
Besov Function Approximation and Binary Classification on Low-Dimensional Manifolds Using Convolutional Residual Networks
Hao Liu, Minshuo Chen, Tuo Zhao +1
Most of existing statistical theories on deep neural networks have sample complexities cursed by the data dimension and therefore cannot well explain the empirical success of deep…
Named Entity Recognition with Small Strongly Labeled and Large Weakly Labeled Data
Haoming Jiang, Danqing Zhang, Tianyu Cao +2
Weak supervision has shown promising results in many natural language processing tasks, such as Named Entity Recognition (NER). Existing work mainly focuses on learning deep NER mo…
COUnty aggRegation mixup AuGmEntation (COURAGE) COVID-19 Prediction
Siawpeng Er, Shihao Yang, Tuo Zhao
The global spread of COVID-19, the disease caused by the novel coronavirus SARS-CoV-2, has cast a significant threat to mankind. As the COVID-19 situation continues to evolve, pred…