9 citations · 16 across the 2 of their papers we have counts for
2 papers
cs.CL2020★ 9 cited
Variance-reduced Language Pretraining via a Mask Proposal Network
Liang Chen
Self-supervised learning, a.k.a., pretraining, is important in natural language processing. Most of the pretraining methods first randomly mask some positions in a sentence and the…
cs.LG2019★ 7 cited
Domain-Aware Dynamic Networks
Tianyuan Zhang, Bichen Wu, Xin Wang +2
Deep neural networks with more parameters and FLOPs have higher capacity and generalize better to diverse domains. But to be deployed on edge devices, the model's complexity has to…