activity
20162021
most citedCoAtNet: Marrying Convolution and Attention for All Data Sizes

742 citations · 1k across the 8 of their papers we have counts for

collaborators

20 papers

cs.LG202128 cited

Combiner: Full Attention Transformer with Sparse Computation Cost

Hongyu Ren, Hanjun Dai, Zihang Dai +4

Transformers provide a class of expressive architectures that are extremely effective for sequence modeling. However, the key limitation of transformers is their quadratic memory a…

cs.CV2021742 cited

CoAtNet: Marrying Convolution and Attention for All Data Sizes

Zihang Dai, Hanxiao Liu, Quoc V. Le +1

Transformers have attracted increasing interests in computer vision, but they still fall behind state-of-the-art convolutional networks. In this work, we show that while Transforme…

cs.LG202133 cited

Pay Attention to MLPs

Hanxiao Liu, Zihang Dai, David R. So +1

Transformers have become one of the most important architectural innovations in deep learning and have enabled many breakthroughs over the past few years. Here we propose a simple…

cs.CL20204 cited

Unsupervised Parallel Corpus Mining on Web Data

Guokun Lai, Zihang Dai, Yiming Yang

With a large amount of parallel data, neural machine translation systems are able to deliver human-level performance for sentence-level translation. However, it is costly to label…

cs.LG2020104 cited

Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing

Zihang Dai, Guokun Lai, Yiming Yang +1

With the success of language pretraining, it is highly desirable to develop more efficient architectures of good scalability that can exploit the abundant unlabeled data at a lower…

cs.LG2020

Meta Pseudo Labels

Hieu Pham, Zihang Dai, Qizhe Xie +2

We present Meta Pseudo Labels, a semi-supervised learning method that achieves a new state-of-the-art top-1 accuracy of 90.2% on ImageNet, which is 1.6% better than the existing st…