activity
20172022
most citedConformer: Convolution-augmented Transformer for Speech Recognition

387 citations · 876 across the 5 of their papers we have counts for

collaborators

15 papers

cs.CL2021

Simple and Efficient ways to Improve REALM

Vidhisha Balachandran, Ashish Vaswani, Yulia Tsvetkov +1

Dense retrieval has been shown to be effective for retrieving relevant documents for Open Domain QA, surpassing popular sparse retrieval methods like BM25. REALM (Guu et al., 2020)…

cs.CV2021

Scaling Local Self-Attention for Parameter Efficient Visual Backbones

Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas +3

Self-attention has the promise of improving computer vision systems due to parameter-independent scaling of receptive fields and content-dependent interactions, in contrast to para…

cs.CV2021

Bottleneck Transformers for Visual Recognition

Aravind Srinivas, Tsung-Yi Lin, Niki Parmar +3

We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, obje…

eess.AS2020387 cited

Conformer: Convolution-augmented Transformer for Speech Recognition

Anmol Gulati, James Qin, Chung-Cheng Chiu +8

Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (…

eess.IV2019

High Resolution Medical Image Analysis with Spatial Partitioning

Le Hou, Youlong Cheng, Noam Shazeer +6

Medical images such as 3D computerized tomography (CT) scans and pathology images, have hundreds of millions or billions of voxels/pixels. It is infeasible to train CNN models dire…

cs.CV2019221 cited

Stand-Alone Self-Attention in Vision Models

Prajit Ramachandran, Niki Parmar, Ashish Vaswani +3

Convolutions are a fundamental building block of modern computer vision systems. Recent approaches have argued for going beyond convolutions in order to capture long-range dependen…