163 citations · 201 across the 12 of their papers we have counts for
4 papers · 1 filter
Learning strides in convolutional neural networks
Rachid Riad, Olivier Teboul, David Grangier +1
Convolutional neural networks typically contain several downsampling operators, such as strided convolutions or pooling layers, that progressively reduce the resolution of intermed…
Auxiliary Task Update Decomposition: The Good, The Bad and The Neutral
Lucio M. Dery, Yann Dauphin, David Grangier
While deep learning has been very beneficial in data-rich settings, tasks with smaller training set often resort to pre-training or multitask learning to leverage data from other t…
Efficient Content-Based Sparse Attention with Routing Transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani +1
Self-attention has recently been adopted for a wide range of sequence modeling problems. Despite its effectiveness, self-attention suffers from quadratic compute and memory require…
Unsupervised Paraphrasing without Translation
Aurko Roy, David Grangier
Paraphrasing exemplifies the ability to abstract semantic content from surface forms. Recent work on automatic paraphrasing is dominated by methods leveraging Machine Translation (…