90 citations · 90 across the 1 of their papers we have counts for
4 papers
Shortformer: Better Language Modeling using Shorter Inputs
Ofir Press, Noah A. Smith, Mike Lewis
Increasing the input length has been a driver of progress in language modeling with transformers. We identify conditions where shorter inputs are not harmful, and achieve perplexit…
Improving Transformer Models by Reordering their Sublayers
Ofir Press, Noah A. Smith, Omer Levy
Multilayer transformer networks consist of interleaved self-attention and feedforward sublayers. Could ordering the sublayers in a different pattern lead to better performance? We…
You May Not Need Attention
Ofir Press, Noah A. Smith
In NMT, how far can we get without attention and without separate encoding and decoding? To answer that question, we introduce a recurrent neural translation model that does not us…
Language Generation with Recurrent Generative Adversarial Networks without Pre-training
Ofir Press, Amir Bar, Ben Bogin +2
Generative Adversarial Networks (GANs) have shown great promise recently in image generation. Training GANs for language generation has proven to be more difficult, because of the…