2 citations · 2 across the 2 of their papers we have counts for
4 papers · 1 filter
How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers
Michael Hassid, Hao Peng, Daniel Rotem +4
The attention mechanism is considered the backbone of the widely-used Transformer architecture. It contextualizes the input by computing input-specific attention matrices. We find…
Sentence Bottleneck Autoencoders from Transformer Language Models
Ivan Montero, Nikolaos Pappas, Noah A. Smith
Representation learning for text via pretraining a language model on a large corpus has become a standard starting point for building NLP systems. This approach stands in contrast…
Pivot Through English: Reliably Answering Multilingual Questions without Document Retrieval
Ivan Montero, Shayne Longpre, Ni Lao +2
Existing methods for open-retrieval question answering in lower resource languages (LRLs) lag significantly behind English. They not only suffer from the shortcomings of non-Englis…
Plug and Play Autoencoders for Conditional Text Generation
Florian Mai, Nikolaos Pappas, Ivan Montero +2
Text autoencoders are commonly used for conditional generation tasks such as style transfer. We propose methods which are plug and play, where any pretrained autoencoder can be use…