1 citations · 2 across the 3 of their papers we have counts for
3 papers
One Wide Feedforward is All You Need
Telmo Pessoa Pires, António V. Lopes, Yannick Assogba +1
The Transformer architecture has two main non-embedding components: Attention and the Feed Forward Network (FFN). Attention captures interdependencies between words regardless of t…
Learning Language-Specific Layers for Multilingual Machine Translation
Telmo Pessoa Pires, Robin M. Schmidt, Yi-Hsiu Liao +1
Multilingual Machine Translation promises to improve translation quality between non-English languages. This is advantageous for several reasons, namely lower latency (no need to t…
State Spaces Aren't Enough: Machine Translation Needs Attention
Ali Vardasbi, Telmo Pessoa Pires, Robin M. Schmidt +1
Structured State Spaces for Sequences (S4) is a recently proposed sequence model with successful applications in various tasks, e.g. vision, language modeling, and audio. Thanks to…