5 citations · 7 across the 8 of their papers we have counts for
3 papers · 1 filter
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
Vukasin Bozic, Danilo Dordevic, Daniele Coppola +2
This work presents an analysis of the effectiveness of using standard shallow feed-forward networks to mimic the behavior of the attention mechanism in the original Transformer mod…
GLOSS: Generative Latent Optimization of Sentence Representations
Sidak Pal Singh, Angela Fan, Michael Auli
We propose a method to learn unsupervised sentence representations in a non-compositional manner based on Generative Latent Optimization. Our approach does not impose any assumptio…
Context Mover's Distance & Barycenters: Optimal Transport of Contexts for Building Representations
Sidak Pal Singh, Andreas Hug, Aymeric Dieuleveut +1
We present a framework for building unsupervised representations of entities and their compositions, where each entity is viewed as a probability distribution rather than a vector…