4 papers · 1 filter
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
Nathan Godey, Éric de la Clergerie, Benoît Sagot
Recent advances in language modeling consist in pretraining highly parameterized neural networks on extremely large web-mined text corpora. Training and inference with such models…
On the Scaling Laws of Geographical Representation in Language Models
Nathan Godey, Éric de la Clergerie, Benoît Sagot
Language models have long been shown to embed geographical information in their hidden representations. This line of work has recently been revisited by extending this result to La…
Anisotropy Is Inherent to Self-Attention in Transformers
Nathan Godey, Éric de la Clergerie, Benoît Sagot
The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotrop…
Headless Language Models: Learning without Predicting with Contrastive Weight Tying
Nathan Godey, Éric de la Clergerie, Benoît Sagot
Self-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative…