13 citations · 18 across the 8 of their papers we have counts for
12 papers
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
Nathan Godey, Alessio Devoto, Yu Zhao +4
Autoregressive language models rely on a Key-Value (KV) Cache, which avoids re-computing past hidden states during generation, making it faster. As model sizes and context lengths…
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
Wissam Antoun, Francis Kulumba, Rian Touchent +3
French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million…
Anisotropy Is Inherent to Self-Attention in Transformers
Nathan Godey, Éric de la Clergerie, Benoît Sagot
The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotrop…
Headless Language Models: Learning without Predicting with Contrastive Weight Tying
Nathan Godey, Éric de la Clergerie, Benoît Sagot
Self-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative…
Is Anisotropy Inherent to Transformers?
Nathan Godey, Éric de la Clergerie, Benoît Sagot
The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotrop…
MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling
Nathan Godey, Roman Castagné, Éric de la Clergerie +1
Static subword tokenization algorithms have been an essential component of recent works on language modeling. However, their static nature results in important flaws that degrade t…