1.6k citations · 2.1k across the 4 of their papers we have counts for
4 papers
Self-training Improves Pre-training for Natural Language Understanding
Jingfei Du, Edouard Grave, Beliz Gunel +5
Unsupervised pre-training has led to much recent progress in natural language understanding. In this paper, we study self-training as another way to leverage unlabeled data through…
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau +4
Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of t…
Cross-lingual Language Model Pretraining
Guillaume Lample, Alexis Conneau
Recent studies have demonstrated the efficiency of generative pretraining for English natural language understanding. In this work, we extend this approach to multiple languages an…
Meta-Prod2Vec - Product Embeddings Using Side-Information for Recommendation
Flavian Vasile, Elena Smirnova, Alexis Conneau
We propose Meta-Prod2vec, a novel method to compute item similarities for recommendation that leverages existing item metadata. Such scenarios are frequently encountered in applica…