activity
20192023
most citedA Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

75 citations · 94 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

14 papers · 1 filter

cs.CL20231 cited

Is Anisotropy Inherent to Transformers?

Nathan Godey, Éric de la Clergerie, Benoît Sagot

The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotrop…

cs.CL20224 cited

SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations

Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong +7

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments…

cs.CL2022

From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French

Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz +4

Language models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historica…

cs.CL2022

Towards a Cleaner Document-Oriented Multilingual Crawled Corpus

Julien Abadji, Pedro Ortiz Suarez, Laurent Romary +1

The need for raw large raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Pr…

cs.CL2020

When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models

Benjamin Muller, Antonis Anastasopoulos, Benoît Sagot +1

Transfer learning based on pretraining language models on a large amount of raw data has become a new norm to reach state-of-the-art performance in NLP. Still, it remains unclear h…

cs.CL2020

Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering

Arij Riabi, Thomas Scialom, Rachel Keraron +3

Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on the Question Answering task. However, most of those datasets are i…