most citedProject PIAF: Building a Native French Question-Answering Dataset

6 citations · 12 across the 2 of their papers we have counts for

collaborators

6 papers

cs.CL20206 cited

Project PIAF: Building a Native French Question-Answering Dataset

Rachel Keraron, Guillaume Lancrenon, Mathilde Bras +5

Motivated by the lack of data for non-English languages, in particular for the evaluation of downstream tasks such as Question Answering, we present a participatory effort to colle…

cs.CL20206 cited

ColdGANs: Taming Language GANs with Cautious Sampling Strategies

Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier +2

Training regimes based on Maximum Likelihood Estimation (MLE) suffer from known limitations, often leading to poorly generated text sequences. At the root of these limitations is t…

cs.CL2020

MLSUM: The Multilingual Summarization Corpus

Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier +2

We present MLSUM, the first large-scale MultiLingual SUMmarization dataset. Obtained from online newspapers, it contains 1.5M+ article/summary pairs in five different languages --…

cs.CL2020

Discriminative Adversarial Search for Abstractive Summarization

Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier +2

We introduce a novel approach for sequence decoding, Discriminative Adversarial Search (DAS), which has the desirable properties of alleviating the effects of exposure bias without…

cs.CL2019

Ask to Learn: A Study on Curiosity-driven Question Generation

Thomas Scialom, Jacopo Staiano

We propose a novel text generation task, namely Curiosity-driven Question Generation. We start from the observation that the Question Generation task has traditionally been conside…

cs.CL2019

Answers Unite! Unsupervised Metrics for Reinforced Summarization Models

Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski +1

Abstractive summarization approaches based on Reinforcement Learning (RL) have recently been proposed to overcome classical likelihood maximization. RL enables to consider complex,…