2 papers
cs.CL2020
On the importance of pre-training data volume for compact language models
Vincent Micheli, Martin d'Hoffschmidt, François Fleuret
Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models. In an effort towards sustainable practices, we study the…
cs.CL2020
FQuAD: French Question Answering Dataset
Martin d'Hoffschmidt, Wacim Belblidia, Tom Brendlé +2
Recent advances in the field of language modeling have improved state-of-the-art results on many Natural Language Processing tasks. Among them, Reading Comprehension has made signi…