65 citations
- Christopher Akiki2 · h 12
- Daniel Alexander van Strien2 · h 10
- F. Toni2 · h 5
- Javier de la Rosa2 · h 9
- Aaron Gokaslan1 · h 9
- Aitor Soroa Etxabe1 · h 33
- Albert Villanova del Moral1 · h 17
- Angelina McMillan-Major1 · h 11
- Anna Rogers1 · h 24
- Chenghao Mou1 · h 7
- Clémentine Fourrier1 · h 10
- David Ifeoluwa Adelani1 · h 33
- British LibraryGB2 papers
- Leipzig UniversityDE2 papers
- The University of Western AustraliaAU2 papers
- Allen Institute for Artificial IntelligenceUS1 paper
- Bavarian State LibraryDE1 paper
- Booz Allen Hamilton (United States)US1 paper
- CentraleSupélecFR1 paper
- Cornell UniversityUS1 paper
- Ferrum CollegeUS1 paper
- Hugging FaceUS1 paper
- Humboldt-Universität zu BerlinDE1 paper
- King Fahd University of Petroleum and MineralsSA1 paper
2 papers
cs.CL2023★ 65 cited
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
cs.CL2022
Entities, Dates, and Languages: Zero-Shot on Historical Texts with T0
Francesco De Toni, Christopher Akiki, Javier de la Rosa +4
In this work, we explore whether the recently demonstrated zero-shot abilities of the T0 model extend to Named Entity Recognition for out-of-distribution languages and time periods…