65 citations · 71 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 6 cited
Semi-automatic staging area for high-quality structured data extraction from scientific literature
Luca Foppiano, Tomoya Mato, Kensei Terashima +7
We propose a semi-automatic staging area for efficiently building an accurate database of experimental physical properties of superconductors from literature, called SuperCon2, to…
cs.CL2023★ 65 cited
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…