65 citations · 77 across the 3 of their papers we have counts for
3 papers
cs.CL2023
ActiveAED: A Human in the Loop Improves Annotation Error Detection
Leon Weber, Barbara Plank
Manually annotated datasets are crucial for training and evaluating Natural Language Processing models. However, recent work has discovered that even widely-used benchmark datasets…
cs.CL2023★ 65 cited
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
cs.CL2022★ 12 cited
BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
Jason Alan Fries, Leon Weber, Natasha Seelam +40
Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompt…