most citedThe BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

65 citations · 78 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2023

Establishing Trustworthiness: Rethinking Tasks and Model Evaluation

Robert Litschko, Max Müller-Eberstein, Rob van der Goot +2

Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing (NLP) community has striven to model computationally for decades. Traditionall…

cs.CL20231 cited

BELB: a Biomedical Entity Linking Benchmark

Samuele Garda, Leon Weber-Genzel, Robert Martin +1

Biomedical entity linking (BEL) is the task of grounding entity mentions to a knowledge base. It plays a vital role in information extraction pipelines for the life sciences litera…

cs.CL2023

ActiveAED: A Human in the Loop Improves Annotation Error Detection

Leon Weber, Barbara Plank

Manually annotated datasets are crucial for training and evaluating Natural Language Processing models. However, recent work has discovered that even widely-used benchmark datasets…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…

cs.CL202212 cited

BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

Jason Alan Fries, Leon Weber, Natasha Seelam +40

Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompt…