65 citations · 67 across the 5 of their papers we have counts for
Showing 2023Show all
3 papers · 1 filter
cs.LG2023
Curriculum Learning with Adam: The Devil Is in the Wrong Details
Lucas Weber, Jaap Jumelet, Paul Michel +2
Curriculum learning (CL) posits that machine learning models -- similar to humans -- may learn more efficiently from data that match their current learning progress. However, CL me…
cs.CL2023
ActiveAED: A Human in the Loop Improves Annotation Error Detection
Leon Weber, Barbara Plank
Manually annotated datasets are crucial for training and evaluating Natural Language Processing models. However, recent work has discovered that even widely-used benchmark datasets…
cs.CL2023★ 65 cited
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…