12 papers
Robustness of transferability estimation metrics for medical imaging
Niclas Claßen, Théo Sourget, Dovile Juodelyte +2
In transfer learning, the choice of source model largely influences the performance on a target dataset. Still, selecting a fitting source remains a challenging task, especially in…
Term-Centric Hierarchy Induction from Heterogeneous Corpora
Elena Senger, Yuri Campbell, Jan-Peter Bergmann +2
Organizing knowledge from diverse text sources into interpretable hierarchies is crucial for tasks such as policy analysis, innovation monitoring, and exploratory domain mapping. E…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
Dataset Diversity Metrics and Impact on Classification Models
Théo Sourget, Niclas ClaÃen, Jack Junchi Xu +2
The diversity of training datasets is usually perceived as an important aspect to obtain a robust model. However, the definition of diversity is often not defined or differs across…
MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages
Weerayut Buaphet, Thanh-Nhi Nguyen, Risa Kondo +8
Social media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automat…
Do Syntactic Categories Help in Developmentally Motivated Curriculum Learning for Language Models?
Arzu Burcu Güven, Anna Rogers, Rob van der Goot
We examine the syntactic properties of BabyLM corpus, and age-groups within CHILDES. While we find that CHILDES does not exhibit strong syntactic differentiation by age, we show th…