3 papers
cs.CL2026
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
cs.CL2023★ 3 cited
Uncertainty-Aware Natural Language Inference with Stochastic Weight Averaging
Aarne Talman, Hande Celikkanat, Sami Virpioja +2
This paper introduces Bayesian uncertainty modeling using Stochastic Weight Averaging-Gaussian (SWAG) in Natural Language Understanding (NLU) tasks. We apply the approach to standa…
cs.CL2019
Predicting Prosodic Prominence from Text with Pre-trained Contextualized Word Representations
Aarne Talman, Antti Suni, Hande Celikkanat +3
In this paper we introduce a new natural language processing dataset and benchmark for predicting prosodic prominence from written text. To our knowledge this will be the largest p…