4 papers
Modeling Orthographic Variation in Occitan's Dialects
Zachary William Hopton, Noëmi Aepli
Effectively normalizing textual data poses a considerable challenge, especially for low-resource languages lacking standardized writing systems. In this study, we fine-tuned a mult…
A Tulu Resource for Machine Translation
Manu Narayanan, Noëmi Aepli
We present the first parallel dataset for English-Tulu translation. Tulu, classified within the South Dravidian linguistic family branch, is predominantly spoken by approximately 2…
Modular Adaptation of Multilingual Encoders to Written Swiss German Dialect
Jannis Vamvas, Noëmi Aepli, Rico Sennrich
Creating neural text encoders for written Swiss German is challenging due to a dearth of training data combined with dialectal variation. In this paper, we build on several existin…
Findings of the VarDial Evaluation Campaign 2023
Noëmi Aepli, Çağrı Çöltekin, Rob Van Der Goot +7
This report presents the results of the shared tasks organized as part of the VarDial Evaluation Campaign 2023. The campaign is part of the tenth workshop on Natural Language Proce…