6 papers · 1 filter
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
Antonia Karamolegkou, Nicolas Angleraud, Benoît Sagot +1
Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on l…
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
Thibault Clérice, Rachel Bawden, Anthony Glaise +2
Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and norma…
How Should We Model the Probability of a Language?
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems exten…
KréyoLID From Language Identification Towards Language Mining
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may b…
Molyé: A Corpus-based Approach to Language Contact in Colonial France
Rasul Dent, Juliette Janès, Thibault Clérice +2
Whether or not several Creole languages which developed during the early modern period can be considered genetic descendants of European languages has been the subject of intense d…